VLC to convert video to audio
This doesn't work for me, for some reason. When I tell it to extract the audio no file gets created.
VLC to convert video to audio
find . -type f -iregex '.*\.\(m4v\|mp4\|mov\|wmv\|avi\|mpg\|mpeg\|rmvb\|rm\|flv\|asf\|mkv\|webm\)' -print0 | parallel -0 ffmpeg -loglevel 0 -i {} -af "dynaudnorm=f=150" -ac 1 -ar 16000 -vn {.}.mp3
I'm going to write out the steps I have to take, and you tell me where it stops working for you.This doesn't work for me, for some reason. When I tell it to extract the audio no file gets created.
Now let's look at even the Medium basic model from Whisper. The only thing it gets wrong is who married Mr. Kazuo. Her mother married Kazuo. I see that mistake a lot, with it incorrectly identifying the subject of a sentence. It is often after these types of intros, and is probably because of the link between sentences-- it is very common for introductions to be talking about themselves, so it is learning that incorrectly here.1
00:00:06,450 --> 00:00:07,500
my name is
2
00:00:08,310 --> 00:00:09,330
Haruka Fujisaki
3
00:00:12,270 --> 00:00:13,680
when i was in high school
4
00:00:14,310 --> 00:00:15,960
I lost his father
5
00:00:16,950 --> 00:00:18,840
to help my mother who was left behind
6
00:00:19,560 --> 00:00:20,730
after graduation
7
00:00:20,940 --> 00:00:22,080
a friend's restaurant
8
00:00:22,080 --> 00:00:23,610
decided to work at
9
00:00:24,990 --> 00:00:25,990
and
10
00:00:26,670 --> 00:00:29,880
With Kazuo who was a customer of the store
11
00:00:30,510 --> 00:00:30,750
3
12
00:00:30,778 --> 00:00:32,820
Married after dating for years
13
00:00:34,020 --> 00:00:35,820
with his sister Miku-chan
14
00:00:36,630 --> 00:00:40,170
I was living a happy life with my mother
15
00:00:41,700 --> 00:00:42,700
However
16
00:00:43,350 --> 00:00:43,470
2
17
00:00:43,500 --> 00:00:44,500
Months ago
18
00:00:45,000 --> 00:00:48,030
Her mother, who had been hospitalized with a mild illness, died suddenly.
19
00:00:50,250 --> 00:00:51,250
I
20
00:00:51,840 --> 00:00:53,850
her mother is in the hospital
21
00:00:54,450 --> 00:00:54,990
medical error
22
00:00:54,990 --> 00:00:56,580
I decided to sue
23
00:00:58,290 --> 00:00:59,290
However
24
00:00:59,760 --> 00:01:00,990
in a recession from time to time
25
00:01:01,650 --> 00:01:02,910
restaurant where I worked
26
00:01:02,910 --> 00:01:04,890
suddenly decided to close the store
27
00:01:05,820 --> 00:01:06,820
I
28
00:01:07,410 --> 00:01:07,650
fat
29
00:01:07,800 --> 00:01:08,250
group
30
00:01:08,250 --> 00:01:08,820
The restaurant in
31
00:01:08,820 --> 00:01:10,410
I started working at
32
00:01:13,980 --> 00:01:14,980
she is
33
00:01:15,300 --> 00:01:16,170
Yabuta Group
34
00:01:16,170 --> 00:01:17,700
Sachiko-san, a talented person of
35
00:01:19,170 --> 00:01:20,730
a stranger to me
36
00:01:21,480 --> 00:01:22,050
manager
37
00:01:22,080 --> 00:01:24,060
I am the benefactor who took me as
38
00:01:25,620 --> 00:01:26,620
and
39
00:01:27,420 --> 00:01:29,010
sitting in
40
00:01:29,310 --> 00:01:31,350
My sister-in-law Miku-chan
41
00:01:32,880 --> 00:01:34,170
Miku-chan is now
42
00:01:35,280 --> 00:01:37,080
I go to Squirting Girls' Academy
43
00:01:38,430 --> 00:01:39,430
Haruka
44
00:01:39,540 --> 00:01:40,540
Miki
45
00:01:41,310 --> 00:01:44,940
The two of you are helping me out at the shop, and I'm really saved.
46
00:01:48,300 --> 00:01:49,440
Even so
47
00:01:50,310 --> 00:01:51,690
a little sooner
48
00:01:52,290 --> 00:01:54,180
If I had met Haruka
49
00:01:55,308 --> 00:01:58,110
I was able to see your mother at the hospital in
I am getting these nice long sentences with punctuation because I set beam_size to 12 or 15. 20 crashes on my computer, but I think the misplaced "I" pronoun is because of where it chose to make the sentence/line break.[00:00.000 --> 00:23.000] My name is Haruka Fujisaki When I was a high school student, I lost my father and decided to work at a restaurant that I knew after graduation to help my mother who was left behind.
[00:23.000 --> 00:40.000] And I married Mr. Kazuo, who was a customer at the restaurant, and his sister Miku, who had been married for three years, and my mother's four children.
[00:40.000 --> 00:56.000] However, two months ago, my mother, who was hospitalized for a mild illness, suddenly died, and I decided to sue the hospital where my mother was hospitalized for a medical mistake.
[00:56.000 --> 01:12.000] However, the restaurant I was working at suddenly closed, and I decided to work at a restaurant in Yabuta Group.
[01:12.000 --> 01:37.000] She is Yukiko, the owner of Yabuta Group. She hired me as her manager. And Miku, my sister, is sitting next to me. Miku is currently going to Kusunoki Jogakuin.
[01:37.000 --> 01:47.000] Haruka-san, Miku-san, you two are really helpful for helping the restaurant.
[01:47.000 --> 01:59.000] Even so, if I had met Haruka-san a little earlier, I would have been able to see my mother at my hospital.
[01:59.000 --> 02:03.000] I'm really sorry.
[02:03.000 --> 02:08.000] Yukiko-san, thank you very much.

It depends what you're going for. The value of Whisper isn't to get the highest quality translation possible, it is about lowering the effort bar to get something good enough. Medium Whisper, with well-chosen parameters and some manual editing from just watching the move and editing as you go, is going to be better and more complete than like 95% of the subtitles that are currently out there for JAV. The places that are good and doing manual translations tend to miss a bunch of the smaller text, but do a better job of making the lines that are supposed to sound dirty be appropriately so.Just a tip on Whisper about translation: Don't use it. Just choose No Translation and do that later in Google translate or DeepL translate (use Tor Browser so you don't get blocked for using up your free ratio (500.000 characters a month). You will need to separate the text from the timing first and then re-attach it as some of the translators introduces glitches in it.
I don't know how good the translator Whisper uses is (you can choose DeepL, too) but if you don't have a basic reference to what is actually transcribed, you can't figure out what the line is if the translator produces garbage.
I can only use the collab page so many of the 'settings' mentioned in posts before this can't be used. I have experimented with VAD-Threshold and have settled on 0.3. What does Chunk_Threshold (3.0) do? Size of the audio parts analyzed? I'll try to lower and raise the number to see if there are improvements. Source separation didn't work for me. I get error messages.
Here are two raw Whisper files for BKD-127 and BKD-153. I don't intend to clean them as I just want to use them to compare to existing EroJapanese subtitles in an attempt to get a better understanding of what Whisper means with dialog like "I'm hungry", I'm sleepy, etc. rather than just relying on my wild ass guess. My goal is to create a spreadsheet to use with the more lewd EroJapanese vocabulary to use when re-interpreting Whisper's dialog.
Sounds interesting. Do you know how the command should look like for command line users like me?ffmpeg can do some processing on audio as it extracts it from a video file. I am currently using this command that will find all video files in a directory and then write out an mp3 for each. It also does a dynamic normalization to increase the volume of quiet parts of the track. (This is a bash shell command.)
Code:find . -type f -iregex '.*\.\(m4v\|mp4\|mov\|wmv\|avi\|mpg\|mpeg\|rmvb\|rm\|flv\|asf\|mkv\|webm\)' -print0 | parallel -0 ffmpeg -loglevel 0 -i {} -af "dynaudnorm=f=150" -ac 1 -ar 16000 -vn {.}.mp3
There are other ffmpeg settings that can do dynamic range compression and other fancy things.
https://ffmpeg.org/ffmpeg-all.html#dynaudnorm
Have had very good luck w/ the above, so not seeing much need to try and isolate voices from the regular audio track,
If you're just using Whisper for the transcribing into Japanese and avoiding the translation, you can drop the .srt into the free Subtitle Edit program and easily auto-translate it. I believe it uses Google to translate languages. I have been doing this for years with the Chinese subtitle files I could find and translating them to English. It was my go-to method before Whisper came around.Just a tip on Whisper about translation: Don't use it. Just choose No Translation and do that later in Google translate or DeepL translate (use Tor Browser so you don't get blocked for using up your free ratio (500.000 characters a month). You will need to separate the text from the timing first and then re-attach it as some of the translators introduces glitches in it.
I don't know how good the translator Whisper uses is (you can choose DeepL, too) but if you don't have a basic reference to what is actually transcribed, you can't figure out what the line is if the translator produces garbage.
I can only use the collab page so many of the 'settings' mentioned in posts before this can't be used. I have experimented with VAD-Threshold and have settled on 0.3. What does Chunk_Threshold (3.0) do? Size of the audio parts analyzed? I'll try to lower and raise the number to see if there are improvements. Source separation didn't work for me. I get error messages.
How do you turn off the translation in Whisper? Does Subtitle Edit also have a limited "Lewd" vocabulary like Whisper?If you're just using Whisper for the transcribing into Japanese and avoiding the translation, you can drop the .srt into the free Subtitle Edit program and easily auto-translate it. I believe it uses Google to translate languages. I have been doing this for years with the Chinese subtitle files I could find and translating them to English. It was my go-to method before Whisper came around.
you change the "task" from "Translate" to "Transcribe"How do you turn off the translation in Whisper? Does Subtitle Edit also have a limited "Lewd" vocabulary like Whisper?
Sounds interesting. Do you know how the command should look like for command line users like me?
