|
b4dc103931
|
I don't know how I did not commit the 'sample from the voices to construct the input prompt for vall-e' change but this helps
|
2023-08-25 04:26:48 +00:00 |
|
|
a657623cbc
|
updated vall-e training template to use path-based speakers because it would just have a batch/epoch size of 1 otherwise; revert hardcoded 'spit processed dataset to this path' from my training rig to spit it out in a sane spot
|
2023-08-24 21:45:50 +00:00 |
|
|
533b73e083
|
fixed the overwrite regression for bark and vall-e backends too
|
2023-08-24 19:46:42 +00:00 |
|
mrq
|
4aa240d48a
|
Merge pull request 'fix filename generation which didn't work and overwrote existing files' (#341) from ben_mkiv/ai-voice-cloning:master into master
Reviewed-on: mrq/ai-voice-cloning#341
|
2023-08-24 12:29:59 +00:00 |
|
|
00b173857d
|
fix filename generation which didn't work and overwrote existing files
|
2023-08-24 09:57:01 +02:00 |
|
|
dc46fdc7d0
|
fixed another issue from haphazardly copying my changes from my training machine
|
2023-08-23 22:09:22 +00:00 |
|
|
29290f574e
|
should fix issue that arises when trying to prepare the dataset without slicing segments
|
2023-08-23 21:49:22 +00:00 |
|
|
2060b6f21c
|
fixed issue with sliced audio being the wrong sample rate
|
2023-08-22 14:22:39 +00:00 |
|
|
eeddd4cb6b
|
forgot the important reason I even started working on AIVC again
|
2023-08-21 03:42:12 +00:00 |
|
|
72a38ff2fc
|
made initialization faster if there's a lot of voice files (because glob fucking sucks), commiting changes buried on my training rig
|
2023-08-21 03:31:49 +00:00 |
|
|
ac645e0a20
|
no longer need to install bark under ./modules/
|
2023-07-11 16:20:28 +00:00 |
|
|
e2a6dc1c0a
|
under bark, properly use transcribed audio if the audio wasn't actually sliced (oops)
|
2023-07-11 14:53:32 +00:00 |
|
|
6c3f48efba
|
uses gitmylo/bark-voice-cloning-HuBERT-quantizer for creating custom voices (it slightly works better over the base method, but still not very good desu)
|
2023-07-03 02:46:10 +00:00 |
|
|
547e1d1277
|
updated bark support, it'll also query for vocos, it actually works (I don't know what specifically was the issue)
|
2023-07-03 01:22:02 +00:00 |
|
|
76ed34ddd2
|
added CLI script (python ./src/cli.py --text=TEXT --voice=VOICE' etc)
|
2023-06-11 04:46:22 +00:00 |
|
|
e227ab8e08
|
updated whisperX integration for use with the latest version (v3) (NOTE: you WILL need to also update whisperx if you pull this commit)
|
2023-06-09 02:41:29 +00:00 |
|
|
805d7d35e8
|
the power of a separate setup for testing
|
2023-05-22 17:36:28 +00:00 |
|
|
2f5486a8d5
|
oops
|
2023-05-21 23:24:13 +00:00 |
|
|
31da215c5f
|
added checkboxes to use the original method for calculating latents (ignores the voice chunk field)
|
2023-05-21 01:47:48 +00:00 |
|
|
cbe21745df
|
I am very smart (need to validate)
|
2023-05-12 17:41:26 +00:00 |
|
|
74bd0f0cdc
|
revert local change that made its way upstream (showing graphs by it instead of epoch)
|
2023-05-11 03:30:54 +00:00 |
|
|
149aaca554
|
fixed the whisperx has no attribute named load_model whatever because I guess whisperx has as stable of an API as I do
|
2023-05-06 10:45:17 +00:00 |
|
|
e416b0fe6f
|
oops
|
2023-05-05 12:36:48 +00:00 |
|
|
5003bc89d3
|
cleaned up brain worms with wrapping around gradio progress by instead just using tqdm directly (slight regressions with some messages not getting pushed)
|
2023-05-04 23:40:33 +00:00 |
|
|
09d849a78f
|
quick hotfix if it actually is a problem in the repo itself
|
2023-05-04 23:01:47 +00:00 |
|
|
853c7fdccf
|
forgot to uncomment the block to transcribe and slice when using transcribe all because I was piece-processing a huge batch of LibriTTS and somehow that leaked over to the repo
|
2023-05-03 21:31:37 +00:00 |
|
|
eddb8aaa9a
|
indentation fix
|
2023-04-28 15:56:57 +00:00 |
|
|
99387920e1
|
backported caching of phonemizer backend from mrq/vall-e
|
2023-04-28 15:31:45 +00:00 |
|
|
b6440091fb
|
Very, very, VERY, barebones integration with Bark (documentation soon)
|
2023-04-26 04:48:09 +00:00 |
|
|
faa8da12d7
|
modified logic to determine valid voice folders, also allows subdirs within the folder (for example: ./voices/SH/james/ will be named SH/james)
|
2023-04-13 21:10:38 +00:00 |
|
|
02beb1dd8e
|
should fix #203
|
2023-04-13 03:14:06 +00:00 |
|
|
d8b996911c
|
a bunch of shit i had uncommited over the past while pertaining to VALL-E
|
2023-04-12 20:02:46 +00:00 |
|
|
4744120be2
|
added VALL-E inference support (very rudimentary, gimped, but it will load a model trained on a config generated through the web UI)
|
2023-03-31 03:26:00 +00:00 |
|
|
9b01377667
|
only include auto in the list of models under setting, nothing else
|
2023-03-29 19:53:23 +00:00 |
|
|
f66281f10c
|
added mixing models (shamelessly inspired from voldy's web ui)
|
2023-03-29 19:29:13 +00:00 |
|
|
c89c648b4a
|
fixes #176
|
2023-03-26 11:05:50 +00:00 |
|
|
41d47c7c2a
|
for real this time show those new vall-e metrics
|
2023-03-26 04:31:50 +00:00 |
|
|
c4ca04cc92
|
added showing reported training accuracy and eval/validation metrics to graph
|
2023-03-26 04:08:45 +00:00 |
|
|
8c647c889d
|
now there should be feature parity between trainers
|
2023-03-25 04:12:03 +00:00 |
|
|
fd9b2e082c
|
x_lim and y_lim for graph
|
2023-03-25 02:34:14 +00:00 |
|
|
9856db5900
|
actually make parsing VALL-E metrics work
|
2023-03-23 15:42:51 +00:00 |
|
|
69d84bb9e0
|
I forget
|
2023-03-23 04:53:31 +00:00 |
|
|
444bcdaf62
|
my sanitizer actually did work, it was just batch sizes leading to problems when transcribing
|
2023-03-23 04:41:56 +00:00 |
|
|
a6daf289bc
|
when the sanitizer thingy works in testing but it doesn't outside of testing, and you have to retranscribe for the fourth time today
|
2023-03-23 02:37:44 +00:00 |
|
|
86589fff91
|
why does this keep happening to me
|
2023-03-23 01:55:16 +00:00 |
|
|
0ea93a7f40
|
more cleanup, use 24KHz for preparing for VALL-E (encodec will resample to 24Khz anyways, makes audio a little nicer), some other things
|
2023-03-23 01:52:26 +00:00 |
|
|
d2a9ab9e41
|
remove redundant phonemize for vall-e (oops), quantize all files and then phonemize all files for cope optimization, load alignment model once instead of for every transcription (speedup with whisperx)
|
2023-03-23 00:22:25 +00:00 |
|
|
19c0854e6a
|
do not write current whisper.json if there's no changes
|
2023-03-22 22:24:07 +00:00 |
|
|
932eaccdf5
|
added whisper transcription 'sanitizing' (collapse very short transcriptions to the previous segment) (I really have to stop having several copies spanning several machines for AIVC, I keep reverting shit)
|
2023-03-22 22:10:01 +00:00 |
|
|
736cdc8926
|
disable diarization for whisperx as it's just a useless performance hit (I don't have anything that's multispeaker within the same audio file at the moment)
|
2023-03-22 20:38:58 +00:00 |
|