James Betker
|
4c01d82265
|
Fix for voxpopuli
|
2021-08-16 22:52:05 -06:00 |
|
James Betker
|
1fede41b7b
|
Audio segmentor
|
2021-08-16 22:51:53 -06:00 |
|
James Betker
|
2d3372054d
|
Add support for voxpopuli to nv_tacotron_dataset
|
2021-08-16 17:13:40 -06:00 |
|
James Betker
|
729c1fd5a9
|
Fix up max lengths to save memory
|
2021-08-15 21:29:28 -06:00 |
|
James Betker
|
9e47e64d5a
|
Add gpt_segmentor model
The idea is to specifically train a model that extracts phrases from
audio clips.
|
2021-08-15 21:23:07 -06:00 |
|
James Betker
|
a826d5f658
|
Mods to dvae
- Add resblock to each layer
- Increase filter size for each layer
- Use SiLU
|
2021-08-15 20:54:10 -06:00 |
|
James Betker
|
b8bec22f1a
|
Fix gpt_asr inference bug
|
2021-08-15 20:53:42 -06:00 |
|
James Betker
|
3580c52eac
|
Fix up wavfile_dataset to be able to provide a full clip
|
2021-08-15 20:53:26 -06:00 |
|
James Betker
|
a523c4f932
|
Auto-normalize wav files by data type
|
2021-08-15 09:09:51 -06:00 |
|
James Betker
|
98057b6516
|
Make lrdvae use quantized mode in eval()
|
2021-08-14 23:43:01 -06:00 |
|
James Betker
|
c28f657ab8
|
Allow usage of pre-rendered mels saved to npy files
|
2021-08-14 23:38:15 -06:00 |
|
James Betker
|
ad3391bd96
|
Fix nan issue when interpolating audio
|
2021-08-14 20:42:01 -06:00 |
|
James Betker
|
769f0acc53
|
Moar fix
|
2021-08-14 17:23:15 -06:00 |
|
James Betker
|
3d2e724083
|
Fix audio ranging problem
|
2021-08-14 17:18:55 -06:00 |
|
James Betker
|
d6a73acaed
|
Allow processing of multiple audio sources at once from nv_tacotron_dataset
|
2021-08-14 16:04:05 -06:00 |
|
James Betker
|
007976082b
|
GPT_asr for inference
|
2021-08-14 14:37:17 -06:00 |
|
James Betker
|
e1bdd3f7c7
|
Fix gpt_asr bug. Initial implementation of beam search
|
2021-08-13 22:47:00 -06:00 |
|
James Betker
|
72622b4d61
|
Allow saving mel strips as files from the dataset implementation
|
2021-08-13 22:46:41 -06:00 |
|
James Betker
|
cfd284f425
|
Fix up some stuff that allows the MEL to be computed on-GPU
|
2021-08-13 18:35:55 -06:00 |
|
James Betker
|
cdee31c60b
|
GPT_ASR
|
2021-08-13 15:02:18 -06:00 |
|
James Betker
|
81e91c99de
|
Misc
|
2021-08-13 13:58:59 -06:00 |
|
James Betker
|
fff1a59e08
|
max/min mel invalid fix
|
2021-08-13 09:36:31 -06:00 |
|
James Betker
|
4b2946e581
|
More fix
|
2021-08-12 15:51:23 -06:00 |
|
James Betker
|
4c76257c71
|
Dont require collation for nv_tacotron
|
2021-08-12 15:44:55 -06:00 |
|
James Betker
|
5b07d3b623
|
Found error that I was trying to fix with reload=True
|
2021-08-12 15:22:34 -06:00 |
|
James Betker
|
430b650a34
|
......
|
2021-08-12 10:31:10 -06:00 |
|
James Betker
|
b35d6ae028
|
Print some metrics from tacotron dataset when it croaks
|
2021-08-12 09:21:12 -06:00 |
|
James Betker
|
0c4d6b1916
|
Just offer generic re-load for nv-tacotron
|
2021-08-12 09:09:12 -06:00 |
|
James Betker
|
154f5aa73c
|
Fix annoying warning and add to requirements
|
2021-08-11 17:32:06 -06:00 |
|
James Betker
|
f5a9b88ef6
|
tacotron cleaners: remove quotation marks
these don't really have relevance for tts or asr
|
2021-08-11 16:18:44 -06:00 |
|
James Betker
|
20586a8edc
|
Fix LRDVAE bug with quantizer integration
|
2021-08-11 16:17:22 -06:00 |
|
James Betker
|
f04a7bdf63
|
Bug fixes for tacotron dataset on mozilla cv
- Support a max mel length (mozilla cv has some tracks that are basically unbounded..)
- Don't fail on low sample rates (mozilla cv has some of those)
|
2021-08-11 16:17:03 -06:00 |
|
James Betker
|
98f37241cf
|
Update idea files
This shouldnt be versioned controlled...
|
2021-08-11 16:15:55 -06:00 |
|
James Betker
|
2d3f0cc33c
|
nv_tacotron_dataset - Allow training on mozilla cv
|
2021-08-11 13:34:31 -06:00 |
|
James Betker
|
d0c74278bf
|
Enable multiple wavfile paths to be specified, fix eps bug in mp3 splitter
|
2021-08-11 08:46:02 -06:00 |
|
James Betker
|
e19c00398e
|
More improvements to random_mp3_splitter
|
2021-08-09 21:31:12 -06:00 |
|
James Betker
|
04d14b3acc
|
No batch factors for eval
|
2021-08-09 16:02:01 -06:00 |
|
James Betker
|
82fc69abfa
|
Add "pure" evaluator
Which simply computes the training loss against an eval dataset
|
2021-08-09 14:58:35 -06:00 |
|
James Betker
|
080bea2f19
|
No, really
|
2021-08-09 12:02:31 -06:00 |
|
James Betker
|
e1ce4671e4
|
Apply dropout to gpt_tts, get rid of min_gpt implementation
|
2021-08-09 12:01:10 -06:00 |
|
James Betker
|
74342b860b
|
Revert "Undo forced text padding"
This reverts commit 83ab5e6a00 .
|
2021-08-09 11:56:34 -06:00 |
|
James Betker
|
1068f53b78
|
Add a sampling beam search
|
2021-08-09 11:56:06 -06:00 |
|
James Betker
|
d4e33bf15f
|
Fixes to the mp3 splitter
|
2021-08-09 11:55:46 -06:00 |
|
James Betker
|
4100469902
|
Add a tool to split mp3 files into arbitrary chunks of wav files
|
2021-08-08 23:23:13 -06:00 |
|
James Betker
|
01cfae28d8
|
Beam search implementation in one pass? Dayyyum
|
2021-08-08 23:22:42 -06:00 |
|
James Betker
|
83ab5e6a00
|
Undo forced text padding
|
2021-08-08 11:42:20 -06:00 |
|
James Betker
|
690d7e86d3
|
Fix nv_tacotron_dataset bug which incorrectly mapped filenames
dammit..
|
2021-08-08 11:38:52 -06:00 |
|
James Betker
|
a2afb25e42
|
Fix inference, always flow full text tokens through transformer
|
2021-08-07 20:11:10 -06:00 |
|
James Betker
|
4c678172d6
|
ugh
|
2021-08-06 22:10:18 -06:00 |
|
James Betker
|
e723137273
|
Make gpttts more configurable
|
2021-08-06 22:08:51 -06:00 |
|