James Betker
|
f3cab45658
|
Revise audio datasets to include interesting statistics in batch
Stats include:
- How many indices were skipped to retrieve a given index
- Whether or not a conditioning input was actually the file itself
|
2022-01-06 11:15:16 -07:00 |
|
James Betker
|
06c1093090
|
Remove collating from paired_voice_audio_dataset
This will now be done at the model level, which is more efficient
|
2022-01-06 10:29:39 -07:00 |
|
James Betker
|
5e1d1da2e9
|
Clean paired_voice
|
2022-01-06 10:26:53 -07:00 |
|
James Betker
|
0fe34f57d1
|
Use torch resampler
|
2022-01-05 15:47:22 -07:00 |
|
James Betker
|
d4a6298658
|
more debugging
|
2022-01-01 14:25:27 -07:00 |
|
James Betker
|
35abefd038
|
More fix
|
2022-01-01 10:31:03 -07:00 |
|
James Betker
|
d5a5111890
|
Fix collating on by default on grand_conjoined
|
2022-01-01 10:30:15 -07:00 |
|
James Betker
|
4d9ba4a48a
|
can i has fix now
|
2022-01-01 00:48:27 -07:00 |
|
James Betker
|
56752f1dbc
|
Fix collator bug
|
2022-01-01 00:33:31 -07:00 |
|
James Betker
|
c28d8770c7
|
fix tensor lengths
|
2022-01-01 00:23:46 -07:00 |
|
James Betker
|
bbacffb790
|
dataset improvements and fix to unified_voice_Bilevel
|
2022-01-01 00:16:30 -07:00 |
|
James Betker
|
17fb934575
|
wer update
|
2021-12-31 16:21:39 -07:00 |
|
James Betker
|
f0c4cd6317
|
Taking another stab at a BPE tokenizer
|
2021-12-30 13:41:24 -07:00 |
|
James Betker
|
f2cd6a7f08
|
For loading conditional clips, default to falling back to loading the clip itself
|
2021-12-30 09:10:14 -07:00 |
|
James Betker
|
51ce1b5007
|
Add conditioning clips features to grand_conjoined
|
2021-12-29 14:44:32 -07:00 |
|
James Betker
|
c6ef0eef0b
|
asdf
|
2021-12-29 10:07:39 -07:00 |
|
James Betker
|
53784ec806
|
grand conjoined dataset: support collating
|
2021-12-29 09:44:37 -07:00 |
|
James Betker
|
07c2b9907c
|
Add voice2voice clip model
|
2021-12-28 16:18:12 -07:00 |
|
James Betker
|
746392f35c
|
Fix DS
|
2021-12-25 15:28:59 -07:00 |
|
James Betker
|
736c2626ee
|
build in character tokenizer
|
2021-12-25 15:21:01 -07:00 |
|
James Betker
|
52410fd9d9
|
256-bpe tokenizer
|
2021-12-25 08:52:08 -07:00 |
|
James Betker
|
ead2a74bf0
|
Add debug_failures flag
|
2021-12-23 16:12:16 -07:00 |
|
James Betker
|
9677f7084c
|
dataset mod
|
2021-12-23 15:21:30 -07:00 |
|
James Betker
|
8b19c37409
|
UnifiedGptVoice!
|
2021-12-23 15:20:26 -07:00 |
|
James Betker
|
5bc9772cb0
|
grand: support validation mode
|
2021-12-23 15:03:20 -07:00 |
|
James Betker
|
e55d949855
|
GrandConjoinedDataset
|
2021-12-23 14:32:33 -07:00 |
|
James Betker
|
b9de8a8eda
|
More fixes
|
2021-12-22 19:21:29 -07:00 |
|
James Betker
|
191e0130ee
|
Another fix
|
2021-12-22 18:30:50 -07:00 |
|
James Betker
|
6c6daa5795
|
Build a bigger, better tokenizer
|
2021-12-22 17:46:18 -07:00 |
|
James Betker
|
c737632eae
|
Train and use a bespoke tokenizer
|
2021-12-22 15:06:14 -07:00 |
|
James Betker
|
a9629f7022
|
Try out using the GPT tokenizer rather than nv_tacotron
This results in a significant compression of the text domain, I'm curious what the
effect on speech quality will be.
|
2021-12-22 14:03:18 -07:00 |
|
James Betker
|
ced81a760b
|
restore nv_tacotron
|
2021-12-22 13:48:53 -07:00 |
|
James Betker
|
7bf4f9f580
|
duplicate nvtacotron
|
2021-12-22 13:48:30 -07:00 |
|
James Betker
|
9e8a9bf6ca
|
Various fixes to gpt_tts_hf
|
2021-12-16 23:28:44 -07:00 |
|
James Betker
|
31fc693a8a
|
dafsdf
|
2021-12-02 22:55:36 -07:00 |
|
James Betker
|
040d998922
|
maasd
|
2021-12-02 22:53:48 -07:00 |
|
James Betker
|
cc10e7e7e8
|
Add tsv loader
|
2021-12-02 22:43:07 -07:00 |
|
James Betker
|
702607556d
|
nv_tacotron_dataset: allow it to load conditioning signals
|
2021-12-02 22:14:44 -07:00 |
|
James Betker
|
0604060580
|
Finish up mods for next version of GptAsrHf
|
2021-11-20 21:33:49 -07:00 |
|
James Betker
|
9b3c3b1227
|
use sets instead of list ops
|
2021-11-07 20:45:57 -07:00 |
|
James Betker
|
722d3dbdc2
|
f
|
2021-11-07 18:52:05 -07:00 |
|
James Betker
|
18b1de9b2c
|
Add exclusion_lists to unsupervised_audio_dataset
|
2021-11-07 18:46:47 -07:00 |
|
James Betker
|
fd14746bf8
|
badtimes
|
2021-11-03 00:33:38 -06:00 |
|
James Betker
|
2fa80486de
|
tacotron_dataset: recover gracefully
|
2021-11-03 00:31:50 -06:00 |
|
James Betker
|
af51d00dee
|
Load wav files from voxpopuli instead of oggs
|
2021-11-02 09:32:26 -06:00 |
|
James Betker
|
f7d0901ce6
|
Decouple MEL from nv_tacotron_dataset
|
2021-10-31 15:01:38 -06:00 |
|
James Betker
|
b8b268b5f6
|
Misc
|
2021-10-31 14:29:23 -06:00 |
|
James Betker
|
579f0a70ee
|
Move UnsupervisedAudioDataset to use my new mp3 loader
|
2021-10-28 22:33:12 -06:00 |
|
James Betker
|
5d714bc566
|
Add deepspeech model and support for decoding with it
|
2021-10-27 13:09:46 -06:00 |
|
James Betker
|
21b6daa0ed
|
Introduce clip resampling
|
2021-10-26 10:42:23 -06:00 |
|
James Betker
|
c3421b7f6d
|
Dataset work for audio quality processor
|
2021-10-24 09:09:34 -06:00 |
|
James Betker
|
06ea6191a9
|
Initial implementation of audio_with_noise dataset
|
2021-10-21 16:45:19 -06:00 |
|
James Betker
|
d016a2fbad
|
Go back to vanilla flavor of diffusion
|
2021-10-17 17:32:46 -06:00 |
|
James Betker
|
6833048bf7
|
Alterations to diffusion_dvae so it can be used directly on spectrograms
|
2021-09-23 15:56:25 -06:00 |
|
James Betker
|
359e9e27a7
|
unsupervised_audio_dataset: try to recover from failures of audio2numpy
|
2021-09-17 15:25:57 -06:00 |
|
James Betker
|
f78ce9d924
|
Get diffusion_dvae ready for prime time!
|
2021-09-16 22:43:10 -06:00 |
|
James Betker
|
1197ae1928
|
Misc
|
2021-09-16 10:53:56 -06:00 |
|
James Betker
|
8d9857f33d
|
More fixes
|
2021-09-14 20:45:05 -06:00 |
|
James Betker
|
9a9c90660f
|
Fixes
|
2021-09-14 18:29:17 -06:00 |
|
James Betker
|
e513052fca
|
Add unsupervised_audio_dataset
|
2021-09-14 17:43:16 -06:00 |
|
James Betker
|
b8f2e0f452
|
mydvae
|
2021-09-06 17:45:30 -06:00 |
|
James Betker
|
30cd33fe44
|
another fix
|
2021-08-31 14:46:46 -06:00 |
|
James Betker
|
8810d3de97
|
fix wavfile_dataset
|
2021-08-31 14:45:29 -06:00 |
|
James Betker
|
dabd87246d
|
Add unet_diffusion_vocoder
|
2021-08-31 14:38:33 -06:00 |
|
James Betker
|
570ed327ed
|
Stop dataset - attempt #2
|
2021-08-18 18:29:38 -06:00 |
|
James Betker
|
8332923f5c
|
Two more tools to test the audio segmentor
|
2021-08-17 09:09:11 -06:00 |
|
James Betker
|
93e903af15
|
Rework wavfile dataset to be usable for things other than augments
|
2021-08-16 22:52:35 -06:00 |
|
James Betker
|
d7f30232c3
|
Oh yeah
|
2021-08-16 22:52:15 -06:00 |
|
James Betker
|
4c01d82265
|
Fix for voxpopuli
|
2021-08-16 22:52:05 -06:00 |
|
James Betker
|
1fede41b7b
|
Audio segmentor
|
2021-08-16 22:51:53 -06:00 |
|
James Betker
|
2d3372054d
|
Add support for voxpopuli to nv_tacotron_dataset
|
2021-08-16 17:13:40 -06:00 |
|
James Betker
|
3580c52eac
|
Fix up wavfile_dataset to be able to provide a full clip
|
2021-08-15 20:53:26 -06:00 |
|
James Betker
|
a523c4f932
|
Auto-normalize wav files by data type
|
2021-08-15 09:09:51 -06:00 |
|
James Betker
|
c28f657ab8
|
Allow usage of pre-rendered mels saved to npy files
|
2021-08-14 23:38:15 -06:00 |
|
James Betker
|
ad3391bd96
|
Fix nan issue when interpolating audio
|
2021-08-14 20:42:01 -06:00 |
|
James Betker
|
769f0acc53
|
Moar fix
|
2021-08-14 17:23:15 -06:00 |
|
James Betker
|
3d2e724083
|
Fix audio ranging problem
|
2021-08-14 17:18:55 -06:00 |
|
James Betker
|
d6a73acaed
|
Allow processing of multiple audio sources at once from nv_tacotron_dataset
|
2021-08-14 16:04:05 -06:00 |
|
James Betker
|
007976082b
|
GPT_asr for inference
|
2021-08-14 14:37:17 -06:00 |
|
James Betker
|
72622b4d61
|
Allow saving mel strips as files from the dataset implementation
|
2021-08-13 22:46:41 -06:00 |
|
James Betker
|
cfd284f425
|
Fix up some stuff that allows the MEL to be computed on-GPU
|
2021-08-13 18:35:55 -06:00 |
|
James Betker
|
fff1a59e08
|
max/min mel invalid fix
|
2021-08-13 09:36:31 -06:00 |
|
James Betker
|
4b2946e581
|
More fix
|
2021-08-12 15:51:23 -06:00 |
|
James Betker
|
4c76257c71
|
Dont require collation for nv_tacotron
|
2021-08-12 15:44:55 -06:00 |
|
James Betker
|
5b07d3b623
|
Found error that I was trying to fix with reload=True
|
2021-08-12 15:22:34 -06:00 |
|
James Betker
|
430b650a34
|
......
|
2021-08-12 10:31:10 -06:00 |
|
James Betker
|
b35d6ae028
|
Print some metrics from tacotron dataset when it croaks
|
2021-08-12 09:21:12 -06:00 |
|
James Betker
|
0c4d6b1916
|
Just offer generic re-load for nv-tacotron
|
2021-08-12 09:09:12 -06:00 |
|
James Betker
|
154f5aa73c
|
Fix annoying warning and add to requirements
|
2021-08-11 17:32:06 -06:00 |
|
James Betker
|
f04a7bdf63
|
Bug fixes for tacotron dataset on mozilla cv
- Support a max mel length (mozilla cv has some tracks that are basically unbounded..)
- Don't fail on low sample rates (mozilla cv has some of those)
|
2021-08-11 16:17:03 -06:00 |
|
James Betker
|
2d3f0cc33c
|
nv_tacotron_dataset - Allow training on mozilla cv
|
2021-08-11 13:34:31 -06:00 |
|
James Betker
|
d0c74278bf
|
Enable multiple wavfile paths to be specified, fix eps bug in mp3 splitter
|
2021-08-11 08:46:02 -06:00 |
|
James Betker
|
e19c00398e
|
More improvements to random_mp3_splitter
|
2021-08-09 21:31:12 -06:00 |
|
James Betker
|
74342b860b
|
Revert "Undo forced text padding"
This reverts commit 83ab5e6a00 .
|
2021-08-09 11:56:34 -06:00 |
|
James Betker
|
d4e33bf15f
|
Fixes to the mp3 splitter
|
2021-08-09 11:55:46 -06:00 |
|
James Betker
|
4100469902
|
Add a tool to split mp3 files into arbitrary chunks of wav files
|
2021-08-08 23:23:13 -06:00 |
|
James Betker
|
83ab5e6a00
|
Undo forced text padding
|
2021-08-08 11:42:20 -06:00 |
|
James Betker
|
690d7e86d3
|
Fix nv_tacotron_dataset bug which incorrectly mapped filenames
dammit..
|
2021-08-08 11:38:52 -06:00 |
|
James Betker
|
a2afb25e42
|
Fix inference, always flow full text tokens through transformer
|
2021-08-07 20:11:10 -06:00 |
|
James Betker
|
b43683b772
|
Add lucidrains_dvae
|
2021-08-06 12:03:46 -06:00 |
|