James Betker
|
0822792d79
|
Fix options.py bug
|
2021-10-29 14:47:31 -06:00 |
|
James Betker
|
928e7026c2
|
Mod STFT injector to be specifiable
|
2021-10-28 22:34:12 -06:00 |
|
James Betker
|
579f0a70ee
|
Move UnsupervisedAudioDataset to use my new mp3 loader
|
2021-10-28 22:33:12 -06:00 |
|
James Betker
|
2afea126d7
|
mod trainer to be very explicit about the fact that loading models and state together dont work, but allow it
|
2021-10-28 22:32:42 -06:00 |
|
James Betker
|
bb0a0c8264
|
classify_into_folders script
|
2021-10-27 14:56:16 -06:00 |
|
James Betker
|
d91dcbd404
|
Make classifier inference script more open
|
2021-10-27 13:18:54 -06:00 |
|
James Betker
|
58494b0888
|
Add support for distilling gpt_asr
|
2021-10-27 13:10:07 -06:00 |
|
James Betker
|
5d714bc566
|
Add deepspeech model and support for decoding with it
|
2021-10-27 13:09:46 -06:00 |
|
James Betker
|
15437b2fc3
|
WER script
|
2021-10-26 13:30:29 -06:00 |
|
James Betker
|
3a9d1c53ea
|
Rework conditioning inputs provided
|
2021-10-26 10:46:33 -06:00 |
|
James Betker
|
21b6daa0ed
|
Introduce clip resampling
|
2021-10-26 10:42:23 -06:00 |
|
James Betker
|
43e389aac6
|
Add time_embed_dim_multiplier
|
2021-10-26 08:55:55 -06:00 |
|
James Betker
|
ba6e46c02a
|
Further simplify diffusion_vocoder and make noise_surfer work
|
2021-10-26 08:54:30 -06:00 |
|
James Betker
|
c3421b7f6d
|
Dataset work for audio quality processor
|
2021-10-24 09:09:34 -06:00 |
|
James Betker
|
0ee1c67ce5
|
Rework how conditioning inputs are applied to DiffusionVocoder
|
2021-10-24 09:08:58 -06:00 |
|
James Betker
|
b1248e7114
|
Get rid of filter_urbansounds
|
2021-10-21 16:46:04 -06:00 |
|
James Betker
|
06ea6191a9
|
Initial implementation of audio_with_noise dataset
|
2021-10-21 16:45:19 -06:00 |
|
James Betker
|
9a3e89ec53
|
Force LR fix
|
2021-10-21 12:01:01 -06:00 |
|
James Betker
|
40cb25292a
|
Fix force_lr logic
|
2021-10-21 11:51:30 -06:00 |
|
James Betker
|
0dee15f875
|
base DVAE & vector_quantizer
|
2021-10-20 21:19:38 -06:00 |
|
James Betker
|
f2a31702b5
|
Clean stuff up, move more things into arch_util
|
2021-10-20 21:19:25 -06:00 |
|
James Betker
|
a6f0f854b9
|
Fix codes when inferring from dvae
|
2021-10-17 22:51:17 -06:00 |
|
James Betker
|
d016a2fbad
|
Go back to vanilla flavor of diffusion
|
2021-10-17 17:32:46 -06:00 |
|
James Betker
|
23da073037
|
Norm decoder outputs now
|
2021-10-16 09:07:10 -06:00 |
|
James Betker
|
0edc98f6c4
|
Throw out the idea of conditioning on discrete codes. Oh well :(
|
2021-10-16 09:02:01 -06:00 |
|
James Betker
|
62c8c5d93e
|
Zero out spectrogram code inputs initially.
|
2021-10-15 12:10:11 -06:00 |
|
James Betker
|
1d0b44ebc2
|
More tweaks to diffusion-vocoder
|
2021-10-15 11:51:17 -06:00 |
|
James Betker
|
3b19581f9a
|
Allow num_resblocks to specified per-level
|
2021-10-14 11:26:04 -06:00 |
|
James Betker
|
83798887a8
|
Mods to support unet diffusion vocoder with conditioning
|
2021-10-13 21:23:18 -06:00 |
|
James Betker
|
c861054218
|
Restore spleeter_splitter
The mods don't help - in TF mode, everything is done on the GPU anyways. Something else
is going to have to be done to fix this.
|
2021-10-09 23:55:42 -06:00 |
|
James Betker
|
32ba496632
|
More fixes
|
2021-10-09 23:27:14 -06:00 |
|
James Betker
|
932ea29a83
|
Add multiprocessing to the spleeter splitter script to try and improve performance further
|
2021-10-09 23:15:36 -06:00 |
|
James Betker
|
b94e587f46
|
Improvements to spleeter_filter_noisy_clips
|
2021-10-07 21:28:00 -06:00 |
|
James Betker
|
33120cb35c
|
Add norming to discretization_loss
|
2021-10-06 17:10:50 -06:00 |
|
James Betker
|
bb891a3a53
|
Add partitioning and improved resuming to the spleeter filtering
|
2021-10-06 17:10:12 -06:00 |
|
James Betker
|
f2977d360c
|
Allow attention_dim in channel attention to be specified, add converter
|
2021-10-05 17:29:38 -06:00 |
|
James Betker
|
9c0d7288ea
|
Discretization loss attempt
|
2021-10-04 20:59:21 -06:00 |
|
James Betker
|
66f99a159c
|
Rev2
|
2021-10-03 15:20:50 -06:00 |
|
James Betker
|
09f373e3b1
|
Add dvae with channel attention
|
2021-10-03 10:52:01 -06:00 |
|
James Betker
|
0396a9d2ca
|
Increase baseline codes recording across all dvae models
|
2021-09-30 08:09:07 -06:00 |
|
James Betker
|
f84ccbdfb2
|
Fix quantizer with balancing_heuristic
|
2021-09-29 14:46:05 -06:00 |
|
James Betker
|
4914c526dc
|
More cleanup
|
2021-09-29 14:24:49 -06:00 |
|
James Betker
|
6e550edfe3
|
Attentive dvae
|
2021-09-29 14:17:29 -06:00 |
|
James Betker
|
fc8ae4679a
|
Work on spleeter filtering script
|
2021-09-29 09:24:56 -06:00 |
|
James Betker
|
55b58fb67f
|
Clean up codebase
Remove stuff that I'm likely not going to use again (or generally failed experiments)
|
2021-09-29 09:21:44 -06:00 |
|
James Betker
|
4d1a42e944
|
Add switchnorm to gumbel_quantizer
|
2021-09-24 18:49:25 -06:00 |
|
James Betker
|
ac57cdc794
|
Add scheduling to quantizer, enable cudnn_benchmarking to be disabled
|
2021-09-24 17:01:36 -06:00 |
|
James Betker
|
3e64e847c2
|
Gumbel quantizer
|
2021-09-23 23:32:03 -06:00 |
|
James Betker
|
c5297ccec6
|
Add dvae balancing heuristic
|
2021-09-23 21:19:36 -06:00 |
|
James Betker
|
e24c619387
|
Fix
|
2021-09-23 16:07:58 -06:00 |
|
James Betker
|
6833048bf7
|
Alterations to diffusion_dvae so it can be used directly on spectrograms
|
2021-09-23 15:56:25 -06:00 |
|
James Betker
|
97ea329a59
|
Make spleeter filter simpler (and hopefully much faster)
|
2021-09-17 15:29:42 -06:00 |
|
James Betker
|
359e9e27a7
|
unsupervised_audio_dataset: try to recover from failures of audio2numpy
|
2021-09-17 15:25:57 -06:00 |
|
James Betker
|
5c8d266d4f
|
chk
|
2021-09-17 09:15:36 -06:00 |
|
James Betker
|
a6544f1684
|
More checkpointing fixes
|
2021-09-16 23:12:43 -06:00 |
|
James Betker
|
94899d88f3
|
Fix overuse of checkpointing
|
2021-09-16 23:00:28 -06:00 |
|
James Betker
|
f78ce9d924
|
Get diffusion_dvae ready for prime time!
|
2021-09-16 22:43:10 -06:00 |
|
James Betker
|
1197ae1928
|
Misc
|
2021-09-16 10:53:56 -06:00 |
|
James Betker
|
6f48674647
|
Support diffusion models with extra return values & inference in diffusion_dvae
|
2021-09-16 10:53:46 -06:00 |
|
James Betker
|
8d9857f33d
|
More fixes
|
2021-09-14 20:45:05 -06:00 |
|
James Betker
|
9a9c90660f
|
Fixes
|
2021-09-14 18:29:17 -06:00 |
|
James Betker
|
4334a67924
|
Spleeter mods
|
2021-09-14 17:43:40 -06:00 |
|
James Betker
|
0382660159
|
Get diffusion_dvae functional
|
2021-09-14 17:43:31 -06:00 |
|
James Betker
|
e513052fca
|
Add unsupervised_audio_dataset
|
2021-09-14 17:43:16 -06:00 |
|
James Betker
|
bc603c3231
|
Script adjustments and fixes
|
2021-09-12 21:26:45 -06:00 |
|
James Betker
|
76e2c497f7
|
Improvements to splitter
|
2021-09-09 23:34:56 -06:00 |
|
James Betker
|
742f9b4010
|
Batch spleeter cleaner using GPU
|
2021-09-09 23:14:32 -06:00 |
|
James Betker
|
73b930c0f6
|
Add diffusion_dvae
Increase split_on_silence interval
|
2021-09-09 16:22:05 -06:00 |
|
James Betker
|
b8f2e0f452
|
mydvae
|
2021-09-06 17:45:30 -06:00 |
|
James Betker
|
92e7e57f81
|
Update diffusion_noise_surfer to support audio
|
2021-09-01 08:34:47 -06:00 |
|
James Betker
|
3e073cff85
|
Set kernel_size in diffusion_vocoder
|
2021-09-01 08:33:46 -06:00 |
|
James Betker
|
30cd33fe44
|
another fix
|
2021-08-31 14:46:46 -06:00 |
|
James Betker
|
8810d3de97
|
fix wavfile_dataset
|
2021-08-31 14:45:29 -06:00 |
|
James Betker
|
dabd87246d
|
Add unet_diffusion_vocoder
|
2021-08-31 14:38:33 -06:00 |
|
James Betker
|
274d352e6f
|
dug
|
2021-08-30 21:45:58 -06:00 |
|
James Betker
|
f1a0c21fb2
|
asr_eval
|
2021-08-30 21:41:34 -06:00 |
|
James Betker
|
ed6eae407f
|
More scripts for splitting and formatting audio
|
2021-08-30 21:20:52 -06:00 |
|
James Betker
|
909754cc27
|
Add find_faulty_files.py
|
2021-08-25 18:00:43 -06:00 |
|
James Betker
|
08b33c8e3a
|
Support silu activation
|
2021-08-25 09:03:14 -06:00 |
|
James Betker
|
67bf7f5219
|
dvae mods
Trying to squeeze as much performance out of this net as possible
|
2021-08-25 08:55:13 -06:00 |
|
James Betker
|
d05cc1f46c
|
Misc
|
2021-08-24 17:12:04 -06:00 |
|
James Betker
|
9dfe936c16
|
Fix ddp for sampler
|
2021-08-19 16:45:34 -06:00 |
|
James Betker
|
b521d94b01
|
Make gpt-asr more configurable
|
2021-08-19 16:33:41 -06:00 |
|
James Betker
|
570ed327ed
|
Stop dataset - attempt #2
|
2021-08-18 18:29:38 -06:00 |
|
James Betker
|
17453ccbe8
|
Revert mods to lrdvae
They didn't really change anything
|
2021-08-17 09:09:29 -06:00 |
|
James Betker
|
8332923f5c
|
Two more tools to test the audio segmentor
|
2021-08-17 09:09:11 -06:00 |
|
James Betker
|
7c086d0c2c
|
libritts - only write on successful check
|
2021-08-16 22:52:55 -06:00 |
|
James Betker
|
93e903af15
|
Rework wavfile dataset to be usable for things other than augments
|
2021-08-16 22:52:35 -06:00 |
|
James Betker
|
d7f30232c3
|
Oh yeah
|
2021-08-16 22:52:15 -06:00 |
|
James Betker
|
4c01d82265
|
Fix for voxpopuli
|
2021-08-16 22:52:05 -06:00 |
|
James Betker
|
1fede41b7b
|
Audio segmentor
|
2021-08-16 22:51:53 -06:00 |
|
James Betker
|
2d3372054d
|
Add support for voxpopuli to nv_tacotron_dataset
|
2021-08-16 17:13:40 -06:00 |
|
James Betker
|
729c1fd5a9
|
Fix up max lengths to save memory
|
2021-08-15 21:29:28 -06:00 |
|
James Betker
|
9e47e64d5a
|
Add gpt_segmentor model
The idea is to specifically train a model that extracts phrases from
audio clips.
|
2021-08-15 21:23:07 -06:00 |
|
James Betker
|
a826d5f658
|
Mods to dvae
- Add resblock to each layer
- Increase filter size for each layer
- Use SiLU
|
2021-08-15 20:54:10 -06:00 |
|
James Betker
|
b8bec22f1a
|
Fix gpt_asr inference bug
|
2021-08-15 20:53:42 -06:00 |
|
James Betker
|
3580c52eac
|
Fix up wavfile_dataset to be able to provide a full clip
|
2021-08-15 20:53:26 -06:00 |
|
James Betker
|
a523c4f932
|
Auto-normalize wav files by data type
|
2021-08-15 09:09:51 -06:00 |
|
James Betker
|
98057b6516
|
Make lrdvae use quantized mode in eval()
|
2021-08-14 23:43:01 -06:00 |
|
James Betker
|
c28f657ab8
|
Allow usage of pre-rendered mels saved to npy files
|
2021-08-14 23:38:15 -06:00 |
|
James Betker
|
ad3391bd96
|
Fix nan issue when interpolating audio
|
2021-08-14 20:42:01 -06:00 |
|
James Betker
|
769f0acc53
|
Moar fix
|
2021-08-14 17:23:15 -06:00 |
|
James Betker
|
3d2e724083
|
Fix audio ranging problem
|
2021-08-14 17:18:55 -06:00 |
|
James Betker
|
d6a73acaed
|
Allow processing of multiple audio sources at once from nv_tacotron_dataset
|
2021-08-14 16:04:05 -06:00 |
|
James Betker
|
007976082b
|
GPT_asr for inference
|
2021-08-14 14:37:17 -06:00 |
|
James Betker
|
e1bdd3f7c7
|
Fix gpt_asr bug. Initial implementation of beam search
|
2021-08-13 22:47:00 -06:00 |
|
James Betker
|
72622b4d61
|
Allow saving mel strips as files from the dataset implementation
|
2021-08-13 22:46:41 -06:00 |
|
James Betker
|
cfd284f425
|
Fix up some stuff that allows the MEL to be computed on-GPU
|
2021-08-13 18:35:55 -06:00 |
|
James Betker
|
cdee31c60b
|
GPT_ASR
|
2021-08-13 15:02:18 -06:00 |
|
James Betker
|
81e91c99de
|
Misc
|
2021-08-13 13:58:59 -06:00 |
|
James Betker
|
fff1a59e08
|
max/min mel invalid fix
|
2021-08-13 09:36:31 -06:00 |
|
James Betker
|
4b2946e581
|
More fix
|
2021-08-12 15:51:23 -06:00 |
|
James Betker
|
4c76257c71
|
Dont require collation for nv_tacotron
|
2021-08-12 15:44:55 -06:00 |
|
James Betker
|
5b07d3b623
|
Found error that I was trying to fix with reload=True
|
2021-08-12 15:22:34 -06:00 |
|
James Betker
|
430b650a34
|
......
|
2021-08-12 10:31:10 -06:00 |
|
James Betker
|
b35d6ae028
|
Print some metrics from tacotron dataset when it croaks
|
2021-08-12 09:21:12 -06:00 |
|
James Betker
|
0c4d6b1916
|
Just offer generic re-load for nv-tacotron
|
2021-08-12 09:09:12 -06:00 |
|
James Betker
|
154f5aa73c
|
Fix annoying warning and add to requirements
|
2021-08-11 17:32:06 -06:00 |
|
James Betker
|
f5a9b88ef6
|
tacotron cleaners: remove quotation marks
these don't really have relevance for tts or asr
|
2021-08-11 16:18:44 -06:00 |
|
James Betker
|
20586a8edc
|
Fix LRDVAE bug with quantizer integration
|
2021-08-11 16:17:22 -06:00 |
|
James Betker
|
f04a7bdf63
|
Bug fixes for tacotron dataset on mozilla cv
- Support a max mel length (mozilla cv has some tracks that are basically unbounded..)
- Don't fail on low sample rates (mozilla cv has some of those)
|
2021-08-11 16:17:03 -06:00 |
|
James Betker
|
2d3f0cc33c
|
nv_tacotron_dataset - Allow training on mozilla cv
|
2021-08-11 13:34:31 -06:00 |
|
James Betker
|
d0c74278bf
|
Enable multiple wavfile paths to be specified, fix eps bug in mp3 splitter
|
2021-08-11 08:46:02 -06:00 |
|
James Betker
|
e19c00398e
|
More improvements to random_mp3_splitter
|
2021-08-09 21:31:12 -06:00 |
|
James Betker
|
04d14b3acc
|
No batch factors for eval
|
2021-08-09 16:02:01 -06:00 |
|
James Betker
|
82fc69abfa
|
Add "pure" evaluator
Which simply computes the training loss against an eval dataset
|
2021-08-09 14:58:35 -06:00 |
|
James Betker
|
080bea2f19
|
No, really
|
2021-08-09 12:02:31 -06:00 |
|
James Betker
|
e1ce4671e4
|
Apply dropout to gpt_tts, get rid of min_gpt implementation
|
2021-08-09 12:01:10 -06:00 |
|
James Betker
|
74342b860b
|
Revert "Undo forced text padding"
This reverts commit 83ab5e6a00 .
|
2021-08-09 11:56:34 -06:00 |
|
James Betker
|
1068f53b78
|
Add a sampling beam search
|
2021-08-09 11:56:06 -06:00 |
|
James Betker
|
d4e33bf15f
|
Fixes to the mp3 splitter
|
2021-08-09 11:55:46 -06:00 |
|
James Betker
|
4100469902
|
Add a tool to split mp3 files into arbitrary chunks of wav files
|
2021-08-08 23:23:13 -06:00 |
|
James Betker
|
01cfae28d8
|
Beam search implementation in one pass? Dayyyum
|
2021-08-08 23:22:42 -06:00 |
|
James Betker
|
83ab5e6a00
|
Undo forced text padding
|
2021-08-08 11:42:20 -06:00 |
|
James Betker
|
690d7e86d3
|
Fix nv_tacotron_dataset bug which incorrectly mapped filenames
dammit..
|
2021-08-08 11:38:52 -06:00 |
|
James Betker
|
a2afb25e42
|
Fix inference, always flow full text tokens through transformer
|
2021-08-07 20:11:10 -06:00 |
|
James Betker
|
4c678172d6
|
ugh
|
2021-08-06 22:10:18 -06:00 |
|
James Betker
|
e723137273
|
Make gpttts more configurable
|
2021-08-06 22:08:51 -06:00 |
|
James Betker
|
a7496b661c
|
combined dvae ftw
|
2021-08-06 22:01:06 -06:00 |
|
James Betker
|
0237e96b34
|
Fix dvae bug
|
2021-08-06 14:17:01 -06:00 |
|
James Betker
|
0799d95af5
|
Use quantizer from rosinality/vqvae with openai dvae
|
2021-08-06 14:06:26 -06:00 |
|
James Betker
|
d3ace153af
|
Add logic for performing inference using gpt_tts with dual-encoder modes
|
2021-08-06 12:04:12 -06:00 |
|
James Betker
|
b43683b772
|
Add lucidrains_dvae
|
2021-08-06 12:03:46 -06:00 |
|
James Betker
|
62c7570512
|
Constrain wav_aug a bit more
|
2021-08-06 08:19:38 -06:00 |
|
James Betker
|
f126040da2
|
Undo noise first
|
2021-08-05 23:24:38 -06:00 |
|
James Betker
|
908ef5495f
|
Add noise first to audio_aug
|
2021-08-05 23:22:44 -06:00 |
|
James Betker
|
d6007c6de1
|
dataset fixes
|
2021-08-05 23:12:59 -06:00 |
|
James Betker
|
3ca51e80b2
|
Only fix weird path bug in windows
|
2021-08-05 22:21:25 -06:00 |
|
James Betker
|
70dcd1107f
|
Fix byol_model_wrapper to function with audio inputs
|
2021-08-05 22:20:22 -06:00 |
|
James Betker
|
f86df53ce0
|
Export extract_byol_model as a function
|
2021-08-05 22:15:26 -06:00 |
|