James Betker
|
3a9d1c53ea
|
Rework conditioning inputs provided
|
2021-10-26 10:46:33 -06:00 |
|
James Betker
|
43e389aac6
|
Add time_embed_dim_multiplier
|
2021-10-26 08:55:55 -06:00 |
|
James Betker
|
ba6e46c02a
|
Further simplify diffusion_vocoder and make noise_surfer work
|
2021-10-26 08:54:30 -06:00 |
|
James Betker
|
0ee1c67ce5
|
Rework how conditioning inputs are applied to DiffusionVocoder
|
2021-10-24 09:08:58 -06:00 |
|
James Betker
|
0dee15f875
|
base DVAE & vector_quantizer
|
2021-10-20 21:19:38 -06:00 |
|
James Betker
|
a6f0f854b9
|
Fix codes when inferring from dvae
|
2021-10-17 22:51:17 -06:00 |
|
James Betker
|
d016a2fbad
|
Go back to vanilla flavor of diffusion
|
2021-10-17 17:32:46 -06:00 |
|
James Betker
|
23da073037
|
Norm decoder outputs now
|
2021-10-16 09:07:10 -06:00 |
|
James Betker
|
0edc98f6c4
|
Throw out the idea of conditioning on discrete codes. Oh well :(
|
2021-10-16 09:02:01 -06:00 |
|
James Betker
|
62c8c5d93e
|
Zero out spectrogram code inputs initially.
|
2021-10-15 12:10:11 -06:00 |
|
James Betker
|
1d0b44ebc2
|
More tweaks to diffusion-vocoder
|
2021-10-15 11:51:17 -06:00 |
|
James Betker
|
3b19581f9a
|
Allow num_resblocks to specified per-level
|
2021-10-14 11:26:04 -06:00 |
|
James Betker
|
83798887a8
|
Mods to support unet diffusion vocoder with conditioning
|
2021-10-13 21:23:18 -06:00 |
|
James Betker
|
33120cb35c
|
Add norming to discretization_loss
|
2021-10-06 17:10:50 -06:00 |
|
James Betker
|
f2977d360c
|
Allow attention_dim in channel attention to be specified, add converter
|
2021-10-05 17:29:38 -06:00 |
|
James Betker
|
9c0d7288ea
|
Discretization loss attempt
|
2021-10-04 20:59:21 -06:00 |
|
James Betker
|
66f99a159c
|
Rev2
|
2021-10-03 15:20:50 -06:00 |
|
James Betker
|
09f373e3b1
|
Add dvae with channel attention
|
2021-10-03 10:52:01 -06:00 |
|
James Betker
|
0396a9d2ca
|
Increase baseline codes recording across all dvae models
|
2021-09-30 08:09:07 -06:00 |
|
James Betker
|
6e550edfe3
|
Attentive dvae
|
2021-09-29 14:17:29 -06:00 |
|
James Betker
|
6833048bf7
|
Alterations to diffusion_dvae so it can be used directly on spectrograms
|
2021-09-23 15:56:25 -06:00 |
|
James Betker
|
a6544f1684
|
More checkpointing fixes
|
2021-09-16 23:12:43 -06:00 |
|
James Betker
|
f78ce9d924
|
Get diffusion_dvae ready for prime time!
|
2021-09-16 22:43:10 -06:00 |
|
James Betker
|
0382660159
|
Get diffusion_dvae functional
|
2021-09-14 17:43:31 -06:00 |
|
James Betker
|
b8f2e0f452
|
mydvae
|
2021-09-06 17:45:30 -06:00 |
|
James Betker
|
dabd87246d
|
Add unet_diffusion_vocoder
|
2021-08-31 14:38:33 -06:00 |
|
James Betker
|
909754cc27
|
Add find_faulty_files.py
|
2021-08-25 18:00:43 -06:00 |
|
James Betker
|
08b33c8e3a
|
Support silu activation
|
2021-08-25 09:03:14 -06:00 |
|
James Betker
|
67bf7f5219
|
dvae mods
Trying to squeeze as much performance out of this net as possible
|
2021-08-25 08:55:13 -06:00 |
|
James Betker
|
b521d94b01
|
Make gpt-asr more configurable
|
2021-08-19 16:33:41 -06:00 |
|
James Betker
|
570ed327ed
|
Stop dataset - attempt #2
|
2021-08-18 18:29:38 -06:00 |
|
James Betker
|
17453ccbe8
|
Revert mods to lrdvae
They didn't really change anything
|
2021-08-17 09:09:29 -06:00 |
|
James Betker
|
8332923f5c
|
Two more tools to test the audio segmentor
|
2021-08-17 09:09:11 -06:00 |
|
James Betker
|
1fede41b7b
|
Audio segmentor
|
2021-08-16 22:51:53 -06:00 |
|
James Betker
|
729c1fd5a9
|
Fix up max lengths to save memory
|
2021-08-15 21:29:28 -06:00 |
|
James Betker
|
9e47e64d5a
|
Add gpt_segmentor model
The idea is to specifically train a model that extracts phrases from
audio clips.
|
2021-08-15 21:23:07 -06:00 |
|
James Betker
|
a826d5f658
|
Mods to dvae
- Add resblock to each layer
- Increase filter size for each layer
- Use SiLU
|
2021-08-15 20:54:10 -06:00 |
|
James Betker
|
b8bec22f1a
|
Fix gpt_asr inference bug
|
2021-08-15 20:53:42 -06:00 |
|
James Betker
|
98057b6516
|
Make lrdvae use quantized mode in eval()
|
2021-08-14 23:43:01 -06:00 |
|
James Betker
|
007976082b
|
GPT_asr for inference
|
2021-08-14 14:37:17 -06:00 |
|
James Betker
|
e1bdd3f7c7
|
Fix gpt_asr bug. Initial implementation of beam search
|
2021-08-13 22:47:00 -06:00 |
|
James Betker
|
cdee31c60b
|
GPT_ASR
|
2021-08-13 15:02:18 -06:00 |
|
James Betker
|
20586a8edc
|
Fix LRDVAE bug with quantizer integration
|
2021-08-11 16:17:22 -06:00 |
|
James Betker
|
82fc69abfa
|
Add "pure" evaluator
Which simply computes the training loss against an eval dataset
|
2021-08-09 14:58:35 -06:00 |
|
James Betker
|
080bea2f19
|
No, really
|
2021-08-09 12:02:31 -06:00 |
|
James Betker
|
e1ce4671e4
|
Apply dropout to gpt_tts, get rid of min_gpt implementation
|
2021-08-09 12:01:10 -06:00 |
|
James Betker
|
1068f53b78
|
Add a sampling beam search
|
2021-08-09 11:56:06 -06:00 |
|
James Betker
|
01cfae28d8
|
Beam search implementation in one pass? Dayyyum
|
2021-08-08 23:22:42 -06:00 |
|
James Betker
|
690d7e86d3
|
Fix nv_tacotron_dataset bug which incorrectly mapped filenames
dammit..
|
2021-08-08 11:38:52 -06:00 |
|
James Betker
|
a2afb25e42
|
Fix inference, always flow full text tokens through transformer
|
2021-08-07 20:11:10 -06:00 |
|