James Betker
|
1609101a42
|
musical gap filler
|
2022-05-05 16:47:08 -06:00 |
|
James Betker
|
d66ab2d28c
|
Remove unused waveform_gens
|
2022-05-04 21:06:54 -06:00 |
|
James Betker
|
47662b9ec5
|
some random crap
|
2022-05-04 20:29:23 -06:00 |
|
James Betker
|
6655f7845a
|
add pixel shuffling for 1d cases
|
2022-05-04 08:03:09 -06:00 |
|
James Betker
|
c42c53e75a
|
Add a trainable network for converting a normal distribution into a latent space
|
2022-05-02 09:47:30 -06:00 |
|
James Betker
|
e402089556
|
abstractify
|
2022-05-02 00:11:26 -06:00 |
|
James Betker
|
ab219fbefb
|
output variance
|
2022-05-02 00:10:33 -06:00 |
|
James Betker
|
3b074aac34
|
add checkpointing
|
2022-05-02 00:07:42 -06:00 |
|
James Betker
|
ae5f934ea1
|
diffwave
|
2022-05-02 00:05:04 -06:00 |
|
James Betker
|
f4254609c1
|
MDF
around and around in circles........
|
2022-05-01 23:04:56 -06:00 |
|
James Betker
|
b712d3b72b
|
break out get_conditioning_latent from unified_voice
|
2022-05-01 23:04:44 -06:00 |
|
James Betker
|
afa2df57c9
|
gen3
|
2022-04-30 10:41:38 -06:00 |
|
James Betker
|
64c7582bf5
|
full pipeline
|
2022-04-28 22:47:26 -06:00 |
|
James Betker
|
8aa6651fc7
|
fix surrogate loss return in waveform_gen2
|
2022-04-28 10:10:11 -06:00 |
|
James Betker
|
e208d9fb80
|
gate augmentations with a flag
|
2022-04-28 10:09:22 -06:00 |
|
James Betker
|
3f67cb2023
|
music diffusion fid adjustments
|
2022-04-28 10:08:55 -06:00 |
|
James Betker
|
ab8176b217
|
audio prep misc
|
2022-04-28 10:08:38 -06:00 |
|
James Betker
|
f02b01bd9d
|
reverse univnet classifier
|
2022-04-20 21:37:55 -06:00 |
|
James Betker
|
9df85c902e
|
New gen2
Which is basically a autoencoder with a giant diffusion appendage attached
|
2022-04-20 21:37:34 -06:00 |
|
James Betker
|
b1c2c48720
|
music diffusion fid
|
2022-04-20 00:28:03 -06:00 |
|
James Betker
|
084b1c1527
|
file splitter
|
2022-04-20 00:27:49 -06:00 |
|
James Betker
|
b4549eed9f
|
uv2 fix
|
2022-04-20 00:27:38 -06:00 |
|
James Betker
|
24fdafd855
|
fix2
|
2022-04-20 00:03:29 -06:00 |
|
James Betker
|
0af0051399
|
fix
|
2022-04-20 00:01:57 -06:00 |
|
James Betker
|
419f4d37bd
|
gen2 music
|
2022-04-19 23:38:37 -06:00 |
|
James Betker
|
c85ab738c5
|
paired fix
|
2022-04-16 23:41:57 -06:00 |
|
James Betker
|
8fe0dff33c
|
support tts typing
|
2022-04-16 23:36:57 -06:00 |
|
James Betker
|
48cb6a5abd
|
misc
|
2022-04-16 20:28:04 -06:00 |
|
James Betker
|
147478a148
|
cvvp
|
2022-04-16 20:27:46 -06:00 |
|
James Betker
|
546ecd5aeb
|
music!
|
2022-04-15 21:21:37 -06:00 |
|
James Betker
|
254357724d
|
gradprop
|
2022-04-15 09:37:20 -06:00 |
|
James Betker
|
fbf1f4f637
|
update
|
2022-04-15 09:34:44 -06:00 |
|
James Betker
|
82aad335ba
|
add distributued logic for loss
|
2022-04-15 09:31:48 -06:00 |
|
James Betker
|
efe12cb816
|
Update clvp to add masking probabilities in conditioning and to support code inputs
|
2022-04-15 09:11:23 -06:00 |
|
James Betker
|
3cad1b8114
|
more fixes
|
2022-04-11 15:18:44 -06:00 |
|
James Betker
|
6dea7da7a8
|
another fix
|
2022-04-11 12:29:43 -06:00 |
|
James Betker
|
f2c172291f
|
fix audio_diffusion_fid for autoregressive latent inputs
|
2022-04-11 12:08:15 -06:00 |
|
James Betker
|
8ea5c307fb
|
Fixes for training the diffusion model on autoregressive inputs
|
2022-04-11 11:02:44 -06:00 |
|
James Betker
|
a3622462c1
|
Change latent_conditioner back
|
2022-04-11 09:00:13 -06:00 |
|
James Betker
|
03d0b90bda
|
fixes
|
2022-04-10 21:02:12 -06:00 |
|
James Betker
|
19ca5b26c1
|
Remove flat0 and move it into flat
|
2022-04-10 21:01:59 -06:00 |
|
James Betker
|
81c952a00a
|
undo relative
|
2022-04-08 16:32:52 -06:00 |
|
James Betker
|
944b4c3335
|
more undos
|
2022-04-08 16:31:08 -06:00 |
|
James Betker
|
032983e2ed
|
fix bug and allow position encodings to be trained separately from the rest of the model
|
2022-04-08 16:26:01 -06:00 |
|
James Betker
|
09ab1aa9bc
|
revert rotary embeddings work
I'm not really sure that this is going to work. I'd rather explore re-using what I've already trained
|
2022-04-08 16:18:35 -06:00 |
|
James Betker
|
2fb9ffb0aa
|
Align autoregressive text using start and stop tokens
|
2022-04-08 09:41:59 -06:00 |
|
James Betker
|
628569af7b
|
Another fix
|
2022-04-08 09:41:18 -06:00 |
|
James Betker
|
423293e518
|
fix xtransformers bug
|
2022-04-08 09:12:46 -06:00 |
|
James Betker
|
048f6f729a
|
remove lightweight_gan
|
2022-04-07 23:12:08 -07:00 |
|
James Betker
|
e634996a9c
|
autoregressive_codegen: support key_value caching for faster inference
|
2022-04-07 23:08:46 -07:00 |
|
James Betker
|
d05e162f95
|
reformat x_transformers
|
2022-04-07 23:08:03 -07:00 |
|
James Betker
|
7c578eb59b
|
Fix inference in new autoregressive_codegen
|
2022-04-07 21:22:46 -06:00 |
|
James Betker
|
3f8d7955ef
|
unified_voice with rotary embeddings
|
2022-04-07 20:11:14 -06:00 |
|
James Betker
|
573e5552b9
|
CLVP v1
|
2022-04-07 20:10:57 -06:00 |
|
James Betker
|
71b73db044
|
clean up
|
2022-04-07 11:34:10 -06:00 |
|
James Betker
|
6fc4f49e86
|
some dumb stuff
|
2022-04-07 11:32:34 -06:00 |
|
James Betker
|
e6387c7613
|
Fix eval logic to not run immediately
|
2022-04-07 11:29:57 -06:00 |
|
James Betker
|
305dc95e4b
|
cg2
|
2022-04-06 21:24:36 -06:00 |
|
James Betker
|
e011166dd6
|
autoregressive_codegen r3
|
2022-04-06 21:04:23 -06:00 |
|
James Betker
|
33ef17e9e5
|
fix context
|
2022-04-06 00:45:42 -06:00 |
|
James Betker
|
37bdfe82b2
|
Modify x_transformers to do checkpointing and use relative positional biases
|
2022-04-06 00:35:29 -06:00 |
|
James Betker
|
09879b434d
|
bring in x_transformers
|
2022-04-06 00:21:58 -06:00 |
|
James Betker
|
3d916e7687
|
Fix evaluation when using multiple batch sizes
|
2022-04-05 07:51:09 -06:00 |
|
James Betker
|
572d137589
|
track iteration rate
|
2022-04-04 12:33:25 -06:00 |
|
James Betker
|
4cdb0169d0
|
update training data encountered when using force_start_step
|
2022-04-04 12:25:00 -06:00 |
|
James Betker
|
cdd12ff46c
|
Add code validation to autoregressive_codegen
|
2022-04-04 09:51:41 -06:00 |
|
James Betker
|
99de63a922
|
man I'm really on it tonight....
|
2022-04-02 22:01:33 -06:00 |
|
James Betker
|
a4bdc80933
|
moikmadsf
|
2022-04-02 21:59:50 -06:00 |
|
James Betker
|
1cf20b7337
|
sdfds
|
2022-04-02 21:58:09 -06:00 |
|
James Betker
|
b6afc4d542
|
dsfa
|
2022-04-02 21:57:00 -06:00 |
|
James Betker
|
4c6bdfc9e2
|
get rid of relative position embeddings, which do not work with DDP & checkpointing
|
2022-04-02 21:55:32 -06:00 |
|
James Betker
|
b6d62aca5d
|
add inference model on top of codegen
|
2022-04-02 21:25:10 -06:00 |
|
James Betker
|
2b6ff09225
|
autoregressive_codegen v1
|
2022-04-02 15:07:39 -06:00 |
|
James Betker
|
00767219fc
|
undo latent converter change
|
2022-04-01 20:46:27 -06:00 |
|
James Betker
|
55c86e02c7
|
Flat fix
|
2022-04-01 19:13:33 -06:00 |
|
James Betker
|
8623c51902
|
fix bug
|
2022-04-01 16:11:34 -06:00 |
|
James Betker
|
035bcd9f6c
|
fwd fix
|
2022-04-01 16:03:07 -06:00 |
|
James Betker
|
f6a8b0a5ca
|
prep flat0 for feeding from autoregressive_latent_converter
|
2022-04-01 15:53:45 -06:00 |
|
James Betker
|
3e97abc8a9
|
update flat0 to break out timestep-independent inference steps
|
2022-04-01 14:38:53 -06:00 |
|
James Betker
|
a6181a489b
|
Fix loss gapping caused by poor gradients into mel_pred
|
2022-03-26 22:49:14 -06:00 |
|
James Betker
|
0070867d0f
|
inference script for diffusion image models
|
2022-03-26 22:48:24 -06:00 |
|
James Betker
|
1feade23ff
|
support x-transformers in text_voice_clip and support relative positional embeddings
|
2022-03-26 22:48:10 -06:00 |
|
James Betker
|
9b90472e15
|
feed direct inputs into gd
|
2022-03-26 08:36:19 -06:00 |
|
James Betker
|
6909f196b4
|
make code pred returns optional
|
2022-03-26 08:33:30 -06:00 |
|
James Betker
|
2a29a71c37
|
attempt to force meaningful codes by adding a surrogate loss
|
2022-03-26 08:31:40 -06:00 |
|
James Betker
|
45804177b8
|
more stuff
|
2022-03-25 00:03:18 -06:00 |
|
James Betker
|
d4218d8443
|
mods
|
2022-03-24 23:31:20 -06:00 |
|
James Betker
|
9c79fec734
|
update adf
|
2022-03-24 21:20:29 -06:00 |
|
James Betker
|
07731d5491
|
Fix ET
|
2022-03-24 21:20:22 -06:00 |
|
James Betker
|
a15970dd97
|
disable checkpointing in conditioning encoder
|
2022-03-24 11:49:04 -06:00 |
|
James Betker
|
cc5fc91562
|
flat0 work
|
2022-03-24 11:46:53 -06:00 |
|
James Betker
|
b0d2827fad
|
flat0
|
2022-03-24 11:30:40 -06:00 |
|
James Betker
|
8707a3e0c3
|
drop full layers in layerdrop, not half layers
|
2022-03-23 17:15:08 -06:00 |
|
James Betker
|
57da6d0ddf
|
more simplifications
|
2022-03-22 11:46:03 -06:00 |
|
James Betker
|
f3f391b372
|
undo sandwich
|
2022-03-22 11:43:24 -06:00 |
|
James Betker
|
927731f3b4
|
tts9: fix position embeddings snafu
|
2022-03-22 11:41:32 -06:00 |
|
James Betker
|
536511fc4b
|
unified_voice: relative position encodings
|
2022-03-22 11:41:13 -06:00 |
|
James Betker
|
be5f052255
|
misc
|
2022-03-22 11:40:56 -06:00 |
|
James Betker
|
963f0e9cee
|
fix unscaler
|
2022-03-22 11:40:02 -06:00 |
|
James Betker
|
5405ce4363
|
fix flat
|
2022-03-22 11:39:39 -06:00 |
|
James Betker
|
e47a759ed8
|
.......
|
2022-03-21 17:22:35 -06:00 |
|
James Betker
|
cc4c9faf9a
|
resolve more issues
|
2022-03-21 17:20:05 -06:00 |
|
James Betker
|
3692c4cae3
|
map vocoder into cpu
|
2022-03-21 17:10:57 -06:00 |
|
James Betker
|
9e97cd800c
|
take the conditioning mean rather than the first element
|
2022-03-21 16:58:03 -06:00 |
|
James Betker
|
9c7598dc9a
|
fix conditioning_free signal
|
2022-03-21 15:29:17 -06:00 |
|
James Betker
|
2a65c982ca
|
dont double nest checkpointing
|
2022-03-21 15:27:51 -06:00 |
|
James Betker
|
723f324eda
|
Make it even better
|
2022-03-21 14:50:59 -06:00 |
|
James Betker
|
e735d8e1fa
|
unified_voice fixes
|
2022-03-21 14:44:00 -06:00 |
|
James Betker
|
1ad18d29a8
|
Flat fixes
|
2022-03-21 14:43:52 -06:00 |
|
James Betker
|
26dcf7f1a2
|
r2 of the flat diffusion
|
2022-03-21 11:40:43 -06:00 |
|
James Betker
|
c5000420f6
|
more arbitrary fixes
|
2022-03-17 17:45:44 -06:00 |
|
James Betker
|
c14fc003ed
|
flat diffusion
|
2022-03-17 17:45:27 -06:00 |
|
James Betker
|
428911cd4d
|
flat diffusion network
|
2022-03-17 10:53:56 -06:00 |
|
James Betker
|
bf08519d71
|
fixes
|
2022-03-17 10:53:39 -06:00 |
|
James Betker
|
95ea0a592f
|
More cleaning
|
2022-03-16 12:05:56 -06:00 |
|
James Betker
|
d186414566
|
More spring cleaning
|
2022-03-16 12:04:00 -06:00 |
|
James Betker
|
735f6e4640
|
Move gen_similarities and rename
|
2022-03-16 11:59:34 -06:00 |
|
James Betker
|
8b376e63d9
|
More improvements
|
2022-03-16 10:16:34 -06:00 |
|
James Betker
|
54202aa099
|
fix mel normalization
|
2022-03-16 09:26:55 -06:00 |
|
James Betker
|
8437bb0c53
|
fixes
|
2022-03-15 23:52:48 -06:00 |
|
James Betker
|
3f244f6a68
|
add mel_norm to std injector
|
2022-03-15 22:16:59 -06:00 |
|
James Betker
|
0fc877cbc8
|
tts9 fix for alignment size
|
2022-03-15 21:43:14 -06:00 |
|
James Betker
|
f563a8dd41
|
fixes
|
2022-03-15 21:43:00 -06:00 |
|
James Betker
|
b754058018
|
Update wav2vec2 wrapper
|
2022-03-15 11:35:38 -06:00 |
|
James Betker
|
1e3a8554a1
|
updates to audio_diffusion_fid
|
2022-03-15 11:35:09 -06:00 |
|
James Betker
|
9c6f776980
|
Add univnet vocoder
|
2022-03-15 11:34:51 -06:00 |
|
James Betker
|
7929fd89de
|
Refactor audio-style models into the audio folder
|
2022-03-15 11:06:25 -06:00 |
|
James Betker
|
f95d3d2b82
|
move waveglow to audio/vocoders
|
2022-03-15 11:03:07 -06:00 |
|
James Betker
|
0419a64107
|
misc
|
2022-03-15 10:36:34 -06:00 |
|
James Betker
|
bb03cbb9fc
|
composable initial checkin
|
2022-03-15 10:35:40 -06:00 |
|
James Betker
|
86b0d76fb9
|
tts8 (incomplete, may be removed)
|
2022-03-15 10:35:31 -06:00 |
|
James Betker
|
eecbc0e678
|
Use wider spectrogram when asked
|
2022-03-15 10:35:11 -06:00 |
|
James Betker
|
9767260c6c
|
tacotron stft - loosen bounds restrictions and clip
|
2022-03-15 10:31:26 -06:00 |
|
James Betker
|
f8631ad4f7
|
Updates to support inputting MELs into the conditioning encoder
|
2022-03-14 17:31:42 -06:00 |
|
James Betker
|
e045fb0ad7
|
fix clip grad norm with scaler
|
2022-03-13 16:28:23 -06:00 |
|
James Betker
|
22c67ce8d3
|
tts9 mods
|
2022-03-13 10:25:55 -06:00 |
|
James Betker
|
08599b4c75
|
fix random_audio_crop injector
|
2022-03-12 20:42:29 -07:00 |
|
James Betker
|
8f130e2b3f
|
add scale_shift_norm back to tts9
|
2022-03-12 20:42:13 -07:00 |
|
James Betker
|
9bbbe26012
|
update audio_with_noise
|
2022-03-12 20:41:47 -07:00 |
|
James Betker
|
e754c4fbbc
|
sweep update
|
2022-03-12 15:33:00 -07:00 |
|
James Betker
|
73bfd4a86d
|
another tts9 update
|
2022-03-12 15:17:06 -07:00 |
|
James Betker
|
0523777ff7
|
add efficient config to tts9
|
2022-03-12 15:10:35 -07:00 |
|
James Betker
|
896accb71f
|
data and prep improvements
|
2022-03-12 15:10:11 -07:00 |
|
James Betker
|
1e87b934db
|
potentially average conditioning inputs
|
2022-03-10 20:37:41 -07:00 |
|
James Betker
|
e6a95f7c11
|
Update tts9: Remove torchscript provisions and add mechanism to train solely on codes
|
2022-03-09 09:43:38 -07:00 |
|
James Betker
|
726e30c4f7
|
Update noise augmentation dataset to include voices that are appended at the end of another clip.
|
2022-03-09 09:43:10 -07:00 |
|
James Betker
|
c4e4cf91a0
|
add support for the original vocoder to audio_diffusion_fid; also add a new "intelligibility" metric
|
2022-03-08 15:53:27 -07:00 |
|
James Betker
|
3e5da71b16
|
add grad scaler scale to metrics
|
2022-03-08 15:52:42 -07:00 |
|
James Betker
|
d2bdeb6f20
|
misc audio support
|
2022-03-08 15:52:26 -07:00 |
|
James Betker
|
d553808d24
|
misc
|
2022-03-08 15:52:16 -07:00 |
|