James Betker
|
26dcf7f1a2
|
r2 of the flat diffusion
|
2022-03-21 11:40:43 -06:00 |
|
James Betker
|
c5000420f6
|
more arbitrary fixes
|
2022-03-17 17:45:44 -06:00 |
|
James Betker
|
c14fc003ed
|
flat diffusion
|
2022-03-17 17:45:27 -06:00 |
|
James Betker
|
428911cd4d
|
flat diffusion network
|
2022-03-17 10:53:56 -06:00 |
|
James Betker
|
bf08519d71
|
fixes
|
2022-03-17 10:53:39 -06:00 |
|
James Betker
|
95ea0a592f
|
More cleaning
|
2022-03-16 12:05:56 -06:00 |
|
James Betker
|
d186414566
|
More spring cleaning
|
2022-03-16 12:04:00 -06:00 |
|
James Betker
|
735f6e4640
|
Move gen_similarities and rename
|
2022-03-16 11:59:34 -06:00 |
|
James Betker
|
8b376e63d9
|
More improvements
|
2022-03-16 10:16:34 -06:00 |
|
James Betker
|
54202aa099
|
fix mel normalization
|
2022-03-16 09:26:55 -06:00 |
|
James Betker
|
8437bb0c53
|
fixes
|
2022-03-15 23:52:48 -06:00 |
|
James Betker
|
3f244f6a68
|
add mel_norm to std injector
|
2022-03-15 22:16:59 -06:00 |
|
James Betker
|
0fc877cbc8
|
tts9 fix for alignment size
|
2022-03-15 21:43:14 -06:00 |
|
James Betker
|
f563a8dd41
|
fixes
|
2022-03-15 21:43:00 -06:00 |
|
James Betker
|
b754058018
|
Update wav2vec2 wrapper
|
2022-03-15 11:35:38 -06:00 |
|
James Betker
|
1e3a8554a1
|
updates to audio_diffusion_fid
|
2022-03-15 11:35:09 -06:00 |
|
James Betker
|
9c6f776980
|
Add univnet vocoder
|
2022-03-15 11:34:51 -06:00 |
|
James Betker
|
7929fd89de
|
Refactor audio-style models into the audio folder
|
2022-03-15 11:06:25 -06:00 |
|
James Betker
|
f95d3d2b82
|
move waveglow to audio/vocoders
|
2022-03-15 11:03:07 -06:00 |
|
James Betker
|
0419a64107
|
misc
|
2022-03-15 10:36:34 -06:00 |
|
James Betker
|
bb03cbb9fc
|
composable initial checkin
|
2022-03-15 10:35:40 -06:00 |
|
James Betker
|
86b0d76fb9
|
tts8 (incomplete, may be removed)
|
2022-03-15 10:35:31 -06:00 |
|
James Betker
|
eecbc0e678
|
Use wider spectrogram when asked
|
2022-03-15 10:35:11 -06:00 |
|
James Betker
|
9767260c6c
|
tacotron stft - loosen bounds restrictions and clip
|
2022-03-15 10:31:26 -06:00 |
|
James Betker
|
f8631ad4f7
|
Updates to support inputting MELs into the conditioning encoder
|
2022-03-14 17:31:42 -06:00 |
|
James Betker
|
e045fb0ad7
|
fix clip grad norm with scaler
|
2022-03-13 16:28:23 -06:00 |
|
James Betker
|
22c67ce8d3
|
tts9 mods
|
2022-03-13 10:25:55 -06:00 |
|
James Betker
|
08599b4c75
|
fix random_audio_crop injector
|
2022-03-12 20:42:29 -07:00 |
|
James Betker
|
8f130e2b3f
|
add scale_shift_norm back to tts9
|
2022-03-12 20:42:13 -07:00 |
|
James Betker
|
9bbbe26012
|
update audio_with_noise
|
2022-03-12 20:41:47 -07:00 |
|
James Betker
|
e754c4fbbc
|
sweep update
|
2022-03-12 15:33:00 -07:00 |
|
James Betker
|
73bfd4a86d
|
another tts9 update
|
2022-03-12 15:17:06 -07:00 |
|
James Betker
|
0523777ff7
|
add efficient config to tts9
|
2022-03-12 15:10:35 -07:00 |
|
James Betker
|
896accb71f
|
data and prep improvements
|
2022-03-12 15:10:11 -07:00 |
|
James Betker
|
1e87b934db
|
potentially average conditioning inputs
|
2022-03-10 20:37:41 -07:00 |
|
James Betker
|
e6a95f7c11
|
Update tts9: Remove torchscript provisions and add mechanism to train solely on codes
|
2022-03-09 09:43:38 -07:00 |
|
James Betker
|
726e30c4f7
|
Update noise augmentation dataset to include voices that are appended at the end of another clip.
|
2022-03-09 09:43:10 -07:00 |
|
James Betker
|
c4e4cf91a0
|
add support for the original vocoder to audio_diffusion_fid; also add a new "intelligibility" metric
|
2022-03-08 15:53:27 -07:00 |
|
James Betker
|
3e5da71b16
|
add grad scaler scale to metrics
|
2022-03-08 15:52:42 -07:00 |
|
James Betker
|
d2bdeb6f20
|
misc audio support
|
2022-03-08 15:52:26 -07:00 |
|
James Betker
|
d553808d24
|
misc
|
2022-03-08 15:52:16 -07:00 |
|
James Betker
|
7dabc17626
|
phase2 filter initial commit
|
2022-03-08 15:51:55 -07:00 |
|
James Betker
|
f56edb2122
|
minicoder with classifier head: spread out probability mass for 0 predictions
|
2022-03-08 15:51:31 -07:00 |
|
James Betker
|
29b2921222
|
move diffusion vocoder
|
2022-03-08 15:51:05 -07:00 |
|
James Betker
|
94222b0216
|
tts9 initial commit
|
2022-03-08 15:50:45 -07:00 |
|
James Betker
|
38fd9fc985
|
Improve efficiency of audio_with_noise_dataset
|
2022-03-08 15:50:13 -07:00 |
|
James Betker
|
b3def182de
|
move processing pipeline to "phase_1"
|
2022-03-08 15:49:51 -07:00 |
|
James Betker
|
30ddac69aa
|
lots of bad entries
|
2022-03-05 23:15:59 -07:00 |
|
James Betker
|
dcf98df0c2
|
++
|
2022-03-05 23:12:34 -07:00 |
|
James Betker
|
64d764ccd7
|
fml
|
2022-03-05 23:11:10 -07:00 |
|