DL-Art-School

Author	SHA1	Message	Date
James Betker	8d01f7685c	Get rid of absolute positional embeddings in unifiedvoice	2021-12-26 00:10:24 -07:00
James Betker	6700f8851d	moar verbosity	2021-12-25 23:23:21 -07:00
James Betker	8acf3b3097	Better dimensional asserting	2021-12-25 23:18:25 -07:00
James Betker	e959541494	Add position embeddings back into unified_voice I think this may be the solution behind the days problems.	2021-12-25 23:10:56 -07:00
James Betker	ab9cafa572	Make tokenization configs more configurable	2021-12-25 12:17:50 -07:00
James Betker	52410fd9d9	256-bpe tokenizer	2021-12-25 08:52:08 -07:00
James Betker	8e26400ce2	Add inference for unified gpt	2021-12-24 13:27:06 -07:00
James Betker	8b19c37409	UnifiedGptVoice!	2021-12-23 15:20:26 -07:00
James Betker	e55d949855	GrandConjoinedDataset	2021-12-23 14:32:33 -07:00
James Betker	c737632eae	Train and use a bespoke tokenizer	2021-12-22 15:06:14 -07:00
James Betker	66bc60aeff	Re-add start_text_token	2021-12-22 14:10:35 -07:00
James Betker	a9629f7022	Try out using the GPT tokenizer rather than nv_tacotron This results in a significant compression of the text domain, I'm curious what the effect on speech quality will be.	2021-12-22 14:03:18 -07:00
James Betker	7ae7d423af	VoiceCLIP model	2021-12-22 13:44:11 -07:00
James Betker	09f7f3e615	Remove obsolete lucidrains DALLE stuff, re-create it in a dedicated folder	2021-12-22 13:44:02 -07:00
James Betker	a42b94ab72	gpt_tts_hf inference fixes	2021-12-22 13:22:15 -07:00
James Betker	48e3ee9a5b	Shuffle conditioning inputs along the positional axis to reduce fitting on prosody and other positional information The mels should still retain some short-range positional information the model can use for tone and frequencies, for example.	2021-12-20 19:05:56 -07:00
James Betker	53858b2055	Fix gpt_tts_hf inference	2021-12-20 17:45:26 -07:00
James Betker	712d746e9b	gpt_tts: format conditioning inputs more for contextual voice clues and less for prosidy also support single conditional inputs	2021-12-19 17:42:29 -07:00
James Betker	c813befd53	Remove dedicated positioning embeddings	2021-12-19 09:01:31 -07:00
James Betker	b4ddcd7111	More inference improvements	2021-12-19 09:01:19 -07:00
James Betker	f9c45d70f0	Fix mel terminator	2021-12-18 17:18:06 -07:00
James Betker	937045cb63	Fixes	2021-12-18 16:45:38 -07:00
James Betker	9b9f7ea61b	GptTtsHf: Make the input/target placement easier to reason about	2021-12-17 10:24:14 -07:00
James Betker	2fb4213a3e	More lossy fixes	2021-12-17 10:01:42 -07:00
James Betker	9e8a9bf6ca	Various fixes to gpt_tts_hf	2021-12-16 23:28:44 -07:00
James Betker	62c8ed9a29	move speech utils	2021-12-16 20:47:37 -07:00
James Betker	4f8c4d130c	gpt_tts_hf: pad mel tokens with an <end_of_sequence> token.	2021-12-12 20:04:50 -07:00
James Betker	8917c02a4d	gpt_tts_hf inference first pass	2021-12-12 19:51:44 -07:00
James Betker	5a664aa56e	misc	2021-12-11 08:17:26 -07:00
James Betker	6ccff3f49f	Record codes more often	2021-12-07 09:22:45 -07:00
James Betker	d0b2f931bf	Add feature to diffusion vocoder where the spectrogram conditioning layers can be re-trained apart from the rest of the model	2021-12-07 09:22:30 -07:00
James Betker	662920bde3	Log codes when simply fetching codebook_indices	2021-12-06 09:21:43 -07:00
James Betker	380a5d5475	gdi..	2021-12-03 08:53:09 -07:00
James Betker	101a01f744	Fix dvae codes issue	2021-12-02 23:28:36 -07:00
James Betker	07b0124712	GptTtsHf!	2021-12-02 21:48:42 -07:00
James Betker	85542ec547	One last fix for gpt_asr_hf2	2021-12-02 21:19:28 -07:00
James Betker	04454ee63a	Add evaluation logic for gpt_asr_hf2	2021-12-02 21:04:36 -07:00
James Betker	5956eb757c	ffffff	2021-11-24 00:19:47 -07:00
James Betker	f1ed0588e3	another fix	2021-11-24 00:11:21 -07:00
James Betker	7a3c4a4fc6	Fix lr quantizer decode	2021-11-24 00:01:26 -07:00
James Betker	3f6ecfe0db	q fix	2021-11-23 23:50:27 -07:00
James Betker	d9747fe623	Integrate with lr_quantizer	2021-11-23 19:48:22 -07:00
James Betker	82d0e7720e	Add choke to lucidrains_dvae	2021-11-23 18:53:37 -07:00
James Betker	934395d4b8	A few fixes for gpt_asr_hf2	2021-11-23 09:29:29 -07:00
James Betker	01e635168b	whoops	2021-11-22 17:24:13 -07:00
James Betker	973f47c525	misc nonfunctional	2021-11-22 17:16:39 -07:00
James Betker	3125ca38f5	Further wandb logs	2021-11-22 16:40:19 -07:00
James Betker	0604060580	Finish up mods for next version of GptAsrHf	2021-11-20 21:33:49 -07:00
James Betker	14f3155ec4	misc	2021-11-20 17:45:14 -07:00
James Betker	555b7e52ad	Add rev2 of GptAsrHf	2021-11-18 20:02:24 -07:00
James Betker	1287915f3c	Fix dvae test failure	2021-11-18 00:58:36 -07:00
James Betker	019acfa4c5	Allow flat dvae	2021-11-18 00:53:42 -07:00
James Betker	f3db41f125	Fix code logging	2021-11-18 00:34:37 -07:00
James Betker	79367f753d	Fix error & add nonfinite warning	2021-11-09 23:58:41 -07:00
James Betker	c584320cf3	Fix gpt_asr_hf distillation	2021-11-07 21:53:21 -07:00
James Betker	a367ea3fda	Add script for computing attention for gpt_asr	2021-11-07 18:42:06 -07:00
James Betker	756b4dad09	Working gpt_asr_hf inference - and it's a beast!	2021-11-06 21:47:15 -06:00
James Betker	596a62fe01	Apply fix to gpt_asr_hf and prep it for inference Fix is that we were predicting two characters in advance, not next character	2021-11-04 10:09:24 -06:00
James Betker	993bd52d42	Add spec_augment injector	2021-11-01 18:43:11 -06:00
James Betker	4cff774b0e	Reduce complexity of the encoder for gpt_asr_hf	2021-11-01 17:02:28 -06:00
James Betker	da55ca0438	gpt_asr using the huggingfaces transformer	2021-11-01 17:00:22 -06:00
James Betker	83cccef9d8	Condition on full signal	2021-10-30 19:58:34 -06:00
James Betker	df45a9dec2	Fix inference mode for lucidrains_gpt	2021-10-30 16:59:18 -06:00
James Betker	92fe8b4dd9	ffffpt2	2021-10-29 17:29:49 -06:00
James Betker	95ca88efce	Fix feedforward	2021-10-29 17:27:51 -06:00
James Betker	b476516340	Check in backing changes (which may have broken something?)	2021-10-29 17:22:33 -06:00
James Betker	986fc9628d	Check in GPT with new inference methods (but not the backing code..)	2021-10-29 17:21:40 -06:00
James Betker	58494b0888	Add support for distilling gpt_asr	2021-10-27 13:10:07 -06:00
James Betker	5d714bc566	Add deepspeech model and support for decoding with it	2021-10-27 13:09:46 -06:00
James Betker	3a9d1c53ea	Rework conditioning inputs provided	2021-10-26 10:46:33 -06:00
James Betker	43e389aac6	Add time_embed_dim_multiplier	2021-10-26 08:55:55 -06:00
James Betker	ba6e46c02a	Further simplify diffusion_vocoder and make noise_surfer work	2021-10-26 08:54:30 -06:00
James Betker	0ee1c67ce5	Rework how conditioning inputs are applied to DiffusionVocoder	2021-10-24 09:08:58 -06:00
James Betker	06ea6191a9	Initial implementation of audio_with_noise dataset	2021-10-21 16:45:19 -06:00
James Betker	0dee15f875	base DVAE & vector_quantizer	2021-10-20 21:19:38 -06:00
James Betker	f2a31702b5	Clean stuff up, move more things into arch_util	2021-10-20 21:19:25 -06:00
James Betker	a6f0f854b9	Fix codes when inferring from dvae	2021-10-17 22:51:17 -06:00
James Betker	d016a2fbad	Go back to vanilla flavor of diffusion	2021-10-17 17:32:46 -06:00
James Betker	23da073037	Norm decoder outputs now	2021-10-16 09:07:10 -06:00
James Betker	0edc98f6c4	Throw out the idea of conditioning on discrete codes. Oh well :(	2021-10-16 09:02:01 -06:00
James Betker	62c8c5d93e	Zero out spectrogram code inputs initially.	2021-10-15 12:10:11 -06:00
James Betker	1d0b44ebc2	More tweaks to diffusion-vocoder	2021-10-15 11:51:17 -06:00
James Betker	3b19581f9a	Allow num_resblocks to specified per-level	2021-10-14 11:26:04 -06:00
James Betker	83798887a8	Mods to support unet diffusion vocoder with conditioning	2021-10-13 21:23:18 -06:00
James Betker	33120cb35c	Add norming to discretization_loss	2021-10-06 17:10:50 -06:00
James Betker	f2977d360c	Allow attention_dim in channel attention to be specified, add converter	2021-10-05 17:29:38 -06:00
James Betker	9c0d7288ea	Discretization loss attempt	2021-10-04 20:59:21 -06:00
James Betker	66f99a159c	Rev2	2021-10-03 15:20:50 -06:00
James Betker	09f373e3b1	Add dvae with channel attention	2021-10-03 10:52:01 -06:00
James Betker	0396a9d2ca	Increase baseline codes recording across all dvae models	2021-09-30 08:09:07 -06:00
James Betker	f84ccbdfb2	Fix quantizer with balancing_heuristic	2021-09-29 14:46:05 -06:00
James Betker	4914c526dc	More cleanup	2021-09-29 14:24:49 -06:00
James Betker	6e550edfe3	Attentive dvae	2021-09-29 14:17:29 -06:00
James Betker	55b58fb67f	Clean up codebase Remove stuff that I'm likely not going to use again (or generally failed experiments)	2021-09-29 09:21:44 -06:00
James Betker	4d1a42e944	Add switchnorm to gumbel_quantizer	2021-09-24 18:49:25 -06:00
James Betker	ac57cdc794	Add scheduling to quantizer, enable cudnn_benchmarking to be disabled	2021-09-24 17:01:36 -06:00
James Betker	3e64e847c2	Gumbel quantizer	2021-09-23 23:32:03 -06:00
James Betker	c5297ccec6	Add dvae balancing heuristic	2021-09-23 21:19:36 -06:00
James Betker	e24c619387	Fix	2021-09-23 16:07:58 -06:00
James Betker	6833048bf7	Alterations to diffusion_dvae so it can be used directly on spectrograms	2021-09-23 15:56:25 -06:00
James Betker	5c8d266d4f	chk	2021-09-17 09:15:36 -06:00
James Betker	a6544f1684	More checkpointing fixes	2021-09-16 23:12:43 -06:00
James Betker	94899d88f3	Fix overuse of checkpointing	2021-09-16 23:00:28 -06:00
James Betker	f78ce9d924	Get diffusion_dvae ready for prime time!	2021-09-16 22:43:10 -06:00
James Betker	6f48674647	Support diffusion models with extra return values & inference in diffusion_dvae	2021-09-16 10:53:46 -06:00
James Betker	0382660159	Get diffusion_dvae functional	2021-09-14 17:43:31 -06:00
James Betker	76e2c497f7	Improvements to splitter	2021-09-09 23:34:56 -06:00
James Betker	742f9b4010	Batch spleeter cleaner using GPU	2021-09-09 23:14:32 -06:00
James Betker	73b930c0f6	Add diffusion_dvae Increase split_on_silence interval	2021-09-09 16:22:05 -06:00
James Betker	b8f2e0f452	mydvae	2021-09-06 17:45:30 -06:00
James Betker	3e073cff85	Set kernel_size in diffusion_vocoder	2021-09-01 08:33:46 -06:00
James Betker	dabd87246d	Add unet_diffusion_vocoder	2021-08-31 14:38:33 -06:00
James Betker	909754cc27	Add find_faulty_files.py	2021-08-25 18:00:43 -06:00
James Betker	08b33c8e3a	Support silu activation	2021-08-25 09:03:14 -06:00
James Betker	67bf7f5219	dvae mods Trying to squeeze as much performance out of this net as possible	2021-08-25 08:55:13 -06:00
James Betker	b521d94b01	Make gpt-asr more configurable	2021-08-19 16:33:41 -06:00
James Betker	570ed327ed	Stop dataset - attempt #2	2021-08-18 18:29:38 -06:00
James Betker	17453ccbe8	Revert mods to lrdvae They didn't really change anything	2021-08-17 09:09:29 -06:00
James Betker	8332923f5c	Two more tools to test the audio segmentor	2021-08-17 09:09:11 -06:00
James Betker	1fede41b7b	Audio segmentor	2021-08-16 22:51:53 -06:00
James Betker	729c1fd5a9	Fix up max lengths to save memory	2021-08-15 21:29:28 -06:00
James Betker	9e47e64d5a	Add gpt_segmentor model The idea is to specifically train a model that extracts phrases from audio clips.	2021-08-15 21:23:07 -06:00
James Betker	a826d5f658	Mods to dvae - Add resblock to each layer - Increase filter size for each layer - Use SiLU	2021-08-15 20:54:10 -06:00
James Betker	b8bec22f1a	Fix gpt_asr inference bug	2021-08-15 20:53:42 -06:00
James Betker	a523c4f932	Auto-normalize wav files by data type	2021-08-15 09:09:51 -06:00
James Betker	98057b6516	Make lrdvae use quantized mode in eval()	2021-08-14 23:43:01 -06:00
James Betker	ad3391bd96	Fix nan issue when interpolating audio	2021-08-14 20:42:01 -06:00
James Betker	d6a73acaed	Allow processing of multiple audio sources at once from nv_tacotron_dataset	2021-08-14 16:04:05 -06:00
James Betker	007976082b	GPT_asr for inference	2021-08-14 14:37:17 -06:00
James Betker	e1bdd3f7c7	Fix gpt_asr bug. Initial implementation of beam search	2021-08-13 22:47:00 -06:00
James Betker	cdee31c60b	GPT_ASR	2021-08-13 15:02:18 -06:00
James Betker	f5a9b88ef6	tacotron cleaners: remove quotation marks these don't really have relevance for tts or asr	2021-08-11 16:18:44 -06:00
James Betker	20586a8edc	Fix LRDVAE bug with quantizer integration	2021-08-11 16:17:22 -06:00
James Betker	82fc69abfa	Add "pure" evaluator Which simply computes the training loss against an eval dataset	2021-08-09 14:58:35 -06:00
James Betker	080bea2f19	No, really	2021-08-09 12:02:31 -06:00
James Betker	e1ce4671e4	Apply dropout to gpt_tts, get rid of min_gpt implementation	2021-08-09 12:01:10 -06:00
James Betker	1068f53b78	Add a sampling beam search	2021-08-09 11:56:06 -06:00
James Betker	01cfae28d8	Beam search implementation in one pass? Dayyyum	2021-08-08 23:22:42 -06:00
James Betker	690d7e86d3	Fix nv_tacotron_dataset bug which incorrectly mapped filenames dammit..	2021-08-08 11:38:52 -06:00
James Betker	a2afb25e42	Fix inference, always flow full text tokens through transformer	2021-08-07 20:11:10 -06:00
James Betker	4c678172d6	ugh	2021-08-06 22:10:18 -06:00
James Betker	e723137273	Make gpttts more configurable	2021-08-06 22:08:51 -06:00
James Betker	a7496b661c	combined dvae ftw	2021-08-06 22:01:06 -06:00
James Betker	0237e96b34	Fix dvae bug	2021-08-06 14:17:01 -06:00
James Betker	0799d95af5	Use quantizer from rosinality/vqvae with openai dvae	2021-08-06 14:06:26 -06:00
James Betker	d3ace153af	Add logic for performing inference using gpt_tts with dual-encoder modes	2021-08-06 12:04:12 -06:00
James Betker	b43683b772	Add lucidrains_dvae	2021-08-06 12:03:46 -06:00
James Betker	70dcd1107f	Fix byol_model_wrapper to function with audio inputs	2021-08-05 22:20:22 -06:00
James Betker	89d15c9e74	Move gpt-tts back to lucidrains implementation Much better performance.	2021-08-05 22:15:13 -06:00
James Betker	d120e1aa99	Add audio augmentation to wavfile_dataset, utility to test audio similary	2021-08-05 22:14:49 -06:00
James Betker	c0f61a2e15	Rework how DVAE tokens are ordered It might make more sense to have top tokens, then bottom tokens with top tokens having different discretized values.	2021-08-05 07:07:17 -06:00
James Betker	4017236ba9	Fix up inference for gpt_tts	2021-08-05 06:46:30 -06:00
James Betker	5037220ac7	Mods to support contrastive learning on audio files	2021-08-05 05:57:04 -06:00
James Betker	341f28dd82	It works!	2021-08-04 20:07:51 -06:00
James Betker	36c7c1fbdb	Fix training flow for NEXT TOKEN prediction instead of same token prediction doh	2021-08-04 10:28:09 -06:00
James Betker	d9936df363	Add gpt_tts dataset and implement inference - Adds a script which preprocesses quantized mels given a DVAE - Adds a dataset which can consume preprocessed qmels - Reworks GPT TTS to consume the outputs of that dataset (removes logic to add padding and start/end tokens) - Adds inference to gpt_tts	2021-08-04 00:44:04 -06:00
James Betker	4c98b9703f	Get dalle-style TTS to "work"	2021-08-03 21:08:27 -06:00
James Betker	2814307eee	Alterations to support VQVAE on mel spectrograms	2021-08-01 07:54:21 -06:00
James Betker	0c9e75bc69	Improvements to GptTts	2021-07-31 15:57:57 -06:00
James Betker	31ee9ae262	Checkin	2021-07-30 23:07:35 -06:00
James Betker	dadc54795c	Add gpt_tts	2021-07-27 20:33:30 -06:00
James Betker	398185e109	More work on wave-diffusion	2021-07-27 05:36:17 -06:00
James Betker	49e3b310ea	Allow audio sample rate interpolation for faster training	2021-07-26 17:44:06 -06:00
James Betker	96e90e7047	Add support for a gaussian-diffusion-based wave tacotron	2021-07-26 16:27:31 -06:00
James Betker	97d7cbbc34	Additional work for audio xformer (which doesnt really do a great job)	2021-07-23 10:58:14 -06:00
James Betker	d81386c1be	Mods to support vqvae in audio mode (1d)	2021-07-20 08:36:46 -06:00
James Betker	5584cfcc7a	tacotron2 work	2021-07-14 21:41:57 -06:00
James Betker	fe0c699ced	Various fixes	2021-07-14 00:08:42 -06:00
James Betker	be2745f42d	Add waveglow & inference capabilities to audio generator	2021-07-08 23:07:36 -06:00
James Betker	1ff434218e	tacotron2, ready for prime time!	2021-07-08 22:13:44 -06:00
James Betker	86fd3ad7fd	Initial checkin of nvidia tacotron model & dataset These two are tested, full support for training to come.	2021-07-06 11:11:35 -06:00
James Betker	afa41f1804	Allow hq color jittering and corruptions that are not included in the corruption factor	2021-06-30 09:44:46 -06:00
James Betker	6fd16ea9c8	Add meta-anomaly detection, colorjitter augmentation	2021-06-29 13:41:55 -06:00
James Betker	46e9f62be0	Add unet with latent guide This is a diffusion network that uses both a LQ image and a reference sample HQ image that is compressed into a latent vector to perform upsampling The hope is that we can steer the upsampling network with sample images.	2021-06-26 11:02:58 -06:00
James Betker	0ded106562	Merge remote-tracking branch 'origin/master'	2021-06-25 13:16:28 -06:00
James Betker	a57ed8e960	Various mods to support better jpeg image filtering	2021-06-25 13:16:15 -06:00
James Betker	a0ef07ddb8	Create unet_latent_guide.py	2021-06-25 11:25:14 -06:00
James Betker	e7890dc0ba	Misc fixes for diffusion nets	2021-06-21 10:38:07 -06:00
James Betker	65c474eecf	Various changes to fix testing	2021-06-11 15:31:10 -06:00
James Betker	220f11a5e4	Half channel sizes in cifar_resnet	2021-06-09 17:06:37 -06:00
James Betker	9b5f4abb91	Add fade in for hard switch	2021-06-07 18:15:09 -06:00
James Betker	108c5d829c	Fix dropout norm	2021-06-07 16:13:23 -06:00
James Betker	438217094c	Also debug distribution of switch	2021-06-07 15:36:07 -06:00
James Betker	44b09e5f20	Amplify dropout rate	2021-06-07 15:20:53 -06:00
James Betker	f0d4eb9182	Fixor	2021-06-07 11:58:36 -06:00
James Betker	c456a60466	Another go at fixing nan	2021-06-07 11:51:43 -06:00
James Betker	1c574c5bd1	Attempt to fix nan	2021-06-07 11:43:42 -06:00
James Betker	eda796985b	Try out dropout norm	2021-06-07 11:33:33 -06:00
James Betker	6c6e82406e	Pass a corruption factor through the dataset into the upsampling network The intuition is this will help guide the network to make better informed decisions about how it performs upsampling based on how it perceives the underlying content. (I'm giving up on letting networks detect their own quality - I'm not convinced it is actually feasible)	2021-06-07 09:13:54 -06:00
James Betker	061dbcd458	Another fix to anorm	2021-06-06 15:09:49 -06:00
James Betker	9a6991e461	Fix switch norm average	2021-06-06 15:04:28 -06:00
James Betker	57e1a6a0f2	cifar: add hard routing Also mods switched_routing to support non-pixular inputs	2021-06-06 14:53:43 -06:00
James Betker	692e9c417b	Support diffusion unet	2021-06-06 13:57:22 -06:00
James Betker	a0158ebc69	Simplify cifar resnet further for faster training	2021-06-06 10:02:24 -06:00
James Betker	75567a9814	Only head norm removed	2021-06-05 23:29:11 -06:00
James Betker	65d0376b90	Re-add normalization at the tail of the RRDB	2021-06-05 23:04:05 -06:00
James Betker	184e887122	Remove rrdb normalization	2021-06-05 21:39:19 -06:00
James Betker	f5e75602b9	Add regular attention to cifar_resnet	2021-06-05 21:34:07 -06:00
James Betker	af52751d6b	Fix device error	2021-06-05 14:21:32 -06:00
James Betker	5f0cc65f3b	Register branched resnet properly	2021-06-05 14:19:03 -06:00
James Betker	fb405d9ef1	CIFAR stuff - Extract coarse labels for the CIFAR dataset - Add simple resnet that branches lower layers based on coarse labels - Some other cleanup	2021-06-05 14:16:02 -06:00
James Betker	80d4404367	A few fixes: - Output better prediction of xstart from eps - Support LossAwareSampler - Support AdamW	2021-06-05 13:40:32 -06:00
James Betker	7c251af7a8	Support cifar100 with resnet	2021-06-04 17:29:07 -06:00
James Betker	bf811f80c1	GD mods & fixes - Report variational loss separately - Report model prediction from injector - Log these things - Use respacing like guided diffusion	2021-06-04 17:13:16 -06:00
James Betker	6084915af8	Support gaussian diffusion models Adds support for GD models, courtesy of some maths from openai. Also: - Fixes requirement for eval{} even when it isn't being used - Adds support for denormalizing an imagenet norm	2021-06-02 21:47:32 -06:00
James Betker	f129eaa39e	Clean up byol a bit - Remove option to aug in dataset (there's really no reason for this now that kornia works on GPU on windows) - Other stufff	2021-05-24 21:35:46 -06:00
James Betker	1a2b9fa130	Get rid of old byol net wrapping Simplifies and makes this usable with DLAS' multi-gpu trainer	2021-04-27 12:48:34 -06:00
James Betker	119f17c808	Add testing capabilities for segformer & contrastive feature	2021-04-27 09:59:50 -06:00
James Betker	9bbe6fc81e	Get segformer to a trainable state	2021-04-25 11:45:20 -06:00
James Betker	fc623d4b5a	Add segformer model. Start work on BYOL adaptation that will support training it.	2021-04-23 17:16:46 -06:00
James Betker	17555e7d07	misc adjustments for stylegan	2021-04-21 18:14:17 -06:00
James Betker	b687ef4cd0	Misc	2021-04-21 18:09:46 -06:00
James Betker	9fc3df3f5b	Switched conv: add conversion function with allowlist	2021-03-13 10:44:56 -07:00
James Betker	cf9a6da889	Fix some bugs, checkin work on vqvae3	2021-03-02 20:56:19 -07:00
James Betker	f89ea5f1c6	Mods to support lightweight_gan model	2021-03-02 20:51:48 -07:00
James Betker	39fd755baa	New benchmark numbers	2021-02-08 08:09:41 -07:00
James Betker	784b96c059	Misc options to add support for training stylegan2-rosinality models: - Allow image_folder_dataset to normalize inbound images - ExtensibleTrainer can denormalize images on the output path - Support .webp - an output from LSUN - Support logistic GAN divergence loss - Support stylegan2 TF weight extraction for discriminator - New injector that produces latent noise (with separated paths) - Modify FID evaluator to be operable with rosinality-style GANs	2021-02-08 08:09:21 -07:00
James Betker	e7be4bdff3	Revert	2021-02-05 08:43:07 -07:00
James Betker	6dec1f5968	Back to groupnorm	2021-02-05 08:42:11 -07:00
James Betker	336f807c8e	lambda2	2021-02-05 00:00:24 -07:00
James Betker	025a5867c4	Use syncbatchnorm instead	2021-02-04 22:26:36 -07:00
James Betker	bb79fafb89	Fix groupnorm specification	2021-02-04 22:15:38 -07:00
James Betker	43da1f9c4b	Convert lambda coupler to use groupnorm instead of batchnorm	2021-02-04 21:59:44 -07:00
James Betker	7070142805	Make vqvae3_hard more configurable	2021-02-04 09:03:22 -07:00
James Betker	b980028ca8	Add get_debug_values for vqvae_3_hardswitch	2021-02-03 14:12:24 -07:00
James Betker	1405ff06b8	Fix SwitchedConvHardRoutingFunction for current cuda router	2021-02-03 14:11:55 -07:00
James Betker	d7bec392dd	...	2021-02-02 23:50:25 -07:00
James Betker	b0a8fa00bc	Visual dbg in vqvae3hs	2021-02-02 23:50:01 -07:00
James Betker	f5f91850fd	hardswitch variant of vqvae3	2021-02-02 21:00:04 -07:00
James Betker	320edbaa3c	Move switched_conv logic around a bit	2021-02-02 20:41:24 -07:00
James Betker	0dca36946f	Hard Routing mods - Turns out my custom convolution was RIDDLED with backwards bugs, which is why the existing implementation wasn't working so well. - Implements the switch logic from both Mixture of Experts and Switch Transformers for testing purposes.	2021-02-02 20:35:58 -07:00
James Betker	29c1c3bede	Register vqvae3	2021-01-29 15:26:28 -07:00
James Betker	bc20b4739e	vqvae3 Changes VQVAE as so: - Reverts back to smaller codebook - Adds an additional conv layer at the highest resolution for both the encoder & decoder - Uses LeakyReLU on trunk	2021-01-29 15:24:26 -07:00
James Betker	96bc80313c	Add switch norm, up dropout rate, detach selector	2021-01-26 09:31:53 -07:00
James Betker	2cdac6bd09	Add PWCNet for human optical flow	2021-01-25 08:25:44 -07:00
James Betker	51b63b2aa6	Add switched_conv with hard routing and make vqvae use it.	2021-01-25 08:25:29 -07:00
James Betker	ae4ff4a1e7	Enable lambda visualization	2021-01-23 15:53:27 -07:00
James Betker	10ec6bda1d	lambda nets in switched_conv and a vqvae to use it	2021-01-23 14:57:57 -07:00
James Betker	b374dcdd46	update vqvae to double codebook size for bottom quantizer	2021-01-23 13:47:07 -07:00
James Betker	1b8a26db93	New switched_conv	2021-01-23 13:46:30 -07:00
James Betker	d919ae7148	Add VQVAE with no Conv2dTranspose	2021-01-18 08:49:59 -07:00
James Betker	587a4f4050	resnet_unet_3 I'm being really lazy here - these nets are not really different from each other except at which layer they terminate. This one terminates at 2x downsampling, which is simply indicative of a direction I want to go for testing these pixpro networks.	2021-01-15 14:51:03 -07:00
James Betker	038b8654b6	Pixpro: unwrap losses	2021-01-13 11:54:25 -07:00
James Betker	8990801a3f	Fix pixpro stochastic sampling bugs	2021-01-13 11:34:24 -07:00
James Betker	19475a072f	Pixpro: Rather than using a latent square for pixpro, use an entirely stochastic sampling of the pixels	2021-01-13 11:26:51 -07:00
James Betker	d1007ccfe7	Adjustments to pixpro to allow training against networks with arbitrarily large structural latents - The pixpro latent now rescales the latent space instead of using a "coordinate vector", which might have performance implications. - The latent against which the pixel loss is computed can now be a small, randomly sampled patch out of the entire latent, allowing further memory/computational discounts. Since the loss computation does not have a receptive field, this should not alter the loss. - The instance projection size can now be separate from the pixel projection size. - PixContrast removed entirely. - ResUnet with full resolution added.	2021-01-12 09:17:45 -07:00
James Betker	34f8c8641f	Support training imagenet classifier	2021-01-11 20:09:16 -07:00
James Betker	f3db381fa1	Allow uresnet to use pretrained resnet50	2021-01-10 12:57:31 -07:00
James Betker	07168ecfb4	Enable vqvae to use a switched_conv variant	2021-01-09 20:53:14 -07:00
James Betker	5a8156026a	Did anyone ask for k-means clustering? This is so cool...	2021-01-07 22:37:41 -07:00
James Betker	de10c7246a	Add injected noise into bypass maps	2021-01-07 16:31:12 -07:00
James Betker	61a86a3c1e	VQVAE	2021-01-07 10:20:15 -07:00
James Betker	01a589e712	Adjustments to pixpro & resnet-unet I'm not really satisfied with what I got out of these networks on round 1. Lets try again..	2021-01-06 15:00:46 -07:00
James Betker	2f2f87bbea	Styled SR fixes	2021-01-05 20:14:39 -07:00
James Betker	9fed90393f	Add lucidrains pixpro trainer	2021-01-05 20:14:22 -07:00
James Betker	ade2732c82	Transfer learning for styleSR This is a concept from "Lifelong Learning GAN", although I'm skeptical of it's novelty - basically you scale and shift the weights for the generator and discriminator of a pretrained GAN to "shift" into new modalities, e.g. faces->birds or whatever. There are some interesting applications of this that I would like to try out.	2021-01-04 20:10:48 -07:00
James Betker	2c65b6b28e	More mods to support styledsr	2021-01-04 11:32:28 -07:00
James Betker	2225fe6ac2	Undo lucidrains changes for new discriminator This "new" code will live in the styledsr directory from now on.	2021-01-04 10:57:09 -07:00
James Betker	40ec71da81	Move styled_sr into its own folder	2021-01-04 10:54:34 -07:00
James Betker	5916f5f7d4	Misc fixes	2021-01-04 10:53:53 -07:00
James Betker	4d8064c32c	Modifications to allow partially trained stylegan discriminators to be used	2021-01-03 16:37:18 -07:00
James Betker	bdbab65082	Allow optimizers to train separate param groups, add higher dimensional VGG discriminator Did this to support training 512x512px networks off of a pretrained 256x256 network.	2021-01-02 15:10:06 -07:00
James Betker	193cdc6636	Move discriminators to the create_model paradigm Also cleans up a lot of old discriminator models that I have no intention of using again.	2021-01-01 15:56:09 -07:00
James Betker	f39179e85a	styled_sr: fix bug when using initial_stride	2021-01-01 12:13:21 -07:00
James Betker	913fc3b75e	Need init to pick up styled_sr	2021-01-01 12:10:32 -07:00
James Betker	e992e18767	Add initial_stride term to style_sr Also fix fid and a networks.py issue.	2021-01-01 11:59:36 -07:00
James Betker	e214e6ce33	Styled SR model	2020-12-31 20:54:18 -07:00
James Betker	b1fb82476b	Add gp debug (fix)	2020-12-30 15:26:54 -07:00
James Betker	63cf3d3126	Injector auto-registration I love it!	2020-12-29 20:58:02 -07:00
James Betker	a777c1e4f9	Misc script fixes	2020-12-29 20:25:09 -07:00
James Betker	ba543d1152	Glean mods - Fixes fixed upscale factor issues - Refines a few ops to decrease computation & parameterization	2020-12-27 12:25:06 -07:00
James Betker	f9be049adb	GLEAN mod to support custom initial strides	2020-12-26 13:51:14 -07:00
James Betker	3fd627fc62	Mods to support image classification & filtering	2020-12-26 13:49:27 -07:00
James Betker	10fdfa1563	Migrate generators to dynamic model registration	2020-12-24 23:02:10 -07:00
James Betker	29db7c7a02	Further mods to BYOL	2020-12-24 09:28:41 -07:00
James Betker	036684893e	Add LARS optimizer & support for BYOL idiosyncrasies - Added LARS and SGD optimizer variants that support turning off certain features for BN and bias layers - Added a variant of pytorch's resnet model that supports gradient checkpointing. - Modify the trainer infrastructure to support above - Fix bug with BYOL (should have been nonfunctional)	2020-12-23 20:33:43 -07:00
James Betker	1bbcb96ee8	Implement a few changes to support training BYOL networks	2020-12-23 10:50:23 -07:00
James Betker	ae666dc520	Fix bugs with srflow after refactor	2020-12-19 10:28:23 -07:00
James Betker	4328c2f713	Change default ReLU slope to .2 BREAKS COMPATIBILITY This conforms my ConvGnLelu implementation with the generally accepted negative_slope=.2. I have no idea where I got .1. This will break backwards compatibility with some older models but will likely improve their performance when freshly trained. I did some auditing to find what these models might be, and I am not actively using any of them, so probably OK.	2020-12-19 08:28:03 -07:00
James Betker	9377d34ac3	glean mods	2020-12-19 08:26:07 -07:00
James Betker	92f9a129f7	GLEAN!	2020-12-18 16:04:19 -07:00
James Betker	c717765bcb	Notes for lucidrains converter.	2020-12-18 09:55:38 -07:00
James Betker	b4720ea377	Move stylegan to new location	2020-12-18 09:52:36 -07:00
James Betker	1708136b55	Commit my attempt at "conforming" the lucidrains stylegan implementation to the reference spec. Not working. will probably be abandoned.	2020-12-18 09:51:48 -07:00
James Betker	209332292a	Rosinality stylegan fix	2020-12-18 09:50:41 -07:00
James Betker	d875ca8342	More refactor changes	2020-12-18 09:24:31 -07:00
James Betker	5640e4efe4	More refactoring	2020-12-18 09:18:34 -07:00
James Betker	b905b108da	Large cleanup Removed a lot of old code that I won't be touching again. Refactored some code elements into more logical places.	2020-12-18 09:10:44 -07:00
James Betker	3074f41877	Get rosinality model converter to work Mostly, just needed to remove the custom cuda ops, not so bueno on Windows.	2020-12-17 16:03:39 -07:00
James Betker	e838c6e75b	Rosinality stylegan2 port	2020-12-17 14:18:46 -07:00
James Betker	49327b99fe	SRFlow outputs RRDB output	2020-12-16 10:28:02 -07:00
James Betker	c25b49bb12	Clean up of SRFlowNet_arch	2020-12-16 10:27:38 -07:00
James Betker	42ac8e3eeb	Remove unnecessary comment from SRFlowNet	2020-12-16 09:43:07 -07:00
James Betker	09de3052ac	Add softmax to spinenet classification head	2020-12-16 09:42:15 -07:00
James Betker	8661207d57	Merge branch 'gan_lab' of https://github.com/neonbjb/DL-Art-School into gan_lab	2020-12-15 17:16:48 -07:00
James Betker	fc376d34b2	Spinenet with logits head	2020-12-15 17:16:19 -07:00
James Betker	0a19e53df0	BYOL mods	2020-12-14 23:59:11 -07:00
James Betker	ef7eabf457	Allow RRDB to upscale 8x	2020-12-14 23:58:52 -07:00
James Betker	ec0ee25f4b	Structural latents checkpoint	2020-12-11 12:01:09 -07:00
James Betker	26ceca68c0	BYOL with structure!	2020-12-10 15:07:35 -07:00
James Betker	c203cee31e	Allow swapping to torch DDP as needed in code	2020-12-09 15:03:59 -07:00
James Betker	97ff25a086	BYOL! Man, is there anything ExtensibleTrainer can't train? :)	2020-12-08 13:07:53 -07:00
James Betker	bca59ed98a	Merge remote-tracking branch 'origin/gan_lab' into gan_lab	2020-12-07 12:51:04 -07:00
James Betker	ea56eb61f0	Fix DDP errors for discriminator - Don't define training_net in define_optimizers - this drops the shell and leads to problems downstream - Get rid of support for multiple training nets per opt. This was half baked and needs a better solution if needed downstream.	2020-12-07 12:50:57 -07:00
James Betker	88fc049c8d	spinenet latent playground!	2020-12-05 20:30:36 -07:00
James Betker	11155aead4	Directly use dataset keys This has been a long time coming. Cleans up messy "GT" nomenclature and simplifies ExtensibleTraner.feed_data	2020-12-04 20:14:53 -07:00
James Betker	8a83b1c716	Go back to apex DDP, fix distributed bugs	2020-12-04 16:39:21 -07:00
James Betker	7a81d4e2f4	Revert gaussian loss changes	2020-12-04 12:49:20 -07:00
James Betker	711780126e	Cleanup	2020-12-03 23:42:51 -07:00
James Betker	ac7256d4a3	Do tqdm reporting when calculating flow_gaussian_nll	2020-12-03 23:42:29 -07:00
James Betker	dc9ff8e05b	Allow the majority of the srflow steps to checkpoint	2020-12-03 23:41:57 -07:00
James Betker	06d1c62c5a	iGPT support! Sweeeeet	2020-12-03 15:32:21 -07:00
James Betker	c18adbd606	Delete mdcn & panet Garbage, all of it.	2020-12-02 22:25:57 -07:00
James Betker	f2880b33c9	Get rid of mean shift from MDCN	2020-12-02 14:18:33 -07:00
James Betker	8a00f15746	Implement FlowGaussianNll evaluator	2020-12-02 14:09:54 -07:00
James Betker	edf408508c	Fix discriminator	2020-12-01 17:45:56 -07:00
James Betker	9a421a41f4	SRFlow: accomodate mismatches between global scale and flow_scale	2020-12-01 11:11:51 -07:00
James Betker	e343722d37	Add stepped rrdb	2020-12-01 11:11:15 -07:00
James Betker	2e0bbda640	Remove unused archs	2020-12-01 11:10:48 -07:00
James Betker	a1c8300052	Add mdcn	2020-11-30 16:14:21 -07:00
James Betker	1e0f69e34b	extra_conv in gn discriminator, multiframe support in rrdb.	2020-11-29 15:39:50 -07:00
James Betker	da604752e6	Misc RRDB changes	2020-11-29 12:21:31 -07:00
James Betker	a1d4c9f83c	multires rrdb work	2020-11-28 14:35:46 -07:00
James Betker	929cd45c05	Fix for RRDB scale	2020-11-27 21:37:10 -07:00
James Betker	71fa532356	Adjustments to how flow networks set size and scale	2020-11-27 21:37:00 -07:00
James Betker	6f958bb150	Maybe this is necessary after all?	2020-11-27 15:21:13 -07:00
James Betker	ef8d5f88c1	Bring split gaussian nll out of split so it can be computed accurately with the rest of the nll component	2020-11-27 13:30:21 -07:00
James Betker	4ab49b0d69	RRDB disc work	2020-11-27 12:03:08 -07:00
James Betker	6de4dabb73	Remove srflow (modified version) Starting from orig and re-working from there.	2020-11-27 12:02:06 -07:00
James Betker	fd356580c0	Play with lambdas	2020-11-26 20:30:55 -07:00
James Betker	cb045121b3	Expose srflow rrdb	2020-11-24 13:20:20 -07:00
James Betker	f6098155cd	Mods to tecogan to allow use of embeddings as input	2020-11-24 09:24:02 -07:00
James Betker	b10bcf6436	Rework stylegan_for_sr to incorporate structure as an adain block	2020-11-23 11:31:11 -07:00
James Betker	519ba6f10c	Support 2x RRDB with 4x srflow	2020-11-21 14:46:15 -07:00
James Betker	cad92bada8	Report logp and logdet for srflow	2020-11-21 10:13:05 -07:00
James Betker	c37d3faa58	More adjustments to srflow_orig	2020-11-20 19:38:33 -07:00
James Betker	d51d12a41a	Adjustments to srflow to (maybe?) fix training	2020-11-20 14:44:24 -07:00
James Betker	6c8c35ac47	Support training RRDB encoder [srflow]	2020-11-20 10:03:06 -07:00
James Betker	5ccdbcefe3	srflow_orig integration	2020-11-19 23:47:24 -07:00
James Betker	2b2d754d8e	Bring in an original SRFlow implementation for reference	2020-11-19 21:42:39 -07:00
James Betker	1e0d7be3ce	"Clean up" SRFlow	2020-11-19 21:42:24 -07:00
James Betker	d7877d0a36	Fixes to teco losses and translational losses	2020-11-19 11:35:05 -07:00
James Betker	5c10264538	Remove pyramid_disc hard dependencies	2020-11-17 18:34:11 -07:00
James Betker	6b679e2b51	Make grad_penalty available to classical discs	2020-11-17 18:31:40 -07:00
James Betker	8a19c9ae15	Add additive mode to rrdb	2020-11-16 20:45:09 -07:00
James Betker	2a507987df	Merge remote-tracking branch 'origin/gan_lab' into gan_lab	2020-11-15 16:16:30 -07:00
James Betker	931ed903c1	Allow combined additive loss	2020-11-15 16:16:18 -07:00
James Betker	4b68116977	import fix	2020-11-15 16:15:42 -07:00
James Betker	98eada1e4c	More circular dependency fixes + unet fixes	2020-11-15 11:53:35 -07:00
James Betker	e587d549f7	Fix circular imports	2020-11-15 11:32:35 -07:00

... 5 6 7 8 9 ...

1212 Commits