vall-e

mrq/vall-e

Author	SHA1	Message	Date
mrq	e50edc3b48	added a flag to convert to a HF compatible model on export by stitching things	2024-06-03 22:34:47 -05:00
mrq	934672252b	feverish cleanup	2024-06-03 21:28:49 -05:00
mrq	7feeb944a0	probably insane with even entertaining going this route	2024-06-03 20:26:27 -05:00
mrq	c2a436d368	somehow between training sessions grad_norm = None even though it worked before	2024-06-02 08:29:27 -05:00
mrq	c1fcd889d5	reverted automatically disabling split loss calc, since it seems that it's actually cacling loss on prom causes the oddities, maybe	2024-06-01 12:34:59 -05:00
mrq	8cf176ab46	ugh	2024-06-01 10:46:42 -05:00
mrq	827cf632e7	report current loss scale and adjust grad norm by loss scale (for deepspeed)	2024-06-01 10:44:32 -05:00
mrq	d0ebce6bac	ugh	2024-06-01 10:30:13 -05:00
mrq	39bc019142	actually save per-rank sampler states	2024-06-01 09:46:32 -05:00
mrq	74df2f5332	split sampler dict by global_rank, also handle splitting dataset paths by global_rank if sampler_type == path (because I do not trust DistributedSampler) (need to test)	2024-06-01 09:29:49 -05:00
mrq	31785f4eeb	actually don't default to compute split losses, test bitnet model doesn't seem to be doing things right (despite debug printouts showing theyre roughly the same logit/loss sequences, could just be bitnet linears being not up to par on actual models)	2024-06-01 09:12:51 -05:00
mrq	e9c87060df	oops	2024-05-31 22:22:28 -05:00
mrq	b482ca19ff	added model config option to set KV head count for MQA/GQA instead of MHA for llama-based models (i think its very negligible both ways on such a small model size)	2024-05-31 19:32:37 -05:00
mrq	e15c6c74c3	correctness	2024-05-30 20:50:45 -05:00
mrq	da473295b7	better way to compute per-segment losses	2024-05-28 19:29:54 -05:00
mrq	6c49ad06a3	forgot to reinclude mult by loss factors	2024-05-27 20:40:21 -05:00
mrq	b82f0d5c0c	finally nailed the issue that caused logging to break on one machine but not another (bitnet includes zetascale which is a parasite that will break logging)	2024-05-27 19:47:58 -05:00
mrq	c0ac84c795	uh	2024-05-27 19:05:56 -05:00
mrq	197d517181	ugh	2024-05-27 17:09:35 -05:00
mrq	5af6f41c94	added loss calcs against prom (requires the right settings for not shit results, disabled by default)	2024-05-27 08:43:00 -05:00
mrq	05cd8b797e	nevermind it breaks training	2024-05-25 18:03:43 -05:00
mrq	85f9684720	some cleanup	2024-05-25 17:46:52 -05:00
mrq	d760924719	added kludgy eval only so I don't have to start training, type eval, stop training, then delete the logs for that session	2024-05-25 17:39:51 -05:00
mrq	ddbacde0d1	DAC just doesn't work well enough......	2024-05-25 11:07:52 -05:00
mrq	e3ef89f5aa	100x better for subtrain/eval to be by group instead	2024-05-19 16:40:14 -05:00
mrq	458b95d196	added option to split between text loss and audio loss (to-do: document this better), because it may or may not be a problem with LLaMA-backed models because my loss hovers around 3.9 / 56% accuracy despite sounding decent at the moment	2024-05-19 11:23:56 -05:00
mrq	74e531d391	ugh	2024-05-18 12:02:56 -05:00
mrq	59ef9461f8	ugh	2024-05-18 10:13:58 -05:00
mrq	4bc7e5a6d1	fix loading without needing an hdf5 dataset already prepped (and some other incidental speedups during dataloader prep)	2024-05-18 07:14:26 -05:00
mrq	d88a5ca183	ugh	2024-05-16 07:25:33 -05:00
mrq	d9aabfa3ae	final tweaks, hopefully, again	2024-05-15 23:04:19 -05:00
mrq	8d79f78e0a	god I need to replace omegaconf	2024-05-12 14:01:52 -05:00
mrq	5eb5db7f7f	just don't use DAC 24Khz, it's bad	2024-05-12 13:41:17 -05:00
mrq	230da8b559	should be the final things to scramble around for, DAC's 24KHz model is unusable for this, but both encodec's 24KHz and DAC's 44KHz work	2024-05-12 13:22:08 -05:00
mrq	2437a86efa	ugh	2024-05-12 13:02:15 -05:00
mrq	4f1593c8db	a bunch of shit to salvage my old encodec-quantized audio because dac-encoded audio just does not want to converge	2024-05-12 10:17:29 -05:00
mrq	917eeb40d2	ughhh	2024-05-12 08:22:39 -05:00
mrq	9910c75d5a	checkpointing for bitnet impl	2024-05-12 07:52:54 -05:00
mrq	14709ac67f	ughh	2024-05-12 07:30:59 -05:00
mrq	3774fcbdee	ugh	2024-05-11 22:58:38 -05:00
mrq	856545f8bb	nan loss detection (should have added it earlier), loss scaling for local backend + fp16	2024-05-11 22:23:29 -05:00
mrq	a755eb3c62	ugh	2024-05-11 17:34:45 -05:00
mrq	88e9b9caff	local ddp fix	2024-05-11 17:29:01 -05:00
mrq	3337c69e5a	leverage between xformers and `torch.backends.cuda.sdp_kernel` for attention	2024-05-11 17:14:05 -05:00
mrq	d33c7bb7cf	ugh	2024-05-11 16:47:19 -05:00
mrq	0b6499601b	sanitizing	2024-05-11 16:31:05 -05:00
mrq	71e373064f	remove redundant loss, tweak readme	2024-05-11 15:02:47 -05:00
mrq	04a80d6b55	maybe it's better to be more explicit in deepspeed configs	2024-05-11 13:57:43 -05:00
mrq	4d93a16ef7	might just be better to explicitly define prompt duration ranges, especially under a "train small contexts then increase it" training paradigm	2024-05-11 09:50:54 -05:00
mrq	bd0a36ba8d	I swear I keep seeing tqdm flicker back a number	2024-05-10 18:36:01 -05:00

1 2 3 4 5 ...

310 Commits