Commit Graph

  • 52d13b321f I rather have it default to non-strict loading instead so I can clean up YAMLs mrq 2024-07-30 22:24:38 -0500
  • d7c6be6f78 fix weird regression in handling checkpoints when backend is local, but deepspeed checkpoints are in (it was handled with LoRA loading but not real loading...) mrq 2024-07-30 22:15:56 -0500
  • 07f8e2ad06 added option to set the causal size (how many tokens to sample per AR step), but requires the model to be trained for this (which explains why recurrent chunk sampling just doesn't work for the retnet tests, obvious in hindsight) mrq 2024-07-30 20:53:51 -0500
  • ebf848d249 possible speedup for samplers that require a list of previous tokens (the DRY sampler made me realize that I should copy the tolist() thing from the rep pen sampler for everything else) mrq 2024-07-29 20:23:26 -0500
  • 55b0121b1a trying (and failing) to nail a weird regression in fancier attentions mrq 2024-07-29 19:53:37 -0500
  • c2f5b916fc added what I think is DRY sampling mrq 2024-07-29 19:15:07 -0500
  • ce8bb1e4f7 sanity cleanups with weird off-by-one-ness, cleaned up and validated vall_e.models.experimental works again mrq 2024-07-27 15:36:05 -0500
  • 06e948aec1 suppress warning on exit about distributed not being cleaned up (because I updated my system) mrq 2024-07-25 16:50:47 -0500
  • 682e4387dc oops (fixed proms being erased from a config oversight) mrq 2024-07-25 12:39:57 -0500
  • 1acb0e9c84 added experimental training setting to perform token dropout to MAYBE compensate for errors from the preceding RVQ level (two types: token error offset, token dropout embedding replace) mrq 2024-07-24 19:35:17 -0500
  • 611a1c4bdc might help mrq 2024-07-22 20:57:01 -0500
  • 188d116222 some weird fixes for an equally weird regression with LoRA loading mrq 2024-07-22 20:47:24 -0500
  • e33c4b0cb1 oops mrq 2024-07-22 19:38:39 -0500
  • 75b04686f8 added prom-less training / inferencing, some other things mrq 2024-07-22 19:36:07 -0500
  • 491ae2a684 some insanity for sanity checks (some phonemes from phonemizing japanese are not in my tokenizer...) mrq 2024-07-22 00:30:40 -0500
  • ad024f400f actually pass language into dataset process script, fix coercing japanese into hiragana because espeak does not like kanji mrq 2024-07-21 23:21:37 -0500
  • 3e5ca3a201 more demo page tweaks mrq 2024-07-21 19:31:13 -0500
  • 7366f36f81 oops mrq 2024-07-21 19:17:25 -0500
  • e19aa643a6 cleaned up demo page creation, added option to pass in RVQ level sampling distribution for training mrq 2024-07-21 19:12:03 -0500
  • ba7ee8c0ee added demo link to readme mrq 2024-07-19 21:22:30 -0500
  • 9ec88d9444 validated passing URI path for assets instead of base64 encoding them mrq 2024-07-19 21:07:17 -0500
  • d87b492295 added rudimentary demo page creator (currently just embeds base64 wavs into the page, need to test not doing that) mrq 2024-07-19 20:49:40 -0500
  • d53038a9e4 actually have split classifiers working mrq 2024-07-19 15:33:31 -0500
  • 692d09f9c1 eval/validation fix for SpeechX tasks mrq 2024-07-19 09:16:37 -0500
  • 28a674e0f1 fixes... mrq 2024-07-18 23:25:32 -0500
  • 39f961abcd test trainer (vall_e.models.ar_nar) tests some SpeechX features mrq 2024-07-18 18:46:45 -0500
  • 83a0954f85 fixes for re-introducing SpeechX tasks (need to actually validate if these all do the right things) mrq 2024-07-18 17:16:32 -0500
  • bccbb77a1a added option to either naively concat codes to concat audio waveforms (prior behavior) or to decode => concat => encode instead (although this only currently happens for prom sampling if an utternace is too small) mrq 2024-07-18 16:48:41 -0500
  • 97e768601c re-introducing SpeechX tasks (need to validate them all, everything works with base tts anyways) mrq 2024-07-18 16:16:14 -0500
  • c2b8035e74 oops, kept forgetting to actually pass in lang/tone tokens (despite not really using these at the moment) mrq 2024-07-18 14:18:34 -0500
  • 22fe53508c added experimental disjointed position IDs (because I *think* this might help because technically a sequence is made up of several parts, and the position embeddings shouldn't be unified) mrq 2024-07-16 19:52:41 -0500
  • fe0f235335 mechanism to store the model config inside the weights and load them, some other things to allow LoRA training on the RetNet (gradient checkpointing will gripe about inputs not having require_grad and nothing seems to remedy it) mrq 2024-07-16 18:23:13 -0500
  • 3acc54df22 allow loading a different model within the web ui (apparently I did not have the web UI in the documentation) mrq 2024-07-15 19:59:48 -0500
  • 7b210d9738 sanity cleanup mrq 2024-07-04 15:58:08 -0500
  • 1ecf2793f4 (commented-out) support for facebookresearch/AudioDec, but support really didn't wow me (so I commented it out until I figure out why my output audio is super crusty with AudioDec) mrq 2024-07-04 15:40:51 -0500
  • db62e55a38 oops, I forgot to use the new thing for audio_backend mrq 2024-07-04 14:54:11 -0500
  • f770467eb3 stuff mrq 2024-07-01 18:13:29 -0500
  • 312a8e3ead add shuffle to samplers that can support it mrq 2024-06-30 11:36:46 -0500
  • 396af541c5 ugh mrq 2024-06-30 11:11:58 -0500
  • dced595391 more cleanup mrq 2024-06-30 11:00:12 -0500
  • bc2a6fa756 sanity cleanup: moved experimental features under its own thing mrq 2024-06-30 10:37:33 -0500
  • b21f74a5c5 added summing of external embeddings (at this point i dont think any amount of cope bandaids will get DAC to train nicely, I think the RVQ levels the NAR tends add too much noise if they're not accurate) mrq 2024-06-29 23:42:30 -0500
  • 793ccb16fb ugh mrq 2024-06-29 22:14:35 -0500
  • 2808f881c8 cleaned up subjugated audio embedding into a flag, flag can also have it include the original, underlying embedding as well (it seems to do better when set to inclusive) mrq 2024-06-29 21:46:35 -0500
  • ec5eaebcbc experimental method of using DACs quantizer ""embeddings"" to see if it helps with model quality mrq 2024-06-29 19:46:11 -0500
  • a8718d35a4 nasty bandaid because some of my DAC dataset only has 8 RVQ levels instead of the full 9 mrq 2024-06-29 10:16:37 -0500
  • c4dd523b6f change from chunk-slicing paths for distributed dataloader to instead interleave mrq 2024-06-29 10:10:35 -0500
  • dd40463803 limit eval size because the training batch size seems to be used for the eval dataloader, somehow (bandaid) mrq 2024-06-29 09:11:28 -0500
  • 591d3ac848 have eval dataloader use eval batch size for batchedordersampler mrq 2024-06-28 22:44:00 -0500
  • 1a392b69f6 local training backend should be a bit more aware of variable batch sizes, maybe mrq 2024-06-28 22:39:05 -0500
  • 83075c1505 sort duration buckets to ensure that paths sorted-by-duration are actually sorted by duration (because i didnt know that python dicts can have non-strings as keys), added batching samples based on total duration to ensure best training throughput mrq 2024-06-28 22:28:54 -0500
  • 5176ced35f readme tweaks mrq 2024-06-28 21:02:54 -0500
  • 8fffb94964 backport fix from tortoise_tts with local trainer + loading state when training lora mrq 2024-06-25 13:41:29 -0500
  • 62a53eed64 fixed deducing tokenizer path, added option to default to naive tokenizer (for old models, like ar+nar-retnet-8) mrq 2024-06-18 22:11:14 -0500
  • 8a986eb480 load exported LoRA weights if exists (to-do: make a better LoRA loading mechanism) mrq 2024-06-18 21:45:46 -0500
  • 2bfe786ebd ban stop token for NAR levels (because sometimes it gets sampled and causes problems) mrq 2024-06-17 22:14:43 -0500
  • 7cfb78fa64 enable LoRA for targetted RVQ levels (to experiment with, seems to help) mrq 2024-06-17 21:45:03 -0500
  • 7047fcc6e2 actually make deepspeed work with LoRAs mrq 2024-06-17 13:55:37 -0500
  • 1d159b1476 updated export routine to split LoRA weights from the state dict (should work with deepspeed) mrq 2024-06-17 13:28:18 -0500
  • 726a4b613f naive, rudimentary DeepSpeed support (just live with the LoRA weights living with the original weights, they can be split later) mrq 2024-06-17 13:17:24 -0500
  • bd0bc10ec0 added LoRA policy to decide what layer of the model gets adapted based on simple inclusion/exclusion terms mrq 2024-06-17 13:05:06 -0500
  • be051d9544 added other LoRA method using parametrization rather than linear injection mrq 2024-06-17 09:58:34 -0500
  • 45a39fb79f very rudimentary lora support (no deepspeed support, tested training and saving but not loading yet) mrq 2024-06-17 00:09:16 -0500
  • 19410a919e ugh mrq 2024-06-15 12:29:03 -0500
  • d343bde09b residual_in_fp32=False for mamba arch backends because it breaks the classifier (output projection / lm head / what-have-you) under AMP mrq 2024-06-15 12:08:03 -0500
  • ccb14c06ef mamba2-hf using vasqu/mamba2-torch because it lets me use mamba2 without triton ops (training with my 4xV100s are not happy with mamba2 because of triton) mrq 2024-06-14 19:42:17 -0500
  • 31f71fa134 sampler update (some brainworm just never actually had a sampler for sample_type=path) mrq 2024-06-14 16:55:40 -0500
  • b3b67f34ac added option to sort paths by durations to better group equally lengthed sequences together (and there was maybe a logic error from creating the samplers and then interleave-reordering paths, desyncing them, maybe) mrq 2024-06-13 22:37:34 -0500
  • 83eab4fa59 actually going for the suggested "2x layers, no intermediate scaling" is wrong for VALL-E, directly copying the normal transformer structure fixes mamba2 performance in the test trainer mrq 2024-06-13 20:08:22 -0500
  • ff97e7480d fixed pip shitting itself on setup mrq 2024-06-13 13:03:36 -0500
  • 26da24fd8d mamba updated to fix that pesky NaN error during training mrq 2024-06-13 12:38:33 -0500
  • bcf3910a17 the NAR only dream is dead (it just won't work) mrq 2024-06-12 19:49:47 -0500
  • a9353cf9fa ugh mrq 2024-06-12 00:14:29 -0500
  • cca542a4c0 ugh mrq 2024-06-11 23:59:28 -0500
  • 65a8960305 option to split classifier per-level instead of sharing one (at this point I'm just scrambling to try and cope with training a DAC model, the NAR is being a pain) mrq 2024-06-11 22:28:59 -0500
  • a7a6e0ac76 validated that inferencing works, changed some defaults (NAR benefits from greedy sampling) mrq 2024-06-09 17:11:38 -0500
  • 234f9efc6e ugh mrq 2024-06-09 11:39:43 -0500
  • 132a02c48b sanity cleanup, backup config yaml for each log file mrq 2024-06-09 11:22:52 -0500
  • 8d92dac829 forgot I renamed this mrq 2024-06-09 11:12:30 -0500
  • 80f9530840 ugh mrq 2024-06-09 01:43:44 -0500
  • 5c732b72ee ugh mrq 2024-06-08 20:34:00 -0500
  • 8d068fa3f9 reticulating splines mrq 2024-06-08 20:30:15 -0500
  • ead3e2f0cb ugh mrq 2024-06-08 16:14:57 -0500
  • b072f9b96b fixes mrq 2024-06-08 16:01:34 -0500
  • 58fb0a84db added experimental NAR only model (inferences text length, need more experimenting), AudioEmbedding logic cleanup (I still think it's being done wrong) mrq 2024-06-08 15:42:02 -0500
  • e35a91c67a ugh mrq 2024-06-07 21:56:14 -0500
  • 7d6fff24f9 un-tensor'd quant_level marker since it doesn't need to be one (I forgot why I had it as one but nothing seems to need it as a tensor that didn't already make it one) mrq 2024-06-07 20:46:22 -0500
  • b0158a61d5 fixed some logic errors with training (grabbing wrong quant level...) mrq 2024-06-07 20:34:36 -0500
  • eafa622be2 I forgot the actual reason I was cleaning things up was to re-include prom loss calculation (I realized the reason I did this was because of an prom embedding oversight, it seems to work now) mrq 2024-06-07 20:29:25 -0500
  • da8242d086 finally got around to removing omegaconf mrq 2024-06-07 20:23:53 -0500
  • 4ade2b60ee ugh mrq 2024-06-06 21:57:11 -0500
  • f9f309281a ugh mrq 2024-06-06 20:55:27 -0500
  • a5c90348d9 head hurt mrq 2024-06-06 20:51:31 -0500
  • 516b0894d7 m mrq 2024-06-06 19:41:26 -0500
  • ee25d2e62e removed the need to supply targ_list + different AudioEmbedding + other things mrq 2024-06-06 18:52:41 -0500
  • fcac9503e2 cleanup mrq 2024-06-06 13:08:02 -0500
  • b2194b859a re-added loading multiple models because I'm now entertaining having split AR/NAR models again (and need a way to load both at once) mrq 2024-06-06 09:48:43 -0500
  • b05a905b95 ugh mrq 2024-06-05 21:02:05 -0500
  • 4073656293 oops mrq 2024-06-05 20:53:10 -0500
  • ff6fe6f1bc cleanup mrq 2024-06-05 20:30:43 -0500