to23oise-tts

Author	SHA1	Message	Date
James Betker	8139afd0e5	Remove CVVP After training a similar model for a different purpose, I realized that this model is faulty: the contrastive loss it uses only pays attention to high-frequency details which do not contribute meaningfully to output quality. I validated this by comparing a no-CVVP output with a baseline using tts-scores and found no differences.	2022-05-17 12:21:25 -06:00
James Betker	aef86d21bf	Add a way to get deterministic behavior from tortoise and add debug states for reporting	2022-05-17 12:11:18 -06:00
James Betker	fda5130819	Add support for multiple output candidates in do_tts.	2022-05-12 11:25:35 -06:00
James Betker	14617f8963	Fix default output path	2022-05-02 21:37:39 -06:00
James Betker	f499d66493	misc fixes	2022-05-02 18:00:57 -06:00
James Betker	39ec1b0db5	Support totally random voices (and make fixes to previous changes)	2022-05-02 15:40:03 -06:00
James Betker	b1fc2b13c9	add support for specifying the model_dir	2022-05-01 17:29:25 -06:00
James Betker	0ffc191408	Add support for extracting and feeding conditioning latents directly into the model - Adds a new script and API endpoints for doing this - Reworks autoregressive and diffusion models so that the conditioning is computed separately (which will actually provide a mild performance boost) - Updates README This is untested. Need to do the following manual tests (and someday write unit tests for this behemoth before it becomes a problem..) 1) Does get_conditioning_latents.py work? 2) Can I feed those latents back into the model by creating a new voice? 3) Can I still mix and match voices (both with conditioning latents and normal voices) with read.py?	2022-05-01 17:25:18 -06:00
James Betker	f7c8decfdb	Move everything into the tortoise/ subdirectory For eventual packaging.	2022-05-01 16:24:24 -06:00

9 Commits