to23oise-tts

Author	SHA1	Message	Date
mrq	c12ada600b	added reset generation settings to default button, revamped utilities tab to double as plain jane voice importer (and runs through voicefixer despite it not really doing anything if your voice samples are already of decent quality anyways), ditched load_wav_to_torch or whatever it was called because it literally exists as torchaudio.load, sample voice is now a combined waveform of all your samples and will always return even if using a latents file	2023-02-14 21:20:04 +00:00
mrq	7471bc209c	Moved voices out of the tortoise folder because it kept being processed for setup.py	2023-02-10 20:11:56 +00:00
mrq	a37546ad99	owari da...	2023-02-09 01:53:25 +00:00
mrq	793515772a	un-hardcoded input output sampling rates (changing them "works" but leads to wrong audio, naturally)	2023-02-07 18:34:29 +00:00
mrq	945136330c	Forgot to rename the cached latents to the new filename	2023-02-05 23:51:52 +00:00
mrq	5bf21fdbe1	modified how conditional latents are computed (before, it just happened to only bother reading the first 102400/24000=4.26 seconds per audio input, now it will chunk it all to compute latents)	2023-02-05 23:25:41 +00:00
mrq	5c876b81f3	Added small optimization with caching latents, dropped Anaconda for just a py3.9 + pip + venv setup, added helper install scripts for such, cleaned up app.py, added flag '--low-vram' to disable minor optimizations	2023-02-04 01:50:57 +00:00
Johan Nordberg	b876a6b32c	Allow running on CPU	2022-06-11 20:03:14 +09:00
Johan Nordberg	d8f98c07b4	Remove some assumptions about working directory This allows cli tool to run when not standing in repository dir	2022-05-29 01:10:19 +00:00
Johan Nordberg	9f6ae0f0b3	Add tortoise_cli.py	2022-05-28 05:25:23 +00:00
Johan Nordberg	e34ffca8fb	Allow passing additional voice directories when loading voices	2022-05-19 21:02:11 +09:00
Danila Berezin	dc3d7b1667	Fix bug in load_voices in audio.py The read.py script did not work with pth latents, so I fix bug in audio.py. It seems that in the elif statement, instead of voice, voices should be clip, clips. And torch stack doesn't work with tuples, so I had to split this operation.	2022-05-17 18:34:54 +03:00
James Betker	75b0e03ab3	Add error message	2022-05-12 20:15:40 -06:00
James Betker	ffd0238a16	v2.2	2022-05-06 00:11:10 -06:00
James Betker	ee6f9b15ce	Use librosa for loading mp3s	2022-05-03 20:44:31 -06:00
James Betker	9acce239d3	fix paths	2022-05-02 20:56:28 -06:00
James Betker	f499d66493	misc fixes	2022-05-02 18:00:57 -06:00
James Betker	cdf44d7506	more fixes	2022-05-02 16:44:47 -06:00
James Betker	39ec1b0db5	Support totally random voices (and make fixes to previous changes)	2022-05-02 15:40:03 -06:00
James Betker	0ffc191408	Add support for extracting and feeding conditioning latents directly into the model - Adds a new script and API endpoints for doing this - Reworks autoregressive and diffusion models so that the conditioning is computed separately (which will actually provide a mild performance boost) - Updates README This is untested. Need to do the following manual tests (and someday write unit tests for this behemoth before it becomes a problem..) 1) Does get_conditioning_latents.py work? 2) Can I feed those latents back into the model by creating a new voice? 3) Can I still mix and match voices (both with conditioning latents and normal voices) with read.py?	2022-05-01 17:25:18 -06:00
James Betker	f7c8decfdb	Move everything into the tortoise/ subdirectory For eventual packaging.	2022-05-01 16:24:24 -06:00

21 Commits