If you have a large dataset, go to your dataset and rename the audio folder so it doesnt get seen by the UI. Select 10-50 audio samples from the DS audio folder, put these in the voices…
I think I fixed this with: cd text-generation-webui source ./venv/bin/activate sudo apt install ffmpeg (and maybe a pip install ffmpeg) deactivate
Fine tune a model,~50-200 epochs. If you have a large dataset, go to your dataset and rename the audio folder so it doesnt get seen by the UI. Select 10-50 audio samples from the DS audio…
Tortoise can laugh and emote if you're able to force it. Try: Lower the sample count, raise iterations, enable condition free tag text with [laughing.] HA-HA-HA-HA-HA-HA-HA-HA-HA-HA-HA-HA,…
Well that formatting looks like shit, and I pasted the wrong code. Let's try again ~Line 2020 src/utils.py
from faster_whisper import WhisperModel
device = "cuda" if get_device_name() ==…