Llais Cymraeg, Welsh text-to-speech in your pocket
As far as we can find, the first Welsh TTS that can speak in a custom voice from a few seconds of reference audio, running locally on a CPU.
Type Welsh text, choose a voice, and press Speak.
This Space runs on a shared free GPU (ZeroGPU), with a daily quota per visitor. The first request may wait a moment for a GPU slot. The model itself is small enough to run faster than real time on an ordinary desktop CPU, with no GPU at all.
Eight reference voices, four to six seconds each.
Using your own reference voice
The model speaks in any voice given a few seconds of reference audio, and that is the point of it. That control is not offered in this Space. It is in the local demo that ships with the project, where the audio never leaves your machine. Run the model yourself from the weights at EryriLabs/pocket-tts-cymraeg, or with llama-tts from the GGUF build at EryriLabs/pocket-tts-GGUF.
Responsible use
Reference voices must be your own, or used with the speaker's permission. The eight voices here are Common Voice speakers, labelled Voice 1 to Voice 8, with no identities or identifiers carried through anywhere. No claim is made that the output sounds exactly like any given person.
If a sentence comes out garbled
Press Speak again. About four draws in ten do not come out cleanly, and this is the model's main weakness. Short sentences fare much better than long ones. Temperature is fixed at 0.3, which is what the model was measured at.
Attribution
The Welsh model was trained from scratch by EryriLabs on 173 hours of openly licensed Welsh speech. The recipe and architecture are Kyutai's Pocket TTS, and the frozen Mimi codec inside the weights is Kyutai's, CC-BY-4.0, used unmodified. The reference voices and most of the training audio come from Mozilla Common Voice, CC0-1.0, with further Welsh corpora from Bangor University and partners. Full data table and licences are on the model card. Model weights CC-BY-4.0.