Clone a voice from one short clip.
Play the original, then the clone. The clone says words that were never recorded; everything else about the voice comes from the 15-second original.
“His prospects of success, in pleading for a favorable reception of his brother's message, were so uncertain that he refrained, in fear of raising hopes which he might not be able to justify, from taking Herbert into his confidence.”
“This voice was cloned from the recording you just heard. The words are new, but the voice, the pace and the character all come from that one short clip.”
Twelve open voices, ready to use.
OpenAI's voice names, each cloned from a LibriVox reader in LibriTTS-R. Every sample reads the same text with the default settings.
fable uses the ballad voice. Point any name at your own clip in
app/openai_voices.json. Credits:
Voices/ATTRIBUTION.md.
Built for your stack.
Runs on your GPU.
Everything an OpenAI client expects from a speech API, with voice cloning on top, running on hardware you control. No account, no per-character fees, and your text never leaves your machine.
OpenAI-compatible API
/v1/audio/speech with every OpenAI output format, speed, and all 13 OpenAI voice
names. Errors follow OpenAI's format.
Voice cloning in seconds
Drop a 3 to 20 second clip in Voices/ and use its name as the voice. Add a transcript for an
even better clone.
OpenAI's custom voice API
Create voices from an audio sample with a consent recording, or from a description such as female, british accent.
Voice design
Describe a voice instead of cloning one: design:female, middle-aged.
600+ languages
OmniVoice speaks over 600 languages, with any voice.
Expressive tags
Write [laughter] or [sigh] in the text and hear it.
Small footprint
About 4 GB of VRAM at peak. Unloads after 10 idle minutes.
Plays well with others
Caps its GPU memory, queues requests and recovers from out-of-memory errors.
Interactive API docs
Try every endpoint at /docs, described in OpenAPI 3.1.
From clone to speech
in four steps.
You need Docker and an NVIDIA GPU, with at least 6 GB of VRAM recommended. The start script does the rest.
Fast on a GPU.
Possible on a CPU.
Measured through the API with the default settings (16 steps, CUDA graphs) on an RTX 5090 and an RTX 3080. A sentence is about 4.6 seconds of speech, a paragraph about 32 seconds.
Time to the complete audio file
- About 4.6 s of speech
- RTX 3080: 0.34 s
- With voice design: 0.12 s (RTX 3080: 0.18 s)
- On 4 CPU cores: 62 s
Time to the complete audio file
- About 32 s of speech
- RTX 3080: 1.5 s
- On 4 CPU cores: 183 s
- A new voice is prepared once: 6 s with a transcript, 14 s without
- Docker Desktop (Windows), or Docker Engine with the NVIDIA Container Toolkit (Linux)
- 16 GB of RAM
- About 40 GB of free disk for the first start
- CPU mode works too, more than 100 times slower
Can I use it commercially?
The code in this repository is MIT licensed, but the OmniVoice model is licensed CC BY-NC: non-commercial use only. Read the model's license before you use speech made with it commercially.
Does my text leave my computer?
No. Text and audio are processed on your machine. The only downloads are the models, from Hugging Face, the first time they are needed.
Which apps can use it?
Any app that supports OpenAI text to speech: Open WebUI, LibreChat, SillyTavern, AnythingLLM and the
openai SDKs. Set the base URL to http://<host>:8008/v1 and the model to
tts-1.
Which GPUs work?
NVIDIA GPUs, with at least 6 GB of VRAM recommended. The start script builds the kernels for your card's architecture.
Tested on an RTX 3080 and an RTX 5090. On an RTX 3080, voice cloning is about 8% faster with CUDA graphs
turned off: set OMNIVOICE_FLASHINFER=1 in .env.
Can I clone my own voice?
Yes. Put a 3 to 20 second clip of one speaker in Voices/, with what it says in a
.txt file next to it, and use the file name as the voice. Only clone voices you have the
right to use.
Does it stream audio?
Each request returns a complete audio file, in about a third of a second for a sentence on a GPU.
stream_format: "sse" is not supported.