// Open source · OpenAI-compatible //

A sentence in 0.3 seconds.

Any voice, 600+ languages, on your own GPU. A self-hosted text-to-speech server with an OpenAI-compatible API, built on k2‑fsa's OmniVoice.

Glossy orange bars rising like a sound wave from white tiles
POST/v1/audio/speech200 OK
{ "model": "tts-1", "voice": "nova",
  "input": "Hi! I am a voice running on your own GPU. Point any OpenAI app at me, and I will speak." }
Works with any OpenAI client: Open WebUILibreChatSillyTavernAnythingLLMopenai SDKs
0.17sto speak a sentence on an RTX 5090
0.34sto speak a sentence on an RTX 3080
4GBof VRAM at peak
600+languages
// Voice cloning //

Clone a voice from one short clip.

Play the original, then the clone. The clone says words that were never recorded; everything else about the voice comes from the 15-second original.

1 · The original recording
“His prospects of success, in pleading for a favorable reception of his brother's message, were so uncertain that he refrained, in fear of raising hopes which he might not be able to justify, from taking Herbert into his confidence.”
A LibriVox reader from LibriTTS-R (CC BY 4.0)
2 · The clone, made by the API
“This voice was cloned from the recording you just heard. The words are new, but the voice, the pace and the character all come from that one short clip.”
Generated by /v1/audio/speech with default settings
// Voices //

Twelve open voices, ready to use.

OpenAI's voice names, each cloned from a LibriVox reader in LibriTTS-R. Every sample reads the same text with the default settings.

fable uses the ballad voice. Point any name at your own clip in app/openai_voices.json. Credits: Voices/ATTRIBUTION.md.

// Key features //

Built for your stack.
Runs on your GPU.

Everything an OpenAI client expects from a speech API, with voice cloning on top, running on hardware you control. No account, no per-character fees, and your text never leaves your machine.

OpenAI-compatible API

/v1/audio/speech with every OpenAI output format, speed, and all 13 OpenAI voice names. Errors follow OpenAI's format.

Voice cloning in seconds

Drop a 3 to 20 second clip in Voices/ and use its name as the voice. Add a transcript for an even better clone.

OpenAI's custom voice API

Create voices from an audio sample with a consent recording, or from a description such as female, british accent.

Voice design

Describe a voice instead of cloning one: design:female, middle-aged.

600+ languages

OmniVoice speaks over 600 languages, with any voice.

Expressive tags

Write [laughter] or [sigh] in the text and hear it.

Small footprint

About 4 GB of VRAM at peak. Unloads after 10 idle minutes.

Plays well with others

Caps its GPU memory, queues requests and recovers from out-of-memory errors.

Interactive API docs

Try every endpoint at /docs, described in OpenAPI 3.1.

// How it works //

From clone to speech
in four steps.

You need Docker and an NVIDIA GPU, with at least 6 GB of VRAM recommended. The start script does the rest.


        
// Performance //

Fast on a GPU.
Possible on a CPU.

Measured through the API with the default settings (16 steps, CUDA graphs) on an RTX 5090 and an RTX 3080. A sentence is about 4.6 seconds of speech, a paragraph about 32 seconds.

One sentence
0.17 s on an RTX 5090

Time to the complete audio file

  • About 4.6 s of speech
  • RTX 3080: 0.34 s
  • With voice design: 0.12 s (RTX 3080: 0.18 s)
  • On 4 CPU cores: 62 s
One paragraph
0.64 s on an RTX 5090

Time to the complete audio file

  • About 32 s of speech
  • RTX 3080: 1.5 s
  • On 4 CPU cores: 183 s
  • A new voice is prepared once: 6 s with a transcript, 14 s without
What you need
6 GB of VRAM, recommended

NVIDIA GPU, driver 580 or newer

See the quick start
  • Docker Desktop (Windows), or Docker Engine with the NVIDIA Container Toolkit (Linux)
  • 16 GB of RAM
  • About 40 GB of free disk for the first start
  • CPU mode works too, more than 100 times slower
// FAQ //

Questions, answered.

Something else? Open an issue on GitHub.

Can I use it commercially?

The code in this repository is MIT licensed, but the OmniVoice model is licensed CC BY-NC: non-commercial use only. Read the model's license before you use speech made with it commercially.

Does my text leave my computer?

No. Text and audio are processed on your machine. The only downloads are the models, from Hugging Face, the first time they are needed.

Which apps can use it?

Any app that supports OpenAI text to speech: Open WebUI, LibreChat, SillyTavern, AnythingLLM and the openai SDKs. Set the base URL to http://<host>:8008/v1 and the model to tts-1.

Which GPUs work?

NVIDIA GPUs, with at least 6 GB of VRAM recommended. The start script builds the kernels for your card's architecture. Tested on an RTX 3080 and an RTX 5090. On an RTX 3080, voice cloning is about 8% faster with CUDA graphs turned off: set OMNIVOICE_FLASHINFER=1 in .env.

Can I clone my own voice?

Yes. Put a 3 to 20 second clip of one speaker in Voices/, with what it says in a .txt file next to it, and use the file name as the voice. Only clone voices you have the right to use.

Does it stream audio?

Each request returns a complete audio file, in about a third of a second for a sentence on a GPU. stream_format: "sse" is not supported.