Engines¶
TTS-Studio ships with three interchangeable engines. All of them share the same CLI surface, chapter splitting, resumability, and output handling — pick one with --engine.
kokoro |
edge |
breeze |
|
|---|---|---|---|
| Default | ✅ | ||
| Backend | Kokoro-82M | Microsoft Edge TTS | Breeze TTS 2 via mlx-audio |
| Hardware | MPS / CUDA / CPU | CPU only | Apple Silicon |
| Network | model download on first run | every run (online service) | model download on first run |
| Output | 24 kHz WAV | MP3 | 24 kHz WAV |
| Voice selection | --voice (e.g. af_heart) |
--voice |
--instruction (design), sample/ref audio (clone) |
| Extra install | --extra kokoro |
included in core | --extra breeze |
kokoro¶
The default engine: a small, high-quality local model that runs anywhere (Metal on Apple Silicon, CUDA, or CPU). It supports real-time streaming with --stream and per-chapter resumability.
1 2 | |
--voice: Kokoro voice ID (defaultaf_heart)--lang:a(en, en-us) orb(en-gb)
edge¶
A thin wrapper around edge-tts: fast, no model downloads, and a large catalogue of voices and languages. Handy when you don't want to pull in torch or MLX — it works with a bare uv sync.
1 | |
--voice: Edge TTS voice name (defaultaf_heart)- Output is MP3 rather than WAV
breeze¶
Breeze TTS 2 on Apple Silicon, with voice design and cloning:
1 2 3 4 5 6 7 8 9 | |
Key options:
--instruction: text voice description or a path to sample audio--ref-audio/--ref-text: reference audio plus its exact transcript--cfg-scale: CFG guidance scale for text instructions (try 4)--breeze-model: checkpoint dir or HF repo id (BREEZE_TTS_MODELenv var overrides)--seed: sampling seed (default 42)--breeze-workers: parallel chunk workers, each with its own model copy (~3 GB RAM)--breeze-depth-mode:cachedorcompileddepth-decoder mode
Licensing
Breeze TTS 2 model weights and self-hosted outputs are licensed for research and non-commercial use only — see the BreezeBlue license.
See Breeze TTS 2 notes on the home page for performance and voice-anchoring details.