Kokoro
Generate speech
in less time.
Kokoro generated the same set of speech prompts in 46% less time on the same CPU. Automated audio checks passed; human listening checks are still pending.
What changed
We moved speech synthesis to a CPU-optimized runtime and changed how text is split into chunks and how audio is assembled. The comparison kept the same model, voice and prompts; no output caching was used.
Audio checks
The comparison used the same model, voice and prompts. Automated audio and perceptual checks passed.
Human listening checks are still pending, so this result does not establish that listeners perceive identical speech quality.
Test setup
- Model
- Kokoro 82M, English af_heart voice, normal speed, 24 kHz mono audio.
- Before
- Kokoro 0.9.4 on PyTorch 2.11.0+cpu, eager FP32 execution.
- With GAISSA
- OpenVINO 2026.2.1 FP32, overlap-add audio processing and semantic re-chunking. Caching disabled.
- Hardware
- AMD Ryzen 9 7950X CPU.
- Workload
- 50 held-out speech prompts, using the same voice in both setups.
- Runs and reporting
- Four repetitions per setup, giving 200 paired prompt evaluations. The headline compares total synthesis time across those evaluations.
The measurement covers speech synthesis. Model loading, warmup, initial text preparation, final audio joining, playback and delivery are excluded. Human listening checks remain pending.
What could improve in your AI?
Tell us what you run and what you want to improve.
