Kokoro

Generate speech
in less time.

Kokoro generated the same set of speech prompts in 46% less time on the same CPU. Automated audio checks passed; human listening checks are still pending.

46%less time to generate speech
Total speech-generation timeseconds
Before251.27
With GAISSA135.52

What changed

We moved speech synthesis to a CPU-optimized runtime and changed how text is split into chunks and how audio is assembled. The comparison kept the same model, voice and prompts; no output caching was used.

Audio checks

The comparison used the same model, voice and prompts. Automated audio and perceptual checks passed.

Human listening checks are still pending, so this result does not establish that listeners perceive identical speech quality.

Test setup
Model
Kokoro 82M, English af_heart voice, normal speed, 24 kHz mono audio.
Before
Kokoro 0.9.4 on PyTorch 2.11.0+cpu, eager FP32 execution.
With GAISSA
OpenVINO 2026.2.1 FP32, overlap-add audio processing and semantic re-chunking. Caching disabled.
Hardware
AMD Ryzen 9 7950X CPU.
Workload
50 held-out speech prompts, using the same voice in both setups.
Runs and reporting
Four repetitions per setup, giving 200 paired prompt evaluations. The headline compares total synthesis time across those evaluations.

The measurement covers speech synthesis. Model loading, warmup, initial text preparation, final audio joining, playback and delivery are excluded. Human listening checks remain pending.

What could improve in your AI?

Tell us what you run and what you want to improve.

Let’s talk about your AI