NVIDIA Parakeet

More audio.
Same GPU.

Parakeet processed 2.4 times as much audio per second while maintaining transcription accuracy. The optimized setup also used less GPU memory.

2.4×as much audio processed per second
Audio processed per secondseconds of audio
Before7.89
With GAISSA18.95

What changed

We changed how Parakeet processes audio so it spends less time on execution overhead and keeps working efficiently as streams finish. The model and GPU stayed the same.

Performance and resource use
Performance and resource use
MeasureBeforeWith GAISSA
Audio processedseconds of audio per second7.8918.95
GPU memorypeak process allocation, MiB1,509.901,282.46
Transcription accuracy

Both setups transcribed the same English recordings. The optimized setup made slightly fewer transcription errors.

Transcription accuracy
MeasureBeforeWith GAISSA
Word error ratelower is better2.97%2.86%
Test setup
Model
NVIDIA Parakeet Unified EN 0.6B
Before
NeMo 3.0.0 and PyTorch 2.11.0, BF16, eager streaming execution.
With GAISSA
Compiled audio encoder with active streams regrouped as recordings finish. BF16 precision and streaming context were retained.
Hardware
NVIDIA GeForce RTX 3090 Ti, 24 GiB.
Workload
16 simultaneous streams; 128 held-out LibriSpeech test-clean recordings, totaling 1,032 seconds of audio.
Runs and reporting
Five paired repetitions, with the run order alternated. Before and after values are the median for each setup.

Timing covers model processing. Network delivery is outside this comparison. GPU memory is peak process allocation.

What could improve in your AI?

Tell us what you run and what you want to improve.

Let’s talk about your AI