NVIDIA Parakeet
More audio.
Same GPU.
Parakeet processed 2.4 times as much audio per second while maintaining transcription accuracy. The optimized setup also used less GPU memory.
Audio processed per secondseconds of audio
What changed
We changed how Parakeet processes audio so it spends less time on execution overhead and keeps working efficiently as streams finish. The model and GPU stayed the same.
Performance and resource use
| Measure | Before | With GAISSA |
|---|---|---|
| Audio processedseconds of audio per second | 7.89 | 18.95 |
| GPU memorypeak process allocation, MiB | 1,509.90 | 1,282.46 |
Transcription accuracy
Both setups transcribed the same English recordings. The optimized setup made slightly fewer transcription errors.
| Measure | Before | With GAISSA |
|---|---|---|
| Word error ratelower is better | 2.97% | 2.86% |
Test setup
- Model
- NVIDIA Parakeet Unified EN 0.6B
- Before
- NeMo 3.0.0 and PyTorch 2.11.0, BF16, eager streaming execution.
- With GAISSA
- Compiled audio encoder with active streams regrouped as recordings finish. BF16 precision and streaming context were retained.
- Hardware
- NVIDIA GeForce RTX 3090 Ti, 24 GiB.
- Workload
- 16 simultaneous streams; 128 held-out LibriSpeech test-clean recordings, totaling 1,032 seconds of audio.
- Runs and reporting
- Five paired repetitions, with the run order alternated. Before and after values are the median for each setup.
Timing covers model processing. Network delivery is outside this comparison. GPU memory is peak process allocation.
What could improve in your AI?
Tell us what you run and what you want to improve.
