Falcon3-7B
More requests.
Faster responses.
Falcon3-7B handled 65% more requests per second, with 36% shorter response times. Reading-comprehension accuracy was maintained across four languages.
What changed
We compressed the model’s weights and the intermediate values it uses during inference. This reduced its memory footprint and let the same GPU complete more requests, with shorter response times.
Performance and resource use
| Measure | Before | With GAISSA |
|---|---|---|
| Requests completedper second | 3.28 | 5.41 |
| Response time95th percentile, seconds | 2.51 | 1.60 |
| Model memoryat loading, GiB | 13.93 | 7.78 |
Accuracy across four languages
Both versions answered 200 Belebele reading-comprehension questions, with 50 questions in each language.
| Measure | Before | With GAISSA |
|---|---|---|
| English | 94% | 94% |
| French | 88% | 86% |
| Spanish | 90% | 90% |
| Portuguese | 80% | 84% |
| Overall | 88% | 88.5% |
97% of answers were unchanged. Each language stayed within the comparison’s quality requirements.
Test setup
- Model
- TII Falcon3-7B Instruct
- Before
- vLLM 0.24.0, BF16 weights and a fixed 512 MiB BF16 KV cache.
- With GAISSA
- W8A8 INT8 weights and activations using SmoothQuant and GPTQ. The BF16 KV cache and serving runtime stayed the same.
- Hardware
- NVIDIA GeForce RTX 4090, 24 GB.
- Workload
- 128 requests per setup and load, with 256 input tokens and 128 output tokens. The featured comparison uses eight simultaneous requests.
- Runs and reporting
- One recorded paired performance comparison at each of three request loads (1, 4 and 8). These figures are not an average of repeated benchmark runs.
Quality was checked separately on the same 200 Belebele questions. Two incorrect answers became correct; one correct answer became incorrect. Model memory is the footprint at loading.
What could improve in your AI?
Tell us what you run and what you want to improve.
