Falcon3-7B

More requests.
Faster responses.

Falcon3-7B handled 65% more requests per second, with 36% shorter response times. Reading-comprehension accuracy was maintained across four languages.

65%more requests completed per second
Requests completed per secondrequests
Before3.28
With GAISSA5.41

What changed

We compressed the model’s weights and the intermediate values it uses during inference. This reduced its memory footprint and let the same GPU complete more requests, with shorter response times.

Performance and resource use
Performance and resource use
MeasureBeforeWith GAISSA
Requests completedper second3.285.41
Response time95th percentile, seconds2.511.60
Model memoryat loading, GiB13.937.78
Accuracy across four languages

Both versions answered 200 Belebele reading-comprehension questions, with 50 questions in each language.

Accuracy across four languages
MeasureBeforeWith GAISSA
English94%94%
French88%86%
Spanish90%90%
Portuguese80%84%
Overall88%88.5%

97% of answers were unchanged. Each language stayed within the comparison’s quality requirements.

Test setup
Model
TII Falcon3-7B Instruct
Before
vLLM 0.24.0, BF16 weights and a fixed 512 MiB BF16 KV cache.
With GAISSA
W8A8 INT8 weights and activations using SmoothQuant and GPTQ. The BF16 KV cache and serving runtime stayed the same.
Hardware
NVIDIA GeForce RTX 4090, 24 GB.
Workload
128 requests per setup and load, with 256 input tokens and 128 output tokens. The featured comparison uses eight simultaneous requests.
Runs and reporting
One recorded paired performance comparison at each of three request loads (1, 4 and 8). These figures are not an average of repeated benchmark runs.

Quality was checked separately on the same 200 Belebele questions. Two incorrect answers became correct; one correct answer became incorrect. Model memory is the footprint at loading.

What could improve in your AI?

Tell us what you run and what you want to improve.

Let’s talk about your AI