Qwen2.5-7B

Less memory.
Same quality.

Qwen2.5-7B used 43% less memory to load the model, while maintaining answer accuracy. That leaves more GPU memory available for running requests.

43%less memory to load the model
Model memoryGiB at loading
Before14.29
With GAISSA8.18

What changed

We compressed the model’s weights to an 8-bit format while keeping its request cache at BF16 precision. The smaller model footprint leaves more GPU memory available for serving requests.

Answer accuracy

Both versions answered the same 200 questions covering algebra, computer security, computer science and logical reasoning.

93% of answers were unchanged. Six incorrect answers became correct, while three correct answers became incorrect.

Answer accuracy
MeasureBeforeWith GAISSA
Correct answers158 / 200161 / 200
Accuracy79.0%80.5%
Memory use across four language models
Memory used to load each model, in GiB.
ModelBeforeWith GAISSAReduction
Qwen2.5-7B14.298.1843%
Mistral-7B v0.313.517.0148%
Phi-3.5 Mini7.143.8047%
OLMo 2 7B13.607.6044%
Test setup
Model
Qwen2.5-7B Instruct
Before
vLLM 0.24.0, BF16 weights and BF16 KV cache.
With GAISSA
Offline FP8 weights and dynamic FP8 activations, with the BF16 KV cache retained.
Hardware
NVIDIA GeForce RTX 4090, 24 GB.
Workload
Runtime-reported model memory at loading. Quality was checked separately on 200 MMLU questions.
Runs and reporting
Recorded model-loading footprint for each setup. No repeated-run aggregate is available for this memory comparison.

The memory comparison measures the model’s footprint at loading, not total GPU memory during use. Accuracy was checked using the same task and answer checks before and after optimization.

What could improve in your AI?

Tell us what you run and what you want to improve.

Let’s talk about your AI