BGE
More searches.
Same hardware.
BGE completed 2.3 times as many searches per second on the same CPU. Search quality stayed within the target range, with small changes in result rankings.
Searches completed per secondsearches
What changed
We moved the model to a CPU-optimized runtime and tuned document preparation and query processing separately. This let the same CPU handle more searches against the prepared document collection.
Search relevance
Both setups searched the same documents using 300 SciFact queries. The quality measures changed slightly and stayed within the target range.
| Measure | Before | With GAISSA |
|---|---|---|
| Recall in the top 10higher is better | 83.62% | 83.29% |
| Ranking qualitynDCG@10, higher is better | 0.7127 | 0.7120 |
Test setup
- Model
- BAAI BGE-small-en-v1.5
- Before
- PyTorch FP32, using the strongest conventional configuration selected on the training data. The runtime version is not included in the retained comparison record.
- With GAISSA
- OpenVINO FP32 with separate execution settings for documents and queries.
- Hardware
- AMD Ryzen 9 7950X, 16 pinned CPU threads.
- Workload
- 300 held-out SciFact queries against the same prepared document collection.
- Runs and reporting
- Five paired repetitions, with the run order alternated. Before and after values are the median for each setup.
Timing covers query encoding and search against prepared documents. Document preparation, model loading and warmup are excluded. This measures batch query throughput, not individual request latency.
What could improve in your AI?
Tell us what you run and what you want to improve.
