BGE

More searches.
Same hardware.

BGE completed 2.3 times as many searches per second on the same CPU. Search quality stayed within the target range, with small changes in result rankings.

2.3×as many searches completed per second
Searches completed per secondsearches
Before433.5
With GAISSA998.1

What changed

We moved the model to a CPU-optimized runtime and tuned document preparation and query processing separately. This let the same CPU handle more searches against the prepared document collection.

Search relevance

Both setups searched the same documents using 300 SciFact queries. The quality measures changed slightly and stayed within the target range.

Search relevance
MeasureBeforeWith GAISSA
Recall in the top 10higher is better83.62%83.29%
Ranking qualitynDCG@10, higher is better0.71270.7120
Test setup
Model
BAAI BGE-small-en-v1.5
Before
PyTorch FP32, using the strongest conventional configuration selected on the training data. The runtime version is not included in the retained comparison record.
With GAISSA
OpenVINO FP32 with separate execution settings for documents and queries.
Hardware
AMD Ryzen 9 7950X, 16 pinned CPU threads.
Workload
300 held-out SciFact queries against the same prepared document collection.
Runs and reporting
Five paired repetitions, with the run order alternated. Before and after values are the median for each setup.

Timing covers query encoding and search against prepared documents. Document preparation, model loading and warmup are excluded. This measures batch query throughput, not individual request latency.

What could improve in your AI?

Tell us what you run and what you want to improve.

Let’s talk about your AI