Vector engine benchmark: eshoponweb
Run started 3 Oct 2026, 18:39 UTC. Data: eShopOnWeb, 524 vectors. Queries: 20 labelled questions.
Machine: CPU Intel(R) Xeon(R) CPU E5-1620 v3 @ 3.50GHz; Logical CPUs 8; RAM (GiB) 62.7; OS Ubuntu 24.04.5 LTS, kernel 6.8.0-142-generic; Governor not recorded; Partition not recorded; Release build.
Warm-up, as this run's notes record it: Latency is client-side wall time around each search (network and driver included), one query at a time, after 20 warm-up queries; at least 200 samples (small query sets are repeated).
| Engine | p50 (ms) | Searches per second, one searcher | Searches per second, eight searchers at once | Exact mode p50 (ms) | Client CPU per search, one searcher (ms) | Recall in hits | Flags |
|---|---|---|---|---|---|---|---|
| Qdrant (HNSW) | 2.51 | 452 | 3,077 | 1.00 | - | 200 of 200 | |
| Qdrant (exact) | 2.86 | 408 | 3,137 | - | - | 200 of 200 | Pm |
| SQL Server 2025 | 5.43 | 218 | 942 | - | - | 200 of 200 | Pm |
Engines are listed in alphabetical order. The table does not rank them.
A dash means the table has no figure there.
The small markers after an engine name are flags. Hover a marker for its evidence, or read the list below the tables.
- Pm
p50-mean-inconsistentThe p50 and the mean from the one-searcher pass disagree by more than the limit named in the evidence.
Evidence behind the flags
- Qdrant (exact)
p50-mean-inconsistentp50 2.86 ms, mean 2.45 ms, p50/mean 1.17; limits: p50 above the mean by more than 1.15 times, or the mean above p50 by more than 1.5 times - SQL Server 2025
p50-mean-inconsistentp50 5.43 ms, mean 4.59 ms, p50/mean 1.18; limits: p50 above the mean by more than 1.15 times, or the mean above p50 by more than 1.5 times
Full results
This report is printed as the run wrote it, except for any sentence a note above it says was left out. This page does not check the engine texts in it. The summary page's facts table gives the label of each fact it uses.
Run 2026-10-03T18:39:55Z (run-all). Command line:
~/ForClaude/GenericVectorBuilder/src/GenericVectorBuilder.Bench/bin/Release/net10.0/GenericVectorBuilder.Bench.dll run-all --pipeline eshoponweb --targets sql,qdrant,qdrant-hnsw --queries golden
- Machine: bench-host, Intel(R) Xeon(R) CPU E5-1620 v3 @ 3.50GHz (8 logical CPUs), 62.7 GiB RAM, GPU NVIDIA GeForce RTX 3060, 12288 MiB, Ubuntu 24.04.5 LTS, kernel 6.8.0-142-generic, .NET 10.0.12, load average at start 2.50 2.60 2.71
- Data: 524 vectors x 1024 dims, collection
gvbbench_eshoponweb. Read READ-ONLY from GenericVectorBuilder.dbo.gvb_eshoponweb in ChunkId order, vectors as native SqlVector<float>, 0.2 s. - Queries: golden: 20 labelled questions from ~/ForClaude/evalkit/questions_golden.json, embedded with qwen3-emb-0.6b (embedded now: 1 probe + 20 query calls to the GPU service). Top 10. Throughput at concurrency 1, 8 for 20 s each.
- Ground truth: brute force over every vector in memory, 0.1 s for all 20 queries.
- nDCG@10 of the exact answer itself (the ceiling for this embedder): 0.518
| engine | index | load rows/s | p50 ms | p95 ms | p99 ms | QPS@1 | QPS@8 | recall@10 | nDCG@10 | RAM | disk |
|---|---|---|---|---|---|---|---|---|---|---|---|
| sql | exact VECTOR_DISTANCE cosine, no vector index (full scan) | 531 | 5.43 | 8.90 | 11.2 | 217.6 | 942.5 | 1.000 | 0.518 | 4.13 GiB | 5.15 MiB |
| qdrant | exact search (the builder's setting), HNSW m=16 ef_construct=100 built but not used | 2,869 | 2.86 | 3.68 | 3.92 | 408.2 | 3136.6 | 1.000 | 0.518 | 4.21 GiB | 328.14 MiB |
| qdrant-hnsw | HNSW m=16 ef_construct=100, hnsw_ef=server default, cosine | 5,388 | 2.51 | 3.13 | 3.82 | 452.3 | 3077.2 | 1.000 | 0.518 | 4.21 GiB | 328.14 MiB |
Details per target
SQL Server 2025 sql
- Engine: Microsoft SQL Server 2025 (RTM-CU9) (KB5122048) - 17.0.5005.3 (X64) (always-on)
- Index: exact VECTOR_DISTANCE cosine, no vector index (full scan)
- Load: 524 rows in batches of 1000, 1.0 s of upserts (531 rows/s); count matched 0.0 s after the last upsert; index step n/a
- Search: 200 latency samples, target held 524 rows, 0 errors
- RAM: 4.13 GiB (whole SQL Server process); disk: 5.15 MiB (table and its indexes, reserved pages)
- Load average 2.38 2.58 2.70 (1/5/15 min) when searching began.
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
Qdrant (exact) qdrant
- Engine: Qdrant 1.17.0 (systemd, local) (always-on)
- Index: exact search (the builder's setting), HNSW m=16 ef_construct=100 built but not used
- Load: 524 rows in batches of 1000, 0.2 s of upserts (2,869 rows/s); count matched 0.0 s after the last upsert; index step 2.0 s (status Green, 0 of 524 vectors in HNSW segments (4 segments))
- Search: 200 latency samples, target held 524 rows, 0 errors
- RAM: 4.21 GiB (whole Qdrant process, every collection); disk: 328.14 MiB (collection folder)
- Load average 3.48 2.81 2.77 (1/5/15 min) when searching began.
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
Qdrant (HNSW) qdrant-hnsw
- Engine: Qdrant 1.17.0 (systemd, local) (always-on)
- Index: HNSW m=16 ef_construct=100, hnsw_ef=server default, cosine
- Load: 524 rows in batches of 1000, 0.1 s of upserts (5,388 rows/s); count matched 0.0 s after the last upsert; index step 2.0 s (status Green, 0 of 524 vectors in HNSW segments (4 segments))
- Search: 200 latency samples, target held 524 rows, 0 errors
- Exact mode: 20 queries, p50 1.00 ms, p95 1.36 ms, recall 1.000
- RAM: 4.21 GiB (whole Qdrant process, every collection); disk: 328.14 MiB (collection folder)
- Load average 4.53 3.10 2.87 (1/5/15 min) when searching began.
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
Notes
- Load rows/s counts only time inside each target's upsert calls: one writer, batches of 1000, rows already in memory, the collection dropped and created fresh first.
- Latency is client-side wall time around each search (network and driver included), one query at a time, after 20 warm-up queries; at least 200 samples (small query sets are repeated).
- QPS: N workers searching back to back for 20 s per level; completed searches divided by elapsed time.
- Recall@10: share of the exact top 10 (brute force in memory) that the engine returned. A hit whose exact similarity ties the 10th best (within 1e-5) also counts, because duplicate rows embed to identical vectors.
- RAM of always-on servers (SQL Server, Qdrant) is the whole process, including every other database or collection it serves; for compose engines it is docker stats of the engine's containers. Disk is the table's reserved pages (SQL), the collection folder (Qdrant), or for other engines what the load added to the engine's data folder (engines that keep data in memory until a snapshot show almost nothing).
- The empty benchmark database GvbBench was dropped at the end.