Vector engine benchmark: eshoponweb
Run started 7 Oct 2026, 23:35 UTC. Data: eShopOnWeb, 524 vectors. Queries: 20 labelled questions.
Used by: published-2026-10-08 (v8, basis)
Machine: CPU Intel(R) Xeon(R) CPU E5-1620 v3 @ 3.50GHz; Logical CPUs 8; RAM (GiB) 62.7; OS Ubuntu 24.04.5 LTS, kernel 6.8.0-142-generic; Governor performance; Partition client 0-1,4-5; engines 2-3,6-7; Release build; Code commit 9b924200abc3f2c3e41c2a9520e5c7bb16748e6c; Exact mode seconds 60.
Warm-up, as this run's notes record it: Warm-up and settle check, untimed, right before every timed pass: the pass's own search at the pass's own number of searchers for at least 15 s and at least 20 searches (at most 120 s), read in windows of at least 2 s and 100 searches; then a 3 s trial of the same pass.
Clock: pinned true; pinnedMhz 3500; ceilingBeforeMhz 3600; noTurbo atEnd=1, before=0, during=1; uncoreMsr620 atEnd=0x1e1e, before=0xc1e, during=0x1e1e; throttle atEnd=coreEvents=0, packageEvents=0, perCpu=coreCount=0, cpu=0, packageCount=0, packageId=0, siblings=0,4,coreCount=0, cpu=1, packageCount=0, packageId=0, siblings=1,5,coreCount=0, cpu=2, packageCount=0, packageId=0, siblings=2,6,coreCount=0, cpu=3, packageCount=0, packageId=0, siblings=3,7,coreCount=0, cpu=4, packageCount=0, packageId=0, siblings=0,4,coreCount=0, cpu=5, packageCount=0, packageId=0, siblings=1,5,coreCount=0, cpu=6, packageCount=0, packageId=0, siblings=2,6,coreCount=0, cpu=7, packageCount=0, packageId=0, siblings=3,7, readUtc=2026-10-08T00:56:40.521Z, atStart=coreEvents=0, packageEvents=0, perCpu=coreCount=0, cpu=0, packageCount=0, packageId=0, siblings=0,4,coreCount=0, cpu=1, packageCount=0, packageId=0, siblings=1,5,coreCount=0, cpu=2, packageCount=0, packageId=0, siblings=2,6,coreCount=0, cpu=3, packageCount=0, packageId=0, siblings=3,7,coreCount=0, cpu=4, packageCount=0, packageId=0, siblings=0,4,coreCount=0, cpu=5, packageCount=0, packageId=0, siblings=1,5,coreCount=0, cpu=6, packageCount=0, packageId=0, siblings=2,6,coreCount=0, cpu=7, packageCount=0, packageId=0, siblings=3,7, readUtc=2026-10-07T23:35:51.750Z; toleranceBp 100; rule every CPU's clock is read from /sys/devices/system/cpu/cpuN/cpufreq/scaling_cur_freq every 250 ms; each CPU's median over a pass is taken, then the median of those medians per CPU group: the engine CPUs (2-3,6-7) and the client CPUs (0-1,4-5) are judged as two groups; a pass is 'clock off' (clockOff) when a group's median differs from the pinned clock by more than 100 bp (1%), and 'clock not read' (clockRead false) when a group has no reading; either counts as not at the pinned clock. Limited power: a CPU's median hides off-clock readings while they are fewer than half of its readings, and the median of 4 CPU medians hides 1 CPU(s) off the clock for the whole pass. This rule reads the kernel's figures only; the independent check of the clock is APERF/MPERF, read by the run's observer outside this tool. Thermal throttle counters (thermal_throttle/core_throttle_count and package_throttle_count of every CPU) are read at the start and the end of the run and at each target's start, first warm-up line and end, never by the clock sampler and never inside a timed pass; the core counter is counted once per physical core (thread_siblings_list) and the package counter once per package (physical_package_id), a rise on any CPU counts, a counter that went down counts as a rise, and only a rise against no rise means anything. They count thermal events only: no power-limit counters are read, so power capping (RAPL) and AVX clock drops are not seen..
| Engine | p50 (ms) | Searches per second, one searcher | Searches per second, eight searchers at once | Exact mode p50 (ms) | Client CPU per search, one searcher (ms) | Recall in hits | Flags |
|---|---|---|---|---|---|---|---|
| Chroma | 2.07 | 478 | 791 | - | 0.47 | 200 of 200 | |
| ClickHouse | 4.86 | 197 | 349 | 6.60 | 0.86 | 199 of 200 | UnNh |
| DuckDB | 3.64 | 272 | 583 | 4.10 | 5.01 | 200 of 200 | |
| Elasticsearch | 1.32 | 742 | 2,694 | 1.26 | 0.45 | 200 of 200 | Ix |
| MariaDB | 0.62 | 1,576 | 5,956 | 1.61 | 0.49 | 200 of 200 | |
| Milvus | 2.24 | 423 | 1,266 | - | 0.54 | 200 of 200 | |
| MongoDB Atlas Local | 1.39 | 705 | 2,098 | 1.21 | 0.30 | 200 of 200 | |
| OpenSearch | 1.98 | 495 | 1,507 | 2.56 | 0.47 | 200 of 200 | |
| Oracle 23ai Free | 0.93 | 1,050 | 2,792 | 2.42 | 0.89 | 200 of 200 | |
| Qdrant (HNSW) | 0.92 | 1,072 | 3,468 | 0.90 | 0.73 | 200 of 200 | |
| Qdrant (exact) | 0.89 | 1,109 | 3,296 | - | 0.72 | 200 of 200 | |
| Redis redis holds its data in memory: its saved docs page says 'Redis is an in-memory but persistent on disk database', and its compose file sets save "300 1" and appendonly no. sources
| 0.39 | 2,559 | 5,258 | 0.36 | 0.56 | 200 of 200 | |
| SQL Server 2025 | 3.93 | 245 | 589 | - | 1.21 | 200 of 200 | |
| SQL Server 2025 + DiskANN | 3.57 | 278 | 694 | 3.63 | 1.00 | 193 of 200 | |
| Typesense | 4.16 | 237 | 572 | 5.63 | 0.52 | 200 of 200 | |
| Vespa | 1.87 | 517 | 1,582 | 1.87 | 0.92 | 200 of 200 | |
| Weaviate | 5.33 | 167 | 306 | - | 0.56 | 200 of 200 | |
| pgvector | 0.85 | 1,147 | 4,019 | 2.19 | 0.68 | 200 of 200 | |
| sqlite-vec | 1.86 | 529 | 1,140 | 1.87 | 1.98 | 200 of 200 |
Engines are listed in alphabetical order. The table does not rank them.
A dash means the table has no figure there.
The small markers after an engine name are flags. Hover a marker for its evidence, or read the list below the tables.
- Un
unsettled-targetThe engine did not pass its own settle check before a timed pass. The evidence names the check. - Ix
index-not-readyThe engine did not report a finished index after the load or after the searches. - Nh
not-heldThe run recorded a timed pass of this engine as NOT HELD: the engine was still changing when it was timed. The evidence gives the figures. The row is shown, not ranked.
Evidence behind the flags
- ClickHouse
unsettled-targetlatency had NOT settled when timing began for default@8 (each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it). exact: warm-up 15.0 s and 2,204 searches (0 failed) at 1 searcher; p50 of the last windows 6.584, 6.565, 6.581 ms (windows of at least 2 s and 100 searches); trial of 443 searches in 3.0 s at 1 searcher: p50 6.622 ms against the settled p50 6.581 ms, 1% apart (limit 10%); the timed pass p50 6.603 ms, 0% from the settled p50 6.581 ms (limit 10% or 0.1 ms; 0.022 ms); default@8: warm-up 15.0 s and 6,591 searches (0 failed) at 8 searchers; QPS of the last windows 472, 465, 405 (windows of at least 2 s and 100 searches) (its last 3 windows differed by more than 5%); trial of 1,295 searches in 3.0 s at 8 searchers: 431 QPS against the settled 465 QPS, 8% apart (limit 10%), so the warm-up was EXTENDED once (30.0 s and 12,492 searches (0 failed) at 8 searchers; QPS of its older 6 windows 427 (windows 436, 399, 412, 462, 462, 391) and of its newer 6 416 (windows 400, 458, 446, 411, 409, 370), 2.7% apart (limit 5%) (windows of at least 2 s and 100 searches)); second trial of 1,078 searches in 3.0 s at 8 searchers: 358 QPS against the settled 416 QPS, 16% apart (limit 10%), outside its newer windows' range 370 to 458 QPS; still disagreeing after the extension; the timed pass 349 QPS, 19% from the settled 416 QPS (limit 10%), NOT HELD: the engine was still changing when it was timed; default@1: warm-up 15.0 s and 2,479 searches (0 failed) at 1 searcher; p50 of the last windows 5.497, 5.449, 5.221 ms (windows of at least 2 s and 100 searches) (its last 3 windows differed by more than 5%); trial of 466 searches in 3.0 s at 1 searcher: p50 6.861 ms against the settled p50 5.449 ms, 21% apart (limit 10%), so the warm-up was EXTENDED once (30.0 s and 5,806 searches (0 failed) at 1 searcher; p50 of its older 6 windows 4.910 ms (windows 4.850, 4.932, 5.369, 4.842, 4.893, 4.847) and of its newer 6 4.855 ms (windows 4.875, 4.849, 4.863, 4.838, 4.868, 4.838), 1.1% apart (limit 5% or 0.1 ms; 0.056 ms) (windows of at least 2 s and 100 searches)); second trial of 595 searches in 3.0 s at 1 searcher: p50 4.851 ms against the settled p50 4.855 ms, 0% apart (limit 10%), inside its newer windows' range 4.838 to 4.875 ms; settled after the extension; the timed pass p50 4.859 ms, 0% from the settled p50 4.855 ms (limit 10% or 0.1 ms; 0.004 ms). Its numbers may still include warm-up or a change the engine was still going through; rerun before quoting them. - ClickHouse
not-helddefault@8: the timed pass 349 QPS, 19% from the settled 416 QPS (limit 10%), NOT HELD - Elasticsearch
index-not-readyafterLoad: the engine's index state read not ready, 0 of 524 vectors indexed - Elasticsearch
index-not-readyafterSearch: the engine's index state read not ready, 0 of 524 vectors indexed
Full results
This report is printed as the run wrote it, except for any sentence a note above it says was left out. This page does not check the engine texts in it. The summary page's facts table gives the label of each fact it uses.
Run 2026-10-07T23:35:51Z (run-all). Command line:
~/gvb-work/v8-run/bin/GenericVectorBuilder.Bench.dll run-all --pipeline eshoponweb --queries golden --seed 803 --out ~/ForClaude/GenericVectorBuilder/bench-results --repo ~/ForClaude/GenericVectorBuilder
- Machine: bench-host, Intel(R) Xeon(R) CPU E5-1620 v3 @ 3.50GHz (8 logical CPUs), 62.7 GiB RAM, GPU NVIDIA GeForce RTX 3060, 12288 MiB, Ubuntu 24.04.5 LTS, kernel 6.8.0-142-generic, .NET 10.0.12, load average at start 2.20 3.63 3.60
- Data: 524 vectors x 1024 dims, collection
gvbbench_eshoponweb. Read READ-ONLY from GenericVectorBuilder.dbo.gvb_eshoponweb in ChunkId order, vectors as native SqlVector<float>, 0.2 s. - Queries: golden: 20 labelled questions from ~/ForClaude/evalkit/questions_golden.json, embedded with qwen3-emb-0.6b (cached in ~/gvb-data/bench-cache/golden-eshoponweb-68df42ca69efa079.json, no embedding calls). Top 10. Throughput at concurrency 1, 8 for 20 s each.
- Request speed on a small collection (524 vectors): Measured end to end through each engine's .NET client [Correction 11]; at this size it reflects per-request cost including the client library, not index scaling.
- Client CPU per search is the CPU time the test's .NET client itself used for each search, measured in the same pass as the figure beside it, summed over every thread, so it can exceed the time per search. For an embedded engine (DuckDB, sqlite-vec) the engine runs inside the client process, so its figure includes the engine's own CPU time.
- Ground truth: brute force over every vector in memory, 0.0 s for all 20 queries.
- nDCG@10 of the exact answer itself (the ceiling for this embedder): 0.518
| engine | index | load rows/s | p50 ms | p95 ms | p99 ms | QPS@1 | QPS@8 | client CPU ms/search@1 | client CPU ms/search@8 | recall@10 | nDCG@10 | RAM | disk |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| sql | exact VECTOR_DISTANCE cosine, no vector index (full scan) | 543 | 3.93 | 4.45 | 5.56 | 245.4 | 589.1 | 1.21 | 1.18 | 1.000 | 0.518 | 529.9 MiB | 5.15 MiB |
| mariadb | VECTOR INDEX (HNSW variant) DISTANCE=cosine, M=16 (no ef_construction setting exists), mhnsw_ef_search=100 per statement (the ef 100 most engines here use, so the search effort matches; MariaDB's own default is 20; recall@10 at ef 100 falls as the set grows (random 1024-dimension vectors, measured 2026-10-04: 0.99 at 524, 0.89 to 0.92 at 2,000; an earlier run gave about 0.09 at 100,000 [Correction 12])), mhnsw_max_cache_size 4G; exact mode = IGNORE INDEX full scan | 1,020 | 0.62 | 0.70 | 0.88 | 1576.0 | 5955.8 | 0.49 | 0.38 | 1.000 | 0.518 | 170.2 MiB | 22.01 MiB |
| chroma | HNSW M=16 ef_construction=128, ef_search=100 (Chroma default), cosine; approximate only (no exact mode) | 793 | 2.07 | 2.34 | 2.58 | 477.9 | 790.5 | 0.47 | 0.46 | 1.000 | 0.518 | 52.3 MiB | 414.16 KiB |
| weaviate | HNSW maxConnections(M)=16 efConstruction=128, ef=-1 (dynamic: limit x 8 clamped 100..500), cosine, no quantization; approximate only (no exact mode) | 994 | 5.33 | 10.7 | 11.8 | 166.9 | 306.1 | 0.56 | 0.61 | 1.000 | 0.518 | 94.7 MiB | 2.96 MiB |
| pgvector | HNSW vector_cosine_ops m=16 ef_construction=128, hnsw.ef_search=100 per query, float32 vector(n), cosine; exact mode = same query with index scans off (sequential scan) | 618 | 0.85 | 1.02 | 1.30 | 1146.6 | 4018.6 | 0.68 | 0.54 | 1.000 | 0.518 | 100.6 MiB | 7.56 MiB |
| mongodb | vectorSearch index, HNSW maxEdges=16 numEdgeCandidates=128, float32 binData, cosine, numCandidates=20x hits (min 100); exact mode = $vectorSearch exact:true | 2,741 | 1.39 | 1.60 | 1.97 | 705.3 | 2098.3 | 0.30 | 0.28 | 1.000 | 0.518 | 647.3 MiB | 10.48 KiB |
| opensearch | faiss HNSW float32, no compression, m=16, ef_construction=128, cosinesimil; search k=top, ef_search=100; 1 shard, 0 replicas; graph built at any segment size (approximate_threshold=0); force-merged to one segment after the load | 314 | 1.98 | 2.21 | 2.76 | 495.0 | 1507.3 | 0.47 | 0.50 | 1.000 | 0.518 | 2.6 GiB | 9.66 MiB |
| vespa | HNSW float32 tensor, prenormalized-angular (cosine), max-links-per-node=16, neighbors-to-explore-at-insert=128; search targetHits=top, ef=100 via exploreAdditionalHits; exact mode = approximate:false; vectors held in memory | 291 | 1.87 | 2.32 | 2.66 | 517.1 | 1581.8 | 0.92 | 0.87 | 1.000 | 0.518 | 2.95 GiB | 235.33 MiB |
| redis | HNSW TYPE FLOAT32 M=16 EF_CONSTRUCTION=128, EF_RUNTIME=100 per query, cosine; exact mode = FLAT index built on first exact query | 11,812 | 0.39 | 0.48 | 0.70 | 2559.3 | 5257.6 | 0.56 | 0.35 | 1.000 | 0.518 | 253.8 MiB | 2.43 MiB |
| duckdb | HNSW (vss extension) FLOAT[n] metric=cosine m=16 ef_construction=128, ef_search=100 per connection, persistent (hnsw_enable_experimental_persistence=true, checkpoint_threshold=256MB); exact mode = array_cosine_similarity sequential scan; score = 1 - cosine distance; searches run concurrently, one connection per searcher (opened as searchers arrive, at most 32), writes run one at a time and never overlap a search | 1,433 | 3.64 | 4.06 | 5.31 | 271.6 | 582.8 | 5.01 | 6.83 | 1.000 | 0.518 | - | 2.9 MiB |
| typesense | HNSW float32 (hnswlib), m=16, ef_construction=128, cosine; search k=top, ef=100; exact mode = filter ordinal:>=0 with flat_search_cutoff; index held in memory | 1,016 | 4.16 | 4.54 | 5.20 | 237.1 | 572.1 | 0.52 | 0.58 | 1.000 | 0.518 | 118.9 MiB | 1.67 MiB |
| sql-diskann | DiskANN (preview) via VECTOR_SEARCH, cosine, build {"StartId":"306", "L":"48", "M":"8", "R":"48"}; the exact mode scans the same table | 619 | 3.57 | 4.00 | 4.52 | 277.8 | 693.8 | 1.00 | 1.08 | 0.965 | 0.526 | 533 MiB | 4.52 MiB |
| qdrant-hnsw | HNSW m=16 ef_construct=100, hnsw_ef=server default, cosine; indexing_threshold_kb 1 and full_scan_threshold_kb 10 (server defaults are 10,000 each) so a small collection builds and walks its graph | 4,677 | 0.92 | 1.03 | 1.50 | 1072.2 | 3467.8 | 0.73 | 0.52 | 1.000 | 0.518 | 37.02 MiB | 164.1 MiB |
| milvus | HNSW M=16 efConstruction=128, ef=100, metric COSINE, Strong consistency searches; approximate only (no exact mode) | 1,158 | 2.24 | 2.82 | 3.81 | 423.1 | 1266.1 | 0.54 | 0.63 | 1.000 | 0.518 | 186.5 MiB | 10.95 MiB |
| clickhouse | vector_similarity HNSW cosineDistance, quantization bf16, M=16 ef_construction=128, hnsw_candidate_list_size_for_search=256, rescoring off; exact mode = full scan with skip indexes off | 2,737 | 4.86 | 6.54 | 7.81 | 196.8 | 348.7 | 0.86 | 1.05 | 0.995 | 0.528 | 1011 MiB | 4.67 GiB |
| sqlitevec | vec0 brute-force scan, no ANN index (exact), float32, cosine distance, default chunk_size=1024; score = 1 - cosine distance; searches run concurrently, one WAL reader connection per searcher (opened as searchers arrive, at most 32), writes run one at a time and may overlap searches | 3,687 | 1.86 | 2.08 | 2.38 | 528.6 | 1140.0 | 1.98 | 3.50 | 1.000 | 0.518 | - | 9.72 MiB |
| elasticsearch | HNSW float32, no quantization, m=16, ef_construction=128, cosine; search k=top, num_candidates=100; 1 shard, 0 replicas; force-merged to one segment after the load (at 1,024 dimensions a segment under 1,043 vectors gets no graph) | 473 | 1.32 | 1.50 | 1.77 | 742.2 | 2693.8 | 0.45 | 0.49 | 1.000 | 0.518 | 2.58 GiB | 2.39 MiB |
| qdrant | exact scan: the builder's sink sends exact=true on every search, so no HNSW graph is used whether or not Qdrant has built one (see the index state) | 5,738 | 0.89 | 1.00 | 1.43 | 1109.2 | 3295.9 | 0.72 | 0.54 | 1.000 | 0.518 | 35.61 MiB | 196.08 MiB |
| oracle | HNSW in-memory neighbor graph NEIGHBORS=16 EFCONSTRUCTION=128, EFSEARCH=100 per query, cosine; exact mode = FETCH EXACT FIRST (full scan); Oracle Free caps itself at 2 CPUs (cpu_count 2 in V$PARAMETER, edition FREE in V$INSTANCE, 8 host CPUs in V$OSSTAT NUM_CPUS; the 2 CPU thread limit is Oracle's documented Free edition limit) | 1,951 | 0.93 | 1.09 | 1.89 | 1049.9 | 2792.4 | 0.89 | 0.64 | 1.000 | 0.518 | 2.14 GiB | 0 B |
Details per target
SQL Server 2025 sql
- Engine: Microsoft SQL Server 2025 (RTM-CU9) (KB5122048) - 17.0.5005.3 (X64), Enterprise Developer Edition (64-bit), in container gvb-mssql (image mcr.microsoft.com/mssql/server:2025-CU9-ubuntu-24.04@sha256:2b5b581621126574f3d1f75e78d3eebe8d05aedb59ad0cfdf9aa42cb0634d726; cpuset 2-3,6-7 from its creation; SQL Server counts 4 CPU(s) and runs 4 visible scheduler(s), affinity AUTO) (compose)
- Index: exact VECTOR_DISTANCE cosine, no vector index (full scan)
- Search settings: description=exact VECTOR_DISTANCE cosine, no vector index (full scan)
- Engine settings: HostConfig.Memory = 8589934592 (read: docker inspect HostConfig.Memory (0 = no limit)); HostConfig.CpusetCpus = 2-3,6-7 (read: docker inspect HostConfig.CpusetCpus (empty = every CPU)); HostConfig.NanoCpus = 0 (read: docker inspect HostConfig.NanoCpus (0 = no limit))
- Image: mcr.microsoft.com/mssql/server:2025-CU9-ubuntu-24.04@sha256:2b5b581621126574f3d1f75e78d3eebe8d05aedb59ad0cfdf9aa42cb0634d726, id sha256:2b5b581621126574f3d1f75e78d3eebe8d05aedb59ad0cfdf9aa42cb0634d726 (the pinned id), tagged on this machine at 2026-10-04T22:02:11.293843649Z
- Data folder ~/gvb-data/engines/mssql: 105.07 MiB at the start [Correction 6], 441.85 MiB at the end
- Load: 524 rows in batches of 1000, 1.0 s of upserts (543 rows/s); count matched 0.0 s after the last upsert; index step n/a
- Search: 4908 latency samples, target held 524 rows, 0 errors
- Client CPU per search: 1 searcher 1.21 ms, 8 searchers 1.18 ms
- RAM: 529.9 MiB (docker stats); disk: 5.15 MiB (table and its indexes, reserved pages)
- Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started mssql.compose.yaml with every container created on CPUs 2-3,6-7: gvb-mssql cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards.
- Engine files (SHA-256, first 12 hex digits): [mssql.compose.yaml b7359663cc15]. This run created the engine from these files.
- Index after the load: ready, 0 of 524 indexed (no index used, exact scan by design: dbo.gvb_gvbbench_eshoponweb in GvbBench holds 524 rows; sys.vector_indexes lists no index on the table; indexes: PK_gvb_gvbbench_eshoponweb CLUSTERED, IX_gvb_gvbbench_eshoponweb_DocKey NONCLUSTERED; a real search under SET STATISTICS XML ON ran: Clustered Index Scan of GvbBench.gvb_gvbbench_eshoponweb through PK_gvb_gvbbench_eshoponweb (IndexKind Clustered), 524 rows read, no vector index operator, 10 hits). Durability: A commit returns after its transaction-log records are written to disk (SQL Server write-ahead logging; delayed durability is DISABLED); a database created here copies the model database: recovery model FULL, page_verify CHECKSUM; no global trace flags are enabled; mssql.conf of container gvb-mssql (/var/opt/mssql/mssql.conf, on the host at ~/gvb-data/engines/mssql/mssql.conf) does not exist, so it sets: nothing; it has no [control] or [traceflag] entry, so SQL Server's own Linux defaults for flushing writes apply (Microsoft's Linux performance guide names trace flag 3982 as that default; read from the guide, not tested here). Not tested by cutting power; whether the disk's own write cache reaches the media was not checked. Container settings from its environment (names only): MSSQL_AGENT_ENABLED, MSSQL_MEMORY_LIMIT_MB, MSSQL_PID, MSSQL_RPC_PORT, MSSQL_SA_PASSWORD..
- WARNING: processes outside the benchmark used 0.62 CPUs on average over the last 3 s and 0.62 over the last few seconds when searching began (limit 0.3): the box was busy as the untimed rehearsal started. Each timed pass still waits for a quiet box before its warm-up, and one that ran busy is flagged 'busy box' on its own. The load average 2.01 3.54 3.57 (1/5/15 min; it also counts this benchmark's own client and engines, so it is recorded, not judged).
- Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@1 7,413 searches in 30.0 s, default@8 18,177 searches in 30.1 s.
- Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it. default@1: warm-up 15.0 s and 3,757 searches (0 failed) at 1 searcher; p50 of the last windows 3.926, 3.920, 3.927 ms (windows of at least 2 s and 100 searches); trial of 749 searches in 3.0 s at 1 searcher: p50 3.932 ms against the settled p50 3.926 ms, 0% apart (limit 10%); the timed pass p50 3.933 ms, 0% from the settled p50 3.926 ms (limit 10% or 0.1 ms; 0.007 ms); default@8: warm-up 15.0 s and 9,243 searches (0 failed) at 8 searchers; QPS of the last windows 606, 624, 625 (windows of at least 2 s and 100 searches); trial of 1,869 searches in 3.0 s at 8 searchers: 621 QPS against the settled 624 QPS, 0% apart (limit 10%); the timed pass 589 QPS, 6% from the settled 624 QPS (limit 10%).
- Pass default@1 after rehearsal: 2026-10-07T23:37:21.727Z to 2026-10-07T23:37:41.730Z (20.0 s), 4,908 searches, 0 failed, p50 3.933 ms, mean 4.072 ms, p99 5.558 ms, 245.4 QPS (1000/QPS 4.076 ms).
- Pass default@8 after default@1: 2026-10-07T23:37:59.812Z to 2026-10-07T23:38:19.823Z (20.0 s), 11,788 searches, 0 failed, p50 12.426 ms, mean 13.572 ms, p99 25.902 ms, 589.1 QPS (1000/QPS 1.698 ms).
- Passes in the order run: default@1, default@8; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 41,208 sent, 0 failed (warmupErrors). Index after the searches: ready, 0 of 524 indexed (no index used, exact scan by design: dbo.gvb_gvbbench_eshoponweb in GvbBench holds 524 rows; sys.vector_indexes lists no index on the table; indexes: PK_gvb_gvbbench_eshoponweb CLUSTERED, IX_gvb_gvbbench_eshoponweb_DocKey NONCLUSTERED; a real search under SET STATISTICS XML ON ran: Clustered Index Scan of GvbBench.gvb_gvbbench_eshoponweb through PK_gvb_gvbbench_eshoponweb (IndexKind Clustered), 524 rows read, no vector index operator, 10 hits).
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
- Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-mssql (e63be1146b72) on network gvb-mssql_default at container-ip:1433, not through docker-proxy localhost:14330. Open after the passes: container-ip:1433 x9 (this target).
- CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-mssql: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-mssql: 86 thread(s) on CPUs 2-3,6-7 (every thread of the container).
- Clock per pass (median MHz of engine CPUs / client CPUs, each the median of its CPUs' medians and every CPU when the CPUs were not split, their average MHz, governor, CPUs busy outside the benchmark, client and engine CPU per search; the clock was pinned at 3500 MHz (turbo off), and a pass whose median on either group is more than 100 bp (1%) off it, or that has no reading on a group, is flagged): default@1 3492/3492 MHz, average 3493.2/3492.3, performance, outside load 0.17 before the warm-up, 0.17 during, client CPU 1.210 ms per search, engine CPU 3.876 ms per search; default@8 3492/3492 MHz, average 3492.0/3492.4, performance, outside load 0.16 before the warm-up, 0.15 during, client CPU 1.176 ms per search, engine CPU 5.933 ms per search.
MariaDB mariadb
- Engine: MariaDB 11.8.9 (InnoDB + VECTOR INDEX) (compose)
- Index: VECTOR INDEX (HNSW variant) DISTANCE=cosine, M=16 (no ef_construction setting exists), mhnsw_ef_search=100 per statement (the ef 100 most engines here use, so the search effort matches; MariaDB's own default is 20; recall@10 at ef 100 falls as the set grows (random 1024-dimension vectors, measured 2026-10-04: 0.99 at 524, 0.89 to 0.92 at 2,000; an earlier run gave about 0.09 at 100,000 [Correction 12])), mhnsw_max_cache_size 4G; exact mode = IGNORE INDEX full scan
- Search settings: DISTANCE=cosine, M=16, mhnsw_ef_search=100
- Engine settings: HostConfig.Memory = 8589934592 (read: docker inspect HostConfig.Memory (0 = no limit)); HostConfig.CpusetCpus = 2-3,6-7 (read: docker inspect HostConfig.CpusetCpus (empty = every CPU)); HostConfig.NanoCpus = 0 (read: docker inspect HostConfig.NanoCpus (0 = no limit)); innodb_buffer_pool_size = 2147483648 (read: docker exec gvbbench-mariadb mariadb SHOW GLOBAL VARIABLES); innodb_flush_log_at_trx_commit = 2 (read: docker exec gvbbench-mariadb mariadb SHOW GLOBAL VARIABLES); mhnsw_max_cache_size = 4294967296 (read: docker exec gvbbench-mariadb mariadb SHOW GLOBAL VARIABLES)
- Image: mariadb:11.8.9, id sha256:6422478cb8e159f080fb1d8ccf65101e26fe51385787fde7d16c3b165a331f15 (the pinned id), tagged on this machine at 2026-10-03T14:10:11.675870329Z
- Data folder ~/gvb-data/engines/mariadb-bench: 166.77 MiB at the start [Correction 4], 188.79 MiB at the end
- Load: 524 rows in batches of 1000, 0.5 s of upserts (1,020 rows/s); count matched 0.0 s after the last upsert; index step 0.5 s (VECTOR INDEX is maintained inside every INSERT, nothing to build; confirmed after 2 poll(s) in 0.5 s (poll 1 of 2 said: not ready: a search at mhnsw_ef_search 100 returned only 383 of 524 requested rows through the index (EXPLAIN key 'vec_idx') (information_schema lists vec_idx as INDEX_TYPE VECTOR, EXPLAIN of the default search uses key vec_idx, the server applied mhnsw_ef_search 100 (asked for 100, read back from the server), a search at that effort returned 383 of 524 requested rows through the index; MariaDB has no per-index row count (the graph is a hidden InnoDB table), so indexed vectors are not reported)): information_schema lists vec_idx as INDEX_TYPE VECTOR, EXPLAIN of the default search uses key vec_idx, the server applied mhnsw_ef_search 100 (asked for 100, read back from the server), a search at that effort returned 517 of 524 requested rows through the index; MariaDB has no per-index row count (the graph is a hidden InnoDB table), so indexed vectors are not reported)
- Search: 31520 latency samples, target held 524 rows, 0 errors
- Client CPU per search: 1 searcher 0.49 ms, 8 searchers 0.38 ms
- Exact mode: 36,567 searches, p50 1.61 ms, p95 1.72 ms, recall 1.000
- RAM: 170.2 MiB (docker stats); disk: 22.01 MiB added by this load (188.79 MiB (whole engine data folder), 166.78 MiB before)
- Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started mariadb-bench.compose.yaml with every container created on CPUs 2-3,6-7: gvbbench-mariadb cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards.
- Engine files (SHA-256, first 12 hex digits): [mariadb-bench.compose.yaml e9a2147e6f91]. This run created the engine from these files.
- Index after the load: ready, ? of 524 indexed (information_schema lists vec_idx as INDEX_TYPE VECTOR, EXPLAIN of the default search uses key vec_idx, the server applied mhnsw_ef_search 100 (asked for 100, read back from the server), a search at that effort returned 517 of 524 requested rows through the index; MariaDB has no per-index row count (the graph is a hidden InnoDB table), so indexed vectors are not reported). Durability: innodb_flush_log_at_trx_commit=2 (set in mariadb-bench.compose.yaml for the benchmark's own container gvbbench-mariadb, and in mariadb.compose.yaml for the daily gvb-mariadb, which has the same server settings): the InnoDB redo log is written to the operating system at every commit but fsynced about once a second, so a crash of the mariadbd process loses nothing, while an operating-system crash or power cut can lose the last second of commits; innodb_doublewrite is on (default), the binary log is off, and the vector graph is an InnoDB table under the same log (settings read from the running server with SHOW VARIABLES; the crash behaviour is InnoDB's documented behaviour for this setting and was not tested on either container).
- Outside load when searching began: 0.24 CPUs on average over the last 1 s and 0.24 over the last few seconds (limit 0.3), a quiet box. The load average 2.96 3.32 3.48 (1/5/15 min; it also counts this benchmark's own client and engines, so it is recorded, not judged).
- Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@1 47,117 searches in 30.0 s, exact 18,234 searches in 30.0 s, default@8 180,140 searches in 30.0 s.
- Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it. default@1: warm-up 15.0 s and 23,434 searches (0 failed) at 1 searcher; p50 of the last windows 0.624, 0.634, 0.642 ms (windows of at least 2 s and 100 searches); trial of 4,626 searches in 3.0 s at 1 searcher: p50 0.639 ms against the settled p50 0.634 ms, 1% apart (limit 10%); the timed pass p50 0.623 ms, 2% from the settled p50 0.634 ms (limit 10% or 0.1 ms; 0.011 ms); exact: warm-up 15.0 s and 9,118 searches (0 failed) at 1 searcher; p50 of the last windows 1.606, 1.610, 1.635 ms (windows of at least 2 s and 100 searches); trial of 1,819 searches in 3.0 s at 1 searcher: p50 1.616 ms against the settled p50 1.610 ms, 0% apart (limit 10%); the timed pass p50 1.615 ms, 0% from the settled p50 1.610 ms (limit 10% or 0.1 ms; 0.005 ms); default@8: warm-up 15.0 s and 89,124 searches (0 failed) at 8 searchers; QPS of the last windows 5,941, 5,972, 5,982 (windows of at least 2 s and 100 searches); trial of 17,877 searches in 3.0 s at 8 searchers: 5,956 QPS against the settled 5,972 QPS, 0% apart (limit 10%); the timed pass 5,956 QPS, 0% from the settled 5,972 QPS (limit 10%).
- Pass default@1 after rehearsal: 2026-10-07T23:40:17.644Z to 2026-10-07T23:40:37.644Z (20.0 s), 31,520 searches, 0 failed, p50 0.623 ms, mean 0.633 ms, p99 0.883 ms, 1576.0 QPS (1000/QPS 0.635 ms).
- Pass exact after default@1: 2026-10-07T23:40:55.683Z to 2026-10-07T23:41:55.685Z (60.0 s), 36,567 searches, 0 failed, p50 1.615 ms, mean 1.639 ms, p99 2.069 ms, 609.4 QPS (1000/QPS 1.641 ms).
- Pass default@8 after exact: 2026-10-07T23:42:13.726Z to 2026-10-07T23:42:33.727Z (20.0 s), 119,123 searches, 0 failed, p50 1.301 ms, mean 1.341 ms, p99 2.446 ms, 5955.8 QPS (1000/QPS 0.168 ms).
- Passes in the order run: default@1, exact, default@8; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 391,489 sent, 0 failed (warmupErrors). Index after the searches: ready, ? of 524 indexed (information_schema lists vec_idx as INDEX_TYPE VECTOR, EXPLAIN of the default search uses key vec_idx, the server applied mhnsw_ef_search 100 (asked for 100, read back from the server), a search at that effort returned 517 of 524 requested rows through the index; MariaDB has no per-index row count (the graph is a hidden InnoDB table), so indexed vectors are not reported).
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
- Connection: container address on its Docker network, not the published port (no docker-proxy): container gvbbench-mariadb (908d9e98739b) on network gvbbench-mariadb_default at container-ip:3306, not through docker-proxy localhost:13306. Open after the passes: container-ip:3306 x8 (this target).
- CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvbbench-mariadb: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvbbench-mariadb: 13 thread(s) on CPUs 2-3,6-7 (every thread of the container).
- Clock per pass (median MHz of engine CPUs / client CPUs, each the median of its CPUs' medians and every CPU when the CPUs were not split, their average MHz, governor, CPUs busy outside the benchmark, client and engine CPU per search; the clock was pinned at 3500 MHz (turbo off), and a pass whose median on either group is more than 100 bp (1%) off it, or that has no reading on a group, is flagged): default@1 3492/3492 MHz, average 3493.2/3488.6, performance, outside load 0.09 before the warm-up, 0.13 during, client CPU 0.485 ms per search, engine CPU 0.449 ms per search; exact 3492/3492 MHz, average 3494.7/3492.6, performance, outside load 0.11 before the warm-up, 0.13 during, client CPU 0.500 ms per search, engine CPU 1.445 ms per search; default@8 3492/3492 MHz, average 3492.0/3492.0, performance, outside load 0.13 before the warm-up, 0.05 during, client CPU 0.377 ms per search, engine CPU 0.620 ms per search.
Chroma chroma
- Engine: Chroma 1.4.4 (single node, REST API v2) (compose)
- Index: HNSW M=16 ef_construction=128, ef_search=100 (Chroma default), cosine; approximate only (no exact mode)
- Search settings: M=16, ef_construction=128, ef_search=100
- Engine settings: HostConfig.Memory = 6442450944 (read: docker inspect HostConfig.Memory (0 = no limit)); HostConfig.CpusetCpus = 2-3,6-7 (read: docker inspect HostConfig.CpusetCpus (empty = every CPU)); HostConfig.NanoCpus = 0 (read: docker inspect HostConfig.NanoCpus (0 = no limit)); chroma_version = 1.0.0 (read: GET /api/v2/version)
- Image: chromadb/chroma:latest, id sha256:1e0b73a187a28757c572acba508c46f48c9e8b0acaf5c20e6d95cdedce1acdf6 (the pinned id), tagged on this machine at 2026-10-03T14:09:56.392230738Z
- Data folder ~/gvb-data/engines/chroma: 2.01 GiB at the start, 2.01 GiB at the end
- Load: 524 rows in batches of 1000, 0.7 s of upserts (793 rows/s); count matched 0.0 s after the last upsert; index step 0.0 s (Chroma has no index build to wait for and exposes no index status; checked in 0.0 s: Chroma exposes no index status (indexing_status endpoint: HTTP 500 {"error":"InternalError","message":"Method scout_logs is not implemented"}), so indexed vectors are not reported and ready only means the two counts agree to within 1 percent: count endpoint 524, a vector search asking for 524 results returned 524)
- Search: 9558 latency samples, target held 524 rows, 0 errors
- Client CPU per search: 1 searcher 0.47 ms, 8 searchers 0.46 ms
- RAM: 52.3 MiB (docker stats); disk: 414.16 KiB added by this load (2.01 GiB (whole engine data folder), 2.01 GiB before)
- Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started chroma.compose.yaml with every container created on CPUs 2-3,6-7: gvb-chroma cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards.
- Engine files (SHA-256, first 12 hex digits): [chroma.compose.yaml f7ed9d3c8ff7]. This run created the engine from these files.
- Index after the load: ready, ? of 524 indexed (Chroma exposes no index status (indexing_status endpoint: HTTP 500 {"error":"InternalError","message":"Method scout_logs is not implemented"}), so indexed vectors are not reported and ready only means the two counts agree to within 1 percent: count endpoint 524, a vector search asking for 524 results returned 524). Durability: SQLite rollback journal with its default synchronous=FULL under the data directory (chroma.compose.yaml sets IS_PERSISTENT=1 and PERSIST_DIRECTORY=/data and no sync setting): every write is committed to chroma.sqlite3 with fsync of the journal, the directory and the database file before the call returns (strace: 15 database fsyncs and 38 journal fsyncs for one create, four 500-vector upserts and one delete), and the HNSW files are written every sync_threshold=1000 vectors and rebuilt from the SQLite log after a crash; measured with kill -9 and a restart: all 50, 300 and 2,000 vectors were still stored and searchable.
- Outside load when searching began: 0.09 CPUs on average over the last 1 s and 0.09 over the last few seconds (limit 0.3), a quiet box. The load average 4.73 3.60 3.52 (1/5/15 min; it also counts this benchmark's own client and engines, so it is recorded, not judged).
- Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@8 23,562 searches in 30.0 s, default@1 14,322 searches in 30.0 s.
- Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it. default@8: warm-up 15.0 s and 11,846 searches (0 failed) at 8 searchers; QPS of the last windows 786, 791, 793 (windows of at least 2 s and 100 searches); trial of 2,376 searches in 3.0 s at 8 searchers: 790 QPS against the settled 791 QPS, 0% apart (limit 10%); the timed pass 791 QPS, 0% from the settled 791 QPS (limit 10%); default@1: warm-up 15.0 s and 7,176 searches (0 failed) at 1 searcher; p50 of the last windows 2.063, 2.074, 2.074 ms (windows of at least 2 s and 100 searches); trial of 1,435 searches in 3.0 s at 1 searcher: p50 2.071 ms against the settled p50 2.074 ms, 0% apart (limit 10%); the timed pass p50 2.067 ms, 0% from the settled p50 2.074 ms (limit 10% or 0.1 ms; 0.007 ms).
- Pass default@8 after rehearsal: 2026-10-07T23:44:00.637Z to 2026-10-07T23:44:20.645Z (20.0 s), 15,817 searches, 0 failed, p50 10.059 ms, mean 10.116 ms, p99 11.635 ms, 790.5 QPS (1000/QPS 1.265 ms).
- Pass default@1 after default@8: 2026-10-07T23:44:38.654Z to 2026-10-07T23:44:58.654Z (20.0 s), 9,558 searches, 0 failed, p50 2.067 ms, mean 2.091 ms, p99 2.576 ms, 477.9 QPS (1000/QPS 2.093 ms).
- Passes in the order run: default@8, default@1; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 60,717 sent, 0 failed (warmupErrors). Index after the searches: ready, ? of 524 indexed (Chroma exposes no index status (indexing_status endpoint: HTTP 500 {"error":"InternalError","message":"Method scout_logs is not implemented"}), so indexed vectors are not reported and ready only means the two counts agree to within 1 percent: count endpoint 524, a vector search asking for 524 results returned 524).
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
- Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-chroma (dbb629a5541e) on network gvb-chroma_default at container-ip:8000, not through docker-proxy localhost:8000. Open after the passes: container-ip:8000 x8 (this target).
- CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-chroma: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-chroma: 9 thread(s) on CPUs 2-3,6-7 (every thread of the container).
- Clock per pass (median MHz of engine CPUs / client CPUs, each the median of its CPUs' medians and every CPU when the CPUs were not split, their average MHz, governor, CPUs busy outside the benchmark, client and engine CPU per search; the clock was pinned at 3500 MHz (turbo off), and a pass whose median on either group is more than 100 bp (1%) off it, or that has no reading on a group, is flagged): default@8 3492/3492 MHz, average 3492.0/3492.6, performance, outside load 0.09 before the warm-up, 0.07 during, client CPU 0.464 ms per search, engine CPU 2.720 ms per search; default@1 3492/3492 MHz, average 3492.0/3493.3, performance, outside load 0.08 before the warm-up, 0.11 during, client CPU 0.467 ms per search, engine CPU 2.454 ms per search.
Weaviate weaviate
- Engine: Weaviate 1.39.8 (single node, REST + GraphQL) (compose)
- Index: HNSW maxConnections(M)=16 efConstruction=128, ef=-1 (dynamic: limit x 8 clamped 100..500), cosine, no quantization; approximate only (no exact mode)
- Search settings: ef=-1, efConstruction=128
- Engine settings: HostConfig.Memory = 8589934592 (read: docker inspect HostConfig.Memory (0 = no limit)); HostConfig.CpusetCpus = 2-3,6-7 (read: docker inspect HostConfig.CpusetCpus (empty = every CPU)); HostConfig.NanoCpus = 0 (read: docker inspect HostConfig.NanoCpus (0 = no limit)); GOMEMLIMIT = 6GiB (read: docker inspect Config.Env, the name GOMEMLIMIT only)
- Image: semitechnologies/weaviate:1.39.8, id sha256:f6f4a5961f99e8718a02c822ca6a0d65ed3f9c567734419a7b2144933014f13a (the pinned id), tagged on this machine at 2026-10-03T14:05:40.948499656Z
- Data folder ~/gvb-data/engines/weaviate: 2.06 MiB at the start, 5.02 MiB at the end
- Load: 524 rows in batches of 1000, 0.5 s of upserts (994 rows/s); count matched 0.0 s after the last upsert; index step 0.0 s (HNSW is updated inside every batch, nothing to build; confirmed in 0.0 s: shard ZoGlkMkMEzvF vectorIndexingStatus READY, vectorQueueLength 0, node status objectCount 0 (refreshed only when the memtable is flushed, so it lags a minute and does not decide readiness), schema shard ZoGlkMkMEzvF status READY, Aggregate count 524, Aggregate nearVector with objectLimit 524 reached 524, Weaviate reports no count of vectors in the HNSW index, so indexed vectors are not reported)
- Search: 3338 latency samples, target held 524 rows, 0 errors
- Client CPU per search: 1 searcher 0.56 ms, 8 searchers 0.61 ms
- RAM: 94.7 MiB (docker stats); disk: 2.96 MiB added by this load (5.02 MiB (whole engine data folder), 2.06 MiB before)
- Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started weaviate.compose.yaml with every container created on CPUs 2-3,6-7: gvb-weaviate cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards.
- Engine files (SHA-256, first 12 hex digits): [weaviate.compose.yaml d96b7399662a]. This run created the engine from these files.
- Index after the load: ready, ? of 524 indexed (shard ZoGlkMkMEzvF vectorIndexingStatus READY, vectorQueueLength 0, node status objectCount 0 (refreshed only when the memtable is flushed, so it lags a minute and does not decide readiness), schema shard ZoGlkMkMEzvF status READY, Aggregate count 524, Aggregate nearVector with objectLimit 524 reached 524, Weaviate reports no count of vectors in the HNSW index, so indexed vectors are not reported). Durability: Weaviate 1.39.8 defaults, weaviate.compose.yaml sets no persistence variable [Correction 1]: every object is appended to the LSM write-ahead log with a plain write and no fsync, and the log is fsynced only when its memtable is flushed, 60 seconds after the last write (PERSISTENCE_MEMTABLES_FLUSH_IDLE_AFTER_SECONDS default; measured with strace: 2,080 writes into the objects log during a 2,000-object load, the first fsync 60 s after the last write), so a power cut can lose the last minute of writes; the HNSW commit log is buffered inside the process (77 writes for 2,000 vectors), so a killed process loses the newest vectors from the vector index while their objects survive (measured with docker kill, which is SIGKILL, one second after the load and a restart, two runs each: 0 to 1 of 50, 264 to 267 of 300 and 1,992 to 1,997 of 2,000 stored objects were still reachable through a vector search).
- Outside load when searching began: 0.19 CPUs on average over the last 1 s and 0.19 over the last few seconds (limit 0.3), a quiet box. The load average 1.98 2.87 3.25 (1/5/15 min; it also counts this benchmark's own client and engines, so it is recorded, not judged).
- Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@1 5,059 searches in 30.0 s, default@8 9,297 searches in 30.0 s.
- Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it. default@1: warm-up 15.0 s and 2,507 searches (0 failed) at 1 searcher; p50 of the last windows 5.329, 5.345, 5.312 ms (windows of at least 2 s and 100 searches); trial of 499 searches in 3.0 s at 1 searcher: p50 5.331 ms against the settled p50 5.329 ms, 0% apart (limit 10%); the timed pass p50 5.327 ms, 0% from the settled p50 5.329 ms (limit 10% or 0.1 ms; 0.001 ms); default@8: warm-up 15.0 s and 4,591 searches (0 failed) at 8 searchers; QPS of the last windows 308, 304, 306 (windows of at least 2 s and 100 searches); trial of 923 searches in 3.0 s at 8 searchers: 306 QPS against the settled 306 QPS, 0% apart (limit 10%); the timed pass 306 QPS, 0% from the settled 306 QPS (limit 10%).
- Pass default@1 after rehearsal: 2026-10-07T23:46:40.799Z to 2026-10-07T23:47:00.802Z (20.0 s), 3,338 searches, 0 failed, p50 5.327 ms, mean 5.989 ms, p99 11.762 ms, 166.9 QPS (1000/QPS 5.993 ms).
- Pass default@8 after default@1: 2026-10-07T23:47:18.842Z to 2026-10-07T23:47:38.865Z (20.0 s), 6,129 searches, 0 failed, p50 24.782 ms, mean 26.123 ms, p99 53.577 ms, 306.1 QPS (1000/QPS 3.267 ms).
- Passes in the order run: default@1, default@8; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 22,876 sent, 0 failed (warmupErrors). Index after the searches: ready, ? of 524 indexed (shard ZoGlkMkMEzvF vectorIndexingStatus READY, vectorQueueLength 0, node status objectCount 524 (refreshed only when the memtable is flushed, so it lags a minute and does not decide readiness), schema shard ZoGlkMkMEzvF status READY, Aggregate count 524, Aggregate nearVector with objectLimit 524 reached 524, Weaviate reports no count of vectors in the HNSW index, so indexed vectors are not reported).
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
- Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-weaviate (dd646149338c) on network gvb-weaviate_default at container-ip:8080, not through docker-proxy localhost:8085. Open after the passes: container-ip:8080 x8 (this target).
- CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-weaviate: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-weaviate: 9 thread(s) on CPUs 2-3,6-7 (every thread of the container).
- Clock per pass (median MHz of engine CPUs / client CPUs, each the median of its CPUs' medians and every CPU when the CPUs were not split, their average MHz, governor, CPUs busy outside the benchmark, client and engine CPU per search; the clock was pinned at 3500 MHz (turbo off), and a pass whose median on either group is more than 100 bp (1%) off it, or that has no reading on a group, is flagged): default@1 3492/3492 MHz, average 3492.1/3493.0, performance, outside load 0.13 before the warm-up, 0.16 during, client CPU 0.562 ms per search, engine CPU 8.240 ms per search; default@8 3492/3492 MHz, average 3492.0/3492.8, performance, outside load 0.14 before the warm-up, 0.13 during, client CPU 0.609 ms per search, engine CPU 12.966 ms per search.
pgvector
- Engine: PostgreSQL 17 + pgvector 0.8.7 (compose)
- Index: HNSW vector_cosine_ops m=16 ef_construction=128, hnsw.ef_search=100 per query, float32 vector(n), cosine; exact mode = same query with index scans off (sequential scan)
- Search settings: ef_construction=128, hnsw.ef_search=100, m=16
- Engine settings: HostConfig.Memory = 8589934592 (read: docker inspect HostConfig.Memory (0 = no limit)); HostConfig.CpusetCpus = 2-3,6-7 (read: docker inspect HostConfig.CpusetCpus (empty = every CPU)); HostConfig.NanoCpus = 0 (read: docker inspect HostConfig.NanoCpus (0 = no limit)); shared_buffers = 2GB (read: docker exec gvb-pgvector psql current_setting); enable_seqscan = off (set: src/GenericVectorBuilder.Engines/Sinks/PgVectorSink.cs#SET LOCAL enable_seqscan = off; SET LOCAL jit = off;); jit = off (set: src/GenericVectorBuilder.Engines/Sinks/PgVectorSink.cs#SET LOCAL enable_seqscan = off; SET LOCAL jit = off;)
- Image: pgvector/pgvector:0.8.7-pg17-trixie, id sha256:7a7e9f22015b67edb4bef5c59daeebcd7e74bfa570df6ce60ae01237c8648a84 (the pinned id), tagged on this machine at 2026-10-03T14:04:52.665253039Z
- Data folder ~/gvb-data/engines/pgvector: 335.01 MiB at the start, 342.57 MiB at the end
- Load: 524 rows in batches of 1000, 0.8 s of upserts (618 rows/s); count matched 0.0 s after the last upsert; index step 0.0 s (HNSW is maintained inside every insert, nothing to build; confirmed in 0.0 s: pg_indexes lists gvb_gvbbench_eshoponweb_hnsw (USING hnsw (embedding vector_cosine_ops) WITH (m='16', ef_construction='128')), pg_index valid and ready = True, EXPLAIN of the default search uses Index Scan on it = True, that search returned 10 of 10 rows; PostgreSQL keeps no entry count for an HNSW index, so indexed vectors are not reported)
- Search: 22932 latency samples, target held 524 rows, 0 errors
- Client CPU per search: 1 searcher 0.68 ms, 8 searchers 0.54 ms
- Exact mode: 26,863 searches, p50 2.19 ms, p95 2.42 ms, recall 1.000
- RAM: 100.6 MiB (docker stats); disk: 7.56 MiB added by this load (342.57 MiB (whole engine data folder), 335.01 MiB before)
- Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started pgvector.compose.yaml with every container created on CPUs 2-3,6-7: gvb-pgvector cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards.
- Engine files (SHA-256, first 12 hex digits): [pgvector.compose.yaml 27549563c622]. This run created the engine from these files.
- Index after the load: ready, ? of 524 indexed (pg_indexes lists gvb_gvbbench_eshoponweb_hnsw (USING hnsw (embedding vector_cosine_ops) WITH (m='16', ef_construction='128')), pg_index valid and ready = True, EXPLAIN of the default search uses Index Scan on it = True, that search returned 10 of 10 rows; PostgreSQL keeps no entry count for an HNSW index, so indexed vectors are not reported). Durability: fsync on, synchronous_commit on, full_page_writes on, wal_sync_method fdatasync (PostgreSQL defaults; pgvector.compose.yaml sets only shared_buffers, maintenance_work_mem and max_wal_size): every commit is flushed to the write-ahead log before it returns and HNSW index changes are WAL-logged, so a crash loses no committed row (max_wal_size 4GB only spaces out checkpoints; measured with docker kill, which is SIGKILL, right after a 2,000-vector load and a restart: all 2,000 rows were there and the HNSW index was valid and used by the default search).
- Outside load when searching began: 0.10 CPUs on average over the last 1 s and 0.10 over the last few seconds (limit 0.3), a quiet box. The load average 3.02 2.91 3.20 (1/5/15 min; it also counts this benchmark's own client and engines, so it is recorded, not judged).
- Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@8 124,431 searches in 30.0 s, exact 13,317 searches in 30.0 s, default@1 34,282 searches in 30.0 s.
- Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it. default@8: warm-up 15.0 s and 60,480 searches (0 failed) at 8 searchers; QPS of the last windows 3,974, 4,056, 4,052 (windows of at least 2 s and 100 searches); trial of 12,037 searches in 3.0 s at 8 searchers: 4,011 QPS against the settled 4,052 QPS, 1% apart (limit 10%); the timed pass 4,019 QPS, 1% from the settled 4,052 QPS (limit 10%); exact: warm-up 15.0 s and 6,705 searches (0 failed) at 1 searcher; p50 of the last windows 2.184, 2.210, 2.199 ms (windows of at least 2 s and 100 searches); trial of 1,335 searches in 3.0 s at 1 searcher: p50 2.199 ms against the settled p50 2.199 ms, 0% apart (limit 10%); the timed pass p50 2.193 ms, 0% from the settled p50 2.199 ms (limit 10% or 0.1 ms; 0.006 ms); default@1: warm-up 15.0 s and 17,002 searches (0 failed) at 1 searcher; p50 of the last windows 0.869, 0.870, 0.866 ms (windows of at least 2 s and 100 searches); trial of 3,414 searches in 3.0 s at 1 searcher: p50 0.861 ms against the settled p50 0.869 ms, 1% apart (limit 10%); the timed pass p50 0.852 ms, 2% from the settled p50 0.869 ms (limit 10% or 0.1 ms; 0.017 ms).
- Pass default@8 after rehearsal: 2026-10-07T23:49:36.446Z to 2026-10-07T23:49:56.448Z (20.0 s), 80,377 searches, 0 failed, p50 1.946 ms, mean 1.989 ms, p99 3.532 ms, 4018.6 QPS (1000/QPS 0.249 ms).
- Pass exact after default@8: 2026-10-07T23:50:14.481Z to 2026-10-07T23:51:14.481Z (60.0 s), 26,863 searches, 0 failed, p50 2.193 ms, mean 2.231 ms, p99 3.296 ms, 447.7 QPS (1000/QPS 2.234 ms).
- Pass default@1 after exact: 2026-10-07T23:51:32.511Z to 2026-10-07T23:51:52.512Z (20.0 s), 22,932 searches, 0 failed, p50 0.852 ms, mean 0.870 ms, p99 1.295 ms, 1146.6 QPS (1000/QPS 0.872 ms).
- Passes in the order run: default@8, exact, default@1; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 273,003 sent, 0 failed (warmupErrors). Index after the searches: ready, ? of 524 indexed (pg_indexes lists gvb_gvbbench_eshoponweb_hnsw (USING hnsw (embedding vector_cosine_ops) WITH (m='16', ef_construction='128')), pg_index valid and ready = True, EXPLAIN of the default search uses Index Scan on it = True, that search returned 10 of 10 rows; PostgreSQL keeps no entry count for an HNSW index, so indexed vectors are not reported).
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
- Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-pgvector (19dd65db0d9a) on network gvb-pgvector_default at container-ip:5432, not through docker-proxy localhost:5432. Open after the passes: container-ip:5432 x8 (this target).
- CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-pgvector: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-pgvector: 6 thread(s) on CPUs 2-3,6-7 (every thread of the container).
- Clock per pass (median MHz of engine CPUs / client CPUs, each the median of its CPUs' medians and every CPU when the CPUs were not split, their average MHz, governor, CPUs busy outside the benchmark, client and engine CPU per search; the clock was pinned at 3500 MHz (turbo off), and a pass whose median on either group is more than 100 bp (1%) off it, or that has no reading on a group, is flagged): default@8 3492/3492 MHz, average 3492.0/3492.0, performance, outside load 0.13 before the warm-up, 0.18 during, client CPU 0.538 ms per search, engine CPU 0.972 ms per search; exact 3492/3492 MHz, average 3493.0/3492.7, performance, outside load 0.15 before the warm-up, 0.13 during, client CPU 0.680 ms per search, engine CPU 2.066 ms per search; default@1 3492/3492 MHz, average 3493.1/3492.1, performance, outside load 0.13 before the warm-up, 0.16 during, client CPU 0.678 ms per search, engine CPU 0.716 ms per search.
MongoDB Atlas Local mongodb
- Engine: MongoDB 8.0.32 Atlas Local (mongod + mongot Vector Search) (compose)
- Index: vectorSearch index, HNSW maxEdges=16 numEdgeCandidates=128, float32 binData, cosine, numCandidates=20x hits (min 100); exact mode = $vectorSearch exact:true
- Search settings: maxEdges=16, numCandidates=20x, numEdgeCandidates=128
- Engine settings: HostConfig.Memory = 8589934592 (read: docker inspect HostConfig.Memory (0 = no limit)); HostConfig.CpusetCpus = 2-3,6-7 (read: docker inspect HostConfig.CpusetCpus (empty = every CPU)); HostConfig.NanoCpus = 0 (read: docker inspect HostConfig.NanoCpus (0 = no limit))
- Image: mongodb/mongodb-atlas-local:8.0.32, id sha256:1985314b0ded756ba965e0f400d17406b7a89bec659d4763a490f42f006e4fbd (the pinned id), tagged on this machine at 2026-10-03T14:11:20.531944474Z
- Data folder ~/gvb-data/engines/mongodb: 1.38 GiB at the start, 1.38 GiB at the end
- Load: 524 rows in batches of 1000, 0.2 s of upserts (2,741 rows/s); count matched 2.3 s after the last upsert; index step 10.2 s (index status READY, queryable True; mongot holds 524 of 524 documents in 4 segment(s) (220 + 216 + 75 + 13 documents), 4 searched through the HNSW graph (Approximate); segment layout unchanged for 10.2 s over 11 readings; wait took 10.2 s)
- Search: 14106 latency samples, target held 524 rows, 0 errors
- Client CPU per search: 1 searcher 0.30 ms, 8 searchers 0.28 ms
- Exact mode: 48,588 searches, p50 1.21 ms, p95 1.38 ms, recall 1.000
- RAM: 647.3 MiB (docker stats); disk: 10.48 KiB added by this load (1.38 GiB (whole engine data folder), 1.38 GiB before)
- Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started mongodb.compose.yaml with every container created on CPUs 2-3,6-7: gvb-mongodb cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards.
- Engine files (SHA-256, first 12 hex digits): [mongodb.compose.yaml 4ee6f6411f1f]. This run created the engine from these files.
- Index after the load: ready, 524 of 524 indexed (index status READY, queryable True; mongot holds 524 of 524 documents in 4 segment(s) (220 + 216 + 75 + 13 documents), 4 searched through the HNSW graph (Approximate)). Durability: Acknowledged writes are journaled before the acknowledgement. The sink sets no write concern, so the server default applies: getDefaultRWConcern gives w majority and the one-member replica set (--replSet gvbmongo in the image) has writeConcernMajorityJournalDefault true, with WiredTiger journaling on (journalCommitInterval 100 ms is only the interval for unacknowledged work). Measured 2026-10-04: 30 acknowledged single-document inserts raised WiredTiger 'log sync operations' by 30 and strace showed fdatasync on /data/db/journal/WiredTigerLog files. A crash loses no acknowledged write. mongot's search index is not part of that promise: it follows the collection asynchronously and is rebuilt from it..
- Outside load when searching began: 0.16 CPUs on average over the last 14 s and 0.14 over the last few seconds (limit 0.3), a quiet box. The load average 2.44 3.31 3.37 (1/5/15 min; it also counts this benchmark's own client and engines, so it is recorded, not judged).
- Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: exact 20,720 searches in 30.0 s, default@8 56,953 searches in 30.0 s, default@1 21,266 searches in 30.0 s.
- Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it. exact: warm-up 15.0 s and 12,075 searches (0 failed) at 1 searcher; p50 of the last windows 1.217, 1.229, 1.213 ms (windows of at least 2 s and 100 searches); trial of 2,416 searches in 3.0 s at 1 searcher: p50 1.211 ms against the settled p50 1.217 ms, 0% apart (limit 10%); the timed pass p50 1.211 ms, 0% from the settled p50 1.217 ms (limit 10% or 0.1 ms; 0.006 ms); default@8: warm-up 15.0 s and 31,448 searches (0 failed) at 8 searchers; QPS of the last windows 2,077, 2,122, 2,112 (windows of at least 2 s and 100 searches); trial of 6,324 searches in 3.0 s at 8 searchers: 2,106 QPS against the settled 2,112 QPS, 0% apart (limit 10%); the timed pass 2,098 QPS, 1% from the settled 2,112 QPS (limit 10%); default@1: warm-up 15.0 s and 10,578 searches (0 failed) at 1 searcher; p50 of the last windows 1.386, 1.392, 1.380 ms (windows of at least 2 s and 100 searches); trial of 2,096 searches in 3.0 s at 1 searcher: p50 1.398 ms against the settled p50 1.386 ms, 1% apart (limit 10%); the timed pass p50 1.389 ms, 0% from the settled p50 1.386 ms (limit 10% or 0.1 ms; 0.003 ms).
- Pass exact after rehearsal: 2026-10-07T23:54:02.447Z to 2026-10-07T23:55:02.448Z (60.0 s), 48,588 searches, 0 failed, p50 1.211 ms, mean 1.233 ms, p99 1.722 ms, 809.8 QPS (1000/QPS 1.235 ms).
- Pass default@8 after exact: 2026-10-07T23:55:20.505Z to 2026-10-07T23:55:40.507Z (20.0 s), 41,971 searches, 0 failed, p50 3.691 ms, mean 3.810 ms, p99 6.771 ms, 2098.3 QPS (1000/QPS 0.477 ms).
- Pass default@1 after default@8: 2026-10-07T23:55:58.522Z to 2026-10-07T23:56:18.522Z (20.0 s), 14,106 searches, 0 failed, p50 1.389 ms, mean 1.416 ms, p99 1.967 ms, 705.3 QPS (1000/QPS 1.418 ms).
- Passes in the order run: exact, default@8, default@1; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 163,876 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (index status READY, queryable True; mongot holds 524 of 524 documents in 4 segment(s) (220 + 216 + 75 + 13 documents), 4 searched through the HNSW graph (Approximate)).
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
- Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-mongodb (6b3e3a92aacb) on network gvb-mongodb_default at container-ip:27017, not through docker-proxy localhost:27017. Open after the passes: container-ip:27017 x10 (this target).
- CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-mongodb: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-mongodb: 211 thread(s) on CPUs 2-3,6-7 (every thread of the container).
- Clock per pass (median MHz of engine CPUs / client CPUs, each the median of its CPUs' medians and every CPU when the CPUs were not split, their average MHz, governor, CPUs busy outside the benchmark, client and engine CPU per search; the clock was pinned at 3500 MHz (turbo off), and a pass whose median on either group is more than 100 bp (1%) off it, or that has no reading on a group, is flagged): exact 3492/3492 MHz, average 3492.4/3492.5, performance, outside load 0.19 before the warm-up, 0.22 during, client CPU 0.296 ms per search, engine CPU 1.141 ms per search; default@8 3492/3492 MHz, average 3492.0/3492.2, performance, outside load 0.22 before the warm-up, 0.14 during, client CPU 0.285 ms per search, engine CPU 1.812 ms per search; default@1 3492/3492 MHz, average 3491.7/3492.6, performance, outside load 0.18 before the warm-up, 0.23 during, client CPU 0.303 ms per search, engine CPU 1.327 ms per search.
OpenSearch opensearch
- Engine: OpenSearch 3.9.0 (k-NN plugin, faiss) (compose)
- Index: faiss HNSW float32, no compression, m=16, ef_construction=128, cosinesimil; search k=top, ef_search=100; 1 shard, 0 replicas; graph built at any segment size (approximate_threshold=0); force-merged to one segment after the load
- Search settings: approximate_threshold=0, ef_construction=128, ef_search=100, k=top, m=16
- Engine settings: HostConfig.Memory = 6442450944 (read: docker inspect HostConfig.Memory (0 = no limit)); HostConfig.CpusetCpus = 2-3,6-7 (read: docker inspect HostConfig.CpusetCpus (empty = every CPU)); HostConfig.NanoCpus = 0 (read: docker inspect HostConfig.NanoCpus (0 = no limit)); heap_init_in_bytes = 2147483648 (read: GET /_nodes/jvm (nodes.<id>.jvm.mem.heap_init_in_bytes and heap_max_in_bytes)); heap_max_in_bytes = 2147483648 (read: GET /_nodes/jvm (nodes.<id>.jvm.mem.heap_init_in_bytes and heap_max_in_bytes))
- Image: opensearchproject/opensearch:3.9.0, id sha256:adfa61f85025d06b4aeb562e7e74fde7e31c437039c93c3862c17e9acebd6c7c (the pinned id), tagged on this machine at 2026-10-03T14:09:29.512875356Z
- Data folder ~/gvb-data/engines/opensearch: 4.23 MiB at the start, 13.89 MiB at the end
- Load: 524 rows in batches of 1000, 1.7 s of upserts (314 rows/s); count matched 0.1 s after the last upsert; index step 0.9 s (force-merged to one segment, graph ready after 0.9 s: every segment searched through its HNSW graph, covering 524 of 524 vectors: 1 segment(s), 0 merge(s) running, profiled probe search: ann_search_count 1 (segments tried through a graph), exact_search_count 0 (of those, scanned instead), approximate_threshold 0; force-merged to one segment on purpose, so one graph answers every search)
- Search: 9900 latency samples, target held 524 rows, 0 errors
- Client CPU per search: 1 searcher 0.47 ms, 8 searchers 0.50 ms
- Exact mode: 22,879 searches, p50 2.56 ms, p95 2.95 ms, recall 1.000
- RAM: 2.6 GiB (docker stats); disk: 9.66 MiB added by this load (13.89 MiB (whole engine data folder), 4.23 MiB before)
- Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started opensearch.compose.yaml with every container created on CPUs 2-3,6-7: gvb-opensearch cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards.
- Engine files (SHA-256, first 12 hex digits): [opensearch.compose.yaml 028b44c40129]. This run created the engine from these files.
- Index after the load: ready, 524 of 524 indexed (every segment searched through its HNSW graph, covering 524 of 524 vectors: 1 segment(s), 0 merge(s) running, profiled probe search: ann_search_count 1 (segments tried through a graph), exact_search_count 0 (of those, scanned instead), approximate_threshold 0; force-merged to one segment on purpose, so one graph answers every search). Durability: Every acknowledged bulk request is fsynced to the translog before the answer: index.translog.durability=request, the OpenSearch default, which neither this sink nor opensearch.compose.yaml overrides (OpenSearchReadinessTests reads it back from the live index). By that setting a process crash or power loss loses no acknowledged write; this is read from the setting, not shown by pulling power. One node and no replicas, so a lost disk loses the data..
- Outside load when searching began: 0.10 CPUs on average over the last 3 s and 0.10 over the last few seconds (limit 0.3), a quiet box. The load average 2.18 2.96 3.24 (1/5/15 min; it also counts this benchmark's own client and engines, so it is recorded, not judged).
- Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: exact 8,317 searches in 30.0 s, default@1 13,531 searches in 30.0 s, default@8 42,336 searches in 30.0 s.
- Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it. exact: warm-up 15.0 s and 5,413 searches (0 failed) at 1 searcher; p50 of the last windows 2.635, 2.640, 2.646 ms (windows of at least 2 s and 100 searches); trial of 1,122 searches in 3.0 s at 1 searcher: p50 2.602 ms against the settled p50 2.640 ms, 1% apart (limit 10%); the timed pass p50 2.562 ms, 3% from the settled p50 2.640 ms (limit 10% or 0.1 ms; 0.077 ms); default@1: warm-up 15.0 s and 7,405 searches (0 failed) at 1 searcher; p50 of the last windows 1.991, 1.980, 1.988 ms (windows of at least 2 s and 100 searches); trial of 1,440 searches in 3.0 s at 1 searcher: p50 1.999 ms against the settled p50 1.988 ms, 1% apart (limit 10%); the timed pass p50 1.983 ms, 0% from the settled p50 1.988 ms (limit 10% or 0.1 ms; 0.005 ms); default@8: warm-up 15.0 s and 22,393 searches (0 failed) at 8 searchers; QPS of the last windows 1,514, 1,504, 1,516 (windows of at least 2 s and 100 searches); trial of 4,544 searches in 3.0 s at 8 searchers: 1,513 QPS against the settled 1,514 QPS, 0% apart (limit 10%); the timed pass 1,507 QPS, 0% from the settled 1,514 QPS (limit 10%).
- Pass exact after rehearsal: 2026-10-07T23:59:38.514Z to 2026-10-08T00:00:38.516Z (60.0 s), 22,879 searches, 0 failed, p50 2.562 ms, mean 2.620 ms, p99 3.660 ms, 381.3 QPS (1000/QPS 2.623 ms).
- Pass default@1 after exact: 2026-10-08T00:00:56.540Z to 2026-10-08T00:01:16.540Z (20.0 s), 9,900 searches, 0 failed, p50 1.983 ms, mean 2.018 ms, p99 2.757 ms, 495.0 QPS (1000/QPS 2.020 ms).
- Pass default@8 after default@1: 2026-10-08T00:01:34.559Z to 2026-10-08T00:01:54.562Z (20.0 s), 30,150 searches, 0 failed, p50 5.090 ms, mean 5.305 ms, p99 10.327 ms, 1507.3 QPS (1000/QPS 0.663 ms).
- Passes in the order run: exact, default@1, default@8; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 106,501 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (every segment searched through its HNSW graph, covering 524 of 524 vectors: 1 segment(s), 0 merge(s) running, profiled probe search: ann_search_count 1 (segments tried through a graph), exact_search_count 0 (of those, scanned instead), approximate_threshold 0; force-merged to one segment on purpose, so one graph answers every search).
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
- Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-opensearch (5e29495635a0) on network engines_default at container-ip:9200, not through docker-proxy localhost:9201. Open after the passes: container-ip:9200 x8 (this target).
- CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-opensearch: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-opensearch: 56 thread(s) on CPUs 2-3,6-7 (every thread of the container).
- Clock per pass (median MHz of engine CPUs / client CPUs, each the median of its CPUs' medians and every CPU when the CPUs were not split, their average MHz, governor, CPUs busy outside the benchmark, client and engine CPU per search; the clock was pinned at 3500 MHz (turbo off), and a pass whose median on either group is more than 100 bp (1%) off it, or that has no reading on a group, is flagged): exact 3492/3492 MHz, average 3492.3/3493.3, performance, outside load 0.12 before the warm-up, 0.16 during, client CPU 0.462 ms per search, engine CPU 2.407 ms per search; default@1 3492/3492 MHz, average 3491.9/3493.2, performance, outside load 0.16 before the warm-up, 0.15 during, client CPU 0.466 ms per search, engine CPU 1.750 ms per search; default@8 3492/3492 MHz, average 3492.0/3492.4, performance, outside load 0.15 before the warm-up, 0.14 during, client CPU 0.500 ms per search, engine CPU 2.554 ms per search.
Vespa vespa
- Engine: Vespa 8.754.14 (tensor attribute + HNSW) (compose)
- Index: HNSW float32 tensor, prenormalized-angular (cosine), max-links-per-node=16, neighbors-to-explore-at-insert=128; search targetHits=top, ef=100 via exploreAdditionalHits; exact mode = approximate:false; vectors held in memory
- Search settings: ef=100, max-links-per-node=16, neighbors-to-explore-at-insert=128, targetHits=top
- Engine settings: HostConfig.Memory = 6442450944 (read: docker inspect HostConfig.Memory (0 = no limit)); HostConfig.CpusetCpus = 2-3,6-7 (read: docker inspect HostConfig.CpusetCpus (empty = every CPU)); HostConfig.NanoCpus = 0 (read: docker inspect HostConfig.NanoCpus (0 = no limit))
- Image: vespaengine/vespa:8.754.14, id sha256:5c30f5c41e7563498c4f925db6a837a3848f04726a3ed26aed4a7c8ab69f18fd (the pinned id), tagged on this machine at 2026-10-03T14:13:52.169232039Z
- Data folder ~/gvb-data/engines/vespa: 1.9 GiB at the start [Correction 9], 2.13 GiB at the end
- Load: 524 rows in batches of 1000, 1.8 s of upserts (291 rows/s); count matched 0.0 s after the last upsert; index step 0.1 s (nothing to build, Vespa inserts into the HNSW graph while writing; ready after 0.1 s: HNSW graph returns 524 of 524 vectors: totalCount 524, coverage full True, nearestNeighbor approximate:true with targetHits above the count: has_index True, algorithm "index top k", top_k_hits 524; the graph is updated while each write is applied, so there is no later build step)
- Search: 10343 latency samples, target held 524 rows, 0 errors
- Client CPU per search: 1 searcher 0.92 ms, 8 searchers 0.87 ms
- Exact mode: 31,102 searches, p50 1.87 ms, p95 2.31 ms, recall 1.000
- RAM: 2.95 GiB (docker stats); disk: 235.33 MiB added by this load [Correction 10] (2.13 GiB (whole engine data folder), 1.9 GiB before)
- Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started vespa.compose.yaml with every container created on CPUs 2-3,6-7: gvb-vespa cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards.
- Engine files (SHA-256, first 12 hex digits): [vespa.compose.yaml 3cfb702cc29f]. This run created the engine from these files.
- Index after the load: ready, 524 of 524 indexed (HNSW graph returns 524 of 524 vectors: totalCount 524, coverage full True, nearestNeighbor approximate:true with targetHits above the count: has_index True, algorithm "index top k", top_k_hits 524; the graph is updated while each write is applied, so there is no later build step). Durability: Vespa's transaction log server fsyncs after each commit: searchlib.translogserver usefsync=true, and an operation is searchable when it is acknowledged: proton documentdb visibilitydelay=0. Neither is overridden by this sink or by the services.xml it deploys; VespaReadinessTests reads both from the live config server. After a crash the node replays the transaction log over its last flushed data and rebuilds the in-memory HNSW graph. By those settings an acknowledged write survives a crash; this is read from the settings, not shown by pulling power. One node and min-redundancy 1, so a lost disk loses the data..
- Outside load when searching began: 0.17 CPUs on average over the last 23 s and 0.16 over the last few seconds (limit 0.3), a quiet box. The load average 6.21 3.81 3.44 (1/5/15 min; it also counts this benchmark's own client and engines, so it is recorded, not judged).
- Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: exact 8,862 searches in 30.0 s, default@1 12,858 searches in 30.0 s, default@8 44,746 searches in 30.0 s.
- Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it. exact: warm-up 15.0 s and 7,257 searches (0 failed) at 1 searcher; p50 of the last windows 1.939, 1.922, 1.891 ms (windows of at least 2 s and 100 searches); trial of 1,547 searches in 3.0 s at 1 searcher: p50 1.878 ms against the settled p50 1.922 ms, 2% apart (limit 10%); the timed pass p50 1.868 ms, 3% from the settled p50 1.922 ms (limit 10% or 0.1 ms; 0.054 ms); default@1: warm-up 15.0 s and 7,781 searches (0 failed) at 1 searcher; p50 of the last windows 1.888, 1.865, 1.877 ms (windows of at least 2 s and 100 searches); trial of 1,575 searches in 3.0 s at 1 searcher: p50 1.867 ms against the settled p50 1.877 ms, 1% apart (limit 10%); the timed pass p50 1.871 ms, 0% from the settled p50 1.877 ms (limit 10% or 0.1 ms; 0.006 ms); default@8: warm-up 15.0 s and 23,496 searches (0 failed) at 8 searchers; QPS of the last windows 1,509, 1,578, 1,566 (windows of at least 2 s and 100 searches); trial of 4,736 searches in 3.0 s at 8 searchers: 1,578 QPS against the settled 1,566 QPS, 1% apart (limit 10%); the timed pass 1,582 QPS, 1% from the settled 1,566 QPS (limit 10%).
- Pass exact after rehearsal: 2026-10-08T00:04:25.199Z to 2026-10-08T00:05:25.200Z (60.0 s), 31,102 searches, 0 failed, p50 1.868 ms, mean 1.926 ms, p99 2.641 ms, 518.4 QPS (1000/QPS 1.929 ms).
- Pass default@1 after exact: 2026-10-08T00:05:43.233Z to 2026-10-08T00:06:03.235Z (20.0 s), 10,343 searches, 0 failed, p50 1.871 ms, mean 1.931 ms, p99 2.657 ms, 517.1 QPS (1000/QPS 1.934 ms).
- Pass default@8 after default@1: 2026-10-08T00:06:21.252Z to 2026-10-08T00:06:41.255Z (20.0 s), 31,640 searches, 0 failed, p50 4.933 ms, mean 5.054 ms, p99 8.372 ms, 1581.8 QPS (1000/QPS 0.632 ms).
- Passes in the order run: exact, default@1, default@8; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 112,858 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (HNSW graph returns 524 of 524 vectors: totalCount 524, coverage full True, nearestNeighbor approximate:true with targetHits above the count: has_index True, algorithm "index top k", top_k_hits 524; the graph is updated while each write is applied, so there is no later build step).
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
- Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-vespa (2461332c111e) on network engines_default at container-ip:8080, not through docker-proxy localhost:8090; container gvb-vespa (2461332c111e) on network engines_default at container-ip:19071, not through docker-proxy localhost:19071. Open after the passes: container-ip:8080 x1 (this target).
- CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-vespa: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-vespa: 605 thread(s) on CPUs 2-3,6-7 (every thread of the container).
- Clock per pass (median MHz of engine CPUs / client CPUs, each the median of its CPUs' medians and every CPU when the CPUs were not split, their average MHz, governor, CPUs busy outside the benchmark, client and engine CPU per search; the clock was pinned at 3500 MHz (turbo off), and a pass whose median on either group is more than 100 bp (1%) off it, or that has no reading on a group, is flagged): exact 3492/3492 MHz, average 3492.0/3492.3, performance, outside load 0.15 before the warm-up, 0.11 during, client CPU 0.915 ms per search, engine CPU 2.263 ms per search; default@1 3492/3492 MHz, average 3492.0/3492.3, performance, outside load 0.11 before the warm-up, 0.12 during, client CPU 0.916 ms per search, engine CPU 2.260 ms per search; default@8 3492/3492 MHz, average 3492.0/3492.1, performance, outside load 0.11 before the warm-up, 0.20 during, client CPU 0.873 ms per search, engine CPU 2.422 ms per search.
Redis redis
- Engine: Redis 8.10.2 (query engine, HASH + vector index) (compose)
- Index: HNSW TYPE FLOAT32 M=16 EF_CONSTRUCTION=128, EF_RUNTIME=100 per query, cosine; exact mode = FLAT index built on first exact query
- Search settings: EF_CONSTRUCTION=128, EF_RUNTIME=100, M=16
- Engine settings: HostConfig.Memory = 12884901888 (read: docker inspect HostConfig.Memory (0 = no limit)); HostConfig.CpusetCpus = 2-3,6-7 (read: docker inspect HostConfig.CpusetCpus (empty = every CPU)); HostConfig.NanoCpus = 0 (read: docker inspect HostConfig.NanoCpus (0 = no limit)); search-workers = 8 (read: docker exec gvb-redis redis-cli CONFIG GET search-workers); save = 300 1 (read: docker exec gvb-redis redis-cli CONFIG GET save); appendonly = no (read: docker exec gvb-redis redis-cli CONFIG GET appendonly); maxmemory = 0 (read: docker exec gvb-redis redis-cli CONFIG GET maxmemory)
- Image: redis:8.10.2, id sha256:6f81e8915c60b065a524e6967e0ad1c639ba6efa84d669f823683ea04d9150ee (the pinned id), tagged on this machine at 2026-10-03T14:09:34.185138197Z
- Data folder ~/gvb-data/engines/redis: 94.22 MiB at the start, 96.65 MiB at the end
- Load: 524 rows in batches of 1000, 0.0 s of upserts (11,812 rows/s); count matched 1.0 s after the last upsert; index step 0.0 s (HNSW index finished after 0.0 s of waiting: FT.INFO indexing 0, percent_indexed 1, hash_indexing_failures 0, flat_buffer_size 0, num_docs 524, hashes counted with SCAN under gvb_gvbbench_eshoponweb: 524)
- Search: 51187 latency samples, target held 524 rows, 0 errors
- Client CPU per search: 1 searcher 0.56 ms, 8 searchers 0.35 ms
- Exact mode: 164,112 searches, p50 0.36 ms, p95 0.46 ms, recall 1.000
- RAM: 253.8 MiB (docker stats); disk: 2.43 MiB added by this load (96.65 MiB (whole engine data folder), 94.22 MiB before)
- Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started redis.compose.yaml with every container created on CPUs 2-3,6-7: gvb-redis cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards.
- Engine files (SHA-256, first 12 hex digits): [redis.compose.yaml 7448634e0d92]. This run created the engine from these files.
- Index after the load: ready, 524 of 524 indexed (FT.INFO indexing 0, percent_indexed 1, hash_indexing_failures 0, flat_buffer_size 0, num_docs 524, hashes counted with SCAN under gvb_gvbbench_eshoponweb: 524). Durability: save "300 1" and appendonly no (redis.compose.yaml): an RDB snapshot is written every 5 minutes if at least one key changed, and on a clean stop; there is no append-only log, so a crash loses every write since the last snapshot (up to 5 minutes plus the time a snapshot takes; measured with docker kill, which is SIGKILL, right after a 2,000-vector load: none of the 2,000 hashes were there after the restart, and the restart loaded an older snapshot that still held keys of a collection that had been dropped since).
- Outside load when searching began: 0.10 CPUs on average over the last 2 s and 0.10 over the last few seconds (limit 0.3), a quiet box. The load average 7.44 5.06 4.00 (1/5/15 min; it also counts this benchmark's own client and engines, so it is recorded, not judged).
- Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@1 78,173 searches in 30.0 s, default@8 153,380 searches in 30.0 s, exact 78,752 searches in 30.0 s.
- Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it. default@1: warm-up 15.0 s and 38,373 searches (0 failed) at 1 searcher; p50 of the last windows 0.380, 0.393, 0.384 ms (windows of at least 2 s and 100 searches); trial of 7,686 searches in 3.0 s at 1 searcher: p50 0.393 ms against the settled p50 0.384 ms, 2% apart (limit 10%); the timed pass p50 0.392 ms, 2% from the settled p50 0.384 ms (limit 10% or 0.1 ms; 0.007 ms); default@8: warm-up 15.0 s and 76,578 searches (0 failed) at 8 searchers; QPS of the last windows 5,193, 5,047, 5,060 (windows of at least 2 s and 100 searches); trial of 15,729 searches in 3.0 s at 8 searchers: 5,241 QPS against the settled 5,060 QPS, 3% apart (limit 10%); the timed pass 5,258 QPS, 4% from the settled 5,060 QPS (limit 10%); exact: warm-up 15.0 s and 41,233 searches (0 failed) at 1 searcher; p50 of the last windows 0.376, 0.358, 0.356 ms (windows of at least 2 s and 100 searches) (its last 3 windows differed by more than 5%); trial of 8,203 searches in 3.0 s at 1 searcher: p50 0.362 ms against the settled p50 0.358 ms, 1% apart (limit 10%), so the warm-up was EXTENDED once (30.0 s and 81,575 searches (0 failed) at 1 searcher; p50 of its older 6 windows 0.366 ms (windows 0.375, 0.376, 0.368, 0.361, 0.343, 0.369) and of its newer 6 0.366 ms (windows 0.375, 0.373, 0.372, 0.353, 0.369, 0.341), 0.0% apart (limit 5% or 0.1 ms; 0.000 ms) (windows of at least 2 s and 100 searches)); second trial of 8,193 searches in 3.0 s at 1 searcher: p50 0.367 ms against the settled p50 0.366 ms, 0% apart (limit 10%), inside its newer windows' range 0.341 to 0.375 ms; settled after the extension; the timed pass p50 0.365 ms, 0% from the settled p50 0.366 ms (limit 10% or 0.1 ms; 0.001 ms).
- Pass default@1 after rehearsal: 2026-10-08T00:08:57.317Z to 2026-10-08T00:09:17.317Z (20.0 s), 51,187 searches, 0 failed, p50 0.392 ms, mean 0.389 ms, p99 0.701 ms, 2559.3 QPS (1000/QPS 0.391 ms).
- Pass default@8 after default@1: 2026-10-08T00:09:35.382Z to 2026-10-08T00:09:55.384Z (20.0 s), 105,161 searches, 0 failed, p50 1.480 ms, mean 1.520 ms, p99 2.154 ms, 5257.6 QPS (1000/QPS 0.190 ms).
- Pass exact after default@8: 2026-10-08T00:11:01.447Z to 2026-10-08T00:12:01.447Z (60.0 s), 164,112 searches, 0 failed, p50 0.365 ms, mean 0.364 ms, p99 0.698 ms, 2735.2 QPS (1000/QPS 0.366 ms).
- Passes in the order run: default@1, default@8, exact; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 627,900 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (FT.INFO indexing 0, percent_indexed 1, hash_indexing_failures 0, flat_buffer_size 0, num_docs 524, hashes counted with SCAN under gvb_gvbbench_eshoponweb: 524).
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
- Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-redis (6b2aef88edc2) on network gvb-redis_default at container-ip:6379, not through docker-proxy localhost:6379. Open after the passes: container-ip:6379 x1 (this target).
- CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-redis: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-redis: 16 thread(s) on CPUs 2-3,6-7 (every thread of the container).
- Clock per pass (median MHz of engine CPUs / client CPUs, each the median of its CPUs' medians and every CPU when the CPUs were not split, their average MHz, governor, CPUs busy outside the benchmark, client and engine CPU per search; the clock was pinned at 3500 MHz (turbo off), and a pass whose median on either group is more than 100 bp (1%) off it, or that has no reading on a group, is flagged): default@1 3492/3492 MHz, average 3492.5/3492.2, performance, outside load 0.11 before the warm-up, 0.12 during, client CPU 0.562 ms per search, engine CPU 0.260 ms per search; default@8 3492/3492 MHz, average 3492.2/3491.9, performance, outside load 0.12 before the warm-up, 0.12 during, client CPU 0.354 ms per search, engine CPU 0.230 ms per search; exact 3492/3492 MHz, average 3492.3/3492.0, performance, outside load 0.12 before the warm-up, 0.14 during, client CPU 0.555 ms per search, engine CPU 0.242 ms per search.
DuckDB duckdb
- Engine: DuckDB 1.5.6 + vss b833341 (HNSW, embedded, in process) (embedded)
- Index: HNSW (vss extension) FLOAT[n] metric=cosine m=16 ef_construction=128, ef_search=100 per connection, persistent (hnsw_enable_experimental_persistence=true, checkpoint_threshold=256MB); exact mode = array_cosine_similarity sequential scan; score = 1 - cosine distance; searches run concurrently, one connection per searcher (opened as searchers arrive, at most 32), writes run one at a time and never overlap a search
- Search settings: checkpoint_threshold=256MB, ef_construction=128, ef_search=100, hnsw_enable_experimental_persistence=true, m=16, metric=cosine
- Engine settings: duckdb_version = v1.5.6 (read: SELECT version() on an in-memory DuckDB connection); threads = 8 (read: SELECT current_setting('threads') on an in-memory DuckDB connection; DuckDB's own thread count, not a CPU cap (it counts all CPUs and ignores affinity)); hnsw_ef_search = 100 (set: src/GenericVectorBuilder.Engines/Sinks/DuckDbSink.cs#SET hnsw_ef_search); hnsw_enable_experimental_persistence = true (set: src/GenericVectorBuilder.Engines/Sinks/DuckDbSink.cs#SET hnsw_enable_experimental_persistence = true); memory_limit = 7.4 GiB (read: SELECT current_setting('memory_limit') on a second connection to the bench file opened with the sink's own connection string); checkpoint_threshold = 244.1 MiB (read: SELECT current_setting('checkpoint_threshold') on a second connection to the bench file opened with the sink's own connection string)
- Data folder ~/gvb-data/engines/duckdb: 1.26 GiB at the start, 1.26 GiB at the end; the benchmark's own database file 2.9 MiB
- Load: 524 rows in batches of 1000, 0.4 s of upserts (1,433 rows/s); count matched 0.0 s after the last upsert; index step 0.0 s (HNSW is maintained inside every transaction, nothing to build; confirmed in 0.01 s: duckdb_indexes() lists gvb_gvbbench_eshoponweb_hnsw, pragma_hnsw_index_info() counts 524 vectors of 524 rows, EXPLAIN of the default search shows HNSW_INDEX_SCAN on it = True)
- Search: 5432 latency samples, target held 524 rows, 0 errors
- Client CPU per search: 1 searcher 5.01 ms, 8 searchers 6.83 ms
- Exact mode: 14,326 searches, p50 4.10 ms, p95 4.82 ms, recall 1.000
- RAM: in this process, not measured; disk: 2.9 MiB added by this load (1.26 GiB (engine data folder), 1.26 GiB before)
- Index after the load: ready, 524 of 524 indexed (duckdb_indexes() lists gvb_gvbbench_eshoponweb_hnsw, pragma_hnsw_index_info() counts 524 vectors of 524 rows, EXPLAIN of the default search shows HNSW_INDEX_SCAN on it = True). Durability: DuckDB's write-ahead log is fsynced at every commit and no DuckDbSink setting changes that (checkpoint_threshold=256MB only spaces out the checkpoints that write the database file; measured with strace: 42 WAL fsyncs for 40 single-record commits plus 2 setup statements), so an operating-system crash or power cut loses no committed row; measured with kill -9 right after the last commit and a reopen: all 50, 300 and 2,000 rows and an HNSW index that counted the same number were recovered from the log; not tested and documented by DuckDB: the HNSW index is file-backed only through hnsw_enable_experimental_persistence = true, which this sink turns on, and WAL recovery for such custom indexes is not complete, so a crash during a checkpoint or a later commit can damage the index while the rows survive.
- Outside load when searching began: not known yet (too little sampled history). Load average 2.63 3.60 3.66 (1/5/15 min; it also counts this benchmark's own client and engines, so it is recorded, not judged).
- Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: exact 7,261 searches in 30.0 s, default@1 8,237 searches in 30.0 s, default@8 17,473 searches in 30.0 s.
- Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it. exact: warm-up 15.0 s and 3,581 searches (0 failed) at 1 searcher; p50 of the last windows 4.092, 4.122, 4.099 ms (windows of at least 2 s and 100 searches); trial of 711 searches in 3.0 s at 1 searcher: p50 4.118 ms against the settled p50 4.099 ms, 0% apart (limit 10%); the timed pass p50 4.101 ms, 0% from the settled p50 4.099 ms (limit 10% or 0.1 ms; 0.002 ms); default@1: warm-up 15.0 s and 4,112 searches (0 failed) at 1 searcher; p50 of the last windows 3.623, 3.608, 3.635 ms (windows of at least 2 s and 100 searches); trial of 816 searches in 3.0 s at 1 searcher: p50 3.635 ms against the settled p50 3.623 ms, 0% apart (limit 10%); the timed pass p50 3.636 ms, 0% from the settled p50 3.623 ms (limit 10% or 0.1 ms; 0.013 ms); default@8: warm-up 15.0 s and 8,681 searches (0 failed) at 8 searchers; QPS of the last windows 581, 574, 576 (windows of at least 2 s and 100 searches); trial of 1,766 searches in 3.0 s at 8 searchers: 587 QPS against the settled 576 QPS, 2% apart (limit 10%); the timed pass 583 QPS, 1% from the settled 576 QPS (limit 10%).
- Pass exact after rehearsal: 2026-10-08T00:13:52.440Z to 2026-10-08T00:14:52.440Z (60.0 s), 14,326 searches, 0 failed, p50 4.101 ms, mean 4.185 ms, p99 5.783 ms, 238.8 QPS (1000/QPS 4.188 ms).
- Pass default@1 after exact: 2026-10-08T00:15:10.461Z to 2026-10-08T00:15:30.461Z (20.0 s), 5,432 searches, 0 failed, p50 3.636 ms, mean 3.679 ms, p99 5.308 ms, 271.6 QPS (1000/QPS 3.682 ms).
- Pass default@8 after default@1: 2026-10-08T00:15:48.483Z to 2026-10-08T00:16:08.491Z (20.0 s), 11,660 searches, 0 failed, p50 12.995 ms, mean 13.719 ms, p99 27.665 ms, 582.8 QPS (1000/QPS 1.716 ms).
- Passes in the order run: exact, default@1, default@8; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 52,638 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (duckdb_indexes() lists gvb_gvbbench_eshoponweb_hnsw, pragma_hnsw_index_info() counts 524 vectors of 524 rows, EXPLAIN of the default search shows HNSW_INDEX_SCAN on it = True).
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
- Connection: in this process (embedded), no network: no network address. No open connection seen (no open TCP connection was found).
- CPU pinning: none: an embedded engine runs inside the benchmark client process, so it shares the client CPUs, CPUs 0-1,4-5.
- Clock per pass (median MHz of engine CPUs / client CPUs, each the median of its CPUs' medians and every CPU when the CPUs were not split, their average MHz, governor, CPUs busy outside the benchmark, client and engine CPU per search; the clock was pinned at 3500 MHz (turbo off), and a pass whose median on either group is more than 100 bp (1%) off it, or that has no reading on a group, is flagged): exact 3495/3492 MHz, average 3495.9/3492.1, performance, outside load 0.13 before the warm-up, 0.15 during, client CPU 6.135 ms per search; default@1 3493/3492 MHz, average 3495.5/3492.2, performance, outside load 0.15 before the warm-up, 0.14 during, client CPU 5.009 ms per search; default@8 3493/3492 MHz, average 3494.4/3492.0, performance, outside load 0.13 before the warm-up, 0.15 during, client CPU 6.833 ms per search.
Typesense typesense
- Engine: Typesense 30.2 (vector search) (compose)
- Index: HNSW float32 (hnswlib), m=16, ef_construction=128, cosine; search k=top, ef=100; exact mode = filter ordinal:>=0 with flat_search_cutoff; index held in memory
- Search settings: ef=100, ef_construction=128, k=top, m=16
- Engine settings: HostConfig.Memory = 6442450944 (read: docker inspect HostConfig.Memory (0 = no limit)); HostConfig.CpusetCpus = 2-3,6-7 (read: docker inspect HostConfig.CpusetCpus (empty = every CPU)); HostConfig.NanoCpus = 0 (read: docker inspect HostConfig.NanoCpus (0 = no limit))
- Image: typesense/typesense:30.2, id sha256:610f2d34b1f93d00762869da2c67736775e5798d19a2c8b91b014b8a0cc1e110 (the pinned id), tagged on this machine at 2026-10-03T14:12:37.929033132Z
- Data folder ~/gvb-data/engines/typesense: 688.52 MiB at the start [Correction 8], 296 MiB at the end
- Load: 524 rows in batches of 1000, 0.5 s of upserts (1,016 rows/s); count matched 0.0 s after the last upsert; index step 0.0 s (nothing to build, Typesense inserts into the HNSW graph while writing; ready after 0.0 s: HNSW graph returns 524 of 524 vectors, no writes queued: num_documents 524, pending_write_batches 0, health ok, unfiltered vector query with k above the count returned 524; the graph is built in memory during each write, so there is no later build step)
- Search: 4742 latency samples, target held 524 rows, 0 errors
- Client CPU per search: 1 searcher 0.52 ms, 8 searchers 0.58 ms
- Exact mode: 10,514 searches, p50 5.63 ms, p95 6.06 ms, recall 1.000
- RAM: 118.9 MiB (docker stats); disk: 1.67 MiB added by this load (296 MiB (whole engine data folder), 294.34 MiB before)
- Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started typesense.compose.yaml with every container created on CPUs 2-3,6-7: gvb-typesense cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards.
- Engine files (SHA-256, first 12 hex digits): [typesense.compose.yaml 96a165f62566]. This run created the engine from these files.
- Index after the load: ready, 524 of 524 indexed (HNSW graph returns 524 of 524 vectors, no writes queued: num_documents 524, pending_write_batches 0, health ok, unfiltered vector query with k above the count returned 524; the graph is built in memory during each write, so there is no later build step). Durability: Every acknowledged write is appended to Typesense's raft log and fsynced before the answer: braft raft_sync=true with raft_sync_policy=0 (sync immediately), read from the running container's brpc /flags page on 2026-10-04. A restart replays the log from the last snapshot (typesense.compose.yaml sets TYPESENSE_SNAPSHOT_INTERVAL_SECONDS=300) and rebuilds the in-memory HNSW graph, so by those settings a crash loses no acknowledged write, but the restart is slow after a big load. This is read from the settings, not shown by pulling power. One node and no replicas, so a lost disk loses the data..
- Outside load when searching began: 0.18 CPUs on average over the last 7 s and 0.17 over the last few seconds (limit 0.3), a quiet box. The load average 6.98 4.60 4.00 (1/5/15 min; it also counts this benchmark's own client and engines, so it is recorded, not judged).
- Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: exact 5,263 searches in 30.0 s, default@1 7,119 searches in 30.0 s, default@8 17,226 searches in 30.0 s.
- Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it. exact: warm-up 15.0 s and 2,603 searches (0 failed) at 1 searcher; p50 of the last windows 5.629, 5.725, 5.631 ms (windows of at least 2 s and 100 searches); trial of 524 searches in 3.0 s at 1 searcher: p50 5.631 ms against the settled p50 5.631 ms, 0% apart (limit 10%); the timed pass p50 5.630 ms, 0% from the settled p50 5.631 ms (limit 10% or 0.1 ms; 0.002 ms); default@1: warm-up 15.0 s and 3,543 searches (0 failed) at 1 searcher; p50 of the last windows 4.157, 4.165, 4.158 ms (windows of at least 2 s and 100 searches); trial of 714 searches in 3.0 s at 1 searcher: p50 4.158 ms against the settled p50 4.158 ms, 0% apart (limit 10%); the timed pass p50 4.160 ms, 0% from the settled p50 4.158 ms (limit 10% or 0.1 ms; 0.002 ms); default@8: warm-up 15.0 s and 8,582 searches (0 failed) at 8 searchers; QPS of the last windows 570, 574, 570 (windows of at least 2 s and 100 searches); trial of 1,718 searches in 3.0 s at 8 searchers: 571 QPS against the settled 570 QPS, 0% apart (limit 10%); the timed pass 572 QPS, 0% from the settled 570 QPS (limit 10%).
- Pass exact after rehearsal: 2026-10-08T00:18:10.293Z to 2026-10-08T00:19:10.293Z (60.0 s), 10,514 searches, 0 failed, p50 5.630 ms, mean 5.704 ms, p99 8.304 ms, 175.2 QPS (1000/QPS 5.707 ms).
- Pass default@1 after exact: 2026-10-08T00:19:28.308Z to 2026-10-08T00:19:48.309Z (20.0 s), 4,742 searches, 0 failed, p50 4.160 ms, mean 4.216 ms, p99 5.196 ms, 237.1 QPS (1000/QPS 4.218 ms).
- Pass default@8 after default@1: 2026-10-08T00:20:06.335Z to 2026-10-08T00:20:26.345Z (20.0 s), 11,447 searches, 0 failed, p50 13.509 ms, mean 13.979 ms, p99 26.022 ms, 572.1 QPS (1000/QPS 1.748 ms).
- Passes in the order run: exact, default@1, default@8; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 47,292 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (HNSW graph returns 524 of 524 vectors, no writes queued: num_documents 524, pending_write_batches 0, health ok, unfiltered vector query with k above the count returned 524; the graph is built in memory during each write, so there is no later build step).
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
- Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-typesense (7a355b59d52c) on network engines_default at container-ip:8108, not through docker-proxy localhost:8108. Open after the passes: container-ip:8108 x8 (this target).
- CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-typesense: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-typesense: 303 thread(s) on CPUs 2-3,6-7 (every thread of the container).
- Clock per pass (median MHz of engine CPUs / client CPUs, each the median of its CPUs' medians and every CPU when the CPUs were not split, their average MHz, governor, CPUs busy outside the benchmark, client and engine CPU per search; the clock was pinned at 3500 MHz (turbo off), and a pass whose median on either group is more than 100 bp (1%) off it, or that has no reading on a group, is flagged): exact 3492/3492 MHz, average 3492.8/3493.7, performance, outside load 0.07 before the warm-up, 0.15 during, client CPU 0.529 ms per search, engine CPU 5.345 ms per search; default@1 3492/3492 MHz, average 3492.8/3493.6, performance, outside load 0.15 before the warm-up, 0.15 during, client CPU 0.523 ms per search, engine CPU 3.897 ms per search; default@8 3492/3492 MHz, average 3492.0/3492.6, performance, outside load 0.15 before the warm-up, -0.04 during, client CPU 0.582 ms per search, engine CPU 6.702 ms per search.
SQL Server 2025 + DiskANN sql-diskann
- Engine: Microsoft SQL Server 2025 (RTM-CU9) (KB5122048) - 17.0.5005.3 (X64), Enterprise Developer Edition (64-bit), in container gvb-mssql (image mcr.microsoft.com/mssql/server:2025-CU9-ubuntu-24.04@sha256:2b5b581621126574f3d1f75e78d3eebe8d05aedb59ad0cfdf9aa42cb0634d726; cpuset 2-3,6-7 from its creation; SQL Server counts 4 CPU(s) and runs 4 visible scheduler(s), affinity AUTO) (compose)
- Index: DiskANN (preview) via VECTOR_SEARCH, cosine, build {"StartId":"306", "L":"48", "M":"8", "R":"48"}; the exact mode scans the same table
- Search settings: L=48, M=8, R=48, StartId=306
- Engine settings: HostConfig.Memory = 8589934592 (read: docker inspect HostConfig.Memory (0 = no limit)); HostConfig.CpusetCpus = 2-3,6-7 (read: docker inspect HostConfig.CpusetCpus (empty = every CPU)); HostConfig.NanoCpus = 0 (read: docker inspect HostConfig.NanoCpus (0 = no limit))
- Image: mcr.microsoft.com/mssql/server:2025-CU9-ubuntu-24.04@sha256:2b5b581621126574f3d1f75e78d3eebe8d05aedb59ad0cfdf9aa42cb0634d726, id sha256:2b5b581621126574f3d1f75e78d3eebe8d05aedb59ad0cfdf9aa42cb0634d726 (the pinned id), tagged on this machine at 2026-10-04T22:02:11.293843649Z
- Data folder ~/gvb-data/engines/mssql: 105.06 MiB at the start [Correction 7], 505.9 MiB at the end
- Load: 524 rows in batches of 1000, 0.8 s of upserts (619 rows/s); count matched 0.0 s after the last upsert; index step 4.0 s (copied 524 rows to dbo.gvb_gvbbench_eshoponweb_ann (INT key) and built DiskANN, parameters {"StartId":"306", "L":"48", "M":"8", "R":"48"}, graph covers 524 rows)
- Search: 5556 latency samples, target held 524 rows, 0 errors
- Client CPU per search: 1 searcher 1.00 ms, 8 searchers 1.08 ms
- Exact mode: 16,154 searches, p50 3.63 ms, p95 4.19 ms, recall 1.000
- RAM: 533 MiB (docker stats); disk: 4.52 MiB (table and its indexes, reserved pages)
- Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started mssql.compose.yaml with every container created on CPUs 2-3,6-7: gvb-mssql cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards.
- Engine files (SHA-256, first 12 hex digits): [mssql.compose.yaml b7359663cc15]. This run created the engine from these files.
- Index after the load: ready, 524 of 524 indexed (DiskANN index built and used. sys.vector_indexes: vix_gvb_gvbbench_eshoponweb on dbo.gvb_gvbbench_eshoponweb_ann in GvbBenchDiskAnn, DiskANN, COSINE, enabled, build {"StartId":"306", "L":"48", "M":"8", "R":"48"}; graph table rows 524 of 524 table rows. Real VECTOR_SEARCH plan: Vector Index Seek on index vix_gvb_gvbbench_eshoponweb of GvbBenchDiskAnn.gvb_gvbbench_eshoponweb_ann (IndexKind DiskANN), 10 rows returned, 10 hits. The exact mode searches the SAME table (dbo.gvb_gvbbench_eshoponweb_ann, database GvbBenchDiskAnn) and its real plan is: Clustered Index Scan of GvbBenchDiskAnn.gvb_gvbbench_eshoponweb_ann through PK_gvb_gvbbench_eshoponweb_ann (IndexKind Clustered), 524 rows read, no vector index operator, 10 hits). Durability: A commit returns after its transaction-log records are written to disk (SQL Server write-ahead logging; delayed durability is DISABLED); a database created here copies the model database: recovery model FULL, page_verify CHECKSUM; no global trace flags are enabled; mssql.conf of container gvb-mssql (/var/opt/mssql/mssql.conf, on the host at ~/gvb-data/engines/mssql/mssql.conf) does not exist, so it sets: nothing; it has no [control] or [traceflag] entry, so SQL Server's own Linux defaults for flushing writes apply (Microsoft's Linux performance guide names trace flag 3982 as that default; read from the guide, not tested here). Not tested by cutting power; whether the disk's own write cache reaches the media was not checked. Container settings from its environment (names only): MSSQL_AGENT_ENABLED, MSSQL_MEMORY_LIMIT_MB, MSSQL_PID, MSSQL_RPC_PORT, MSSQL_SA_PASSWORD..
- Outside load when searching began: 0.16 CPUs on average over the last 6 s and 0.15 over the last few seconds (limit 0.3), a quiet box. The load average 4.22 3.85 3.79 (1/5/15 min; it also counts this benchmark's own client and engines, so it is recorded, not judged).
- Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@1 8,138 searches in 30.0 s, default@8 20,493 searches in 30.0 s, exact 7,967 searches in 30.0 s.
- Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it. default@1: warm-up 15.0 s and 4,161 searches (0 failed) at 1 searcher; p50 of the last windows 3.566, 3.545, 3.571 ms (windows of at least 2 s and 100 searches); trial of 833 searches in 3.0 s at 1 searcher: p50 3.567 ms against the settled p50 3.566 ms, 0% apart (limit 10%); the timed pass p50 3.568 ms, 0% from the settled p50 3.566 ms (limit 10% or 0.1 ms; 0.002 ms); default@8: warm-up 15.0 s and 10,404 searches (0 failed) at 8 searchers; QPS of the last windows 701, 705, 680 (windows of at least 2 s and 100 searches); trial of 2,065 searches in 3.1 s at 8 searchers: 675 QPS against the settled 701 QPS, 4% apart (limit 10%); the timed pass 694 QPS, 1% from the settled 701 QPS (limit 10%); exact: warm-up 15.0 s and 4,009 searches (0 failed) at 1 searcher; p50 of the last windows 3.642, 3.632, 3.631 ms (windows of at least 2 s and 100 searches); trial of 810 searches in 3.0 s at 1 searcher: p50 3.621 ms against the settled p50 3.632 ms, 0% apart (limit 10%); the timed pass p50 3.631 ms, 0% from the settled p50 3.632 ms (limit 10% or 0.1 ms; 0.001 ms).
- Pass default@1 after rehearsal: 2026-10-08T00:22:31.176Z to 2026-10-08T00:22:51.180Z (20.0 s), 5,556 searches, 0 failed, p50 3.568 ms, mean 3.598 ms, p99 4.525 ms, 277.8 QPS (1000/QPS 3.600 ms).
- Pass default@8 after default@1: 2026-10-08T00:23:09.255Z to 2026-10-08T00:23:29.263Z (20.0 s), 13,882 searches, 0 failed, p50 11.318 ms, mean 11.525 ms, p99 19.574 ms, 693.8 QPS (1000/QPS 1.441 ms).
- Pass exact after default@8: 2026-10-08T00:23:47.271Z to 2026-10-08T00:24:47.271Z (60.0 s), 16,154 searches, 0 failed, p50 3.631 ms, mean 3.711 ms, p99 5.225 ms, 269.2 QPS (1000/QPS 3.714 ms).
- Passes in the order run: default@1, default@8, exact; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 58,880 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (DiskANN index built and used. sys.vector_indexes: vix_gvb_gvbbench_eshoponweb on dbo.gvb_gvbbench_eshoponweb_ann in GvbBenchDiskAnn, DiskANN, COSINE, enabled, build {"StartId":"306", "L":"48", "M":"8", "R":"48"}; graph table rows 524 of 524 table rows. Real VECTOR_SEARCH plan: Vector Index Seek on index vix_gvb_gvbbench_eshoponweb of GvbBenchDiskAnn.gvb_gvbbench_eshoponweb_ann (IndexKind DiskANN), 10 rows returned, 10 hits. The exact mode searches the SAME table (dbo.gvb_gvbbench_eshoponweb_ann, database GvbBenchDiskAnn) and its real plan is: Clustered Index Scan of GvbBenchDiskAnn.gvb_gvbbench_eshoponweb_ann through PK_gvb_gvbbench_eshoponweb_ann (IndexKind Clustered), 524 rows read, no vector index operator, 10 hits).
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
- Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-mssql (0a67a06ed461) on network gvb-mssql_default at container-ip:1433, not through docker-proxy localhost:14330. Open after the passes: container-ip:1433 x9 (this target).
- CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-mssql: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-mssql: 85 thread(s) on CPUs 2-3,6-7 (every thread of the container).
- Clock per pass (median MHz of engine CPUs / client CPUs, each the median of its CPUs' medians and every CPU when the CPUs were not split, their average MHz, governor, CPUs busy outside the benchmark, client and engine CPU per search; the clock was pinned at 3500 MHz (turbo off), and a pass whose median on either group is more than 100 bp (1%) off it, or that has no reading on a group, is flagged): default@1 3492/3492 MHz, average 3493.2/3492.4, performance, outside load 0.16 before the warm-up, 0.13 during, client CPU 1.001 ms per search, engine CPU 3.481 ms per search; default@8 3492/3492 MHz, average 3492.0/3492.3, performance, outside load 0.16 before the warm-up, 0.11 during, client CPU 1.077 ms per search, engine CPU 5.631 ms per search; exact 3492/3492 MHz, average 3493.0/3492.4, performance, outside load 0.13 before the warm-up, 0.15 during, client CPU 1.127 ms per search, engine CPU 3.584 ms per search.
Qdrant (HNSW) qdrant-hnsw
- Engine: Qdrant 1.17.0 in container gvb-qdrant (image qdrant/qdrant:v1.17.0@sha256:f1c7272cdac52b38c1a0e89313922d940ba50afd90d593a1605dbbc214e66ffb; cpuset 2-3,6-7 from its creation; 3 search thread(s)) (compose)
- Index: HNSW m=16 ef_construct=100, hnsw_ef=server default, cosine; indexing_threshold_kb 1 and full_scan_threshold_kb 10 (server defaults are 10,000 each) so a small collection builds and walks its graph
- Search settings: ef_construct=100, hnsw_ef=server, m=16
- Engine settings: HostConfig.Memory = 8589934592 (read: docker inspect HostConfig.Memory (0 = no limit)); HostConfig.CpusetCpus = 2-3,6-7 (read: docker inspect HostConfig.CpusetCpus (empty = every CPU)); HostConfig.NanoCpus = 0 (read: docker inspect HostConfig.NanoCpus (0 = no limit)); search_threads = 3 (read: threads named search-* of the qdrant process in container gvb-qdrant (/proc))
- Image: qdrant/qdrant:v1.17.0@sha256:f1c7272cdac52b38c1a0e89313922d940ba50afd90d593a1605dbbc214e66ffb, id sha256:f1c7272cdac52b38c1a0e89313922d940ba50afd90d593a1605dbbc214e66ffb (the pinned id), tagged on this machine at 2026-10-04T22:03:16.973657523Z
- Data folder ~/gvb-data/engines/qdrant: 357 B at the start, 164.1 MiB at the end
- Load: 524 rows in batches of 1000, 0.1 s of upserts (4,677 rows/s); count matched 0.0 s after the last upsert; index step 2.0 s (indexing_threshold_kb 1, full_scan_threshold_kb 10; status Green, 524 of 524 vectors in HNSW segments (2 segments), waited 2.0 s)
- Search: 21444 latency samples, target held 524 rows, 0 errors
- Client CPU per search: 1 searcher 0.73 ms, 8 searchers 0.52 ms
- Exact mode: 65,813 searches, p50 0.90 ms, p95 1.01 ms, recall 1.000
- RAM: 37.02 MiB (docker stats); disk: 164.1 MiB (collection folder)
- Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started qdrant.compose.yaml with every container created on CPUs 2-3,6-7: gvb-qdrant cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards.
- Engine files (SHA-256, first 12 hex digits): [qdrant.compose.yaml eb57e55e93c1]. This run created the engine from these files.
- Index after the load: ready, 524 of 524 indexed (status Green, optimizer ok, indexed_vectors_count 524 of 524 points, 2 segments, indexing_threshold_kb 1, full_scan_threshold_kb 10; every one of 1 non-empty segments has a complete HNSW graph above its full-scan threshold; no search has run yet, so the walk is proven by the segment settings only; searches counted since the collection was created: unfiltered_hnsw 0, unfiltered_plain 0, unfiltered_exact 0). Durability: What was measured (strace -f on the Qdrant server process, 30 single-point REST upserts per setting, every sync call timed against the request that caused it): with wait=true each upsert had exactly one msync(MS_SYNC) of the write-ahead-log segment inside the request, before the reply (30 of 30); with wait=false the 30 upserts were acknowledged within 0.26 s with no sync call, and the first WAL msync (one call covering all 30 records) came 2.7 s after the last reply, followed by the segment-file flushes. So the log is flushed to disk under both settings, but with wait=false the flush comes after the acknowledgement: an operating-system crash or power cut in that gap loses acknowledged writes, while a crash of the Qdrant process alone should not, because the bytes are already in the kernel's page cache (inferred, not tested). Every upsert here is sent with wait=true in batches of 256 over gRPC; the strace used one point per request, so one flush per batch is inferred, not measured. Segment files are flushed every 5 s; log segments 32 MB, 0 created ahead. /qdrant/config/config.yaml in container gvb-qdrant (the image's own file; its storage keys are listed) sets: storage.collection.quantization = null, storage.collection.replication_factor = 1, storage.collection.vectors.on_disk = null, storage.collection.write_consistency_factor = 1, storage.hnsw_index.ef_construct = 100, storage.hnsw_index.full_scan_threshold_kb = 10000, storage.hnsw_index.m = 16, storage.hnsw_index.max_indexing_threads = 0, storage.hnsw_index.on_disk = false, storage.hnsw_index.payload_m = null, storage.max_collections = null, storage.node_type = Normal, storage.on_disk_payload = true, storage.optimizers.default_segment_number = 0, storage.optimizers.deleted_threshold = 0.2, storage.optimizers.flush_interval_sec = 5, storage.optimizers.indexing_threshold_kb = 10000, storage.optimizers.max_optimization_threads = null, storage.optimizers.max_segment_size_kb = null, storage.optimizers.vacuum_min_vector_number = 1000, storage.performance.max_search_threads = 0, storage.performance.optimizer_cpu_budget = 0, storage.performance.update_rate_limit = null, storage.shard_transfer_method = null, storage.snapshots_config.snapshots_storage = local, storage.snapshots_path = ./snapshots, storage.storage_path = ./storage, storage.temp_path = null, storage.update_concurrency = null, storage.wal.wal_capacity_mb = 32, storage.wal.wal_segments_ahead = 0; no QDRANT__ environment overrides in the server process. Not tested by cutting power; whether the disk's own write cache reaches the media was not checked..
- Outside load when searching began: 0.17 CPUs on average over the last 3 s and 0.17 over the last few seconds (limit 0.3), a quiet box. The load average 1.68 2.89 3.42 (1/5/15 min; it also counts this benchmark's own client and engines, so it is recorded, not judged).
- Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: exact 32,794 searches in 30.0 s, default@8 103,725 searches in 30.0 s, default@1 32,075 searches in 30.0 s.
- Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it. exact: warm-up 15.0 s and 16,437 searches (0 failed) at 1 searcher; p50 of the last windows 0.897, 0.902, 0.904 ms (windows of at least 2 s and 100 searches); trial of 3,286 searches in 3.0 s at 1 searcher: p50 0.901 ms against the settled p50 0.902 ms, 0% apart (limit 10%); the timed pass p50 0.899 ms, 0% from the settled p50 0.902 ms (limit 10% or 0.1 ms; 0.003 ms); default@8: warm-up 15.0 s and 51,929 searches (0 failed) at 8 searchers; QPS of the last windows 3,464, 3,450, 3,433 (windows of at least 2 s and 100 searches); trial of 10,396 searches in 3.0 s at 8 searchers: 3,464 QPS against the settled 3,450 QPS, 0% apart (limit 10%); the timed pass 3,468 QPS, 1% from the settled 3,450 QPS (limit 10%); default@1: warm-up 15.0 s and 16,064 searches (0 failed) at 1 searcher; p50 of the last windows 0.927, 0.921, 0.926 ms (windows of at least 2 s and 100 searches); trial of 3,201 searches in 3.0 s at 1 searcher: p50 0.925 ms against the settled p50 0.926 ms, 0% apart (limit 10%); the timed pass p50 0.922 ms, 0% from the settled p50 0.926 ms (limit 10% or 0.1 ms; 0.004 ms).
- Pass exact after rehearsal: 2026-10-08T00:26:45.996Z to 2026-10-08T00:27:45.996Z (60.0 s), 65,813 searches, 0 failed, p50 0.899 ms, mean 0.910 ms, p99 1.425 ms, 1096.9 QPS (1000/QPS 0.912 ms).
- Pass default@8 after exact: 2026-10-08T00:28:04.073Z to 2026-10-08T00:28:24.074Z (20.0 s), 69,361 searches, 0 failed, p50 2.220 ms, mean 2.304 ms, p99 3.973 ms, 3467.8 QPS (1000/QPS 0.288 ms).
- Pass default@1 after default@8: 2026-10-08T00:28:42.102Z to 2026-10-08T00:29:02.102Z (20.0 s), 21,444 searches, 0 failed, p50 0.922 ms, mean 0.931 ms, p99 1.503 ms, 1072.2 QPS (1000/QPS 0.933 ms).
- Passes in the order run: exact, default@8, default@1; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 269,907 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (status Green, optimizer ok, indexed_vectors_count 524 of 524 points, 2 segments, indexing_threshold_kb 1, full_scan_threshold_kb 10; every one of 1 non-empty segments has a complete HNSW graph above its full-scan threshold; 308,195 segment searches walked the graph and none scanned; searches counted since the collection was created: unfiltered_hnsw 308,195, unfiltered_plain 0, unfiltered_exact 118,330).
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
- Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-qdrant (a48d51b0dc64) on network gvb-qdrant_default at container-ip:6334, not through docker-proxy localhost:16334; container gvb-qdrant (a48d51b0dc64) on network gvb-qdrant_default at container-ip:6333, not through docker-proxy localhost:16333. Open after the passes: container-ip:6334 x1 (this target).
- CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-qdrant: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-qdrant: 24 thread(s) on CPUs 2-3,6-7 (every thread of the container).
- Clock per pass (median MHz of engine CPUs / client CPUs, each the median of its CPUs' medians and every CPU when the CPUs were not split, their average MHz, governor, CPUs busy outside the benchmark, client and engine CPU per search; the clock was pinned at 3500 MHz (turbo off), and a pass whose median on either group is more than 100 bp (1%) off it, or that has no reading on a group, is flagged): exact 3492/3492 MHz, average 3492.0/3492.0, performance, outside load 0.12 before the warm-up, 0.11 during, client CPU 0.733 ms per search, engine CPU 0.714 ms per search; default@8 3492/3492 MHz, average 3492.0/3492.0, performance, outside load 0.11 before the warm-up, 0.13 during, client CPU 0.523 ms per search, engine CPU 0.933 ms per search; default@1 3492/3492 MHz, average 3492.0/3492.0, performance, outside load 0.13 before the warm-up, 0.09 during, client CPU 0.734 ms per search, engine CPU 0.731 ms per search.
Milvus milvus
- Engine: Milvus 2.6.25 (standalone, embedded etcd, local storage, REST API v2) (compose)
- Index: HNSW M=16 efConstruction=128, ef=100, metric COSINE, Strong consistency searches; approximate only (no exact mode)
- Search settings: M=16, ef=100, efConstruction=128
- Engine settings: HostConfig.Memory = 8589934592 (read: docker inspect HostConfig.Memory (0 = no limit)); HostConfig.CpusetCpus = 2-3,6-7 (read: docker inspect HostConfig.CpusetCpus (empty = every CPU)); HostConfig.NanoCpus = 0 (read: docker inspect HostConfig.NanoCpus (0 = no limit))
- Image: milvusdb/milvus:v2.6.25, id sha256:29f7668e64df1c6d5cdadbfbee99f38c72e4001054f149a298725acae57230ad (the pinned id), tagged on this machine at 2026-10-03T14:05:28.412151323Z
- Data folder ~/gvb-data/engines/milvus: 235.67 MiB at the start [Correction 5], 246.62 MiB at the end
- Load: 524 rows in batches of 1000, 0.5 s of upserts (1,158 rows/s); count matched 1.7 s after the last upsert; index step 138.5 s (index Finished, type HNSW (COSINE, {"M":16,"efConstruction":128}), indexedRows 524 of 524 sealed rows, pendingRows 0; stored rows 524; LoadStateLoaded; query node: 1 sealed segment(s) with the index loaded covering 524 rows against 524 stored rows, 0 sealed without it, 0 rows in growing segments; datacoord: 1 live segment(s), 1 flushed and indexed covering 524 rows, 0 delete-log (L0) segment(s) waiting for a compaction, 3 segment(s) compacted away so far; flush, index build and load took 14.7 s, then compaction: 1 job(s) accepted by Milvus, layout ready and unchanged for 75 s, 123.9 s for the whole step; search ledger started, later state reads count the searches against the query node's own counters)
- Search: 8462 latency samples, target held 524 rows, 0 errors
- Client CPU per search: 1 searcher 0.54 ms, 8 searchers 0.63 ms
- RAM: 186.5 MiB (docker stats); disk: 10.95 MiB added by this load (246.62 MiB (whole engine data folder), 235.67 MiB before)
- Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started milvus.compose.yaml with every container created on CPUs 2-3,6-7: gvb-milvus cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards.
- Engine files (SHA-256, first 12 hex digits): [milvus.compose.yaml eaefecdc3671]. This run created the engine from these files.
- Index after the load: ready, 524 of 524 indexed (index Finished, type HNSW (COSINE, {"M":16,"efConstruction":128}), indexedRows 524 of 524 sealed rows, pendingRows 0; stored rows 524; LoadStateLoaded; query node: 1 sealed segment(s) with the index loaded covering 524 rows against 524 stored rows, 0 sealed without it, 0 rows in growing segments; datacoord: 1 live segment(s), 1 flushed and indexed covering 524 rows, 0 delete-log (L0) segment(s) waiting for a compaction, 3 segment(s) compacted away so far; search ledger since the index finished: the sink sent 0 searches (searchParams.params.ef=100, Strong consistency), the query node counted 0 search requests for this collection, segment searches Sealed +0 and Growing +0, the same sealed segments with the same loaded index builds at both readings, segments compacted away +0). Durability: Writes are not fsynced. Standalone uses the default message queue rocksmq (RocksDB, /var/lib/milvus/rdb_data, mq.type default) and flushed segments go to local-disk object storage (COMMON_STORAGETYPE=local in milvus.compose.yaml). Measured 2026-10-04 with strace over 30 acknowledged single-row upserts and one flush: fdatasync ran only on the embedded etcd files (member/wal, member/snap/db), never on rdb_data or the segment files. A host power loss or kernel crash can lose acknowledged rows that are still in the OS page cache; a Milvus process crash alone should not (inferred, not tested). Metadata in embedded etcd is fdatasynced on every commit..
- Outside load when searching began: 0.13 CPUs on average over the last 60 s and 0.13 over the last few seconds (limit 0.3), a quiet box. The load average 0.46 1.78 2.79 (1/5/15 min; it also counts this benchmark's own client and engines, so it is recorded, not judged).
- Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@1 12,668 searches in 30.0 s, default@8 38,100 searches in 30.0 s.
- Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it. default@1: warm-up 15.0 s and 6,336 searches (0 failed) at 1 searcher; p50 of the last windows 2.243, 2.252, 2.241 ms (windows of at least 2 s and 100 searches); trial of 1,276 searches in 3.0 s at 1 searcher: p50 2.248 ms against the settled p50 2.243 ms, 0% apart (limit 10%); the timed pass p50 2.243 ms, 0% from the settled p50 2.243 ms (limit 10% or 0.1 ms; 0.001 ms); default@8: warm-up 15.0 s and 18,942 searches (0 failed) at 8 searchers; QPS of the last windows 1,269, 1,278, 1,274 (windows of at least 2 s and 100 searches); trial of 3,803 searches in 3.0 s at 8 searchers: 1,266 QPS against the settled 1,274 QPS, 1% apart (limit 10%); the timed pass 1,266 QPS, 1% from the settled 1,274 QPS (limit 10%).
- Pass default@1 after rehearsal: 2026-10-08T00:32:50.899Z to 2026-10-08T00:33:10.899Z (20.0 s), 8,462 searches, 0 failed, p50 2.243 ms, mean 2.361 ms, p99 3.807 ms, 423.1 QPS (1000/QPS 2.364 ms).
- Pass default@8 after default@1: 2026-10-08T00:33:28.915Z to 2026-10-08T00:33:48.922Z (20.0 s), 25,331 searches, 0 failed, p50 5.767 ms, mean 6.315 ms, p99 14.603 ms, 1266.1 QPS (1000/QPS 0.790 ms).
- Passes in the order run: default@1, default@8; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 81,125 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (index Finished, type HNSW (COSINE, {"M":16,"efConstruction":128}), indexedRows 524 of 524 sealed rows, pendingRows 0; stored rows 524; LoadStateLoaded; query node: 1 sealed segment(s) with the index loaded covering 524 rows against 524 stored rows, 0 sealed without it, 0 rows in growing segments; datacoord: 1 live segment(s), 1 flushed and indexed covering 524 rows, 0 delete-log (L0) segment(s) waiting for a compaction, 3 segment(s) compacted away so far; search ledger since the index finished: the sink sent 114918 searches (searchParams.params.ef=100, Strong consistency), the query node counted 114918 search requests for this collection, segment searches Sealed +112922 and Growing +0, the same sealed segments with the same loaded index builds at both readings, segments compacted away +0).
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
- Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-milvus (15cbe83fa711) on network gvb-milvus_default at container-ip:19530, not through docker-proxy localhost:19530; container gvb-milvus (15cbe83fa711) on network gvb-milvus_default at container-ip:9091, not through docker-proxy localhost:9091. Open after the passes: container-ip:19530 x8 (this target).
- CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-milvus: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-milvus: 24 thread(s) on CPUs 2-3,6-7 (every thread of the container).
- Clock per pass (median MHz of engine CPUs / client CPUs, each the median of its CPUs' medians and every CPU when the CPUs were not split, their average MHz, governor, CPUs busy outside the benchmark, client and engine CPU per search; the clock was pinned at 3500 MHz (turbo off), and a pass whose median on either group is more than 100 bp (1%) off it, or that has no reading on a group, is flagged): default@1 3492/3492 MHz, average 3492.0/3492.7, performance, outside load 0.15 before the warm-up, 0.13 during, client CPU 0.541 ms per search, engine CPU 2.743 ms per search; default@8 3492/3492 MHz, average 3492.0/3492.8, performance, outside load 0.14 before the warm-up, 0.18 during, client CPU 0.630 ms per search, engine CPU 2.958 ms per search.
ClickHouse clickhouse
- Engine: ClickHouse 26.3.39.7 (MergeTree + vector_similarity index; image-default system logging, each log table with a TTL of 1 hour) (compose)
- Index: vector_similarity HNSW cosineDistance, quantization bf16, M=16 ef_construction=128, hnsw_candidate_list_size_for_search=256, rescoring off; exact mode = full scan with skip indexes off
- Search settings: M=16, ef_construction=128, hnsw_candidate_list_size_for_search=256
- Engine settings: HostConfig.Memory = 8589934592 (read: docker inspect HostConfig.Memory (0 = no limit)); HostConfig.CpusetCpus = 2-3,6-7 (read: docker inspect HostConfig.CpusetCpus (empty = every CPU)); HostConfig.NanoCpus = 0 (read: docker inspect HostConfig.NanoCpus (0 = no limit)); max_threads = \'auto(4)\' (read: ClickHouse system.settings, name max_threads, over HTTP)
- Image: clickhouse/clickhouse-server:26.3.39.7, id sha256:3a91276f066905da0edbbd622d3fc2a2df632c87ea7fe32ef4e74fdf8c9567b0 (the pinned id), tagged on this machine at 2026-10-03T14:12:01.743437616Z
- Data folder ~/gvb-data/engines/clickhouse: 6.92 GiB at the start, 2.09 GiB after the reset, 6.76 GiB at the end; reset at the start: truncated 9 MergeTree log tables in database system (asynchronous_insert_log, asynchronous_metric_log, background_schedule_pool_log, metric_log, part_log, processors_profile_log, query_log, text_log, trace_log); ClickHouse's own active bytes of those tables 185.06 KiB before and 0 B after; data folder 6.92 GiB before and 2.09 GiB after (the folder was steady) [Correction 3]; not truncated, name does not end in _log: none
- Load: 524 rows in batches of 1000, 0.2 s of upserts (2,737 rows/s); count matched 0.0 s after the last upsert; index step 6.2 s (1 data part(s), index files on 1 of them, 524 of 524 live rows in indexed parts; 0 merge(s) running, 0 unfinished mutation(s); plan of the default search uses the Skip index vec_idx on 1 of 1 parts; wait took 6.2 s)
- Search: 3937 latency samples, target held 524 rows, 0 errors
- Client CPU per search: 1 searcher 0.86 ms, 8 searchers 1.05 ms
- Exact mode: 8,761 searches, p50 6.60 ms, p95 8.63 ms, recall 1.000
- RAM: 1011 MiB (docker stats); disk: 4.67 GiB added by this load [Correction 2] (6.76 GiB (whole engine data folder), 2.09 GiB before)
- Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started clickhouse.compose.yaml with every container created on CPUs 2-3,6-7: gvb-clickhouse cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards.
- Engine files (SHA-256, first 12 hex digits): [clickhouse-config/gvb-system-logs.xml 02cc53f0f171; clickhouse.compose.yaml 560e0ae487c6]. This run created the engine from these files.
- Index after the load: ready, 524 of 524 indexed (1 data part(s), index files on 1 of them, 524 of 524 live rows in indexed parts; 0 merge(s) running, 0 unfinished mutation(s); plan of the default search uses the Skip index vec_idx on 1 of 1 parts). Durability: Acknowledged inserts are not fsynced. The sink creates plain MergeTree tables, so the MergeTree settings are the defaults: fsync_after_insert 0 and fsync_part_directory 0 (system.merge_tree_settings), and async_insert 1 with wait_for_async_insert 1 is the server default for the gvb login (system.settings), so an insert is acknowledged once its part is written, not once it is on disk. Measured 2026-10-04 with strace over 30 acknowledged single-row inserts (30 Ok rows in system.asynchronous_insert_log): zero fsync, fdatasync or sync_file_range calls, while the control, a table created with fsync_after_insert=1, made 61 fdatasync calls for 5 inserts. A host power loss or kernel crash can lose acknowledged rows that are still in the OS page cache; a ClickHouse process crash alone should not (inferred, not tested)..
- Outside load when searching began: 0.17 CPUs on average over the last 28 s and 0.17 over the last few seconds (limit 0.3), a quiet box. The load average 2.79 2.69 2.99 (1/5/15 min; it also counts this benchmark's own client and engines, so it is recorded, not judged).
- Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: exact 4,412 searches in 30.0 s, default@8 13,038 searches in 30.0 s, default@1 5,727 searches in 30.0 s.
- WARNING: latency had NOT settled when timing began for default@8 (each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it). exact: warm-up 15.0 s and 2,204 searches (0 failed) at 1 searcher; p50 of the last windows 6.584, 6.565, 6.581 ms (windows of at least 2 s and 100 searches); trial of 443 searches in 3.0 s at 1 searcher: p50 6.622 ms against the settled p50 6.581 ms, 1% apart (limit 10%); the timed pass p50 6.603 ms, 0% from the settled p50 6.581 ms (limit 10% or 0.1 ms; 0.022 ms); default@8: warm-up 15.0 s and 6,591 searches (0 failed) at 8 searchers; QPS of the last windows 472, 465, 405 (windows of at least 2 s and 100 searches) (its last 3 windows differed by more than 5%); trial of 1,295 searches in 3.0 s at 8 searchers: 431 QPS against the settled 465 QPS, 8% apart (limit 10%), so the warm-up was EXTENDED once (30.0 s and 12,492 searches (0 failed) at 8 searchers; QPS of its older 6 windows 427 (windows 436, 399, 412, 462, 462, 391) and of its newer 6 416 (windows 400, 458, 446, 411, 409, 370), 2.7% apart (limit 5%) (windows of at least 2 s and 100 searches)); second trial of 1,078 searches in 3.0 s at 8 searchers: 358 QPS against the settled 416 QPS, 16% apart (limit 10%), outside its newer windows' range 370 to 458 QPS; still disagreeing after the extension; the timed pass 349 QPS, 19% from the settled 416 QPS (limit 10%), NOT HELD: the engine was still changing when it was timed; default@1: warm-up 15.0 s and 2,479 searches (0 failed) at 1 searcher; p50 of the last windows 5.497, 5.449, 5.221 ms (windows of at least 2 s and 100 searches) (its last 3 windows differed by more than 5%); trial of 466 searches in 3.0 s at 1 searcher: p50 6.861 ms against the settled p50 5.449 ms, 21% apart (limit 10%), so the warm-up was EXTENDED once (30.0 s and 5,806 searches (0 failed) at 1 searcher; p50 of its older 6 windows 4.910 ms (windows 4.850, 4.932, 5.369, 4.842, 4.893, 4.847) and of its newer 6 4.855 ms (windows 4.875, 4.849, 4.863, 4.838, 4.868, 4.838), 1.1% apart (limit 5% or 0.1 ms; 0.056 ms) (windows of at least 2 s and 100 searches)); second trial of 595 searches in 3.0 s at 1 searcher: p50 4.851 ms against the settled p50 4.855 ms, 0% apart (limit 10%), inside its newer windows' range 4.838 to 4.875 ms; settled after the extension; the timed pass p50 4.859 ms, 0% from the settled p50 4.855 ms (limit 10% or 0.1 ms; 0.004 ms). Its numbers may still include warm-up or a change the engine was still going through; rerun before quoting them.
- Pass exact after rehearsal: 2026-10-08T00:36:14.552Z to 2026-10-08T00:37:14.554Z (60.0 s), 8,761 searches, 0 failed, p50 6.603 ms, mean 6.844 ms, p99 10.417 ms, 146.0 QPS (1000/QPS 6.849 ms).
- Pass default@8 after exact: 2026-10-08T00:38:20.628Z to 2026-10-08T00:38:40.644Z (20.0 s), 6,979 searches, 0 failed, p50 22.359 ms, mean 22.935 ms, p99 41.937 ms, 348.7 QPS (1000/QPS 2.868 ms).
- Pass default@1 after default@8: 2026-10-08T00:39:46.662Z to 2026-10-08T00:40:06.663Z (20.0 s), 3,937 searches, 0 failed, p50 4.859 ms, mean 5.077 ms, p99 7.807 ms, 196.8 QPS (1000/QPS 5.080 ms).
- Passes in the order run: exact, default@8, default@1; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 65,539 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (1 data part(s), index files on 1 of them, 524 of 524 live rows in indexed parts; 0 merge(s) running, 0 unfinished mutation(s); plan of the default search uses the Skip index vec_idx on 1 of 1 parts).
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
- Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-clickhouse (3d35aca5a71e) on network gvb-clickhouse_default at container-ip:8123, not through docker-proxy localhost:8123. Open after the passes: container-ip:8123 x1 (this target).
- CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-clickhouse: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-clickhouse: 724 thread(s) on CPUs 2-3,6-7 (every thread of the container).
- Clock per pass (median MHz of engine CPUs / client CPUs, each the median of its CPUs' medians and every CPU when the CPUs were not split, their average MHz, governor, CPUs busy outside the benchmark, client and engine CPU per search; the clock was pinned at 3500 MHz (turbo off), and a pass whose median on either group is more than 100 bp (1%) off it, or that has no reading on a group, is flagged): exact 3492/3492 MHz, average 3492.0/3492.9, performance, outside load 0.15 before the warm-up, 0.17 during, client CPU 0.876 ms per search, engine CPU 9.309 ms per search; default@8 3492/3492 MHz, average 3492.0/3492.3, performance, outside load 0.15 before the warm-up, 0.14 during, client CPU 1.054 ms per search, engine CPU 11.374 ms per search; default@1 3492/3492 MHz, average 3492.0/3492.9, performance, outside load 0.16 before the warm-up, 0.16 during, client CPU 0.858 ms per search, engine CPU 7.168 ms per search.
sqlite-vec sqlitevec
- Engine: SQLite 3.53.3 + sqlite-vec 0.1.7-alpha.2.1 (vec0, embedded, in process) (embedded)
- Index: vec0 brute-force scan, no ANN index (exact), float32, cosine distance, default chunk_size=1024; score = 1 - cosine distance; searches run concurrently, one WAL reader connection per searcher (opened as searchers arrive, at most 32), writes run one at a time and may overlap searches
- Search settings: chunk_size=1024
- Engine settings: sqlite_version = 3.53.3 (read: SELECT sqlite_version() on an in-memory SQLite connection); synchronous = NORMAL (set: src/GenericVectorBuilder.Engines/Sinks/SqliteVecSink.cs#PRAGMA synchronous=NORMAL); journal_mode = wal (read: PRAGMA journal_mode on a second, read-only connection to the bench file)
- Data folder ~/gvb-data/engines/sqlitevec: 292.93 MiB at the start, 302.65 MiB at the end; the benchmark's own database file 9.72 MiB
- Load: 524 rows in batches of 1000, 0.1 s of upserts (3,687 rows/s); count matched 0.0 s after the last upsert; index step 0.0 s (nothing to build, every search scans all vectors; checked in 0.00 s: no index, exact scan by design: gvb_gvbbench_eshoponweb is a vec0 virtual table holding 524 vectors and every search compares all of them (EXPLAIN QUERY PLAN: SCAN gvb_gvbbench_eshoponweb VIRTUAL TABLE INDEX 0:3{___}___ | USE TEMP B-TREE FOR ORDER BY))
- Search: 10572 latency samples, target held 524 rows, 0 errors
- Client CPU per search: 1 searcher 1.98 ms, 8 searchers 3.50 ms
- Exact mode: 31,572 searches, p50 1.87 ms, p95 2.08 ms, recall 1.000
- RAM: in this process, not measured; disk: 9.72 MiB added by this load (302.65 MiB (engine data folder), 292.93 MiB before)
- Index after the load: ready, ? of 524 indexed (no index, exact scan by design: gvb_gvbbench_eshoponweb is a vec0 virtual table holding 524 vectors and every search compares all of them (EXPLAIN QUERY PLAN: SCAN gvb_gvbbench_eshoponweb VIRTUAL TABLE INDEX 0:3{___}___ | USE TEMP B-TREE FOR ORDER BY)). Durability: PRAGMA journal_mode=WAL and PRAGMA synchronous=NORMAL (set by this sink when it opens the file): each commit is appended to the -wal file and handed to the operating system, but the WAL is fsynced only when SQLite checkpoints it (measured with strace: 40 single-record commits caused 3 WAL fsyncs), so a crash of the process loses nothing (measured with kill -9 and a reopen: all 50, 300 and 2,000 rows were there), while an operating-system crash or power cut can lose the newest commits since the last checkpoint and leaves the database consistent.
- Outside load when searching began: not known yet (too little sampled history). Load average 4.25 4.62 3.77 (1/5/15 min; it also counts this benchmark's own client and engines, so it is recorded, not judged).
- Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: exact 15,919 searches in 30.0 s, default@8 34,161 searches in 30.0 s, default@1 15,817 searches in 30.0 s.
- Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it. exact: warm-up 15.0 s and 7,886 searches (0 failed) at 1 searcher; p50 of the last windows 1.884, 1.886, 1.880 ms (windows of at least 2 s and 100 searches); trial of 1,541 searches in 3.0 s at 1 searcher: p50 1.886 ms against the settled p50 1.884 ms, 0% apart (limit 10%); the timed pass p50 1.874 ms, 1% from the settled p50 1.884 ms (limit 10% or 0.1 ms; 0.011 ms); default@8: warm-up 15.0 s and 17,079 searches (0 failed) at 8 searchers; QPS of the last windows 1,132, 1,139, 1,138 (windows of at least 2 s and 100 searches); trial of 3,429 searches in 3.0 s at 8 searchers: 1,142 QPS against the settled 1,138 QPS, 0% apart (limit 10%); the timed pass 1,140 QPS, 0% from the settled 1,138 QPS (limit 10%); default@1: warm-up 15.0 s and 7,955 searches (0 failed) at 1 searcher; p50 of the last windows 1.864, 1.870, 1.851 ms (windows of at least 2 s and 100 searches); trial of 1,606 searches in 3.0 s at 1 searcher: p50 1.841 ms against the settled p50 1.864 ms, 1% apart (limit 10%); the timed pass p50 1.863 ms, 0% from the settled p50 1.864 ms (limit 10% or 0.1 ms; 0.001 ms).
- Pass exact after rehearsal: 2026-10-08T00:42:02.672Z to 2026-10-08T00:43:02.673Z (60.0 s), 31,572 searches, 0 failed, p50 1.874 ms, mean 1.898 ms, p99 2.287 ms, 526.2 QPS (1000/QPS 1.900 ms).
- Pass default@8 after exact: 2026-10-08T00:43:20.712Z to 2026-10-08T00:43:40.715Z (20.0 s), 22,803 searches, 0 failed, p50 6.730 ms, mean 7.009 ms, p99 13.655 ms, 1140.0 QPS (1000/QPS 0.877 ms).
- Pass default@1 after default@8: 2026-10-08T00:43:58.726Z to 2026-10-08T00:44:18.727Z (20.0 s), 10,572 searches, 0 failed, p50 1.863 ms, mean 1.889 ms, p99 2.375 ms, 528.6 QPS (1000/QPS 1.892 ms).
- Passes in the order run: exact, default@8, default@1; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 105,393 sent, 0 failed (warmupErrors). Index after the searches: ready, ? of 524 indexed (no index, exact scan by design: gvb_gvbbench_eshoponweb is a vec0 virtual table holding 524 vectors and every search compares all of them (EXPLAIN QUERY PLAN: SCAN gvb_gvbbench_eshoponweb VIRTUAL TABLE INDEX 0:3{___}___ | USE TEMP B-TREE FOR ORDER BY)).
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
- Connection: in this process (embedded), no network: no network address. No open connection seen (no open TCP connection was found).
- CPU pinning: none: an embedded engine runs inside the benchmark client process, so it shares the client CPUs, CPUs 0-1,4-5.
- Clock per pass (median MHz of engine CPUs / client CPUs, each the median of its CPUs' medians and every CPU when the CPUs were not split, their average MHz, governor, CPUs busy outside the benchmark, client and engine CPU per search; the clock was pinned at 3500 MHz (turbo off), and a pass whose median on either group is more than 100 bp (1%) off it, or that has no reading on a group, is flagged): exact 3498/3492 MHz, average 3495.9/3494.0, performance, outside load 0.15 before the warm-up, 0.15 during, client CPU 1.998 ms per search; default@8 3492/3492 MHz, average 3493.7/3492.0, performance, outside load 0.15 before the warm-up, 0.14 during, client CPU 3.502 ms per search; default@1 3499/3492 MHz, average 3495.7/3494.0, performance, outside load 0.14 before the warm-up, 0.15 during, client CPU 1.983 ms per search.
Elasticsearch elasticsearch
- Engine: Elasticsearch 9.5.3 (dense_vector) (compose)
- Index: HNSW float32, no quantization, m=16, ef_construction=128, cosine; search k=top, num_candidates=100; 1 shard, 0 replicas; force-merged to one segment after the load (at 1,024 dimensions a segment under 1,043 vectors gets no graph)
- Search settings: ef_construction=128, k=top, m=16, num_candidates=100
- Engine settings: HostConfig.Memory = 6442450944 (read: docker inspect HostConfig.Memory (0 = no limit)); HostConfig.CpusetCpus = 2-3,6-7 (read: docker inspect HostConfig.CpusetCpus (empty = every CPU)); HostConfig.NanoCpus = 0 (read: docker inspect HostConfig.NanoCpus (0 = no limit)); heap_init_in_bytes = 2147483648 (read: GET /_nodes/jvm (nodes.<id>.jvm.mem.heap_init_in_bytes and heap_max_in_bytes)); heap_max_in_bytes = 2147483648 (read: GET /_nodes/jvm (nodes.<id>.jvm.mem.heap_init_in_bytes and heap_max_in_bytes))
- Image: elasticsearch:9.5.3, id sha256:e23d4758358a4e356cc2ef3259a5a6f345cacc1722c10dffe6d7150a1a78e52f (the pinned id), tagged on this machine at 2026-10-03T14:07:11.337389556Z
- Data folder ~/gvb-data/engines/elasticsearch: 198.92 KiB at the start, 2.58 MiB at the end
- Load: 524 rows in batches of 1000, 1.1 s of upserts (473 rows/s); count matched 0.1 s after the last upsert; index step 0.4 s (FAILED after 0.4 s: NO HNSW GRAPH, searches scan all 524 vectors: 1 segment(s), 524 vectors, total_vex_size_bytes 0 (_stats dense_vector), index_options hnsw m=16 ef_construction=128; Lucene builds no graph for a segment this small (measured at 1,024 dimensions: 1,042 vectors none, 1,043 vectors a graph), so default search equals exact search)
- Search: 14844 latency samples, target held 524 rows, 0 errors
- Client CPU per search: 1 searcher 0.45 ms, 8 searchers 0.49 ms
- Exact mode: 46,883 searches, p50 1.26 ms, p95 1.45 ms, recall 1.000
- RAM: 2.58 GiB (docker stats); disk: 2.39 MiB added by this load (2.58 MiB (whole engine data folder), 198.92 KiB before)
- Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started elasticsearch.compose.yaml with every container created on CPUs 2-3,6-7: gvb-elasticsearch cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards.
- Engine files (SHA-256, first 12 hex digits): [elasticsearch.compose.yaml 582b9a97b8de]. This run created the engine from these files.
- WARNING: the index step after the load FAILED (NO HNSW GRAPH, searches scan all 524 vectors: 1 segment(s), 524 vectors, total_vex_size_bytes 0 (_stats dense_vector), index_options hnsw m=16 ef_construction=128; Lucene builds no graph for a segment this small (measured at 1,024 dimensions: 1,042 vectors none, 1,043 vectors a graph), so default search equals exact search). Searched anyway; see the index state.
- WARNING: index not ready after the load: NOT ready, 0 of 524 indexed (NO HNSW GRAPH, searches scan all 524 vectors: 1 segment(s), 524 vectors, total_vex_size_bytes 0 (_stats dense_vector), index_options hnsw m=16 ef_construction=128; Lucene builds no graph for a segment this small (measured at 1,024 dimensions: 1,042 vectors none, 1,043 vectors a graph), so default search equals exact search). Measured anyway, so its search numbers may come from a scan or a half-built index. Durability: Every acknowledged bulk request is fsynced to the translog before the answer: index.translog.durability=request, the Elasticsearch default, which neither this sink nor elasticsearch.compose.yaml overrides (ElasticsearchReadinessTests reads it back from the live index). By that setting a process crash or power loss loses no acknowledged write; this is read from the setting, not shown by pulling power. One node and no replicas, so a lost disk loses the data..
- Outside load when searching began: 0.14 CPUs on average over the last 2 s and 0.14 over the last few seconds (limit 0.3), a quiet box. The load average 2.77 3.74 3.62 (1/5/15 min; it also counts this benchmark's own client and engines, so it is recorded, not judged).
- Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@1 13,644 searches in 30.0 s, default@8 74,239 searches in 30.0 s, exact 22,284 searches in 30.0 s.
- Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it. default@1: warm-up 15.0 s and 11,036 searches (0 failed) at 1 searcher; p50 of the last windows 1.332, 1.322, 1.330 ms (windows of at least 2 s and 100 searches); trial of 2,217 searches in 3.0 s at 1 searcher: p50 1.331 ms against the settled p50 1.330 ms, 0% apart (limit 10%); the timed pass p50 1.325 ms, 0% from the settled p50 1.330 ms (limit 10% or 0.1 ms; 0.005 ms); default@8: warm-up 15.0 s and 40,268 searches (0 failed) at 8 searchers; QPS of the last windows 2,706, 2,676, 2,703 (windows of at least 2 s and 100 searches); trial of 8,063 searches in 3.0 s at 8 searchers: 2,685 QPS against the settled 2,703 QPS, 1% apart (limit 10%); the timed pass 2,694 QPS, 0% from the settled 2,703 QPS (limit 10%); exact: warm-up 15.0 s and 11,653 searches (0 failed) at 1 searcher; p50 of the last windows 1.268, 1.255, 1.259 ms (windows of at least 2 s and 100 searches); trial of 2,339 searches in 3.0 s at 1 searcher: p50 1.259 ms against the settled p50 1.259 ms, 0% apart (limit 10%); the timed pass p50 1.258 ms, 0% from the settled p50 1.259 ms (limit 10% or 0.1 ms; 0.001 ms).
- Pass default@1 after rehearsal: 2026-10-08T00:46:40.107Z to 2026-10-08T00:47:00.108Z (20.0 s), 14,844 searches, 0 failed, p50 1.325 ms, mean 1.345 ms, p99 1.772 ms, 742.2 QPS (1000/QPS 1.347 ms).
- Pass default@8 after default@1: 2026-10-08T00:47:18.128Z to 2026-10-08T00:47:38.130Z (20.0 s), 53,882 searches, 0 failed, p50 2.839 ms, mean 2.967 ms, p99 5.345 ms, 2693.8 QPS (1000/QPS 0.371 ms).
- Pass exact after default@8: 2026-10-08T00:47:56.149Z to 2026-10-08T00:48:56.149Z (60.0 s), 46,883 searches, 0 failed, p50 1.258 ms, mean 1.278 ms, p99 1.696 ms, 781.4 QPS (1000/QPS 1.280 ms).
- Passes in the order run: default@1, default@8, exact; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 185,743 sent, 0 failed (warmupErrors). Index after the searches: NOT ready, 0 of 524 indexed (NO HNSW GRAPH, searches scan all 524 vectors: 1 segment(s), 524 vectors, total_vex_size_bytes 0 (_stats dense_vector), index_options hnsw m=16 ef_construction=128; Lucene builds no graph for a segment this small (measured at 1,024 dimensions: 1,042 vectors none, 1,043 vectors a graph), so default search equals exact search).
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
- Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-elasticsearch (27ca3548f8ae) on network engines_default at container-ip:9200, not through docker-proxy localhost:9200. Open after the passes: container-ip:9200 x1 (this target).
- CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-elasticsearch: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-elasticsearch: 75 thread(s) on CPUs 2-3,6-7 (every thread of the container).
- Clock per pass (median MHz of engine CPUs / client CPUs, each the median of its CPUs' medians and every CPU when the CPUs were not split, their average MHz, governor, CPUs busy outside the benchmark, client and engine CPU per search; the clock was pinned at 3500 MHz (turbo off), and a pass whose median on either group is more than 100 bp (1%) off it, or that has no reading on a group, is flagged): default@1 3492/3492 MHz, average 3491.8/3491.1, performance, outside load 0.14 before the warm-up, 0.14 during, client CPU 0.453 ms per search, engine CPU 1.066 ms per search; default@8 3492/3492 MHz, average 3492.0/3492.1, performance, outside load 0.15 before the warm-up, 0.16 during, client CPU 0.490 ms per search, engine CPU 1.393 ms per search; exact 3492/3492 MHz, average 3492.2/3492.9, performance, outside load 0.16 before the warm-up, 0.14 during, client CPU 0.452 ms per search, engine CPU 0.974 ms per search.
Qdrant (exact) qdrant
- Engine: Qdrant 1.17.0 in container gvb-qdrant (image qdrant/qdrant:v1.17.0@sha256:f1c7272cdac52b38c1a0e89313922d940ba50afd90d593a1605dbbc214e66ffb; cpuset 2-3,6-7 from its creation; 3 search thread(s)) (compose)
- Index: exact scan: the builder's sink sends exact=true on every search, so no HNSW graph is used whether or not Qdrant has built one (see the index state)
- Search settings: exact=true
- Engine settings: HostConfig.Memory = 8589934592 (read: docker inspect HostConfig.Memory (0 = no limit)); HostConfig.CpusetCpus = 2-3,6-7 (read: docker inspect HostConfig.CpusetCpus (empty = every CPU)); HostConfig.NanoCpus = 0 (read: docker inspect HostConfig.NanoCpus (0 = no limit)); search_threads = 3 (read: threads named search-* of the qdrant process in container gvb-qdrant (/proc))
- Image: qdrant/qdrant:v1.17.0@sha256:f1c7272cdac52b38c1a0e89313922d940ba50afd90d593a1605dbbc214e66ffb, id sha256:f1c7272cdac52b38c1a0e89313922d940ba50afd90d593a1605dbbc214e66ffb (the pinned id), tagged on this machine at 2026-10-04T22:03:16.973657523Z
- Data folder ~/gvb-data/engines/qdrant: 357 B at the start, 196.08 MiB at the end
- Load: 524 rows in batches of 1000, 0.1 s of upserts (5,738 rows/s); count matched 0.0 s after the last upsert; index step 1.0 s (status Green, 0 of 524 vectors in HNSW segments (2 segments), waited 1.0 s)
- Search: 22185 latency samples, target held 524 rows, 0 errors
- Client CPU per search: 1 searcher 0.72 ms, 8 searchers 0.54 ms
- RAM: 35.61 MiB (docker stats); disk: 196.08 MiB (collection folder)
- Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started qdrant.compose.yaml with every container created on CPUs 2-3,6-7: gvb-qdrant cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards.
- Engine files (SHA-256, first 12 hex digits): [qdrant.compose.yaml eb57e55e93c1]. This run created the engine from these files.
- Index after the load: ready, 0 of 524 indexed (status Green, optimizer ok, indexed_vectors_count 0 of 524 points, 2 segments, indexing_threshold_kb 10,000, full_scan_threshold_kb 10,000; no index used, exact scan by design (the builder's sink sends exact=true on every search); no graph was built; searches counted since the collection was created: unfiltered_hnsw 0, unfiltered_plain 0, unfiltered_exact 0). Durability: What was measured (strace -f on the Qdrant server process, 30 single-point REST upserts per setting, every sync call timed against the request that caused it): with wait=true each upsert had exactly one msync(MS_SYNC) of the write-ahead-log segment inside the request, before the reply (30 of 30); with wait=false the 30 upserts were acknowledged within 0.26 s with no sync call, and the first WAL msync (one call covering all 30 records) came 2.7 s after the last reply, followed by the segment-file flushes. So the log is flushed to disk under both settings, but with wait=false the flush comes after the acknowledgement: an operating-system crash or power cut in that gap loses acknowledged writes, while a crash of the Qdrant process alone should not, because the bytes are already in the kernel's page cache (inferred, not tested). Every upsert here is sent with wait=true in batches of 256 over gRPC; the strace used one point per request, so one flush per batch is inferred, not measured. Segment files are flushed every 5 s; log segments 32 MB, 0 created ahead. /qdrant/config/config.yaml in container gvb-qdrant (the image's own file; its storage keys are listed) sets: storage.collection.quantization = null, storage.collection.replication_factor = 1, storage.collection.vectors.on_disk = null, storage.collection.write_consistency_factor = 1, storage.hnsw_index.ef_construct = 100, storage.hnsw_index.full_scan_threshold_kb = 10000, storage.hnsw_index.m = 16, storage.hnsw_index.max_indexing_threads = 0, storage.hnsw_index.on_disk = false, storage.hnsw_index.payload_m = null, storage.max_collections = null, storage.node_type = Normal, storage.on_disk_payload = true, storage.optimizers.default_segment_number = 0, storage.optimizers.deleted_threshold = 0.2, storage.optimizers.flush_interval_sec = 5, storage.optimizers.indexing_threshold_kb = 10000, storage.optimizers.max_optimization_threads = null, storage.optimizers.max_segment_size_kb = null, storage.optimizers.vacuum_min_vector_number = 1000, storage.performance.max_search_threads = 0, storage.performance.optimizer_cpu_budget = 0, storage.performance.update_rate_limit = null, storage.shard_transfer_method = null, storage.snapshots_config.snapshots_storage = local, storage.snapshots_path = ./snapshots, storage.storage_path = ./storage, storage.temp_path = null, storage.update_concurrency = null, storage.wal.wal_capacity_mb = 32, storage.wal.wal_segments_ahead = 0; no QDRANT__ environment overrides in the server process. Not tested by cutting power; whether the disk's own write cache reaches the media was not checked..
- Outside load when searching began: 0.13 CPUs on average over the last 1 s and 0.13 over the last few seconds (limit 0.3), a quiet box. The load average 2.09 3.40 3.55 (1/5/15 min; it also counts this benchmark's own client and engines, so it is recorded, not judged).
- Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@8 99,055 searches in 30.0 s, default@1 33,257 searches in 30.0 s.
- Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it. default@8: warm-up 15.0 s and 49,295 searches (0 failed) at 8 searchers; QPS of the last windows 3,302, 3,296, 3,240 (windows of at least 2 s and 100 searches); trial of 9,880 searches in 3.0 s at 8 searchers: 3,291 QPS against the settled 3,296 QPS, 0% apart (limit 10%); the timed pass 3,296 QPS, 0% from the settled 3,296 QPS (limit 10%); default@1: warm-up 15.0 s and 16,685 searches (0 failed) at 1 searcher; p50 of the last windows 0.887, 0.891, 0.884 ms (windows of at least 2 s and 100 searches); trial of 3,306 searches in 3.0 s at 1 searcher: p50 0.892 ms against the settled p50 0.887 ms, 1% apart (limit 10%); the timed pass p50 0.890 ms, 0% from the settled p50 0.887 ms (limit 10% or 0.1 ms; 0.003 ms).
- Pass default@8 after rehearsal: 2026-10-08T00:50:26.107Z to 2026-10-08T00:50:46.109Z (20.0 s), 65,925 searches, 0 failed, p50 2.343 ms, mean 2.425 ms, p99 4.191 ms, 3295.9 QPS (1000/QPS 0.303 ms).
- Pass default@1 after default@8: 2026-10-08T00:51:04.136Z to 2026-10-08T00:51:24.137Z (20.0 s), 22,185 searches, 0 failed, p50 0.890 ms, mean 0.900 ms, p99 1.430 ms, 1109.2 QPS (1000/QPS 0.902 ms).
- Passes in the order run: default@8, default@1; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 211,478 sent, 0 failed (warmupErrors). Index after the searches: ready, 0 of 524 indexed (status Green, optimizer ok, indexed_vectors_count 0 of 524 points, 2 segments, indexing_threshold_kb 10,000, full_scan_threshold_kb 10,000; no index used, exact scan by design (the builder's sink sends exact=true on every search); no graph was built; searches counted since the collection was created: unfiltered_hnsw 0, unfiltered_plain 599,176, unfiltered_exact 0).
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
- Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-qdrant (34f09e7ee805) on network gvb-qdrant_default at container-ip:6334, not through docker-proxy localhost:16334; container gvb-qdrant (34f09e7ee805) on network gvb-qdrant_default at container-ip:6333, not through docker-proxy localhost:16333. Open after the passes: container-ip:6334 x1 (this target).
- CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-qdrant: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-qdrant: 24 thread(s) on CPUs 2-3,6-7 (every thread of the container).
- Clock per pass (median MHz of engine CPUs / client CPUs, each the median of its CPUs' medians and every CPU when the CPUs were not split, their average MHz, governor, CPUs busy outside the benchmark, client and engine CPU per search; the clock was pinned at 3500 MHz (turbo off), and a pass whose median on either group is more than 100 bp (1%) off it, or that has no reading on a group, is flagged): default@8 3492/3492 MHz, average 3492.0/3492.0, performance, outside load 0.10 before the warm-up, 0.14 during, client CPU 0.535 ms per search, engine CPU 1.002 ms per search; default@1 3492/3492 MHz, average 3492.0/3491.9, performance, outside load 0.13 before the warm-up, 0.11 during, client CPU 0.722 ms per search, engine CPU 0.787 ms per search.
Oracle 23ai Free oracle
- Engine: Oracle AI Database Free 23.26 (23ai line, VECTOR FLOAT32) (compose)
- Index: HNSW in-memory neighbor graph NEIGHBORS=16 EFCONSTRUCTION=128, EFSEARCH=100 per query, cosine; exact mode = FETCH EXACT FIRST (full scan); Oracle Free caps itself at 2 CPUs (cpu_count 2 in V$PARAMETER, edition FREE in V$INSTANCE, 8 host CPUs in V$OSSTAT NUM_CPUS; the 2 CPU thread limit is Oracle's documented Free edition limit)
- Search settings: EFCONSTRUCTION=128, EFSEARCH=100, NEIGHBORS=16
- Engine settings: HostConfig.Memory = 6442450944 (read: docker inspect HostConfig.Memory (0 = no limit)); HostConfig.CpusetCpus = 2-3,6-7 (read: docker inspect HostConfig.CpusetCpus (empty = every CPU)); HostConfig.NanoCpus = 0 (read: docker inspect HostConfig.NanoCpus (0 = no limit)); cpu_count = 2 (read: docker exec gvb-oracle sqlplus / as sysdba, V$PARAMETER and V$INSTANCE at the instance); sga_target = 1610612736 (read: docker exec gvb-oracle sqlplus / as sysdba, V$PARAMETER and V$INSTANCE at the instance); oracle_version = 23.26.3.0.0 (read: docker exec gvb-oracle sqlplus / as sysdba, V$PARAMETER and V$INSTANCE at the instance)
- Image: gvenzl/oracle-free:23-slim-faststart, id sha256:f5ff19033860d662c821cb04eb10483fa94f14f78eae252d054291ea07028093 (the pinned id), tagged on this machine at 2026-10-03T14:17:05.132133694Z
- Data folder ~/gvb-data/engines/oracle: 6.9 GiB at the start, 6.9 GiB at the end
- Load: 524 rows in batches of 1000, 0.3 s of upserts (1,951 rows/s); count matched 0.1 s after the last upsert; index step 3.5 s (HNSW graph holds 524 of 524 rows, change log waiting: 0 inserts and 0 deletes, USER_INDEXES status VALID, plan of the default search: VECTOR INDEX HNSW SCAN > TABLE ACCESS BY INDEX ROWID, index used by 0 queries so far; Oracle Free caps itself at 2 CPUs (cpu_count 2 in V$PARAMETER, edition FREE in V$INSTANCE, 8 host CPUs in V$OSSTAT NUM_CPUS; the 2 CPU thread limit is Oracle's documented Free edition limit); graph rebuilt in 2.3 s)
- Search: 20998 latency samples, target held 524 rows, 0 errors
- Client CPU per search: 1 searcher 0.89 ms, 8 searchers 0.64 ms
- Exact mode: 24,106 searches, p50 2.42 ms, p95 2.91 ms, recall 1.000
- RAM: 2.14 GiB (docker stats); disk: 0 B added by this load (6.9 GiB (whole engine data folder), 6.9 GiB before)
- Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started oracle.compose.yaml with every container created on CPUs 2-3,6-7: gvb-oracle cpuset 2-3,6-7, gvb-oracle-seed cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards.
- Engine files (SHA-256, first 12 hex digits): [oracle-init/01-vector-memory.sh 6452e4e1953e; oracle.compose.yaml 2a7fd367012f]. This run created the engine from these files.
- Index after the load: ready, 524 of 524 indexed (HNSW graph holds 524 of 524 rows, change log waiting: 0 inserts and 0 deletes, USER_INDEXES status VALID, plan of the default search: VECTOR INDEX HNSW SCAN > TABLE ACCESS BY INDEX ROWID, index used by 0 queries so far; Oracle Free caps itself at 2 CPUs (cpu_count 2 in V$PARAMETER, edition FREE in V$INSTANCE, 8 host CPUs in V$OSSTAT NUM_CPUS; the 2 CPU thread limit is Oracle's documented Free edition limit)). Durability: By these settings a committed row should survive a crash or power loss. The sink uses a plain COMMIT and commit_logging, commit_wait and commit_write are unset in the database, so every commit waits for its redo to be written. Measured 2026-10-04: 50 separate client commits raised V$SYSSTAT 'redo synch writes' by 55 (the 50 commits plus the CREATE and DROP of the probe table), and the log writer and the datafile writer hold their files open with O_DSYNC (open flags 02110002, filesystemio_options none). The database runs NOARCHIVELOG (V$DATABASE.LOG_MODE), so redo serves crash recovery only and there is no point-in-time restore. The HNSW graph lives in the 768 MB vector memory pool (oracle-init/01-vector-memory.sh) and is not the durable copy; the table is. Not tested by cutting power; whether the disk's own write cache reaches the media was not checked. CPU: Oracle Free caps itself at 2 CPUs (measured 2026-10-04: cpu_count 2 in V$PARAMETER, edition FREE in V$INSTANCE, 8 host CPUs in V$OSSTAT NUM_CPUS; the 2 CPU thread limit is Oracle's documented Free edition limit; this sink reads the live value when the index state is read)..
- Outside load when searching began: 0.18 CPUs on average over the last 6 s and 0.20 over the last few seconds (limit 0.3), a quiet box. The load average 3.13 3.67 3.65 (1/5/15 min; it also counts this benchmark's own client and engines, so it is recorded, not judged).
- Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@8 85,783 searches in 30.0 s, default@1 31,641 searches in 30.0 s, exact 12,032 searches in 30.0 s.
- Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure, and the timed pass itself within 10% of it. default@8: warm-up 15.0 s and 46,395 searches (0 failed) at 8 searchers; QPS of the last windows 3,377, 2,408, 3,252 (windows of at least 2 s and 100 searches) (its last 3 windows differed by more than 5%); trial of 7,281 searches in 3.2 s at 8 searchers: 2,304 QPS against the settled 3,252 QPS, 41% apart (limit 10%), so the warm-up was EXTENDED once (30.5 s and 86,462 searches (0 failed) at 8 searchers; QPS of its older 7 windows 2,836 (windows 3,407, 2,408, 3,180, 2,626, 2,759, 3,070, 2,404) and of its newer 7 2,841 (windows 3,436, 2,380, 3,140, 2,756, 2,647, 3,172, 2,358), 0.2% apart (limit 5%) (windows of at least 2 s and 100 searches)); second trial of 9,127 searches in 3.0 s at 8 searchers: 3,041 QPS against the settled 2,841 QPS, 7% apart (limit 10%), inside its newer windows' range 2,358 to 3,436 QPS; settled after the extension; the timed pass 2,792 QPS, 2% from the settled 2,841 QPS (limit 10%); default@1: warm-up 15.0 s and 15,721 searches (0 failed) at 1 searcher; p50 of the last windows 0.941, 0.924, 0.913 ms (windows of at least 2 s and 100 searches); trial of 3,171 searches in 3.0 s at 1 searcher: p50 0.925 ms against the settled p50 0.924 ms, 0% apart (limit 10%); the timed pass p50 0.928 ms, 0% from the settled p50 0.924 ms (limit 10% or 0.1 ms; 0.004 ms); exact: warm-up 15.0 s and 6,004 searches (0 failed) at 1 searcher; p50 of the last windows 2.438, 2.435, 2.437 ms (windows of at least 2 s and 100 searches); trial of 1,209 searches in 3.0 s at 1 searcher: p50 2.426 ms against the settled p50 2.437 ms, 0% apart (limit 10%); the timed pass p50 2.422 ms, 1% from the settled p50 2.437 ms (limit 10% or 0.1 ms; 0.014 ms).
- Pass default@8 after rehearsal: 2026-10-08T00:54:20.344Z to 2026-10-08T00:54:40.467Z (20.1 s), 56,192 searches, 0 failed, p50 1.614 ms, mean 2.862 ms, p99 3.555 ms, 2792.4 QPS (1000/QPS 0.358 ms).
- Pass default@1 after default@8: 2026-10-08T00:54:58.490Z to 2026-10-08T00:55:18.490Z (20.0 s), 20,998 searches, 0 failed, p50 0.928 ms, mean 0.950 ms, p99 1.889 ms, 1049.9 QPS (1000/QPS 0.953 ms).
- Pass exact after default@1: 2026-10-08T00:55:36.515Z to 2026-10-08T00:56:36.515Z (60.0 s), 24,106 searches, 0 failed, p50 2.422 ms, mean 2.487 ms, p99 3.871 ms, 401.8 QPS (1000/QPS 2.489 ms).
- Passes in the order run: default@8, default@1, exact; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 347,110 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (HNSW graph holds 524 of 524 rows, change log waiting: 0 inserts and 0 deletes, USER_INDEXES status VALID, plan of the default search: VECTOR INDEX HNSW SCAN > TABLE ACCESS BY INDEX ROWID, index used by 405055 queries so far; Oracle Free caps itself at 2 CPUs (cpu_count 2 in V$PARAMETER, edition FREE in V$INSTANCE, 8 host CPUs in V$OSSTAT NUM_CPUS; the 2 CPU thread limit is Oracle's documented Free edition limit)).
- Benchmark copy gvbbench_eshoponweb dropped afterwards.
- Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-oracle (308d802d0653) on network gvb-oracle_default at container-ip:1521, not through docker-proxy localhost:1521. Open after the passes: container-ip:1521 x15 (this target).
- CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-oracle: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-oracle: 86 thread(s) on CPUs 2-3,6-7 (every thread of the container).
- Clock per pass (median MHz of engine CPUs / client CPUs, each the median of its CPUs' medians and every CPU when the CPUs were not split, their average MHz, governor, CPUs busy outside the benchmark, client and engine CPU per search; the clock was pinned at 3500 MHz (turbo off), and a pass whose median on either group is more than 100 bp (1%) off it, or that has no reading on a group, is flagged): default@8 3492/3492 MHz, average 3102.8/3262.5, performance, outside load 0.14 before the warm-up, 0.13 during, client CPU 0.643 ms per search, engine CPU 0.720 ms per search; default@1 3492/3492 MHz, average 3492.1/3492.0, performance, outside load 0.14 before the warm-up, 0.15 during, client CPU 0.885 ms per search, engine CPU 0.575 ms per search; exact 3492/3492 MHz, average 3492.4/3492.3, performance, outside load 0.14 before the warm-up, 0.15 during, client CPU 0.931 ms per search, engine CPU 2.116 ms per search.
Notes
- Load average at start: 2.20 3.63 3.60 on 8 logical CPUs (1/5/15 min). It counts this benchmark's own earlier work and the engines it started, so it is recorded here and never judged; each timed pass is judged by the outside load machine control measures (conditions.passes).
- Order: targets ran one at a time in a random order from seed 803 (runSeed; --seed 803 repeats it, targetOrder lists it). Inside each target the timed passes also ran in a random order from the same seed and the target's name (passOrder): default@1 is one searcher for 20 s (every search's latency gives p50/p95/p99, the completed searches give QPS@1, its first answer to each query gives recall and nDCG), default@N is N searchers for 20 s (QPS@N), exact is the engine's exact mode, one searcher for 60 s cycling the queries.
- Preparation, untimed, before any timed pass: a rehearsal of every pass type at its own concurrency for 30 s each, through the same code the passes use.
- Warm-up and settle check, untimed, right before every timed pass: the pass's own search at the pass's own number of searchers for at least 15 s and at least 20 searches (at most 120 s), read in windows of at least 2 s and 100 searches; then a 3 s trial of the same pass. The trial's figure (p50 with one searcher, QPS with several) must lie within 10% of the warm-up's settled figure (the median of its last 3 windows, which must agree within 5%); if not, the warm-up is extended once (at least 30 s, until its windows agree, at most 120 s), where an extension's windows agree when its level test passes: the older and the newer half of its latest windows, 5 to 10 windows a half, each half read as the pass reads it (the p50 of all its searches with one searcher, its searches per second with several), agree within 5%, judged on the windows that stopped it (an extension whose cap runs out with fewer than 10 windows is judged by its last 3 instead); then a second trial is taken, which must lie inside the range of the newer half's windows or within 10% of the newer half's figure, and the warm-up runs again before the pass. After the pass its own figure is held against the settled figure, within 10%. In the level test, the second trial and the hold, two one-searcher p50s no more than 0.1 ms apart agree whatever their percentage (below what this box resolves at one searcher). Each target's notes give every check, and a pass that still disagrees, or whose own figure did not hold, flags its target as unsettled. With machine control on, the check for a quiet box is made when the warm-up is announced, before it starts (a wait between the warm-up and the timed pass let the engine go cold), so by the time the clock opens that check is as old as the warm-up and the trial (about 18 s, more after an extension); each pass's conditions record that lead (quietCheckLeadSeconds). Failed untimed searches are counted per target (warmupErrors) and are not in the timed error counts.
- Load rows/s counts only time inside each target's upsert calls: one writer, batches of 1000, rows already in memory, the collection dropped and created fresh first. Engines that build or finish their index after the writes do it in a separate timed index step (the load's index seconds), and the run waits for it before searching.
- Index proof: each engine's own report of its index (indexState) is read after the load and again after the last pass. A target whose index was not ready after the load was still measured and carries a WARNING; an engine that reports nothing counts as not ready.
- Search settings: each searched target records the settings its own index description states (searchSettings: the build parameters and the effort per query, such as m, ef_construction and ef_search), read after the load; consolidating never averages runs whose settings differ.
- Durability: each target's crash-safety setting as configured here (durability). Engines that do not force writes to disk on every commit load faster for that reason.
- Engine record: right after each engine was bound, its settings (engineSettings: the container's limits and what the engine itself reports, each entry saying how it was obtained), its image against the id pinned in deploy/bench/image-pins.json (image; another id ends that target with an error) and the size of its data folder (dataFolder: at the start, after any reset, and at the end) were recorded. In run-all, ClickHouse's system log tables are truncated before its load, and the reset is written into its dataFolder.
- Latency is client-side wall time around each search (network and driver included), every search of the default@1 window, one searcher, the queries cycled; a window with fewer than 200 searches, a p50 above 1.25 x its mean, or a mean above its p99 is flagged.
- QPS: N searchers back to back for 20 s per level; completed searches divided by the window's elapsed time.
- Recall@10: share of the exact top 10 (brute force in memory) that the engine returned. A hit whose exact similarity ties the 10th best (within 1e-5) also counts, because duplicate rows embed to identical vectors.
- Routes: a container engine (the benchmark's own SQL Server and Qdrant containers included) is reached at its container's own address on its Docker network, never through the published localhost port (docker-proxy); the native comparison targets (sql-native, qdrant-native) are native services reached directly over loopback. Each target's addresses and the connections the client held open after its passes are in its notes and in conditions.connections.
- RAM of compose engines, the benchmark's own SQL Server and Qdrant containers (sql, sql-diskann, qdrant, qdrant-hnsw) included, is docker stats of the engine's containers; for the native comparison targets (sql-native, qdrant-native) it is the whole native process, including every other database or collection it serves. Disk is the table's reserved pages (SQL), the collection folder (Qdrant), or for other engines what the load added to the engine's data folder (engines that keep data in memory until a snapshot show almost nothing); when the folder was smaller after the load than before it, no figure is given.
- Engines: run-all starts an engine that is down and stops it afterwards only if it was not running when the run began; an engine that was already running is left running.
- WARNING: the engine did not report a ready index after the load for: elasticsearch. They were measured anyway; their numbers may come from a scan or a half-built index (see each target's index state).
- Machine control on: governor performance on every CPU during the run (before: schedutil; at the end: performance; after putting it back: schedutil). CPU clock pinned for the run: turbo off (intel_pstate/no_turbo 0 -> 1), so every CPU's clock is held at its ceiling of 3500 MHz whatever the engine runs [Correction 13]; the uncore (L3 and memory) clock is held at 3000 MHz (min 3000 MHz, max 3000 MHz (MSR 0x620 = 0x1e1e); was min 1200 MHz, max 3000 MHz (MSR 0x620 = 0xc1e)) (before: turbo on (intel_pstate/no_turbo 0), ceiling 3600 MHz on CPUs 0-7, floor 1200 MHz on CPUs 0-7; while pinned: turbo off (intel_pstate/no_turbo 1), ceiling 3500 MHz on CPUs 0-7, floor 1200 MHz on CPUs 0-7; at the end: turbo off (intel_pstate/no_turbo 1), ceiling 3500 MHz on CPUs 0-7, floor 1200 MHz on CPUs 0-7; after putting it back: turbo on (intel_pstate/no_turbo 0), ceiling 3600 MHz on CPUs 0-7, floor 1200 MHz on CPUs 0-7). Uncore limit at the end: min 3000 MHz, max 3000 MHz (MSR 0x620 = 0x1e1e); after putting it back: min 1200 MHz, max 3000 MHz (MSR 0x620 = 0xc1e). Each target's clock note gives every pass's median MHz on the engine CPUs and on the client CPUs (the median of each CPU's median); a pass whose median on either is more than 100 bp (1%) off the pinned 3500 MHz, or that has no reading on one, is flagged (conditions.clock.rule). Figures from runs made with turbo on are not comparable with these in absolute terms. engine CPUs 2-3,6-7 (cores 2,6 and 3,7), client CPUs 0-1,4-5 (cores 0,4 and 1,5); the client process was pinned; each engine was pinned to the engine CPUs for its turn and put back after (conditions.engines); an engine run-all started (and stopped) was asked to be created on them, and its notes say whether the host did so or it was moved there after the start. Busy box: right before each pass's warm-up the run waits, up to 10 min, while processes outside the benchmark (everything but this client and the engine under test's cgroups) use more than 0.3 CPUs on average over the last 60 s or the last 5 s, neither window reaching back past the start of the target or of its engine (so the engine's own start-up is not outside work); if the box does not clear the pass runs anyway, flagged 'busy box', as is a pass whose own outside load is above the limit. Outside load counts busy = user + nice + system + irq + softirq + steal (guest time is already in user); CONFIG_IRQ_TIME_ACCOUNTING is not set, so task and cgroup run time include the interrupt and softirq time that hit them and it is added; CONFIG_PARAVIRT_TIME_ACCOUNTING is not set, so steal is added (/boot/config-6.8.0-142-generic) (conditions.cpuAccounting); each pass also records this client's own CPU time per search and the engine's (conditions.passes[].clientCpuMsPerSearch and engineCpuMsPerSearch; how in conditions.engineCpu.rule). CPU clocks were sampled every 250 ms; each pass's min/median/max per CPU is in conditions.passes. CPU idle states, recorded and left as found: driver intel_idle, governor menu, intel_idle max_cstate 9; POLL on, C1 on, C1E on, C3 OFF on CPUs 0-7 (default disabled), C6 on (conditions.cpuIdle). How each target was reached: conditions.connections. Client build Release, .NET 10.0.12.
- duckdb is embedded: it ran inside the client process on the client CPUs 0-1,4-5, sharing them with the client.
- sqlitevec is embedded: it ran inside the client process on the client CPUs 0-1,4-5, sharing them with the client.
Raw files
The raw files below hold the recorded statements that the corrections above refer to, as written.