# Vector engine benchmark: eshoponweb Run 2026-10-05T08:54:48Z (run-all). Command line: ``` ~/gvb-work/lanes/v6-final/src/GenericVectorBuilder.Bench/bin/Release/net10.0/GenericVectorBuilder.Bench.dll run-all --pipeline eshoponweb --queries golden --seed 602 --out ~/ForClaude/GenericVectorBuilder/bench-results ``` - Machine: bench-host, Intel(R) Xeon(R) CPU E5-1620 v3 @ 3.50GHz (8 logical CPUs), 62.7 GiB RAM, GPU NVIDIA GeForce RTX 3060, 12288 MiB, Ubuntu 24.04.5 LTS, kernel 6.8.0-142-generic, .NET 10.0.12, load average at start 3.32 2.91 3.31 - Data: 524 vectors x 1024 dims, collection `gvbbench_eshoponweb`. Read READ-ONLY from GenericVectorBuilder.dbo.gvb_eshoponweb in ChunkId order, vectors as native SqlVector, 0.2 s. - Queries: golden: 20 labelled questions from ~/ForClaude/evalkit/questions_golden.json, embedded with qwen3-emb-0.6b (cached in ~/gvb-data/bench-cache/golden-eshoponweb-68df42ca69efa079.json, no embedding calls). Top 10. Throughput at concurrency 1, 8 for 20 s each. - Request speed on a small collection (524 vectors): Measured end to end through each engine's .NET client; at this size it reflects per-request cost including the client library, not index scaling. - Client CPU per search is the CPU time the test's .NET client itself used for each search, measured in the same pass as the figure beside it. Where it is close to the latency, the client library is a large part of what is measured. For an embedded engine (DuckDB, sqlite-vec) the engine runs inside the client process, so its figure is the engine's own CPU time, not client overhead. - Ground truth: brute force over every vector in memory, 0.0 s for all 20 queries. - nDCG@10 of the exact answer itself (the ceiling for this embedder): 0.518 | engine | index | load rows/s | p50 ms | p95 ms | p99 ms | QPS@1 | QPS@8 | client CPU ms/search@1 | client CPU ms/search@8 | recall@10 | nDCG@10 | RAM | disk | |---|---|---|---|---|---|---|---|---|---|---|---|---|---| | sql | exact VECTOR_DISTANCE cosine, no vector index (full scan) | 575 | 3.82 | 4.35 | 5.69 | 255.6 | 613.9 | 1.18 | 1.18 | 1.000 | 0.518 | 530.8 MiB | 5.15 MiB | | opensearch | faiss HNSW float32, no compression, m=16, ef_construction=128, cosinesimil; search k=top, ef_search=100; 1 shard, 0 replicas; graph built at any segment size (approximate_threshold=0); force-merged to one segment after the load | 421 | 2.00 | 2.25 | 2.78 | 494.1 | 1511.9 | 0.46 | 0.51 | 1.000 | 0.518 | 2.56 GiB | 9.55 MiB | | pgvector | HNSW vector_cosine_ops m=16 ef_construction=128, hnsw.ef_search=100 per query, float32 vector(n), cosine; exact mode = same query with index scans off (sequential scan) | 611 | 0.82 | 0.95 | 1.22 | 1200.4 | 4142.5 | 0.66 | 0.53 | 1.000 | 0.518 | 100.8 MiB | 7.58 MiB | | duckdb | HNSW (vss extension) FLOAT[n] metric=cosine m=16 ef_construction=128, ef_search=100 per connection, persistent (hnsw_enable_experimental_persistence=true, checkpoint_threshold=256MB); exact mode = array_cosine_similarity sequential scan; score = 1 - cosine distance; searches run concurrently, one connection per searcher (opened as searchers arrive, at most 32), writes run one at a time and never overlap a search | 1,559 | 3.52 | 3.92 | 4.99 | 280.9 | 598.5 | 4.84 | 6.64 | 1.000 | 0.518 | - | 2.9 MiB | | milvus | HNSW M=16 efConstruction=128, ef=100, metric COSINE, Strong consistency searches; approximate only (no exact mode) | 1,175 | 2.22 | 2.74 | 3.65 | 430.1 | 1281.5 | 0.55 | 0.63 | 1.000 | 0.518 | 203.6 MiB | 10.87 MiB | | sqlitevec | vec0 brute-force scan, no ANN index (exact), float32, cosine distance, default chunk_size=1024; score = 1 - cosine distance; searches run concurrently, one WAL reader connection per searcher (opened as searchers arrive, at most 32), writes run one at a time and may overlap searches | 4,690 | 1.85 | 2.12 | 2.45 | 529.8 | 1150.7 | 1.97 | 3.47 | 1.000 | 0.518 | - | 9.72 MiB | | qdrant-hnsw | HNSW m=16 ef_construct=100, hnsw_ef=server default, cosine; indexing_threshold_kb 1 and full_scan_threshold_kb 10 (server defaults are 10,000 each) so a small collection builds and walks its graph | 5,446 | 0.92 | 1.03 | 1.38 | 1074.8 | 3453.8 | 0.73 | 0.52 | 1.000 | 0.518 | 39.84 MiB | 164.1 MiB | | typesense | HNSW float32 (hnswlib), m=16, ef_construction=128, cosine; search k=top, ef=100; exact mode = filter ordinal:>=0 with flat_search_cutoff; index held in memory | 1,040 | 4.04 | 4.38 | 5.02 | 243.8 | 590.9 | 0.50 | 0.56 | 1.000 | 0.518 | 114.2 MiB | 0 B | | qdrant | exact scan: the builder's sink sends exact=true on every search, so no HNSW graph is used whether or not Qdrant has built one (see the index state) | 8,597 | 0.88 | 0.99 | 1.38 | 1122.0 | 3291.4 | 0.74 | 0.53 | 1.000 | 0.518 | 38.17 MiB | 196.08 MiB | | oracle | HNSW in-memory neighbor graph NEIGHBORS=16 EFCONSTRUCTION=128, EFSEARCH=100 per query, cosine; exact mode = FETCH EXACT FIRST (full scan); Oracle Free caps itself at 2 CPUs (cpu_count 2 in V$PARAMETER, edition FREE in V$INSTANCE, 8 host CPUs in V$OSSTAT NUM_CPUS; the 2 CPU thread limit is Oracle's documented Free edition limit) | 2,007 | 0.90 | 1.06 | 1.79 | 1089.1 | 2896.5 | 0.87 | 0.64 | 1.000 | 0.518 | 2.17 GiB | 0 B | | weaviate | HNSW maxConnections(M)=16 efConstruction=128, ef=-1 (dynamic: limit x 8 clamped 100..500), cosine, no quantization; approximate only (no exact mode) | 998 | 5.18 | 10.4 | 11.3 | 171.6 | 312.1 | 0.57 | 0.60 | 1.000 | 0.518 | 70.7 MiB | 2.96 MiB | | sql-diskann | DiskANN (preview) via VECTOR_SEARCH, cosine, build {"StartId":"306", "L":"48", "M":"8", "R":"48"}; the exact mode scans the same table | 637 | 3.56 | 4.06 | 5.51 | 276.4 | 690.7 | 1.00 | 1.09 | 0.965 | 0.526 | 529.7 MiB | 4.52 MiB | | chroma | HNSW M=16 ef_construction=128, ef_search=100 (Chroma default), cosine; approximate only (no exact mode) | 860 | 2.03 | 2.33 | 2.54 | 486.3 | 806.9 | 0.45 | 0.46 | 1.000 | 0.518 | 46.11 MiB | 414.16 KiB | | redis | HNSW TYPE FLOAT32 M=16 EF_CONSTRUCTION=128, EF_RUNTIME=100 per query, cosine; exact mode = FLAT index built on first exact query | 11,414 | 0.37 | 0.47 | 0.68 | 2620.4 | 5132.7 | 0.55 | 0.36 | 1.000 | 0.518 | 21.46 MiB | 2.43 MiB | | vespa | HNSW float32 tensor, prenormalized-angular (cosine), max-links-per-node=16, neighbors-to-explore-at-insert=128; search targetHits=top, ef=100 via exploreAdditionalHits; exact mode = approximate:false; vectors held in memory | 288 | 1.86 | 2.26 | 2.53 | 522.7 | 1560.0 | 0.94 | 0.88 | 1.000 | 0.518 | 2.77 GiB | 6.98 MiB | | mariadb | VECTOR INDEX (HNSW variant) DISTANCE=cosine, M=16 (no ef_construction setting exists), mhnsw_ef_search=100 per statement (the ef 100 most engines here use, so the search effort matches; MariaDB's own default is 20; recall@10 at ef 100 falls as the set grows (random 1024-dimension vectors, measured 2026-10-04: 0.99 at 524, 0.89 to 0.92 at 2,000; an earlier run gave about 0.09 at 100,000)), mhnsw_max_cache_size 4G; exact mode = IGNORE INDEX full scan | 1,020 | 0.62 | 0.70 | 0.85 | 1581.5 | 5940.1 | 0.49 | 0.38 | 0.995 | 0.518 | 170.6 MiB | 22.01 MiB | | elasticsearch | HNSW float32, no quantization, m=16, ef_construction=128, cosine; search k=top, num_candidates=100; 1 shard, 0 replicas; force-merged to one segment after the load (at 1,024 dimensions a segment under 1,043 vectors gets no graph) | 543 | 1.31 | 1.49 | 1.73 | 751.0 | 2713.2 | 0.46 | 0.50 | 1.000 | 0.518 | 2.6 GiB | 2.39 MiB | | clickhouse | vector_similarity HNSW cosineDistance, quantization bf16, M=16 ef_construction=128, hnsw_candidate_list_size_for_search=256, rescoring off; exact mode = full scan with skip indexes off | 2,977 | 4.51 | 5.75 | 6.75 | 214.7 | 492.8 | 0.82 | 1.02 | 0.995 | 0.528 | 199.5 MiB | 3.26 MiB | | mongodb | vectorSearch index, HNSW maxEdges=16 numEdgeCandidates=128, float32 binData, cosine, numCandidates=20x hits (min 100); exact mode = $vectorSearch exact:true | 2,969 | 1.41 | 1.64 | 1.99 | 693.7 | 2088.1 | 0.31 | 0.29 | 1.000 | 0.518 | 656.1 MiB | 42.39 KiB | ## Details per target ### sql - Engine: Microsoft SQL Server 2025 (RTM-CU9) (KB5122048) - 17.0.5005.3 (X64), Enterprise Developer Edition (64-bit), in container gvb-mssql (image mcr.microsoft.com/mssql/server:2025-CU9-ubuntu-24.04@sha256:2b5b581621126574f3d1f75e78d3eebe8d05aedb59ad0cfdf9aa42cb0634d726; cpuset 2-3,6-7 from its creation; SQL Server counts 4 CPU(s) and runs 4 visible scheduler(s), affinity AUTO) (compose) - Index: exact VECTOR_DISTANCE cosine, no vector index (full scan) - Search settings: description=exact VECTOR_DISTANCE cosine, no vector index (full scan) - Load: 524 rows in batches of 1000, 0.9 s of upserts (575 rows/s); count matched 0.0 s after the last upsert; index step n/a - Search: 5113 latency samples, target held 524 rows, 0 errors - Client CPU per search: 1 searcher 1.18 ms, 8 searchers 1.18 ms - RAM: 530.8 MiB (docker stats); disk: 5.15 MiB (table and its indexes, reserved pages) - Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started mssql.compose.yaml with every container created on CPUs 2-3,6-7: gvb-mssql cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards. - Index after the load: ready, 0 of 524 indexed (no index used, exact scan by design: dbo.gvb_gvbbench_eshoponweb in GvbBench holds 524 rows; sys.vector_indexes lists no index on the table; indexes: PK_gvb_gvbbench_eshoponweb CLUSTERED, IX_gvb_gvbbench_eshoponweb_DocKey NONCLUSTERED; a real search under SET STATISTICS XML ON ran: Clustered Index Scan of GvbBench.gvb_gvbbench_eshoponweb through PK_gvb_gvbbench_eshoponweb (IndexKind Clustered), 524 rows read, no vector index operator, 10 hits). Durability: A commit returns after its transaction-log records are written to disk (SQL Server write-ahead logging; delayed durability is DISABLED); a database created here copies the model database: recovery model FULL, page_verify CHECKSUM; no global trace flags are enabled; mssql.conf of container gvb-mssql (/var/opt/mssql/mssql.conf, on the host at ~/gvb-data/engines/mssql/mssql.conf) does not exist, so it sets: nothing; it has no [control] or [traceflag] entry, so SQL Server's own Linux defaults for flushing writes apply (Microsoft's Linux performance guide names trace flag 3982 as that default; read from the guide, not tested here). Not tested by cutting power; whether the disk's own write cache reaches the media was not checked. Container settings from its environment (names only): MSSQL_AGENT_ENABLED, MSSQL_MEMORY_LIMIT_MB, MSSQL_PID, MSSQL_RPC_PORT, MSSQL_SA_PASSWORD.. - Load average 2.97 2.84 3.28 (1/5/15 min) when searching began. - Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@8 16,950 searches in 30.0 s, default@1 7,581 searches in 30.0 s. - Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure. default@8: warm-up 15.0 s and 8,905 searches (0 failed) at 8 searchers; QPS of the last windows 569, 562, 585 (windows of at least 2 s and 100 searches); trial of 1,780 searches in 3.0 s at 8 searchers: 592 QPS against the settled 569 QPS, 4% apart (limit 10%); default@1: warm-up 15.0 s and 3,768 searches (0 failed) at 1 searcher; p50 of the last windows 3.827, 3.847, 3.855 ms (windows of at least 2 s and 100 searches); trial of 772 searches in 3.0 s at 1 searcher: p50 3.831 ms against the settled p50 3.847 ms, 0% apart (limit 10%). - Pass default@8 after rehearsal: 2026-10-05T08:56:17.718Z to 2026-10-05T08:56:37.729Z (20.0 s), 12,285 searches, 0 failed, p50 11.764 ms, mean 13.025 ms, p99 25.690 ms, 613.9 QPS (1000/QPS 1.629 ms). - Pass default@1 after default@8: 2026-10-05T08:56:55.773Z to 2026-10-05T08:57:15.774Z (20.0 s), 5,113 searches, 0 failed, p50 3.823 ms, mean 3.909 ms, p99 5.690 ms, 255.6 QPS (1000/QPS 3.912 ms). - Passes in the order run: default@8, default@1; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 39,756 sent, 0 failed (warmupErrors). Index after the searches: ready, 0 of 524 indexed (no index used, exact scan by design: dbo.gvb_gvbbench_eshoponweb in GvbBench holds 524 rows; sys.vector_indexes lists no index on the table; indexes: PK_gvb_gvbbench_eshoponweb CLUSTERED, IX_gvb_gvbbench_eshoponweb_DocKey NONCLUSTERED; a real search under SET STATISTICS XML ON ran: Clustered Index Scan of GvbBench.gvb_gvbbench_eshoponweb through PK_gvb_gvbbench_eshoponweb (IndexKind Clustered), 524 rows read, no vector index operator, 10 hits). - Benchmark copy gvbbench_eshoponweb dropped afterwards. - Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-mssql (2f3c42731827) on network gvb-mssql_default at container-ip:1433, not through docker-proxy localhost:14330. Open after the passes: container-ip:1433 x9 (this target). - CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-mssql: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-mssql: 86 thread(s) on CPUs 2-3,6-7 (every thread of the container). - Clock per pass (median MHz of engine CPUs / client CPUs, governor, CPUs busy outside the benchmark, client CPU per search): default@8 3492/3492 MHz, performance, outside load 0.16 before the warm-up, 0.15 during, client CPU 1.181 ms per search; default@1 3503/3494 MHz, performance, outside load 0.15 before the warm-up, 0.15 during, client CPU 1.182 ms per search. ### opensearch - Engine: OpenSearch 3.9.0 (k-NN plugin, faiss) (compose) - Index: faiss HNSW float32, no compression, m=16, ef_construction=128, cosinesimil; search k=top, ef_search=100; 1 shard, 0 replicas; graph built at any segment size (approximate_threshold=0); force-merged to one segment after the load - Search settings: approximate_threshold=0, ef_construction=128, ef_search=100, k=top, m=16 - Load: 524 rows in batches of 1000, 1.2 s of upserts (421 rows/s); count matched 0.1 s after the last upsert; index step 0.9 s (force-merged to one segment, graph ready after 0.9 s: every segment searched through its HNSW graph, covering 524 of 524 vectors: 1 segment(s), 0 merge(s) running, profiled probe search: ann_search_count 1 (segments tried through a graph), exact_search_count 0 (of those, scanned instead), approximate_threshold 0; force-merged to one segment on purpose, so one graph answers every search) - Search: 9882 latency samples, target held 524 rows, 0 errors - Client CPU per search: 1 searcher 0.46 ms, 8 searchers 0.51 ms - Exact mode: 23,128 searches, p50 2.56 ms, p95 2.87 ms, recall 1.000 - RAM: 2.56 GiB (docker stats); disk: 9.55 MiB added by this load (11.71 MiB (whole engine data folder), 2.16 MiB before) - Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started opensearch.compose.yaml with every container created on CPUs 2-3,6-7: gvb-opensearch cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards. - Index after the load: ready, 524 of 524 indexed (every segment searched through its HNSW graph, covering 524 of 524 vectors: 1 segment(s), 0 merge(s) running, profiled probe search: ann_search_count 1 (segments tried through a graph), exact_search_count 0 (of those, scanned instead), approximate_threshold 0; force-merged to one segment on purpose, so one graph answers every search). Durability: Every acknowledged bulk request is fsynced to the translog before the answer: index.translog.durability=request, the OpenSearch default, which neither this sink nor opensearch.compose.yaml overrides (OpenSearchReadinessTests reads it back from the live index). By that setting a process crash or power loss loses no acknowledged write; this is read from the setting, not shown by pulling power. One node and no replicas, so a lost disk loses the data.. - Load average 2.75 2.93 3.25 (1/5/15 min) when searching began. - Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@8 27,032 searches in 30.0 s, exact 11,089 searches in 30.0 s, default@1 14,815 searches in 30.0 s. - Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure. default@8: warm-up 15.0 s and 21,928 searches (0 failed) at 8 searchers; QPS of the last windows 1,514, 1,514, 1,515 (windows of at least 2 s and 100 searches); trial of 4,545 searches in 3.0 s at 8 searchers: 1,513 QPS against the settled 1,514 QPS, 0% apart (limit 10%); exact: warm-up 15.0 s and 5,714 searches (0 failed) at 1 searcher; p50 of the last windows 2.573, 2.569, 2.560 ms (windows of at least 2 s and 100 searches); trial of 1,151 searches in 3.0 s at 1 searcher: p50 2.565 ms against the settled p50 2.569 ms, 0% apart (limit 10%); default@1: warm-up 15.0 s and 7,417 searches (0 failed) at 1 searcher; p50 of the last windows 2.001, 1.993, 2.010 ms (windows of at least 2 s and 100 searches); trial of 1,472 searches in 3.0 s at 1 searcher: p50 2.000 ms against the settled p50 2.001 ms, 0% apart (limit 10%). - Pass default@8 after rehearsal: 2026-10-05T08:59:35.153Z to 2026-10-05T08:59:55.157Z (20.0 s), 30,244 searches, 0 failed, p50 5.087 ms, mean 5.289 ms, p99 10.321 ms, 1511.9 QPS (1000/QPS 0.661 ms). - Pass exact after default@8: 2026-10-05T09:00:13.171Z to 2026-10-05T09:01:13.171Z (60.0 s), 23,128 searches, 0 failed, p50 2.561 ms, mean 2.592 ms, p99 3.506 ms, 385.5 QPS (1000/QPS 2.594 ms). - Pass default@1 after exact: 2026-10-05T09:01:31.196Z to 2026-10-05T09:01:51.198Z (20.0 s), 9,882 searches, 0 failed, p50 1.998 ms, mean 2.022 ms, p99 2.783 ms, 494.1 QPS (1000/QPS 2.024 ms). - Passes in the order run: default@8, exact, default@1; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 95,163 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (every segment searched through its HNSW graph, covering 524 of 524 vectors: 1 segment(s), 0 merge(s) running, profiled probe search: ann_search_count 1 (segments tried through a graph), exact_search_count 0 (of those, scanned instead), approximate_threshold 0; force-merged to one segment on purpose, so one graph answers every search). - Benchmark copy gvbbench_eshoponweb dropped afterwards. - Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-opensearch (70247ce66ad8) on network engines_default at container-ip:9200, not through docker-proxy localhost:9201. Open after the passes: container-ip:9200 x1 (this target). - CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-opensearch: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-opensearch: 54 thread(s) on CPUs 2-3,6-7 (every thread of the container). - Clock per pass (median MHz of engine CPUs / client CPUs, governor, CPUs busy outside the benchmark, client CPU per search): default@8 3492/3492 MHz, performance, outside load 0.16 before the warm-up, 0.13 during, client CPU 0.507 ms per search; exact 3492/3492 MHz, performance, outside load 0.14 before the warm-up, 0.15 during, client CPU 0.462 ms per search; default@1 3492/3492 MHz, performance, outside load 0.15 before the warm-up, 0.16 during, client CPU 0.456 ms per search. ### pgvector - Engine: PostgreSQL 17 + pgvector 0.8.7 (compose) - Index: HNSW vector_cosine_ops m=16 ef_construction=128, hnsw.ef_search=100 per query, float32 vector(n), cosine; exact mode = same query with index scans off (sequential scan) - Search settings: ef_construction=128, hnsw.ef_search=100, m=16 - Load: 524 rows in batches of 1000, 0.9 s of upserts (611 rows/s); count matched 0.0 s after the last upsert; index step 0.0 s (HNSW is maintained inside every insert, nothing to build; confirmed in 0.0 s: pg_indexes lists gvb_gvbbench_eshoponweb_hnsw (USING hnsw (embedding vector_cosine_ops) WITH (m='16', ef_construction='128')), pg_index valid and ready = True, EXPLAIN of the default search uses Index Scan on it = True, that search returned 10 of 10 rows; PostgreSQL keeps no entry count for an HNSW index, so indexed vectors are not reported) - Search: 24009 latency samples, target held 524 rows, 0 errors - Client CPU per search: 1 searcher 0.66 ms, 8 searchers 0.53 ms - Exact mode: 27,106 searches, p50 2.17 ms, p95 2.42 ms, recall 1.000 - RAM: 100.8 MiB (docker stats); disk: 7.58 MiB added by this load (438.66 MiB (whole engine data folder), 431.08 MiB before) - Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started pgvector.compose.yaml with every container created on CPUs 2-3,6-7: gvb-pgvector cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards. - Index after the load: ready, ? of 524 indexed (pg_indexes lists gvb_gvbbench_eshoponweb_hnsw (USING hnsw (embedding vector_cosine_ops) WITH (m='16', ef_construction='128')), pg_index valid and ready = True, EXPLAIN of the default search uses Index Scan on it = True, that search returned 10 of 10 rows; PostgreSQL keeps no entry count for an HNSW index, so indexed vectors are not reported). Durability: fsync on, synchronous_commit on, full_page_writes on, wal_sync_method fdatasync (PostgreSQL defaults; pgvector.compose.yaml sets only shared_buffers, maintenance_work_mem and max_wal_size): every commit is flushed to the write-ahead log before it returns and HNSW index changes are WAL-logged, so a crash loses no committed row (max_wal_size 4GB only spaces out checkpoints; measured with docker kill, which is SIGKILL, right after a 2,000-vector load and a restart: all 2,000 rows were there and the HNSW index was valid and used by the default search). - Load average 1.56 3.03 3.30 (1/5/15 min) when searching began. - Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@1 38,664 searches in 30.0 s, exact 13,937 searches in 30.0 s, default@8 126,604 searches in 30.0 s. - Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure. default@1: warm-up 15.0 s and 18,019 searches (0 failed) at 1 searcher; p50 of the last windows 0.810, 0.840, 0.797 ms (windows of at least 2 s and 100 searches) (its last 3 windows differed by more than 5%); trial of 3,599 searches in 3.0 s at 1 searcher: p50 0.816 ms against the settled p50 0.810 ms, 1% apart (limit 10%), so the warm-up was EXTENDED once (30.0 s and 35,790 searches (0 failed) at 1 searcher; p50 of the last windows 0.827, 0.837, 0.827 ms (windows of at least 2 s and 100 searches)); second trial of 3,566 searches in 3.0 s at 1 searcher: p50 0.822 ms against the settled p50 0.827 ms, 1% apart (limit 10%); settled after the extension; exact: warm-up 15.0 s and 6,819 searches (0 failed) at 1 searcher; p50 of the last windows 2.167, 2.161, 2.158 ms (windows of at least 2 s and 100 searches); trial of 1,370 searches in 3.0 s at 1 searcher: p50 2.151 ms against the settled p50 2.161 ms, 0% apart (limit 10%); default@8: warm-up 15.0 s and 62,144 searches (0 failed) at 8 searchers; QPS of the last windows 4,111, 4,162, 4,109 (windows of at least 2 s and 100 searches); trial of 12,448 searches in 3.0 s at 8 searchers: 4,148 QPS against the settled 4,111 QPS, 1% apart (limit 10%). - Pass default@1 after rehearsal: 2026-10-05T09:04:36.709Z to 2026-10-05T09:04:56.709Z (20.0 s), 24,009 searches, 0 failed, p50 0.817 ms, mean 0.831 ms, p99 1.221 ms, 1200.4 QPS (1000/QPS 0.833 ms). - Pass exact after default@1: 2026-10-05T09:05:14.735Z to 2026-10-05T09:06:14.737Z (60.0 s), 27,106 searches, 0 failed, p50 2.168 ms, mean 2.211 ms, p99 3.627 ms, 451.8 QPS (1000/QPS 2.214 ms). - Pass default@8 after exact: 2026-10-05T09:06:32.760Z to 2026-10-05T09:06:52.761Z (20.0 s), 82,853 searches, 0 failed, p50 1.886 ms, mean 1.929 ms, p99 3.402 ms, 4142.5 QPS (1000/QPS 0.241 ms). - Passes in the order run: default@1, exact, default@8; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 340,798 sent, 0 failed (warmupErrors). Index after the searches: ready, ? of 524 indexed (pg_indexes lists gvb_gvbbench_eshoponweb_hnsw (USING hnsw (embedding vector_cosine_ops) WITH (m='16', ef_construction='128')), pg_index valid and ready = True, EXPLAIN of the default search uses Index Scan on it = True, that search returned 10 of 10 rows; PostgreSQL keeps no entry count for an HNSW index, so indexed vectors are not reported). - Benchmark copy gvbbench_eshoponweb dropped afterwards. - Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-pgvector (92697c1224d2) on network gvb-pgvector_default at container-ip:5432, not through docker-proxy localhost:5432. Open after the passes: container-ip:5432 x8 (this target). - CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-pgvector: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-pgvector: 6 thread(s) on CPUs 2-3,6-7 (every thread of the container). - Clock per pass (median MHz of engine CPUs / client CPUs, governor, CPUs busy outside the benchmark, client CPU per search): default@1 3492/3492 MHz, performance, outside load 0.14 before the warm-up, 0.14 during, client CPU 0.665 ms per search; exact 3492/3492 MHz, performance, outside load 0.14 before the warm-up, 0.14 during, client CPU 0.684 ms per search; default@8 3492/3492 MHz, performance, outside load 0.14 before the warm-up, 0.14 during, client CPU 0.531 ms per search. ### duckdb - Engine: DuckDB 1.5.6 + vss b833341 (HNSW, embedded, in process) (embedded) - Index: HNSW (vss extension) FLOAT[n] metric=cosine m=16 ef_construction=128, ef_search=100 per connection, persistent (hnsw_enable_experimental_persistence=true, checkpoint_threshold=256MB); exact mode = array_cosine_similarity sequential scan; score = 1 - cosine distance; searches run concurrently, one connection per searcher (opened as searchers arrive, at most 32), writes run one at a time and never overlap a search - Search settings: checkpoint_threshold=256MB, ef_construction=128, ef_search=100, hnsw_enable_experimental_persistence=true, m=16, metric=cosine - Load: 524 rows in batches of 1000, 0.3 s of upserts (1,559 rows/s); count matched 0.0 s after the last upsert; index step 0.0 s (HNSW is maintained inside every transaction, nothing to build; confirmed in 0.01 s: duckdb_indexes() lists gvb_gvbbench_eshoponweb_hnsw, pragma_hnsw_index_info() counts 524 vectors of 524 rows, EXPLAIN of the default search shows HNSW_INDEX_SCAN on it = True) - Search: 5618 latency samples, target held 524 rows, 0 errors - Client CPU per search: 1 searcher 4.84 ms, 8 searchers 6.64 ms - Exact mode: 14,805 searches, p50 3.99 ms, p95 4.65 ms, recall 1.000 - RAM: in this process, not measured; disk: 2.9 MiB added by this load (413.83 MiB (engine data folder), 410.93 MiB before) - Index after the load: ready, 524 of 524 indexed (duckdb_indexes() lists gvb_gvbbench_eshoponweb_hnsw, pragma_hnsw_index_info() counts 524 vectors of 524 rows, EXPLAIN of the default search shows HNSW_INDEX_SCAN on it = True). Durability: DuckDB's write-ahead log is fsynced at every commit and no DuckDbSink setting changes that (checkpoint_threshold=256MB only spaces out the checkpoints that write the database file; measured with strace: 42 WAL fsyncs for 40 single-record commits plus 2 setup statements), so an operating-system crash or power cut loses no committed row; measured with kill -9 right after the last commit and a reopen: all 50, 300 and 2,000 rows and an HNSW index that counted the same number were recovered from the log; not tested and documented by DuckDB: the HNSW index is file-backed only through hnsw_enable_experimental_persistence = true, which this sink turns on, and WAL recovery for such custom indexes is not complete, so a crash during a checkpoint or a later commit can damage the index while the rows survive. - Load average 5.54 3.70 3.46 (1/5/15 min) when searching began. - Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: exact 7,528 searches in 30.0 s, default@8 18,134 searches in 30.0 s, default@1 8,506 searches in 30.0 s. - Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure. exact: warm-up 15.0 s and 3,665 searches (0 failed) at 1 searcher; p50 of the last windows 4.045, 3.996, 4.030 ms (windows of at least 2 s and 100 searches); trial of 747 searches in 3.0 s at 1 searcher: p50 3.922 ms against the settled p50 4.030 ms, 3% apart (limit 10%); default@8: warm-up 15.0 s and 9,003 searches (0 failed) at 8 searchers; QPS of the last windows 602, 600, 606 (windows of at least 2 s and 100 searches); trial of 1,806 searches in 3.0 s at 8 searchers: 601 QPS against the settled 602 QPS, 0% apart (limit 10%); default@1: warm-up 15.0 s and 4,240 searches (0 failed) at 1 searcher; p50 of the last windows 3.505, 3.508, 3.502 ms (windows of at least 2 s and 100 searches); trial of 847 searches in 3.0 s at 1 searcher: p50 3.515 ms against the settled p50 3.505 ms, 0% apart (limit 10%). - Pass exact after rehearsal: 2026-10-05T09:08:42.933Z to 2026-10-05T09:09:42.934Z (60.0 s), 14,805 searches, 0 failed, p50 3.985 ms, mean 4.050 ms, p99 5.472 ms, 246.7 QPS (1000/QPS 4.053 ms). - Pass default@8 after exact: 2026-10-05T09:10:00.960Z to 2026-10-05T09:10:20.965Z (20.0 s), 11,973 searches, 0 failed, p50 12.747 ms, mean 13.357 ms, p99 27.478 ms, 598.5 QPS (1000/QPS 1.671 ms). - Pass default@1 after default@8: 2026-10-05T09:10:38.974Z to 2026-10-05T09:10:58.975Z (20.0 s), 5,618 searches, 0 failed, p50 3.522 ms, mean 3.558 ms, p99 4.993 ms, 280.9 QPS (1000/QPS 3.560 ms). - Passes in the order run: exact, default@8, default@1; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 54,476 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (duckdb_indexes() lists gvb_gvbbench_eshoponweb_hnsw, pragma_hnsw_index_info() counts 524 vectors of 524 rows, EXPLAIN of the default search shows HNSW_INDEX_SCAN on it = True). - Benchmark copy gvbbench_eshoponweb dropped afterwards. - Connection: in this process (embedded), no network: no network address. No open connection seen (no open TCP connection was found). - CPU pinning: none: an embedded engine runs inside the benchmark client process, so it shares the client CPUs, CPUs 0-1,4-5. - Clock per pass (median MHz of engine CPUs / client CPUs, governor, CPUs busy outside the benchmark, client CPU per search): exact 3593/3592 MHz, performance, outside load 0.14 before the warm-up, 0.14 during, client CPU 5.973 ms per search; default@8 3593/3592 MHz, performance, outside load 0.14 before the warm-up, 0.15 during, client CPU 6.638 ms per search; default@1 3596/3592 MHz, performance, outside load 0.15 before the warm-up, 0.13 during, client CPU 4.837 ms per search. ### milvus - Engine: Milvus 2.6.25 (standalone, embedded etcd, local storage, REST API v2) (compose) - Index: HNSW M=16 efConstruction=128, ef=100, metric COSINE, Strong consistency searches; approximate only (no exact mode) - Search settings: M=16, ef=100, efConstruction=128 - Load: 524 rows in batches of 1000, 0.4 s of upserts (1,175 rows/s); count matched 1.7 s after the last upsert; index step 138.8 s (index Finished, type HNSW (COSINE, {"M":16,"efConstruction":128}), indexedRows 524 of 524 sealed rows, pendingRows 0; stored rows 524; LoadStateLoaded; query node: 1 sealed segment(s) with the index loaded covering 524 rows against 524 stored rows, 0 sealed without it, 0 rows in growing segments; datacoord: 1 live segment(s), 1 flushed and indexed covering 524 rows, 0 delete-log (L0) segment(s) waiting for a compaction, 3 segment(s) compacted away so far; flush, index build and load took 14.7 s, then compaction: 1 job(s) accepted by Milvus, layout ready and unchanged for 75 s, 124.1 s for the whole step; search ledger started, later state reads count the searches against the query node's own counters) - Search: 8602 latency samples, target held 524 rows, 0 errors - Client CPU per search: 1 searcher 0.55 ms, 8 searchers 0.63 ms - RAM: 203.6 MiB (docker stats); disk: 10.87 MiB added by this load (184.52 MiB (whole engine data folder), 173.65 MiB before) - Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started milvus.compose.yaml with every container created on CPUs 2-3,6-7: gvb-milvus cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards. - Index after the load: ready, 524 of 524 indexed (index Finished, type HNSW (COSINE, {"M":16,"efConstruction":128}), indexedRows 524 of 524 sealed rows, pendingRows 0; stored rows 524; LoadStateLoaded; query node: 1 sealed segment(s) with the index loaded covering 524 rows against 524 stored rows, 0 sealed without it, 0 rows in growing segments; datacoord: 1 live segment(s), 1 flushed and indexed covering 524 rows, 0 delete-log (L0) segment(s) waiting for a compaction, 3 segment(s) compacted away so far; search ledger since the index finished: the sink sent 0 searches (searchParams.params.ef=100, Strong consistency), the query node counted 0 search requests for this collection, segment searches Sealed +0 and Growing +0, the same sealed segments with the same loaded index builds at both readings, segments compacted away +0). Durability: Writes are not fsynced. Standalone uses the default message queue rocksmq (RocksDB, /var/lib/milvus/rdb_data, mq.type default) and flushed segments go to local-disk object storage (COMMON_STORAGETYPE=local in milvus.compose.yaml). Measured 2026-10-04 with strace over 30 acknowledged single-row upserts and one flush: fdatasync ran only on the embedded etcd files (member/wal, member/snap/db), never on rdb_data or the segment files. A host power loss or kernel crash can lose acknowledged rows that are still in the OS page cache; a Milvus process crash alone should not (inferred, not tested). Metadata in embedded etcd is fdatasynced on every commit.. - Load average 0.57 2.80 3.29 (1/5/15 min) when searching began. - Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@1 12,841 searches in 30.0 s, default@8 38,524 searches in 30.0 s. - Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure. default@1: warm-up 15.0 s and 6,428 searches (0 failed) at 1 searcher; p50 of the last windows 2.217, 2.216, 2.207 ms (windows of at least 2 s and 100 searches); trial of 1,288 searches in 3.0 s at 1 searcher: p50 2.218 ms against the settled p50 2.216 ms, 0% apart (limit 10%); default@8: warm-up 15.0 s and 19,198 searches (0 failed) at 8 searchers; QPS of the last windows 1,282, 1,288, 1,283 (windows of at least 2 s and 100 searches); trial of 3,860 searches in 3.0 s at 8 searchers: 1,285 QPS against the settled 1,283 QPS, 0% apart (limit 10%). - Pass default@1 after rehearsal: 2026-10-05T09:14:47.283Z to 2026-10-05T09:15:07.284Z (20.0 s), 8,602 searches, 0 failed, p50 2.218 ms, mean 2.323 ms, p99 3.650 ms, 430.1 QPS (1000/QPS 2.325 ms). - Pass default@8 after default@1: 2026-10-05T09:15:25.300Z to 2026-10-05T09:15:45.303Z (20.0 s), 25,635 searches, 0 failed, p50 5.693 ms, mean 6.239 ms, p99 14.522 ms, 1281.5 QPS (1000/QPS 0.780 ms). - Passes in the order run: default@1, default@8; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 82,139 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (index Finished, type HNSW (COSINE, {"M":16,"efConstruction":128}), indexedRows 524 of 524 sealed rows, pendingRows 0; stored rows 524; LoadStateLoaded; query node: 1 sealed segment(s) with the index loaded covering 524 rows against 524 stored rows, 0 sealed without it, 0 rows in growing segments; datacoord: 1 live segment(s), 1 flushed and indexed covering 524 rows, 0 delete-log (L0) segment(s) waiting for a compaction, 3 segment(s) compacted away so far; search ledger since the index finished: the sink sent 116376 searches (searchParams.params.ef=100, Strong consistency), the query node counted 116376 search requests for this collection, segment searches Sealed +114365 and Growing +0, the same sealed segments with the same loaded index builds at both readings, segments compacted away +0). - Benchmark copy gvbbench_eshoponweb dropped afterwards. - Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-milvus (c04bb361e643) on network gvb-milvus_default at container-ip:19530, not through docker-proxy localhost:19530; container gvb-milvus (c04bb361e643) on network gvb-milvus_default at container-ip:9091, not through docker-proxy localhost:9091. Open after the passes: container-ip:19530 x8 (this target). - CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-milvus: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-milvus: 24 thread(s) on CPUs 2-3,6-7 (every thread of the container). - Clock per pass (median MHz of engine CPUs / client CPUs, governor, CPUs busy outside the benchmark, client CPU per search): default@1 3541/3503 MHz, performance, outside load 0.15 before the warm-up, 0.13 during, client CPU 0.547 ms per search; default@8 3495/3495 MHz, performance, outside load 0.14 before the warm-up, 0.19 during, client CPU 0.630 ms per search. ### sqlitevec - Engine: SQLite 3.53.3 + sqlite-vec 0.1.7-alpha.2.1 (vec0, embedded, in process) (embedded) - Index: vec0 brute-force scan, no ANN index (exact), float32, cosine distance, default chunk_size=1024; score = 1 - cosine distance; searches run concurrently, one WAL reader connection per searcher (opened as searchers arrive, at most 32), writes run one at a time and may overlap searches - Search settings: chunk_size=1024 - Load: 524 rows in batches of 1000, 0.1 s of upserts (4,690 rows/s); count matched 0.0 s after the last upsert; index step 0.0 s (nothing to build, every search scans all vectors; checked in 0.00 s: no index, exact scan by design: gvb_gvbbench_eshoponweb is a vec0 virtual table holding 524 vectors and every search compares all of them (EXPLAIN QUERY PLAN: SCAN gvb_gvbbench_eshoponweb VIRTUAL TABLE INDEX 0:3{___}___ | USE TEMP B-TREE FOR ORDER BY)) - Search: 10597 latency samples, target held 524 rows, 0 errors - Client CPU per search: 1 searcher 1.97 ms, 8 searchers 3.47 ms - Exact mode: 31,923 searches, p50 1.85 ms, p95 2.08 ms, recall 1.000 - RAM: in this process, not measured; disk: 9.72 MiB added by this load (9.72 MiB (engine data folder), 0 B before) - Index after the load: ready, ? of 524 indexed (no index, exact scan by design: gvb_gvbbench_eshoponweb is a vec0 virtual table holding 524 vectors and every search compares all of them (EXPLAIN QUERY PLAN: SCAN gvb_gvbbench_eshoponweb VIRTUAL TABLE INDEX 0:3{___}___ | USE TEMP B-TREE FOR ORDER BY)). Durability: PRAGMA journal_mode=WAL and PRAGMA synchronous=NORMAL (set by this sink when it opens the file): each commit is appended to the -wal file and handed to the operating system, but the WAL is fsynced only when SQLite checkpoints it (measured with strace: 40 single-record commits caused 3 WAL fsyncs), so a crash of the process loses nothing (measured with kill -9 and a reopen: all 50, 300 and 2,000 rows were there), while an operating-system crash or power cut can lose the newest commits since the last checkpoint and leaves the database consistent. - Load average 5.59 3.68 3.53 (1/5/15 min) when searching began. - Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: exact 16,003 searches in 30.0 s, default@8 34,541 searches in 30.0 s, default@1 15,926 searches in 30.0 s. - Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure. exact: warm-up 15.0 s and 7,977 searches (0 failed) at 1 searcher; p50 of the last windows 1.843, 1.847, 1.830 ms (windows of at least 2 s and 100 searches); trial of 1,578 searches in 3.0 s at 1 searcher: p50 1.856 ms against the settled p50 1.843 ms, 1% apart (limit 10%); default@8: warm-up 15.0 s and 17,245 searches (0 failed) at 8 searchers; QPS of the last windows 1,146, 1,143, 1,152 (windows of at least 2 s and 100 searches); trial of 3,455 searches in 3.0 s at 8 searchers: 1,151 QPS against the settled 1,146 QPS, 0% apart (limit 10%); default@1: warm-up 15.0 s and 8,006 searches (0 failed) at 1 searcher; p50 of the last windows 1.851, 1.843, 1.854 ms (windows of at least 2 s and 100 searches); trial of 1,575 searches in 3.0 s at 1 searcher: p50 1.856 ms against the settled p50 1.851 ms, 0% apart (limit 10%). - Pass exact after rehearsal: 2026-10-05T09:17:37.064Z to 2026-10-05T09:18:37.064Z (60.0 s), 31,923 searches, 0 failed, p50 1.850 ms, mean 1.877 ms, p99 2.273 ms, 532.0 QPS (1000/QPS 1.880 ms). - Pass default@8 after exact: 2026-10-05T09:18:55.098Z to 2026-10-05T09:19:15.101Z (20.0 s), 23,017 searches, 0 failed, p50 6.764 ms, mean 6.941 ms, p99 13.335 ms, 1150.7 QPS (1000/QPS 0.869 ms). - Pass default@1 after default@8: 2026-10-05T09:19:33.109Z to 2026-10-05T09:19:53.110Z (20.0 s), 10,597 searches, 0 failed, p50 1.847 ms, mean 1.885 ms, p99 2.451 ms, 529.8 QPS (1000/QPS 1.887 ms). - Passes in the order run: exact, default@8, default@1; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 106,306 sent, 0 failed (warmupErrors). Index after the searches: ready, ? of 524 indexed (no index, exact scan by design: gvb_gvbbench_eshoponweb is a vec0 virtual table holding 524 vectors and every search compares all of them (EXPLAIN QUERY PLAN: SCAN gvb_gvbbench_eshoponweb VIRTUAL TABLE INDEX 0:3{___}___ | USE TEMP B-TREE FOR ORDER BY)). - Benchmark copy gvbbench_eshoponweb dropped afterwards. - Connection: in this process (embedded), no network: no network address. No open connection seen (no open TCP connection was found). - CPU pinning: none: an embedded engine runs inside the benchmark client process, so it shares the client CPUs, CPUs 0-1,4-5. - Clock per pass (median MHz of engine CPUs / client CPUs, governor, CPUs busy outside the benchmark, client CPU per search): exact 3592/3592 MHz, performance, outside load 0.16 before the warm-up, 0.17 during, client CPU 1.971 ms per search; default@8 3593/3592 MHz, performance, outside load 0.17 before the warm-up, 0.16 during, client CPU 3.466 ms per search; default@1 3594/3592 MHz, performance, outside load 0.16 before the warm-up, 0.22 during, client CPU 1.974 ms per search. ### qdrant-hnsw - Engine: Qdrant 1.17.0 in container gvb-qdrant (image qdrant/qdrant:v1.17.0@sha256:f1c7272cdac52b38c1a0e89313922d940ba50afd90d593a1605dbbc214e66ffb; cpuset 2-3,6-7 from its creation; 3 search thread(s)) (compose) - Index: HNSW m=16 ef_construct=100, hnsw_ef=server default, cosine; indexing_threshold_kb 1 and full_scan_threshold_kb 10 (server defaults are 10,000 each) so a small collection builds and walks its graph - Search settings: ef_construct=100, hnsw_ef=server, m=16 - Load: 524 rows in batches of 1000, 0.1 s of upserts (5,446 rows/s); count matched 0.0 s after the last upsert; index step 2.0 s (indexing_threshold_kb 1, full_scan_threshold_kb 10; status Green, 524 of 524 vectors in HNSW segments (2 segments), waited 2.0 s) - Search: 21497 latency samples, target held 524 rows, 0 errors - Client CPU per search: 1 searcher 0.73 ms, 8 searchers 0.52 ms - Exact mode: 65,992 searches, p50 0.90 ms, p95 1.01 ms, recall 1.000 - RAM: 39.84 MiB (docker stats); disk: 164.1 MiB (collection folder) - Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started qdrant.compose.yaml with every container created on CPUs 2-3,6-7: gvb-qdrant cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards. - Index after the load: ready, 524 of 524 indexed (status Green, optimizer ok, indexed_vectors_count 524 of 524 points, 2 segments, indexing_threshold_kb 1, full_scan_threshold_kb 10; every one of 1 non-empty segments has a complete HNSW graph above its full-scan threshold; no search has run yet, so the walk is proven by the segment settings only; searches counted since the collection was created: unfiltered_hnsw 0, unfiltered_plain 0, unfiltered_exact 0). Durability: What was measured (strace -f on the Qdrant server process, 30 single-point REST upserts per setting, every sync call timed against the request that caused it): with wait=true each upsert had exactly one msync(MS_SYNC) of the write-ahead-log segment inside the request, before the reply (30 of 30); with wait=false the 30 upserts were acknowledged within 0.26 s with no sync call, and the first WAL msync (one call covering all 30 records) came 2.7 s after the last reply, followed by the segment-file flushes. So the log is flushed to disk under both settings, but with wait=false the flush comes after the acknowledgement: an operating-system crash or power cut in that gap loses acknowledged writes, while a crash of the Qdrant process alone should not, because the bytes are already in the kernel's page cache (inferred, not tested). Every upsert here is sent with wait=true in batches of 256 over gRPC; the strace used one point per request, so one flush per batch is inferred, not measured. Segment files are flushed every 5 s; log segments 32 MB, 0 created ahead. /qdrant/config/config.yaml in container gvb-qdrant (the image's own file; its storage keys are listed) sets: storage.collection.quantization = null, storage.collection.replication_factor = 1, storage.collection.vectors.on_disk = null, storage.collection.write_consistency_factor = 1, storage.hnsw_index.ef_construct = 100, storage.hnsw_index.full_scan_threshold_kb = 10000, storage.hnsw_index.m = 16, storage.hnsw_index.max_indexing_threads = 0, storage.hnsw_index.on_disk = false, storage.hnsw_index.payload_m = null, storage.max_collections = null, storage.node_type = Normal, storage.on_disk_payload = true, storage.optimizers.default_segment_number = 0, storage.optimizers.deleted_threshold = 0.2, storage.optimizers.flush_interval_sec = 5, storage.optimizers.indexing_threshold_kb = 10000, storage.optimizers.max_optimization_threads = null, storage.optimizers.max_segment_size_kb = null, storage.optimizers.vacuum_min_vector_number = 1000, storage.performance.max_search_threads = 0, storage.performance.optimizer_cpu_budget = 0, storage.performance.update_rate_limit = null, storage.shard_transfer_method = null, storage.snapshots_config.snapshots_storage = local, storage.snapshots_path = ./snapshots, storage.storage_path = ./storage, storage.temp_path = null, storage.update_concurrency = null, storage.wal.wal_capacity_mb = 32, storage.wal.wal_segments_ahead = 0; no QDRANT__ environment overrides in the server process. Not tested by cutting power; whether the disk's own write cache reaches the media was not checked.. - Load average 4.12 3.77 3.58 (1/5/15 min) when searching began. - Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: exact 33,345 searches in 30.0 s, default@1 32,801 searches in 30.0 s, default@8 103,933 searches in 30.0 s. - Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure. exact: warm-up 15.0 s and 16,498 searches (0 failed) at 1 searcher; p50 of the last windows 0.899, 0.897, 0.898 ms (windows of at least 2 s and 100 searches); trial of 3,299 searches in 3.0 s at 1 searcher: p50 0.900 ms against the settled p50 0.898 ms, 0% apart (limit 10%); default@1: warm-up 15.0 s and 16,120 searches (0 failed) at 1 searcher; p50 of the last windows 0.921, 0.918, 0.922 ms (windows of at least 2 s and 100 searches); trial of 3,228 searches in 3.0 s at 1 searcher: p50 0.922 ms against the settled p50 0.921 ms, 0% apart (limit 10%); default@8: warm-up 15.0 s and 51,900 searches (0 failed) at 8 searchers; QPS of the last windows 3,496, 3,465, 3,422 (windows of at least 2 s and 100 searches); trial of 10,409 searches in 3.0 s at 8 searchers: 3,468 QPS against the settled 3,465 QPS, 0% apart (limit 10%). - Pass exact after rehearsal: 2026-10-05T09:21:49.892Z to 2026-10-05T09:22:49.892Z (60.0 s), 65,992 searches, 0 failed, p50 0.898 ms, mean 0.907 ms, p99 1.397 ms, 1099.9 QPS (1000/QPS 0.909 ms). - Pass default@1 after exact: 2026-10-05T09:23:07.952Z to 2026-10-05T09:23:27.952Z (20.0 s), 21,497 searches, 0 failed, p50 0.920 ms, mean 0.929 ms, p99 1.385 ms, 1074.8 QPS (1000/QPS 0.930 ms). - Pass default@8 after default@1: 2026-10-05T09:23:45.976Z to 2026-10-05T09:24:05.977Z (20.0 s), 69,080 searches, 0 failed, p50 2.227 ms, mean 2.314 ms, p99 4.012 ms, 3453.8 QPS (1000/QPS 0.290 ms). - Passes in the order run: exact, default@1, default@8; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 271,533 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (status Green, optimizer ok, indexed_vectors_count 524 of 524 points, 2 segments, indexing_threshold_kb 1, full_scan_threshold_kb 10; every one of 1 non-empty segments has a complete HNSW graph above its full-scan threshold; 308,968 segment searches walked the graph and none scanned; searches counted since the collection was created: unfiltered_hnsw 308,968, unfiltered_plain 0, unfiltered_exact 119,134). - Benchmark copy gvbbench_eshoponweb dropped afterwards. - Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-qdrant (f46b28c18b0b) on network gvb-qdrant_default at container-ip:6334, not through docker-proxy localhost:16334; container gvb-qdrant (f46b28c18b0b) on network gvb-qdrant_default at container-ip:6333, not through docker-proxy localhost:16333. Open after the passes: container-ip:6334 x1 (this target). - CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-qdrant: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-qdrant: 24 thread(s) on CPUs 2-3,6-7 (every thread of the container). - Clock per pass (median MHz of engine CPUs / client CPUs, governor, CPUs busy outside the benchmark, client CPU per search): exact 3492/3492 MHz, performance, outside load 0.14 before the warm-up, 0.12 during, client CPU 0.734 ms per search; default@1 3492/3492 MHz, performance, outside load 0.12 before the warm-up, 0.11 during, client CPU 0.729 ms per search; default@8 3492/3492 MHz, performance, outside load 0.11 before the warm-up, 0.18 during, client CPU 0.524 ms per search. ### typesense - Engine: Typesense 30.2 (vector search) (compose) - Index: HNSW float32 (hnswlib), m=16, ef_construction=128, cosine; search k=top, ef=100; exact mode = filter ordinal:>=0 with flat_search_cutoff; index held in memory - Search settings: ef=100, ef_construction=128, k=top, m=16 - Load: 524 rows in batches of 1000, 0.5 s of upserts (1,040 rows/s); count matched 0.0 s after the last upsert; index step 0.0 s (nothing to build, Typesense inserts into the HNSW graph while writing; ready after 0.0 s: HNSW graph returns 524 of 524 vectors, no writes queued: num_documents 524, pending_write_batches 0, health ok, unfiltered vector query with k above the count returned 524; the graph is built in memory during each write, so there is no later build step) - Search: 4877 latency samples, target held 524 rows, 0 errors - Client CPU per search: 1 searcher 0.50 ms, 8 searchers 0.56 ms - Exact mode: 10,816 searches, p50 5.48 ms, p95 5.86 ms, recall 1.000 - RAM: 114.2 MiB (docker stats); disk: 0 B added by this load (238.82 MiB (whole engine data folder), 258.97 MiB before) - Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started typesense.compose.yaml with every container created on CPUs 2-3,6-7: gvb-typesense cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards. - Index after the load: ready, 524 of 524 indexed (HNSW graph returns 524 of 524 vectors, no writes queued: num_documents 524, pending_write_batches 0, health ok, unfiltered vector query with k above the count returned 524; the graph is built in memory during each write, so there is no later build step). Durability: Every acknowledged write is appended to Typesense's raft log and fsynced before the answer: braft raft_sync=true with raft_sync_policy=0 (sync immediately), read from the running container's brpc /flags page on 2026-10-04. A restart replays the log from the last snapshot (typesense.compose.yaml sets TYPESENSE_SNAPSHOT_INTERVAL_SECONDS=300) and rebuilds the in-memory HNSW graph, so by those settings a crash loses no acknowledged write, but the restart is slow after a big load. This is read from the settings, not shown by pulling power. One node and no replicas, so a lost disk loses the data.. - Load average 5.33 3.81 3.58 (1/5/15 min) when searching began. - Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@1 7,313 searches in 30.0 s, default@8 17,709 searches in 30.0 s, exact 5,408 searches in 30.0 s. - Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure. default@1: warm-up 15.0 s and 3,657 searches (0 failed) at 1 searcher; p50 of the last windows 4.049, 4.048, 4.052 ms (windows of at least 2 s and 100 searches); trial of 733 searches in 3.0 s at 1 searcher: p50 4.051 ms against the settled p50 4.049 ms, 0% apart (limit 10%); default@8: warm-up 15.0 s and 8,849 searches (0 failed) at 8 searchers; QPS of the last windows 588, 589, 592 (windows of at least 2 s and 100 searches); trial of 1,779 searches in 3.0 s at 8 searchers: 592 QPS against the settled 589 QPS, 0% apart (limit 10%); exact: warm-up 15.0 s and 2,709 searches (0 failed) at 1 searcher; p50 of the last windows 5.486, 5.479, 5.484 ms (windows of at least 2 s and 100 searches); trial of 541 searches in 3.0 s at 1 searcher: p50 5.472 ms against the settled p50 5.484 ms, 0% apart (limit 10%). - Pass default@1 after rehearsal: 2026-10-05T09:26:05.104Z to 2026-10-05T09:26:25.105Z (20.0 s), 4,877 searches, 0 failed, p50 4.044 ms, mean 4.099 ms, p99 5.018 ms, 243.8 QPS (1000/QPS 4.101 ms). - Pass default@8 after default@1: 2026-10-05T09:26:43.126Z to 2026-10-05T09:27:03.130Z (20.0 s), 11,820 searches, 0 failed, p50 13.005 ms, mean 13.536 ms, p99 25.455 ms, 590.9 QPS (1000/QPS 1.692 ms). - Pass exact after default@8: 2026-10-05T09:27:21.136Z to 2026-10-05T09:28:21.138Z (60.0 s), 10,816 searches, 0 failed, p50 5.484 ms, mean 5.545 ms, p99 6.653 ms, 180.3 QPS (1000/QPS 5.547 ms). - Passes in the order run: default@1, default@8, exact; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 48,698 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (HNSW graph returns 524 of 524 vectors, no writes queued: num_documents 524, pending_write_batches 0, health ok, unfiltered vector query with k above the count returned 524; the graph is built in memory during each write, so there is no later build step). - Benchmark copy gvbbench_eshoponweb dropped afterwards. - Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-typesense (7027e0020b32) on network engines_default at container-ip:8108, not through docker-proxy localhost:8108. Open after the passes: container-ip:8108 x1 (this target). - CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-typesense: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-typesense: 303 thread(s) on CPUs 2-3,6-7 (every thread of the container). - Clock per pass (median MHz of engine CPUs / client CPUs, governor, CPUs busy outside the benchmark, client CPU per search): default@1 3592/3592 MHz, performance, outside load 0.07 before the warm-up, 0.14 during, client CPU 0.505 ms per search; default@8 3592/3592 MHz, performance, outside load 0.15 before the warm-up, -0.02 during, client CPU 0.564 ms per search; exact 3592/3592 MHz, performance, outside load 0.04 before the warm-up, 0.16 during, client CPU 0.521 ms per search. ### qdrant - Engine: Qdrant 1.17.0 in container gvb-qdrant (image qdrant/qdrant:v1.17.0@sha256:f1c7272cdac52b38c1a0e89313922d940ba50afd90d593a1605dbbc214e66ffb; cpuset 2-3,6-7 from its creation; 3 search thread(s)) (compose) - Index: exact scan: the builder's sink sends exact=true on every search, so no HNSW graph is used whether or not Qdrant has built one (see the index state) - Search settings: exact=true - Load: 524 rows in batches of 1000, 0.1 s of upserts (8,597 rows/s); count matched 0.0 s after the last upsert; index step 1.0 s (status Green, 0 of 524 vectors in HNSW segments (2 segments), waited 1.0 s) - Search: 22441 latency samples, target held 524 rows, 0 errors - Client CPU per search: 1 searcher 0.74 ms, 8 searchers 0.53 ms - RAM: 38.17 MiB (docker stats); disk: 196.08 MiB (collection folder) - Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started qdrant.compose.yaml with every container created on CPUs 2-3,6-7: gvb-qdrant cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards. - Index after the load: ready, 0 of 524 indexed (status Green, optimizer ok, indexed_vectors_count 0 of 524 points, 2 segments, indexing_threshold_kb 10,000, full_scan_threshold_kb 10,000; no index used, exact scan by design (the builder's sink sends exact=true on every search); no graph was built; searches counted since the collection was created: unfiltered_hnsw 0, unfiltered_plain 0, unfiltered_exact 0). Durability: What was measured (strace -f on the Qdrant server process, 30 single-point REST upserts per setting, every sync call timed against the request that caused it): with wait=true each upsert had exactly one msync(MS_SYNC) of the write-ahead-log segment inside the request, before the reply (30 of 30); with wait=false the 30 upserts were acknowledged within 0.26 s with no sync call, and the first WAL msync (one call covering all 30 records) came 2.7 s after the last reply, followed by the segment-file flushes. So the log is flushed to disk under both settings, but with wait=false the flush comes after the acknowledgement: an operating-system crash or power cut in that gap loses acknowledged writes, while a crash of the Qdrant process alone should not, because the bytes are already in the kernel's page cache (inferred, not tested). Every upsert here is sent with wait=true in batches of 256 over gRPC; the strace used one point per request, so one flush per batch is inferred, not measured. Segment files are flushed every 5 s; log segments 32 MB, 0 created ahead. /qdrant/config/config.yaml in container gvb-qdrant (the image's own file; its storage keys are listed) sets: storage.collection.quantization = null, storage.collection.replication_factor = 1, storage.collection.vectors.on_disk = null, storage.collection.write_consistency_factor = 1, storage.hnsw_index.ef_construct = 100, storage.hnsw_index.full_scan_threshold_kb = 10000, storage.hnsw_index.m = 16, storage.hnsw_index.max_indexing_threads = 0, storage.hnsw_index.on_disk = false, storage.hnsw_index.payload_m = null, storage.max_collections = null, storage.node_type = Normal, storage.on_disk_payload = true, storage.optimizers.default_segment_number = 0, storage.optimizers.deleted_threshold = 0.2, storage.optimizers.flush_interval_sec = 5, storage.optimizers.indexing_threshold_kb = 10000, storage.optimizers.max_optimization_threads = null, storage.optimizers.max_segment_size_kb = null, storage.optimizers.vacuum_min_vector_number = 1000, storage.performance.max_search_threads = 0, storage.performance.optimizer_cpu_budget = 0, storage.performance.update_rate_limit = null, storage.shard_transfer_method = null, storage.snapshots_config.snapshots_storage = local, storage.snapshots_path = ./snapshots, storage.storage_path = ./storage, storage.temp_path = null, storage.update_concurrency = null, storage.wal.wal_capacity_mb = 32, storage.wal.wal_segments_ahead = 0; no QDRANT__ environment overrides in the server process. Not tested by cutting power; whether the disk's own write cache reaches the media was not checked.. - Load average 1.90 3.29 3.46 (1/5/15 min) when searching began. - Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@1 34,094 searches in 30.0 s, default@8 98,864 searches in 30.0 s. - Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure. default@1: warm-up 15.0 s and 16,795 searches (0 failed) at 1 searcher; p50 of the last windows 0.878, 0.885, 0.877 ms (windows of at least 2 s and 100 searches); trial of 3,376 searches in 3.0 s at 1 searcher: p50 0.875 ms against the settled p50 0.878 ms, 0% apart (limit 10%); default@8: warm-up 15.0 s and 49,625 searches (0 failed) at 8 searchers; QPS of the last windows 3,332, 3,304, 3,312 (windows of at least 2 s and 100 searches); trial of 9,881 searches in 3.0 s at 8 searchers: 3,292 QPS against the settled 3,312 QPS, 1% apart (limit 10%). - Pass default@1 after rehearsal: 2026-10-05T09:29:49.132Z to 2026-10-05T09:30:09.132Z (20.0 s), 22,441 searches, 0 failed, p50 0.879 ms, mean 0.889 ms, p99 1.380 ms, 1122.0 QPS (1000/QPS 0.891 ms). - Pass default@8 after default@1: 2026-10-05T09:30:27.156Z to 2026-10-05T09:30:47.158Z (20.0 s), 65,834 searches, 0 failed, p50 2.344 ms, mean 2.428 ms, p99 4.207 ms, 3291.4 QPS (1000/QPS 0.304 ms). - Passes in the order run: default@1, default@8; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 212,635 sent, 0 failed (warmupErrors). Index after the searches: ready, 0 of 524 indexed (status Green, optimizer ok, indexed_vectors_count 0 of 524 points, 2 segments, indexing_threshold_kb 10,000, full_scan_threshold_kb 10,000; no index used, exact scan by design (the builder's sink sends exact=true on every search); no graph was built; searches counted since the collection was created: unfiltered_hnsw 0, unfiltered_plain 601,820, unfiltered_exact 0). - Benchmark copy gvbbench_eshoponweb dropped afterwards. - Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-qdrant (a0fd2df401ff) on network gvb-qdrant_default at container-ip:6334, not through docker-proxy localhost:16334; container gvb-qdrant (a0fd2df401ff) on network gvb-qdrant_default at container-ip:6333, not through docker-proxy localhost:16333. Open after the passes: container-ip:6334 x1 (this target). - CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-qdrant: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-qdrant: 25 thread(s) on CPUs 2-3,6-7 (every thread of the container). - Clock per pass (median MHz of engine CPUs / client CPUs, governor, CPUs busy outside the benchmark, client CPU per search): default@1 3492/3492 MHz, performance, outside load 0.14 before the warm-up, 0.12 during, client CPU 0.736 ms per search; default@8 3492/3492 MHz, performance, outside load 0.13 before the warm-up, 0.15 during, client CPU 0.535 ms per search. ### oracle - Engine: Oracle AI Database Free 23.26 (23ai line, VECTOR FLOAT32) (compose) - Index: HNSW in-memory neighbor graph NEIGHBORS=16 EFCONSTRUCTION=128, EFSEARCH=100 per query, cosine; exact mode = FETCH EXACT FIRST (full scan); Oracle Free caps itself at 2 CPUs (cpu_count 2 in V$PARAMETER, edition FREE in V$INSTANCE, 8 host CPUs in V$OSSTAT NUM_CPUS; the 2 CPU thread limit is Oracle's documented Free edition limit) - Search settings: EFCONSTRUCTION=128, EFSEARCH=100, NEIGHBORS=16 - Load: 524 rows in batches of 1000, 0.3 s of upserts (2,007 rows/s); count matched 0.1 s after the last upsert; index step 3.6 s (HNSW graph holds 524 of 524 rows, change log waiting: 0 inserts and 0 deletes, USER_INDEXES status VALID, plan of the default search: VECTOR INDEX HNSW SCAN > TABLE ACCESS BY INDEX ROWID, index used by 0 queries so far; Oracle Free caps itself at 2 CPUs (cpu_count 2 in V$PARAMETER, edition FREE in V$INSTANCE, 8 host CPUs in V$OSSTAT NUM_CPUS; the 2 CPU thread limit is Oracle's documented Free edition limit); graph rebuilt in 2.3 s) - Search: 21783 latency samples, target held 524 rows, 0 errors - Client CPU per search: 1 searcher 0.87 ms, 8 searchers 0.64 ms - Exact mode: 24,528 searches, p50 2.38 ms, p95 2.88 ms, recall 1.000 - RAM: 2.17 GiB (docker stats); disk: 0 B added by this load (6.88 GiB (whole engine data folder), 6.88 GiB before) - Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started oracle.compose.yaml with every container created on CPUs 2-3,6-7: gvb-oracle cpuset 2-3,6-7, gvb-oracle-seed cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards. - Index after the load: ready, 524 of 524 indexed (HNSW graph holds 524 of 524 rows, change log waiting: 0 inserts and 0 deletes, USER_INDEXES status VALID, plan of the default search: VECTOR INDEX HNSW SCAN > TABLE ACCESS BY INDEX ROWID, index used by 0 queries so far; Oracle Free caps itself at 2 CPUs (cpu_count 2 in V$PARAMETER, edition FREE in V$INSTANCE, 8 host CPUs in V$OSSTAT NUM_CPUS; the 2 CPU thread limit is Oracle's documented Free edition limit)). Durability: By these settings a committed row should survive a crash or power loss. The sink uses a plain COMMIT and commit_logging, commit_wait and commit_write are unset in the database, so every commit waits for its redo to be written. Measured 2026-10-04: 50 separate client commits raised V$SYSSTAT 'redo synch writes' by 55 (the 50 commits plus the CREATE and DROP of the probe table), and the log writer and the datafile writer hold their files open with O_DSYNC (open flags 02110002, filesystemio_options none). The database runs NOARCHIVELOG (V$DATABASE.LOG_MODE), so redo serves crash recovery only and there is no point-in-time restore. The HNSW graph lives in the 768 MB vector memory pool (oracle-init/01-vector-memory.sh) and is not the durable copy; the table is. Not tested by cutting power; whether the disk's own write cache reaches the media was not checked. CPU: Oracle Free caps itself at 2 CPUs (measured 2026-10-04: cpu_count 2 in V$PARAMETER, edition FREE in V$INSTANCE, 8 host CPUs in V$OSSTAT NUM_CPUS; the 2 CPU thread limit is Oracle's documented Free edition limit; this sink reads the live value when the index state is read).. - Load average 4.25 3.86 3.65 (1/5/15 min) when searching began. - Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: exact 12,278 searches in 30.0 s, default@1 36,169 searches in 30.0 s, default@8 95,747 searches in 30.0 s. - WARNING: latency had NOT settled when timing began for default@1, default@8 (each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure). exact: warm-up 15.0 s and 5,949 searches (0 failed) at 1 searcher; p50 of the last windows 2.363, 2.376, 2.356 ms (windows of at least 2 s and 100 searches); trial of 1,229 searches in 3.0 s at 1 searcher: p50 2.387 ms against the settled p50 2.363 ms, 1% apart (limit 10%); default@1: warm-up 15.0 s and 16,397 searches (0 failed) at 1 searcher; p50 of the last windows 0.860, 0.892, 0.913 ms (windows of at least 2 s and 100 searches) (its last 3 windows differed by more than 5%); trial of 3,246 searches in 3.0 s at 1 searcher: p50 0.900 ms against the settled p50 0.892 ms, 1% apart (limit 10%), so the warm-up was EXTENDED once (30.0 s and 32,617 searches (0 failed) at 1 searcher; p50 of the last windows 0.883, 0.857, 0.926 ms (windows of at least 2 s and 100 searches), its last 3 windows differed by more than 5%); second trial of 3,341 searches in 3.0 s at 1 searcher: p50 0.874 ms against the settled p50 0.883 ms, 1% apart (limit 10%); still disagreeing after the extension; default@8: warm-up 15.0 s and 48,930 searches (0 failed) at 8 searchers; QPS of the last windows 3,514, 2,364, 3,100 (windows of at least 2 s and 100 searches) (its last 3 windows differed by more than 5%); trial of 7,829 searches in 3.1 s at 8 searchers: 2,519 QPS against the settled 3,100 QPS, 23% apart (limit 10%), so the warm-up was EXTENDED once (120.0 s and 315,148 searches (0 failed) at 8 searchers; QPS of the last windows 2,207, 2,452, 2,930 (windows of at least 2 s and 100 searches), the 120 s cap ran out); second trial of 7,269 searches in 3.3 s at 8 searchers: 2,230 QPS against the settled 2,452 QPS, 10% apart (limit 10%); still disagreeing after the extension. Its numbers may still include warm-up; rerun before quoting them. - Pass exact after rehearsal: 2026-10-05T09:32:54.351Z to 2026-10-05T09:33:54.353Z (60.0 s), 24,528 searches, 0 failed, p50 2.380 ms, mean 2.444 ms, p99 3.841 ms, 408.8 QPS (1000/QPS 2.446 ms). - Pass default@1 after exact: 2026-10-05T09:35:00.380Z to 2026-10-05T09:35:20.380Z (20.0 s), 21,783 searches, 0 failed, p50 0.896 ms, mean 0.916 ms, p99 1.789 ms, 1089.1 QPS (1000/QPS 0.918 ms). - Pass default@8 after default@1: 2026-10-05T09:37:56.991Z to 2026-10-05T09:38:17.357Z (20.4 s), 58,993 searches, 0 failed, p50 1.592 ms, mean 2.759 ms, p99 3.435 ms, 2896.5 QPS (1000/QPS 0.345 ms). - Passes in the order run: exact, default@1, default@8; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 645,981 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (HNSW graph holds 524 of 524 rows, change log waiting: 0 inserts and 0 deletes, USER_INDEXES status VALID, plan of the default search: VECTOR INDEX HNSW SCAN > TABLE ACCESS BY INDEX ROWID, index used by 707301 queries so far; Oracle Free caps itself at 2 CPUs (cpu_count 2 in V$PARAMETER, edition FREE in V$INSTANCE, 8 host CPUs in V$OSSTAT NUM_CPUS; the 2 CPU thread limit is Oracle's documented Free edition limit)). - Benchmark copy gvbbench_eshoponweb dropped afterwards. - Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-oracle (dc5ec01531ae) on network gvb-oracle_default at container-ip:1521, not through docker-proxy localhost:1521. Open after the passes: container-ip:1521 x14 (this target). - CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-oracle: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-oracle: 86 thread(s) on CPUs 2-3,6-7 (every thread of the container). - Clock per pass (median MHz of engine CPUs / client CPUs, governor, CPUs busy outside the benchmark, client CPU per search): exact 3492/3492 MHz, performance, outside load 0.16 before the warm-up, 0.15 during, client CPU 0.933 ms per search; default@1 3492/3492 MHz, performance, outside load 0.16 before the warm-up, 0.16 during, client CPU 0.868 ms per search; default@8 3492/3492 MHz, performance, outside load 0.16 before the warm-up, 0.17 during, client CPU 0.644 ms per search. ### weaviate - Engine: Weaviate 1.39.8 (single node, REST + GraphQL) (compose) - Index: HNSW maxConnections(M)=16 efConstruction=128, ef=-1 (dynamic: limit x 8 clamped 100..500), cosine, no quantization; approximate only (no exact mode) - Search settings: ef=-1, efConstruction=128 - Load: 524 rows in batches of 1000, 0.5 s of upserts (998 rows/s); count matched 0.0 s after the last upsert; index step 0.0 s (HNSW is updated inside every batch, nothing to build; confirmed in 0.0 s: shard VyfgMkQG0NE1 vectorIndexingStatus READY, vectorQueueLength 0, node status objectCount 0 (refreshed only when the memtable is flushed, so it lags a minute and does not decide readiness), schema shard VyfgMkQG0NE1 status READY, Aggregate count 524, Aggregate nearVector with objectLimit 524 reached 524, Weaviate reports no count of vectors in the HNSW index, so indexed vectors are not reported) - Search: 3432 latency samples, target held 524 rows, 0 errors - Client CPU per search: 1 searcher 0.57 ms, 8 searchers 0.60 ms - RAM: 70.7 MiB (docker stats); disk: 2.96 MiB added by this load (5.02 MiB (whole engine data folder), 2.06 MiB before) - Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started weaviate.compose.yaml with every container created on CPUs 2-3,6-7: gvb-weaviate cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards. - Index after the load: ready, ? of 524 indexed (shard VyfgMkQG0NE1 vectorIndexingStatus READY, vectorQueueLength 0, node status objectCount 0 (refreshed only when the memtable is flushed, so it lags a minute and does not decide readiness), schema shard VyfgMkQG0NE1 status READY, Aggregate count 524, Aggregate nearVector with objectLimit 524 reached 524, Weaviate reports no count of vectors in the HNSW index, so indexed vectors are not reported). Durability: Weaviate 1.39.8 defaults, weaviate.compose.yaml sets no persistence variable: every object is appended to the LSM write-ahead log with a plain write and no fsync, and the log is fsynced only when its memtable is flushed, 60 seconds after the last write (PERSISTENCE_MEMTABLES_FLUSH_IDLE_AFTER_SECONDS default; measured with strace: 2,080 writes into the objects log during a 2,000-object load, the first fsync 60 s after the last write), so a power cut can lose the last minute of writes; the HNSW commit log is buffered inside the process (77 writes for 2,000 vectors), so a killed process loses the newest vectors from the vector index while their objects survive (measured with docker kill, which is SIGKILL, one second after the load and a restart, two runs each: 0 to 1 of 50, 264 to 267 of 300 and 1,992 to 1,997 of 2,000 stored objects were still reachable through a vector search). - Load average 4.98 4.11 3.66 (1/5/15 min) when searching began. - Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@8 9,494 searches in 30.0 s, default@1 5,201 searches in 30.0 s. - Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure. default@8: warm-up 15.0 s and 4,703 searches (0 failed) at 8 searchers; QPS of the last windows 312, 312, 312 (windows of at least 2 s and 100 searches); trial of 940 searches in 3.0 s at 8 searchers: 312 QPS against the settled 312 QPS, 0% apart (limit 10%); default@1: warm-up 15.0 s and 2,566 searches (0 failed) at 1 searcher; p50 of the last windows 5.165, 5.207, 5.178 ms (windows of at least 2 s and 100 searches); trial of 516 searches in 3.0 s at 1 searcher: p50 5.182 ms against the settled p50 5.178 ms, 0% apart (limit 10%). - Pass default@8 after rehearsal: 2026-10-05T09:40:02.675Z to 2026-10-05T09:40:22.701Z (20.0 s), 6,251 searches, 0 failed, p50 24.362 ms, mean 25.621 ms, p99 53.759 ms, 312.1 QPS (1000/QPS 3.204 ms). - Pass default@1 after default@8: 2026-10-05T09:40:40.712Z to 2026-10-05T09:41:00.716Z (20.0 s), 3,432 searches, 0 failed, p50 5.185 ms, mean 5.826 ms, p99 11.291 ms, 171.6 QPS (1000/QPS 5.829 ms). - Passes in the order run: default@8, default@1; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 23,420 sent, 0 failed (warmupErrors). Index after the searches: ready, ? of 524 indexed (shard VyfgMkQG0NE1 vectorIndexingStatus READY, vectorQueueLength 0, node status objectCount 524 (refreshed only when the memtable is flushed, so it lags a minute and does not decide readiness), schema shard VyfgMkQG0NE1 status READY, Aggregate count 524, Aggregate nearVector with objectLimit 524 reached 524, Weaviate reports no count of vectors in the HNSW index, so indexed vectors are not reported). - Benchmark copy gvbbench_eshoponweb dropped afterwards. - Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-weaviate (caf050903bb9) on network gvb-weaviate_default at container-ip:8080, not through docker-proxy localhost:8085. Open after the passes: container-ip:8080 x1 (this target). - CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-weaviate: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-weaviate: 9 thread(s) on CPUs 2-3,6-7 (every thread of the container). - Clock per pass (median MHz of engine CPUs / client CPUs, governor, CPUs busy outside the benchmark, client CPU per search): default@8 3574/3567 MHz, performance, outside load 0.15 before the warm-up, 0.15 during, client CPU 0.603 ms per search; default@1 3580/3548 MHz, performance, outside load 0.15 before the warm-up, 0.15 during, client CPU 0.573 ms per search. ### sql-diskann - Engine: Microsoft SQL Server 2025 (RTM-CU9) (KB5122048) - 17.0.5005.3 (X64), Enterprise Developer Edition (64-bit), in container gvb-mssql (image mcr.microsoft.com/mssql/server:2025-CU9-ubuntu-24.04@sha256:2b5b581621126574f3d1f75e78d3eebe8d05aedb59ad0cfdf9aa42cb0634d726; cpuset 2-3,6-7 from its creation; SQL Server counts 4 CPU(s) and runs 4 visible scheduler(s), affinity AUTO) (compose) - Index: DiskANN (preview) via VECTOR_SEARCH, cosine, build {"StartId":"306", "L":"48", "M":"8", "R":"48"}; the exact mode scans the same table - Search settings: L=48, M=8, R=48, StartId=306 - Load: 524 rows in batches of 1000, 0.8 s of upserts (637 rows/s); count matched 0.0 s after the last upsert; index step 3.9 s (copied 524 rows to dbo.gvb_gvbbench_eshoponweb_ann (INT key) and built DiskANN, parameters {"StartId":"306", "L":"48", "M":"8", "R":"48"}, graph covers 524 rows) - Search: 5528 latency samples, target held 524 rows, 0 errors - Client CPU per search: 1 searcher 1.00 ms, 8 searchers 1.09 ms - Exact mode: 16,405 searches, p50 3.56 ms, p95 4.12 ms, recall 1.000 - RAM: 529.7 MiB (docker stats); disk: 4.52 MiB (table and its indexes, reserved pages) - Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started mssql.compose.yaml with every container created on CPUs 2-3,6-7: gvb-mssql cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards. - Index after the load: ready, 524 of 524 indexed (DiskANN index built and used. sys.vector_indexes: vix_gvb_gvbbench_eshoponweb on dbo.gvb_gvbbench_eshoponweb_ann in GvbBenchDiskAnn, DiskANN, COSINE, enabled, build {"StartId":"306", "L":"48", "M":"8", "R":"48"}; graph table rows 524 of 524 table rows. Real VECTOR_SEARCH plan: Vector Index Seek on index vix_gvb_gvbbench_eshoponweb of GvbBenchDiskAnn.gvb_gvbbench_eshoponweb_ann (IndexKind DiskANN), 10 rows returned, 10 hits. The exact mode searches the SAME table (dbo.gvb_gvbbench_eshoponweb_ann, database GvbBenchDiskAnn) and its real plan is: Clustered Index Scan of GvbBenchDiskAnn.gvb_gvbbench_eshoponweb_ann through PK_gvb_gvbbench_eshoponweb_ann (IndexKind Clustered), 524 rows read, no vector index operator, 10 hits). Durability: A commit returns after its transaction-log records are written to disk (SQL Server write-ahead logging; delayed durability is DISABLED); a database created here copies the model database: recovery model FULL, page_verify CHECKSUM; no global trace flags are enabled; mssql.conf of container gvb-mssql (/var/opt/mssql/mssql.conf, on the host at ~/gvb-data/engines/mssql/mssql.conf) does not exist, so it sets: nothing; it has no [control] or [traceflag] entry, so SQL Server's own Linux defaults for flushing writes apply (Microsoft's Linux performance guide names trace flag 3982 as that default; read from the guide, not tested here). Not tested by cutting power; whether the disk's own write cache reaches the media was not checked. Container settings from its environment (names only): MSSQL_AGENT_ENABLED, MSSQL_MEMORY_LIMIT_MB, MSSQL_PID, MSSQL_RPC_PORT, MSSQL_SA_PASSWORD.. - Load average 3.17 3.78 3.61 (1/5/15 min) when searching began. - Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: exact 8,108 searches in 30.0 s, default@8 20,452 searches in 30.0 s, default@1 8,354 searches in 30.0 s. - Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure. exact: warm-up 15.0 s and 4,096 searches (0 failed) at 1 searcher; p50 of the last windows 3.599, 3.584, 3.574 ms (windows of at least 2 s and 100 searches); trial of 827 searches in 3.0 s at 1 searcher: p50 3.567 ms against the settled p50 3.584 ms, 0% apart (limit 10%); default@8: warm-up 15.0 s and 10,379 searches (0 failed) at 8 searchers; QPS of the last windows 676, 700, 698 (windows of at least 2 s and 100 searches); trial of 2,120 searches in 3.0 s at 8 searchers: 704 QPS against the settled 698 QPS, 1% apart (limit 10%); default@1: warm-up 15.0 s and 4,157 searches (0 failed) at 1 searcher; p50 of the last windows 3.567, 3.569, 3.574 ms (windows of at least 2 s and 100 searches); trial of 836 searches in 3.0 s at 1 searcher: p50 3.570 ms against the settled p50 3.569 ms, 0% apart (limit 10%). - Pass exact after rehearsal: 2026-10-05T09:43:03.764Z to 2026-10-05T09:44:03.765Z (60.0 s), 16,405 searches, 0 failed, p50 3.563 ms, mean 3.654 ms, p99 5.601 ms, 273.4 QPS (1000/QPS 3.657 ms). - Pass default@8 after exact: 2026-10-05T09:44:21.796Z to 2026-10-05T09:44:41.805Z (20.0 s), 13,821 searches, 0 failed, p50 11.335 ms, mean 11.576 ms, p99 19.899 ms, 690.7 QPS (1000/QPS 1.448 ms). - Pass default@1 after default@8: 2026-10-05T09:44:59.813Z to 2026-10-05T09:45:19.814Z (20.0 s), 5,528 searches, 0 failed, p50 3.564 ms, mean 3.616 ms, p99 5.513 ms, 276.4 QPS (1000/QPS 3.618 ms). - Passes in the order run: exact, default@8, default@1; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 59,329 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (DiskANN index built and used. sys.vector_indexes: vix_gvb_gvbbench_eshoponweb on dbo.gvb_gvbbench_eshoponweb_ann in GvbBenchDiskAnn, DiskANN, COSINE, enabled, build {"StartId":"306", "L":"48", "M":"8", "R":"48"}; graph table rows 524 of 524 table rows. Real VECTOR_SEARCH plan: Vector Index Seek on index vix_gvb_gvbbench_eshoponweb of GvbBenchDiskAnn.gvb_gvbbench_eshoponweb_ann (IndexKind DiskANN), 10 rows returned, 10 hits. The exact mode searches the SAME table (dbo.gvb_gvbbench_eshoponweb_ann, database GvbBenchDiskAnn) and its real plan is: Clustered Index Scan of GvbBenchDiskAnn.gvb_gvbbench_eshoponweb_ann through PK_gvb_gvbbench_eshoponweb_ann (IndexKind Clustered), 524 rows read, no vector index operator, 10 hits). - Benchmark copy gvbbench_eshoponweb dropped afterwards. - Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-mssql (c90b9e36521f) on network gvb-mssql_default at container-ip:1433, not through docker-proxy localhost:14330. Open after the passes: container-ip:1433 x9 (this target). - CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-mssql: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-mssql: 84 thread(s) on CPUs 2-3,6-7 (every thread of the container). - Clock per pass (median MHz of engine CPUs / client CPUs, governor, CPUs busy outside the benchmark, client CPU per search): exact 3502/3494 MHz, performance, outside load 0.14 before the warm-up, 0.16 during, client CPU 1.124 ms per search; default@8 3492/3492 MHz, performance, outside load 0.16 before the warm-up, 0.13 during, client CPU 1.088 ms per search; default@1 3497/3493 MHz, performance, outside load 0.14 before the warm-up, 0.15 during, client CPU 0.998 ms per search. ### chroma - Engine: Chroma 1.4.4 (single node, REST API v2) (compose) - Index: HNSW M=16 ef_construction=128, ef_search=100 (Chroma default), cosine; approximate only (no exact mode) - Search settings: M=16, ef_construction=128, ef_search=100 - Load: 524 rows in batches of 1000, 0.6 s of upserts (860 rows/s); count matched 0.0 s after the last upsert; index step 0.0 s (Chroma has no index build to wait for and exposes no index status; checked in 0.0 s: Chroma exposes no index status (indexing_status endpoint: HTTP 500 {"error":"InternalError","message":"Method scout_logs is not implemented"}), so indexed vectors are not reported and ready only means the two counts agree to within 1 percent: count endpoint 524, a vector search asking for 524 results returned 524) - Search: 9726 latency samples, target held 524 rows, 0 errors - Client CPU per search: 1 searcher 0.45 ms, 8 searchers 0.46 ms - RAM: 46.11 MiB (docker stats); disk: 414.16 KiB added by this load (2.01 GiB (whole engine data folder), 2.01 GiB before) - Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started chroma.compose.yaml with every container created on CPUs 2-3,6-7: gvb-chroma cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards. - Index after the load: ready, ? of 524 indexed (Chroma exposes no index status (indexing_status endpoint: HTTP 500 {"error":"InternalError","message":"Method scout_logs is not implemented"}), so indexed vectors are not reported and ready only means the two counts agree to within 1 percent: count endpoint 524, a vector search asking for 524 results returned 524). Durability: SQLite rollback journal with its default synchronous=FULL under the data directory (chroma.compose.yaml sets IS_PERSISTENT=1 and PERSIST_DIRECTORY=/data and no sync setting): every write is committed to chroma.sqlite3 with fsync of the journal, the directory and the database file before the call returns (strace: 15 database fsyncs and 38 journal fsyncs for one create, four 500-vector upserts and one delete), and the HNSW files are written every sync_threshold=1000 vectors and rebuilt from the SQLite log after a crash; measured with kill -9 and a restart: all 50, 300 and 2,000 vectors were still stored and searchable. - Load average 2.78 3.36 3.47 (1/5/15 min) when searching began. - Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@1 14,651 searches in 30.0 s, default@8 24,247 searches in 30.0 s. - Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure. default@1: warm-up 15.0 s and 7,303 searches (0 failed) at 1 searcher; p50 of the last windows 2.031, 2.024, 2.027 ms (windows of at least 2 s and 100 searches); trial of 1,462 searches in 3.0 s at 1 searcher: p50 2.026 ms against the settled p50 2.027 ms, 0% apart (limit 10%); default@8: warm-up 15.0 s and 12,102 searches (0 failed) at 8 searchers; QPS of the last windows 806, 804, 808 (windows of at least 2 s and 100 searches); trial of 2,441 searches in 3.0 s at 8 searchers: 812 QPS against the settled 806 QPS, 1% apart (limit 10%). - Pass default@1 after rehearsal: 2026-10-05T09:46:46.603Z to 2026-10-05T09:47:06.605Z (20.0 s), 9,726 searches, 0 failed, p50 2.032 ms, mean 2.055 ms, p99 2.541 ms, 486.3 QPS (1000/QPS 2.057 ms). - Pass default@8 after default@1: 2026-10-05T09:47:24.627Z to 2026-10-05T09:47:44.635Z (20.0 s), 16,144 searches, 0 failed, p50 9.832 ms, mean 9.911 ms, p99 12.077 ms, 806.9 QPS (1000/QPS 1.239 ms). - Passes in the order run: default@1, default@8; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 62,206 sent, 0 failed (warmupErrors). Index after the searches: ready, ? of 524 indexed (Chroma exposes no index status (indexing_status endpoint: HTTP 500 {"error":"InternalError","message":"Method scout_logs is not implemented"}), so indexed vectors are not reported and ready only means the two counts agree to within 1 percent: count endpoint 524, a vector search asking for 524 results returned 524). - Benchmark copy gvbbench_eshoponweb dropped afterwards. - Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-chroma (fa5c916b4914) on network gvb-chroma_default at container-ip:8000, not through docker-proxy localhost:8000. Open after the passes: container-ip:8000 x8 (this target). - CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-chroma: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-chroma: 9 thread(s) on CPUs 2-3,6-7 (every thread of the container). - Clock per pass (median MHz of engine CPUs / client CPUs, governor, CPUs busy outside the benchmark, client CPU per search): default@1 3592/3592 MHz, performance, outside load 0.09 before the warm-up, 0.08 during, client CPU 0.454 ms per search; default@8 3592/3592 MHz, performance, outside load 0.08 before the warm-up, 0.09 during, client CPU 0.458 ms per search. ### redis - Engine: Redis 8.10.2 (query engine, HASH + vector index) (compose) - Index: HNSW TYPE FLOAT32 M=16 EF_CONSTRUCTION=128, EF_RUNTIME=100 per query, cosine; exact mode = FLAT index built on first exact query - Search settings: EF_CONSTRUCTION=128, EF_RUNTIME=100, M=16 - Load: 524 rows in batches of 1000, 0.0 s of upserts (11,414 rows/s); count matched 0.0 s after the last upsert; index step 0.0 s (HNSW index finished after 0.0 s of waiting: FT.INFO indexing 0, percent_indexed 1, hash_indexing_failures 0, flat_buffer_size 0, num_docs 524, hashes counted with SCAN under gvb_gvbbench_eshoponweb: 524) - Search: 52409 latency samples, target held 524 rows, 0 errors - Client CPU per search: 1 searcher 0.55 ms, 8 searchers 0.36 ms - Exact mode: 168,258 searches, p50 0.35 ms, p95 0.45 ms, recall 1.000 - RAM: 21.46 MiB (docker stats); disk: 2.43 MiB added by this load (2.43 MiB (whole engine data folder), 89 B before) - Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started redis.compose.yaml with every container created on CPUs 2-3,6-7: gvb-redis cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards. - Index after the load: ready, 524 of 524 indexed (FT.INFO indexing 0, percent_indexed 1, hash_indexing_failures 0, flat_buffer_size 0, num_docs 524, hashes counted with SCAN under gvb_gvbbench_eshoponweb: 524). Durability: save "300 1" and appendonly no (redis.compose.yaml): an RDB snapshot is written every 5 minutes if at least one key changed, and on a clean stop; there is no append-only log, so a crash loses every write since the last snapshot (up to 5 minutes plus the time a snapshot takes; measured with docker kill, which is SIGKILL, right after a 2,000-vector load: none of the 2,000 hashes were there after the restart, and the restart loaded an older snapshot that still held keys of a collection that had been dropped since). - Load average 1.97 2.83 3.25 (1/5/15 min) when searching began. - Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@1 79,753 searches in 30.0 s, exact 82,291 searches in 30.0 s, default@8 150,860 searches in 30.0 s. - WARNING: latency had NOT settled when timing began for default@8 (each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure). default@1: warm-up 15.0 s and 39,692 searches (0 failed) at 1 searcher; p50 of the last windows 0.352, 0.350, 0.369 ms (windows of at least 2 s and 100 searches) (its last 3 windows differed by more than 5%); trial of 8,035 searches in 3.0 s at 1 searcher: p50 0.355 ms against the settled p50 0.352 ms, 1% apart (limit 10%), so the warm-up was EXTENDED once (30.0 s and 77,908 searches (0 failed) at 1 searcher; p50 of the last windows 0.379, 0.384, 0.376 ms (windows of at least 2 s and 100 searches)); second trial of 7,691 searches in 3.0 s at 1 searcher: p50 0.388 ms against the settled p50 0.379 ms, 3% apart (limit 10%); settled after the extension; exact: warm-up 15.0 s and 41,325 searches (0 failed) at 1 searcher; p50 of the last windows 0.367, 0.376, 0.359 ms (windows of at least 2 s and 100 searches); trial of 8,543 searches in 3.0 s at 1 searcher: p50 0.342 ms against the settled p50 0.367 ms, 7% apart (limit 10%); default@8: warm-up 15.0 s and 75,695 searches (0 failed) at 8 searchers; QPS of the last windows 4,921, 5,016, 5,277 (windows of at least 2 s and 100 searches) (its last 3 windows differed by more than 5%); trial of 15,098 searches in 3.0 s at 8 searchers: 5,030 QPS against the settled 5,016 QPS, 0% apart (limit 10%), so the warm-up was EXTENDED once (30.0 s and 153,806 searches (0 failed) at 8 searchers; QPS of the last windows 5,068, 5,484, 4,940 (windows of at least 2 s and 100 searches), its last 3 windows differed by more than 5%); second trial of 15,130 searches in 3.0 s at 8 searchers: 5,041 QPS against the settled 5,068 QPS, 1% apart (limit 10%); still disagreeing after the extension. Its numbers may still include warm-up; rerun before quoting them. - Pass default@1 after rehearsal: 2026-10-05T09:50:28.827Z to 2026-10-05T09:50:48.827Z (20.0 s), 52,409 searches, 0 failed, p50 0.373 ms, mean 0.380 ms, p99 0.676 ms, 2620.4 QPS (1000/QPS 0.382 ms). - Pass exact after default@1: 2026-10-05T09:51:06.883Z to 2026-10-05T09:52:06.883Z (60.0 s), 168,258 searches, 0 failed, p50 0.350 ms, mean 0.355 ms, p99 0.644 ms, 2804.3 QPS (1000/QPS 0.357 ms). - Pass default@8 after exact: 2026-10-05T09:53:13.042Z to 2026-10-05T09:53:33.044Z (20.0 s), 102,661 searches, 0 failed, p50 1.528 ms, mean 1.557 ms, p99 2.209 ms, 5132.7 QPS (1000/QPS 0.195 ms). - Passes in the order run: default@1, exact, default@8; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 872,560 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (FT.INFO indexing 0, percent_indexed 1, hash_indexing_failures 0, flat_buffer_size 0, num_docs 524, hashes counted with SCAN under gvb_gvbbench_eshoponweb: 524). - Benchmark copy gvbbench_eshoponweb dropped afterwards. - Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-redis (6f8328664602) on network gvb-redis_default at container-ip:6379, not through docker-proxy localhost:6379. Open after the passes: container-ip:6379 x1 (this target). - CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-redis: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-redis: 6 thread(s) on CPUs 2-3,6-7 (every thread of the container). - Clock per pass (median MHz of engine CPUs / client CPUs, governor, CPUs busy outside the benchmark, client CPU per search): default@1 3492/3492 MHz, performance, outside load 0.15 before the warm-up, 0.13 during, client CPU 0.547 ms per search; exact 3492/3492 MHz, performance, outside load 0.16 before the warm-up, 0.13 during, client CPU 0.549 ms per search; default@8 3492/3492 MHz, performance, outside load 0.13 before the warm-up, 0.13 during, client CPU 0.361 ms per search. ### vespa - Engine: Vespa 8.754.14 (tensor attribute + HNSW) (compose) - Index: HNSW float32 tensor, prenormalized-angular (cosine), max-links-per-node=16, neighbors-to-explore-at-insert=128; search targetHits=top, ef=100 via exploreAdditionalHits; exact mode = approximate:false; vectors held in memory - Search settings: ef=100, max-links-per-node=16, neighbors-to-explore-at-insert=128, targetHits=top - Load: 524 rows in batches of 1000, 1.8 s of upserts (288 rows/s); count matched 0.0 s after the last upsert; index step 0.1 s (nothing to build, Vespa inserts into the HNSW graph while writing; ready after 0.1 s: HNSW graph returns 524 of 524 vectors: totalCount 524, coverage full True, nearestNeighbor approximate:true with targetHits above the count: has_index True, algorithm "index top k", top_k_hits 524; the graph is updated while each write is applied, so there is no later build step) - Search: 10455 latency samples, target held 524 rows, 0 errors - Client CPU per search: 1 searcher 0.94 ms, 8 searchers 0.88 ms - Exact mode: 31,461 searches, p50 1.85 ms, p95 2.28 ms, recall 1.000 - RAM: 2.77 GiB (docker stats); disk: 6.98 MiB added by this load (2.12 GiB (whole engine data folder), 2.11 GiB before) - Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started vespa.compose.yaml with every container created on CPUs 2-3,6-7: gvb-vespa cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards. - Index after the load: ready, 524 of 524 indexed (HNSW graph returns 524 of 524 vectors: totalCount 524, coverage full True, nearestNeighbor approximate:true with targetHits above the count: has_index True, algorithm "index top k", top_k_hits 524; the graph is updated while each write is applied, so there is no later build step). Durability: Vespa's transaction log server fsyncs after each commit: searchlib.translogserver usefsync=true, and an operation is searchable when it is acknowledged: proton documentdb visibilitydelay=0. Neither is overridden by this sink or by the services.xml it deploys; VespaReadinessTests reads both from the live config server. After a crash the node replays the transaction log over its last flushed data and rebuilds the in-memory HNSW graph. By those settings an acknowledged write survives a crash; this is read from the settings, not shown by pulling power. One node and min-redundancy 1, so a lost disk loses the data.. - Load average 5.42 3.44 3.29 (1/5/15 min) when searching began. - Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@8 18,205 searches in 30.0 s, exact 14,779 searches in 30.0 s, default@1 15,496 searches in 30.0 s. - Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure. default@8: warm-up 15.0 s and 21,698 searches (0 failed) at 8 searchers; QPS of the last windows 1,528, 1,547, 1,553 (windows of at least 2 s and 100 searches); trial of 4,760 searches in 3.0 s at 8 searchers: 1,585 QPS against the settled 1,547 QPS, 2% apart (limit 10%); exact: warm-up 15.0 s and 7,847 searches (0 failed) at 1 searcher; p50 of the last windows 1.850, 1.864, 1.847 ms (windows of at least 2 s and 100 searches); trial of 1,569 searches in 3.0 s at 1 searcher: p50 1.852 ms against the settled p50 1.850 ms, 0% apart (limit 10%); default@1: warm-up 15.0 s and 7,823 searches (0 failed) at 1 searcher; p50 of the last windows 1.853, 1.872, 1.866 ms (windows of at least 2 s and 100 searches); trial of 1,576 searches in 3.0 s at 1 searcher: p50 1.855 ms against the settled p50 1.866 ms, 1% apart (limit 10%). - Pass default@8 after rehearsal: 2026-10-05T09:56:15.357Z to 2026-10-05T09:56:35.360Z (20.0 s), 31,205 searches, 0 failed, p50 4.974 ms, mean 5.125 ms, p99 8.636 ms, 1560.0 QPS (1000/QPS 0.641 ms). - Pass exact after default@8: 2026-10-05T09:56:53.371Z to 2026-10-05T09:57:53.371Z (60.0 s), 31,461 searches, 0 failed, p50 1.848 ms, mean 1.904 ms, p99 2.605 ms, 524.3 QPS (1000/QPS 1.907 ms). - Pass default@1 after exact: 2026-10-05T09:58:11.396Z to 2026-10-05T09:58:31.398Z (20.0 s), 10,455 searches, 0 failed, p50 1.858 ms, mean 1.910 ms, p99 2.533 ms, 522.7 QPS (1000/QPS 1.913 ms). - Passes in the order run: default@8, exact, default@1; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 93,753 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (HNSW graph returns 524 of 524 vectors: totalCount 524, coverage full True, nearestNeighbor approximate:true with targetHits above the count: has_index True, algorithm "index top k", top_k_hits 524; the graph is updated while each write is applied, so there is no later build step). - Benchmark copy gvbbench_eshoponweb dropped afterwards. - Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-vespa (5c9e5ada9c72) on network engines_default at container-ip:8080, not through docker-proxy localhost:8090; container gvb-vespa (5c9e5ada9c72) on network engines_default at container-ip:19071, not through docker-proxy localhost:19071. Open after the passes: container-ip:8080 x1 (this target). - CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-vespa: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-vespa: 213 thread(s) on CPUs 2-3,6-7 (every thread of the container). - Clock per pass (median MHz of engine CPUs / client CPUs, governor, CPUs busy outside the benchmark, client CPU per search): default@8 3492/3492 MHz, performance, outside load 0.12 before the warm-up, 0.22 during, client CPU 0.877 ms per search; exact 3521/3498 MHz, performance, outside load 0.18 before the warm-up, 0.12 during, client CPU 0.934 ms per search; default@1 3524/3499 MHz, performance, outside load 0.12 before the warm-up, 0.11 during, client CPU 0.936 ms per search. ### mariadb - Engine: MariaDB 11.8.9 (InnoDB + VECTOR INDEX) (compose) - Index: VECTOR INDEX (HNSW variant) DISTANCE=cosine, M=16 (no ef_construction setting exists), mhnsw_ef_search=100 per statement (the ef 100 most engines here use, so the search effort matches; MariaDB's own default is 20; recall@10 at ef 100 falls as the set grows (random 1024-dimension vectors, measured 2026-10-04: 0.99 at 524, 0.89 to 0.92 at 2,000; an earlier run gave about 0.09 at 100,000)), mhnsw_max_cache_size 4G; exact mode = IGNORE INDEX full scan - Search settings: DISTANCE=cosine, M=16, mhnsw_ef_search=100 - Load: 524 rows in batches of 1000, 0.5 s of upserts (1,020 rows/s); count matched 0.0 s after the last upsert; index step 0.5 s (VECTOR INDEX is maintained inside every INSERT, nothing to build; confirmed after 2 poll(s) in 0.5 s (poll 1 of 2 said: not ready: a search at mhnsw_ef_search 100 returned only 397 of 524 requested rows through the index (EXPLAIN key 'vec_idx') (information_schema lists vec_idx as INDEX_TYPE VECTOR, EXPLAIN of the default search uses key vec_idx, the server applied mhnsw_ef_search 100 (asked for 100, read back from the server), a search at that effort returned 397 of 524 requested rows through the index; MariaDB has no per-index row count (the graph is a hidden InnoDB table), so indexed vectors are not reported)): information_schema lists vec_idx as INDEX_TYPE VECTOR, EXPLAIN of the default search uses key vec_idx, the server applied mhnsw_ef_search 100 (asked for 100, read back from the server), a search at that effort returned 515 of 524 requested rows through the index; MariaDB has no per-index row count (the graph is a hidden InnoDB table), so indexed vectors are not reported) - Search: 31631 latency samples, target held 524 rows, 0 errors - Client CPU per search: 1 searcher 0.49 ms, 8 searchers 0.38 ms - Exact mode: 37,340 searches, p50 1.58 ms, p95 1.68 ms, recall 1.000 - RAM: 170.6 MiB (docker stats); disk: 22.01 MiB added by this load (182.79 MiB (whole engine data folder), 160.78 MiB before) - Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started mariadb-bench.compose.yaml with every container created on CPUs 2-3,6-7: gvbbench-mariadb cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards. - Index after the load: ready, ? of 524 indexed (information_schema lists vec_idx as INDEX_TYPE VECTOR, EXPLAIN of the default search uses key vec_idx, the server applied mhnsw_ef_search 100 (asked for 100, read back from the server), a search at that effort returned 515 of 524 requested rows through the index; MariaDB has no per-index row count (the graph is a hidden InnoDB table), so indexed vectors are not reported). Durability: innodb_flush_log_at_trx_commit=2 (set in mariadb-bench.compose.yaml for the benchmark's own container gvbbench-mariadb, and in mariadb.compose.yaml for the daily gvb-mariadb, which has the same server settings): the InnoDB redo log is written to the operating system at every commit but fsynced about once a second, so a crash of the mariadbd process loses nothing, while an operating-system crash or power cut can lose the last second of commits; innodb_doublewrite is on (default), the binary log is off, and the vector graph is an InnoDB table under the same log (settings read from the running server with SHOW VARIABLES; the crash behaviour is InnoDB's documented behaviour for this setting and was not tested on either container). - Load average 2.32 4.05 3.71 (1/5/15 min) when searching began. - Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@1 46,769 searches in 30.0 s, exact 18,604 searches in 30.0 s, default@8 179,772 searches in 30.0 s. - Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure. default@1: warm-up 15.0 s and 23,655 searches (0 failed) at 1 searcher; p50 of the last windows 0.630, 0.624, 0.609 ms (windows of at least 2 s and 100 searches); trial of 4,697 searches in 3.0 s at 1 searcher: p50 0.627 ms against the settled p50 0.624 ms, 0% apart (limit 10%); exact: warm-up 15.0 s and 9,312 searches (0 failed) at 1 searcher; p50 of the last windows 1.585, 1.583, 1.590 ms (windows of at least 2 s and 100 searches); trial of 1,863 searches in 3.0 s at 1 searcher: p50 1.589 ms against the settled p50 1.585 ms, 0% apart (limit 10%); default@8: warm-up 15.0 s and 89,243 searches (0 failed) at 8 searchers; QPS of the last windows 5,915, 5,973, 5,962 (windows of at least 2 s and 100 searches); trial of 17,911 searches in 3.0 s at 8 searchers: 5,968 QPS against the settled 5,962 QPS, 0% apart (limit 10%). - Pass default@1 after rehearsal: 2026-10-05T10:00:36.729Z to 2026-10-05T10:00:56.729Z (20.0 s), 31,631 searches, 0 failed, p50 0.622 ms, mean 0.630 ms, p99 0.854 ms, 1581.5 QPS (1000/QPS 0.632 ms). - Pass exact after default@1: 2026-10-05T10:01:14.758Z to 2026-10-05T10:02:14.759Z (60.0 s), 37,340 searches, 0 failed, p50 1.584 ms, mean 1.605 ms, p99 2.020 ms, 622.3 QPS (1000/QPS 1.607 ms). - Pass default@8 after exact: 2026-10-05T10:02:32.789Z to 2026-10-05T10:02:52.790Z (20.0 s), 118,807 searches, 0 failed, p50 1.300 ms, mean 1.345 ms, p99 2.477 ms, 5940.1 QPS (1000/QPS 0.168 ms). - Passes in the order run: default@1, exact, default@8; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 391,826 sent, 0 failed (warmupErrors). Index after the searches: ready, ? of 524 indexed (information_schema lists vec_idx as INDEX_TYPE VECTOR, EXPLAIN of the default search uses key vec_idx, the server applied mhnsw_ef_search 100 (asked for 100, read back from the server), a search at that effort returned 515 of 524 requested rows through the index; MariaDB has no per-index row count (the graph is a hidden InnoDB table), so indexed vectors are not reported). - Benchmark copy gvbbench_eshoponweb dropped afterwards. - Connection: container address on its Docker network, not the published port (no docker-proxy): container gvbbench-mariadb (b70c21d9151f) on network gvbbench-mariadb_default at container-ip:3306, not through docker-proxy localhost:13306. Open after the passes: container-ip:3306 x8 (this target). - CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvbbench-mariadb: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvbbench-mariadb: 12 thread(s) on CPUs 2-3,6-7 (every thread of the container). - Clock per pass (median MHz of engine CPUs / client CPUs, governor, CPUs busy outside the benchmark, client CPU per search): default@1 3492/3492 MHz, performance, outside load 0.12 before the warm-up, 0.12 during, client CPU 0.486 ms per search; exact 3592/3592 MHz, performance, outside load 0.11 before the warm-up, 0.14 during, client CPU 0.489 ms per search; default@8 3492/3492 MHz, performance, outside load 0.14 before the warm-up, 0.11 during, client CPU 0.378 ms per search. ### elasticsearch - Engine: Elasticsearch 9.5.3 (dense_vector) (compose) - Index: HNSW float32, no quantization, m=16, ef_construction=128, cosine; search k=top, num_candidates=100; 1 shard, 0 replicas; force-merged to one segment after the load (at 1,024 dimensions a segment under 1,043 vectors gets no graph) - Search settings: ef_construction=128, k=top, m=16, num_candidates=100 - Load: 524 rows in batches of 1000, 1.0 s of upserts (543 rows/s); count matched 0.2 s after the last upsert; index step 0.4 s (FAILED after 0.4 s: NO HNSW GRAPH, searches scan all 524 vectors: 1 segment(s), 524 vectors, total_vex_size_bytes 0 (_stats dense_vector), index_options hnsw m=16 ef_construction=128; Lucene builds no graph for a segment this small (measured at 1,024 dimensions: 1,042 vectors none, 1,043 vectors a graph), so default search equals exact search) - Search: 15020 latency samples, target held 524 rows, 0 errors - Client CPU per search: 1 searcher 0.46 ms, 8 searchers 0.50 ms - Exact mode: 47,098 searches, p50 1.25 ms, p95 1.44 ms, recall 1.000 - RAM: 2.6 GiB (docker stats); disk: 2.39 MiB added by this load (2.54 MiB (whole engine data folder), 154.47 KiB before) - Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started elasticsearch.compose.yaml with every container created on CPUs 2-3,6-7: gvb-elasticsearch cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards. - WARNING: the index step after the load FAILED (NO HNSW GRAPH, searches scan all 524 vectors: 1 segment(s), 524 vectors, total_vex_size_bytes 0 (_stats dense_vector), index_options hnsw m=16 ef_construction=128; Lucene builds no graph for a segment this small (measured at 1,024 dimensions: 1,042 vectors none, 1,043 vectors a graph), so default search equals exact search). Searched anyway; see the index state. - WARNING: index not ready after the load: NOT ready, 0 of 524 indexed (NO HNSW GRAPH, searches scan all 524 vectors: 1 segment(s), 524 vectors, total_vex_size_bytes 0 (_stats dense_vector), index_options hnsw m=16 ef_construction=128; Lucene builds no graph for a segment this small (measured at 1,024 dimensions: 1,042 vectors none, 1,043 vectors a graph), so default search equals exact search). Measured anyway, so its search numbers may come from a scan or a half-built index. Durability: Every acknowledged bulk request is fsynced to the translog before the answer: index.translog.durability=request, the Elasticsearch default, which neither this sink nor elasticsearch.compose.yaml overrides (ElasticsearchReadinessTests reads it back from the live index). By that setting a process crash or power loss loses no acknowledged write; this is read from the setting, not shown by pulling power. One node and no replicas, so a lost disk loses the data.. - Load average 4.08 3.84 3.67 (1/5/15 min) when searching began. - Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: exact 15,224 searches in 30.0 s, default@1 19,838 searches in 30.0 s, default@8 77,725 searches in 30.0 s. - Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure. exact: warm-up 15.0 s and 11,677 searches (0 failed) at 1 searcher; p50 of the last windows 1.250, 1.251, 1.250 ms (windows of at least 2 s and 100 searches); trial of 2,363 searches in 3.0 s at 1 searcher: p50 1.253 ms against the settled p50 1.250 ms, 0% apart (limit 10%); default@1: warm-up 15.0 s and 11,265 searches (0 failed) at 1 searcher; p50 of the last windows 1.311, 1.303, 1.300 ms (windows of at least 2 s and 100 searches); trial of 2,251 searches in 3.0 s at 1 searcher: p50 1.310 ms against the settled p50 1.303 ms, 1% apart (limit 10%); default@8: warm-up 15.0 s and 40,564 searches (0 failed) at 8 searchers; QPS of the last windows 2,694, 2,713, 2,707 (windows of at least 2 s and 100 searches); trial of 8,144 searches in 3.0 s at 8 searchers: 2,712 QPS against the settled 2,707 QPS, 0% apart (limit 10%). - Pass exact after rehearsal: 2026-10-05T10:05:15.620Z to 2026-10-05T10:06:15.621Z (60.0 s), 47,098 searches, 0 failed, p50 1.252 ms, mean 1.272 ms, p99 1.665 ms, 785.0 QPS (1000/QPS 1.274 ms). - Pass default@1 after exact: 2026-10-05T10:06:33.660Z to 2026-10-05T10:06:53.660Z (20.0 s), 15,020 searches, 0 failed, p50 1.309 ms, mean 1.330 ms, p99 1.734 ms, 751.0 QPS (1000/QPS 1.332 ms). - Pass default@8 after default@1: 2026-10-05T10:07:11.677Z to 2026-10-05T10:07:31.679Z (20.0 s), 54,270 searches, 0 failed, p50 2.818 ms, mean 2.946 ms, p99 5.333 ms, 2713.2 QPS (1000/QPS 0.369 ms). - Passes in the order run: exact, default@1, default@8; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 189,051 sent, 0 failed (warmupErrors). Index after the searches: NOT ready, 0 of 524 indexed (NO HNSW GRAPH, searches scan all 524 vectors: 1 segment(s), 524 vectors, total_vex_size_bytes 0 (_stats dense_vector), index_options hnsw m=16 ef_construction=128; Lucene builds no graph for a segment this small (measured at 1,024 dimensions: 1,042 vectors none, 1,043 vectors a graph), so default search equals exact search). - Benchmark copy gvbbench_eshoponweb dropped afterwards. - Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-elasticsearch (6cb22625f140) on network engines_default at container-ip:9200, not through docker-proxy localhost:9200. Open after the passes: container-ip:9200 x8 (this target). - CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-elasticsearch: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-elasticsearch: 73 thread(s) on CPUs 2-3,6-7 (every thread of the container). - Clock per pass (median MHz of engine CPUs / client CPUs, governor, CPUs busy outside the benchmark, client CPU per search): exact 3492/3492 MHz, performance, outside load 0.16 before the warm-up, 0.15 during, client CPU 0.456 ms per search; default@1 3492/3492 MHz, performance, outside load 0.16 before the warm-up, 0.15 during, client CPU 0.461 ms per search; default@8 3492/3492 MHz, performance, outside load 0.15 before the warm-up, 0.16 during, client CPU 0.496 ms per search. ### clickhouse - Engine: ClickHouse 26.3.39.7 (MergeTree + vector_similarity index) (compose) - Index: vector_similarity HNSW cosineDistance, quantization bf16, M=16 ef_construction=128, hnsw_candidate_list_size_for_search=256, rescoring off; exact mode = full scan with skip indexes off - Search settings: M=16, ef_construction=128, hnsw_candidate_list_size_for_search=256 - Load: 524 rows in batches of 1000, 0.2 s of upserts (2,977 rows/s); count matched 0.0 s after the last upsert; index step 6.2 s (1 data part(s), index files on 1 of them, 524 of 524 live rows in indexed parts; 0 merge(s) running, 0 unfinished mutation(s); plan of the default search uses the Skip index vec_idx on 1 of 1 parts; wait took 6.2 s) - Search: 4294 latency samples, target held 524 rows, 0 errors - Client CPU per search: 1 searcher 0.82 ms, 8 searchers 1.02 ms - Exact mode: 9,549 searches, p50 6.10 ms, p95 7.27 ms, recall 1.000 - RAM: 199.5 MiB (docker stats); disk: 3.26 MiB added by this load (3.37 MiB (whole engine data folder), 116.74 KiB before) - Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started clickhouse.compose.yaml with every container created on CPUs 2-3,6-7: gvb-clickhouse cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards. - Index after the load: ready, 524 of 524 indexed (1 data part(s), index files on 1 of them, 524 of 524 live rows in indexed parts; 0 merge(s) running, 0 unfinished mutation(s); plan of the default search uses the Skip index vec_idx on 1 of 1 parts). Durability: Acknowledged inserts are not fsynced. The sink creates plain MergeTree tables, so the MergeTree settings are the defaults: fsync_after_insert 0 and fsync_part_directory 0 (system.merge_tree_settings), and async_insert 1 with wait_for_async_insert 1 is the server default for the gvb login (system.settings), so an insert is acknowledged once its part is written, not once it is on disk. Measured 2026-10-04 with strace over 30 acknowledged single-row inserts (30 Ok rows in system.asynchronous_insert_log): zero fsync, fdatasync or sync_file_range calls, while the control, a table created with fsync_after_insert=1, made 61 fdatasync calls for 5 inserts. A host power loss or kernel crash can lose acknowledged rows that are still in the OS page cache; a ClickHouse process crash alone should not (inferred, not tested).. - Load average 3.97 3.71 3.64 (1/5/15 min) when searching began. - Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: default@8 14,802 searches in 30.0 s, exact 4,802 searches in 30.0 s, default@1 6,492 searches in 30.0 s. - WARNING: latency had NOT settled when timing began for default@8 (each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure). default@8: warm-up 15.0 s and 7,401 searches (0 failed) at 8 searchers; QPS of the last windows 458, 442, 444 (windows of at least 2 s and 100 searches); trial of 1,581 searches in 3.0 s at 8 searchers: 525 QPS against the settled 444 QPS, 16% apart (limit 10%), so the warm-up was EXTENDED once (30.0 s and 14,706 searches (0 failed) at 8 searchers; QPS of the last windows 530, 525, 452 (windows of at least 2 s and 100 searches), its last 3 windows differed by more than 5%); second trial of 1,311 searches in 3.0 s at 8 searchers: 435 QPS against the settled 525 QPS, 21% apart (limit 10%); still disagreeing after the extension; exact: warm-up 15.0 s and 2,370 searches (0 failed) at 1 searcher; p50 of the last windows 6.089, 6.058, 6.067 ms (windows of at least 2 s and 100 searches); trial of 487 searches in 3.0 s at 1 searcher: p50 6.077 ms against the settled p50 6.067 ms, 0% apart (limit 10%); default@1: warm-up 15.0 s and 3,267 searches (0 failed) at 1 searcher; p50 of the last windows 4.489, 4.504, 4.493 ms (windows of at least 2 s and 100 searches); trial of 657 searches in 3.0 s at 1 searcher: p50 4.468 ms against the settled p50 4.493 ms, 1% apart (limit 10%). Its numbers may still include warm-up; rerun before quoting them. - Pass default@8 after rehearsal: 2026-10-05T10:10:24.737Z to 2026-10-05T10:10:44.750Z (20.0 s), 9,862 searches, 0 failed, p50 15.809 ms, mean 16.228 ms, p99 29.075 ms, 492.8 QPS (1000/QPS 2.029 ms). - Pass exact after default@8: 2026-10-05T10:11:02.759Z to 2026-10-05T10:12:02.760Z (60.0 s), 9,549 searches, 0 failed, p50 6.103 ms, mean 6.280 ms, p99 9.167 ms, 159.1 QPS (1000/QPS 6.284 ms). - Pass default@1 after exact: 2026-10-05T10:12:20.775Z to 2026-10-05T10:12:40.776Z (20.0 s), 4,294 searches, 0 failed, p50 4.508 ms, mean 4.655 ms, p99 6.753 ms, 214.7 QPS (1000/QPS 4.658 ms). - Passes in the order run: default@8, exact, default@1; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 65,235 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (1 data part(s), index files on 1 of them, 524 of 524 live rows in indexed parts; 0 merge(s) running, 0 unfinished mutation(s); plan of the default search uses the Skip index vec_idx on 1 of 1 parts). - Benchmark copy gvbbench_eshoponweb dropped afterwards. - Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-clickhouse (c0152c2a4823) on network gvb-clickhouse_default at container-ip:8123, not through docker-proxy localhost:8123. Open after the passes: container-ip:8123 x1 (this target). - CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-clickhouse: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-clickhouse: 659 thread(s) on CPUs 2-3,6-7 (every thread of the container). - Clock per pass (median MHz of engine CPUs / client CPUs, governor, CPUs busy outside the benchmark, client CPU per search): default@8 3537/3559 MHz, performance, outside load 0.13 before the warm-up, 0.17 during, client CPU 1.019 ms per search; exact 3592/3592 MHz, performance, outside load 0.15 before the warm-up, 0.17 during, client CPU 0.847 ms per search; default@1 3591/3589 MHz, performance, outside load 0.17 before the warm-up, 0.16 during, client CPU 0.815 ms per search. ### mongodb - Engine: MongoDB 8.0.32 Atlas Local (mongod + mongot Vector Search) (compose) - Index: vectorSearch index, HNSW maxEdges=16 numEdgeCandidates=128, float32 binData, cosine, numCandidates=20x hits (min 100); exact mode = $vectorSearch exact:true - Search settings: maxEdges=16, numCandidates=20x, numEdgeCandidates=128 - Load: 524 rows in batches of 1000, 0.2 s of upserts (2,969 rows/s); count matched 1.8 s after the last upsert; index step 10.2 s (index status READY, queryable True; mongot holds 524 of 524 documents in 4 segment(s) (268 + 254 + 1 + 1 documents), 4 searched through the HNSW graph (Approximate); segment layout unchanged for 10.2 s over 11 readings; wait took 10.2 s) - Search: 13874 latency samples, target held 524 rows, 0 errors - Client CPU per search: 1 searcher 0.31 ms, 8 searchers 0.29 ms - Exact mode: 47,167 searches, p50 1.25 ms, p95 1.44 ms, recall 1.000 - RAM: 656.1 MiB (docker stats); disk: 42.39 KiB added by this load (1.36 GiB (whole engine data folder), 1.36 GiB before) - Started by run-all for this measurement with every container created on CPUs 2-3,6-7 (started mongodb.compose.yaml with every container created on CPUs 2-3,6-7: gvb-mongodb cpuset 2-3,6-7 (read back from docker inspect)); stopped afterwards. - Index after the load: ready, 524 of 524 indexed (index status READY, queryable True; mongot holds 524 of 524 documents in 4 segment(s) (268 + 254 + 1 + 1 documents), 4 searched through the HNSW graph (Approximate)). Durability: Acknowledged writes are journaled before the acknowledgement. The sink sets no write concern, so the server default applies: getDefaultRWConcern gives w majority and the one-member replica set (--replSet gvbmongo in the image) has writeConcernMajorityJournalDefault true, with WiredTiger journaling on (journalCommitInterval 100 ms is only the interval for unacknowledged work). Measured 2026-10-04: 30 acknowledged single-document inserts raised WiredTiger 'log sync operations' by 30 and strace showed fdatasync on /data/db/journal/WiredTigerLog files. A crash loses no acknowledged write. mongot's search index is not part of that promise: it follows the collection asynchronously and is rebuilt from it.. - Load average 2.02 4.07 3.97 (1/5/15 min) when searching began. - Rehearsal before any timed pass, untimed, every pass type at its own concurrency for at least 30 s through the same code the timed passes use: exact 20,429 searches in 30.0 s, default@8 55,729 searches in 30.0 s, default@1 20,825 searches in 30.0 s. - Settle: settled before every timed pass, each by its own warm-up at its own concurrency and a trial of the same pass within 10% of the warm-up's settled figure. exact: warm-up 15.0 s and 11,783 searches (0 failed) at 1 searcher; p50 of the last windows 1.243, 1.230, 1.240 ms (windows of at least 2 s and 100 searches); trial of 2,355 searches in 3.0 s at 1 searcher: p50 1.251 ms against the settled p50 1.240 ms, 1% apart (limit 10%); default@8: warm-up 15.0 s and 31,357 searches (0 failed) at 8 searchers; QPS of the last windows 2,070, 2,105, 2,103 (windows of at least 2 s and 100 searches); trial of 6,266 searches in 3.0 s at 8 searchers: 2,086 QPS against the settled 2,103 QPS, 1% apart (limit 10%); default@1: warm-up 15.0 s and 10,383 searches (0 failed) at 1 searcher; p50 of the last windows 1.415, 1.423, 1.420 ms (windows of at least 2 s and 100 searches); trial of 2,093 searches in 3.0 s at 1 searcher: p50 1.412 ms against the settled p50 1.420 ms, 1% apart (limit 10%). - Pass exact after rehearsal: 2026-10-05T10:14:55.665Z to 2026-10-05T10:15:55.665Z (60.0 s), 47,167 searches, 0 failed, p50 1.248 ms, mean 1.270 ms, p99 1.789 ms, 786.1 QPS (1000/QPS 1.272 ms). - Pass default@8 after exact: 2026-10-05T10:16:13.708Z to 2026-10-05T10:16:33.711Z (20.0 s), 41,767 searches, 0 failed, p50 3.707 ms, mean 3.829 ms, p99 6.795 ms, 2088.1 QPS (1000/QPS 0.479 ms). - Pass default@1 after default@8: 2026-10-05T10:16:51.725Z to 2026-10-05T10:17:11.725Z (20.0 s), 13,874 searches, 0 failed, p50 1.412 ms, mean 1.439 ms, p99 1.991 ms, 693.7 QPS (1000/QPS 1.442 ms). - Passes in the order run: exact, default@8, default@1; each after its own warm-up. Untimed searches (rehearsals, per-pass warm-ups, settle checks and extensions): 161,220 sent, 0 failed (warmupErrors). Index after the searches: ready, 524 of 524 indexed (index status READY, queryable True; mongot holds 524 of 524 documents in 4 segment(s) (268 + 254 + 1 + 1 documents), 4 searched through the HNSW graph (Approximate)). - Benchmark copy gvbbench_eshoponweb dropped afterwards. - Connection: container address on its Docker network, not the published port (no docker-proxy): container gvb-mongodb (c00d33276051) on network gvb-mongodb_default at container-ip:27017, not through docker-proxy localhost:27017. Open after the passes: container-ip:27017 x10 (this target). - CPU pinning: cpuset from the container's creation when run-all started it on the engine CPUs, else docker update --cpuset-cpus (each container's change line says which), CPUs 2-3,6-7. Changed: container gvb-mongodb: already on CPUs 2-3,6-7 (its cpuset since creation), not changed. Read back: container gvb-mongodb: 210 thread(s) on CPUs 2-3,6-7 (every thread of the container). - Clock per pass (median MHz of engine CPUs / client CPUs, governor, CPUs busy outside the benchmark, client CPU per search): exact 3492/3492 MHz, performance, outside load 0.20 before the warm-up, 0.25 during, client CPU 0.300 ms per search; default@8 3492/3492 MHz, performance, outside load 0.25 before the warm-up, 0.15 during, client CPU 0.293 ms per search; default@1 3492/3492 MHz, performance, outside load 0.20 before the warm-up, 0.25 during, client CPU 0.308 ms per search. ## Notes - Order: targets ran one at a time in a random order from seed 602 (runSeed; --seed 602 repeats it, targetOrder lists it). Inside each target the timed passes also ran in a random order from the same seed and the target's name (passOrder): default@1 is one searcher for 20 s (every search's latency gives p50/p95/p99, the completed searches give QPS@1, its first answer to each query gives recall and nDCG), default@N is N searchers for 20 s (QPS@N), exact is the engine's exact mode, one searcher for 60 s cycling the queries. - Preparation, untimed, before any timed pass: a rehearsal of every pass type at its own concurrency for 30 s each, through the same code the passes use. - Warm-up and settle check, untimed, right before every timed pass: the pass's own search at the pass's own number of searchers for at least 15 s and at least 20 searches (at most 120 s), read in windows of at least 2 s and 100 searches; then a 3 s trial of the same pass. The trial's figure (p50 with one searcher, QPS with several) must lie within 10% of the warm-up's settled figure (the median of its last 3 windows, which must agree within 5%); if not, the warm-up is extended once (at least 30 s, until its windows agree, at most 120 s), a second trial is taken and the warm-up runs again before the pass. Each target's notes give every check, and a pass that still disagrees flags its target as unsettled. With machine control on, the check for a quiet box is made when the warm-up is announced, before it starts (a wait between the warm-up and the timed pass let the engine go cold), so by the time the clock opens that check is as old as the warm-up and the trial (about 18 s, more after an extension); each pass's conditions record that lead (quietCheckLeadSeconds). Failed untimed searches are counted per target (warmupErrors) and are not in the timed error counts. - Load rows/s counts only time inside each target's upsert calls: one writer, batches of 1000, rows already in memory, the collection dropped and created fresh first. Engines that build or finish their index after the writes do it in a separate timed index step (the load's index seconds), and the run waits for it before searching. - Index proof: each engine's own report of its index (indexState) is read after the load and again after the last pass. A target whose index was not ready after the load was still measured and carries a WARNING; an engine that reports nothing counts as not ready. - Search settings: each searched target records the settings its own index description states (searchSettings: the build parameters and the effort per query, such as m, ef_construction and ef_search), read after the load; consolidating never averages runs whose settings differ. - Durability: each target's crash-safety setting as configured here (durability). Engines that do not force writes to disk on every commit load faster for that reason. - Latency is client-side wall time around each search (network and driver included), every search of the default@1 window, one searcher, the queries cycled; a window with fewer than 200 searches, a p50 above 1.25 x its mean, or a mean above its p99 is flagged. - QPS: N searchers back to back for 20 s per level; completed searches divided by the window's elapsed time. - Recall@10: share of the exact top 10 (brute force in memory) that the engine returned. A hit whose exact similarity ties the 10th best (within 1e-5) also counts, because duplicate rows embed to identical vectors. - Routes: a container engine (the benchmark's own SQL Server and Qdrant containers included) is reached at its container's own address on its Docker network, never through the published localhost port (docker-proxy); the native comparison targets (sql-native, qdrant-native) are native services reached directly over loopback. Each target's addresses and the connections the client held open after its passes are in its notes and in conditions.connections. - RAM of compose engines, the benchmark's own SQL Server and Qdrant containers (sql, sql-diskann, qdrant, qdrant-hnsw) included, is docker stats of the engine's containers; for the native comparison targets (sql-native, qdrant-native) it is the whole native process, including every other database or collection it serves. Disk is the table's reserved pages (SQL), the collection folder (Qdrant), or for other engines what the load added to the engine's data folder (engines that keep data in memory until a snapshot show almost nothing). - Engines: run-all starts an engine that is down and stops it afterwards only if it was not running when the run began; an engine that was already running is left running. - WARNING: the engine did not report a ready index after the load for: elasticsearch. They were measured anyway; their numbers may come from a scan or a half-built index (see each target's index state). - Machine control on: governor performance on every CPU during the run (before: schedutil; at the end: performance; after putting it back: schedutil). engine CPUs 2-3,6-7 (cores 2,6 and 3,7), client CPUs 0-1,4-5 (cores 0,4 and 1,5); the client process was pinned; each engine was pinned to the engine CPUs for its turn and put back after (conditions.engines); an engine run-all started (and stopped) was asked to be created on them, and its notes say whether the host did so or it was moved there after the start. Busy box: right before each pass's warm-up the run waits, up to 10 min, while processes outside the benchmark (everything but this client and the engine under test's cgroups) use more than 0.3 CPUs on average over the last 60 s or the last 5 s, neither window reaching back past the start of the target or of its engine (so the engine's own start-up is not outside work); if the box does not clear the pass runs anyway, flagged 'busy box', as is a pass whose own outside load is above the limit. Outside load counts busy = user + nice + system + irq + softirq + steal (guest time is already in user); CONFIG_IRQ_TIME_ACCOUNTING is not set, so task and cgroup run time include the interrupt and softirq time that hit them and it is added; CONFIG_PARAVIRT_TIME_ACCOUNTING is not set, so steal is added (/boot/config-6.8.0-142-generic) (conditions.cpuAccounting); each pass also records this client's own CPU time per search (conditions.passes[].clientCpuMsPerSearch). CPU clocks were sampled every 250 ms; each pass's min/median/max per CPU is in conditions.passes. CPU idle states, recorded and left as found: driver intel_idle, governor menu, intel_idle max_cstate 9; POLL on, C1 on, C1E on, C3 OFF on CPUs 0-7 (default disabled), C6 on (conditions.cpuIdle). How each target was reached: conditions.connections. Client build Release, .NET 10.0.12. - duckdb is embedded: it ran inside the client process on the client CPUs 0-1,4-5, sharing them with the client. - sqlitevec is embedded: it ran inside the client process on the client CPUs 0-1,4-5, sharing them with the client.