m12p1 (measurement truth): TidalDb::vector_search_items pure k-NN probe + POST /vector_search (standalone + region node, merge-by-distance) + tidal-stress --verify-recall (deterministic id-keyed corpus, in-RAM brute-force cosine oracle, open-loop ramp → recall@k + true p99 + read-knee + JSON/gate exit). Repaired fabricated p99 columns (mean-as-p99) in social-scale.md / scale.rs. Verified real: recall@10=0.9997 at 20k/1536-D vs brute-force. m12p2 (G1 unblock): ANN candidate-gen wired into RETRIEVE — for_you=preference vector, related=seed embedding (similar_to), graceful scan-fallback. Cached per-signal-type top-K (signals/ledger/hot_top_k.rs, decay-order-invariant) so trending serves O(K). related over HTTP (FeedQuery.similar_to). Harness gains --feed-profile / --seed-preferences. Verified: trending retrieve p99 3.5-7.7ms. m12p3 (G2): per-query ef_search now honored (RwLock epoch-guard with_expansion, shared guard for same-ef concurrency) + dimension-aware brute→HNSW crossover usearch_min_vectors(dim) + memory_usage() + examples/ann_grid_search.rs. Measured 1536-D/100k clustered: default M=16/ef_c=400/F16/ef_s=200 clears G1+G2 (recall 0.997, p99 1.4ms); F16 -0.25% vs F32; Int8 rejected (-28%). Recall corpus is now clustered (Gaussian mixture) in grid + harness.
9.7 KiB
USearch Parameter Tuning
Summary
The production default (VectorIndexConfig::default()) is M=16,
ef_construction=400, ef_search=200, F16. m12p3 re-ran the grid search at the
production shape (1536-D) — not the historical 128-D — with a real exact
brute-force oracle, and validated the F16/Int8 recall and memory frontier there.
What changed in m12p3. The earlier version of this doc reported an extrapolated 128-D table (values from published ANN-Benchmarks, not measured here). That is replaced below by measured 1536-D numbers from
cargo run --release --example ann_grid_search. The example builds aUsearchIndexand aBruteForceIndexoracle over the same deterministic, id-keyed corpus and reports recall@10 vs the exact oracle, single-thread mean/p99 search latency, build time, and the truememory_usage()footprint.
Grid Search Setup (m12p3, measured)
| Parameter | Values |
|---|---|
| M (connectivity) | 16, 24, 32 |
| ef_construction | 400 |
| ef_search (per-query) | 128, 200, 400, 600 |
| Quantization | F32, F16, Int8 |
| Dataset size | 100,000 vectors (and 1M on the k3s node) |
| Dimensionality | 1536D (production embedding width) |
| Corpus shape | clustered (Gaussian mixture) — see below for why not uniform-random |
| Distance metric | L2 (L2-normalized → equivalent to cosine) |
| Recall metric | recall@10 vs exact brute-force cosine (200-query average) |
| Hardware | Apple Silicon laptop (release build) |
Method
- Build a deterministic, id-keyed clustered corpus (SplitMix64 Gaussian mixture — see below) so the run is reproducible, the recall numbers are comparable across machines, and recall@k is a meaningful metric at scale.
- Build a
BruteForceIndex(F32) over the same vectors and compute the exact top-10 for every query once — the ground truth the HNSW approximates. - For each
(M, ef_construction)build the graph once and sweepef_search(a per-query knob — no rebuild). For quantization, build F32/F16/Int8 at the recall-frontier graph. - Record recall@10, mean + p99 search latency (µs), build time (s), and the true in-memory footprint, extrapolated to 1M vectors.
The corpus must be clustered, not uniform-random
A subtle but decisive measurement point: uniform-random high-dimensional vectors are a pathological ANN benchmark. By concentration of measure, every pair of random unit vectors in 1536-D sits at cosine ≈ 0, so beyond a tiny perturbation a query's true top-10 is an arbitrary draw from a thick equidistant shell — recall@10 then measures impossible tie-breaking, not index quality, and it gets worse as the corpus grows (the shell thickens). Measured on uniform-random data, recall@10 fell from ~0.97 at 10k to ~0.54 at 100k — a corpus-size artifact, not an index regression.
Real text/image embeddings instead live on a low-dimensional manifold with
clusters: a point's neighbours are its cluster-mates, distinctly closer than the
bulk. The harness therefore builds a Gaussian mixture (--clusters,
--spread-milli; default ~100 points/cluster, spread 0.5 ⇒ intra-cluster cosine
≈ 0.89, inter ≈ 0) — the shape related/for_you queries actually run against,
where recall@10 is a meaningful, scale-stable metric. All numbers below use it.
Results — HNSW parameter sweep (1536D, F16, 100k clustered corpus, measured)
| M / ef_c / ef_s | recall@10 | mean (µs) | p99 (µs) | build (s) | mem/1M (GB) |
|---|---|---|---|---|---|
| 16 / 400 / 128 | 0.9970 | 670 | 1018 | 19.7 | 5.21 |
| 16 / 400 / 200 | 0.9970 | 1129 | 1394 | 19.7 | 5.21 |
| 16 / 400 / 400 | 0.9970 | 2292 | 2719 | 19.7 | 5.21 |
| 16 / 400 / 600 | 0.9970 | 3477 | 4255 | 19.7 | 5.21 |
| 24 / 400 / 128 | 0.9985 | 1074 | 1299 | 39.0 | 5.53 |
| 24 / 400 / 200 | 0.9985 | 1812 | 2239 | 39.0 | 5.53 |
| 24 / 400 / 400 | 0.9985 | 3536 | 4088 | 39.0 | 5.53 |
| 32 / 400 / 200 | 0.9975 | 2331 | 2857 | 42.5 | 5.53 |
Every point clears recall@10 ≥ 0.95 and p99 ≤ 10ms (raw single-thread ANN cost). Two readings:
- ef_search is the latency lever, and recall saturates early. At M=16, recall is already 0.997 at ef_s=128 (p99 1.0ms) and does not improve with a wider beam — only latency grows (ef_s=600 → p99 4.3ms). So the cheap operating point is ef_s≈128; the per-query override lets a latency-sensitive caller pick it without rebuilding.
- M=24 buys ~0.15% recall for ~1.6× latency and 2× build — worth it only when the last fraction of recall matters. M=32 is not better than M=24 here.
Results — quantization sweep (1536D, M=24, ef_c=400, ef_s=400, 100k, measured)
| quantization | recall@10 | mean (µs) | p99 (µs) | mem/1M (GB) |
|---|---|---|---|---|
| F32 | 0.9985 | 3642 | 5126 | 10.53 |
| F16 | 0.9960 | 3569 | 4680 | 5.53 |
| Int8 | 0.7150 | 1389 | 1678 | 3.03 |
- F16 is the right default: it costs only 0.25% recall vs F32 (0.9960 vs 0.9985) for half the memory — validating the "<1% recall loss" claim at 1536-D, not just the 128-D it was previously checked at.
- Int8 is NOT viable at 1536-D as a drop-in: recall collapses to 0.715 (a 28% loss). The memory saving (3.0 vs 5.5 GB/1M) is real but does not justify losing a quarter of the neighbours; Int8 would need quantization-aware scaling before it is usable. The roadmap's "Int8 ≈ 1.8 GB/1M as a RAM fallback" option is therefore rejected at this dim on accuracy grounds.
Memory note. The per-1M figures are the true
USearchfootprint (memory_usage(), graph links + quantized vectors), not a vectors-only estimate. F16 ≈ 5.2 GB/1M, not the ~3.4 GB the roadmap projected from1536 × 2 bytesalone — the ~1.8 GB difference is the HNSW proximity graph (M=16). Size pods at ≈ 5.5 GB/1M for F16, ≈ 10.5 GB/1M for F32.
Live cross-check (real engine, HTTP)
The standalone grid above measures the index in isolation. The authoritative
end-to-end number is tidal-stress --verify-recall against a real server: at
20k/1536-D (clustered), the engine's /vector_search returned recall@10 =
1.0000 vs the brute-force cosine oracle at the default beam and at
--recall-ef-search 400 — confirming both the index quality and that the
per-query ef_search is honored over the wire. (The harness's end-to-end p99 of
~19 ms there is HTTP round-trip + 1536-float (de)serialization on one laptop, not
ANN — the raw search is ~1 ms per the grid; it is a single-node wire cost, not an
index limit.)
Decision
Default stays M=16, ef_construction=400, ef_search=200, F16 — the grid validates it at the production shape: recall@10 0.997, p99 ≈ 1.4 ms raw ANN at 100k/1536-D clustered, 5.2 GB/1M. No change is warranted.
- Latency-sensitive reads can drop to
ef_search=128per request (recall still 0.997, p99 ≈ 1.0 ms) via the per-query override. - Maximum recall (e.g. an offline eval) can use M=24, recall 0.9985, at ~1.6× latency — but M=24 is not the default because M=16 already clears G2 (≥0.95) and G1 (≤10 ms) with wide margin.
- F32 only if an application cannot tolerate the 0.25% F16 gap; doubles RAM.
- Int8 rejected at 1536-D (recall 0.715).
Per-query ef_search (m12p3: now honored)
M and ef_construction are graph properties fixed at build time. ef_search is
a per-query knob: a low-latency query can pass ef_search=64 and a
high-recall query ef_search=400 against the same shared index, concurrently.
Before m12p3, UsearchIndex accepted ef_search for trait compliance but
ignored it (logging a warning) because USearch 2.24 has no per-call beam
argument — the only knob is the index-global change_expansion_search. m12p3
makes the override real and race-free via an RwLock epoch guard
(UsearchIndex::with_expansion): searches that agree on ef_search run in
parallel under a shared guard; only a query that changes the live beam width
takes the exclusive guard for its (set, search) window. The override reaches the
wire through vector_search_items(.., ef_search) and the /vector_search
ef_search field, and tidal-stress --verify-recall --recall-ef-search sweeps it
open-loop.
The ef_search axis in the table above is the recall/latency trade this knob now
controls per request — higher ef_search recovers more true neighbours at a
latency cost, exactly as the HNSW theory predicts (spec
07-vector-retrieval.md).
Brute-force → HNSW crossover (m12p3: dimension-aware)
The brute-force backend (exact, used for small slots) scans every vector under a
read lock, at count × dim cost. A single fixed crossover is wrong: 10,000
vectors is a 1.3M-FMA scan at 128-D (sub-ms) but a 15.4M-FMA scan at 1536-D (tens
of ms, under a lock that also blocks writers). usearch_min_vectors(dim) now
scales the crossover so a brute-force scan stays within ~4M FMAs: ≈10,000 at/under
128-D (unchanged), ≈2,600 at 1536-D. Above the crossover a slot is the production
HNSW; below it stays exact. See storage/vector/registry.rs.
Regression Guard
tidal/tests/vector_usearch.rs:
recall_at_10_above_threshold— default config (M=16, ef_c=400) recall@10 > 0.95.usearch_per_query_ef_search_is_honored— a wide per-queryef_searchrecovers strictly more true neighbours than a starved beam (proves the override is live).
Reproduce
# Production shape on one laptop (oracle ≈ 0.6 GB at 100k/1536D):
cargo run --release --example ann_grid_search -- \
--corpus 100000 --dim 1536 --queries 200 --k 10
# 1M shape (oracle ≈ 6 GB; run on the k3s node):
cargo run --release --example ann_grid_search -- \
--corpus 1000000 --dim 1536 --queries 200 --k 10