m12p1 (measurement truth): TidalDb::vector_search_items pure k-NN probe + POST /vector_search (standalone + region node, merge-by-distance) + tidal-stress --verify-recall (deterministic id-keyed corpus, in-RAM brute-force cosine oracle, open-loop ramp → recall@k + true p99 + read-knee + JSON/gate exit). Repaired fabricated p99 columns (mean-as-p99) in social-scale.md / scale.rs. Verified real: recall@10=0.9997 at 20k/1536-D vs brute-force. m12p2 (G1 unblock): ANN candidate-gen wired into RETRIEVE — for_you=preference vector, related=seed embedding (similar_to), graceful scan-fallback. Cached per-signal-type top-K (signals/ledger/hot_top_k.rs, decay-order-invariant) so trending serves O(K). related over HTTP (FeedQuery.similar_to). Harness gains --feed-profile / --seed-preferences. Verified: trending retrieve p99 3.5-7.7ms. m12p3 (G2): per-query ef_search now honored (RwLock epoch-guard with_expansion, shared guard for same-ef concurrency) + dimension-aware brute→HNSW crossover usearch_min_vectors(dim) + memory_usage() + examples/ann_grid_search.rs. Measured 1536-D/100k clustered: default M=16/ef_c=400/F16/ef_s=200 clears G1+G2 (recall 0.997, p99 1.4ms); F16 -0.25% vs F32; Int8 rejected (-28%). Recall corpus is now clustered (Gaussian mixture) in grid + harness.
14 KiB
Scale Benchmarks: 1M-Item Baselines
Hardware
macOS Darwin 23.6.0 (Apple Silicon / x86-64 — see run date below).
Run command:
cargo bench --manifest-path tidal/Cargo.toml --bench scale
Date: 2026-02-23
Dataset
| Parameter | Value |
|---|---|
| Items | 1,000,000 |
| Creators | 10,000 (100 items/creator) |
| Categories | 20 |
| Embedding dim | 128 (not 1536 — reduced for bench RAM) |
| Signal coverage | 10% view, 5% like |
| Bench tool | Criterion (sample_size=10, 30s measurement, Flat mode) — closed-loop, single-threaded; reports the mean, not the tail |
Measurement contract (read before trusting any number below)
The
time:figures Criterion reports are single-threaded mean per-op cost under a closed loop — the[lower mean upper]triple is a confidence interval on the mean, not a latency distribution. They are regression tripwires, not tail-SLO evidence.A
p99/p999/tail SLO can only be honored by an open-loop, coordinated-omission-corrected measurement under real concurrency (thetidal-stressramp). A closed-loop mean cannot observe the tail it hides: when the system stalls, a closed-loop harness simply stops sending, so the queue that would inflate p99 never forms. Reporting a mean under ap99label understates the tail by 10–100× under load.Therefore every row below is labelled isolated per-op cost (mean). Tail SLOs are validated separately — see
tidal-stress(open-loop) for the authoritative p99/p999 under the production workload.
Acceptance Criteria (isolated per-op cost — regression tripwires, NOT tail SLOs)
| Benchmark | Tail SLO (validated open-loop) | Isolated per-op cost (mean, closed-loop) | Tripwire |
|---|---|---|---|
| RETRIEVE | < 50ms p99 | 152 µs (for_you) | ✅ well under |
| SEARCH | < 100ms p99 | 28.9 ms (text_only) | ✅ well under |
| Signal write | < 100µs p99 | 82 ns (rotating 1K) | ✅ well under |
The mean per-op costs sit far under the tail targets — a necessary but not
sufficient condition for the p99 SLO. "Well under" means the mean clears the
target with headroom; the p99 itself is signed off only by the open-loop
tidal-stress run, never by this table.
Benchmark Results
RETRIEVE (1M items)
retrieve_1m/for_you time: [151.88 µs 152.13 µs 152.40 µs]
retrieve_1m/trending time: [127.96 µs 128.25 µs 128.52 µs]
retrieve_1m/new_filtered time: [ 7.5636 µs 7.5855 µs 7.6058 µs]
All RETRIEVE queries < 200µs. The 50ms target is beaten by 3 orders of magnitude.
for_you: signal-scored ranking over full 1M-item universe — 152µstrending: windowed view count ranking — 128µsnew_filtered: category filter at ~5% selectivity — 7.6µs (bitmap pre-filter eliminates 95% of candidates)
SEARCH (1M items)
search_1m/text_only time: [28.844 ms 28.934 ms 29.021 ms]
search_1m/text_filtered time: [ 1.8972 ms 1.9104 ms 1.9220 ms]
Both SEARCH queries < 30ms. The 100ms target is beaten by 3-50×.
text_only: BM25 over 1M documents — 28.9ms (most expensive path; dominated by Tantivy posting list traversal)text_filtered: BM25 with category filter reduces candidate set — 1.9ms
Signal Write (1M-item DB, rotating 1K entities)
signal_write_1m/write_rotating_1k_entities time: [82.033 ns 82.286 ns 82.535 ns]
82 ns per write. The 100µs target is beaten by 1,200×. DashMap hot-path write amortises to sub-100ns across 1K rotating entity IDs.
Setup Notes
The LazyLock<TidalDb> pattern ensures the 1M-item database is built exactly once per bench run. Build time ~30s on the reference hardware above. The text syncer waits 3s after ingestion.
Database Build Time
Approximately 30 seconds on reference hardware (observed from [scale bench] Database ready log line).
Analysis
tidalDB's isolated per-op mean cost sits well within all three
acceptance-criteria targets at 1M items. The dominant cost is SEARCH text_only at ~29ms — driven by Tantivy posting list traversal across 1M documents. The LogMergePolicy tuning (< 20 segments at steady state) keeps this below the 100ms target with headroom. The p99 tail SLOs themselves are signed off by the open-loop tidal-stress ramp, not by these closed-loop means (see the measurement contract above).
Signal writes at 82ns confirm the DashMap hot-path is not a bottleneck at this scale. The 5M-entry LRU trimming threshold (DEFAULT_MAX_SIGNAL_ENTRIES) provides ample headroom for the 100K-item signal coverage in this benchmark (~200K entries = ~218MB).
Read-recall harness — recall@k + true p99 (m12p1, open-loop)
Every number above is an isolated per-op mean and says nothing about recall — the fraction of the true nearest neighbours an ANN query actually returns. m12p1 adds the open-loop harness that measures recall@k AND true p99 at the production shape, so G1 (p99 ≤ 10ms) and G2 (recall@10 ≥ 0.95) can be steered against real numbers instead of absent ones.
How it works
tidal-stress --verify-recallseeds the corpus with deterministic, id-keyed embeddings, so the generator reconstructs the exact indexed vectors and computes a brute-force cosine ground truth in RAM (the exact answer the HNSW index approximates).- It ramps
POST /vector_searchprobes open-loop (coordinated-omission corrected) and scores each response'srecall@kagainst the precomputed ground truth, recording the true p99 (a genuine tail, not a closed-loop mean). - The verdict reports the read-knee: the highest sustained QPS at which
p99 ≤ target AND recall@k ≥ targetboth hold, plus a machine-readable JSON summary and a--fail-on-kneePASS/FAIL exit for the soak gate.
/vector_search is a pure k-NN probe (TidalDb::vector_search_items): raw
HNSW nearest neighbours with NO profile scoring, fusion, or diversity, so the
number is the ANN index quality in isolation — exactly the G2 metric. It is a
single-index measurement: run it against a standalone node or a cluster at the
S=1 shape (every replica holds the full corpus). RAM for the brute-force oracle
is corpus × dim × 4 bytes (~0.6 GB at 100k/1536D, ~6 GB at 1M/1536D).
Measured (local, standalone, 1536-dim)
Verified end-to-end against a real standalone server (--schema content_vector =
1536, release build, Apple Silicon laptop):
| Corpus / dim | recall@10 (vs brute-force cosine) | per-query cost | p99 @ 100 rps |
|---|---|---|---|
| 20,000 / 1536D | 0.9997 | ~16 ms (single request) | ~37 ms |
The HNSW index (M=16, ef_construction=400, ef_search=200, F16) returns essentially
the exact neighbour set — recall@10 ≈ 0.9997, far above the 0.95 G2 target. The
~16–19 ms per-query cost on a single laptop node is well over the 10 ms G1 target
and saturates past ~400 rps; G1 is the work of m12p2 (ANN candidate-gen in
RETRIEVE) + m12p3 (index tuning + per-request ef_search), measured by this same
harness at scale.
Running the exit-gate shape (100k AND 1M)
# 100k corpus, 1536-dim, against a single-index node:
tidal-stress --target http://<node>:9500 --verify-recall \
--corpus 100000 --embedding-dim 1536 --recall-queries 1000 \
--ramp "200:30,500:30,1000:30,2000:30" \
--read-p99-target-ms 10 --recall-target 0.95 \
--json-summary recall-100k.json --fail-on-knee
# 1M corpus: same, --corpus 1000000 (oracle needs ~6 GB RAM on the generator).
The 1M run requires the k3s cluster (RAM + a seeded 1M/1536D corpus) and is tracked as the cluster-side step of the m12p1 exit gate; the harness, the probe endpoint, and the verdict are proven locally at 20k/1536D above.
RETRIEVE with ANN candidate generation (m12p2)
m12p2 replaced the arbitrary low-id scan slice in candidate generation with
relevance-bounded sources: for_you/related source candidates by ANN
nearest-neighbour (over the user's preference vector / the seed item's
embedding), and trending by a cached per-signal-type top-K. Measured against a
real 1536-dim standalone server (release, Apple Silicon laptop):
| Profile | Candidate source | Retrieve feed p99 | G1 (≤10ms) |
|---|---|---|---|
trending |
cached SignalRanked(view) top-K |
3.5–7.7 ms (600–1200 rps, concurrent writes) | ✅ met |
for_you |
ANN over the user preference vector | ~24 ms (100–400 rps) | ⚠ see below |
trendingholds p99 ≤ 10ms because the cached top-K serves candidates in O(K) (the O(N) rebuild is throttled off the read hot path), so it ranks the actually-viewed corpus at any id while staying fast even under a concurrent write stream.for_youis ANN-backed (the candidate recall is the m12p1/vector_searchprobe, 0.9997). Its ~24ms p99 on this single laptop node is dominated NOT by the ANN search (~1.5ms raw at 10k/1536D — seeann_recall_at_10_1536dinbenches/vector.rs) but by the Stage-3 preference boost, which reads each of the ~240 candidates' embeddings from storage and recomputes cosine to the preference vector. That O(K) per-candidate recompute is the m12p3 target — the risk register's "materialized-score layer" / reuse the ANN distance already computed in candidate-gen. The absolute 10ms at 1M is an m12p3 (ef_searchtuning) + production-hardware result; m12p2 delivered the algorithmic change (relevance-bounded candidate generation) and surfaced the remaining bottleneck.
Reproduce the per-profile retrieve p99 with the harness:
# for_you (ANN): build preference vectors, force the profile, gate at 10ms.
tidal-stress --target http://<node> --mix "feed=1" \
--feed-profile for_you --seed-preferences \
--corpus 100000 --embedding-dim 1536 --users 10000 --max-p99-ms 10
# trending (cached SignalRanked): peach mix writes views; force the profile.
tidal-stress --target http://<node> --mix peach --feed-profile trending --skip-seed \
--corpus 100000 --embedding-dim 1536 --max-p99-ms 10
Index tuning + recall/memory frontier at the production shape (m12p3)
m12p3 tuned the HNSW at the production shape (1536-D) — earlier tuning was
only validated at 128-D — and characterized the recall/latency/memory frontier
with a real exact oracle. The full method, the chosen M/ef, and the
reproduce commands live in usearch-tuning.md; the headline
recall + per-1M memory sizing is below.
The numbers come from cargo run --release --example ann_grid_search: a
deterministic id-keyed corpus, an exact BruteForceIndex (F32) oracle, recall@10
scored per query, and the index's true memory_usage() extrapolated to 1M.
The corpus is a Gaussian mixture (clustered), not uniform-random: uniform high-dimensional vectors are pathological for recall@k (every pair ≈ orthogonal, so top-k is an arbitrary draw from an equidistant shell — recall@10 measured ~0.97 at 10k but ~0.54 at 100k, a corpus-size artifact, not an index regression). Real embeddings cluster on a manifold; the clustered corpus reflects that, so recall@k is meaningful and scale-stable. Details in usearch-tuning.md.
Recall@10 at 1536-D (measured, 100k clustered corpus, vs exact brute-force cosine)
| Config (F16) | recall@10 | p99 (raw ANN) | G2 (≥0.95) | G1 (≤10ms) |
|---|---|---|---|---|
| M=16, ef_c=400, ef_s=128 | 0.9970 | 1.0 ms | ✅ | ✅ |
| M=16, ef_c=400, ef_s=200 (default) | 0.9970 | 1.4 ms | ✅ | ✅ |
| M=24, ef_c=400, ef_s=200 | 0.9985 | 2.2 ms | ✅ | ✅ |
| M=32, ef_c=400, ef_s=200 | 0.9975 | 2.9 ms | ✅ | ✅ |
The production default (M=16, ef_c=400, F16) clears both G1 and G2 at
100k/1536-D with wide margin. ef_search is the latency lever and recall
saturates by ef_s=128, so the per-query override (m12p3) lets a latency-sensitive
read drop to ef_s=128 (p99 ≈ 1.0 ms) without losing recall. Live cross-check:
tidal-stress --verify-recall against a real server returned recall@10 = 1.0000
at 20k/1536-D (clustered), at the default beam and at --recall-ef-search 400.
Per-1M memory sizing by quantization (true measured footprint, extrapolated from 100k)
| Quantization | recall@10 (M=24, ef_c=400, ef_s=400) | mem/1M (GB) | Verdict |
|---|---|---|---|
| F32 | 0.9985 | 10.53 | exact, 2× RAM |
| F16 (default) | 0.9960 | 5.53 | −0.25% recall, half RAM — chosen |
| Int8 | 0.7150 | 3.03 | rejected: −28% recall at 1536-D |
The footprint is USearch's true memory_usage() (graph links + quantized
vectors), so F16 is ≈ 5.2–5.5 GB/1M, not the ~3.4 GB a vectors-only estimate
(1536 × 2 B) gives — the ~1.8 GB difference is the HNSW graph. Size pods at
≈ 5.5 GB/1M for F16. Int8's 1.8 GB option is rejected: it loses a quarter of
the neighbours at 1536-D and would need quantization-aware scaling first.
Per-query ef_search is now honored
The ef_search column in the frontier table is a per-request knob as of
m12p3 (it was silently ignored before). A low-latency surface can request a
narrow beam and a high-recall surface a wide one against the same shared index,
concurrently — see usearch-tuning.md.
tidal-stress --verify-recall --recall-ef-search <N> sweeps it open-loop.
The 1M run
The 100k numbers above fit one laptop (the F32 oracle is ≈ 0.6 GB). The 1M exit-gate run needs ≈ 6 GB for the oracle plus the HNSW, so it runs on the k3s node:
cargo run --release --example ann_grid_search -- \
--corpus 1000000 --dim 1536 --queries 200 --k 10