# Scale Benchmarks: 1M-Item Baselines ## Hardware macOS Darwin 23.6.0 (Apple Silicon / x86-64 — see run date below). **Run command:** ```bash cargo bench --manifest-path tidal/Cargo.toml --bench scale ``` **Date:** 2026-02-23 ## Dataset | Parameter | Value | |-----------|-------| | Items | 1,000,000 | | Creators | 10,000 (100 items/creator) | | Categories | 20 | | Embedding dim | 128 (not 1536 — reduced for bench RAM) | | Signal coverage | 10% view, 5% like | | Bench tool | Criterion (sample_size=10, 30s measurement, Flat mode) — **closed-loop, single-threaded; reports the mean, not the tail** | ## Measurement contract (read before trusting any number below) > The `time:` figures Criterion reports are **single-threaded mean per-op cost > under a closed loop** — the `[lower mean upper]` triple is a *confidence > interval on the mean*, **not** a latency distribution. They are > **regression tripwires**, not tail-SLO evidence. > > A `p99`/`p999`/tail SLO can only be honored by an **open-loop, > coordinated-omission-corrected** measurement under real concurrency (the > `tidal-stress` ramp). A closed-loop mean cannot observe the tail it hides: > when the system stalls, a closed-loop harness simply *stops sending*, so the > queue that would inflate p99 never forms. Reporting a mean under a `p99` label > understates the tail by 10–100× under load. > > Therefore every row below is labelled **isolated per-op cost (mean)**. Tail > SLOs are validated separately — see `tidal-stress` (open-loop) for the > authoritative p99/p999 under the production workload. ## Acceptance Criteria (isolated per-op cost — regression tripwires, NOT tail SLOs) | Benchmark | Tail SLO (validated open-loop) | Isolated per-op cost (mean, closed-loop) | Tripwire | |-----------|--------------------------------|------------------------------------------|----------| | RETRIEVE | < 50ms p99 | **152 µs** (for_you) | ✅ well under | | SEARCH | < 100ms p99 | **28.9 ms** (text_only) | ✅ well under | | Signal write | < 100µs p99 | **82 ns** (rotating 1K) | ✅ well under | The mean per-op costs sit far under the tail targets — a necessary but **not sufficient** condition for the p99 SLO. "Well under" means the *mean* clears the target with headroom; the p99 itself is signed off only by the open-loop `tidal-stress` run, never by this table. ## Benchmark Results ### RETRIEVE (1M items) ``` retrieve_1m/for_you time: [151.88 µs 152.13 µs 152.40 µs] retrieve_1m/trending time: [127.96 µs 128.25 µs 128.52 µs] retrieve_1m/new_filtered time: [ 7.5636 µs 7.5855 µs 7.6058 µs] ``` **All RETRIEVE queries < 200µs.** The 50ms target is beaten by 3 orders of magnitude. - `for_you`: signal-scored ranking over full 1M-item universe — 152µs - `trending`: windowed view count ranking — 128µs - `new_filtered`: category filter at ~5% selectivity — 7.6µs (bitmap pre-filter eliminates 95% of candidates) ### SEARCH (1M items) ``` search_1m/text_only time: [28.844 ms 28.934 ms 29.021 ms] search_1m/text_filtered time: [ 1.8972 ms 1.9104 ms 1.9220 ms] ``` **Both SEARCH queries < 30ms.** The 100ms target is beaten by 3-50×. - `text_only`: BM25 over 1M documents — 28.9ms (most expensive path; dominated by Tantivy posting list traversal) - `text_filtered`: BM25 with category filter reduces candidate set — 1.9ms ### Signal Write (1M-item DB, rotating 1K entities) ``` signal_write_1m/write_rotating_1k_entities time: [82.033 ns 82.286 ns 82.535 ns] ``` **82 ns per write.** The 100µs target is beaten by 1,200×. DashMap hot-path write amortises to sub-100ns across 1K rotating entity IDs. ## Setup Notes The `LazyLock` pattern ensures the 1M-item database is built exactly once per bench run. Build time ~30s on the reference hardware above. The text syncer waits 3s after ingestion. ## Database Build Time Approximately **30 seconds** on reference hardware (observed from `[scale bench] Database ready` log line). ## Analysis tidalDB's **isolated per-op mean cost** sits well within all three acceptance-criteria targets at 1M items. The dominant cost is SEARCH text_only at ~29ms — driven by Tantivy posting list traversal across 1M documents. The LogMergePolicy tuning (< 20 segments at steady state) keeps this below the 100ms target with headroom. **The p99 tail SLOs themselves are signed off by the open-loop `tidal-stress` ramp, not by these closed-loop means** (see the measurement contract above). Signal writes at 82ns confirm the DashMap hot-path is not a bottleneck at this scale. The 5M-entry LRU trimming threshold (DEFAULT_MAX_SIGNAL_ENTRIES) provides ample headroom for the 100K-item signal coverage in this benchmark (~200K entries = ~218MB). ## Read-recall harness — recall@k + true p99 (m12p1, open-loop) Every number above is an *isolated per-op mean* and says nothing about **recall** — the fraction of the true nearest neighbours an ANN query actually returns. m12p1 adds the open-loop harness that measures recall@k AND true p99 at the production shape, so G1 (p99 ≤ 10ms) and G2 (recall@10 ≥ 0.95) can be steered against real numbers instead of absent ones. ### How it works 1. `tidal-stress --verify-recall` seeds the corpus with **deterministic, id-keyed** embeddings, so the generator reconstructs the exact indexed vectors and computes a **brute-force cosine ground truth** in RAM (the exact answer the HNSW index approximates). 2. It ramps `POST /vector_search` probes **open-loop** (coordinated-omission corrected) and scores each response's `recall@k` against the precomputed ground truth, recording the **true p99** (a genuine tail, not a closed-loop mean). 3. The verdict reports the **read-knee**: the highest sustained QPS at which `p99 ≤ target AND recall@k ≥ target` both hold, plus a machine-readable JSON summary and a `--fail-on-knee` PASS/FAIL exit for the soak gate. `/vector_search` is a **pure k-NN probe** (`TidalDb::vector_search_items`): raw HNSW nearest neighbours with NO profile scoring, fusion, or diversity, so the number is the ANN index quality in isolation — exactly the G2 metric. It is a single-index measurement: run it against a standalone node or a cluster at the S=1 shape (every replica holds the full corpus). RAM for the brute-force oracle is `corpus × dim × 4` bytes (~0.6 GB at 100k/1536D, ~6 GB at 1M/1536D). ### Measured (local, standalone, 1536-dim) Verified end-to-end against a real standalone server (`--schema` content_vector = 1536, release build, Apple Silicon laptop): | Corpus / dim | recall@10 (vs brute-force cosine) | per-query cost | p99 @ 100 rps | |--------------|-----------------------------------|----------------|---------------| | 20,000 / 1536D | **0.9997** | ~16 ms (single request) | ~37 ms | The HNSW index (M=16, ef_construction=400, ef_search=200, F16) returns essentially the exact neighbour set — recall@10 ≈ 0.9997, far above the 0.95 G2 target. The ~16–19 ms per-query cost on a single laptop node is well over the 10 ms G1 target and saturates past ~400 rps; G1 is the work of m12p2 (ANN candidate-gen in RETRIEVE) + m12p3 (index tuning + per-request `ef_search`), measured by this same harness at scale. ### Running the exit-gate shape (100k AND 1M) ```bash # 100k corpus, 1536-dim, against a single-index node: tidal-stress --target http://:9500 --verify-recall \ --corpus 100000 --embedding-dim 1536 --recall-queries 1000 \ --ramp "200:30,500:30,1000:30,2000:30" \ --read-p99-target-ms 10 --recall-target 0.95 \ --json-summary recall-100k.json --fail-on-knee # 1M corpus: same, --corpus 1000000 (oracle needs ~6 GB RAM on the generator). ``` The 1M run requires the k3s cluster (RAM + a seeded 1M/1536D corpus) and is tracked as the cluster-side step of the m12p1 exit gate; the harness, the probe endpoint, and the verdict are proven locally at 20k/1536D above. ## RETRIEVE with ANN candidate generation (m12p2) m12p2 replaced the arbitrary low-id scan slice in candidate generation with relevance-bounded sources: `for_you`/`related` source candidates by ANN nearest-neighbour (over the user's preference vector / the seed item's embedding), and `trending` by a cached per-signal-type top-K. Measured against a real 1536-dim standalone server (release, Apple Silicon laptop): | Profile | Candidate source | Retrieve feed p99 | G1 (≤10ms) | |---------|------------------|-------------------|------------| | `trending` | cached `SignalRanked(view)` top-K | **3.5–7.7 ms** (600–1200 rps, concurrent writes) | ✅ met | | `for_you` | ANN over the user preference vector | ~24 ms (100–400 rps) | ⚠ see below | - **`trending`** holds p99 ≤ 10ms because the cached top-K serves candidates in O(K) (the O(N) rebuild is throttled off the read hot path), so it ranks the actually-viewed corpus at any id while staying fast even under a concurrent write stream. - **`for_you`** is ANN-backed (the candidate recall is the m12p1 `/vector_search` probe, 0.9997). Its ~24ms p99 on this single laptop node is dominated NOT by the ANN search (~1.5ms raw at 10k/1536D — see `ann_recall_at_10_1536d` in `benches/vector.rs`) but by the **Stage-3 preference boost**, which reads each of the ~240 candidates' embeddings from storage and recomputes cosine to the preference vector. That O(K) per-candidate recompute is the m12p3 target — the risk register's "materialized-score layer" / reuse the ANN distance already computed in candidate-gen. The absolute 10ms at 1M is an m12p3 (`ef_search` tuning) + production-hardware result; m12p2 delivered the algorithmic change (relevance-bounded candidate generation) and surfaced the remaining bottleneck. Reproduce the per-profile retrieve p99 with the harness: ```bash # for_you (ANN): build preference vectors, force the profile, gate at 10ms. tidal-stress --target http:// --mix "feed=1" \ --feed-profile for_you --seed-preferences \ --corpus 100000 --embedding-dim 1536 --users 10000 --max-p99-ms 10 # trending (cached SignalRanked): peach mix writes views; force the profile. tidal-stress --target http:// --mix peach --feed-profile trending --skip-seed \ --corpus 100000 --embedding-dim 1536 --max-p99-ms 10 ``` ## Index tuning + recall/memory frontier at the production shape (m12p3) m12p3 tuned the HNSW at the **production shape (1536-D)** — earlier tuning was only validated at 128-D — and characterized the recall/latency/memory frontier with a real exact oracle. The full method, the chosen `M`/`ef`, and the reproduce commands live in [usearch-tuning.md](usearch-tuning.md); the headline recall + per-1M memory sizing is below. The numbers come from `cargo run --release --example ann_grid_search`: a deterministic id-keyed corpus, an exact `BruteForceIndex` (F32) oracle, recall@10 scored per query, and the index's true `memory_usage()` extrapolated to 1M. The corpus is a **Gaussian mixture** (clustered), not uniform-random: uniform high-dimensional vectors are pathological for recall@k (every pair ≈ orthogonal, so top-k is an arbitrary draw from an equidistant shell — recall@10 measured ~0.97 at 10k but ~0.54 at 100k, a corpus-size artifact, not an index regression). Real embeddings cluster on a manifold; the clustered corpus reflects that, so recall@k is meaningful and scale-stable. Details in [usearch-tuning.md](usearch-tuning.md). ### Recall@10 at 1536-D (measured, 100k clustered corpus, vs exact brute-force cosine) | Config (F16) | recall@10 | p99 (raw ANN) | G2 (≥0.95) | G1 (≤10ms) | |---|---|---|---|---| | M=16, ef_c=400, ef_s=128 | 0.9970 | 1.0 ms | ✅ | ✅ | | **M=16, ef_c=400, ef_s=200 (default)** | **0.9970** | **1.4 ms** | ✅ | ✅ | | M=24, ef_c=400, ef_s=200 | 0.9985 | 2.2 ms | ✅ | ✅ | | M=32, ef_c=400, ef_s=200 | 0.9975 | 2.9 ms | ✅ | ✅ | The production default (M=16, ef_c=400, F16) **clears both G1 and G2 at 100k/1536-D** with wide margin. `ef_search` is the latency lever and recall saturates by ef_s=128, so the per-query override (m12p3) lets a latency-sensitive read drop to ef_s=128 (p99 ≈ 1.0 ms) without losing recall. **Live cross-check:** `tidal-stress --verify-recall` against a real server returned recall@10 = 1.0000 at 20k/1536-D (clustered), at the default beam and at `--recall-ef-search 400`. ### Per-1M memory sizing by quantization (true measured footprint, extrapolated from 100k) | Quantization | recall@10 (M=24, ef_c=400, ef_s=400) | mem/1M (GB) | Verdict | |---|---|---|---| | F32 | 0.9985 | 10.53 | exact, 2× RAM | | **F16 (default)** | **0.9960** | **5.53** | −0.25% recall, half RAM — **chosen** | | Int8 | 0.7150 | 3.03 | **rejected**: −28% recall at 1536-D | The footprint is `USearch`'s true `memory_usage()` (graph links + quantized vectors), so F16 is **≈ 5.2–5.5 GB/1M**, not the ~3.4 GB a vectors-only estimate (`1536 × 2 B`) gives — the ~1.8 GB difference is the HNSW graph. Size pods at ≈ 5.5 GB/1M for F16. **Int8's 1.8 GB option is rejected**: it loses a quarter of the neighbours at 1536-D and would need quantization-aware scaling first. ### Per-query `ef_search` is now honored The `ef_search` column in the frontier table is a **per-request** knob as of m12p3 (it was silently ignored before). A low-latency surface can request a narrow beam and a high-recall surface a wide one against the same shared index, concurrently — see [usearch-tuning.md](usearch-tuning.md#per-query-ef_search-m12p3-now-honored). `tidal-stress --verify-recall --recall-ef-search ` sweeps it open-loop. ### The 1M run The 100k numbers above fit one laptop (the F32 oracle is ≈ 0.6 GB). The 1M exit-gate run needs ≈ 6 GB for the oracle plus the HNSW, so it runs on the k3s node: ```bash cargo run --release --example ann_grid_search -- \ --corpus 1000000 --dim 1536 --queries 200 --k 10 ```