tidaldb/docs/planning/milestone-12/phase-3.md
jordan fe8d0c87e7 harden: restore CI verification, remove four wire-level fabrications, instrument the 401 path
Implements tmp/tidaldb-fleet-hardening (20 planned tasks + 2 found by measurement).

Ring 0 — restore verification. .woodpecker.yaml step pods ran at the namespace
default of 1500m/2Gi, which OOMKilled a prior pipeline and starved the release
gate past its budget. Both push-path steps now declare
backend_options.kubernetes.resources as two YAML anchors declared once on their
first consuming step. The values are CALIBRATED against measured free node
capacity, not against the LimitRange max: `requests: cpu 2` (this roadmap's
original figure) fits on NO node and would sit Pending forever, because
`ci-build-bounds` grants permission and the nodes supply capacity, and those are
not the same thing.

The `nightly` cron described in this file for 216 days was never created, so
tier-3 chaos, the fault classes, mTLS and the PITR test produced exactly zero
signal while reading like standing coverage. nightly-chaos and
nightly-security-ops now alias the anchors and have budgets matching the gate
(their 120/90 were TIGHTER on the same runner, so they would have failed
nightly for a budget reason, not a correctness one). nightly-soak is REMOVED,
not scheduled: it drives 1000 rps for 600s gating on p99 <= 250ms, and the best
node has 1700m free CPU, so it would fail on starvation rather than regression —
manufacturing a nightly false alarm. Its commands move verbatim to
docs/runbooks/nightly-soak.md.

Ring 1 — four fabrications removed from the wire.
- scatter_merge sorted and truncated without re-stamping rank, so /feed and
  /search returned 1,1,2 under full placement. Reuses merge_cross_shard's
  existing stamp; asserted on BOTH the multi-group merge path and the
  single-group [only] fast path that bypasses it.
- aggregate_region_row's None arm invented `applied_events: 0` plus a deficit
  derived from it. applied_events/lag_events are now Option<u64>, null on the
  wire. leader_last_seq was also unwrap_or(0), so a node that could not reach
  the LEADER computed 0 - applied = 0 for every region and reported a converged
  cluster it had never measured — a fabrication pointing the dangerous way.
- tidalctl inferred NO REPORT from `applied == 0 && lag > 0`. That heuristic was
  actively hiding the PVC-wipe shape: a measured zero with a real deficit
  rendered as "no report" instead of BEHIND. Now read off the wire; converged
  exits 0, partitioned still exits nonzero.
- /sharded/* answered 201/204 for single-copy writes with nothing anywhere
  saying so. Now requires `x-tidal-ack: local`, rejecting with 400 via the
  existing invalid_input path. Six call sites migrated, not the two this
  roadmap predicted — including docs/runbooks/cluster.md §16.3, which told
  operators to run a quorum-write probe via POST /sharded/items. That probe
  cannot verify quorum: the surface applies locally with no WAL append. It was
  used as the safety check between every step of a staged deploy earlier today.

Ring 2 — observability. JSON_LOGS was already implemented and the deployment
simply never asked for it; the StatefulSet now sets it, plus
TIDAL_SERVICE_NAME=tidaldb because enabling it silently renames the
VictoriaLogs `service` stream field and would have blinded every query keyed on
it. Adds tidaldb_usearch_replicated_vectors_total, incremented on BOTH the
origin (wal_blob_first -> Ok(Some)) and the follower apply path — counting only
the origin would mean each vector lands on exactly one node, replicas never
agree, and the alert built on it pages forever.

Found by measurement, not planned: the 401 path discarded every fact about
every rejection. Traefik has served 101,858 rejected requests to the public
ingress — 87.6% of all its traffic — with no record of who or why anywhere.
unauthorized_response now emits reason (missing_token vs invalid_token, the
distinction that separates a scanner from a rotation that missed a consumer)
and the forwarded client. The token is never logged.

Also: scripts/restore-fleet.sh --cluster started the soak monitor while
deliberately leaving its gate suspended, orphaning a watcher that has reported
"0/30 green nights" for 13 days. The pair now moves together. Doc-guard's
three-warning backlog is cleared with real backfill for M4/M6/M12.

Verified: fmt clean; clippy 5 crates 0 new warnings (74 vs 74 baseline,
counted in a detached worktree at HEAD); lib 2110 passed; cluster_sharding 5;
cluster_runbook 10; tidalctl 38; doc-guard 0 warnings. Playwright 32/34 with
the two remaining failures asserting the rank fix against the not-yet-rolled
image — they are the post-deploy proof.
2026-08-30 20:55:58 -06:00

3.6 KiB
Raw Permalink Blame History

m12p3 — Index tuning + recall/latency/memory frontier ( COMPLETE 2026-06-14)

Landed in bb21e69. Changelog: CHANGELOG.md. Milestone index: README.md. Backfilled record. This is the G2 work.

What shipped

  1. Per-query ef_search — now actually honored. UsearchIndex::search / filtered_search respect a per-request ef_search (tidal/src/storage/vector/usearch_index.rs). Pre-m12p3 the parameter was accepted for trait compliance, logged a warning, and ignored, because USearch 2.24 has no per-call beam argument. The override is race-free via an RwLock epoch guard (with_expansion): searches that agree on ef_search run in parallel under a shared guard, and only a query that changes the live beam width takes the exclusive guard for its (set, search) window — not a per-search mutex. ef_search = 0 selects the slot default. The knob had been plumbed end-to-end in m12p1; m12p3 is what makes it move recall.
  2. Dimension-aware brute-force → HNSW crossover. The exact BruteForceIndex scans every vector under a read lock at count × dim cost, so a fixed 10,000 crossover meant a 15.4M-FMA scan at 1536-D — tens of milliseconds, blocking writers. usearch_min_vectors(dim) (tidal/src/storage/vector/registry.rs) now keeps a brute-force scan within ~4M FMAs: ≈10,000 at or under 128-D (byte-compatible with pre-m12p3 behaviour), ≈2,600 at 1536-D. High-dimension mid-size slots flip to HNSW before the scan blows the SLA rather than after.
  3. memory_usage() on UsearchIndex — the true graph + vector footprint reported by USearch, not the index_stats lower bound, so pod sizing uses a real number.
  4. A grid-search harness. cargo run --release --example ann_grid_search (tidal/examples/ann_grid_search.rs) builds a UsearchIndex plus an exact BruteForceIndex oracle over the same deterministic id-keyed corpus and reports, per (M, ef_construction, ef_search, quantization) point: measured recall@10 vs the oracle, mean and p99 search latency, build time, and true footprint.

Measured frontier (1536-D, 100k clustered corpus, real exact oracle)

The production default M=16, ef_construction=400, F16 clears G1 and G2: recall@10 0.997 at p99 ≈ 1.4 ms raw ANN. ef_search is the latency lever — recall saturates by ef_s=128, where p99 ≈ 1.0 ms. F16 costs 0.25% recall vs F32 for half the RAM (≈5.2 GB per 1M true footprint including the graph). Int8 is rejected at 1536-D: recall 0.715, a 28% loss. Live tidal-stress --verify-recall against a real server measured recall@10 = 1.0000 at 20k/1536-D at both the default beam and --recall-ef-search 400.

Full table: docs/profiling/usearch-tuning.md. 1M command and baselines: docs/profiling/scale-baselines.md.

The measurement artifact this phase caught

The recall corpus is now clustered (a Gaussian mixture) in both the grid harness and tidal-stress (recall::embedding_for). Uniform-random high-dim vectors are pathological for recall@k: under uniform data recall@10 fell from ≈0.97 at 10k to ≈0.54 at 100k. That was a measurement artifact, not an index regression — at high dimension in a uniform shell, recall@10 measures impossible tie-breaking rather than index quality. Every number above uses the clustered corpus.

Spec follow-through

docs/specs/07-vector-retrieval.md was updated in the same wave: per-query ef_search is IMPLEMENTED, not deferred.