Implements tmp/tidaldb-fleet-hardening (20 planned tasks + 2 found by measurement). Ring 0 — restore verification. .woodpecker.yaml step pods ran at the namespace default of 1500m/2Gi, which OOMKilled a prior pipeline and starved the release gate past its budget. Both push-path steps now declare backend_options.kubernetes.resources as two YAML anchors declared once on their first consuming step. The values are CALIBRATED against measured free node capacity, not against the LimitRange max: `requests: cpu 2` (this roadmap's original figure) fits on NO node and would sit Pending forever, because `ci-build-bounds` grants permission and the nodes supply capacity, and those are not the same thing. The `nightly` cron described in this file for 216 days was never created, so tier-3 chaos, the fault classes, mTLS and the PITR test produced exactly zero signal while reading like standing coverage. nightly-chaos and nightly-security-ops now alias the anchors and have budgets matching the gate (their 120/90 were TIGHTER on the same runner, so they would have failed nightly for a budget reason, not a correctness one). nightly-soak is REMOVED, not scheduled: it drives 1000 rps for 600s gating on p99 <= 250ms, and the best node has 1700m free CPU, so it would fail on starvation rather than regression — manufacturing a nightly false alarm. Its commands move verbatim to docs/runbooks/nightly-soak.md. Ring 1 — four fabrications removed from the wire. - scatter_merge sorted and truncated without re-stamping rank, so /feed and /search returned 1,1,2 under full placement. Reuses merge_cross_shard's existing stamp; asserted on BOTH the multi-group merge path and the single-group [only] fast path that bypasses it. - aggregate_region_row's None arm invented `applied_events: 0` plus a deficit derived from it. applied_events/lag_events are now Option<u64>, null on the wire. leader_last_seq was also unwrap_or(0), so a node that could not reach the LEADER computed 0 - applied = 0 for every region and reported a converged cluster it had never measured — a fabrication pointing the dangerous way. - tidalctl inferred NO REPORT from `applied == 0 && lag > 0`. That heuristic was actively hiding the PVC-wipe shape: a measured zero with a real deficit rendered as "no report" instead of BEHIND. Now read off the wire; converged exits 0, partitioned still exits nonzero. - /sharded/* answered 201/204 for single-copy writes with nothing anywhere saying so. Now requires `x-tidal-ack: local`, rejecting with 400 via the existing invalid_input path. Six call sites migrated, not the two this roadmap predicted — including docs/runbooks/cluster.md §16.3, which told operators to run a quorum-write probe via POST /sharded/items. That probe cannot verify quorum: the surface applies locally with no WAL append. It was used as the safety check between every step of a staged deploy earlier today. Ring 2 — observability. JSON_LOGS was already implemented and the deployment simply never asked for it; the StatefulSet now sets it, plus TIDAL_SERVICE_NAME=tidaldb because enabling it silently renames the VictoriaLogs `service` stream field and would have blinded every query keyed on it. Adds tidaldb_usearch_replicated_vectors_total, incremented on BOTH the origin (wal_blob_first -> Ok(Some)) and the follower apply path — counting only the origin would mean each vector lands on exactly one node, replicas never agree, and the alert built on it pages forever. Found by measurement, not planned: the 401 path discarded every fact about every rejection. Traefik has served 101,858 rejected requests to the public ingress — 87.6% of all its traffic — with no record of who or why anywhere. unauthorized_response now emits reason (missing_token vs invalid_token, the distinction that separates a scanner from a rotation that missed a consumer) and the forwarded client. The token is never logged. Also: scripts/restore-fleet.sh --cluster started the soak monitor while deliberately leaving its gate suspended, orphaning a watcher that has reported "0/30 green nights" for 13 days. The pair now moves together. Doc-guard's three-warning backlog is cleared with real backfill for M4/M6/M12. Verified: fmt clean; clippy 5 crates 0 new warnings (74 vs 74 baseline, counted in a detached worktree at HEAD); lib 2110 passed; cluster_sharding 5; cluster_runbook 10; tidalctl 38; doc-guard 0 warnings. Playwright 32/34 with the two remaining failures asserting the rank fix against the not-yet-rolled image — they are the post-deploy proof.
54 lines
2.8 KiB
Markdown
54 lines
2.8 KiB
Markdown
# m12p1 — Read-recall harness: measurement truth (✅ COMPLETE 2026-06-14)
|
|
|
|
Landed in `bb21e69`. Changelog: [CHANGELOG.md](../../../CHANGELOG.md).
|
|
Milestone index: [README.md](README.md). Backfilled record.
|
|
|
|
## Why this came first
|
|
|
|
G1 and G2 are numeric gates. Before m12p1 the project had no way to measure
|
|
either: no pure ANN probe (every read path mixed profile scoring, fusion, and
|
|
diversity into the result), and no ground truth to compare against. A gate you
|
|
cannot measure is a gate you cannot pass — so the harness shipped before the
|
|
tuning it exists to judge.
|
|
|
|
## What shipped
|
|
|
|
1. **A pure k-NN probe.** `TidalDb::vector_search_items(query, k, ef_search)`
|
|
returns raw HNSW nearest neighbours over the item content embedding slot with
|
|
**no** profile scoring, fusion, or diversity, so the result measures index
|
|
quality in isolation — the G2 metric. Exposed as `POST /vector_search` on the
|
|
standalone router and on the multi-process region node, which merges by
|
|
distance across hosted shard groups. A dimension-mismatched query is a **400,
|
|
not a 500**: caller error, not server fault.
|
|
2. **The oracle.** `tidal-stress --verify-recall`
|
|
(`tidal-stress/src/recall.rs`) seeds the corpus with deterministic, id-keyed
|
|
embeddings (reproducible with `--skip-seed`), holds a brute-force cosine
|
|
ground truth in RAM, and ramps `/vector_search` probes **open-loop** with
|
|
coordinated-omission correction.
|
|
3. **A verdict, not a number dump.** It reports per-stage **true p99** and mean
|
|
recall@k, plus the **read-knee** — the highest sustained QPS at which
|
|
`p99 ≤ target AND recall@k ≥ target` both hold — with a machine-readable JSON
|
|
summary (`tidal-stress/src/summary.rs`) and a `--fail-on-knee` PASS/FAIL exit
|
|
code. Knobs: `--recall-k`, `--recall-queries`, `--read-p99-target-ms`,
|
|
`--recall-target`, `--recall-ef-search`.
|
|
4. **Repaired fabricated numbers.** The p99 column in
|
|
`docs/profiling/social-scale.md` was mean-as-p99 (the `social` bench is
|
|
Criterion, which reports mean only). That column and the `scale.rs` /
|
|
scale-baselines framing were corrected: every closed-loop number is now
|
|
labelled an *isolated per-op mean (regression tripwire)*, and only the
|
|
open-loop harness signs off p99/recall SLOs.
|
|
|
|
## Evidence
|
|
|
|
- **recall@10 = 0.9997** at 20k items / 1536-D against a real standalone server
|
|
(HNSW M=16 / ef=400 / F16 vs brute-force cosine) — far above the 0.95 G2 target.
|
|
- `tidal/tests/m12p1_vector_recall.rs` (3), `tidal-server/tests/vector_search.rs` (8).
|
|
- Harness section in [docs/profiling/scale-baselines.md](../../profiling/scale-baselines.md).
|
|
- A 1536-D HNSW-vs-brute recall@10 bench added to `tidal/benches/vector.rs`.
|
|
|
|
## Carried constraint
|
|
|
|
The brute-force oracle needs roughly 6 GB RAM at 1M vectors, which is why the
|
|
100k/1M exit-gate runs execute the same harness on the k3s cluster rather than a
|
|
laptop.
|