tidaldb/docs/planning/milestone-12/README.md
jordan fe8d0c87e7 harden: restore CI verification, remove four wire-level fabrications, instrument the 401 path
Implements tmp/tidaldb-fleet-hardening (20 planned tasks + 2 found by measurement).

Ring 0 — restore verification. .woodpecker.yaml step pods ran at the namespace
default of 1500m/2Gi, which OOMKilled a prior pipeline and starved the release
gate past its budget. Both push-path steps now declare
backend_options.kubernetes.resources as two YAML anchors declared once on their
first consuming step. The values are CALIBRATED against measured free node
capacity, not against the LimitRange max: `requests: cpu 2` (this roadmap's
original figure) fits on NO node and would sit Pending forever, because
`ci-build-bounds` grants permission and the nodes supply capacity, and those are
not the same thing.

The `nightly` cron described in this file for 216 days was never created, so
tier-3 chaos, the fault classes, mTLS and the PITR test produced exactly zero
signal while reading like standing coverage. nightly-chaos and
nightly-security-ops now alias the anchors and have budgets matching the gate
(their 120/90 were TIGHTER on the same runner, so they would have failed
nightly for a budget reason, not a correctness one). nightly-soak is REMOVED,
not scheduled: it drives 1000 rps for 600s gating on p99 <= 250ms, and the best
node has 1700m free CPU, so it would fail on starvation rather than regression —
manufacturing a nightly false alarm. Its commands move verbatim to
docs/runbooks/nightly-soak.md.

Ring 1 — four fabrications removed from the wire.
- scatter_merge sorted and truncated without re-stamping rank, so /feed and
  /search returned 1,1,2 under full placement. Reuses merge_cross_shard's
  existing stamp; asserted on BOTH the multi-group merge path and the
  single-group [only] fast path that bypasses it.
- aggregate_region_row's None arm invented `applied_events: 0` plus a deficit
  derived from it. applied_events/lag_events are now Option<u64>, null on the
  wire. leader_last_seq was also unwrap_or(0), so a node that could not reach
  the LEADER computed 0 - applied = 0 for every region and reported a converged
  cluster it had never measured — a fabrication pointing the dangerous way.
- tidalctl inferred NO REPORT from `applied == 0 && lag > 0`. That heuristic was
  actively hiding the PVC-wipe shape: a measured zero with a real deficit
  rendered as "no report" instead of BEHIND. Now read off the wire; converged
  exits 0, partitioned still exits nonzero.
- /sharded/* answered 201/204 for single-copy writes with nothing anywhere
  saying so. Now requires `x-tidal-ack: local`, rejecting with 400 via the
  existing invalid_input path. Six call sites migrated, not the two this
  roadmap predicted — including docs/runbooks/cluster.md §16.3, which told
  operators to run a quorum-write probe via POST /sharded/items. That probe
  cannot verify quorum: the surface applies locally with no WAL append. It was
  used as the safety check between every step of a staged deploy earlier today.

Ring 2 — observability. JSON_LOGS was already implemented and the deployment
simply never asked for it; the StatefulSet now sets it, plus
TIDAL_SERVICE_NAME=tidaldb because enabling it silently renames the
VictoriaLogs `service` stream field and would have blinded every query keyed on
it. Adds tidaldb_usearch_replicated_vectors_total, incremented on BOTH the
origin (wal_blob_first -> Ok(Some)) and the follower apply path — counting only
the origin would mean each vector lands on exactly one node, replicas never
agree, and the alert built on it pages forever.

Found by measurement, not planned: the 401 path discarded every fact about
every rejection. Traefik has served 101,858 rejected requests to the public
ingress — 87.6% of all its traffic — with no record of who or why anywhere.
unauthorized_response now emits reason (missing_token vs invalid_token, the
distinction that separates a scanner from a rotation that missed a consumer)
and the forwarded client. The token is never logged.

Also: scripts/restore-fleet.sh --cluster started the soak monitor while
deliberately leaving its gate suspended, orphaning a watcher that has reported
"0/30 green nights" for 13 days. The pair now moves together. Doc-guard's
three-warning backlog is cleared with real backfill for M4/M6/M12.

Verified: fmt clean; clippy 5 crates 0 new warnings (74 vs 74 baseline,
counted in a detached worktree at HEAD); lib 2110 passed; cluster_sharding 5;
cluster_runbook 10; tidalctl 38; doc-guard 0 warnings. Playwright 32/34 with
the two remaining failures asserting the rank fix against the not-yet-rolled
image — they are the post-deploy proof.
2026-08-30 20:55:58 -06:00

4.5 KiB
Raw Permalink Blame History

Milestone 12 · Vector Retrieval at Production Shape ( COMPLETE 2026-06-14 / 2026-06-23)

Milestone summary row: ROADMAP · Milestone Summary, M12. Primary changelog record: CHANGELOG.md ([Unreleased]).

Backfilled index, not a plan. M12 has no ## Milestone 12 prose section in the ROADMAP — only a (detailed) summary-table row — and shipped without a planning directory. These records were assembled after the fact (2026-08-30) from the CHANGELOG, the M12 commits, the docs/profiling/ evidence files, and the shipped test suite. Unlike M4 and M6, M12's numbers are recorded: this milestone's whole thesis was that ranking-at-scale must steer on measured numbers, so it left measurements behind.

What the milestone proves

The feed stays relevant and bounded as the corpus grows past the scan cap. Before M12, for_you and related fell back to scanning an arbitrary low-id slice of the universe, per-query ef_search was accepted and silently ignored, and the recall/latency/memory frontier at the production embedding shape (1536-D) was unmeasured. M12 closed G1 (bounded, relevant candidate generation) and G2 (measured ANN quality), then took the result to a real sharded, mTLS, elastic cluster.

Phases

Phase Name Record Landed
m12p1 Read-recall harness — measurement truth phase-1.md bb21e69 2026-06-14
m12p2 ANN candidate generation in RETRIEVE phase-2.md bb21e69 2026-06-14, cache fix da5d2d4
m12p3 Index tuning + recall/memory frontier phase-3.md bb21e69 2026-06-14
m12p4 Sharded ingestion phase-4.md 31ee612 2026-06-14
m12p5 Idle-readiness convergence phase-5.md aa94fd9 2026-06-14
m12p6 TLS scale-up + HNSW graph persistence phase-6.md 8e39ee1, 4db3f1e, a039955, 727fbfc, 44ec878 2026-06-14→16
Multi-vector user preference modelling multi-vector-preference.md 6a937fc 2026-06-23

The last item is not numbered m12p7: the ROADMAP row lists it as a peer of the six phases without a phase number, and no commit or doc ever called it m12p7. Inventing one would put a fake identifier into the planning record.

Evidence files this milestone left behind

Evidence File
HNSW parameter sweep + F32/F16/Int8 frontier at 1536-D docs/profiling/usearch-tuning.md
Scale baselines, recall-harness section, honest labelling of mean-vs-p99 docs/profiling/scale-baselines.md
Sharded-ingestion real-cluster run (T5) docs/profiling/m12p4-t5-sharded-throughput.md
Idle-readiness fix + elasticity under load (T4) docs/profiling/m12p5-idle-readiness-elasticity.md
k3s deploy + recall findings, the 3-shard bug chain docs/profiling/m12-cluster-deploy-findings.md
Spec updated to shipped reality (per-query ef_search IMPLEMENTED, not deferred) docs/specs/07-vector-retrieval.md
Multi-vector preference design docs/research/multi-vector-preference.md

Tests

tidal/tests/m12p1_vector_recall.rs (3), m12p2_ann_retrieve.rs (5), m12p6_graph_persistence.rs (1), m12_reseed_term_marker.rs (4), m12_preference_event_time.rs (1), plus tidal-server/tests/vector_search.rs (8) and tidal-server/tests/cluster_region.rs (14).

What M12 did NOT close — stated, not buried

The ROADMAP records the honest residue, and it should not be read as complete:

  • The ≥2.5× write-scaling / ≥5,000 quorum-writes-per-second gate remains Ref-A/k3s-pending. m12p4 proved with data why: at fixed per-pod CPU, full-placement sharding scales failover, not write throughput. The ≥2.5× goal belongs to a ≥5-node partitioned-placement topology (Ref-B).
  • The 30-day nightly gate is parked: 23 PASS / 32 FAIL at 200 rps with repeated 4 GiB OOMKills. That is not a production-readiness signal yet.
  • The Ref-A 100k/1536-D memory growth is still unisolated under the corrected 4 GiB request / 6 GiB canary limit.
  • Deferred follow-ups: writer_agent u16 interning on the WAL v3 envelope, and the offline medoid-recluster tier for multi-vector preference.