Implements tmp/tidaldb-fleet-hardening (20 planned tasks + 2 found by measurement). Ring 0 — restore verification. .woodpecker.yaml step pods ran at the namespace default of 1500m/2Gi, which OOMKilled a prior pipeline and starved the release gate past its budget. Both push-path steps now declare backend_options.kubernetes.resources as two YAML anchors declared once on their first consuming step. The values are CALIBRATED against measured free node capacity, not against the LimitRange max: `requests: cpu 2` (this roadmap's original figure) fits on NO node and would sit Pending forever, because `ci-build-bounds` grants permission and the nodes supply capacity, and those are not the same thing. The `nightly` cron described in this file for 216 days was never created, so tier-3 chaos, the fault classes, mTLS and the PITR test produced exactly zero signal while reading like standing coverage. nightly-chaos and nightly-security-ops now alias the anchors and have budgets matching the gate (their 120/90 were TIGHTER on the same runner, so they would have failed nightly for a budget reason, not a correctness one). nightly-soak is REMOVED, not scheduled: it drives 1000 rps for 600s gating on p99 <= 250ms, and the best node has 1700m free CPU, so it would fail on starvation rather than regression — manufacturing a nightly false alarm. Its commands move verbatim to docs/runbooks/nightly-soak.md. Ring 1 — four fabrications removed from the wire. - scatter_merge sorted and truncated without re-stamping rank, so /feed and /search returned 1,1,2 under full placement. Reuses merge_cross_shard's existing stamp; asserted on BOTH the multi-group merge path and the single-group [only] fast path that bypasses it. - aggregate_region_row's None arm invented `applied_events: 0` plus a deficit derived from it. applied_events/lag_events are now Option<u64>, null on the wire. leader_last_seq was also unwrap_or(0), so a node that could not reach the LEADER computed 0 - applied = 0 for every region and reported a converged cluster it had never measured — a fabrication pointing the dangerous way. - tidalctl inferred NO REPORT from `applied == 0 && lag > 0`. That heuristic was actively hiding the PVC-wipe shape: a measured zero with a real deficit rendered as "no report" instead of BEHIND. Now read off the wire; converged exits 0, partitioned still exits nonzero. - /sharded/* answered 201/204 for single-copy writes with nothing anywhere saying so. Now requires `x-tidal-ack: local`, rejecting with 400 via the existing invalid_input path. Six call sites migrated, not the two this roadmap predicted — including docs/runbooks/cluster.md §16.3, which told operators to run a quorum-write probe via POST /sharded/items. That probe cannot verify quorum: the surface applies locally with no WAL append. It was used as the safety check between every step of a staged deploy earlier today. Ring 2 — observability. JSON_LOGS was already implemented and the deployment simply never asked for it; the StatefulSet now sets it, plus TIDAL_SERVICE_NAME=tidaldb because enabling it silently renames the VictoriaLogs `service` stream field and would have blinded every query keyed on it. Adds tidaldb_usearch_replicated_vectors_total, incremented on BOTH the origin (wal_blob_first -> Ok(Some)) and the follower apply path — counting only the origin would mean each vector lands on exactly one node, replicas never agree, and the alert built on it pages forever. Found by measurement, not planned: the 401 path discarded every fact about every rejection. Traefik has served 101,858 rejected requests to the public ingress — 87.6% of all its traffic — with no record of who or why anywhere. unauthorized_response now emits reason (missing_token vs invalid_token, the distinction that separates a scanner from a rotation that missed a consumer) and the forwarded client. The token is never logged. Also: scripts/restore-fleet.sh --cluster started the soak monitor while deliberately leaving its gate suspended, orphaning a watcher that has reported "0/30 green nights" for 13 days. The pair now moves together. Doc-guard's three-warning backlog is cleared with real backfill for M4/M6/M12. Verified: fmt clean; clippy 5 crates 0 new warnings (74 vs 74 baseline, counted in a detached worktree at HEAD); lib 2110 passed; cluster_sharding 5; cluster_runbook 10; tidalctl 38; doc-guard 0 warnings. Playwright 32/34 with the two remaining failures asserting the rank fix against the not-yet-rolled image — they are the post-deploy proof.
65 lines
3.4 KiB
Markdown
65 lines
3.4 KiB
Markdown
# m6p6 — Notification Capping + Adaptive Preferences + Creator Profile Modes + M6 UAT (✅ COMPLETE 2026-02-23)
|
|
|
|
Phase spec and acceptance criteria: [ROADMAP · Milestone 6 · Phase 6](../ROADMAP.md).
|
|
Milestone index: [README.md](README.md). Backfilled record.
|
|
|
|
## What shipped
|
|
|
|
1. **Notification capping** (`tidal/src/db/notification_tracker.rs`).
|
|
`NotificationCaps { max_per_creator_per_day, max_total_per_day }` attaches to
|
|
a query via `RetrieveBuilder::notification_caps(caps)` and is enforced as a
|
|
post-diversity pass, with per-`(user, creator, date)` delivery counts tracked
|
|
so the cap spans queries rather than only trimming one result page.
|
|
2. **Adaptive preference learning rate** (`tidal/src/entities/preference.rs`).
|
|
The EMA alpha decays logarithmically with a user's update count:
|
|
|
|
```
|
|
alpha = base_alpha / (1 + ln(update_count + 1)) // base_alpha default 0.1
|
|
```
|
|
|
|
At `count = 0` this is exactly `base_alpha`. Early signals move a cold user's
|
|
vector hard; later signals refine it gently, which is what keeps a
|
|
well-established taste profile from being yanked by one outlier interaction.
|
|
Unit coverage for the decay curve lives beside the implementation
|
|
(`preference.rs` tests, incl. the explicit `count=0 ⇒ alpha=0.1` case).
|
|
3. **Creator profile modes.** `RetrieveBuilder::for_creator(creator_id)` adds the
|
|
creator filter and restricts candidate generation to that creator's items, so
|
|
`for_creator(x) + for_you` ranks x's catalogue by the *querying user's*
|
|
preferences while `for_creator(x) + hot` ranks x's catalogue by heat.
|
|
4. **The M6 UAT** (`tidal/tests/m6_uat.rs`) — 9 `#[test]` functions over a shared
|
|
fixture, exercising the full use-case surface.
|
|
|
|
## Evidence
|
|
|
|
- `tidal/tests/m6_uat.rs` (9) and `tidal/tests/m6p6_creator_profile.rs` (6).
|
|
- Recorded at close in the ROADMAP status row: 1,082 total (835 lib + 247
|
|
integration), 9 `m6_uat` passing; the same row records that all prior milestone
|
|
UATs (m2, m3, m4, m5, m5p4) continued to pass — the phase's "no regression"
|
|
criterion.
|
|
|
|
## Divergence from the plan
|
|
|
|
**The "update counts are not persisted" limitation is stale — and it was closed
|
|
by M12, not M7.** `tidal/src/entities/preference.rs:37-45` still carries a
|
|
`# Known Limitation` block saying `update_counts` is in-memory only, that every
|
|
user resets to `count = 0` on restart, and that persisting it is "deferred to M7
|
|
(Production Hardening)". Both halves are now wrong:
|
|
|
|
- The count **is** persisted. The multi-vector preference checkpoint format
|
|
serializes it — `[update_count:8 LE][importance_at_anchor:4 LE][anchor_ts:8 LE]`
|
|
(`tidal/src/entities/multi_preference.rs:127,149`) — with
|
|
`MultiPreferenceVectors::checkpoint` / `restore`
|
|
(`multi_preference.rs:642,703`) writing and reading it, and
|
|
`PreferenceVectors::insert_restored` (`preference.rs:239`) loading a legacy
|
|
single-vector row back as a cold-start user with its count intact.
|
|
Round-trip and legacy-row tests exist
|
|
(`checkpoint_restore_roundtrip_multi_cluster`,
|
|
`restore_reads_legacy_single_vector_rows_as_cold_start`).
|
|
- The milestone that closed it was **M12** (multi-vector user preference
|
|
modelling), not M7. See
|
|
[milestone-12/multi-vector-preference.md](../milestone-12/multi-vector-preference.md).
|
|
|
|
The stale doc comment is in Rust source and is out of scope for this backfill;
|
|
recorded here so the next reader of that comment does not re-plan work that is
|
|
already done.
|