Commit Graph

11 Commits

Author SHA1 Message Date
jordan
44820d4f81 ci: gate the image on deterministic suites; the rolling upgrade becomes pre-release
All checks were successful
ci/woodpecker/push/woodpecker Pipeline was successful
Four data points settle this. mp_rolling_upgrade_no_loss_no_stall passed in
pipelines #5 and #6 and failed in #9 and #10, all on the same 3 CPU / 6Gi step and
all four ending identically:

  cluster_lifecycle.rs:362  timed out: WAL relay alone must reconverge all three
                            nodes to 1e-6 after the rolling upgrade

The same test on the same commit passes locally in 19.61s. It spawns three real
tidal-server processes — each with its own WAL, HNSW index, gRPC transport and
tokio runtime — and asks them to reconverge to 1e-6 inside 180s, on a node with
~1700m free CPU shared with the production cluster. Two passes and two failures is
a coin flip, and a coin flip that blocks image builds teaches everyone to re-run
until it goes green, which is how a gate stops being one.

I did not raise the budget again. That would be loosening a measured threshold to
hide the hardware, and it is the third time this session that the honest answer
was "the number is right, the environment is the finding".

This is the escalation task 01 prescribed verbatim: a tier-3 three-process test
does not belong on a 4-CPU shared node, so it becomes a documented pre-release
step run where it demonstrably passes. It is in
docs/runbooks/deploy-verification.md with the command and the expected 20s, and it
still runs nightly inside cluster_lifecycle where a flake costs a re-read of the
morning report instead of a blocked release.

The image is now gated by `fast-suites`: 51 tests across nine deterministic
in-process suites, no spawned processes, no convergence budget to starve — so its
verdict means the same thing on a loaded shared node as on a workstation. It
catches the class of regression that actually reached main today: cluster_routes
asserting a wire fabrication deleted hours earlier.

Also removes the last resource anchor that made step order load-bearing. Moving a
step broke an anchor defined on it three times in one session (resources-light,
cargo_env, resources-heavy); with the gate gone the heavy shape has exactly one
consumer, so it is written inline with the reason recorded.
2026-08-31 02:24:43 -06:00
jordan
320d640d13 ci: move the in-process suites to the nightly; the gate keeps its measured budget
Some checks failed
ci/woodpecker/push/woodpecker Pipeline failed
Pipelines #8 and #9 both failed, and the second one told me exactly why once the
nested build's stderr was no longer discarded:

  cluster_lifecycle.rs:362  timed out: WAL relay alone must reconverge all three
                            nodes to 1e-6 after the rolling upgrade

That is verbatim the budget sensitivity this file already documents at the gate:
"at the default budget this test times out on 'WAL relay alone must reconverge all
three nodes to 1e-6' on a loaded machine, and passes in 21s with the raised
budget — the convergence itself is fast, the default just leaves no slack."

So it was not a code regression and not disk. `fast-suites` burned ~872s of
CPU-saturating rustc immediately before the gate, and the gate's 300s/180s budgets
are calibrated for a node that is NOT fresh off fifteen minutes of parallel
compilation. I invalidated the calibration by adding load in front of it.

Raising the budget would be loosening a measured threshold to hide load I
introduced — the exact move this project forbids. Blocking the image build is the
gate's job and it outranks fifteen-minutes-faster feedback on in-process suites,
so the eight suites moved into `nightly-security-ops` (already the light
in-process nightly step) and the gate kept its calibration and its proven shape
from pipelines #5 and #6.

Coverage still goes from NEVER to nightly for all fourteen previously-unscheduled
suites; every one of the 23 now has a runner. CARGO_INCREMENTAL=0 stays: it is
correct in CI regardless, since each workflow gets a fresh 10Gi workspace and
incremental artifacts measured 11G of a 27G target tree.

Also documents an anchoring cost discovered the hard way: the resource shapes are
declared on their first consuming step (to avoid a schema-risky top-level key),
and moving the step that held `&resources-light` broke its aliases. The note now
says to check for anchor definitions before removing a step.
2026-08-31 01:57:27 -06:00
jordan
25361bb660 ci: disable incremental compilation, surface the nested build's error
Some checks failed
ci/woodpecker/push/woodpecker Pipeline failed
Pipeline #8 failed and told us nothing. Both fixes here address that.

WHY IT FAILED: Woodpecker gives each workflow a fresh 10Gi workspace PVC. The
release gate alone fit (pipelines #5 and #6 passed), but the new `fast-suites`
step builds eight test binaries ahead of it, and the gate's nested
`cargo build --features fault-injection` then had no room. CARGO_INCREMENTAL=0
now applies to every Rust step: incremental artifacts are pure waste in CI since
nothing is ever reused across pipelines, and they measured 11G of a 27G target
tree locally — 41%.

WHY IT SAID NOTHING: tests/support/multiproc.rs:1686 built with
`stderr(Stdio::null())`, so the assert fired with "cargo build ... failed" and no
compiler error, no ENOSPC, no exit code. A build failure whose reason is discarded
costs more than the build. stderr is now captured and included in the panic
message; the output is only read on the failure path.

The env block is anchored on its first consuming step rather than a top-level
`variables:` key. I initially used the top-level form and reverted it: the file is
schema-validated and `when: branch: main` means only main triggers a pipeline, so
a rejected key could not be caught on a throwaway branch and would break every
push until reverted. Same rule the resource shapes already follow.
2026-08-31 01:31:03 -06:00
jordan
a6f663f002 harden: validate embeddings before the WAL, fix the reseed-latch leak, run every test suite
Some checks failed
ci/woodpecker/push/woodpecker Pipeline failed
Fixes the two defects a malformed probe exposed on the live cluster, plus the
coverage gap that let a stale assertion survive the same day it was falsified.

TASK 17 — validate before the WAL append. A 128-dim vector against a 1536-dim
slot was appended to the WAL FIRST, then validated, then answered 500 — so an
already-durable, unapplicable record shipped to both followers, halted both
receivers, and put shard 1 into a quorum-write outage. Validation now runs before
the append and returns 400 via invalid_input; nothing enters the log.
`storage::vector::validate_dimensions` is now the single comparison, replacing an
inline duplicate of the same rule in lifecycle/ops.rs:57-62 — two copies of a
dimension check drift, and the apply-path copy is the one that halts replication
when it disagrees.

The receiver's halt-vs-skip decision is now explicit instead of "halt on
anything". A record whose failure is deterministic and node-independent (schema
width) is skipped, counted on blobs_apply_failed_total and ERROR-logged, so the
frontier advances; a record that could become applicable after a binary upgrade
(unknown batch kind, capability skew) still halts, because skipping那 would
silently drop replicated data. Both branches are proven reachable by tests.

TASK 18 — the reseed latch outlived its discharge. A node hosting 3 shard groups
latched a marker per group but discharged on a single seqno, so two latches meant
permanent 503 on a node whose every shard read lag 0 — it hit all three pods
during the roll and each needed a manual delete. Gaps are now tracked per group
in a ReseedGapSet and cleared on evidence about themselves; a REFUSED
reseed_self_restart re-evaluates every 15s instead of waiting for a latch that
never arrives. /health's cause ladder was also lying: it printed "joiner boot not
yet converged" for a node whose groups had all converged, because the fallback
asserted a state it never tested. It now names the outstanding gaps, gained the
decommissioned-by-signal arm that is_ready checked but the ladder did not, and
its terminal arm says "reason unavailable" rather than inventing one.

COVERAGE — 14 of 23 integration suites were run by NO pipeline. Not theoretical:
cluster_routes still asserted the wire fabrication removed hours earlier
(applied_events == 0 with a lag derived from it) and nothing caught it because
nothing ran it. cluster_sharding (dense-rank, /sharded/* opt-in), vector_search
(distance contract) and cluster_poison_embedding (task 17's own gate) were in the
same position, so those guards would have rotted identically. Every suite now has
a runner: 8 in-process ones in a new `fast-suites` push step (measured 71s, runs
FIRST so a cheap failure precedes the 6.5-min gate), 6 multiproc ones in the
nightly. All 23 scheduled; all 4 never-before-run heavy suites verified passing
before being scheduled.

Also fixes cluster_chaos.rs:329, which the nightly's FIRST EVER run caught 13
minutes in — it demanded an unreachable peer report worst-case lag, i.e. it
required the fabrication task 04a deleted.

Verified: fmt clean; clippy 72 vs 73 baseline (one FEWER, zero added, measured on
touched trees at 431340f); lib 2115 passed; all 8 fast suites green;
cluster_chaos 5, cluster_sharding 5, cluster_poison_embedding 1,
cluster_cross_shard_reads 2, cluster_graph_persistence 1, cluster_multiproc 5,
cluster_e2e 2; doc-guard OK.
2026-08-31 00:46:00 -06:00
jordan
fe8d0c87e7 harden: restore CI verification, remove four wire-level fabrications, instrument the 401 path
Implements tmp/tidaldb-fleet-hardening (20 planned tasks + 2 found by measurement).

Ring 0 — restore verification. .woodpecker.yaml step pods ran at the namespace
default of 1500m/2Gi, which OOMKilled a prior pipeline and starved the release
gate past its budget. Both push-path steps now declare
backend_options.kubernetes.resources as two YAML anchors declared once on their
first consuming step. The values are CALIBRATED against measured free node
capacity, not against the LimitRange max: `requests: cpu 2` (this roadmap's
original figure) fits on NO node and would sit Pending forever, because
`ci-build-bounds` grants permission and the nodes supply capacity, and those are
not the same thing.

The `nightly` cron described in this file for 216 days was never created, so
tier-3 chaos, the fault classes, mTLS and the PITR test produced exactly zero
signal while reading like standing coverage. nightly-chaos and
nightly-security-ops now alias the anchors and have budgets matching the gate
(their 120/90 were TIGHTER on the same runner, so they would have failed
nightly for a budget reason, not a correctness one). nightly-soak is REMOVED,
not scheduled: it drives 1000 rps for 600s gating on p99 <= 250ms, and the best
node has 1700m free CPU, so it would fail on starvation rather than regression —
manufacturing a nightly false alarm. Its commands move verbatim to
docs/runbooks/nightly-soak.md.

Ring 1 — four fabrications removed from the wire.
- scatter_merge sorted and truncated without re-stamping rank, so /feed and
  /search returned 1,1,2 under full placement. Reuses merge_cross_shard's
  existing stamp; asserted on BOTH the multi-group merge path and the
  single-group [only] fast path that bypasses it.
- aggregate_region_row's None arm invented `applied_events: 0` plus a deficit
  derived from it. applied_events/lag_events are now Option<u64>, null on the
  wire. leader_last_seq was also unwrap_or(0), so a node that could not reach
  the LEADER computed 0 - applied = 0 for every region and reported a converged
  cluster it had never measured — a fabrication pointing the dangerous way.
- tidalctl inferred NO REPORT from `applied == 0 && lag > 0`. That heuristic was
  actively hiding the PVC-wipe shape: a measured zero with a real deficit
  rendered as "no report" instead of BEHIND. Now read off the wire; converged
  exits 0, partitioned still exits nonzero.
- /sharded/* answered 201/204 for single-copy writes with nothing anywhere
  saying so. Now requires `x-tidal-ack: local`, rejecting with 400 via the
  existing invalid_input path. Six call sites migrated, not the two this
  roadmap predicted — including docs/runbooks/cluster.md §16.3, which told
  operators to run a quorum-write probe via POST /sharded/items. That probe
  cannot verify quorum: the surface applies locally with no WAL append. It was
  used as the safety check between every step of a staged deploy earlier today.

Ring 2 — observability. JSON_LOGS was already implemented and the deployment
simply never asked for it; the StatefulSet now sets it, plus
TIDAL_SERVICE_NAME=tidaldb because enabling it silently renames the
VictoriaLogs `service` stream field and would have blinded every query keyed on
it. Adds tidaldb_usearch_replicated_vectors_total, incremented on BOTH the
origin (wal_blob_first -> Ok(Some)) and the follower apply path — counting only
the origin would mean each vector lands on exactly one node, replicas never
agree, and the alert built on it pages forever.

Found by measurement, not planned: the 401 path discarded every fact about
every rejection. Traefik has served 101,858 rejected requests to the public
ingress — 87.6% of all its traffic — with no record of who or why anywhere.
unauthorized_response now emits reason (missing_token vs invalid_token, the
distinction that separates a scanner from a rotation that missed a consumer)
and the forwarded client. The token is never logged.

Also: scripts/restore-fleet.sh --cluster started the soak monitor while
deliberately leaving its gate suspended, orphaning a watcher that has reported
"0/30 green nights" for 13 days. The pair now moves together. Doc-guard's
three-warning backlog is cleared with real backfill for M4/M6/M12.

Verified: fmt clean; clippy 5 crates 0 new warnings (74 vs 74 baseline,
counted in a detached worktree at HEAD); lib 2110 passed; cluster_sharding 5;
cluster_runbook 10; tidalctl 38; doc-guard 0 warnings. Playwright 32/34 with
the two remaining failures asserting the rank fix against the not-yet-rolled
image — they are the post-deploy proof.
2026-08-30 20:55:58 -06:00
jordan
488aa515c5 ci: give the release gate the budget headroom every nightly step already has
Some checks failed
ci/woodpecker/push/woodpecker Pipeline failed
The push-path release gate (mp_rolling_upgrade_no_loss_no_stall) ran with the
compiled-in defaults of 60s boot / 30s convergence
(tidal-server/tests/support/multiproc.rs:54,62), which are tuned for a developer
machine. It spawns three real OS processes, drives a graceful SIGTERM ->
version-tagged restart -> heal cycle, then waits for three-way feed parity to 1e-6
over loopback gRPC.

Measured today: at the default budget it times out on "WAL relay alone must
reconverge all three nodes to 1e-6 after the rolling upgrade". With
TIDAL_TEST_BOOT_BUDGET_SECS=300 / TIDAL_TEST_CONVERGENCE_BUDGET_SECS=180 it passes
in 21s. So convergence is fast; the default simply leaves no slack. Reproduced
identically on a pre-change baseline (53c345e) in a separate worktree, so this is
budget sensitivity, not a regression from the vector-search or e2e work.

This was the only push-path step without headroom, while every nightly step already
sets it with the comment "Boot / convergence budgets are raised for a shared CI
runner" - and this is the step whose failure BLOCKS the Kaniko image build, so its
flake cost is the highest in the file.

Budgets are overrides, not weakened assertions: the test still demands exact
three-way parity to 1e-6 with no reconcile, and still fails if convergence stalls.
2026-08-30 15:49:44 -06:00
jx12n
1265140e28 feat(m11): continuous correctness (m11p9) — fault classes, invariant checkers, soak gates, nightly pipeline
- fault-injection cargo feature (compiled OUT of prod): slow-fsync + disk-full
  WAL hooks in tidal/src/fault.rs, inert until armed, tier-3 builds with feature
- first-class invariant checkers (tests/support/invariants.rs): AckLedger
  no-acked-loss (now consumed by m11p3 gate), feed parity, single-leader-per-term,
  monotonic frontiers
- cluster_faults.rs tier-3 suite 4/4: disk-full degrade+recover, slow-fsync
  lag+converge, both-slow quorum 503, asymmetric partition no-split-brain
- tidal-stress soak gates: --json-summary + --max-p99-ms/--max-error-pct/
  --fail-on-knee → non-zero exit on regression
- Woodpecker cron nightly flow (chaos + gated soak), event-routed, not GH Actions
- guarantee-traceability.md: roadmap §2 guarantees → named tests (closes G-C
  apparatus; 30-day-green is a calendar criterion)
2026-06-13 15:23:59 -06:00
jx12n
005e292cbb fix(m11): review remediation + tidal-stress perf sweep + perf wave 2
Resolve all BLOCKER/CRITICAL/WARNING findings from the m11p7/p8 review:
- tidalctl restore: safe_join path-traversal/Zip-Slip guard + fsync on write
- corrupt-WAL checkpoint_seq guard; PITR archive-before-delete
- cluster: x-tidal-relayed audit-dedup marker; forward_failures counts 5xx
- mTLS/HTTP-TLS handshake hardening; accept-loop EMFILE backoff
- per-principal rate-limit + node-token marker-pinning tests
- self-heal tier-3 coverage; 5 router-auth tests

tidal-stress: measurement-fidelity fixes (schedule-lag p99/max, exact
feed-over-SLO verdict, shed annotation) + typed Body, workload.next
184ns->68ns, RoundRobin len==1 short-circuit, HeaderValue cache;
new benches/hotpath.rs + lib.rs.

perf wave 2: signal_snapshot SmallVec/SignalKey carrier; one-get-per-type
ranking pre-pass.
2026-06-13 12:28:04 -06:00
jx12n
d5d1e7d81a feat(m11): observability+ops (m11p8) + perf-sweep wave 2 T2
m11p8 closes G-O + §1.4-3:
- Cluster metrics: breaker state, forwards, self-heal on /metrics; multi-shard sibling render (shard="N")
- Grafana cluster row + 8-rule Prometheus alert group
- Request-id / TraceLayer on both cluster routers; id rides forward hop
- Truthful status: flushed leader applied_events frontier; post-promote ShardId(0) keying fix
- Self-driving heal: tick_self_heal re-arms stuck-peer backlog every ~3s
- WAL PITR: wal.archive_dir, archive-before-delete gap-free
- tidalctl backup/restore with BLAKE3 content-hash verification
- Rolling-upgrade build_version handshake (N/N+1, never rejects) + Woodpecker release gate

perf-sweep wave 2 T2: one-get-per-type pre-pass in ranking executor
- signal_values.rs pre-fetches all signal kinds before scoring loop
- Eliminates per-item repeated DashMap lookups: −18.8% for_you, −31% under writes
- Byte-identical output verified with A/B test harness
2026-06-13 09:17:49 -06:00
jx12n
cdbaf4b475 ci(woodpecker): build-only pipeline (drop auto standalone deploy)
The deploy step did 'kubectl set image deployment/tidaldb' on every image
build, coupling image builds to a standalone roll. Deployment is manual via
kustomize from the orchard9-k3sf ops repo (its contract: no CI/CD deploy). The
one tidal-server binary serves both standalone and multi-process 'cluster
--region' modes, so this image covers the m8p10 3/3 cluster deploy too. Adds an
explicit 'm8p10' image tag for deterministic manifest pinning.
2026-06-10 14:33:29 -06:00
jordan
3cc798fc15 ci: add Woodpecker pipeline for tidal-server image build and k8s deploy
Some checks failed
ci/woodpecker/push/woodpecker Pipeline failed
2026-02-28 10:21:37 -07:00