Implements tmp/tidaldb-fleet-hardening (20 planned tasks + 2 found by measurement). Ring 0 — restore verification. .woodpecker.yaml step pods ran at the namespace default of 1500m/2Gi, which OOMKilled a prior pipeline and starved the release gate past its budget. Both push-path steps now declare backend_options.kubernetes.resources as two YAML anchors declared once on their first consuming step. The values are CALIBRATED against measured free node capacity, not against the LimitRange max: `requests: cpu 2` (this roadmap's original figure) fits on NO node and would sit Pending forever, because `ci-build-bounds` grants permission and the nodes supply capacity, and those are not the same thing. The `nightly` cron described in this file for 216 days was never created, so tier-3 chaos, the fault classes, mTLS and the PITR test produced exactly zero signal while reading like standing coverage. nightly-chaos and nightly-security-ops now alias the anchors and have budgets matching the gate (their 120/90 were TIGHTER on the same runner, so they would have failed nightly for a budget reason, not a correctness one). nightly-soak is REMOVED, not scheduled: it drives 1000 rps for 600s gating on p99 <= 250ms, and the best node has 1700m free CPU, so it would fail on starvation rather than regression — manufacturing a nightly false alarm. Its commands move verbatim to docs/runbooks/nightly-soak.md. Ring 1 — four fabrications removed from the wire. - scatter_merge sorted and truncated without re-stamping rank, so /feed and /search returned 1,1,2 under full placement. Reuses merge_cross_shard's existing stamp; asserted on BOTH the multi-group merge path and the single-group [only] fast path that bypasses it. - aggregate_region_row's None arm invented `applied_events: 0` plus a deficit derived from it. applied_events/lag_events are now Option<u64>, null on the wire. leader_last_seq was also unwrap_or(0), so a node that could not reach the LEADER computed 0 - applied = 0 for every region and reported a converged cluster it had never measured — a fabrication pointing the dangerous way. - tidalctl inferred NO REPORT from `applied == 0 && lag > 0`. That heuristic was actively hiding the PVC-wipe shape: a measured zero with a real deficit rendered as "no report" instead of BEHIND. Now read off the wire; converged exits 0, partitioned still exits nonzero. - /sharded/* answered 201/204 for single-copy writes with nothing anywhere saying so. Now requires `x-tidal-ack: local`, rejecting with 400 via the existing invalid_input path. Six call sites migrated, not the two this roadmap predicted — including docs/runbooks/cluster.md §16.3, which told operators to run a quorum-write probe via POST /sharded/items. That probe cannot verify quorum: the surface applies locally with no WAL append. It was used as the safety check between every step of a staged deploy earlier today. Ring 2 — observability. JSON_LOGS was already implemented and the deployment simply never asked for it; the StatefulSet now sets it, plus TIDAL_SERVICE_NAME=tidaldb because enabling it silently renames the VictoriaLogs `service` stream field and would have blinded every query keyed on it. Adds tidaldb_usearch_replicated_vectors_total, incremented on BOTH the origin (wal_blob_first -> Ok(Some)) and the follower apply path — counting only the origin would mean each vector lands on exactly one node, replicas never agree, and the alert built on it pages forever. Found by measurement, not planned: the 401 path discarded every fact about every rejection. Traefik has served 101,858 rejected requests to the public ingress — 87.6% of all its traffic — with no record of who or why anywhere. unauthorized_response now emits reason (missing_token vs invalid_token, the distinction that separates a scanner from a rotation that missed a consumer) and the forwarded client. The token is never logged. Also: scripts/restore-fleet.sh --cluster started the soak monitor while deliberately leaving its gate suspended, orphaning a watcher that has reported "0/30 green nights" for 13 days. The pair now moves together. Doc-guard's three-warning backlog is cleared with real backfill for M4/M6/M12. Verified: fmt clean; clippy 5 crates 0 new warnings (74 vs 74 baseline, counted in a detached worktree at HEAD); lib 2110 passed; cluster_sharding 5; cluster_runbook 10; tidalctl 38; doc-guard 0 warnings. Playwright 32/34 with the two remaining failures asserting the rank fix against the not-yet-rolled image — they are the post-deploy proof.
4.1 KiB
Soak: pre-release load and regression gate
Run this before tagging a release. It was a nightly-soak step in
.woodpecker.yaml until 2026-08-30 and moved here unchanged — the commands
below are the step's commands verbatim, so the coverage survives in a runnable
form rather than only in git history.
Why this is not a CI step
The step never actually ran: the nightly cron it was gated on was never
created, so for 216 days it produced zero signal while reading like standing
coverage.
When the cron was finally configured, this step was deliberately excluded, and the reason is a measurement rather than a preference. The soak drives 1000 rps for 600s and fails the build if p99 > 250ms or errors > 1%. Measured free CPU on the k3s cluster, 2026-08-30:
| node | allocatable CPU | free CPU (requests) | free memory |
|---|---|---|---|
k3s-agent-1 |
4000m | 1700m | 4193Mi |
k3s-server-1 |
3000m | 435m | 2309Mi |
k3s-server-2 |
3000m | 550m | 2257Mi |
A p99 gate of 250ms cannot be met from 1700m of contended CPU shared with production tidalDB. The step would fail nightly on starvation, not regression — a false alarm every morning, which is the fake-coverage defect inverted rather than fixed. A gate that cannot distinguish its own failure mode from the thing it is watching for is not a gate.
So it runs here, by hand, on hardware where its numbers mean something.
Where to run it
Anywhere with ≥ 2 dedicated cores and no co-tenant under load. A developer workstation qualifies; the shared k3s nodes do not. If you only have the k3s cluster, the honest options are to raise its thresholds to match the hardware (and say so in the output) or to skip it and record that you skipped it — not to run it and read the result as meaningful.
Run
export TIDAL_SOAK_RPS=1000
export TIDAL_SOAK_SECS=600
export TIDAL_SOAK_MAX_P99_MS=250
export TIDAL_SOAK_MAX_ERROR_PCT=1
# Soak target port — an UNCLAIMED slot in the project's reserved dev band
# (59520-59529): 59520=site, 59521=iknowyou, so the soak uses 59526.
export TIDAL_SOAK_PORT=59526
cargo build -p tidal-server -p tidal-stress
PORT="${TIDAL_SOAK_PORT}"
TARGET="${TIDAL_SOAK_TARGET:-http://127.0.0.1:$PORT}"
if [ -z "$TIDAL_SOAK_TARGET" ]; then
mkdir -p /tmp/soak-data
./target/debug/tidal-server standalone --listen "127.0.0.1:$PORT" \
--schema tidal-server/config/default-schema.yaml --data-dir /tmp/soak-data &
SRV=$!
# Reap the background server + scratch dir on ANY exit (success, gate failure,
# or the boot-failure exit below) so a re-run is idempotent.
trap 'kill "$SRV" 2>/dev/null; rm -rf /tmp/soak-data' EXIT
up=0
for i in $(seq 1 100); do
curl -sf "$TARGET/health/startup" >/dev/null 2>&1 && { up=1; break; } || sleep 0.3
done
# Fail FAST and unambiguously on a boot failure (bad schema, port already
# bound) rather than soaking a dead target and mislabeling it as an error-rate
# regression on the trend line.
if [ "$up" != "1" ]; then echo "soak target failed to start at $TARGET"; exit 1; fi
fi
./target/debug/tidal-stress --target "$TARGET" \
--ramp "${TIDAL_SOAK_RPS}:${TIDAL_SOAK_SECS}" --corpus 5000 --mix peach \
--json-summary soak-summary.json \
--max-error-pct "${TIDAL_SOAK_MAX_ERROR_PCT}" --max-p99-ms "${TIDAL_SOAK_MAX_P99_MS}" --fail-on-knee
cat soak-summary.json
The
$${VAR}escaping in the original CI step is not needed here. Woodpecker preprocesses a bare${VAR}before the shell sees it, so the step had to write$${VAR}to pass a literal through. In a plain shell the single${VAR}form above is correct.
What to expect
soak-summary.json is the artifact — keep it with the release notes so there is
a trend line rather than a single opinion. A --fail-on-knee failure means
throughput stopped scaling before the target rps, which is a different finding
from a p99 breach and worth reporting separately.
Targeting the live cluster
Setting TIDAL_SOAK_TARGET to the live Ref-A cluster and TIDAL_SOAK_SECS=3600
gives the GA-bar 1-hour 100k-DAU soak. This soaks production at 1000 rps.
Do not set it casually, and never as part of an unattended run.