The push-path release gate (mp_rolling_upgrade_no_loss_no_stall) ran with the
compiled-in defaults of 60s boot / 30s convergence
(tidal-server/tests/support/multiproc.rs:54,62), which are tuned for a developer
machine. It spawns three real OS processes, drives a graceful SIGTERM ->
version-tagged restart -> heal cycle, then waits for three-way feed parity to 1e-6
over loopback gRPC.
Measured today: at the default budget it times out on "WAL relay alone must
reconverge all three nodes to 1e-6 after the rolling upgrade". With
TIDAL_TEST_BOOT_BUDGET_SECS=300 / TIDAL_TEST_CONVERGENCE_BUDGET_SECS=180 it passes
in 21s. So convergence is fast; the default simply leaves no slack. Reproduced
identically on a pre-change baseline (53c345e) in a separate worktree, so this is
budget sensitivity, not a regression from the vector-search or e2e work.
This was the only push-path step without headroom, while every nightly step already
sets it with the comment "Boot / convergence budgets are raised for a shared CI
runner" - and this is the step whose failure BLOCKS the Kaniko image build, so its
flake cost is the highest in the file.
Budgets are overrides, not weakened assertions: the test still demands exact
three-way parity to 1e-6 with no reconcile, and still fails if convergence stalls.
The deploy step did 'kubectl set image deployment/tidaldb' on every image
build, coupling image builds to a standalone roll. Deployment is manual via
kustomize from the orchard9-k3sf ops repo (its contract: no CI/CD deploy). The
one tidal-server binary serves both standalone and multi-process 'cluster
--region' modes, so this image covers the m8p10 3/3 cluster deploy too. Adds an
explicit 'm8p10' image tag for deterministic manifest pinning.