From 431340fc3417d84af9a0ea95920bcb0d376df369 Mon Sep 17 00:00:00 2001 From: jordan Date: Sun, 30 Aug 2026 22:41:02 -0600 Subject: [PATCH] docs: never deploy inside the Velero backup window Found by violating it. The staged roll landed at 03:50/03:55/03:56 UTC; velero-fleet-daily fires at 03:30 and takes 9-25 min. Restarting tidaldb-0 cancelled its own in-flight volume backup (podvolumebackup ...-wgtkt, pod=tidaldb-0 volume=data) at 2.7 GB of 5.4 GB, and the parent Backup froze at 3529/3907 items. It did not fail - it sat InProgress for 67+ minutes heading for the 240-minute timeout that produces PartiallyFailed, and a stalled Backup blocks the next scheduled run. That is almost certainly the explanation for the PartiallyFailed runs on 2026-08-17/18/19 and 08-25: the namespace holds 4 Canceled and 4 Failed PVBs, all clustered on exactly those dates. The hazard is invisible from both sides - nothing in the deploy path mentions Velero and nothing in the Velero config mentions deploys - so the warning goes at the TOP of the deploy runbook rather than in a section nobody reaches. Note what caught it: the recalibrated 60-minute in-flight bound from the previous commit. The old binary "any InProgress fails" assertion would have been red every day during the normal window, so a real stall would have looked like the usual noise. Resolved by confirming 20260830033034 was Completed at 3707/3707 and all four PVCs Bound, then deleting the stalled Backup. Suite 34/34. --- docs/runbooks/deploy-verification.md | 17 +++++++++++++++++ 1 file changed, 17 insertions(+) diff --git a/docs/runbooks/deploy-verification.md b/docs/runbooks/deploy-verification.md index fc98b95..c61b114 100644 --- a/docs/runbooks/deploy-verification.md +++ b/docs/runbooks/deploy-verification.md @@ -1,5 +1,22 @@ # Deploy verification +> ## Do NOT deploy between 03:30 and 04:00 UTC +> +> `velero-fleet-daily` fires at **03:30 UTC** and takes **9–25 min** (measured over 15 successful runs). +> A rolling restart inside that window **cancels the in-flight volume backup for every pod it restarts**, and +> the parent Velero `Backup` then stalls indefinitely — it does not fail, it sits `InProgress` and blocks the +> next scheduled run. +> +> Observed 2026-08-31: a roll at 03:50/03:55/03:56 cancelled +> `podvolumebackup …-wgtkt` (`pod=tidaldb-0 volume=data`) at 2.7 GB of 5.4 GB. The Backup froze at +> **3529/3907 items** and was still frozen 67 minutes later, heading for the 240-minute timeout that produces +> `PartiallyFailed`. The same fingerprint — Canceled/Failed PVBs clustered on one date — explains the +> `PartiallyFailed` runs on 2026-08-17, 08-18, 08-19 and 08-25. +> +> If you have already done it: confirm the previous day's backup is `Completed` with full item coverage, then +> delete the stalled Backup so the schedule unblocks. Do not wait for the timeout. + + Walk this top to bottom. Every command here was executed against the live `orchard9-k3sf` deployment and its output recorded — nothing is aspirational.