The rc4 live deploy converged tidaldb-2 but through ~4 needless self-restarts: a
caught-up shard whose persisted frontier is briefly behind the leader's ADVANCED
baseline (the leader kept writing while the node was down) latches a
snapshot_required marker on the first heartbeat's decide_join, which ARMS a
self-restart. The shard then catches up via the stream — note_term_joined
journals the durable term marker and clear_stale_reseed_marker_if_caught_up
clears the marker — but the already-armed self-restart still fires (Fix 3 defers
it 5s, then exits). The reboot reseeds NOTHING (the leader answers needed=false
for a caught-up shard), so it is futile and flaps readiness.
Fix: gate the self-restart on the marker still being LATCHED at the fire point —
re-check after the (slow) quorum poll and again in the Fix 3 deferred timer
(where the catch-up actually completes within the grace). A marker that healed
via stream catch-up aborts the restart; only a marker that CANNOT self-heal (a
genuine compacted gap, still latched) proceeds to reseed at the next boot. This
converges a caught-up shard IN PLACE (no reboot), while preserving the real
reseed for a genuinely-behind shard.
Verified: tidal-server clippy -D warnings clean; the cluster_reseed e2e
(quarantine reseed still fires, rolling restart still 0-reseed) and the live
rollout confirm the genuine-reseed path is unaffected.