tidaldb/tidal/src/replication/relay.rs
jx12n 5ed2edb211 feat(m11): quorum-acked writes — ack=leader|quorum, commit index, durable frontier reports (m11p3)
ack=quorum gates replicated writes on a majority of the replica set durably
holding them: followers push their durably-applied frontier (ReportApplied,
once per apply round, decoupled from ship acks), the leader folds frontier
reports + ship-ack hints + heal resumes into a leadership-scoped CommitIndex
(k-th-largest durable mark), and handlers await it through an async
watch-channel bridge (zero parked threads per waiter). Honest timeouts:
retryable 503 naming the laggards; x-tidal-seq on every cluster write.
Follower blob applies are batched under group-commit fsyncs (22x seeding).
Exit gate: 167/167 leader-SIGKILL kill points, zero acked-write loss.

Seven-dimension review pass (all confirmed findings fixed):
- WAL blob drain now ABORTS on the first write failure instead of reusing
  the failed seqno mid-drain (a torn record buried mid-segment would
  truncate every later acked record on replay)
- apply_replicated_blobs waits every staged append even after a mid-batch
  failure, parses metadata once, and moves records into Arcs shared with
  the WAL writer (no deep clone per record on the follower apply path)
- CommitIndex: zero-peer fast path now respects demotion (active checked
  under lock before the single-replica return), k-th-largest uses
  select_nth over a reused scratch buffer
- await_quorum: re-reads the index once after the deadline fires (no false
  503 for a write that committed in the race window), warns when the
  commit-watch bridge dies outside shutdown, zero-peer path checks active
- notify_applied report failures: WARN on the first failure of a streak,
  INFO on recovery (a silently stalling frontier reads as unexplained
  quorum 503s); receiver skips re-notifying unadvanced frontiers
- x-tidal-deduplicated: 1 marks dedup-suppressed signal writes (relayed
  through forwards) so durability cursors can tell dedup from no-seqno
- docs: 167/167 kill-point record corrected in CHANGELOG; rolling-upgrade
  order (leader first — a pre-m11p3 leader silently downgrades quorum
  requests to leader-ack) in CHANGELOG + runbook §8; monitoring note for
  report-loss diagnosis on the quorum-timeout alert

Verified: workspace clippy -D warnings (incl. cluster-e2e targets), full
tidaldb/tidal-net/tidal-server/tidalctl suites green, tier-3 multi-process
quorum suite green (8/8 kill points, zero acked loss, partition gate/recover).
2026-06-11 13:28:08 -06:00

1282 lines
50 KiB
Rust

//! Signal-relay primitives shared by the simulated cluster harness and the
//! multi-process region node.
//!
//! A leader's replication stream has four load-bearing pieces that must agree
//! between the in-process test harness ([`crate::testing::cluster::SimulatedCluster`])
//! and the real `tidal-server` region node:
//!
//! 1. a monotonic, gap-free **seqno** counter scoped to the leader's shard;
//! 2. an **event log** that records every staged event so followers can be
//! shipped (eagerly, batched, or re-driven after a partition heals);
//! 3. a **durable frontier**: the highest seqno N such that every event
//! `1..=N` is fsynced on the leader — the only prefix that may ever be
//! shipped (shipping past it could replicate an event the leader then
//! fails to persist: silent leader/follower divergence);
//! 4. the **payload shape** for replicated batches ([`range_payload`] /
//! [`single_event_payload`]).
//!
//! # Write paths (m11p1)
//!
//! The original `write_and_ship` coupled everything: it held the seqno lock
//! across the WAL group-commit fsync (serializing all writers onto solo
//! fsyncs — the measured ~90 writes/s ceiling) and shipped to every peer
//! inline on the request path. The staged path splits it:
//!
//! ```text
//! stage_write → seqno lock { bump; WAL-stage; log push } (microseconds)
//! complete_write → group-commit fsync wait (shared across writers)
//! → in-memory fold → mark_durable(seqno)
//! ShipQueue → per-peer sender threads batch contiguous runs
//! ≤ durable frontier and ship them off the request path
//! ```
//!
//! `write_and_ship` remains for the harness and is now stage + complete +
//! inline eager ship — byte-identical payloads, same invariants, one code
//! path for the atomicity rules.
//!
//! # Durability-failure poisoning
//!
//! A staging failure rolls the seqno back inside the lock (no burned seqno).
//! A *completion* failure cannot roll back — later writers already took higher
//! seqnos — so the relay is **poisoned**: new writes are rejected and the
//! durable frontier stops advancing (the ship queue therefore never ships the
//! failed-or-later events). This is the CockroachDB/Postgres fsync-failure
//! posture: a leader that cannot persist its log must stop leading, not
//! guess. (Post-fsyncgate, retrying a failed fsync proves nothing.)
use std::{
collections::{BTreeSet, HashSet},
sync::{
Mutex, PoisonError,
atomic::{AtomicBool, AtomicU64, Ordering},
},
};
use crate::{
db::{StagedSignal, TidalDb},
replication::{
WalSegmentId,
shard::{RegionId, ShardId},
transport::{Transport, WalSegmentPayload},
},
schema::EntityId,
wal::format::batch::{EventRecord, MAX_EVENTS_PER_BATCH, encode_batch},
};
// ── Relay event log ─────────────────────────────────────────────────────────
/// One staged event recorded in the relay log for shipping and re-delivery.
///
/// Stores the raw [`EventRecord`] (32 bytes) rather than pre-encoded batch
/// bytes so per-peer senders can coalesce contiguous runs into multi-event
/// batches (one RPC per run instead of one per event). Encoding is
/// deterministic (`encode_batch` is a pure function of events + `first_seq` +
/// timestamp), so re-encoding at ship/re-delivery time yields byte-identical
/// payloads to an eager ship of the same run.
#[derive(Clone)]
pub struct RelayEvent {
/// 1-indexed sequence number scoped to `source_shard` (contiguous: the
/// entry at log index `i` always has `seqno == i + 1`).
pub seqno: u64,
/// The replicated event.
pub event: EventRecord,
/// Batch timestamp stamped into the encoded header (the event's own
/// timestamp; informational — WAL ordering is by seqno).
pub batch_ts: u64,
}
// ── Payload construction ───────────────────────────────────────────────────
/// Build the [`WalSegmentPayload`] for a contiguous seqno range
/// `[first, last]` of one source shard.
///
/// The payload's `id.seqno` carries the range's authoritative FIRST seqno and
/// `leader_last_seq` its LAST — exactly the `[first, last]` coverage range the
/// receiver's gap-aware `apply_segment` gates on. Every ship path (eager,
/// batched, re-delivery) builds payloads through here so the
/// `id.seqno == first` / `leader_last_seq == last` invariants cannot drift.
#[must_use]
pub const fn range_payload(
source_shard: ShardId,
first_seq: u64,
last_seq: u64,
event_count: u64,
bytes: Vec<u8>,
) -> WalSegmentPayload {
WalSegmentPayload {
id: WalSegmentId::new(RegionId::SINGLE, source_shard, first_seq),
bytes,
event_count,
leader_last_seq: last_seq,
stream_baseline: 0,
}
}
/// Build the single-event [`WalSegmentPayload`] for one replicated batch
/// (`first == last == seqno`): the eager-ship / re-delivery payload shape.
#[must_use]
pub const fn single_event_payload(
source_shard: ShardId,
seqno: u64,
bytes: Vec<u8>,
) -> WalSegmentPayload {
range_payload(source_shard, seqno, seqno, 1, bytes)
}
/// Encode a contiguous run of relay events into one WAL batch.
///
/// `entries` must be non-empty, seqno-contiguous, and no longer than
/// [`MAX_EVENTS_PER_BATCH`] — guaranteed by construction for slices taken from
/// the relay log (contiguous seqnos) by callers that honor the batch cap.
///
/// # Errors
///
/// Propagates `encode_batch`'s corruption error (empty/oversized run) as
/// `TidalError::Internal`; unreachable for well-formed slices.
pub fn encode_run(entries: &[RelayEvent]) -> crate::Result<Vec<u8>> {
debug_assert!(!entries.is_empty(), "encode_run requires a non-empty run");
debug_assert!(
entries.windows(2).all(|w| w[1].seqno == w[0].seqno + 1),
"encode_run requires seqno-contiguous entries"
);
let events: Vec<EventRecord> = entries.iter().map(|e| e.event.clone()).collect();
encode_batch(&events, entries[0].seqno, entries[0].batch_ts)
.map_err(|e| crate::TidalError::internal("relay_encode_run", e.to_string()))
}
// ── Re-delivery helper ────────────────────────────────────────────────────
/// Re-deliver all log entries not yet applied to `db`, sending through the
/// given transport.
///
/// This is the *local* call shape: the sender owns the follower's [`TidalDb`],
/// so it reads the applied high-water-mark directly. The multi-process server
/// does NOT own the follower's `TidalDb` (the follower is a separate process),
/// so it uses [`SignalRelay::redeliver_to`] with a remotely-reported applied
/// seqno instead. Both share the `seqno > applied` gate and the destination
/// contract below.
///
/// # Load-bearing invariants
///
/// Re-delivery correctness rests on two invariants that hold today but are easy
/// to break silently:
///
/// 1. **Per-source FIFO ordering.** `log` is appended to in strictly increasing
/// `seqno` order per `source_shard` ([`SignalRelay::stage_write`] bumps the
/// seqno and pushes the [`RelayEvent`] inside the same critical section),
/// and we iterate `log` in that same order. The receiver applies via
/// `ReplicationState::apply_range`, which is gap-aware per shard, so
/// re-shipping in seqno order closes any hole the eager/batched path left.
/// The `entry.seqno > applied` gate below is what makes the scan safe to run
/// repeatedly (every poll) against the *same* growing log.
///
/// 2. **Idempotent re-application.** The same batch may be shipped more than
/// once: eagerly/batched by the ship path and again here on every
/// convergence poll and on heal. Re-applying an already-applied `seqno` must
/// be a no-op. We enforce that on the *send* side with the
/// `entry.seqno > applied` check, and the receiver enforces it again on the
/// *apply* side because already-applied ranges are gated out. Both guards
/// are required: dropping the send-side check would flood the transport with
/// redundant payloads; relying on it alone (without the receiver's gate)
/// would double-apply on any race between the `applied_seqno` read here and
/// the receiver thread's advance.
///
/// 3. **Destination is the follower's own shard.** `send_segment`'s first
/// argument is the DESTINATION peer, and a follower's transport only knows
/// its own shard as a peer (the gRPC self-loop wiring in
/// `tidal-server::cluster`). The eager path ships with
/// `send_segment(ShardId(region.0), …)`; re-delivery must match it via
/// `transport.local_shard()` — the follower this transport belongs to.
/// Shipping to `entry.source_shard` (the LEADER's shard) fails the peer
/// lookup on every `GrpcTransport` and recovery silently never happens.
/// `ChannelTransport` ignores the destination, which is exactly why the
/// in-process suite cannot catch that regression — the unit test below
/// locks it instead. The payload itself still carries the SOURCE shard:
/// the receiver advances `applied_seqno` PER SOURCE shard.
pub fn redeliver_missed(
transport: &dyn Transport,
db: &TidalDb,
source_shard: ShardId,
log: &[RelayEvent],
) {
redeliver_to_dest(transport, transport.local_shard(), source_shard, log, |s| {
db.replication_state().applied_seqno(s).unwrap_or(0)
});
}
/// The shared body of both redelivery call shapes.
///
/// `dest` is the destination peer (the follower's own shard); `applied` reports
/// the follower's per-source high-water-mark — read from a local `TidalDb`
/// ([`redeliver_missed`]) or from a remotely-reported seqno
/// ([`SignalRelay::redeliver_to`]). Keeping the gate and the WARN here means the
/// two shapes cannot drift on idempotency, FIFO ordering, or the destination
/// contract.
fn redeliver_to_dest(
transport: &dyn Transport,
dest: ShardId,
source_shard: ShardId,
log: &[RelayEvent],
applied: impl Fn(ShardId) -> u64,
) {
for entry in log {
// Read the follower's high-water-mark for the source shard. The
// `seqno > applied` gate enforces both invariants documented above:
// it skips already-applied batches (idempotency) and, combined with the
// FIFO append order of `log`, ships missing batches in seqno order.
if entry.seqno > applied(source_shard) {
let bytes = match encode_run(std::slice::from_ref(entry)) {
Ok(bytes) => bytes,
Err(e) => {
// Unreachable for a well-formed single-entry run; never
// silently drop a re-delivery if it somehow fires.
tracing::error!(seqno = entry.seqno, error = %e, "re-delivery encode failed");
continue;
}
};
let payload = single_event_payload(source_shard, entry.seqno, bytes);
// Mirror the eager path's observability contract: recovery is
// best-effort, but a swallowed ship error leaves an operator
// staring at a follower that never catches up. WARN, don't drop.
if let Err(e) = transport.send_segment(dest, payload) {
tracing::warn!(
dest_shard = dest.0,
source_shard = source_shard.0,
seqno = entry.seqno,
error = %e,
"re-delivery ship failed; will retry on next heal/convergence pass"
);
}
}
}
}
// ── SignalRelay ────────────────────────────────────────────────────────────
/// A write staged through [`SignalRelay::stage_write`].
///
/// Its seqno is committed and its event is in the relay log, but leader
/// durability (and the in-memory fold) are pending until
/// [`SignalRelay::complete_write`].
#[derive(Debug)]
pub struct StagedRelayWrite {
/// The committed seqno for this write.
pub seqno: u64,
staged: StagedSignal,
}
/// Contiguous durable-prefix tracker: seqnos complete out of relay order (the
/// WAL assigns its own batch order), so completions park in `completed` until
/// the run from `frontier + 1` is contiguous.
#[derive(Default)]
struct DurableFrontier {
frontier: u64,
completed: BTreeSet<u64>,
}
/// One leader's replication stream: monotonic seqno, event log, durable
/// frontier, eager ship, idempotent redelivery.
///
/// Single source of truth for the write/ship atomicity invariants shared by
/// [`SimulatedCluster`] and `tidal-server`'s multi-process region node.
///
/// # Non-negotiable invariants
///
/// 1. **Atomic bump + stage + log-push with rollback.** [`stage_write`] holds
/// the `seqno` lock across the seqno bump, the WAL staging, and the log
/// push. If staging fails, the seqno is rolled back inside the lock, so no
/// other writer ever observes a burned seqno and the log order always
/// equals the seqno order (gap-free, FIFO by construction).
/// 2. **Payloads only via [`range_payload`]/[`single_event_payload`].** Every
/// ship path builds its `WalSegmentPayload` through those helpers, so
/// `id.seqno == first` and `leader_last_seq == last` cannot drift.
/// 3. **Only the durable prefix ships.** The ship queue (and any other
/// batched sender) must not ship past [`durable_seq`]: a not-yet-fsynced
/// event on a follower but not the leader is silent divergence.
/// 4. **Best-effort eager ship.** A failed ship is logged at WARN and never
/// fails the write — the write is already durable on the leader and
/// recorded in the log for re-delivery.
/// 5. **Idempotent, per-source-FIFO redelivery.** [`redeliver_to`] re-ships
/// only `seqno > applied` batches, in log (seqno) order, so re-running it
/// is a no-op once the follower has caught up.
/// 6. **Poison on completion failure.** A durability failure after staging
/// cannot be rolled back; the relay rejects all further writes and freezes
/// the durable frontier (see the module docs).
///
/// [`SimulatedCluster`]: crate::testing::cluster::SimulatedCluster
/// [`stage_write`]: Self::stage_write
/// [`durable_seq`]: Self::durable_seq
/// [`redeliver_to`]: Self::redeliver_to
pub struct SignalRelay {
/// Leader shard this stream is scoped to. Every batch carries it as the
/// source shard and stamps it into the payload.
source_shard: ShardId,
/// 1-indexed monotonic seqno for this leader's stream. Guarded so the
/// bump/stage/log-push region is atomic with respect to concurrent writers.
seqno: Mutex<u64>,
/// Every staged event, in seqno order (entry `i` has `seqno == i + 1`).
log: Mutex<Vec<RelayEvent>>,
/// Mirror of the committed seqno for lock-free reads (status, ship queue).
last_seq_atomic: AtomicU64,
/// Contiguous durable prefix: every seqno `<= durable` is leader-fsynced.
durable: Mutex<DurableFrontier>,
/// Lock-free mirror of `durable.frontier`.
durable_atomic: AtomicU64,
/// Latched on a completion (durability) failure: the stream is dead.
poisoned: AtomicBool,
}
impl SignalRelay {
/// Create an empty relay for the given leader shard (seqno starts at 0).
#[must_use]
pub fn new(source_shard: ShardId) -> Self {
Self {
source_shard,
seqno: Mutex::new(0),
log: Mutex::new(Vec::new()),
last_seq_atomic: AtomicU64::new(0),
durable: Mutex::new(DurableFrontier::default()),
durable_atomic: AtomicU64::new(0),
poisoned: AtomicBool::new(false),
}
}
/// The leader shard this stream ships under.
#[must_use]
pub const fn source_shard(&self) -> ShardId {
self.source_shard
}
/// Whether the relay has been poisoned by a durability failure.
#[must_use]
pub fn is_poisoned(&self) -> bool {
self.poisoned.load(Ordering::Acquire)
}
/// Stage one signal: atomically commit a seqno, submit the event to the
/// leader WAL, and record it in the relay log.
///
/// Holds the seqno lock across the bump, the WAL staging
/// ([`TidalDb::signal_staged`]), and the log push (invariant 1). On a
/// staging failure the seqno is rolled back inside the lock and the error
/// is returned with no state change. The critical section is
/// **microseconds** (validation + one bounded-channel send) — the
/// group-commit fsync wait happens in [`complete_write`](Self::complete_write),
/// outside the lock, so concurrent writers coalesce into shared fsyncs.
///
/// `signal_type_id` is the resolved `u8` WAL id; the relay stays
/// schema-agnostic (the caller resolves name → id and keeps the
/// u8-overflow guard). `signal_type` is the schema name passed through to
/// the WAL staging.
///
/// # Errors
///
/// - `TidalError::Internal` if the relay is poisoned.
/// - Any admission error from `signal_staged` (unknown signal type,
/// invalid weight, WAL backpressure). The seqno is rolled back in every
/// error case, so a retry never burns a seqno.
// The seqno guard is deliberately held across the whole bump + stage +
// log-push region so no other writer observes an in-flight seqno before a
// rollback (atomicity); clippy's drop-tightening suggestion would break
// that invariant.
#[allow(clippy::significant_drop_tightening)]
pub fn stage_write(
&self,
db: &TidalDb,
signal_type_id: u8,
signal_type: &str,
entity_id: EntityId,
weight: f64,
timestamp: crate::schema::Timestamp,
) -> crate::Result<StagedRelayWrite> {
if self.is_poisoned() {
return Err(crate::TidalError::internal(
"relay_stage_write",
"relay poisoned by a leader WAL durability failure; \
the node must be restarted (or leadership moved) before \
accepting replicated writes",
));
}
let event = EventRecord::signal(
entity_id.as_u64(),
signal_type_id,
weight as f32,
timestamp.as_nanos(),
);
let mut slot = self.seqno.lock().unwrap_or_else(PoisonError::into_inner);
*slot += 1;
let seqno = *slot;
let staged = match db.signal_staged(signal_type, entity_id, weight, timestamp) {
Ok(staged) => staged,
Err(e) => {
*slot -= 1;
return Err(e);
}
};
// Log push inside the same critical section: log order == seqno order
// by construction (invariant 1; redelivery FIFO depends on it).
self.log
.lock()
.unwrap_or_else(PoisonError::into_inner)
.push(RelayEvent {
seqno,
event,
batch_ts: timestamp.as_nanos(),
});
self.last_seq_atomic.store(seqno, Ordering::Release);
Ok(StagedRelayWrite { seqno, staged })
}
/// Complete a staged write: block until the event is durable on the
/// leader (group-commit fsync), fold it into the in-memory aggregate, and
/// advance the durable frontier.
///
/// # Errors
///
/// A durability failure **poisons the relay** (invariant 6): the staged
/// seqno is already woven into the stream and cannot be rolled back, so
/// the stream must stop rather than ship a hole or an event the leader
/// lost. Returns the underlying `TidalError::Durability`.
pub fn complete_write(&self, db: &TidalDb, write: StagedRelayWrite) -> crate::Result<u64> {
let StagedRelayWrite { seqno, staged } = write;
match staged.wait(db) {
Ok(_wal_seq) => {
self.mark_durable(seqno);
Ok(seqno)
}
Err(e) => {
self.poison(seqno, &e);
Err(e)
}
}
}
/// Latch the poisoned flag (idempotent) and log the cause loudly.
fn poison(&self, seqno: u64, cause: &crate::TidalError) {
if !self.poisoned.swap(true, Ordering::AcqRel) {
tracing::error!(
seqno,
error = %cause,
"relay POISONED: leader WAL durability failed for a staged \
replicated write; rejecting further writes and freezing the \
ship frontier (restart the node or move leadership)"
);
}
}
/// Record `seqno` as leader-durable and advance the contiguous durable
/// frontier across any parked completions.
fn mark_durable(&self, seqno: u64) {
let mut durable = self.durable.lock().unwrap_or_else(PoisonError::into_inner);
durable.completed.insert(seqno);
while durable
.completed
.first()
.is_some_and(|&next| next == durable.frontier + 1)
{
durable.frontier += 1;
let advanced = durable.frontier;
durable.completed.remove(&advanced);
}
self.durable_atomic
.store(durable.frontier, Ordering::Release);
}
/// The contiguous durable prefix: every seqno `<= durable_seq()` is
/// fsynced on the leader. The ship queue dispatches only up to here
/// (invariant 3).
#[must_use]
pub fn durable_seq(&self) -> u64 {
self.durable_atomic.load(Ordering::Acquire)
}
/// Atomically commit one signal to the local leader db and ship it.
///
/// The eager (harness) path: [`stage_write`](Self::stage_write) +
/// [`complete_write`](Self::complete_write) + an inline best-effort ship
/// to every peer in `peers` except the leader's own shard and any shard in
/// `skip` (e.g. partitioned regions). Returns the committed seqno.
///
/// # Errors
///
/// Any staging error (seqno rolled back, no state change) or completion
/// error (relay poisoned — see [`complete_write`](Self::complete_write)).
#[allow(clippy::too_many_arguments)]
pub fn write_and_ship(
&self,
db: &TidalDb,
signal_type_id: u8,
signal_type: &str,
entity_id: EntityId,
weight: f64,
timestamp: crate::schema::Timestamp,
transport: &dyn Transport,
peers: &[ShardId],
skip: &HashSet<ShardId>,
) -> crate::Result<u64> {
let staged = self.stage_write(
db,
signal_type_id,
signal_type,
entity_id,
weight,
timestamp,
)?;
let seqno = staged.seqno;
let entry = RelayEvent {
seqno,
event: EventRecord::signal(
entity_id.as_u64(),
signal_type_id,
weight as f32,
timestamp.as_nanos(),
),
batch_ts: timestamp.as_nanos(),
};
self.complete_write(db, staged)?;
// Best-effort eager ship to every non-skipped peer (invariant 4),
// strictly after leader durability (invariant 3).
for &peer in peers {
if peer == self.source_shard || skip.contains(&peer) {
continue;
}
let bytes = match encode_run(std::slice::from_ref(&entry)) {
Ok(bytes) => bytes,
Err(e) => {
tracing::error!(seqno, error = %e, "eager ship encode failed");
break;
}
};
// Payload via the shared helper (invariant 2).
let payload = single_event_payload(self.source_shard, seqno, bytes);
if let Err(e) = transport.send_segment(peer, payload) {
tracing::warn!(
peer_shard = peer.0,
source_shard = self.source_shard.0,
seqno,
error = %e,
"eager follower ship failed; leader write is durable and queued for \
re-delivery (redeliver_to)"
);
}
}
Ok(seqno)
}
/// Re-deliver unapplied batches to ONE peer (heal / convergence path).
///
/// `dest` is the destination peer's shard (where `send_segment` routes the
/// payload); the payload itself keeps this relay's source shard so the
/// receiver advances `applied_seqno` for the right origin. `peer_applied`
/// reports the destination's per-source applied seqno — read from a local
/// `TidalDb` in the simulated cluster, or fetched over HTTP from a follower
/// process by the multi-process server. Both shapes flow through the same
/// `seqno > applied` gate (invariant 5), so redelivery is idempotent and
/// per-source FIFO regardless of how `peer_applied` is sourced.
///
/// Only the **durable prefix** is re-delivered (invariant 3): an event the
/// leader has not yet fsynced must never reach a follower.
///
/// Snapshots the log under the lock and releases it before the (possibly
/// blocking) `send_segment` calls, so a slow follower cannot stall a
/// concurrent `stage_write` that must take the same log lock.
pub fn redeliver_to(
&self,
transport: &dyn Transport,
dest: ShardId,
peer_applied: impl Fn(ShardId) -> u64,
) {
let durable = self.durable_seq();
let log_snapshot: Vec<RelayEvent> = self
.log
.lock()
.unwrap_or_else(PoisonError::into_inner)
.iter()
.filter(|e| e.seqno <= durable)
.cloned()
.collect();
redeliver_to_dest(
transport,
dest,
self.source_shard,
&log_snapshot,
peer_applied,
);
}
/// Clone the durable prefix of the log (for harness-side redelivery).
#[must_use]
pub fn snapshot_log(&self) -> Vec<RelayEvent> {
let durable = self.durable_seq();
self.log
.lock()
.unwrap_or_else(PoisonError::into_inner)
.iter()
.filter(|e| e.seqno <= durable)
.cloned()
.collect()
}
/// Clone a contiguous run of **durable** log entries starting at
/// `from_seq`, capped at `max` entries and the wire-format batch limit.
///
/// Returns an empty vector when `from_seq` is past the durable frontier.
/// The relay log is seqno-contiguous from 1 (invariant 1), so the slice is
/// an O(1) index computation.
#[must_use]
pub fn snapshot_range(&self, from_seq: u64, max: usize) -> Vec<RelayEvent> {
let durable = self.durable_seq();
if from_seq == 0 || from_seq > durable {
return Vec::new();
}
let cap = max.min(usize::from(MAX_EVENTS_PER_BATCH)).max(1);
let log = self.log.lock().unwrap_or_else(PoisonError::into_inner);
let start = (from_seq - 1) as usize;
debug_assert!(
log.get(start).is_none_or(|e| e.seqno == from_seq),
"relay log must be seqno-contiguous from 1"
);
let end_seq = durable.min(from_seq + cap as u64 - 1);
let end = end_seq as usize; // exclusive index = end_seq (seqno end_seq is at index end_seq-1)
log.get(start..end)
.map_or_else(Vec::new, <[RelayEvent]>::to_vec)
}
/// The last committed seqno (0 before the first write). Lock-free.
#[must_use]
pub fn last_seq(&self) -> u64 {
self.last_seq_atomic.load(Ordering::Acquire)
}
/// Number of events recorded in the log.
#[must_use]
pub fn log_len(&self) -> usize {
self.log
.lock()
.unwrap_or_else(PoisonError::into_inner)
.len()
}
}
#[cfg(test)]
#[allow(clippy::unwrap_used)]
mod tests {
use std::sync::Mutex as StdMutex;
use super::*;
use crate::{
replication::transport::TransportError,
schema::{DecaySpec, EntityKind, Schema, SchemaBuilder, Timestamp, Window},
};
// ── Moved invariant tests (verbatim contract) ──────────────────────────
/// Locks the load-bearing payload-shape invariants of the single source of
/// truth that every ship path shares: `id.seqno == first`,
/// `leader_last_seq == last`, and the single-event form's
/// `first == last == seqno`. The receiver's gap-aware `apply_segment`
/// gates on exactly these fields; a future edit that breaks either
/// invariant fails here instead of silently diverging between the
/// eager-ship, batched, and recovery sites.
#[test]
fn payload_helpers_hold_replication_invariants() {
let shard = ShardId(3);
let seqno = 42;
let bytes = vec![1u8, 2, 3, 4];
let payload = single_event_payload(shard, seqno, bytes.clone());
assert_eq!(payload.event_count, 1, "must be a single-event batch");
assert_eq!(
payload.leader_last_seq, seqno,
"receiver gating relies on leader_last_seq == seqno"
);
assert_eq!(payload.id.region_id, RegionId::SINGLE);
assert_eq!(payload.id.shard_id, shard);
assert_eq!(payload.id.seqno, seqno);
assert_eq!(payload.bytes, bytes);
let ranged = range_payload(shard, 10, 14, 5, vec![9]);
assert_eq!(ranged.id.seqno, 10, "id.seqno carries the range FIRST");
assert_eq!(
ranged.leader_last_seq, 14,
"leader_last_seq carries the range LAST"
);
assert_eq!(ranged.event_count, 5);
}
/// A transport that records every `(dest, payload-source)` pair. Unlike
/// `ChannelTransport` (which IGNORES the destination — the reason the
/// in-process suite cannot catch a wrong-destination regression), this
/// mock asserts on it.
struct RecordingTransport {
local: ShardId,
sent: StdMutex<Vec<(ShardId, ShardId, u64)>>, // (dest, payload source, seqno)
}
impl Transport for RecordingTransport {
fn send_segment(
&self,
to: ShardId,
payload: WalSegmentPayload,
) -> Result<(), TransportError> {
self.sent
.lock()
.unwrap()
.push((to, payload.id.shard_id, payload.id.seqno));
Ok(())
}
fn recv_segment(&self) -> Option<WalSegmentPayload> {
None
}
fn local_shard(&self) -> ShardId {
self.local
}
}
fn relay_event(seqno: u64) -> RelayEvent {
RelayEvent {
seqno,
event: EventRecord::signal(seqno, 0, 1.0, seqno * 1_000_000),
batch_ts: seqno * 1_000_000,
}
}
/// Re-delivery must ship to the FOLLOWER'S OWN shard (`local_shard`) —
/// never to the source shard (the leader). A follower's gRPC transport
/// only knows itself as a peer (self-loop wiring), so the wrong
/// destination fails the peer lookup and partition recovery silently
/// never happens (the m8p9 heal bug). The payload, by contrast, must keep
/// the LEADER's shard: the receiver advances `applied_seqno` per source.
// The `sent` guard is held across the assertion block on purpose (the test
// reads through it repeatedly); clippy's drop-tightening suggestion would
// change the verbatim moved test body that locks the m8p9 regression.
#[test]
#[allow(clippy::significant_drop_tightening)]
fn redeliver_ships_to_local_shard_with_source_payload() {
let leader_shard = ShardId(0);
let follower_shard = ShardId(2);
let transport = RecordingTransport {
local: follower_shard,
sent: StdMutex::new(Vec::new()),
};
let db = crate::TidalDb::builder()
.ephemeral()
.open()
.expect("open ephemeral db");
let log = vec![relay_event(1), relay_event(2)];
redeliver_missed(&transport, &db, leader_shard, &log);
let sent = transport.sent.lock().unwrap();
assert_eq!(sent.len(), 2, "both unapplied entries must re-ship");
for (dest, payload_source, _seqno) in sent.iter() {
assert_eq!(
*dest, follower_shard,
"re-delivery destination must be the follower's own shard \
(transport.local_shard()), not the leader's"
);
assert_eq!(
*payload_source, leader_shard,
"payload must carry the SOURCE shard so the receiver advances \
applied_seqno for the right origin"
);
}
assert_eq!(
(sent[0].2, sent[1].2),
(1, 2),
"per-source FIFO order must be preserved"
);
}
// ── SignalRelay unit tests ─────────────────────────────────────────────
/// A recording transport that always succeeds and counts ships per
/// destination, used to assert the eager-ship skip set and FIFO order.
struct CountingTransport {
local: ShardId,
sent: StdMutex<Vec<(ShardId, ShardId, u64)>>,
fail: bool,
}
impl CountingTransport {
fn new(local: ShardId) -> Self {
Self {
local,
sent: StdMutex::new(Vec::new()),
fail: false,
}
}
fn failing(local: ShardId) -> Self {
Self {
local,
sent: StdMutex::new(Vec::new()),
fail: true,
}
}
fn ships(&self) -> Vec<(ShardId, ShardId, u64)> {
self.sent.lock().unwrap().clone()
}
}
impl Transport for CountingTransport {
fn send_segment(
&self,
to: ShardId,
payload: WalSegmentPayload,
) -> Result<(), TransportError> {
if self.fail {
return Err(TransportError::Closed);
}
self.sent
.lock()
.unwrap()
.push((to, payload.id.shard_id, payload.id.seqno));
Ok(())
}
fn recv_segment(&self) -> Option<WalSegmentPayload> {
None
}
fn local_shard(&self) -> ShardId {
self.local
}
}
fn view_schema() -> Schema {
let mut b = SchemaBuilder::new();
let _ = b
.signal(
"view",
EntityKind::Item,
DecaySpec::Exponential {
half_life: std::time::Duration::from_secs(7 * 24 * 3600),
},
)
.windows(&[Window::AllTime])
.velocity(false)
.add();
b.build().expect("schema builds")
}
fn leader_db() -> TidalDb {
TidalDb::builder()
.ephemeral()
.with_schema(view_schema())
.open()
.expect("ephemeral leader db opens")
}
/// A clean write bumps the seqno, records exactly one event, advances the
/// durable frontier, and ships to every peer except the source shard and
/// the skip set.
#[test]
fn write_and_ship_commits_and_ships_to_peers() {
let relay = SignalRelay::new(ShardId(0));
let db = leader_db();
let transport = CountingTransport::new(ShardId(0));
let peers = [ShardId(0), ShardId(1), ShardId(2)];
let mut skip: HashSet<ShardId> = HashSet::new();
skip.insert(ShardId(2));
let seqno = relay
.write_and_ship(
&db,
0,
"view",
EntityId::new(7),
1.0,
Timestamp::now(),
&transport,
&peers,
&skip,
)
.expect("clean write succeeds");
assert_eq!(seqno, 1, "first commit is seqno 1");
assert_eq!(relay.last_seq(), 1);
assert_eq!(relay.log_len(), 1, "exactly one event recorded");
assert_eq!(relay.durable_seq(), 1, "completed write is durable");
let ships = transport.ships();
// Source shard 0 and skipped shard 2 excluded → only shard 1 shipped.
assert_eq!(ships.len(), 1, "only the one eligible peer is shipped");
assert_eq!(ships[0].0, ShardId(1), "destination is the eligible peer");
assert_eq!(ships[0].1, ShardId(0), "payload carries the source shard");
assert_eq!(ships[0].2, 1, "payload seqno matches the commit");
}
/// A staging failure (unknown signal type) rolls the seqno back to its
/// prior value and records no event — no burned seqno, no
/// leader-ahead-of-log, no poisoning.
#[test]
fn stage_failure_rolls_back_seqno() {
let relay = SignalRelay::new(ShardId(0));
let db = leader_db();
let transport = CountingTransport::new(ShardId(0));
let peers = [ShardId(1)];
let skip = HashSet::new();
// One good write to advance to seqno 1.
relay
.write_and_ship(
&db,
0,
"view",
EntityId::new(1),
1.0,
Timestamp::now(),
&transport,
&peers,
&skip,
)
.expect("first write succeeds");
assert_eq!(relay.last_seq(), 1);
// Now an unknown signal type fails the staging → rollback.
let err = relay.write_and_ship(
&db,
9, // resolved id is irrelevant; the name is what staging rejects
"does-not-exist",
EntityId::new(2),
1.0,
Timestamp::now(),
&transport,
&peers,
&skip,
);
assert!(err.is_err(), "unknown signal type must fail the write");
assert_eq!(
relay.last_seq(),
1,
"seqno must roll back to its pre-write value (no burned seqno)"
);
assert_eq!(
relay.log_len(),
1,
"a failed write records no event in the log"
);
assert!(
!relay.is_poisoned(),
"a STAGING failure is fully rolled back, never poisons"
);
}
/// A failed eager ship never fails the write: the leader commit is durable
/// and the event is recorded for re-delivery.
#[test]
fn write_and_ship_eager_ship_failure_does_not_fail_write() {
let relay = SignalRelay::new(ShardId(0));
let db = leader_db();
let transport = CountingTransport::failing(ShardId(0));
let peers = [ShardId(1)];
let skip = HashSet::new();
let seqno = relay
.write_and_ship(
&db,
0,
"view",
EntityId::new(3),
1.0,
Timestamp::now(),
&transport,
&peers,
&skip,
)
.expect("write succeeds even when the eager ship fails");
assert_eq!(seqno, 1);
assert_eq!(relay.log_len(), 1, "event is recorded for re-delivery");
assert_eq!(relay.durable_seq(), 1);
}
/// Redelivery is idempotent: once the follower reports the batch applied, a
/// second `redeliver_to` ships nothing.
#[test]
fn redeliver_to_is_idempotent_against_reported_applied() {
let relay = SignalRelay::new(ShardId(0));
let db = leader_db();
// Use a failing eager transport so the only ships come from redelivery.
let eager = CountingTransport::failing(ShardId(0));
let peers = [ShardId(1)];
let skip = HashSet::new();
for i in 1..=3u64 {
relay
.write_and_ship(
&db,
0,
"view",
EntityId::new(i),
1.0,
Timestamp::now(),
&eager,
&peers,
&skip,
)
.expect("write succeeds");
}
let redeliver = CountingTransport::new(ShardId(1));
// Follower has applied nothing yet → all three re-ship.
relay.redeliver_to(&redeliver, ShardId(1), |_| 0);
assert_eq!(redeliver.ships().len(), 3, "all unapplied batches re-ship");
// Follower now reports all three applied → second pass ships nothing.
relay.redeliver_to(&redeliver, ShardId(1), |_| 3);
assert_eq!(
redeliver.ships().len(),
3,
"no additional ships once the follower has caught up (idempotent)"
);
}
/// Redelivery honors a partially-applied follower: only `seqno > applied`
/// batches re-ship, in FIFO order.
#[test]
fn redeliver_to_ships_only_missing_in_fifo_order() {
let relay = SignalRelay::new(ShardId(0));
let db = leader_db();
let eager = CountingTransport::failing(ShardId(0));
let peers = [ShardId(1)];
let skip = HashSet::new();
for i in 1..=4u64 {
relay
.write_and_ship(
&db,
0,
"view",
EntityId::new(i),
1.0,
Timestamp::now(),
&eager,
&peers,
&skip,
)
.expect("write succeeds");
}
let redeliver = CountingTransport::new(ShardId(1));
// Follower applied through seqno 2 → only 3 and 4 re-ship, in order.
relay.redeliver_to(&redeliver, ShardId(1), |_| 2);
let seqs: Vec<u64> = redeliver.ships().iter().map(|s| s.2).collect();
assert_eq!(seqs, vec![3, 4], "only missing batches, in FIFO order");
// Destination is the follower's shard; payload carries the source.
for ship in redeliver.ships() {
assert_eq!(ship.0, ShardId(1), "destination is the peer shard");
assert_eq!(ship.1, ShardId(0), "payload carries the source shard");
}
}
/// Staged writes complete out of relay order without breaking the
/// contiguous durable frontier, and `snapshot_range` only ever exposes the
/// durable prefix in batch-capped contiguous runs.
#[test]
fn durable_frontier_advances_contiguously_and_bounds_snapshots() {
let relay = SignalRelay::new(ShardId(0));
let db = leader_db();
// Stage three writes without completing any.
let w1 = relay
.stage_write(&db, 0, "view", EntityId::new(1), 1.0, Timestamp::now())
.expect("stage 1");
let w2 = relay
.stage_write(&db, 0, "view", EntityId::new(2), 1.0, Timestamp::now())
.expect("stage 2");
let w3 = relay
.stage_write(&db, 0, "view", EntityId::new(3), 1.0, Timestamp::now())
.expect("stage 3");
assert_eq!(relay.last_seq(), 3);
assert_eq!(relay.durable_seq(), 0, "nothing durable before completion");
assert!(
relay.snapshot_range(1, 256).is_empty(),
"snapshot_range must not expose staged-but-not-durable events"
);
// Complete out of order: 2 first — frontier must NOT advance past the
// missing 1.
relay.complete_write(&db, w2).expect("complete 2");
assert_eq!(relay.durable_seq(), 0, "gap at 1 holds the frontier");
relay.complete_write(&db, w1).expect("complete 1");
assert_eq!(relay.durable_seq(), 2, "1 then parked 2 drain together");
relay.complete_write(&db, w3).expect("complete 3");
assert_eq!(relay.durable_seq(), 3);
let run = relay.snapshot_range(1, 2);
assert_eq!(run.len(), 2, "cap honored");
assert_eq!((run[0].seqno, run[1].seqno), (1, 2), "contiguous from 1");
let rest = relay.snapshot_range(3, 256);
assert_eq!(rest.len(), 1);
assert_eq!(rest[0].seqno, 3);
assert!(relay.snapshot_range(4, 256).is_empty(), "past the frontier");
}
/// `encode_run` produces bytes the receiver decodes to the same `[first,
/// last]` range, and is deterministic (re-delivery byte-equals eager ship).
#[test]
fn encode_run_is_deterministic_and_range_correct() {
let entries: Vec<RelayEvent> = (5..=8).map(relay_event).collect();
let a = encode_run(&entries).expect("encode");
let b = encode_run(&entries).expect("encode again");
assert_eq!(a, b, "encoding must be deterministic");
let (header, events) =
crate::wal::format::batch::decode_batch(&a).expect("round-trips through decode");
assert_eq!(header.first_seq, 5);
assert_eq!(header.event_count, 4);
assert_eq!(events.len(), 4);
assert_eq!(events[0].entity_id, 5);
}
// ── Property test: interleaved write_and_ship + redeliver_to ────────────
use proptest::prelude::*;
/// Model the receiver as a monotonic per-source max: applying a payload
/// advances the high-water-mark to `seqno` only when `seqno == applied + 1`
/// (gap-free), and is a no-op for any `seqno <= applied` (idempotent) or a
/// gap (`seqno > applied + 1`, which the real receiver parks).
#[derive(Default)]
struct ModelReceiver {
applied: u64,
}
impl ModelReceiver {
fn apply(&mut self, seqno: u64) {
if seqno == self.applied + 1 {
self.applied = seqno;
}
}
}
/// A transport that drives a shared model receiver and counts double-applies
/// (a payload whose seqno is `<= applied` at receive time but is not a clean
/// re-ship). The model receiver is monotonic, so a correct relay never makes
/// it regress or apply a seqno twice.
struct ModelTransport {
local: ShardId,
receiver: StdMutex<ModelReceiver>,
delivered: StdMutex<Vec<u64>>, // every accepted (advancing) apply
}
impl ModelTransport {
/// Read the follower's applied seqno, releasing the guard immediately.
fn applied(&self) -> u64 {
self.receiver.lock().unwrap().applied
}
}
impl Transport for ModelTransport {
fn send_segment(
&self,
_to: ShardId,
payload: WalSegmentPayload,
) -> Result<(), TransportError> {
let mut rx = self.receiver.lock().unwrap();
let before = rx.applied;
rx.apply(payload.id.seqno);
let advanced = rx.applied > before;
let now = rx.applied;
drop(rx);
if advanced {
self.delivered.lock().unwrap().push(now);
}
Ok(())
}
fn recv_segment(&self) -> Option<WalSegmentPayload> {
None
}
fn local_shard(&self) -> ShardId {
self.local
}
}
proptest! {
// Modest case count keeps the suite fast (project rule: slow tests are
// bugs). Each case interleaves writes and redelivers.
#![proptest_config(ProptestConfig::with_cases(128))]
/// Any interleaving of `write_and_ship` (eager ship failing, so the only
/// successful delivery is via redelivery) and `redeliver_to` preserves
/// per-source FIFO and never double-applies: the model receiver advances
/// 1, 2, …, N exactly once each, ending at the relay's last_seq.
#[test]
fn prop_interleaving_preserves_fifo_no_double_apply(
ops in prop::collection::vec(any::<bool>(), 0..40),
) {
let relay = SignalRelay::new(ShardId(0));
let db = leader_db();
// Eager ship always fails so deliveries come only from redeliver_to,
// letting us interleave the two paths under test deterministically.
let eager = ModelTransport {
local: ShardId(0),
receiver: StdMutex::new(ModelReceiver::default()),
delivered: StdMutex::new(Vec::new()),
};
// The real delivery transport drives the model receiver.
let deliver = ModelTransport {
local: ShardId(1),
receiver: StdMutex::new(ModelReceiver::default()),
delivered: StdMutex::new(Vec::new()),
};
let peers = [ShardId(1)];
let skip = HashSet::new();
let mut writes = 0u64;
for op in ops {
if op {
// write: bump seqno + record event. Eager ship fails (model
// receiver on `eager` never advances) so nothing is delivered
// here — delivery is exclusively via redeliver_to below.
writes += 1;
relay
.write_and_ship(
&db,
0,
"view",
EntityId::new(writes),
1.0,
Timestamp::now(),
&eager,
&peers,
&skip,
)
.expect("write succeeds");
} else {
// redeliver against the live model applied seqno.
let applied = deliver.applied();
relay.redeliver_to(&deliver, ShardId(1), |_| applied);
}
}
// Final drain so every written batch reaches the follower.
loop {
let applied = deliver.applied();
if applied >= writes {
break;
}
relay.redeliver_to(&deliver, ShardId(1), |_| applied);
let after = deliver.applied();
// A redeliver pass that cannot advance the follower would loop
// forever — that is itself a FIFO/gap bug, so assert progress.
prop_assert!(
after > applied || writes == applied,
"redeliver made no progress: applied={applied} writes={writes}"
);
}
// The model receiver advanced 1..=writes exactly once each (FIFO, no
// double-apply): every accepted apply is strictly increasing by 1.
let delivered = deliver.delivered.lock().unwrap().clone();
let expected: Vec<u64> = (1..=writes).collect();
prop_assert_eq!(delivered, expected, "FIFO, gap-free, no double-apply");
prop_assert_eq!(relay.last_seq(), writes, "last_seq tracks commits");
}
}
}