tidaldb/tidal/src/wal/writer.rs
jx12n 5ed2edb211 feat(m11): quorum-acked writes — ack=leader|quorum, commit index, durable frontier reports (m11p3)
ack=quorum gates replicated writes on a majority of the replica set durably
holding them: followers push their durably-applied frontier (ReportApplied,
once per apply round, decoupled from ship acks), the leader folds frontier
reports + ship-ack hints + heal resumes into a leadership-scoped CommitIndex
(k-th-largest durable mark), and handlers await it through an async
watch-channel bridge (zero parked threads per waiter). Honest timeouts:
retryable 503 naming the laggards; x-tidal-seq on every cluster write.
Follower blob applies are batched under group-commit fsyncs (22x seeding).
Exit gate: 167/167 leader-SIGKILL kill points, zero acked-write loss.

Seven-dimension review pass (all confirmed findings fixed):
- WAL blob drain now ABORTS on the first write failure instead of reusing
  the failed seqno mid-drain (a torn record buried mid-segment would
  truncate every later acked record on replay)
- apply_replicated_blobs waits every staged append even after a mid-batch
  failure, parses metadata once, and moves records into Arcs shared with
  the WAL writer (no deep clone per record on the follower apply path)
- CommitIndex: zero-peer fast path now respects demotion (active checked
  under lock before the single-replica return), k-th-largest uses
  select_nth over a reused scratch buffer
- await_quorum: re-reads the index once after the deadline fires (no false
  503 for a write that committed in the race window), warns when the
  commit-watch bridge dies outside shutdown, zero-peer path checks active
- notify_applied report failures: WARN on the first failure of a streak,
  INFO on recovery (a silently stalling frontier reads as unexplained
  quorum 503s); receiver skips re-notifying unadvanced frontiers
- x-tidal-deduplicated: 1 marks dedup-suppressed signal writes (relayed
  through forwards) so durability cursors can tell dedup from no-seqno
- docs: 167/167 kill-point record corrected in CHANGELOG; rolling-upgrade
  order (leader first — a pre-m11p3 leader silently downgrades quorum
  requests to leader-ack) in CHANGELOG + runbook §8; monitoring note for
  report-loss diagnosis on the quorum-timeout alert

Verified: workspace clippy -D warnings (incl. cluster-e2e targets), full
tidaldb/tidal-net/tidal-server/tidalctl suites green, tier-3 multi-process
quorum suite green (8/8 kill points, zero acked loss, partition gate/recover).
2026-06-11 13:28:08 -06:00

959 lines
42 KiB
Rust

use std::{
path::PathBuf,
time::{Duration, Instant},
};
use crossbeam::channel::Receiver;
use std::sync::Arc;
use super::{
dedup::DedupWindow,
error::WalError,
feed::{FlushedBatch, WalShipFeed},
format::{self, BlobRecord, EventRecord, MAX_EVENTS_PER_BATCH, SessionSeqNo, SessionWalEvent},
segment::{self, SegmentWriter},
session_journal::SessionJournal,
};
use crate::replication::{RegionId, ShardId};
/// A single queued append: the event plus the caller's reply channel.
///
/// The reply channel receives the assigned sequence number once the batch
/// containing this event is durably fsynced, the dedup sentinel `0` if the
/// event was suppressed as a duplicate, or the write error if the batch failed
/// to persist. It is critical that *every* queued append eventually resolves
/// its reply — dropping the sender silently surfaces as `Closed` to a caller
/// that is in fact still durable-or-not, so both the steady-state loop and the
/// shutdown drain funnel through the same [`flush_batch`] routine.
type QueuedAppend = (
EventRecord,
crossbeam::channel::Sender<Result<u64, WalError>>,
);
/// A queued blob (kind-1/2) append: the record plus its reply channel.
///
/// Same reply contract as [`QueuedAppend`] — every queued blob eventually
/// resolves its reply with its assigned seqno or the flush error. Blobs skip
/// the dedup window (they are idempotent upserts; a duplicate apply is
/// harmless) and never return the dedup sentinel. The record rides in an
/// `Arc` so batch stagers (the follower's replicated blob apply) share one
/// buffer with the writer by refcount instead of deep-cloning every
/// metadata/embedding payload across the channel.
type QueuedBlob = (
Arc<BlobRecord>,
crossbeam::channel::Sender<Result<u64, WalError>>,
);
// Narrow test-only fault hook: when set, the next `flush_batch` write fails.
//
// Lets a unit test drive the *error* arm of `flush_batch` deterministically
// (without a real I/O fault) to prove that both the steady-state loop and the
// shutdown-drain tail notify every waiting caller with an error instead of
// dropping their reply channels. Thread-local so concurrent tests don't
// interfere. Always compiled out of production builds.
#[cfg(test)]
thread_local! {
static FAIL_NEXT_FLUSH: std::cell::Cell<bool> = const { std::cell::Cell::new(false) };
}
/// Arm the test-only flush-failure hook for the current thread.
#[cfg(test)]
pub(crate) fn arm_flush_failure() {
FAIL_NEXT_FLUSH.with(|c| c.set(true));
}
/// Returns and clears the test-only flush-failure flag for the current thread.
#[cfg(test)]
fn take_flush_failure() -> bool {
FAIL_NEXT_FLUSH.with(|c| c.replace(false))
}
/// Commands sent from `WalHandle` to the writer thread.
pub enum WalCommand {
/// Append a signal event. The reply channel receives the assigned
/// sequence number (or an error) once the batch containing this
/// event has been durably fsynced.
Append {
event: EventRecord,
reply: crossbeam::channel::Sender<Result<u64, WalError>>,
},
/// Append a blob (item-metadata / embedding) record as its own
/// single-seqno batch (m11p2: items and embeddings ride the one
/// replicated log). The reply receives the assigned seqno once the blob
/// batch is durably fsynced.
AppendBlob {
record: Arc<BlobRecord>,
reply: crossbeam::channel::Sender<Result<u64, WalError>>,
},
/// Delete segments whose first sequence number is less than `before_seq`.
/// Runs inside the writer thread to avoid racing with concurrent writes.
TruncateBefore {
before_seq: u64,
reply: crossbeam::channel::Sender<Result<(), WalError>>,
},
/// Graceful shutdown: flush remaining events and exit.
Shutdown,
// ── Session lifecycle commands ────────────────────────────────────────
// These are fire-and-forget (no reply channel). They bypass the signal
// batch system and write directly to the session journal with fsync.
/// Record that a session was started.
SessionStart {
session_id: u64,
user_id: u64,
started_at_ns: u64,
agent_id: String,
policy_name: String,
},
/// Record that a signal was written within a session.
SessionSignal {
session_id: u64,
entity_id: u64,
weight: f64,
ts_ns: u64,
signal_name: String,
annotation: Option<String>,
/// Monotonic sequence number for this write (used for idempotent replay).
/// `None` for legacy writes that predate the seqno mechanism.
session_seqno: Option<u64>,
/// BLAKE3-derived idempotency key for duplicate suppression.
/// `None` for legacy writes.
idempotency_key: Option<u128>,
},
/// Record that a session was closed.
SessionClose { session_id: u64 },
}
/// Configuration for the group commit writer.
pub struct WriterConfig {
pub dir: PathBuf,
pub segment_size: u64,
pub batch_size: usize,
pub batch_timeout: Duration,
pub dedup_window: Duration,
/// Path for the session journal file (optional; `None` in ephemeral mode).
pub session_journal_path: Option<PathBuf>,
/// Shard identity for this writer. Written into every batch header so
/// `WalShipper` and receivers can identify the source shard.
/// Defaults to `ShardId::SINGLE` for single-node deployments.
pub shard_id: ShardId,
/// Region identity for this writer. Written into every batch header.
/// Defaults to `RegionId::SINGLE` for single-node deployments.
pub region_id: RegionId,
/// Optional per-flush observer (fsync wall time + batch event count);
/// see [`crate::wal::config::SyncObserver`].
pub sync_observer: Option<crate::wal::config::SyncObserver>,
/// Optional flushed-batch ship feed (m11p2): after every successful
/// fsync the writer hands the already-encoded batch bytes here for the
/// replication ship queue. `None` outside cluster mode.
pub ship_feed: Option<Arc<WalShipFeed>>,
}
/// Validate a writer configuration before the writer thread is spawned.
///
/// The wire format caps a single batch at [`MAX_EVENTS_PER_BATCH`] events
/// (`format::encode_batch_with_shard` rejects anything larger). The writer
/// drains up to `batch_size` events into one batch, so a `batch_size` above the
/// cap would make the *first* full batch fail to encode — and because that
/// failure propagates out of `run_writer`, it would terminate the writer thread
/// and wedge every subsequent append forever. We reject it here so the caller of
/// `WalHandle::open` fails loudly at startup instead of silently after the first
/// full batch.
///
/// # Errors
///
/// Returns [`WalError::InvalidConfig`] if `batch_size` is `0` or exceeds
/// [`MAX_EVENTS_PER_BATCH`].
pub(crate) fn validate_writer_config(batch_size: usize) -> Result<(), WalError> {
let max = usize::from(MAX_EVENTS_PER_BATCH);
if batch_size == 0 {
return Err(WalError::InvalidConfig {
message: "batch_size must be at least 1".to_string(),
});
}
if batch_size > max {
return Err(WalError::InvalidConfig {
message: format!("batch_size {batch_size} exceeds MAX_EVENTS_PER_BATCH ({max})"),
});
}
Ok(())
}
/// The largest batch the wire format can encode in a single fsync.
///
/// `run_writer` clamps its drain limit to this value as a defence-in-depth
/// belt-and-braces guard: `validate_writer_config` already rejects an oversized
/// `batch_size` at open time, but if a future caller constructs a `WriterConfig`
/// directly (bypassing validation) the clamp still prevents the thread from ever
/// building an unencodable batch and crashing. Correctness is preserved either
/// way — clamping only means more, smaller batches, never a dropped event.
fn effective_batch_size(batch_size: usize) -> usize {
batch_size.clamp(1, usize::from(MAX_EVENTS_PER_BATCH))
}
/// Encode, write, and fsync one batch of non-duplicate events, notifying every
/// caller of the outcome.
///
/// This is the single shared commit routine used by both the steady-state loop
/// in [`run_writer`] and its shutdown-drain tail, so the two paths cannot
/// diverge in how they handle a write failure (the divergence that previously
/// let the drain path drop reply channels on error). Responsibilities:
///
/// 1. Encode the batch at `batch_seq` with the writer's shard/region identity.
/// 2. Rotate the segment first if it has reached its size threshold.
/// 3. Write the encoded bytes and fsync.
/// 4. On success: notify each caller of its assigned sequence number and return
/// the next free sequence number.
/// 5. On failure: notify *every* waiting caller with the underlying error before
/// propagating it, so no caller is ever left blocked on a dropped channel.
///
/// `kept_events` and `kept_replies` are 1:1 and must be non-empty (the callers
/// only invoke this when there is at least one event to persist).
///
/// # Errors
///
/// Returns the underlying [`WalError`] from encode/rotate/write/sync. On error,
/// all reply channels have already been notified with an equivalent error.
///
/// Never panics on a clock anomaly: the batch timestamp is sourced from
/// [`crate::schema::Timestamp::now`], which saturates a pre-Unix-epoch clock to
/// the epoch and logs a warning rather than panicking. WAL ordering is by
/// sequence number, not timestamp, so a clamped timestamp is informational only.
fn flush_batch(
segment: &mut SegmentWriter,
config: &WriterConfig,
batch_seq: u64,
kept_events: &[EventRecord],
kept_replies: Vec<crossbeam::channel::Sender<Result<u64, WalError>>>,
) -> Result<u64, WalError> {
debug_assert_eq!(kept_events.len(), kept_replies.len());
// Source the batch timestamp from the canonical clock-anomaly-safe helper:
// a pre-Unix-epoch wall clock (NTP step-back, dead/uninitialized RTC at boot)
// saturates to the epoch and logs a warning instead of panicking the writer
// thread. The timestamp is informational metadata only — the WAL is ordered
// by sequence number — so a clamped value is safe.
let batch_ts = crate::schema::Timestamp::now().as_nanos();
let write_result = (|| -> Result<Arc<Vec<u8>>, WalError> {
// Test-only deterministic fault: exercise the caller-notification error
// arm without a real I/O failure. Compiled out of production builds.
#[cfg(test)]
if take_flush_failure() {
return Err(WalError::Io(std::io::Error::other(
"injected flush failure",
)));
}
let encoded = format::encode_batch_with_shard(
kept_events,
batch_seq,
batch_ts,
config.shard_id,
config.region_id,
)?;
if segment.needs_rotation() {
segment.rotate(batch_seq)?;
}
segment.write_batch_bytes(&encoded)?;
sync_segment_observed(segment, config, kept_events.len())?;
Ok(Arc::new(encoded))
})();
match write_result {
Ok(encoded) => {
let event_count = kept_events.len() as u64;
tracing::debug!(seq = batch_seq, events = event_count, "wal: batch appended");
// Hand the fsynced batch to the ship feed BEFORE notifying callers:
// a caller's "durable" ack must imply the batch is shippable, so a
// follower can never be told about a seqno the feed has not seen.
if let Some(feed) = &config.ship_feed {
feed.push(FlushedBatch {
bytes: encoded,
first_seq: batch_seq,
last_seq: batch_seq + event_count - 1,
event_count,
});
}
for (i, reply) in kept_replies.into_iter().enumerate() {
let _ = reply.send(Ok(batch_seq + i as u64));
}
Ok(batch_seq + event_count)
}
Err(err) => {
// Notify all waiting callers with the actual error before
// propagating. We cannot clone WalError, so we send a synthetic
// I/O error carrying the same description; dropping the channels
// would surface as a generic `Closed` and hide the real cause.
let err_msg = err.to_string();
for reply in kept_replies {
let _ = reply.send(Err(WalError::Io(std::io::Error::other(err_msg.clone()))));
}
Err(err)
}
}
}
/// Durably sync the segment, timing the fsync for the cluster metrics when an
/// observer is wired (m11p1 "profile first": fsync cost on the deployment's
/// volume is the load-bearing unknown behind group-commit tuning). The clock
/// reads are skipped entirely when no observer is wired. The observer runs
/// behind a panic guard: it is observability-only, and a panicking observer
/// closure must degrade to a logged error — never kill the writer thread,
/// which would wedge every future write in the database.
fn sync_segment_observed(
segment: &SegmentWriter,
config: &WriterConfig,
batch_events: usize,
) -> Result<(), WalError> {
if let Some(observer) = &config.sync_observer {
let sync_start = Instant::now();
segment.sync()?;
let elapsed = sync_start.elapsed();
if std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| {
(observer.0)(elapsed, batch_events);
}))
.is_err()
{
tracing::error!(
"WAL sync observer panicked; fsync metrics for this batch dropped \
(the writer thread continues — fix the observer closure)"
);
}
} else {
segment.sync()?;
}
Ok(())
}
/// Encode and write ONE blob (kind-1/2) record as its own single-seqno
/// batch — WITHOUT syncing. The caller groups consecutive blob writes under
/// one fsync ([`flush_pending_blobs`]); rotation is still safe mid-group
/// because [`SegmentWriter::rotate`] syncs the outgoing segment first.
fn write_blob(
segment: &mut SegmentWriter,
config: &WriterConfig,
seq: u64,
record: &BlobRecord,
) -> Result<Arc<Vec<u8>>, WalError> {
#[cfg(test)]
if take_flush_failure() {
return Err(WalError::Io(std::io::Error::other(
"injected flush failure",
)));
}
let batch_ts = crate::schema::Timestamp::now().as_nanos();
let encoded = record.encode(seq, batch_ts, config.shard_id, config.region_id)?;
if segment.needs_rotation() {
segment.rotate(seq)?;
}
segment.write_batch_bytes(&encoded)?;
Ok(Arc::new(encoded))
}
/// Flush every queued blob in arrival order under ONE group fsync (m11p3):
/// a sync per blob capped item/embedding apply throughput at the fsync
/// floor (~100/s on macOS `F_FULLFSYNC`) — the follower's batched blob
/// apply stages a whole round of records into one drain window, and this
/// is the half that turns that window into one disk round trip.
///
/// Failure contract: a failed WRITE aborts the drain — the failed blob and
/// every blob still queued behind it are notified with the error (their
/// seqnos were never consumed; retries re-stage), while the records written
/// BEFORE the failure are intact and still group-sync + ack below. The drain
/// must not keep writing past a failed write: `write_all` may have left
/// partial bytes at that offset, and appending another record after them
/// would bury a torn record MID-segment (replay stops at the first tear, so
/// every later record — though acked — would vanish on recovery) while
/// handing the same seqno to two records. The freed seqno is reused by the
/// NEXT drain — the same write-failure contract the steady-state event path
/// ([`flush_batch`]) has always had; tail-quarantine for that shared
/// residual (rotating away a suspect tail) is a WAL-wide change tracked in
/// the roadmap, not a per-path patch. A failed group SYNC notifies every
/// caller whose write it covered, and nothing unsynced reaches the ship
/// feed.
fn flush_pending_blobs(
segment: &mut SegmentWriter,
config: &WriterConfig,
mut next_seq: u64,
blobs: Vec<QueuedBlob>,
) -> u64 {
type BlobReply = crossbeam::channel::Sender<Result<u64, WalError>>;
/// One written-but-unsynced blob batch: (seq, encoded, kind, entity, reply).
type WrittenBlob = (u64, Arc<Vec<u8>>, u8, u64, BlobReply);
let mut written: Vec<WrittenBlob> = Vec::new();
let mut queue = blobs.into_iter();
for (record, reply) in queue.by_ref() {
match write_blob(segment, config, next_seq, &record) {
Ok(encoded) => {
written.push((next_seq, encoded, record.kind(), record.entity_id(), reply));
next_seq += 1;
}
Err(e) => {
tracing::error!(
error = %e,
seq = next_seq,
kind = record.kind(),
"wal: blob write failed; drain aborted, this and queued \
callers notified to retry, writer continuing"
);
let err_msg = e.to_string();
let _ = reply.send(Err(WalError::Io(std::io::Error::other(err_msg.clone()))));
for (_, queued_reply) in queue.by_ref() {
let _ = queued_reply
.send(Err(WalError::Io(std::io::Error::other(err_msg.clone()))));
}
break;
}
}
}
if written.is_empty() {
return next_seq;
}
match sync_segment_observed(segment, config, written.len()) {
Ok(()) => {
for (seq, encoded, kind, entity, reply) in written {
tracing::debug!(seq, kind, entity, "wal: blob batch appended");
// Feed BEFORE the caller ack — same ordering contract as
// flush_batch (durable implies shippable).
if let Some(feed) = &config.ship_feed {
feed.push(FlushedBatch {
bytes: encoded,
first_seq: seq,
last_seq: seq,
event_count: 1,
});
}
let _ = reply.send(Ok(seq));
}
}
Err(e) => {
tracing::error!(
error = %e,
blobs = written.len(),
"wal: blob group fsync failed; callers notified, writer continuing"
);
let err_msg = e.to_string();
for (_, _, _, _, reply) in written {
let _ = reply.send(Err(WalError::Io(std::io::Error::other(err_msg.clone()))));
}
}
}
next_seq
}
/// Outcome of dispatching one [`WalCommand`] through [`handle_aux_command`].
///
/// The three command loops in [`run_writer`] (blocking recv, deadline drain,
/// shutdown drain) share their handling of the side-effecting *auxiliary*
/// commands (`TruncateBefore`, `Session*`) but differ only in their
/// continue/break control flow. Routing every command through one helper that
/// returns this enum means each loop matches on the enum — never on the raw
/// command bodies — so the auxiliary-command logic (including the
/// active-segment clamp in [`handle_aux_command`]) lives in exactly one place
/// and cannot drift between the three sites.
enum CommandOutcome {
/// An `Append` was received: the loop must push `(event, reply)` onto its
/// pending batch.
Pushed(QueuedAppend),
/// An `AppendBlob` was received: the loop must queue it for a blob flush
/// after the pending signal batch commits (blobs never mix into a signal
/// batch — each flushes as its own single-seqno batch).
Blob(QueuedBlob),
/// A side-effecting command (`TruncateBefore`, `Session*`) was fully
/// handled inside the helper; the loop should keep going.
Handled,
/// `Shutdown` was received (or the channel disconnected): the loop must
/// stop draining.
Shutdown,
}
/// Dispatch a single received [`WalCommand`], executing any side effect.
///
/// `Append` is returned as [`CommandOutcome::Pushed`] for the caller to batch;
/// every other variant is fully handled here and reported as
/// [`CommandOutcome::Handled`] or [`CommandOutcome::Shutdown`]. Sharing this
/// dispatch across all three loops in [`run_writer`] keeps the auxiliary-command
/// bodies byte-identical by construction.
///
/// # Active-segment protection (`TruncateBefore`)
///
/// `TruncateBefore` must NOT unlink the segment the live writer is appending to.
/// After any write burst the active segment's `first_seq` sits below the
/// materialized checkpoint, so a naive `delete_segments_before(checkpoint_seq)`
/// would `remove_file` the very inode this writer still holds open — and on
/// Linux the writer would keep appending to the now-unlinked inode, silently
/// losing every post-checkpoint, already-fsync'd, acknowledged write on the next
/// open. We therefore clamp the deletion floor to the live segment's
/// `first_seq`, exactly as [`crate::wal::compaction::compact_wal_online`] does,
/// guaranteeing the active segment always survives. The writer thread already
/// owns `active_first_seq`, so the clamp is a single `min` with no extra I/O.
fn handle_aux_command(
cmd: WalCommand,
config: &WriterConfig,
active_first_seq: u64,
session_journal: &mut Option<SessionJournal>,
) -> CommandOutcome {
match cmd {
WalCommand::Append { event, reply } => CommandOutcome::Pushed((event, reply)),
WalCommand::AppendBlob { record, reply } => CommandOutcome::Blob((record, reply)),
WalCommand::TruncateBefore { before_seq, reply } => {
// Clamp the deletion floor so the live segment (the maximum-first_seq
// segment, which this writer holds open) is never unlinked out from
// under our FD. See the function-level rustdoc and
// `compaction::compact_wal_online` for the full hazard analysis.
let floor = before_seq.min(active_first_seq);
let result = segment::delete_segments_before(&config.dir, floor);
let _ = reply.send(result.map(|_| ()));
CommandOutcome::Handled
}
cmd @ (WalCommand::SessionStart { .. }
| WalCommand::SessionSignal { .. }
| WalCommand::SessionClose { .. }) => {
handle_session_command(cmd, session_journal);
CommandOutcome::Handled
}
WalCommand::Shutdown => CommandOutcome::Shutdown,
}
}
/// Split a drained batch into kept events (1:1 with their replies) and duplicate
/// replies, notifying duplicates immediately with the dedup sentinel `0`.
///
/// Shared by the steady-state loop and the shutdown drain so dedup accounting
/// cannot diverge between them.
fn partition_dedup(
dedup: &mut DedupWindow,
batch: impl IntoIterator<Item = QueuedAppend>,
) -> (
Vec<EventRecord>,
Vec<crossbeam::channel::Sender<Result<u64, WalError>>>,
) {
let mut kept_events: Vec<EventRecord> = Vec::new();
let mut kept_replies: Vec<crossbeam::channel::Sender<Result<u64, WalError>>> = Vec::new();
// Hashes kept *within this batch*, so two identical events in the same drained
// batch still dedup against each other even though the durable dedup window is
// only updated AFTER a successful flush (see `run_writer`). This preserves
// intra-batch suppression without violating the "mark as seen only once
// durable" invariant.
let mut batch_seen: std::collections::HashSet<u128> = std::collections::HashSet::new();
for (event, reply) in batch {
let hash = format::event_content_hash(&event);
// `dedup.contains` CHECKS membership without recording; `batch_seen.insert`
// returns false when this exact event already appeared earlier in this
// batch. Recording into `dedup` happens in `run_writer`, only after the
// batch is durably persisted.
if dedup.contains(&event) || !batch_seen.insert(hash) {
// Duplicate: notify with the dedup sentinel (seq=0) immediately.
let _ = reply.send(Ok(0));
} else {
kept_events.push(event);
kept_replies.push(reply);
}
}
(kept_events, kept_replies)
}
/// The group commit writer loop.
///
/// Runs on a dedicated thread. Receives events via crossbeam channel,
/// accumulates them into batches, writes batches to the WAL segment,
/// and fsyncs once per batch. Callers are notified of their sequence
/// numbers via per-event reply channels.
///
/// # Batch formation
///
/// 1. Block until the first event arrives.
/// 2. Drain additional events from the channel up to `batch_size` or
/// until `batch_timeout` elapses (whichever comes first).
/// 3. Deduplicate events, encode the batch, write to segment, fsync.
/// 4. Send sequence numbers back to all waiting callers.
///
/// # Resilience
///
/// A flush failure in the **steady-state loop** does NOT terminate the writer.
/// `flush_batch` notifies every waiting caller with the error (so they can
/// retry), the batch's events are left un-recorded in the dedup window (so a
/// retry is accepted, not suppressed as a phantom duplicate), and the loop keeps
/// serving — a transient I/O fault (ENOSPC, EINTR, NFS blip) must never convert
/// into a permanent write outage by dropping the command channel.
///
/// # Errors
///
/// Only the shutdown drain / final fsync propagate an error out of this function
/// (the WAL is closing anyway, and callers were already notified). The encoding
/// path cannot fail under normal operation — `effective_batch_size` clamps the
/// drain limit to `MAX_EVENTS_PER_BATCH` so an oversized `batch_size` can never
/// produce an unencodable batch.
///
/// Never panics on a clock anomaly: [`flush_batch`] sources its batch timestamp
/// from the clock-anomaly-safe [`crate::schema::Timestamp::now`], so a
/// pre-Unix-epoch wall clock cannot kill the writer thread.
// The encode/dedup/write/sync logic is extracted into `flush_batch` and
// `partition_dedup`, and the per-command dispatch into `handle_aux_command`;
// what remains is the steady-state loop plus the shutdown drain, each matching
// only on the `CommandOutcome` enum. Splitting the control flow further would
// obscure the single read-recv/drain structure.
#[allow(clippy::too_many_lines)]
pub fn run_writer(
rx: &Receiver<WalCommand>,
config: &WriterConfig,
mut segment: SegmentWriter,
start_seq: u64,
mut dedup: DedupWindow,
) -> Result<(), WalError> {
let mut next_seq = start_seq;
// Clamp the drain limit to what the wire format can encode in one batch.
// `validate_writer_config` rejects an oversized `batch_size` at open time;
// this is the defence-in-depth guard for any caller that builds a
// `WriterConfig` directly. See `effective_batch_size`.
let max_batch = effective_batch_size(config.batch_size);
let mut batch: Vec<QueuedAppend> = Vec::with_capacity(max_batch);
// Blob (kind-1/2) appends queued during this iteration's drain window.
// Flushed AFTER the signal batch commits, each as its own single-seqno
// batch — seqnos are assigned in flush order, so the relative ordering of
// concurrent signal and blob submissions within one window is arbitrary
// (exactly as it is between any two concurrent writers).
let mut pending_blobs: Vec<QueuedBlob> = Vec::new();
let mut shutdown_requested = false;
// Open the session journal if a path was provided (persistent mode).
let mut session_journal: Option<SessionJournal> = config
.session_journal_path
.as_ref()
.and_then(|p| match SessionJournal::open(p) {
Ok(j) => Some(j),
Err(e) => {
tracing::error!(error = %e, "failed to open session journal; session WAL writes will be skipped");
None
}
});
loop {
// Block until the first event arrives (or shutdown/disconnect). All
// side-effecting commands route through `handle_aux_command` so the
// `TruncateBefore` active-segment clamp and the `Session*` delegation
// live in exactly one place; this loop only steers control flow.
match rx.recv() {
Ok(cmd) => {
match handle_aux_command(cmd, config, segment.first_seq(), &mut session_journal) {
CommandOutcome::Pushed(queued) => batch.push(queued),
CommandOutcome::Blob(queued) => pending_blobs.push(queued),
CommandOutcome::Handled => continue,
CommandOutcome::Shutdown => break,
}
}
Err(_) => break,
}
// Drain up to the (clamped) batch limit with a deadline.
let deadline = Instant::now() + config.batch_timeout;
while batch.len() < max_batch {
match rx.recv_deadline(deadline) {
Ok(cmd) => {
match handle_aux_command(cmd, config, segment.first_seq(), &mut session_journal)
{
CommandOutcome::Pushed(queued) => batch.push(queued),
CommandOutcome::Blob(queued) => pending_blobs.push(queued),
// Side-effecting commands bypass the batch; keep draining.
CommandOutcome::Handled => {}
CommandOutcome::Shutdown => {
shutdown_requested = true;
break;
}
}
}
Err(crossbeam::channel::RecvTimeoutError::Disconnected) => {
shutdown_requested = true;
break;
}
Err(crossbeam::channel::RecvTimeoutError::Timeout) => break,
}
}
// Deduplicate, then commit through the shared flush routine so this
// path cannot diverge from the shutdown drain below. `partition_dedup`
// notifies duplicate senders with the dedup sentinel; `flush_batch`
// notifies every kept caller (success *or* error) before returning.
// drain(..) reuses batch's heap allocation across loop iterations.
#[allow(clippy::iter_with_drain)]
let (kept_events, kept_replies) = partition_dedup(&mut dedup, batch.drain(..));
if !kept_events.is_empty() {
match flush_batch(&mut segment, config, next_seq, &kept_events, kept_replies) {
Ok(seq) => {
next_seq = seq;
// Record events as seen ONLY now that the batch is durably
// persisted. Recording earlier (in `partition_dedup`) would
// suppress a legitimate retry of a batch that failed to flush.
for event in &kept_events {
dedup.record(event);
}
}
Err(e) => {
// A transient I/O fault (a brief ENOSPC that an operator then
// clears, an EINTR, an NFS/network-storage blip) must NOT tear
// down the writer thread — that would drop the command channel
// and wedge every future write forever, since the WAL is the
// source of truth for all entity/signal/relationship writes.
// `flush_batch` has already notified each waiting caller with
// the error, so they can retry; we keep `next_seq` unchanged
// (the failed batch's sequence range is free for the retry),
// leave the events UNRECORDED in the dedup window so the retry
// is accepted, and keep serving.
tracing::error!(
error = %e,
seq = next_seq,
events = kept_events.len(),
"wal: batch flush failed; callers notified to retry, writer continuing"
);
}
}
}
// Blob appends queued during this window flush after the signal batch
// (each as its own single-seqno batch through the same resilience
// contract: a failed flush notifies its caller and frees the seqno).
if !pending_blobs.is_empty() {
next_seq = flush_pending_blobs(
&mut segment,
config,
next_seq,
std::mem::take(&mut pending_blobs),
);
}
if shutdown_requested {
break;
}
}
// Drain any remaining commands that arrived before senders observed
// the shutdown. This ensures in-flight append() calls are not silently
// dropped, which would cause callers to block forever or receive
// WalError::Closed instead of a real sequence number.
let mut final_batch: Vec<QueuedAppend> = Vec::new();
let mut final_blobs: Vec<QueuedBlob> = Vec::new();
// Same shared dispatch as the steady loop: `Append` queues into the final
// batch, `AppendBlob` into the final blob list, `TruncateBefore`/`Session*`
// are handled (with the active-segment clamp), and a duplicate `Shutdown`
// is a no-op. The loop ends when the channel is empty or disconnected.
while let Ok(cmd) = rx.try_recv() {
match handle_aux_command(cmd, config, segment.first_seq(), &mut session_journal) {
CommandOutcome::Pushed(queued) => final_batch.push(queued),
CommandOutcome::Blob(queued) => final_blobs.push(queued),
CommandOutcome::Handled | CommandOutcome::Shutdown => {}
}
}
// Flush the final drain batch through the SAME shared routine the steady
// loop uses. Routing both paths through `flush_batch` is what guarantees the
// drain path notifies waiting callers on a write error instead of dropping
// their reply channels (the divergence this consolidation fixes). A drained
// batch can exceed `max_batch` if many appends queued during shutdown, so we
// commit it in `max_batch`-sized chunks the wire format can encode.
// `next_seq` is advanced per chunk so sequence numbers stay monotonic
// across the (rare) multi-chunk drain. It is not propagated past the
// writer's exit, but is read on each subsequent chunk/blob iteration.
//
// CRITICAL: do NOT early-return (`?`) on a chunk failure here.
// `flush_batch` notifies only the CURRENT chunk's callers before
// returning Err, so propagating immediately would drop the reply channels
// of every NOT-YET-PROCESSED chunk (and every drained blob) — those
// callers would observe a misleading `Closed` instead of the real I/O
// fault, violating the module invariant that every queued append
// eventually resolves its reply. Instead we remember the first error,
// keep draining, notify each remaining caller with the same error class,
// and propagate the original error only after every queued append has
// been resolved.
let mut drain_err: Option<WalError> = None;
if !final_batch.is_empty() {
let (kept_events, kept_replies) = partition_dedup(&mut dedup, final_batch);
let mut events = kept_events.into_iter();
let mut replies = kept_replies.into_iter();
loop {
let chunk_events: Vec<EventRecord> = events.by_ref().take(max_batch).collect();
if chunk_events.is_empty() {
break;
}
let chunk_replies: Vec<_> = replies.by_ref().take(chunk_events.len()).collect();
if let Some(ref e) = drain_err {
// A prior chunk already failed; flush_batch only notified that
// chunk's callers. Resolve THESE callers' replies with the same
// error class rather than dropping their channels.
let msg = e.to_string();
for reply in chunk_replies {
let _ = reply.send(Err(WalError::Io(std::io::Error::other(msg.clone()))));
}
continue;
}
match flush_batch(&mut segment, config, next_seq, &chunk_events, chunk_replies) {
Ok(seq) => next_seq = seq,
Err(e) => drain_err = Some(e),
}
}
}
// Drained blobs flush through the same contract: healthy → one group
// flush in order (write errors notify their own callers inside); after
// any drain failure → notify the remaining callers with the same error
// class instead of dropping their channels.
if let Some(ref e) = drain_err {
for (_, reply) in final_blobs {
let _ = reply.send(Err(WalError::Io(std::io::Error::other(e.to_string()))));
}
} else if !final_blobs.is_empty() {
// The post-blob seqno is final here — the writer exits after this drain.
let _ = flush_pending_blobs(&mut segment, config, next_seq, final_blobs);
}
if let Some(e) = drain_err {
return Err(e);
}
// Final sync before exit
segment.sync()?;
Ok(())
}
/// Write a session lifecycle command to the session journal.
///
/// This function is called from the writer thread. Session commands bypass the
/// signal batch system entirely. Errors are logged and swallowed -- session WAL
/// writes are best-effort; the in-memory session state is the source of truth.
fn handle_session_command(cmd: WalCommand, journal: &mut Option<SessionJournal>) {
let Some(journal) = journal.as_mut() else {
// No session journal open (should not happen in persistent mode, but
// log defensively).
return;
};
let event = match cmd {
WalCommand::SessionStart {
session_id,
user_id,
started_at_ns,
agent_id,
policy_name,
} => SessionWalEvent::Start {
session_id,
user_id,
started_at_ns,
agent_id,
policy_name,
},
WalCommand::SessionSignal {
session_id,
entity_id,
weight,
ts_ns,
signal_name,
annotation,
session_seqno,
idempotency_key,
} => SessionWalEvent::Signal {
session_id,
entity_id,
weight,
ts_ns,
signal_name,
annotation,
session_seqno: session_seqno.map(SessionSeqNo),
idempotency_key,
},
WalCommand::SessionClose { session_id } => SessionWalEvent::Close { session_id },
// Other commands are not handled here.
_ => return,
};
if let Err(e) = journal.append(&event) {
tracing::warn!(error = %e, "session journal write failed");
}
}
#[cfg(test)]
#[allow(clippy::unwrap_used, clippy::similar_names)]
#[path = "writer_tests.rs"]
mod tests;
#[cfg(test)]
#[allow(clippy::unwrap_used)]
mod rotation_failure_tests {
use crossbeam::channel::bounded;
use super::*;
use crate::wal::segment::{SegmentWriter, segment_filename};
fn make_event(id: u64) -> EventRecord {
EventRecord::signal(id, 1, 1.0, 1_000_000_000)
}
/// wal-write SUG: cover a flush failure that lands specifically inside
/// `rotate()` (old segment synced, new segment file fails to open), not just
/// the pre-encode fault the `FAIL_NEXT_FLUSH` hook injects.
///
/// We force the failure deterministically with real filesystem semantics:
/// `max_size = 0` makes `needs_rotation()` true on the first flush, and we
/// pre-create a *directory* at the exact path the new segment file would take
/// (`wal-{batch_seq:020}.seg`) so `rotate()`'s `OpenOptions::open` fails. The
/// contract under test is the same as every other flush-error path: the
/// caller's reply channel must carry the error rather than being dropped (a
/// dropped channel would surface as a misleading `Closed`).
#[test]
fn flush_batch_error_inside_rotate_notifies_callers() {
let dir = tempfile::tempdir().expect("tempdir creation should succeed");
// max_size = 0 -> needs_rotation() is true immediately, so the very first
// flush attempts a rotate before writing.
let mut segment =
SegmentWriter::open(dir.path(), ShardId::SINGLE, 1, 0).expect("open should succeed");
let config = WriterConfig {
dir: dir.path().to_path_buf(),
segment_size: 0,
batch_size: 100,
batch_timeout: Duration::from_millis(10),
dedup_window: Duration::from_secs(30),
session_journal_path: None,
shard_id: ShardId::SINGLE,
region_id: RegionId::SINGLE,
sync_observer: None,
ship_feed: None,
};
// flush_batch rotates to `segment_filename(SINGLE, batch_seq)` where
// batch_seq = 1 (matching SegmentWriter::open's first_seq). Pre-create a
// DIRECTORY at that path so opening it as a file fails — a real, not
// mocked, I/O fault landing inside rotate().
let collision = dir.path().join(segment_filename(ShardId::SINGLE, 1));
std::fs::remove_file(&collision).expect("seed segment file should exist to remove");
std::fs::create_dir(&collision).expect("create blocking directory should succeed");
let (reply_tx, reply_rx) = bounded(1);
let kept_events = vec![make_event(1)];
let kept_replies = vec![reply_tx];
let result = flush_batch(&mut segment, &config, 1, &kept_events, kept_replies);
assert!(
result.is_err(),
"a rotate() failure inside flush must propagate, got {result:?}"
);
// The caller must be notified with the error, not left on a dropped channel.
let reply = reply_rx
.recv()
.expect("reply channel must NOT be dropped when rotate() fails");
assert!(
matches!(reply, Err(WalError::Io(_))),
"caller must receive the rotate I/O error, got {reply:?}"
);
}
}