Implements tmp/tidaldb-fleet-hardening (20 planned tasks + 2 found by measurement). Ring 0 — restore verification. .woodpecker.yaml step pods ran at the namespace default of 1500m/2Gi, which OOMKilled a prior pipeline and starved the release gate past its budget. Both push-path steps now declare backend_options.kubernetes.resources as two YAML anchors declared once on their first consuming step. The values are CALIBRATED against measured free node capacity, not against the LimitRange max: `requests: cpu 2` (this roadmap's original figure) fits on NO node and would sit Pending forever, because `ci-build-bounds` grants permission and the nodes supply capacity, and those are not the same thing. The `nightly` cron described in this file for 216 days was never created, so tier-3 chaos, the fault classes, mTLS and the PITR test produced exactly zero signal while reading like standing coverage. nightly-chaos and nightly-security-ops now alias the anchors and have budgets matching the gate (their 120/90 were TIGHTER on the same runner, so they would have failed nightly for a budget reason, not a correctness one). nightly-soak is REMOVED, not scheduled: it drives 1000 rps for 600s gating on p99 <= 250ms, and the best node has 1700m free CPU, so it would fail on starvation rather than regression — manufacturing a nightly false alarm. Its commands move verbatim to docs/runbooks/nightly-soak.md. Ring 1 — four fabrications removed from the wire. - scatter_merge sorted and truncated without re-stamping rank, so /feed and /search returned 1,1,2 under full placement. Reuses merge_cross_shard's existing stamp; asserted on BOTH the multi-group merge path and the single-group [only] fast path that bypasses it. - aggregate_region_row's None arm invented `applied_events: 0` plus a deficit derived from it. applied_events/lag_events are now Option<u64>, null on the wire. leader_last_seq was also unwrap_or(0), so a node that could not reach the LEADER computed 0 - applied = 0 for every region and reported a converged cluster it had never measured — a fabrication pointing the dangerous way. - tidalctl inferred NO REPORT from `applied == 0 && lag > 0`. That heuristic was actively hiding the PVC-wipe shape: a measured zero with a real deficit rendered as "no report" instead of BEHIND. Now read off the wire; converged exits 0, partitioned still exits nonzero. - /sharded/* answered 201/204 for single-copy writes with nothing anywhere saying so. Now requires `x-tidal-ack: local`, rejecting with 400 via the existing invalid_input path. Six call sites migrated, not the two this roadmap predicted — including docs/runbooks/cluster.md §16.3, which told operators to run a quorum-write probe via POST /sharded/items. That probe cannot verify quorum: the surface applies locally with no WAL append. It was used as the safety check between every step of a staged deploy earlier today. Ring 2 — observability. JSON_LOGS was already implemented and the deployment simply never asked for it; the StatefulSet now sets it, plus TIDAL_SERVICE_NAME=tidaldb because enabling it silently renames the VictoriaLogs `service` stream field and would have blinded every query keyed on it. Adds tidaldb_usearch_replicated_vectors_total, incremented on BOTH the origin (wal_blob_first -> Ok(Some)) and the follower apply path — counting only the origin would mean each vector lands on exactly one node, replicas never agree, and the alert built on it pages forever. Found by measurement, not planned: the 401 path discarded every fact about every rejection. Traefik has served 101,858 rejected requests to the public ingress — 87.6% of all its traffic — with no record of who or why anywhere. unauthorized_response now emits reason (missing_token vs invalid_token, the distinction that separates a scanner from a rotation that missed a consumer) and the forwarded client. The token is never logged. Also: scripts/restore-fleet.sh --cluster started the soak monitor while deliberately leaving its gate suspended, orphaning a watcher that has reported "0/30 green nights" for 13 days. The pair now moves together. Doc-guard's three-warning backlog is cleared with real backfill for M4/M6/M12. Verified: fmt clean; clippy 5 crates 0 new warnings (74 vs 74 baseline, counted in a detached worktree at HEAD); lib 2110 passed; cluster_sharding 5; cluster_runbook 10; tidalctl 38; doc-guard 0 warnings. Playwright 32/34 with the two remaining failures asserting the rank fix against the not-yet-rolled image — they are the post-deploy proof.
9716 lines
443 KiB
Rust
9716 lines
443 KiB
Rust
//! `ClusterNode` (the process) hosting `ShardReplica`s (the shard groups), m11p6.
|
||
//!
|
||
//! Two types, one file:
|
||
//!
|
||
//! * [`ClusterNode`] is the OS process and the axum `State`. It owns node
|
||
//! identity, the shared forward client, the gateway [`ShardRouter`], and a
|
||
//! `BTreeMap<ShardId, Arc<ShardReplica>>` of the groups this node replicates.
|
||
//! Every entity write hash-routes through it to the owning group's leader
|
||
//! (applying locally or forwarding); corpus reads scatter over its hosted
|
||
//! groups. For the legacy single group (`shards:` absent) it holds exactly
|
||
//! ONE `ShardReplica` spanning every region — byte-for-byte the pre-m11p6
|
||
//! one-region-per-process node.
|
||
//! * [`ShardReplica`] is ONE shard group's full replication machinery (formerly
|
||
//! `RegionClusterState`, when a process WAS exactly one group). A node hosts N
|
||
//! of them, not one. Each owns:
|
||
//! * one [`TidalDb`] (`NodeRole::Single`, `shard = ShardId(region)`,
|
||
//! `peer_shards = the group's other replica nodes`) — direct-writable and
|
||
//! promotable. The in-group identity stays region-id based; the data-shard
|
||
//! only namespaces its data subdir + gRPC port;
|
||
//! * one [`GrpcTransport`] whose server binds THIS group's resolved gRPC
|
||
//! address and whose peers are the group's OTHER replica nodes;
|
||
//! * an **always-on** segment receiver (`db.start_replication`) — leadership
|
||
//! can move, so an inbound segment is legal on any node after a promote;
|
||
//! * a [`SignalRelay`]-backed leader write path with a leader check.
|
||
//!
|
||
//! Without `--region` the existing single-process [`ClusterState`] runs
|
||
//! byte-for-byte unchanged.
|
||
//!
|
||
//! # Replication model (m11p2 — one replicated log)
|
||
//!
|
||
//! EVERY replicated mutation rides the leader's WAL: signals as kind-0 event
|
||
//! batches, item metadata and embeddings as kind-1/2 blob records. The WAL
|
||
//! writer hands each fsynced batch to the ship feed; the [`ShipQueue`]'s
|
||
//! per-peer senders push the hot tail, and a follower that detects a gap
|
||
//! pulls history itself via the `StreamSegments` catch-up stream over the
|
||
//! leader's durable segments. The m8p10-era HTTP item/embedding broadcast —
|
||
//! the marker-gated fan-out and the O(items) heal re-broadcast, source of
|
||
//! both 2026-06-10 live bugs — is DELETED; those bug classes are now
|
||
//! impossible by construction. Hard negatives remain CRDT (user-scoped,
|
||
//! commutative; `/cluster/reconcile`): log = totally-ordered global data,
|
||
//! CRDT = per-user convergent data.
|
||
//!
|
||
//! Multi-process cluster mode therefore REQUIRES a persistent `--data-dir`:
|
||
//! the durable WAL is the replication stream, so a node without one has
|
||
//! nothing to ship or serve catch-up from.
|
||
|
||
use std::{
|
||
collections::{BTreeMap, HashMap, HashSet},
|
||
sync::{
|
||
Arc, RwLock, Weak,
|
||
atomic::{AtomicBool, AtomicU64, Ordering},
|
||
},
|
||
time::{Duration, Instant},
|
||
};
|
||
|
||
use axum::{
|
||
Json, Router,
|
||
extract::{Path, Query, RawQuery, Request, State},
|
||
http::{HeaderMap, StatusCode},
|
||
middleware::{self, Next},
|
||
response::{IntoResponse, Response},
|
||
routing::{get, post},
|
||
};
|
||
use serde::{Deserialize, Serialize};
|
||
use tidal_net::{
|
||
GrpcTransport,
|
||
config::GrpcTransportConfig,
|
||
sources::{AppliedSource, SegmentChunk, SegmentReadError, SegmentSource, ServingSources},
|
||
};
|
||
use tidaldb::{
|
||
TidalDb,
|
||
db::{
|
||
StagedSignal,
|
||
config::{NodeConfig, NodeRole},
|
||
},
|
||
query::{retrieve::Retrieve, search::Search},
|
||
replication::{
|
||
CommitIndex, ReseedMarker, ReseedMarkerStore, ReseedReason, ShipQueue, Transport,
|
||
WalFeedSource,
|
||
shard::{RegionId, ShardId, ShardRouter},
|
||
},
|
||
schema::{EntityId, Schema, Timestamp},
|
||
wal::feed::WalShipFeed,
|
||
};
|
||
use tower::{ServiceBuilder, limit::ConcurrencyLimitLayer};
|
||
use tower_http::timeout::TimeoutLayer;
|
||
use utoipa::ToSchema;
|
||
|
||
use super::{
|
||
forward::{
|
||
self, broadcast_marked, forward_json_with_headers, forwarded_auth, is_internal, is_relayed,
|
||
peer_url,
|
||
},
|
||
reseed,
|
||
routes::{ClusterAppError, ScatterGatherInfo, ShardedFeedResponse, ShardedSearchResponse},
|
||
topology::{ResolvedShardGroup, TopologySpec, shard_of_region},
|
||
transport::{GRPC_READY_TIMEOUT, grpc_server_ready, resolve_grpc_bind_addr},
|
||
};
|
||
use crate::{
|
||
dto::{
|
||
EmbeddingRequest, FeedQuery, FeedResponse, ItemRequest, SearchQueryParams, SearchResponse,
|
||
SignalRequest, VectorSearchRequest, VectorSearchResponse, feed_items, search_items,
|
||
vector_matches,
|
||
},
|
||
error::{Result, ServerError},
|
||
offload::{ClusterWritePool, offload_read},
|
||
scatter_gather::{
|
||
HttpShardContext, entity_shard, scatter_gather_retrieve_http, scatter_gather_search_http,
|
||
},
|
||
};
|
||
|
||
/// Write-acknowledgment mode for replicated cluster writes (m11p3).
|
||
///
|
||
/// `Leader` (the default) succeeds at leader group-commit fsync — the m0-m11p2
|
||
/// contract. `Quorum` additionally blocks until a majority of the replica set
|
||
/// (leader + peers) durably holds the write (the commit index passes its
|
||
/// seqno), trading tail latency for failover-survivable durability. Set the
|
||
/// deployment default with `replication.ack` in the topology; override per
|
||
/// request with the `x-tidal-ack` header.
|
||
///
|
||
/// Adding a mode fans out beyond this enum: `parse`/`as_str` here, the
|
||
/// topology validation table (`topology.rs`), the stress CLI's `--ack`, and
|
||
/// the runbook §8 table — grep `x-tidal-ack` for the full set.
|
||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||
pub enum AckMode {
|
||
/// Success = durable on the leader (its WAL group-commit fsync).
|
||
Leader,
|
||
/// Success = durable on a majority of the replica set.
|
||
Quorum,
|
||
}
|
||
|
||
impl AckMode {
|
||
/// Parse a topology/header value.
|
||
fn parse(value: &str) -> Option<Self> {
|
||
match value {
|
||
"leader" => Some(Self::Leader),
|
||
"quorum" => Some(Self::Quorum),
|
||
_ => None,
|
||
}
|
||
}
|
||
|
||
const fn as_str(self) -> &'static str {
|
||
match self {
|
||
Self::Leader => "leader",
|
||
Self::Quorum => "quorum",
|
||
}
|
||
}
|
||
}
|
||
|
||
/// Persisted stream-baseline filename inside the node's data dir.
|
||
///
|
||
/// The baseline is the WAL seqno at which THIS node's outbound stream started
|
||
/// (0 for the topology's original leader; the promote-time flushed frontier
|
||
/// after a promotion). It must survive restarts: a restarted leader that
|
||
/// forgot its baseline would serve catch-up from pre-stream history and
|
||
/// double-apply it on followers.
|
||
const STREAM_BASELINE_FILE: &str = "stream_baseline";
|
||
|
||
/// Durable file holding the [`ShardReplica::leader_acked`] frontier (m12
|
||
/// election-divergence-fix). Raw 8-byte LE `u64`, mirroring [`STREAM_BASELINE_FILE`].
|
||
const LEADER_ACKED_FILE: &str = "leader_acked";
|
||
|
||
/// How long a leader-sanctioned transfer waits for the target to hold the
|
||
/// full flushed prefix before `TimeoutNow` (the drain, m11p4).
|
||
const TRANSFER_CATCHUP_WAIT: Duration = Duration::from_secs(5);
|
||
|
||
/// How long `/cluster/promote` waits for the transfer election to take
|
||
/// (target leads at a higher term) before reporting failure.
|
||
const TRANSFER_TAKEOVER_WAIT: Duration = Duration::from_secs(10);
|
||
|
||
/// How long a gracefully-shutting-down LEADER waits for its flushed tail (its
|
||
/// last entries AND the term marker it journaled when it won) to commit to a
|
||
/// quorum before it steps down. Bounded well inside a k8s SIGTERM grace period so
|
||
/// shutdown never hangs; if the deadline passes the node steps down anyway (no
|
||
/// worse than the pre-fix behavior). This is the graceful leadership hand-off:
|
||
/// committing the tail before step-down means the next leader (a quorum member,
|
||
/// so caught-up by the vote restriction) holds the full prefix, and THIS node
|
||
/// rejoins as a clean follower instead of carrying a divergent suffix the new
|
||
/// term never saw — the root of the rolling-restart reseed/quarantine churn.
|
||
// m12 election-divergence-fix: 3s, not 10s. With the write-quiesce freezing the
|
||
// flushed frontier at shutdown, the committed index catches up within a heartbeat
|
||
// or two (sub-second), so 3s is ample for a real hand-off — and the old leader
|
||
// keeps HEARTBEATING until it steps down, so a longer wait only DELAYS the
|
||
// survivors' election (the "shard leaders did not converge" stall under load).
|
||
// The durable `leader_acked` frontier is the safety net: any un-replicated
|
||
// ack=leader tail the drain did not flush still quarantines on rejoin.
|
||
const SHUTDOWN_HANDOFF_WAIT: Duration = Duration::from_secs(3);
|
||
|
||
/// Cap on one commit-watch bridge condvar wait: the longest the bridge
|
||
/// thread can go without re-checking its stop flag, i.e. the worst-case
|
||
/// shutdown latency the bridge adds. Deliberately a constant, not a
|
||
/// topology knob — it is an internal poll bound, invisible to clients
|
||
/// (commit-index CHANGES wake the bridge immediately regardless).
|
||
const COMMIT_BRIDGE_WAKE_INTERVAL: Duration = Duration::from_secs(1);
|
||
|
||
/// The removal-delivery grace (m11p5 §3.3): how long the leader keeps a removed
|
||
/// peer's ship cell + gRPC entry open, still shipping it the `Removed` record,
|
||
/// before giving up and retiring the cell anyway. The cell retires EARLY the
|
||
/// instant the removed peer's applied mark covers the record (it learned of its
|
||
/// removal through the log); this deadline only bounds the wait for a DEAD or
|
||
/// unreachable removed peer so it cannot pin the cell forever. Env-overridable
|
||
/// for tests via `TIDAL_REMOVE_DELIVERY_GRACE_MS`.
|
||
const REMOVE_DELIVERY_GRACE_DEFAULT_MS: u64 = 30_000;
|
||
|
||
/// Election ticks (50 ms each) between self-driving heal passes (m11p8). ~3 s —
|
||
/// frequent enough to re-arm a stuck peer's backlog re-ship well within a
|
||
/// circuit-breaker reset window (30 s) without churning cursors every tick.
|
||
const SELF_HEAL_TICKS: u64 = 60;
|
||
|
||
/// The election driver's boot bundle: prepared in [`ShardReplica::new`]
|
||
/// (where the durable classification runs), consumed by
|
||
/// [`ShardReplica::start_election_driver`] once the node is in its
|
||
/// final `Arc`.
|
||
struct ElectionBoot {
|
||
config: tidaldb::replication::ElectionConfig,
|
||
hard: tidaldb::replication::HardState,
|
||
boots_as_leader: bool,
|
||
store: tidaldb::replication::ElectionStore,
|
||
topology_leader: RegionId,
|
||
}
|
||
|
||
/// One shard group's full replication machinery, hosted inside a [`ClusterNode`].
|
||
///
|
||
/// A node holds one `ShardReplica` per group it replicates. It owns one region's
|
||
/// in-group identity — the data-shard `group_shard` only namespaces its dir/port.
|
||
// The bool fields (`multi_shard`, `reseed_self_restart`, `seed_joiner`,
|
||
// `install_boot`) are INDEPENDENT boot/identity facts, not a state machine that
|
||
// would read better as an enum — an enum would force false either/or relations
|
||
// between orthogonal flags. Allow the bool count rather than contort the model.
|
||
#[allow(clippy::struct_excessive_bools)]
|
||
pub struct ShardReplica {
|
||
/// This process's region id (index into the topology declaration order).
|
||
region: RegionId,
|
||
/// This region's human-readable name.
|
||
region_name: String,
|
||
/// The data-shard group this replica serves (m11p6): the gateway hash
|
||
/// output, the `ClusterNode` `shards` map key, and the metrics/status
|
||
/// shard label. The in-group replication identity stays `region`-based
|
||
/// (`shard_of_region`) — `group_shard` only distinguishes co-hosted groups.
|
||
group_shard: ShardId,
|
||
/// Whether this node co-hosts more than one shard group (m11p6 S>1). When
|
||
/// true, this replica's intra-group admin forwards/broadcasts carry a
|
||
/// `?shard=<group_shard>` selector so the receiving sibling targets the SAME
|
||
/// group, and `NotLeader` names the group. `false` for the legacy single
|
||
/// group keeps the S=1 wire format byte-for-byte (no selector, no shard in
|
||
/// the error body) — set once at construction from the resolved group count.
|
||
multi_shard: bool,
|
||
/// `Some` for the server's lifetime; cleared on shutdown so the `TidalDb` is
|
||
/// dropped (checkpoint + WAL fsync + HNSW-graph checkpoint + thread join)
|
||
/// deterministically.
|
||
///
|
||
/// Held in an [`arc_swap::ArcSwapOption`] (not a plain `Option<Arc<…>>`) so
|
||
/// the deterministic close can run through a SHARED `&self` (m12p6 SIGTERM
|
||
/// fix). On a k8s SIGTERM the graceful HTTP drain can be blocked by stuck
|
||
/// peer keep-alive connection tasks that still hold an `Arc<ClusterNode>`, so
|
||
/// the post-serve `Arc::try_unwrap` in `serve_state` can fail and we never
|
||
/// regain `&mut self`. A lock-free `swap(None)` lets [`shutdown`] take and
|
||
/// drop the db from `&self`, firing the HNSW-graph checkpoint inside the 60s
|
||
/// grace window regardless. Reads (`db`/`db_arc`) are wait-free `load_full`
|
||
/// clones.
|
||
///
|
||
/// [`shutdown`]: Self::shutdown
|
||
db: arc_swap::ArcSwapOption<TidalDb>,
|
||
/// gRPC transport: server on this region's `grpc_addr`; peers = siblings.
|
||
transport: Arc<GrpcTransport>,
|
||
/// The WAL's flushed-batch feed (m11p2): the one replicated log's
|
||
/// in-memory tail, populated by the engine's group-commit writer.
|
||
ship_feed: Arc<WalShipFeed>,
|
||
/// Per-peer windowed batch senders draining the ship feed off the request
|
||
/// path: `/signals`//`/items`//`/embeddings` ack at leader group-commit
|
||
/// fsync; these threads push the durable tail to followers, and a
|
||
/// follower that detects a gap pulls history via `StreamSegments`.
|
||
/// Dispatch is gated on leadership (an inactive follower queue parks).
|
||
ship_queue: ShipQueue,
|
||
/// This node's outbound stream baseline (see [`STREAM_BASELINE_FILE`]).
|
||
/// Shared with the gRPC `SegmentSource` so catch-up never serves
|
||
/// pre-stream history.
|
||
stream_baseline: Arc<AtomicU64>,
|
||
/// Durable "leader-acked frontier" (m12 election-divergence-fix): the highest
|
||
/// WAL seqno this node ACKED to a client under `ack=leader` (journal-only, no
|
||
/// quorum wait) and may therefore hold UN-REPLICATED. The election-divergence
|
||
/// classifier ([`decide_join`]) quarantines a rejoining node iff this exceeds
|
||
/// the new leadership's election baseline — exactly "leader-acked writes the
|
||
/// cluster elected past" (the divergent-suffix definition). Reset to 0 on a
|
||
/// clean term join (caught up ⇒ nothing un-replicated), advanced on each
|
||
/// `ack=leader` write, persisted on graceful shutdown, and falls back to the
|
||
/// durable WAL tail (conservative) when a hard kill leaves no persisted value.
|
||
/// This is what distinguishes a genuine `ack=leader` divergent suffix
|
||
/// (`mp_quarantined`) from a benign `ack=quorum` uncommitted tail
|
||
/// (rolling-restart-under-load): the latter never advances this frontier.
|
||
leader_acked: Arc<AtomicU64>,
|
||
/// The node's data dir (baseline persistence). Multi-process cluster
|
||
/// mode requires one — validated in [`Self::new`].
|
||
data_dir: std::path::PathBuf,
|
||
/// The topology's term-0 leader (the stream owner for pre-election logs
|
||
/// in the vote restriction's frontier comparison).
|
||
boot_topology_leader: RegionId,
|
||
/// This node's ELECTION-TIME log position, captured at leadership
|
||
/// activation BEFORE the term marker bumps the tail term (m11p4):
|
||
/// announced in every heartbeat so a joining follower can run the
|
||
/// divergence check in a comparable numbering. ONE lock: the pair is a
|
||
/// single logical value (a torn term/seq read would mis-judge
|
||
/// divergence), it is written once per won election and read a few
|
||
/// times per second. Stored only AFTER the activation's term marker is
|
||
/// durable, so readers never see values for an aborted activation.
|
||
activation_prev: std::sync::Mutex<tidaldb::replication::LogPosition>,
|
||
/// Shared `tidaldb_cluster_*` metrics cell (ship path, write pool, relay
|
||
/// frontiers), rendered by this node's `/metrics` listener.
|
||
cluster_metrics: Arc<tidaldb::db::metrics::cluster::ClusterMetrics>,
|
||
/// Current leadership view: which region this node believes leads, or
|
||
/// `None` during an election (leaderless windows are real and reported
|
||
/// honestly — forwards then return a retryable 503, m11p4).
|
||
leader: RwLock<Option<RegionId>>,
|
||
/// The election hooks cell shared with the gRPC transport (late-bound by
|
||
/// [`Self::start_election_driver`], like the applied sink).
|
||
election_hooks_cell: Arc<std::sync::OnceLock<Arc<dyn tidal_net::ElectionHooks>>>,
|
||
/// The election runtime (m11p4), set by [`Self::start_election_driver`].
|
||
election_runtime: std::sync::OnceLock<Arc<super::election_driver::ElectionRuntime>>,
|
||
/// Everything the driver build needs, prepared at construction and taken
|
||
/// once by [`Self::start_election_driver`].
|
||
election_boot: std::sync::Mutex<Option<ElectionBoot>>,
|
||
/// Leader-side ship-skip set: peers we are partitioned from do not receive
|
||
/// eager ships until healed.
|
||
partitioned: RwLock<HashSet<RegionId>>,
|
||
/// Serializes the partition/heal admin verbs so a `partition` arriving
|
||
/// mid-heal cannot interleave with the heal's remove → re-ship →
|
||
/// re-broadcast sequence (the verbs run start-to-finish in arrival order).
|
||
/// Deliberately NOT the `partitioned` `RwLock` itself: heal holds this for
|
||
/// its whole O(items) HTTP backfill, and holding the `RwLock` that long
|
||
/// would stall every signal write's skip-set read. Both verbs run on the
|
||
/// write pool, so a waiter blocks a pool worker, never the reactor.
|
||
admin_op: std::sync::Mutex<()>,
|
||
/// Region id → name, from topology declaration order. The era-0 inverse of
|
||
/// the name→id map; name→id resolution now goes through the effective roster
|
||
/// ([`Self::resolve_region`]), so only the id→name display path keeps this.
|
||
id_to_name: HashMap<RegionId, String>,
|
||
/// Peer region id → public HTTP address (for task-03 forwarding/aggregation).
|
||
peer_http: HashMap<RegionId, String>,
|
||
/// Shared async forwarding client (connection pooling across requests) for
|
||
/// write/read forwarding, item/embedding broadcast, promote fan-out, status
|
||
/// aggregation, and the reconcile snapshot exchange. See [`crate::cluster::forward`].
|
||
client: reqwest::Client,
|
||
/// m11p7 reloadable cluster credentials, shared with the owning `ClusterNode`
|
||
/// (one instance). The per-group forward/broadcast path mints a per-node
|
||
/// internal token from the cluster key so the receiving sibling's marker
|
||
/// guard sees a VERIFIED node. `None` cluster key ⇒ token dormant.
|
||
creds: Arc<crate::cluster::security::ClusterCreds>,
|
||
/// Per-peer budget for the item/embedding broadcast and promote fan-out,
|
||
/// resolved from the topology's optional `timeouts.broadcast_peer_secs`
|
||
/// (default [`forward::BROADCAST_PEER_TIMEOUT`]).
|
||
broadcast_peer_timeout: Duration,
|
||
/// Blocking client for the `/sharded/*` scatter-gather remote fetch (those
|
||
/// workers are detached OS threads with no tokio runtime, so they cannot use
|
||
/// the async client). HTTP is the only inter-node read transport this phase;
|
||
/// if a second transport (gRPC reads, native protocol) lands, extract a
|
||
/// `ShardTransport` trait in `scatter_gather` rather than branching here.
|
||
blocking_client: reqwest::blocking::Client,
|
||
/// Fixed-size runtime-free OS-thread pool for blocking write admission
|
||
/// (`/signals` staging, the cluster admin verbs). Shared by every write
|
||
/// request.
|
||
write_pool: ClusterWritePool,
|
||
/// Quorum commit index over the ship queue's peers (m11p3); leadership-
|
||
/// gated alongside the queue. `ack=quorum` writes await it through
|
||
/// `commit_watch` (async — never a parked thread per waiter).
|
||
commit: Arc<CommitIndex>,
|
||
/// Async mirror of `commit`: `(epoch, commit_index, active)`, published
|
||
/// by one dedicated bridge thread. Handlers `await` changes on a clone —
|
||
/// a thread-per-wait design exhausts the blocking pool under open-loop
|
||
/// load and starves the completions that advance the index (measured:
|
||
/// total quorum collapse at 1k rps).
|
||
commit_watch: tokio::sync::watch::Receiver<(u64, u64, bool)>,
|
||
/// Stops the commit-watch bridge thread on shutdown.
|
||
commit_bridge_stop: Arc<AtomicBool>,
|
||
/// Deployment-default write acknowledgment mode (`replication.ack`).
|
||
ack_default: AckMode,
|
||
/// Budget an `ack=quorum` write waits for the commit index before the
|
||
/// retryable 503 (`replication.quorum_timeout_ms`, default 2s).
|
||
quorum_timeout: Duration,
|
||
/// Flipped on shutdown so `/health` reports not-ready while draining.
|
||
shutting_down: AtomicBool,
|
||
/// The durable reseed-marker store (m11p5 §2.4): a running follower that hits
|
||
/// a typed `snapshot-required` refusal (or the divergence quarantine) latches
|
||
/// it here, then keeps serving degraded; the reseed runs on the next boot.
|
||
reseed_marker_store: ReseedMarkerStore,
|
||
/// `replication.reseed_self_restart` (m11p5 §2.4): when true, a marker latch
|
||
/// triggers a graceful drain + clean exit(0) — gated by the §2.4 quorum
|
||
/// refusal (never exit while the remaining voters cannot sustain quorum).
|
||
reseed_self_restart: bool,
|
||
/// `replication.learner_promote_lag` (m11p5 §3.3/§4): the readiness-
|
||
/// convergence hysteresis threshold AND the learner-promotion distance.
|
||
learner_promote_lag: u64,
|
||
/// Whether THIS node booted as a seed-join LEARNER (m11p5 §3.4/§4): a joiner
|
||
/// whose role in the recovered/boot roster is `Learner`. Readiness is 503
|
||
/// until it first-converges (lag ≤ `learner_promote_lag`), sticky-ready after
|
||
/// — even once the leader's auto-promotion duty flips it to `Voter`. Derived
|
||
/// at boot from the effective roster, no new constructor arg (§4).
|
||
seed_joiner: bool,
|
||
/// Whether THIS boot installed a snapshot (an install-boot joiner): readiness
|
||
/// is sticky-503 until it first converges (§4). Plain restarts are `false`
|
||
/// and keep today's readiness.
|
||
install_boot: bool,
|
||
/// Sticky readiness latch (m11p5 §4): set once this install-boot node's
|
||
/// catch-up lag first falls at or below `learner_promote_lag`; ready forever
|
||
/// after this process (hysteresis — never `lag == 0`, which an open-loop load
|
||
/// keeps perpetually false). Non-install boots ignore it (ready on today's
|
||
/// terms).
|
||
converged: AtomicBool,
|
||
/// m12 reseed-loop-fix (readiness gating): mirrors whether a reseed marker
|
||
/// (`SnapshotRequired` or `Quarantine`) is currently latched. `is_ready`
|
||
/// returns 503 while set, so a node that needs a reseed — INCLUDING a plain
|
||
/// restart (`install_boot == false`) that re-latched `SnapshotRequired` while
|
||
/// merely behind — is drained from the client VIP until it heals: it reseeds
|
||
/// on the next boot, or it catches up via the stream (which clears the marker
|
||
/// through `clear_stale_reseed_marker_if_caught_up`). Maintained by
|
||
/// `latch_reseed_marker` (set) and the marker-clear paths (cleared);
|
||
/// initialized at boot from the durable store. Readiness is thus BOUNDED
|
||
/// STALENESS (lag ≤ `learner_promote_lag` once converged, no unhealed reseed
|
||
/// marker), not "ready the instant the process is up".
|
||
reseed_marker_latched: AtomicBool,
|
||
/// Whether a `reseed_self_restart` was REFUSED by the §2.4 quorum check —
|
||
/// surfaced in status so a stuck self-restart is diagnosable. Set when the
|
||
/// node would have exited but the remaining voters cannot sustain quorum.
|
||
self_restart_refused: AtomicBool,
|
||
/// The transport's `snapshot-required` refusal sink cell (m11p5 §2.4),
|
||
/// late-bound by [`Self::start_election_driver`] with a sink that latches
|
||
/// this node's reseed marker.
|
||
snapshot_required_cell:
|
||
Arc<std::sync::OnceLock<Arc<dyn tidal_net::sources::SnapshotRequiredSink>>>,
|
||
/// Late-bound sink for COMPLETED catch-up pulls — the positive evidence that
|
||
/// discharges this node's reseed marker (a frontier comparison cannot; see
|
||
/// [`ShardReplica::discharge_reseed_marker_if_served`]).
|
||
catchup_served_cell: Arc<std::sync::OnceLock<Arc<dyn tidal_net::sources::CatchupServedSink>>>,
|
||
/// The effective roster (m11p5 §3): the single source of roster truth.
|
||
/// Derived at boot from the WAL-recovered `ClusterMembership` cell when
|
||
/// non-`None` (the membership era — RECORD ids), else from the topology
|
||
/// positional tables (era 0 — byte-for-byte today's behavior). A conf-change
|
||
/// swaps it through the one fenced apply path. `Arc` so the join adapter and
|
||
/// the (future) auto-promotion duty share it.
|
||
membership: Arc<super::membership::MembershipView>,
|
||
/// The transport's `JoinCluster` hooks cell (m11p5 §3.3), late-bound by
|
||
/// [`Self::start_election_driver`] with a leader-side join adapter holding a
|
||
/// `Weak<ShardReplica>`.
|
||
join_hooks_cell: Arc<std::sync::OnceLock<Arc<dyn tidal_net::JoinHooks>>>,
|
||
/// Guards the auto-promotion duty (m11p5 §3.3) so at most ONE promotion
|
||
/// evaluation runs at a time: the leader tick fires every 50ms, but a single
|
||
/// promotion's append + bounded same-term commit wait can outlast a tick. The
|
||
/// flag is set before spawning the duty and cleared when it finishes, so a
|
||
/// fast-ticking leader never piles up promotion threads while one is in flight.
|
||
promote_inflight: AtomicBool,
|
||
/// Self-driving heal (m11p8): coarse-cadence counter so the heal reconcile
|
||
/// pass runs roughly every `SELF_HEAL_TICKS` election ticks, not every 50ms.
|
||
heal_tick: AtomicU64,
|
||
/// Self-driving heal (m11p8): the set of peers this leader drove last pass
|
||
/// (breaker not closed AND trailing past the convergence threshold). Used to
|
||
/// count convergence transitions for `tidaldb_cluster_heal_successes_total`
|
||
/// and to clear the healing gauge on demotion.
|
||
healing: std::sync::Mutex<std::collections::HashSet<ShardId>>,
|
||
/// The durable membership cache (`data_dir/membership`, m11p5 §3.6). The
|
||
/// WAL-recovered `ClusterMembership` cell is the AUTHORITATIVE roster (it wins
|
||
/// at open — the view is built from it); this cache exists only for the
|
||
/// pre-open seed-join boot loop (seed lists + leader discovery). It is
|
||
/// rewritten from the cell after every applied record so a stale cache from an
|
||
/// older boot never out-of-dates the next pre-open loop.
|
||
membership_store: tidaldb::replication::MembershipStore,
|
||
/// The typed-removed-signal latch (m11p5 §3.3): set once any voter's
|
||
/// heartbeat / vote reply tells this node its own region is a `Removed`
|
||
/// member (the MISSED-RECORD delivery path — the node was down during the
|
||
/// removal-delivery grace and never folded the `Removed` record). Read by
|
||
/// [`Self::is_ready`] to flip readiness 503 even when the local cell never
|
||
/// learned of the removal. One-way: a removed node stays decommissioned
|
||
/// until the operator tears it down (runbook §8). Explicitly NOT a reseed
|
||
/// marker (a remove is not a reseed).
|
||
decommissioned_by_signal: AtomicBool,
|
||
/// Pending removed-peer ship-cell retirements (m11p5 §3.3 removal-delivery
|
||
/// grace): the leader defers a removed peer's `ShipQueue`/gRPC retirement
|
||
/// until the `Removed` record is same-term quorum-committed AND the removed
|
||
/// peer's applied mark covers the record's seq (it learned of its removal
|
||
/// through the log) — or a bounded give-up elapses. Polled on the election
|
||
/// driver tick ([`Self::poll_deferred_retires`]). The removed node must
|
||
/// learn of its removal BEFORE its ship cell retires; otherwise it never
|
||
/// receives the record, never stops campaigning, and zombie-serves as Ready.
|
||
deferred_retires: std::sync::Mutex<Vec<DeferredRetire>>,
|
||
}
|
||
|
||
/// One pending removed-peer ship-cell retirement (m11p5 §3.3). The leader holds
|
||
/// the removed peer's gRPC + ship cell open until the peer has DEMONSTRABLY
|
||
/// learned of its removal (its applied mark covers the `Removed` record's seq)
|
||
/// or a bounded deadline forces the retirement (a give-up the operator can see
|
||
/// in the metric/log).
|
||
#[derive(Debug, Clone)]
|
||
struct DeferredRetire {
|
||
/// The removed peer's shard (the ship cell + gRPC entry to retire).
|
||
shard: ShardId,
|
||
/// The removed peer's region name (for the log line).
|
||
name: String,
|
||
/// The `Removed` record's WAL seq: the peer's applied mark must reach this
|
||
/// (it then holds the record itself) before the retirement is safe.
|
||
record_seq: u64,
|
||
/// The hard give-up deadline: past this, retire anyway and WARN (a dead /
|
||
/// unreachable removed peer must not pin the ship cell forever).
|
||
deadline: Instant,
|
||
}
|
||
|
||
const RESEED_TERMINATION_LOG: &str = "/dev/termination-log";
|
||
|
||
fn write_reseed_termination_message(
|
||
path: &std::path::Path,
|
||
region: &str,
|
||
shard: ShardId,
|
||
) -> std::io::Result<()> {
|
||
let message = serde_json::json!({
|
||
"reason": "reseed_self_restart",
|
||
"region": region,
|
||
"shard": shard.0,
|
||
});
|
||
std::fs::write(path, message.to_string())
|
||
}
|
||
|
||
#[cfg(test)]
|
||
#[allow(clippy::unwrap_used)]
|
||
mod termination_message_tests {
|
||
use super::*;
|
||
|
||
#[test]
|
||
fn reseed_self_restart_writes_machine_readable_termination_reason() {
|
||
let suffix = std::time::SystemTime::now()
|
||
.duration_since(std::time::UNIX_EPOCH)
|
||
.unwrap()
|
||
.as_nanos();
|
||
let path = std::env::temp_dir().join(format!(
|
||
"tidaldb-termination-message-{}-{suffix}",
|
||
std::process::id()
|
||
));
|
||
|
||
write_reseed_termination_message(&path, "us-east", ShardId(7)).unwrap();
|
||
let message: serde_json::Value =
|
||
serde_json::from_slice(&std::fs::read(&path).unwrap()).unwrap();
|
||
std::fs::remove_file(path).unwrap();
|
||
|
||
assert_eq!(message["reason"], "reseed_self_restart");
|
||
assert_eq!(message["region"], "us-east");
|
||
assert_eq!(message["shard"], 7);
|
||
}
|
||
}
|
||
|
||
impl ShardReplica {
|
||
/// Build the single region named `region_name` from `topology`.
|
||
///
|
||
/// `RegionId`s are assigned by topology declaration order — the IDENTICAL
|
||
/// rule `ClusterState::new` uses — so every process that parses the same
|
||
/// topology file agrees on the region → id mapping.
|
||
///
|
||
/// `data_dir` selects a persistent `TidalDb` (else ephemeral); `hlc_offset_ms`
|
||
/// injects a wall-clock skew into this node's HLC (the clock-skew UAT surface).
|
||
///
|
||
/// # Blocking / threading
|
||
///
|
||
/// MUST be called from a non-async thread: [`GrpcTransport::new`] blocks on
|
||
/// its own tokio runtime and asserts it is not inside another runtime, so the
|
||
/// caller (`main::run_cluster`) hops to a dedicated `std::thread` — exactly
|
||
/// like `ClusterState::new`.
|
||
///
|
||
/// # Errors
|
||
///
|
||
/// Returns [`ServerError::SchemaConfig`] when the topology is invalid for
|
||
/// multi-process mode (missing addresses, unknown `region_name`, duplicate
|
||
/// names, undeclared leader), or [`ServerError::Cluster`] when the gRPC
|
||
/// transport cannot be built or its server does not become ready in time,
|
||
/// or [`ServerError::Tidal`] when the `TidalDb` fails to open.
|
||
///
|
||
/// # Panics
|
||
///
|
||
/// Does not panic on any caller input: the two internal `expect`s on the
|
||
/// region/leader name lookups are guarded by [`validate_multiproc`], which
|
||
/// runs first and proves both names are declared.
|
||
// Linear construction sequence (validate → db → sources → transport →
|
||
// receiver → queue); splitting it would scatter the ordering invariants.
|
||
#[allow(clippy::too_many_lines, clippy::too_many_arguments)]
|
||
pub fn new(
|
||
topology: &TopologySpec,
|
||
region_name: &str,
|
||
schema: Schema,
|
||
profiles: Vec<tidaldb::ranking::profile::RankingProfile>,
|
||
data_dir: Option<std::path::PathBuf>,
|
||
hlc_offset_ms: i64,
|
||
group: &ResolvedShardGroup,
|
||
enable_metrics: bool,
|
||
multi_shard: bool,
|
||
creds: Arc<crate::cluster::security::ClusterCreds>,
|
||
) -> Result<Self> {
|
||
super::topology::validate_multiproc(topology, region_name)?;
|
||
|
||
// m11p2: the durable WAL IS the replicated log — it is what ships to
|
||
// peers and what serves their catch-up streams. A node without a data
|
||
// dir has no log to lead with, so refuse loudly at startup.
|
||
let Some(data_dir) = data_dir else {
|
||
return Err(ServerError::SchemaConfig(
|
||
"multi-process cluster mode requires --data-dir: the durable WAL is \
|
||
the replication stream (m11p2)"
|
||
.into(),
|
||
));
|
||
};
|
||
|
||
// RegionIds by topology declaration order (identical rule to
|
||
// ClusterState::new), so all processes agree on the same ids.
|
||
let mut name_to_id = HashMap::new();
|
||
let mut id_to_name = HashMap::new();
|
||
for (i, region) in topology.regions.iter().enumerate() {
|
||
let id = RegionId(u16::try_from(i).map_err(|_| {
|
||
ServerError::SchemaConfig("topology declares more than 65535 regions".into())
|
||
})?);
|
||
name_to_id.insert(region.name.clone(), id);
|
||
id_to_name.insert(id, region.name.clone());
|
||
}
|
||
|
||
let region = *name_to_id
|
||
.get(region_name)
|
||
.expect("validate_multiproc proved region_name is declared");
|
||
// m11p6: leadership is per shard GROUP — the term-0 / preferred leader of
|
||
// THIS group, not a single cluster-wide leader. For the legacy single
|
||
// group this is `topology.leader` (the resolver synthesizes it), so S=1
|
||
// is byte-for-byte. In-group peer identity stays region-id based.
|
||
let leader = group.leader;
|
||
let my_shard = shard_of_region(region);
|
||
|
||
// m11p4 boot classification (phase-4.md §1): durable election state
|
||
// decides the boot role — NEVER the topology file alone. A clean
|
||
// rejoin always boots a follower (the §1.4-1 restart-amnesia fix);
|
||
// only a genuinely fresh node takes the topology's term-0 roles.
|
||
let election_store = tidaldb::replication::ElectionStore::new(&data_dir);
|
||
let wal_exists = data_dir
|
||
.join("wal")
|
||
.read_dir()
|
||
.map(|mut it| it.next().is_some())
|
||
.unwrap_or(false);
|
||
let boot_state = election_store.load(wal_exists).map_err(|e| {
|
||
ServerError::Cluster(format!(
|
||
"refusing to boot: {e} (a node that cannot prove its term must not guess term 0)"
|
||
))
|
||
})?;
|
||
let (hard_state, is_leader_at_boot) = match boot_state {
|
||
tidaldb::replication::BootState::Rejoin(h) => {
|
||
tracing::info!(
|
||
term = h.current_term,
|
||
"election state recovered; booting as a FOLLOWER (leadership is learned, never assumed from a restart)"
|
||
);
|
||
(h, false)
|
||
}
|
||
tidaldb::replication::BootState::Fresh => {
|
||
let h = tidaldb::replication::HardState {
|
||
current_term: 0,
|
||
voted_for: None,
|
||
};
|
||
election_store.persist(h).map_err(|e| {
|
||
ServerError::Cluster(format!("persist boot election state: {e}"))
|
||
})?;
|
||
(h, leader == region)
|
||
}
|
||
tidaldb::replication::BootState::StateFileLost => {
|
||
tracing::error!(
|
||
"election_state is MISSING but a WAL exists: the file was deleted out from under a node that has run before. Booting as a follower at term 0 (never re-assuming topology leadership on a guess); terms recover from the first leader contact"
|
||
);
|
||
let h = tidaldb::replication::HardState {
|
||
current_term: 0,
|
||
voted_for: None,
|
||
};
|
||
election_store.persist(h).map_err(|e| {
|
||
ServerError::Cluster(format!("persist recovered election state: {e}"))
|
||
})?;
|
||
(h, false)
|
||
}
|
||
};
|
||
// The boot leadership view: a fresh term-0 node trusts the topology;
|
||
// a rejoining node trusts only contact. Critically, a rejoining
|
||
// TERM-0 EX-LEADER gets `None`, never `Some(self)` — a self-leading
|
||
// view would re-open the §1.4-1 write-acceptance hole with the ship
|
||
// queue parked (locally-durable, never-shipped writes).
|
||
let initial_view: Option<RegionId> = if is_leader_at_boot {
|
||
Some(region)
|
||
} else if hard_state.current_term == 0 && leader != region {
|
||
Some(leader)
|
||
} else {
|
||
None
|
||
};
|
||
|
||
// This group's OTHER replica nodes + their gRPC/HTTP addresses (m11p6).
|
||
let PeerTables {
|
||
peer_shards,
|
||
peer_grpc,
|
||
peer_http,
|
||
my_grpc_spec,
|
||
my_grpc_bind_spec,
|
||
my_tls,
|
||
} = build_group_peer_tables(group, topology, region);
|
||
|
||
// One TidalDb owning this region (NodeRole::Single, shard = this region,
|
||
// peer_shards = siblings). Persistent (validated above); the HLC
|
||
// offset flows from the env surface into reconcile-time HLC stamping.
|
||
let db = open_region_db(
|
||
topology,
|
||
region_name,
|
||
schema,
|
||
profiles,
|
||
data_dir.clone(),
|
||
hlc_offset_ms,
|
||
my_shard,
|
||
&peer_shards,
|
||
enable_metrics,
|
||
)?;
|
||
let db = Arc::new(db);
|
||
|
||
// The WAL ship feed exists by construction (cluster peers + persistent
|
||
// mode); its absence would mean the engine gate drifted.
|
||
let ship_feed = db.wal_ship_feed().ok_or_else(|| {
|
||
ServerError::Cluster(
|
||
"engine did not create a WAL ship feed for this cluster node \
|
||
(persistent + peer_shards should guarantee one)"
|
||
.into(),
|
||
)
|
||
})?;
|
||
|
||
// This node's outbound stream baseline: 0 for the topology's original
|
||
// leader, the promote-time flushed frontier after a promotion.
|
||
// Persisted so a restarted leader never serves pre-stream history.
|
||
let stream_baseline = Arc::new(AtomicU64::new(load_stream_baseline(&data_dir)));
|
||
|
||
// One GrpcTransport: server on THIS region's grpc_addr (explicit — tried
|
||
// once, no port reallocation), peers = sibling regions. TLS comes from
|
||
// the region's optional `grpc_tls` topology block; absent ⇒ plaintext
|
||
// (the loopback/VPC posture every shipped topology uses). The serving
|
||
// sources (m11p2) hold WEAK db handles so the gRPC layer can never keep
|
||
// the database alive past this node's shutdown.
|
||
let sources = ServingSources {
|
||
election: Arc::default(),
|
||
applied: Some(Arc::new(NodeAppliedSource {
|
||
db: Arc::downgrade(&db),
|
||
})),
|
||
segments: Some(Arc::new(NodeSegmentSource {
|
||
db: Arc::downgrade(&db),
|
||
shard: my_shard,
|
||
baseline: Arc::clone(&stream_baseline),
|
||
feed: Arc::clone(&ship_feed),
|
||
})),
|
||
// Late-bound below (m11p3): the commit index only exists once
|
||
// the ship queue is spawned, which needs this transport first.
|
||
applied_sink: Arc::default(),
|
||
// m11p5: the snapshot source and the `snapshot-required` refusal
|
||
// sink are wired by stage B (the node-side install path); unset
|
||
// here makes `FetchSnapshot` answer Unimplemented and a
|
||
// snapshot-required trailer fall back to log + retry.
|
||
..ServingSources::default()
|
||
};
|
||
let applied_sink_cell = Arc::clone(&sources.applied_sink);
|
||
let election_hooks_cell = Arc::clone(&sources.election);
|
||
// m11p5 §2.4: the snapshot-required refusal sink cell, late-bound by
|
||
// `start_election_driver` (the sink holds a `Weak<ShardReplica>`,
|
||
// so it can only be built once the node is in its final Arc). Until set,
|
||
// a snapshot-required catch-up trailer falls back to log + retry.
|
||
let snapshot_required_cell = Arc::clone(&sources.snapshot_required);
|
||
// The positive-evidence counterpart, same late-bound discipline.
|
||
let catchup_served_cell = Arc::clone(&sources.catchup_served);
|
||
// m11p5: the snapshot-source cell, late-bound below once `cluster_metrics`
|
||
// exists (the source reports the staged/fetch/force-drop series). Until
|
||
// set, `FetchSnapshot` answers Unimplemented.
|
||
let snapshot_source_cell = Arc::clone(&sources.snapshots);
|
||
// m11p5 §3.3: the `JoinCluster` hooks cell, late-bound by
|
||
// `start_election_driver` (the adapter holds a `Weak<ShardReplica>`
|
||
// for the same reason as the snapshot-required sink). Until set,
|
||
// `JoinCluster` answers Unimplemented.
|
||
let join_hooks_cell = Arc::clone(&sources.join);
|
||
// m11p5 bind/advertise split: the local socket binds the DERIVED bind
|
||
// address (explicit `grpc_bind`, the literal `grpc_addr`, or
|
||
// `0.0.0.0:<port>` for a DNS `grpc_addr`), NOT the advertised address
|
||
// siblings dial. `my_grpc_spec` is the advertised value (it feeds the
|
||
// peer tables of OTHER nodes); here it only contributes its port.
|
||
let listen_addr = resolve_grpc_bind_addr(
|
||
my_grpc_bind_spec.as_deref(),
|
||
my_grpc_spec.as_deref(),
|
||
region_name,
|
||
)?;
|
||
// Keep a copy for the era-0 membership view (the transport consumes the
|
||
// original). The clone is one-time boot cost, not a hot path.
|
||
let peer_grpc_for_view = peer_grpc.clone();
|
||
let transport_defaults = GrpcTransportConfig::default();
|
||
let transport = GrpcTransport::new_with_sources(
|
||
GrpcTransportConfig {
|
||
local_shard: my_shard,
|
||
listen_addr,
|
||
peers: peer_grpc,
|
||
insecure: my_tls.is_none(),
|
||
// Clone so `my_tls` survives for the m11p7 inter-node HTTP
|
||
// forwarding clients below (they trust the same cluster CA).
|
||
tls: my_tls.clone(),
|
||
// m11p4: how long a FAILED catch-up pull waits before the
|
||
// timer re-pulls (the idle-cluster self-heal). Topology knob
|
||
// `replication.catchup_retry_ms`; default 30s.
|
||
catchup_retry_interval: topology.replication.catchup_retry_ms.map_or(
|
||
transport_defaults.catchup_retry_interval,
|
||
Duration::from_millis,
|
||
),
|
||
..transport_defaults
|
||
},
|
||
sources,
|
||
)
|
||
.map_err(|e| {
|
||
ServerError::Cluster(format!(
|
||
"build gRPC transport for region '{region_name}' on {listen_addr}: {e}"
|
||
))
|
||
})?;
|
||
if !grpc_server_ready(listen_addr) {
|
||
return Err(ServerError::Cluster(format!(
|
||
"gRPC server {listen_addr} for region '{region_name}' did not become ready \
|
||
within {GRPC_READY_TIMEOUT:?}"
|
||
)));
|
||
}
|
||
let transport = Arc::new(transport);
|
||
|
||
// m11p5 §2.6: post-open seed for an install boot. When the boot-time
|
||
// install swapped a snapshot in, the new data dir carries the
|
||
// `reseed-install-pending` sentinel — and it records the DISCOVERED
|
||
// leader's region (the snapshot was fetched from the CURRENT elected
|
||
// leader, which for a reseeded ex-leader is NOT the boot topology
|
||
// leader — that is the node itself). BEFORE the receiver, transport
|
||
// serving, or any pull, seed the DISCOVERED-leader-shard frontier to the
|
||
// installed artifact's captured stream seq (the sentinel's `snapshot_seq`;
|
||
// NOT `last_wal_seq()`, which a checkpoint-based restore leaves at 0 — see
|
||
// the seed below), persist the replication-state checkpoint SYNCHRONOUSLY,
|
||
// then delete
|
||
// the sentinel. A crash anywhere before the sentinel delete finds the
|
||
// WAL unchanged (no receiver ran) and re-derives the same seed
|
||
// idempotently. The discovered shard + tail are carried forward to issue
|
||
// the post-install catch-up pull once the receiver is up (below).
|
||
let install_target: Option<(ShardId, u64)> = match reseed::read_install_sentinel(&data_dir)
|
||
{
|
||
Ok(Some(sentinel)) => {
|
||
// The discovered leader's shard (the snapshot source). Fall back
|
||
// to the boot topology leader's shard only for a legacy 8-byte
|
||
// sentinel that predates the recorded region.
|
||
let leader_shard = sentinel
|
||
.leader_region
|
||
.map_or_else(|| shard_of_region(leader), |r| shard_of_region(RegionId(r)));
|
||
// The artifact's captured STREAM seq, recorded authoritatively in the
|
||
// install sentinel, IS the recovered frontier. Do NOT derive it from
|
||
// `db.last_wal_seq()`: a checkpoint-based restore leaves the engine's
|
||
// own WAL EMPTY (the data lives in the restored checkpoint/keyspaces,
|
||
// not a replayed WAL), so `last_wal_seq()` reads 0 — which seeds frontier
|
||
// 0, makes the post-install pull request from seqno 1, which the leader's
|
||
// long-compacted WAL cannot serve → `snapshot_required` re-latch → reseed
|
||
// loop. (Observed in prod on an 805 MB / seq-1.04M artifact whose WAL
|
||
// restored empty; small-corpus e2e artifacts restore a non-empty WAL so
|
||
// they never tripped it.) `snapshot_seq` is the leader's stream position
|
||
// the artifact represents — exactly what the catch-up pull below
|
||
// (`recovered_tail + 1`) must resume from.
|
||
let recovered_tail = sentinel.snapshot_seq;
|
||
db.replication_state().advance(leader_shard, recovered_tail);
|
||
// m12 reseed-loop-fix: durably synthesize the term marker the
|
||
// installed artifact represents. A checkpoint-based restore leaves
|
||
// the engine's own WAL EMPTY (the data lives in the restored
|
||
// keyspaces, not a replayed WAL), so `wal_term_mark()` reports
|
||
// tail_term 0 — and `decide_join` compares `(tail_term, frontier)`
|
||
// lexically, tail_term FIRST, so `0 < leader_term` classifies this
|
||
// shard `ReseedRequired` on EVERY boot regardless of the correctly-
|
||
// seeded frontier → an infinite reseed loop (observed on tidaldb-2,
|
||
// the 805 MB / seq-1.04M artifact whose WAL restored empty). Writing
|
||
// a real kind-3 TERM_MARKER record at the artifact's captured term +
|
||
// the reseed-leader's region makes tail_term truthful on this boot
|
||
// AND every reboot (blob records are deliberately NOT checkpoint-
|
||
// filtered on recovery — wal/reader.rs), so `election_log_position`
|
||
// reads the frontier in the DISCOVERED-leader stream (the shard just
|
||
// seeded) and `decide_join` returns Clean. Truthful, not a bypass:
|
||
// the artifact IS the leader's authoritative state at (term, seq) and
|
||
// the install discarded any prior suffix, so a genuinely-divergent
|
||
// node (which has NO install sentinel) still surfaces tail_term >
|
||
// term → Quarantine. Crash-idempotent: once a prior attempt's marker
|
||
// is folded by recovery the guard (`< artifact_term`, monotonic-by-
|
||
// term) skips the re-append. Skipped for a legacy sentinel that
|
||
// carries no term (cannot fabricate one) and for the term-0 topology
|
||
// era (which never journals a marker).
|
||
if let (Some(artifact_term), Some(marker_region)) =
|
||
(sentinel.artifact_term, sentinel.leader_region)
|
||
&& artifact_term > 0
|
||
&& db.wal_term_mark().0 < artifact_term
|
||
{
|
||
db.append_term_marker(artifact_term, marker_region)
|
||
.map_err(|e| {
|
||
ServerError::Cluster(format!(
|
||
"reseed post-open seed: synthesize term marker (term \
|
||
{artifact_term}, region {marker_region}) failed: {e}"
|
||
))
|
||
})?;
|
||
tracing::info!(
|
||
region = region_name,
|
||
artifact_term,
|
||
marker_region,
|
||
recovered_tail,
|
||
"reseed install boot: durably synthesized the artifact's term marker so \
|
||
wal_term_mark reports its term (not 0) on this and every reboot — \
|
||
decide_join now classifies the reseeded shard on a truthful (term, \
|
||
frontier) instead of looping ReseedRequired (§2.6 m12 reseed-loop-fix)"
|
||
);
|
||
}
|
||
db.persist_replication_checkpoint().map_err(|e| {
|
||
ServerError::Cluster(format!(
|
||
"reseed post-open seed: persist replication checkpoint failed: {e}"
|
||
))
|
||
})?;
|
||
reseed::clear_install_sentinel(&data_dir)?;
|
||
tracing::info!(
|
||
region = region_name,
|
||
leader_shard = leader_shard.0,
|
||
recovered_tail,
|
||
discovered_leader = sentinel.leader_region,
|
||
"reseed install boot: seeded the DISCOVERED-leader-shard frontier to the \
|
||
artifact's captured stream seq (sentinel snapshot_seq) and persisted the \
|
||
checkpoint (§2.6); a catch-up pull toward that shard resumes the stream \
|
||
after the receiver starts"
|
||
);
|
||
Some((leader_shard, recovered_tail))
|
||
}
|
||
Ok(None) => None,
|
||
Err(e) => return Err(e),
|
||
};
|
||
let install_boot = install_target.is_some();
|
||
|
||
// m12 election-divergence-fix: the durable leader-acked frontier. Each
|
||
// `ack=leader` write persists it on advance (so a HARD KILL recovers the
|
||
// exact un-replicated suffix — `mp_quarantined`); a clean join persists 0.
|
||
// Absent ⇒ 0: a node that never ack=leader-wrote (a follower, or a leader
|
||
// that only served `ack=quorum`) holds nothing un-replicated. Crucially we
|
||
// do NOT fall back to the WAL tail — a follower's `flushed` is
|
||
// applied-from-leader (replicated) data, and treating it as leader-acked
|
||
// would false-quarantine a clean former leader whose tail was quorum-replicated.
|
||
let leader_acked = Arc::new(AtomicU64::new(load_leader_acked(&data_dir).unwrap_or(0)));
|
||
|
||
// Always-on receiver: leadership can move, so EVERY node runs a receiver
|
||
// (an inbound segment is legal on the current leader after a promote
|
||
// elsewhere). The blob-applier wiring routes replicated kind-1/2
|
||
// records through this db's WAL-first item write paths (m11p2).
|
||
db.start_replication_with_blobs(Arc::clone(&transport))
|
||
.map_err(ServerError::Tidal)?;
|
||
|
||
let pool_config = topology.write_pool_config();
|
||
tracing::info!(
|
||
region = region_name,
|
||
%listen_addr,
|
||
peers = peer_shards.len(),
|
||
leader = %topology.leader,
|
||
workers = pool_config.workers,
|
||
"region cluster node started (one region, gRPC peers at sibling addrs)"
|
||
);
|
||
let cluster_metrics = db.cluster_metrics();
|
||
|
||
// m11p5: wire the leader-side snapshot source for `FetchSnapshot`. A
|
||
// `Weak<TidalDb>` (like `NodeSegmentSource`) so the transport can never
|
||
// keep the database alive past shutdown. The artifact-reuse TTL comes
|
||
// from `replication.snapshot_artifact_ttl_ms` (default 600000); it
|
||
// governs reuse only — never the pin of an active consumer (§2.1).
|
||
let snapshot_ttl = Duration::from_millis(
|
||
topology
|
||
.replication
|
||
.snapshot_artifact_ttl_ms
|
||
.unwrap_or(600_000),
|
||
);
|
||
let _ = snapshot_source_cell.set(Arc::new(super::snapshot::NodeSnapshotSource::new(
|
||
&db,
|
||
data_dir.clone(),
|
||
Arc::clone(&stream_baseline),
|
||
snapshot_ttl,
|
||
Arc::clone(&cluster_metrics),
|
||
)) as Arc<dyn tidal_net::sources::SnapshotSource>);
|
||
|
||
let write_pool =
|
||
ClusterWritePool::with_metrics(pool_config, Some(Arc::clone(&cluster_metrics)));
|
||
// m11p7: inter-node forwarding clients trust the cluster CA when TLS is
|
||
// configured, so `https://` forwards verify the peer's server cert.
|
||
let (client, blocking_client) = build_forwarding_clients(my_tls.as_ref())?;
|
||
|
||
// Per-peer senders over the WAL ship feed. Dispatch is leadership-
|
||
// gated: a follower's queue parks until a promote activates it.
|
||
let ship_queue = ShipQueue::spawn(
|
||
Arc::new(WalFeedSource::new(my_shard, Arc::clone(&ship_feed))),
|
||
Arc::clone(&transport) as Arc<dyn Transport>,
|
||
&peer_shards,
|
||
topology.ship_queue_config().with_leader_region(region.0),
|
||
is_leader_at_boot,
|
||
Some(Arc::clone(&cluster_metrics)),
|
||
);
|
||
|
||
// Follower boot self-heal (m11p2): proactively pull anything this
|
||
// node missed while down. If it is already caught up the stream opens
|
||
// and closes empty — one cheap RPC; if it is behind, catch-up starts
|
||
// immediately instead of waiting for the next live ship to expose the
|
||
// gap. (`request_catchup` is non-blocking and single-flight.)
|
||
//
|
||
// m11p5 §2.6 (install-boot catch-up correction): after a snapshot
|
||
// install, the pull must target the DISCOVERED leader's shard from the
|
||
// seeded artifact tail + 1 — NOT the boot topology leader, which for a
|
||
// reseeded ex-leader is the node itself (a self-pull never converges).
|
||
// The plain self-heal targets the boot topology leader — but ONLY in the
|
||
// genuine topology era (durable term 0). Post-election that field is dead
|
||
// config, and pulling against it is actively harmful:
|
||
//
|
||
// `applied_seqno` is keyed per LEADER REGION, so a group retains a position
|
||
// for every leadership it has followed. Live tidaldb-0's group 1 held
|
||
// keys [[0, 13540659], [1, 13540652], [2, 13540661]] — current leader
|
||
// tidaldb-2 (key 2) fully converged at 13540661, alongside a STALE key 1
|
||
// from when tidaldb-1 led the group. The old code pulled the topology
|
||
// leader's key regardless, i.e. key 1 from 13540652 + 1 = 13540653, which
|
||
// tidaldb-1's WAL had compacted below (earliest 13540657) → a permanent
|
||
// `snapshot-required` refusal → marker latch → `reseed_self_restart` loop.
|
||
// The comment here used to claim "term fencing + later election traffic
|
||
// rescue it"; they do not, because the refusal re-latches faster than the
|
||
// rescue converges.
|
||
//
|
||
// Post-election, convergence is driven by the heartbeat path
|
||
// (`note_leader_frontier_for_readiness`, which carries the CURRENT leader's
|
||
// frontier and works on an idle cluster since m12p5) and by the receiver's
|
||
// gap detection on real ship traffic — both keyed to the leader that is
|
||
// actually shipping, never a historical one.
|
||
if let Some((discovered_shard, recovered_tail)) = install_target {
|
||
transport.request_catchup(discovered_shard, recovered_tail + 1);
|
||
} else if !is_leader_at_boot && hard_state.current_term == 0 {
|
||
let leader_shard = shard_of_region(leader);
|
||
// A node can NEVER pull its own stream from itself: only the node whose
|
||
// `source_shard == S` serves shard S's stream (server.rs returns
|
||
// NOT_FOUND otherwise), and in full placement the group's TOPOLOGY
|
||
// leader IS self for the group this node leads — so the old
|
||
// unconditional pull emitted `PeerUnreachable(my_shard)` forever
|
||
// (handle_for(my_shard) is None; self is never a registered peer). For a
|
||
// self-led group the receiver gap path (from real ship source shards)
|
||
// and live election traffic drive convergence instead.
|
||
if leader_shard != my_shard {
|
||
let applied = db
|
||
.replication_state()
|
||
.applied_seqno(leader_shard)
|
||
.unwrap_or(0);
|
||
transport.request_catchup(leader_shard, applied + 1);
|
||
}
|
||
}
|
||
|
||
let commit = ship_queue.commit_index();
|
||
// Wire follower frontier reports (gRPC `ReportApplied`) into the
|
||
// quorum commit index: the transport's server folds each report's
|
||
// durable mark through this sink. Decoupled from ship acks by
|
||
// design — see `Transport::notify_applied`.
|
||
let _ = applied_sink_cell.set(Arc::new(CommitIndexSink {
|
||
commit: Arc::clone(&commit),
|
||
}) as Arc<dyn tidal_net::sources::AppliedSink>);
|
||
let ack_default = topology
|
||
.replication
|
||
.ack
|
||
.as_deref()
|
||
.and_then(AckMode::parse)
|
||
.unwrap_or(AckMode::Leader);
|
||
let quorum_timeout =
|
||
Duration::from_millis(topology.replication.quorum_timeout_ms.unwrap_or(2_000));
|
||
|
||
// The async commit bridge: ONE thread blocks on the index's condvar
|
||
// and republishes every change into a watch channel that any number
|
||
// of request handlers await for free.
|
||
let (commit_tx, commit_watch) = tokio::sync::watch::channel(commit.snapshot());
|
||
let commit_bridge_stop = Arc::new(AtomicBool::new(false));
|
||
{
|
||
let commit = Arc::clone(&commit);
|
||
let stop = Arc::clone(&commit_bridge_stop);
|
||
std::thread::Builder::new()
|
||
.name("tidal-commit-watch".into())
|
||
.spawn(move || {
|
||
let (mut epoch, mut idx, _) = commit.snapshot();
|
||
while !stop.load(Ordering::Acquire) {
|
||
// The wake cap bounds shutdown latency; an unchanged
|
||
// index republishes nothing (send_if_modified).
|
||
let next = commit.wait_change(epoch, idx, COMMIT_BRIDGE_WAKE_INTERVAL);
|
||
(epoch, idx) = (next.0, next.1);
|
||
commit_tx.send_if_modified(|cur| {
|
||
if *cur == next {
|
||
false
|
||
} else {
|
||
*cur = next;
|
||
true
|
||
}
|
||
});
|
||
}
|
||
})
|
||
.expect("spawn commit-watch bridge thread");
|
||
}
|
||
|
||
// Prepared for the election driver (started once the node is in its
|
||
// final Arc — see `start_election_driver`).
|
||
let election_boot = ElectionBoot {
|
||
config: super::election_driver::election_config(
|
||
&topology.election,
|
||
region,
|
||
// m11p6: the election's voter set is THIS group's replica nodes
|
||
// (excluding self), not every region. For the legacy single
|
||
// group this is every other region — byte-for-byte.
|
||
group
|
||
.replicas
|
||
.iter()
|
||
.map(|r| r.region)
|
||
.filter(|&r| r != region)
|
||
.collect(),
|
||
true,
|
||
),
|
||
hard: hard_state,
|
||
boots_as_leader: is_leader_at_boot,
|
||
store: election_store,
|
||
topology_leader: leader,
|
||
};
|
||
|
||
// The effective roster (m11p5 §3): precedence is the WAL-recovered
|
||
// `ClusterMembership` cell over the topology (§3.6) — the cell wins
|
||
// because it is the replicated truth; the topology is only the era-0
|
||
// bootstrap. A cell present means a kind-4 record is in this node's log
|
||
// (the membership era has begun); absent means era 0, and the view is
|
||
// built byte-for-byte from the same positional tables the rest of boot
|
||
// used, so era-0 behavior is unchanged.
|
||
let membership = match db.cluster_membership() {
|
||
Some((version, mterm, members)) => {
|
||
tracing::info!(
|
||
version,
|
||
term = mterm,
|
||
members = members.len(),
|
||
"membership recovered from the WAL (kind-4 era); roster is record-derived"
|
||
);
|
||
Arc::new(super::membership::MembershipView::from_record(
|
||
region, version, mterm, members,
|
||
))
|
||
}
|
||
None => Arc::new(super::membership::MembershipView::era0(
|
||
region,
|
||
group,
|
||
topology,
|
||
&name_to_id,
|
||
&id_to_name,
|
||
&peer_http,
|
||
&peer_grpc_for_view,
|
||
)),
|
||
};
|
||
|
||
// m11p5 §3.4/§4: a seed-join LEARNER boot is a joiner — 503 until
|
||
// first-converged, sticky after. Derived from the effective roster
|
||
// (the snapshot install carried the leader's WAL with this node's
|
||
// kind-4 Learner record, so `from_record` boots `self_role() ==
|
||
// Learner`; the small-cluster `needed=false` join boots era-0, but the
|
||
// synthesized topology marks the joiner a voter only if the join roster
|
||
// did — a fresh learner reads `Learner` here either way once the cell
|
||
// applies). No new constructor arg.
|
||
let seed_joiner = membership.self_role() == Some(tidaldb::wal::format::MemberRole::Learner);
|
||
|
||
// The reseed-marker store binds the same data dir the struct moves into
|
||
// its `data_dir` field; clone the path before the move.
|
||
let data_dir_for_state = data_dir.clone();
|
||
|
||
// §3.6 boot-time cache reconcile: when the view booted FROM THE CELL (a
|
||
// kind-4 record is in the WAL), rewrite the durable cache from the cell so
|
||
// a STALE cache (an older boot's roster, e.g. pre-promotion) is corrected
|
||
// at open — the cell is authoritative and wins, the cache only feeds the
|
||
// pre-open boot loop. Best-effort (a persist failure only slows a future
|
||
// discovery, never correctness). Era 0 (no cell) leaves any cache alone.
|
||
{
|
||
let roster = membership.roster();
|
||
if roster.from_record {
|
||
let store = tidaldb::replication::MembershipStore::new(&data_dir_for_state);
|
||
let snapshot = tidaldb::replication::MembershipSnapshot {
|
||
version: roster.version,
|
||
term: roster.term,
|
||
members: roster.members.clone(),
|
||
};
|
||
if let Err(e) = store.persist(&snapshot) {
|
||
tracing::warn!(
|
||
version = roster.version,
|
||
error = %e,
|
||
"boot: failed to reconcile the membership cache from the WAL cell (§3.6); \
|
||
the cell is authoritative so this is non-fatal"
|
||
);
|
||
} else {
|
||
tracing::info!(
|
||
version = roster.version,
|
||
"boot: reconciled the durable membership cache from the WAL-recovered \
|
||
cell (§3.6 precedence — the cell wins, the cache is rewritten)"
|
||
);
|
||
}
|
||
}
|
||
}
|
||
|
||
// m12 reseed-loop-fix (readiness gating): pre-load the durable reseed
|
||
// marker so `is_ready` reflects a still-latched marker at boot — a
|
||
// degraded install-fallback that kept its marker (the reseed did NOT
|
||
// complete) must boot 503, not serve stale data. Maintained live by the
|
||
// latch/clear paths thereafter.
|
||
let reseed_marker_store = ReseedMarkerStore::new(&data_dir_for_state);
|
||
let reseed_marker_latched_init = matches!(reseed_marker_store.load(), Ok(Some(_)));
|
||
|
||
Ok(Self {
|
||
region,
|
||
region_name: region_name.to_string(),
|
||
group_shard: group.shard,
|
||
multi_shard,
|
||
db: arc_swap::ArcSwapOption::new(Some(db)),
|
||
transport,
|
||
ship_feed,
|
||
ship_queue,
|
||
stream_baseline,
|
||
leader_acked,
|
||
data_dir,
|
||
boot_topology_leader: leader,
|
||
activation_prev: std::sync::Mutex::new(tidaldb::replication::LogPosition {
|
||
tail_term: 0,
|
||
frontier: 0,
|
||
}),
|
||
cluster_metrics,
|
||
leader: RwLock::new(initial_view),
|
||
election_hooks_cell,
|
||
election_runtime: std::sync::OnceLock::new(),
|
||
election_boot: std::sync::Mutex::new(Some(election_boot)),
|
||
partitioned: RwLock::new(HashSet::new()),
|
||
admin_op: std::sync::Mutex::new(()),
|
||
id_to_name,
|
||
peer_http,
|
||
client,
|
||
creds,
|
||
broadcast_peer_timeout: topology.broadcast_peer_timeout(),
|
||
blocking_client,
|
||
write_pool,
|
||
commit,
|
||
commit_watch,
|
||
commit_bridge_stop,
|
||
ack_default,
|
||
quorum_timeout,
|
||
shutting_down: AtomicBool::new(false),
|
||
reseed_marker_store,
|
||
reseed_marker_latched: AtomicBool::new(reseed_marker_latched_init),
|
||
reseed_self_restart: topology.replication.reseed_self_restart.unwrap_or(false),
|
||
learner_promote_lag: topology.replication.learner_promote_lag.unwrap_or(1024),
|
||
seed_joiner,
|
||
install_boot,
|
||
converged: AtomicBool::new(false),
|
||
self_restart_refused: AtomicBool::new(false),
|
||
snapshot_required_cell,
|
||
catchup_served_cell,
|
||
membership,
|
||
join_hooks_cell,
|
||
promote_inflight: AtomicBool::new(false),
|
||
heal_tick: AtomicU64::new(0),
|
||
healing: std::sync::Mutex::new(std::collections::HashSet::new()),
|
||
membership_store: tidaldb::replication::MembershipStore::new(&data_dir_for_state),
|
||
decommissioned_by_signal: AtomicBool::new(false),
|
||
deferred_retires: std::sync::Mutex::new(Vec::new()),
|
||
})
|
||
}
|
||
|
||
// ── Lifecycle ────────────────────────────────────────────────────────────
|
||
|
||
/// Flip `/health` to not-ready so the load balancer drains this node.
|
||
pub fn set_shutting_down(&self) {
|
||
self.shutting_down.store(true, Ordering::Release);
|
||
}
|
||
|
||
/// True once shutdown began (drain in progress or db already closed).
|
||
#[must_use]
|
||
pub fn is_shutting_down(&self) -> bool {
|
||
self.shutting_down.load(Ordering::Acquire)
|
||
}
|
||
|
||
/// Drop this region's `TidalDb` (checkpoint + WAL fsync + HNSW-graph
|
||
/// checkpoint + thread join) and signal the segment receiver to exit.
|
||
/// Idempotent.
|
||
///
|
||
/// Takes `&self` (m12p6): the db lives in an [`arc_swap::ArcSwapOption`], so
|
||
/// the deterministic close runs even when only a SHARED reference is reachable
|
||
/// — the case `serve_state` hits when a stuck peer connection blocks the
|
||
/// graceful drain and the post-serve `Arc::try_unwrap` fails. `swap(None)`
|
||
/// removes the db handle; dropping the returned `Arc` fires `TidalDb::Drop`
|
||
/// (and thus the HNSW-graph checkpoint) when it is the last strong reference,
|
||
/// which it is once the request-scoped clones have drained.
|
||
pub fn shutdown(&self) {
|
||
self.set_shutting_down();
|
||
if let Some(rt) = self.election_runtime.get() {
|
||
// Graceful leadership hand-off (rolling-restart churn fix): if this
|
||
// node leads an ELECTED term, drain its flushed tail to a quorum
|
||
// BEFORE stepping down. Without this, a gracefully-restarted leader
|
||
// abandons its just-journaled term marker (and any leader-durable
|
||
// tail) un-replicated; the survivors elect a new leader WITHOUT it,
|
||
// and the deposed leader rejoins carrying a divergent suffix the new
|
||
// term never saw → quarantine + full snapshot reseed on EVERY rolling
|
||
// deploy/upgrade/reboot. Committing the tail first means the next
|
||
// leader holds the full prefix (vote restriction) and this node
|
||
// rejoins clean. Bounded; steps down anyway past the deadline.
|
||
//
|
||
// m12 election-divergence-fix: drain whenever LEADING, including the
|
||
// term-0 TOPOLOGY leader. Handing off its committed prefix lets the
|
||
// survivors elect cleanly and converge fast (empirically tighter than
|
||
// leaving an un-handed-off tail). The membership-propagation timing this
|
||
// shifted is now absorbed by a polled roster assertion (the seed-join
|
||
// test). `needed_peers() == 0` still short-circuits a single replica.
|
||
if self.is_leader() {
|
||
self.drain_committed_before_stepdown();
|
||
}
|
||
rt.stop();
|
||
}
|
||
// m12 election-divergence-fix: persist the leader-acked frontier on a
|
||
// GRACEFUL shutdown (writes are already quiesced by `set_shutting_down`,
|
||
// so the value is final). A hard kill skips this and the next boot falls
|
||
// back to the conservative WAL tail. After a clean drain the value is
|
||
// already 0 (reset on the last clean join) or fully committed, so a
|
||
// gracefully-restarted node rejoins clean rather than re-quarantining.
|
||
self.persist_leader_acked_now();
|
||
self.commit_bridge_stop.store(true, Ordering::Release);
|
||
// Quiesce the ship-queue senders FIRST so no batch ship races the
|
||
// transport/db teardown below (an in-flight send_segment against a
|
||
// half-down transport would only add noisy shutdown errors). `deactivate`
|
||
// is `&self` (parks dispatch, deactivates the commit index); the sender
|
||
// threads are then JOINED by `ShipQueue::Drop` when this `ShardReplica`
|
||
// is finally dropped. We do not call `ShipQueue::shutdown` (which would
|
||
// join here) because it requires `&mut self`, and m12p6's shared-`&self`
|
||
// shutdown must reach the db close even when only an `Arc` is held; the
|
||
// join is non-durable bookkeeping, the db close below is what matters.
|
||
self.ship_queue.deactivate();
|
||
// Wake the always-on receiver so its thread can exit and join, then drop
|
||
// the db (the receiver handle lives inside TidalDb).
|
||
self.transport.shutdown_receivers();
|
||
// `swap(None)` is the take: the first caller gets the db `Arc` and drops
|
||
// it (firing the deterministic close); a second call swaps `None`→`None`
|
||
// and drops nothing — idempotent.
|
||
if let Some(db) = self.db.swap(None) {
|
||
// m12p6 FIX: run the deterministic close on the SHARED handle EXPLICITLY
|
||
// rather than relying on `drop(db)` reaching `TidalDb::Drop`. Under load
|
||
// a request-scoped clone (or a background task) often still holds the
|
||
// `Arc` at this point, so `drop` here is NOT the last reference, `Drop`
|
||
// never fires, and `checkpoint_embedding_graphs` never runs — the
|
||
// observed cluster gap where a follower restart rebuilt the HNSW graph
|
||
// because no `{data_dir}/vector` file was written on SIGTERM. `close_shared`
|
||
// takes `&self` and is idempotent (the `closed` CAS), so the trailing
|
||
// `drop(db)` and any later `Drop` from a lingering clone are no-ops.
|
||
// m12p6 7th-edge: when this shard is RESEED-PENDING (a divergent /
|
||
// quarantined node about to be re-baselined — its marker latched), skip
|
||
// the HNSW-graph save. The in-memory index reflects suspect data the
|
||
// next boot DISCARDS via the snapshot install, so saving it is
|
||
// pointless and risks committing a graph the next open fails to read
|
||
// ("Failed to read vectors" → a slow rebuild). The durable checkpoints
|
||
// + WAL flush still run.
|
||
let reseed_pending = matches!(self.reseed_marker_store.load(), Ok(Some(_)));
|
||
if let Err(e) = db.close_shared(!reseed_pending) {
|
||
tracing::error!(
|
||
region = %self.region_name,
|
||
error = %e,
|
||
"region cluster node shutdown: deterministic close reported a \
|
||
durable-flush error (state may be partially flushed)"
|
||
);
|
||
}
|
||
drop(db);
|
||
tracing::info!(
|
||
region = %self.region_name,
|
||
reseed_pending,
|
||
"region cluster node shutdown: database closed (checkpoint + WAL fsync + HNSW graph unless reseed-pending)"
|
||
);
|
||
}
|
||
}
|
||
|
||
/// Wait (bounded by [`SHUTDOWN_HANDOFF_WAIT`]) until this leader's flushed WAL
|
||
/// tail is COMMITTED to a quorum, so a graceful step-down leaves no divergent
|
||
/// suffix for the deposed leader to reseed over. Reuses the exact signal the
|
||
/// commit path already maintains — `commit.committed()` (the quorum-acked
|
||
/// frontier, fed by peers' `ReportApplied`) vs `ship_feed.flushed_seq()` (this
|
||
/// leader's durable WAL tail). No-op when no quorum is needed (single replica)
|
||
/// or the tail is already committed; logs and proceeds past the deadline so
|
||
/// shutdown never hangs. Called only while still leading and before the ship
|
||
/// queue is deactivated, so the in-flight tail can still ship and be acked.
|
||
fn drain_committed_before_stepdown(&self) {
|
||
if self.commit.needed_peers() == 0 {
|
||
return; // no quorum to wait on (single replica) — nothing to divergence-proof
|
||
}
|
||
let flushed = self.ship_feed.flushed_seq();
|
||
if flushed == 0 || self.commit.committed() >= flushed {
|
||
return; // empty stream or already fully committed
|
||
}
|
||
let deadline = std::time::Instant::now() + SHUTDOWN_HANDOFF_WAIT;
|
||
loop {
|
||
let committed = self.commit.committed();
|
||
if committed >= flushed {
|
||
tracing::info!(
|
||
region = %self.region_name,
|
||
flushed,
|
||
"graceful shutdown: leader tail committed to quorum before step-down (clean hand-off)"
|
||
);
|
||
return;
|
||
}
|
||
if std::time::Instant::now() >= deadline {
|
||
tracing::warn!(
|
||
region = %self.region_name,
|
||
committed,
|
||
flushed,
|
||
"graceful shutdown: leader tail NOT fully committed before the hand-off \
|
||
deadline; stepping down anyway (a follower may briefly reconcile the small \
|
||
uncommitted tail — no acked-write loss, the WAL is durable)"
|
||
);
|
||
return;
|
||
}
|
||
std::thread::sleep(std::time::Duration::from_millis(50));
|
||
}
|
||
}
|
||
|
||
// ── Accessors ──────────────────────────────────────────────────────────
|
||
|
||
/// Clone the node's `TidalDb` handle, or `Unavailable` once shutdown took it.
|
||
///
|
||
/// Returns an owned `Arc` (a wait-free `ArcSwapOption::load_full`): every
|
||
/// caller does `let db = self.db()?; db.method(…)`, and `Arc<TidalDb>` derefs
|
||
/// to `&TidalDb`, so the owned handle is drop-in for the old `&Arc` borrow.
|
||
fn db(&self) -> Result<Arc<TidalDb>> {
|
||
self.db
|
||
.load_full()
|
||
.ok_or_else(|| ServerError::Unavailable("server shutting down".into()))
|
||
}
|
||
|
||
/// Clone the `TidalDb` handle (for `move` into an offloaded closure), or
|
||
/// `Unavailable` once shutdown took it. Identical to [`Self::db`] now that the
|
||
/// handle is `ArcSwapOption`-backed; kept as a distinct name for call-site
|
||
/// intent (this one's result is moved into a `'static` worker).
|
||
fn db_arc(&self) -> Result<Arc<TidalDb>> {
|
||
self.db
|
||
.load_full()
|
||
.ok_or_else(|| ServerError::Unavailable("server shutting down".into()))
|
||
}
|
||
|
||
/// Map a region name to its id; `BadRequest` for an undeclared name (the
|
||
/// 400 every region-targeted route surfaces). Reads through the effective
|
||
/// roster (§3): era 0 = the topology positional ids, byte-for-byte; the
|
||
/// membership era = the record ids (a joined region resolves, a never-joined
|
||
/// name is unknown).
|
||
fn resolve_region(&self, name: &str) -> Result<RegionId> {
|
||
self.membership
|
||
.name_to_id(name)
|
||
.ok_or_else(|| ServerError::BadRequest(format!("unknown region '{name}'")))
|
||
}
|
||
|
||
/// The region's declared name (`"unknown"` for an id outside the topology,
|
||
/// which only display paths ever see). The id→name display path keeps the
|
||
/// topology-order table; roster-aware name resolution (which a conf-change
|
||
/// can change) goes through the effective roster ([`Self::resolve_region`]
|
||
/// and the `/cluster/members` surface).
|
||
fn region_name_of(&self, id: RegionId) -> &str {
|
||
self.id_to_name.get(&id).map_or("unknown", String::as_str)
|
||
}
|
||
|
||
/// This replica's data-shard for admin surfaces (m11p6): `Some(group_shard)`
|
||
/// when the node co-hosts several groups, `None` for the legacy single group.
|
||
/// The seam every per-shard admin surface (the `?shard=` selector, the
|
||
/// `NotLeader` body, the intra-group forward suffix) reads, so S=1 stays
|
||
/// byte-for-byte by construction (one place returns `None`).
|
||
const fn admin_shard(&self) -> Option<ShardId> {
|
||
if self.multi_shard {
|
||
Some(self.group_shard)
|
||
} else {
|
||
None
|
||
}
|
||
}
|
||
|
||
/// An intra-group admin path with this replica's `?shard=` selector appended
|
||
/// when the node co-hosts several groups (m11p6). A forward/broadcast to a
|
||
/// sibling NODE must target the SAME group — the receiver's handler resolves
|
||
/// `replica_for(?shard)`. Returns `base` verbatim for the legacy single group
|
||
/// (the receiver's `replica_for(None)` resolves its sole group), keeping the
|
||
/// S=1 forward wire byte-for-byte. The query rides through [`peer_url`], which
|
||
/// appends the path verbatim.
|
||
fn admin_path(&self, base: &str) -> String {
|
||
self.admin_shard().map_or_else(
|
||
|| base.to_string(),
|
||
|shard| format!("{base}?shard={}", shard.0),
|
||
)
|
||
}
|
||
|
||
/// This node's current leadership view (`None` during an election).
|
||
#[must_use]
|
||
fn current_leader(&self) -> Option<RegionId> {
|
||
*read_recovered(&self.leader, "leader")
|
||
}
|
||
|
||
/// True iff this node believes it leads.
|
||
#[must_use]
|
||
fn is_leader(&self) -> bool {
|
||
self.current_leader() == Some(self.region)
|
||
}
|
||
|
||
/// The HTTP address of the current leader, for a `NotLeader` body. Reads
|
||
/// through the effective roster (§3).
|
||
fn leader_http(&self) -> Option<String> {
|
||
match self.current_leader() {
|
||
Some(leader) if leader != self.region => self.membership.peer_http(leader),
|
||
_ => None,
|
||
}
|
||
}
|
||
|
||
/// This node's current election term (0 = the topology era / driver not
|
||
/// started).
|
||
#[must_use]
|
||
pub(crate) fn election_term(&self) -> u64 {
|
||
self.election_runtime
|
||
.get()
|
||
.map_or(0, |rt| rt.current_term())
|
||
}
|
||
|
||
/// A typed `NotLeader` error naming the leader (or the in-progress
|
||
/// election), its HTTP address, and this node's term so the forwarder
|
||
/// can tell a stale answer from a fresh one (m11p4).
|
||
fn not_leader(&self) -> ServerError {
|
||
ServerError::NotLeader {
|
||
leader: self.current_leader().map_or_else(
|
||
|| "none (election in progress)".to_string(),
|
||
|l| self.region_name_of(l).to_string(),
|
||
),
|
||
http_addr: self.leader_http(),
|
||
term: self.election_term(),
|
||
// m11p6: name the group when this node co-hosts several (S>1). `None`
|
||
// for the legacy single group keeps the S=1 body byte-for-byte.
|
||
shard: self.admin_shard().map(|s| s.0),
|
||
}
|
||
}
|
||
|
||
/// Every PEER region as `(name, http_addr)`, in id order, for broadcast /
|
||
/// promote fan-out. Excludes this node's own region.
|
||
///
|
||
/// Reads through the effective roster (m11p5 §3): in era 0 the view is
|
||
/// byte-for-byte the topology positional tables (proven by
|
||
/// `membership::tests::era0_view_matches_topology_tables`); in the membership
|
||
/// era it reflects every applied conf-change, so a join/remove is visible
|
||
/// here with no separate table to update.
|
||
fn peers_named(&self) -> Vec<(String, String)> {
|
||
self.membership.peers_named()
|
||
}
|
||
|
||
/// Mint a fresh per-node internal token (m11p7) naming THIS node. `None` when
|
||
/// no cluster key is configured (token dormant; marker stays hint-only).
|
||
fn mint_node_token(&self) -> Option<String> {
|
||
self.creds.mint_node_token(&self.region_name)
|
||
}
|
||
|
||
/// The per-node internal token (m11p7) as a forward passthrough header entry,
|
||
/// minted fresh from this node's cluster key + region name. Attached whenever
|
||
/// this replica sets the `x-tidal-internal` marker (forward to leader,
|
||
/// item/embedding broadcast, promote fan-out) so the receiving sibling's
|
||
/// marker guard sees a VERIFIED node. Empty when no cluster key is configured
|
||
/// (the marker keeps hint-only behavior — backward compatible).
|
||
fn node_token_passthrough(&self) -> Vec<(&'static str, String)> {
|
||
self.mint_node_token()
|
||
.map(|t| vec![(crate::cluster::security::NODE_TOKEN_HEADER, t)])
|
||
.unwrap_or_default()
|
||
}
|
||
|
||
/// Forward headers for a RELAYED operator hop (m11p7): the node token (so the
|
||
/// receiving sibling honors the marker) plus the [`RELAY_MARKER`] that tells
|
||
/// the receiver to run the full protocol but NOT re-audit the verb — the
|
||
/// operator's entry node already recorded it once. Used by `/cluster/promote`
|
||
/// when a node forwards an operator promote to the target/leader that performs
|
||
/// the fenced transfer.
|
||
///
|
||
/// [`RELAY_MARKER`]: crate::cluster::forward::RELAY_MARKER
|
||
fn relay_passthrough(&self) -> Vec<(&'static str, String)> {
|
||
let mut headers = self.node_token_passthrough();
|
||
headers.push((
|
||
super::forward::RELAY_MARKER,
|
||
super::forward::RELAY_MARKER_VALUE.to_string(),
|
||
));
|
||
headers
|
||
}
|
||
|
||
/// Every region (self + peers) as `(RegionId, name, http_addr_or_none)`, in
|
||
/// id order, for status aggregation. Self's `http_addr` is `None` (served
|
||
/// in-process). A peer with no declared `http_addr` is `None` too
|
||
/// (unreachable). Reads through the effective roster (§3).
|
||
fn all_regions_for_status(&self) -> Vec<(RegionId, String, Option<String>)> {
|
||
self.membership.all_regions_for_status()
|
||
}
|
||
|
||
/// The current leader's HTTP base address, or `None` when THIS node leads,
|
||
/// no leader is known (election in progress), or the leader has no
|
||
/// declared `http_addr`.
|
||
fn leader_http_addr(&self) -> Option<String> {
|
||
match self.current_leader() {
|
||
Some(leader) if leader != self.region => self.membership.peer_http(leader),
|
||
_ => None,
|
||
}
|
||
}
|
||
|
||
/// The leader's region name (for forwarding/error bodies).
|
||
fn leader_name(&self) -> String {
|
||
self.current_leader().map_or_else(
|
||
|| "none (election in progress)".to_string(),
|
||
|l| self.region_name_of(l).to_string(),
|
||
)
|
||
}
|
||
|
||
/// Heal a peer (m11p2): clear the partition, resume the ship queue past
|
||
/// everything the follower reports applied, and nudge the follower to
|
||
/// PULL any history that has rotated out of the leader's ship tail via
|
||
/// its `StreamSegments` catch-up stream.
|
||
///
|
||
/// There is no O(items) HTTP re-broadcast anymore: items and embeddings
|
||
/// ride the same WAL stream as signals, so log catch-up IS the full heal.
|
||
/// The verb is idempotent — re-running it against a converged peer
|
||
/// resumes an already-resumed queue and requests a stream that opens and
|
||
/// closes empty.
|
||
///
|
||
/// `auth` is the healing operator's verbatim `Authorization` header,
|
||
/// threaded onto the catch-up nudge so the inter-node POST carries the
|
||
/// same credential the operator presented (the bearer middleware on the
|
||
/// peer runs first; an unauthenticated nudge would 401 — the exact
|
||
/// side-POST bug class m11p2 exists to eliminate).
|
||
///
|
||
/// Runs ON the write pool (blocking HTTP fetch + blocking nudge POST).
|
||
fn heal_peer_with_reported_applied(&self, region_name: &str, auth: Option<&str>) -> Result<()> {
|
||
if !self.is_leader() {
|
||
return Err(self.not_leader());
|
||
}
|
||
let id = self.resolve_region(region_name)?;
|
||
// Hold the admin-verb lock across the whole remove → resume → nudge
|
||
// sequence: a concurrent `partition` must observe the heal as
|
||
// completed (and then re-partition), never interleave inside it.
|
||
let _admin = self
|
||
.admin_op
|
||
.lock()
|
||
.unwrap_or_else(std::sync::PoisonError::into_inner);
|
||
write_recovered(&self.partitioned, "partitioned").remove(&id);
|
||
|
||
// Ask the follower how far it has applied (over HTTP) and resume the
|
||
// ship queue past it: retries of data the follower already holds
|
||
// prune, and dispatch continues from applied + 1 (or skips to the
|
||
// tail floor, leaving a gap the follower's pull closes). A failed
|
||
// fetch degrades to a plain resume — the ack piggyback re-learns the
|
||
// follower's applied on the first successful ship.
|
||
let remote = self.fetch_remote_applied(id);
|
||
let my_term = self.election_term();
|
||
let reported_applied = match remote {
|
||
// The fold into the acked frontier / commit index is term-gated
|
||
// (m11p4, design-review C13): a follower still on another term
|
||
// resumes dispatch but contributes no quorum mark until it joins
|
||
// this leadership (its next ReportApplied after joining does).
|
||
Some((applied, remote_term)) if remote_term == my_term => {
|
||
self.ship_queue.resume_from(shard_of_region(id), applied);
|
||
Some(applied)
|
||
}
|
||
Some((applied, remote_term)) => {
|
||
tracing::warn!(
|
||
healed = region_name,
|
||
remote_term,
|
||
my_term,
|
||
"heal: follower is on another term; resuming dispatch without folding its mark (it re-reports after joining this term)"
|
||
);
|
||
self.ship_queue.resume(shard_of_region(id));
|
||
Some(applied)
|
||
}
|
||
None => {
|
||
self.ship_queue.resume(shard_of_region(id));
|
||
None
|
||
}
|
||
};
|
||
|
||
// Nudge the follower to pull catch-up NOW rather than waiting for the
|
||
// next live ship to expose its gap (a quiet cluster would otherwise
|
||
// leave it behind until the next write). Best-effort: the follower
|
||
// also self-triggers on the next shipped batch.
|
||
let leader_flushed = self.ship_feed.flushed_seq();
|
||
let behind = reported_applied.is_none_or(|a| a < leader_flushed);
|
||
let mut nudged = false;
|
||
if behind && let Some(http_addr) = self.peer_http.get(&id) {
|
||
// m11p6: the nudge must hit the follower's replica of THIS group when
|
||
// it co-hosts several (`?shard=`); `admin_path` is a no-op for S=1.
|
||
let url = super::forward::peer_url(http_addr, &self.admin_path("/cluster/catchup"));
|
||
let body = serde_json::json!({
|
||
"shard": shard_of_region(self.region).0,
|
||
"from_seqno": reported_applied.unwrap_or(0) + 1,
|
||
});
|
||
let mut req = self
|
||
.blocking_client
|
||
.post(&url)
|
||
.timeout(super::forward::STATUS_PEER_TIMEOUT)
|
||
.header(
|
||
super::forward::INTERNAL_MARKER,
|
||
super::forward::INTERNAL_MARKER_VALUE,
|
||
)
|
||
.json(&body);
|
||
if let Some(auth) = auth {
|
||
req = req.header(axum::http::header::AUTHORIZATION, auth);
|
||
}
|
||
// m11p7: prove sibling identity so the follower's marker guard honors
|
||
// this internal catch-up nudge.
|
||
if let Some(token) = self.mint_node_token() {
|
||
req = req.header(crate::cluster::security::NODE_TOKEN_HEADER, token);
|
||
}
|
||
nudged = req.send().map(|r| r.status().is_success()).unwrap_or(false);
|
||
if !nudged {
|
||
tracing::warn!(
|
||
healed = region_name,
|
||
"heal: catch-up nudge did not land; the follower will \
|
||
self-trigger on the next shipped batch"
|
||
);
|
||
}
|
||
}
|
||
|
||
tracing::info!(
|
||
region = %self.region_name,
|
||
healed = region_name,
|
||
reported_applied,
|
||
leader_flushed,
|
||
nudged,
|
||
"heal: ship queue resumed; follower pulls any rotated-out history \
|
||
via the catch-up stream"
|
||
);
|
||
Ok(())
|
||
}
|
||
|
||
/// Blocking fetch of a peer's `/cluster/status/local`: its
|
||
/// `applied_events` and election `term` (0 for a pre-m11p4 peer).
|
||
/// `None` on any transport/parse failure (caller falls back to full redeliver).
|
||
fn fetch_remote_applied(&self, peer: RegionId) -> Option<(u64, u64)> {
|
||
let http_addr = self.peer_http.get(&peer)?;
|
||
let url = super::forward::peer_url(http_addr, "/cluster/status/local");
|
||
let resp = self
|
||
.blocking_client
|
||
.get(&url)
|
||
.timeout(super::forward::STATUS_PEER_TIMEOUT)
|
||
.send()
|
||
.ok()?;
|
||
if !resp.status().is_success() {
|
||
return None;
|
||
}
|
||
let body: serde_json::Value = resp.json().ok()?;
|
||
let applied = body
|
||
.get("applied_events")
|
||
.and_then(serde_json::Value::as_u64)?;
|
||
let term = body
|
||
.get("term")
|
||
.and_then(serde_json::Value::as_u64)
|
||
.unwrap_or(0);
|
||
Some((applied, term))
|
||
}
|
||
|
||
// ── Write path ──────────────────────────────────────────────────────────
|
||
|
||
/// Resolve a write's acknowledgment mode: the `x-tidal-ack` request
|
||
/// header when present (the caller's explicit choice), else the
|
||
/// deployment default (`replication.ack`).
|
||
fn ack_mode_for(&self, headers: &HeaderMap) -> Result<AckMode> {
|
||
let Some(value) = headers.get(forward::ACK_HEADER) else {
|
||
return Ok(self.ack_default);
|
||
};
|
||
let value = value
|
||
.to_str()
|
||
.map_err(|_| ServerError::BadRequest("x-tidal-ack must be ASCII".into()))?;
|
||
AckMode::parse(value).ok_or_else(|| {
|
||
ServerError::BadRequest(format!(
|
||
"x-tidal-ack must be \"leader\" or \"quorum\", got {value:?}"
|
||
))
|
||
})
|
||
}
|
||
|
||
/// The retryable quorum-timeout 503 for `seq`, built from the index's
|
||
/// current marks (names the laggard regions).
|
||
fn quorum_timeout_error(&self, seq: u64) -> ServerError {
|
||
self.cluster_metrics.incr_quorum_timeouts();
|
||
let marks = self.commit.peer_marks();
|
||
let confirmed = marks.iter().filter(|&&(_, mark)| mark >= seq).count();
|
||
ServerError::QuorumTimeout {
|
||
seq,
|
||
needed: self.commit.needed_peers(),
|
||
confirmed,
|
||
committed: self.commit.committed(),
|
||
laggards: marks
|
||
.iter()
|
||
.filter(|&&(_, mark)| mark < seq)
|
||
.map(|&(shard, _)| self.region_name_of(RegionId(shard.0)).to_string())
|
||
.collect(),
|
||
}
|
||
}
|
||
|
||
/// Stage a signal write: submit it to the leader WAL (microseconds — one
|
||
/// bounded-channel send; the WAL writer assigns the stream seqno at
|
||
/// flush, m11p2).
|
||
///
|
||
/// The expensive halves are elsewhere by design (m11p1): the group-commit
|
||
/// fsync wait + in-memory fold happen in
|
||
/// [`complete_signal_write`](Self::complete_signal_write) (off the write
|
||
/// pool, so concurrent writers coalesce into shared fsyncs), and follower
|
||
/// shipping happens on the ship queue's dedicated sender threads, fed by
|
||
/// the WAL's flush feed (never on the request path).
|
||
///
|
||
/// Runs ON the write pool (bounded queue → 429 admission control).
|
||
/// Returns [`ServerError::NotLeader`] when this node does not lead.
|
||
fn stage_signal_local(
|
||
&self,
|
||
signal: &str,
|
||
entity: EntityId,
|
||
weight: f64,
|
||
) -> Result<StagedSignal> {
|
||
if !self.is_leader() {
|
||
return Err(self.not_leader());
|
||
}
|
||
let db = self.db()?;
|
||
db.signal_staged(signal, entity, weight, Timestamp::now())
|
||
.map_err(ServerError::Tidal)
|
||
}
|
||
|
||
/// Complete a staged signal write: block until the leader group-commit
|
||
/// fsync covers it (the 204's durability promise) and fold it into the
|
||
/// in-memory aggregate. Shipping needs no wake — the WAL writer's flush
|
||
/// feed notifies the ship queue's senders directly.
|
||
///
|
||
/// Runs on the blocking pool.
|
||
///
|
||
/// Deliberately does NOT re-check leadership: the staged event is already
|
||
/// submitted to THIS node's WAL, so a promote landing mid-write must not
|
||
/// abort completion — the leader's own aggregate would silently miss an
|
||
/// event its durable log carries. The dual-stream window a mid-write
|
||
/// demotion opens is the same one the old eager path had; term fencing
|
||
/// (m11p4) is what closes it.
|
||
/// Returns the write's assigned WAL seqno (`0` = dedup-suppressed; an
|
||
/// identical record is already durable, so there is nothing new to gate
|
||
/// quorum on).
|
||
fn complete_signal_write(&self, staged: StagedSignal) -> Result<u64> {
|
||
let db = self.db()?;
|
||
let seq = staged.wait(&db).map_err(ServerError::Tidal)?;
|
||
self.set_frontier_gauges();
|
||
Ok(seq)
|
||
}
|
||
|
||
/// Publish the stream-frontier gauges: `relay_last_seq` = the WAL flushed
|
||
/// frontier, `relay_durable_seq` = the quorum commit index (m11p3 — the
|
||
/// gap between them is the cluster's quorum lag). With zero peers the
|
||
/// leader alone is the majority, so the gauges coincide.
|
||
///
|
||
/// Publishes BOTH halves or neither. `CommitIndex::committed` returns 0
|
||
/// as a sentinel for "no quorum information yet", not as a seqno, so a
|
||
/// replicated node that has not yet satisfied a quorum wait in this term
|
||
/// has no commit index to report. Publishing the sentinel against a live
|
||
/// flushed frontier claimed the entire relay log as lag: on the fleet's
|
||
/// three-voter cluster that read as 13.5M events behind and held
|
||
/// `TidalDBClusterQuorumLag` critical while every region sat at lag 0.
|
||
pub(crate) fn set_frontier_gauges(&self) {
|
||
let flushed = self.ship_feed.flushed_seq();
|
||
let commit = if self.commit.needed_peers() == 0 {
|
||
flushed
|
||
} else {
|
||
match self.commit.committed() {
|
||
0 => return,
|
||
commit => commit,
|
||
}
|
||
};
|
||
self.cluster_metrics.set_relay_frontiers(flushed, commit);
|
||
}
|
||
|
||
/// Apply an item write on the LEADER: the engine journals it as a kind-1
|
||
/// WAL record first (the one replicated log — this is what replicates it
|
||
/// and serves catch-up), then upserts storage (m11p2).
|
||
/// Returns the kind-1 record's WAL seqno (`None` only when blob
|
||
/// journaling is off, which cluster mode never is).
|
||
fn apply_item_local(
|
||
db: &TidalDb,
|
||
entity: EntityId,
|
||
metadata: &HashMap<String, String>,
|
||
) -> Result<Option<u64>> {
|
||
db.write_item_with_metadata(entity, metadata)
|
||
.map_err(ServerError::Tidal)
|
||
}
|
||
|
||
/// Apply an embedding write on the LEADER (kind-2 WAL record first; see
|
||
/// [`apply_item_local`]).
|
||
fn apply_embedding_local(
|
||
db: &TidalDb,
|
||
entity: EntityId,
|
||
values: &[f32],
|
||
) -> Result<Option<u64>> {
|
||
db.write_item_embedding(entity, values)
|
||
.map_err(ServerError::Tidal)
|
||
}
|
||
|
||
/// Record a hide hard-negative for `(user, item)` on THIS node.
|
||
///
|
||
/// Node-local by design; convergence across nodes is the CRDT reconcile path
|
||
/// (task 03). Routes through the engine's user-scoped signal path so the
|
||
/// durable `Tag::HardNeg` row `take_crdt_snapshot` reads from is written.
|
||
fn record_hardneg_local(&self, user_id: u64, item_id: u64) -> Result<()> {
|
||
let db = self.db()?;
|
||
db.signal_with_context(
|
||
"hide",
|
||
EntityId::new(item_id),
|
||
1.0,
|
||
Timestamp::now(),
|
||
Some(user_id),
|
||
None,
|
||
)
|
||
.map_err(ServerError::Tidal)
|
||
}
|
||
|
||
// ── Read path ──────────────────────────────────────────────────────────
|
||
|
||
/// Resolve the read region: default is the LOCAL region; a `?region=` that
|
||
/// names a DIFFERENT region returns [`ServerError::NotLocal`] (task 03
|
||
/// forwards). An unknown name is a 400.
|
||
fn read_region(&self, region_name: Option<&str>) -> Result<()> {
|
||
match region_name {
|
||
None => Ok(()),
|
||
Some(name) => {
|
||
let id = self.resolve_region(name)?;
|
||
if id == self.region {
|
||
Ok(())
|
||
} else {
|
||
Err(ServerError::NotLocal {
|
||
region: name.to_string(),
|
||
})
|
||
}
|
||
}
|
||
}
|
||
}
|
||
|
||
// ── Management ───────────────────────────────────────────────────────────
|
||
|
||
/// Apply a leadership change on THIS node (m11p2). Returns the stream
|
||
/// baseline when this node is the promotion target (it starts leading).
|
||
///
|
||
/// - **This node is the new leader:** the outbound stream starts at its
|
||
/// current WAL flushed frontier — everything below is pre-stream
|
||
/// history (replicated applies from the previous leader) that must
|
||
/// never push to peers or serve from the catch-up stream. The baseline
|
||
/// persists to disk so a restart cannot forget it, then the ship queue
|
||
/// activates from it.
|
||
/// - **Another node is the new leader:** deactivate the ship queue (we
|
||
/// may have been leading) and, when the fan-out carries the new
|
||
/// stream's baseline, jump this node's applied frontier for the new
|
||
/// leader's shard to it (seqnos at or below the baseline are not data).
|
||
///
|
||
/// 400 on an unknown region.
|
||
fn promote_local(&self, region_name: &str, baseline: Option<u64>) -> Result<Option<u64>> {
|
||
// The legacy fan-out promote is a TOPOLOGY-ERA verb (m11p4): once
|
||
// this node has JOINED an elected leadership, leadership moves only
|
||
// through elections / fenced transfers — a topology-era fan-out must
|
||
// not depose an elected leader. The gate is the JOINED term, not the
|
||
// node's own durable term: a failed candidacy inflates the latter
|
||
// without any elected leadership existing (the isolated
|
||
// operator-override case).
|
||
let joined = self.election_runtime.get().map_or(0, |rt| rt.joined_term());
|
||
if joined >= 1 {
|
||
return Err(ServerError::Cluster(format!(
|
||
"legacy promote is fenced: this node has joined an elected \
|
||
leadership (term {joined}); use /cluster/promote without the \
|
||
internal marker (the fenced transfer path)"
|
||
)));
|
||
}
|
||
let id = self.resolve_region(region_name)?;
|
||
*write_recovered(&self.leader, "leader") = Some(id);
|
||
// Keep the election machine's view in step so its term-0 heartbeats
|
||
// do not fight the verb.
|
||
if let Some(rt) = self.election_runtime.get() {
|
||
rt.force_term0_view(id);
|
||
}
|
||
|
||
if id == self.region {
|
||
let baseline = self.ship_feed.flushed_seq();
|
||
persist_stream_baseline(&self.data_dir, baseline);
|
||
self.stream_baseline.store(baseline, Ordering::Release);
|
||
// Legacy fan-out promote stays in the topology era (term 0); the
|
||
// election driver's transitions stamp real terms (m11p4).
|
||
self.ship_queue.activate_from(baseline, 0);
|
||
tracing::info!(
|
||
region = %self.region_name,
|
||
baseline,
|
||
"promote: this node now leads; ship queue activated"
|
||
);
|
||
return Ok(Some(baseline));
|
||
}
|
||
|
||
self.ship_queue.deactivate();
|
||
if let Some(baseline) = baseline
|
||
&& baseline > 0
|
||
{
|
||
let db = self.db()?;
|
||
let new_leader_shard = shard_of_region(id);
|
||
db.replication_state().advance(new_leader_shard, baseline);
|
||
}
|
||
tracing::info!(
|
||
region = %self.region_name,
|
||
new_leader = region_name,
|
||
baseline,
|
||
"promote: leadership view updated"
|
||
);
|
||
Ok(None)
|
||
}
|
||
|
||
/// Insert a peer into the leader-side ship-skip set. Meaningful only while
|
||
/// this node leads; a non-leader returns [`ServerError::NotLeader`].
|
||
///
|
||
/// Runs ON the write pool: the admin-verb lock serializes it against an
|
||
/// in-flight heal (which holds the lock for its whole backfill), and a
|
||
/// blocked waiter must occupy a pool worker, never a reactor thread.
|
||
fn partition_peer(&self, region_name: &str) -> Result<()> {
|
||
if !self.is_leader() {
|
||
return Err(self.not_leader());
|
||
}
|
||
let id = self.resolve_region(region_name)?;
|
||
let _admin = self
|
||
.admin_op
|
||
.lock()
|
||
.unwrap_or_else(std::sync::PoisonError::into_inner);
|
||
write_recovered(&self.partitioned, "partitioned").insert(id);
|
||
// Hold the ship queue's dispatch to this peer; new batches stop
|
||
// (in-flight sends complete — same race window as the old eager path).
|
||
self.ship_queue.pause(shard_of_region(id));
|
||
tracing::info!(region = %self.region_name, partitioned = region_name, "partition: peer added to ship-skip set");
|
||
Ok(())
|
||
}
|
||
|
||
/// Take this node's CRDT snapshot off the engine (blocking).
|
||
fn take_snapshot(&self) -> Result<tidaldb::replication::reconcile::StateSnapshot> {
|
||
self.db()?.take_crdt_snapshot().map_err(ServerError::Tidal)
|
||
}
|
||
|
||
/// Apply a remote CRDT snapshot via the engine's reconcile path (blocking).
|
||
/// Returns the number of operations applied.
|
||
fn reconcile_remote(
|
||
&self,
|
||
remote: &tidaldb::replication::reconcile::StateSnapshot,
|
||
) -> Result<usize> {
|
||
self.db()?
|
||
.reconcile_with(remote)
|
||
.map_err(ServerError::Tidal)
|
||
}
|
||
|
||
// ── Election integration (m11p4) ────────────────────────────────────────
|
||
|
||
/// This node's region id (for the election driver).
|
||
pub(crate) const fn self_region(&self) -> RegionId {
|
||
self.region
|
||
}
|
||
|
||
/// The term the quorum commit index is currently activated with (m11p4):
|
||
/// the reference for the frontier-report gate. Equals the machine's term
|
||
/// whenever an ELECTED leadership activated the queue, and 0 for the
|
||
/// topology era / legacy activations — using the MACHINE term here
|
||
/// instead would drop legitimate reports whenever a node's own failed
|
||
/// candidacy inflated its durable term past its activation.
|
||
pub(crate) fn commit_active_term(&self) -> u64 {
|
||
self.commit.active_term()
|
||
}
|
||
|
||
/// The effective roster (m11p5 §3): the single source of roster truth.
|
||
pub(crate) const fn membership(&self) -> &Arc<super::membership::MembershipView> {
|
||
&self.membership
|
||
}
|
||
|
||
/// The applied membership conf version (0 = the topology era / no kind-4
|
||
/// record). Surfaced in `/cluster/status/local`.
|
||
#[must_use]
|
||
pub(crate) fn membership_version(&self) -> u64 {
|
||
self.membership.version()
|
||
}
|
||
|
||
/// This node's role string in the effective roster (m11p5 §3): `"voter"`,
|
||
/// `"learner"`, or `"removed"`. Era-0 nodes are always `"voter"`.
|
||
#[must_use]
|
||
pub(crate) fn role_in_roster(&self) -> &'static str {
|
||
use tidaldb::wal::format::MemberRole;
|
||
match self.membership.self_role() {
|
||
Some(MemberRole::Voter) | None => "voter",
|
||
Some(MemberRole::Learner) => "learner",
|
||
Some(MemberRole::Removed) => "removed",
|
||
}
|
||
}
|
||
|
||
/// The one fenced membership-apply path (m11p5 §3.3): swap the effective
|
||
/// roster, then drive the four other surfaces (`PeerPool`, `ShipQueue`,
|
||
/// `CommitIndex`, `ElectionState`) from the diff the view computed — under
|
||
/// the apply lock the view holds for the swap, so the five copies can never
|
||
/// disagree about the roster.
|
||
///
|
||
/// Leader-side this is called synchronously at append (with the record's WAL
|
||
/// seq, so a removal arms the delivery grace); follower-side the election
|
||
/// driver's tick compares the engine cell's version to the applied view
|
||
/// version and calls this on change (no seq, removals retire immediately —
|
||
/// a follower does not ship to the removed peer). It is idempotent-by-
|
||
/// version: a replayed record yields a no-op plan and touches nothing.
|
||
///
|
||
/// `record_seq` is `Some` only on the leader's synchronous-append path. When
|
||
/// `Some` AND this node leads, a removed peer's ship-cell + gRPC retirement
|
||
/// is DEFERRED (m11p5 §3.3 removal-delivery grace): the leader must keep
|
||
/// shipping the `Removed` record to the removed peer until the peer learns
|
||
/// of its own removal through the log (its applied mark covers `record_seq`)
|
||
/// or a bounded give-up elapses — otherwise the removed node never receives
|
||
/// the record, never stops campaigning, and zombie-serves as Ready.
|
||
pub(crate) fn apply_membership(
|
||
&self,
|
||
record: &tidaldb::wal::format::MembershipRecord,
|
||
record_seq: Option<u64>,
|
||
) {
|
||
let plan = self.membership.apply_record(record);
|
||
|
||
// PeerPool + ShipQueue: add per-peer channels and sender threads.
|
||
for (shard, addr) in &plan.added_peers {
|
||
if let Err(e) = self.transport.add_peer(*shard, addr) {
|
||
tracing::warn!(shard = shard.0, %addr, error = %e, "membership apply: add_peer failed");
|
||
} else {
|
||
self.ship_queue.add_peer(*shard);
|
||
}
|
||
}
|
||
// Remove per-peer channels and sender threads. On the LEADER's append
|
||
// path (`record_seq.is_some()`), DEFER the removed peer's ship-cell +
|
||
// gRPC retirement behind the delivery grace (§3.3) — the gRPC entry is
|
||
// exactly what the ship cell needs to deliver the `Removed` record to
|
||
// the removed node. Every other path (follower apply, boot reconcile,
|
||
// activation re-append) retires immediately: a follower never ships to
|
||
// the removed peer, so holding its cell open would only leak threads.
|
||
let leader_defer = record_seq.is_some() && self.is_leader();
|
||
for shard in &plan.removed_peers {
|
||
if leader_defer {
|
||
self.arm_deferred_retire(*shard, record_seq.unwrap_or(0));
|
||
} else {
|
||
self.ship_queue.remove_peer(*shard);
|
||
self.transport.remove_peer(*shard);
|
||
}
|
||
}
|
||
|
||
// CommitIndex: resize voter/learner marks + `needed` under its own lock;
|
||
// in-flight `ack=quorum` waiters are re-evaluated, not failed (§3.3).
|
||
self.commit
|
||
.reconfigure(&plan.commit_voters, &plan.commit_learners);
|
||
|
||
// ElectionState: recompute majority over the new voter set + the campaign
|
||
// gate (self voter-ness), via the driver's reconfigure entry point. The
|
||
// heartbeat fan-out reaches voters AND learners (a learner joins the term
|
||
// and reports — the auto-promotion input); only voters count to quorum.
|
||
if let Some(rt) = self.election_runtime.get() {
|
||
rt.reconfigure_voters(
|
||
plan.election_voters.clone(),
|
||
&plan.election_learners,
|
||
plan.self_is_voter,
|
||
);
|
||
}
|
||
|
||
// §3.6: rewrite the durable cache from the just-applied record so the
|
||
// pre-open seed-join boot loop (a future restart) reads the current
|
||
// roster, never a stale one. The cell is the authoritative source (it
|
||
// wins at open); this cache is a best-effort discovery aid, so a persist
|
||
// failure is logged but never fatal to the apply.
|
||
self.rewrite_membership_cache(record.version, record.term, &record.members);
|
||
|
||
tracing::info!(
|
||
version = record.version,
|
||
term = record.term,
|
||
voters = plan.voter_count,
|
||
added = plan.added_peers.len(),
|
||
removed = plan.removed_peers.len(),
|
||
"membership applied (one fenced path)"
|
||
);
|
||
}
|
||
|
||
/// Arm the removal-delivery grace for a removed peer (m11p5 §3.3): the
|
||
/// leader holds the peer's ship cell + gRPC entry open until the peer has
|
||
/// learned of its removal through the log (its applied mark reaches
|
||
/// `record_seq`) or the bounded deadline forces the retirement. Idempotent
|
||
/// per shard — a re-armed shard (a replayed apply) refreshes nothing if it
|
||
/// is already pending.
|
||
fn arm_deferred_retire(&self, shard: ShardId, record_seq: u64) {
|
||
let grace = std::env::var("TIDAL_REMOVE_DELIVERY_GRACE_MS")
|
||
.ok()
|
||
.and_then(|v| v.trim().parse::<u64>().ok())
|
||
.unwrap_or(REMOVE_DELIVERY_GRACE_DEFAULT_MS);
|
||
let deadline = Instant::now() + Duration::from_millis(grace);
|
||
let name = self.region_name_of(RegionId(shard.0)).to_string();
|
||
let mut pending = self
|
||
.deferred_retires
|
||
.lock()
|
||
.unwrap_or_else(std::sync::PoisonError::into_inner);
|
||
if pending.iter().any(|d| d.shard == shard) {
|
||
return; // already pending — keep the original deadline.
|
||
}
|
||
tracing::info!(
|
||
shard = shard.0,
|
||
region = %name,
|
||
record_seq,
|
||
grace_ms = grace,
|
||
"removal-delivery grace armed (§3.3): the removed peer's ship cell stays open \
|
||
until it learns of its removal through the log (or the grace elapses) — the \
|
||
removed node must receive the Removed record before its cell retires"
|
||
);
|
||
pending.push(DeferredRetire {
|
||
shard,
|
||
name,
|
||
record_seq,
|
||
deadline,
|
||
});
|
||
}
|
||
|
||
/// Poll the pending removal-delivery graces (m11p5 §3.3), retiring each
|
||
/// removed peer's ship cell + gRPC entry once the peer has DEMONSTRABLY
|
||
/// learned of its removal (its applied mark — read from the transport's
|
||
/// per-peer applied hint, which a `ReportApplied` advances and a
|
||
/// `reconfigure` does NOT clear — covers the `Removed` record's seq) OR the
|
||
/// bounded grace deadline elapses (a dead/unreachable removed peer must not
|
||
/// pin the ship cell forever; that case WARNs + bumps a counter).
|
||
///
|
||
/// Called on the election driver's tick (the leader is the only node with a
|
||
/// pending grace, since only its apply path arms one). A no-op when nothing
|
||
/// is pending. Cheap: a brief lock + an atomic-map read per pending entry.
|
||
pub(crate) fn poll_deferred_retires(&self) {
|
||
// Fast path: avoid the lock churn when there is nothing pending (the
|
||
// common case — the leader arms a grace only on a remove verb).
|
||
if self
|
||
.deferred_retires
|
||
.lock()
|
||
.unwrap_or_else(std::sync::PoisonError::into_inner)
|
||
.is_empty()
|
||
{
|
||
return;
|
||
}
|
||
let now = Instant::now();
|
||
// Collect the entries ready to retire under the lock, then act outside it
|
||
// (`remove_peer` enters the transport runtime — never hold our mutex
|
||
// across that). Retained entries stay pending for the next tick.
|
||
let mut to_retire: Vec<DeferredRetire> = Vec::new();
|
||
{
|
||
let mut pending = self
|
||
.deferred_retires
|
||
.lock()
|
||
.unwrap_or_else(std::sync::PoisonError::into_inner);
|
||
pending.retain(|d| {
|
||
let delivered = self.transport.peer_applied_hint(d.shard) >= d.record_seq;
|
||
let expired = now >= d.deadline;
|
||
if delivered || expired {
|
||
to_retire.push(d.clone());
|
||
false // drop from pending; it retires below.
|
||
} else {
|
||
true // still waiting on delivery within the grace.
|
||
}
|
||
});
|
||
}
|
||
for d in to_retire {
|
||
let delivered = self.transport.peer_applied_hint(d.shard) >= d.record_seq;
|
||
if delivered {
|
||
tracing::info!(
|
||
shard = d.shard.0,
|
||
region = %d.name,
|
||
record_seq = d.record_seq,
|
||
"removal-delivery grace satisfied (§3.3): the removed peer's applied mark \
|
||
covers the Removed record — it learned of its removal through the log; \
|
||
retiring its ship cell + gRPC entry"
|
||
);
|
||
} else {
|
||
// The give-up: a dead/unreachable removed peer. Retire anyway so
|
||
// the cell does not leak, but make it loud + metered.
|
||
self.cluster_metrics.incr_remove_delivery_giveup();
|
||
tracing::warn!(
|
||
shard = d.shard.0,
|
||
region = %d.name,
|
||
record_seq = d.record_seq,
|
||
"removal-delivery grace EXPIRED (§3.3): the removed peer never acked the \
|
||
Removed record within the grace (it is down or unreachable). Retiring its \
|
||
ship cell anyway. If/when it restarts it learns of its removal via the \
|
||
typed removed signal in any voter's heartbeat/vote refusal (runbook §8)"
|
||
);
|
||
}
|
||
self.ship_queue.remove_peer(d.shard);
|
||
self.transport.remove_peer(d.shard);
|
||
}
|
||
}
|
||
|
||
/// Rewrite the durable membership cache (`data_dir/membership`, §3.6) from a
|
||
/// roster. Best-effort: the WAL-recovered cell is the authoritative roster,
|
||
/// so a persist failure is logged (a stale cache only slows the next pre-open
|
||
/// boot loop's discovery, never affects correctness — the cell still wins).
|
||
fn rewrite_membership_cache(
|
||
&self,
|
||
version: u64,
|
||
term: u64,
|
||
members: &[tidaldb::wal::format::MemberEntry],
|
||
) {
|
||
let snapshot = tidaldb::replication::MembershipSnapshot {
|
||
version,
|
||
term,
|
||
members: members.to_vec(),
|
||
};
|
||
if let Err(e) = self.membership_store.persist(&snapshot) {
|
||
tracing::warn!(
|
||
version,
|
||
error = %e,
|
||
"failed to rewrite the durable membership cache (§3.6); the WAL cell remains \
|
||
authoritative — the next pre-open boot loop falls back to the seed list"
|
||
);
|
||
}
|
||
}
|
||
|
||
/// Follower-side membership-apply trigger (m11p5 §3.3): the engine folds
|
||
/// kind-4 records into its `ClusterMembership` cell as they stream in; this
|
||
/// compares the cell's version to the applied view's and drives the one
|
||
/// fenced apply path on change. Called from the election driver's 50ms tick
|
||
/// (a cheap version compare). A no-op on the leader (it applies
|
||
/// synchronously at append, so the cell and view already agree).
|
||
pub(crate) fn maybe_apply_membership_from_cell(&self) {
|
||
let Some(db) = self.db.load_full() else {
|
||
return;
|
||
};
|
||
let Some((version, term, members)) = db.cluster_membership() else {
|
||
return; // still the topology era; nothing to apply.
|
||
};
|
||
if version <= self.membership.version() {
|
||
return; // already applied (or a stale re-fold).
|
||
}
|
||
let record = tidaldb::wal::format::MembershipRecord {
|
||
version,
|
||
term,
|
||
members,
|
||
};
|
||
// Follower / leader-via-cell path: no record seq, removals retire
|
||
// immediately (a follower does not ship the `Removed` record to the
|
||
// removed peer — only the leader's synchronous-append path does, and it
|
||
// arms the delivery grace there).
|
||
self.apply_membership(&record, None);
|
||
}
|
||
|
||
/// Leader-side conf-change gates (m11p5 §3.1, §3.3): refuse a join/remove
|
||
/// unless this node leads, the era's first conf-change is capability-clean
|
||
/// (every voter reports kind-4), and the prior conf-change record is
|
||
/// same-term quorum-committed (one-at-a-time). Returns the leader's term on
|
||
/// success (the record's term stamp).
|
||
fn conf_change_gate(&self) -> Result<u64, super::membership::ConfChangeRefusal> {
|
||
use super::membership::ConfChangeRefusal;
|
||
if !self.is_leader() {
|
||
return Err(ConfChangeRefusal::NotLeader {
|
||
leader: self
|
||
.current_leader()
|
||
.map(|l| self.region_name_of(l).to_string()),
|
||
});
|
||
}
|
||
let term = self.commit_active_term();
|
||
let roster = self.membership.roster();
|
||
// §3.1 capability gate: every current voter must report kind-4. The
|
||
// leader's own binary is capable by definition.
|
||
if let Err(incapable) = super::membership::capability_gate(&roster, self.region, |rid| {
|
||
self.transport.peer_capabilities(shard_of_region(rid))
|
||
}) {
|
||
return Err(ConfChangeRefusal::IncapableVoters {
|
||
voters: incapable
|
||
.into_iter()
|
||
.map(|r| self.region_name_of(r).to_string())
|
||
.collect(),
|
||
});
|
||
}
|
||
Ok(term)
|
||
}
|
||
|
||
/// Append a kind-4 record and wait (bounded) for its SAME-TERM quorum commit
|
||
/// (§3.2/§3.3 — never the activation-reset `committed()`). The one-at-a-time
|
||
/// gate is exactly this wait: the next conf-change cannot append until the
|
||
/// prior record's seq is `committed_in_term`. Applies the record locally on
|
||
/// success (the leader applies synchronously at append).
|
||
///
|
||
/// Returns the seq the record landed at, or a refusal when the commit does
|
||
/// not land within the quorum timeout (retryable).
|
||
fn append_and_commit_membership(
|
||
&self,
|
||
record: &tidaldb::wal::format::MembershipRecord,
|
||
) -> std::result::Result<u64, super::membership::ConfChangeRefusal> {
|
||
use super::membership::ConfChangeRefusal;
|
||
let Ok(db) = self.db() else {
|
||
return Err(ConfChangeRefusal::NotLeader { leader: None });
|
||
};
|
||
let seq = match db.append_membership_record(record.clone()) {
|
||
Ok(seq) => seq,
|
||
Err(e) => {
|
||
tracing::error!(error = %e, "membership record append failed");
|
||
return Err(ConfChangeRefusal::PriorChangeUncommitted { awaiting_seq: 0 });
|
||
}
|
||
};
|
||
// Leader applies synchronously at append (the cell already advanced in
|
||
// the engine; drive the four other surfaces now). Pass the record's seq
|
||
// so a `Removed` record arms the removal-delivery grace (§3.3) rather
|
||
// than retiring the removed peer's ship cell out from under the very
|
||
// delivery that teaches it of its removal.
|
||
self.apply_membership(record, Some(seq));
|
||
// Bounded wait for same-term quorum commit (the one-at-a-time predicate).
|
||
let deadline = std::time::Instant::now() + self.quorum_timeout;
|
||
loop {
|
||
if self.commit.committed_in_term(seq) {
|
||
return Ok(seq);
|
||
}
|
||
if std::time::Instant::now() >= deadline {
|
||
return Err(ConfChangeRefusal::PriorChangeUncommitted { awaiting_seq: seq });
|
||
}
|
||
// The commit index advances on folds; poll briefly. (A watch bridge
|
||
// exists for the ack path; conf-changes are rare enough that a short
|
||
// poll is the simpler, deadlock-free choice here.)
|
||
std::thread::sleep(std::time::Duration::from_millis(10));
|
||
}
|
||
}
|
||
|
||
/// Leader-side `JoinCluster` handler (m11p5 §3.3): the membership runtime's
|
||
/// answer to a joiner. Idempotent by name (a known member returns its
|
||
/// id/role, appends nothing); else assigns `max(all ids ever) + 1`, appends
|
||
/// a Learner record, waits for same-term commit, and answers with the roster.
|
||
/// A non-leader refuses with a leader hint (the joiner re-targets).
|
||
///
|
||
/// BLOCKING — called via `spawn_blocking` from the gRPC handler.
|
||
fn handle_join(&self, ask: &tidal_net::JoinAsk) -> tidal_net::JoinOutcome {
|
||
use super::membership::{ConfChangeRefusal, JoinPlan};
|
||
let refuse = |reason: String| tidal_net::JoinOutcome {
|
||
accepted: false,
|
||
refusal_reason: reason,
|
||
assigned_id: 0,
|
||
term: self.election_term(),
|
||
leader_region: self
|
||
.current_leader()
|
||
.map(|l| self.region_name_of(l).to_string())
|
||
.unwrap_or_default(),
|
||
leader_grpc_addr: String::new(),
|
||
leader_http_addr: self.leader_http().unwrap_or_default(),
|
||
members: Vec::new(),
|
||
membership_version: self.membership.version(),
|
||
};
|
||
|
||
// §3.3: a non-leader refuses with a leader hint; gates check capability.
|
||
let term = match self.conf_change_gate() {
|
||
Ok(term) => term,
|
||
Err(ConfChangeRefusal::NotLeader { leader }) => {
|
||
let mut out = refuse(
|
||
ConfChangeRefusal::NotLeader {
|
||
leader: leader.clone(),
|
||
}
|
||
.to_string(),
|
||
);
|
||
if let Some(l) = leader {
|
||
out.leader_region = l;
|
||
}
|
||
return out;
|
||
}
|
||
Err(other) => return refuse(other.to_string()),
|
||
};
|
||
|
||
let roster = self.membership.roster();
|
||
// The era only BEGINS via this capability-gated first conf-change, so a
|
||
// record can never reach a pre-p5 follower (§3.1). next_version = the
|
||
// current applied version + 1.
|
||
let next_version = self.membership.version() + 1;
|
||
match super::membership::plan_join(
|
||
&roster,
|
||
&ask.name,
|
||
&ask.grpc_addr,
|
||
&ask.http_addr,
|
||
next_version,
|
||
term,
|
||
) {
|
||
JoinPlan::Existing { id, role } => {
|
||
// Idempotent: append nothing, answer with the current roster.
|
||
tracing::info!(name = %ask.name, id, ?role, "join is idempotent (known member)");
|
||
self.join_success_outcome(id, term)
|
||
}
|
||
JoinPlan::IdSpaceExhausted => {
|
||
refuse("the member id space is exhausted (65 536 ids assigned)".to_string())
|
||
}
|
||
JoinPlan::Append { id, record } => match self.append_and_commit_membership(&record) {
|
||
Ok(_seq) => {
|
||
tracing::info!(name = %ask.name, id, "joiner accepted as a Learner");
|
||
self.join_success_outcome(id, term)
|
||
}
|
||
Err(e) => refuse(e.to_string()),
|
||
},
|
||
}
|
||
}
|
||
|
||
/// Build a successful `JoinOutcome` from the CURRENT (post-apply) roster.
|
||
fn join_success_outcome(&self, assigned_id: u16, term: u64) -> tidal_net::JoinOutcome {
|
||
let roster = self.membership.roster();
|
||
let members = roster
|
||
.members
|
||
.iter()
|
||
.map(|m| tidal_net::MemberInfo {
|
||
id: m.id,
|
||
name: m.name.clone(),
|
||
grpc_addr: m.grpc_addr.clone(),
|
||
http_addr: m.http_addr.clone(),
|
||
role: m.role.to_byte(),
|
||
})
|
||
.collect();
|
||
// The leader is self here (the gate proved it).
|
||
let leader_grpc = roster
|
||
.members
|
||
.iter()
|
||
.find(|m| m.id == self.region.0)
|
||
.map(|m| m.grpc_addr.clone())
|
||
.unwrap_or_default();
|
||
tidal_net::JoinOutcome {
|
||
accepted: true,
|
||
refusal_reason: String::new(),
|
||
assigned_id,
|
||
term,
|
||
leader_region: self.region_name.clone(),
|
||
leader_grpc_addr: leader_grpc,
|
||
leader_http_addr: String::new(),
|
||
members,
|
||
membership_version: roster.version,
|
||
}
|
||
}
|
||
|
||
/// Leader-side member removal (m11p5 §3.3): append a `Removed` kind-4 record
|
||
/// (gated), burning the member's id. Idempotent (an unknown or already-
|
||
/// removed name is `NotPresent`). The fenced apply path retires the peer's
|
||
/// ship cell; the removed node learns via the typed `removed` signal.
|
||
///
|
||
/// BLOCKING — called on the write pool.
|
||
fn handle_remove(&self, name: &str) -> RemoveOutcome {
|
||
let term = match self.conf_change_gate() {
|
||
Ok(term) => term,
|
||
Err(refusal) => return RemoveOutcome::Refused(refusal.to_string()),
|
||
};
|
||
let roster = self.membership.roster();
|
||
let next_version = self.membership.version() + 1;
|
||
let Some(record) = super::membership::plan_remove(&roster, name, next_version, term) else {
|
||
return RemoveOutcome::NotPresent;
|
||
};
|
||
match self.append_and_commit_membership(&record) {
|
||
Ok(_seq) => RemoveOutcome::Removed {
|
||
version: next_version,
|
||
},
|
||
Err(e) => RemoveOutcome::Refused(e.to_string()),
|
||
}
|
||
}
|
||
|
||
/// Auto-promotion: a STANDING LEADER DUTY (m11p5 §3.3) derived solely from
|
||
/// the applied `ClusterMembership`. Evaluated on every leader tick (the
|
||
/// driver re-arms it on activation and on every membership apply by simply
|
||
/// calling it each tick), it promotes a caught-up learner to Voter so a
|
||
/// seed-joined node becomes a quorum member with ZERO operator verbs.
|
||
///
|
||
/// A learner is promotion-eligible once its durable mark (the side-map mark
|
||
/// the commit index tracks, never counted toward quorum) is within
|
||
/// `learner_promote_lag` of the leader's flushed frontier. One promotion at a
|
||
/// time (the same one-at-a-time conf-change gate every change rides): if the
|
||
/// prior conf-change is not yet same-term quorum-committed, this tick's
|
||
/// promotion is held and retried next tick.
|
||
///
|
||
/// Non-blocking-by-design caveat: `append_and_commit_membership` BLOCKS on
|
||
/// the bounded same-term commit wait, so this runs on the WRITE POOL (a
|
||
/// detached job the driver tick submits), never inline on the election
|
||
/// driver thread. A no-op when this node is not the leader, the era has not
|
||
/// begun, or no learner is eligible.
|
||
///
|
||
/// BLOCKING — called on the write pool.
|
||
fn maybe_auto_promote(&self) {
|
||
use tidaldb::wal::format::MemberRole;
|
||
if !self.is_leader() {
|
||
return;
|
||
}
|
||
let roster = self.membership.roster();
|
||
// Pick ONE eligible learner (lowest id for determinism). A learner is
|
||
// eligible when its durable side-map mark is within `learner_promote_lag`
|
||
// of the leader's flushed frontier. The leader's own flushed_seq is the
|
||
// catch-up target; a learner within the threshold is durably holding
|
||
// essentially the whole log.
|
||
let flushed = self.ship_feed.flushed_seq();
|
||
let mut eligible: Option<u16> = None;
|
||
for m in &roster.members {
|
||
if m.role != MemberRole::Learner {
|
||
continue;
|
||
}
|
||
let mark = self.commit.learner_mark(shard_of_region(RegionId(m.id)));
|
||
if flushed.saturating_sub(mark) <= self.learner_promote_lag {
|
||
eligible = Some(m.id);
|
||
break;
|
||
}
|
||
}
|
||
let Some(learner_id) = eligible else {
|
||
return;
|
||
};
|
||
|
||
// Gate the conf-change (one-at-a-time + capability). If the gate refuses
|
||
// (e.g. the prior change is not yet committed), retry next tick.
|
||
let Ok(term) = self.conf_change_gate() else {
|
||
return;
|
||
};
|
||
let next_version = self.membership.version() + 1;
|
||
let Some(record) = super::membership::plan_promote(&roster, learner_id, next_version, term)
|
||
else {
|
||
return; // already a voter / not a learner (raced with another apply).
|
||
};
|
||
match self.append_and_commit_membership(&record) {
|
||
Ok(seq) => {
|
||
tracing::info!(
|
||
learner_id,
|
||
version = next_version,
|
||
seq,
|
||
flushed,
|
||
"auto-promoted a caught-up learner to Voter (§3.3 standing leader duty)"
|
||
);
|
||
}
|
||
Err(e) => {
|
||
// The commit did not land within the quorum timeout — the next
|
||
// tick re-evaluates and retries. Loud so a stuck promotion is
|
||
// diagnosable (the `promotion_pending` story).
|
||
tracing::warn!(learner_id, error = %e, "auto-promotion held; retrying next tick");
|
||
}
|
||
}
|
||
}
|
||
|
||
/// Run the auto-promotion duty on a detached thread (the append + bounded
|
||
/// commit wait BLOCK, so it must not run on the election driver thread).
|
||
/// Called once per leader tick. At most ONE promotion runs at a time (a
|
||
/// 50ms tick can outpace a promotion's commit wait); the `promote_inflight`
|
||
/// flag serializes them, and the duty is a no-op when no learner exists, so
|
||
/// a steady-state leader never spawns a thread.
|
||
pub(crate) fn submit_auto_promote(self: &Arc<Self>) {
|
||
// Only the leader has a duty, only in the membership era.
|
||
if !self.is_leader() || !self.membership.era_begun() {
|
||
return;
|
||
}
|
||
// Steady-state fast path: no learner → nothing to promote, no thread.
|
||
if self.membership.roster().learner_ids().is_empty() {
|
||
return;
|
||
}
|
||
// One promotion at a time: if a prior evaluation is still in flight, let
|
||
// it finish (it re-evaluates on its own; the next tick re-arms).
|
||
if self
|
||
.promote_inflight
|
||
.compare_exchange(false, true, Ordering::AcqRel, Ordering::Acquire)
|
||
.is_err()
|
||
{
|
||
return;
|
||
}
|
||
let node = Arc::clone(self);
|
||
if std::thread::Builder::new()
|
||
.name("tidal-auto-promote".into())
|
||
.spawn(move || {
|
||
node.maybe_auto_promote();
|
||
node.promote_inflight.store(false, Ordering::Release);
|
||
})
|
||
.is_err()
|
||
{
|
||
// The spawn failed — clear the guard so the next tick retries.
|
||
self.promote_inflight.store(false, Ordering::Release);
|
||
}
|
||
}
|
||
|
||
/// Self-driving heal (m11p8 §4): a STANDING LEADER DUTY re-armed every tick
|
||
/// at a coarse cadence (~`SELF_HEAL_TICKS` × the 50 ms election tick). It
|
||
/// closes incident §1.4-3 — "the breaker eats the first heal; re-issue until
|
||
/// lag 0" — by driving convergence itself instead of the operator.
|
||
///
|
||
/// Each pass refreshes the per-peer breaker gauge, then for every peer that
|
||
/// is (a) NOT operator-partitioned, (b) has a non-closed ship breaker
|
||
/// (replication impaired), and (c) trails the leader's flushed frontier by
|
||
/// more than the convergence threshold, re-arms the backlog re-ship from the
|
||
/// peer's durable mark. The moment the breaker half-opens, the leader pushes
|
||
/// the WHOLE gap — no operator verb. A closed-breaker peer that is merely a
|
||
/// little behind self-heals through the normal ship path and is left alone
|
||
/// (no cursor churn). NON-BLOCKING (cursor updates + atomic gauge stores), so
|
||
/// it runs inline on the election tick rather than a detached thread.
|
||
///
|
||
/// Operator-partitioned peers are intentionally ship-skipped (maintenance);
|
||
/// self-heal never auto-undoes a `/cluster/partition`. The manual
|
||
/// `/cluster/heal` verb still exists as an immediate nudge, but is no longer
|
||
/// required for convergence.
|
||
pub(crate) fn tick_self_heal(&self) {
|
||
// Coarse cadence: act ~every SELF_HEAL_TICKS election ticks.
|
||
if !self
|
||
.heal_tick
|
||
.fetch_add(1, Ordering::Relaxed)
|
||
.is_multiple_of(SELF_HEAL_TICKS)
|
||
{
|
||
return;
|
||
}
|
||
if !self.is_leader() {
|
||
// Off the leader there is no ship-driving duty: clear the healing
|
||
// gauge + tracked set so a demoted node stops reporting stale state.
|
||
self.cluster_metrics.set_healing_peers(0);
|
||
self.healing
|
||
.lock()
|
||
.unwrap_or_else(std::sync::PoisonError::into_inner)
|
||
.clear();
|
||
return;
|
||
}
|
||
|
||
let flushed = self.ship_feed.flushed_seq();
|
||
let partitioned = read_recovered(&self.partitioned, "partitioned").clone();
|
||
let marks: std::collections::HashMap<ShardId, u64> =
|
||
self.commit.peer_marks().into_iter().collect();
|
||
|
||
let mut behind_now: std::collections::HashSet<ShardId> = std::collections::HashSet::new();
|
||
for shard in self.ship_queue.peers() {
|
||
// Refresh the breaker gauge for EVERY configured peer (including
|
||
// partitioned ones) so the dashboard is truthful.
|
||
let bstate = self.ship_queue.peer_breaker_state(shard);
|
||
self.cluster_metrics.set_peer_breaker_state(shard, bstate);
|
||
|
||
// In-group shard id == region id, so the partition (ship-skip) set,
|
||
// keyed by region, is checked with the shard's numeric id.
|
||
if partitioned.contains(&RegionId(shard.0)) {
|
||
continue;
|
||
}
|
||
let mark = marks.get(&shard).copied().unwrap_or(0);
|
||
// Stuck = breaker not closed AND trailing past the convergence
|
||
// threshold. A closed breaker means the normal ship path is flowing.
|
||
if bstate != 0 && flushed.saturating_sub(mark) > self.learner_promote_lag {
|
||
behind_now.insert(shard);
|
||
self.cluster_metrics.incr_heal_attempts();
|
||
self.ship_queue.resume_from(shard, mark);
|
||
}
|
||
}
|
||
|
||
// Convergence transitions: peers driven last pass that are no longer
|
||
// behind recovered — count them and update the tracked set.
|
||
let mut prev = self
|
||
.healing
|
||
.lock()
|
||
.unwrap_or_else(std::sync::PoisonError::into_inner);
|
||
for s in prev.iter() {
|
||
if !behind_now.contains(s) {
|
||
self.cluster_metrics.incr_heal_successes();
|
||
}
|
||
}
|
||
let healing_count = behind_now.len() as u64;
|
||
*prev = behind_now;
|
||
drop(prev);
|
||
|
||
if healing_count == 0 {
|
||
// Liveness heartbeat: the loop ran and found everything converged.
|
||
self.cluster_metrics.incr_heal_noops();
|
||
}
|
||
self.cluster_metrics.set_healing_peers(healing_count);
|
||
}
|
||
|
||
/// The shared cluster metrics cell (election gauges live here too).
|
||
pub(crate) const fn cluster_metrics(
|
||
&self,
|
||
) -> &Arc<tidaldb::db::metrics::cluster::ClusterMetrics> {
|
||
&self.cluster_metrics
|
||
}
|
||
|
||
/// The election runtime, when the driver has started.
|
||
pub(crate) fn election_runtime(&self) -> Option<&Arc<super::election_driver::ElectionRuntime>> {
|
||
self.election_runtime.get()
|
||
}
|
||
|
||
/// This node's log position for the vote restriction:
|
||
/// `(wal_tail_term, durable_frontier)`.
|
||
///
|
||
/// The frontier is read in the LAST JOINED TERM'S STREAM numbering — the
|
||
/// only numbering comparable across nodes. A node that LED that term
|
||
/// reads its own flushed frontier (its WAL IS the stream); a follower
|
||
/// reads its durable applied frontier for that leader's shard. The
|
||
/// node's own WAL numbering would lie after a reseed (the baseline jump
|
||
/// makes local seqnos diverge from stream seqnos). Term-0 logs compare
|
||
/// in the topology leader's stream.
|
||
pub(crate) fn election_log_position(&self) -> tidaldb::replication::LogPosition {
|
||
let Some(db) = self.db.load_full() else {
|
||
return tidaldb::replication::LogPosition {
|
||
tail_term: 0,
|
||
frontier: 0,
|
||
};
|
||
};
|
||
let (tail_term, _, marker_region) = db.wal_term_mark();
|
||
let stream_region = if tail_term == 0 {
|
||
self.boot_topology_leader
|
||
} else {
|
||
RegionId(marker_region)
|
||
};
|
||
let frontier = if stream_region == self.region {
|
||
self.ship_feed.flushed_seq()
|
||
} else {
|
||
db.replication_state()
|
||
.applied_seqno(shard_of_region(stream_region))
|
||
.unwrap_or(0)
|
||
};
|
||
tidaldb::replication::LogPosition {
|
||
tail_term,
|
||
frontier,
|
||
}
|
||
}
|
||
|
||
/// The durable leader-acked frontier (m12 election-divergence-fix): the
|
||
/// highest seqno this node acked under `ack=leader` and may hold
|
||
/// un-replicated. See [`Self::leader_acked`].
|
||
pub(crate) fn leader_acked_frontier(&self) -> u64 {
|
||
self.leader_acked.load(Ordering::Acquire)
|
||
}
|
||
|
||
/// Advance the leader-acked frontier for an `ack=leader` write that just
|
||
/// succeeded at `seq`, and PERSIST it durably on advance. The durability is
|
||
/// load-bearing: an `ack=leader` write is acked to the client on journal alone
|
||
/// (no quorum), so if this node is HARD-KILLED before the watermark is durable,
|
||
/// the next boot must still recover the un-replicated suffix to quarantine
|
||
/// (the `mp_quarantined` invariant). The write already fsynced the WAL, so this
|
||
/// is one extra small fsync on the ack=leader path only — `ack=quorum` writes
|
||
/// never call this. Monotonic: a stale/duplicate `seq` neither advances nor
|
||
/// re-persists.
|
||
pub(crate) fn note_leader_acked(&self, seq: u64) {
|
||
let prev = self.leader_acked.fetch_max(seq, Ordering::AcqRel);
|
||
if seq > prev {
|
||
persist_leader_acked(&self.data_dir, seq);
|
||
}
|
||
}
|
||
|
||
/// Reset the leader-acked frontier to 0 and persist it: a CLEAN term join
|
||
/// proves the node is caught up to the new leadership, so it holds nothing
|
||
/// un-replicated the new term does not subsume. Persisting 0 (not just
|
||
/// clearing in memory) is what makes a later graceful restart read 0 instead
|
||
/// of conservatively falling back to the WAL tail.
|
||
pub(crate) fn reset_leader_acked(&self) {
|
||
self.leader_acked.store(0, Ordering::Release);
|
||
persist_leader_acked(&self.data_dir, 0);
|
||
}
|
||
|
||
/// Persist the current leader-acked frontier (graceful-shutdown durability).
|
||
fn persist_leader_acked_now(&self) {
|
||
persist_leader_acked(&self.data_dir, self.leader_acked.load(Ordering::Acquire));
|
||
}
|
||
|
||
/// The transport's election fan-out handle.
|
||
pub(crate) fn election_net(&self) -> tidal_net::ElectionNet {
|
||
self.transport.election_net()
|
||
}
|
||
|
||
/// Build one leader heartbeat: the lease assertion carrying
|
||
/// `(term, leadership, the term's activation baseline)`.
|
||
pub(crate) fn election_heartbeat(&self, term: u64) -> tidal_net::proto::HeartbeatRequest {
|
||
let prev = *self
|
||
.activation_prev
|
||
.lock()
|
||
.unwrap_or_else(std::sync::PoisonError::into_inner);
|
||
tidal_net::proto::HeartbeatRequest {
|
||
shard_id: u32::from(shard_of_region(self.region).0),
|
||
region_id: u32::from(self.region.0),
|
||
term,
|
||
leader_region: u32::from(self.region.0),
|
||
stream_baseline: self.stream_baseline.load(Ordering::Acquire),
|
||
prev_log_term: prev.tail_term,
|
||
prev_log_seq: prev.frontier,
|
||
// m12p5: the leader's LIVE flushed frontier (not the immutable term
|
||
// baseline) so a follower converges readiness from the heartbeat —
|
||
// which flows on an idle cluster — instead of waiting for ship
|
||
// traffic to seed its lag gauge (the idle-readiness stall).
|
||
leader_last_seq: self.ship_feed.flushed_seq(),
|
||
..Default::default()
|
||
}
|
||
}
|
||
|
||
/// The leadership activation sequence for an ELECTED term (phase-4.md
|
||
/// §3): baseline = own flushed frontier (persisted), queue activated
|
||
/// term-scoped, then the term marker journaled as the term's FIRST log
|
||
/// entry (it lands at baseline+1 and ships to every follower), and only
|
||
/// then the leader view flips so the write path opens.
|
||
///
|
||
/// A marker failure aborts the leadership: a leader whose term marker is
|
||
/// not durable cannot prove its term.
|
||
pub(crate) fn become_leader_for_term(&self, term: u64) -> Result<()> {
|
||
// Capture the election-time position BEFORE the marker bumps the
|
||
// tail term: this is the heartbeat-announced reference a joining
|
||
// follower compares against for divergence (same numbering as the
|
||
// vote restriction). It is PUBLISHED only after the marker is
|
||
// durable (below), so an aborted activation never exposes it.
|
||
let prev = self.election_log_position();
|
||
let baseline = self.ship_feed.flushed_seq();
|
||
persist_stream_baseline(&self.data_dir, baseline);
|
||
self.stream_baseline.store(baseline, Ordering::Release);
|
||
self.ship_queue.activate_from(baseline, term);
|
||
let db = self.db()?;
|
||
if let Err(e) = db.append_term_marker(term, self.region.0) {
|
||
self.ship_queue.deactivate();
|
||
return Err(ServerError::Tidal(e));
|
||
}
|
||
*self
|
||
.activation_prev
|
||
.lock()
|
||
.unwrap_or_else(std::sync::PoisonError::into_inner) = prev;
|
||
*write_recovered(&self.leader, "leader") = Some(self.region);
|
||
// Record that we JOINED this elected term by WINNING it. A node never
|
||
// processes its own heartbeats, so without this the leader's `joined_term`
|
||
// stays 0 and `cluster_promote` mis-reads the elected shard as topology-era,
|
||
// routing a leader rebalance (`/cluster/shards/{id}/transfer`) down the
|
||
// LEGACY fan-out promote — which an elected target correctly fences with a
|
||
// 500. The won term marker is durable above, so this is a clean self-join.
|
||
if let Some(rt) = self.election_runtime.get() {
|
||
rt.note_self_won_term(term);
|
||
}
|
||
tracing::info!(
|
||
term,
|
||
baseline,
|
||
region = %self.region_name,
|
||
"election won: term marker journaled, ship queue active, writes open"
|
||
);
|
||
|
||
// m12p5 idle-readiness: a joiner that WINS leadership is, by construction,
|
||
// caught up to its own log — but it will never receive a heartbeat to
|
||
// drive `note_leader_frontier_for_readiness`, so converge the sticky latch
|
||
// here. Without this a promoted-then-elected joiner would stay 503 forever
|
||
// on an idle cluster (`is_ready` gates on `converged` for a joiner boot).
|
||
if (self.install_boot || self.seed_joiner) && !self.converged.swap(true, Ordering::AcqRel) {
|
||
tracing::info!(
|
||
term,
|
||
region = %self.region_name,
|
||
"joiner won leadership → trivially converged; readiness sticky-ready (m12p5)"
|
||
);
|
||
}
|
||
|
||
// NO reseed-marker discharge here, deliberately. Winning an election proves
|
||
// this node's log is at least as up-to-date as a QUORUM (the vote
|
||
// restriction) — it does NOT prove the node's log is CONTIGUOUS. A node
|
||
// whose applied frontier was re-based across a compacted gap satisfies the
|
||
// vote restriction while still missing history, so discharging here would
|
||
// promote a leader with a hole in its log. Earlier revisions cleared the
|
||
// marker at this point (first unconditionally, then on a frontier
|
||
// comparison); both could strand a divergent node as leader.
|
||
//
|
||
// A genuine false-alarm marker is discharged by
|
||
// `discharge_reseed_marker_if_served` when a catch-up pull actually serves
|
||
// the latching range, which is the same evidence a leader would need. A
|
||
// marker that survives to leadership is real: let the reseed run.
|
||
|
||
// §3.2 — the activation membership record (the linchpin): once the
|
||
// kind-3 marker is durable AND the membership era has begun (a kind-4
|
||
// record is in the log), re-append the CURRENT roster as a fresh kind-4
|
||
// record (version+1) stamped with this term. This is what makes
|
||
// committed membership reach every joiner of the term: it sits ABOVE the
|
||
// new baseline, so a follower behind at the transfer receives the
|
||
// current roster in-stream before any same-term traffic, and the
|
||
// one-at-a-time gate's predicate becomes the SAME-TERM commit of this
|
||
// record (never the activation-reset `committed()`).
|
||
//
|
||
// Mixed-version safety: the era only BEGINS via the capability-gated
|
||
// first conf-change, so an activation re-append never reaches a pre-p5
|
||
// follower (it would not be a voter that could elect this leader, and
|
||
// the gate refused the era-begin until every voter was kind-4 capable).
|
||
if self.membership.era_begun() {
|
||
let roster = self.membership.roster();
|
||
let record = super::membership::plan_activation_reappend(&roster, term);
|
||
match db.append_membership_record(record.clone()) {
|
||
Ok(seq) => {
|
||
// The activation re-append carries the CURRENT roster, so its
|
||
// diff against the just-applied view has no NEW removals to
|
||
// grace (removed members are already tombstoned, their peer
|
||
// entries already gone) — apply with no deferral.
|
||
self.apply_membership(&record, None);
|
||
tracing::info!(
|
||
term,
|
||
version = record.version,
|
||
seq,
|
||
"activation membership record re-appended (§3.2)"
|
||
);
|
||
}
|
||
Err(e) => {
|
||
// Non-fatal to the activation itself (the term is already
|
||
// won and durable). A failed re-append leaves the prior
|
||
// record as the latest; the next conf-change re-establishes
|
||
// the gate. Loud so a stuck membership is diagnosable.
|
||
tracing::error!(term, error = %e, "activation membership re-append failed");
|
||
}
|
||
}
|
||
}
|
||
Ok(())
|
||
}
|
||
|
||
/// Step down (or learn a new leader): the VIEW flips first so no new
|
||
/// write is accepted, THEN the ship queue deactivates (failing in-flight
|
||
/// `ack=quorum` waits) — design-review C3's ordering.
|
||
pub(crate) fn step_down_view(&self, term: u64, leader: Option<RegionId>) {
|
||
*write_recovered(&self.leader, "leader") = leader;
|
||
self.ship_queue.deactivate();
|
||
tracing::info!(
|
||
term,
|
||
leader = leader.map(|l| self.region_name_of(l).to_string()),
|
||
region = %self.region_name,
|
||
"leadership view updated (follower)"
|
||
);
|
||
}
|
||
|
||
/// A clean term-join (phase-4.md §5): jump this node's applied frontier
|
||
/// for the new leader's stream to the term's activation baseline —
|
||
/// seqnos at or below it are pre-stream history this node already holds.
|
||
pub(crate) fn note_term_joined(&self, term: u64, leader: RegionId, baseline: u64) {
|
||
if baseline > 0
|
||
&& let Ok(db) = self.db()
|
||
{
|
||
// m12 election-divergence-fix: advance PAST the term marker (baseline+1),
|
||
// not just to baseline. The marker is the term's first entry and has NO
|
||
// storage effect, so JOINING the term IS applying it. Advancing past it
|
||
// closes the gap a gap-gated marker leaves (and DRAINS any data parked
|
||
// behind that gap), curing the stuck `lag=1` convergence stall (Agent A:
|
||
// the marker rides the same seqno-gated segment path as data and a
|
||
// resume-from-`applied+1` skips it). Folding the term cell keeps
|
||
// `election_log_position` in the joined leader's numbering.
|
||
db.replication_state()
|
||
.advance(shard_of_region(leader), baseline + 1);
|
||
// m12 reseed-loop-fix: durably repair a STALE WAL-tail term on a clean
|
||
// join — the checkpoint-restore empty-WAL anomaly where `wal_term_mark()`
|
||
// reads a term BELOW `term` while the replication frontier already covers
|
||
// the baseline. A CAUGHT-UP shard that never installs a snapshot (the
|
||
// leader answers needed=false) is not reached by the install-boot term-
|
||
// marker synthesis, so without this it would re-read tail_term 0 on every
|
||
// reboot and self-restart-loop through a futile reseed (observed live on
|
||
// tidaldb-2). Writing a real kind-3 record makes tail_term truthful across
|
||
// reboots, reaching decide_join's `tail_term == term` clean-join fast path.
|
||
// Guarded so it fires once per stale term — a steady clean re-join just
|
||
// folds in memory (cheap), never a WAL append per heartbeat. On append
|
||
// failure, fall back to the in-memory fold: decide_join's frontier-covers-
|
||
// baseline arm still classifies the node Clean, so the loop stays broken.
|
||
if db.wal_term_mark().0 < term {
|
||
if let Err(e) = db.append_term_marker(term, leader.0) {
|
||
tracing::warn!(
|
||
term,
|
||
leader = leader.0,
|
||
error = %e,
|
||
"reseed-loop-fix: durable term-marker repair failed on clean join; \
|
||
folded in memory (tail-term is still truthful this boot)"
|
||
);
|
||
db.fold_term_marker(term, baseline + 1, leader.0);
|
||
}
|
||
} else {
|
||
db.fold_term_marker(term, baseline + 1, leader.0);
|
||
}
|
||
}
|
||
// m12 election-divergence-fix: a CLEAN join proves this node is caught up
|
||
// to (subsumed by) the new leadership, so it holds nothing un-replicated
|
||
// that term does not cover — reset the durable leader-acked frontier to 0.
|
||
// This keeps the frontier in the CURRENT era's numbering (it never carries
|
||
// a stale prior-era acked seqno into a later term's divergence compare) and
|
||
// is what lets a gracefully-restarted, previously-clean node read 0 instead
|
||
// of conservatively falling back to its WAL tail.
|
||
self.reset_leader_acked();
|
||
tracing::info!(
|
||
term,
|
||
leader = %self.region_name_of(leader),
|
||
baseline,
|
||
"joined leadership term"
|
||
);
|
||
}
|
||
|
||
/// Whether `region` is a `Removed` tombstone in this node's APPLIED roster
|
||
/// (m11p5 §3.3 typed removed signal). Read by the heartbeat / vote handlers
|
||
/// to stamp the `removed` bit on a refusal so the removed peer learns of its
|
||
/// decommission even if it missed the `Removed` record. An era-0 / pre-
|
||
/// membership node has no `Removed` members → always false.
|
||
pub(crate) fn is_member_removed(&self, region: RegionId) -> bool {
|
||
self.membership.roster().role_of(region) == Some(tidaldb::wal::format::MemberRole::Removed)
|
||
}
|
||
|
||
/// Latch the typed-removed-signal decommission (m11p5 §3.3): a voter told
|
||
/// this node (via a heartbeat / vote reply) that its own region is a
|
||
/// `Removed` member. This is the MISSED-RECORD delivery path — the node was
|
||
/// down during the removal-delivery grace and never folded the `Removed`
|
||
/// record, so its local cell still shows it as a member. Flip readiness 503
|
||
/// + log the decommission runbook pointer. Idempotent (one-way latch).
|
||
///
|
||
/// Explicitly NOT a reseed: a remove is not a reseed (§3.3), so this never
|
||
/// touches the reseed marker. Campaign suppression is applied by the caller
|
||
/// under the machine lock (the driver's lock discipline).
|
||
pub(crate) fn note_removed_by_peer_signal(&self) {
|
||
// One-way latch: only log the FIRST transition (every subsequent
|
||
// heartbeat would otherwise re-log on every round).
|
||
if self.decommissioned_by_signal.swap(true, Ordering::Release) {
|
||
return;
|
||
}
|
||
tracing::error!(
|
||
region = %self.region_name,
|
||
"DECOMMISSIONED: a voter's heartbeat/vote refusal reports this node as a \
|
||
Removed member, but this node never folded the Removed record (it was down \
|
||
during the removal-delivery grace). Flipping readiness to 503 and suppressing \
|
||
campaigning. This is NOT a reseed — the node stays decommissioned until torn \
|
||
down. Recovery: delete this node (runbook §8 'decommission')"
|
||
);
|
||
}
|
||
|
||
/// The divergence quarantine (phase-4.md §5): loud, actionable, metered.
|
||
/// Additionally latches the durable reseed marker (reason `quarantine`,
|
||
/// m11p5 §2.4) so the reseed runs on the next boot — and a successful reseed
|
||
/// boot clears the quarantine latch (the WAL becomes the leader's copy).
|
||
pub(crate) fn note_quarantined(
|
||
self: &Arc<Self>,
|
||
term: u64,
|
||
tail_term: u64,
|
||
frontier: u64,
|
||
baseline: u64,
|
||
) {
|
||
self.cluster_metrics.set_divergence_quarantined(true);
|
||
// m12p6: re-baseline a DIVERGENT node (frontier ABOVE the new term's
|
||
// baseline) from `baseline` — `frontier + 1` would land above the leader's
|
||
// WAL tail → `wal_covers` answers needed=false → the install no-ops and the
|
||
// divergent suffix self-restart-loops. A `from_seqno <= baseline` forces the
|
||
// snapshot that discards the suffix. (When the frontier is at/below the
|
||
// baseline — not the divergent case — `frontier + 1` lets the leader pick a
|
||
// cheap catch-up vs snapshot, avoiding a needless reseed cascade.)
|
||
let from_seqno = if frontier > baseline {
|
||
baseline
|
||
} else {
|
||
frontier.saturating_add(1)
|
||
};
|
||
self.latch_reseed_marker(ReseedReason::Quarantine, from_seqno);
|
||
tracing::error!(
|
||
term,
|
||
tail_term,
|
||
frontier,
|
||
baseline,
|
||
region = %self.region_name,
|
||
"DIVERGENT SUFFIX: this node's WAL extends past the elected term's baseline with pre-term data (leader-acked writes the cluster elected past). QUARANTINED from the data plane — it still votes, but applies and reports nothing. Recovery: reseed this node from the leader (runbook §8); m11p5's snapshot transfer automates this"
|
||
);
|
||
}
|
||
|
||
/// Durably latch the reseed marker (m11p5 §2.4), idempotently.
|
||
///
|
||
/// Sets the `tidaldb_cluster_reseed_required` gauge and persists the marker
|
||
/// (reason + the seqno the next-boot fetch resumes from). Idempotent: a
|
||
/// re-latch from the same or a different reason simply rewrites the marker.
|
||
/// On `reseed_self_restart`, additionally evaluates the §2.4 quorum refusal
|
||
/// and either drains+exits or refuses (loudly, in status).
|
||
///
|
||
/// A failed persist is logged at ERROR but never panics — the node keeps
|
||
/// serving degraded; the runtime retry timer keeps its standing wake-up, so
|
||
/// a transient fsync failure does not strand the node (the next refusal
|
||
/// re-latches).
|
||
pub(crate) fn latch_reseed_marker(self: &Arc<Self>, reason: ReseedReason, from_seqno: u64) {
|
||
// Readiness gating (m12 reseed-loop-fix): the node is operationally in the
|
||
// reseed-required state the instant it latches, regardless of whether the
|
||
// durable persist below succeeds — flip readiness to 503 now so it drains
|
||
// from the client VIP. Cleared by `clear_stale_reseed_marker_if_caught_up`
|
||
// once it catches up via the stream, or consumed by the next-boot reseed.
|
||
self.reseed_marker_latched.store(true, Ordering::Release);
|
||
let marker = ReseedMarker { reason, from_seqno };
|
||
match self.reseed_marker_store.persist(marker) {
|
||
Ok(()) => {
|
||
self.cluster_metrics.set_reseed_required(true);
|
||
tracing::warn!(
|
||
region = %self.region_name,
|
||
reason = reason.as_str(),
|
||
from_seqno,
|
||
"reseed marker latched durably (m11p5 §2.4); the reseed runs on the next \
|
||
boot. Serving stays degraded until then."
|
||
);
|
||
}
|
||
Err(e) => {
|
||
tracing::error!(
|
||
region = %self.region_name,
|
||
reason = reason.as_str(),
|
||
error = %e,
|
||
"failed to durably latch the reseed marker; the node keeps serving degraded \
|
||
and the standing retry timer re-latches on the next refusal"
|
||
);
|
||
// Reflect the latch intent in the gauge even on a persist miss
|
||
// (the node IS in the reseed-required state operationally).
|
||
self.cluster_metrics.set_reseed_required(true);
|
||
return;
|
||
}
|
||
}
|
||
if self.reseed_self_restart {
|
||
self.spawn_self_restart_eval();
|
||
}
|
||
}
|
||
|
||
/// Discharge a reseed marker on POSITIVE EVIDENCE that the stream served the
|
||
/// range which latched it: a `StreamSegments` pull that began at or below the
|
||
/// marker's `from_seqno` and ran to completion.
|
||
///
|
||
/// This exists for the genuine false-alarm case (the `own < prev_log →
|
||
/// ReseedRequired` join arm latches for a node merely BEHIND by a shippable
|
||
/// tail, which the stream then serves), without the unsoundness of the two
|
||
/// frontier comparisons that preceded it.
|
||
///
|
||
/// # Why no frontier comparison can work here
|
||
///
|
||
/// The first version asked `applied >= leader_last_seq`, asserting "a node
|
||
/// genuinely behind a COMPACTED gap never reaches caught_up". The second asked
|
||
/// `applied >= marker.from_seqno`. BOTH are unsound for the same reason: the
|
||
/// applied frontier is a HIGH-WATER-MARK, and a term join re-bases it onto the
|
||
/// new leader's stream (`replication_state().advance(.., baseline + 1)`), so it
|
||
/// leaps over history the node never received.
|
||
///
|
||
/// Measured, not theorised — `mp_follower_reseeds_via_snapshot_after_compaction`
|
||
/// stops a follower at frontier 9, compacts the leader so it retains only from
|
||
/// 15722, and the follower's frontier is re-based to 16810. Both predicates
|
||
/// discharge the marker there. The node then skips its reseed and reports
|
||
/// `lag_events: 0` while missing 10..15721 — a silent hole, served to readers.
|
||
/// Production showed the same shape: `applied 13540660` against a marker
|
||
/// resuming at 13540653 that no live WAL could serve.
|
||
///
|
||
/// So the only sound discharges are: a completed pull covering the gap (here),
|
||
/// or a snapshot install (which replaces the data dir and the marker with it).
|
||
fn discharge_reseed_marker_if_served(&self, served_from: u64) {
|
||
if self
|
||
.election_runtime
|
||
.get()
|
||
.is_some_and(|rt| rt.is_quarantined())
|
||
{
|
||
return;
|
||
}
|
||
let Ok(Some(marker)) = self.reseed_marker_store.load() else {
|
||
return;
|
||
};
|
||
if !marker.discharged_by_served_range(served_from) {
|
||
// Refuse quietly. The undischarged state is already observable: the
|
||
// latch logged WARN, `tidaldb_cluster_reseed_required` stays 1 (which
|
||
// is what makes the 10m alert reachable), and `/cluster/status/local`
|
||
// reports `reseed_required: true`.
|
||
return;
|
||
}
|
||
match self.reseed_marker_store.clear() {
|
||
Ok(()) => {
|
||
self.cluster_metrics.set_reseed_required(false);
|
||
// Readiness gating (m12 reseed-loop-fix): the marker is healed —
|
||
// clear the readiness latch so `is_ready` can return 200 again.
|
||
self.reseed_marker_latched.store(false, Ordering::Release);
|
||
tracing::info!(
|
||
region = %self.region_name,
|
||
reason = marker.reason.as_str(),
|
||
from_seqno = marker.from_seqno,
|
||
served_from,
|
||
"reseed marker discharged — a catch-up pull completed from at/below the \
|
||
seqno that latched it, so the stream genuinely closed the gap and no \
|
||
reseed is needed"
|
||
);
|
||
}
|
||
Err(e) => tracing::warn!(
|
||
region = %self.region_name,
|
||
error = %e,
|
||
"failed to clear a discharged reseed marker; retried on the next served pull"
|
||
),
|
||
}
|
||
}
|
||
|
||
/// Offload the §2.4 quorum-refusal evaluation to a detached thread.
|
||
///
|
||
/// `latch_reseed_marker` may be invoked from a tonic handler thread (the
|
||
/// quarantine path runs inside `on_heartbeat`); the quorum poll is a
|
||
/// BLOCKING HTTP fan-out, so it must NOT run inline on an async worker. The
|
||
/// thread holds an `Arc` upgrade of `self` so it outlives the caller; it is a
|
||
/// rare path (only `reseed_self_restart: true` + a fresh latch).
|
||
fn spawn_self_restart_eval(self: &Arc<Self>) {
|
||
let node = Arc::clone(self);
|
||
if let Err(e) = std::thread::Builder::new()
|
||
.name("tidal-reseed-quorum-eval".into())
|
||
.spawn(move || node.maybe_self_restart())
|
||
{
|
||
tracing::error!(error = %e, "failed to spawn reseed quorum-eval thread");
|
||
}
|
||
}
|
||
|
||
/// Evaluate the §2.4 quorum refusal and, when safe, trigger a graceful
|
||
/// drain + clean exit(0) so the node restarts and reseeds.
|
||
///
|
||
/// The refusal poll asks the OTHER voters' `/cluster/status/local` (short
|
||
/// timeout, the same status client) how many are alive. If
|
||
/// `alive_voters_excluding_self < majority(total_voters)` the self-restart is
|
||
/// REFUSED (the marker stays latched, an ERROR is logged, a refusal note is
|
||
/// set in status) — exiting during a 2-voter window is total write
|
||
/// unavailability. Otherwise it flips readiness to 503 and exits via the
|
||
/// existing graceful-shutdown path (NOT `process::abort`).
|
||
fn maybe_self_restart(self: &Arc<Self>) {
|
||
// Poll the OTHER regions' local status synchronously on a throwaway
|
||
// blocking client (this runs on the write pool, never the reactor).
|
||
// The voter count from the effective roster (§3): era 0 = the full
|
||
// topology, the membership era = the live voters only (tombstones and
|
||
// learners excluded — a quorum is over voters).
|
||
let total_voters = self.membership.roster().voter_ids().len();
|
||
let alive_others = self.count_alive_other_voters();
|
||
if !reseed::self_restart_quorum_safe(alive_others, total_voters) {
|
||
self.self_restart_refused.store(true, Ordering::Release);
|
||
tracing::error!(
|
||
region = %self.region_name,
|
||
alive_voters_excluding_self = alive_others,
|
||
total_voters,
|
||
majority = reseed::majority(total_voters),
|
||
"reseed_self_restart REFUSED: the remaining voters cannot sustain quorum without \
|
||
this node (alive_others < majority). Keeping the marker latched and serving \
|
||
degraded — exiting now would be total write unavailability (§2.4). The reseed \
|
||
runs once quorum is safe (a future latch re-evaluates)."
|
||
);
|
||
return;
|
||
}
|
||
self.self_restart_refused.store(false, Ordering::Release);
|
||
// m12 reseed-loop-fix: re-check the marker before committing to the exit.
|
||
// A snapshot_required latch from a TRANSIENT classification — the leader's
|
||
// baseline advanced past this node's persisted frontier while it was down,
|
||
// so the first heartbeat's decide_join saw it briefly behind — self-heals
|
||
// via the stream: `clear_stale_reseed_marker_if_caught_up` clears the marker
|
||
// once caught up, journaling the durable term marker on the clean join. The
|
||
// quorum poll above took time (a blocking peer fan-out); if the marker
|
||
// healed meanwhile, a reboot would reseed NOTHING (the leader answers
|
||
// needed=false for a caught-up shard), so self-restarting is futile and
|
||
// loops. Abort. Only a marker that CANNOT self-heal (a genuine compacted
|
||
// gap, still latched here) proceeds to the restart.
|
||
if !self.reseed_marker_latched.load(Ordering::Acquire) {
|
||
tracing::info!(
|
||
region = %self.region_name,
|
||
"reseed_self_restart: the reseed marker healed via stream catch-up during the \
|
||
quorum check — the shard is caught up, a reboot would reseed nothing. Aborting \
|
||
the self-restart (m12 reseed-loop-fix)."
|
||
);
|
||
return;
|
||
}
|
||
// m12 reseed-loop-fix (Fix 3): the process-wide exit is owned by the
|
||
// node-level coordinator, not this single shard. It fires once, only
|
||
// after every hosted shard has also requested a restart or a bounded
|
||
// grace has elapsed — so this shard's exit never aborts a co-hosted
|
||
// sibling's in-flight install. A single-shard host (production) fires
|
||
// immediately (FireNow), so this is a behavioral no-op there.
|
||
match super::reseed_restart::request_restart(self.group_shard) {
|
||
super::reseed_restart::Decision::FireNow => {
|
||
tracing::warn!(
|
||
region = %self.region_name,
|
||
alive_voters_excluding_self = alive_others,
|
||
total_voters,
|
||
"reseed_self_restart: quorum is safe and every hosted shard has settled — \
|
||
draining and exiting cleanly so the next boot reseeds via snapshot (§2.4). \
|
||
Readiness flips to 503 before drain; the graceful-shutdown path, never abort."
|
||
);
|
||
self.fire_graceful_self_restart();
|
||
}
|
||
super::reseed_restart::Decision::Defer { deadline } => {
|
||
tracing::warn!(
|
||
region = %self.region_name,
|
||
grace_secs = super::reseed_restart::grace().as_secs(),
|
||
"reseed_self_restart: quorum is safe, but deferring the process-wide exit so a \
|
||
co-hosted sibling shard's in-flight install can finish first (Fix 3); the exit \
|
||
fires once all hosted shards settle or the grace elapses."
|
||
);
|
||
let node = Arc::clone(self);
|
||
if let Err(e) = std::thread::Builder::new()
|
||
.name("tidal-reseed-restart-defer".into())
|
||
.spawn(move || {
|
||
let now = Instant::now();
|
||
if deadline > now {
|
||
std::thread::sleep(deadline - now);
|
||
}
|
||
// m12 reseed-loop-fix: the grace window is exactly when a
|
||
// transient snapshot_required latch self-heals via stream
|
||
// catch-up. Re-check before firing — if THIS shard's marker
|
||
// healed, a reboot would reseed nothing (caught up), so abort
|
||
// the futile self-restart. Only a still-latched marker (a
|
||
// genuine compacted gap) fires the exit.
|
||
if !node.reseed_marker_latched.load(Ordering::Acquire) {
|
||
tracing::info!(
|
||
region = %node.region_name,
|
||
"reseed_self_restart: the reseed marker healed via stream \
|
||
catch-up during the Fix 3 grace — aborting the self-restart \
|
||
(the shard is caught up; a reboot would reseed nothing)."
|
||
);
|
||
return;
|
||
}
|
||
if super::reseed_restart::fire_due() {
|
||
node.fire_graceful_self_restart();
|
||
}
|
||
})
|
||
{
|
||
tracing::error!(error = %e, "failed to spawn reseed self-restart defer thread");
|
||
}
|
||
}
|
||
super::reseed_restart::Decision::AlreadyFired => {
|
||
tracing::info!(
|
||
region = %self.region_name,
|
||
"reseed_self_restart: a co-hosted shard already triggered the process-wide \
|
||
exit; this shard's marker stays latched and rides the same graceful restart."
|
||
);
|
||
}
|
||
}
|
||
}
|
||
|
||
/// Flip readiness to 503 and trigger the graceful self-exit (drain →
|
||
/// checkpoint → clean exit(0)), after a brief grace so an in-flight admin
|
||
/// reseed response flushes. The SINGLE firing path for both the immediate
|
||
/// (`FireNow`) and the deferred (Fix 3 grace-deadline) restart decisions. We
|
||
/// use the existing shutdown-signal path (SIGTERM-equivalent) so
|
||
/// `serve_state`'s graceful shutdown fires — never `process::abort`.
|
||
fn fire_graceful_self_restart(&self) {
|
||
if let Err(error) = write_reseed_termination_message(
|
||
std::path::Path::new(RESEED_TERMINATION_LOG),
|
||
&self.region_name,
|
||
self.group_shard,
|
||
) {
|
||
tracing::warn!(
|
||
path = RESEED_TERMINATION_LOG,
|
||
region = %self.region_name,
|
||
shard = self.group_shard.0,
|
||
error = %error,
|
||
"failed to write Kubernetes termination message for reseed self-restart"
|
||
);
|
||
}
|
||
self.set_shutting_down();
|
||
if let Err(e) = std::thread::Builder::new()
|
||
.name("tidal-reseed-restart".into())
|
||
.spawn(|| {
|
||
// A brief grace so an in-flight admin response (the 202 from
|
||
// /cluster/reseed) flushes before the drain begins.
|
||
std::thread::sleep(Duration::from_millis(200));
|
||
trigger_graceful_self_exit();
|
||
})
|
||
{
|
||
tracing::error!(error = %e, "failed to spawn reseed self-restart thread");
|
||
}
|
||
}
|
||
|
||
/// Count how many OTHER voters report reachable + alive via
|
||
/// `/cluster/status/local` (a short-timeout blocking poll). Used by the §2.4
|
||
/// quorum refusal. Excludes this node.
|
||
fn count_alive_other_voters(&self) -> usize {
|
||
// m12p6 FIX: use the SHARED `blocking_client`, which is built with the
|
||
// cluster CA as a trust anchor (`build_forwarding_clients`). A freshly
|
||
// built bare `reqwest::blocking::Client` trusts only the SYSTEM roots, so
|
||
// on a TLS cluster every `https://peer:9500/cluster/status/local` poll
|
||
// fails certificate verification → this returned 0 → the reseed
|
||
// self-restart was ALWAYS refused (it believed it would break quorum) →
|
||
// a node that needed a reseed wedged forever instead of self-healing.
|
||
// The per-request timeout keeps the tight status budget.
|
||
let mut alive = 0usize;
|
||
for (rid, _name, http) in self.all_regions_for_status() {
|
||
if rid == self.region {
|
||
continue;
|
||
}
|
||
let Some(http_addr) = http else { continue };
|
||
let url = peer_url(&http_addr, "/cluster/status/local");
|
||
let mut req = self
|
||
.blocking_client
|
||
.get(&url)
|
||
.timeout(forward::STATUS_PEER_TIMEOUT);
|
||
if let Some(key) = crate::cluster::security::bearer_from_env() {
|
||
req = req.bearer_auth(key);
|
||
}
|
||
if let Ok(resp) = req.send()
|
||
&& resp.status().is_success()
|
||
{
|
||
alive += 1;
|
||
}
|
||
}
|
||
alive
|
||
}
|
||
|
||
/// Record convergence progress for the sticky readiness latch (§4): an
|
||
/// install-boot node becomes ready once its catch-up lag first falls at or
|
||
/// below `learner_promote_lag` (never `lag == 0`). Called from the status
|
||
/// path (which already computes lag) so no new polling thread is needed.
|
||
/// This follower's applied frontier for the CURRENT leader's source shard
|
||
/// (`shard_of_region(leader)`), or 0 if the shard is untracked or the db is
|
||
/// gone. The single reader for "how far have I applied the leader's stream",
|
||
/// shared by the heartbeat readiness drive ([`note_leader_frontier_for_readiness`])
|
||
/// and `local_status` so the per-source-shard keying (BUG 1) lives in one place.
|
||
///
|
||
/// [`note_leader_frontier_for_readiness`]: Self::note_leader_frontier_for_readiness
|
||
fn applied_for_leader_shard(&self, leader_shard: ShardId) -> u64 {
|
||
self.db()
|
||
.ok()
|
||
.and_then(|db| db.replication_state().applied_seqno(leader_shard))
|
||
.unwrap_or(0)
|
||
}
|
||
|
||
/// Record convergence progress for the sticky readiness latch (§4).
|
||
///
|
||
/// `leader_seqno` is this node's KNOWLEDGE of the followed leader's frontier.
|
||
/// It is required and must be non-zero: `lag_events` is
|
||
/// `leader_seqno - applied`, an unsigned subtraction, so a node that has
|
||
/// learned nothing computes `0 - 0 = 0` and would latch "converged" while
|
||
/// holding no data at all. That is not hypothetical — in the multi-group
|
||
/// reproduction all three groups reported `lag_events: 0` with
|
||
/// `applied_events` of 24, 14 and 0 against 5600 written items, and in
|
||
/// production a PVC-wiped tidaldb-0 entered the client VIP with an EMPTY
|
||
/// corpus. Convergence must rest on a frontier we actually learned.
|
||
fn note_lag_for_readiness(&self, leader_seqno: u64, lag_events: u64) {
|
||
if leader_seqno == 0 {
|
||
return; // no information yet — absence of lag is not convergence
|
||
}
|
||
if !self.converged.load(Ordering::Acquire) && lag_events <= self.learner_promote_lag {
|
||
self.converged.store(true, Ordering::Release);
|
||
tracing::info!(
|
||
region = %self.region_name,
|
||
leader_seqno,
|
||
lag_events,
|
||
threshold = self.learner_promote_lag,
|
||
seed_joiner = self.seed_joiner,
|
||
install_boot = self.install_boot,
|
||
"first-converged against a KNOWN leader frontier (lag <= learner_promote_lag); \
|
||
readiness is now sticky-ready for this process (§4 hysteresis)"
|
||
);
|
||
}
|
||
}
|
||
|
||
/// m12p5 idle-readiness drive: fold a leader heartbeat's live flushed
|
||
/// frontier into the lag gauge and the sticky readiness latch.
|
||
///
|
||
/// The pre-m12p5 readiness latch (`note_lag_for_readiness`) only fired when
|
||
/// something recomputed lag — observed ship traffic seeding the gauge, or an
|
||
/// external `/cluster/status/local` poll calling `local_status`. On an IDLE
|
||
/// cluster neither happens, so a freshly caught-up joiner could sit 503 for
|
||
/// hours (WORKLOG 2026-06-13: an 11.5h stall) and never join the VIP. The
|
||
/// heartbeat, by contrast, flows every heartbeat interval regardless of write
|
||
/// traffic and now carries `leader_last_seq` — the leader's live frontier in
|
||
/// the SAME stream numbering as a follower's per-shard `applied_seqno`.
|
||
///
|
||
/// Seeding the gauge (monotonic) keeps the lag metric/status truthful on an
|
||
/// idle cluster for every follower; the readiness note then converges a
|
||
/// caught-up joiner using a REAL leader frontier (never the uninitialized-0
|
||
/// gauge reading, which would false-converge a still-behind joiner).
|
||
/// `leader_last_seq == 0` = a pre-m12p5 leader conveyed nothing → no-op
|
||
/// (the status-poll path still applies).
|
||
pub(crate) fn note_leader_frontier_for_readiness(
|
||
&self,
|
||
leader_region: RegionId,
|
||
leader_last_seq: u64,
|
||
) {
|
||
if leader_last_seq == 0 {
|
||
return;
|
||
}
|
||
let Ok(db) = self.db() else { return };
|
||
let leader_shard = shard_of_region(leader_region);
|
||
// Keep the lag gauge (hence the `lag_segments` metric and `local_status`)
|
||
// fresh on an idle cluster, not only after ship traffic. Monotonic per
|
||
// shard, so a stale/lower heartbeat can never walk it back.
|
||
db.control_plane()
|
||
.lag_gauge()
|
||
.update_leader_seqno_for(leader_shard, leader_last_seq);
|
||
let applied = self.applied_for_leader_shard(leader_shard);
|
||
// m12 seed-join promotion fix: report this follower's caught-up frontier
|
||
// BACK to the leader on the heartbeat path. The leader's durable per-peer
|
||
// `learner_mark` advances ONLY from a follower frontier-report, which the
|
||
// receiver emits after APPLYING a streamed event. A node that converged via
|
||
// a SNAPSHOT INSTALL (frontier seeded at boot, the leader's WAL already
|
||
// covering it) has nothing to stream-apply, so it never tells the leader it
|
||
// is caught up — and the auto-promotion gate (`flushed - learner_mark <=
|
||
// learner_promote_lag`) strands a seed-joined Learner forever (regression:
|
||
// `mp_seed_join_snapshot_catchup`; before Fix 1's frontier-seed the joiner
|
||
// re-pulled from seqno 1 and THOSE applies emitted the reports that promoted
|
||
// it). The heartbeat flows on an IDLE cluster and carries the CURRENT term,
|
||
// so this report is both recurring (survives the join/registration race) and
|
||
// term-correct (the term-checked `update_peer_for_term` fold accepts it,
|
||
// unlike a boot-time report stamped term 0). `notify_applied` dedups per
|
||
// shard (only an ADVANCED frontier pushes), so a steady follower never spams.
|
||
// Safe: it reports this node's TRUE applied frontier (never above what it
|
||
// durably holds, so it cannot over-credit a not-caught-up node), and learner
|
||
// marks never feed the quorum commit index — they gate promotion only.
|
||
// SCOPED to a LEARNER: a Voter's frontier already reaches the leader via
|
||
// the ship-ack path, and a voter mark DOES feed `compute_commit`, so folding
|
||
// one off the heartbeat could perturb the same-term commit gate (Raft fig-8);
|
||
// a learner mark never feeds `compute_commit`, so this is provably
|
||
// commit-safe and is exactly the signal the auto-promotion gate consumes.
|
||
if applied > 0
|
||
&& self.membership.self_role() == Some(tidaldb::wal::format::MemberRole::Learner)
|
||
{
|
||
self.transport.notify_applied(leader_shard, applied);
|
||
}
|
||
// Drive the convergence latch for EVERY boot, not just install/seed-join.
|
||
// A plain restarted voter used to skip this entirely and be Ready on
|
||
// arrival, so it joined the client VIP before it knew whether it held the
|
||
// data — which is how a PVC-wiped tidaldb-0 served an EMPTY corpus. The
|
||
// heartbeat carries the leader's live frontier, so it is the signal that
|
||
// makes convergence knowable on an idle cluster too (m12p5).
|
||
if !self.converged.load(Ordering::Acquire) {
|
||
self.note_lag_for_readiness(leader_last_seq, leader_last_seq.saturating_sub(applied));
|
||
}
|
||
// NO reseed-marker discharge on the heartbeat path. This is where both
|
||
// unsound predicates lived (`applied >= leader_last_seq`, then `applied >=
|
||
// marker.from_seqno`). A heartbeat carries frontier numbers only, and a
|
||
// frontier is a high-water-mark that a term join re-bases across
|
||
// un-received history — see `discharge_reseed_marker_if_served`, which is
|
||
// driven by a COMPLETED catch-up pull instead.
|
||
}
|
||
|
||
/// Whether this node is READY to serve (m11p5 §4 readiness predicate).
|
||
///
|
||
/// 503 while shutting down, quarantined, REMOVED, or an install/seed-join
|
||
/// boot has not yet first-converged. A restarted NON-install (PVC-retained)
|
||
/// voter keeps today's behavior (ready on today's terms).
|
||
fn is_ready(&self) -> bool {
|
||
if self.is_shutting_down() {
|
||
return false;
|
||
}
|
||
if self
|
||
.election_runtime
|
||
.get()
|
||
.is_some_and(|rt| rt.is_quarantined())
|
||
{
|
||
return false;
|
||
}
|
||
// m11p5 §3.3: a node a `Removed` record retired flips to 503 (and stops
|
||
// campaigning via the campaign gate) once the record reaches it — the
|
||
// removal-delivery grace guarantees the removed peer is delivered the
|
||
// record before its ship cell retires, so this view IS applied here.
|
||
if self.membership.self_role() == Some(tidaldb::wal::format::MemberRole::Removed) {
|
||
return false;
|
||
}
|
||
// m11p5 §3.3 MISSED-RECORD path: a node that was down during the
|
||
// removal-delivery grace never folded the `Removed` record, so its cell
|
||
// still shows it as a member — but a voter's heartbeat/vote refusal told
|
||
// it (the typed removed signal) it is decommissioned. Honor that latch.
|
||
if self.decommissioned_by_signal.load(Ordering::Acquire) {
|
||
return false;
|
||
}
|
||
// m12 reseed-loop-fix (readiness gating): an unhealed reseed marker
|
||
// (snapshot-required or quarantine) means this node holds stale data it is
|
||
// about to discard — drain it from the client VIP until it heals. The latch
|
||
// clears when a completed catch-up pull proves the stream served the gap
|
||
// (`discharge_reseed_marker_if_served`) or the next boot reseeds.
|
||
if self.reseed_marker_latched.load(Ordering::Acquire) {
|
||
return false;
|
||
}
|
||
// POSITIVE EVIDENCE, every boot. This was
|
||
// `if (install_boot || seed_joiner) && !converged`, so a plain restarted
|
||
// voter fell straight through to ready — admitted to the client VIP before
|
||
// it had learned the leader's frontier, let alone caught up. Combined with
|
||
// `lag_events` reading `0 - 0 = 0` on an uninitialized gauge, that is how a
|
||
// PVC-wiped tidaldb-0 served an EMPTY corpus, and how a node missing 15,000
|
||
// entries reported itself converged.
|
||
//
|
||
// `converged` is STICKY for the process (§4 hysteresis), so this costs a
|
||
// restarted voter only the time to receive one heartbeat carrying the
|
||
// leader's frontier — heartbeats flow on an idle cluster (m12p5) — and a
|
||
// later transient leader loss never un-readies it. A node that genuinely
|
||
// cannot reach a leader stays 503, which is the honest answer: it does not
|
||
// know whether it holds the data.
|
||
if !self.converged.load(Ordering::Acquire) {
|
||
// A LEADER is trivially converged: it WRITES the log rather than
|
||
// applying someone else's, so there is no frontier to catch up to. This
|
||
// arm is load-bearing for bootstrap — a fresh cluster's leader has
|
||
// `last_seq == 0`, so requiring a non-zero learned frontier would leave
|
||
// it permanently 503 and the cluster would never come up.
|
||
//
|
||
// It MUST test ESTABLISHED leadership from the election runtime, never
|
||
// `current_leader()`. That view is seeded from the TOPOLOGY FILE, and in
|
||
// a sharded topology group `s` names node `s` as its term-0 leader — so
|
||
// a booting node would self-certify convergence for the group it merely
|
||
// believes it leads, while holding none of that group's data. The
|
||
// durable §1.4-1 rule is that a restart always boots a FOLLOWER, so the
|
||
// runtime role is the only honest source here.
|
||
let established_leader = self
|
||
.election_runtime
|
||
.get()
|
||
.is_some_and(|rt| matches!(rt.role(), tidaldb::replication::Role::Leader));
|
||
if established_leader {
|
||
self.converged.store(true, Ordering::Release);
|
||
} else {
|
||
return false;
|
||
}
|
||
}
|
||
true
|
||
}
|
||
|
||
/// Start the election driver (m11p4). Called once the node is in its
|
||
/// final `Arc` (the driver holds a `Weak` back-reference).
|
||
pub fn start_election_driver(self: &Arc<Self>) {
|
||
let Some(boot) = self
|
||
.election_boot
|
||
.lock()
|
||
.unwrap_or_else(std::sync::PoisonError::into_inner)
|
||
.take()
|
||
else {
|
||
return;
|
||
};
|
||
// Seed the randomized election timeouts from wall-clock entropy +
|
||
// the region id, so simultaneous boots draw different timeouts.
|
||
let seed = u64::from(boot.config.self_region.0)
|
||
^ std::time::SystemTime::now()
|
||
.duration_since(std::time::UNIX_EPOCH)
|
||
.map_or(0x5EED, |d| d.as_nanos() as u64);
|
||
let machine = tidaldb::replication::ElectionState::new(
|
||
boot.config,
|
||
boot.hard.current_term,
|
||
boot.hard.voted_for,
|
||
boot.boots_as_leader,
|
||
boot.topology_leader,
|
||
seed,
|
||
std::time::Instant::now(),
|
||
);
|
||
let runtime =
|
||
super::election_driver::start(self, machine, boot.store, &self.election_hooks_cell);
|
||
let _ = self.election_runtime.set(runtime);
|
||
|
||
// m11p5 §3.3/§3.4: reconcile the election surfaces to the effective
|
||
// roster when the view booted FROM THE CELL (a kind-4 record is in the
|
||
// WAL — a membership-era restart OR a seed-join snapshot install). The
|
||
// election config was built from the synthesized/topology positional
|
||
// tables (all-voters), but the cell is authoritative: a node whose
|
||
// recovered role is LEARNER must boot CAMPAIGN-SUPPRESSED (it is not in
|
||
// the voter set and never self-starts an election), and the voter set +
|
||
// majority must match the cell — NOT the topology. Without this, a
|
||
// seed-joined learner boots believing it is a voter and campaigns,
|
||
// disrupting the cluster it just joined. The runtime apply path
|
||
// (`maybe_apply_membership_from_cell`) is a no-op here because the cell
|
||
// version already equals the applied view version, so the reconcile must
|
||
// happen ONCE at startup.
|
||
if self.membership.era_begun()
|
||
&& let Some(rt) = self.election_runtime.get()
|
||
{
|
||
use tidaldb::wal::format::MemberRole;
|
||
let roster = self.membership.roster();
|
||
let election_voters: Vec<RegionId> = roster
|
||
.voter_ids()
|
||
.into_iter()
|
||
.filter(|&r| r != self.region)
|
||
.collect();
|
||
let election_learners: Vec<RegionId> = roster
|
||
.learner_ids()
|
||
.into_iter()
|
||
.filter(|&r| r != self.region)
|
||
.collect();
|
||
let self_is_voter = roster.role_of(self.region) == Some(MemberRole::Voter);
|
||
// The election surface: a learner boots NOT in the voter set (campaign
|
||
// gated) until a Voter record applies; voters AND learners are in the
|
||
// heartbeat fan-out so this node both campaigns toward voters (if it is
|
||
// one) and ships heartbeats to learners (if it leads).
|
||
rt.reconfigure_voters(election_voters, &election_learners, self_is_voter);
|
||
// The commit-index surface: split voters/learners per the cell (boot
|
||
// built it all-voter from the synthesized topology). The PeerPool /
|
||
// ShipQueue peers are already correct (the synthesized topology lists
|
||
// every live region), so only the quorum math needs the cell's split.
|
||
let commit_voters: Vec<ShardId> = roster
|
||
.voter_ids()
|
||
.into_iter()
|
||
.filter(|&r| r != self.region)
|
||
.map(shard_of_region)
|
||
.collect();
|
||
let commit_learners: Vec<ShardId> = roster
|
||
.learner_ids()
|
||
.into_iter()
|
||
.filter(|&r| r != self.region)
|
||
.map(shard_of_region)
|
||
.collect();
|
||
self.commit.reconfigure(&commit_voters, &commit_learners);
|
||
tracing::info!(
|
||
version = roster.version,
|
||
self_is_voter,
|
||
voters = commit_voters.len(),
|
||
learners = commit_learners.len(),
|
||
"boot: reconciled election + commit surfaces to the WAL cell (§3.4 — a learner \
|
||
boots campaign-suppressed; the quorum math matches the cell, not the topology)"
|
||
);
|
||
}
|
||
|
||
// m11p5 §2.4: late-bind the snapshot-required refusal sink now that the
|
||
// node is in its final Arc (the sink holds a Weak back-reference). When a
|
||
// catch-up pull is refused with the `snapshot-required` trailer, the
|
||
// transport invokes this so the node durably latches its reseed marker.
|
||
let _ = self
|
||
.snapshot_required_cell
|
||
.set(Arc::new(NodeSnapshotRequiredSink {
|
||
node: Arc::downgrade(self),
|
||
})
|
||
as Arc<dyn tidal_net::sources::SnapshotRequiredSink>);
|
||
|
||
// The positive-evidence counterpart: a COMPLETED catch-up pull is what
|
||
// discharges the marker the sink above latches. Bound here for the same
|
||
// Weak-back-reference reason.
|
||
let _ = self
|
||
.catchup_served_cell
|
||
.set(Arc::new(NodeCatchupServedSink {
|
||
node: Arc::downgrade(self),
|
||
})
|
||
as Arc<dyn tidal_net::sources::CatchupServedSink>);
|
||
|
||
// m11p5 §3.3: late-bind the `JoinCluster` adapter (same Weak-back-ref
|
||
// discipline). Until set, `JoinCluster` answers Unimplemented.
|
||
let _ = self.join_hooks_cell.set(Arc::new(NodeJoinHooks {
|
||
node: Arc::downgrade(self),
|
||
}) as Arc<dyn tidal_net::JoinHooks>);
|
||
}
|
||
|
||
/// Build the `GET /cluster/status/local` body.
|
||
fn local_status(&self) -> Result<LocalStatusResponse> {
|
||
let db = self.db()?;
|
||
let leader = self.current_leader();
|
||
let is_leader = leader == Some(self.region);
|
||
// Leaderless (mid-election): compute applied/lag against this node's
|
||
// own shard — a transient zero-lag self-view that the next heartbeat
|
||
// replaces. Honest fields below (`leader: null`, `role`) carry the
|
||
// real story.
|
||
let leader_shard = shard_of_region(leader.unwrap_or(self.region));
|
||
|
||
let last_seq = if is_leader {
|
||
self.ship_feed.flushed_seq()
|
||
} else {
|
||
0
|
||
};
|
||
// The leader does NOT apply its own stream through the receiver — it
|
||
// writes its WAL directly — so `applied_seqno(own_shard)` never advances
|
||
// on a leader and would read a stale follower-era value (0 on a node
|
||
// elected without first following). The leader's true applied frontier
|
||
// IS its durable flushed WAL frontier (it is, by construction, caught up
|
||
// to its own log). Followers read how far they have applied the CURRENT
|
||
// leader's per-source-shard stream (BUG 1: keyed by the live leader's
|
||
// shard, never a stale single scalar — see below).
|
||
let applied_events = if is_leader {
|
||
last_seq
|
||
} else {
|
||
self.applied_for_leader_shard(leader_shard)
|
||
};
|
||
// lag = the CURRENT leader's per-source-shard high-water-mark − applied
|
||
// for that same shard (BUG 1). The gauge tracks the leader HWM PER SOURCE
|
||
// SHARD, so a `/cluster/promote` that moves leadership to a different shard
|
||
// computes lag against the NEW leader's stream — not a stale single scalar
|
||
// that still holds a previously-followed leader's HWM (which left a
|
||
// converged node reporting a permanent phantom lag). A leader has zero lag
|
||
// against itself.
|
||
// This node's KNOWLEDGE of the leader's frontier for the group it follows.
|
||
// Surfaced (as `leader_seqno`) because `lag_events` alone is ambiguous: it
|
||
// is an unsigned subtraction against this gauge, so a freshly-booted node
|
||
// that has learned NOTHING reports `0 - 0 = 0` — indistinguishable from
|
||
// genuinely caught up. That ambiguity is what made the 2026-08-20 incident
|
||
// unreadable: every group answered `lag_events: 0` while holding
|
||
// `applied_events` of 24, 14 and 0 against 5600 written items. With this
|
||
// exposed, `lag_events: 0 AND leader_seqno: 0` reads as "no information",
|
||
// which is the truth.
|
||
let leader_seqno = if is_leader {
|
||
last_seq
|
||
} else {
|
||
db.control_plane()
|
||
.lag_gauge()
|
||
.leader_seqno_for(leader_shard)
|
||
};
|
||
let lag_events = if is_leader {
|
||
0
|
||
} else {
|
||
leader_seqno.saturating_sub(applied_events)
|
||
};
|
||
|
||
let partitioned: Vec<String> = read_recovered(&self.partitioned, "partitioned")
|
||
.iter()
|
||
.map(|r| self.region_name_of(*r).to_string())
|
||
.collect();
|
||
|
||
let commit_index = if is_leader {
|
||
if self.commit.needed_peers() == 0 {
|
||
last_seq
|
||
} else {
|
||
self.commit.committed()
|
||
}
|
||
} else {
|
||
0
|
||
};
|
||
|
||
let status_prev = *self
|
||
.activation_prev
|
||
.lock()
|
||
.unwrap_or_else(std::sync::PoisonError::into_inner);
|
||
// The LIVE election position the vote/join paths read (NOT the frozen
|
||
// `status_prev`): exposed for the divergence-consistency oracle.
|
||
let election_pos = self.election_log_position();
|
||
let (term, role, quarantined) = self.election_runtime.get().map_or_else(
|
||
|| (0, "unknown".to_string(), false),
|
||
|rt| {
|
||
let role = match rt.role() {
|
||
tidaldb::replication::Role::Leader => "leader",
|
||
tidaldb::replication::Role::Follower => "follower",
|
||
tidaldb::replication::Role::PreCandidate => "pre-candidate",
|
||
tidaldb::replication::Role::Candidate => "candidate",
|
||
};
|
||
(rt.current_term(), role.to_string(), rt.is_quarantined())
|
||
},
|
||
);
|
||
|
||
// m11p5 §4: feed the sticky readiness latch from the lag we just
|
||
// computed (no separate polling thread) and report the durable reseed
|
||
// state.
|
||
self.note_lag_for_readiness(leader_seqno, lag_events);
|
||
let reseed_required = self.reseed_marker_store.exists();
|
||
// `reseeding` = this node has not yet first-converged against a KNOWN
|
||
// leader frontier. No longer scoped to joiner boots: a plain restart is
|
||
// equally un-converged until it learns where the leader is, and reporting
|
||
// it as settled is what let a blind voter into the VIP.
|
||
let reseeding = !self.converged.load(Ordering::Acquire);
|
||
|
||
Ok(LocalStatusResponse {
|
||
region: self.region_name.clone(),
|
||
is_leader,
|
||
leader: leader.map(|l| self.region_name_of(l).to_string()),
|
||
last_seq,
|
||
applied_events,
|
||
lag_events,
|
||
partitioned,
|
||
commit_index,
|
||
ack: self.ack_default.as_str().to_string(),
|
||
reachable: true,
|
||
term,
|
||
role,
|
||
quarantined,
|
||
prev_log_term: status_prev.tail_term,
|
||
prev_log_seq: status_prev.frontier,
|
||
election_tail_term: election_pos.tail_term,
|
||
election_frontier: election_pos.frontier,
|
||
leader_acked: self.leader_acked_frontier(),
|
||
leader_seqno,
|
||
// Every tracked stream key, not just the current leader's. See
|
||
// `ReplicationState::applied_by_key` — a key retained from a PREVIOUS
|
||
// leadership is what the receiver's gap check can keep chasing, and it
|
||
// was invisible in every other field.
|
||
applied_by_key: db
|
||
.replication_state()
|
||
.applied_by_key()
|
||
.into_iter()
|
||
.map(|(k, v)| (u32::from(k.0), v))
|
||
.collect(),
|
||
reseed_required,
|
||
reseeding,
|
||
self_restart_refused: self.self_restart_refused.load(Ordering::Acquire),
|
||
membership_version: self.membership_version(),
|
||
membership_term: self.membership.term(),
|
||
membership_role: self.role_in_roster().to_string(),
|
||
version: node_build_version(),
|
||
// Populated by `ClusterNode::status_local` (it owns the group set);
|
||
// a bare per-replica status carries only its own row implicitly.
|
||
shards: Vec::new(),
|
||
})
|
||
}
|
||
}
|
||
|
||
/// This node's build version + hash for status/observability (m11p8):
|
||
/// `"<cargo-version>+<build-hash>"` (e.g. `"0.1.0+dev"`). All workspace crates
|
||
/// share the Cargo version, so the server version is the binary version.
|
||
fn node_build_version() -> String {
|
||
format!("{}+{}", env!("CARGO_PKG_VERSION"), tidaldb::BUILD_HASH)
|
||
}
|
||
|
||
impl Drop for ShardReplica {
|
||
fn drop(&mut self) {
|
||
self.shutdown();
|
||
}
|
||
}
|
||
|
||
/// Read a shared-state lock, recovering from (and WARN-logging) poison.
|
||
///
|
||
/// A poisoned lock means a thread panicked while holding it. Every value
|
||
/// guarded this way here (the leadership view, the ship-skip set) is written
|
||
/// by-value inside a short critical section, so the last-written state is
|
||
/// always internally consistent and recovery is safe — but the panic itself is
|
||
/// a bug, so the recovery must be visible in logs, never silent.
|
||
fn read_recovered<'a, T>(
|
||
lock: &'a RwLock<T>,
|
||
what: &'static str,
|
||
) -> std::sync::RwLockReadGuard<'a, T> {
|
||
lock.read().unwrap_or_else(|poisoned| {
|
||
tracing::warn!(
|
||
lock = what,
|
||
"lock poisoned by a panicked thread; recovering last-written state"
|
||
);
|
||
poisoned.into_inner()
|
||
})
|
||
}
|
||
|
||
/// Write-lock counterpart of [`read_recovered`].
|
||
fn write_recovered<'a, T>(
|
||
lock: &'a RwLock<T>,
|
||
what: &'static str,
|
||
) -> std::sync::RwLockWriteGuard<'a, T> {
|
||
lock.write().unwrap_or_else(|poisoned| {
|
||
tracing::warn!(
|
||
lock = what,
|
||
"lock poisoned by a panicked thread; recovering last-written state"
|
||
);
|
||
poisoned.into_inner()
|
||
})
|
||
}
|
||
|
||
/// The sibling-region address tables resolved from the topology for THIS region.
|
||
struct PeerTables {
|
||
/// Sibling shard ids (the relay's eager-ship peer list).
|
||
peer_shards: Vec<ShardId>,
|
||
/// Sibling shard id → ADVERTISED gRPC `host:port` (the transport peers).
|
||
///
|
||
/// Kept as a String, NOT a pre-resolved `SocketAddr` (m11p5): when the
|
||
/// advertised address is a DNS name, the tonic channel re-resolves it on
|
||
/// every reconnect, so a peer rescheduled onto a new IP becomes reachable
|
||
/// again without restarting THIS process. An IP literal re-resolves to
|
||
/// itself, so literal-IP topologies are unaffected.
|
||
peer_grpc: HashMap<ShardId, String>,
|
||
/// Sibling region id → public HTTP address (forwarding / aggregation).
|
||
peer_http: HashMap<RegionId, String>,
|
||
/// This region's own ADVERTISED `grpc_addr` spec (may be a DNS name); the
|
||
/// local bind is derived from it (and `my_grpc_bind_spec`) per the
|
||
/// bind/advertise split.
|
||
my_grpc_spec: Option<String>,
|
||
/// This region's own explicit `grpc_bind` spec, if any (a concrete
|
||
/// `SocketAddr`). `None` ⇒ the bind is derived from `my_grpc_spec`.
|
||
my_grpc_bind_spec: Option<String>,
|
||
/// This region's own TLS material (from the optional `grpc_tls` block);
|
||
/// `None` ⇒ plaintext transport.
|
||
my_tls: Option<tidal_net::config::TlsConfig>,
|
||
}
|
||
|
||
/// Resolve the address tables for THIS region's replica of one shard `group`
|
||
/// (m11p6). Peers are the group's OTHER replica nodes (not all topology
|
||
/// regions): their gRPC comes from the group's resolved replica addresses
|
||
/// (explicit or port-derived), their HTTP from each node's single
|
||
/// `RegionSpec.http_addr` (one HTTP server per node serves all its shards).
|
||
/// THIS replica's own gRPC advertise/bind comes from the group's resolved
|
||
/// self-replica; TLS from this node's `RegionSpec`.
|
||
///
|
||
/// In-group peer identity stays region-id based (`shard_of_region(rid)` =
|
||
/// `ShardId(rid.0)`), exactly as today's single-group cluster — the data-shard
|
||
/// `group.shard` only namespaces the directory/port, never the wire identity.
|
||
/// For the legacy single group (RF = all regions, port offset 0) this yields
|
||
/// byte-for-byte the pre-m11p6 tables.
|
||
fn build_group_peer_tables(
|
||
group: &ResolvedShardGroup,
|
||
topology: &TopologySpec,
|
||
region: RegionId,
|
||
) -> PeerTables {
|
||
let mut peer_shards = Vec::new();
|
||
let mut peer_grpc: HashMap<ShardId, String> = HashMap::new();
|
||
let mut peer_http: HashMap<RegionId, String> = HashMap::new();
|
||
let mut my_grpc_spec: Option<String> = None;
|
||
let mut my_grpc_bind_spec: Option<String> = None;
|
||
let mut my_tls: Option<tidal_net::config::TlsConfig> = None;
|
||
for replica in &group.replicas {
|
||
let rid = replica.region;
|
||
// This node's own TLS comes from its RegionSpec (one cert per node).
|
||
let region_spec = &topology.regions[usize::from(rid.0)];
|
||
if rid == region {
|
||
my_grpc_spec = Some(replica.grpc_addr.clone());
|
||
my_grpc_bind_spec.clone_from(&replica.grpc_bind);
|
||
my_tls = region_spec
|
||
.grpc_tls
|
||
.as_ref()
|
||
.map(super::topology::GrpcTlsSpec::to_tls_config);
|
||
continue;
|
||
}
|
||
peer_shards.push(shard_of_region(rid));
|
||
peer_grpc.insert(shard_of_region(rid), replica.grpc_addr.clone());
|
||
if let Some(http) = ®ion_spec.http_addr {
|
||
peer_http.insert(rid, http.clone());
|
||
}
|
||
}
|
||
PeerTables {
|
||
peer_shards,
|
||
peer_grpc,
|
||
peer_http,
|
||
my_grpc_spec,
|
||
my_grpc_bind_spec,
|
||
my_tls,
|
||
}
|
||
}
|
||
|
||
/// Open the one `TidalDb` a region node owns, applying the topology's m11p1
|
||
/// knobs: the optional `wal:` group-commit overrides and the per-region
|
||
/// `metrics_addr` (cluster mode previously had NO `/metrics` listener at all —
|
||
/// the roadmap's headline observability gap).
|
||
#[allow(clippy::too_many_arguments)]
|
||
fn open_region_db(
|
||
topology: &TopologySpec,
|
||
region_name: &str,
|
||
schema: Schema,
|
||
profiles: Vec<tidaldb::ranking::profile::RankingProfile>,
|
||
data_dir: std::path::PathBuf,
|
||
hlc_offset_ms: i64,
|
||
my_shard: ShardId,
|
||
peer_shards: &[ShardId],
|
||
enable_metrics: bool,
|
||
) -> Result<TidalDb> {
|
||
let mut builder = TidalDb::builder()
|
||
.with_schema(schema)
|
||
.with_profiles(profiles)
|
||
.with_cluster(NodeConfig {
|
||
role: NodeRole::Single,
|
||
shard_id: my_shard,
|
||
peer_shards: peer_shards.to_vec(),
|
||
..NodeConfig::default()
|
||
})
|
||
.with_hlc_offset_ms(hlc_offset_ms)
|
||
// Persistent by requirement (m11p2): the durable WAL is the
|
||
// replication stream; `Self::new` rejected a missing data dir.
|
||
.with_data_dir(data_dir);
|
||
if let Some(batch_size) = topology.wal.batch_size {
|
||
builder = builder.wal_batch_size(batch_size);
|
||
}
|
||
if let Some(timeout_ms) = topology.wal.batch_timeout_ms {
|
||
builder = builder.wal_batch_timeout(Duration::from_millis(timeout_ms));
|
||
}
|
||
// m11p8 PITR: archive sealed WAL segments before compaction deletes them.
|
||
// Segment filenames encode the shard id, so co-located groups share one
|
||
// archive dir without collision.
|
||
if let Some(ref archive_dir) = topology.wal.archive_dir {
|
||
builder = builder.wal_archive_dir(archive_dir);
|
||
}
|
||
// m11p6: at most ONE hosted shard per node binds the engine's `/metrics`
|
||
// server (N TidalDb instances would otherwise fight for one `metrics_addr`).
|
||
// ClusterNode designates the metrics-owning shard; the others open without.
|
||
if enable_metrics && let Some(metrics_addr) = topology.metrics_addr_of(region_name) {
|
||
builder = builder.enable_metrics(metrics_addr);
|
||
}
|
||
let db = builder.open().map_err(ServerError::Tidal)?;
|
||
if let Some(addr) = db.metrics_addr() {
|
||
// This shard owns the listener, so its MetricsState is the one actually
|
||
// scraped: register the process's HTTP series against it. Without this,
|
||
// cluster mode (the production shape) would expose the engine's domain
|
||
// metrics but nothing about requests, statuses, or HTTP latency.
|
||
crate::http_metrics::publish_to(db.metrics());
|
||
tracing::info!(
|
||
region = region_name,
|
||
"cluster metrics endpoint listening on http://{addr}/metrics"
|
||
);
|
||
}
|
||
Ok(db)
|
||
}
|
||
|
||
/// A staged signal write whose completion is guaranteed even if the HTTP
|
||
/// request that created it is cancelled (client disconnect, timeout layer).
|
||
///
|
||
/// A staged write's event is already submitted to the leader WAL — it WILL be
|
||
/// flushed, replicated, and replayed regardless of the caller. What
|
||
/// completion guarantees is the LEADER'S OWN in-memory fold: if nothing ever
|
||
/// calls `complete_signal_write`, the leader's aggregate silently misses an
|
||
/// event its durable log carries (and its followers apply) — a
|
||
/// leader-vs-follower divergence on the leader itself. Every axum await
|
||
/// point between staging and completion is a cancellation window, so the
|
||
/// guarantee cannot live in the handler: it lives here, in `Drop`.
|
||
///
|
||
/// The happy path calls [`complete`](Self::complete) (disarming the guard);
|
||
/// an orphaned ticket completes on a freshly spawned detached thread — rare
|
||
/// (only on mid-request cancellation), bounded (one group-commit wait), and
|
||
/// off the async runtime so `Drop` never blocks a reactor.
|
||
struct StagedWriteTicket {
|
||
staged: Option<StagedSignal>,
|
||
node: Arc<ShardReplica>,
|
||
}
|
||
|
||
impl StagedWriteTicket {
|
||
const fn new(staged: StagedSignal, node: Arc<ShardReplica>) -> Self {
|
||
Self {
|
||
staged: Some(staged),
|
||
node,
|
||
}
|
||
}
|
||
|
||
/// Complete the write on the current (blocking-pool) thread. Returns the
|
||
/// write's WAL seqno (`0` = dedup-suppressed).
|
||
fn complete(mut self) -> Result<u64> {
|
||
// `staged` is always `Some` until consumed here or in Drop; this is the
|
||
// only consuming method and it takes `self` by value.
|
||
let staged = self
|
||
.staged
|
||
.take()
|
||
.ok_or_else(|| ServerError::Cluster("staged write ticket already consumed".into()))?;
|
||
self.node.complete_signal_write(staged)
|
||
}
|
||
}
|
||
|
||
impl Drop for StagedWriteTicket {
|
||
fn drop(&mut self) {
|
||
let Some(staged) = self.staged.take() else {
|
||
return; // completed normally
|
||
};
|
||
let node = Arc::clone(&self.node);
|
||
tracing::warn!(
|
||
"request cancelled with a staged signal write in flight; \
|
||
completing it on a detached thread (the leader's aggregate must \
|
||
fold every event its WAL carries)"
|
||
);
|
||
let spawned = std::thread::Builder::new()
|
||
.name("tidal-orphan-complete".into())
|
||
.spawn(move || {
|
||
if let Err(e) = node.complete_signal_write(staged) {
|
||
tracing::error!(
|
||
error = %e,
|
||
"orphaned staged write failed to complete; if this is a \
|
||
durability failure the relay is poisoned (see runbook §8)"
|
||
);
|
||
}
|
||
});
|
||
if let Err(e) = spawned {
|
||
// Thread exhaustion at this point is an OOM-class host failure;
|
||
// the durable frontier may stall until restart. Say so loudly.
|
||
tracing::error!(
|
||
error = %e,
|
||
"could not spawn orphan-completion thread; the relay's durable \
|
||
frontier may stall until this node restarts"
|
||
);
|
||
}
|
||
}
|
||
}
|
||
|
||
/// Read the persisted stream baseline from the data dir (0 when absent or
|
||
/// unreadable — the fresh-cluster / never-promoted value).
|
||
fn load_stream_baseline(data_dir: &std::path::Path) -> u64 {
|
||
let path = data_dir.join(STREAM_BASELINE_FILE);
|
||
match std::fs::read(&path) {
|
||
Ok(bytes) if bytes.len() == 8 => {
|
||
let mut buf = [0u8; 8];
|
||
buf.copy_from_slice(&bytes);
|
||
u64::from_le_bytes(buf)
|
||
}
|
||
Ok(_) => {
|
||
tracing::error!(
|
||
path = %path.display(),
|
||
"stream baseline file is malformed; treating as 0 — if this \
|
||
node was promoted before, its catch-up stream may serve \
|
||
pre-stream history (re-promote to repair)"
|
||
);
|
||
0
|
||
}
|
||
Err(e) if e.kind() == std::io::ErrorKind::NotFound => 0,
|
||
Err(e) => {
|
||
tracing::error!(
|
||
path = %path.display(),
|
||
error = %e,
|
||
"stream baseline file unreadable; treating as 0"
|
||
);
|
||
0
|
||
}
|
||
}
|
||
}
|
||
|
||
/// Persist the stream baseline durably (write + fsync via a temp rename).
|
||
/// Best-effort with a loud error: a lost baseline only matters across a
|
||
/// restart-after-promote, and the operator can re-promote to repair.
|
||
fn persist_stream_baseline(data_dir: &std::path::Path, baseline: u64) {
|
||
let path = data_dir.join(STREAM_BASELINE_FILE);
|
||
let tmp = data_dir.join(format!("{STREAM_BASELINE_FILE}.tmp"));
|
||
let result = (|| -> std::io::Result<()> {
|
||
std::fs::write(&tmp, baseline.to_le_bytes())?;
|
||
let f = std::fs::File::open(&tmp)?;
|
||
f.sync_all()?;
|
||
std::fs::rename(&tmp, &path)?;
|
||
Ok(())
|
||
})();
|
||
if let Err(e) = result {
|
||
tracing::error!(
|
||
path = %path.display(),
|
||
baseline,
|
||
error = %e,
|
||
"failed to persist the stream baseline; a restart before the next \
|
||
promote may serve pre-stream history from the catch-up stream"
|
||
);
|
||
}
|
||
}
|
||
|
||
/// Read the persisted leader-acked frontier (m12 election-divergence-fix).
|
||
/// `None` when ABSENT (so the caller can pick a conservative fallback —
|
||
/// distinct from a persisted 0, which a clean join writes deliberately). A
|
||
/// malformed file is treated as absent (loud error → fallback).
|
||
fn load_leader_acked(data_dir: &std::path::Path) -> Option<u64> {
|
||
let path = data_dir.join(LEADER_ACKED_FILE);
|
||
match std::fs::read(&path) {
|
||
Ok(bytes) if bytes.len() == 8 => {
|
||
let mut buf = [0u8; 8];
|
||
buf.copy_from_slice(&bytes);
|
||
Some(u64::from_le_bytes(buf))
|
||
}
|
||
Ok(_) => {
|
||
tracing::error!(
|
||
path = %path.display(),
|
||
"leader-acked file is malformed; treating as absent (conservative \
|
||
fallback to the WAL tail)"
|
||
);
|
||
None
|
||
}
|
||
Err(e) if e.kind() == std::io::ErrorKind::NotFound => None,
|
||
Err(e) => {
|
||
tracing::error!(
|
||
path = %path.display(),
|
||
error = %e,
|
||
"leader-acked file unreadable; treating as absent"
|
||
);
|
||
None
|
||
}
|
||
}
|
||
}
|
||
|
||
/// Persist the leader-acked frontier durably (write + fsync via a temp rename),
|
||
/// mirroring [`persist_stream_baseline`]. Best-effort with a loud error: a lost
|
||
/// value falls back to the conservative WAL tail on the next boot.
|
||
fn persist_leader_acked(data_dir: &std::path::Path, value: u64) {
|
||
let path = data_dir.join(LEADER_ACKED_FILE);
|
||
let tmp = data_dir.join(format!("{LEADER_ACKED_FILE}.tmp"));
|
||
let result = (|| -> std::io::Result<()> {
|
||
std::fs::write(&tmp, value.to_le_bytes())?;
|
||
let f = std::fs::File::open(&tmp)?;
|
||
f.sync_all()?;
|
||
std::fs::rename(&tmp, &path)?;
|
||
Ok(())
|
||
})();
|
||
if let Err(e) = result {
|
||
tracing::error!(
|
||
path = %path.display(),
|
||
value,
|
||
error = %e,
|
||
"failed to persist the leader-acked frontier; the next boot falls back \
|
||
to the conservative WAL tail (may quarantine + reseed a clean node)"
|
||
);
|
||
}
|
||
}
|
||
|
||
/// The gRPC layer's applied-seqno reader (ack piggyback): a WEAK view over
|
||
/// this node's replication state, so the transport can never keep the
|
||
/// database alive past shutdown.
|
||
struct NodeAppliedSource {
|
||
db: Weak<TidalDb>,
|
||
}
|
||
|
||
impl AppliedSource for NodeAppliedSource {
|
||
fn applied_seqno(&self, source_shard: ShardId) -> u64 {
|
||
self.db.upgrade().map_or(0, |db| {
|
||
db.replication_state()
|
||
.applied_seqno(source_shard)
|
||
.unwrap_or(0)
|
||
})
|
||
}
|
||
}
|
||
|
||
/// The gRPC layer's WAL read-back source for `StreamSegments` (m11p2): serves
|
||
/// this node's durable batches to a pulling follower, clamped above the
|
||
/// stream baseline.
|
||
struct NodeSegmentSource {
|
||
db: Weak<TidalDb>,
|
||
shard: ShardId,
|
||
baseline: Arc<AtomicU64>,
|
||
feed: Arc<WalShipFeed>,
|
||
}
|
||
|
||
impl SegmentSource for NodeSegmentSource {
|
||
fn source_shard(&self) -> ShardId {
|
||
self.shard
|
||
}
|
||
|
||
fn stream_baseline(&self) -> u64 {
|
||
self.baseline.load(Ordering::Acquire)
|
||
}
|
||
|
||
fn flushed_seq(&self) -> u64 {
|
||
self.feed.flushed_seq()
|
||
}
|
||
|
||
fn collect_from(
|
||
&self,
|
||
from_seq: u64,
|
||
max_events: u64,
|
||
max_bytes: usize,
|
||
) -> std::result::Result<Vec<SegmentChunk>, SegmentReadError> {
|
||
let db = self.db.upgrade().ok_or_else(|| SegmentReadError::Failed {
|
||
detail: "database closed".to_string(),
|
||
})?;
|
||
let batches = db
|
||
.read_wal_batches(from_seq, max_events, max_bytes)
|
||
.map_err(|e| match e {
|
||
// This log can NEVER serve the range: its segments carry a
|
||
// format this binary cannot read (rolling-upgrade residue).
|
||
// The service maps this to FAILED_PRECONDITION ("snapshot
|
||
// required") so the follower knows retrying alone won't heal
|
||
// it (m11p4).
|
||
tidaldb::wal::error::WalError::SegmentFormatUnknown { .. } => {
|
||
SegmentReadError::Unavailable {
|
||
detail: e.to_string(),
|
||
}
|
||
}
|
||
other => SegmentReadError::Failed {
|
||
detail: other.to_string(),
|
||
},
|
||
})?;
|
||
if batches.is_empty() {
|
||
return Ok(Vec::new());
|
||
}
|
||
// Concatenate the contiguous batches into ONE chunk (already bounded
|
||
// by the caps): fewer stream messages, identical bytes.
|
||
let first_seq = batches[0].first_seq;
|
||
let last_seq = batches.last().map_or(first_seq, |b| b.last_seq);
|
||
let event_count: u64 = batches.iter().map(|b| b.event_count).sum();
|
||
let mut bytes = Vec::with_capacity(batches.iter().map(|b| b.bytes.len()).sum());
|
||
for batch in batches {
|
||
bytes.extend_from_slice(&batch.bytes);
|
||
}
|
||
Ok(vec![SegmentChunk {
|
||
bytes,
|
||
first_seq,
|
||
last_seq,
|
||
event_count,
|
||
}])
|
||
}
|
||
}
|
||
|
||
/// Build the node's two shared forwarding clients off the reactor.
|
||
///
|
||
/// Both build their own runtime internally and assert they are NOT inside one;
|
||
/// `ShardReplica::new` runs on a dedicated `std::thread`, so building them
|
||
/// here is sound. The async client carries the standard connect/request timeouts;
|
||
/// the blocking client (for the `/sharded/*` detached-thread fetch) carries NO
|
||
/// global timeout — the per-shard deadline is applied per request so a slow shard
|
||
/// degrades within budget.
|
||
/// `tls` (m11p7): when `Some`, the cluster CA is read from `tls.ca_cert` and added
|
||
/// to BOTH clients as a trust anchor so `https://` inter-node forwards/fetches
|
||
/// verify the peer's cluster-signed server cert. `None` ⇒ plaintext clients (the
|
||
/// trusted-loopback opt-out).
|
||
fn build_forwarding_clients(
|
||
tls: Option<&tidal_net::config::TlsConfig>,
|
||
) -> Result<(reqwest::Client, reqwest::blocking::Client)> {
|
||
let ca_pem = match tls {
|
||
Some(t) => Some(
|
||
std::fs::read(&t.ca_cert)
|
||
.map_err(|e| ServerError::Cluster(format!("read inter-node CA cert: {e}")))?,
|
||
),
|
||
None => None,
|
||
};
|
||
let client = forward::build_client(ca_pem.as_deref())
|
||
.map_err(|e| ServerError::Cluster(format!("build cluster forwarding client: {e}")))?;
|
||
let mut blocking =
|
||
reqwest::blocking::Client::builder().connect_timeout(forward::CONNECT_TIMEOUT);
|
||
if let Some(pem) = &ca_pem {
|
||
let ca = reqwest::Certificate::from_pem(pem)
|
||
.map_err(|e| ServerError::Cluster(format!("parse inter-node CA cert: {e}")))?;
|
||
blocking = blocking.add_root_certificate(ca);
|
||
}
|
||
let blocking_client = blocking
|
||
.build()
|
||
.map_err(|e| ServerError::Cluster(format!("build cluster sharded-fetch client: {e}")))?;
|
||
Ok((client, blocking_client))
|
||
}
|
||
|
||
// ── ClusterNode: the process handle hosting N shard-group replicas (m11p6) ───
|
||
|
||
/// The on-disk data subdir for shard group `shard` (zero-padded to 5 digits).
|
||
/// The ONE authoritative formatter for the per-group directory layout — a
|
||
/// durability contract any ops tooling (backup, `tidalctl`, the L3 rebalance
|
||
/// stager) must format identically to find an existing group's data.
|
||
///
|
||
/// `pub(crate)` so the boot-time reseed (`cluster::reseed::run_boot_install_for_region`)
|
||
/// resolves each hosted group's marker-bearing subdir via the SAME formatter —
|
||
/// the parent-dir-only install silently skipped per-group markers in S>1.
|
||
//
|
||
// `pub(crate)` is the deliberately-narrow, intent-expressing visibility for a
|
||
// crate-internal durability contract; the nursery `redundant_pub_crate` lint
|
||
// only flags it because the enclosing `node` module is private (so pub(crate)
|
||
// and pub reach identically) — keep the narrower, documented visibility.
|
||
#[allow(clippy::redundant_pub_crate)]
|
||
pub(crate) fn shard_subdir(shard: ShardId) -> String {
|
||
format!("shard-{:05}", shard.0)
|
||
}
|
||
|
||
/// Where an entity-routed write must go: a locally-hosted replica of the
|
||
/// entity's shard group, or a remote node hosting it (this node hosts no
|
||
/// replica of that group).
|
||
enum EntityRoute {
|
||
/// This node hosts the entity's shard group — apply/forward via this
|
||
/// replica's own leadership view (the existing per-group write path).
|
||
Local(Arc<ShardReplica>),
|
||
/// This node hosts no replica of the entity's group — forward to a replica
|
||
/// node, which routes to the group's leader.
|
||
Remote {
|
||
/// The target shard group (for the log/error).
|
||
shard: ShardId,
|
||
/// Ordered HTTP bases to try: the group's believed leader first, then
|
||
/// its other replicas (deduped). The gateway walks them on a CONNECT
|
||
/// failure so a single dead replica does not 503 a write the rest of the
|
||
/// group's quorum can still serve. Empty ⇒ no replica is addressable.
|
||
candidates: Vec<String>,
|
||
},
|
||
}
|
||
|
||
/// The cluster process node (m11p6).
|
||
///
|
||
/// One OS process hosting a replica of each shard group it is assigned to. It
|
||
/// owns the shared write/forward clients and node identity, holds one
|
||
/// [`ShardReplica`] per hosted group, and routes each request by entity hash to
|
||
/// the right group's leader. For the legacy single group (`shards:` absent) it
|
||
/// holds exactly one replica spanning every region — byte-for-byte the pre-m11p6
|
||
/// process.
|
||
pub struct ClusterNode {
|
||
/// Every shard group in the cluster, resolved (for routing a write to a
|
||
/// group this node does NOT host).
|
||
placement: BTreeMap<ShardId, ResolvedShardGroup>,
|
||
/// The shard groups THIS node hosts a replica of, keyed by data-shard id.
|
||
groups: BTreeMap<ShardId, Arc<ShardReplica>>,
|
||
/// Entity → data-shard router (`Single` for one group, `Hash(S)` otherwise).
|
||
router: ShardRouter,
|
||
/// Node region id → public HTTP base (cross-node forward to a shard leader).
|
||
node_http: HashMap<RegionId, String>,
|
||
/// Shared async forwarding client for the cross-shard gateway hop.
|
||
client: reqwest::Client,
|
||
/// m11p7 reloadable credentials: the bearer key (gates this node's HTTP
|
||
/// surface) AND the cluster key (mints/verifies per-node internal tokens).
|
||
/// Handlers reach it to mint a node token on every outbound forward/broadcast
|
||
/// and to resolve the request principal for the admin audit log.
|
||
creds: Arc<crate::cluster::security::ClusterCreds>,
|
||
/// m11p7: this node's gRPC TLS material (from the region's `grpc_tls` block),
|
||
/// derived once from the topology. The inter-node HTTP listener serves with
|
||
/// the SAME files. `None` ⇒ plaintext on both planes.
|
||
tls_files: Option<tidal_net::config::TlsConfig>,
|
||
/// m11p7 admin-verb audit sink (tracing + optional JSONL file).
|
||
audit: crate::cluster::audit::AuditSink,
|
||
/// m11p8: the metrics-owner shard's cluster-metrics handle — the cell this
|
||
/// node's `/metrics` listener renders. Node-level gateway events (cross-shard
|
||
/// write forwards) increment it so they surface on the single per-node scrape.
|
||
cluster_metrics: Arc<tidaldb::db::metrics::cluster::ClusterMetrics>,
|
||
/// Flipped on shutdown so `/health` reports not-ready while draining.
|
||
shutting_down: AtomicBool,
|
||
}
|
||
|
||
impl ClusterNode {
|
||
/// Build the cluster node: resolve the shard-group assignment, open one
|
||
/// [`ShardReplica`] per group this node hosts (each in its own data subdir
|
||
/// and gRPC port for `S > 1`; the node data dir verbatim for `S == 1`), and
|
||
/// wire the gateway router. Same argument shape as the pre-m11p6
|
||
/// `ShardReplica::new` so the boot paths and tests need no extra plumbing.
|
||
///
|
||
/// # Errors
|
||
///
|
||
/// Propagates [`ShardReplica::new`] failures, or [`ServerError`] when the
|
||
/// topology is invalid, this region hosts no group, or a per-group data
|
||
/// subdir cannot be created.
|
||
///
|
||
/// # Panics
|
||
///
|
||
/// Does not panic on caller input: the one internal `expect` on the region
|
||
/// name is guarded by [`validate_multiproc`], which runs first and proves
|
||
/// the name is declared.
|
||
#[allow(clippy::needless_pass_by_value, clippy::too_many_lines)] // schema/profiles cloned per group; m11p7 added creds/TLS/audit wiring
|
||
pub fn new(
|
||
topology: &TopologySpec,
|
||
region_name: &str,
|
||
schema: Schema,
|
||
profiles: Vec<tidaldb::ranking::profile::RankingProfile>,
|
||
data_dir: Option<std::path::PathBuf>,
|
||
hlc_offset_ms: i64,
|
||
) -> Result<Self> {
|
||
super::topology::validate_multiproc(topology, region_name)?;
|
||
|
||
let mut name_to_id = HashMap::new();
|
||
for (i, r) in topology.regions.iter().enumerate() {
|
||
let id = RegionId(u16::try_from(i).map_err(|_| {
|
||
ServerError::SchemaConfig("topology declares more than 65535 regions".into())
|
||
})?);
|
||
name_to_id.insert(r.name.clone(), id);
|
||
}
|
||
let region = *name_to_id
|
||
.get(region_name)
|
||
.expect("validate_multiproc proved region_name is declared");
|
||
|
||
let resolved = topology.resolve_shard_groups()?;
|
||
let single = resolved.len() == 1;
|
||
|
||
// m11p7: build the reloadable credentials ONCE and share the SAME instance
|
||
// with every hosted ShardReplica (forward/broadcast token minting) AND the
|
||
// node's router middleware + rotation poller, so a rotation is seen
|
||
// everywhere.
|
||
let creds = Arc::new(crate::cluster::security::ClusterCreds::from_env());
|
||
|
||
let mut groups: BTreeMap<ShardId, Arc<ShardReplica>> = BTreeMap::new();
|
||
let mut metrics_owner: Option<ShardId> = None;
|
||
for group in &resolved {
|
||
if !group.replicas.iter().any(|r| r.region == region) {
|
||
continue; // this node does not host this group
|
||
}
|
||
// S=1 uses the node data dir verbatim (existing clusters restart
|
||
// unchanged); S>1 isolates each group under `shard-<id>/`.
|
||
let group_dir = match &data_dir {
|
||
Some(d) if single => Some(d.clone()),
|
||
Some(d) => {
|
||
let sub = d.join(shard_subdir(group.shard));
|
||
std::fs::create_dir_all(&sub).map_err(|e| ServerError::io(&sub, e))?;
|
||
Some(sub)
|
||
}
|
||
None => None,
|
||
};
|
||
// At most ONE hosted group binds the engine's `/metrics` server
|
||
// (N TidalDb instances would otherwise fight for one `metrics_addr`).
|
||
// The other co-located groups register their cluster series with the
|
||
// owner below (m11p8) so the single per-node listener still exposes
|
||
// every group's replication metrics, `shard`-labeled.
|
||
let enable_metrics = metrics_owner.is_none();
|
||
let replica = ShardReplica::new(
|
||
topology,
|
||
region_name,
|
||
schema.clone(),
|
||
profiles.clone(),
|
||
group_dir,
|
||
hlc_offset_ms,
|
||
group,
|
||
enable_metrics,
|
||
!single,
|
||
Arc::clone(&creds),
|
||
)?;
|
||
if enable_metrics {
|
||
metrics_owner = Some(group.shard);
|
||
}
|
||
groups.insert(group.shard, Arc::new(replica));
|
||
}
|
||
if groups.is_empty() {
|
||
return Err(ServerError::Cluster(format!(
|
||
"region '{region_name}' is not a replica of any shard group"
|
||
)));
|
||
}
|
||
// m12 reseed-loop-fix (Fix 3): tell the node-level reseed-restart
|
||
// coordinator how many shard groups this process hosts, so a self-restart
|
||
// from one group fires the single process-wide exit only after every
|
||
// hosted group has also requested it or a bounded grace elapsed — never
|
||
// aborting a co-hosted sibling's in-flight install. A no-op on the S=1
|
||
// topology (one group → the gate fires immediately, today's behavior).
|
||
super::reseed_restart::register_hosted(groups.len());
|
||
// m11p8: when multiple shard groups co-locate on this node, the metrics
|
||
// owner exposes the siblings' `tidaldb_cluster_*` series under their own
|
||
// `shard="N"` label so one `/metrics` scrape covers every hosted group.
|
||
// No-op on the S=1 topology (a single hosted group is the owner).
|
||
if groups.len() > 1
|
||
&& let Some(owner_shard) = metrics_owner
|
||
&& let Some(owner_replica) = groups.get(&owner_shard)
|
||
&& let Ok(owner_db) = owner_replica.db()
|
||
{
|
||
for (shard, replica) in &groups {
|
||
if *shard == owner_shard {
|
||
continue;
|
||
}
|
||
if let Ok(sib_db) = replica.db() {
|
||
owner_db.register_metrics_sibling(shard.0, &sib_db);
|
||
}
|
||
}
|
||
}
|
||
|
||
let router = if single {
|
||
ShardRouter::single()
|
||
} else {
|
||
let n = u16::try_from(resolved.len())
|
||
.map_err(|_| ServerError::SchemaConfig("more than 65535 shard groups".into()))?;
|
||
ShardRouter::hash(n).map_err(|e| {
|
||
ServerError::Cluster(format!("build shard router for {n} groups: {e}"))
|
||
})?
|
||
};
|
||
|
||
// Node id → public HTTP base for the cross-node forward. Reuses the
|
||
// already-validated `name_to_id` (the one overflow-checked region→id
|
||
// map) rather than re-deriving the positional id a second time.
|
||
let mut node_http = HashMap::new();
|
||
for r in &topology.regions {
|
||
if let Some(h) = &r.http_addr {
|
||
node_http.insert(name_to_id[&r.name], h.clone());
|
||
}
|
||
}
|
||
let placement: BTreeMap<ShardId, ResolvedShardGroup> =
|
||
resolved.into_iter().map(|g| (g.shard, g)).collect();
|
||
// The router's shard space and the placement map are two derivations of
|
||
// the same resolved set; `resolve_shard_groups` guarantees dense ids in
|
||
// `[0, len)`, so every `route()` output is a placement key. Assert it so
|
||
// an L3 change that mutates the group set (split/merge) cannot let
|
||
// `route()` return a ShardId absent from `placement`.
|
||
debug_assert!(
|
||
router
|
||
.all_shards()
|
||
.iter()
|
||
.all(|s| placement.contains_key(s)),
|
||
"router shard space must be covered by placement"
|
||
);
|
||
// m11p7: this node's TLS files (the region's `grpc_tls` block, derived the
|
||
// same way the gRPC transport did). The inter-node HTTP listener serves
|
||
// with the SAME cert. Works for both the declared topology and a
|
||
// seed-join's synthesized one (the joiner's grpc_tls rides §3.5 knobs).
|
||
let tls_files = topology
|
||
.regions
|
||
.iter()
|
||
.find(|r| r.name == region_name)
|
||
.and_then(|r| r.grpc_tls.as_ref())
|
||
.map(super::topology::GrpcTlsSpec::to_tls_config);
|
||
|
||
// m11p7: declare the inter-node HTTP scheme for this process. With TLS
|
||
// configured, every `peer_url` (forward/broadcast/scatter/status/seed)
|
||
// emits `https://` and the forwarding clients trust the cluster CA.
|
||
super::forward::set_inter_node_https(tls_files.is_some());
|
||
let (client, _blocking) = build_forwarding_clients(tls_files.as_ref())?;
|
||
|
||
tracing::info!(
|
||
region = region_name,
|
||
hosted_groups = groups.len(),
|
||
total_groups = placement.len(),
|
||
tls = tls_files.is_some(),
|
||
version = %node_build_version(),
|
||
"cluster node started (m11p6: one replica per hosted shard group)"
|
||
);
|
||
|
||
// m11p8: the metrics-owner's cluster-metrics handle (the cell `/metrics`
|
||
// renders) for node-level gateway events. Falls back to the lowest-shard
|
||
// group if no owner was flagged; `groups` is non-empty (checked above).
|
||
let cluster_metrics = metrics_owner
|
||
.and_then(|s| groups.get(&s))
|
||
.or_else(|| groups.values().next())
|
||
.map(|r| Arc::clone(&r.cluster_metrics))
|
||
.expect("groups is non-empty");
|
||
|
||
Ok(Self {
|
||
placement,
|
||
groups,
|
||
router,
|
||
node_http,
|
||
client,
|
||
creds,
|
||
tls_files,
|
||
audit: crate::cluster::audit::AuditSink::from_env(),
|
||
cluster_metrics,
|
||
shutting_down: AtomicBool::new(false),
|
||
})
|
||
}
|
||
|
||
/// Record an admin-verb audit event (m11p7) on the operator-originated leg.
|
||
/// A no-op on an internal forwarded re-apply (the operator leg recorded it).
|
||
fn audit_admin(
|
||
&self,
|
||
headers: &HeaderMap,
|
||
verb: &str,
|
||
target: &str,
|
||
term: u64,
|
||
result: &std::result::Result<Response, ClusterAppError>,
|
||
) {
|
||
if crate::cluster::audit::is_operator_request(headers) {
|
||
self.audit_admin_outcome(
|
||
headers,
|
||
verb,
|
||
target,
|
||
term,
|
||
&crate::cluster::audit::outcome_of(result),
|
||
);
|
||
}
|
||
}
|
||
|
||
/// Record an admin-verb audit event with an already-resolved outcome string
|
||
/// (m11p7). The caller has established the operator leg; this resolves the
|
||
/// principal and emits the record. Used by `cluster_promote`, whose
|
||
/// `Json`-typed result does not fit [`audit::outcome_of`].
|
||
fn audit_admin_outcome(
|
||
&self,
|
||
headers: &HeaderMap,
|
||
verb: &str,
|
||
target: &str,
|
||
term: u64,
|
||
outcome: &str,
|
||
) {
|
||
let principal = crate::cluster::audit::principal_of(&self.creds, headers);
|
||
self.audit.record(&principal, verb, target, term, outcome);
|
||
}
|
||
|
||
/// The node's reloadable credentials (m11p7). Shared (cheap `Arc` clone) with
|
||
/// the router's auth layer + the rotation poller so all read the same live
|
||
/// bearer/cluster keys. Public so the binary's `serve_state` passes this one
|
||
/// instance to both the router middleware and the rotation poller.
|
||
#[must_use]
|
||
pub fn creds(&self) -> Arc<crate::cluster::security::ClusterCreds> {
|
||
Arc::clone(&self.creds)
|
||
}
|
||
|
||
/// This node's gRPC TLS material (m11p7), shared across every hosted group
|
||
/// (one cert per node). The inter-node HTTP listener serves with the SAME
|
||
/// files, so one rotation covers both planes. `None` ⇒ plaintext (the
|
||
/// trusted-loopback opt-out): the HTTP listener then stays plaintext too.
|
||
/// Public so the binary's `serve_state` can derive the HTTP TLS material.
|
||
#[must_use]
|
||
pub fn grpc_tls_files(&self) -> Option<tidal_net::config::TlsConfig> {
|
||
self.tls_files.clone()
|
||
}
|
||
|
||
/// Start the per-group election driver on every hosted replica (called once
|
||
/// the node is in its final `Arc`, like the single-node `started` hook).
|
||
pub fn start_all_elections(&self) {
|
||
for replica in self.groups.values() {
|
||
replica.start_election_driver();
|
||
}
|
||
}
|
||
|
||
/// Route an entity-scoped write to the replica that owns its shard group.
|
||
fn route_entity(&self, entity_id: u64) -> EntityRoute {
|
||
let shard = self.router.route(EntityId::new(entity_id));
|
||
if let Some(replica) = self.groups.get(&shard) {
|
||
return EntityRoute::Local(Arc::clone(replica));
|
||
}
|
||
EntityRoute::Remote {
|
||
shard,
|
||
candidates: self.forward_candidates(shard),
|
||
}
|
||
}
|
||
|
||
/// Ordered HTTP bases to forward a write for a group this node does NOT host:
|
||
/// the group's believed leader first (placement's term-0/preferred leader),
|
||
/// then its other replicas, deduped and skipping any without a known HTTP
|
||
/// base. The receiver re-routes to the CURRENT leader, so any live replica
|
||
/// suffices — leader-first just minimizes the extra hop, and the ordered list
|
||
/// lets the gateway fail over on a connect error instead of pinning a write
|
||
/// to one dead replica.
|
||
fn forward_candidates(&self, shard: ShardId) -> Vec<String> {
|
||
let Some(group) = self.placement.get(&shard) else {
|
||
return Vec::new();
|
||
};
|
||
let mut seen = HashSet::new();
|
||
std::iter::once(group.leader)
|
||
.chain(group.replicas.iter().map(|r| r.region))
|
||
.filter(|r| seen.insert(*r))
|
||
.filter_map(|r| self.node_http.get(&r).cloned())
|
||
.collect()
|
||
}
|
||
|
||
/// Resolve the [`ShardReplica`] for an admin/read/status surface. `None`
|
||
/// returns the FIRST hosted group — exact for `S=1` (the sole group), and the
|
||
/// gateway's default for surfaces that do not yet take a `?shard=` selector
|
||
/// (the L3 per-shard admin work passes `Some(shard)`). `Some(s)` returns the
|
||
/// keyed group or a 400 if this node hosts no replica of it. This is the one
|
||
/// seam the deferred `?shard=` selector threads through — handlers never reach
|
||
/// into `self.groups` directly.
|
||
fn replica_for(
|
||
&self,
|
||
shard: Option<ShardId>,
|
||
) -> std::result::Result<Arc<ShardReplica>, ClusterAppError> {
|
||
if let Some(s) = shard {
|
||
return self.groups.get(&s).cloned().ok_or_else(|| {
|
||
ClusterAppError(ServerError::BadRequest(format!(
|
||
"shard {} is not hosted by this node",
|
||
s.0
|
||
)))
|
||
});
|
||
}
|
||
self.groups.values().next().cloned().ok_or_else(|| {
|
||
ClusterAppError(ServerError::Cluster(
|
||
"node hosts no shard group (new() rejects empty)".into(),
|
||
))
|
||
})
|
||
}
|
||
|
||
/// Borrow every hosted replica (status aggregation, health, shutdown).
|
||
fn hosted(&self) -> impl Iterator<Item = &Arc<ShardReplica>> {
|
||
self.groups.values()
|
||
}
|
||
|
||
/// The `TidalDb` of every hosted group, for an in-process read scatter
|
||
/// (m11p6: each group holds a disjoint entity subset, so a corpus-wide
|
||
/// `/feed`//`/search` queries every local group and merges). A group that is
|
||
/// shutting down is skipped (its `db_arc` errors). For `S=1` this is the one
|
||
/// db — byte-for-byte today's single read.
|
||
fn hosted_dbs(&self) -> Vec<Arc<TidalDb>> {
|
||
self.groups
|
||
.values()
|
||
.filter_map(|r| r.db_arc().ok())
|
||
.collect()
|
||
}
|
||
|
||
/// A per-shard status row for every hosted group (m11p6 status surface).
|
||
fn shard_status_rows(&self) -> Vec<ShardStatusRow> {
|
||
self.groups
|
||
.values()
|
||
.filter_map(|r| {
|
||
let s = r.local_status().ok()?;
|
||
Some(ShardStatusRow {
|
||
shard: r.group_shard.0,
|
||
is_leader: s.is_leader,
|
||
leader: s.leader,
|
||
term: s.term,
|
||
role: s.role,
|
||
applied_events: s.applied_events,
|
||
leader_seqno: s.leader_seqno,
|
||
lag_events: s.lag_events,
|
||
commit_index: s.commit_index,
|
||
// PER-GROUP reseed state. The flat `reseed_required` /
|
||
// `reseeding` fields above describe ONE group — whichever
|
||
// `replica_for(None)` resolves (the lowest id) — so on a
|
||
// multi-group node they say nothing about the others. That is
|
||
// how the 2026-08-20 incident hid: tidaldb-0 answered
|
||
// `reseed_required: false, lag_events: 0` while a different
|
||
// hosted group was stuck behind a compacted leader, and every
|
||
// operator (and every diagnosis) read it as converged.
|
||
reseed_required: s.reseed_required,
|
||
reseeding: s.reseeding,
|
||
})
|
||
})
|
||
.collect()
|
||
}
|
||
|
||
/// Resolve an entity write to its local replica, or forward it to a node
|
||
/// hosting the group. The single gateway-routing seam for every entity-write
|
||
/// handler: `Continue(replica)` ⇒ apply on the local per-group path,
|
||
/// `Break(resp)` ⇒ the relayed cross-node response to return verbatim.
|
||
async fn route_or_forward<B: serde::Serialize + Sync + ?Sized>(
|
||
&self,
|
||
key: u64,
|
||
path: &str,
|
||
body: &B,
|
||
headers: &HeaderMap,
|
||
) -> std::ops::ControlFlow<std::result::Result<Response, ClusterAppError>, Arc<ShardReplica>>
|
||
{
|
||
match self.route_entity(key) {
|
||
EntityRoute::Local(replica) => std::ops::ControlFlow::Continue(replica),
|
||
EntityRoute::Remote { shard, candidates } => std::ops::ControlFlow::Break(
|
||
self.forward_to_group_node(&candidates, shard, path, body, headers)
|
||
.await,
|
||
),
|
||
}
|
||
}
|
||
|
||
/// Forward an entity write to a remote node that hosts its group (this node
|
||
/// hosts no replica). `internal=false` so the receiver re-routes to its
|
||
/// group's CURRENT leader; the caller's `x-tidal-ack` and the verdict headers
|
||
/// relay. Walks `candidates` (leader-first) on a CONNECT failure so one dead
|
||
/// replica does not 503 a write the group's surviving quorum can serve; a
|
||
/// real verdict (a 2xx/NotLeader/QuorumTimeout STATUS) comes back as `Ok` and
|
||
/// is relayed immediately, never retried.
|
||
async fn forward_to_group_node<B: serde::Serialize + Sync + ?Sized>(
|
||
&self,
|
||
candidates: &[String],
|
||
shard: ShardId,
|
||
path: &str,
|
||
body: &B,
|
||
headers: &HeaderMap,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
if candidates.is_empty() {
|
||
return Err(ClusterAppError(ServerError::Unavailable(format!(
|
||
"shard {} has no reachable replica to route the write to",
|
||
shard.0
|
||
))));
|
||
}
|
||
let auth = forwarded_auth(headers);
|
||
let passthrough = forward::ack_passthrough(headers);
|
||
self.cluster_metrics.incr_forwards();
|
||
let mut last_err = String::new();
|
||
for http in candidates {
|
||
let url = peer_url(http, path);
|
||
match forward_json_with_headers(
|
||
&self.client,
|
||
&url,
|
||
body,
|
||
auth.as_deref(),
|
||
false,
|
||
&passthrough,
|
||
)
|
||
.await
|
||
{
|
||
// A real verdict (status received): relay it, do not try another
|
||
// replica. The receiver already re-routed to its current leader.
|
||
Ok(resp) => return Ok(forward::relay_forwarded(resp)),
|
||
// A connect/transport failure: this replica is unreachable —
|
||
// try the next candidate (the group's quorum may still be live).
|
||
Err(e) => last_err = format!("{url}: {e}"),
|
||
}
|
||
}
|
||
// Every candidate replica was unreachable — the forward could not be
|
||
// delivered to the group at all (m11p8 forward-failure signal).
|
||
self.cluster_metrics.incr_forward_failures();
|
||
Err(ClusterAppError(ServerError::Unavailable(format!(
|
||
"shard {} unreachable: all {} replica candidates failed (last {last_err})",
|
||
shard.0,
|
||
candidates.len()
|
||
))))
|
||
}
|
||
|
||
/// Ready iff not draining AND every hosted group is ready (m11p6: a node
|
||
/// serves traffic only when all its shard replicas can).
|
||
fn is_ready(&self) -> bool {
|
||
!self.shutting_down.load(Ordering::Acquire) && self.hosted().all(|r| r.is_ready())
|
||
}
|
||
|
||
/// Flip `/health` to not-ready and propagate to every hosted replica.
|
||
pub fn set_shutting_down(&self) {
|
||
self.shutting_down.store(true, Ordering::Release);
|
||
for replica in self.hosted() {
|
||
replica.set_shutting_down();
|
||
}
|
||
}
|
||
|
||
/// True once shutdown began.
|
||
#[must_use]
|
||
pub fn is_shutting_down(&self) -> bool {
|
||
self.shutting_down.load(Ordering::Acquire)
|
||
}
|
||
|
||
/// Deterministically shut down every hosted replica (checkpoint, WAL fsync,
|
||
/// HNSW-graph checkpoint, and thread join per group).
|
||
///
|
||
/// Takes `&self` (m12p6): each [`ShardReplica::shutdown`] is itself `&self`
|
||
/// (its db is an `ArcSwapOption`), and idempotent, so we drive the close
|
||
/// through the shared `Arc<ShardReplica>` directly — no `try_unwrap`. This is
|
||
/// what lets `serve_state` run the full deterministic close even when a stuck
|
||
/// peer connection has left an `Arc<ClusterNode>` alive past the graceful
|
||
/// drain (the SIGTERM path that previously skipped the HNSW-graph save). The
|
||
/// election drivers and transport sources hold only `Weak`, so the
|
||
/// per-replica `swap(None)` drop is the last `Arc<TidalDb>` reference once the
|
||
/// request-scoped clones have drained, and the checkpoint runs synchronously.
|
||
pub fn shutdown(&self) {
|
||
self.shutting_down.store(true, Ordering::Release);
|
||
let hosted: Vec<&Arc<ShardReplica>> = self.hosted().collect();
|
||
// m12p6: when this node hosts MULTIPLE shard groups, save their HNSW
|
||
// graphs CONCURRENTLY. Each `ShardReplica::shutdown` serializes + fsyncs
|
||
// its slot graphs (~100 MB+ per slot at the 1536-dim production shape);
|
||
// run serially across S groups that is S × the slowest save, which on a
|
||
// 3-shard node routinely overran the SIGTERM grace window → k8s SIGKILL
|
||
// mid-save → the next boot rebuilt instead of loading. Each replica's
|
||
// close is `&self` and touches only its own db/shard, so the saves are
|
||
// independent and safe to run in parallel; `thread::scope` joins them all
|
||
// before returning (the process must not exit until every graph is durable).
|
||
if hosted.len() <= 1 {
|
||
for replica in &hosted {
|
||
replica.shutdown();
|
||
}
|
||
} else {
|
||
std::thread::scope(|s| {
|
||
for replica in &hosted {
|
||
let replica = Arc::clone(replica);
|
||
s.spawn(move || replica.shutdown());
|
||
}
|
||
});
|
||
}
|
||
}
|
||
}
|
||
|
||
// ── Router ──────────────────────────────────────────────────────────────────
|
||
|
||
/// Build the multi-process region router.
|
||
///
|
||
/// Same public/protected split, load-shedding stack, and shared constants as
|
||
/// [`build_cluster_router`](super::build_cluster_router). Task 03 makes the node
|
||
/// a coherent cluster gateway: writes forward to the leader, items/embeddings
|
||
/// broadcast leader→peers, `/cluster/status` aggregates every region, the
|
||
/// `/cluster/reconcile*` pair exchanges CRDT snapshots, and `/sharded/*` fans
|
||
/// out across processes.
|
||
pub fn build_region_router(
|
||
node: Arc<ClusterNode>,
|
||
creds: Arc<crate::cluster::security::ClusterCreds>,
|
||
) -> Router {
|
||
// ONLY the probe contract and the OpenAPI document are unauthenticated.
|
||
//
|
||
// `/cluster/status` and `/cluster/status/local` used to live here. They report
|
||
// leader identity, membership, term, and per-shard applied/lag/commit seqnos,
|
||
// which is reconnaissance rather than a probe, so they moved to `protected`.
|
||
// Every internal caller already authenticates: seed-join and reseed leader
|
||
// discovery and the status fan-out all send the bearer (via
|
||
// `security::bearer_from_env`, which honours the `*_FILE` form too).
|
||
let public = Router::new()
|
||
.route("/health", get(region_health))
|
||
.route("/health/startup", get(crate::health::health_startup))
|
||
.route("/health/live", get(crate::health::health_live))
|
||
.route("/openapi.json", get(crate::openapi::serve_region))
|
||
.with_state(Arc::clone(&node));
|
||
|
||
// The DESTRUCTIVE operator verbs. Separated from the data surface because
|
||
// they used to share its credential: one bearer let any client key remove a
|
||
// member, force a partition, or move a shard. `admin_gate` requires the
|
||
// admin key (or a verified sibling) once one is configured.
|
||
//
|
||
// Peer-callable verbs deliberately DO NOT live here - `/cluster/catchup`
|
||
// (self-heal nudge), `/cluster/join` + `/cluster/members` (seed-join), and
|
||
// the `/cluster/reconcile*` pair are all dialled node-to-node with the plain
|
||
// bearer, so gating them on the admin key would break replication and
|
||
// joining.
|
||
let admin_creds = Arc::clone(&creds);
|
||
let admin = Router::new()
|
||
.route("/cluster/promote", post(cluster_promote))
|
||
.route("/cluster/partition", post(cluster_partition))
|
||
.route("/cluster/heal", post(cluster_heal))
|
||
.route("/cluster/reseed", post(cluster_reseed))
|
||
.route("/cluster/members/remove", post(cluster_member_remove))
|
||
// m11p6 L3 rebalancing verbs (per-group, reusing the m11p5 machinery).
|
||
.route("/cluster/shards/{id}/replicas", post(shard_replicas))
|
||
.route("/cluster/shards/{id}/transfer", post(shard_transfer))
|
||
.layer(middleware::from_fn(move |req: Request, next: Next| {
|
||
admin_gate(Arc::clone(&admin_creds), req, next)
|
||
}))
|
||
.with_state(Arc::clone(&node));
|
||
|
||
let protected = Router::new()
|
||
.route("/items", post(create_item))
|
||
.route("/embeddings", post(write_embedding))
|
||
.route("/signals", post(write_signal))
|
||
.route("/hardnegs", post(write_hardneg))
|
||
.route("/feed", get(feed))
|
||
.route("/search", get(search))
|
||
.route("/vector_search", post(vector_search))
|
||
.route("/cluster/catchup", post(cluster_catchup))
|
||
.route("/cluster/reconcile", post(cluster_reconcile))
|
||
.route("/cluster/members", get(cluster_members))
|
||
.route("/cluster/join", post(cluster_join))
|
||
// Moved out of `public`: topology reconnaissance, not a probe.
|
||
.route("/cluster/status/local", get(status_local))
|
||
.route("/cluster/status", get(cluster_status))
|
||
// The snapshot body is corpus-sized, so it gets its own cap BEFORE the
|
||
// group's data-surface limit applies (an inner layer wins). See
|
||
// `RECONCILE_BODY_LIMIT_BYTES` for why the shared 2 MiB made
|
||
// divergence unhealable on the live cluster.
|
||
.route(
|
||
"/cluster/reconcile/snapshot",
|
||
post(cluster_reconcile_snapshot).layer(axum::extract::DefaultBodyLimit::max(
|
||
crate::router::RECONCILE_BODY_LIMIT_BYTES,
|
||
)),
|
||
)
|
||
.route("/sharded/items", post(sharded_create_item))
|
||
.route("/sharded/embeddings", post(sharded_write_embedding))
|
||
.route("/sharded/signals", post(sharded_write_signal))
|
||
.route("/sharded/feed", get(sharded_feed))
|
||
.route("/sharded/search", get(sharded_search))
|
||
.with_state(node)
|
||
.merge(admin)
|
||
.layer(axum::extract::DefaultBodyLimit::max(
|
||
crate::router::BODY_LIMIT_BYTES,
|
||
));
|
||
|
||
// m11p7 cluster auth, in one layer so `next` runs at most once. Extracted to
|
||
// `cluster_auth_middleware` so the assembled-router composition (bearer ->
|
||
// marker-pinning -> rate limit) is exercised by tests, not just the pure
|
||
// `ClusterCreds` helpers.
|
||
let protected = protected.layer(middleware::from_fn(move |req: Request, next: Next| {
|
||
cluster_auth_middleware(Arc::clone(&creds), req, next)
|
||
}));
|
||
|
||
let protected = protected.layer(
|
||
ServiceBuilder::new()
|
||
.layer(TimeoutLayer::with_status_code(
|
||
StatusCode::REQUEST_TIMEOUT,
|
||
Duration::from_secs(crate::router::REQUEST_TIMEOUT_SECS),
|
||
))
|
||
.layer(ConcurrencyLimitLayer::new(crate::router::MAX_CONCURRENCY)),
|
||
);
|
||
|
||
// m11p8: assign/echo `x-request-id` and open a per-request span on the
|
||
// multi-process region router too. Forwarded writes carry the originating
|
||
// gateway's id across the leader hop (see `cluster::forward`).
|
||
crate::router::with_request_id_tracing(public.merge(protected))
|
||
}
|
||
|
||
/// The m11p7 cluster-auth middleware, applied to every protected region route in
|
||
/// one layer so `next` runs at most once. Three gates in order:
|
||
///
|
||
/// 1. **Bearer**, read PER REQUEST from `creds` (rotatable; open when unset).
|
||
/// 2. **Marker-pinning**: when a cluster key is configured, a request that sets a
|
||
/// sibling-only marker (`x-tidal-internal` or `x-tidal-relayed`) WITHOUT a
|
||
/// valid node token is rejected (403). The markers stay routing/audit hints —
|
||
/// only a verified sibling may set them — so neither is a standalone bypass.
|
||
/// 3. **Per-principal rate limit**: verified sibling NODES are exempt (replication
|
||
/// must not be throttled); external principals consume their bucket → 429.
|
||
///
|
||
/// Extracted from [`build_region_router`] so this composition is testable through
|
||
/// the real layer (a layer-order regression would otherwise compile and pass the
|
||
/// pure-`ClusterCreds` unit tests while reopening the marker bypass).
|
||
async fn cluster_auth_middleware(
|
||
creds: Arc<crate::cluster::security::ClusterCreds>,
|
||
req: Request,
|
||
next: Next,
|
||
) -> Response {
|
||
if !creds.authenticated(req.headers()) {
|
||
return crate::router::unauthorized_response(req.headers());
|
||
}
|
||
let marked =
|
||
super::forward::is_internal(req.headers()) || super::forward::is_relayed(req.headers());
|
||
if creds.marker_without_node_identity(req.headers(), marked) {
|
||
return (
|
||
StatusCode::FORBIDDEN,
|
||
Json(serde_json::json!({
|
||
"error": "x-tidal-internal / x-tidal-relayed require a valid \
|
||
x-tidal-node-token; these markers are honored only from a \
|
||
verified cluster sibling"
|
||
})),
|
||
)
|
||
.into_response();
|
||
}
|
||
let principal = creds.principal(req.headers());
|
||
if let Err((retry_after_ms, limit)) = creds.check_rate(&principal) {
|
||
return crate::router::too_many_requests(retry_after_ms, limit);
|
||
}
|
||
next.run(req).await
|
||
}
|
||
|
||
/// Authorization gate for the DESTRUCTIVE cluster verbs, layered INSIDE
|
||
/// [`cluster_auth_middleware`] so the bearer/marker/rate gates run first and this
|
||
/// only decides authority.
|
||
///
|
||
/// Rejects with 403 (authenticated but not permitted) rather than 401 — the
|
||
/// caller's bearer was valid, it simply is not an operator credential. Passes
|
||
/// through untouched when no admin key is configured, so an existing deployment
|
||
/// behaves exactly as before (with a startup WARN from `ClusterCreds::from_env`).
|
||
async fn admin_gate(
|
||
creds: Arc<crate::cluster::security::ClusterCreds>,
|
||
req: Request,
|
||
next: Next,
|
||
) -> Response {
|
||
if creds.admin_ok(req.headers()) {
|
||
return next.run(req).await;
|
||
}
|
||
crate::router::admin_forbidden_response()
|
||
}
|
||
|
||
// ── Health ──────────────────────────────────────────────────────────────────
|
||
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
async fn region_health(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
) -> std::result::Result<(StatusCode, Json<serde_json::Value>), ClusterAppError> {
|
||
let state = node.replica_for(None)?;
|
||
// m11p5 §4 readiness: 503 while shutting down, quarantined, or an install
|
||
// boot has not yet first-converged (sticky-ready after). m11p6: the node is
|
||
// ready only when EVERY hosted shard group is ready (aggregate).
|
||
if !node.is_ready() {
|
||
// Inspect the group ACTUALLY keeping the node unready (S>1: not
|
||
// necessarily the first hosted group) so the 503 cause is accurate;
|
||
// fall back to the default group if only the node-level drain flag is set.
|
||
let unready = node
|
||
.hosted()
|
||
.find(|r| !r.is_ready())
|
||
.cloned()
|
||
.unwrap_or_else(|| Arc::clone(&state));
|
||
let cause = if unready.is_shutting_down() {
|
||
"shutting down"
|
||
} else if unready
|
||
.election_runtime
|
||
.get()
|
||
.is_some_and(|rt| rt.is_quarantined())
|
||
{
|
||
"quarantined (divergent suffix); reseeds on next boot"
|
||
} else if unready.membership().self_role()
|
||
== Some(tidaldb::wal::format::MemberRole::Removed)
|
||
{
|
||
"removed from the cluster (decommissioned)"
|
||
} else {
|
||
"joiner boot not yet converged"
|
||
};
|
||
return Ok((
|
||
StatusCode::SERVICE_UNAVAILABLE,
|
||
Json(serde_json::json!({
|
||
"ok": false,
|
||
"service": "tidaldb",
|
||
"cause": cause,
|
||
})),
|
||
));
|
||
}
|
||
let leader = state.current_leader();
|
||
Ok((
|
||
StatusCode::OK,
|
||
Json(serde_json::json!({
|
||
"ok": true,
|
||
"service": "tidaldb",
|
||
"mode": "cluster",
|
||
"process": "single-region",
|
||
"region": state.region_name,
|
||
"leader": leader.map(|l| state.region_name_of(l).to_string()),
|
||
})),
|
||
))
|
||
}
|
||
|
||
/// A process-wide "graceful self-exit requested" latch (m11p5 §2.4 reseed
|
||
/// self-restart). The serve loop's [`crate::self_exit::wait_self_exit`] selects
|
||
/// on it alongside SIGTERM/ctrl-c, so a self-triggered exit runs the IDENTICAL
|
||
/// graceful path (readiness 503 → drain → checkpoint) — never `process::abort`.
|
||
///
|
||
/// Triggered here on a detached thread (so the admin response flushes first).
|
||
fn trigger_graceful_self_exit() {
|
||
crate::self_exit::request_self_exit();
|
||
}
|
||
|
||
// ── Cluster management ──────────────────────────────────────────────────────
|
||
|
||
/// `GET /cluster/status/local` response body.
|
||
// The status DTO is intentionally a flat bag of independent boolean flags
|
||
// (is_leader, reachable, quarantined, reseed_required, reseeding,
|
||
// self_restart_refused) — each a distinct, orthogonal condition an operator
|
||
// reads directly; collapsing them into an enum would lose the
|
||
// multiple-simultaneously-true reality (e.g. quarantined AND reseed_required).
|
||
#[allow(clippy::struct_excessive_bools)]
|
||
#[derive(Clone, Serialize, ToSchema)]
|
||
pub struct LocalStatusResponse {
|
||
/// This node's region name.
|
||
region: String,
|
||
/// Whether this node currently believes it is the leader.
|
||
is_leader: bool,
|
||
/// The region this node believes leads (`null` during an election).
|
||
leader: Option<String>,
|
||
/// The leader's relay seqno (only meaningful when `is_leader`).
|
||
last_seq: u64,
|
||
/// Replication events applied on this node from the current leader.
|
||
applied_events: u64,
|
||
/// Events this node lags the leader by (0 when leading).
|
||
lag_events: u64,
|
||
/// Peers this leader is currently partitioned from (ship-skipped).
|
||
partitioned: Vec<String>,
|
||
/// The quorum commit index (m11p3): the highest seqno a majority of the
|
||
/// replica set durably holds. Only meaningful when `is_leader`;
|
||
/// `last_seq - commit_index` is the cluster's quorum lag.
|
||
commit_index: u64,
|
||
/// This node's write acknowledgment default (`leader` or `quorum`).
|
||
ack: String,
|
||
/// Always true (this node is serving its own status request).
|
||
reachable: bool,
|
||
/// This node's current election term (m11p4; 0 = the topology era).
|
||
term: u64,
|
||
/// This node's election role: `leader`, `follower`, `pre-candidate`,
|
||
/// `candidate` (or `unknown` before the driver starts).
|
||
role: String,
|
||
/// While leading: this node's ELECTION-TIME log position
|
||
/// `(term, frontier)` in the previous stream's numbering — the value the
|
||
/// vote restriction compared, and the zero-acked-loss frontier proof's
|
||
/// reference (the leader's own `last_seq` is in its NEW stream's
|
||
/// numbering and is not comparable to pre-election seqs).
|
||
prev_log_term: u64,
|
||
prev_log_seq: u64,
|
||
/// This node's LIVE election log position `(tail_term, frontier)` — the
|
||
/// exact `(LogPosition)` the Raft vote restriction (`rpc.log >= my_log`) and
|
||
/// `decide_join` read RIGHT NOW, distinct from the frozen `prev_log_*`
|
||
/// (published only while leading). Exposed so the election-divergence
|
||
/// consistency invariant is observable: for the same committed state a
|
||
/// caught-up node's `election_frontier` must equal its applied frontier in
|
||
/// the leader's stream (`applied_events`); a divergence here is the
|
||
/// cross-numbering tear (an ex-leader reading its own `flushed_seq` instead
|
||
/// of the leader-stream `applied_seqno`). `#[serde(default)]` so a
|
||
/// mixed-version peer's status still deserializes.
|
||
#[serde(default)]
|
||
election_tail_term: u64,
|
||
#[serde(default)]
|
||
election_frontier: u64,
|
||
/// The durable leader-acked frontier (m12 election-divergence-fix): the
|
||
/// highest seqno acked under `ack=leader` that may be un-replicated. The
|
||
/// election-divergence classifier quarantines a rejoin iff this exceeds the
|
||
/// new leadership's election baseline. Exposed for divergence diagnosis.
|
||
#[serde(default)]
|
||
leader_acked: u64,
|
||
/// This node's KNOWLEDGE of the followed group's leader frontier — the value
|
||
/// `lag_events` is subtracted from.
|
||
///
|
||
/// Publish it because `lag_events` alone cannot be read: it is an unsigned
|
||
/// subtraction against this number, so a node that has learned nothing reports
|
||
/// `0 - 0 = 0` and looks perfectly caught up. In the 2026-08-20 incident every
|
||
/// group answered `lag_events: 0` while holding `applied_events` of 24, 14 and
|
||
/// 0 against 5600 written items. `lag_events: 0` together with
|
||
/// `leader_seqno: 0` means NO INFORMATION, not converged.
|
||
#[serde(default)]
|
||
leader_seqno: u64,
|
||
/// Every tracked replication stream key with its applied frontier, as
|
||
/// `(key, applied)` pairs ascending by key.
|
||
///
|
||
/// A key is a per-LEADER-REGION stream id (`shard_of_region`), not a shard
|
||
/// group, so a group accumulates one key per leadership it has followed. Every
|
||
/// other field reports only the CURRENT leader's key, which made a position
|
||
/// retained from a previous leadership unobservable — while the receiver's gap
|
||
/// check (`receiver.rs`: `request_catchup(key, applied + 1)`) will chase ANY
|
||
/// key that received data. In the 2026-08-20 incident a node pulled
|
||
/// `from_seqno=13540653` (a key at 13540652) while the group it reported on sat
|
||
/// at 13540661, and nothing in the status could say which key that was.
|
||
#[serde(default)]
|
||
applied_by_key: Vec<(u32, u64)>,
|
||
/// Whether this node is quarantined with a divergent suffix (m11p4):
|
||
/// fenced from the data plane until reseeded.
|
||
quarantined: bool,
|
||
/// Whether this node has durably latched a `reseed_required` marker (m11p5
|
||
/// §2.4): a snapshot reseed runs on its next boot; it serves degraded until
|
||
/// then. Mirrors the `tidaldb_cluster_reseed_required` gauge.
|
||
reseed_required: bool,
|
||
/// Whether this node is mid-reseed: an install boot whose post-snapshot
|
||
/// catch-up has not yet first-converged (m11p5 §4). Readiness is 503 while
|
||
/// this is true.
|
||
reseeding: bool,
|
||
/// Whether a `reseed_self_restart` was REFUSED by the §2.4 quorum check
|
||
/// (the remaining voters cannot sustain quorum without this node). A `true`
|
||
/// here means a reseed is pending but the node is deliberately NOT exiting.
|
||
self_restart_refused: bool,
|
||
/// The applied membership conf version (m11p5 §3): 0 = the topology era (no
|
||
/// kind-4 record), else the highest applied record's version.
|
||
membership_version: u64,
|
||
/// The leadership term that authored the current roster (m11p5 §3): 0 in the
|
||
/// topology era.
|
||
membership_term: u64,
|
||
/// This node's role in the effective roster (m11p5 §3): `voter`, `learner`,
|
||
/// or `removed`. Distinct from `role` (the election role); a learner is a
|
||
/// `follower` here.
|
||
membership_role: String,
|
||
/// This node's build version + hash (m11p8 rolling-upgrade visibility):
|
||
/// `"<cargo-version>+<build-hash>"`. The aggregating gateway's status
|
||
/// fan-out collects every node's version, so an operator can confirm the
|
||
/// cluster is within the supported N/N+1 skew before/during a rolling
|
||
/// upgrade without shelling into each pod. `#[serde(default)]` so the
|
||
/// gateway can still deserialize a pre-m11p8 peer's status (empty version)
|
||
/// during a mixed-version upgrade window — the exact compat this phase needs.
|
||
#[serde(default)]
|
||
version: String,
|
||
/// Per-shard-group status for every group THIS node hosts (m11p6). For the
|
||
/// legacy single group this is one row mirroring the flat fields above; with
|
||
/// sharding it carries one row per hosted group, so an operator (and the
|
||
/// kill-node exit gate) can see which shard-leaderships this node holds and
|
||
/// confirm only a dead node's leaderships move.
|
||
#[serde(default)]
|
||
shards: Vec<ShardStatusRow>,
|
||
}
|
||
|
||
/// One hosted shard group's status within a [`LocalStatusResponse`] (m11p6).
|
||
#[derive(Clone, Serialize, ToSchema)]
|
||
pub struct ShardStatusRow {
|
||
/// The data-shard group id.
|
||
shard: u16,
|
||
/// Whether this node currently leads this group.
|
||
is_leader: bool,
|
||
/// The region this node believes leads this group (`null` mid-election).
|
||
leader: Option<String>,
|
||
/// This group's current election term.
|
||
term: u64,
|
||
/// This node's election role in this group (`leader`/`follower`/…).
|
||
role: String,
|
||
/// Replication events this node has applied for this group.
|
||
applied_events: u64,
|
||
/// This node's knowledge of THIS GROUP's leader frontier — the value
|
||
/// `lag_events` subtracts from. `lag_events: 0` with `leader_seqno: 0` means
|
||
/// NO INFORMATION (a node that has learned nothing), not converged.
|
||
#[serde(default)]
|
||
leader_seqno: u64,
|
||
/// Events this node lags this group's leader by (0 when leading).
|
||
lag_events: u64,
|
||
/// This group's quorum commit index (meaningful when leading).
|
||
commit_index: u64,
|
||
/// Whether THIS GROUP has a durable reseed marker latched.
|
||
///
|
||
/// Load-bearing on a multi-group node: the flat `reseed_required` on
|
||
/// [`LocalStatusResponse`] reports only the group `replica_for(None)` picks
|
||
/// (the lowest hosted id), so a marker on any OTHER group was invisible in the
|
||
/// status surface. That is how the 2026-08-20 incident hid — tidaldb-0 answered
|
||
/// `reseed_required: false, lag_events: 0` while another hosted group sat behind
|
||
/// a compacted leader. `#[serde(default)]` so a gateway can still deserialize a
|
||
/// pre-fix peer during a mixed-version window.
|
||
#[serde(default)]
|
||
reseed_required: bool,
|
||
/// Whether THIS GROUP is mid-reseed (install/seed-join not yet converged).
|
||
#[serde(default)]
|
||
reseeding: bool,
|
||
}
|
||
|
||
/// Local replication / leadership status for THIS region.
|
||
#[utoipa::path(
|
||
get,
|
||
path = "/cluster/status/local",
|
||
tag = "cluster",
|
||
responses(
|
||
(status = 200, description = "This region's local status", body = LocalStatusResponse),
|
||
(status = 401, description = "Missing or invalid credential"),
|
||
(status = 503, description = "Server shutting down"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn status_local(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
Query(sel): Query<ShardSelector>,
|
||
) -> std::result::Result<Json<LocalStatusResponse>, ClusterAppError> {
|
||
// m11p6: the flat fields mirror the selected group (the first hosted group by
|
||
// default — S=1 byte-for-byte); the `shards` array always carries every
|
||
// hosted group so an operator/the kill-node gate sees which shard-leaderships
|
||
// this node holds.
|
||
let mut status = node
|
||
.replica_for(sel.shard_id())?
|
||
.local_status()
|
||
.map_err(ClusterAppError)?;
|
||
status.shards = node.shard_status_rows();
|
||
Ok(Json(status))
|
||
}
|
||
|
||
// ── Aggregated cluster status ─────────────────────────────────────────────────
|
||
|
||
/// `GET /cluster/status` aggregated response body.
|
||
#[derive(Serialize, ToSchema)]
|
||
pub struct AggregatedStatusResponse {
|
||
/// This node's view of the current leader. **For `S=1` this is THE cluster
|
||
/// leader; for `S>1` it is the DEFAULT (lowest-id hosted) group's leader** —
|
||
/// a sharded cluster has one leader per group, so read `shards` for the full
|
||
/// picture. The flat `leader`/`relay_log_len`/`regions` block is the
|
||
/// default group's cross-region view, kept verbatim for `S=1` back-compat.
|
||
leader: String,
|
||
/// The default group's leader relay seqno (high-water-mark): the count every
|
||
/// region of that group should converge to.
|
||
relay_log_len: u64,
|
||
/// Per-region replication status for the default group, in topology id order.
|
||
regions: Vec<AggregatedRegionStatus>,
|
||
/// Per-hosted-group leadership rows from THIS node (m11p6): one row per shard
|
||
/// group this node replicates, so an operator (and the kill-node exit gate)
|
||
/// can see which shard-leaderships this node holds across all its groups —
|
||
/// the multi-leader truth the flat `leader` field cannot express. For `S=1`
|
||
/// this is a single row mirroring the flat fields. NB: it is THIS node's view
|
||
/// of its hosted groups; a cross-cluster per-shard merge is the L4 surface.
|
||
#[serde(default)]
|
||
shards: Vec<ShardStatusRow>,
|
||
}
|
||
|
||
/// One region's aggregated status within an [`AggregatedStatusResponse`].
|
||
///
|
||
/// `applied_events` / `lag_events` are `Option<u64>` because the aggregating node
|
||
/// frequently CANNOT know a peer's frontier: its status probe can fail, and only
|
||
/// the leader holds an ack record to fall back on. Before this was
|
||
/// representable the `None` arm of [`aggregate_region_row`] returned
|
||
/// `applied_events: 0, lag_events: leader_last_seq` — and since `lag` is
|
||
/// `leader_hwm − applied`, "the probe timed out" was rendered as a 13.3-million
|
||
/// event deficit on a cluster whose own `shards[]` rows showed every replica
|
||
/// converged at an identical frontier. `null` is the honest answer; `0` is the
|
||
/// most alarming possible lie.
|
||
#[derive(Serialize, ToSchema)]
|
||
pub struct AggregatedRegionStatus {
|
||
/// Region name.
|
||
name: String,
|
||
/// Replication events applied on this region (from its own local status; the
|
||
/// leader row mirrors today's semantics — its replication-from-others count),
|
||
/// or **`null` when this node has no report for it**. `null` means UNKNOWN —
|
||
/// never read it as 0.
|
||
applied_events: Option<u64>,
|
||
/// Events this region lags the leader by (`leader_last_seq − applied_events`),
|
||
/// or **`null` when either side of that subtraction is unknown** (no frontier
|
||
/// report for the region, or the leader itself unreachable so there is no
|
||
/// high-water-mark to subtract from). `null` is NOT `0`: it means this node
|
||
/// cannot say, and an operator must query the region directly.
|
||
lag_events: Option<u64>,
|
||
/// Whether this region is currently partitioned from the leader.
|
||
partitioned: bool,
|
||
/// Whether this node could reach the region's `/cluster/status/local` within
|
||
/// the per-peer budget. An unreachable region reports `applied_events: null`,
|
||
/// `lag_events: null`, `partitioned: true` — the honest unknown, never a
|
||
/// fabricated worst case.
|
||
reachable: bool,
|
||
/// The region's reported build version (m11p8): `/cluster/status` is the
|
||
/// single pane an operator reads to confirm the whole cluster is within the
|
||
/// supported N/N+1 skew before a rolling upgrade. Empty for an unreachable
|
||
/// region or a pre-m11p8 peer.
|
||
#[serde(default)]
|
||
version: String,
|
||
}
|
||
|
||
/// Aggregated replication status across EVERY region of the default shard group.
|
||
///
|
||
/// Queries every region's `/cluster/status/local` (own region in-process, peers
|
||
/// over HTTP) concurrently with a 500ms per-peer budget, then assembles the
|
||
/// cluster-wide view: `relay_log_len` is the leader's `last_seq`, and each
|
||
/// region's `lag_events = leader_last_seq − applied` when BOTH are known.
|
||
///
|
||
/// An unreachable peer is reported as `reachable: false`, `partitioned: true`,
|
||
/// `applied_events: null`, `lag_events: null` — an explicit UNKNOWN. It used to
|
||
/// report `applied 0` / `lag = leader_last_seq`, which turned a 500ms probe
|
||
/// timeout into a fabricated multi-million-event deficit; nothing here
|
||
/// reconstructs a guess when the value is not known.
|
||
///
|
||
/// m11p6: the flat fields describe the DEFAULT (lowest-id hosted) group only —
|
||
/// exact for `S=1`. `shards` carries this node's per-group leadership so a
|
||
/// multi-shard operator/the kill-node gate is not blind to the other groups; a
|
||
/// full cross-cluster per-shard aggregation is the L4 surface.
|
||
#[utoipa::path(
|
||
get,
|
||
path = "/cluster/status",
|
||
tag = "cluster",
|
||
responses(
|
||
(status = 200, description = "Cluster-wide replication status", body = AggregatedStatusResponse),
|
||
(status = 401, description = "Missing or invalid credential"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn cluster_status(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
) -> std::result::Result<Json<AggregatedStatusResponse>, ClusterAppError> {
|
||
let state = node.replica_for(None)?;
|
||
let leader_name = state.leader_name();
|
||
let regions = state.all_regions_for_status();
|
||
|
||
// Query every region's local status concurrently. The own region is served
|
||
// in-process (no HTTP hop); peers over the async client with the per-peer
|
||
// budget. Each entry: (RegionId, name, Option<local-status-json>).
|
||
let client = state.client.clone();
|
||
let own_status = state.local_status().ok();
|
||
let futures = regions.into_iter().map(|(rid, name, http)| {
|
||
let client = client.clone();
|
||
let own_status = if rid == state.region {
|
||
own_status.clone()
|
||
} else {
|
||
None
|
||
};
|
||
async move {
|
||
// Own region: use the in-process status, no HTTP. A
|
||
// LocalStatusResponse that fails to serialize is a programming
|
||
// error; degrade this row to unreachable (None ⇒ `reachable:
|
||
// false`) rather than aggregating a null row as healthy.
|
||
if let Some(local) = own_status {
|
||
let json = serde_json::to_value(local)
|
||
.map_err(|e| tracing::error!(error = %e, "serialize own region status"))
|
||
.ok();
|
||
return (rid, name, json);
|
||
}
|
||
let Some(http_addr) = http else {
|
||
return (rid, name, None); // unreachable: no declared addr
|
||
};
|
||
let url = peer_url(&http_addr, "/cluster/status/local");
|
||
let body = client
|
||
.get(&url)
|
||
.timeout(forward::STATUS_PEER_TIMEOUT)
|
||
.send()
|
||
.await
|
||
.ok();
|
||
let json = match body {
|
||
Some(resp) if resp.status().is_success() => {
|
||
resp.json::<serde_json::Value>().await.ok()
|
||
}
|
||
_ => None,
|
||
};
|
||
(rid, name, json)
|
||
}
|
||
});
|
||
let results = futures_util::future::join_all(futures).await;
|
||
|
||
// The leader's last_seq is the high-water-mark every region converges to.
|
||
// Source it from the leader's local status row (or this node's own when it
|
||
// leads). `None` when the leader itself was unreachable — there is then NO
|
||
// high-water-mark, so every row's `lag_events` is `null` rather than a
|
||
// deficit measured against a zero nobody reported.
|
||
let leader_last_seq: Option<u64> = results
|
||
.iter()
|
||
.find(|(_, name, _)| name == &leader_name)
|
||
.and_then(|(_, _, json)| json.as_ref())
|
||
.and_then(|j| j.get("last_seq").and_then(serde_json::Value::as_u64));
|
||
|
||
// Leader-side gRPC liveness, for the false-partition fix: under a sustained
|
||
// quorum write-burst (large 1536-dim HNSW inserts), a follower's HTTP
|
||
// control-plane can be momentarily starved while it keeps APPLYING and
|
||
// ACKING replication. The old probe-only check then misread the
|
||
// `/cluster/status/local` timeout as a partition, stalled quorum writes, and
|
||
// 408'd reads — yet the follower was alive. A peer the leader has had a gRPC
|
||
// round-trip with inside this window (an accepted ship OR a backpressure
|
||
// reply, both proving the peer's gRPC server is up) is SLOW-but-alive, not
|
||
// partitioned. A genuinely partitioned/dead peer never refreshes that stamp
|
||
// (a real TCP severance kills both the HTTP probe and the gRPC ship), so it
|
||
// still falls through to `reachable: false` — the chaos-suite contract holds.
|
||
let leader_view = state.is_leader();
|
||
// The leader's per-peer durable mark map (its quorum-fold view of how far
|
||
// each follower has acked), so an HTTP-unreachable-but-gRPC-alive peer
|
||
// reports an HONEST lag from the leader's own ack record rather than a
|
||
// worst-case `lag = leader_last_seq`.
|
||
let peer_marks: std::collections::HashMap<ShardId, u64> = if leader_view {
|
||
state.commit.peer_marks().into_iter().collect()
|
||
} else {
|
||
std::collections::HashMap::new()
|
||
};
|
||
|
||
let region_rows = results
|
||
.into_iter()
|
||
.map(|(rid, name, json)| {
|
||
// A peer the leader has had fresh gRPC contact with is alive even if
|
||
// its HTTP status probe timed out (the apply-burst false-partition
|
||
// guard). Computed once per row so `aggregate_region_row` stays a pure
|
||
// assembler.
|
||
let grpc_fresh = leader_view
|
||
&& state
|
||
.transport
|
||
.peer_grpc_fresh(shard_of_region(rid), forward::GRPC_LIVENESS_WINDOW);
|
||
// `None` when the leader holds no mark for this peer yet (and always
|
||
// off the leader, where the map is empty): an absent mark is an
|
||
// unknown frontier, never `applied 0`.
|
||
let leader_mark = peer_marks.get(&shard_of_region(rid)).copied();
|
||
aggregate_region_row(
|
||
name,
|
||
json.as_ref(),
|
||
&leader_name,
|
||
leader_last_seq,
|
||
grpc_fresh,
|
||
leader_mark,
|
||
)
|
||
})
|
||
.collect::<Vec<_>>();
|
||
|
||
// Authoritative partition view: a region is partitioned iff the LEADER's
|
||
// ship-skip set names it. Re-derive every row's `partitioned` from the
|
||
// leader's reported set so a follower's own (empty) set cannot mask it.
|
||
let region_rows = apply_leader_partition_view(&state, region_rows);
|
||
|
||
Ok(Json(AggregatedStatusResponse {
|
||
leader: leader_name,
|
||
// `0` here means the leader's own status probe failed, so its
|
||
// high-water-mark is unknown. Nothing DERIVES from that zero: every
|
||
// region's `lag_events` is `null` in that case (see
|
||
// `aggregate_region_row`), which is the fabrication that mattered.
|
||
relay_log_len: leader_last_seq.unwrap_or(0),
|
||
regions: region_rows,
|
||
// Per-hosted-group leadership for THIS node — the multi-shard view the
|
||
// flat `leader` cannot express (S=1 ⇒ one row mirroring the flat fields).
|
||
shards: node.shard_status_rows(),
|
||
}))
|
||
}
|
||
|
||
/// Assemble one region's aggregated status row from its `/cluster/status/local`
|
||
/// reply (`json`), or from the leader's own view when that probe failed.
|
||
///
|
||
/// Three cases:
|
||
/// 1. **`json` present** — the peer answered ABOUT ITSELF, which is an honest
|
||
/// source whichever node is aggregating: copy its applied/version, derive lag
|
||
/// from the leader HWM, and mirror its self-reported partition flag. A reply
|
||
/// that omits `applied_events` yields `None`, not `0`.
|
||
/// 2. **`json` absent but `grpc_fresh`** — the HTTP probe timed out yet the leader
|
||
/// has had a recent gRPC round-trip (accepted ship OR backpressure) with the
|
||
/// peer: it is SLOW-but-alive (its HTTP control-plane is starved under an apply
|
||
/// burst), so report `reachable: true`, `partitioned: false`, with an HONEST
|
||
/// frontier from the leader's own ack mark (`leader_mark`) — and `None` when the
|
||
/// leader holds no mark for it. Only the leader ever has marks, so a FOLLOWER
|
||
/// never reaches this arm (`grpc_fresh` is gated on `leader_view`): a follower
|
||
/// has no ack record and therefore never guesses from one.
|
||
/// 3. **`json` absent and not `grpc_fresh`** — no reply and no recent gRPC contact:
|
||
/// a genuine partition/dead peer, or simply a peer this node cannot see. Report
|
||
/// `reachable: false`, `partitioned: true`, and `applied_events`/`lag_events`
|
||
/// **`null`** — the honest unknown. This arm used to return
|
||
/// `applied_events: 0, lag_events: leader_last_seq`, which rendered a 500ms
|
||
/// probe timeout as a multi-million-event deficit on a converged cluster.
|
||
///
|
||
/// `lag_events` is `Some` only when BOTH sides of the subtraction are known: an
|
||
/// unknown frontier or an unreachable leader (no HWM) both yield `null`, never a
|
||
/// number derived from a value nobody reported.
|
||
///
|
||
/// `apply_leader_partition_view` runs after this and re-stamps `partitioned` from
|
||
/// the authoritative ship-skip set, so an operator `/cluster/partition` always wins
|
||
/// over the case-2 `partitioned: false`.
|
||
fn aggregate_region_row(
|
||
name: String,
|
||
json: Option<&serde_json::Value>,
|
||
leader_name: &str,
|
||
leader_last_seq: Option<u64>,
|
||
grpc_fresh: bool,
|
||
leader_mark: Option<u64>,
|
||
) -> AggregatedRegionStatus {
|
||
/// `leader_hwm − frontier`, or `None` when either side is unknown.
|
||
fn lag_from(leader_hwm: Option<u64>, frontier: Option<u64>) -> Option<u64> {
|
||
frontier
|
||
.zip(leader_hwm)
|
||
.map(|(applied, hwm)| hwm.saturating_sub(applied))
|
||
}
|
||
|
||
match json {
|
||
Some(j) => {
|
||
// The peer's report about its OWN frontier. Absent ⇒ unknown.
|
||
let applied = j.get("applied_events").and_then(serde_json::Value::as_u64);
|
||
let is_leader = name == leader_name;
|
||
let lag = if is_leader {
|
||
// The leader is by definition at its own high-water-mark.
|
||
Some(0)
|
||
} else {
|
||
lag_from(leader_last_seq, applied)
|
||
};
|
||
let partitioned = j
|
||
.get("partitioned")
|
||
.and_then(|p| p.as_array())
|
||
.is_some_and(|a| {
|
||
// A region is "partitioned" in the aggregate iff the LEADER
|
||
// lists it in its ship-skip set. The leader's own local status
|
||
// carries that set; mirror it.
|
||
a.iter().any(|v| v.as_str() == Some(name.as_str()))
|
||
});
|
||
let version = j
|
||
.get("version")
|
||
.and_then(serde_json::Value::as_str)
|
||
.unwrap_or("")
|
||
.to_owned();
|
||
AggregatedRegionStatus {
|
||
name,
|
||
applied_events: applied,
|
||
lag_events: lag,
|
||
partitioned,
|
||
reachable: true,
|
||
version,
|
||
}
|
||
}
|
||
// HTTP probe failed but the leader has fresh gRPC contact: alive-but-slow.
|
||
None if grpc_fresh => AggregatedRegionStatus {
|
||
name,
|
||
applied_events: leader_mark,
|
||
lag_events: lag_from(leader_last_seq, leader_mark),
|
||
partitioned: false,
|
||
reachable: true,
|
||
version: String::new(),
|
||
},
|
||
// No reply, no recent gRPC contact: this node cannot see the peer. It says
|
||
// so, and says NOTHING about the peer's frontier.
|
||
None => AggregatedRegionStatus {
|
||
name,
|
||
applied_events: None,
|
||
lag_events: None,
|
||
partitioned: true,
|
||
reachable: false,
|
||
version: String::new(),
|
||
},
|
||
}
|
||
}
|
||
|
||
/// Override each region row's `partitioned` flag with the LEADER's ship-skip set
|
||
/// (the single authority for explicit partition state), UNION the
|
||
/// implicitly-partitioned set of unreachable regions.
|
||
///
|
||
/// An UNREACHABLE region is partitioned regardless of the ship-skip set (its
|
||
/// row already carries `reachable: false` from the aggregation), so the leader's
|
||
/// empty skip set must not clear it. When this node is NOT the leader, the
|
||
/// per-row flag (copied from each region's own local status, plus the
|
||
/// unreachable override) stands.
|
||
fn apply_leader_partition_view(
|
||
state: &Arc<ShardReplica>,
|
||
mut rows: Vec<AggregatedRegionStatus>,
|
||
) -> Vec<AggregatedRegionStatus> {
|
||
if !state.is_leader() {
|
||
// A follower aggregator trusts each row's own partition flag plus the
|
||
// unreachable override already applied during aggregation.
|
||
return rows;
|
||
}
|
||
let skip: std::collections::HashSet<String> = read_recovered(&state.partitioned, "partitioned")
|
||
.iter()
|
||
.map(|r| state.region_name_of(*r).to_string())
|
||
.collect();
|
||
for row in &mut rows {
|
||
// Explicitly ship-skipped OR unreachable ⇒ partitioned. The unreachable
|
||
// case must survive an empty ship-skip set.
|
||
row.partitioned = skip.contains(&row.name) || !row.reachable;
|
||
}
|
||
rows
|
||
}
|
||
|
||
/// Body for the region management routes.
|
||
///
|
||
/// `Serialize` so promote can fan it out and partition/heal can forward it.
|
||
#[derive(Serialize, Deserialize, ToSchema)]
|
||
pub struct RegionRequest {
|
||
/// Target region name.
|
||
#[schema(example = "us-east")]
|
||
region: String,
|
||
/// The new leader's stream baseline (m11p2), set only on the promote
|
||
/// fan-out legs: peers jump their applied frontier for the new leader's
|
||
/// shard to it so the new stream is contiguous from its first batch.
|
||
#[serde(default, skip_serializing_if = "Option::is_none")]
|
||
baseline: Option<u64>,
|
||
}
|
||
|
||
/// The `?shard=` selector (m11p6) shared by every per-shard admin surface
|
||
/// (`/cluster/promote`, `/heal`, `/partition`, `/catchup`, `/reseed`,
|
||
/// `/members`, `/members/remove`, `/join`, `/status/local`).
|
||
///
|
||
/// Absent ⇒ the FIRST hosted group — exact for the legacy single group (S=1,
|
||
/// byte-for-byte) and the gateway default. `?shard=N` targets that group via
|
||
/// [`ClusterNode::replica_for`] (400 if this node hosts no replica of it). An
|
||
/// intra-group forward carries the selector forward ([`ShardReplica::admin_path`])
|
||
/// so the receiving sibling resolves the SAME group.
|
||
#[derive(Debug, Default, Deserialize)]
|
||
pub struct ShardSelector {
|
||
/// The data-shard group id to target. Absent for the legacy single group.
|
||
#[serde(default)]
|
||
shard: Option<u16>,
|
||
}
|
||
|
||
impl ShardSelector {
|
||
/// The selected shard as a [`ShardId`], or `None` for the default group.
|
||
fn shard_id(&self) -> Option<ShardId> {
|
||
self.shard.map(ShardId)
|
||
}
|
||
}
|
||
|
||
/// Promote a region to leader across the cluster (m11p2).
|
||
///
|
||
/// * **internal marker present**: apply the leadership change locally only
|
||
/// (this is how the fan-out terminates on each peer). When this node IS the
|
||
/// target it activates its ship queue and returns its stream `baseline`;
|
||
/// otherwise it deactivates and (when the body carries a `baseline`) jumps
|
||
/// its applied frontier for the new leader's shard.
|
||
/// * **external request**: resolve the TARGET's baseline first — locally when
|
||
/// this node is the target, else via a marked promote to the target — then
|
||
/// fan the promote (with the baseline) out to every other peer. Returns
|
||
/// `{ok, leader, baseline, acked: [..], failed: [..]}`. A peer that misses
|
||
/// the fan-out self-corrects: its first parked batch from the new stream
|
||
/// triggers a catch-up pull whose chunks announce the baseline.
|
||
#[utoipa::path(
|
||
post,
|
||
path = "/cluster/promote",
|
||
tag = "cluster",
|
||
request_body = RegionRequest,
|
||
responses(
|
||
(status = 200, description = "Leadership updated (and fanned out to peers)"),
|
||
(status = 400, description = "Unknown region"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
(status = 503, description = "Promotion target unreachable"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
// One linear protocol pass (marked leg -> fenced transfer -> takeover wait ->
|
||
// legacy fallback); splitting it would scatter the transfer's ordering rules.
|
||
#[allow(clippy::too_many_lines)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn cluster_promote(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
Query(sel): Query<ShardSelector>,
|
||
Json(req): Json<RegionRequest>,
|
||
) -> std::result::Result<Json<serde_json::Value>, ClusterAppError> {
|
||
let state = node.replica_for(sel.shard_id())?;
|
||
if is_internal(&headers) {
|
||
// Marked fan-out leg (the LEGACY term-0 protocol): apply locally and
|
||
// terminate. promote_local fences this once the cluster is
|
||
// term-governed (m11p4).
|
||
let baseline = state
|
||
.promote_local(&req.region, req.baseline)
|
||
.map_err(ClusterAppError)?;
|
||
return Ok(Json(serde_json::json!({
|
||
"ok": true,
|
||
"leader": req.region,
|
||
"baseline": baseline,
|
||
})));
|
||
}
|
||
|
||
// m11p7: the operator-originated promote is wrapped so its many `return`s
|
||
// become this block's value, which we audit (principal/target/term/outcome)
|
||
// before returning. The internal fan-out leg above is NOT audited (the
|
||
// operator leg already recorded the action once).
|
||
let term = state.election_term();
|
||
let target_name = req.region.clone();
|
||
let result: std::result::Result<Json<serde_json::Value>, ClusterAppError> = async {
|
||
// m11p4: promote is a FENCED TRANSFER — leadership moves through an
|
||
// election (term+1), never an un-fenced view flip. The legacy term-0
|
||
// fan-out remains only as the mixed-version fallback below.
|
||
let target = state.resolve_region(&req.region).map_err(ClusterAppError)?;
|
||
// TOPOLOGY ERA: no node has JOINED an elected term ≥ 1 yet (the topology file
|
||
// names the leader at term 0, `joined_term == 0`). In this era the LEADER
|
||
// does not fence-transfer leadership AWAY — it falls through to the LEGACY
|
||
// fan-out (apply the target's baseline, fan it out, return the documented
|
||
// `{baseline, acked, failed}` shape; the split-brain-heal collapse step
|
||
// depends on that fan-out reaching every peer). A node campaigning FOR
|
||
// ITSELF (`target == self`) still campaigns — that is the only way a
|
||
// restarted-as-follower node regains leadership when auto-election is off —
|
||
// and a non-leader/non-target node still FORWARDS to the target. Only the
|
||
// leader-driven `transfer_to` is era-gated.
|
||
let topology_era = state
|
||
.election_runtime()
|
||
.is_none_or(|rt| rt.joined_term() == 0);
|
||
// Topology-era re-assert: promoting the node that ALREADY leads, before any
|
||
// election has ever been joined, is the legacy verb's re-broadcast shape —
|
||
// skip the transfer machinery and fall through to the legacy fan-out.
|
||
let legacy_reassert = state.current_leader() == Some(target) && topology_era;
|
||
if let Some(rt) = state.election_runtime().filter(|_| !legacy_reassert) {
|
||
let prior_term = rt.current_term();
|
||
// Promoting the CURRENT leader once an elected leadership exists is
|
||
// a no-op success, whether this node IS that leader or merely knows
|
||
// it — the operator's "assert who leads" shape.
|
||
if state.current_leader() == Some(target) {
|
||
return Ok(Json(serde_json::json!({
|
||
"ok": true,
|
||
"leader": req.region,
|
||
"term": prior_term,
|
||
"transfer": "already-leader",
|
||
})));
|
||
}
|
||
|
||
// A topology-era LEADER does not fence-transfer leadership away: no
|
||
// elected term has ever existed (there is nothing to fence, and the
|
||
// target cannot acknowledge a TimeoutNow at a joined term it does not
|
||
// hold). It falls straight through to the LEGACY fan-out below, which
|
||
// moves the term-0 view and returns the documented `{baseline, acked,
|
||
// failed}` shape. This is the m11p4 contract — the leader's topology-era
|
||
// transfer never "took", so it always reached the legacy fan-out;
|
||
// m11p5's faster election made the transfer take, silently changing the
|
||
// response shape, and this gate restores it. The self-campaign and
|
||
// forward legs (a restarted-as-follower node regaining leadership, or a
|
||
// non-leader/non-target relaying the verb) still run in the topology era.
|
||
let leader_topology_passthrough = state.is_leader() && topology_era;
|
||
if leader_topology_passthrough {
|
||
// intentionally empty — fall through to the legacy fan-out below.
|
||
} else if state.is_leader() {
|
||
// Leader-sanctioned transfer: the catch-up wait IS the drain —
|
||
// the target must hold the full flushed prefix before TimeoutNow.
|
||
// The signal is the COMMIT INDEX's per-peer durable mark (fed by
|
||
// the target's own ReportApplied pushes): unlike the ship
|
||
// queue's acked frontier it cannot stall behind an open circuit
|
||
// breaker after the target's restart — the target's catch-up
|
||
// pull and frontier reports flow regardless of this leader's
|
||
// outbound breaker state.
|
||
let target_shard = shard_of_region(target);
|
||
let target_mark = |state: &ShardReplica| {
|
||
state
|
||
.commit
|
||
.peer_marks()
|
||
.into_iter()
|
||
.find(|&(p, _)| p == target_shard)
|
||
.map_or(0, |(_, m)| m)
|
||
};
|
||
let deadline = std::time::Instant::now() + TRANSFER_CATCHUP_WAIT;
|
||
loop {
|
||
let flushed = state.ship_feed.flushed_seq();
|
||
if target_mark(&state) >= flushed {
|
||
break;
|
||
}
|
||
if std::time::Instant::now() >= deadline {
|
||
return Err(ClusterAppError(ServerError::Cluster(format!(
|
||
"transfer target '{}' lags the flushed frontier ({} < {}); \
|
||
heal it first, then retry the promote",
|
||
req.region,
|
||
target_mark(&state),
|
||
state.ship_feed.flushed_seq()
|
||
))));
|
||
}
|
||
tokio::time::sleep(std::time::Duration::from_millis(50)).await;
|
||
}
|
||
rt.transfer_to(target);
|
||
} else if target == state.region {
|
||
// Target-side request. Prefer the leader-sanctioned path when a
|
||
// live leader is known; otherwise campaign directly (the dead-
|
||
// leader failover drill).
|
||
let sanctioned = if let Some(addr) = state.leader_http_addr() {
|
||
let url = peer_url(&addr, &state.admin_path("/cluster/promote"));
|
||
let auth = forwarded_auth(&headers);
|
||
let body = serde_json::json!({ "region": req.region });
|
||
// A relayed operator hop: the leader runs the full fenced
|
||
// transfer but does NOT re-audit (this node audits once below).
|
||
forward_json_with_headers(
|
||
&state.client,
|
||
&url,
|
||
&body,
|
||
auth.as_deref(),
|
||
false,
|
||
&state.relay_passthrough(),
|
||
)
|
||
.await
|
||
.map(|resp| resp.status.is_success())
|
||
.unwrap_or(false)
|
||
} else {
|
||
false
|
||
};
|
||
if !sanctioned {
|
||
rt.campaign_now();
|
||
}
|
||
} else {
|
||
// Neither the leader nor the target: hand the request to the
|
||
// target (it sanctions through its leader or campaigns).
|
||
let Some(addr) = state.peer_http.get(&target).cloned() else {
|
||
return Err(ClusterAppError(ServerError::BadRequest(format!(
|
||
"region '{}' has no http_addr to promote",
|
||
req.region
|
||
))));
|
||
};
|
||
let url = peer_url(&addr, &state.admin_path("/cluster/promote"));
|
||
let auth = forwarded_auth(&headers);
|
||
let body = serde_json::json!({ "region": req.region });
|
||
// A relayed operator hop: the target runs the full protocol
|
||
// (sanction-through-leader or campaign) but does NOT re-audit.
|
||
if let Err(e) = forward_json_with_headers(
|
||
&state.client,
|
||
&url,
|
||
&body,
|
||
auth.as_deref(),
|
||
false,
|
||
&state.relay_passthrough(),
|
||
)
|
||
.await
|
||
{
|
||
return Err(ClusterAppError(ServerError::RegionUnreachable {
|
||
region: req.region,
|
||
cause: e,
|
||
}));
|
||
}
|
||
}
|
||
|
||
// The topology-era leader took the empty passthrough leg above: it did
|
||
// not initiate a transfer, so do not wait for one — fall through to the
|
||
// legacy fan-out (the m11p4 shape).
|
||
if !leader_topology_passthrough {
|
||
// Wait for the transfer to take: the target leads. A genuine
|
||
// takeover bumps the term; a forwarded request that discovers the
|
||
// target ALREADY led resolves at the same term — both are success
|
||
// (hence >=, not >).
|
||
let deadline = std::time::Instant::now() + TRANSFER_TAKEOVER_WAIT;
|
||
while std::time::Instant::now() < deadline {
|
||
if state.current_leader() == Some(target) && state.election_term() >= prior_term
|
||
{
|
||
return Ok(Json(serde_json::json!({
|
||
"ok": true,
|
||
"leader": req.region,
|
||
"term": state.election_term(),
|
||
"transfer": "elected",
|
||
})));
|
||
}
|
||
tokio::time::sleep(std::time::Duration::from_millis(50)).await;
|
||
}
|
||
|
||
// The election did not take. Once an elected leadership has been
|
||
// joined (`joined_term >= 1`) there is no safe fallback — report
|
||
// honestly and keep the current leader. In the TOPOLOGY ERA a
|
||
// self-campaign or relayed promote that did not take instead falls
|
||
// through to the LEGACY fan-out below (m11p4's mixed-version /
|
||
// operator-override path), with a loud WARN.
|
||
if !topology_era {
|
||
return Err(ClusterAppError(ServerError::Cluster(format!(
|
||
"leadership transfer to '{}' did not complete within {:?} \
|
||
(term {} -> {}); the cluster keeps its current leader — retry, \
|
||
or check the target's health",
|
||
req.region,
|
||
TRANSFER_TAKEOVER_WAIT,
|
||
prior_term,
|
||
state.election_term()
|
||
))));
|
||
}
|
||
tracing::warn!(
|
||
target = %req.region,
|
||
"topology-era promote did not take via election (auto-election off, \
|
||
or a mixed-version / isolated-operator override); falling back to \
|
||
the LEGACY fan-out promote"
|
||
);
|
||
}
|
||
}
|
||
|
||
// Legacy term-0 fan-out (pre-m11p4 protocol; also the mixed-version
|
||
// fallback): the TARGET's baseline must exist before the fan-out
|
||
// (peers need it to jump their frontiers). Resolve it locally when this
|
||
// node is the target; otherwise ask the target first via a marked promote.
|
||
let target = state.resolve_region(&req.region).map_err(ClusterAppError)?;
|
||
let baseline = if target == state.region {
|
||
state
|
||
.promote_local(&req.region, None)
|
||
.map_err(ClusterAppError)?
|
||
} else {
|
||
let Some(http_addr) = state.peer_http.get(&target).cloned() else {
|
||
return Err(ClusterAppError(ServerError::BadRequest(format!(
|
||
"region '{}' has no http_addr to promote",
|
||
req.region
|
||
))));
|
||
};
|
||
let url = peer_url(&http_addr, &state.admin_path("/cluster/promote"));
|
||
let auth = forwarded_auth(&headers);
|
||
let body = serde_json::json!({ "region": req.region });
|
||
let baseline = match forward_json_with_headers(
|
||
&state.client,
|
||
&url,
|
||
&body,
|
||
auth.as_deref(),
|
||
true,
|
||
&state.node_token_passthrough(),
|
||
)
|
||
.await
|
||
{
|
||
Ok(resp) if resp.status.is_success() => resp
|
||
.body
|
||
.get("baseline")
|
||
.and_then(serde_json::Value::as_u64),
|
||
Ok(resp) => {
|
||
return Err(ClusterAppError(ServerError::RegionUnreachable {
|
||
region: req.region,
|
||
cause: format!("promotion target returned {}", resp.status),
|
||
}));
|
||
}
|
||
Err(e) => {
|
||
return Err(ClusterAppError(ServerError::RegionUnreachable {
|
||
region: req.region,
|
||
cause: e,
|
||
}));
|
||
}
|
||
};
|
||
// Apply locally AFTER the target accepted (this node demotes itself /
|
||
// updates its view and jumps its frontier to the announced baseline).
|
||
let _ = state
|
||
.promote_local(&req.region, baseline)
|
||
.map_err(ClusterAppError)?;
|
||
baseline
|
||
};
|
||
|
||
// Fan out to every OTHER peer with the marker + baseline set. The target
|
||
// already applied (above); re-applying there would be harmless (its
|
||
// promote_local is idempotent) but is skipped by the fan-out filter
|
||
// below being all-peers — the target simply re-activates from the same
|
||
// baseline, which is a no-op for its peers' frontiers.
|
||
let fan_body = serde_json::json!({ "region": req.region, "baseline": baseline });
|
||
let outcome = broadcast_to_peers(
|
||
&state,
|
||
&state.admin_path("/cluster/promote"),
|
||
&fan_body,
|
||
&headers,
|
||
)
|
||
.await;
|
||
Ok(Json(serde_json::json!({
|
||
"ok": true,
|
||
"leader": req.region,
|
||
"baseline": baseline,
|
||
"acked": outcome.acked,
|
||
"failed": outcome.failed,
|
||
})))
|
||
}
|
||
.await;
|
||
let outcome_str = match &result {
|
||
Ok(_) => "applied".to_string(),
|
||
// `ClusterAppError` is a newtype over the `Display` `ServerError`.
|
||
Err(e) => format!("error: {}", e.0),
|
||
};
|
||
// Audit EXACTLY ONCE, at the node the operator's request first hit. The marked
|
||
// fan-out leg already returned above (never reaches here). A RELAYED hop (a
|
||
// sanction/forward to the target or leader, `x-tidal-relayed`) ran the full
|
||
// fenced protocol but must NOT re-audit — the operator's entry node records
|
||
// it, with the operator's principal, not the relaying node's.
|
||
if !is_relayed(&headers) {
|
||
node.audit_admin_outcome(&headers, "promote", &target_name, term, &outcome_str);
|
||
}
|
||
result
|
||
}
|
||
|
||
/// `POST /cluster/catchup` request body (internal).
|
||
#[derive(Serialize, Deserialize, ToSchema)]
|
||
pub struct CatchupRequest {
|
||
/// Source shard whose stream to pull from.
|
||
shard: u16,
|
||
/// First seqno to request.
|
||
from_seqno: u64,
|
||
}
|
||
|
||
/// Internal: trigger this node's catch-up pull from a source shard (m11p2).
|
||
///
|
||
/// MARKER REQUIRED. The leader's heal verb posts this to a healed follower so
|
||
/// catch-up starts immediately on a quiet cluster instead of waiting for the
|
||
/// next live ship to expose the gap. Idempotent and rate-limited: the
|
||
/// transport keeps at most one in-flight pull per source shard.
|
||
#[utoipa::path(
|
||
post,
|
||
path = "/cluster/catchup",
|
||
tag = "cluster",
|
||
request_body = CatchupRequest,
|
||
responses(
|
||
(status = 202, description = "Catch-up pull requested"),
|
||
(status = 400, description = "Missing internal marker"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn cluster_catchup(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
Query(sel): Query<ShardSelector>,
|
||
Json(req): Json<CatchupRequest>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let state = node.replica_for(sel.shard_id())?;
|
||
if !is_internal(&headers) {
|
||
return Err(ClusterAppError(ServerError::BadRequest(
|
||
"/cluster/catchup is internal; the x-tidal-internal marker is required \
|
||
(it is driven by the leader's /cluster/heal)"
|
||
.into(),
|
||
)));
|
||
}
|
||
state
|
||
.transport
|
||
.request_catchup(ShardId(req.shard), req.from_seqno.max(1));
|
||
Ok(StatusCode::ACCEPTED.into_response())
|
||
}
|
||
|
||
/// Operator-requested reseed (m11p5 §2.4): durably latch the `reseed_required`
|
||
/// marker so the snapshot reseed runs on the next boot.
|
||
///
|
||
/// This is a PER-NODE verb (a reseed is node-local — it does NOT forward to the
|
||
/// leader): the operator targets the specific node that must reseed. The marker
|
||
/// is latched with reason `operator`; the seqno the next-boot fetch resumes from
|
||
/// is this node's current applied frontier + 1 (the honest resume point). With
|
||
/// `replication.reseed_self_restart: true`, the latch may trigger a graceful
|
||
/// drain + exit (gated by the §2.4 quorum refusal). The response reports whether
|
||
/// a self-restart was initiated.
|
||
#[utoipa::path(
|
||
post,
|
||
path = "/cluster/reseed",
|
||
tag = "cluster",
|
||
responses(
|
||
(status = 202, description = "Reseed marker latched; runs on next boot"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
(status = 503, description = "Server shutting down"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn cluster_reseed(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
Query(sel): Query<ShardSelector>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let state = node.replica_for(sel.shard_id())?;
|
||
let term = state.election_term();
|
||
// The resume seqno: this node's applied frontier (against the current
|
||
// leader's shard) + 1 — the first seqno past what it has durably applied.
|
||
let from_seqno = {
|
||
let leader_shard = shard_of_region(state.current_leader().unwrap_or(state.region));
|
||
let db = state.db().map_err(ClusterAppError)?;
|
||
db.replication_state()
|
||
.applied_seqno(leader_shard)
|
||
.unwrap_or(0)
|
||
.saturating_add(1)
|
||
};
|
||
let self_restart = state.reseed_self_restart;
|
||
// Latch on the write pool: `latch_reseed_marker` may evaluate the §2.4
|
||
// quorum refusal (a blocking status poll) and must not run on the reactor.
|
||
let state_for_job = Arc::clone(&state);
|
||
let result = state
|
||
.write_pool
|
||
.submit(move || {
|
||
state_for_job.latch_reseed_marker(ReseedReason::Operator, from_seqno);
|
||
Ok::<(), ServerError>(())
|
||
})
|
||
.await
|
||
.map_err(ClusterAppError)
|
||
.map(|()| {
|
||
(
|
||
StatusCode::ACCEPTED,
|
||
Json(serde_json::json!({
|
||
"reseed_required": true,
|
||
"self_restart": self_restart,
|
||
})),
|
||
)
|
||
.into_response()
|
||
});
|
||
node.audit_admin(&headers, "reseed", "self", term, &result);
|
||
result
|
||
}
|
||
|
||
/// The outcome of [`ShardReplica::handle_remove`].
|
||
enum RemoveOutcome {
|
||
/// The member was removed; the `Removed` record committed at `version`.
|
||
Removed { version: u64 },
|
||
/// The name is not a current member (unknown or already removed) — a no-op.
|
||
NotPresent,
|
||
/// A gate refused the conf-change (not leader / incapable voters /
|
||
/// prior-change uncommitted). Retryable.
|
||
Refused(String),
|
||
}
|
||
|
||
/// One member of the cluster roster in the `/cluster/members` response.
|
||
#[derive(Serialize, Deserialize, ToSchema)]
|
||
pub struct MemberRosterEntry {
|
||
/// Permanent member id (never renumbered, never reused).
|
||
id: u16,
|
||
/// Region name.
|
||
name: String,
|
||
/// Advertised gRPC address.
|
||
grpc_addr: String,
|
||
/// Advertised HTTP address.
|
||
http_addr: String,
|
||
/// `voter`, `learner`, or `removed`.
|
||
role: String,
|
||
}
|
||
|
||
/// The `/cluster/members` response (m11p5 §3.3): the effective roster + the
|
||
/// applied conf version + per-learner promotion lag.
|
||
#[derive(Serialize, Deserialize, ToSchema)]
|
||
pub struct MembersResponse {
|
||
/// The applied membership conf version (0 = the topology era).
|
||
membership_version: u64,
|
||
/// The full roster (voters, learners, removed tombstones).
|
||
members: Vec<MemberRosterEntry>,
|
||
}
|
||
|
||
/// The cluster membership roster (m11p5 §3.3): voters, learners, and removed
|
||
/// tombstones, with the applied conf version. Served from any node's effective
|
||
/// roster (the view is record-derived in the membership era, topology-derived in
|
||
/// era 0).
|
||
#[utoipa::path(
|
||
get,
|
||
path = "/cluster/members",
|
||
tag = "cluster",
|
||
responses(
|
||
(status = 200, description = "The cluster membership roster", body = MembersResponse),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn cluster_members(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
Query(sel): Query<ShardSelector>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let state = node.replica_for(sel.shard_id())?;
|
||
let roster = state.membership().roster();
|
||
let members = roster
|
||
.members
|
||
.iter()
|
||
.map(|m| {
|
||
use tidaldb::wal::format::MemberRole;
|
||
MemberRosterEntry {
|
||
id: m.id,
|
||
name: m.name.clone(),
|
||
grpc_addr: m.grpc_addr.clone(),
|
||
http_addr: m.http_addr.clone(),
|
||
role: match m.role {
|
||
MemberRole::Voter => "voter",
|
||
MemberRole::Learner => "learner",
|
||
MemberRole::Removed => "removed",
|
||
}
|
||
.to_string(),
|
||
}
|
||
})
|
||
.collect();
|
||
Ok((
|
||
StatusCode::OK,
|
||
Json(MembersResponse {
|
||
membership_version: roster.version,
|
||
members,
|
||
}),
|
||
)
|
||
.into_response())
|
||
}
|
||
|
||
/// Remove a member from the cluster (m11p5 §3.3, leader-only). Appends a
|
||
/// `Removed` kind-4 record (gated: one-at-a-time + capability), burning the
|
||
/// member's id. A non-leader forwards to the leader. The removed peer's ship
|
||
/// cell is retired by the fenced apply path; the node learns it is removed via
|
||
/// the typed `removed` signal and flips to readiness 503.
|
||
#[utoipa::path(
|
||
post,
|
||
path = "/cluster/members/remove",
|
||
tag = "cluster",
|
||
request_body = RegionRequest,
|
||
responses(
|
||
(status = 200, description = "Member removed (Removed record committed)"),
|
||
(status = 400, description = "Unknown region or already removed"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
(status = 503, description = "Not the leader / conf-change held; retry"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn cluster_member_remove(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
Query(sel): Query<ShardSelector>,
|
||
Json(req): Json<RegionRequest>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let state = node.replica_for(sel.shard_id())?;
|
||
let term = state.election_term();
|
||
let target = req.region.clone();
|
||
// Non-leader: forward to the leader (the conf-change must run there). The
|
||
// forward carries the `?shard=` selector so the leader resolves the SAME
|
||
// group (m11p6; `admin_path` is a no-op for S=1).
|
||
if !is_internal(&headers) && !state.is_leader() {
|
||
let result = forward_write(
|
||
&state,
|
||
&state.admin_path("/cluster/members/remove"),
|
||
&req,
|
||
&headers,
|
||
)
|
||
.await;
|
||
node.audit_admin(&headers, "member_remove", &target, term, &result);
|
||
return result;
|
||
}
|
||
let result = do_remove(&state, &req.region).await;
|
||
node.audit_admin(&headers, "member_remove", &target, term, &result);
|
||
result
|
||
}
|
||
|
||
/// Leader-side member removal on a resolved group (m11p6): the write-pool
|
||
/// submit + response shape shared by `/cluster/members/remove` and the
|
||
/// `/cluster/shards/{id}/replicas` (remove) rebalance verb. The caller has
|
||
/// already resolved leadership/forwarding and audits the outcome.
|
||
async fn do_remove(
|
||
state: &Arc<ShardReplica>,
|
||
name: &str,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let name_owned = name.to_string();
|
||
let state_for_job = Arc::clone(state);
|
||
// Runs on the write pool: the append + bounded same-term commit wait blocks.
|
||
let outcome = state
|
||
.write_pool
|
||
.submit(move || Ok::<_, ServerError>(state_for_job.handle_remove(&name_owned)))
|
||
.await
|
||
.map_err(ClusterAppError)?;
|
||
match outcome {
|
||
RemoveOutcome::Removed { version } => Ok((
|
||
StatusCode::OK,
|
||
Json(serde_json::json!({ "removed": name, "membership_version": version })),
|
||
)
|
||
.into_response()),
|
||
RemoveOutcome::NotPresent => Err(ClusterAppError(ServerError::Cluster(format!(
|
||
"region '{name}' is not a current member (unknown or already removed)"
|
||
)))),
|
||
RemoveOutcome::Refused(reason) => Err(ClusterAppError(ServerError::Cluster(reason))),
|
||
}
|
||
}
|
||
|
||
/// Join the cluster over HTTP (m11p5 §3.3): the operator/manual wrapper over the
|
||
/// same leader-side logic the gRPC `JoinCluster` verb uses. A non-leader
|
||
/// forwards to the leader; the leader assigns an id, appends a Learner record,
|
||
/// waits for same-term commit, and answers with the roster.
|
||
#[utoipa::path(
|
||
post,
|
||
path = "/cluster/join",
|
||
tag = "cluster",
|
||
request_body = JoinHttpRequest,
|
||
responses(
|
||
(status = 200, description = "Joined as a Learner (or idempotent re-join)"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
(status = 503, description = "Not the leader / conf-change held; retry"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn cluster_join(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
Query(sel): Query<ShardSelector>,
|
||
Json(req): Json<JoinHttpRequest>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let state = node.replica_for(sel.shard_id())?;
|
||
let term = state.election_term();
|
||
let target = req.name.clone();
|
||
if !is_internal(&headers) && !state.is_leader() {
|
||
let result =
|
||
forward_write(&state, &state.admin_path("/cluster/join"), &req, &headers).await;
|
||
node.audit_admin(&headers, "join", &target, term, &result);
|
||
return result;
|
||
}
|
||
let ask = tidal_net::JoinAsk {
|
||
name: req.name.clone(),
|
||
grpc_addr: req.grpc_addr.clone(),
|
||
http_addr: req.http_addr.clone(),
|
||
capabilities: tidal_net::CAP_KIND4_MEMBERSHIP,
|
||
};
|
||
let result = do_join(&state, ask).await;
|
||
node.audit_admin(&headers, "join", &target, term, &result);
|
||
result
|
||
}
|
||
|
||
/// Leader-side join on a resolved group (m11p6): the write-pool submit +
|
||
/// response shape shared by `/cluster/join` and the
|
||
/// `/cluster/shards/{id}/replicas` (add) rebalance verb. The caller has already
|
||
/// resolved leadership/forwarding and audits the outcome.
|
||
async fn do_join(
|
||
state: &Arc<ShardReplica>,
|
||
ask: tidal_net::JoinAsk,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let state_for_job = Arc::clone(state);
|
||
let outcome = state
|
||
.write_pool
|
||
.submit(move || Ok::<_, ServerError>(state_for_job.handle_join(&ask)))
|
||
.await
|
||
.map_err(ClusterAppError)?;
|
||
if outcome.accepted {
|
||
Ok((
|
||
StatusCode::OK,
|
||
Json(serde_json::json!({
|
||
"accepted": true,
|
||
"assigned_id": outcome.assigned_id,
|
||
"term": outcome.term,
|
||
"membership_version": outcome.membership_version,
|
||
})),
|
||
)
|
||
.into_response())
|
||
} else {
|
||
Err(ClusterAppError(ServerError::Cluster(format!(
|
||
"join refused: {} (leader: {})",
|
||
outcome.refusal_reason, outcome.leader_region
|
||
))))
|
||
}
|
||
}
|
||
|
||
/// `POST /cluster/join` request body (the HTTP wrapper).
|
||
#[derive(Serialize, Deserialize, ToSchema)]
|
||
pub struct JoinHttpRequest {
|
||
/// The joining region's name.
|
||
#[schema(example = "ap-south")]
|
||
name: String,
|
||
/// Its advertised gRPC address.
|
||
#[schema(example = "tidaldb-3.peers.svc:9500")]
|
||
grpc_addr: String,
|
||
/// Its advertised HTTP address.
|
||
#[schema(example = "http://tidaldb-3.peers.svc:9501")]
|
||
http_addr: String,
|
||
}
|
||
|
||
/// Partition a peer from the leader (leader-side ship-skip).
|
||
///
|
||
/// Partition mutates leader-side state, so an external request on a NON-leader is
|
||
/// forwarded to the current leader (the marker terminates the forward). The
|
||
/// leader applies it on the write pool — the admin-verb lock serializes it
|
||
/// against an in-flight heal, and a blocked waiter must occupy a pool worker,
|
||
/// never a reactor thread.
|
||
#[utoipa::path(
|
||
post,
|
||
path = "/cluster/partition",
|
||
tag = "cluster",
|
||
request_body = RegionRequest,
|
||
responses(
|
||
(status = 200, description = "Peer partitioned from the leader"),
|
||
(status = 400, description = "Unknown region"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
(status = 429, description = "Write pool saturated"),
|
||
(status = 503, description = "Leader unreachable while forwarding"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn cluster_partition(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
Query(sel): Query<ShardSelector>,
|
||
Json(req): Json<RegionRequest>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let state = node.replica_for(sel.shard_id())?;
|
||
let term = state.election_term();
|
||
let target = req.region.clone();
|
||
if !is_internal(&headers) && !state.is_leader() {
|
||
let result = forward_write(
|
||
&state,
|
||
&state.admin_path("/cluster/partition"),
|
||
&req,
|
||
&headers,
|
||
)
|
||
.await;
|
||
node.audit_admin(&headers, "partition", &target, term, &result);
|
||
return result;
|
||
}
|
||
let region = req.region.clone();
|
||
let state_for_job = Arc::clone(&state);
|
||
let result = state
|
||
.write_pool
|
||
.submit(move || state_for_job.partition_peer(®ion))
|
||
.await
|
||
.map_err(ClusterAppError)
|
||
.map(|()| {
|
||
Json(serde_json::json!({ "ok": true, "partitioned": req.region })).into_response()
|
||
});
|
||
node.audit_admin(&headers, "partition", &target, term, &result);
|
||
result
|
||
}
|
||
|
||
/// Heal a peer and redeliver its missed segments over gRPC.
|
||
///
|
||
/// Heal mutates leader-side state, so an external request on a NON-leader is
|
||
/// forwarded to the current leader. The leader fetches the follower's
|
||
/// HTTP-reported applied seqno and redelivers only the batches it has not applied
|
||
/// (idempotent — a full-log fallback over-ships harmlessly).
|
||
#[utoipa::path(
|
||
post,
|
||
path = "/cluster/heal",
|
||
tag = "cluster",
|
||
request_body = RegionRequest,
|
||
responses(
|
||
(status = 200, description = "Peer healed (missed segments re-shipped)"),
|
||
(status = 400, description = "Unknown region"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
(status = 429, description = "Write pool saturated"),
|
||
(status = 503, description = "Leader unreachable while forwarding"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn cluster_heal(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
Query(sel): Query<ShardSelector>,
|
||
Json(req): Json<RegionRequest>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let state = node.replica_for(sel.shard_id())?;
|
||
let term = state.election_term();
|
||
let target = req.region.clone();
|
||
if !is_internal(&headers) && !state.is_leader() {
|
||
let result =
|
||
forward_write(&state, &state.admin_path("/cluster/heal"), &req, &headers).await;
|
||
node.audit_admin(&headers, "heal", &target, term, &result);
|
||
return result;
|
||
}
|
||
// Reject non-leader BEFORE touching the write pool so the typed NotLeader
|
||
// surfaces without consuming a worker slot.
|
||
if !state.is_leader() {
|
||
let result = Err(ClusterAppError(state.not_leader()));
|
||
node.audit_admin(&headers, "heal", &target, term, &result);
|
||
return result;
|
||
}
|
||
let region = req.region.clone();
|
||
let auth = forwarded_auth(&headers);
|
||
let state_for_job = Arc::clone(&state);
|
||
let result = state
|
||
.write_pool
|
||
.submit(move || state_for_job.heal_peer_with_reported_applied(®ion, auth.as_deref()))
|
||
.await
|
||
.map_err(ClusterAppError)
|
||
.map(|()| Json(serde_json::json!({ "ok": true, "healed": req.region })).into_response());
|
||
node.audit_admin(&headers, "heal", &target, term, &result);
|
||
result
|
||
}
|
||
|
||
// ── Rebalancing verbs (m11p6 L3) ───────────────────────────────────────────────
|
||
|
||
/// `POST /cluster/shards/{id}/replicas` request body (m11p6 rebalancing).
|
||
#[derive(Serialize, Deserialize, ToSchema)]
|
||
pub struct ShardReplicaChange {
|
||
/// `"add"` — seed a node into this group's roster as a Learner (the m11p5
|
||
/// join flow per-group; the joiner catches up via snapshot+stream and the
|
||
/// leader auto-promotes it to Voter). `"remove"` — the m11p5 fenced removal
|
||
/// of a replica from this group.
|
||
#[schema(example = "add")]
|
||
action: String,
|
||
/// The region/node name to add to or remove from this shard group.
|
||
#[schema(example = "ap-south")]
|
||
name: String,
|
||
/// The joining node's advertised gRPC address (required for `add`).
|
||
#[serde(default)]
|
||
grpc_addr: Option<String>,
|
||
/// The joining node's advertised HTTP address (required for `add`).
|
||
#[serde(default)]
|
||
http_addr: Option<String>,
|
||
}
|
||
|
||
/// Add or remove a replica of one shard group (m11p6 rebalancing).
|
||
///
|
||
/// The roadmap's `POST /cluster/shards/{id}/replicas` operator verb. It is the
|
||
/// per-group projection of the m11p5 membership flow: `add` runs the leader-side
|
||
/// join (Learner → snapshot+stream catch-up → auto-promotion to Voter) against
|
||
/// group `{id}`; `remove` runs the fenced removal on group `{id}`'s own log.
|
||
/// Both reuse the existing `/cluster/join` / `/cluster/members/remove` handlers
|
||
/// with the `?shard={id}` selector, so the non-leader forward, the conf-change
|
||
/// gate, the audit record, and the snapshot+stream cutover are the proven m11p5
|
||
/// machinery — unchanged, just scoped to one group.
|
||
#[utoipa::path(
|
||
post,
|
||
path = "/cluster/shards/{id}/replicas",
|
||
tag = "cluster",
|
||
params(("id" = u16, Path, description = "Target shard group id")),
|
||
request_body = ShardReplicaChange,
|
||
responses(
|
||
(status = 200, description = "Replica added (Learner) or removed (Removed record committed)"),
|
||
(status = 400, description = "Unknown action/region, or this node hosts no replica of the group"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
(status = 503, description = "Not the group leader / conf-change held; retry"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
pub async fn shard_replicas(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
Path(shard): Path<u16>,
|
||
headers: HeaderMap,
|
||
Json(req): Json<ShardReplicaChange>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let sel = ShardSelector { shard: Some(shard) };
|
||
match req.action.as_str() {
|
||
"add" => {
|
||
let (Some(grpc_addr), Some(http_addr)) = (req.grpc_addr, req.http_addr) else {
|
||
return Err(ClusterAppError(ServerError::BadRequest(
|
||
"shard replica 'add' requires grpc_addr and http_addr".into(),
|
||
)));
|
||
};
|
||
cluster_join(
|
||
State(node),
|
||
headers,
|
||
Query(sel),
|
||
Json(JoinHttpRequest {
|
||
name: req.name,
|
||
grpc_addr,
|
||
http_addr,
|
||
}),
|
||
)
|
||
.await
|
||
}
|
||
"remove" => {
|
||
cluster_member_remove(
|
||
State(node),
|
||
headers,
|
||
Query(sel),
|
||
Json(RegionRequest {
|
||
region: req.name,
|
||
baseline: None,
|
||
}),
|
||
)
|
||
.await
|
||
}
|
||
other => Err(ClusterAppError(ServerError::BadRequest(format!(
|
||
"unknown shard replica action '{other}' (expected 'add' or 'remove')"
|
||
)))),
|
||
}
|
||
}
|
||
|
||
/// Transfer one shard group's leadership to a named replica (m11p6 rebalancing).
|
||
///
|
||
/// The roadmap's `POST /cluster/shards/{id}/transfer` operator verb — the
|
||
/// rebalance-back-to-preferred-leader path (design §5/§6). It is the `REST`
|
||
/// projection of `/cluster/promote` scoped to group `{id}`: the same fenced
|
||
/// transfer (catch-up drain → `TimeoutNow` → term+1 election), forwarding, and
|
||
/// audit, reached by delegating to [`cluster_promote`] with the `?shard={id}`
|
||
/// selector so none of that delicate election logic is duplicated.
|
||
#[utoipa::path(
|
||
post,
|
||
path = "/cluster/shards/{id}/transfer",
|
||
tag = "cluster",
|
||
params(("id" = u16, Path, description = "Target shard group id")),
|
||
request_body = RegionRequest,
|
||
responses(
|
||
(status = 200, description = "Leadership transferred to the named replica"),
|
||
(status = 400, description = "Unknown region, or this node hosts no replica of the group"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
(status = 503, description = "Transfer did not complete; the group keeps its current leader"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
pub async fn shard_transfer(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
Path(shard): Path<u16>,
|
||
headers: HeaderMap,
|
||
Json(req): Json<RegionRequest>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let sel = ShardSelector { shard: Some(shard) };
|
||
cluster_promote(State(node), headers, Query(sel), Json(req))
|
||
.await
|
||
.map(IntoResponse::into_response)
|
||
}
|
||
|
||
// ── Data routes ──────────────────────────────────────────────────────────────
|
||
|
||
/// Create an item on the cluster (m11p2: items ride the one replicated log).
|
||
///
|
||
/// Routing:
|
||
/// * **non-leader, external request**: forward to the leader and relay its
|
||
/// status.
|
||
/// * **leader (external OR forwarded)**: the engine journals the item as a
|
||
/// kind-1 WAL record FIRST — that record is what replicates it to every
|
||
/// follower (live push + catch-up stream) — then upserts local storage.
|
||
/// There is NO HTTP fan-out anymore: the marker-gated broadcast and its
|
||
/// bug class are gone.
|
||
/// * **non-leader with the marker**: a stale forward — surfaces the typed
|
||
/// `NotLeader` so the client re-targets (followers never apply items from
|
||
/// HTTP; they apply them from the log).
|
||
#[utoipa::path(
|
||
post,
|
||
path = "/items",
|
||
tag = "data",
|
||
request_body = ItemRequest,
|
||
responses(
|
||
(status = 201, description = "Item durably journaled and applied on the leader"),
|
||
(status = 400, description = "Invalid request"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
(status = 503, description = "Leader unreachable while forwarding"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn create_item(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
Json(req): Json<ItemRequest>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
// m11p6: hash-route to the replica that owns this entity's shard group.
|
||
// Continue ⇒ the existing leader/forward path below; Break ⇒ this node hosts
|
||
// no replica of the group, so the relayed cross-node response is returned.
|
||
let state = match node
|
||
.route_or_forward(req.entity_id, "/items", &req, &headers)
|
||
.await
|
||
{
|
||
std::ops::ControlFlow::Continue(replica) => replica,
|
||
std::ops::ControlFlow::Break(resp) => return resp,
|
||
};
|
||
// Non-leader external request → forward to the leader and relay verbatim.
|
||
if !is_internal(&headers) && !state.is_leader() {
|
||
return forward_write(&state, "/items", &req, &headers).await;
|
||
}
|
||
// A marked (forwarded) write must land on the leader; a stale cluster
|
||
// view surfaces honestly instead of applying off-log. A leader that has begun
|
||
// a graceful shutdown also stops accepting writes (typed `NotLeader`, so the
|
||
// client re-targets): this FREEZES the flushed frontier so the step-down drain
|
||
// (`drain_committed_before_stepdown`) converges instead of chasing a tail that
|
||
// keeps growing under load — the rolling-restart-under-load divergence fix.
|
||
if !state.is_leader() || state.is_shutting_down() {
|
||
return Err(ClusterAppError(state.not_leader()));
|
||
}
|
||
let ack = state.ack_mode_for(&headers).map_err(ClusterAppError)?;
|
||
let db = state.db_arc().map_err(ClusterAppError)?;
|
||
let entity = EntityId::new(req.entity_id);
|
||
let metadata = req.metadata.clone();
|
||
// Admit through the bounded write pool, NOT the unbounded `spawn_blocking`
|
||
// pool (m11p5 forward-stall fix). The kind-1 WAL journal + storage upsert is
|
||
// the same slow, runtime-free CPU/IO work the `/signals` staging path admits
|
||
// here: a 1536-dim HNSW apply at ef_construction=400 on a 4-vCPU node is
|
||
// ~tens of ms, so under a sustained forwarded write-burst an UNBOUNDED queue
|
||
// grows past the follower's 15s forward budget and the forward times out into
|
||
// a 503 (declaring a slow-but-alive leader "unreachable"). The bounded pool
|
||
// sheds a sustained overload as a fast `Backpressure` (→ 429) the forwarder
|
||
// retries with backoff, so ingest slows but COMPLETES instead of failing.
|
||
// Correctness is unchanged: this is the identical leader-side `apply_item_local`
|
||
// (kind-1 WAL-first journal) the gateway forwarded here — only the thread it
|
||
// runs on changes; the quorum gate below is untouched.
|
||
let seq = state
|
||
.write_pool
|
||
.submit(move || ShardReplica::apply_item_local(&db, entity, &metadata))
|
||
.await
|
||
.map_err(ClusterAppError)?;
|
||
if ack == AckMode::Quorum
|
||
&& let Some(seq) = seq
|
||
{
|
||
await_quorum(&state, seq).await?;
|
||
} else if let Some(seq) = seq {
|
||
// ack=leader: the write is acked on journal alone (no quorum), so it may
|
||
// be un-replicated — record it on the durable leader-acked frontier (the
|
||
// election-divergence divergent-suffix signal).
|
||
state.note_leader_acked(seq);
|
||
}
|
||
Ok(with_seq_header(StatusCode::CREATED, seq))
|
||
}
|
||
|
||
/// Write an item embedding on the cluster (m11p2: kind-2 WAL record — same
|
||
/// routing and replication model as [`create_item`]).
|
||
#[utoipa::path(
|
||
post,
|
||
path = "/embeddings",
|
||
tag = "data",
|
||
request_body = EmbeddingRequest,
|
||
responses(
|
||
(status = 204, description = "Embedding durably journaled and applied on the leader"),
|
||
(status = 400, description = "Invalid request"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
(status = 503, description = "Leader unreachable while forwarding"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn write_embedding(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
Json(req): Json<EmbeddingRequest>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let state = match node
|
||
.route_or_forward(req.entity_id, "/embeddings", &req, &headers)
|
||
.await
|
||
{
|
||
std::ops::ControlFlow::Continue(replica) => replica,
|
||
std::ops::ControlFlow::Break(resp) => return resp,
|
||
};
|
||
if !is_internal(&headers) && !state.is_leader() {
|
||
return forward_write(&state, "/embeddings", &req, &headers).await;
|
||
}
|
||
// A leader mid-graceful-shutdown stops accepting writes (freezes flushed so the
|
||
// step-down drain converges — see `create_item`).
|
||
if !state.is_leader() || state.is_shutting_down() {
|
||
return Err(ClusterAppError(state.not_leader()));
|
||
}
|
||
let ack = state.ack_mode_for(&headers).map_err(ClusterAppError)?;
|
||
let db = state.db_arc().map_err(ClusterAppError)?;
|
||
let entity = EntityId::new(req.entity_id);
|
||
let values = req.values.clone();
|
||
// Admit through the bounded write pool (m11p5 forward-stall fix). The HNSW
|
||
// insert (ef_construction=400, 1536-dim) is the slowest apply on the hot
|
||
// path; running it on the UNBOUNDED `spawn_blocking` pool let a sustained
|
||
// forwarded write-burst queue without limit until the follower's 15s forward
|
||
// budget elapsed (→ 503 "leader unreachable" on a slow-but-alive leader).
|
||
// The bounded pool degrades sustained overload to a fast `Backpressure`
|
||
// (→ 429) the forwarder retries with backoff. Correctness is unchanged: this
|
||
// is the identical leader-side `apply_embedding_local` (kind-2 WAL-first
|
||
// journal) the gateway forwarded here, only the thread it runs on changes;
|
||
// the quorum gate below is untouched.
|
||
let seq = state
|
||
.write_pool
|
||
.submit(move || ShardReplica::apply_embedding_local(&db, entity, &values))
|
||
.await
|
||
.map_err(ClusterAppError)?;
|
||
if ack == AckMode::Quorum
|
||
&& let Some(seq) = seq
|
||
{
|
||
await_quorum(&state, seq).await?;
|
||
} else if let Some(seq) = seq {
|
||
// ack=leader: journal-only ack (may be un-replicated) — record it on the
|
||
// durable leader-acked frontier (the election-divergence signal).
|
||
state.note_leader_acked(seq);
|
||
}
|
||
Ok(with_seq_header(StatusCode::NO_CONTENT, seq))
|
||
}
|
||
|
||
/// Record a signal on the cluster.
|
||
///
|
||
/// * **non-leader, external**: forward to the leader and relay its status/body.
|
||
/// * **leader (or marked forward)**: durably apply on the leader (storage + WAL
|
||
/// fsync) and best-effort eager-ship to siblings over gRPC. The signal is the
|
||
/// replicated stream, so there is NO HTTP broadcast — followers receive it via
|
||
/// the WAL relay. A 204 asserts leader durability only.
|
||
///
|
||
/// The 204 carries `x-tidal-seq` (the write's replicated-log seqno) — or
|
||
/// `x-tidal-deduplicated: 1` on the rare write suppressed by the WAL's
|
||
/// content-hash dedup window (an identical record is already durable; no new
|
||
/// log entry exists to name).
|
||
#[utoipa::path(
|
||
post,
|
||
path = "/signals",
|
||
tag = "data",
|
||
request_body = SignalRequest,
|
||
responses(
|
||
(status = 204, description = "Signal durably applied on the leader"),
|
||
(status = 400, description = "Invalid request"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
(status = 429, description = "Write pool saturated"),
|
||
(status = 503, description = "Leader unreachable while forwarding"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn write_signal(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
Json(req): Json<SignalRequest>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
// m11p6: hash-route the signal to the replica owning its entity's group.
|
||
let state = match node
|
||
.route_or_forward(req.entity_id, "/signals", &req, &headers)
|
||
.await
|
||
{
|
||
std::ops::ControlFlow::Continue(replica) => replica,
|
||
std::ops::ControlFlow::Break(resp) => return resp,
|
||
};
|
||
// A non-leader external request forwards to the leader. A marked request
|
||
// (forwarded here) must be on the leader, or the cluster has a stale view —
|
||
// fall through to the leader check, which surfaces NotLeader honestly.
|
||
if !is_internal(&headers) && !state.is_leader() {
|
||
return forward_write(&state, "/signals", &req, &headers).await;
|
||
}
|
||
// Reject non-leader BEFORE the write pool so NotLeader does not consume a slot.
|
||
// A leader mid-graceful-shutdown also stops accepting writes (freezes the
|
||
// flushed frontier so the step-down drain converges — see `create_item`).
|
||
if !state.is_leader() || state.is_shutting_down() {
|
||
return Err(ClusterAppError(state.not_leader()));
|
||
}
|
||
let signal = req.signal;
|
||
let entity = EntityId::new(req.entity_id);
|
||
let weight = req.weight;
|
||
// Two-phase write (m11p1): STAGE on the write pool (microseconds; the
|
||
// bounded queue keeps the 429 admission semantics), then COMPLETE — the
|
||
// group-commit fsync wait — on the blocking pool, freeing the pool worker
|
||
// so concurrent writers coalesce into shared fsyncs instead of each
|
||
// serializing a solo batch-timeout. Follower shipping is fully off this
|
||
// path (the ship queue's sender threads own it); the 204 still asserts
|
||
// exactly what it always did: leader durability.
|
||
//
|
||
// The staged write travels inside a `StagedWriteTicket` constructed in the
|
||
// pool job itself, so EVERY cancellation window is covered: if this
|
||
// request future is dropped at either await (the pool worker's oneshot
|
||
// send fails, or the completion future is never polled), the ticket's
|
||
// Drop completes the write on a detached thread — an orphaned staged
|
||
// seqno would otherwise stall the durable frontier (and all shipping)
|
||
// forever.
|
||
let ack = state.ack_mode_for(&headers).map_err(ClusterAppError)?;
|
||
let state_for_job = Arc::clone(&state);
|
||
let ticket = state
|
||
.write_pool
|
||
.submit(move || {
|
||
let staged = state_for_job.stage_signal_local(&signal, entity, weight)?;
|
||
Ok(StagedWriteTicket::new(staged, Arc::clone(&state_for_job)))
|
||
})
|
||
.await
|
||
.map_err(ClusterAppError)?;
|
||
let seq = offload_region_read(move || ticket.complete()).await?;
|
||
// Quorum gate (m11p3): the write is leader-durable; now block until a
|
||
// majority of the replica set durably holds it. The dedup sentinel (0)
|
||
// skips the gate — an identical record is already durable, and ITS ack
|
||
// covered quorum (this request created no new log entry to gate on).
|
||
if ack == AckMode::Quorum && seq > 0 {
|
||
await_quorum(&state, seq).await?;
|
||
} else if ack == AckMode::Leader && seq > 0 {
|
||
// ack=leader: journal-only ack (may be un-replicated) — record it on the
|
||
// durable leader-acked frontier (the election-divergence signal).
|
||
state.note_leader_acked(seq);
|
||
}
|
||
Ok(with_seq_header(StatusCode::NO_CONTENT, Some(seq)))
|
||
}
|
||
|
||
/// `POST /hardnegs` request body.
|
||
#[derive(Serialize, Deserialize, ToSchema)]
|
||
pub struct HardNegRequest {
|
||
/// User who is hiding the item.
|
||
#[schema(example = 42)]
|
||
user_id: u64,
|
||
/// Item being hidden.
|
||
#[schema(example = 7)]
|
||
item_id: u64,
|
||
}
|
||
|
||
/// Record a hide hard-negative for `(user, item)`.
|
||
///
|
||
/// * **non-leader, external**: forward to the leader (so the durable
|
||
/// `Tag::HardNeg` row lands on the leader's store, where the reconcile snapshot
|
||
/// reads it). Cross-node convergence is then the `/cluster/reconcile` CRDT path.
|
||
/// * **leader (or marked forward)**: record locally. There is NO broadcast — hard
|
||
/// negatives converge via the LWW-resolved reconcile snapshot, not a fan-out.
|
||
#[utoipa::path(
|
||
post,
|
||
path = "/hardnegs",
|
||
tag = "data",
|
||
request_body = HardNegRequest,
|
||
responses(
|
||
(status = 204, description = "Hard-negative recorded (on the leader)"),
|
||
(status = 400, description = "Invalid request"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
(status = 503, description = "Leader unreachable while forwarding"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn write_hardneg(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
Json(req): Json<HardNegRequest>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
// m11p6: a hard-negative co-locates with the ITEM it hides (entity-sharded),
|
||
// so the item's shard ranks-and-filters it; route by item_id.
|
||
let state = match node
|
||
.route_or_forward(req.item_id, "/hardnegs", &req, &headers)
|
||
.await
|
||
{
|
||
std::ops::ControlFlow::Continue(replica) => replica,
|
||
std::ops::ControlFlow::Break(resp) => return resp,
|
||
};
|
||
if !is_internal(&headers) && !state.is_leader() {
|
||
return forward_write(&state, "/hardnegs", &req, &headers).await;
|
||
}
|
||
let state_for_job = Arc::clone(&state);
|
||
let user_id = req.user_id;
|
||
let item_id = req.item_id;
|
||
// The durable Tag::HardNeg write is a blocking storage call; offload it.
|
||
offload_region_read(move || state_for_job.record_hardneg_local(user_id, item_id)).await?;
|
||
Ok(StatusCode::NO_CONTENT.into_response())
|
||
}
|
||
|
||
// ── Forwarding / broadcast helpers ───────────────────────────────────────────
|
||
|
||
/// Forward an external write to the current leader and relay its status + body.
|
||
///
|
||
/// Passes the caller's `Authorization` through verbatim and sets the internal
|
||
/// marker so the leader applies the write locally and does NOT re-forward. A
|
||
/// leader that cannot be reached degrades to a 503 JSON body naming the leader
|
||
/// and the connect error (never a hang, never a silent drop).
|
||
/// Await the commit index covering `seq` (m11p3 `ack=quorum`): a majority of
|
||
/// the replica set has durably applied the write. Returns the commit index
|
||
/// that satisfied the wait.
|
||
///
|
||
/// Fully async — each waiter holds a watch receiver, never a thread. The
|
||
/// write is already durable on THIS leader; a timeout means the durability
|
||
/// claim could not be confirmed in budget — the retryable 503 names the
|
||
/// laggards, and the write MAY still commit (at-least-once on retry; see
|
||
/// runbook §8). A leadership change mid-wait (epoch bump / deactivation)
|
||
/// fails with `NotLeader`: a demoted leader must never claim quorum.
|
||
async fn await_quorum(
|
||
state: &Arc<ShardReplica>,
|
||
seq: u64,
|
||
) -> std::result::Result<u64, ClusterAppError> {
|
||
if state.commit.needed_peers() == 0 {
|
||
// Single-replica: the leader alone is the majority — but only while
|
||
// it still leads (the demoted-leader invariant has no replica-count
|
||
// exception; mirrors `CommitIndex::wait_for`).
|
||
let (_, _, active) = state.commit.snapshot();
|
||
if !active {
|
||
return Err(ClusterAppError(state.not_leader()));
|
||
}
|
||
state.set_frontier_gauges();
|
||
return Ok(seq);
|
||
}
|
||
let mut watch = state.commit_watch.clone();
|
||
let deadline = tokio::time::Instant::now() + state.quorum_timeout;
|
||
let entry_epoch = {
|
||
let (epoch, _, active) = *watch.borrow_and_update();
|
||
if !active {
|
||
return Err(ClusterAppError(state.not_leader()));
|
||
}
|
||
epoch
|
||
};
|
||
loop {
|
||
{
|
||
let (epoch, commit, active) = *watch.borrow_and_update();
|
||
if epoch != entry_epoch || !active {
|
||
return Err(ClusterAppError(state.not_leader()));
|
||
}
|
||
if commit >= seq {
|
||
// A satisfied wait is the one moment both halves of the
|
||
// frontier pair are known: publish them together. Signal
|
||
// writes are not the only writers, so leaving this to
|
||
// `complete_signal_write` left the pair stale on an
|
||
// item-or-embedding workload.
|
||
state.set_frontier_gauges();
|
||
return Ok(commit);
|
||
}
|
||
}
|
||
match tokio::time::timeout_at(deadline, watch.changed()).await {
|
||
Ok(Ok(())) => {}
|
||
Ok(Err(_)) => {
|
||
// The bridge sender is gone. Expected on a clean shutdown;
|
||
// anything else means the 'tidal-commit-watch' thread died —
|
||
// say so loudly, or its quorum 503s get chased as a network
|
||
// problem at 3am.
|
||
if !state.is_shutting_down() {
|
||
tracing::warn!(
|
||
seq,
|
||
"commit-watch bridge disconnected outside shutdown; \
|
||
check for a panic in the 'tidal-commit-watch' thread"
|
||
);
|
||
}
|
||
return Err(ClusterAppError(ServerError::Unavailable(
|
||
"server shutting down".into(),
|
||
)));
|
||
}
|
||
Err(_) => {
|
||
// Deadline expired — but re-read the INDEX (not the watch
|
||
// mirror: the index is fresher by the bridge's republish
|
||
// latency) once before erroring: the commit can pass `seq`
|
||
// between the timer firing and this arm running, and a 503
|
||
// for a write that IS quorum-committed would be a false
|
||
// negative the caller then retries at-least-once. The same
|
||
// final read keeps the error's laggard list as fresh as a
|
||
// snapshot can be.
|
||
let (epoch, commit, active) = state.commit.snapshot();
|
||
if epoch == entry_epoch && active && commit >= seq {
|
||
return Ok(commit);
|
||
}
|
||
return Err(ClusterAppError(state.quorum_timeout_error(seq)));
|
||
}
|
||
}
|
||
}
|
||
}
|
||
|
||
/// Folds follower `ReportApplied` durable marks into the quorum commit
|
||
/// index (m11p3). The index itself enforces monotonicity, unknown-peer
|
||
/// rejection, and leadership epochs.
|
||
struct CommitIndexSink {
|
||
commit: Arc<CommitIndex>,
|
||
}
|
||
|
||
impl tidal_net::sources::AppliedSink for CommitIndexSink {
|
||
fn peer_applied(&self, peer: ShardId, applied: u64, reporter_term: u64) {
|
||
// m11p4 (design-review C4/C8/C12): fold ONLY a report stamped with
|
||
// the term this index was activated with — the check runs under the
|
||
// index's own lock, so a leadership change can never race the fold.
|
||
// Term-0 reports against a term-0 activation pass (the topology era
|
||
// and pre-m11p4 reporters).
|
||
let _ = self
|
||
.commit
|
||
.update_peer_for_term(peer, applied, reporter_term);
|
||
}
|
||
}
|
||
|
||
/// Node-side consumer of a typed `snapshot-required` catch-up refusal (m11p5
|
||
/// §2.4): latch the durable reseed marker so the reseed runs on the next boot.
|
||
///
|
||
/// Holds a `Weak<ShardReplica>` (set once the node is in its final Arc),
|
||
/// so the transport's catch-up path can never keep the node alive past
|
||
/// shutdown. The retry timer keeps its standing wake-up regardless (the m11p4
|
||
/// re-arm-on-skip liveness fix) — this only latches the marker; it does not
|
||
/// touch the running engine.
|
||
/// The leader-side `JoinCluster` adapter (m11p5 §3.3): bridges the gRPC
|
||
/// `JoinHooks` trait to [`ShardReplica::handle_join`]. A `Weak` so the
|
||
/// runtime cannot leak the node.
|
||
struct NodeJoinHooks {
|
||
node: Weak<ShardReplica>,
|
||
}
|
||
|
||
impl tidal_net::JoinHooks for NodeJoinHooks {
|
||
fn join(&self, ask: tidal_net::JoinAsk) -> tidal_net::JoinOutcome {
|
||
let Some(node) = self.node.upgrade() else {
|
||
// The node is shutting down; refuse with no leader hint (the joiner
|
||
// retries the next seed).
|
||
return tidal_net::JoinOutcome {
|
||
accepted: false,
|
||
refusal_reason: "node shutting down".to_string(),
|
||
assigned_id: 0,
|
||
term: 0,
|
||
leader_region: String::new(),
|
||
leader_grpc_addr: String::new(),
|
||
leader_http_addr: String::new(),
|
||
members: Vec::new(),
|
||
membership_version: 0,
|
||
};
|
||
};
|
||
node.handle_join(&ask)
|
||
}
|
||
}
|
||
|
||
struct NodeSnapshotRequiredSink {
|
||
node: Weak<ShardReplica>,
|
||
}
|
||
|
||
impl tidal_net::sources::SnapshotRequiredSink for NodeSnapshotRequiredSink {
|
||
fn snapshot_required(&self, _shard: ShardId, from_seqno: u64) {
|
||
let Some(node) = self.node.upgrade() else {
|
||
return; // node shutting down; nothing to latch
|
||
};
|
||
// The stream refused us at `from_seqno`, so that is exactly the seqno the
|
||
// next-boot snapshot fetch must request from (§2.7 honesty — never
|
||
// re-derive a frontier from a half-open engine). Latching is idempotent.
|
||
node.latch_reseed_marker(ReseedReason::SnapshotRequired, from_seqno);
|
||
}
|
||
}
|
||
|
||
struct NodeCatchupServedSink {
|
||
node: Weak<ShardReplica>,
|
||
}
|
||
|
||
impl tidal_net::sources::CatchupServedSink for NodeCatchupServedSink {
|
||
fn catchup_served(&self, _shard: ShardId, from_seqno: u64) {
|
||
let Some(node) = self.node.upgrade() else {
|
||
return; // node shutting down; nothing to discharge
|
||
};
|
||
// A pull that STARTED at `from_seqno` ran to completion, so the source
|
||
// streamed that range and the receiver applied it. This is the positive
|
||
// evidence a reseed marker needs — the only alternative being a snapshot
|
||
// install. Never discharge from a frontier comparison; see
|
||
// `ShardReplica::discharge_reseed_marker_if_served`.
|
||
node.discharge_reseed_marker_if_served(from_seqno);
|
||
}
|
||
}
|
||
|
||
/// Build a write-success response carrying the `x-tidal-seq` header (the
|
||
/// write's replicated-log seqno, m11p3). The dedup sentinel (`Some(0)` — an
|
||
/// identical record is already durable, no new log entry exists) carries
|
||
/// `x-tidal-deduplicated: 1` instead, so a caller tracking durability
|
||
/// cursors can tell "suppressed as duplicate" from "no seqno surface"
|
||
/// (`None` — not journaled, e.g. outside cluster mode).
|
||
fn with_seq_header(status: StatusCode, seq: Option<u64>) -> Response {
|
||
let mut resp = status.into_response();
|
||
match seq {
|
||
Some(seq) if seq > 0 => {
|
||
if let Ok(value) = axum::http::HeaderValue::from_str(&seq.to_string()) {
|
||
resp.headers_mut().insert(
|
||
axum::http::HeaderName::from_static(forward::SEQ_HEADER),
|
||
value,
|
||
);
|
||
}
|
||
}
|
||
Some(_) => {
|
||
resp.headers_mut().insert(
|
||
axum::http::HeaderName::from_static(forward::DEDUP_HEADER),
|
||
axum::http::HeaderValue::from_static(forward::DEDUP_HEADER_VALUE),
|
||
);
|
||
}
|
||
None => {}
|
||
}
|
||
resp
|
||
}
|
||
|
||
/// Total number of forward attempts (1 initial + retries) before a slow or
|
||
/// backpressured leader is surfaced as a 503 to the client (m11p5 forward-stall
|
||
/// fix).
|
||
///
|
||
/// Under a sustained 1536-dim quorum write-burst the leader's apply plane runs
|
||
/// hot: a single forward can either time out at [`forward::FORWARD_REQUEST_TIMEOUT`]
|
||
/// (15s — a slow-but-alive leader) OR return a fast `429` (the bounded write
|
||
/// pool shedding a momentary overload). The OLD path declared either one
|
||
/// "leader unreachable" on the FIRST miss → a 503 the targeted-follower client
|
||
/// saw as a hard failure, stalling ingest. Retrying lets a slow leader be
|
||
/// WAITED-ON and a backpressured leader be retried-after-backoff, so a sustained
|
||
/// load directed at a follower DEGRADES (slower) rather than FAILS.
|
||
///
|
||
/// Retry is correctness-safe for the data-write forwards this guards
|
||
/// (`/items`, `/embeddings`): those are idempotent by-`entity_id` upserts (blobs
|
||
/// skip the WAL content-hash dedup window — a re-applied identical write
|
||
/// overwrites the same entity slot with the same value on every replica, so it
|
||
/// neither loses an acked write nor double-counts), and the forward carries the
|
||
/// internal marker so the leader applies locally without re-fanning out.
|
||
const FORWARD_MAX_ATTEMPTS: u32 = 3;
|
||
|
||
/// Backoff after a leader-side `429` (bounded write pool saturated). Honored
|
||
/// from the leader's `retry_after_ms` hint when present; this is the floor when
|
||
/// it is absent. Short by design — the pool drains in worker-thread time, so a
|
||
/// brief park then retry lands the write as soon as a slot frees.
|
||
const FORWARD_BACKPRESSURE_BACKOFF_MS: u64 = 50;
|
||
|
||
/// Backoff after a forward TIMEOUT (slow-but-alive leader). Larger than the
|
||
/// backpressure backoff: a timeout means the leader's apply plane is genuinely
|
||
/// saturated for seconds, so a longer park before re-dialing avoids piling more
|
||
/// concurrent forwards onto an already-starved leader.
|
||
const FORWARD_TIMEOUT_BACKOFF_MS: u64 = 250;
|
||
|
||
/// Wall-clock budget for the WHOLE forward (across all attempts), kept just
|
||
/// under the protected-route [`crate::router::REQUEST_TIMEOUT_SECS`] (30s) so the
|
||
/// gateway surfaces its OWN typed 503 ("leader unreachable") rather than letting
|
||
/// the outer `TimeoutLayer` cut the request into an opaque 408 mid-retry. A
|
||
/// fast-429 retry loop never approaches this; it only bites when consecutive
|
||
/// 15s timeouts would otherwise overrun the route budget — at which point we
|
||
/// stop retrying and surface the honest 503.
|
||
const FORWARD_TOTAL_BUDGET: std::time::Duration = std::time::Duration::from_secs(28);
|
||
|
||
/// Parse the `retry_after_ms` hint from a leader's `429` body, falling back to
|
||
/// [`FORWARD_BACKPRESSURE_BACKOFF_MS`]. Capped so a hostile/garbled hint cannot
|
||
/// park the forward longer than a single forward-timeout budget.
|
||
fn backpressure_backoff_ms(body: &serde_json::Value) -> u64 {
|
||
body.get("retry_after_ms")
|
||
.and_then(serde_json::Value::as_u64)
|
||
.unwrap_or(FORWARD_BACKPRESSURE_BACKOFF_MS)
|
||
.clamp(FORWARD_BACKPRESSURE_BACKOFF_MS, 1_000)
|
||
}
|
||
|
||
/// What the forward retry loop should do after one attempt — the pure policy,
|
||
/// extracted so it is unit-testable without standing up a real cluster.
|
||
#[derive(Debug, PartialEq, Eq)]
|
||
enum ForwardStep {
|
||
/// Hand the peer's response straight back to the client (2xx, or a terminal
|
||
/// 4xx/5xx the leader owns — including a final-attempt 429 relayed honestly).
|
||
Relay,
|
||
/// Retry after parking `backoff_ms`: a transient leader-backpressure (429)
|
||
/// or a slow-but-alive leader (transport timeout) with budget remaining.
|
||
Retry { backoff_ms: u64 },
|
||
/// Give up: every attempt is exhausted, or the retry budget cannot fit
|
||
/// another full attempt. Surface the typed 503 naming the leader.
|
||
Fail,
|
||
}
|
||
|
||
/// Decide the next action after one forward attempt (pure; see [`ForwardStep`]).
|
||
///
|
||
/// * A relayed `429` with attempts AND budget left ⇒ `Retry` (the leader shed
|
||
/// the write fast; no log entry was created, so re-forwarding is safe).
|
||
/// * Any other `Ok(status)` ⇒ `Relay` (the leader produced a verdict — 2xx, a
|
||
/// terminal error, or a final-attempt 429 the client should see as retryable).
|
||
/// * A transport error (timeout) with attempts AND room for another full
|
||
/// attempt+backoff inside [`FORWARD_TOTAL_BUDGET`] ⇒ `Retry`.
|
||
/// * Otherwise ⇒ `Fail` (genuinely unreachable / starved past the budget).
|
||
fn classify_forward_attempt(
|
||
outcome: &std::result::Result<forward::ForwardedResponse, String>,
|
||
attempt: u32,
|
||
elapsed: std::time::Duration,
|
||
) -> ForwardStep {
|
||
let attempts_left = attempt < FORWARD_MAX_ATTEMPTS;
|
||
match outcome {
|
||
Ok(resp) if resp.status == StatusCode::TOO_MANY_REQUESTS && attempts_left => {
|
||
ForwardStep::Retry {
|
||
backoff_ms: backpressure_backoff_ms(&resp.body),
|
||
}
|
||
}
|
||
Ok(_) => ForwardStep::Relay,
|
||
Err(_)
|
||
if attempts_left
|
||
&& elapsed
|
||
+ forward::FORWARD_REQUEST_TIMEOUT
|
||
+ std::time::Duration::from_millis(FORWARD_TIMEOUT_BACKOFF_MS)
|
||
<= FORWARD_TOTAL_BUDGET =>
|
||
{
|
||
ForwardStep::Retry {
|
||
backoff_ms: FORWARD_TIMEOUT_BACKOFF_MS,
|
||
}
|
||
}
|
||
Err(_) => ForwardStep::Fail,
|
||
}
|
||
}
|
||
|
||
async fn forward_write<B: serde::Serialize + Sync + ?Sized>(
|
||
state: &Arc<ShardReplica>,
|
||
path: &str,
|
||
body: &B,
|
||
headers: &HeaderMap,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let Some(leader_http) = state.leader_http_addr() else {
|
||
// No leader address: this node thinks it leads (race) or the leader has
|
||
// no http_addr. Surface the typed NotLeader so the client re-targets.
|
||
return Err(ClusterAppError(state.not_leader()));
|
||
};
|
||
let url = peer_url(&leader_http, path);
|
||
let auth = forwarded_auth(headers);
|
||
// The caller's ack-mode override travels WITH the write (m11p3): the
|
||
// leader honors the caller's choice, not this gateway's default. The m11p7
|
||
// node token rides alongside so the leader's marker guard sees a verified
|
||
// sibling (this forward sets the internal marker).
|
||
let mut passthrough = forward::ack_passthrough(headers);
|
||
passthrough.extend(state.node_token_passthrough());
|
||
state.cluster_metrics.incr_forwards();
|
||
|
||
// Bounded retry loop (m11p5): a slow leader (forward timeout) or a
|
||
// momentarily-saturated leader (relayed 429) is RETRIED with backoff before
|
||
// it is declared unreachable, so a sustained forwarded write-burst completes
|
||
// (slower) instead of stalling on the first miss. See FORWARD_MAX_ATTEMPTS
|
||
// for the correctness/idempotency argument.
|
||
let started = std::time::Instant::now();
|
||
let mut last_err: Option<String> = None;
|
||
for attempt in 1..=FORWARD_MAX_ATTEMPTS {
|
||
let outcome = forward_json_with_headers(
|
||
&state.client,
|
||
&url,
|
||
body,
|
||
auth.as_deref(),
|
||
true,
|
||
&passthrough,
|
||
)
|
||
.await;
|
||
match classify_forward_attempt(&outcome, attempt, started.elapsed()) {
|
||
// Relay the leader's seq/dedup headers so the original caller sees
|
||
// the write's replicated-log verdict through the forward (shared with
|
||
// the cross-shard gateway hop via `forward::relay_forwarded`). A 429
|
||
// on the FINAL attempt is relayed verbatim too — the client gets a
|
||
// retryable 429 (honest backpressure), not a misleading 503.
|
||
ForwardStep::Relay => {
|
||
// Safe: `Relay` is only chosen for `Ok(resp)` outcomes.
|
||
let resp = outcome.expect("classify_forward_attempt: Relay implies Ok");
|
||
return Ok(forward::relay_forwarded(resp));
|
||
}
|
||
// A slow leader (transport timeout) or a momentarily-saturated leader
|
||
// (relayed 429, write shed before any log entry — safe to retry):
|
||
// park for the chosen backoff, then re-forward.
|
||
ForwardStep::Retry { backoff_ms } => {
|
||
match &outcome {
|
||
Ok(_) => tracing::debug!(
|
||
leader = %state.leader_name(), %url, attempt, backoff_ms,
|
||
"forward hit leader backpressure (429); retrying after backoff"
|
||
),
|
||
Err(e) => {
|
||
tracing::debug!(
|
||
leader = %state.leader_name(), %url, attempt, error = %e,
|
||
"forward to leader timed out / failed; retrying after backoff"
|
||
);
|
||
last_err = Some(e.clone());
|
||
}
|
||
}
|
||
tokio::time::sleep(std::time::Duration::from_millis(backoff_ms)).await;
|
||
}
|
||
// Genuinely unreachable (or starved past the whole retry budget):
|
||
// remember the cause and fall through to the typed 503 below.
|
||
ForwardStep::Fail => {
|
||
if let Err(e) = outcome {
|
||
last_err = Some(e);
|
||
}
|
||
}
|
||
}
|
||
}
|
||
|
||
// Every attempt exhausted on transport failure: the leader is genuinely
|
||
// unreachable (or starved past the whole retry budget). Surface the typed
|
||
// 503 naming the leader, its address, and the last connect error.
|
||
state.cluster_metrics.incr_forward_failures();
|
||
let leader = state.leader_name();
|
||
let cause = last_err.unwrap_or_else(|| "forward exhausted retries".to_owned());
|
||
tracing::warn!(%leader, %url, error = %cause, "forward to leader failed; leader unreachable");
|
||
Err(ClusterAppError(ServerError::LeaderUnreachable {
|
||
leader,
|
||
http_addr: leader_http,
|
||
cause,
|
||
}))
|
||
}
|
||
|
||
/// Best-effort broadcast of an item/embedding write to every peer with the
|
||
/// internal marker set. Returns the per-peer success/failure outcome.
|
||
async fn broadcast_to_peers<B: serde::Serialize + Sync>(
|
||
state: &Arc<ShardReplica>,
|
||
path: &str,
|
||
body: &B,
|
||
headers: &HeaderMap,
|
||
) -> forward::BroadcastOutcome {
|
||
let peers = state.peers_named();
|
||
let auth = forwarded_auth(headers);
|
||
let node_token = state.mint_node_token();
|
||
broadcast_marked(
|
||
&state.client,
|
||
peers,
|
||
path,
|
||
body,
|
||
auth.as_deref(),
|
||
node_token.as_deref(),
|
||
state.broadcast_peer_timeout,
|
||
)
|
||
.await
|
||
}
|
||
|
||
/// Scatter a corpus-wide read over the hosted shard groups and merge: run
|
||
/// `per_db` on each group's [`TidalDb`], **SUM** `total_candidates` (entity-
|
||
/// sharded groups own disjoint key subsets — no dedup-by-max would undercount),
|
||
/// then score-sort descending and truncate to `limit`. The ONE place the
|
||
/// `/feed`//`/search` merge contract lives, so the two surfaces cannot drift.
|
||
///
|
||
/// `S=1` short-circuits to the single group's result, which the engine already
|
||
/// returns score-sorted and limited — no redundant re-sort, byte-for-byte the
|
||
/// pre-m11p6 single read. The cross-group merge keeps each group's own diversity
|
||
/// pass but does NOT re-diversify across groups (a cross-shard re-rank is the L4
|
||
/// follow-up; disjoint groups make the score-merge sound for cardinality).
|
||
///
|
||
/// `set_rank` re-stamps the 1-based page rank over the merged order (see the
|
||
/// stamp site below). It mirrors [`merge_cross_shard`]'s parameter of the same
|
||
/// name so there is exactly ONE rank-stamping mechanism in the file, and it is a
|
||
/// no-op for result types with no rank field. The `S=1` fast path deliberately
|
||
/// bypasses it: the engine's own result is already densely ranked `1..k`.
|
||
async fn scatter_merge<T, F>(
|
||
dbs: Vec<Arc<TidalDb>>,
|
||
limit: usize,
|
||
score: impl Fn(&T) -> f64 + Send,
|
||
set_rank: impl Fn(&mut T, usize) + Send,
|
||
per_db: F,
|
||
) -> std::result::Result<(Vec<T>, usize), ServerError>
|
||
where
|
||
T: Send + 'static,
|
||
// `per_db` runs on a `spawn_blocking` thread, once per shard, possibly
|
||
// concurrently — it must be `Clone + Send + Sync + 'static` and own a clone
|
||
// of its captured query state. Each invocation gets its own `Arc<TidalDb>`.
|
||
F: Fn(Arc<TidalDb>) -> crate::offload::ShardSearch<T> + Clone + Send + Sync + 'static,
|
||
{
|
||
// S=1: the one group's result is already ranked + limited by the engine. A
|
||
// single group's error IS the read's error — nothing to degrade to. Still
|
||
// gated + offloaded (one search, one permit) so it cannot pin the reactor.
|
||
if let [only] = dbs.as_slice() {
|
||
let only = Arc::clone(only);
|
||
let per_db = per_db.clone();
|
||
return crate::offload::offload_search(move || per_db(only)).await;
|
||
}
|
||
// Fan every hosted group's search out CONCURRENTLY. `offload_search` gates
|
||
// each on the process-wide SEARCH_GATE (~= core count), so a wide fan-out
|
||
// never oversubscribes the cores; a shed shard returns Backpressure here and
|
||
// is degraded over exactly like any other per-shard error.
|
||
let futures = dbs.into_iter().map(|db| {
|
||
let per_db = per_db.clone();
|
||
crate::offload::offload_search(move || per_db(db))
|
||
});
|
||
let results = futures_util::future::join_all(futures).await;
|
||
|
||
let mut merged: Vec<T> = Vec::new();
|
||
let mut total = 0usize;
|
||
let mut ok_groups = 0usize;
|
||
let mut last_err: Option<ServerError> = None;
|
||
for (idx, result) in results.into_iter().enumerate() {
|
||
match result {
|
||
Ok((items, candidates)) => {
|
||
total = total.saturating_add(candidates);
|
||
merged.extend(items);
|
||
ok_groups += 1;
|
||
}
|
||
// m12p6: a SINGLE shard's probe error (e.g. a shard still converging
|
||
// after a restart) must NOT abort the whole corpus-wide read with a
|
||
// 500 — the old `?` turned one churning shard into a request-wide error
|
||
// storm. Record it and serve the surviving groups as an honest
|
||
// (reduced-recall) partial; the recall harness measures actual recall,
|
||
// so a silently-degraded shard cannot pass the gate.
|
||
Err(e) => {
|
||
tracing::warn!(
|
||
group_index = idx,
|
||
error = %e,
|
||
"cross-shard read: a local shard probe failed; serving the \
|
||
remaining groups as a degraded partial"
|
||
);
|
||
last_err = Some(e);
|
||
}
|
||
}
|
||
}
|
||
// Only when EVERY group failed is there no answer to return — surface it.
|
||
if ok_groups == 0 {
|
||
return Err(last_err.unwrap_or_else(|| {
|
||
ServerError::Cluster("cross-shard read: every local shard probe failed".into())
|
||
}));
|
||
}
|
||
merged.sort_by(|a, b| {
|
||
score(b)
|
||
.partial_cmp(&score(a))
|
||
.unwrap_or(std::cmp::Ordering::Equal)
|
||
});
|
||
merged.truncate(limit);
|
||
// Re-stamp the 1-based page rank over the MERGED order. Each hosted group
|
||
// ranks its own slice locally, so without this the wire `rank` is per-group
|
||
// counters concatenated (observed live: 1,1,2 for a 3-item /search).
|
||
// `merge_cross_shard` does the same thing for the PARTIAL-placement path;
|
||
// full placement returns from HERE, which is why it was missed. `set_rank`
|
||
// is a no-op for result types that carry no rank field (e.g. vector
|
||
// matches), exactly as in `merge_cross_shard`.
|
||
for (i, item) in merged.iter_mut().enumerate() {
|
||
set_rank(item, i + 1);
|
||
}
|
||
Ok((merged, total))
|
||
}
|
||
|
||
// ── m12p4: cross-shard unified reads (the m11p6 L4 follow-up) ──────────────────
|
||
|
||
/// One remote shard group's contribution to a cross-shard read: its parsed items
|
||
/// (already in the engine merge type `T`) and the per-group `total_candidates`.
|
||
struct RemoteGroupSlice<T> {
|
||
items: Vec<T>,
|
||
total_candidates: usize,
|
||
}
|
||
|
||
/// The merged outcome of a corpus-wide cluster read assembled from the local
|
||
/// scatter PLUS every missing group's remote slice.
|
||
struct CrossShardRead<T> {
|
||
/// Local + remote items, score-sorted descending and truncated to `limit`.
|
||
items: Vec<T>,
|
||
/// `total_candidates` SUMMED across the disjoint sources (local groups + each
|
||
/// remote group), mirroring [`scatter_merge`]'s entity-sharded rule — the
|
||
/// groups own disjoint key subsets, so the candidate universes add up.
|
||
total_candidates: usize,
|
||
/// Names of the missing groups whose every forward target was unreachable
|
||
/// (the honest degraded contract — never a silent truncation). Empty ⇒ a
|
||
/// fully-covered corpus read.
|
||
unavailable_shards: Vec<String>,
|
||
}
|
||
|
||
impl<T> CrossShardRead<T> {
|
||
/// Whether any missing group could not be reached on this read.
|
||
const fn degraded(&self) -> bool {
|
||
!self.unavailable_shards.is_empty()
|
||
}
|
||
}
|
||
|
||
/// Score-merge the local scatter result with every missing group's remote slice,
|
||
/// SUM the disjoint `total_candidates`, sort descending, and truncate to `limit`.
|
||
///
|
||
/// `T` is the engine merge type (`RetrieveResult` / `SearchResultItem`); `score`
|
||
/// reads its ranking score. The disjoint-source sum matches [`scatter_merge`]'s
|
||
/// entity-sharded rule — each group owns a distinct key subset, so the candidate
|
||
/// universes add rather than dedup-by-max.
|
||
fn merge_cross_shard<T>(
|
||
local_items: Vec<T>,
|
||
local_total: usize,
|
||
remote: Vec<RemoteGroupSlice<T>>,
|
||
unavailable_shards: Vec<String>,
|
||
limit: usize,
|
||
score: impl Fn(&T) -> f64,
|
||
set_rank: impl Fn(&mut T, usize),
|
||
) -> CrossShardRead<T> {
|
||
let mut merged = local_items;
|
||
let mut total = local_total;
|
||
for slice in remote {
|
||
total = total.saturating_add(slice.total_candidates);
|
||
merged.extend(slice.items);
|
||
}
|
||
merged.sort_by(|a, b| {
|
||
score(b)
|
||
.partial_cmp(&score(a))
|
||
.unwrap_or(std::cmp::Ordering::Equal)
|
||
});
|
||
merged.truncate(limit);
|
||
// Re-stamp the 1-based page rank over the MERGED order. Each group ranks its
|
||
// own slice locally (and remote slices arrive with rank 0), so without this
|
||
// the wire `rank` would be incoherent across the merge — local items keep
|
||
// their per-group 1,2,3… while every remote item reports 0. `set_rank` is a
|
||
// no-op for result types that carry no rank field (e.g. vector matches).
|
||
for (i, item) in merged.iter_mut().enumerate() {
|
||
set_rank(item, i + 1);
|
||
}
|
||
CrossShardRead {
|
||
items: merged,
|
||
total_candidates: total,
|
||
unavailable_shards,
|
||
}
|
||
}
|
||
|
||
/// The data-shard groups in the cluster this node does NOT host a replica of
|
||
/// (`placement.keys() − groups.keys()`), in ascending id order. Empty ⇒ full
|
||
/// placement (or `S=1`): the local scatter already covers the whole corpus and
|
||
/// the gateway returns it UNCHANGED — the byte-for-byte short-circuit.
|
||
impl ClusterNode {
|
||
fn missing_groups(&self) -> Vec<ShardId> {
|
||
self.placement
|
||
.keys()
|
||
.copied()
|
||
.filter(|s| !self.groups.contains_key(s))
|
||
.collect()
|
||
}
|
||
|
||
/// Build the internal per-group cross-shard read hop for group `shard`:
|
||
/// `GET {target}{path}?{base_query}&shard={shard}` carrying the internal
|
||
/// marker, this node's per-node token, and the forwarded bearer auth. Tries
|
||
/// `candidates` (leader-first) in order, failing over on a CONNECT-class error
|
||
/// so one dead replica does not strand a group; a real HTTP response (2xx or
|
||
/// otherwise) ends the walk. Returns the parsed `{items, total_candidates}`
|
||
/// slice, or `None` when EVERY target was unreachable (caller marks the group
|
||
/// degraded — never silently dropped).
|
||
///
|
||
/// `parse_item` maps one `serde_json::Value` element of the response `items`
|
||
/// array into the engine merge type `T`, so RETRIEVE/SEARCH/vector each keep
|
||
/// their own wire field set (`score` vs. `distance`).
|
||
async fn fetch_remote_group<T>(
|
||
&self,
|
||
shard: ShardId,
|
||
path: &str,
|
||
base_query: &str,
|
||
headers: &HeaderMap,
|
||
parse_item: impl Fn(&serde_json::Value) -> Option<T>,
|
||
) -> Option<RemoteGroupSlice<T>> {
|
||
let candidates = self.forward_candidates(shard);
|
||
if candidates.is_empty() {
|
||
tracing::warn!(
|
||
shard = shard.0,
|
||
"cross-shard read: missing group has no reachable forward target"
|
||
);
|
||
return None;
|
||
}
|
||
let auth = forwarded_auth(headers);
|
||
// Mint a fresh per-node token off this node's local replica so the remote
|
||
// node's marker guard sees a verified sibling and serves the internal hop.
|
||
let node_token = self
|
||
.replica_for(None)
|
||
.ok()
|
||
.and_then(|r| r.mint_node_token());
|
||
|
||
let sep = if base_query.is_empty() { "" } else { "&" };
|
||
for http_addr in &candidates {
|
||
let url = format!(
|
||
"{}?{base_query}{sep}shard={}",
|
||
peer_url(http_addr, path),
|
||
shard.0
|
||
);
|
||
let mut req = self
|
||
.client
|
||
.get(&url)
|
||
.header(forward::INTERNAL_MARKER, forward::INTERNAL_MARKER_VALUE);
|
||
if let Some(auth) = &auth {
|
||
req = req.header(axum::http::header::AUTHORIZATION, auth);
|
||
}
|
||
if let Some(token) = &node_token {
|
||
req = req.header(crate::cluster::security::NODE_TOKEN_HEADER, token.clone());
|
||
}
|
||
match req.send().await {
|
||
Ok(resp) if resp.status().is_success() => {
|
||
let body: serde_json::Value =
|
||
resp.json().await.unwrap_or(serde_json::Value::Null);
|
||
return Some(parse_group_slice(&body, &parse_item));
|
||
}
|
||
Ok(resp) => {
|
||
// A real verdict from a live target (e.g. 400/500): the group
|
||
// is reachable but errored. Do NOT fail over to another
|
||
// replica on a non-connect status — surface it as degraded.
|
||
tracing::warn!(
|
||
shard = shard.0, %url, status = %resp.status(),
|
||
"cross-shard read: remote group returned non-success"
|
||
);
|
||
return None;
|
||
}
|
||
Err(e) => {
|
||
// Connect-class failure: try the next replica before degrading.
|
||
tracing::warn!(shard = shard.0, %url, error = %e, "cross-shard read: target unreachable, trying next");
|
||
}
|
||
}
|
||
}
|
||
None
|
||
}
|
||
}
|
||
|
||
/// Parse a remote group's `{items:[...], total_candidates}` JSON into a typed
|
||
/// slice. `total_candidates` falls back to the item count when absent, so a
|
||
/// partial-but-present response never under-reports below what it returned.
|
||
fn parse_group_slice<T>(
|
||
body: &serde_json::Value,
|
||
parse_item: impl Fn(&serde_json::Value) -> Option<T>,
|
||
) -> RemoteGroupSlice<T> {
|
||
let items: Vec<T> = body
|
||
.get("items")
|
||
.and_then(|v| v.as_array())
|
||
.map(|arr| arr.iter().filter_map(&parse_item).collect())
|
||
.unwrap_or_default();
|
||
let total = body
|
||
.get("total_candidates")
|
||
.and_then(serde_json::Value::as_u64)
|
||
.map_or(items.len(), |t| t as usize);
|
||
RemoteGroupSlice {
|
||
items,
|
||
total_candidates: total,
|
||
}
|
||
}
|
||
|
||
/// Fan out a cross-shard read to every `missing` group CONCURRENTLY and gather
|
||
/// the slices, marking any group whose every target was unreachable as degraded.
|
||
/// The merge itself is the caller's ([`merge_cross_shard`]); this owns only the
|
||
/// scatter + honest-degraded accounting.
|
||
async fn fetch_missing_groups<T>(
|
||
node: &Arc<ClusterNode>,
|
||
missing: &[ShardId],
|
||
path: &str,
|
||
base_query: &str,
|
||
headers: &HeaderMap,
|
||
parse_item: impl Fn(&serde_json::Value) -> Option<T> + Clone,
|
||
) -> (Vec<RemoteGroupSlice<T>>, Vec<String>) {
|
||
let futures: Vec<_> = missing
|
||
.iter()
|
||
.map(|&shard| {
|
||
let node = Arc::clone(node);
|
||
let path = path.to_string();
|
||
let base_query = base_query.to_string();
|
||
let headers = headers.clone();
|
||
let parse_item = parse_item.clone();
|
||
async move {
|
||
let slice = node
|
||
.fetch_remote_group(shard, &path, &base_query, &headers, parse_item)
|
||
.await;
|
||
(shard, slice)
|
||
}
|
||
})
|
||
.collect();
|
||
|
||
let mut slices = Vec::with_capacity(missing.len());
|
||
let mut unavailable = Vec::new();
|
||
for (shard, slice) in futures_util::future::join_all(futures).await {
|
||
match slice {
|
||
Some(s) => slices.push(s),
|
||
None => unavailable.push(format!("s{}", shard.0)),
|
||
}
|
||
}
|
||
(slices, unavailable)
|
||
}
|
||
|
||
/// Ranked feed. Default read region is LOCAL; a `?region=` that names a DIFFERENT
|
||
/// region is forwarded to that region's process (region-aware reads). An internal
|
||
/// (marked) request always serves locally, so a forwarded read never loops.
|
||
#[utoipa::path(
|
||
get,
|
||
path = "/feed",
|
||
tag = "data",
|
||
params(FeedQuery),
|
||
responses(
|
||
(status = 200, description = "Ranked feed from the target region", body = FeedResponse),
|
||
(status = 400, description = "Unknown region or invalid request"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
(status = 503, description = "Target region unreachable while forwarding"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
// m12p4: region-aware forward + cross-shard internal branch + local scatter +
|
||
// missing-group fan-out + merge is one linear read assembly; splitting it would
|
||
// scatter the byte-for-byte full-placement short-circuit from the partial path.
|
||
#[allow(clippy::too_many_lines)]
|
||
pub async fn feed(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
RawQuery(raw_query): RawQuery,
|
||
Query(query): Query<FeedQuery>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let state = node.replica_for(None)?;
|
||
// Region-aware read: forward a foreign `?region=` to its owner unless this is
|
||
// already an internal (forwarded) request, which serves locally.
|
||
if is_internal(&headers) {
|
||
// A marked read must resolve to THIS region; an unknown name is still a 400.
|
||
state
|
||
.read_region(query.region.as_deref())
|
||
.map_err(ClusterAppError)?;
|
||
} else if let Some(resp) = maybe_forward_region_read(
|
||
&state,
|
||
query.region.as_deref(),
|
||
"/feed",
|
||
raw_query.as_deref(),
|
||
&headers,
|
||
)
|
||
.await?
|
||
{
|
||
return Ok(resp);
|
||
}
|
||
|
||
let limit = query.clamped_limit() as usize;
|
||
let build_retrieve = || {
|
||
let mut builder = Retrieve::builder().profile(&query.profile).limit(limit);
|
||
if let Some(user_id) = query.user_id {
|
||
builder = builder.for_user(user_id);
|
||
}
|
||
// m12p2: "more like this" seed for `profile=related` ANN candidate-gen.
|
||
if let Some(seed) = query.similar_to {
|
||
builder = builder.similar_to(EntityId::new(seed));
|
||
}
|
||
builder
|
||
.build()
|
||
.map_err(|e| ClusterAppError(ServerError::Tidal(e.into())))
|
||
};
|
||
|
||
// m12p4 cross-shard internal hop: an internal request with a `?shard=g`
|
||
// selector reads ONLY group `g` (the gateway's per-group fan-out leg) and
|
||
// serves it locally — single-db, NO further `hosted_dbs` scatter and NO
|
||
// re-fan-out, so the gateway's remote hop never loops.
|
||
//
|
||
// Security: gating on `is_internal` is sufficient and deliberate. The m11p7
|
||
// marker-pinning middleware (`cluster_auth_middleware`, a blanket layer on
|
||
// every protected route incl. this one) 403s ANY request carrying the
|
||
// internal marker WITHOUT a valid node token whenever a cluster key is
|
||
// configured — so an external caller can NEVER reach this single-group branch
|
||
// in a secured deployment. With NO cluster key the marker is hint-only by the
|
||
// documented trusted-network model, the SAME model the sibling fan-out leg
|
||
// relies on (it cannot present a token either), so a stricter `Principal::Node`
|
||
// gate here would break no-cluster-key partial-placement fan-out, not harden it.
|
||
if is_internal(&headers)
|
||
&& let Some(shard) = query.shard.map(ShardId)
|
||
{
|
||
let db = node
|
||
.replica_for(Some(shard))?
|
||
.db_arc()
|
||
.map_err(ClusterAppError)?;
|
||
let retrieve = build_retrieve()?;
|
||
let (items, total_candidates) = offload_region_read(move || {
|
||
let r = db.retrieve(&retrieve).map_err(ServerError::Tidal)?;
|
||
Ok((r.items, r.total_candidates))
|
||
})
|
||
.await?;
|
||
return Ok(Json(FeedResponse {
|
||
items: feed_items(&items),
|
||
total_candidates,
|
||
region: query.region,
|
||
unavailable_shards: None, // single-region serve: complete
|
||
})
|
||
.into_response());
|
||
}
|
||
|
||
// m11p6: scatter the corpus-wide read over the LOCALLY hosted shard groups
|
||
// and merge (see `scatter_merge`).
|
||
let retrieve = std::sync::Arc::new(build_retrieve()?);
|
||
let dbs = node.hosted_dbs();
|
||
let (local_items, local_total) = scatter_merge(
|
||
dbs,
|
||
limit,
|
||
|it: &tidaldb::query::RetrieveResult| it.score,
|
||
|it: &mut tidaldb::query::RetrieveResult, rank| it.rank = rank,
|
||
move |db: Arc<TidalDb>| {
|
||
let r = db.retrieve(&retrieve).map_err(ServerError::Tidal)?;
|
||
Ok((r.items, r.total_candidates))
|
||
},
|
||
)
|
||
.await
|
||
.map_err(ClusterAppError)?;
|
||
|
||
// m12p4: full placement / `S=1` → the local scatter already covers the whole
|
||
// corpus; return it UNCHANGED (byte-for-byte the pre-m12p4 response). Under
|
||
// PARTIAL placement, fan out to the groups this node does not host and merge.
|
||
let missing = node.missing_groups();
|
||
if missing.is_empty() {
|
||
return Ok(Json(FeedResponse {
|
||
items: feed_items(&local_items),
|
||
total_candidates: local_total,
|
||
region: query.region,
|
||
unavailable_shards: None, // full placement / S=1: corpus-complete locally
|
||
})
|
||
.into_response());
|
||
}
|
||
|
||
let base_query = feed_base_query(&query);
|
||
let (remote, unavailable) =
|
||
fetch_missing_groups(&node, &missing, "/feed", &base_query, &headers, |v| {
|
||
parse_scored_item(v).map(|(entity_id, score)| tidaldb::query::RetrieveResult {
|
||
entity_id,
|
||
score,
|
||
rank: 0,
|
||
signals: Vec::new(),
|
||
reasons: Vec::new(),
|
||
})
|
||
})
|
||
.await;
|
||
let merged = merge_cross_shard(
|
||
local_items,
|
||
local_total,
|
||
remote,
|
||
unavailable,
|
||
limit,
|
||
|it: &tidaldb::query::RetrieveResult| it.score,
|
||
|it: &mut tidaldb::query::RetrieveResult, rank| it.rank = rank,
|
||
);
|
||
let unavailable_shards = merged.degraded().then(|| merged.unavailable_shards.clone());
|
||
if let Some(shards) = &unavailable_shards {
|
||
tracing::warn!(
|
||
unavailable = ?shards,
|
||
"cross-shard /feed served degraded (some groups unreachable)"
|
||
);
|
||
}
|
||
|
||
Ok(Json(FeedResponse {
|
||
items: feed_items(&merged.items),
|
||
total_candidates: merged.total_candidates,
|
||
region: query.region,
|
||
unavailable_shards,
|
||
})
|
||
.into_response())
|
||
}
|
||
|
||
/// Build the internal cross-shard `/feed` query string for the per-group hop:
|
||
/// `profile`/`limit`/`user_id`/`similar_to` only — NOT `region` (the hop is
|
||
/// region-agnostic) and NOT `shard` (the fan-out appends the per-group selector
|
||
/// itself).
|
||
fn feed_base_query(query: &FeedQuery) -> String {
|
||
use std::fmt::Write as _;
|
||
let mut q = format!(
|
||
"profile={}&limit={}",
|
||
cross_shard_urlencode(&query.profile),
|
||
query.clamped_limit()
|
||
);
|
||
if let Some(uid) = query.user_id {
|
||
let _ = write!(q, "&user_id={uid}");
|
||
}
|
||
if let Some(seed) = query.similar_to {
|
||
let _ = write!(q, "&similar_to={seed}");
|
||
}
|
||
q
|
||
}
|
||
|
||
/// Parse one `{entity_id, score}` element of a remote read's `items` array into
|
||
/// an `(EntityId, f64)` pair. A malformed element is skipped (`None`), never a
|
||
/// hard error — the merge proceeds with what parsed.
|
||
fn parse_scored_item(v: &serde_json::Value) -> Option<(EntityId, f64)> {
|
||
let entity_id = v.get("entity_id").and_then(serde_json::Value::as_u64)?;
|
||
let score = v.get("score").and_then(serde_json::Value::as_f64)?;
|
||
Some((EntityId::new(entity_id), score))
|
||
}
|
||
|
||
/// Percent-encode a cross-shard read's query-string value (space + the reserved
|
||
/// set), enough for the `profile` / `query` params the per-group hop carries.
|
||
/// The hop URL is `?{base_query}&shard={g}`, so an unencoded `&`/`=`/space in a
|
||
/// `profile`/`query` value would corrupt the receiver's parse.
|
||
fn cross_shard_urlencode(s: &str) -> String {
|
||
use std::fmt::Write as _;
|
||
let mut out = String::with_capacity(s.len());
|
||
for b in s.bytes() {
|
||
match b {
|
||
b'A'..=b'Z' | b'a'..=b'z' | b'0'..=b'9' | b'-' | b'_' | b'.' | b'~' => {
|
||
out.push(b as char);
|
||
}
|
||
_ => {
|
||
let _ = write!(out, "%{b:02X}");
|
||
}
|
||
}
|
||
}
|
||
out
|
||
}
|
||
|
||
/// Ranked search. Region-aware reads as [`feed`]: a foreign `?region=` forwards
|
||
/// to its owner unless the request is internal (marked), which serves locally.
|
||
#[utoipa::path(
|
||
get,
|
||
path = "/search",
|
||
tag = "data",
|
||
params(SearchQueryParams),
|
||
responses(
|
||
(status = 200, description = "Ranked search from the target region", body = SearchResponse),
|
||
(status = 400, description = "Unknown region or invalid request"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
(status = 503, description = "Target region unreachable while forwarding"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
// m12p4: same linear read assembly as `feed` (region forward + cross-shard
|
||
// internal branch + local scatter + missing-group fan-out + merge).
|
||
#[allow(clippy::too_many_lines)]
|
||
pub async fn search(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
RawQuery(raw_query): RawQuery,
|
||
Query(query): Query<SearchQueryParams>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let state = node.replica_for(None)?;
|
||
if is_internal(&headers) {
|
||
state
|
||
.read_region(query.region.as_deref())
|
||
.map_err(ClusterAppError)?;
|
||
} else if let Some(resp) = maybe_forward_region_read(
|
||
&state,
|
||
query.region.as_deref(),
|
||
"/search",
|
||
raw_query.as_deref(),
|
||
&headers,
|
||
)
|
||
.await?
|
||
{
|
||
return Ok(resp);
|
||
}
|
||
|
||
let limit = query.clamped_limit();
|
||
let build_search = || {
|
||
let mut builder = Search::builder().query(&query.query).limit(limit);
|
||
if let Some(user_id) = query.user_id {
|
||
builder = builder.for_user(user_id);
|
||
}
|
||
builder
|
||
.build()
|
||
.map_err(|e| ClusterAppError(ServerError::Tidal(e.into())))
|
||
};
|
||
|
||
// m12p4 cross-shard internal hop: an internal `?shard=g` request searches
|
||
// ONLY group `g` (the gateway's per-group fan-out leg) and serves it locally,
|
||
// single-db, NO `hosted_dbs` scatter and NO re-fan-out.
|
||
if is_internal(&headers)
|
||
&& let Some(shard) = query.shard.map(ShardId)
|
||
{
|
||
let db = node
|
||
.replica_for(Some(shard))?
|
||
.db_arc()
|
||
.map_err(ClusterAppError)?;
|
||
let search_query = build_search()?;
|
||
let (items, total_candidates) = offload_region_read(move || {
|
||
db.reload_text_index().map_err(ServerError::Tidal)?;
|
||
let r = db.search(&search_query).map_err(ServerError::Tidal)?;
|
||
Ok((r.items, r.total_candidates))
|
||
})
|
||
.await?;
|
||
return Ok(Json(SearchResponse {
|
||
items: search_items(&items),
|
||
total_candidates,
|
||
region: query.region,
|
||
unavailable_shards: None, // single-region serve: complete
|
||
})
|
||
.into_response());
|
||
}
|
||
|
||
// m11p6: scatter the search over the LOCALLY hosted shard groups and merge.
|
||
let search_query = std::sync::Arc::new(build_search()?);
|
||
let dbs = node.hosted_dbs();
|
||
let (local_items, local_total) = scatter_merge(
|
||
dbs,
|
||
limit as usize,
|
||
|it: &tidaldb::query::SearchResultItem| it.score,
|
||
|it: &mut tidaldb::query::SearchResultItem, rank| it.rank = rank,
|
||
move |db: Arc<TidalDb>| {
|
||
db.reload_text_index().map_err(ServerError::Tidal)?;
|
||
let r = db.search(&search_query).map_err(ServerError::Tidal)?;
|
||
Ok((r.items, r.total_candidates))
|
||
},
|
||
)
|
||
.await
|
||
.map_err(ClusterAppError)?;
|
||
|
||
// m12p4: full placement / `S=1` → unchanged. Partial placement → fan out.
|
||
let missing = node.missing_groups();
|
||
if missing.is_empty() {
|
||
return Ok(Json(SearchResponse {
|
||
items: search_items(&local_items),
|
||
total_candidates: local_total,
|
||
region: query.region,
|
||
unavailable_shards: None, // full placement / S=1: corpus-complete locally
|
||
})
|
||
.into_response());
|
||
}
|
||
|
||
let base_query = search_base_query(&query);
|
||
let (remote, unavailable) =
|
||
fetch_missing_groups(&node, &missing, "/search", &base_query, &headers, |v| {
|
||
parse_scored_item(v).map(|(entity_id, score)| {
|
||
tidaldb::query::search::SearchResultItem {
|
||
entity_id,
|
||
score,
|
||
rank: 0,
|
||
bm25_score: None,
|
||
semantic_score: None,
|
||
signals: Vec::new(),
|
||
metadata: None,
|
||
reasons: Vec::new(),
|
||
}
|
||
})
|
||
})
|
||
.await;
|
||
let merged = merge_cross_shard(
|
||
local_items,
|
||
local_total,
|
||
remote,
|
||
unavailable,
|
||
limit as usize,
|
||
|it: &tidaldb::query::search::SearchResultItem| it.score,
|
||
|it: &mut tidaldb::query::search::SearchResultItem, rank| it.rank = rank,
|
||
);
|
||
let unavailable_shards = merged.degraded().then(|| merged.unavailable_shards.clone());
|
||
if let Some(shards) = &unavailable_shards {
|
||
tracing::warn!(
|
||
unavailable = ?shards,
|
||
"cross-shard /search served degraded (some groups unreachable)"
|
||
);
|
||
}
|
||
|
||
Ok(Json(SearchResponse {
|
||
items: search_items(&merged.items),
|
||
total_candidates: merged.total_candidates,
|
||
region: query.region,
|
||
unavailable_shards,
|
||
})
|
||
.into_response())
|
||
}
|
||
|
||
/// Build the internal cross-shard `/search` query string for the per-group hop:
|
||
/// `query`/`limit`/`user_id` only — NOT `region` and NOT `shard` (the fan-out
|
||
/// appends the per-group selector itself).
|
||
fn search_base_query(query: &SearchQueryParams) -> String {
|
||
use std::fmt::Write as _;
|
||
let mut q = format!(
|
||
"query={}&limit={}",
|
||
cross_shard_urlencode(&query.query),
|
||
query.clamped_limit()
|
||
);
|
||
if let Some(uid) = query.user_id {
|
||
let _ = write!(q, "&user_id={uid}");
|
||
}
|
||
q
|
||
}
|
||
|
||
/// Pure k-NN vector search (the m12p1 recall probe). Serves LOCALLY from this
|
||
/// node's hosted shard groups — no `?region=` forwarding: it is a measurement
|
||
/// surface, and at the S=1 exit-gate shape every region replica holds the full
|
||
/// corpus, so any node answers the whole-corpus nearest set. With S>1 the probe
|
||
/// merges each hosted group's local nearest by distance; whole-corpus recall is
|
||
/// then bounded by the cross-shard merge (the m12p4 cross-shard read follow-up).
|
||
#[utoipa::path(
|
||
post,
|
||
path = "/vector_search",
|
||
tag = "data",
|
||
request_body = VectorSearchRequest,
|
||
responses(
|
||
(status = 200, description = "Nearest items by vector distance, closest-first", body = VectorSearchResponse),
|
||
(status = 400, description = "Empty/dimension-mismatched query vector, or no embedding slot"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
pub async fn vector_search(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
Json(req): Json<VectorSearchRequest>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
if req.vector.is_empty() {
|
||
return Err(ClusterAppError(ServerError::BadRequest(
|
||
"vector_search requires a non-empty query vector".into(),
|
||
)));
|
||
}
|
||
let k = req.clamped_k();
|
||
let ef_search = req.ef_search();
|
||
|
||
// m12p4 cross-shard internal hop: an internal request carrying `shard=g`
|
||
// probes ONLY group `g` (the gateway's per-group fan-out leg) and serves it
|
||
// locally, single-db, NO `hosted_dbs` scatter and NO re-fan-out.
|
||
if is_internal(&headers)
|
||
&& let Some(shard) = req.shard.map(ShardId)
|
||
{
|
||
let db = node
|
||
.replica_for(Some(shard))?
|
||
.db_arc()
|
||
.map_err(ClusterAppError)?;
|
||
let vector = req.vector;
|
||
let items = offload_region_read(move || {
|
||
db.vector_search_items(&vector, k, ef_search)
|
||
.map_err(ServerError::Tidal)
|
||
})
|
||
.await?;
|
||
return Ok(Json(VectorSearchResponse {
|
||
items: vector_matches(&items),
|
||
region: None,
|
||
unavailable_shards: None, // single-region serve: complete
|
||
})
|
||
.into_response());
|
||
}
|
||
|
||
// m11p6: scatter the probe over the LOCALLY hosted shard groups and merge.
|
||
let vector = std::sync::Arc::new(req.vector.clone());
|
||
let dbs = node.hosted_dbs();
|
||
let (local_items, local_total) = scatter_merge(
|
||
dbs,
|
||
k,
|
||
// Distance is "lower = better"; scatter_merge ranks by "higher =
|
||
// better", so the merge key is the negated distance.
|
||
|r: &tidaldb::storage::vector::VectorSearchResult| -f64::from(r.distance),
|
||
// Vector matches carry no rank field (ordered by distance on the wire),
|
||
// so the stamp is a genuine no-op — never a panic.
|
||
|_r: &mut tidaldb::storage::vector::VectorSearchResult, _rank| {},
|
||
move |db: Arc<TidalDb>| {
|
||
let r = db
|
||
.vector_search_items(&vector, k, ef_search)
|
||
.map_err(ServerError::Tidal)?;
|
||
let n = r.len();
|
||
Ok((r, n))
|
||
},
|
||
)
|
||
.await
|
||
.map_err(ClusterAppError)?;
|
||
|
||
// m12p4: full placement / `S=1` → unchanged. Partial placement → fan out the
|
||
// probe to each missing group (POST body carrying `shard=g`) and merge by
|
||
// ascending distance so whole-corpus recall is no longer local-shard-only.
|
||
let missing = node.missing_groups();
|
||
if missing.is_empty() {
|
||
return Ok(Json(VectorSearchResponse {
|
||
items: vector_matches(&local_items),
|
||
region: None,
|
||
unavailable_shards: None, // full placement / S=1: corpus-complete locally
|
||
})
|
||
.into_response());
|
||
}
|
||
|
||
let (remote, unavailable) = node
|
||
.fetch_missing_groups_vector(&missing, &req, k, ef_search, &headers)
|
||
.await;
|
||
let merged = merge_cross_shard(
|
||
local_items,
|
||
local_total,
|
||
remote,
|
||
unavailable,
|
||
k,
|
||
// Closest-first: negate distance so the shared descending merge keeps the
|
||
// nearest neighbours.
|
||
|r: &tidaldb::storage::vector::VectorSearchResult| -f64::from(r.distance),
|
||
// Vector matches carry no rank field (ordered by distance on the wire).
|
||
|_r: &mut tidaldb::storage::vector::VectorSearchResult, _rank| {},
|
||
);
|
||
let unavailable_shards = merged.degraded().then(|| merged.unavailable_shards.clone());
|
||
if let Some(shards) = &unavailable_shards {
|
||
tracing::warn!(
|
||
unavailable = ?shards,
|
||
"cross-shard /vector_search served degraded (some groups unreachable)"
|
||
);
|
||
}
|
||
|
||
Ok(Json(VectorSearchResponse {
|
||
items: vector_matches(&merged.items),
|
||
region: None,
|
||
unavailable_shards,
|
||
})
|
||
.into_response())
|
||
}
|
||
|
||
impl ClusterNode {
|
||
/// Fan out the vector probe to every missing group CONCURRENTLY (POST body
|
||
/// carrying the per-group `shard` selector), gather each group's nearest
|
||
/// slice, and mark any group whose every target was unreachable as degraded
|
||
/// — the POST analogue of [`fetch_missing_groups`] for the body-carried
|
||
/// query vector.
|
||
async fn fetch_missing_groups_vector(
|
||
self: &Arc<Self>,
|
||
missing: &[ShardId],
|
||
req: &VectorSearchRequest,
|
||
k: usize,
|
||
ef_search: Option<usize>,
|
||
headers: &HeaderMap,
|
||
) -> (
|
||
Vec<RemoteGroupSlice<tidaldb::storage::vector::VectorSearchResult>>,
|
||
Vec<String>,
|
||
) {
|
||
let auth = forwarded_auth(headers);
|
||
let node_token = self
|
||
.replica_for(None)
|
||
.ok()
|
||
.and_then(|r| r.mint_node_token());
|
||
let futures: Vec<_> = missing
|
||
.iter()
|
||
.map(|&shard| {
|
||
let node = Arc::clone(self);
|
||
let auth = auth.clone();
|
||
let node_token = node_token.clone();
|
||
// The per-group hop body: the same vector + knobs, with `shard`
|
||
// set so the remote serves ONLY this group and never re-fans-out.
|
||
let body = VectorSearchRequest {
|
||
vector: req.vector.clone(),
|
||
k: u32::try_from(k).unwrap_or(u32::MAX),
|
||
ef_search: ef_search.map(|e| u32::try_from(e).unwrap_or(u32::MAX)),
|
||
shard: Some(shard.0),
|
||
};
|
||
async move {
|
||
let slice = node
|
||
.fetch_remote_group_vector(
|
||
shard,
|
||
&body,
|
||
auth.as_deref(),
|
||
node_token.as_deref(),
|
||
)
|
||
.await;
|
||
(shard, slice)
|
||
}
|
||
})
|
||
.collect();
|
||
|
||
let mut slices = Vec::with_capacity(missing.len());
|
||
let mut unavailable = Vec::new();
|
||
for (shard, slice) in futures_util::future::join_all(futures).await {
|
||
match slice {
|
||
Some(s) => slices.push(s),
|
||
None => unavailable.push(format!("s{}", shard.0)),
|
||
}
|
||
}
|
||
(slices, unavailable)
|
||
}
|
||
|
||
/// POST the vector probe to one missing group's forward targets (leader-first,
|
||
/// failing over on a connect error). Returns the parsed `{items:[{entity_id,
|
||
/// distance}], …}` slice, or `None` when every target was unreachable.
|
||
async fn fetch_remote_group_vector(
|
||
&self,
|
||
shard: ShardId,
|
||
body: &VectorSearchRequest,
|
||
auth: Option<&str>,
|
||
node_token: Option<&str>,
|
||
) -> Option<RemoteGroupSlice<tidaldb::storage::vector::VectorSearchResult>> {
|
||
let candidates = self.forward_candidates(shard);
|
||
if candidates.is_empty() {
|
||
tracing::warn!(
|
||
shard = shard.0,
|
||
"cross-shard vector probe: missing group has no reachable forward target"
|
||
);
|
||
return None;
|
||
}
|
||
for http_addr in &candidates {
|
||
let url = peer_url(http_addr, "/vector_search");
|
||
let mut req = self
|
||
.client
|
||
.post(&url)
|
||
.header(forward::INTERNAL_MARKER, forward::INTERNAL_MARKER_VALUE)
|
||
.json(body);
|
||
if let Some(auth) = auth {
|
||
req = req.header(axum::http::header::AUTHORIZATION, auth);
|
||
}
|
||
if let Some(token) = node_token {
|
||
req = req.header(crate::cluster::security::NODE_TOKEN_HEADER, token);
|
||
}
|
||
match req.send().await {
|
||
Ok(resp) if resp.status().is_success() => {
|
||
let json: serde_json::Value =
|
||
resp.json().await.unwrap_or(serde_json::Value::Null);
|
||
return Some(parse_group_slice(&json, parse_vector_match));
|
||
}
|
||
Ok(resp) => {
|
||
tracing::warn!(
|
||
shard = shard.0, %url, status = %resp.status(),
|
||
"cross-shard vector probe: remote group returned non-success"
|
||
);
|
||
return None;
|
||
}
|
||
Err(e) => {
|
||
tracing::warn!(shard = shard.0, %url, error = %e, "cross-shard vector probe: target unreachable, trying next");
|
||
}
|
||
}
|
||
}
|
||
None
|
||
}
|
||
}
|
||
|
||
/// Parse one `{entity_id, distance}` element of a `/vector_search` response into
|
||
/// an engine `VectorSearchResult`. A malformed element is skipped (`None`).
|
||
fn parse_vector_match(
|
||
v: &serde_json::Value,
|
||
) -> Option<tidaldb::storage::vector::VectorSearchResult> {
|
||
let id = v.get("entity_id").and_then(serde_json::Value::as_u64)?;
|
||
let distance = v.get("distance").and_then(serde_json::Value::as_f64)?;
|
||
Some(tidaldb::storage::vector::VectorSearchResult {
|
||
id,
|
||
distance: distance as f32,
|
||
})
|
||
}
|
||
|
||
/// If `region` names a DIFFERENT region than this node owns, forward the read
|
||
/// (verbatim query string, marker set, auth passed through) to that region's
|
||
/// process and relay its response. Returns `Ok(None)` when the read should be
|
||
/// served locally (no `?region=`, or it names THIS region). A 400 for an unknown
|
||
/// region, a 503 for an unreachable owner.
|
||
///
|
||
/// `raw_query` is the originating request's verbatim query string (without the
|
||
/// `?`), so the owner runs the IDENTICAL query (`profile`/`limit`/`user_id`/`region`).
|
||
async fn maybe_forward_region_read(
|
||
state: &Arc<ShardReplica>,
|
||
region: Option<&str>,
|
||
path: &str,
|
||
raw_query: Option<&str>,
|
||
headers: &HeaderMap,
|
||
) -> std::result::Result<Option<Response>, ClusterAppError> {
|
||
let Some(name) = region else {
|
||
return Ok(None); // local read
|
||
};
|
||
let id = state.resolve_region(name).map_err(ClusterAppError)?;
|
||
if id == state.region {
|
||
return Ok(None); // already local
|
||
}
|
||
let Some(http_addr) = state.peer_http.get(&id).cloned() else {
|
||
return Err(ClusterAppError(ServerError::NotLocal {
|
||
region: name.to_string(),
|
||
}));
|
||
};
|
||
// Relay the original query string verbatim to the owner's same path.
|
||
let url = match raw_query {
|
||
Some(q) if !q.is_empty() => format!("{}?{q}", peer_url(&http_addr, path)),
|
||
_ => peer_url(&http_addr, path),
|
||
};
|
||
let auth = forwarded_auth(headers);
|
||
let mut req = state
|
||
.client
|
||
.get(&url)
|
||
.header(forward::INTERNAL_MARKER, forward::INTERNAL_MARKER_VALUE);
|
||
if let Some(auth) = auth {
|
||
req = req.header(axum::http::header::AUTHORIZATION, auth);
|
||
}
|
||
// m11p7: the node token proves a verified sibling so the owner's marker guard
|
||
// honors this internal region-pinned read forward.
|
||
if let Some(token) = state.mint_node_token() {
|
||
req = req.header(crate::cluster::security::NODE_TOKEN_HEADER, token);
|
||
}
|
||
match req.send().await {
|
||
Ok(resp) => {
|
||
let status = resp.status();
|
||
let bytes = resp.bytes().await.unwrap_or_default();
|
||
// Honor the SAME bodyless-status contract as the write-relay helper
|
||
// (`forward::relay_forwarded`): a 204/304/1xx must carry NO body, or
|
||
// an HTTP/2 client RST_STREAMs it. These read forwards return 200 or
|
||
// a body-bearing error today, but routing through the shared guard
|
||
// keeps the read- and write-relay paths from drifting if a forwarded
|
||
// read ever learns to return a bodyless status (e.g. a 304 ETag).
|
||
if forward::status_forbids_body(status) {
|
||
return Ok(Some(status.into_response()));
|
||
}
|
||
let body: serde_json::Value =
|
||
serde_json::from_slice(&bytes).unwrap_or(serde_json::Value::Null);
|
||
Ok(Some((status, Json(body)).into_response()))
|
||
}
|
||
Err(e) => {
|
||
tracing::warn!(region = name, %url, error = %e, "region read forward failed");
|
||
Err(ClusterAppError(ServerError::RegionUnreachable {
|
||
region: name.to_string(),
|
||
cause: e.to_string(),
|
||
}))
|
||
}
|
||
}
|
||
}
|
||
|
||
// ── Cross-process CRDT reconciliation ─────────────────────────────────────────
|
||
|
||
/// `POST /cluster/reconcile/snapshot` request/response: the wire CRDT snapshot.
|
||
///
|
||
/// The body is the remote node's [`StateSnapshot`](tidaldb::replication::reconcile::StateSnapshot),
|
||
/// which serializes to JSON via its wire form (task 01). This route is INTERNAL
|
||
/// (marker required) — an operator never calls it directly; it is the snapshot
|
||
/// exchange the operator-facing `/cluster/reconcile` drives.
|
||
#[derive(Serialize, ToSchema)]
|
||
pub struct ReconcileSnapshotResponse {
|
||
/// This node's PRE-merge snapshot (so the caller can apply it on its side and
|
||
/// both sides converge to the same CRDT state).
|
||
#[schema(value_type = Object)]
|
||
pre_merge_snapshot: tidaldb::replication::reconcile::StateSnapshot,
|
||
/// Merge+apply time on this node, in milliseconds (NOT the HTTP round-trip).
|
||
elapsed_ms: u64,
|
||
}
|
||
|
||
/// Internal snapshot exchange for cross-process reconciliation.
|
||
///
|
||
/// MARKER REQUIRED: a request without `x-tidal-internal: 1` is rejected (this is
|
||
/// the snapshot-exchange leg the operator-facing `/cluster/reconcile` drives, not
|
||
/// a public endpoint). This node takes its OWN snapshot FIRST (the pre-merge
|
||
/// state it returns), then `reconcile_with(remote)` to fold the caller's snapshot
|
||
/// into its state, and reports the merge+apply `elapsed_ms`. CRDT merge
|
||
/// determinism guarantees both sides converge to identical state.
|
||
#[utoipa::path(
|
||
post,
|
||
path = "/cluster/reconcile/snapshot",
|
||
tag = "cluster",
|
||
request_body(content = Object, description = "Wire StateSnapshot"),
|
||
responses(
|
||
(status = 200, description = "Reconciled; returns pre-merge snapshot + elapsed_ms", body = ReconcileSnapshotResponse),
|
||
(status = 400, description = "Missing internal marker"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn cluster_reconcile_snapshot(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
Json(remote): Json<tidaldb::replication::reconcile::StateSnapshot>,
|
||
) -> std::result::Result<Json<ReconcileSnapshotResponse>, ClusterAppError> {
|
||
let state = node.replica_for(None)?;
|
||
if !is_internal(&headers) {
|
||
return Err(ClusterAppError(ServerError::BadRequest(
|
||
"/cluster/reconcile/snapshot is internal; the x-tidal-internal marker is required \
|
||
(drive it via POST /cluster/reconcile)"
|
||
.into(),
|
||
)));
|
||
}
|
||
let state_for_job = Arc::clone(&state);
|
||
// Snapshot + reconcile are blocking engine calls; offload them. Take OUR
|
||
// pre-merge snapshot first (the value we return), then fold in the remote.
|
||
let (pre_merge_snapshot, elapsed_ms) = offload_region_read(move || {
|
||
let pre = state_for_job.take_snapshot()?;
|
||
let start = std::time::Instant::now();
|
||
state_for_job.reconcile_remote(&remote)?;
|
||
let elapsed_ms = start.elapsed().as_millis() as u64;
|
||
Ok((pre, elapsed_ms))
|
||
})
|
||
.await?;
|
||
Ok(Json(ReconcileSnapshotResponse {
|
||
pre_merge_snapshot,
|
||
elapsed_ms,
|
||
}))
|
||
}
|
||
|
||
/// `POST /cluster/reconcile` response body.
|
||
#[derive(Serialize, ToSchema)]
|
||
pub struct ReconcileResponse {
|
||
/// Always true on success.
|
||
ok: bool,
|
||
/// The region this node reconciled with.
|
||
region: String,
|
||
/// Merge+apply time on THIS node, in milliseconds.
|
||
local_elapsed_ms: u64,
|
||
/// Merge+apply time the REMOTE node reported, in milliseconds.
|
||
remote_elapsed_ms: u64,
|
||
/// Operations this node applied folding in the remote's pre-merge snapshot.
|
||
ops_applied: usize,
|
||
}
|
||
|
||
/// Operator-facing cross-process reconciliation with a target region.
|
||
///
|
||
/// This node snapshots itself, POSTs that snapshot to the target's
|
||
/// `/cluster/reconcile/snapshot` (internal, marker set), receives the target's
|
||
/// pre-merge snapshot, and applies it via its own `reconcile_with`. CRDT merge
|
||
/// determinism guarantees both sides converge to the SAME state. Idempotent: a
|
||
/// second reconcile with an already-converged peer applies a no-op plan.
|
||
#[utoipa::path(
|
||
post,
|
||
path = "/cluster/reconcile",
|
||
tag = "cluster",
|
||
request_body = RegionRequest,
|
||
responses(
|
||
(status = 200, description = "Both sides converged", body = ReconcileResponse),
|
||
(status = 400, description = "Unknown region or no http_addr"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
(status = 503, description = "Target region unreachable"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn cluster_reconcile(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
Json(req): Json<RegionRequest>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let state = node.replica_for(None)?;
|
||
let id = state.resolve_region(&req.region).map_err(ClusterAppError)?;
|
||
if id == state.region {
|
||
return Err(ClusterAppError(ServerError::BadRequest(
|
||
"cannot reconcile a region with itself".into(),
|
||
)));
|
||
}
|
||
let Some(http_addr) = state.peer_http.get(&id).cloned() else {
|
||
return Err(ClusterAppError(ServerError::BadRequest(format!(
|
||
"region '{}' has no http_addr to reconcile with",
|
||
req.region
|
||
))));
|
||
};
|
||
|
||
// Take our snapshot off the reactor.
|
||
let snap_state = Arc::clone(&state);
|
||
let local_snapshot = offload_region_read(move || snap_state.take_snapshot()).await?;
|
||
|
||
// Exchange snapshots with the target (internal marker so it does not re-fan).
|
||
let url = peer_url(&http_addr, "/cluster/reconcile/snapshot");
|
||
let auth = forwarded_auth(&headers);
|
||
let exchanged = match forward_json_with_headers(
|
||
&state.client,
|
||
&url,
|
||
&local_snapshot,
|
||
auth.as_deref(),
|
||
true,
|
||
&state.node_token_passthrough(),
|
||
)
|
||
.await
|
||
{
|
||
Ok(resp) if resp.status.is_success() => resp.body,
|
||
// A 413 is NOT unreachability: the peer answered, and said the payload
|
||
// is too big. Reported as such it sent an operator hunting TLS and
|
||
// NetworkPolicy while the cluster was perfectly connected. Name the
|
||
// measured size, the cap, and the fix.
|
||
Ok(resp) if resp.status == reqwest::StatusCode::PAYLOAD_TOO_LARGE => {
|
||
let bytes = serde_json::to_vec(&local_snapshot).map_or(0, |v| v.len());
|
||
return Err(ClusterAppError(ServerError::Cluster(format!(
|
||
"reconcile snapshot is {bytes} bytes; peer '{}' capped it at {} bytes. The \
|
||
corpus outgrew the single-shot reconcile — the peer is reachable and healthy. \
|
||
Raise RECONCILE_BODY_LIMIT_BYTES on BOTH nodes, or chunk the exchange.",
|
||
req.region,
|
||
crate::router::RECONCILE_BODY_LIMIT_BYTES,
|
||
))));
|
||
}
|
||
Ok(resp) => {
|
||
return Err(ClusterAppError(ServerError::RegionUnreachable {
|
||
region: req.region,
|
||
cause: format!("reconcile peer returned {}", resp.status),
|
||
}));
|
||
}
|
||
Err(e) => {
|
||
return Err(ClusterAppError(ServerError::RegionUnreachable {
|
||
region: req.region,
|
||
cause: e,
|
||
}));
|
||
}
|
||
};
|
||
|
||
let remote_elapsed_ms = exchanged
|
||
.get("elapsed_ms")
|
||
.and_then(serde_json::Value::as_u64)
|
||
.unwrap_or(0);
|
||
let remote_snapshot: tidaldb::replication::reconcile::StateSnapshot = exchanged
|
||
.get("pre_merge_snapshot")
|
||
.cloned()
|
||
.and_then(|v| serde_json::from_value(v).ok())
|
||
.ok_or_else(|| {
|
||
ClusterAppError(ServerError::Cluster(
|
||
"reconcile peer returned no pre_merge_snapshot".into(),
|
||
))
|
||
})?;
|
||
|
||
// Apply the remote's pre-merge snapshot on our side (blocking; offload).
|
||
let apply_state = Arc::clone(&state);
|
||
let (ops_applied, local_elapsed_ms) = offload_region_read(move || {
|
||
let start = std::time::Instant::now();
|
||
let ops = apply_state.reconcile_remote(&remote_snapshot)?;
|
||
Ok((ops, start.elapsed().as_millis() as u64))
|
||
})
|
||
.await?;
|
||
|
||
Ok(Json(ReconcileResponse {
|
||
ok: true,
|
||
region: req.region,
|
||
local_elapsed_ms,
|
||
remote_elapsed_ms,
|
||
ops_applied,
|
||
})
|
||
.into_response())
|
||
}
|
||
|
||
// ── Sharded (cross-process scatter-gather) routes ─────────────────────────────
|
||
|
||
/// Build the engine `ShardRouter` shard list (every region, in id order) used
|
||
/// for `/sharded/*` hash-partitioning and read fan-out.
|
||
fn sharded_region_ids(state: &Arc<ShardReplica>) -> Vec<RegionId> {
|
||
let mut ids: Vec<RegionId> = state.id_to_name.keys().copied().collect();
|
||
ids.sort_by_key(|r| r.0);
|
||
ids
|
||
}
|
||
|
||
/// Region id → name map for the scatter-gather metadata.
|
||
fn sharded_region_names(state: &Arc<ShardReplica>) -> HashMap<RegionId, String> {
|
||
state.id_to_name.clone()
|
||
}
|
||
|
||
/// Build the HTTP scatter-gather context for this gateway.
|
||
fn http_shard_context(
|
||
state: &Arc<ShardReplica>,
|
||
auth: Option<String>,
|
||
) -> Result<Arc<HttpShardContext>> {
|
||
let db = state.db_arc()?;
|
||
Ok(Arc::new(HttpShardContext::new(
|
||
state.region,
|
||
db,
|
||
state.peer_http.clone(),
|
||
state.blocking_client.clone(),
|
||
auth,
|
||
state.mint_node_token(),
|
||
)))
|
||
}
|
||
|
||
/// The `x-tidal-ack` value a `/sharded/*` WRITE must carry to opt in to
|
||
/// single-copy durability.
|
||
///
|
||
/// Deliberately NOT an [`AckMode`] variant: `local` describes a surface that
|
||
/// never appends to the WAL, so it is meaningless on the replicating routes and
|
||
/// `AckMode::parse` must keep rejecting it there ("must be leader or quorum").
|
||
/// Same header, one name, no second durability knob.
|
||
pub(super) const ACK_LOCAL: &str = "local";
|
||
|
||
/// Gate a `/sharded/*` WRITE on an EXPLICIT single-copy opt-in.
|
||
///
|
||
/// The surface hash-partitions and applies to the owning region's LOCAL store
|
||
/// with no WAL append (see [`sharded_write_route`]), so its data has redundancy
|
||
/// 1 no matter what the replication factor is. It answered `201`/`204` with
|
||
/// nothing at the call site, in the response, or in the `OpenAPI` saying so — and a
|
||
/// caller on a cluster configured `ack: quorum` with RF3 reasonably assumes their
|
||
/// write replicated. An operator probing with `/sharded/embeddings` found each
|
||
/// one on exactly one of three nodes and filed a durability incident that had to
|
||
/// be retracted.
|
||
///
|
||
/// So: reject (400) unless the caller said `x-tidal-ack: local`. Never silently
|
||
/// accept, and never silently reroute to the replicating path — a reroute would
|
||
/// change the write's performance characteristics under the caller's feet, which
|
||
/// is its own bandaid.
|
||
///
|
||
/// The error names the header AND the replicating alternative, derived from the
|
||
/// route path so the two cannot drift.
|
||
///
|
||
/// `pub(super)` so `cluster::routes` (the single-process cluster's copy of these
|
||
/// three routes) enforces the SAME gate from the SAME definition. One URL, one
|
||
/// contract: a client must not have to know the server's process topology to know
|
||
/// whether its write replicated.
|
||
pub(super) fn require_local_ack(headers: &HeaderMap, path: &str) -> Result<()> {
|
||
let header = forward::ACK_HEADER;
|
||
let value = headers.get(header).map(|v| v.to_str());
|
||
let got = match value {
|
||
Some(Ok(ACK_LOCAL)) => return Ok(()),
|
||
None => "the header was absent".to_owned(),
|
||
Some(Ok(v)) => format!("got {v:?}"),
|
||
Some(Err(_)) => "got a non-ASCII value".to_owned(),
|
||
};
|
||
let replicating = path.strip_prefix("/sharded").unwrap_or(path);
|
||
Err(ServerError::BadRequest(format!(
|
||
"{path} applies the write to the owning region's LOCAL store with no WAL \
|
||
append, so it is SINGLE-COPY regardless of the replication factor. Send \
|
||
\"{header}: {ACK_LOCAL}\" to opt in to that, or POST {replicating} \
|
||
instead for a replicated write (leader WAL relay; \
|
||
\"{header}: leader|quorum\"). Rejected: {got}."
|
||
)))
|
||
}
|
||
|
||
/// Forward a sharded WRITE to the owning region (marker set) when this node is
|
||
/// not the owner; else apply locally. Returns the relayed response, or the local
|
||
/// status on a local apply.
|
||
///
|
||
/// The ONE funnel every `/sharded/*` write passes through, and therefore where
|
||
/// the single-copy opt-in is enforced ([`require_local_ack`]) — once, for all
|
||
/// three routes, before anything is applied or forwarded.
|
||
///
|
||
/// An INTERNAL request is exempt: the marker means a verified cluster sibling
|
||
/// already forwarded this write, and `cluster_auth_middleware` rejects the marker
|
||
/// from anyone without a valid node token, so an external caller cannot use it to
|
||
/// slip past the gate. Gating the forwarded leg too would reject the owner's own
|
||
/// hop, since the forward carries the auth + node token, not the caller's headers.
|
||
async fn sharded_write_route<B: serde::Serialize + Sync>(
|
||
state: &Arc<ShardReplica>,
|
||
headers: &HeaderMap,
|
||
entity_id: u64,
|
||
path: &str,
|
||
body: &B,
|
||
local_apply: impl FnOnce() -> Result<()> + Send + 'static,
|
||
success: StatusCode,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let internal = is_internal(headers);
|
||
if !internal {
|
||
require_local_ack(headers, path).map_err(ClusterAppError)?;
|
||
}
|
||
let shards = sharded_region_ids(state);
|
||
let owner = entity_shard(EntityId::new(entity_id), &shards);
|
||
// Owner is self, or this is an internal (already-forwarded) write → apply local.
|
||
if owner == state.region || internal {
|
||
offload_region_read(local_apply).await?;
|
||
return Ok(success.into_response());
|
||
}
|
||
// Forward to the owner's /sharded/* with the marker set.
|
||
let Some(http_addr) = state.peer_http.get(&owner).cloned() else {
|
||
return Err(ClusterAppError(ServerError::Cluster(format!(
|
||
"no http_addr for sharded owner region '{}'",
|
||
state.region_name_of(owner)
|
||
))));
|
||
};
|
||
let url = peer_url(&http_addr, path);
|
||
let auth = forwarded_auth(headers);
|
||
match forward_json_with_headers(
|
||
&state.client,
|
||
&url,
|
||
body,
|
||
auth.as_deref(),
|
||
true,
|
||
&state.node_token_passthrough(),
|
||
)
|
||
.await
|
||
{
|
||
// Relay through the shared helper so a bodyless status (204 from a
|
||
// signal/embedding write) carries NO body — an HTTP/2 client RST_STREAMs
|
||
// a 204+body — and the peer's x-tidal-seq/dedup verdict headers ride back.
|
||
Ok(resp) => Ok(forward::relay_forwarded(resp)),
|
||
Err(e) => Err(ClusterAppError(ServerError::RegionUnreachable {
|
||
region: state.region_name_of(owner).to_string(),
|
||
cause: e,
|
||
})),
|
||
}
|
||
}
|
||
|
||
/// `POST /sharded/items` — route to the owning region (engine `ShardRouter` hash).
|
||
///
|
||
/// **SINGLE-COPY.** The write is applied to the owning region's LOCAL store with
|
||
/// NO WAL append, so it does not ride the leader relay and is not replicated —
|
||
/// redundancy 1 regardless of the replication factor. That is by design (parallel
|
||
/// write throughput across shard owners), which is why the surface requires
|
||
/// `x-tidal-ack: local` as an explicit acknowledgement of the tradeoff. For a
|
||
/// replicated write use `POST /items`.
|
||
#[utoipa::path(
|
||
post,
|
||
path = "/sharded/items",
|
||
tag = "sharded",
|
||
request_body = ItemRequest,
|
||
responses(
|
||
(status = 201, description = "Item written SINGLE-COPY to its owning region's local store (no WAL append, not replicated)"),
|
||
(status = 400, description = "Invalid request, or the required `x-tidal-ack: local` single-copy opt-in is missing (use POST /items for a replicated write)"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
(status = 503, description = "Owning region unreachable"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn sharded_create_item(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
Json(req): Json<ItemRequest>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let state = node.replica_for(None)?;
|
||
let db = state.db_arc().map_err(ClusterAppError)?;
|
||
let entity = EntityId::new(req.entity_id);
|
||
let metadata = req.metadata.clone();
|
||
sharded_write_route(
|
||
&state,
|
||
&headers,
|
||
req.entity_id,
|
||
"/sharded/items",
|
||
&req,
|
||
move || ShardReplica::apply_item_local(&db, entity, &metadata).map(|_seq| ()),
|
||
StatusCode::CREATED,
|
||
)
|
||
.await
|
||
}
|
||
|
||
/// `POST /sharded/embeddings` — route to the owning region.
|
||
///
|
||
/// **SINGLE-COPY.** Applied to the owning region's LOCAL store via
|
||
/// `ShardReplica::apply_embedding_local`, which performs no WAL append and so
|
||
/// ships nothing to peers — redundancy 1 regardless of the replication factor.
|
||
/// Requires `x-tidal-ack: local` as an explicit acknowledgement of the tradeoff.
|
||
/// For a replicated write use `POST /embeddings`.
|
||
#[utoipa::path(
|
||
post,
|
||
path = "/sharded/embeddings",
|
||
tag = "sharded",
|
||
request_body = EmbeddingRequest,
|
||
responses(
|
||
(status = 204, description = "Embedding written SINGLE-COPY to its owning region's local store (no WAL append, not replicated)"),
|
||
(status = 400, description = "Invalid request, or the required `x-tidal-ack: local` single-copy opt-in is missing (use POST /embeddings for a replicated write)"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
(status = 503, description = "Owning region unreachable"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn sharded_write_embedding(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
Json(req): Json<EmbeddingRequest>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let state = node.replica_for(None)?;
|
||
let db = state.db_arc().map_err(ClusterAppError)?;
|
||
let entity = EntityId::new(req.entity_id);
|
||
let values = req.values.clone();
|
||
sharded_write_route(
|
||
&state,
|
||
&headers,
|
||
req.entity_id,
|
||
"/sharded/embeddings",
|
||
&req,
|
||
move || ShardReplica::apply_embedding_local(&db, entity, &values).map(|_seq| ()),
|
||
StatusCode::NO_CONTENT,
|
||
)
|
||
.await
|
||
}
|
||
|
||
/// `POST /sharded/signals` — route to the owning region.
|
||
///
|
||
/// **SINGLE-COPY.** The signal is applied to the owner's LOCAL store (the
|
||
/// `/sharded/*` surface hash-partitions; it does NOT ride the leader WAL relay —
|
||
/// that is the non-sharded `/signals` surface), so it has redundancy 1 regardless
|
||
/// of the replication factor. Requires `x-tidal-ack: local` as an explicit
|
||
/// acknowledgement of the tradeoff. For a replicated write use `POST /signals`.
|
||
#[utoipa::path(
|
||
post,
|
||
path = "/sharded/signals",
|
||
tag = "sharded",
|
||
request_body = SignalRequest,
|
||
responses(
|
||
(status = 204, description = "Signal written SINGLE-COPY to its owning region's local store (no WAL append, not replicated)"),
|
||
(status = 400, description = "Invalid request, or the required `x-tidal-ack: local` single-copy opt-in is missing (use POST /signals for a replicated write)"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
(status = 503, description = "Owning region unreachable"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn sharded_write_signal(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
Json(req): Json<SignalRequest>,
|
||
) -> std::result::Result<Response, ClusterAppError> {
|
||
let state = node.replica_for(None)?;
|
||
let db = state.db_arc().map_err(ClusterAppError)?;
|
||
let entity = EntityId::new(req.entity_id);
|
||
let signal = req.signal.clone();
|
||
let weight = req.weight;
|
||
sharded_write_route(
|
||
&state,
|
||
&headers,
|
||
req.entity_id,
|
||
"/sharded/signals",
|
||
&req,
|
||
move || {
|
||
db.signal(&signal, entity, weight, Timestamp::now())
|
||
.map_err(ServerError::Tidal)
|
||
},
|
||
StatusCode::NO_CONTENT,
|
||
)
|
||
.await
|
||
}
|
||
|
||
/// `GET /sharded/feed` — scatter-gather a ranked feed across every region.
|
||
/// Degraded semantics preserved verbatim: a timed-out / unreachable / erroring
|
||
/// shard lands in `unavailable_shards` with `degraded: true`, never an error.
|
||
#[utoipa::path(
|
||
get,
|
||
path = "/sharded/feed",
|
||
tag = "sharded",
|
||
params(crate::cluster::routes::ShardedFeedQuery),
|
||
responses(
|
||
(status = 200, description = "Scatter-gather ranked feed", body = ShardedFeedResponse),
|
||
(status = 400, description = "Invalid request"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn sharded_feed(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
Query(query): Query<crate::cluster::routes::ShardedFeedQuery>,
|
||
) -> std::result::Result<Json<ShardedFeedResponse>, ClusterAppError> {
|
||
let state = node.replica_for(None)?;
|
||
let mut builder = Retrieve::builder()
|
||
.profile(query.profile())
|
||
.limit(query.clamped_limit());
|
||
if let Some(user_id) = query.user_id() {
|
||
builder = builder.for_user(user_id);
|
||
}
|
||
let retrieve = builder
|
||
.build()
|
||
.map_err(|e| ClusterAppError(ServerError::Tidal(e.into())))?;
|
||
|
||
let auth = forwarded_auth(&headers);
|
||
let ctx = http_shard_context(&state, auth).map_err(ClusterAppError)?;
|
||
let shards = sharded_region_ids(&state);
|
||
let names = sharded_region_names(&state);
|
||
let deadline_ms = query.deadline_ms();
|
||
let (result, meta) = offload_region_read(move || {
|
||
scatter_gather_retrieve_http(&ctx, &retrieve, &shards, &names, deadline_ms)
|
||
})
|
||
.await?;
|
||
|
||
Ok(Json(ShardedFeedResponse {
|
||
items: feed_items(&result.items),
|
||
total_candidates: result.total_candidates,
|
||
scatter_gather: ScatterGatherInfo::from(meta),
|
||
}))
|
||
}
|
||
|
||
/// `GET /sharded/search` — scatter-gather ranked search across every region.
|
||
/// Degraded semantics preserved verbatim (see [`sharded_feed`]).
|
||
#[utoipa::path(
|
||
get,
|
||
path = "/sharded/search",
|
||
tag = "sharded",
|
||
params(crate::cluster::routes::ShardedSearchQuery),
|
||
responses(
|
||
(status = 200, description = "Scatter-gather ranked search", body = ShardedSearchResponse),
|
||
(status = 400, description = "Invalid request"),
|
||
(status = 401, description = "Missing or invalid API key"),
|
||
),
|
||
security(("bearerAuth" = [])),
|
||
)]
|
||
#[allow(clippy::significant_drop_tightening)]
|
||
pub async fn sharded_search(
|
||
State(node): State<Arc<ClusterNode>>,
|
||
headers: HeaderMap,
|
||
Query(query): Query<crate::cluster::routes::ShardedSearchQuery>,
|
||
) -> std::result::Result<Json<ShardedSearchResponse>, ClusterAppError> {
|
||
let state = node.replica_for(None)?;
|
||
let mut builder = Search::builder()
|
||
.query(query.query_text())
|
||
.limit(query.clamped_limit_u32());
|
||
if let Some(user_id) = query.user_id() {
|
||
builder = builder.for_user(user_id);
|
||
}
|
||
let search_query = builder
|
||
.build()
|
||
.map_err(|e| ClusterAppError(ServerError::Tidal(e.into())))?;
|
||
|
||
let auth = forwarded_auth(&headers);
|
||
let ctx = http_shard_context(&state, auth).map_err(ClusterAppError)?;
|
||
let shards = sharded_region_ids(&state);
|
||
let names = sharded_region_names(&state);
|
||
let deadline_ms = query.deadline_ms();
|
||
let (result, meta) = offload_region_read(move || {
|
||
scatter_gather_search_http(&ctx, &search_query, &shards, &names, deadline_ms)
|
||
})
|
||
.await?;
|
||
|
||
Ok(Json(ShardedSearchResponse {
|
||
items: search_items(&result.items),
|
||
total_candidates: result.total_candidates,
|
||
scatter_gather: ScatterGatherInfo::from(meta),
|
||
}))
|
||
}
|
||
|
||
/// Run a blocking read/write `TidalDb` call off the reactor (see
|
||
/// [`crate::offload::offload_read`]) and map its error into a [`ClusterAppError`].
|
||
async fn offload_region_read<F, T>(f: F) -> std::result::Result<T, ClusterAppError>
|
||
where
|
||
F: FnOnce() -> Result<T> + Send + 'static,
|
||
T: Send + 'static,
|
||
{
|
||
offload_read(f).await.map_err(ClusterAppError)
|
||
}
|
||
|
||
#[cfg(test)]
|
||
#[allow(clippy::unwrap_used)]
|
||
mod sharded_optin_tests {
|
||
//! The `/sharded/*` single-copy opt-in gate. A caller must not be able to
|
||
//! write redundancy-1 data by accident: the surface answered `201`/`204` with
|
||
//! nothing saying so, and a live probe through it was filed as a durability
|
||
//! incident that had to be retracted.
|
||
use super::{ACK_LOCAL, require_local_ack};
|
||
use crate::cluster::forward::ACK_HEADER;
|
||
use reqwest::header::{HeaderMap, HeaderValue};
|
||
|
||
fn ack(value: &str) -> HeaderMap {
|
||
let mut h = HeaderMap::new();
|
||
h.insert(ACK_HEADER, HeaderValue::from_str(value).unwrap());
|
||
h
|
||
}
|
||
|
||
#[test]
|
||
fn the_explicit_optin_is_accepted() {
|
||
assert!(require_local_ack(&ack(ACK_LOCAL), "/sharded/items").is_ok());
|
||
}
|
||
|
||
#[test]
|
||
fn a_missing_header_is_rejected_naming_the_header_and_the_alternative() {
|
||
let err = require_local_ack(&HeaderMap::new(), "/sharded/embeddings")
|
||
.expect_err("a write with no opt-in must be rejected");
|
||
let msg = err.to_string();
|
||
assert!(msg.contains(ACK_HEADER), "must name the header: {msg}");
|
||
assert!(msg.contains(ACK_LOCAL), "must name the value: {msg}");
|
||
assert!(
|
||
msg.contains("/embeddings"),
|
||
"must name the REPLICATING alternative: {msg}"
|
||
);
|
||
assert!(
|
||
msg.contains("SINGLE-COPY"),
|
||
"must say what the caller was about to get: {msg}"
|
||
);
|
||
}
|
||
|
||
/// The replicating ack modes are not an opt-in to single-copy. A caller who
|
||
/// asked for `quorum` most emphatically did not ask for redundancy 1.
|
||
#[test]
|
||
fn a_replicating_ack_mode_is_not_an_optin() {
|
||
for mode in ["leader", "quorum", "", "LOCAL"] {
|
||
let err = require_local_ack(&ack(mode), "/sharded/signals")
|
||
.expect_err("only the exact value `local` opts in");
|
||
assert!(err.to_string().contains("/signals"), "mode {mode:?}");
|
||
}
|
||
}
|
||
|
||
/// Every write route derives its replicating alternative from its own path,
|
||
/// so the message cannot drift from the routing table.
|
||
#[test]
|
||
fn the_alternative_is_derived_from_the_route() {
|
||
for (sharded, replicating) in [
|
||
("/sharded/items", "/items"),
|
||
("/sharded/embeddings", "/embeddings"),
|
||
("/sharded/signals", "/signals"),
|
||
] {
|
||
let msg = require_local_ack(&HeaderMap::new(), sharded)
|
||
.expect_err("no opt-in")
|
||
.to_string();
|
||
assert!(
|
||
msg.contains(&format!("POST {replicating} ")),
|
||
"{sharded} must point at {replicating}: {msg}"
|
||
);
|
||
}
|
||
}
|
||
}
|
||
|
||
#[cfg(test)]
|
||
#[allow(clippy::unwrap_used)]
|
||
mod aggregate_region_row_tests {
|
||
//! The apply-burst false-partition fix, at the row-assembly seam: a
|
||
//! slow-but-alive peer (HTTP probe failed, gRPC contact fresh) stays
|
||
//! `reachable: true`; a genuinely silent peer (no reply, no gRPC contact)
|
||
//! is still flagged `reachable: false, partitioned: true` — and now reports
|
||
//! its frontier as UNKNOWN rather than a fabricated worst case.
|
||
use super::aggregate_region_row;
|
||
|
||
const HWM: Option<u64> = Some(1_000);
|
||
|
||
#[test]
|
||
fn peer_that_answered_is_reachable_with_real_lag() {
|
||
let json = serde_json::json!({
|
||
"applied_events": 940u64,
|
||
"version": "0.1.0+dev",
|
||
"partitioned": [],
|
||
});
|
||
// grpc_fresh / leader_mark are irrelevant when the peer answered.
|
||
let row = aggregate_region_row("eu-west".into(), Some(&json), "us-east", HWM, false, None);
|
||
assert!(row.reachable);
|
||
assert!(!row.partitioned);
|
||
assert_eq!(row.applied_events, Some(940));
|
||
assert_eq!(row.lag_events, Some(60), "lag = HWM - applied");
|
||
assert_eq!(row.version, "0.1.0+dev");
|
||
}
|
||
|
||
#[test]
|
||
fn slow_but_alive_peer_stays_reachable_with_honest_lag() {
|
||
// THE FIX: the HTTP probe timed out (json: None) but the leader has fresh
|
||
// gRPC contact, so the peer is alive — its HTTP control-plane is just
|
||
// starved under the apply burst. It must NOT be flagged partitioned, and
|
||
// its lag is the HONEST gap from the leader's ack mark, not worst-case.
|
||
let row = aggregate_region_row("eu-west".into(), None, "us-east", HWM, true, Some(980));
|
||
assert!(
|
||
row.reachable,
|
||
"a peer with fresh gRPC contact must stay reachable despite an HTTP-probe timeout"
|
||
);
|
||
assert!(
|
||
!row.partitioned,
|
||
"a slow-but-alive peer must NOT be marked partitioned"
|
||
);
|
||
assert_eq!(
|
||
row.applied_events,
|
||
Some(980),
|
||
"reports the leader's ack mark"
|
||
);
|
||
assert_eq!(
|
||
row.lag_events,
|
||
Some(20),
|
||
"honest lag = HWM - ack mark, NOT worst-case leader_last_seq"
|
||
);
|
||
}
|
||
|
||
/// A leader with fresh gRPC contact but NO ack mark for the peer knows the
|
||
/// peer is alive and nothing about its frontier. Reporting `applied 0` there
|
||
/// was the same fabrication one arm down.
|
||
#[test]
|
||
fn alive_peer_without_an_ack_mark_reports_an_unknown_frontier() {
|
||
let row = aggregate_region_row("eu-west".into(), None, "us-east", HWM, true, None);
|
||
assert!(row.reachable, "gRPC contact proves the peer is alive");
|
||
assert!(!row.partitioned);
|
||
assert_eq!(
|
||
row.applied_events, None,
|
||
"no ack mark ⇒ unknown frontier, never 0"
|
||
);
|
||
assert_eq!(
|
||
row.lag_events, None,
|
||
"a lag cannot be derived from an unknown frontier"
|
||
);
|
||
}
|
||
|
||
#[test]
|
||
fn genuinely_silent_peer_is_flagged_unreachable_with_an_unknown_frontier() {
|
||
// No HTTP reply AND no recent gRPC contact (a real partition / dead peer):
|
||
// the honest unreachable verdict stands — the chaos-suite contract. A real
|
||
// TCP severance kills both the HTTP probe and the gRPC ship, landing here.
|
||
let row = aggregate_region_row("ap-south".into(), None, "us-east", HWM, false, None);
|
||
assert!(
|
||
!row.reachable,
|
||
"a peer with no reply and no gRPC contact is genuinely unreachable"
|
||
);
|
||
assert!(
|
||
row.partitioned,
|
||
"a genuinely unreachable peer is partitioned"
|
||
);
|
||
// THE 04a FIX: this used to be `Some(HWM)` — a 500ms probe timeout
|
||
// rendered as the leader's entire history as a deficit. Live, that was
|
||
// 13.3M events against a cluster whose `shards[]` rows all read lag 0.
|
||
assert_eq!(
|
||
row.applied_events, None,
|
||
"an unreachable peer's frontier is UNKNOWN, not 0"
|
||
);
|
||
assert_eq!(
|
||
row.lag_events, None,
|
||
"no worst-case lag may be manufactured from an unknown frontier"
|
||
);
|
||
}
|
||
|
||
/// When the LEADER's own probe failed there is no high-water-mark, so no lag
|
||
/// can be computed even for a peer that answered about itself. Reporting the
|
||
/// old `hwm.unwrap_or(0) - applied = 0` would have said "converged" about a
|
||
/// cluster nobody had measured — the dangerous direction of the same bug.
|
||
#[test]
|
||
fn unknown_leader_hwm_yields_an_unknown_lag_not_zero() {
|
||
let json = serde_json::json!({ "applied_events": 940u64, "partitioned": [] });
|
||
let row = aggregate_region_row("eu-west".into(), Some(&json), "us-east", None, false, None);
|
||
assert!(row.reachable);
|
||
assert_eq!(
|
||
row.applied_events,
|
||
Some(940),
|
||
"the peer's own report is still honest"
|
||
);
|
||
assert_eq!(
|
||
row.lag_events, None,
|
||
"lag needs BOTH sides; an unknown leader frontier ⇒ unknown lag"
|
||
);
|
||
}
|
||
|
||
/// The leader row is at its own high-water-mark by definition, so its lag is
|
||
/// a known zero rather than an unknown.
|
||
#[test]
|
||
fn leader_row_reports_a_known_zero_lag() {
|
||
let json = serde_json::json!({ "applied_events": 1_000u64, "partitioned": [] });
|
||
let row = aggregate_region_row("us-east".into(), Some(&json), "us-east", HWM, false, None);
|
||
assert_eq!(row.lag_events, Some(0));
|
||
}
|
||
}
|
||
|
||
#[cfg(test)]
|
||
#[allow(clippy::unwrap_used)]
|
||
mod auth_middleware_tests {
|
||
use super::*;
|
||
use crate::cluster::forward::{INTERNAL_MARKER, RELAY_MARKER};
|
||
use crate::cluster::security::{ClusterCreds, NODE_TOKEN_HEADER};
|
||
use axum::body::Body;
|
||
use axum::routing::get;
|
||
use tower::ServiceExt;
|
||
|
||
fn app(creds: Arc<ClusterCreds>) -> Router {
|
||
Router::new()
|
||
.route("/cluster/promote", get(|| async { StatusCode::OK }))
|
||
.layer(middleware::from_fn(move |req: Request, next: Next| {
|
||
cluster_auth_middleware(Arc::clone(&creds), req, next)
|
||
}))
|
||
}
|
||
|
||
fn req(headers: &[(&str, String)]) -> Request<Body> {
|
||
let mut b = Request::builder().method("GET").uri("/cluster/promote");
|
||
for (k, v) in headers {
|
||
b = b.header(*k, v.clone());
|
||
}
|
||
b.body(Body::empty()).unwrap()
|
||
}
|
||
|
||
#[tokio::test]
|
||
async fn sibling_markers_without_token_are_403_when_key_configured() {
|
||
// Both the internal-propagation and relayed-operator-hop markers are
|
||
// pinned: a request that sets one WITHOUT a valid node token is rejected.
|
||
for marker in [INTERNAL_MARKER, RELAY_MARKER] {
|
||
let creds = Arc::new(ClusterCreds::with_keys(None, Some("cluster-secret")));
|
||
let resp = app(creds)
|
||
.oneshot(req(&[(marker, "1".to_string())]))
|
||
.await
|
||
.unwrap();
|
||
assert_eq!(
|
||
resp.status(),
|
||
StatusCode::FORBIDDEN,
|
||
"marker {marker} without a node token must be 403"
|
||
);
|
||
}
|
||
}
|
||
|
||
#[tokio::test]
|
||
async fn sibling_markers_with_valid_token_pass() {
|
||
let creds = Arc::new(ClusterCreds::with_keys(None, Some("cluster-secret")));
|
||
let token = creds
|
||
.mint_node_token("region-a")
|
||
.expect("token mints with a key");
|
||
for marker in [INTERNAL_MARKER, RELAY_MARKER] {
|
||
let resp = app(Arc::clone(&creds))
|
||
.oneshot(req(&[
|
||
(marker, "1".to_string()),
|
||
(NODE_TOKEN_HEADER, token.clone()),
|
||
]))
|
||
.await
|
||
.unwrap();
|
||
assert_eq!(
|
||
resp.status(),
|
||
StatusCode::OK,
|
||
"marker {marker} with a valid node token is honored"
|
||
);
|
||
}
|
||
}
|
||
|
||
#[tokio::test]
|
||
async fn markers_are_hint_only_without_a_cluster_key() {
|
||
// No cluster key ⇒ pre-m11p7 hint-only behavior (no 403), backward compat.
|
||
let creds = Arc::new(ClusterCreds::unauthenticated());
|
||
let resp = app(creds)
|
||
.oneshot(req(&[(RELAY_MARKER, "1".to_string())]))
|
||
.await
|
||
.unwrap();
|
||
assert_eq!(resp.status(), StatusCode::OK);
|
||
}
|
||
|
||
#[tokio::test]
|
||
async fn external_over_budget_is_429_with_retry_after() {
|
||
// 1 rps / burst 1: the first external request passes, the second (issued
|
||
// immediately, before any refill) is denied with a Retry-After header.
|
||
let creds = Arc::new(ClusterCreds::with_rate_limit(1.0, 1.0));
|
||
let app = app(creds);
|
||
let first = app.clone().oneshot(req(&[])).await.unwrap();
|
||
assert_eq!(first.status(), StatusCode::OK);
|
||
let second = app.oneshot(req(&[])).await.unwrap();
|
||
assert_eq!(second.status(), StatusCode::TOO_MANY_REQUESTS);
|
||
assert!(
|
||
second.headers().get("retry-after").is_some(),
|
||
"a 429 must carry Retry-After"
|
||
);
|
||
}
|
||
|
||
#[tokio::test]
|
||
async fn verified_node_is_exempt_from_the_rate_limit() {
|
||
// A cluster key (mint/verify tokens) AND a 1-rps bucket: a verified sibling
|
||
// NODE is never throttled, even well past the external budget.
|
||
let creds = Arc::new(ClusterCreds::with_cluster_key_and_rate_limit(
|
||
"cluster-secret",
|
||
1.0,
|
||
1.0,
|
||
));
|
||
let token = creds.mint_node_token("region-a").expect("token");
|
||
let app = app(creds);
|
||
for i in 0..5 {
|
||
let resp = app
|
||
.clone()
|
||
.oneshot(req(&[(NODE_TOKEN_HEADER, token.clone())]))
|
||
.await
|
||
.unwrap();
|
||
assert_eq!(
|
||
resp.status(),
|
||
StatusCode::OK,
|
||
"request {i}: a verified node is exempt from the external budget"
|
||
);
|
||
}
|
||
}
|
||
}
|
||
|
||
/// m12p4 cross-shard unified reads — the merge + wire-parse logic that the three
|
||
/// read handlers share. These prove the disjoint-source contract (SUM totals,
|
||
/// score-merge, truncate) and the honest-degraded accounting WITHOUT a live
|
||
/// cluster; the multi-process partial-placement fan-out is covered end-to-end by
|
||
/// `tests/cluster_cross_shard_reads.rs` (the `cluster-e2e` integration suite).
|
||
#[cfg(test)]
|
||
#[allow(clippy::unwrap_used, clippy::float_cmp)]
|
||
mod cross_shard_tests {
|
||
use super::*;
|
||
|
||
fn rr(entity: u64, score: f64) -> tidaldb::query::RetrieveResult {
|
||
tidaldb::query::RetrieveResult {
|
||
entity_id: EntityId::new(entity),
|
||
score,
|
||
rank: 0,
|
||
signals: Vec::new(),
|
||
reasons: Vec::new(),
|
||
}
|
||
}
|
||
|
||
#[test]
|
||
fn merge_sums_disjoint_totals_and_score_sorts() {
|
||
// Local group returned 2 items / saw 10 candidates; one remote group
|
||
// returned 2 items / saw 7 candidates. The groups own disjoint keys, so
|
||
// the merged total SUMS (17), and the page is score-sorted descending.
|
||
let local = vec![rr(1, 0.9), rr(2, 0.4)];
|
||
let remote = vec![RemoteGroupSlice {
|
||
items: vec![rr(3, 0.7), rr(4, 0.2)],
|
||
total_candidates: 7,
|
||
}];
|
||
let merged = merge_cross_shard(
|
||
local,
|
||
10,
|
||
remote,
|
||
Vec::new(),
|
||
3,
|
||
|it| it.score,
|
||
|it, r| {
|
||
it.rank = r;
|
||
},
|
||
);
|
||
assert_eq!(merged.total_candidates, 17, "disjoint groups SUM totals");
|
||
assert!(!merged.degraded(), "all groups reachable ⇒ not degraded");
|
||
let ids: Vec<u64> = merged.items.iter().map(|i| i.entity_id.as_u64()).collect();
|
||
assert_eq!(ids, vec![1, 3, 2], "score-sorted then truncated to limit=3");
|
||
// Rank is re-stamped over the MERGED order (1-based), so a remote item
|
||
// (id 3, which arrived with rank 0) gets a coherent global page rank, not 0.
|
||
let ranks: Vec<usize> = merged.items.iter().map(|i| i.rank).collect();
|
||
assert_eq!(
|
||
ranks,
|
||
vec![1, 2, 3],
|
||
"merged page rank is global 1-based, no rank-0 hole"
|
||
);
|
||
}
|
||
|
||
#[test]
|
||
fn merge_marks_unreachable_group_degraded_never_hard_fails() {
|
||
// One missing group came back (1 item), another was unreachable. The read
|
||
// returns the partial page AND surfaces the unreachable group by name —
|
||
// never a silent truncation, never an error.
|
||
let local = vec![rr(1, 0.5)];
|
||
let remote = vec![RemoteGroupSlice {
|
||
items: vec![rr(2, 0.8)],
|
||
total_candidates: 4,
|
||
}];
|
||
let merged = merge_cross_shard(
|
||
local,
|
||
3,
|
||
remote,
|
||
vec!["s2".to_string()],
|
||
10,
|
||
|it| it.score,
|
||
|it, r| {
|
||
it.rank = r;
|
||
},
|
||
);
|
||
assert!(merged.degraded(), "an unreachable group degrades the read");
|
||
assert_eq!(merged.unavailable_shards, vec!["s2".to_string()]);
|
||
// The reachable items still merge and rank (degraded ≠ empty).
|
||
let ids: Vec<u64> = merged.items.iter().map(|i| i.entity_id.as_u64()).collect();
|
||
assert_eq!(ids, vec![2, 1]);
|
||
assert_eq!(merged.total_candidates, 7, "local 3 + remote 4");
|
||
}
|
||
|
||
#[test]
|
||
fn parse_group_slice_reads_items_and_total_with_fallback() {
|
||
let body = serde_json::json!({
|
||
"items": [{"entity_id": 7, "score": 0.6}, {"entity_id": 8, "score": 0.3}],
|
||
"total_candidates": 42
|
||
});
|
||
let slice = parse_group_slice(&body, |v| {
|
||
parse_scored_item(v).map(|(e, s)| rr(e.as_u64(), s))
|
||
});
|
||
assert_eq!(slice.total_candidates, 42);
|
||
assert_eq!(slice.items.len(), 2);
|
||
assert_eq!(slice.items[0].entity_id.as_u64(), 7);
|
||
|
||
// Missing total_candidates falls back to the item count, never below it.
|
||
let body2 = serde_json::json!({ "items": [{"entity_id": 1, "score": 0.1}] });
|
||
let slice2 = parse_group_slice(&body2, |v| {
|
||
parse_scored_item(v).map(|(e, s)| rr(e.as_u64(), s))
|
||
});
|
||
assert_eq!(slice2.total_candidates, 1, "fallback = item count");
|
||
|
||
// A malformed element is skipped, never a panic / hard error.
|
||
let body3 = serde_json::json!({
|
||
"items": [{"entity_id": 1, "score": 0.1}, {"oops": true}],
|
||
"total_candidates": 5
|
||
});
|
||
let slice3 = parse_group_slice(&body3, |v| {
|
||
parse_scored_item(v).map(|(e, s)| rr(e.as_u64(), s))
|
||
});
|
||
assert_eq!(slice3.items.len(), 1, "the bad element is dropped");
|
||
}
|
||
|
||
#[test]
|
||
fn parse_vector_match_reads_distance_field() {
|
||
// The vector wire shape carries `distance`, not `score` — the probe merges
|
||
// by ascending distance (negated in the handler's merge key).
|
||
let v = serde_json::json!({ "entity_id": 9, "distance": 0.25 });
|
||
let m = parse_vector_match(&v).unwrap();
|
||
assert_eq!(m.id, 9);
|
||
assert_eq!(m.distance, 0.25_f32);
|
||
assert!(parse_vector_match(&serde_json::json!({"entity_id": 9})).is_none());
|
||
}
|
||
|
||
#[test]
|
||
fn vector_merge_keeps_nearest_first_across_groups() {
|
||
// Two groups' nearest sets merge so the globally-closest (smallest
|
||
// distance) come first after the negated-distance descending sort.
|
||
fn vr(id: u64, distance: f32) -> tidaldb::storage::vector::VectorSearchResult {
|
||
tidaldb::storage::vector::VectorSearchResult { id, distance }
|
||
}
|
||
let local = vec![vr(1, 0.5), vr(2, 1.5)];
|
||
let remote = vec![RemoteGroupSlice {
|
||
items: vec![vr(3, 0.2), vr(4, 2.0)],
|
||
total_candidates: 2,
|
||
}];
|
||
let merged = merge_cross_shard(
|
||
local,
|
||
2,
|
||
remote,
|
||
Vec::new(),
|
||
3,
|
||
|r| -f64::from(r.distance),
|
||
|_r, _rank| {},
|
||
);
|
||
let ids: Vec<u64> = merged.items.iter().map(|r| r.id).collect();
|
||
assert_eq!(ids, vec![3, 1, 2], "ascending distance, truncated to k=3");
|
||
}
|
||
}
|
||
|
||
/// The follower→leader forward retry policy (m11p5 forward-stall fix), tested at
|
||
/// its pure decision seam so the live `forward_write` loop is the exact policy
|
||
/// these tests pin — no cluster, no real HTTP needed.
|
||
///
|
||
/// The bug these guard against: under a sustained 1536-dim quorum write-burst
|
||
/// directed at a FOLLOWER, the leader's apply plane runs hot, so a single forward
|
||
/// either times out (slow-but-alive leader) or returns a fast 429 (bounded write
|
||
/// pool shed). The OLD path declared either one "leader unreachable" on the FIRST
|
||
/// miss → a 503 that stalled ingest. The fix RETRIES with backoff so the load
|
||
/// degrades (slower) rather than fails.
|
||
#[cfg(test)]
|
||
#[allow(clippy::unwrap_used)]
|
||
mod forward_retry_tests {
|
||
use std::time::Duration;
|
||
|
||
use super::{
|
||
FORWARD_BACKPRESSURE_BACKOFF_MS, FORWARD_MAX_ATTEMPTS, FORWARD_TIMEOUT_BACKOFF_MS,
|
||
ForwardStep, backpressure_backoff_ms, classify_forward_attempt, forward,
|
||
};
|
||
use axum::http::StatusCode;
|
||
|
||
fn forwarded(status: StatusCode, body: serde_json::Value) -> forward::ForwardedResponse {
|
||
forward::ForwardedResponse {
|
||
seq: None,
|
||
deduplicated: false,
|
||
status,
|
||
body,
|
||
}
|
||
}
|
||
|
||
#[test]
|
||
fn success_relays_immediately() {
|
||
// A 2xx leader verdict is handed straight back — never retried.
|
||
let outcome = Ok::<_, String>(forwarded(StatusCode::CREATED, serde_json::Value::Null));
|
||
assert_eq!(
|
||
classify_forward_attempt(&outcome, 1, Duration::ZERO),
|
||
ForwardStep::Relay
|
||
);
|
||
}
|
||
|
||
#[test]
|
||
fn leader_429_with_attempts_left_retries_with_hinted_backoff() {
|
||
// The leader's bounded write pool shed this write (429 + retry_after_ms):
|
||
// it created no log entry, so re-forwarding is safe. We retry, honoring
|
||
// the leader's hint.
|
||
let outcome = Ok::<_, String>(forwarded(
|
||
StatusCode::TOO_MANY_REQUESTS,
|
||
serde_json::json!({ "retry_after_ms": 80 }),
|
||
));
|
||
assert_eq!(
|
||
classify_forward_attempt(&outcome, 1, Duration::ZERO),
|
||
ForwardStep::Retry { backoff_ms: 80 }
|
||
);
|
||
}
|
||
|
||
#[test]
|
||
fn final_attempt_429_relays_honest_retryable_not_misleading_503() {
|
||
// On the LAST attempt a 429 is relayed verbatim: the targeted-follower
|
||
// client sees a retryable 429 (honest backpressure), NOT a 503 that would
|
||
// wrongly say the leader is unreachable.
|
||
let outcome = Ok::<_, String>(forwarded(
|
||
StatusCode::TOO_MANY_REQUESTS,
|
||
serde_json::Value::Null,
|
||
));
|
||
assert_eq!(
|
||
classify_forward_attempt(&outcome, FORWARD_MAX_ATTEMPTS, Duration::ZERO),
|
||
ForwardStep::Relay
|
||
);
|
||
}
|
||
|
||
#[test]
|
||
fn timeout_with_budget_retries_then_fails_when_exhausted() {
|
||
// A slow-but-alive leader (transport timeout) is WAITED-ON: retried while
|
||
// attempts AND wall-clock budget remain, then surfaced as the typed 503
|
||
// only once the budget cannot fit another full attempt.
|
||
let timeout: Result<forward::ForwardedResponse, String> = Err("operation timed out".into());
|
||
assert_eq!(
|
||
classify_forward_attempt(&timeout, 1, Duration::ZERO),
|
||
ForwardStep::Retry {
|
||
backoff_ms: FORWARD_TIMEOUT_BACKOFF_MS
|
||
},
|
||
"first timeout with full budget must retry, not fail"
|
||
);
|
||
// Last attempt: never retry regardless of budget.
|
||
assert_eq!(
|
||
classify_forward_attempt(&timeout, FORWARD_MAX_ATTEMPTS, Duration::ZERO),
|
||
ForwardStep::Fail,
|
||
"the final attempt's timeout must fail (no attempts left)"
|
||
);
|
||
// Budget nearly spent: another full forward + backoff would overrun
|
||
// FORWARD_TOTAL_BUDGET, so stop and emit the typed 503 ourselves rather
|
||
// than let the outer route TimeoutLayer cut an opaque 408.
|
||
let nearly_spent = super::FORWARD_TOTAL_BUDGET - forward::FORWARD_REQUEST_TIMEOUT;
|
||
assert_eq!(
|
||
classify_forward_attempt(&timeout, 1, nearly_spent),
|
||
ForwardStep::Fail,
|
||
"a timeout with no room for another full attempt must fail, not retry"
|
||
);
|
||
}
|
||
|
||
#[test]
|
||
fn backpressure_backoff_clamps_absent_and_hostile_hints() {
|
||
// Absent hint → the short floor (pool drains in worker-thread time).
|
||
assert_eq!(
|
||
backpressure_backoff_ms(&serde_json::Value::Null),
|
||
FORWARD_BACKPRESSURE_BACKOFF_MS
|
||
);
|
||
// Below-floor hint → floor (never busy-spin).
|
||
assert_eq!(
|
||
backpressure_backoff_ms(&serde_json::json!({ "retry_after_ms": 1 })),
|
||
FORWARD_BACKPRESSURE_BACKOFF_MS
|
||
);
|
||
// Hostile/huge hint → capped so a garbled leader cannot park the forward
|
||
// for an unbounded time.
|
||
assert_eq!(
|
||
backpressure_backoff_ms(&serde_json::json!({ "retry_after_ms": 9_999_999u64 })),
|
||
1_000
|
||
);
|
||
}
|
||
}
|