tidaldb/tidal-server/src/offload.rs
jx12n a946c6128c fix(m12-rc13): read-SLA collapse + WAL_RETENTION_SEGMENTS 16 + tidalctl S3 DR
Read-SLA fix (rc12→rc13 — cpu-cgroup starvation → multi-second p99 + churning
elections):
- offload.rs: add SEARCH_GATE semaphore (core_count+1 permits, 50ms shed to 429)
  so per-shard searches gate on CPU, not reactor threads; concurrent scatter_merge
  fan-out (join_all) replaces the serial blocking offload_region_read loop
- node.rs: scatter_merge → async; per-shard futures run via offload_search
  (each acquires one SEARCH_GATE permit, moves it into spawn_blocking so the
  permit is held for the search's full CPU lifetime)
- main.rs: explicit tokio runtime with worker_threads floored at 4, independent
  of the cgroup quota — keeps the control plane (heartbeat/election/apply) on its
  own workers even when quota < 4
- k8s statefulset: CPU limit 2→3 (was: available_parallelism()=2 → only 2 async
  workers; search burst starved the reactor)
- tidal/wal/compaction.rs: WAL_RETENTION_SEGMENTS 4→16 (64 MiB→256 MiB per-shard
  catch-up window; a briefly-down follower across a rolling restart streams up
  instead of forcing snapshot reseed; disk floor 768 MiB/pod, self-trimming)
- cluster_reseed.rs: OFFLINE_ITEMS 1800→5600 to exceed the new 16-segment
  retention window (19 segs > 17); fix sequential quarantine/reseed race via
  await_status_bool

tidalctl S3/R2 backup DR:
- tidalctl/Cargo.toml: aws-config, aws-sdk-s3, aws-credential-types, tokio, tempfile
- commands/s3.rs: S3Target + export_dir (upload every file, manifest last as
  atomicity marker) + import_to_dir (download prefix into temp staging dir)
- commands/backup.rs: run_backup/run_restore accept Option<&S3Target>; S3 export
  is additive after local fsync barrier; S3 import stages into TempDir then runs
  the unchanged verified restore on it
- main.rs: --s3-endpoint / --s3-bucket / --s3-prefix flags; all-or-nothing
  endpoint+bucket validation; usage updated

tidal-stress/k8s: recall-rc12-spread-job, soak-nightly-cronjob, soak-monitor,
soak-results-pvc, t5-readtput-job manifests
2026-06-17 15:47:37 -06:00

523 lines
22 KiB
Rust

//! Blocking-work offload primitives shared by the standalone and cluster HTTP
//! surfaces.
//!
//! Two distinct hazards, two distinct tools:
//!
//! * **Reads** (`RETRIEVE` / `SEARCH` / text-index reload) are synchronous,
//! CPU/IO-bound `TidalDb` calls. Running them inside an axum handler pins the
//! reactor worker for the whole query, so a burst can starve every other
//! in-flight request. [`offload_read`] hands them to tokio's purpose-built
//! blocking pool via `spawn_blocking`. Reads never touch the gRPC transport,
//! so a `spawn_blocking` thread (which carries an ambient runtime handle) is
//! fine.
//!
//! * **Cluster writes** (`/signals`, `/cluster/heal`) ship WAL segments to
//! followers through [`tidal_net::GrpcTransport`], whose `send_segment` blocks
//! on its own runtime and *asserts* it is not invoked from within another
//! runtime (`Handle::try_current().is_err()`). That rules out both the axum
//! reactor and `spawn_blocking` (its threads carry a runtime handle). The
//! previous design spawned a fresh OS thread per request — unbounded growth on
//! the hottest cluster path. [`ClusterWritePool`] replaces that with a fixed
//! set of runtime-free OS threads draining a bounded queue, so write
//! concurrency is capped and threads are reused. A full queue degrades to
//! [`TidalError::Backpressure`] (→ 429), matching the engine's own
//! flow-control semantics, instead of a hard 500.
//!
//! Both surfaces route through this module so the standalone and cluster paths
//! cannot drift in how they treat blocking work.
use std::sync::OnceLock;
use crossbeam::channel::{Receiver, Sender, TrySendError};
use tidaldb::TidalError;
use tokio::sync::Semaphore;
use crate::error::{Result, ServerError};
/// Retry-after hint (ms) returned when the read-admission gate is saturated. Short
/// by design (mirrors the write pool's `WRITE_BACKPRESSURE_RETRY_AFTER_MS`): a
/// blocking read drains in worker-thread time, so a quick retry once a slot frees
/// beats parking the request.
const READ_BACKPRESSURE_RETRY_AFTER_MS: u64 = 50;
/// Process-wide cap on CONCURRENT blocking reads, so a read storm sheds as a fast
/// 429 instead of piling unboundedly onto tokio's shared 512-thread blocking pool
/// (where it would also starve the leader's WAL-serve `spawn_blocking`) and
/// climbing into a multi-second p99. Lazily sized from the core count with
/// generous headroom, capped well below the blocking-pool limit.
static READ_GATE: OnceLock<Semaphore> = OnceLock::new();
/// Process-wide cap on CONCURRENT single-db blocking reads admitted by
/// [`offload_read`] (standalone reads + the single-group internal `?shard` hops).
/// Each such request runs exactly ONE engine search, so a per-request permit is
/// the right unit here. Re-sized DOWN from the m12p6 `*16` fan-out (which let a
/// 2-quota pod admit 32 reads that, multiplied by the 3-shard serial scatter,
/// became ~96 concurrent CPU-bound searches and a multi-second p99): a small
/// multiple of the core count buffers a brief burst, then sheds as a fast 429.
fn read_inflight_limit() -> usize {
core_parallelism()
.saturating_mul(READ_INFLIGHT_PER_CORE)
.clamp(MIN_READ_INFLIGHT, MAX_READ_INFLIGHT)
}
/// Buffer factor for [`read_inflight_limit`]: a few queued single-db reads per
/// core absorb a burst without letting the blocking pool climb into a backlog.
const READ_INFLIGHT_PER_CORE: usize = 4;
const MIN_READ_INFLIGHT: usize = 4;
const MAX_READ_INFLIGHT: usize = 64;
/// The quota-aware core count. `available_parallelism` honours the cgroup CPU
/// quota (e.g. `limits.cpu="3"` -> 3 even on a 4-core node), so this is the
/// number of CPU-bound shard-searches the pod may run truly in parallel.
fn core_parallelism() -> usize {
std::thread::available_parallelism().map_or(2, std::num::NonZeroUsize::get)
}
/// Process-wide cap on CONCURRENT CPU-bound shard-searches across ALL in-flight
/// cross-shard reads. Sized to the core count (with a +1 of slack so one
/// in-flight search blocked on the embedding-registry read-lock cannot idle a
/// core), so the async reactor and the election/heartbeat/apply loops always
/// retain CPU. EVERY per-shard search in [`offload_search`] acquires one
/// permit; a 3-shard read therefore consumes up to 3 permits, and the gate — not
/// a per-read counter — is what bounds total search parallelism. Excess sheds as
/// a fast 429 (see `SEARCH_ADMIT_TIMEOUT_MS`).
static SEARCH_GATE: OnceLock<Semaphore> = OnceLock::new();
fn search_inflight_limit() -> usize {
core_parallelism().saturating_add(1).clamp(2, 16)
}
/// Max time a per-shard search waits for a [`SEARCH_GATE`] permit before shedding
/// as backpressure. Same 50ms budget as the read/write gates: a search drains in
/// worker-thread time, so a quick retry beats parking the request for seconds.
const SEARCH_ADMIT_TIMEOUT_MS: u64 = READ_BACKPRESSURE_RETRY_AFTER_MS;
/// Run a blocking READ-only query (RETRIEVE / SEARCH / text-index reload) on
/// tokio's blocking pool, off the async reactor, and await its result.
///
/// `TidalDb::retrieve`/`search`/`reload_text_index` are synchronous and
/// CPU/IO-bound; running them directly in an axum handler blocks the reactor
/// worker for the whole query, so a burst of feed/search requests can pin every
/// worker and stall unrelated requests. Reads never call
/// [`tidal_net::GrpcTransport::send_segment`], so unlike cluster writes they do
/// NOT need a runtime-free OS thread — `spawn_blocking` is the right tool.
///
/// # Errors
///
/// Returns the closure's own [`ServerError`], or [`ServerError::Cluster`] if the
/// blocking task panicked or was cancelled (mapped to a 500 by
/// [`crate::router::status_from_error`]).
pub async fn offload_read<F, T>(f: F) -> Result<T>
where
F: FnOnce() -> Result<T> + Send + 'static,
T: Send + 'static,
{
// m12p6 read admission: bound concurrent blocking reads. A healthy burst waits
// briefly for a slot (buffer-then-shed, matching the ClusterWritePool
// contract); a sustained read overload sheds as engine-native backpressure
// (429) instead of an unbounded climb into a 36s p99 on the shared blocking
// pool. A genuinely-closed gate never happens (the static lives for the
// process), so the closed arm is a defensive 500.
let gate = READ_GATE.get_or_init(|| Semaphore::new(read_inflight_limit()));
let _permit = match tokio::time::timeout(
std::time::Duration::from_millis(READ_BACKPRESSURE_RETRY_AFTER_MS),
gate.acquire(),
)
.await
{
Ok(Ok(permit)) => permit,
Ok(Err(_closed)) => {
return Err(ServerError::Cluster("read admission gate closed".into()));
}
Err(_elapsed) => {
return Err(ServerError::Tidal(TidalError::Backpressure {
retry_after_ms: READ_BACKPRESSURE_RETRY_AFTER_MS,
}));
}
};
tokio::task::spawn_blocking(f)
.await
// A JoinError means the blocking task panicked or was cancelled — the
// query produced no answer, so surface it as a server error (500).
.map_err(|e| ServerError::Cluster(format!("blocking read worker failed: {e}")))?
// `_permit` drops here, releasing the read slot.
}
/// Outcome of one shard's CPU-bound search inside a cross-shard scatter.
///
/// Either the engine's `(items, total_candidates)` slice, or a per-shard error
/// to be degraded over (NOT to abort the whole read). A 429 here is a real
/// per-shard shed, surfaced like any other shard error so the merge serves
/// survivors.
pub type ShardSearch<T> = std::result::Result<(Vec<T>, usize), ServerError>;
/// Run one CPU-bound per-shard search under the process-wide [`SEARCH_GATE`].
///
/// Off the reactor: acquires a permit (50ms timeout -> fast `Backpressure` 429,
/// never a hang), then `spawn_blocking`s the closure. The permit is moved INTO
/// the blocking closure and dropped only when the search finishes, so the gate
/// reflects searches actually burning a core, not merely admitted ones.
///
/// # Errors
/// * `TidalError::Backpressure` (429) when the gate is saturated for 50ms.
/// * `ServerError::Cluster` (500) if the gate is closed (defensive; the static
/// lives for the process) or the blocking task panicked.
pub async fn offload_search<F, T>(f: F) -> Result<T>
where
F: FnOnce() -> Result<T> + Send + 'static,
T: Send + 'static,
{
let gate = SEARCH_GATE.get_or_init(|| Semaphore::new(search_inflight_limit()));
let permit = match tokio::time::timeout(
std::time::Duration::from_millis(SEARCH_ADMIT_TIMEOUT_MS),
gate.acquire(),
)
.await
{
Ok(Ok(permit)) => permit,
Ok(Err(_closed)) => {
return Err(ServerError::Cluster("search admission gate closed".into()));
}
Err(_elapsed) => {
return Err(ServerError::Tidal(TidalError::Backpressure {
retry_after_ms: SEARCH_ADMIT_TIMEOUT_MS,
}));
}
};
// Move the permit into the blocking task so it is held for the search's whole
// CPU lifetime and released on the blocking thread when the search returns.
tokio::task::spawn_blocking(move || {
let _permit = permit;
f()
})
.await
.map_err(|e| ServerError::Cluster(format!("blocking search worker failed: {e}")))?
}
/// Configuration for the cluster write worker pool.
///
/// Defaults are derived once at startup: `workers` from
/// [`std::thread::available_parallelism`] (clamped to a sane floor/ceiling),
/// `queue_depth` to a small multiple of `workers` so brief bursts queue rather
/// than 429 while a sustained overload still sheds load promptly.
#[derive(Debug, Clone, Copy)]
pub struct ClusterWritePoolConfig {
/// Number of runtime-free OS worker threads draining the queue.
pub workers: usize,
/// Maximum number of queued (not-yet-started) write closures before new
/// submissions are rejected with backpressure.
pub queue_depth: usize,
}
/// Floor on worker count: a single-core host still gets parallel shipping.
const MIN_WRITE_WORKERS: usize = 2;
/// Ceiling on worker count: cluster writes are gRPC-ship bound, not CPU bound,
/// so a large core count does not warrant an unbounded thread set.
const MAX_WRITE_WORKERS: usize = 8;
/// Queued closures permitted per worker before backpressure trips.
const QUEUE_DEPTH_PER_WORKER: usize = 8;
/// Retry-after hint (milliseconds) returned to clients when the cluster write
/// queue is saturated. Short by design: the queue drains in worker-thread time
/// (a single gRPC ship), not in seconds, so a 50ms backoff lets a client retry
/// almost immediately once a slot frees rather than parking it needlessly. Kept
/// in line with the engine's own short backpressure hints so the HTTP surface
/// and the engine advertise consistent retry semantics.
const WRITE_BACKPRESSURE_RETRY_AFTER_MS: u64 = 50;
impl ClusterWritePoolConfig {
/// Build a config for an explicit worker count, deriving `queue_depth`
/// proportionally so a custom size still buffers brief bursts before 429.
#[must_use]
pub fn with_workers(workers: usize) -> Self {
let workers = workers.max(1);
Self {
workers,
queue_depth: workers.saturating_mul(QUEUE_DEPTH_PER_WORKER),
}
}
}
impl Default for ClusterWritePoolConfig {
fn default() -> Self {
let workers = std::thread::available_parallelism()
.map_or(MIN_WRITE_WORKERS, std::num::NonZeroUsize::get)
.clamp(MIN_WRITE_WORKERS, MAX_WRITE_WORKERS);
Self::with_workers(workers)
}
}
/// A unit of work for the cluster write pool: a boxed closure plus the oneshot
/// used to bridge its result back to the awaiting async handler.
type WriteJob = Box<dyn FnOnce() + Send + 'static>;
/// A fixed-size, runtime-free OS-thread pool for cluster write work.
///
/// Created once at cluster startup and shared via the cluster state. Each worker
/// is a plain `std::thread` with NO ambient tokio runtime, so closures may call
/// [`tidal_net::GrpcTransport::send_segment`] (which asserts it is outside a
/// runtime). Submissions exceeding [`ClusterWritePoolConfig::queue_depth`] are
/// rejected with [`TidalError::Backpressure`] rather than spawning unbounded
/// threads.
///
/// Workers drain the queue until the [`Sender`] is dropped (on
/// [`ClusterWritePool::Drop`]), at which point `recv` returns `Err` and each
/// worker exits; the threads are then joined so no work is abandoned.
pub struct ClusterWritePool {
sender: Option<Sender<WriteJob>>,
workers: Vec<std::thread::JoinHandle<()>>,
/// Optional `tidaldb_cluster_write_pool_*` series (queue depth gauge +
/// backpressure-rejection counter), shared with the engine's `/metrics`.
metrics: Option<std::sync::Arc<tidaldb::db::metrics::cluster::ClusterMetrics>>,
}
impl ClusterWritePool {
/// Build the pool and start its worker threads.
///
/// `workers` and `queue_depth` are clamped to at least 1 so the pool always
/// has a live consumer and a bounded queue.
///
/// # Panics
///
/// Panics if a worker OS thread cannot be spawned. This runs once at server
/// startup (before any request is served), so a thread-exhausted host fails
/// loudly at boot rather than degrading every later request — the opposite
/// of the old per-request spawn that turned the same failure into a hot-path
/// 500.
#[must_use]
pub fn new(config: ClusterWritePoolConfig) -> Self {
Self::with_metrics(config, None)
}
/// [`new`](Self::new), wiring the pool's queue-depth gauge and
/// backpressure counter into the engine's cluster metrics (m11p1).
///
/// # Panics
///
/// Panics if a worker OS thread cannot be spawned (startup-time failure,
/// before any request is served — see [`new`](Self::new)).
#[must_use]
pub fn with_metrics(
config: ClusterWritePoolConfig,
metrics: Option<std::sync::Arc<tidaldb::db::metrics::cluster::ClusterMetrics>>,
) -> Self {
let workers = config.workers.max(1);
let queue_depth = config.queue_depth.max(1);
// Bounded so a sustained burst sheds load (429) instead of growing the
// queue without limit. crossbeam's MPMC channel lets every worker pull
// from the same queue without a shared Mutex<Receiver>.
let (sender, receiver) = crossbeam::channel::bounded::<WriteJob>(queue_depth);
let handles = (0..workers)
.map(|i| {
let rx: Receiver<WriteJob> = receiver.clone();
std::thread::Builder::new()
.name(format!("cluster-write-{i}"))
.spawn(move || worker_loop(&rx))
// A thread that cannot start at construction time is a hard
// startup failure, not a per-request degrade. Surfacing it
// as a panic here (at process start, before serving) is
// acceptable and far better than the old per-request spawn.
.expect("spawn cluster write worker thread")
})
.collect();
Self {
sender: Some(sender),
workers: handles,
metrics,
}
}
/// Submit a blocking, runtime-free write closure to the pool and await its
/// result.
///
/// The closure runs on a pool worker thread with no ambient tokio runtime,
/// so it may ship WAL segments over gRPC. Its `Result` is bridged back
/// through a oneshot the caller awaits, so the reactor is never blocked.
///
/// # Errors
///
/// * [`TidalError::Backpressure`] (→ 429) when the queue is full — the work
/// was never enqueued, so the caller may safely retry after backing off.
/// * [`ServerError::Cluster`] (→ 500) if the pool has been shut down, or a
/// worker dropped the job without responding (e.g. mid-shutdown).
pub async fn submit<F, T>(&self, f: F) -> Result<T>
where
F: FnOnce() -> Result<T> + Send + 'static,
T: Send + 'static,
{
let sender = self
.sender
.as_ref()
.ok_or_else(|| ServerError::Cluster("cluster write pool is shut down".into()))?;
let (tx, rx) = tokio::sync::oneshot::channel();
let job: WriteJob = Box::new(move || {
// Ignore send errors: the receiver is only gone if the request
// future was cancelled, in which case nobody is waiting.
let _ = tx.send(f());
});
match sender.try_send(job) {
Ok(()) => {
// O(1) channel-depth read; a slightly stale gauge is fine.
if let Some(m) = &self.metrics {
m.set_write_pool_depth(sender.len() as u64);
}
}
Err(TrySendError::Full(_)) => {
// Bounded queue saturated: degrade to engine-native backpressure
// (429) instead of growing threads/queue without limit. The work
// was never enqueued, so it is safe to retry.
if let Some(m) = &self.metrics {
m.incr_write_pool_rejections();
}
return Err(ServerError::Tidal(TidalError::Backpressure {
retry_after_ms: WRITE_BACKPRESSURE_RETRY_AFTER_MS,
}));
}
Err(TrySendError::Disconnected(_)) => {
return Err(ServerError::Cluster(
"cluster write pool has no live workers".into(),
));
}
}
rx.await.map_err(|_| {
ServerError::Cluster("cluster write worker dropped the job without responding".into())
})?
}
}
/// Drain jobs until the [`Sender`] is dropped, then exit so the thread can be
/// joined.
fn worker_loop(rx: &Receiver<WriteJob>) {
// `recv` blocks until a job arrives, and returns `Err` only once every
// `Sender` has been dropped — that is the shutdown signal.
while let Ok(job) = rx.recv() {
job();
}
}
impl Drop for ClusterWritePool {
/// Drop the [`Sender`] so workers see a disconnected queue and exit, then
/// join every worker so no in-flight job is abandoned mid-ship.
fn drop(&mut self) {
// Dropping the sender disconnects the channel; each worker's `recv`
// returns `Err` after draining what it already holds, and the loop exits.
self.sender = None;
for handle in self.workers.drain(..) {
// A worker only panics if a submitted closure panics; that is the
// closure author's bug, not the pool's. Log and continue joining the
// rest so shutdown still completes.
if handle.join().is_err() {
tracing::error!("cluster write worker thread panicked during shutdown");
}
}
}
}
#[cfg(test)]
#[allow(clippy::unwrap_used)]
mod tests {
use std::sync::{
Arc,
atomic::{AtomicUsize, Ordering},
};
use super::*;
#[tokio::test]
async fn pool_processes_submitted_work() {
let pool = ClusterWritePool::new(ClusterWritePoolConfig {
workers: 2,
queue_depth: 8,
});
let counter = Arc::new(AtomicUsize::new(0));
let mut results = Vec::new();
for i in 0..16u64 {
let c = Arc::clone(&counter);
results.push(
pool.submit(move || {
c.fetch_add(1, Ordering::SeqCst);
Ok::<u64, ServerError>(i * 2)
})
.await,
);
}
for (i, r) in results.into_iter().enumerate() {
assert_eq!(r.unwrap(), (i as u64) * 2);
}
assert_eq!(counter.load(Ordering::SeqCst), 16);
}
#[tokio::test]
async fn pool_runs_off_any_tokio_runtime() {
// The whole point of the pool: closures execute with NO ambient runtime
// handle, which is what `GrpcTransport::send_segment` asserts.
let pool = ClusterWritePool::new(ClusterWritePoolConfig {
workers: 1,
queue_depth: 4,
});
let outside = pool
.submit(|| Ok::<bool, ServerError>(tokio::runtime::Handle::try_current().is_err()))
.await
.unwrap();
assert!(
outside,
"pool worker must not carry an ambient tokio runtime"
);
}
#[tokio::test]
async fn full_queue_yields_backpressure() {
// One worker, queue depth of 1: occupy the worker with a job that parks
// until released, fill the single queue slot directly, then prove the
// next submit is rejected with Backpressure (→ 429) — not a 500 and not
// an unbounded thread/queue grow.
let pool = ClusterWritePool::new(ClusterWritePoolConfig {
workers: 1,
queue_depth: 1,
});
// Block the sole worker on a channel until the test releases it. We
// submit the parking job directly through the pool's own sender so the
// worker picks it up, then wait for an explicit "started" signal — no
// sleep-based timing.
let sender = pool.sender.as_ref().unwrap().clone();
let (started_tx, started_rx) = std::sync::mpsc::channel::<()>();
let (release_tx, release_rx) = std::sync::mpsc::channel::<()>();
let park: WriteJob = Box::new(move || {
// Tell the test the worker is now busy, then block until released.
started_tx.send(()).unwrap();
release_rx.recv().unwrap();
});
sender
.try_send(park)
.expect("worker accepts the parking job");
// Deterministic handoff: the worker has dequeued and started `park`, so
// the single queue slot is empty again and the worker is occupied.
started_rx
.recv_timeout(std::time::Duration::from_secs(5))
.expect("worker started the parking job");
// Job B fills the one free queue slot (worker is busy on `park`).
let job_b: WriteJob = Box::new(|| {});
sender.try_send(job_b).expect("queue slot accepts job B");
// Job C: queue is now full -> Backpressure (429), via the real submit path.
let c = pool.submit(|| Ok::<(), ServerError>(())).await;
match c {
Err(ServerError::Tidal(TidalError::Backpressure { .. })) => {}
other => panic!("expected Backpressure, got {other:?}"),
}
// Release the worker so the pool drains and shuts down cleanly.
release_tx.send(()).unwrap();
}
}