fleet remediation: make the workspace gate runnable, then fix what it caught
`cargo test --workspace` could not run at all: dependency resolution failed with "aws-types@1.3.16 requires rustc 1.91.1" on the 1.91.0 default toolchain, so the gate the project documents was dead. Making it run exposed a compile break and two wrong tests that had been invisible for months. Now green end to end: 143 suites, 3155 tests, exit 0. Toolchain - rust-toolchain.toml pins the DEV toolchain to 1.91.1. The published MSRV stays `rust-version = "1.91"` (the engine builds on 1.91.0); only tidalctl's AWS SDK chain needs the patch release, and it now declares that itself. Consumer crates migrated to the current engine API (clean cutover) - iknowyou-engine: `AgentPolicy` gained five m10 read/profile-override fields; the literal now spreads `..AgentPolicy::default()` as the engine's own doc example does, so future fields do not break it again. - forage-engine: `RetrieveResult` gained p1 `reasons`. The app builds its own candidate pool, so it now tags what it knows: PreferenceMatch for the preference-vector blend, SemanticMatch (with the seed item) for similar-to-saved, ExplorationBudget for pinned discoveries. - forage-engine: `url_to_item_id` folded into the u32 item universe. The engine narrows item IDs to a u32 slot in durable per-user state and rejects anything above u32::MAX rather than alias two items forever, so every add_item with a 64-bit FNV hash failed. 9 of 28 smoke tests were failing on this alone. - forage-engine: bridge items read the top-2 preference CLUSTERS via `query_vectors`, not the single centroid from `preference_vectors().get()`. Since m12 that accessor returns only the strongest cluster, so a tech+jazz user whose interests split into two clusters looked single-interest and never bridged. Falls back to top-2 dimensions when a user has one cluster. Reconcile tests corrected to the shipped contract - tidal/tests/m8p3_reconcile_production.rs asserted `3 + 5 == 8` for a windowed count after heal. `take_crdt_snapshot` deliberately keys signal contributions to ONE canonical contributor (ShardId::SINGLE) because signals are relayed from a single writer, so per-node attribution double-counted every replicated event on every reconcile. Merge is therefore LWW on (last_update_ns, score) plus PN-counter per-node max: nodes converge on the more complete accumulator. The old expectation was asserting the bug that fix removed. - Rewrote to assert convergence, count survival (not 0), and no inflation, and added `repeated_reconcile_of_converged_nodes_does_not_creep` - the regression guard for the creep itself, which nothing covered. Pre-commit hook unified - hooks/pre-commit dropped `-D warnings`: each crate's `[lints]` table is the source of truth (`clippy::all`/`unwrap_used` deny, `pedantic` warn), and the flag promoted ~58 deliberate pedantic warnings in integration tests to errors, making every Rust commit impossible. - It now lints all five tidal crates instead of path-matching `tidal/`, which silently skipped tidal-server, tidal-net, tidal-stress, tidalctl and applications/ - the rot above lived in exactly those crates. Ported the CODING_GUIDELINES file-length, println, and unsafe-SAFETY checks from the divergent untracked copy that this replaces. - CONTRIBUTING.md now documents the real commands and the toolchain/MSRV split. Fleet recovery and soak - scripts/restore-fleet.sh: the fail-closed selective restore, promoted out of an ignored tmp/ directory into the repository. Preflights retained storage, digest-pinned images, parked state, and aggregate plus per-PV-node scheduler headroom before the first scale; writes a durable transcript under tmp/restore-logs/ with structured start/error/rollback/complete events. - k8s manifests park the standalone store, the RF3 cluster, and the soak monitor at zero replicas with restore-fleet.sh as the only supported scale-up path. - soak-eval/soak-watch and the nightly CronJob fail closed on stale or missing restart evidence instead of silently skipping the restart-aware half of the gate. - docs/ops/capacity-planning.md corrects the RAM envelope to the real hot-tier formula and separates analytic totals from the measured process envelope.
This commit is contained in:
parent
cdbe9cb453
commit
c97aaa8e5b
@ -14,6 +14,15 @@ cargo test --doc --manifest-path tidal/Cargo.toml
|
||||
cargo test --examples --manifest-path tidal/Cargo.toml
|
||||
```
|
||||
|
||||
### Toolchain
|
||||
|
||||
`rust-toolchain.toml` pins the development toolchain to **1.91.1**; rustup honors
|
||||
it automatically. That pin is not the library's MSRV — `tidaldb` and its siblings
|
||||
publish `rust-version = "1.91"` and build on 1.91.0. The extra patch release is
|
||||
required by `tidalctl`'s AWS SDK chain (`aws-types` declares 1.91.1), and without
|
||||
it every workspace-wide command fails during dependency resolution instead of
|
||||
compiling, which silently disables the gates below.
|
||||
|
||||
## Run Samples Checklist
|
||||
|
||||
Before opening a PR that touches public API or examples, verify all samples still work:
|
||||
@ -43,20 +52,27 @@ tidalDB opened, verified, and closed. M0 complete.
|
||||
|
||||
## Full Quality Gate
|
||||
|
||||
The pre-commit hook enforces these automatically on staged Rust files:
|
||||
The pre-commit hook (`hooks/pre-commit`, activated by `scripts/install-hooks.sh`)
|
||||
enforces these on staged Rust files:
|
||||
|
||||
```bash
|
||||
cargo fmt --manifest-path tidal/Cargo.toml -- --check
|
||||
cargo clippy --manifest-path tidal/Cargo.toml -- -D warnings
|
||||
cargo test --manifest-path tidal/Cargo.toml --lib
|
||||
cargo fmt
|
||||
cargo clippy -p tidaldb -p tidal-net -p tidal-server -p tidal-stress -p tidalctl --all-targets
|
||||
cargo test -p tidaldb --lib
|
||||
```
|
||||
|
||||
Run the complete gate manually:
|
||||
Clippy runs **without** `-D warnings` on purpose. Each crate's `[lints]` table is
|
||||
the single source of truth: `clippy::all` and `unwrap_used` are `deny` (so they
|
||||
fail the build), while `pedantic` and `nursery` are `warn` (advisory). Passing
|
||||
`-D warnings` on the command line overrides that and turns ~58 deliberate
|
||||
pedantic warnings in the integration tests into hard errors.
|
||||
|
||||
Run the complete gate manually, across the whole workspace:
|
||||
|
||||
```bash
|
||||
cargo fmt --manifest-path tidal/Cargo.toml
|
||||
cargo clippy --manifest-path tidal/Cargo.toml -- -D warnings
|
||||
cargo test --manifest-path tidal/Cargo.toml
|
||||
cargo fmt --check
|
||||
cargo clippy --workspace --all-targets
|
||||
cargo test --workspace
|
||||
cargo bench --manifest-path tidal/Cargo.toml --no-run # ensure benches compile
|
||||
```
|
||||
|
||||
@ -75,4 +91,4 @@ Key rules:
|
||||
- `Result<T, TidalError>` everywhere — no panics on recoverable failures
|
||||
- `#![forbid(unsafe_code)]` — relaxed only at explicit FFI boundaries with `// SAFETY:` comment
|
||||
- Property tests for invariants, criterion benchmarks for performance claims
|
||||
- `cargo clippy -D warnings` must pass with zero warnings
|
||||
- Deny-level clippy (`clippy::all`, `unwrap_used`) must pass; pedantic/nursery are advisory
|
||||
|
||||
1
Cargo.lock
generated
1
Cargo.lock
generated
@ -4282,6 +4282,7 @@ dependencies = [
|
||||
"serde",
|
||||
"serde_json",
|
||||
"thiserror 2.0.18",
|
||||
"time",
|
||||
"tokio",
|
||||
"tracing",
|
||||
"tracing-subscriber",
|
||||
|
||||
@ -14,5 +14,8 @@ resolver = "2"
|
||||
|
||||
[workspace.package]
|
||||
edition = "2024"
|
||||
# Published MSRV for the engine and its siblings; they compile on 1.91.0. The
|
||||
# dev toolchain is pinned separately in rust-toolchain.toml because `tidalctl`
|
||||
# needs a patch release above this floor - see that file and tidalctl/Cargo.toml.
|
||||
rust-version = "1.91"
|
||||
license = "MIT"
|
||||
|
||||
@ -32,6 +32,7 @@ const BRIDGE_MIN_CATEGORY_SCORE: f32 = 0.05;
|
||||
const EMBEDDER_TIMEOUT_SECS: u64 = 10;
|
||||
|
||||
use tidaldb::TidalDb;
|
||||
use tidaldb::ranking::reason::{ReasonCode, ReasonLabel};
|
||||
use tidaldb::schema::{EntityId, Timestamp};
|
||||
|
||||
pub use labels::{ForageItem, ItemLabel};
|
||||
@ -797,11 +798,23 @@ impl ForageEngine {
|
||||
1.0
|
||||
};
|
||||
let score = boost * semantic_sim + (1.0 - boost) * rank_score;
|
||||
// The pool is app-built, so the engine cannot tag it:
|
||||
// record why each item is here. `weight` is the share of
|
||||
// the blended score that semantic similarity supplied.
|
||||
let semantic_share = if score > 0.0 {
|
||||
(boost * semantic_sim / score).clamp(0.0, 1.0)
|
||||
} else {
|
||||
f64::from(self.mab.semantic_boost)
|
||||
};
|
||||
RetrieveResult {
|
||||
entity_id: r.entity_id,
|
||||
score,
|
||||
rank: r.rank,
|
||||
signals: vec![],
|
||||
reasons: vec![ReasonLabel::new(
|
||||
ReasonCode::PreferenceMatch,
|
||||
semantic_share,
|
||||
)],
|
||||
}
|
||||
})
|
||||
.collect();
|
||||
@ -874,6 +887,12 @@ impl ForageEngine {
|
||||
score,
|
||||
rank: r.rank,
|
||||
signals: vec![],
|
||||
reasons: vec![ReasonLabel::with_context(
|
||||
ReasonCode::SemanticMatch,
|
||||
score,
|
||||
"similar_to",
|
||||
&saved_id.to_string(),
|
||||
)],
|
||||
});
|
||||
}
|
||||
}
|
||||
@ -902,6 +921,8 @@ impl ForageEngine {
|
||||
score: 0.5,
|
||||
rank: 0,
|
||||
signals: vec![],
|
||||
// Pinned by the discovery mechanism, not by ranking.
|
||||
reasons: vec![ReasonLabel::new(ReasonCode::ExplorationBudget, 1.0)],
|
||||
});
|
||||
}
|
||||
}
|
||||
@ -1137,14 +1158,49 @@ impl ForageEngine {
|
||||
"health",
|
||||
];
|
||||
|
||||
let pref = self.db.preference_vectors().get(user_id)?;
|
||||
// Since m12 the engine may model one user as several preference clusters,
|
||||
// and `preference_vectors().get()` returns only the strongest cluster's
|
||||
// centroid. Through that accessor a tech+jazz user whose interests split
|
||||
// into two clusters looks single-interest and never bridges. Ask for the
|
||||
// top-2 cluster vectors instead; the accessor's uniform fallback returns
|
||||
// one vector for cold-start users.
|
||||
let vectors =
|
||||
self.db
|
||||
.preference_vectors()
|
||||
.query_vectors(user_id, Timestamp::now().as_nanos(), 2);
|
||||
let pref = vectors.first()?.clone();
|
||||
|
||||
// Build the midpoint vector and extract category names.
|
||||
// Two cases: category-axis 8-dim embeddings (P0–P3) vs real embedder.
|
||||
let (cat_a, cat_b, midpoint) = if pref.len() == CATS.len() {
|
||||
// 8-dim: find the top-2 dimensions by preference score.
|
||||
let dominant_dim = |v: &[f32]| -> Option<(usize, f32)> {
|
||||
v.iter()
|
||||
.copied()
|
||||
.enumerate()
|
||||
.max_by(|a, b| a.1.partial_cmp(&b.1).unwrap_or(std::cmp::Ordering::Equal))
|
||||
};
|
||||
|
||||
// Two distinct clusters are two interests the user actually holds -
|
||||
// the honest pair to bridge. Fall back to the two strongest
|
||||
// dimensions inside one cluster when there is only one.
|
||||
let across_clusters = match vectors.get(1) {
|
||||
Some(second) => {
|
||||
let (dim_a, score_a) = dominant_dim(&pref)?;
|
||||
let (dim_b, score_b) = dominant_dim(second)?;
|
||||
(dim_a != dim_b
|
||||
&& score_a >= BRIDGE_MIN_CATEGORY_SCORE
|
||||
&& score_b >= BRIDGE_MIN_CATEGORY_SCORE)
|
||||
.then_some((dim_a, dim_b))
|
||||
}
|
||||
None => None,
|
||||
};
|
||||
|
||||
let (dim_a, dim_b) = match across_clusters {
|
||||
Some(pair) => pair,
|
||||
None => {
|
||||
let mut indexed: Vec<(usize, f32)> = pref.iter().copied().enumerate().collect();
|
||||
indexed.sort_by(|a, b| b.1.partial_cmp(&a.1).unwrap_or(std::cmp::Ordering::Equal));
|
||||
indexed
|
||||
.sort_by(|a, b| b.1.partial_cmp(&a.1).unwrap_or(std::cmp::Ordering::Equal));
|
||||
let (dim_a, score_a) = indexed[0];
|
||||
let (dim_b, score_b) = indexed[1];
|
||||
|
||||
@ -1152,6 +1208,9 @@ impl ForageEngine {
|
||||
if score_a < BRIDGE_MIN_CATEGORY_SCORE || score_b < BRIDGE_MIN_CATEGORY_SCORE {
|
||||
return None;
|
||||
}
|
||||
(dim_a, dim_b)
|
||||
}
|
||||
};
|
||||
|
||||
let cat_a = CATS.get(dim_a).copied()?.to_string();
|
||||
let cat_b = CATS.get(dim_b).copied()?.to_string();
|
||||
|
||||
@ -14,22 +14,27 @@ pub struct SeedItem {
|
||||
pub url: String,
|
||||
}
|
||||
|
||||
/// Map a URL to a stable item ID using FNV-1a.
|
||||
/// Map a URL to a stable item ID using FNV-1a, folded into the engine's item
|
||||
/// universe.
|
||||
///
|
||||
/// Produces a deterministic, collision-resistant u64 that is always > 100_000,
|
||||
/// keeping it well clear of the seed corpus range (1–100).
|
||||
/// Produces a deterministic ID that is always `> 100_000` (clear of the seed
|
||||
/// corpus range 1-100) and always `<= u32::MAX`. The upper bound is a hard
|
||||
/// engine contract, not a style choice: durable per-user state (seen / saved /
|
||||
/// liked / hard-negative rows) narrows an item ID to a `u32` slot, so
|
||||
/// `write_item_with_metadata` and every `for_user` signal reject an ID above
|
||||
/// `u32::MAX` rather than alias two items forever. The fold xors both halves of
|
||||
/// the hash so entropy from the high word is not discarded.
|
||||
pub fn url_to_item_id(url: &str) -> u64 {
|
||||
/// Highest ID reserved for the seed corpus and its headroom.
|
||||
const SEED_CEILING: u32 = 100_000;
|
||||
|
||||
let mut hash: u64 = 14_695_981_039_346_656_037;
|
||||
for byte in url.bytes() {
|
||||
hash ^= u64::from(byte);
|
||||
hash = hash.wrapping_mul(1_099_511_628_211);
|
||||
}
|
||||
// Map tiny hashes out of the seed range without clustering.
|
||||
if hash <= 100_000 {
|
||||
hash.wrapping_add(100_001)
|
||||
} else {
|
||||
hash
|
||||
}
|
||||
let folded = ((hash >> 32) ^ (hash & 0xffff_ffff)) as u32;
|
||||
u64::from(SEED_CEILING + 1 + folded % (u32::MAX - SEED_CEILING))
|
||||
}
|
||||
|
||||
const CATEGORIES: &[&str] = &[
|
||||
|
||||
@ -752,6 +752,9 @@ fn build_schema() -> std::result::Result<Schema, TidalError> {
|
||||
denied_signals: vec![],
|
||||
max_session_duration: Duration::from_secs(60 * 60),
|
||||
max_signals_per_session: 1_000,
|
||||
// Read-path and profile-override controls (m10) stay at their
|
||||
// defaults: this app writes signals and never overrides profiles.
|
||||
..AgentPolicy::default()
|
||||
},
|
||||
);
|
||||
|
||||
|
||||
@ -21,6 +21,8 @@ RUN cargo build -p tidal-stress --release --locked
|
||||
|
||||
FROM debian:trixie-slim
|
||||
ARG DEBIAN_FRONTEND=noninteractive
|
||||
ARG GIT_HASH=unknown
|
||||
LABEL org.opencontainers.image.revision="${GIT_HASH}"
|
||||
WORKDIR /srv
|
||||
|
||||
# ca-certificates only matters if a target is https:// (in-cluster targets are
|
||||
@ -34,6 +36,9 @@ COPY --from=builder /app/target/release/tidal-stress /usr/local/bin/tidal-stress
|
||||
# tidal-stress/src/bin/soak-eval.rs). The soak monitor's evaluator container runs
|
||||
# it to join ledger.tsv + restarts.tsv into the honest streak verdict.
|
||||
COPY --from=builder /app/target/release/soak-eval /usr/local/bin/soak-eval
|
||||
# soak-watch: direct in-cluster Kubernetes restart evidence recorder. Keeping
|
||||
# it in this image removes the mutable shell+kubectl sidecar dependency.
|
||||
COPY --from=builder /app/target/release/soak-watch /usr/local/bin/soak-watch
|
||||
|
||||
USER stress
|
||||
|
||||
|
||||
@ -10,17 +10,17 @@ All estimates assume a single-node deployment with default configuration (30-sec
|
||||
|
||||
tidalDB is an in-memory-first database. USearch HNSW indexes, the signal ledger hot tier, and Tantivy reader segments all reside in RAM during operation. There is no swap tolerance for USearch -- if the process is swapped, ANN query latency degrades from microseconds to seconds.
|
||||
|
||||
| Items | Embedding Dims | USearch RAM | Signal Ledger RAM (10 signals) | Tantivy RAM | Total Estimate |
|
||||
|------:|---------------:|------------:|-------------------------------:|------------:|---------------:|
|
||||
| 100K | 128D | ~26 MB | ~110 MB | ~50 MB | ~200 MB |
|
||||
| 100K | 768D | ~154 MB | ~110 MB | ~50 MB | ~320 MB |
|
||||
| 100K | 1536D | ~307 MB | ~110 MB | ~50 MB | ~470 MB |
|
||||
| 1M | 128D | ~256 MB | ~1.1 GB | ~200 MB | ~1.6 GB |
|
||||
| 1M | 768D | ~1.5 GB | ~1.1 GB | ~200 MB | ~2.8 GB |
|
||||
| 1M | 1536D | ~3.1 GB | ~1.1 GB | ~200 MB | ~4.4 GB |
|
||||
| 10M | 128D | ~2.6 GB | ~11 GB | ~500 MB | ~14 GB |
|
||||
| 10M | 768D | ~15 GB | ~11 GB | ~500 MB | ~27 GB |
|
||||
| 10M | 1536D | ~31 GB | ~11 GB | ~500 MB | ~43 GB |
|
||||
| Items | Embedding Dims | USearch RAM | Signal Ledger RAM (10 active signals) | Tantivy RAM | Analytic Total |
|
||||
|------:|---------------:|------------:|--------------------------------------:|------------:|---------------:|
|
||||
| 100K | 128D | ~31 MB | ~1.09 GB | ~50 MB | ~1.17 GB |
|
||||
| 100K | 768D | ~184 MB | ~1.09 GB | ~50 MB | ~1.32 GB |
|
||||
| 100K | 1536D | ~369 MB | ~1.09 GB | ~50 MB | ~1.51 GB |
|
||||
| 1M | 128D | ~307 MB | ~5.44 GB (hot-tier cap) | ~200 MB | ~5.95 GB |
|
||||
| 1M | 768D | ~1.84 GB | ~5.44 GB (hot-tier cap) | ~200 MB | ~7.48 GB |
|
||||
| 1M | 1536D | ~3.69 GB | ~5.44 GB (hot-tier cap) | ~200 MB | ~9.33 GB |
|
||||
| 10M | 128D | ~3.07 GB | ~5.44 GB (hot-tier cap) | ~500 MB | ~9.01 GB |
|
||||
| 10M | 768D | ~18.43 GB | ~5.44 GB (hot-tier cap) | ~500 MB | ~24.37 GB |
|
||||
| 10M | 1536D | ~36.86 GB | ~5.44 GB (hot-tier cap) | ~500 MB | ~42.80 GB |
|
||||
|
||||
### Formulas
|
||||
|
||||
@ -35,12 +35,22 @@ The 20% graph overhead accounts for HNSW neighbor lists (M=16 default, two layer
|
||||
**Signal ledger hot tier:**
|
||||
|
||||
```
|
||||
items * signal_count * ~1,088 bytes/entry
|
||||
min(items * active_signal_types_per_item, DEFAULT_MAX_SIGNAL_ENTRIES)
|
||||
* ~1,088 bytes/entry
|
||||
```
|
||||
|
||||
Each `(entity_id, signal_type_id)` entry in the DashMap holds the running decay score, windowed counters (BucketedCounter with minute and hour buckets), velocity state, and the DashMap per-shard overhead. The 1,088 bytes/entry figure was measured in the m7p3 scale benchmarks.
|
||||
Each `(entity_id, signal_type_id)` entry in the DashMap holds the running decay
|
||||
score, windowed counters (BucketedCounter with minute and hour buckets),
|
||||
velocity state, and the DashMap per-shard overhead. The 1,088 bytes/entry figure
|
||||
was measured in the m7p3 scale benchmarks. The table uses ten active signal
|
||||
types per item and the default 5M-entry hot-tier ceiling. Without trimming, the
|
||||
signal column would be ~10.88 GB at 1M items and ~108.8 GB at 10M items.
|
||||
|
||||
The signal ledger has a memory budget of 5M entries (`DEFAULT_MAX_SIGNAL_ENTRIES`). When exceeded, the checkpoint thread evicts cold entries (oldest `last_update` timestamp). If your workload has more than 5M active `(entity, signal_type)` pairs, cold entries will be served from fjall checkpoints (slower, but correct).
|
||||
The signal ledger has a memory budget of 5M entries
|
||||
(`DEFAULT_MAX_SIGNAL_ENTRIES`, ~5.44 GB). When exceeded, the checkpoint thread
|
||||
evicts cold entries (oldest `last_update` timestamp). If your workload has more
|
||||
than 5M active `(entity, signal_type)` pairs, cold entries will be served from
|
||||
fjall checkpoints (slower, but correct).
|
||||
|
||||
**Tantivy text index:**
|
||||
|
||||
@ -49,8 +59,14 @@ Tantivy's RAM usage depends on the number of indexed documents, average document
|
||||
### Notes
|
||||
|
||||
- Signal ledger RAM is for the in-memory hot tier only. The WAL and fjall checkpoints add disk usage, not RAM.
|
||||
- The "10 signals" column assumes 10 distinct signal types per entity. Scale linearly for more signal types.
|
||||
- USearch RAM is the dominant cost at high dimensionality. If you use 1536D embeddings (e.g., OpenAI text-embedding-3-large), plan for USearch to consume 70%+ of total RAM at 10M items.
|
||||
- The table assumes ten active signal types per item. Below the 5M-entry hot-tier
|
||||
cap it scales linearly; beyond the cap, additional cold entries trade RAM for
|
||||
checkpoint-backed lookup cost.
|
||||
- USearch RAM is the dominant index cost at high dimensionality. If you use 1536D embeddings (e.g., OpenAI text-embedding-3-large), budget from the complete analytic total and measured process envelope, not the vector column alone.
|
||||
- The analytic table excludes process/runtime overhead, replication and catch-up
|
||||
buffers, allocator fragmentation, and concurrent query/write working memory.
|
||||
It is a sizing input, not a container limit. Use the measured envelope below
|
||||
for the multi-process cluster.
|
||||
|
||||
---
|
||||
|
||||
@ -175,8 +191,8 @@ The tables above are single-node, analytic estimates. This section is the **meas
|
||||
|
||||
| metric | value | note |
|
||||
|--------|-------|------|
|
||||
| write knee | **~250 rps** | the peach mix is WRITE-heavy (view/like/skip ≈ 90% of ops) |
|
||||
| 30-night soak | **200 rps** | sustained with margin |
|
||||
| write knee | **~250 rps in the clean sweep** | one benchmark, not a sustained-safe rate |
|
||||
| nightly 200 rps soak | **23 PASS / 32 FAIL (42% green)** | not sustained with margin; gate parked |
|
||||
| write scaling vs node count | **~1.0×** | does NOT scale at full-placement RF3 |
|
||||
|
||||
- Write tput does **NOT** scale with node count at full-placement RF3: every per-shard quorum spans all 3 nodes, so **every follower applies every 1536-D write**. Adding nodes adds replication work, not write capacity (~1.0×, not 2.5×).
|
||||
@ -184,13 +200,20 @@ The tables above are single-node, analytic estimates. This section is the **meas
|
||||
|
||||
### Per-pod memory at 1536-D
|
||||
|
||||
| corpus | per-pod RSS | pod mem limit | verdict |
|
||||
|--------|-------------|---------------|---------|
|
||||
| 100k × 1536-D | **~1.9 GiB** (HNSW load peak) | 4 GiB | fits |
|
||||
| 1M × 1536-D | **~7–8 GB** | 4 GiB | **OOMs the 16 GiB nodes** |
|
||||
| corpus / workload | observed per-pod working set | configured limit at observation | verdict |
|
||||
|-------------------|------------------------------|---------------------------------|---------|
|
||||
| 100k idle baseline | **~3.58 GiB** | 4 GiB | only ~420 MiB load/recovery headroom |
|
||||
| 100k / 200 rps soak | **3.97-4.00 GiB before OOMKill** | 4 GiB | limit is insufficient |
|
||||
| 1M × 1536-D estimate | **~7-8 GiB before runtime overhead** | — | requires Ref-B measurement |
|
||||
|
||||
- Full placement → **each pod holds the WHOLE corpus** (no per-shard sharding of memory).
|
||||
- 1M × 1536-D ≈ 7–8 GB/pod RSS overruns the ~13 GiB allocatable (already ~7 GiB of co-tenants) → the **1M production-read gate needs nodes >16 GiB** (a Ref-B requirement).
|
||||
- Full placement means each pod holds the whole corpus.
|
||||
- Four independent historical windows show `OOMKilled` at 3.97-4.00 GiB, followed
|
||||
by snapshot-required reseed self-restarts. The exact internal growth source is
|
||||
not yet isolated; signal hot-entry count stayed flat, so do not label it a leak
|
||||
without allocation/heap evidence.
|
||||
- The corrected canary envelope is a 4 GiB request and 6 GiB limit (measured peak
|
||||
plus 50% recovery/profiling headroom). Profile the exact 100k/1536-D/200-rps
|
||||
trajectory before treating 6 GiB as a production ceiling.
|
||||
|
||||
### Startup / boot
|
||||
|
||||
@ -198,12 +221,12 @@ The tables above are single-node, analytic estimates. This section is the **meas
|
||||
- `startupProbe` budget is **~20 min** (`failureThreshold` 240 × 5 s).
|
||||
- Graceful shutdown **saves the graphs** (grace **600 s**), so a clean restart **skips the rebuild** (load, not rebuild). A SIGKILL/crash skips the save → next boot rebuilds.
|
||||
|
||||
### Pod resources (live)
|
||||
### Pod resources (corrected canary manifest)
|
||||
|
||||
| resource | request | limit |
|
||||
|----------|---------|-------|
|
||||
| CPU | 500m | 3 |
|
||||
| memory | 1 GiB | 4 GiB |
|
||||
| CPU | 2 | 3 |
|
||||
| memory | 4 GiB | 6 GiB |
|
||||
| PVC | — | 5 GiB/pod, `local-path` |
|
||||
|
||||
- PVC is **local NVMe** (`local-path`) — longhorn's fsync overhead was unacceptable for the WAL path.
|
||||
|
||||
@ -1,39 +1,19 @@
|
||||
# =============================================================================
|
||||
# DESIGN-REFERENCE RULE SET — NOT LOADED BY ANY ALERTMANAGER TODAY.
|
||||
# TIDALDB ALERT CONTRACT
|
||||
# =============================================================================
|
||||
# This file is reference/design config. It is NOT wired into any running
|
||||
# alerting pipeline:
|
||||
# - vmalert (pilot/dev) loads ONLY ops/vmalert/rules/*.yaml
|
||||
# - prod CRDs live ONLY in infra/k8s/prom/prometheus-rules/*.yaml
|
||||
# Nothing mounts or imports docs/ops/prometheus-alerts.yaml.
|
||||
# Do not read these as live pager rules.
|
||||
# The orchard9 fleet runs these expressions through vmalert from:
|
||||
# ../k3s-fleet/deployments/k8s/base/observability/alerting-rules.yaml
|
||||
#
|
||||
# Provenance: every metric referenced below (tidaldb_health_ok,
|
||||
# tidaldb_checkpoint_age_seconds, tidaldb_checkpoint_failures_total,
|
||||
# tidaldb_wal_lag_bytes, tidaldb_signal_hot_entries, tidaldb_degradation_level,
|
||||
# tidaldb_active_sessions, tidaldb_rate_limited_total,
|
||||
# tidaldb_tantivy_segment_count, tidaldb_retrieve_latency_us_bucket,
|
||||
# tidaldb_search_latency_us_bucket) IS really emitted today by
|
||||
# tidal/src/db/metrics/mod.rs. So these rules are promotable — they
|
||||
# are not orphaned against phantom metrics.
|
||||
# This product-side copy documents the emitted metrics and alert thresholds.
|
||||
# Keep every expression and threshold aligned with the fleet source when the
|
||||
# contract changes. The fleet source adds routing labels (`capability`,
|
||||
# `surface`) and runbook URLs required by the shared Alertmanager pipeline.
|
||||
#
|
||||
# Promotion path (the right long-term fix — do it deliberately, do NOT
|
||||
# hand-copy these as live pager rules without the steps below):
|
||||
# 1. Add a vmalert rule file: ops/vmalert/rules/tidaldb.yaml, translating
|
||||
# each rule and stamping the canonical labels every routed rule carries —
|
||||
# severity + capability + surface (+ optional component). Map per the
|
||||
# canonical severity taxonomy in ops/alerting.md:
|
||||
# labels {severity: critical} = pager class (TidalDBDown only here),
|
||||
# {severity: warning} = non-paging / business-hours,
|
||||
# {severity: info} = pure-FYI.
|
||||
# Add capability: tidaldb and a surface label per rule, and a
|
||||
# runbook_url annotation:
|
||||
# https://github.com/orchard9/tidaldb/blob/main/docs/runbooks/<slug>.md
|
||||
# 2. Add the prod CRD twin: infra/k8s/prom/prometheus-rules/tidaldb.yaml,
|
||||
# kept byte-aligned in expr/threshold with the vmalert file.
|
||||
# 3. Verify the alerts fire against real metrics (scrape tidaldb, force a
|
||||
# degraded state) before declaring them on-call-ready.
|
||||
# Until steps 1–3 land, this file stays a design reference only.
|
||||
# Every metric referenced below is emitted by tidalDB today. Standalone rules
|
||||
# stay inactive when no standalone instance is running; cluster-only series
|
||||
# activate only in cluster mode. The shared fleet rules additionally cover
|
||||
# target loss, unavailable replicas, image-pull failures, restarts, PVCs, and
|
||||
# Kubernetes Jobs.
|
||||
# =============================================================================
|
||||
groups:
|
||||
- name: tidaldb
|
||||
|
||||
@ -156,7 +156,7 @@ The roadmap now has two tracks:
|
||||
|
||||
**Engine status:** M0–M12 **COMPLETE**. M9/M10 shipped 2026-06-06; M11 (Enterprise-Grade Cluster, all nine phases) closed 2026-06-13; M12 (Vector Retrieval at production shape) closed 2026-06-14 / 2026-06-23 — see the M11/M12 milestone rows above and *Implementation Status* below.
|
||||
|
||||
**Next (engine):** close the v1.0 bar — the 30-day-green nightly chaos+soak calendar (m11p9) and the standing Ref-A/k3s throughput re-runs (≥5,000/s + ≥2.5× single-shard scaling). Deferred follow-ups: the `writer_agent` u16 interning on the WAL v3 envelope and the offline medoid-recluster tier for multi-vector preference.
|
||||
**Next (engine):** isolate the Ref-A 100k/1536-D memory growth under the corrected 4 GiB request / 6 GiB canary limit. The 30-day nightly gate remains parked: its 200 rps history is 23 PASS / 32 FAIL with repeated 4 GiB OOMKills, so it is not a production-readiness signal yet. Resume only after the controlled profile passes and restart evidence is fail-closed. The old ≥2.5× write-scaling goal belongs to Ref-B (≥5 nodes **with partitioned placement**); adding nodes to full-placement RF3 does not scale writes. Deferred follow-ups: the `writer_agent` u16 interning on the WAL v3 envelope and the offline medoid-recluster tier for multi-vector preference.
|
||||
**Next (product):** iknowyou M5 acceptance pass, then M6 Closed Loop (session lifecycle + preference drift validation).
|
||||
|
||||
---
|
||||
|
||||
@ -140,22 +140,24 @@ The bar is NOT "30 green soak Job exits". It is **30 consecutive nights where th
|
||||
soak passed its SLO gates AND no cluster pod restarted under load**. The two
|
||||
halves come from two streams on the shared result PVC:
|
||||
|
||||
- `ledger.tsv` — one row per night from the nightly CronJob: `PASS`/`FAIL` on the
|
||||
armed gates (`--fail-on-knee`, `--max-p99-ms 150`, `--max-error-pct 1`).
|
||||
- `restarts.tsv` — the monitor's 5-minute snapshot of each `tidaldb-{0,1,2}` pod's
|
||||
cumulative `restartCount`.
|
||||
- `ledger.tsv` — one row per UTC date from the nightly CronJob: `PASS`/`FAIL`
|
||||
on the armed gates (`--fail-on-knee`, `--max-p99-ms 150`,
|
||||
`--max-error-pct 1`), plus exact `start_utc` / `end_utc` load bounds.
|
||||
- `restarts.tsv` — the monitor's one-minute snapshot of each
|
||||
`tidaldb-{0,1,2}` pod UID and cumulative `restartCount`.
|
||||
|
||||
`soak-eval` (the proven core in `tidal_stress::soak_eval`, table-driven-tested;
|
||||
the binary runs in the monitor's evaluator container and once per night in the
|
||||
CronJob) joins them into `streak.tsv`:
|
||||
`soak-eval` (the proven core in `tidal_stress::soak_eval`,
|
||||
table-driven-tested; the binary runs in the monitor's evaluator container and
|
||||
once per night in the CronJob) joins them into `streak.tsv`:
|
||||
|
||||
- a night is **GREEN** iff its ledger row is `PASS` **and** no pod's cumulative
|
||||
restart count rose within that night's window;
|
||||
- the **streak** is the trailing run of consecutive green nights;
|
||||
- **any** non-green night resets the streak to **0**. A night resets it when
|
||||
EITHER the soak breached an SLO gate (Job `FAIL`) **OR** a pod restarted under
|
||||
load (the zero-unrecovered-restart half — a green ledger row with a restart
|
||||
in-window is a FAIL, not a PASS: that was the false-confidence hole).
|
||||
- a night is **GREEN** iff its ledger row is `PASS`, every expected pod has a
|
||||
fresh sample immediately before and after the measured load window, and
|
||||
neither its pod UID nor restart count changes between those samples;
|
||||
- the **streak** is the trailing run of unique, consecutive UTC calendar dates;
|
||||
- a gate breach, restart, pod replacement, missing/stale boundary sample,
|
||||
duplicate date, or date gap cannot advance the streak. This closes both
|
||||
fail-open paths: a green load exit with a recovery event and multiple rows
|
||||
masquerading as multiple nights.
|
||||
|
||||
The evaluator emits exactly one alert per non-green transition (a durable
|
||||
`ALERT-<date>.txt` the http surface serves, plus an optional `NOTIFY_WEBHOOK`
|
||||
|
||||
@ -1512,29 +1512,35 @@ writes with `ENOSPC` after `n` cumulative bytes this process lifetime (arm it on
|
||||
a node at restart to fail after `n` bytes of post-restart writes). NEVER set
|
||||
these on a production node.
|
||||
|
||||
**Soak with regression gates** (`tidal-stress`):
|
||||
**Fleet soak with regression gates** (`tidal-stress`):
|
||||
|
||||
```
|
||||
tidal-stress --target https://<gateway> --ramp "200:3600" --mix peach \
|
||||
--json-summary soak.json --max-error-pct 1 --max-p99-ms 250 --fail-on-knee
|
||||
tidal-stress --target https://<cluster-node>:9500 \
|
||||
--ca-cert /etc/tidaldb/tls/ca.crt --skip-seed \
|
||||
--corpus 100000 --embedding-dim 1536 \
|
||||
--ramp "200:3600" --mix peach --json-summary soak.json \
|
||||
--max-error-pct 1 --max-p99-ms 150 --fail-on-knee
|
||||
```
|
||||
|
||||
> **Soak runs at 200 rps (re-scoped 2026-06-19).** The peach mix is write-heavy
|
||||
> and the write knee on this 3-node fleet is **~250 rps** (1536-D `ack=quorum`
|
||||
> ingest), so the soak ramp is a single **measured-sustainable 200-rps** stage
|
||||
> (`--ramp "200:3600"`) rather than the retired `3900:3600` constant (a pre-m12
|
||||
> signal-write figure that does not hold at the 1536-D production shape). Point
|
||||
> `--target` at an `https://` gateway — the `:9500` plane serves TLS.
|
||||
> **The Ref-A gate is 200 rps (re-scoped 2026-06-19).** The peach mix is
|
||||
> write-heavy and the measured write knee on this full-placement three-node
|
||||
> fleet is about 250 rps. The one-hour gate therefore uses the measured 200-rps
|
||||
> stage, not the retired pre-m12 `3900:3600` figure. The `:9500` plane is TLS.
|
||||
|
||||
`--fail-on-knee` (built-in SLO), `--max-p99-ms`, and `--max-error-pct` make the
|
||||
run exit non-zero on a regression; `--json-summary` writes a machine-readable
|
||||
per-stage roll-up for trend lines. A bounded version runs nightly; the GA-bar
|
||||
100k-DAU soak points `--target` at the live Ref-A cluster at 200 rps.
|
||||
`--fail-on-knee`, `--max-p99-ms`, and `--max-error-pct` make the load process
|
||||
exit non-zero on regression. The in-cluster `tidal-soak-nightly` CronJob adds
|
||||
the recovery half of the gate: one ledger row per UTC date, one-minute pod UID
|
||||
and restart-count evidence immediately before and after the exact load window,
|
||||
and an atomic `streak.tsv`. A date advances the streak only when the load passes,
|
||||
all three pods have fresh boundary samples, and no UID or counter changes.
|
||||
Duplicate dates, gaps, stale evidence, pod replacement, and load failure cannot
|
||||
inflate the 30-night streak.
|
||||
|
||||
**Nightly CI** (`.woodpecker.yaml`, cron `nightly` — Woodpecker, never GitHub
|
||||
Actions): the chaos suites with elevated kill-points (`TIDAL_QUORUM_KILLPOINTS`,
|
||||
`TIDAL_ELECTION_KILLPOINTS`) then the gated soak. A nightly failure flags the
|
||||
day's correctness or performance regression. The guarantee→test map is
|
||||
Actions) remains the bounded code-regression signal: tier-3 chaos with elevated
|
||||
kill-points, then a local gated soak. It is not the Ref-A calendar gate unless
|
||||
`TIDAL_SOAK_TARGET` and its duration/rate are explicitly pointed at that fleet.
|
||||
The guarantee→test map is
|
||||
[docs/planning/milestone-11/guarantee-traceability.md](../planning/milestone-11/guarantee-traceability.md).
|
||||
|
||||
## Cross-references
|
||||
|
||||
@ -4,8 +4,22 @@
|
||||
# git config core.hooksPath hooks
|
||||
# or run scripts/install-hooks.sh once.
|
||||
#
|
||||
# Gates: Rust fmt/clippy/test (only when Rust is staged), the documentation
|
||||
# consolidation guard (always), and site eslint (when node_modules exist).
|
||||
# Gates: Rust fmt/clippy/test (only when Rust is staged), the CODING_GUIDELINES
|
||||
# source checks, the documentation consolidation guard (always), and site eslint
|
||||
# (when node_modules exist).
|
||||
#
|
||||
# Two rules this file learned the hard way, 2026-08-16:
|
||||
#
|
||||
# * Clippy runs WITHOUT `-D warnings`. Every tidal crate declares its own
|
||||
# posture in `[lints]` (`clippy::all = deny`, `unwrap_used = deny`,
|
||||
# `pedantic = warn`), and that table is the single source of truth. Adding
|
||||
# `-D warnings` here promoted ~58 deliberate pedantic warnings in the
|
||||
# integration tests to errors and made every Rust commit impossible.
|
||||
# * The Rust gates key off ALL crates, not a `tidal/` path match. An earlier
|
||||
# untracked copy of this hook tested `grep -q 'tidal/'`, which silently
|
||||
# skipped `tidal-server/`, `tidal-net/`, `tidal-stress/`, `tidalctl/`, and
|
||||
# `applications/` - the workspace was left to rot until a workspace-wide
|
||||
# `cargo test` was attempted months later.
|
||||
|
||||
set -uo pipefail
|
||||
ROOT="$(git rev-parse --show-toplevel)"
|
||||
@ -15,16 +29,64 @@ staged() { git diff --cached --name-only --diff-filter=ACM; }
|
||||
rust_staged=$(staged | grep -E '\.rs$' || true)
|
||||
site_staged=$(staged | grep -E '^site/.*\.(ts|tsx|js|jsx|mjs)$' || true)
|
||||
|
||||
# --- Rust (engine crate) ----------------------------------------------------
|
||||
# --- Rust -------------------------------------------------------------------
|
||||
if [ -n "$rust_staged" ]; then
|
||||
echo "pre-commit: cargo fmt"
|
||||
cargo fmt -p tidaldb || { echo "cargo fmt failed" >&2; exit 1; }
|
||||
cargo fmt || { echo "cargo fmt failed" >&2; exit 1; }
|
||||
# re-stage any files fmt rewrote
|
||||
echo "$rust_staged" | while IFS= read -r f; do [ -f "$f" ] && git add "$f"; done
|
||||
echo "pre-commit: cargo clippy -p tidaldb -D warnings"
|
||||
cargo clippy -p tidaldb --all-targets -- -D warnings || exit 1
|
||||
|
||||
echo "pre-commit: cargo clippy -p tidaldb -p tidal-net -p tidal-server -p tidal-stress -p tidalctl"
|
||||
cargo clippy -p tidaldb -p tidal-net -p tidal-server -p tidal-stress -p tidalctl \
|
||||
--all-targets || exit 1
|
||||
|
||||
echo "pre-commit: cargo test -p tidaldb --lib"
|
||||
cargo test -p tidaldb --lib || exit 1
|
||||
|
||||
# ── CODING_GUIDELINES source checks (engine sources only) ─────────────────
|
||||
engine_src=$(echo "$rust_staged" | grep -E '^tidal/src/' || true)
|
||||
if [ -n "$engine_src" ]; then
|
||||
fail=0
|
||||
|
||||
# §9 file length: 600 lines max.
|
||||
while IFS= read -r f; do
|
||||
[ -z "$f" ] && continue
|
||||
lines=$(wc -l < "$f" 2>/dev/null || echo 0)
|
||||
if [ "$lines" -gt 600 ]; then
|
||||
echo " FAIL: $f is $lines lines (max 600) - split it (CODING_GUIDELINES §9)" >&2
|
||||
fail=1
|
||||
fi
|
||||
done <<< "$engine_src"
|
||||
|
||||
# §11 no println!/eprintln! in engine sources - use tracing::.
|
||||
while IFS= read -r f; do
|
||||
[ -z "$f" ] && continue
|
||||
case "$f" in */benches/*) continue ;; esac
|
||||
hits=$(grep -n 'println!\|eprintln!' "$f" 2>/dev/null | grep -vE '^[0-9]+:[[:space:]]*//' || true)
|
||||
if [ -n "$hits" ]; then
|
||||
echo " FAIL: println!/eprintln! in $f - use tracing:: (CODING_GUIDELINES §11)" >&2
|
||||
awk '{print " " $0}' <<< "$hits" >&2
|
||||
fail=1
|
||||
fi
|
||||
done <<< "$engine_src"
|
||||
|
||||
# §10 every unsafe block carries a // SAFETY: comment.
|
||||
while IFS= read -r f; do
|
||||
[ -z "$f" ] && continue
|
||||
hits=$(awk '
|
||||
/\/\/ SAFETY:/ { had_safety=1; next }
|
||||
/unsafe \{/ { if (!had_safety) print NR": "$0; had_safety=0; next }
|
||||
{ had_safety=0 }
|
||||
' "$f" 2>/dev/null || true)
|
||||
if [ -n "$hits" ]; then
|
||||
echo " FAIL: unsafe block without // SAFETY: in $f (CODING_GUIDELINES §10)" >&2
|
||||
awk '{print " " $0}' <<< "$hits" >&2
|
||||
fail=1
|
||||
fi
|
||||
done <<< "$engine_src"
|
||||
|
||||
[ "$fail" -eq 0 ] || exit 1
|
||||
fi
|
||||
fi
|
||||
|
||||
# --- Documentation consolidation guard (always) -----------------------------
|
||||
|
||||
@ -1,10 +1,9 @@
|
||||
# The tidalDB CLUSTER: ONE StatefulSet, every pod a region (m11p5 §4).
|
||||
#
|
||||
# MUTUALLY EXCLUSIVE with the standalone set in k8s/ (namespace `tidaldb`,
|
||||
# replicas: 1, `standalone` subcommand). This is namespace `tidaldb-cluster`,
|
||||
# replicas: 3, the `cluster --region` subcommand, real quorum-ack writes,
|
||||
# automatic election, and elastic membership. Deploy ONE or the OTHER per
|
||||
# namespace — never both.
|
||||
# MUTUALLY EXCLUSIVE with the standalone set in k8s/. Both source workloads are
|
||||
# parked at 0. `scripts/restore-fleet.sh` restores one selected data plane:
|
||||
# cardinality 1 in namespace `tidaldb`, or three `cluster --region` processes in
|
||||
# namespace `tidaldb-cluster` with real quorum-ack writes. Never run both.
|
||||
#
|
||||
# WHY ONE StatefulSet (not one-per-region): the m11p5 bind/advertise split lets
|
||||
# every pod mount the SAME topology ConfigMap (peers are advertised by per-pod
|
||||
@ -22,7 +21,7 @@ metadata:
|
||||
app.kubernetes.io/component: cluster-node
|
||||
spec:
|
||||
serviceName: tidaldb-peers # the headless peer Service — stable per-pod DNS
|
||||
replicas: 3 # the initial voter set; scale up/down per the runbook
|
||||
replicas: 0 # parked; scripts/restore-fleet.sh restores three voters explicitly
|
||||
# Parallel: bring all pods up at once. There is no ordered-bootstrap
|
||||
# dependency — siblings boot in any order (an unreachable-at-startup peer is
|
||||
# normal; the election + catch-up timer converge them). Ordered start would
|
||||
@ -72,7 +71,7 @@ spec:
|
||||
type: RuntimeDefault
|
||||
initContainers:
|
||||
- name: init-datadir
|
||||
image: busybox:1.36
|
||||
image: busybox@sha256:73aaf090f3d85aa34ee199857f03fa3a95c8ede2ffd4cc2cdb5b94e566b11662
|
||||
imagePullPolicy: IfNotPresent
|
||||
command: ["sh", "-c", "mkdir -p /data/db && chown -R 10001:10001 /data/db"]
|
||||
securityContext:
|
||||
@ -84,7 +83,7 @@ spec:
|
||||
mountPath: /data
|
||||
containers:
|
||||
- name: tidaldb
|
||||
image: registry.threesix.ai/tidal/server@sha256:171505745b801dcf231b531de6167dbc309a7182957811cbc2228f0a302572b1 # m12-writeburst-rc7 (LIVE; tidaldb commit 1b5bcba). Adds the write-burst false-partition fix on top of rc6 (bd211e7338): a follower whose 1536-D HNSW apply momentarily starves its transport runtime makes a leader ship RPC miss the 10s deadline -> tonic DeadlineExceeded was counted as a transport failure (record_failure) -> both followers' breakers latched Open -> commit stalled -> ack=quorum 503-stormed with no self-heal. Fix (tidal-net, heuristic set only; commit.rs/election.rs/vote path UNTOUCHED): record_timeout opens the breaker ONLY when there's no recent proof of life (last_contact stale -> genuine blackhole still detected); DeadlineExceeded|Cancelled route to it, genuine Unavailable still opens. Verified: 6 unit + 2 real-gRPC integration + slow-fsync multi-process regression. Inherits rc6's seed-join + reseed-loop + election-divergence fixes and rc13's WAL_RETENTION_SEGMENTS=16. amd64 manifest (6 layers, imagetools-verified).
|
||||
image: registry.threesix.ai/tidal/server:m12-fleet-remediation-20260813@sha256:2e4baaf974ad2cf650609b1843689122bf4fd2751bcb8749906df8089223f4a3
|
||||
imagePullPolicy: IfNotPresent
|
||||
# The image ENTRYPOINT is the bare binary. We override the command with
|
||||
# a tiny /bin/sh wrapper (the bookworm-slim runtime HAS a shell) so we
|
||||
@ -220,9 +219,13 @@ spec:
|
||||
path: /health/live
|
||||
port: http
|
||||
scheme: HTTPS
|
||||
# A transiently saturated query runtime must shed readiness before
|
||||
# kubelet turns load into a cascading restart. Six 10s failures give
|
||||
# the process roughly one minute to recover while still detecting a
|
||||
# genuinely wedged runtime.
|
||||
periodSeconds: 10
|
||||
timeoutSeconds: 3
|
||||
failureThreshold: 3
|
||||
timeoutSeconds: 5
|
||||
failureThreshold: 6
|
||||
readinessProbe:
|
||||
httpGet:
|
||||
path: /health # cluster-aware: 503 joiner/quarantined/draining
|
||||
@ -233,18 +236,22 @@ spec:
|
||||
failureThreshold: 3
|
||||
resources:
|
||||
requests:
|
||||
cpu: "500m"
|
||||
memory: 1Gi
|
||||
# Measured Ref-A soak envelope: the busiest leader sustained
|
||||
# ~1.7 cores and reached ~2.5, so 2 cores is the honest scheduler
|
||||
# reservation. A fleet that cannot place this request cannot run
|
||||
# the 200 rps gate without request-level contention.
|
||||
cpu: "2"
|
||||
# Baseline working set was ~3.6 GiB before load. Reserving 4 GiB
|
||||
# prevents the scheduler from hiding that resident footprint.
|
||||
memory: 4Gi
|
||||
limits:
|
||||
# m12 read-SLA fix: 2->3. The cgroup cpu quota is what the engine's
|
||||
# available_parallelism() reads (SEARCH_GATE / worker_threads sizing).
|
||||
# At "2" a cross-shard search burst starved the async reactor + the
|
||||
# election/heartbeat/apply loops (reads hung to the 30s route timeout;
|
||||
# the starved control plane churned elections -> reseed self-exit).
|
||||
# "3" leaves ~1 core for kubelet/system on the 4-core nodes; requests
|
||||
# stay at 500m so the pod still schedules (server nodes alloc=3).
|
||||
# Three cores preserves one core for kubelet/system on the Ref-A
|
||||
# four-core node while allowing the measured query burst.
|
||||
cpu: "3"
|
||||
memory: 4Gi # m12/1536: 100k×1536 HNSW load peaks ~1.9Gi; 1M needs more headroom (nodes have 13Gi allocatable)
|
||||
# Four independent OOMKills occurred at 3.97-4.00 GiB. Six GiB is
|
||||
# measured peak plus 50% recovery/profiling headroom; the exact
|
||||
# internal growth source still requires heap/allocation profiling.
|
||||
memory: 6Gi
|
||||
securityContext:
|
||||
allowPrivilegeEscalation: false
|
||||
readOnlyRootFilesystem: true # writes only /data (PVC) and /tmp (emptyDir)
|
||||
@ -297,6 +304,8 @@ spec:
|
||||
name: data
|
||||
labels:
|
||||
app.kubernetes.io/name: tidaldb
|
||||
backup.orchard9.ai/class: expendable
|
||||
backup.orchard9.ai/method: tidal-stress-reseed
|
||||
spec:
|
||||
accessModes: ["ReadWriteOnce"]
|
||||
storageClassName: local-path
|
||||
|
||||
@ -3,11 +3,11 @@
|
||||
# WHY A STATEFULSET (not a Deployment): tidalDB is single-node-first and
|
||||
# embeddable — the server wraps ONE engine instance whose state (WAL +
|
||||
# checkpoints + indexes) lives on a durable data dir. It scales VERTICALLY
|
||||
# (bigger node), not by adding replicas. `replicas: 1` is intentional and load-
|
||||
# bearing: there is no shared-storage multi-writer mode. If you need multi-node
|
||||
# HA, that is a DIFFERENT deployment: the multi-process `cluster` subcommand
|
||||
# (one process per region, quorum-acked writes, automatic failover) ships as its
|
||||
# own StatefulSet under k8s/cluster/ — it does not change this manifest.
|
||||
# (bigger node), not by adding active replicas. The parked source state keeps
|
||||
# `replicas: 0`; `scripts/restore-fleet.sh --standalone` restores cardinality 1.
|
||||
# For HA, use the DIFFERENT multi-process `cluster` deployment: one process per
|
||||
# region, quorum-acked writes, and automatic failover. It ships as its own
|
||||
# StatefulSet under k8s/cluster/ and does not change this manifest.
|
||||
# See docs/runbooks/kubernetes.md and docs/runbooks/cluster.md.
|
||||
apiVersion: apps/v1
|
||||
kind: StatefulSet
|
||||
@ -19,7 +19,7 @@ metadata:
|
||||
app.kubernetes.io/component: server
|
||||
spec:
|
||||
serviceName: tidaldb # the headless Service in service.yaml — stable network id
|
||||
replicas: 1 # single-node-first: scale up, not out (see header)
|
||||
replicas: 0 # parked; scripts/restore-fleet.sh is the only supported scale-up path
|
||||
selector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: tidaldb
|
||||
|
||||
14
rust-toolchain.toml
Normal file
14
rust-toolchain.toml
Normal file
@ -0,0 +1,14 @@
|
||||
# Development toolchain for the whole workspace.
|
||||
#
|
||||
# This is NOT the library's MSRV. `tidaldb` and its siblings publish
|
||||
# `rust-version = "1.91"` (see Cargo.toml) and build on 1.91.0. The pin exists
|
||||
# because `tidalctl`'s AWS SDK dependency chain requires 1.91.1: on 1.91.0 every
|
||||
# workspace-wide command (`cargo test --workspace`, `cargo clippy --workspace`)
|
||||
# fails during resolution with "aws-types@1.3.16 requires rustc 1.91.1" and no
|
||||
# gate can run at all. Pinning the dev toolchain keeps the published MSRV honest
|
||||
# while making the workspace gates executable from a fresh clone.
|
||||
#
|
||||
# Raise this only alongside the `rust-version` fields it is meant to exceed.
|
||||
[toolchain]
|
||||
channel = "1.91.1"
|
||||
components = ["rustfmt", "clippy"]
|
||||
@ -26,26 +26,61 @@ BUILDER="${TIDAL_BUILDX_BUILDER:-amd64builder}"
|
||||
die() { echo "build-release: $*" >&2; exit 1; }
|
||||
|
||||
[ -n "$TAG" ] || die "usage: build-release.sh <tag> [server|dr|stress|all]"
|
||||
[ "$TAG" != "latest" ] || die "refusing mutable release tag 'latest'"
|
||||
[[ "$TAG" =~ ^[a-z0-9][a-z0-9._-]{0,127}$ ]] \
|
||||
|| die "tag must match ^[a-z0-9][a-z0-9._-]{0,127}$"
|
||||
case "$COMPONENT" in server|dr|stress|all) ;; *) die "component must be server|dr|stress|all" ;; esac
|
||||
|
||||
# ── Toolchain preflight (fail fast, no half-built image) ─────────────────────
|
||||
command -v cargo >/dev/null || die "cargo not on PATH"
|
||||
command -v docker >/dev/null || die "docker not on PATH"
|
||||
command -v "${TARGET}-gcc" >/dev/null || die "missing ${TARGET}-gcc (brew install ${TARGET})"
|
||||
PROTOC_BIN="${PROTOC:-$(command -v protoc || true)}"
|
||||
[ -n "$PROTOC_BIN" ] || die "protoc not found (brew install protobuf), or set PROTOC"
|
||||
rustup target list --installed 2>/dev/null | grep -qx "$TARGET" \
|
||||
command -v git >/dev/null || die "git not on PATH"
|
||||
command -v jq >/dev/null || die "jq not on PATH"
|
||||
if [ "$COMPONENT" != "stress" ]; then
|
||||
command -v cargo >/dev/null || die "cargo not on PATH"
|
||||
command -v "${TARGET}-gcc" >/dev/null \
|
||||
|| die "missing ${TARGET}-gcc (brew install ${TARGET})"
|
||||
PROTOC_BIN="${PROTOC:-$(command -v protoc || true)}"
|
||||
[ -n "$PROTOC_BIN" ] || die "protoc not found (brew install protobuf), or set PROTOC"
|
||||
rustup target list --installed 2>/dev/null | grep -qx "$TARGET" \
|
||||
|| die "rust target $TARGET not installed (rustup target add $TARGET)"
|
||||
fi
|
||||
|
||||
REPO_ROOT="$(cd "$(dirname "$0")/.." && pwd)"
|
||||
cd "$REPO_ROOT"
|
||||
|
||||
# Every image must have one reproducible source identity. Clean releases use
|
||||
# HEAD. An explicitly allowed canary hashes HEAD, both tracked diffs, and every
|
||||
# untracked non-ignored file (path + Git blob hash), so two distinct source
|
||||
# trees cannot silently share the same `-dirty` identity. This check also covers
|
||||
# stress: its Docker build consumes the repository source directly.
|
||||
UNTRACKED="$(git ls-files --others --exclude-standard)"
|
||||
if git diff --quiet && git diff --cached --quiet && [ -z "$UNTRACKED" ]; then
|
||||
export GIT_HASH="${TIDAL_BUILD_HASH:-$(git rev-parse --verify HEAD)}"
|
||||
else
|
||||
[ "${TIDAL_ALLOW_DIRTY:-0}" = "1" ] \
|
||||
|| die "source changes present; commit them or set TIDAL_ALLOW_DIRTY=1 for an explicit canary"
|
||||
DIRTY_HASH="$(
|
||||
{
|
||||
git rev-parse --verify HEAD
|
||||
git diff --binary
|
||||
git diff --cached --binary
|
||||
git ls-files --others --exclude-standard -z |
|
||||
while IFS= read -r -d '' file; do
|
||||
printf 'untracked\0%s\0' "$file"
|
||||
git hash-object -- "$file"
|
||||
done
|
||||
} | shasum -a 256 | cut -c1-16
|
||||
)"
|
||||
export GIT_HASH="${TIDAL_BUILD_HASH:-${DIRTY_HASH}-dirty}"
|
||||
fi
|
||||
echo "==> source identity $GIT_HASH"
|
||||
|
||||
# ── Pinned cross-compile environment (the documented recipe) ─────────────────
|
||||
export CC_x86_64_unknown_linux_gnu="${TARGET}-gcc"
|
||||
export CXX_x86_64_unknown_linux_gnu="${TARGET}-g++"
|
||||
export AR_x86_64_unknown_linux_gnu="${TARGET}-ar"
|
||||
export CARGO_TARGET_X86_64_UNKNOWN_LINUX_GNU_LINKER="${TARGET}-gcc"
|
||||
export PROTOC="$PROTOC_BIN"
|
||||
export PROTOC="${PROTOC_BIN:-}"
|
||||
|
||||
# Which crates to cross-compile for the selected component(s).
|
||||
PKGS=()
|
||||
@ -78,8 +113,9 @@ build_image() { # $1=image_name $2=dockerfile $3=stage_subdir
|
||||
docker buildx build --builder "$BUILDER" --platform linux/amd64 \
|
||||
-f "$dockerfile" -t "$image" --push "$ctx"
|
||||
local digest
|
||||
digest="$(docker buildx imagetools inspect "$image" --format '{{.Manifest.Digest}}' 2>/dev/null || true)"
|
||||
echo "RELEASE $1: $REGISTRY/$1@${digest:-<digest-unavailable>} (tag $TAG)"
|
||||
digest="$(docker buildx imagetools inspect "$image" --format '{{json .Manifest}}' | jq -er '.digest')"
|
||||
[ -n "$digest" ] || die "registry returned no digest for $image"
|
||||
echo "RELEASE $1: $image@$digest"
|
||||
}
|
||||
|
||||
if [ "$COMPONENT" = server ] || [ "$COMPONENT" = all ]; then
|
||||
@ -98,13 +134,16 @@ if [ "$COMPONENT" = dr ] || [ "$COMPONENT" = all ]; then
|
||||
fi
|
||||
|
||||
if [ "$COMPONENT" = stress ] || [ "$COMPONENT" = all ]; then
|
||||
# Pure-Rust, in-container build (includes the soak-eval binary). Context = repo
|
||||
# root so the workspace manifests resolve; the .dockerignore prunes it.
|
||||
echo "==> building $REGISTRY/stress:$TAG (in-container, includes soak-eval)"
|
||||
# Pure-Rust, in-container build (includes soak-eval + soak-watch). Context =
|
||||
# repo root so workspace manifests resolve; the .dockerignore prunes it. The
|
||||
# OCI revision label makes the source identity inspectable after the build.
|
||||
echo "==> building $REGISTRY/stress:$TAG (in-container, includes soak-eval + soak-watch)"
|
||||
docker buildx build --builder "$BUILDER" --platform linux/amd64 \
|
||||
--build-arg "GIT_HASH=$GIT_HASH" \
|
||||
-f docker/stress/Dockerfile -t "$REGISTRY/stress:$TAG" --push .
|
||||
d="$(docker buildx imagetools inspect "$REGISTRY/stress:$TAG" --format '{{.Manifest.Digest}}' 2>/dev/null || true)"
|
||||
echo "RELEASE stress: $REGISTRY/stress@${d:-<digest-unavailable>} (tag $TAG)"
|
||||
d="$(docker buildx imagetools inspect "$REGISTRY/stress:$TAG" --format '{{json .Manifest}}' | jq -er '.digest')"
|
||||
[ -n "$d" ] || die "registry returned no digest for $REGISTRY/stress:$TAG"
|
||||
echo "RELEASE stress: $REGISTRY/stress:$TAG@$d"
|
||||
fi
|
||||
|
||||
echo "==> done. Pin the printed @sha256 digest(s) in k8s/cluster/statefulset.yaml and the DR/soak manifests."
|
||||
echo "==> done. Deploy only the printed tag@sha256 references."
|
||||
|
||||
523
scripts/restore-fleet.sh
Executable file
523
scripts/restore-fleet.sh
Executable file
@ -0,0 +1,523 @@
|
||||
#!/usr/bin/env bash
|
||||
# Restore selected tidalDB workloads parked after the 2026-08-12 fleet review.
|
||||
#
|
||||
# This script never recreates the retired public Ingresses or DNS records. It
|
||||
# fails before mutation unless retained storage, immutable images, and scheduler
|
||||
# headroom satisfy the selected workload's current resource contract.
|
||||
|
||||
set -euo pipefail
|
||||
umask 077
|
||||
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
|
||||
REPO_ROOT="${SCRIPT_DIR%/scripts}"
|
||||
RESTORE_LOG_DIR="${TIDAL_RESTORE_LOG_DIR:-$REPO_ROOT/tmp/restore-logs}"
|
||||
mkdir -p "$RESTORE_LOG_DIR"
|
||||
RESTORE_STARTED_AT="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
|
||||
RESTORE_LOG_FILE="$RESTORE_LOG_DIR/restore-${RESTORE_STARTED_AT//:/}-$$.log"
|
||||
exec > >(tee -a "$RESTORE_LOG_FILE") 2>&1
|
||||
printf 'event=restore_start ts=%s pid=%s argv=' "$RESTORE_STARTED_AT" "$$"
|
||||
printf '%q ' "$@"
|
||||
printf 'log_file=%q\n' "${RESTORE_LOG_FILE##*/}"
|
||||
|
||||
export KUBECONFIG="${KUBECONFIG:-$HOME/.kube/orchard9-k3sf.yaml}"
|
||||
|
||||
RESTORE_CLUSTER=0
|
||||
RESTORE_STANDALONE=0
|
||||
RESTORE_SOAK=0
|
||||
ACCEPT_FAILED_GATE=0
|
||||
CHECK_ONLY=0
|
||||
STANDALONE_DATA_NODE=""
|
||||
CLUSTER_DATA_NODE_0=""
|
||||
CLUSTER_DATA_NODE_1=""
|
||||
CLUSTER_DATA_NODE_2=""
|
||||
PREFLIGHT_POD=""
|
||||
|
||||
usage() {
|
||||
cat <<'USAGE'
|
||||
Usage:
|
||||
scripts/restore-fleet.sh --cluster [--check]
|
||||
scripts/restore-fleet.sh --standalone [--check]
|
||||
scripts/restore-fleet.sh --cluster --soak --accept-failed-gate [--check]
|
||||
|
||||
Options:
|
||||
--cluster Restore the three-node RF3 cluster and soak monitor.
|
||||
--standalone Restore the internal standalone store. Public routing
|
||||
remains retired and must be provisioned separately.
|
||||
--soak Unsuspend the nightly endurance job after restoring the
|
||||
cluster. Requires --accept-failed-gate.
|
||||
--accept-failed-gate Acknowledge that the unchanged 200 rps gate passed only
|
||||
23/55 measured nights and is not a release signal.
|
||||
--check Run every preflight and print capacity; mutate nothing.
|
||||
-h, --help Show this help.
|
||||
|
||||
Exactly one of --cluster or --standalone is required.
|
||||
|
||||
No-argument execution is intentionally a no-op. The old script could restore a
|
||||
6-core cluster into a fleet with only ~2 cores free and then claim success via
|
||||
public routes that no longer exist.
|
||||
USAGE
|
||||
}
|
||||
|
||||
die() {
|
||||
printf 'event=restore_error ts=%s message=%q\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" "$*" >&2
|
||||
printf 'ERROR: %s\n' "$*" >&2
|
||||
exit 1
|
||||
}
|
||||
|
||||
for arg in "$@"; do
|
||||
case "$arg" in
|
||||
--cluster) RESTORE_CLUSTER=1 ;;
|
||||
--standalone) RESTORE_STANDALONE=1 ;;
|
||||
--soak) RESTORE_SOAK=1 ;;
|
||||
--accept-failed-gate) ACCEPT_FAILED_GATE=1 ;;
|
||||
--check) CHECK_ONLY=1 ;;
|
||||
-h|--help) usage; exit 0 ;;
|
||||
*) usage >&2; die "unknown option: $arg" ;;
|
||||
esac
|
||||
done
|
||||
|
||||
if [ "$RESTORE_CLUSTER" -eq 0 ] && [ "$RESTORE_STANDALONE" -eq 0 ]; then
|
||||
usage >&2
|
||||
exit 2
|
||||
fi
|
||||
if [ "$RESTORE_CLUSTER" -eq 1 ] && [ "$RESTORE_STANDALONE" -eq 1 ]; then
|
||||
die "choose exactly one data plane: --cluster or --standalone"
|
||||
fi
|
||||
if [ "$RESTORE_SOAK" -eq 1 ] && [ "$RESTORE_CLUSTER" -eq 0 ]; then
|
||||
die "--soak requires --cluster"
|
||||
fi
|
||||
if [ "$RESTORE_SOAK" -eq 1 ] && [ "$ACCEPT_FAILED_GATE" -eq 0 ]; then
|
||||
die "--soak requires --accept-failed-gate; the unchanged gate passed only 23/55 measured nights"
|
||||
fi
|
||||
|
||||
for command in kubectl jq grep; do
|
||||
command -v "$command" >/dev/null 2>&1 || die "$command is required"
|
||||
done
|
||||
kubectl version --client >/dev/null
|
||||
kubectl get namespace tidaldb tidaldb-cluster >/dev/null
|
||||
|
||||
KUBECONFIG_DISPLAY="$KUBECONFIG"
|
||||
if [[ "$KUBECONFIG_DISPLAY" == "$HOME/"* ]]; then
|
||||
# Display a literal tilde so the runtime log never records workstation identity.
|
||||
# shellcheck disable=SC2088
|
||||
printf -v KUBECONFIG_DISPLAY '~/%s' "${KUBECONFIG_DISPLAY#"$HOME"/}"
|
||||
fi
|
||||
printf '==> kubeconfig: %s\n' "$KUBECONFIG_DISPLAY"
|
||||
printf '==> context: %s\n' "$(kubectl config current-context)"
|
||||
printf '==> selected: cluster=%s standalone=%s soak=%s check_only=%s\n' \
|
||||
"$RESTORE_CLUSTER" "$RESTORE_STANDALONE" "$RESTORE_SOAK" "$CHECK_ONLY"
|
||||
|
||||
check_claim() {
|
||||
local namespace="$1" claim="$2" phase volume policy node
|
||||
IFS=$'\t' read -r phase volume <<<"$(
|
||||
kubectl -n "$namespace" get "pvc/$claim" \
|
||||
-o jsonpath='{.status.phase}{"\t"}{.spec.volumeName}'
|
||||
)"
|
||||
[ "$phase" = "Bound" ] || die "$namespace/$claim is $phase, expected Bound"
|
||||
[ -n "$volume" ] || die "$namespace/$claim has no backing PV"
|
||||
policy="$(kubectl get "pv/$volume" -o jsonpath='{.spec.persistentVolumeReclaimPolicy}')"
|
||||
[ "$policy" = "Retain" ] || die "$namespace/$claim backing PV $volume uses $policy, expected Retain"
|
||||
node="$(
|
||||
kubectl get "pv/$volume" -o json |
|
||||
jq -r '[
|
||||
try (
|
||||
.spec.nodeAffinity.required.nodeSelectorTerms[].matchExpressions[]
|
||||
| select(.key == "kubernetes.io/hostname" and .operator == "In")
|
||||
| .values[]
|
||||
) catch empty
|
||||
] | unique | if length <= 1 then (.[]? // "") else error("PV has ambiguous hostname affinity") end'
|
||||
)"
|
||||
case "$namespace/$claim" in
|
||||
tidaldb/tidaldb-data) STANDALONE_DATA_NODE="$node" ;;
|
||||
tidaldb-cluster/data-tidaldb-0) CLUSTER_DATA_NODE_0="$node" ;;
|
||||
tidaldb-cluster/data-tidaldb-1) CLUSTER_DATA_NODE_1="$node" ;;
|
||||
tidaldb-cluster/data-tidaldb-2) CLUSTER_DATA_NODE_2="$node" ;;
|
||||
esac
|
||||
printf ' %-18s %-28s Bound / Retain (node=%s)\n' \
|
||||
"$namespace" "$claim" "${node:-network-backed}"
|
||||
}
|
||||
|
||||
printf '==> retained storage\n'
|
||||
if [ "$RESTORE_STANDALONE" -eq 1 ]; then
|
||||
check_claim tidaldb tidaldb-data
|
||||
fi
|
||||
if [ "$RESTORE_CLUSTER" -eq 1 ]; then
|
||||
check_claim tidaldb-cluster data-tidaldb-0
|
||||
check_claim tidaldb-cluster data-tidaldb-1
|
||||
check_claim tidaldb-cluster data-tidaldb-2
|
||||
check_claim tidaldb-cluster tidal-soak-results
|
||||
fi
|
||||
|
||||
check_pod_template_images() {
|
||||
local namespace="$1" resource="$2" image images
|
||||
images="$(
|
||||
kubectl -n "$namespace" get "$resource" -o jsonpath='{range .spec.template.spec.initContainers[*]}{.image}{"\n"}{end}{range .spec.template.spec.containers[*]}{.image}{"\n"}{end}'
|
||||
)"
|
||||
[ -n "$images" ] || die "$namespace/$resource has no images"
|
||||
for image in $images; do
|
||||
case "$image" in
|
||||
*@sha256:*) ;;
|
||||
*) die "$namespace/$resource uses mutable image $image" ;;
|
||||
esac
|
||||
done
|
||||
}
|
||||
|
||||
check_cronjob_images() {
|
||||
local namespace="$1" name="$2" image images
|
||||
images="$(
|
||||
kubectl -n "$namespace" get "cronjob/$name" -o jsonpath='{range .spec.jobTemplate.spec.template.spec.initContainers[*]}{.image}{"\n"}{end}{range .spec.jobTemplate.spec.template.spec.containers[*]}{.image}{"\n"}{end}'
|
||||
)"
|
||||
[ -n "$images" ] || die "$namespace/cronjob/$name has no images"
|
||||
for image in $images; do
|
||||
case "$image" in
|
||||
*@sha256:*) ;;
|
||||
*) die "$namespace/cronjob/$name uses mutable image $image" ;;
|
||||
esac
|
||||
done
|
||||
}
|
||||
|
||||
printf '==> immutable workload images\n'
|
||||
if [ "$RESTORE_STANDALONE" -eq 1 ]; then
|
||||
check_pod_template_images tidaldb deployment/tidaldb
|
||||
fi
|
||||
if [ "$RESTORE_CLUSTER" -eq 1 ]; then
|
||||
check_pod_template_images tidaldb-cluster statefulset/tidaldb
|
||||
check_pod_template_images tidaldb-cluster deployment/tidal-soak-monitor
|
||||
fi
|
||||
if [ "$RESTORE_SOAK" -eq 1 ]; then
|
||||
check_cronjob_images tidaldb-cluster tidal-soak-nightly
|
||||
fi
|
||||
printf ' all selected restore dependencies are digest-pinned\n'
|
||||
|
||||
require_parked() {
|
||||
local namespace="$1" resource="$2" replicas
|
||||
replicas="$(kubectl -n "$namespace" get "$resource" -o jsonpath='{.spec.replicas}')"
|
||||
[ "${replicas:-0}" -eq 0 ] || \
|
||||
die "$namespace/$resource already requests $replicas replicas; this restore only accepts the recorded parked state"
|
||||
}
|
||||
|
||||
if [ "$RESTORE_STANDALONE" -eq 1 ]; then
|
||||
require_parked tidaldb deployment/tidaldb
|
||||
fi
|
||||
if [ "$RESTORE_CLUSTER" -eq 1 ]; then
|
||||
require_parked tidaldb-cluster statefulset/tidaldb
|
||||
require_parked tidaldb-cluster deployment/tidal-soak-monitor
|
||||
SOAK_SUSPENDED="$(kubectl -n tidaldb-cluster get cronjob/tidal-soak-nightly -o jsonpath='{.spec.suspend}')"
|
||||
[ "$SOAK_SUSPENDED" = "true" ] || \
|
||||
die "tidaldb-cluster/cronjob/tidal-soak-nightly must be suspended before restore"
|
||||
fi
|
||||
|
||||
# One quantity implementation serves both live pod accounting and workload
|
||||
# templates. The restore requirement therefore follows the manifest instead of
|
||||
# duplicating CPU/memory constants that can silently drift.
|
||||
# The `$q` names below are jq variables and must remain literal shell input.
|
||||
# shellcheck disable=SC2016
|
||||
RESOURCE_FILTERS='
|
||||
def cpu_m:
|
||||
tostring as $q |
|
||||
if $q == "" or $q == "null" then 0
|
||||
elif ($q | endswith("m")) then ($q | rtrimstr("m") | tonumber)
|
||||
elif ($q | endswith("u")) then (($q | rtrimstr("u") | tonumber) / 1000)
|
||||
elif ($q | endswith("n")) then (($q | rtrimstr("n") | tonumber) / 1000000)
|
||||
else (($q | tonumber) * 1000)
|
||||
end;
|
||||
def memory_bytes:
|
||||
tostring as $q |
|
||||
($q | capture("^(?<n>[0-9]+(?:\\.[0-9]+)?)(?<u>[A-Za-z]*)$")) as $p |
|
||||
($p.n | tonumber) as $n |
|
||||
if $p.u == "Ki" then $n * 1024
|
||||
elif $p.u == "Mi" then $n * 1048576
|
||||
elif $p.u == "Gi" then $n * 1073741824
|
||||
elif $p.u == "Ti" then $n * 1099511627776
|
||||
elif $p.u == "K" then $n * 1000
|
||||
elif $p.u == "M" then $n * 1000000
|
||||
elif $p.u == "G" then $n * 1000000000
|
||||
elif $p.u == "T" then $n * 1000000000000
|
||||
elif $p.u == "" then $n
|
||||
else error("unsupported memory quantity: " + $q)
|
||||
end;
|
||||
def pod_cpu:
|
||||
([.containers[]? | (.resources.requests.cpu // "0") | cpu_m] | add // 0) as $regular |
|
||||
([.initContainers[]? | (.resources.requests.cpu // "0") | cpu_m] | max // 0) as $init |
|
||||
([ $regular, $init ] | max) + ((.overhead.cpu // "0") | cpu_m);
|
||||
def pod_memory:
|
||||
([.containers[]? | (.resources.requests.memory // "0") | memory_bytes] | add // 0) as $regular |
|
||||
([.initContainers[]? | (.resources.requests.memory // "0") | memory_bytes] | max // 0) as $init |
|
||||
([ $regular, $init ] | max) + ((.overhead.memory // "0") | memory_bytes);
|
||||
'
|
||||
|
||||
workload_request() {
|
||||
local namespace="$1" resource="$2" template_kind="$3"
|
||||
kubectl -n "$namespace" get "$resource" -o json |
|
||||
jq -c --arg template_kind "$template_kind" "$RESOURCE_FILTERS"'
|
||||
(
|
||||
if $template_kind == "cronjob"
|
||||
then .spec.jobTemplate.spec.template.spec
|
||||
else .spec.template.spec
|
||||
end
|
||||
)
|
||||
| {cpu_m: (pod_cpu | ceil), memory_bytes: (pod_memory | ceil)}
|
||||
'
|
||||
}
|
||||
|
||||
capacity_snapshot() {
|
||||
jq -n \
|
||||
--slurpfile nodes <(kubectl get nodes -o json) \
|
||||
--slurpfile pods <(kubectl get pods -A -o json) "$RESOURCE_FILTERS"'
|
||||
($pods[0].items
|
||||
| map(select(.spec.nodeName != null and .status.phase != "Succeeded" and .status.phase != "Failed"))
|
||||
| group_by(.spec.nodeName)
|
||||
| map({
|
||||
key: .[0].spec.nodeName,
|
||||
value: {
|
||||
cpu_m: (map(.spec | pod_cpu) | add // 0),
|
||||
memory_bytes: (map(.spec | pod_memory) | add // 0)
|
||||
}
|
||||
})
|
||||
| from_entries) as $used |
|
||||
$nodes[0].items
|
||||
| map({
|
||||
node: .metadata.name,
|
||||
ready: any(.status.conditions[]?; .type == "Ready" and .status == "True"),
|
||||
schedulable: ((.spec.unschedulable // false) | not),
|
||||
alloc_cpu_m: (.status.allocatable.cpu | cpu_m),
|
||||
alloc_memory_bytes: (.status.allocatable.memory | memory_bytes),
|
||||
used_cpu_m: ($used[.metadata.name].cpu_m // 0),
|
||||
used_memory_bytes: ($used[.metadata.name].memory_bytes // 0)
|
||||
}
|
||||
| . + {
|
||||
free_cpu_m: (.alloc_cpu_m - .used_cpu_m),
|
||||
free_memory_bytes: (.alloc_memory_bytes - .used_memory_bytes)
|
||||
})
|
||||
'
|
||||
}
|
||||
|
||||
REQUIRED_CPU_M=0
|
||||
REQUIRED_MEMORY_BYTES=0
|
||||
DB_CPU_M=0
|
||||
DB_MEMORY_BYTES=0
|
||||
|
||||
add_requirement() {
|
||||
local request="$1" replicas="$2" cpu memory
|
||||
cpu="$(printf '%s\n' "$request" | jq -r '.cpu_m')"
|
||||
memory="$(printf '%s\n' "$request" | jq -r '.memory_bytes')"
|
||||
REQUIRED_CPU_M=$((REQUIRED_CPU_M + cpu * replicas))
|
||||
REQUIRED_MEMORY_BYTES=$((REQUIRED_MEMORY_BYTES + memory * replicas))
|
||||
}
|
||||
|
||||
if [ "$RESTORE_CLUSTER" -eq 1 ]; then
|
||||
DB_REQUEST="$(workload_request tidaldb-cluster statefulset/tidaldb pod)"
|
||||
MONITOR_REQUEST="$(workload_request tidaldb-cluster deployment/tidal-soak-monitor pod)"
|
||||
DB_CPU_M="$(printf '%s\n' "$DB_REQUEST" | jq -r '.cpu_m')"
|
||||
DB_MEMORY_BYTES="$(printf '%s\n' "$DB_REQUEST" | jq -r '.memory_bytes')"
|
||||
add_requirement "$DB_REQUEST" 3
|
||||
add_requirement "$MONITOR_REQUEST" 1
|
||||
fi
|
||||
if [ "$RESTORE_STANDALONE" -eq 1 ]; then
|
||||
DB_REQUEST="$(workload_request tidaldb deployment/tidaldb pod)"
|
||||
DB_CPU_M="$(printf '%s\n' "$DB_REQUEST" | jq -r '.cpu_m')"
|
||||
DB_MEMORY_BYTES="$(printf '%s\n' "$DB_REQUEST" | jq -r '.memory_bytes')"
|
||||
add_requirement "$DB_REQUEST" 1
|
||||
fi
|
||||
if [ "$RESTORE_SOAK" -eq 1 ]; then
|
||||
SOAK_REQUEST="$(workload_request tidaldb-cluster cronjob/tidal-soak-nightly cronjob)"
|
||||
add_requirement "$SOAK_REQUEST" 1
|
||||
fi
|
||||
|
||||
CAPACITY="$(capacity_snapshot)"
|
||||
printf '==> scheduler request headroom\n'
|
||||
printf '%s\n' "$CAPACITY" | jq -r '.[] | " \(.node): cpu=\(.free_cpu_m | floor)m free, memory=\((.free_memory_bytes / 1048576) | floor)Mi free, ready=\(.ready), schedulable=\(.schedulable)"'
|
||||
|
||||
FREE_CPU_M="$(printf '%s\n' "$CAPACITY" | jq '[.[] | select(.ready and .schedulable) | .free_cpu_m] | add // 0 | floor')"
|
||||
FREE_MEMORY_BYTES="$(printf '%s\n' "$CAPACITY" | jq '[.[] | select(.ready and .schedulable) | .free_memory_bytes] | add // 0 | floor')"
|
||||
|
||||
[ "$FREE_CPU_M" -ge "$REQUIRED_CPU_M" ] || \
|
||||
die "selected restore needs ${REQUIRED_CPU_M}m CPU requests; only ${FREE_CPU_M}m is free"
|
||||
[ "$FREE_MEMORY_BYTES" -ge "$REQUIRED_MEMORY_BYTES" ] || \
|
||||
die "selected restore needs $((REQUIRED_MEMORY_BYTES / 1048576))Mi memory requests; only $((FREE_MEMORY_BYTES / 1048576))Mi is free"
|
||||
|
||||
node_has_capacity() {
|
||||
local node="$1"
|
||||
printf '%s\n' "$CAPACITY" |
|
||||
jq -e --arg node "$node" --argjson cpu "$DB_CPU_M" --argjson memory "$DB_MEMORY_BYTES" \
|
||||
'any(.[]; .node == $node and .ready and .schedulable and .free_cpu_m >= $cpu and .free_memory_bytes >= $memory)' \
|
||||
>/dev/null
|
||||
}
|
||||
|
||||
if [ "$RESTORE_CLUSTER" -eq 1 ]; then
|
||||
if [ -n "$CLUSTER_DATA_NODE_0" ] && [ -n "$CLUSTER_DATA_NODE_1" ] && [ -n "$CLUSTER_DATA_NODE_2" ]; then
|
||||
UNIQUE_DATA_NODES="$(
|
||||
jq -n --arg n0 "$CLUSTER_DATA_NODE_0" --arg n1 "$CLUSTER_DATA_NODE_1" --arg n2 "$CLUSTER_DATA_NODE_2" \
|
||||
'[$n0, $n1, $n2] | unique | length'
|
||||
)"
|
||||
[ "$UNIQUE_DATA_NODES" -eq 3 ] || die "RF3 data volumes are not pinned to three distinct nodes"
|
||||
for node in "$CLUSTER_DATA_NODE_0" "$CLUSTER_DATA_NODE_1" "$CLUSTER_DATA_NODE_2"; do
|
||||
node_has_capacity "$node" || \
|
||||
die "RF3 data volume node $node lacks ${DB_CPU_M}m CPU or $((DB_MEMORY_BYTES / 1048576))Mi memory"
|
||||
done
|
||||
elif [ -z "$CLUSTER_DATA_NODE_0$CLUSTER_DATA_NODE_1$CLUSTER_DATA_NODE_2" ]; then
|
||||
CLUSTER_NODES="$(
|
||||
printf '%s\n' "$CAPACITY" |
|
||||
jq --argjson cpu "$DB_CPU_M" --argjson memory "$DB_MEMORY_BYTES" \
|
||||
'[.[] | select(.ready and .schedulable and .free_cpu_m >= $cpu and .free_memory_bytes >= $memory)] | length'
|
||||
)"
|
||||
[ "$CLUSTER_NODES" -ge 3 ] || \
|
||||
die "RF3 topology needs three nodes with ${DB_CPU_M}m CPU and $((DB_MEMORY_BYTES / 1048576))Mi memory free each; only $CLUSTER_NODES qualify"
|
||||
else
|
||||
die "RF3 data volumes have mixed local and network-backed placement; manual scheduling proof required"
|
||||
fi
|
||||
fi
|
||||
|
||||
if [ "$RESTORE_STANDALONE" -eq 1 ]; then
|
||||
if [ -n "$STANDALONE_DATA_NODE" ]; then
|
||||
node_has_capacity "$STANDALONE_DATA_NODE" || \
|
||||
die "standalone data volume node $STANDALONE_DATA_NODE lacks ${DB_CPU_M}m CPU or $((DB_MEMORY_BYTES / 1048576))Mi memory"
|
||||
else
|
||||
STANDALONE_TARGET_NODE="$(
|
||||
printf '%s\n' "$CAPACITY" |
|
||||
jq -r --argjson cpu "$DB_CPU_M" --argjson memory "$DB_MEMORY_BYTES" \
|
||||
'first(.[] | select(.ready and .schedulable and .free_cpu_m >= $cpu and .free_memory_bytes >= $memory) | .node) // ""'
|
||||
)"
|
||||
[ -n "$STANDALONE_TARGET_NODE" ] || \
|
||||
die "no node can place standalone request ${DB_CPU_M}m/$((DB_MEMORY_BYTES / 1048576))Mi"
|
||||
node_has_capacity "$STANDALONE_TARGET_NODE" || \
|
||||
die "selected standalone node $STANDALONE_TARGET_NODE lost capacity during preflight"
|
||||
fi
|
||||
fi
|
||||
|
||||
printf '==> preflight passed: manifest-derived request %sm CPU and %sMi memory\n' \
|
||||
"$REQUIRED_CPU_M" "$((REQUIRED_MEMORY_BYTES / 1048576))"
|
||||
|
||||
if [ "$CHECK_ONLY" -eq 1 ]; then
|
||||
printf 'VERDICT: RESTORE PREFLIGHT PASSED (no resources changed)\n'
|
||||
printf 'event=restore_complete ts=%s mode=check verdict=passed\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)"
|
||||
exit 0
|
||||
fi
|
||||
MUTATION_STARTED=1
|
||||
rollback_on_exit() {
|
||||
local rc=$?
|
||||
trap - EXIT
|
||||
set +e
|
||||
|
||||
if [ -n "$PREFLIGHT_POD" ]; then
|
||||
kubectl -n tidaldb-cluster delete "pod/$PREFLIGHT_POD" \
|
||||
--ignore-not-found --wait=true >/dev/null 2>&1 ||
|
||||
printf 'WARNING: could not delete evaluator preflight pod %s\n' "$PREFLIGHT_POD" >&2
|
||||
fi
|
||||
if [ "$rc" -eq 0 ] || [ "$MUTATION_STARTED" -eq 0 ]; then
|
||||
exit "$rc"
|
||||
fi
|
||||
|
||||
printf 'ERROR: restore failed rc=%s; returning selected workloads to the parked state\n' "$rc" >&2
|
||||
printf 'event=restore_rollback ts=%s rc=%s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" "$rc" >&2
|
||||
if [ "$RESTORE_SOAK" -eq 1 ]; then
|
||||
kubectl -n tidaldb-cluster patch cronjob/tidal-soak-nightly \
|
||||
--type=merge -p '{"spec":{"suspend":true}}' >/dev/null ||
|
||||
printf 'ERROR: rollback could not suspend tidal-soak-nightly\n' >&2
|
||||
fi
|
||||
if [ "$RESTORE_CLUSTER" -eq 1 ]; then
|
||||
kubectl -n tidaldb-cluster scale deployment/tidal-soak-monitor --replicas=0 >/dev/null ||
|
||||
printf 'ERROR: rollback could not park tidal-soak-monitor\n' >&2
|
||||
kubectl -n tidaldb-cluster scale statefulset/tidaldb --replicas=0 >/dev/null ||
|
||||
printf 'ERROR: rollback could not park tidaldb StatefulSet\n' >&2
|
||||
kubectl -n tidaldb-cluster wait --for=delete pod \
|
||||
-l app.kubernetes.io/name=tidaldb --timeout=300s >/dev/null 2>&1 ||
|
||||
printf 'WARNING: rollback did not observe all cluster pods deleted within 300s\n' >&2
|
||||
kubectl -n tidaldb-cluster wait --for=delete pod \
|
||||
-l app.kubernetes.io/name=tidal-soak-monitor --timeout=300s >/dev/null 2>&1 ||
|
||||
printf 'WARNING: rollback did not observe soak monitor deletion within 300s\n' >&2
|
||||
fi
|
||||
if [ "$RESTORE_STANDALONE" -eq 1 ]; then
|
||||
kubectl -n tidaldb scale deployment/tidaldb --replicas=0 >/dev/null ||
|
||||
printf 'ERROR: rollback could not park standalone tidaldb\n' >&2
|
||||
kubectl -n tidaldb wait --for=delete pod \
|
||||
-l app.kubernetes.io/name=tidaldb --timeout=300s >/dev/null 2>&1 ||
|
||||
printf 'WARNING: rollback did not observe standalone pod deletion within 300s\n' >&2
|
||||
fi
|
||||
printf '==> rollback requested; inspect workload state before retrying\n' >&2
|
||||
exit "$rc"
|
||||
}
|
||||
trap rollback_on_exit EXIT
|
||||
|
||||
if [ "$RESTORE_CLUSTER" -eq 1 ]; then
|
||||
printf '==> restoring RF3 cluster\n'
|
||||
kubectl -n tidaldb-cluster scale statefulset/tidaldb --replicas=3
|
||||
kubectl -n tidaldb-cluster rollout status statefulset/tidaldb --timeout=900s
|
||||
kubectl -n tidaldb-cluster wait --for=condition=Ready \
|
||||
pod -l app.kubernetes.io/name=tidaldb --timeout=300s
|
||||
|
||||
printf '==> restoring soak monitor (nightly job stays suspended)\n'
|
||||
kubectl -n tidaldb-cluster scale deployment/tidal-soak-monitor --replicas=1
|
||||
kubectl -n tidaldb-cluster rollout status deployment/tidal-soak-monitor --timeout=300s
|
||||
fi
|
||||
|
||||
if [ "$RESTORE_STANDALONE" -eq 1 ]; then
|
||||
printf '==> restoring internal standalone store\n'
|
||||
kubectl -n tidaldb scale deployment/tidaldb --replicas=1
|
||||
kubectl -n tidaldb rollout status deployment/tidaldb --timeout=600s
|
||||
printf ' public Ingress and DNS remain retired by design\n'
|
||||
fi
|
||||
|
||||
if [ "$RESTORE_SOAK" -eq 1 ]; then
|
||||
SOAK_IMAGE="$(
|
||||
kubectl -n tidaldb-cluster get cronjob/tidal-soak-nightly \
|
||||
-o jsonpath='{.spec.jobTemplate.spec.template.spec.containers[0].image}'
|
||||
)"
|
||||
PREFLIGHT_POD=tidal-soak-eval-preflight
|
||||
kubectl -n tidaldb-cluster delete "pod/$PREFLIGHT_POD" \
|
||||
--ignore-not-found --wait=true >/dev/null
|
||||
PREFLIGHT_OVERRIDES="$(
|
||||
jq -nc --arg image "$SOAK_IMAGE" '{
|
||||
spec: {
|
||||
automountServiceAccountToken: false,
|
||||
securityContext: {
|
||||
runAsNonRoot: true,
|
||||
runAsUser: 1000,
|
||||
runAsGroup: 1000,
|
||||
seccompProfile: {type: "RuntimeDefault"}
|
||||
},
|
||||
containers: [{
|
||||
name: "preflight",
|
||||
image: $image,
|
||||
imagePullPolicy: "IfNotPresent",
|
||||
command: ["soak-eval", "--help"],
|
||||
resources: {
|
||||
requests: {cpu: "10m", memory: "16Mi"},
|
||||
limits: {cpu: "100m", memory: "64Mi"}
|
||||
},
|
||||
securityContext: {
|
||||
allowPrivilegeEscalation: false,
|
||||
readOnlyRootFilesystem: true,
|
||||
capabilities: {drop: ["ALL"]}
|
||||
}
|
||||
}]
|
||||
}
|
||||
}'
|
||||
)"
|
||||
kubectl -n tidaldb-cluster run "$PREFLIGHT_POD" \
|
||||
--restart=Never --image="$SOAK_IMAGE" --overrides="$PREFLIGHT_OVERRIDES"
|
||||
kubectl -n tidaldb-cluster wait --for=jsonpath='{.status.phase}'=Succeeded \
|
||||
"pod/$PREFLIGHT_POD" --timeout=180s
|
||||
PREFLIGHT_HELP="$(kubectl -n tidaldb-cluster logs "$PREFLIGHT_POD")"
|
||||
printf '%s\n' "$PREFLIGHT_HELP"
|
||||
grep -q -- '--ledger-file' <<<"$PREFLIGHT_HELP" ||
|
||||
die "stress image soak-eval lacks --ledger-file; refusing the non-atomic nightly contract"
|
||||
grep -q -- '--no-write-streak' <<<"$PREFLIGHT_HELP" ||
|
||||
die "stress image soak-eval lacks --no-write-streak; refusing the non-atomic nightly contract"
|
||||
kubectl -n tidaldb-cluster delete "pod/$PREFLIGHT_POD" --wait=true >/dev/null
|
||||
PREFLIGHT_POD=""
|
||||
|
||||
printf '==> unsuspending nightly soak after evaluator proof\n'
|
||||
kubectl -n tidaldb-cluster patch cronjob/tidal-soak-nightly \
|
||||
--type=merge -p '{"spec":{"suspend":false}}'
|
||||
fi
|
||||
|
||||
printf '==> final state\n'
|
||||
if [ "$RESTORE_CLUSTER" -eq 1 ]; then
|
||||
kubectl -n tidaldb-cluster get \
|
||||
statefulset/tidaldb deployment/tidal-soak-monitor cronjob/tidal-soak-nightly
|
||||
fi
|
||||
if [ "$RESTORE_STANDALONE" -eq 1 ]; then
|
||||
kubectl -n tidaldb get deployment/tidaldb
|
||||
fi
|
||||
printf 'event=restore_complete ts=%s mode=restore verdict=passed\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)"
|
||||
printf 'VERDICT: SELECTED TIDALDB WORKLOADS RESTORED\n'
|
||||
@ -497,6 +497,48 @@ struct DeferredRetire {
|
||||
deadline: Instant,
|
||||
}
|
||||
|
||||
const RESEED_TERMINATION_LOG: &str = "/dev/termination-log";
|
||||
|
||||
fn write_reseed_termination_message(
|
||||
path: &std::path::Path,
|
||||
region: &str,
|
||||
shard: ShardId,
|
||||
) -> std::io::Result<()> {
|
||||
let message = serde_json::json!({
|
||||
"reason": "reseed_self_restart",
|
||||
"region": region,
|
||||
"shard": shard.0,
|
||||
});
|
||||
std::fs::write(path, message.to_string())
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
#[allow(clippy::unwrap_used)]
|
||||
mod termination_message_tests {
|
||||
use super::*;
|
||||
|
||||
#[test]
|
||||
fn reseed_self_restart_writes_machine_readable_termination_reason() {
|
||||
let suffix = std::time::SystemTime::now()
|
||||
.duration_since(std::time::UNIX_EPOCH)
|
||||
.unwrap()
|
||||
.as_nanos();
|
||||
let path = std::env::temp_dir().join(format!(
|
||||
"tidaldb-termination-message-{}-{suffix}",
|
||||
std::process::id()
|
||||
));
|
||||
|
||||
write_reseed_termination_message(&path, "us-east", ShardId(7)).unwrap();
|
||||
let message: serde_json::Value =
|
||||
serde_json::from_slice(&std::fs::read(&path).unwrap()).unwrap();
|
||||
std::fs::remove_file(path).unwrap();
|
||||
|
||||
assert_eq!(message["reason"], "reseed_self_restart");
|
||||
assert_eq!(message["region"], "us-east");
|
||||
assert_eq!(message["shard"], 7);
|
||||
}
|
||||
}
|
||||
|
||||
impl ShardReplica {
|
||||
/// Build the single region named `region_name` from `topology`.
|
||||
///
|
||||
@ -3259,6 +3301,19 @@ impl ShardReplica {
|
||||
/// use the existing shutdown-signal path (SIGTERM-equivalent) so
|
||||
/// `serve_state`'s graceful shutdown fires — never `process::abort`.
|
||||
fn fire_graceful_self_restart(&self) {
|
||||
if let Err(error) = write_reseed_termination_message(
|
||||
std::path::Path::new(RESEED_TERMINATION_LOG),
|
||||
&self.region_name,
|
||||
self.group_shard,
|
||||
) {
|
||||
tracing::warn!(
|
||||
path = RESEED_TERMINATION_LOG,
|
||||
region = %self.region_name,
|
||||
shard = self.group_shard.0,
|
||||
error = %error,
|
||||
"failed to write Kubernetes termination message for reseed self-restart"
|
||||
);
|
||||
}
|
||||
self.set_shutting_down();
|
||||
if let Err(e) = std::thread::Builder::new()
|
||||
.name("tidal-reseed-restart".into())
|
||||
@ -8617,10 +8672,6 @@ where
|
||||
offload_read(f).await.map_err(ClusterAppError)
|
||||
}
|
||||
|
||||
/// m11p7 — the cluster auth middleware exercised through the real layer
|
||||
/// composition (`middleware::from_fn`), not just the pure `ClusterCreds`
|
||||
/// helpers. A layer-order or marker-set regression in [`cluster_auth_middleware`]
|
||||
/// is caught HERE (the helper unit tests in `security.rs` would still pass).
|
||||
#[cfg(test)]
|
||||
#[allow(clippy::unwrap_used)]
|
||||
mod aggregate_region_row_tests {
|
||||
@ -8983,22 +9034,19 @@ mod forward_retry_tests {
|
||||
};
|
||||
use axum::http::StatusCode;
|
||||
|
||||
fn ok(
|
||||
status: StatusCode,
|
||||
body: serde_json::Value,
|
||||
) -> Result<forward::ForwardedResponse, String> {
|
||||
Ok(forward::ForwardedResponse {
|
||||
fn forwarded(status: StatusCode, body: serde_json::Value) -> forward::ForwardedResponse {
|
||||
forward::ForwardedResponse {
|
||||
seq: None,
|
||||
deduplicated: false,
|
||||
status,
|
||||
body,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn success_relays_immediately() {
|
||||
// A 2xx leader verdict is handed straight back — never retried.
|
||||
let outcome = ok(StatusCode::CREATED, serde_json::Value::Null);
|
||||
let outcome = Ok::<_, String>(forwarded(StatusCode::CREATED, serde_json::Value::Null));
|
||||
assert_eq!(
|
||||
classify_forward_attempt(&outcome, 1, Duration::ZERO),
|
||||
ForwardStep::Relay
|
||||
@ -9010,10 +9058,10 @@ mod forward_retry_tests {
|
||||
// The leader's bounded write pool shed this write (429 + retry_after_ms):
|
||||
// it created no log entry, so re-forwarding is safe. We retry, honoring
|
||||
// the leader's hint.
|
||||
let outcome = ok(
|
||||
let outcome = Ok::<_, String>(forwarded(
|
||||
StatusCode::TOO_MANY_REQUESTS,
|
||||
serde_json::json!({ "retry_after_ms": 80 }),
|
||||
);
|
||||
));
|
||||
assert_eq!(
|
||||
classify_forward_attempt(&outcome, 1, Duration::ZERO),
|
||||
ForwardStep::Retry { backoff_ms: 80 }
|
||||
@ -9025,7 +9073,10 @@ mod forward_retry_tests {
|
||||
// On the LAST attempt a 429 is relayed verbatim: the targeted-follower
|
||||
// client sees a retryable 429 (honest backpressure), NOT a 503 that would
|
||||
// wrongly say the leader is unreachable.
|
||||
let outcome = ok(StatusCode::TOO_MANY_REQUESTS, serde_json::Value::Null);
|
||||
let outcome = Ok::<_, String>(forwarded(
|
||||
StatusCode::TOO_MANY_REQUESTS,
|
||||
serde_json::Value::Null,
|
||||
));
|
||||
assert_eq!(
|
||||
classify_forward_attempt(&outcome, FORWARD_MAX_ATTEMPTS, Duration::ZERO),
|
||||
ForwardStep::Relay
|
||||
|
||||
@ -9,8 +9,8 @@
|
||||
//! (`axum::serve(...).with_graceful_shutdown(...)`) blocks until every in-flight
|
||||
//! connection closes. Sibling region nodes hold long-lived keep-alive HTTP
|
||||
//! connections that do not close promptly on our SIGTERM, so the drain — and thus
|
||||
//! `axum::serve(...).await` — never returned inside the 60s grace; k8s then
|
||||
//! SIGKILLed the process, and SIGKILL cannot run `Drop`/`shutdown_inner` (where
|
||||
//! `axum::serve(...).await` — never returned inside the 60s grace; k8s then sent
|
||||
//! `SIGKILL`. That signal cannot run `Drop`/`shutdown_inner` (where
|
||||
//! the WAL checkpoint marker AND the HNSW-graph checkpoint are written).
|
||||
//!
|
||||
//! The fix (in `main.rs` + `cluster/node.rs`):
|
||||
|
||||
@ -54,6 +54,7 @@ rand = "0.9"
|
||||
# tidal-net — drops the OpenSSL/native-tls system dependency for a clean static
|
||||
# build. `json` for the typed request/response bodies.
|
||||
reqwest = { version = "0.12", default-features = false, features = ["json", "rustls-tls"] }
|
||||
time = { version = "0.3", features = ["formatting", "parsing"] }
|
||||
|
||||
[dev-dependencies]
|
||||
# Honest before/after for the generator's OWN per-request cost (no server, no
|
||||
|
||||
@ -4,7 +4,7 @@
|
||||
# writes its JSON summaries to, so EVERYTHING the operator needs to judge the
|
||||
# 30-night streak lives in one place and survives the laptop session ending:
|
||||
#
|
||||
# 1. Restart watch: every 5 min, snapshot the tidaldb-{0,1,2} pod restart
|
||||
# 1. Restart watch: every minute, snapshot the tidaldb-{0,1,2} pod identity,
|
||||
# counts + phase into /results/restarts.tsv. The GA bar is not just
|
||||
# "30 green soak verdicts" — it is ZERO unrecovered failures over the
|
||||
# window. An under-load pod restart during a soak night must be visible
|
||||
@ -15,9 +15,9 @@
|
||||
# and read the ledger / nightly summaries / restart log from a browser at
|
||||
# any time, from any machine, without exec'ing into a pod.
|
||||
#
|
||||
# The monitor does NOT generate load and does NOT gate anything — it is a passive
|
||||
# recorder. The pass/fail signal is the CronJob's Job exit codes; this just makes
|
||||
# the 30-night picture observable and durable.
|
||||
# The monitor does not generate load. The nightly Job is the authoritative gate;
|
||||
# this deployment maintains the evidence stream and continuously recomputes the
|
||||
# durable streak/status surface.
|
||||
#
|
||||
# Apply: kubectl apply -f tidal-stress/k8s/soak-monitor.yaml
|
||||
# Read: kubectl port-forward deploy/tidal-soak-monitor 8080:8080 -n tidaldb-cluster
|
||||
@ -76,7 +76,7 @@ metadata:
|
||||
app.kubernetes.io/name: tidal-soak
|
||||
app.kubernetes.io/part-of: tidaldb
|
||||
spec:
|
||||
replicas: 1
|
||||
replicas: 0 # parked; scripts/restore-fleet.sh starts it only after capacity checks
|
||||
selector:
|
||||
matchLabels:
|
||||
app.kubernetes.io/name: tidal-soak-monitor
|
||||
@ -95,28 +95,22 @@ spec:
|
||||
seccompProfile:
|
||||
type: RuntimeDefault
|
||||
containers:
|
||||
# ── Restart-watch sidecar: kubectl snapshot loop ──────────────────────
|
||||
# ── Restart watcher: direct Kubernetes API sampler ───────────────────
|
||||
- name: restart-watch
|
||||
image: bitnami/kubectl:latest
|
||||
image: registry.threesix.ai/tidal/stress:m12-fleet-remediation-20260815-r2@sha256:86bfa95f63ba01ed4afa120e41684cd8694743ec39aca72ef0bd68e5d9ef7f8b
|
||||
imagePullPolicy: IfNotPresent
|
||||
command: ["/bin/sh", "-c"]
|
||||
command: ["soak-watch"]
|
||||
args:
|
||||
- |
|
||||
echo "restart-watch up $(date -u +%FT%TZ)"
|
||||
# Header once (only if the file is new/empty).
|
||||
if [ ! -s /results/restarts.tsv ]; then
|
||||
printf 'ts_utc\tpod\trestarts\tphase\tready\n' >> /results/restarts.tsv
|
||||
fi
|
||||
while true; do
|
||||
TS="$(date -u +%FT%TZ)"
|
||||
kubectl get pods -n tidaldb-cluster \
|
||||
-l app.kubernetes.io/name=tidaldb \
|
||||
-o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.status.containerStatuses[0].restartCount}{"\t"}{.status.phase}{"\t"}{.status.containerStatuses[0].ready}{"\n"}{end}' 2>/dev/null \
|
||||
| while IFS="$(printf '\t')" read -r POD RC PH RD; do
|
||||
[ -n "$POD" ] && printf '%s\t%s\t%s\t%s\t%s\n' "$TS" "$POD" "$RC" "$PH" "$RD" >> /results/restarts.tsv
|
||||
done
|
||||
sleep 300
|
||||
done
|
||||
- --namespace
|
||||
- tidaldb-cluster
|
||||
- --label-selector
|
||||
- app.kubernetes.io/name=tidaldb
|
||||
- --container
|
||||
- tidaldb
|
||||
- --results-file
|
||||
- /results/restarts.tsv
|
||||
- --interval-seconds
|
||||
- "60"
|
||||
resources:
|
||||
requests: { cpu: 10m, memory: 32Mi }
|
||||
limits: { cpu: 100m, memory: 128Mi }
|
||||
@ -127,18 +121,12 @@ spec:
|
||||
volumeMounts:
|
||||
- name: results
|
||||
mountPath: /results
|
||||
# ── Streak evaluator: the GA-bar verdict + alerting ───────────────────
|
||||
# Joins ledger.tsv (the nightly Job verdicts) with restarts.tsv (the
|
||||
# cumulative per-pod restart counts the sidecar above records) into the
|
||||
# HONEST streak: a night is green iff the soak passed AND no pod restarted
|
||||
# under load. Runs the proven `soak-eval` core (table-driven-tested in
|
||||
# tidal_stress::soak_eval), writes streak.tsv, and emits EXACTLY ONE alert
|
||||
# per non-green night transition (a durable ALERT-<date>.txt the http
|
||||
# surface serves, plus an optional webhook POST when NOTIFY_WEBHOOK is set
|
||||
# — reuse the cluster's notify service by pointing that at it). It does NOT
|
||||
# gate load; it makes the 30-night picture trustworthy and observable.
|
||||
# ── Streak evaluator: durable GA-bar status ────────────────────────────
|
||||
# Joins the nightly verdict ledger with the restart evidence above.
|
||||
# The nightly Job fails closed on this same evaluator; this sidecar keeps
|
||||
# streak.tsv and one durable ALERT-<date>.txt current for operators.
|
||||
- name: evaluator
|
||||
image: registry.threesix.ai/tidal/stress:m12-soak-eval
|
||||
image: registry.threesix.ai/tidal/stress:m12-fleet-remediation-20260815-r2@sha256:86bfa95f63ba01ed4afa120e41684cd8694743ec39aca72ef0bd68e5d9ef7f8b
|
||||
imagePullPolicy: IfNotPresent
|
||||
command: ["/bin/sh", "-c"]
|
||||
args:
|
||||
@ -159,8 +147,6 @@ spec:
|
||||
MSG="tidalDB soak: night $LAST NON-GREEN — $REASON (streak reset)"
|
||||
echo "ALERT: $MSG"
|
||||
printf '%s\t%s\n' "$LAST" "$REASON" > "/results/ALERT-$LAST.txt"
|
||||
[ -n "${NOTIFY_WEBHOOK:-}" ] && \
|
||||
wget -q -O- --post-data="{\"text\":\"$MSG\"}" --header='Content-Type: application/json' "$NOTIFY_WEBHOOK" >/dev/null 2>&1 || true
|
||||
echo "$LAST" > /results/.last-alert
|
||||
fi
|
||||
fi
|
||||
@ -171,11 +157,6 @@ spec:
|
||||
find /results -maxdepth 1 -name 'soak-*.json' -mtime +31 -delete 2>/dev/null || true
|
||||
sleep 300
|
||||
done
|
||||
env:
|
||||
# Point at the cluster's notify service to deliver alerts off-cluster;
|
||||
# unset = log + durable ALERT-<date>.txt only (still status-surfaced).
|
||||
- name: NOTIFY_WEBHOOK
|
||||
value: ""
|
||||
resources:
|
||||
requests: { cpu: 10m, memory: 32Mi }
|
||||
limits: { cpu: 200m, memory: 128Mi }
|
||||
@ -188,7 +169,7 @@ spec:
|
||||
mountPath: /results
|
||||
# ── HTTP read surface: serve the durable result dir ───────────────────
|
||||
- name: http
|
||||
image: busybox:1.36
|
||||
image: busybox@sha256:73aaf090f3d85aa34ee199857f03fa3a95c8ede2ffd4cc2cdb5b94e566b11662
|
||||
imagePullPolicy: IfNotPresent
|
||||
# busybox httpd: one-shot static file server rooted at /results.
|
||||
command: ["/bin/sh", "-c"]
|
||||
@ -199,12 +180,11 @@ spec:
|
||||
ports:
|
||||
- name: http
|
||||
containerPort: 8080
|
||||
# Readiness gates the Service endpoint on the http surface actually
|
||||
# serving — a monitor whose node is rebooting must drop out of endpoints
|
||||
# rather than silently serve nothing while a restart it should be
|
||||
# recording slips by unobserved.
|
||||
# Readiness gates the Service endpoint on the HTTP listener. The
|
||||
# result directory intentionally has no index file, so GET / returns
|
||||
# 404 even while the static server is healthy.
|
||||
readinessProbe:
|
||||
httpGet: { path: /, port: 8080 }
|
||||
tcpSocket: { port: http }
|
||||
initialDelaySeconds: 5
|
||||
periodSeconds: 10
|
||||
failureThreshold: 3
|
||||
|
||||
@ -22,17 +22,14 @@
|
||||
# The streak passes when 30 CONSECUTIVE nightly Jobs have exited 0 with zero
|
||||
# under-load pod restarts (the monitor records restarts; see soak-monitor.yaml).
|
||||
#
|
||||
# SUSTAINED RATE: 200 rps mixed (peach: feed/search/view/like/skip/item/embed).
|
||||
# RE-SCOPED 2026-06-19 from 500 rps. The original 500 rps cited the rc12 READ-SLA
|
||||
# gate, but the peach mix is WRITE-heavy (view/like/skip ≈ 90% of ops) and the
|
||||
# write path is structurally capped on this 3-node RF3 full-placement cluster
|
||||
# (every follower applies every 1536-D write; T5). At 500 rps the soak FAILED
|
||||
# every night (2026-06-17/18/19: p99 198-305ms, error 12-33%). A 2026-06-19
|
||||
# capacity sweep on a settled cluster measured the peach mix CLEAN (0.00% error,
|
||||
# p99 schedule-lag <5ms, 0 follower restarts) at 100/150/200/250 rps; the knee is
|
||||
# between 250 and 500. 200 rps is the measured sustainable rate with margin for a
|
||||
# 1-hour × 30-consecutive-night endurance run. Raise only with more write capacity
|
||||
# (≥5 nodes / partitioned placement) — see orchard9-k3sf cluster-state.yaml.
|
||||
# HISTORICAL RATE: 200 rps mixed (peach: feed/search/view/like/skip/item/embed).
|
||||
# The 2026-06-19 one-shot sweep found 200 rps clean, but the subsequent nightly
|
||||
# history was only 23 PASS / 32 FAIL and included repeated 4 GiB OOMKills. This
|
||||
# CronJob is therefore parked (`suspend: true`). Resume it only after a controlled
|
||||
# 100k/1536-D/200-rps memory profile passes under the corrected resource envelope,
|
||||
# restart evidence is fail-closed, and the exact immutable images below exist.
|
||||
# Adding nodes to RF3 full placement does not increase write capacity; the old
|
||||
# ≥2.5x scaling goal requires Ref-B partitioned placement.
|
||||
#
|
||||
# WHY A CRONJOB, NOT ONE LONG JOB: the GA bar is literally "30 CONSECUTIVE DAYS"
|
||||
# of an independent nightly verdict — a CronJob produces exactly that audit trail
|
||||
@ -57,6 +54,7 @@ spec:
|
||||
# morning activity. concurrencyPolicy Forbid: never overlap two soaks (they
|
||||
# would contend for the same 3 nodes and mutually depress p99 → false FAIL).
|
||||
schedule: "0 2 * * *"
|
||||
suspend: true
|
||||
concurrencyPolicy: Forbid
|
||||
startingDeadlineSeconds: 3600
|
||||
successfulJobsHistoryLimit: 30 # keep all 30 nights of PASS verdicts
|
||||
@ -87,7 +85,7 @@ spec:
|
||||
type: RuntimeDefault
|
||||
containers:
|
||||
- name: soak
|
||||
image: registry.threesix.ai/tidal/stress:m12-rc7-seedretry
|
||||
image: registry.threesix.ai/tidal/stress:m12-fleet-remediation-20260815-r2@sha256:86bfa95f63ba01ed4afa120e41684cd8694743ec39aca72ef0bd68e5d9ef7f8b
|
||||
imagePullPolicy: IfNotPresent
|
||||
# entrypoint is tidal-stress; wrap it so we can stamp the result file
|
||||
# name with the date and append a ledger line regardless of verdict.
|
||||
@ -96,8 +94,16 @@ spec:
|
||||
- |
|
||||
set -u
|
||||
DATE="$(date -u +%Y-%m-%d)"
|
||||
START_TS="$(date -u +%FT%TZ)"
|
||||
OUT="/results/soak-$DATE.json"
|
||||
echo "soak $DATE start $(date -u +%H:%M:%SZ) target=cluster image=m12-rc7-seedretry-200rps"
|
||||
OUT_TMP="/results/.soak-$DATE-$$.json"
|
||||
TAB="$(printf '\t')"
|
||||
if [ -f /results/ledger.tsv ] && grep -q "^${DATE}${TAB}" /results/ledger.tsv; then
|
||||
echo "FATAL: ledger already contains $DATE; one calendar night cannot count twice" >&2
|
||||
exit 124
|
||||
fi
|
||||
rm -f "$OUT_TMP"
|
||||
echo "soak $DATE start $START_TS target=cluster image=m12-fleet-remediation-200rps"
|
||||
tidal-stress \
|
||||
--target https://tidaldb-0.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500 \
|
||||
--target https://tidaldb-1.tidaldb-peers.tidaldb-cluster.svc.cluster.local:9500 \
|
||||
@ -109,36 +115,115 @@ spec:
|
||||
--mix peach \
|
||||
--users 50000 \
|
||||
--ramp "200:3600" \
|
||||
--json-summary "$OUT" \
|
||||
--json-summary "$OUT_TMP" \
|
||||
--max-p99-ms 150 \
|
||||
--max-error-pct 1 \
|
||||
--fail-on-knee
|
||||
RC=$?
|
||||
# Extract the verdict + headline p99/error from the JSON summary
|
||||
# for a one-line ledger entry (grep, no jq dependency in the image).
|
||||
# Fields are the real summary.rs keys: top-level "passed" (bool),
|
||||
# per-stage "overall_p99_ms" (ms) and "error_rate" (FRACTION 0-1).
|
||||
# Worst stage = max p99 over the run; we report the LAST stage's
|
||||
# numbers (the sustained-rate stage, since the soak ramp is single-
|
||||
# stage 500:3600 anyway) and trust "passed" for the verdict.
|
||||
VERD="PASS"; [ "$RC" -ne 0 ] && VERD="FAIL"
|
||||
PASSED="$(grep -o '"passed"[: ]*[a-z]*' "$OUT" 2>/dev/null | head -1 | grep -o '[a-z]*$')"
|
||||
P99="$(grep -o '"overall_p99_ms"[: ]*[0-9.]*' "$OUT" 2>/dev/null | tail -1 | grep -o '[0-9.]*$')"
|
||||
ERRF="$(grep -o '"error_rate"[: ]*[0-9.]*' "$OUT" 2>/dev/null | tail -1 | grep -o '[0-9.]*$')"
|
||||
printf '%s\t%s\trc=%s\tpassed=%s\tp99_ms=%s\terr_frac=%s\timage=m12-rc7-seedretry-200rps\n' \
|
||||
"$DATE" "$VERD" "$RC" "${PASSED:-?}" "${P99:-?}" "${ERRF:-?}" >> /results/ledger.tsv
|
||||
echo "soak $DATE end $(date -u +%H:%M:%SZ) verdict=$VERD rc=$RC p99=${P99:-?} err=${ERRF:-?}"
|
||||
# Refresh the streak immediately, CONSULTING restarts.tsv: a night
|
||||
# whose Job passed but whose window carried a pod restart is recorded
|
||||
# NON-GREEN by soak-eval (the zero-under-load-restart half of the GA
|
||||
# bar). This makes the per-night verdict restart-aware without giving
|
||||
# the soak Job pod-read RBAC — the monitor records restarts; soak-eval
|
||||
# only reads the two files. The Job's OWN exit stays the SLO verdict.
|
||||
if command -v soak-eval >/dev/null 2>&1; then
|
||||
soak-eval --results-dir /results --target 30 || \
|
||||
echo "soak $DATE: streak NON-GREEN per soak-eval (see /results/streak.tsv)"
|
||||
LOAD_RC=$?
|
||||
END_TS="$(date -u +%FT%TZ)"
|
||||
|
||||
# Extract the sustained stage's headline values for the ledger.
|
||||
# The process exit remains authoritative, but missing/malformed
|
||||
# evidence is independently fatal: an empty summary can never
|
||||
# turn a completed process into a green night.
|
||||
PASSED="$(grep -o '"passed"[: ]*[a-z]*' "$OUT_TMP" 2>/dev/null | head -1 | grep -o '[a-z]*$')"
|
||||
P99="$(grep -o '"overall_p99_ms"[: ]*[0-9.]*' "$OUT_TMP" 2>/dev/null | tail -1 | grep -o '[0-9.]*$')"
|
||||
ERRF="$(grep -o '"error_rate"[: ]*[0-9.]*' "$OUT_TMP" 2>/dev/null | tail -1 | grep -o '[0-9.]*$')"
|
||||
SUMMARY_RC=0
|
||||
case "$PASSED" in true|false) ;; *) SUMMARY_RC=126 ;; esac
|
||||
printf '%s\n' "$P99" | grep -Eq '^[0-9]+([.][0-9]+)?$' || SUMMARY_RC=126
|
||||
printf '%s\n' "$ERRF" | grep -Eq '^[0-9]+([.][0-9]+)?$' || SUMMARY_RC=126
|
||||
if [ "$LOAD_RC" -eq 0 ] && [ "$PASSED" != "true" ]; then
|
||||
SUMMARY_RC=126
|
||||
fi
|
||||
exit "$RC"
|
||||
if [ "$SUMMARY_RC" -eq 0 ]; then
|
||||
if ! mv "$OUT_TMP" "$OUT"; then
|
||||
echo "FATAL: failed to publish $OUT atomically" >&2
|
||||
SUMMARY_RC=125
|
||||
fi
|
||||
else
|
||||
echo "FATAL: malformed or inconsistent JSON summary (passed=${PASSED:-?} p99=${P99:-?} err=${ERRF:-?})" >&2
|
||||
mv "$OUT_TMP" "/results/soak-$DATE-invalid-$$.json" 2>/dev/null || true
|
||||
fi
|
||||
|
||||
RC="$LOAD_RC"
|
||||
if [ "$RC" -eq 0 ] && [ "$SUMMARY_RC" -ne 0 ]; then
|
||||
RC="$SUMMARY_RC"
|
||||
fi
|
||||
SOAK_VERD="PASS"; [ "$RC" -ne 0 ] && SOAK_VERD="FAIL"
|
||||
# Stage the candidate ledger separately. Nothing reads this
|
||||
# path except the one-shot evaluator, so a killed Job cannot
|
||||
# leave a provisional PASS in the durable public ledger.
|
||||
PENDING_LEDGER_NAME=".ledger-$DATE-pending.tsv"
|
||||
PENDING_LEDGER="/results/$PENDING_LEDGER_NAME"
|
||||
if [ -f /results/ledger.tsv ]; then
|
||||
cp /results/ledger.tsv "$PENDING_LEDGER" || {
|
||||
echo "FATAL: failed to stage existing /results/ledger.tsv" >&2
|
||||
exit 125
|
||||
}
|
||||
: > "$PENDING_LEDGER" || {
|
||||
echo "FATAL: failed to create candidate ledger" >&2
|
||||
exit 125
|
||||
}
|
||||
fi
|
||||
if ! printf '%s\t%s\tsoak_rc=%s\tsummary_rc=%s\tgate_rc=pending\tpassed=%s\tp99_ms=%s\terr_frac=%s\timage=m12-fleet-remediation-200rps\tstart_utc=%s\tend_utc=%s\n' \
|
||||
"$DATE" "$SOAK_VERD" "$RC" "$SUMMARY_RC" "${PASSED:-?}" "${P99:-?}" "${ERRF:-?}" "$START_TS" "$END_TS" \
|
||||
>> "$PENDING_LEDGER"; then
|
||||
echo "FATAL: failed to stage candidate ledger" >&2
|
||||
rm -f "$PENDING_LEDGER"
|
||||
exit 125
|
||||
fi
|
||||
echo "soak $DATE load end $END_TS verdict=$SOAK_VERD load_rc=$LOAD_RC summary_rc=$SUMMARY_RC p99=${P99:-?} err=${ERRF:-?}"
|
||||
|
||||
# Fail closed on the other half of the GA contract. The
|
||||
# monitor samples once per minute; wait one full interval,
|
||||
# then require fresh evidence bracketing both load boundaries
|
||||
# for every expected cluster pod.
|
||||
FINAL_RC="$RC"
|
||||
if ! command -v soak-eval >/dev/null 2>&1; then
|
||||
echo "FATAL: soak-eval missing from stress image" >&2
|
||||
EVAL_RC=127
|
||||
else
|
||||
sleep 70
|
||||
soak-eval \
|
||||
--ledger-file "$PENDING_LEDGER_NAME" \
|
||||
--no-write-streak \
|
||||
--results-dir /results \
|
||||
--target 30 \
|
||||
--require-date "$DATE" \
|
||||
--require-samples-through "$END_TS" \
|
||||
--expected-pods 3
|
||||
EVAL_RC=$?
|
||||
fi
|
||||
FINAL_VERD="$SOAK_VERD"
|
||||
if [ "$EVAL_RC" -ne 0 ]; then
|
||||
echo "soak $DATE: recovery-evidence gate failed rc=$EVAL_RC (see this Job log and /results/restarts.tsv)" >&2
|
||||
FINAL_RC="$EVAL_RC"
|
||||
FINAL_VERD="FAIL"
|
||||
fi
|
||||
|
||||
# Convert the private candidate's last row into the combined
|
||||
# load + recovery verdict. CronJob concurrency is Forbid, so
|
||||
# this is the only writer; rename publishes one atomic ledger.
|
||||
LEDGER_TMP="/results/.ledger-$DATE.tmp"
|
||||
if ! {
|
||||
sed '$d' "$PENDING_LEDGER"
|
||||
printf '%s\t%s\tsoak_rc=%s\tsummary_rc=%s\tgate_rc=%s\tpassed=%s\tp99_ms=%s\terr_frac=%s\timage=m12-fleet-remediation-200rps\tstart_utc=%s\tend_utc=%s\n' \
|
||||
"$DATE" "$FINAL_VERD" "$RC" "$SUMMARY_RC" "$EVAL_RC" "${PASSED:-?}" "${P99:-?}" "${ERRF:-?}" "$START_TS" "$END_TS"
|
||||
} > "$LEDGER_TMP" || ! mv "$LEDGER_TMP" /results/ledger.tsv; then
|
||||
echo "FATAL: failed to finalize /results/ledger.tsv" >&2
|
||||
rm -f "$LEDGER_TMP" "$PENDING_LEDGER"
|
||||
exit 125
|
||||
fi
|
||||
rm -f "$PENDING_LEDGER"
|
||||
|
||||
# Recompute streak.tsv from the finalized ledger. A non-zero
|
||||
# result is expected on a failed night and must not mask FINAL_RC.
|
||||
if command -v soak-eval >/dev/null 2>&1; then
|
||||
soak-eval --results-dir /results --target 30 >/dev/null 2>&1 || true
|
||||
fi
|
||||
echo "soak $DATE final verdict=$FINAL_VERD load_rc=$LOAD_RC summary_rc=$SUMMARY_RC gate_rc=$EVAL_RC"
|
||||
exit "$FINAL_RC"
|
||||
env:
|
||||
- name: TIDAL_API_KEY
|
||||
valueFrom:
|
||||
|
||||
@ -23,6 +23,8 @@ metadata:
|
||||
labels:
|
||||
app.kubernetes.io/name: tidal-soak
|
||||
app.kubernetes.io/part-of: tidaldb
|
||||
backup.orchard9.ai/class: protected
|
||||
backup.orchard9.ai/method: velero-kopia
|
||||
spec:
|
||||
accessModes:
|
||||
- ReadWriteMany
|
||||
|
||||
@ -27,31 +27,118 @@ struct Cli {
|
||||
#[arg(long, default_value = "/results")]
|
||||
results_dir: PathBuf,
|
||||
|
||||
/// Ledger filename within the result dir. The nightly Job points this at an
|
||||
/// unpublished candidate so a killed Job cannot expose a provisional PASS.
|
||||
#[arg(long, default_value = "ledger.tsv")]
|
||||
ledger_file: PathBuf,
|
||||
|
||||
/// Evaluate and return the verdict without replacing `streak.tsv`.
|
||||
#[arg(long)]
|
||||
no_write_streak: bool,
|
||||
|
||||
/// Consecutive green nights required for GA (the exit gate).
|
||||
#[arg(long, default_value_t = 30)]
|
||||
target: usize,
|
||||
|
||||
/// Require the newest ledger row to be this date. The nightly Job passes
|
||||
/// its own UTC date so missing/stale inputs are fatal rather than streak 0.
|
||||
#[arg(long)]
|
||||
require_date: Option<String>,
|
||||
|
||||
/// Require the newest ledger row to end at this ISO-8601 timestamp. Restart
|
||||
/// evaluation then requires fresh per-pod samples immediately before the
|
||||
/// row's start and after this end bound.
|
||||
#[arg(long, requires = "require_date")]
|
||||
require_samples_through: Option<String>,
|
||||
|
||||
/// Number of cluster pods whose restart evidence must be present.
|
||||
#[arg(long, default_value_t = 3)]
|
||||
expected_pods: usize,
|
||||
}
|
||||
|
||||
fn main() -> ExitCode {
|
||||
let cli = Cli::parse();
|
||||
let ledger_path = cli.results_dir.join("ledger.tsv");
|
||||
if cli.target == 0 || cli.expected_pods == 0 {
|
||||
eprintln!("soak-eval: --target and --expected-pods must both be greater than zero");
|
||||
return ExitCode::from(2);
|
||||
}
|
||||
let ledger_path = cli.results_dir.join(&cli.ledger_file);
|
||||
let restarts_path = cli.results_dir.join("restarts.tsv");
|
||||
|
||||
// A missing ledger means no night has run yet — streak 0, not an error.
|
||||
let ledger = std::fs::read_to_string(&ledger_path).unwrap_or_default();
|
||||
let restarts = std::fs::read_to_string(&restarts_path).unwrap_or_default();
|
||||
let read_input = |path: &std::path::Path| match std::fs::read_to_string(path) {
|
||||
Ok(input) => Ok(input),
|
||||
Err(error)
|
||||
if error.kind() == std::io::ErrorKind::NotFound && cli.require_date.is_none() =>
|
||||
{
|
||||
Ok(String::new())
|
||||
}
|
||||
Err(error) => Err(error),
|
||||
};
|
||||
let ledger = match read_input(&ledger_path) {
|
||||
Ok(input) => input,
|
||||
Err(error) => {
|
||||
eprintln!(
|
||||
"soak-eval: failed to read {}: {error}",
|
||||
ledger_path.display()
|
||||
);
|
||||
return ExitCode::from(2);
|
||||
}
|
||||
};
|
||||
let restarts = match read_input(&restarts_path) {
|
||||
Ok(input) => input,
|
||||
Err(error) => {
|
||||
eprintln!(
|
||||
"soak-eval: failed to read {}: {error}",
|
||||
restarts_path.display()
|
||||
);
|
||||
return ExitCode::from(2);
|
||||
}
|
||||
};
|
||||
|
||||
let nights = parse_ledger(&ledger);
|
||||
if let Some(required) = cli.require_date.as_deref()
|
||||
&& nights.last().map(|night| night.date.as_str()) != Some(required)
|
||||
{
|
||||
eprintln!(
|
||||
"soak-eval: newest ledger date is {:?}, expected {required}",
|
||||
nights.last().map(|night| night.date.as_str())
|
||||
);
|
||||
return ExitCode::from(2);
|
||||
}
|
||||
let samples = parse_restarts(&restarts);
|
||||
let verdicts = evaluate(&nights, &samples);
|
||||
if let (Some(required_date), Some(required_timestamp)) = (
|
||||
cli.require_date.as_deref(),
|
||||
cli.require_samples_through.as_deref(),
|
||||
) && nights
|
||||
.last()
|
||||
.and_then(|night| night.end_timestamp.as_deref())
|
||||
!= Some(required_timestamp)
|
||||
{
|
||||
eprintln!(
|
||||
"soak-eval: ledger end timestamp for {required_date} is {:?}, expected {required_timestamp}",
|
||||
nights
|
||||
.last()
|
||||
.and_then(|night| night.end_timestamp.as_deref())
|
||||
);
|
||||
return ExitCode::from(2);
|
||||
}
|
||||
let verdicts = evaluate(&nights, &samples, cli.expected_pods);
|
||||
let streak = current_streak(&verdicts);
|
||||
|
||||
if !cli.no_write_streak {
|
||||
let tsv = render_streak_tsv(&verdicts, cli.target);
|
||||
let out_path = cli.results_dir.join("streak.tsv");
|
||||
if let Err(e) = std::fs::write(&out_path, &tsv) {
|
||||
eprintln!("soak-eval: failed to write {}: {e}", out_path.display());
|
||||
let tmp_path = cli
|
||||
.results_dir
|
||||
.join(format!(".streak-{}.tmp", std::process::id()));
|
||||
if let Err(error) =
|
||||
std::fs::write(&tmp_path, &tsv).and_then(|()| std::fs::rename(&tmp_path, &out_path))
|
||||
{
|
||||
let _ = std::fs::remove_file(&tmp_path);
|
||||
eprintln!("soak-eval: failed to write {}: {error}", out_path.display());
|
||||
return ExitCode::from(2);
|
||||
}
|
||||
}
|
||||
|
||||
println!(
|
||||
"soak-eval: {streak}/{} consecutive green nights ({} nights evaluated)",
|
||||
@ -59,15 +146,14 @@ fn main() -> ExitCode {
|
||||
verdicts.len()
|
||||
);
|
||||
if streak >= cli.target {
|
||||
println!("soak-eval: GA STREAK MET ({streak} ≥ {})", cli.target);
|
||||
println!("soak-eval: GA STREAK MET ({streak} >= {})", cli.target);
|
||||
}
|
||||
|
||||
// Exit non-zero iff the MOST RECENT night is non-green — the alert trigger.
|
||||
match verdicts.last() {
|
||||
Some(v) if !v.green => {
|
||||
Some(verdict) if !verdict.green => {
|
||||
eprintln!(
|
||||
"soak-eval: ALERT — last night {} is NON-GREEN: {}",
|
||||
v.date, v.reason
|
||||
verdict.date, verdict.reason
|
||||
);
|
||||
ExitCode::from(1)
|
||||
}
|
||||
|
||||
304
tidal-stress/src/bin/soak-watch.rs
Normal file
304
tidal-stress/src/bin/soak-watch.rs
Normal file
@ -0,0 +1,304 @@
|
||||
//! `soak-watch` records Kubernetes pod generations and restart evidence.
|
||||
//!
|
||||
//! It talks directly to the in-cluster Kubernetes API with the pod's service
|
||||
//! account. This keeps the durable soak evidence independent of a mutable
|
||||
//! `kubectl` utility image.
|
||||
|
||||
use std::{
|
||||
error::Error,
|
||||
fs::OpenOptions,
|
||||
io::{self, Write},
|
||||
path::{Path, PathBuf},
|
||||
process::ExitCode,
|
||||
time::Duration,
|
||||
};
|
||||
|
||||
use clap::Parser;
|
||||
use serde::Deserialize;
|
||||
use time::{OffsetDateTime, format_description::well_known::Rfc3339};
|
||||
|
||||
#[derive(Parser)]
|
||||
#[command(version, about = "Record Kubernetes pod restart evidence as TSV")]
|
||||
struct Cli {
|
||||
#[arg(long, default_value = "tidaldb-cluster")]
|
||||
namespace: String,
|
||||
|
||||
#[arg(long, default_value = "app.kubernetes.io/name=tidaldb")]
|
||||
label_selector: String,
|
||||
|
||||
#[arg(long, default_value = "tidaldb")]
|
||||
container: String,
|
||||
|
||||
#[arg(long, default_value = "/results/restarts.tsv")]
|
||||
results_file: PathBuf,
|
||||
|
||||
#[arg(long, default_value_t = 60)]
|
||||
interval_seconds: u64,
|
||||
}
|
||||
|
||||
#[derive(Deserialize)]
|
||||
struct PodList {
|
||||
items: Vec<Pod>,
|
||||
}
|
||||
|
||||
#[derive(Deserialize)]
|
||||
struct Pod {
|
||||
metadata: Metadata,
|
||||
#[serde(default)]
|
||||
status: PodStatus,
|
||||
}
|
||||
|
||||
#[derive(Deserialize)]
|
||||
struct Metadata {
|
||||
name: String,
|
||||
uid: String,
|
||||
}
|
||||
|
||||
#[derive(Default, Deserialize)]
|
||||
#[serde(rename_all = "camelCase")]
|
||||
struct PodStatus {
|
||||
#[serde(default)]
|
||||
phase: String,
|
||||
#[serde(default)]
|
||||
container_statuses: Vec<ContainerStatus>,
|
||||
}
|
||||
|
||||
#[derive(Deserialize)]
|
||||
#[serde(rename_all = "camelCase")]
|
||||
struct ContainerStatus {
|
||||
name: String,
|
||||
restart_count: u64,
|
||||
ready: bool,
|
||||
#[serde(default)]
|
||||
last_state: ContainerState,
|
||||
}
|
||||
|
||||
#[derive(Default, Deserialize)]
|
||||
struct ContainerState {
|
||||
terminated: Option<TerminatedState>,
|
||||
}
|
||||
|
||||
#[derive(Deserialize)]
|
||||
#[serde(rename_all = "camelCase")]
|
||||
struct TerminatedState {
|
||||
#[serde(default)]
|
||||
reason: String,
|
||||
exit_code: i32,
|
||||
#[serde(default)]
|
||||
message: String,
|
||||
}
|
||||
|
||||
struct KubernetesClient {
|
||||
client: reqwest::Client,
|
||||
pods_url: String,
|
||||
token: String,
|
||||
}
|
||||
|
||||
impl KubernetesClient {
|
||||
fn in_cluster(namespace: &str) -> Result<Self, Box<dyn Error>> {
|
||||
let host = std::env::var("KUBERNETES_SERVICE_HOST").map_err(|error| {
|
||||
io::Error::new(
|
||||
io::ErrorKind::NotFound,
|
||||
format!("KUBERNETES_SERVICE_HOST is unavailable: {error}"),
|
||||
)
|
||||
})?;
|
||||
let port =
|
||||
std::env::var("KUBERNETES_SERVICE_PORT_HTTPS").unwrap_or_else(|_| String::from("443"));
|
||||
let service_account = Path::new("/var/run/secrets/kubernetes.io/serviceaccount");
|
||||
let token = std::fs::read_to_string(service_account.join("token"))?
|
||||
.trim()
|
||||
.to_owned();
|
||||
let ca = std::fs::read(service_account.join("ca.crt"))?;
|
||||
let certificate = reqwest::Certificate::from_pem(&ca)?;
|
||||
let client = reqwest::Client::builder()
|
||||
.add_root_certificate(certificate)
|
||||
.https_only(true)
|
||||
.timeout(Duration::from_secs(15))
|
||||
.build()?;
|
||||
|
||||
Ok(Self {
|
||||
client,
|
||||
pods_url: format!("https://{host}:{port}/api/v1/namespaces/{namespace}/pods"),
|
||||
token,
|
||||
})
|
||||
}
|
||||
|
||||
async fn pods(&self, label_selector: &str) -> Result<PodList, reqwest::Error> {
|
||||
self.client
|
||||
.get(&self.pods_url)
|
||||
.bearer_auth(&self.token)
|
||||
.query(&[("labelSelector", label_selector)])
|
||||
.send()
|
||||
.await?
|
||||
.error_for_status()?
|
||||
.json()
|
||||
.await
|
||||
}
|
||||
}
|
||||
|
||||
fn timestamp() -> Result<String, time::error::Format> {
|
||||
OffsetDateTime::now_utc().format(&Rfc3339)
|
||||
}
|
||||
|
||||
fn sanitize_tsv(value: &str) -> String {
|
||||
value
|
||||
.chars()
|
||||
.map(|character| match character {
|
||||
'\t' | '\n' | '\r' => ' ',
|
||||
other => other,
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
fn append_sample(
|
||||
path: &Path,
|
||||
sampled_at: &str,
|
||||
pods: &PodList,
|
||||
container_name: &str,
|
||||
) -> io::Result<usize> {
|
||||
let write_header = std::fs::metadata(path).map_or(true, |metadata| metadata.len() == 0);
|
||||
let mut output = OpenOptions::new().create(true).append(true).open(path)?;
|
||||
if write_header {
|
||||
writeln!(
|
||||
output,
|
||||
"ts_utc\tpod\tuid\trestarts\tphase\tready\tlast_reason\tlast_exit\tlast_message"
|
||||
)?;
|
||||
}
|
||||
|
||||
let mut rows = 0;
|
||||
for pod in &pods.items {
|
||||
let Some(container) = pod
|
||||
.status
|
||||
.container_statuses
|
||||
.iter()
|
||||
.find(|status| status.name == container_name)
|
||||
else {
|
||||
continue;
|
||||
};
|
||||
let (reason, exit_code, message) = container.last_state.terminated.as_ref().map_or(
|
||||
("", String::new(), ""),
|
||||
|terminated| {
|
||||
(
|
||||
terminated.reason.as_str(),
|
||||
terminated.exit_code.to_string(),
|
||||
terminated.message.as_str(),
|
||||
)
|
||||
},
|
||||
);
|
||||
writeln!(
|
||||
output,
|
||||
"{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}",
|
||||
sampled_at,
|
||||
sanitize_tsv(&pod.metadata.name),
|
||||
sanitize_tsv(&pod.metadata.uid),
|
||||
container.restart_count,
|
||||
sanitize_tsv(&pod.status.phase),
|
||||
container.ready,
|
||||
sanitize_tsv(reason),
|
||||
exit_code,
|
||||
sanitize_tsv(message),
|
||||
)?;
|
||||
rows += 1;
|
||||
}
|
||||
output.flush()?;
|
||||
Ok(rows)
|
||||
}
|
||||
|
||||
async fn run() -> Result<(), Box<dyn Error>> {
|
||||
let cli = Cli::parse();
|
||||
if cli.interval_seconds == 0 {
|
||||
return Err(io::Error::new(
|
||||
io::ErrorKind::InvalidInput,
|
||||
"--interval-seconds must be greater than zero",
|
||||
)
|
||||
.into());
|
||||
}
|
||||
let kubernetes = KubernetesClient::in_cluster(&cli.namespace)?;
|
||||
println!(
|
||||
"soak-watch started: namespace={} selector={} container={} interval={}s output={}",
|
||||
cli.namespace,
|
||||
cli.label_selector,
|
||||
cli.container,
|
||||
cli.interval_seconds,
|
||||
cli.results_file.display()
|
||||
);
|
||||
|
||||
loop {
|
||||
let sampled_at = timestamp()?;
|
||||
match kubernetes.pods(&cli.label_selector).await {
|
||||
Ok(pods) => {
|
||||
match append_sample(&cli.results_file, &sampled_at, &pods, &cli.container) {
|
||||
Ok(rows) => println!("soak-watch sample: ts={sampled_at} rows={rows}"),
|
||||
Err(error) => eprintln!(
|
||||
"soak-watch write failed: ts={sampled_at} path={} error={error}",
|
||||
cli.results_file.display()
|
||||
),
|
||||
}
|
||||
}
|
||||
Err(error) => {
|
||||
eprintln!("soak-watch Kubernetes sample failed: ts={sampled_at} error={error}");
|
||||
}
|
||||
}
|
||||
tokio::time::sleep(Duration::from_secs(cli.interval_seconds)).await;
|
||||
}
|
||||
}
|
||||
|
||||
#[tokio::main]
|
||||
async fn main() -> ExitCode {
|
||||
match run().await {
|
||||
Ok(()) => ExitCode::SUCCESS,
|
||||
Err(error) => {
|
||||
eprintln!("soak-watch fatal: {error}");
|
||||
ExitCode::from(1)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
#[allow(clippy::unwrap_used)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
#[test]
|
||||
fn append_sample_records_uid_and_termination_message_without_multiline_tsv() {
|
||||
let path = std::env::temp_dir().join(format!(
|
||||
"soak-watch-{}-{}.tsv",
|
||||
std::process::id(),
|
||||
OffsetDateTime::now_utc().unix_timestamp_nanos()
|
||||
));
|
||||
let pods = PodList {
|
||||
items: vec![Pod {
|
||||
metadata: Metadata {
|
||||
name: String::from("tidaldb-0"),
|
||||
uid: String::from("uid-1"),
|
||||
},
|
||||
status: PodStatus {
|
||||
phase: String::from("Running"),
|
||||
container_statuses: vec![ContainerStatus {
|
||||
name: String::from("tidaldb"),
|
||||
restart_count: 4,
|
||||
ready: true,
|
||||
last_state: ContainerState {
|
||||
terminated: Some(TerminatedState {
|
||||
reason: String::from("Completed"),
|
||||
exit_code: 0,
|
||||
message: String::from("{\"reason\":\"reseed\"}\nnext"),
|
||||
}),
|
||||
},
|
||||
}],
|
||||
},
|
||||
}],
|
||||
};
|
||||
|
||||
assert_eq!(
|
||||
append_sample(&path, "2026-08-13T02:00:00Z", &pods, "tidaldb").unwrap(),
|
||||
1
|
||||
);
|
||||
let output = std::fs::read_to_string(&path).unwrap();
|
||||
std::fs::remove_file(path).unwrap();
|
||||
|
||||
assert_eq!(output.lines().count(), 2);
|
||||
assert!(output.contains("tidaldb-0\tuid-1\t4\tRunning\ttrue\tCompleted\t0"));
|
||||
assert!(output.contains("{\"reason\":\"reseed\"} next"));
|
||||
}
|
||||
}
|
||||
@ -18,15 +18,20 @@
|
||||
//! The logic is pure and table-driven-tested so the k8s evaluator container is a
|
||||
//! thin shell around a proven core, not fragile inline `awk`.
|
||||
|
||||
use std::collections::BTreeMap;
|
||||
use std::collections::BTreeSet;
|
||||
|
||||
/// One night's ledger row: the date (`YYYY-MM-DD`) and whether the soak Job
|
||||
/// passed its armed SLO gates.
|
||||
use time::{Date, Duration, Month, OffsetDateTime, format_description::well_known::Rfc3339};
|
||||
|
||||
/// One night's ledger row: date, load window, and armed-gate result.
|
||||
///
|
||||
/// A trustworthy recovery verdict requires monitor samples immediately
|
||||
/// outside both window bounds.
|
||||
#[derive(Debug, Clone, PartialEq, Eq)]
|
||||
pub struct NightResult {
|
||||
pub date: String,
|
||||
pub passed: bool,
|
||||
pub start_timestamp: Option<String>,
|
||||
pub end_timestamp: Option<String>,
|
||||
}
|
||||
|
||||
/// One restart-watch sample: the calendar date of the sample, the pod, and its
|
||||
@ -34,8 +39,15 @@ pub struct NightResult {
|
||||
/// rises until the pod object is recreated).
|
||||
#[derive(Debug, Clone, PartialEq, Eq)]
|
||||
pub struct RestartSample {
|
||||
/// Full ISO-8601 sample timestamp. `restarts.tsv` is append-only, so this
|
||||
/// also documents the ordering contract used by restart detection.
|
||||
pub timestamp: String,
|
||||
pub date: String,
|
||||
pub pod: String,
|
||||
/// Kubernetes pod UID when the monitor emits the current schema. Legacy
|
||||
/// rows omit it; those rows still detect counter increases but cannot prove
|
||||
/// a restart that happened before the first sample of a recreated pod.
|
||||
pub pod_uid: Option<String>,
|
||||
pub restarts: u64,
|
||||
}
|
||||
|
||||
@ -50,10 +62,10 @@ pub struct NightVerdict {
|
||||
|
||||
/// Parse `ledger.tsv` (the nightly CronJob's append-only verdict log).
|
||||
///
|
||||
/// Columns (tab-separated, as the cronjob writes them):
|
||||
/// `DATE VERD rc=N passed=X p99_ms=X err_frac=X image=...`. The verdict is
|
||||
/// taken from the `VERD` column (`PASS`/`FAIL`) — the authoritative Job exit
|
||||
/// translation. Blank lines and a leading header (if any) are skipped.
|
||||
/// Columns (tab-separated, as the CronJob writes them):
|
||||
/// `DATE VERD ... start_utc=TS end_utc=TS`. The verdict is taken from the
|
||||
/// `VERD` column (`PASS`/`FAIL`) and the timestamps delimit the load window.
|
||||
/// Blank lines and a leading header (if any) are skipped.
|
||||
#[must_use]
|
||||
pub fn parse_ledger(tsv: &str) -> Vec<NightResult> {
|
||||
let mut out = Vec::new();
|
||||
@ -67,24 +79,31 @@ pub fn parse_ledger(tsv: &str) -> Vec<NightResult> {
|
||||
continue;
|
||||
}
|
||||
let date = cols[0].trim();
|
||||
// A header row or anything whose first cell is not a date is skipped.
|
||||
// A header row or anything whose first cell is not a real date is skipped.
|
||||
if !is_date(date) {
|
||||
continue;
|
||||
}
|
||||
let verd = cols[1].trim();
|
||||
let field = |name: &str| {
|
||||
cols.iter()
|
||||
.skip(2)
|
||||
.find_map(|column| column.trim().strip_prefix(name))
|
||||
.map(str::to_string)
|
||||
};
|
||||
out.push(NightResult {
|
||||
date: date.to_string(),
|
||||
passed: verd.eq_ignore_ascii_case("PASS"),
|
||||
passed: cols[1].trim().eq_ignore_ascii_case("PASS"),
|
||||
start_timestamp: field("start_utc="),
|
||||
end_timestamp: field("end_utc="),
|
||||
});
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
/// Parse `restarts.tsv` (the monitor's 5-minute restart-count snapshots).
|
||||
/// Parse `restarts.tsv` (the monitor's restart-count snapshots).
|
||||
///
|
||||
/// Columns: `ts_utc pod restarts phase ready` with an ISO `ts_utc` whose
|
||||
/// leading `YYYY-MM-DD` is the night key. The header row and malformed rows are
|
||||
/// skipped.
|
||||
/// Legacy columns are `ts_utc pod restarts phase ready`; current rows insert
|
||||
/// the Kubernetes pod UID before `restarts`. Invalid timestamps and malformed
|
||||
/// rows are skipped.
|
||||
#[must_use]
|
||||
pub fn parse_restarts(tsv: &str) -> Vec<RestartSample> {
|
||||
let mut out = Vec::new();
|
||||
@ -97,263 +116,511 @@ pub fn parse_restarts(tsv: &str) -> Vec<RestartSample> {
|
||||
if cols.len() < 3 {
|
||||
continue;
|
||||
}
|
||||
let date = cols[0].get(0..10).unwrap_or("");
|
||||
if !is_date(date) {
|
||||
let timestamp = cols[0].trim();
|
||||
let date = timestamp.get(0..10).unwrap_or("");
|
||||
if !is_date(date) || parse_utc_timestamp(timestamp).is_none() {
|
||||
continue; // header ("ts_utc") or junk
|
||||
}
|
||||
let pod = cols[1].trim();
|
||||
let Ok(restarts) = cols[2].trim().parse::<u64>() else {
|
||||
// A recreated pod starts a fresh counter generation. Legacy rows lack
|
||||
// the UID but remain useful because any bounded-window counter change
|
||||
// is still a recovery event.
|
||||
let (pod_uid, restarts) = if let Ok(restarts) = cols[2].trim().parse::<u64>() {
|
||||
(None, restarts)
|
||||
} else {
|
||||
let Some(restarts) = cols.get(3).and_then(|v| v.trim().parse::<u64>().ok()) else {
|
||||
continue;
|
||||
};
|
||||
let uid = cols[2].trim();
|
||||
let pod_uid = if uid.is_empty() || uid == "<none>" {
|
||||
None
|
||||
} else {
|
||||
Some(uid.to_string())
|
||||
};
|
||||
(pod_uid, restarts)
|
||||
};
|
||||
out.push(RestartSample {
|
||||
timestamp: timestamp.to_string(),
|
||||
date: date.to_string(),
|
||||
pod: pod.to_string(),
|
||||
pod_uid,
|
||||
restarts,
|
||||
});
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
/// `YYYY-MM-DD` shape check (cheap, no chrono dependency).
|
||||
fn is_date(s: &str) -> bool {
|
||||
let b = s.as_bytes();
|
||||
b.len() == 10
|
||||
&& b[4] == b'-'
|
||||
&& b[7] == b'-'
|
||||
&& b[0].is_ascii_digit()
|
||||
&& b[1].is_ascii_digit()
|
||||
&& b[2].is_ascii_digit()
|
||||
&& b[3].is_ascii_digit()
|
||||
&& b[5].is_ascii_digit()
|
||||
&& b[6].is_ascii_digit()
|
||||
&& b[8].is_ascii_digit()
|
||||
&& b[9].is_ascii_digit()
|
||||
fn parse_date(value: &str) -> Option<Date> {
|
||||
let bytes = value.as_bytes();
|
||||
if bytes.len() != 10 || bytes[4] != b'-' || bytes[7] != b'-' {
|
||||
return None;
|
||||
}
|
||||
let year = value.get(0..4)?.parse::<i32>().ok()?;
|
||||
let month = Month::try_from(value.get(5..7)?.parse::<u8>().ok()?).ok()?;
|
||||
let day = value.get(8..10)?.parse::<u8>().ok()?;
|
||||
Date::from_calendar_date(year, month, day).ok()
|
||||
}
|
||||
|
||||
/// Whether any pod restarted within night `date`, given the per-pod cumulative
|
||||
/// restart counts observed up to and including each night.
|
||||
fn is_date(value: &str) -> bool {
|
||||
parse_date(value).is_some()
|
||||
}
|
||||
|
||||
fn parse_utc_timestamp(value: &str) -> Option<OffsetDateTime> {
|
||||
value
|
||||
.ends_with('Z')
|
||||
.then(|| OffsetDateTime::parse(value, &Rfc3339).ok())
|
||||
.flatten()
|
||||
}
|
||||
|
||||
fn window_timestamp(value: Option<&str>, date: &str) -> Option<OffsetDateTime> {
|
||||
let timestamp = parse_utc_timestamp(value?)?;
|
||||
(timestamp.date() == parse_date(date)?).then_some(timestamp)
|
||||
}
|
||||
|
||||
#[derive(Debug, Default)]
|
||||
struct RestartEvidence {
|
||||
dirty: bool,
|
||||
complete: bool,
|
||||
}
|
||||
|
||||
/// Maximum tolerated distance between a load boundary and its monitor sample.
|
||||
/// The watcher runs every minute and the Job waits 70 seconds after load, so
|
||||
/// two minutes admits ordinary scheduling jitter without accepting a stale
|
||||
/// watcher as evidence.
|
||||
const MAX_BOUNDARY_SKEW: Duration = Duration::seconds(120);
|
||||
|
||||
/// Classify restart evidence for exactly one measured load window.
|
||||
///
|
||||
/// A restart shows as the cumulative count RISING. For night D and pod P:
|
||||
/// `restarted = max_count(P, D) > floor`, where `floor` is P's highest count on
|
||||
/// any night strictly before D (its carried-forward baseline), or — for the
|
||||
/// first night P is seen — the MINIMUM count that night (so a restart *within*
|
||||
/// the very first observed night, e.g. 0 → 1, still registers).
|
||||
fn nights_with_restart(samples: &[RestartSample]) -> BTreeSet<String> {
|
||||
// pod -> (date -> (min, max)) of that date's samples.
|
||||
let mut by_pod: BTreeMap<&str, BTreeMap<&str, (u64, u64)>> = BTreeMap::new();
|
||||
for s in samples {
|
||||
let e = by_pod
|
||||
.entry(&s.pod)
|
||||
.or_default()
|
||||
.entry(&s.date)
|
||||
.or_insert((s.restarts, s.restarts));
|
||||
e.0 = e.0.min(s.restarts);
|
||||
e.1 = e.1.max(s.restarts);
|
||||
/// Every expected pod needs a sample no more than two minutes before the load
|
||||
/// starts and another no more than two minutes after it ends. Counter changes
|
||||
/// or pod-UID changes between those bracketing samples make the night dirty.
|
||||
/// This avoids both prior false positives from unrelated daytime restarts and
|
||||
/// false greens from two arbitrary samples somewhere on the same calendar day.
|
||||
fn restart_evidence(
|
||||
night: &NightResult,
|
||||
samples: &[RestartSample],
|
||||
expected_pods: usize,
|
||||
) -> RestartEvidence {
|
||||
let (Some(start), Some(end)) = (
|
||||
window_timestamp(night.start_timestamp.as_deref(), &night.date),
|
||||
window_timestamp(night.end_timestamp.as_deref(), &night.date),
|
||||
) else {
|
||||
return RestartEvidence::default();
|
||||
};
|
||||
if end <= start {
|
||||
return RestartEvidence::default();
|
||||
}
|
||||
let mut dirty = BTreeSet::new();
|
||||
for dates in by_pod.values() {
|
||||
let mut prev_max: Option<u64> = None;
|
||||
for (date, &(min, max)) in dates {
|
||||
let floor = prev_max.unwrap_or(min);
|
||||
if max > floor {
|
||||
dirty.insert((*date).to_string());
|
||||
|
||||
let mut dated_samples: Vec<_> = samples
|
||||
.iter()
|
||||
.filter(|sample| sample.date == night.date)
|
||||
.filter_map(|sample| parse_utc_timestamp(&sample.timestamp).map(|at| (sample, at)))
|
||||
.collect();
|
||||
dated_samples.sort_by_key(|(_, sampled_at)| *sampled_at);
|
||||
|
||||
let pods: BTreeSet<&str> = dated_samples
|
||||
.iter()
|
||||
.map(|(sample, _)| sample.pod.as_str())
|
||||
.collect();
|
||||
let mut evidence = RestartEvidence {
|
||||
dirty: false,
|
||||
complete: expected_pods > 0 && pods.len() == expected_pods,
|
||||
};
|
||||
|
||||
for pod in pods {
|
||||
let pod_samples: Vec<_> = dated_samples
|
||||
.iter()
|
||||
.copied()
|
||||
.filter(|(sample, _)| sample.pod == pod)
|
||||
.collect();
|
||||
let before = pod_samples
|
||||
.iter()
|
||||
.rposition(|(_, sampled_at)| *sampled_at <= start);
|
||||
let after = pod_samples
|
||||
.iter()
|
||||
.position(|(_, sampled_at)| *sampled_at >= end);
|
||||
let (Some(before), Some(after)) = (before, after) else {
|
||||
evidence.complete = false;
|
||||
continue;
|
||||
};
|
||||
let before_at = pod_samples[before].1;
|
||||
let after_at = pod_samples[after].1;
|
||||
if start - before_at > MAX_BOUNDARY_SKEW || after_at - end > MAX_BOUNDARY_SKEW {
|
||||
evidence.complete = false;
|
||||
continue;
|
||||
}
|
||||
prev_max = Some(prev_max.map_or(max, |p| p.max(max)));
|
||||
|
||||
for pair in pod_samples[before..=after].windows(2) {
|
||||
let previous = pair[0].0;
|
||||
let current = pair[1].0;
|
||||
let uid_changed = matches!(
|
||||
(previous.pod_uid.as_deref(), current.pod_uid.as_deref()),
|
||||
(Some(previous_uid), Some(current_uid)) if previous_uid != current_uid
|
||||
);
|
||||
if uid_changed || previous.restarts != current.restarts {
|
||||
evidence.dirty = true;
|
||||
}
|
||||
}
|
||||
dirty
|
||||
}
|
||||
|
||||
evidence
|
||||
}
|
||||
|
||||
/// Join the ledger and the restart samples into a per-night verdict list, in
|
||||
/// ledger order. A night is green iff it passed AND no pod restarted in its
|
||||
/// window.
|
||||
/// Join the ledger and restart samples into per-night verdicts, in ledger order.
|
||||
///
|
||||
/// A night is green only when the load passed and complete bracketing evidence
|
||||
/// proves that every expected pod remained in the same generation and restart
|
||||
/// count throughout the measured window.
|
||||
#[must_use]
|
||||
pub fn evaluate(nights: &[NightResult], samples: &[RestartSample]) -> Vec<NightVerdict> {
|
||||
let dirty = nights_with_restart(samples);
|
||||
pub fn evaluate(
|
||||
nights: &[NightResult],
|
||||
samples: &[RestartSample],
|
||||
expected_pods: usize,
|
||||
) -> Vec<NightVerdict> {
|
||||
nights
|
||||
.iter()
|
||||
.map(|n| {
|
||||
let restarted = dirty.contains(&n.date);
|
||||
let (green, reason) = match (n.passed, restarted) {
|
||||
(true, false) => (true, String::from("pass; no under-load restart")),
|
||||
(false, false) => (false, String::from("soak SLO gate breach (Job FAIL)")),
|
||||
(true, true) => (
|
||||
false,
|
||||
String::from(
|
||||
"under-load pod restart in window (Job passed but streak-breaking)",
|
||||
),
|
||||
),
|
||||
(false, true) => (false, String::from("soak FAIL and under-load pod restart")),
|
||||
};
|
||||
.map(|night| {
|
||||
let evidence = restart_evidence(night, samples, expected_pods);
|
||||
let mut reasons = Vec::with_capacity(3);
|
||||
if !night.passed {
|
||||
reasons.push("soak SLO gate breach (Job FAIL)");
|
||||
}
|
||||
if evidence.dirty {
|
||||
reasons.push("under-load pod restart in window");
|
||||
}
|
||||
if !evidence.complete {
|
||||
reasons.push(
|
||||
"restart evidence missing or stale (need every expected pod at both load boundaries)",
|
||||
);
|
||||
}
|
||||
|
||||
let green = reasons.is_empty();
|
||||
NightVerdict {
|
||||
date: n.date.clone(),
|
||||
date: night.date.clone(),
|
||||
green,
|
||||
reason,
|
||||
reason: if green {
|
||||
String::from("pass; complete restart evidence; no under-load restart")
|
||||
} else {
|
||||
reasons.join("; ")
|
||||
},
|
||||
}
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// The trailing run of consecutive green nights (the GA streak). Any non-green
|
||||
/// night resets it; only the most recent unbroken tail counts.
|
||||
fn dates_are_consecutive(previous: &str, current: &str) -> bool {
|
||||
parse_date(previous).and_then(Date::next_day) == parse_date(current)
|
||||
}
|
||||
|
||||
fn advance_streak(previous_date: Option<&str>, current: &NightVerdict, run: usize) -> usize {
|
||||
if !current.green {
|
||||
return 0;
|
||||
}
|
||||
if previous_date.is_some_and(|previous| dates_are_consecutive(previous, ¤t.date)) {
|
||||
run + 1
|
||||
} else {
|
||||
1
|
||||
}
|
||||
}
|
||||
|
||||
/// The trailing run of unique, consecutive calendar dates with green verdicts.
|
||||
/// A failure, duplicate date, out-of-order row, or date gap starts a new run.
|
||||
#[must_use]
|
||||
pub fn current_streak(verdicts: &[NightVerdict]) -> usize {
|
||||
verdicts.iter().rev().take_while(|v| v.green).count()
|
||||
let mut run = 0;
|
||||
let mut previous_date = None;
|
||||
for verdict in verdicts {
|
||||
run = advance_streak(previous_date, verdict, run);
|
||||
previous_date = Some(verdict.date.as_str());
|
||||
}
|
||||
run
|
||||
}
|
||||
|
||||
/// Render `streak.tsv`: one row per night (`date green streak_after reason`)
|
||||
/// plus a trailing `# streak=<n>/<target>` summary line for an at-a-glance read.
|
||||
#[must_use]
|
||||
pub fn render_streak_tsv(verdicts: &[NightVerdict], target: usize) -> String {
|
||||
let mut s = String::from("date\tgreen\tstreak\treason\n");
|
||||
let mut output = String::from("date\tgreen\tstreak\treason\n");
|
||||
let mut run = 0usize;
|
||||
for v in verdicts {
|
||||
run = if v.green { run + 1 } else { 0 };
|
||||
s.push_str(&format!("{}\t{}\t{}\t{}\n", v.date, v.green, run, v.reason));
|
||||
let mut previous_date = None;
|
||||
for verdict in verdicts {
|
||||
run = advance_streak(previous_date, verdict, run);
|
||||
output.push_str(&format!(
|
||||
"{}\t{}\t{}\t{}\n",
|
||||
verdict.date, verdict.green, run, verdict.reason
|
||||
));
|
||||
previous_date = Some(verdict.date.as_str());
|
||||
}
|
||||
let streak = current_streak(verdicts);
|
||||
s.push_str(&format!("# streak={streak}/{target}\n"));
|
||||
s
|
||||
output.push_str(&format!("# streak={run}/{target}\n"));
|
||||
output
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
const EXPECTED_PODS: usize = 3;
|
||||
|
||||
fn pass_nights(n: usize) -> Vec<NightResult> {
|
||||
(1..=n)
|
||||
.map(|i| NightResult {
|
||||
date: format!("2026-06-{i:02}"),
|
||||
.map(|i| {
|
||||
let date = format!("2026-06-{i:02}");
|
||||
NightResult {
|
||||
start_timestamp: Some(format!("{date}T02:00:00Z")),
|
||||
end_timestamp: Some(format!("{date}T03:00:00Z")),
|
||||
date,
|
||||
passed: true,
|
||||
}
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// A constant-restart-count sample stream (no restarts) for every night —
|
||||
/// pods sitting at a steady `restartCount` (e.g. 2) all window.
|
||||
fn steady_samples(nights: usize, count: u64) -> Vec<RestartSample> {
|
||||
let mut v = Vec::new();
|
||||
for i in 1..=nights {
|
||||
for pod in ["tidaldb-0", "tidaldb-1", "tidaldb-2"] {
|
||||
v.push(RestartSample {
|
||||
date: format!("2026-06-{i:02}"),
|
||||
fn sample(
|
||||
date: &str,
|
||||
time: &str,
|
||||
pod: &str,
|
||||
pod_uid: Option<&str>,
|
||||
restarts: u64,
|
||||
) -> RestartSample {
|
||||
RestartSample {
|
||||
timestamp: format!("{date}T{time}"),
|
||||
date: date.to_string(),
|
||||
pod: pod.to_string(),
|
||||
restarts: count,
|
||||
});
|
||||
pod_uid: pod_uid.map(str::to_string),
|
||||
restarts,
|
||||
}
|
||||
}
|
||||
v
|
||||
|
||||
/// A constant-restart-count stream that tightly brackets every load window.
|
||||
fn steady_samples(nights: usize, count: u64) -> Vec<RestartSample> {
|
||||
let mut samples = Vec::new();
|
||||
for i in 1..=nights {
|
||||
let date = format!("2026-06-{i:02}");
|
||||
for pod in ["tidaldb-0", "tidaldb-1", "tidaldb-2"] {
|
||||
samples.push(sample(&date, "01:59:00Z", pod, None, count));
|
||||
samples.push(sample(&date, "03:01:00Z", pod, None, count));
|
||||
}
|
||||
}
|
||||
samples
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn thirty_clean_nights_make_a_full_streak() {
|
||||
fn thirty_clean_consecutive_nights_make_a_full_streak() {
|
||||
let nights = pass_nights(30);
|
||||
let samples = steady_samples(30, 2);
|
||||
let verdicts = evaluate(&nights, &samples);
|
||||
assert!(verdicts.iter().all(|v| v.green), "all 30 must be green");
|
||||
assert_eq!(current_streak(&verdicts), 30, "30 clean nights → streak 30");
|
||||
let verdicts = evaluate(&nights, &samples, EXPECTED_PODS);
|
||||
assert!(verdicts.iter().all(|verdict| verdict.green));
|
||||
assert_eq!(current_streak(&verdicts), 30);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn an_under_load_restart_breaks_the_streak_even_when_the_job_passed() {
|
||||
let nights = pass_nights(30);
|
||||
let mut samples = steady_samples(30, 2);
|
||||
// Night 15: tidaldb-1's cumulative restartCount rises 2 → 3 within the
|
||||
// window (an under-load restart) while the soak Job still PASSED.
|
||||
samples.push(RestartSample {
|
||||
date: "2026-06-15".into(),
|
||||
pod: "tidaldb-1".into(),
|
||||
restarts: 3,
|
||||
});
|
||||
let verdicts = evaluate(&nights, &samples);
|
||||
let n15 = verdicts.iter().find(|v| v.date == "2026-06-15").unwrap();
|
||||
assert!(
|
||||
!n15.green,
|
||||
"a night with an under-load restart is NOT green"
|
||||
);
|
||||
assert!(n15.reason.contains("restart"), "reason names the restart");
|
||||
// The streak is only the clean tail after night 15: nights 16..=30 = 15.
|
||||
assert_eq!(
|
||||
current_streak(&verdicts),
|
||||
15,
|
||||
"streak resets at the restart night"
|
||||
);
|
||||
samples.push(sample("2026-06-15", "02:30:00Z", "tidaldb-1", None, 3));
|
||||
let verdicts = evaluate(&nights, &samples, EXPECTED_PODS);
|
||||
let n15 = verdicts
|
||||
.iter()
|
||||
.find(|verdict| verdict.date == "2026-06-15")
|
||||
.expect("night 15 exists");
|
||||
assert!(!n15.green);
|
||||
assert!(n15.reason.contains("restart"));
|
||||
assert_eq!(current_streak(&verdicts), 15);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_gate_breach_night_resets_the_streak() {
|
||||
let mut nights = pass_nights(30);
|
||||
nights[19].passed = false; // night 20 breached an SLO gate (Job FAIL)
|
||||
nights[19].passed = false;
|
||||
let samples = steady_samples(30, 2);
|
||||
let verdicts = evaluate(&nights, &samples);
|
||||
let n20 = verdicts.iter().find(|v| v.date == "2026-06-20").unwrap();
|
||||
assert!(!n20.green, "a gate-breach night is not green");
|
||||
assert!(n20.reason.contains("gate"), "reason names the gate breach");
|
||||
assert_eq!(
|
||||
current_streak(&verdicts),
|
||||
10,
|
||||
"tail after night 20 = nights 21..=30"
|
||||
);
|
||||
let verdicts = evaluate(&nights, &samples, EXPECTED_PODS);
|
||||
let n20 = &verdicts[19];
|
||||
assert!(!n20.green);
|
||||
assert!(n20.reason.contains("gate"));
|
||||
assert_eq!(current_streak(&verdicts), 10);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_restart_within_the_first_observed_night_registers() {
|
||||
// No carried-forward baseline: a 0 → 1 rise on the first night must still
|
||||
// count as a restart (min < max that night).
|
||||
let nights = pass_nights(1);
|
||||
let samples = vec![
|
||||
RestartSample {
|
||||
date: "2026-06-01".into(),
|
||||
pod: "tidaldb-0".into(),
|
||||
restarts: 0,
|
||||
},
|
||||
RestartSample {
|
||||
date: "2026-06-01".into(),
|
||||
pod: "tidaldb-0".into(),
|
||||
restarts: 1,
|
||||
},
|
||||
sample("2026-06-01", "01:59:00Z", "tidaldb-0", None, 0),
|
||||
sample("2026-06-01", "03:01:00Z", "tidaldb-0", None, 1),
|
||||
];
|
||||
let verdicts = evaluate(&nights, &samples);
|
||||
assert!(
|
||||
!verdicts[0].green,
|
||||
"0→1 within the first night is a restart"
|
||||
let verdicts = evaluate(&nights, &samples, 1);
|
||||
assert!(!verdicts[0].green);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn pod_recreated_during_window_already_at_one_restart_registers() {
|
||||
let nights = pass_nights(1);
|
||||
let samples = vec![
|
||||
sample(
|
||||
"2026-06-01",
|
||||
"01:59:00Z",
|
||||
"tidaldb-0",
|
||||
Some("pod-before"),
|
||||
0,
|
||||
),
|
||||
sample("2026-06-01", "03:01:00Z", "tidaldb-0", Some("pod-after"), 1),
|
||||
];
|
||||
let verdicts = evaluate(&nights, &samples, 1);
|
||||
assert!(!verdicts[0].green);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn pod_recreation_with_zero_container_restarts_breaks_the_night() {
|
||||
let nights = pass_nights(1);
|
||||
let samples = vec![
|
||||
sample(
|
||||
"2026-06-01",
|
||||
"01:59:00Z",
|
||||
"tidaldb-0",
|
||||
Some("pod-before"),
|
||||
0,
|
||||
),
|
||||
sample("2026-06-01", "03:01:00Z", "tidaldb-0", Some("pod-after"), 0),
|
||||
];
|
||||
let verdicts = evaluate(&nights, &samples, 1);
|
||||
assert!(!verdicts[0].green);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn missing_restart_evidence_fails_closed() {
|
||||
let nights = pass_nights(1);
|
||||
let samples = vec![sample(
|
||||
"2026-06-01",
|
||||
"03:01:00Z",
|
||||
"tidaldb-0",
|
||||
Some("pod-uid"),
|
||||
0,
|
||||
)];
|
||||
let verdicts = evaluate(&nights, &samples, 1);
|
||||
assert!(!verdicts[0].green);
|
||||
assert!(verdicts[0].reason.contains("evidence"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn stale_boundary_samples_fail_closed() {
|
||||
let nights = pass_nights(1);
|
||||
let samples = vec![
|
||||
sample("2026-06-01", "01:55:00Z", "tidaldb-0", None, 0),
|
||||
sample("2026-06-01", "03:05:00Z", "tidaldb-0", None, 0),
|
||||
];
|
||||
let verdicts = evaluate(&nights, &samples, 1);
|
||||
assert!(!verdicts[0].green);
|
||||
assert!(verdicts[0].reason.contains("stale"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_legacy_counter_reset_inside_the_window_is_a_recovery_event() {
|
||||
let nights = pass_nights(2);
|
||||
let mut samples = steady_samples(1, 20);
|
||||
samples.extend([
|
||||
sample("2026-06-02", "01:59:00Z", "tidaldb-0", None, 20),
|
||||
sample("2026-06-02", "02:30:00Z", "tidaldb-0", None, 0),
|
||||
sample("2026-06-02", "03:01:00Z", "tidaldb-0", None, 1),
|
||||
]);
|
||||
for pod in ["tidaldb-1", "tidaldb-2"] {
|
||||
samples.push(sample("2026-06-02", "01:59:00Z", pod, None, 20));
|
||||
samples.push(sample("2026-06-02", "03:01:00Z", pod, None, 20));
|
||||
}
|
||||
let verdicts = evaluate(&nights, &samples, EXPECTED_PODS);
|
||||
assert!(!verdicts[1].green);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_restart_before_the_load_window_does_not_dirty_the_night() {
|
||||
let nights = pass_nights(1);
|
||||
let samples = vec![
|
||||
sample("2026-06-01", "00:01:00Z", "tidaldb-0", None, 0),
|
||||
sample("2026-06-01", "00:02:00Z", "tidaldb-0", None, 1),
|
||||
sample("2026-06-01", "01:59:00Z", "tidaldb-0", None, 1),
|
||||
sample("2026-06-01", "03:01:00Z", "tidaldb-0", None, 1),
|
||||
];
|
||||
let verdicts = evaluate(&nights, &samples, 1);
|
||||
assert!(verdicts[0].green);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_future_extra_pod_does_not_contaminate_prior_night_coverage() {
|
||||
let nights = pass_nights(2);
|
||||
let mut samples = steady_samples(2, 0);
|
||||
samples.push(sample("2026-06-02", "01:59:00Z", "tidaldb-3", None, 0));
|
||||
samples.push(sample("2026-06-02", "03:01:00Z", "tidaldb-3", None, 0));
|
||||
let verdicts = evaluate(&nights, &samples, EXPECTED_PODS);
|
||||
assert!(verdicts[0].green);
|
||||
assert!(!verdicts[1].green);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn missing_load_window_timestamps_fail_closed() {
|
||||
let mut nights = pass_nights(1);
|
||||
nights[0].start_timestamp = None;
|
||||
let verdicts = evaluate(&nights, &steady_samples(1, 0), EXPECTED_PODS);
|
||||
assert!(!verdicts[0].green);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn date_gaps_and_duplicates_cannot_inflate_the_streak() {
|
||||
let green = |date: &str| NightVerdict {
|
||||
date: date.to_string(),
|
||||
green: true,
|
||||
reason: String::from("pass"),
|
||||
};
|
||||
let gap = vec![
|
||||
green("2026-06-01"),
|
||||
green("2026-06-02"),
|
||||
green("2026-06-04"),
|
||||
];
|
||||
let duplicate = vec![
|
||||
green("2026-06-01"),
|
||||
green("2026-06-02"),
|
||||
green("2026-06-02"),
|
||||
];
|
||||
assert_eq!(current_streak(&gap), 1);
|
||||
assert_eq!(current_streak(&duplicate), 1);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn parse_ledger_reads_window_fields_and_rejects_impossible_dates() {
|
||||
let tsv = "\
|
||||
2026-06-01\tPASS\tsoak_rc=0\tstart_utc=2026-06-01T02:00:00Z\tend_utc=2026-06-01T03:00:00Z\n\
|
||||
2026-06-02\tFAIL\tsoak_rc=1\tstart_utc=2026-06-02T02:00:00Z\tend_utc=2026-06-02T03:00:00Z\n\
|
||||
2026-02-30\tPASS\tjunk\n";
|
||||
let rows = parse_ledger(tsv);
|
||||
assert_eq!(rows.len(), 2);
|
||||
assert!(rows[0].passed);
|
||||
assert!(!rows[1].passed);
|
||||
assert_eq!(
|
||||
rows[0].start_timestamp.as_deref(),
|
||||
Some("2026-06-01T02:00:00Z")
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn parse_ledger_skips_headers_and_reads_the_verdict() {
|
||||
fn parse_restarts_accepts_legacy_and_uid_rows() {
|
||||
let tsv = "\
|
||||
2026-06-01\tPASS\trc=0\tpassed=true\tp99_ms=9.3\terr_frac=0.0001\timage=m12\n\
|
||||
2026-06-02\tFAIL\trc=1\tpassed=false\tp99_ms=210\terr_frac=0.04\timage=m12\n\
|
||||
not-a-date\tPASS\tjunk\n";
|
||||
let rows = parse_ledger(tsv);
|
||||
assert_eq!(rows.len(), 2, "the junk/header row is skipped");
|
||||
assert!(rows[0].passed);
|
||||
assert!(!rows[1].passed);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn parse_restarts_skips_header_and_keys_by_date() {
|
||||
let tsv = "\
|
||||
ts_utc\tpod\trestarts\tphase\tready\n\
|
||||
ts_utc\tpod\tuid\trestarts\tphase\tready\n\
|
||||
2026-06-01T02:05:00Z\ttidaldb-0\t2\tRunning\ttrue\n\
|
||||
2026-06-01T02:10:00Z\ttidaldb-1\t2\tRunning\ttrue\n";
|
||||
2026-06-01T02:10:00Z\ttidaldb-1\tuid-1\t3\tRunning\ttrue\n\
|
||||
not-a-timestamp\ttidaldb-2\t4\tRunning\ttrue\n";
|
||||
let rows = parse_restarts(tsv);
|
||||
assert_eq!(rows.len(), 2, "header skipped, two data rows parsed");
|
||||
assert_eq!(rows.len(), 2);
|
||||
assert_eq!(rows[0].date, "2026-06-01");
|
||||
assert_eq!(rows[0].restarts, 2);
|
||||
assert_eq!(rows[0].pod_uid, None);
|
||||
assert_eq!(rows[1].restarts, 3);
|
||||
assert_eq!(rows[1].pod_uid.as_deref(), Some("uid-1"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn render_streak_tsv_carries_running_count_and_summary() {
|
||||
let nights = pass_nights(3);
|
||||
let samples = steady_samples(3, 0);
|
||||
let verdicts = evaluate(&nights, &samples);
|
||||
let verdicts = evaluate(&nights, &samples, EXPECTED_PODS);
|
||||
let tsv = render_streak_tsv(&verdicts, 30);
|
||||
assert!(tsv.contains("date\tgreen\tstreak\treason"), "has a header");
|
||||
assert!(
|
||||
tsv.trim_end().ends_with("# streak=3/30"),
|
||||
"summary line present: {tsv}"
|
||||
);
|
||||
assert!(tsv.contains("date\tgreen\tstreak\treason"));
|
||||
assert!(tsv.trim_end().ends_with("# streak=3/30"));
|
||||
}
|
||||
}
|
||||
|
||||
@ -3,13 +3,25 @@
|
||||
//!
|
||||
//! These tests close the "CRDT engine is never invoked in production" gap: they
|
||||
//! drive the LIVE production entry points on real `TidalDb` instances (not the
|
||||
//! `ReconciliationEngine` directly), partition two nodes, write divergently on
|
||||
//! each side, exchange snapshots, reconcile, and assert the merged signal decay
|
||||
//! score AND the windowed (`AllTime`) bucket count both survive end to end.
|
||||
//! `ReconciliationEngine` directly), diverge two nodes, exchange snapshots,
|
||||
//! reconcile, and assert what the heal actually guarantees.
|
||||
//!
|
||||
//! Each `TidalDb` runs as a distinct shard, so `take_crdt_snapshot` attributes
|
||||
//! its contributions to a distinct node in the CRDT — exactly the multi-node
|
||||
//! divergence the reconciliation engine exists to merge.
|
||||
//! # The contract these tests hold (read before changing an expectation)
|
||||
//!
|
||||
//! `take_crdt_snapshot` keys every signal contribution to ONE canonical
|
||||
//! contributor ([`ShardId::SINGLE`]), not to the local shard. Signals are
|
||||
//! replicated from a single writer through the WAL relay, so each node's hot
|
||||
//! accumulator ALREADY contains the other nodes' relayed events. Attributing it
|
||||
//! per-node would fabricate N disjoint contributions for one logical stream,
|
||||
//! which `CrdtSignalState::merge` then sums - double-counting every replicated
|
||||
//! event on every reconcile (the creep a UAT caught). See the rationale on
|
||||
//! `TidalDb::take_crdt_snapshot`.
|
||||
//!
|
||||
//! With one contributor the merge is therefore deterministic convergence, not
|
||||
//! addition: the decay score is last-writer-wins on `(last_update_ns, score)`,
|
||||
//! and the windowed bucket count is the PN-counter per-node max. Both nodes
|
||||
//! converge on the MORE COMPLETE accumulator and stay there under repeated
|
||||
//! exchange. A test that asserts `3 + 5 == 8` here is asserting the bug.
|
||||
#![allow(clippy::unwrap_used, clippy::float_cmp)]
|
||||
|
||||
use std::time::Duration;
|
||||
@ -66,19 +78,18 @@ fn node(shard: ShardId) -> TidalDb {
|
||||
.expect("ephemeral node opens")
|
||||
}
|
||||
|
||||
/// Two partitioned nodes write divergent signals to the SAME entity, exchange
|
||||
/// snapshots, and reconcile. After healing, BOTH nodes report the merged decay
|
||||
/// score (sum of both contributions) AND the merged windowed count (sum of both
|
||||
/// event counts) — the per-node bucket count round-trips, not dropping to 0.
|
||||
/// Two diverged nodes exchange snapshots and reconcile. Both converge on the
|
||||
/// more complete accumulator, the windowed count survives the round trip (it
|
||||
/// does not drop to 0 - the original finding), and neither side inflates past
|
||||
/// the true event total.
|
||||
#[test]
|
||||
fn two_node_partition_heals_signals_and_windowed_counts() {
|
||||
fn two_node_divergence_converges_on_the_more_complete_accumulator() {
|
||||
let node_a = node(ShardId(0));
|
||||
let node_b = node(ShardId(1));
|
||||
let item = EntityId::new(42);
|
||||
|
||||
// ── Partition: each node accumulates independent events for `item`. ──
|
||||
// Node A: 3 views. Node B: 5 views. Same fixed recent timestamp so decay is
|
||||
// negligible and the arithmetic is exact-ish.
|
||||
// Node A saw 3 views, node B saw 5 of the same logical stream (B is the
|
||||
// more complete replica). Same fixed timestamp so decay is negligible.
|
||||
let ts = Timestamp::now();
|
||||
for _ in 0..3 {
|
||||
node_a.signal("view", item, 1.0, ts).unwrap();
|
||||
@ -87,7 +98,6 @@ fn two_node_partition_heals_signals_and_windowed_counts() {
|
||||
node_b.signal("view", item, 1.0, ts).unwrap();
|
||||
}
|
||||
|
||||
// Pre-reconcile each node sees ONLY its own events.
|
||||
let a_count_before = node_a
|
||||
.read_windowed_count(item, "view", Window::AllTime)
|
||||
.unwrap();
|
||||
@ -97,10 +107,6 @@ fn two_node_partition_heals_signals_and_windowed_counts() {
|
||||
assert_eq!(a_count_before, 3, "node A sees only its 3 events pre-heal");
|
||||
assert_eq!(b_count_before, 5, "node B sees only its 5 events pre-heal");
|
||||
|
||||
let a_score_before = node_a
|
||||
.read_decay_score(item, "view", 0)
|
||||
.unwrap()
|
||||
.unwrap_or(0.0);
|
||||
let b_score_before = node_b
|
||||
.read_decay_score(item, "view", 0)
|
||||
.unwrap()
|
||||
@ -115,26 +121,28 @@ fn two_node_partition_heals_signals_and_windowed_counts() {
|
||||
assert!(ops_a >= 1, "reconcile must apply at least the signal merge");
|
||||
assert!(ops_b >= 1);
|
||||
|
||||
// ── Post-heal: both nodes converge to the merged truth. ──
|
||||
// Windowed count: 3 + 5 = 8 on BOTH nodes (finding 6 — the bucket count is
|
||||
// carried through from_node_contribution and survives merge, not 0).
|
||||
let a_count_after = node_a
|
||||
.read_windowed_count(item, "view", Window::AllTime)
|
||||
.unwrap();
|
||||
let b_count_after = node_b
|
||||
.read_windowed_count(item, "view", Window::AllTime)
|
||||
.unwrap();
|
||||
|
||||
// Convergence is the property that matters: both sides agree.
|
||||
assert_eq!(
|
||||
a_count_after, 8,
|
||||
"node A windowed count must be merged total (3+5), not its local 3 or 0"
|
||||
a_count_after, b_count_after,
|
||||
"both nodes must converge on one windowed count after the exchange"
|
||||
);
|
||||
// ... on the more complete accumulator (PN-counter per-node max of 3 and 5),
|
||||
// never 0 (the count survives the snapshot round trip) and never 8 (summing
|
||||
// one logical stream twice is the double-count bug).
|
||||
assert_eq!(
|
||||
b_count_after, 8,
|
||||
"node B windowed count must be merged total (3+5), not its local 5 or 0"
|
||||
a_count_after, 5,
|
||||
"converged count is the more complete accumulator, not a sum"
|
||||
);
|
||||
|
||||
// Decay score: the merged score is the sum of both nodes' contributions.
|
||||
let expected = a_score_before + b_score_before;
|
||||
// Decay score: LWW on (last_update_ns, score) with a single contributor, so
|
||||
// both sides hold node B's larger accumulator.
|
||||
let a_score_after = node_a
|
||||
.read_decay_score(item, "view", 0)
|
||||
.unwrap()
|
||||
@ -145,27 +153,78 @@ fn two_node_partition_heals_signals_and_windowed_counts() {
|
||||
.unwrap_or(0.0);
|
||||
// Tolerance: decay over the few-ms reconcile window is negligible but nonzero.
|
||||
assert!(
|
||||
(a_score_after - expected).abs() < 1e-3,
|
||||
"node A merged score {a_score_after} should be ~sum {expected}"
|
||||
(a_score_after - b_score_after).abs() < 1e-3,
|
||||
"both nodes must converge on one score: A {a_score_after} vs B {b_score_after}"
|
||||
);
|
||||
assert!(
|
||||
(b_score_after - expected).abs() < 1e-3,
|
||||
"node B merged score {b_score_after} should be ~sum {expected}"
|
||||
(a_score_after - b_score_before).abs() < 1e-3,
|
||||
"converged score {a_score_after} should be node B's {b_score_before}"
|
||||
);
|
||||
}
|
||||
|
||||
/// Reconciliation against a remote snapshot that this node has already fully
|
||||
/// absorbed (its events are a subset of what the local side counted) is a no-op
|
||||
/// for the windowed count — it never shrinks a locally-durable count nor
|
||||
/// double-counts an already-merged remote contribution.
|
||||
/// Repeated exchange between already-converged nodes changes nothing.
|
||||
///
|
||||
/// This is the idempotency that the heal path actually guarantees: re-merging a
|
||||
/// remote snapshot whose every contribution is `<=` what the local node already
|
||||
/// holds leaves the count unchanged. (Re-*snapshotting* after a merge and
|
||||
/// re-merging is NOT idempotent — `apply_crdt_state` force-sets a single warm
|
||||
/// counter and loses per-node provenance — so the heal is a one-shot exchange,
|
||||
/// not a repeatedly-applied stream; we assert the property that genuinely
|
||||
/// holds.)
|
||||
/// This is the regression guard for the creep that per-node attribution caused:
|
||||
/// with the whole accumulator attributed per shard, every reconcile re-summed
|
||||
/// the same relayed events and the count grew without any new signal being
|
||||
/// written. Under the canonical-contributor keying the second and third rounds
|
||||
/// are exact no-ops.
|
||||
#[test]
|
||||
fn repeated_reconcile_of_converged_nodes_does_not_creep() {
|
||||
let node_a = node(ShardId(0));
|
||||
let node_b = node(ShardId(1));
|
||||
let item = EntityId::new(99);
|
||||
|
||||
let ts = Timestamp::now();
|
||||
for _ in 0..4 {
|
||||
node_a.signal("view", item, 1.0, ts).unwrap();
|
||||
}
|
||||
for _ in 0..4 {
|
||||
node_b.signal("view", item, 1.0, ts).unwrap();
|
||||
}
|
||||
|
||||
// First exchange converges the pair.
|
||||
let snap_b = node_b.take_crdt_snapshot().unwrap();
|
||||
node_a.reconcile_with(&snap_b).unwrap();
|
||||
let converged = node_a
|
||||
.read_windowed_count(item, "view", Window::AllTime)
|
||||
.unwrap();
|
||||
assert_eq!(converged, 4, "converged on the shared 4-event accumulator");
|
||||
|
||||
// Two more rounds with FRESH snapshots taken after the merge - the shape
|
||||
// anti-entropy would actually run - must not move the count or the score.
|
||||
let score_after_first = node_a
|
||||
.read_decay_score(item, "view", 0)
|
||||
.unwrap()
|
||||
.unwrap_or(0.0);
|
||||
for round in 1..=2 {
|
||||
let fresh_a = node_a.take_crdt_snapshot().unwrap();
|
||||
let fresh_b = node_b.take_crdt_snapshot().unwrap();
|
||||
node_b.reconcile_with(&fresh_a).unwrap();
|
||||
node_a.reconcile_with(&fresh_b).unwrap();
|
||||
|
||||
let count = node_a
|
||||
.read_windowed_count(item, "view", Window::AllTime)
|
||||
.unwrap();
|
||||
assert_eq!(
|
||||
count, converged,
|
||||
"round {round}: repeated reconcile must not inflate the count"
|
||||
);
|
||||
let score = node_a
|
||||
.read_decay_score(item, "view", 0)
|
||||
.unwrap()
|
||||
.unwrap_or(0.0);
|
||||
assert!(
|
||||
(score - score_after_first).abs() < 1e-3,
|
||||
"round {round}: repeated reconcile must not inflate the score \
|
||||
({score} vs {score_after_first})"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
/// Reconciliation against a remote snapshot the local node already covers is a
|
||||
/// no-op for the windowed count: it never shrinks a locally-durable count, and
|
||||
/// re-applying the same snapshot never grows one.
|
||||
#[test]
|
||||
fn reconcile_with_already_absorbed_remote_is_noop_for_count() {
|
||||
let node_a = node(ShardId(0));
|
||||
@ -180,30 +239,28 @@ fn reconcile_with_already_absorbed_remote_is_noop_for_count() {
|
||||
node_b.signal("view", item, 1.0, ts).unwrap();
|
||||
}
|
||||
|
||||
// Snapshot B (node 1, 4 events) BEFORE any merge.
|
||||
// Snapshot B (4 events) BEFORE any merge. A is the more complete side.
|
||||
let snap_b = node_b.take_crdt_snapshot().unwrap();
|
||||
|
||||
// First reconcile: 6 (node 0) + 4 (node 1) = 10.
|
||||
// Merging a strictly smaller remote accumulator must not shrink A.
|
||||
node_a.reconcile_with(&snap_b).unwrap();
|
||||
let after_first = node_a
|
||||
.read_windowed_count(item, "view", Window::AllTime)
|
||||
.unwrap();
|
||||
assert_eq!(after_first, 10, "first reconcile merges to 10");
|
||||
assert_eq!(
|
||||
after_first, 6,
|
||||
"merging a smaller remote accumulator keeps the local 6"
|
||||
);
|
||||
|
||||
// Re-merge the SAME pre-merge snapshot B. The merged per-node total
|
||||
// (node 0 = 10 absorbed, node 1 = 4) is 14 > 10, so the warm tier would
|
||||
// grow — confirming the heal is intentionally one-shot. We instead assert
|
||||
// the never-shrink guarantee by re-merging a snapshot whose node-1
|
||||
// contribution (4) the local count (10) already covers: applying snap_b's
|
||||
// node-1 bucket of 4 against a local node-1 bucket that is now part of the
|
||||
// 10 must not drop below 10.
|
||||
// Re-merging the same snapshot is idempotent in both directions: no shrink
|
||||
// below 6 and no growth toward 6 + 4.
|
||||
node_a.reconcile_with(&snap_b).unwrap();
|
||||
let after_second = node_a
|
||||
.read_windowed_count(item, "view", Window::AllTime)
|
||||
.unwrap();
|
||||
assert!(
|
||||
after_second >= 10,
|
||||
"re-merge must never shrink the locally-durable count below 10 (got {after_second})"
|
||||
assert_eq!(
|
||||
after_second, 6,
|
||||
"re-merge must neither shrink nor inflate the locally-durable count"
|
||||
);
|
||||
}
|
||||
|
||||
|
||||
@ -2,7 +2,11 @@
|
||||
name = "tidalctl"
|
||||
version = "0.1.0"
|
||||
edition.workspace = true
|
||||
rust-version.workspace = true
|
||||
# Higher than the workspace floor: aws-config -> aws-types 1.3.16 declares
|
||||
# rustc 1.91.1, and resolution fails workspace-wide below it. Declared here so
|
||||
# the requirement is visible where it originates rather than only in a lockfile
|
||||
# error.
|
||||
rust-version = "1.91.1"
|
||||
description = "Command-line inspector for embedded tidalDB instances"
|
||||
license.workspace = true
|
||||
|
||||
|
||||
Loading…
Reference in New Issue
Block a user