Compare commits

..

No commits in common. "e117333feca5b89a92b5f9416540555ae04295cc" and "19316227c5daa5910fed8d1b87aec7aab1dc8d9f" have entirely different histories.

View File

@ -84,67 +84,6 @@
# reports a mode other than `standalone`. That guard existed BECAUSE of the
# context-dropping bug; relax it in the same change that repoints the URL, not
# before, so a stale image cannot be adopted silently.
#
# The consumer also needs the CA, because `:9500` serves TLS from the cluster's
# OWN cert-manager root and no public root can verify it. Kubernetes Secrets do
# not cross namespaces, so `ca.crt` from `tidaldb-cluster-ca` has to be copied
# into the consumer's namespace (peach reads it via
# `TidalClient::from_env_with_ca`).
#
# A copied certificate deserves suspicion — k3s-fleet's own history records an
# incident where "the expired central SDLC wildcard had been copied manually
# and never renewed into its consumer namespaces". This copy is materially
# different and the numbers are why:
#
# tidaldb-cluster-ca duration 87600h (10y), renewBefore 8760h (1y) isCA
# tidaldb-cluster-tls duration 2160h (90d), renewBefore 720h (30d)
#
# The 90-day LEAF is what rotates, roughly every 60 days, and it rotates UNDER
# this same root — so the copied CA keeps verifying it and needs no attention.
# The fleet incident was an expiring leaf; this is a decade-lived root.
#
# The coupling that DOES bite is recreation, not expiry: `tidaldb-cluster-ca` is
# self-signed and generated fresh on first apply, so destroying and re-applying
# this namespace mints a NEW root and every copy goes stale at once — every
# consumer request then fails TLS verification with nothing wrong on either
# side. Treat "refresh the CA copy" as a mandatory step of any recreate, and
# prefer `kubectl rollout restart` over deleting the namespace.
#
# ── STATUS 2026-09-16: applied, then PARKED at replicas: 0 ───────────────────
# Two blockers, in the order they must be fixed. The namespace, credentials,
# both certs and all three PVCs are retained, so resuming is `kubectl scale
# sts/tidaldb --replicas=3`. It was NOT deleted, precisely because deleting the
# namespace destroys the self-signed root and staleness every copy of it (see
# above).
#
# 1. THE IMAGE CAME FROM THE WRONG BUILD PATH — fix this first.
# `docs/runbooks/kubernetes.md` §"Deploy the cluster" is explicit: use
# `./scripts/build-release.sh <tag> server`, NOT a bare docker build, because
# the binary needs `libmvec.so.1` which is ABSENT on bookworm, so the runtime
# must be `debian:trixie-slim`. The Woodpecker/Kaniko build uses
# `docker/standalone/Dockerfile`, whose runtime stage is
# `FROM debian:bookworm-slim`. The digest pinned above is that CI build.
# The pods DO boot on it (WAL recovery, elections and the HTTPS listener all
# work), so nothing fails loudly — libmvec is vectorized math, reached on the
# vector paths rather than at startup, which is exactly why this is worth
# writing down instead of trusting "it started".
# The runbook also requires pinning the linux/amd64 PLATFORM manifest digest,
# never the OCI index or the `unknown/unknown` attestation manifest; verify
# with `docker buildx imagetools inspect ... --raw` before pinning.
#
# 2. A FRESH CLUSTER DID NOT REACH READY, and this must be re-tested on a
# correctly built image before being treated as a bug. Measured: all three
# pods Running and `/health/live` + `/health/startup` 200, but `/health` 503
# for 15 minutes with
# "shard group N has not yet first-converged against a KNOWN leader frontier"
# `/cluster/status` from tidaldb-0 showed its own groups at `leader_seqno: 1`
# (the term marker) but group 1 at `leader_seqno: 0, reseeding: true`, while
# tidaldb-1 and tidaldb-2 both logged "leadership view updated (follower)"
# under leader tidaldb-0 term 1 and applied that term marker. Readiness needs
# `converged`, which `note_lag_for_readiness` refuses to set while
# `leader_seqno == 0` ("absence of lag is not convergence",
# cluster/node.rs:3840).
# Do not debug this against a bookworm image. Blocker 1 first.
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
@ -153,21 +92,6 @@ namespace: tidaldb-cluster
resources:
- ../cluster
# GKE fact 3: the image. Pinned HERE and not in the base, because the k3s fleet
# applies `k8s/cluster/` directly — editing the base pin would silently repoint
# the fleet's production cluster as a side effect of a GKE change. Same
# isolation reason `cluster-local-kind` retargets here.
#
# By DIGEST, not tag: this is the first build carrying the signal-context fix
# (tidalDB 6ad8c51, image created 2026-09-16T04:28:15Z). The base's pin is
# m12-agesort-20260901 / sha256:9234aacb…, created 2026-09-01 — two weeks
# BEFORE the fix — so deploying the base unmodified would run the binary that
# silently discards user_id/creator_id behind a 204, which is the whole reason
# this deployment exists.
images:
- name: registry.threesix.ai/tidal/server
digest: sha256:96696a662c3a3c1bcf2a5c9bdcd0455cd77ca8d876f2d5fa7cacccee137d780f
patches:
# GKE fact 1 — see header.
- target: