diff --git a/k8s/cluster-gke-peach/kustomization.yaml b/k8s/cluster-gke-peach/kustomization.yaml index 30e0d80..ce6eb68 100644 --- a/k8s/cluster-gke-peach/kustomization.yaml +++ b/k8s/cluster-gke-peach/kustomization.yaml @@ -109,6 +109,42 @@ # consumer request then fails TLS verification with nothing wrong on either # side. Treat "refresh the CA copy" as a mandatory step of any recreate, and # prefer `kubectl rollout restart` over deleting the namespace. +# +# ── STATUS 2026-09-16: applied, then PARKED at replicas: 0 ─────────────────── +# Two blockers, in the order they must be fixed. The namespace, credentials, +# both certs and all three PVCs are retained, so resuming is `kubectl scale +# sts/tidaldb --replicas=3`. It was NOT deleted, precisely because deleting the +# namespace destroys the self-signed root and staleness every copy of it (see +# above). +# +# 1. THE IMAGE CAME FROM THE WRONG BUILD PATH — fix this first. +# `docs/runbooks/kubernetes.md` §"Deploy the cluster" is explicit: use +# `./scripts/build-release.sh server`, NOT a bare docker build, because +# the binary needs `libmvec.so.1` which is ABSENT on bookworm, so the runtime +# must be `debian:trixie-slim`. The Woodpecker/Kaniko build uses +# `docker/standalone/Dockerfile`, whose runtime stage is +# `FROM debian:bookworm-slim`. The digest pinned above is that CI build. +# The pods DO boot on it (WAL recovery, elections and the HTTPS listener all +# work), so nothing fails loudly — libmvec is vectorized math, reached on the +# vector paths rather than at startup, which is exactly why this is worth +# writing down instead of trusting "it started". +# The runbook also requires pinning the linux/amd64 PLATFORM manifest digest, +# never the OCI index or the `unknown/unknown` attestation manifest; verify +# with `docker buildx imagetools inspect ... --raw` before pinning. +# +# 2. A FRESH CLUSTER DID NOT REACH READY, and this must be re-tested on a +# correctly built image before being treated as a bug. Measured: all three +# pods Running and `/health/live` + `/health/startup` 200, but `/health` 503 +# for 15 minutes with +# "shard group N has not yet first-converged against a KNOWN leader frontier" +# `/cluster/status` from tidaldb-0 showed its own groups at `leader_seqno: 1` +# (the term marker) but group 1 at `leader_seqno: 0, reseeding: true`, while +# tidaldb-1 and tidaldb-2 both logged "leadership view updated (follower)" +# under leader tidaldb-0 term 1 and applied that term marker. Readiness needs +# `converged`, which `note_lag_for_readiness` refuses to set while +# `leader_seqno == 0` ("absence of lag is not convergence", +# cluster/node.rs:3840). +# Do not debug this against a bookworm image. Blocker 1 first. apiVersion: kustomize.config.k8s.io/v1beta1 kind: Kustomization @@ -117,6 +153,21 @@ namespace: tidaldb-cluster resources: - ../cluster +# GKE fact 3: the image. Pinned HERE and not in the base, because the k3s fleet +# applies `k8s/cluster/` directly — editing the base pin would silently repoint +# the fleet's production cluster as a side effect of a GKE change. Same +# isolation reason `cluster-local-kind` retargets here. +# +# By DIGEST, not tag: this is the first build carrying the signal-context fix +# (tidalDB 6ad8c51, image created 2026-09-16T04:28:15Z). The base's pin is +# m12-agesort-20260901 / sha256:9234aacb…, created 2026-09-01 — two weeks +# BEFORE the fix — so deploying the base unmodified would run the binary that +# silently discards user_id/creator_id behind a 204, which is the whole reason +# this deployment exists. +images: + - name: registry.threesix.ai/tidal/server + digest: sha256:96696a662c3a3c1bcf2a5c9bdcd0455cd77ca8d876f2d5fa7cacccee137d780f + patches: # GKE fact 1 — see header. - target: