thepeach's discover corpus should run the CLUSTER, not the standalone, and it
should run on GKE. Both halves are now evidenced rather than assumed.
Cluster, because the reason for the standalone is gone. The standalone was
chosen to work around cluster-mode `POST /signals` silently discarding
user_id/creator_id behind a 204 — fixed at the source in 6ad8c51. Inheriting a
workaround whose defect is fixed is how a temporary choice becomes permanent.
GKE, because that is where the consumer is. k3s-fleet's own history records the
fleet instance holding 1,720m of CPU requests — 17% of fleet allocatable —
while serving NOTHING, with tidaldb-staging at `items: 0` after 53 days,
because "the only in-source consumers ... run on GKE". Cluster-local DNS does
not cross clusters, so the consumer's
`tidaldb.tidaldb-cluster.svc.cluster.local:9500` can only resolve here.
An overlay, not a fork — precedent is ../cluster-local-kind, which patches the
same storage field. Two GKE facts the base cannot know:
1. `storageClassName: local-path` is the k3s provisioner. GKE has no such
class, so the volumeClaimTemplate would leave every pod Pending
indefinitely with nothing wrong on the StatefulSet itself. Patched to
`standard-rwo` (pd.csi, WaitForFirstConsumer — the binding mode a
StatefulSet wants).
2. The base publishes `tidaldb.threesix.ai`. That is a FLEET hostname with no
DNS record here that must not get one, and the base's own comment says
"Remove for internal-only". Deleted rather than left inert, which keeps DNS,
certificates and load-balancer topology entirely out of this change. Peach
reaches the service over cluster-local DNS; nothing needs exposing.
Worth recording because it nearly shipped: my first delete patch targeted a
`traefik.io` IngressRoute. The base publishes a plain `networking.k8s.io/v1`
Ingress carrying Traefik ANNOTATIONS, so the patch matched nothing, kustomize
rendered happily, and the Ingress survived — GKE would have published that
hostname. Rendering the overlay and reading the object inventory is what caught
it, so that check is now preflight step 1 with its expected output written
down, not left to memory.
The NetworkPolicy is kept verbatim and needs no widening: its `:9500` rule
carries no `from` selector (deliberate in the base, because kubelet probes come
from the node), which is exactly what lets thepeach-staging reach this
namespace with no cross-namespace rule. I had earlier claimed cross-namespace
was BLOCKED here; that was wrong, from reading a grep-filtered view instead of
the file.
Validated clean against GKE 2026-09-16: all 12 objects pass server dry-run,
warnings only. NOT applied — the base's image pin predates the signal-context
fix, so applying now would deploy the buggy binary. Preflight step 3 is the pin.
147 lines
6.9 KiB
YAML
147 lines
6.9 KiB
YAML
# The tidalDB RF3 CLUSTER on GKE, for thepeach's discover post corpus.
|
|
#
|
|
# Same shape as the k3s fleet deployment — ONE StatefulSet, 3 replicas, every
|
|
# pod a region, quorum-acked writes, inter-node mTLS from the cluster's own
|
|
# cert-manager CA. This is an overlay and not a fork: the base owns the
|
|
# topology, and the two things below are the only GKE facts the base cannot
|
|
# know. Precedent is `../cluster-local-kind`, which patches the same field.
|
|
#
|
|
# WHY A CLUSTER AND NOT THE STANDALONE
|
|
# ------------------------------------
|
|
# The standalone was chosen for staging ONLY to work around a real defect:
|
|
# cluster-mode `POST /signals` applied (signal, entity, weight) and silently
|
|
# discarded `user_id`/`creator_id` while still answering 204, so a clustered
|
|
# deployment accepted every behavioural signal and learned nothing. That is
|
|
# fixed at the source in 6ad8c51 (`signal_with_context_staged`), so the
|
|
# workaround is retired rather than inherited.
|
|
#
|
|
# WHY GKE AND NOT THE FLEET
|
|
# -------------------------
|
|
# The consumer is here. `k3s-fleet/deployments/history/tidaldb.md` measured the
|
|
# fleet instance holding 1,720m of CPU requests — 17% of fleet allocatable —
|
|
# while serving NOTHING, with `tidaldb-staging` reporting `items: 0` after 53
|
|
# days, because "the only in-source consumers ... run on GKE". thepeach's api
|
|
# and companion-worker are those consumers, and cluster-local DNS does not
|
|
# cross clusters.
|
|
#
|
|
# ── GKE fact 1: storage class ────────────────────────────────────────────────
|
|
# The base asks for `local-path`, the k3s local-path provisioner. GKE has no
|
|
# such class, so the volumeClaimTemplate would leave every pod Pending
|
|
# indefinitely with no error on the StatefulSet itself. `standard-rwo` is the
|
|
# GKE default (pd.csi.storage.gke.io, WaitForFirstConsumer), which is also the
|
|
# binding mode a StatefulSet wants.
|
|
#
|
|
# ── GKE fact 2: no public ingress ────────────────────────────────────────────
|
|
# The base's `ingress.yaml` publishes `tidaldb.threesix.ai` via Traefik + a
|
|
# Let's Encrypt cert. That hostname is a FLEET name; it has no DNS record here
|
|
# and must not get one. Its own comment says "Remove for internal-only", and
|
|
# internal-only is exactly right: peach reaches this over cluster-local DNS at
|
|
# `tidaldb.tidaldb-cluster.svc.cluster.local:9500`, so nothing needs to be
|
|
# exposed. Deleting the three objects rather than leaving them inert keeps DNS,
|
|
# certificates and load-balancer topology entirely out of this change.
|
|
#
|
|
# The NetworkPolicy is kept verbatim and needs no widening: its `:9500` rule
|
|
# carries no `from` selector, so any namespace may reach the client port. That
|
|
# is deliberate in the base ("`:9500` IS INTENTIONALLY LEFT OPEN") because
|
|
# kubelet probes originate from the node, and it is what lets `thepeach-staging`
|
|
# pods reach this namespace with no cross-namespace rule.
|
|
#
|
|
# ── Preflight, in order ──────────────────────────────────────────────────────
|
|
# 1. Prove nothing public is rendered. This is the check that caught the
|
|
# surviving Ingress; it is cheap and it is the difference between an
|
|
# internal service and a published hostname:
|
|
#
|
|
# kubectl kustomize k8s/cluster-gke-peach | grep -E '^kind:' | sort -u
|
|
#
|
|
# Expect exactly: Certificate, ConfigMap, Issuer, Namespace, NetworkPolicy,
|
|
# PodDisruptionBudget, Service, StatefulSet. Any Ingress, IngressRoute,
|
|
# ServersTransport or Middleware means a delete patch stopped matching —
|
|
# STOP and fix the patch, never apply past it.
|
|
#
|
|
# 2. Create `tidaldb-credentials`. It is NOT in this overlay and never should
|
|
# be — no credential belongs in a manifest. `TIDAL_API_KEY` must be
|
|
# byte-identical to the consumer's `TIDALDB_POSTS_API_KEY` (GSM
|
|
# `thepeach-staging-tidaldb-discover-api-key`) or the bearer check refuses
|
|
# every request; also needs `TIDAL_CLUSTER_KEY`. See
|
|
# `../cluster/secret.example.yaml`.
|
|
#
|
|
# 3. Pin the image digest that carries the signal-context fix. The base's pin
|
|
# predates it, and anything older reintroduces the silent drop this
|
|
# deployment exists to avoid.
|
|
#
|
|
# 4. Server dry-run, which validates admission without creating anything:
|
|
#
|
|
# kubectl apply -k k8s/cluster-gke-peach --dry-run=server
|
|
#
|
|
# Namespaced objects report `NotFound` until the namespace exists; create it
|
|
# first and re-run. Validated clean against GKE 2026-09-16 — all 12 objects,
|
|
# warnings only (the `kubectl create`-vs-`apply` annotation, and
|
|
# cert-manager's rotationPolicy default change).
|
|
#
|
|
# ── After applying ───────────────────────────────────────────────────────────
|
|
# The consumer refuses a cluster instance until its own guard is relaxed:
|
|
# `thepeach/services/api/src/state.rs` disables the client when `/health`
|
|
# reports a mode other than `standalone`. That guard existed BECAUSE of the
|
|
# context-dropping bug; relax it in the same change that repoints the URL, not
|
|
# before, so a stale image cannot be adopted silently.
|
|
apiVersion: kustomize.config.k8s.io/v1beta1
|
|
kind: Kustomization
|
|
|
|
namespace: tidaldb-cluster
|
|
|
|
resources:
|
|
- ../cluster
|
|
|
|
patches:
|
|
# GKE fact 1 — see header.
|
|
- target:
|
|
kind: StatefulSet
|
|
name: tidaldb
|
|
patch: |-
|
|
- op: replace
|
|
path: /spec/volumeClaimTemplates/0/spec/storageClassName
|
|
value: standard-rwo
|
|
|
|
# GKE fact 2 — drop the fleet-only public exposure. Deleted by exact kind and
|
|
# name so a base change that adds a fourth object fails loudly here instead
|
|
# of quietly publishing it.
|
|
#
|
|
# The public object is a plain `networking.k8s.io/v1` Ingress carrying Traefik
|
|
# ANNOTATIONS — NOT a `traefik.io` IngressRoute. An earlier version of this
|
|
# file targeted IngressRoute, which matched nothing: kustomize rendered
|
|
# happily and the Ingress survived, so `tidaldb.threesix.ai` would have been
|
|
# published from GKE. Rendering the overlay and reading the object inventory
|
|
# is what caught it, which is why that check is a runbook step below and not
|
|
# left to someone's memory.
|
|
- target:
|
|
group: networking.k8s.io
|
|
version: v1
|
|
kind: Ingress
|
|
name: tidaldb
|
|
patch: |-
|
|
$patch: delete
|
|
apiVersion: networking.k8s.io/v1
|
|
kind: Ingress
|
|
metadata:
|
|
name: tidaldb
|
|
- target:
|
|
group: traefik.io
|
|
kind: ServersTransport
|
|
name: tidaldb-internal
|
|
patch: |-
|
|
$patch: delete
|
|
apiVersion: traefik.io/v1alpha1
|
|
kind: ServersTransport
|
|
metadata:
|
|
name: tidaldb-internal
|
|
- target:
|
|
group: traefik.io
|
|
kind: Middleware
|
|
name: tidaldb-ratelimit
|
|
patch: |-
|
|
$patch: delete
|
|
apiVersion: traefik.io/v1alpha1
|
|
kind: Middleware
|
|
metadata:
|
|
name: tidaldb-ratelimit
|