Four consecutive Kaniko pushes failed with
`dial tcp 208.122.204.172:443: connect: connection refused` while zot and
Traefik were both healthy and my laptop reached the same URL fine.
Measured from a pod in the same namespace, with the same labels the build Job
carries: 14 of 24 requests to https://registry.threesix.ai/v2/ succeeded, and
8 of 8 succeeded against Traefik's ClusterIP with the same SNI. So the flaky
leg is pod -> the cluster's own public address, and the build was taking it
for both the clone and the push.
The Job now carries hostAliases pinning git.threesix.ai and
registry.threesix.ai to Traefik's ClusterIP, read from the Service at release
time. The image is still named registry.threesix.ai/hush/api:SHA — the route
changes, the reference does not, and the kubelet pull path is untouched.
No retry loop: a retry would have made a 40% failure rate invisible instead of
fixed.
docs/DEPLOY.md no longer duplicates the Job YAML. That duplicate is why the
push guard drifted onto the wrong remote, and a second copy of a spec nobody
runs is worth less than a pointer to the one that runs.
Measured on the previous commit's release: the Kaniko Job reached Complete in
156s and pushed registry.threesix.ai/hush/api:d7cd57f3, and
`kubectl wait --for=condition=complete` sat on its watch until the 900s
timeout, then reported "build FAILED" and exited before the rollout. The
image existed; the deploy did not happen; the message said the opposite of
what had happened.
One long watch against a cluster across a WAN link is the wrong instrument. A
fresh short GET every 5s costs one poll when a connection drops, distinguishes
`.status.failed` from "not finished yet", and reports the deadline as "may
still be running" with the two commands to check — never as a failure.
Both branches exercised against the cluster with the loop as committed: a Job
that exits 1 is reported as FAILED with its last lines, and the already
complete d7cd57f3 Job breaks the loop immediately.
Using hush from an agent needed a clone and docs/MCP.md. It now needs one
command, and the instructions are served by the deployment itself.
`go install github.com/orchard9/hush/cmd/hush-mcp@latest` is the whole
install: cmd/hush-mcp imports only the standard library, so module graph
pruning never reaches the private go-chassis dependency cmd/hushd needs.
Verified against an empty module cache and the public proxy, then create ->
reveal end to end against production with the resulting binary.
The page carries the per-client configuration for Claude Code, Codex CLI,
Gemini CLI, VS Code, Claude Desktop, Cursor and omp. Each command was run
against the installed client rather than copied from documentation, which is
how the differences on it are there at all: VS Code's wrapper key is
`servers`, not `mcpServers`; gemini defaults to project scope, not user;
Claude Code rejects `--env` immediately before the server name.
The shared browser crypto moves from base.html into templates/crypto.html,
which the two pages that encrypt parse and this one does not. An empty
`{{define}}` cannot replace a non-empty one — text/template reads an empty
body as no definition — so the shell holds the call and the partial holds the
code, and the docs page ships no script at all.
Three things this exposed, fixed here:
- The public Ingress enumerates paths, so a handler without one 404s at the
edge while working in `make dev`. The Ingress is now its own manifest:
hush.yaml pins a `:bootstrap` image that does not exist, so re-applying it
to publish a path would roll the workload onto an unpullable image.
`make deploy-ingress` applies the route alone.
- release.sh guarded HEAD against `@{upstream}`, which is the GitHub mirror
here, while Kaniko clones Gitea. A commit pushed to one and not the other
would have built the previous commit silently. It now fetches and compares
the branch that actually gets built.
- smoke.sh checks that /mcp serves the install command, so a stale rollout or
an unexecutable template fails the release instead of being found later.
Confirmed it fails: against production before this deploy it reported 404.
The Woodpecker claim in DEPLOY.md and release.sh was wrong within an hour of
being written. Corrected at the source rather than annotated:
* jordan/hush is active in Woodpecker (repo 139) with the Gitea webhook
installed, so a push to main builds and deploys.
* Activation takes the NUMERIC gitea repo id in forge_remote_id, not
owner/name — worth recording, because passing the wrong shape and passing a
dead token both come back 401 and look identical. GET /api/user separates
them: it is auth-only, so 401 there is the token and 200 there means the
request was the problem.
* The credential is $THREE_SIX_WOODPECKER (and $THREE_SIX_GITEA), in the
operator's environment.
* The stale copy that caused the original 401 is fixed where it lives:
k3sf-rdev-admin-key in GCP Secret Manager, property WOODPECKER_API_TOKEN,
every other property preserved. ESO resynced and the token read out of
rdev/rdev-credentials now answers 200. Documented alongside it: do not patch
that k8s Secret directly, it is ESO-owned and a direct edit is reverted on
the next refresh.
make release stays, with its reason updated — it is now the hotfix/rollback path
and the answer to "CI is down", rather than the only way to deploy.
Woodpecker is not activated for this repo — the WOODPECKER_API_TOKEN in
rdev-credentials returns 401 on /api/user, so it is the token and not the
request shape, and minting a new one needs a browser login I cannot do. That
left `git push` silently not deploying, which is a trap for whoever pushes next.
So `make release` does what the pipeline's build and deploy steps do: a Kaniko
Job for an amd64 image from the pushed git ref, `kubectl set image`, rollout,
then the production smoke. When Woodpecker is activated this becomes redundant,
and stays useful as the manual path for a hotfix or a rollback when CI is down.
Two refusals in it are the interesting part, both for failures that are
otherwise silent:
* A dirty or unpushed tree is refused. Kaniko builds from the GIT CONTEXT, not
the working tree, so uncommitted work would produce an image that does not
contain it while every log line says success.
* After the rollout it asserts the live image equals the one just built.
`kubectl set image` matching no container is silent, and the rollout then
"succeeds" against the old pod.