Four consecutive Kaniko pushes failed with
`dial tcp 208.122.204.172:443: connect: connection refused` while zot and
Traefik were both healthy and my laptop reached the same URL fine.
Measured from a pod in the same namespace, with the same labels the build Job
carries: 14 of 24 requests to https://registry.threesix.ai/v2/ succeeded, and
8 of 8 succeeded against Traefik's ClusterIP with the same SNI. So the flaky
leg is pod -> the cluster's own public address, and the build was taking it
for both the clone and the push.
The Job now carries hostAliases pinning git.threesix.ai and
registry.threesix.ai to Traefik's ClusterIP, read from the Service at release
time. The image is still named registry.threesix.ai/hush/api:SHA — the route
changes, the reference does not, and the kubelet pull path is untouched.
No retry loop: a retry would have made a 40% failure rate invisible instead of
fixed.
docs/DEPLOY.md no longer duplicates the Job YAML. That duplicate is why the
push guard drifted onto the wrong remote, and a second copy of a spec nobody
runs is worth less than a pointer to the one that runs.
Measured on the previous commit's release: the Kaniko Job reached Complete in
156s and pushed registry.threesix.ai/hush/api:d7cd57f3, and
`kubectl wait --for=condition=complete` sat on its watch until the 900s
timeout, then reported "build FAILED" and exited before the rollout. The
image existed; the deploy did not happen; the message said the opposite of
what had happened.
One long watch against a cluster across a WAN link is the wrong instrument. A
fresh short GET every 5s costs one poll when a connection drops, distinguishes
`.status.failed` from "not finished yet", and reports the deadline as "may
still be running" with the two commands to check — never as a failure.
Both branches exercised against the cluster with the loop as committed: a Job
that exits 1 is reported as FAILED with its last lines, and the already
complete d7cd57f3 Job breaks the loop immediately.
Using hush from an agent needed a clone and docs/MCP.md. It now needs one
command, and the instructions are served by the deployment itself.
`go install github.com/orchard9/hush/cmd/hush-mcp@latest` is the whole
install: cmd/hush-mcp imports only the standard library, so module graph
pruning never reaches the private go-chassis dependency cmd/hushd needs.
Verified against an empty module cache and the public proxy, then create ->
reveal end to end against production with the resulting binary.
The page carries the per-client configuration for Claude Code, Codex CLI,
Gemini CLI, VS Code, Claude Desktop, Cursor and omp. Each command was run
against the installed client rather than copied from documentation, which is
how the differences on it are there at all: VS Code's wrapper key is
`servers`, not `mcpServers`; gemini defaults to project scope, not user;
Claude Code rejects `--env` immediately before the server name.
The shared browser crypto moves from base.html into templates/crypto.html,
which the two pages that encrypt parse and this one does not. An empty
`{{define}}` cannot replace a non-empty one — text/template reads an empty
body as no definition — so the shell holds the call and the partial holds the
code, and the docs page ships no script at all.
Three things this exposed, fixed here:
- The public Ingress enumerates paths, so a handler without one 404s at the
edge while working in `make dev`. The Ingress is now its own manifest:
hush.yaml pins a `:bootstrap` image that does not exist, so re-applying it
to publish a path would roll the workload onto an unpullable image.
`make deploy-ingress` applies the route alone.
- release.sh guarded HEAD against `@{upstream}`, which is the GitHub mirror
here, while Kaniko clones Gitea. A commit pushed to one and not the other
would have built the previous commit silently. It now fetches and compares
the branch that actually gets built.
- smoke.sh checks that /mcp serves the install command, so a stale rollout or
an unexecutable template fails the release instead of being found later.
Confirmed it fails: against production before this deploy it reported 404.
docs/ARCHITECTURE.md records why internal/web.render replaces the chassis
JSON-API policy with a per-response nonce policy, and the failure each of the
three decisions prevents: a header-only policy because two policies on one
response intersect, a nonce instead of 'unsafe-inline' because the guarantee is
that only the reviewed same-document script reaches the fragment key, and a
fresh url-alphabet value because a reused nonce is worth 'unsafe-inline' to
anyone who waits for the next load and + or / would make enforcement depend on
entity decoding.
README.md now names the buttons the page actually renders and says outright
that there is no lifetime picker — the server's 24h default applies and
ttl_seconds is where a caller chooses.
base.html drops the opacity transition; nothing animates opacity.
.dockerignore is an allowlist, because the build stage COPYs only go.mod,
go.sum, vendor/, cmd/ and internal/. A blocklist forgets the file nobody
predicted, and for this service that file is a secret. .gitignore grows the
same protection for the working tree.
The Woodpecker claim in DEPLOY.md and release.sh was wrong within an hour of
being written. Corrected at the source rather than annotated:
* jordan/hush is active in Woodpecker (repo 139) with the Gitea webhook
installed, so a push to main builds and deploys.
* Activation takes the NUMERIC gitea repo id in forge_remote_id, not
owner/name — worth recording, because passing the wrong shape and passing a
dead token both come back 401 and look identical. GET /api/user separates
them: it is auth-only, so 401 there is the token and 200 there means the
request was the problem.
* The credential is $THREE_SIX_WOODPECKER (and $THREE_SIX_GITEA), in the
operator's environment.
* The stale copy that caused the original 401 is fixed where it lives:
k3sf-rdev-admin-key in GCP Secret Manager, property WOODPECKER_API_TOKEN,
every other property preserved. ESO resynced and the token read out of
rdev/rdev-credentials now answers 200. Documented alongside it: do not patch
that k8s Secret directly, it is ESO-owned and a direct edit is reverted on
the next refresh.
make release stays, with its reason updated — it is now the hotfix/rollback path
and the answer to "CI is down", rather than the only way to deploy.
Woodpecker is not activated for this repo — the WOODPECKER_API_TOKEN in
rdev-credentials returns 401 on /api/user, so it is the token and not the
request shape, and minting a new one needs a browser login I cannot do. That
left `git push` silently not deploying, which is a trap for whoever pushes next.
So `make release` does what the pipeline's build and deploy steps do: a Kaniko
Job for an amd64 image from the pushed git ref, `kubectl set image`, rollout,
then the production smoke. When Woodpecker is activated this becomes redundant,
and stays useful as the manual path for a hotfix or a rollback when CI is down.
Two refusals in it are the interesting part, both for failures that are
otherwise silent:
* A dirty or unpushed tree is refused. Kaniko builds from the GIT CONTEXT, not
the working tree, so uncommitted work would produce an image that does not
contain it while every log line says success.
* After the rollout it asserts the live image equals the one just built.
`kubectl set image` matching no container is silent, and the rollout then
"succeeds" against the old pod.
Written after the service was live, so every command and every number here was
run against the real deployment rather than assumed:
* DEPLOY.md records the things a rebuild needs and git does not hold — the
Redis ACL user (and why +getdel is the one to notice), the GCP Secret
Manager entry, the DNS record, and the three coordinated edits vmalert
needs because it has no ConfigMap auto-discovery.
* It also records two blockers rather than hiding them: Woodpecker is NOT
activated (the token in rdev-credentials returns 401), so pushes do not
deploy yet and the Kaniko Job is the interim path; and the host is
hush.threesix.ai rather than hush.orchard9.ai because orchard9.ai is on
GoDaddy and no GoDaddy credential exists anywhere I can reach.
* OPERATIONS.md is one section per alert, plus the failure modes that are not
alerts — chiefly that "gone" cannot distinguish already-revealed from
expired from LRU-evicted, on purpose, so the operator's default reading of
an unexpected "gone" is that the secret is compromised and should be
rotated.
* scripts/logs.sh and alerts-check.sh verify rather than assert:
alerts-check asks vmalert what it actually loaded AND checks each rule's
series exists, because a rule reading a metric nothing exports can never
fire and looks exactly like a healthy service.
* scripts/smoke.sh is a real client — it generates a key, encrypts, posts only
ciphertext, reveals, decrypts, then asserts the second reveal is 410, that
three GETs did not consume the secret, that missing and malformed ids are
indistinguishable, and that a plaintext field is refused.
install-mcp.sh proves the MCP handshake before writing any config, backs up
mcp.json, and rewrites only hush's entry — a config pointing at a broken server
surfaces as an opaque host-side connect failure, which is worth one extra check
to avoid.
Paste a secret, get a link, send it. The first person to open it and press
Reveal sees the secret; the link dies at that moment. The recipient needs a
browser and nothing else — no account, no client, no installed tooling.
The server cannot read what it stores. AES-256-GCM happens in the browser and
the key lives in the URL fragment, which browsers never transmit, so hushd
holds ciphertext and no key material. That is a property of where the key sits
rather than a promise about our conduct, which is why there is deliberately no
endpoint accepting a plaintext secret and no server-side-encryption fallback:
two guarantees behind one URL would be worse than one honest guarantee.
Three decisions carry the design:
* GET /s/{id} touches NO storage, not even to check existence. Slack, Teams,
WhatsApp, iMessage and Outlook Safe Links all fetch a URL before a human
sees it, so destroying on GET would destroy most secrets in transit and the
recipient's "already used" would be indistinguishable from interception.
Only POST /reveal consumes. Bot user-agent detection is an arms race;
removing the side effect from GET is not. Pinned by
TestGettingTheRevealPageNeverConsumesTheSecret.
* Destruction is one Redis GETDEL, which is atomic. GET-then-DEL has a window
where two simultaneous readers both win, and for a one-time secret that
window is the product. The store contract demands atomicity and the same
concurrency test runs against both implementations.
* Missing, already-revealed, expired and evicted are ONE indistinguishable
410. Separating them would confirm to a prober that a given link was real.
The secret id IS the capability, so secret.ID is a struct whose every
accidental path — %v, %s, String(), slog, json.Marshal — emits a redacted
handle or refuses, and the raw value needs an explicit Value(). The first
version tried to prevent leaks by implementing no String() at all; its own test
caught that Go's fmt prints unexported fields anyway, so forbidding the method
had removed the control rather than the leak.
Operationally: structured JSON on stdout in the fleet's wire format, which
Vector already collects with no annotation; six hush_* metrics on the chassis
registry with no id, IP or path in any label; five alert rules wired into
vmalert. The public Ingress enumerates /, /s/ and /api/ so /metrics, /healthz
and /readyz share the port but are unreachable from the internet — no
basic-auth middleware to maintain and get wrong.
Dependencies are vendored because go-chassis is private: the Woodpecker test
step and the in-cluster Kaniko build both run -mod=vendor with GOPROXY=off and
hold no git credential.
cmd/hush-mcp is a stdio MCP server doing the same client-side crypto locally,
so using hush from an agent preserves the same guarantee as using it from a
browser.