CI / lint (pull_request) Successful in 1m45s
CI / k8s (pull_request) Successful in 9s
CI / build (pull_request) Successful in 1m41s
CI / unit (pull_request) Successful in 2m3s
CI / frontend (pull_request) Successful in 2m23s
CI / mutation (pull_request) Successful in 5m39s
CI / verify-stack (pull_request) Skipped
Records the decision #177 asked for, the other way round: the hypervisor has no inbound path and the labs Caddy already holds 80/443 and the wildcard cert, so the portals go through it over a reverse SSH tunnel instead of an in-cluster edge. Covers keycloakUrl, KC_PROXY_HEADERS and the optional demo OTP autofill. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
5.9 KiB
5.9 KiB
ADR-0035: The deployed stack is published through the existing labs Caddy
- Status: Accepted
- Date: 2026-09-25
- Deciders: Respellion engineering
- Slice: #177 — that issue proposed the opposite (an in-cluster Caddy edge); this ADR records why the host-side option won. Implemented in #179, #180 and #181.
Context
The stack deploys to a single-node Talos VM (ADR-0033, #175). Until now it was only usable
through five SSH port-forwards: the portals' OIDC flow uses PKCE, PKCE needs
crypto.subtle, and browsers expose that only in a secure context, meaning HTTPS or a
localhost origin. A NodePort on the VM's address is neither. We want a URL a demo
audience can simply open.
Three facts about where things run shape the answer:
- The Talos VM is a libvirt guest on a Fedora hypervisor in the office, behind NAT
with no public address. The only way in from outside is an existing reverse SSH tunnel
(
autossh-reverse-tunnel.service) into anopenssh-servercontainer on the labs server. - The labs server (public IP) already runs Caddy for
*.labs.respellion.tech, with the wildcard certificate (DNS-01 via Cloudflare) and ports 80/443. Every other labs service is published there (repoInfra,infra/development/). - #177 proposed a Caddy inside the cluster, fed by a layer-4 forward on the host, so that routing and certificates would be cluster state. That assumes the public IP is on the hypervisor. It isn't: the hypervisor has no inbound path, and 80/443 on the labs server are already taken by the labs Caddy.
Decision
Publish the portals and Keycloak through the existing labs Caddy. Carry the traffic to the cluster over a second reverse SSH tunnel from the hypervisor.
browser ─https─▶ labs Caddy ─▶ openssh-server:3014x/30180
─reverse SSH tunnel─▶ Fedora hypervisor ─▶ Talos NodePorts
- Hostnames under the existing wildcard:
big-register(openbaar),big-mijn(self-service),big-behandel,big-beheer, andbig-auth(Keycloak, with/admin*answered 404). - Tunnel:
big-portals-tunnel.serviceon the hypervisor (repoInfra) reverse-forwards the five browser-facing NodePorts intoopenssh-server. It is separate from the access tunnel on:6667, so a failed forward can't cut SSH access. Caddy joins theopenssh_defaultnetwork to reach the tunnel ends. - Keycloak's issuer is the public origin. The chart value
keycloakUrlreplaceshost+ NodePort in one helper,big.keycloakUrl, which feeds bothKC_HOSTNAMEand the portals'config.jsonauthority, so the two cannot drift (ADR-0010). The deploy workflow sets it from theKEYCLOAK_URLrepository variable. KC_PROXY_HEADERS=xforwarded:KC_HOSTNAME_BACKCHANNEL_DYNAMICbuilds the token, userinfo and certs URLs from the request. That request reaches Keycloak as plain HTTP, so the URLs came outhttp://and browsers blocked them as mixed content. Trusting Caddy'sX-Forwarded-Protokeeps them HTTPS. In-cluster calls send no such header and still usekeycloak:8080.- Demo MFA (optional):
demo.otpAutofill(OTP_AUTOFILL) makes thebig-demotheme (infra/keycloak/themes/big-demo) Keycloak's default. Its script fills in and submits the medewerker OTP from the fixture secret (ADR-0031), so the step is visibly enforced without an authenticator. It is off by default.
Alternatives considered
- In-cluster Caddy edge (#177, PR #178). It would keep routes and certificates in cluster state. But it needs a public inbound path to the hypervisor that doesn't exist, plus a second certificate authority beside the labs Caddy, which already holds the wildcard. Closed unmerged.
- Port-forward on the office router to the hypervisor. This opens the office network itself to the internet. Rejected.
- Move the cluster to a host with a public IP. It would remove the tunnel, but it's a bigger change than publishing one demo. It remains the natural step if the stack outgrows a lab VM.
- Keep the SSH port-forwards. Fine for one developer, but not something you can send to someone.
Consequences
Positive
- Real hostnames and HTTPS, so PKCE works in any browser with no client-side setup.
- No new certificate handling: the labs Caddy's wildcard covers the new hosts.
- The chart stays edge-agnostic. With
keycloakUrlempty it renders exactly as before, so compose, CI and thelocalhostworkflow are untouched.
Negative / costs
- Routing lives outside the cluster, in the Infra repo's Caddyfile. That is exactly what #177 wanted to avoid. Adding a portal means changing three places: a NodePort in the chart, a forward in the tunnel unit, and a host in the Caddyfile.
- Two SSH hops in the data path. If the hypervisor or the tunnel is down, the portals return 502 even though the cluster is healthy.
- One issuer string. With
keycloakUrlset, thelocalhostport-forward workflow (runbook §5) can no longer log in. - Keycloak trusts
X-Forwarded-*from anything that reaches it. Today that is only in-cluster callers and the tunnel.KC_PROXY_TRUSTED_ADDRESSEScan narrow it if the NodePort is ever exposed more widely. - The portals are public. Anyone with the link can log in with the committed test
credentials, and with
OTP_AUTOFILLon, no second factor stands in the way. That is acceptable for synthetic data. Put the labs Caddy's Azureauthorizein front of thebig-*hosts if the audience must be restricted.
Follow-up
- Runbook:
docs/runbooks/kubernetes-talos.md, "Publishing through the labs Caddy". - Dev-mode Keycloak generates new signing keys on every restart, and the BFF re-fetches them at most every 5 minutes, so expect a few minutes of 401s after a Keycloak restart. Persisting Keycloak's database (runbook §6) would remove that.