Files
register-referentie/docs/architecture/adr-0035-public-access-through-the-labs-caddy.md
T
notandClaude Opus 5.5 84cea6267e
CI / lint (pull_request) Successful in 1m45s
CI / k8s (pull_request) Successful in 9s
CI / build (pull_request) Successful in 1m41s
CI / unit (pull_request) Successful in 2m3s
CI / frontend (pull_request) Successful in 2m23s
CI / mutation (pull_request) Successful in 5m39s
CI / verify-stack (pull_request) Skipped
docs(arch): ADR-0035 — publish the stack through the existing labs Caddy (closes #177)
Records the decision #177 asked for, the other way round: the hypervisor has no
inbound path and the labs Caddy already holds 80/443 and the wildcard cert, so
the portals go through it over a reverse SSH tunnel instead of an in-cluster
edge. Covers keycloakUrl, KC_PROXY_HEADERS and the optional demo OTP autofill.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-25 15:05:38 +02:00

5.9 KiB

ADR-0035: The deployed stack is published through the existing labs Caddy

  • Status: Accepted
  • Date: 2026-09-25
  • Deciders: Respellion engineering
  • Slice: #177 — that issue proposed the opposite (an in-cluster Caddy edge); this ADR records why the host-side option won. Implemented in #179, #180 and #181.

Context

The stack deploys to a single-node Talos VM (ADR-0033, #175). Until now it was only usable through five SSH port-forwards: the portals' OIDC flow uses PKCE, PKCE needs crypto.subtle, and browsers expose that only in a secure context, meaning HTTPS or a localhost origin. A NodePort on the VM's address is neither. We want a URL a demo audience can simply open.

Three facts about where things run shape the answer:

  • The Talos VM is a libvirt guest on a Fedora hypervisor in the office, behind NAT with no public address. The only way in from outside is an existing reverse SSH tunnel (autossh-reverse-tunnel.service) into an openssh-server container on the labs server.
  • The labs server (public IP) already runs Caddy for *.labs.respellion.tech, with the wildcard certificate (DNS-01 via Cloudflare) and ports 80/443. Every other labs service is published there (repo Infra, infra/development/).
  • #177 proposed a Caddy inside the cluster, fed by a layer-4 forward on the host, so that routing and certificates would be cluster state. That assumes the public IP is on the hypervisor. It isn't: the hypervisor has no inbound path, and 80/443 on the labs server are already taken by the labs Caddy.

Decision

Publish the portals and Keycloak through the existing labs Caddy. Carry the traffic to the cluster over a second reverse SSH tunnel from the hypervisor.

browser ─https─▶ labs Caddy ─▶ openssh-server:3014x/30180
        ─reverse SSH tunnel─▶ Fedora hypervisor ─▶ Talos NodePorts
  • Hostnames under the existing wildcard: big-register (openbaar), big-mijn (self-service), big-behandel, big-beheer, and big-auth (Keycloak, with /admin* answered 404).
  • Tunnel: big-portals-tunnel.service on the hypervisor (repo Infra) reverse-forwards the five browser-facing NodePorts into openssh-server. It is separate from the access tunnel on :6667, so a failed forward can't cut SSH access. Caddy joins the openssh_default network to reach the tunnel ends.
  • Keycloak's issuer is the public origin. The chart value keycloakUrl replaces host + NodePort in one helper, big.keycloakUrl, which feeds both KC_HOSTNAME and the portals' config.json authority, so the two cannot drift (ADR-0010). The deploy workflow sets it from the KEYCLOAK_URL repository variable.
  • KC_PROXY_HEADERS=xforwarded: KC_HOSTNAME_BACKCHANNEL_DYNAMIC builds the token, userinfo and certs URLs from the request. That request reaches Keycloak as plain HTTP, so the URLs came out http:// and browsers blocked them as mixed content. Trusting Caddy's X-Forwarded-Proto keeps them HTTPS. In-cluster calls send no such header and still use keycloak:8080.
  • Demo MFA (optional): demo.otpAutofill (OTP_AUTOFILL) makes the big-demo theme (infra/keycloak/themes/big-demo) Keycloak's default. Its script fills in and submits the medewerker OTP from the fixture secret (ADR-0031), so the step is visibly enforced without an authenticator. It is off by default.

Alternatives considered

  • In-cluster Caddy edge (#177, PR #178). It would keep routes and certificates in cluster state. But it needs a public inbound path to the hypervisor that doesn't exist, plus a second certificate authority beside the labs Caddy, which already holds the wildcard. Closed unmerged.
  • Port-forward on the office router to the hypervisor. This opens the office network itself to the internet. Rejected.
  • Move the cluster to a host with a public IP. It would remove the tunnel, but it's a bigger change than publishing one demo. It remains the natural step if the stack outgrows a lab VM.
  • Keep the SSH port-forwards. Fine for one developer, but not something you can send to someone.

Consequences

Positive

  • Real hostnames and HTTPS, so PKCE works in any browser with no client-side setup.
  • No new certificate handling: the labs Caddy's wildcard covers the new hosts.
  • The chart stays edge-agnostic. With keycloakUrl empty it renders exactly as before, so compose, CI and the localhost workflow are untouched.

Negative / costs

  • Routing lives outside the cluster, in the Infra repo's Caddyfile. That is exactly what #177 wanted to avoid. Adding a portal means changing three places: a NodePort in the chart, a forward in the tunnel unit, and a host in the Caddyfile.
  • Two SSH hops in the data path. If the hypervisor or the tunnel is down, the portals return 502 even though the cluster is healthy.
  • One issuer string. With keycloakUrl set, the localhost port-forward workflow (runbook §5) can no longer log in.
  • Keycloak trusts X-Forwarded-* from anything that reaches it. Today that is only in-cluster callers and the tunnel. KC_PROXY_TRUSTED_ADDRESSES can narrow it if the NodePort is ever exposed more widely.
  • The portals are public. Anyone with the link can log in with the committed test credentials, and with OTP_AUTOFILL on, no second factor stands in the way. That is acceptable for synthetic data. Put the labs Caddy's Azure authorize in front of the big-* hosts if the audience must be restricted.

Follow-up

  • Runbook: docs/runbooks/kubernetes-talos.md, "Publishing through the labs Caddy".
  • Dev-mode Keycloak generates new signing keys on every restart, and the BFF re-fetches them at most every 5 minutes, so expect a few minutes of 401s after a Keycloak restart. Persisting Keycloak's database (runbook §6) would remove that.