CI / k8s (push) Successful in 8s
CI / build (push) Successful in 1m38s
CI / lint (push) Successful in 1m55s
CI / mutation (push) Canceled after 0s
CI / verify-stack (push) Canceled after 0s
CI / frontend (push) Canceled after 12s
CI / unit (push) Canceled after 18s
Deploy to Talos / deploy (push) Successful in 2m30s
## What & why For the public demo on `big-behandel` / `big-beheer`, visitors should see MFA being enforced without needing an authenticator app. This adds an opt-in Keycloak theme that fills in and submits the medewerker OTP code itself. - **Theme as real files in `infra/keycloak/themes/big-demo/`**, next to the realms: - `login/theme.properties`: `keycloak.v2` plus `scripts=js/otp-autofill.js`. I checked the 26.1 source: `keycloak.v2` loads theme `scripts` and sets none of its own. - `login/resources/js/otp-autofill.js`: on the OTP page, computes the code (RFC 6238, Keycloak's default policy) from the fixture secret `BIGMEDEWERKEROTPSEED` and submits it. - `account`, `admin`, `email`: plain children of Keycloak 26's defaults. Without them the account console returns 500 (see notes). - **Seeded like every other file input:** `infra/helm/seed-configmaps.sh` creates the `rr-kc-theme` ConfigMap, and the chart mounts it as a directory. The podspec gains `items` so flat ConfigMap keys map to theme paths. Keycloak runs `start-dev` (no theme cache), so edits show up about a minute after a reseed. - **Switch:** `demo.otpAutofill` only decides whether `KC_SPI_THEME_DEFAULT=big-demo` is set. `big.env` now skips env values that render empty, and no existing env var is empty. **Off, the render is identical to main except for that one missing variable,** so Keycloak keeps its stock theme. The realm JSONs are untouched, so compose and the e2e tests still require a code. - **Single-use codes:** a second login in the same 30 s window spends the next counter, as `nextUnusedCounter` does in the e2e. Past that it only fills in the field and doesn't submit, so a rejected code can't loop. - **Deploy workflow:** repo variable `OTP_AUTOFILL=true` → `--set demo.otpAutofill=true`. Flipping it changes the pod's env, so Keycloak restarts. Refs #177 ## Definition of Done - [x] Linked Gitea issue (above). - [ ] Failing test committed before the implementation. *(Not done; checks below.)* - [x] Conventional Commits referencing the issue (`refs #NN`). - [ ] CI green - [x] `docker compose up` unaffected (chart only). - [x] Docs updated (Talos runbook, "Publishing through the labs Caddy"). - [ ] ADR. The fixture-secret trade-off is ADR-0031's; this only automates typing it in. ## Notes for reviewers - **Tested on the live cluster.** I patched the running Keycloak with the rendered theme (autofill on) and ran real headless Chromium logins against the public hosts: - `merel-behandelaar` on big-behandel: only username and password typed. The OTP page loaded the script, submitted by itself, and the user landed in the Werkbak. - `jan-burger` on big-mijn still logs in (regression check). - `/realms/medewerker/account/` returns 200. - **Account console 500, found live and fixed in the second commit.** `KC_SPI_THEME_DEFAULT` applies to every theme type, and Keycloak does *not* fall back for a type the theme lacks (`NullPointerException ... "theme" is null`). `big-demo` now declares login, account, admin and email, each a plain child of Keycloak 26's default. It's one ConfigMap mounted as a directory; the podspec gains `items` for that. - **Keycloak restarts cause about 5 minutes of BFF 401s.** This is not caused by this PR, but you'll see it whenever Keycloak restarts. Dev-mode Keycloak makes new signing keys on each boot, and the BFF refreshes its cached keys at most every 5 minutes. Seen live: 401 right after the restart, 204 about 4½ minutes later. Flipping `OTP_AUTOFILL` restarts Keycloak, so expect this briefly. - `make k8s-lint` and `make k8s-drift` pass. The rendered script's code matches `infra/keycloak/check_realms.py otp`. - **Security:** with it on, the public behandel and beheer portals are protected only by the committed password `test123`. That's intentional for synthetic demo data. Never enable it anywhere real. 🤖 Generated with [Claude Code](https://claude.com/claude-code)Reviewed-on: #181
122 lines
5.5 KiB
YAML
122 lines
5.5 KiB
YAML
name: Deploy to Talos
|
|
|
|
# A merge to main ships the stack to the Talos cluster on the lab server
|
|
# (docs/runbooks/kubernetes-talos.md §9). PR CI is the merge gate, so main is
|
|
# green by construction — this workflow only deploys.
|
|
on:
|
|
push:
|
|
branches: [main]
|
|
workflow_dispatch:
|
|
|
|
permissions:
|
|
contents: read
|
|
|
|
# Queue deploys, never cancel one: a helm upgrade killed half-way leaves the
|
|
# release in `pending-upgrade` and the next run has to be unwedged by hand.
|
|
concurrency:
|
|
group: deploy-talos
|
|
cancel-in-progress: false
|
|
|
|
jobs:
|
|
deploy:
|
|
runs-on: ubuntu-latest
|
|
env:
|
|
# The Talos VM as seen from the Fedora host (libvirt guest IP), and the
|
|
# address a browser uses to reach the cluster. `localhost` is deliberate:
|
|
# the portals' PKCE needs a secure context, so they are reached over
|
|
# `kubectl port-forward` — runbook §5. Override with repo variables.
|
|
TALOS_VM_IP: ${{ vars.TALOS_VM_IP }}
|
|
TALOS_HOST: ${{ vars.TALOS_HOST }}
|
|
# Set it when the labs Caddy publishes the portals: Keycloak's public https
|
|
# origin, e.g. https://big-auth.labs.respellion.tech (runbook, "Publishing
|
|
# through the labs Caddy").
|
|
KEYCLOAK_URL: ${{ vars.KEYCLOAK_URL }}
|
|
# `true` fills in the medewerker OTP step for the public demo (chart value
|
|
# demo.otpAutofill). The fixture secret is committed: demo only.
|
|
OTP_AUTOFILL: ${{ vars.OTP_AUTOFILL }}
|
|
steps:
|
|
- uses: https://github.com/actions/checkout@v4
|
|
|
|
# Pinned static binaries, the same URLs the Talos runbook §0 gives a
|
|
# developer and the same helm the `k8s` CI job uses — no action to vet.
|
|
- name: Install kubectl, helm and crane
|
|
run: |
|
|
set -euo pipefail
|
|
bin="$HOME/.local/bin"; mkdir -p "$bin"
|
|
curl -sSLo "$bin/kubectl" https://dl.k8s.io/release/v1.37.0/bin/linux/amd64/kubectl
|
|
curl -sSL https://get.helm.sh/helm-v3.16.4-linux-amd64.tar.gz | tar xz -O linux-amd64/helm > "$bin/helm"
|
|
curl -sSL https://github.com/google/go-containerregistry/releases/download/v0.20.2/go-containerregistry_Linux_x86_64.tar.gz | tar xz -O crane > "$bin/crane"
|
|
chmod +x "$bin"/{kubectl,helm,crane}
|
|
echo "$bin" >> "$GITHUB_PATH"
|
|
|
|
# The cluster's API and its registry are only reachable through the Fedora
|
|
# host, so forward both to the runner. 30141 is the openbaar portal, for
|
|
# the smoke at the end.
|
|
- name: Tunnel the Talos API + registry through the Fedora host
|
|
env:
|
|
SSH_KEY: ${{ secrets.TALOS_SSH_KEY }}
|
|
run: |
|
|
set -euo pipefail
|
|
: "${TALOS_VM_IP:=192.168.122.173}"
|
|
umask 077
|
|
printf '%s\n' "$SSH_KEY" > ~/.ssh_talos
|
|
ssh -i ~/.ssh_talos -o StrictHostKeyChecking=no -o IdentitiesOnly=yes \
|
|
-o ExitOnForwardFailure=yes -p 6667 -f -N \
|
|
-L 6443:$TALOS_VM_IP:6443 \
|
|
-L 30500:$TALOS_VM_IP:30500 \
|
|
-L 30141:$TALOS_VM_IP:30141 \
|
|
user@labs.respellion.tech
|
|
|
|
# The kubeconfig's server must be https://127.0.0.1:6443 — Talos puts
|
|
# 127.0.0.1 in the apiserver cert SANs, so TLS verification still holds
|
|
# through the tunnel.
|
|
- name: Write the kubeconfig
|
|
env:
|
|
KUBECONFIG_B64: ${{ secrets.TALOS_KUBECONFIG }}
|
|
run: |
|
|
set -euo pipefail
|
|
umask 077
|
|
base64 -d <<< "$KUBECONFIG_B64" > "$RUNNER_TEMP/kubeconfig"
|
|
echo "KUBECONFIG=$RUNNER_TEMP/kubeconfig" >> "$GITHUB_ENV"
|
|
kubectl --kubeconfig "$RUNNER_TEMP/kubeconfig" get nodes
|
|
|
|
# Idempotent; also makes a first deploy onto a bare cluster work. The
|
|
# registry's storage is an emptyDir, so a replaced pod loses the images —
|
|
# which the push in the next step puts back anyway.
|
|
- name: Ensure the in-cluster registry
|
|
run: make k8s-registry
|
|
|
|
# Push through the tunnel (localhost), pull from the node's own NodePort
|
|
# (the address in the Talos registry-mirror patch) — same registry, two
|
|
# names, so the two `make` calls get different K8S_REGISTRY values.
|
|
- name: Build and push the images
|
|
run: make k8s-images K8S_REGISTRY=localhost:30500
|
|
|
|
# k8s-reseed = seed configmaps + helm upgrade + re-run the bootstrap jobs.
|
|
# The jobs are idempotent, and deleting them first is what keeps a changed
|
|
# Job template from wedging the upgrade (`cannot patch … with kind Job`).
|
|
- name: Deploy the chart
|
|
run: |
|
|
make k8s-reseed \
|
|
TALOS_HOST=${TALOS_HOST:-localhost} \
|
|
K8S_REGISTRY=${TALOS_VM_IP:-192.168.122.173}:30500 \
|
|
K8S_SET="${KEYCLOAK_URL:+--set keycloakUrl=$KEYCLOAK_URL} --set demo.otpAutofill=${OTP_AUTOFILL:-false}"
|
|
|
|
# `dev` is a mutable tag and helm sees an unchanged pod template, so the
|
|
# new images only land on a restart (pullPolicy is already Always).
|
|
- name: Roll the services onto the new images
|
|
run: |
|
|
set -euo pipefail
|
|
svcs="acl domain bff event-subscriber projection-api self-service openbaar behandel beheer"
|
|
kubectl -n big rollout restart deploy $svcs
|
|
kubectl -n big rollout status --timeout=300s deploy $svcs
|
|
|
|
# Proves portal → Caddy → BFF → projection end to end. An empty register is
|
|
# a pass; a 502 or a timeout is not.
|
|
- name: Smoke the public register
|
|
run: curl -fsS --retry 10 --retry-delay 6 --retry-all-errors http://localhost:30141/openbaar/register
|
|
|
|
- name: Pods on failure
|
|
if: failure()
|
|
run: kubectl -n big get pods,jobs || true
|