## What & why Two changes, made and verified together on a real cluster. **S-24 / #25 — a Helm chart for the platform.** One chart, `infra/helm/big-reference`, whose `values.yaml` is a near-literal transcription of `infra/docker-compose.yml`, rendered by three generic templates (Deployment, Job, Service) over a `workloads` map. Adding a service is a values edit. `make k8s-lint` renders and schema-checks the whole stack without a cluster. The issue asked for a *sketch*; this is deployed and verified end to end (see below), which is more than it asked for — the part it asked for that is **not** here is the production-posture write-up (HA, secrets, backup), see Known gaps. **#166 — Caddy replaces nginx in the portals.** nginx resolves a variable `proxy_pass` upstream itself, using only the `resolver` directive and never `/etc/resolv.conf`'s search domains. That had cost two workarounds in one script: rewriting the resolver address for rootless podman, and injecting a full FQDN so the bare `bff` name could resolve on Kubernetes. Caddy dials per request through the system resolver, so `reverse_proxy bff:8080` works on every engine unchanged; `apps/portal-nginx-resolver.sh` and the chart's `BFF_HOST` env are deleted. Closes #25 Closes #166 ## Definition of Done - [x] Linked Gitea issue (above). - [x] Failing test committed before the implementation — twice: the Caddyfile contract test before the Caddyfiles, `make k8s-lint` before the chart. - [x] Implementation makes the test pass. - [x] Conventional Commits referencing the issues (`refs #25` / `refs #166`). - [ ] CI green — awaiting the run on this PR (`make k8s-lint`, `dotnet format` and the new unit self-check pass locally; the compose e2e and mutation lanes are CI's). - [ ] `docker compose up` from a fresh clone reaches green health checks within 3 minutes — the portal images were rebuilt and verified standalone, but a full `make up` run has not been done on this branch. Please confirm in review or let CI's smoke test speak. - [x] Docs updated — `docs/runbooks/kubernetes-talos.md` (new), `frontend-decisions.md`, `demo-script.md`, and the docs that named nginx. - [x] ADR added — ADR-0033 (chart) and ADR-0034 (Caddy). - [ ] Demo note in `docs/demo-script.md` — not added: the deployment target is not a user-visible slice, and the Caddy swap is invisible to the demo script beyond the wording fix included here. ## How it was verified Brought up from scratch on a single-node Talos v1.14.0 VM (6 vCPU / 10 GB, virtio disk) under virt-manager: **29 pods ready and four bootstrap Jobs complete in under three minutes, zero restarts**, using ~4.4 GB of the VM's 10 GB. - Full Common Ground path: portal Caddy → BFF → domain → Flowable → ACL → OpenZaak + Objecten → NRC → event-subscriber → projection → public register (`INGEDIEND`, reference matching the submitted registration). - Werkbak read with an MFA'd medewerker token → 200. - The browser flow driven with Playwright against `http://localhost:30140`: secure context, `crypto.subtle` present, Keycloak form reached, login completed, **no console errors**. - Routing checked against a stub BFF: SPA fallback serves deep links, each portal proxies its own groups, and a portal does *not* proxy a neighbour's group. ## Notes for reviewers Three bugs this shook out, each fixed at the cause rather than the symptom: 1. **`command` vs `args`.** Compose's `command:` replaces the image CMD; Kubernetes' replaces the ENTRYPOINT. Transcribing one to the other broke every upstream image that relies on its entrypoint — postgres refused to run as root, Keycloak tried to exec `start-dev`. The chart now `fail`s at render time on `command`. 2. **Concurrent migrations.** Both `/setup_configuration.sh` and `/start.sh` run `manage.py migrate`; compose serialises them with `depends_on`, Kubernetes has no such edge, so the init Job and its web pod raced (`relation "zgw_consumers_service" already exists`). The four Django services now do both steps in order in the web pod — which also deletes four workloads. 3. **`emptyDir` databases are wiped by any pod-template change.** `make k8s-reseed` now also restarts `event-subscriber` and `projection-api`, which create the projection schema on start and otherwise keep writing to a schema-less database. Known gaps / follow-ups: - **Secrets.** `values.yaml` carries the dev credentials in plain text (`admin/admin`, the ZGW client secret, the two Objecten tokens) and the chart has no `Secret` objects. Fine for a laptop demo, and exactly what #25's "production posture" ADR should address — I suggest a follow-up issue rather than stretching this PR. - **No CI gate for the chart yet.** `make k8s-lint` exists but is not wired into `.gitea/workflows/ci.yaml`, and nothing enforces that the chart and the compose file stay in step. Worth a small follow-up. - **This is two slices in one PR.** They were built and verified together and the diff is entangled (the chart was written against Caddy from the start), so splitting now would mean re-creating an nginx-shaped chart to throw away. Happy to split if you'd rather. - **Rebased onto #161** (merged as #165) rather than merged, to keep the history linear. One conflict, in the `unit:` target where both branches add a self-check line — resolved by keeping both. #161's `infra/host-browser.yml` arrived with `/usr/share/nginx/html/config.json` and is fixed to `/usr/share/caddy/` inside the `feat(portals)` commit, so no commit on this branch leaves that overlay pointing at a path the images no longer have.Reviewed-on: #167
This commit was merged in pull request #167.
This commit is contained in:
@@ -43,7 +43,7 @@ export DOCKER_HOST := unix://$(PODMAN_SOCK)
|
||||
endif
|
||||
endif
|
||||
|
||||
.PHONY: ci lint build unit mutation frontend integration verify verify-up verify-acl verify-nrc verify-projection verify-bff verify-domain verify-observability verify-tracing verify-metrics verify-objecttypen verify-objecten verify-registerrecord verify-objecten-notifications verify-notifications smoke up down local verify-local local-down changelog openzaak-up openzaak-smoke openzaak-seed openzaak-down stack-up stack-smoke stack-down keycloak-up keycloak-smoke keycloak-down flowable-up flowable-smoke flowable-down help
|
||||
.PHONY: ci lint build unit mutation frontend integration verify verify-up verify-acl verify-nrc verify-projection verify-bff verify-domain verify-observability verify-tracing verify-metrics verify-objecttypen verify-objecten verify-registerrecord verify-objecten-notifications verify-notifications smoke up down local verify-local local-down changelog openzaak-up openzaak-smoke openzaak-seed openzaak-down stack-up stack-smoke stack-down keycloak-up keycloak-smoke keycloak-down flowable-up flowable-smoke flowable-down k8s-lint k8s-registry k8s-images k8s-seed k8s-up k8s-reseed k8s-portals k8s-down k8s-purge help
|
||||
|
||||
## ci: run the full pipeline — lint, build, unit, mutation, frontend, verify (mirrors Gitea Actions)
|
||||
## `verify` is the live-stack stage (full stack up once → ACL + notification checks).
|
||||
@@ -76,6 +76,7 @@ build:
|
||||
unit:
|
||||
dotnet test $(SLN) -c Release --filter "Category!=Integration" --logger trx --results-directory TestResults
|
||||
python3 infra/test_playwright_summary.py
|
||||
python3 infra/test_portal_caddyfiles.py
|
||||
|
||||
## mutation: run the Stryker.NET ratchet on each service with branching logic (fails below baseline)
|
||||
# Stryker is pinned as a local dotnet tool (.config/dotnet-tools.json); `tool restore`
|
||||
@@ -329,6 +330,91 @@ flowable-down:
|
||||
docker compose -f $(FL_COMPOSE) down --volumes
|
||||
-docker volume rm -f rr-fl-bpmn
|
||||
|
||||
# ── Kubernetes (single-node Talos) ─────────────────────────────────────────────
|
||||
# The Helm chart in infra/helm/big-reference is a port of infra/docker-compose.yml
|
||||
# (ADR-0033). Full walkthrough: docs/runbooks/kubernetes-talos.md.
|
||||
# TALOS_HOST the address the BROWSER uses — pins Keycloak's issuer and the portals'
|
||||
# OIDC authority. Use `localhost` with `make k8s-portals`: the OIDC
|
||||
# library needs crypto.subtle, which browsers only expose on a secure
|
||||
# context (https, or localhost) — see docs/runbooks/kubernetes-talos.md §5
|
||||
# K8S_REGISTRY the registry both sides use for this repo's images (see k8s-registry)
|
||||
K8S_NS ?= big
|
||||
K8S_CHART := infra/helm/big-reference
|
||||
K8S_REGISTRY ?=
|
||||
TALOS_HOST ?=
|
||||
# The images built from this repo — compose service name == image name == chart workload.
|
||||
K8S_IMAGES := acl domain bff event-subscriber projection-api self-service openbaar behandel beheer
|
||||
|
||||
## k8s-lint: render + schema-check the Helm chart (no cluster needed)
|
||||
k8s-lint:
|
||||
helm lint $(K8S_CHART)
|
||||
helm template big $(K8S_CHART) -n $(K8S_NS) --set images.registry=registry.invalid:5000 >/dev/null
|
||||
|
||||
## k8s-registry: deploy the in-cluster image registry (NodePort 30500)
|
||||
k8s-registry:
|
||||
kubectl apply -f infra/helm/registry.yaml
|
||||
kubectl -n registry rollout status deploy/registry --timeout=180s
|
||||
|
||||
## k8s-images: build this repo's images (via compose) and push them to $(K8S_REGISTRY)
|
||||
# `docker save | crane push` rather than `docker push`: the registry speaks plain
|
||||
# HTTP, which the Docker daemon refuses without a root-level insecure-registries
|
||||
# entry, while crane just takes --insecure. Install: see docs/runbooks/kubernetes-talos.md.
|
||||
k8s-images:
|
||||
@command -v crane >/dev/null || { echo "crane not found — see docs/runbooks/kubernetes-talos.md §0" >&2; exit 2; }
|
||||
@test -n "$(K8S_REGISTRY)" || { echo "set K8S_REGISTRY=<registry host:port>" >&2; exit 2; }
|
||||
docker compose -f $(COMPOSE) build $(K8S_IMAGES)
|
||||
@tar=$$(mktemp -t rr-img-XXXX.tar); \
|
||||
for i in $(K8S_IMAGES); do \
|
||||
docker save register-referentie/$$i:dev -o $$tar; \
|
||||
crane push --insecure $$tar $(K8S_REGISTRY)/register-referentie/$$i:dev; \
|
||||
done; rm -f $$tar
|
||||
|
||||
## k8s-seed: create the ConfigMaps the chart mounts (upstream config + bootstrap scripts)
|
||||
k8s-seed:
|
||||
bash infra/helm/seed-configmaps.sh $(K8S_NS)
|
||||
|
||||
## k8s-up: seed the config and install/upgrade the release
|
||||
k8s-up: k8s-seed
|
||||
@test -n "$(TALOS_HOST)" || { echo "set TALOS_HOST=<node ip>" >&2; exit 2; }
|
||||
@test -n "$(K8S_REGISTRY)" || { echo "set K8S_REGISTRY=<registry the node can pull from>" >&2; exit 2; }
|
||||
helm upgrade --install big $(K8S_CHART) -n $(K8S_NS) --create-namespace \
|
||||
--set host=$(TALOS_HOST) --set images.registry=$(K8S_REGISTRY) $(K8S_SET)
|
||||
kubectl -n $(K8S_NS) get pods
|
||||
|
||||
## k8s-reseed: re-run the bootstrap jobs (after a database was wiped, or after
|
||||
## changing a Job in the chart — Job pod templates are immutable, so a plain
|
||||
## `helm upgrade` is rejected)
|
||||
k8s-reseed:
|
||||
kubectl -n $(K8S_NS) delete job -l app.kubernetes.io/component=init --ignore-not-found
|
||||
$(MAKE) k8s-up
|
||||
# The projection's schema is created on service start (Projection.ReadModel migrates in a
|
||||
# hosted service), so a wiped database also needs these two restarted — otherwise they keep
|
||||
# writing to a schema-less DB and fail with `relation "processed_notifications" does not exist`.
|
||||
kubectl -n $(K8S_NS) rollout restart deploy/event-subscriber deploy/projection-api
|
||||
kubectl -n $(K8S_NS) rollout status deploy/event-subscriber deploy/projection-api --timeout=180s
|
||||
|
||||
## k8s-portals: forward the browser-facing services to localhost (Ctrl-C stops them all)
|
||||
# The portals' OIDC flow needs a *secure context* for crypto.subtle (PKCE), and browsers
|
||||
# only grant that to https or localhost — a NodePort on the VM's IP is neither. Forwarding
|
||||
# to localhost on the same port numbers keeps Keycloak's pinned issuer valid. Deploy with
|
||||
# TALOS_HOST=localhost for this to line up.
|
||||
k8s-portals:
|
||||
@echo "self-service http://localhost:30140 · openbaar :30141 · behandel :30142 · beheer :30143 · keycloak :30180"
|
||||
@trap 'kill 0' INT TERM; \
|
||||
for f in self-service:30140:80 openbaar:30141:80 behandel:30142:80 beheer:30143:80 keycloak:30180:8080; do \
|
||||
svc=$${f%%:*}; rest=$${f#*:}; lport=$${rest%%:*}; rport=$${rest#*:}; \
|
||||
kubectl -n $(K8S_NS) port-forward --address 127.0.0.1 svc/$$svc $$lport:$$rport >/dev/null & \
|
||||
done; wait
|
||||
|
||||
## k8s-down: uninstall the release (database PVCs are kept)
|
||||
k8s-down:
|
||||
helm uninstall big -n $(K8S_NS)
|
||||
|
||||
## k8s-purge: uninstall AND drop the namespace, including the database volumes
|
||||
k8s-purge:
|
||||
-helm uninstall big -n $(K8S_NS)
|
||||
kubectl delete namespace $(K8S_NS) --ignore-not-found
|
||||
|
||||
## help: list available targets
|
||||
help:
|
||||
@grep -E '^## ' $(MAKEFILE_LIST) | sed 's/^## //'
|
||||
|
||||
Reference in New Issue
Block a user