docs(k8s): ADR-0033 + the Talos deployment runbook (refs #25)
CI / build (pull_request) Successful in 1m10s
CI / lint (pull_request) Successful in 1m27s
CI / unit (pull_request) Successful in 1m31s
CI / frontend (pull_request) Successful in 3m21s
CI / mutation (pull_request) Successful in 6m23s
CI / verify-stack (pull_request) Canceled after 14m2s
CI / build (pull_request) Successful in 1m10s
CI / lint (pull_request) Successful in 1m27s
CI / unit (pull_request) Successful in 1m31s
CI / frontend (pull_request) Successful in 3m21s
CI / mutation (pull_request) Successful in 6m23s
CI / verify-stack (pull_request) Canceled after 14m2s
ADR-0033 records why one values-driven chart rather than 30 subcharts, the four platform-forced deviations from compose, and the alternatives (kompose, bitnami subcharts, ingress-nginx, Helm hooks for ordering, a laptop-side registry). The runbook is the walkthrough as actually performed on a single-node Talos v1.14 VM under virt-manager, including the parts that bite: virt-manager ejecting the install ISO on first shutdown, Talos 1.14 moving the install disk into its own config document, the control-plane taint, and why the portals must be reached over localhost (crypto.subtle needs a secure context for PKCE).
This commit is contained in:
@@ -0,0 +1,161 @@
|
|||||||
|
# ADR-0033: Kubernetes deployment is one values-driven Helm chart, not a chart per service
|
||||||
|
|
||||||
|
- **Status:** Accepted
|
||||||
|
- **Date:** 2026-09-04
|
||||||
|
- **Deciders:** Respellion engineering
|
||||||
|
- **Slice:** _(none yet — raised directly as a deployment-target request; see
|
||||||
|
"Process note" at the end)_
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
The stack is defined once, in `infra/docker-compose.yml`: 30-odd containers made of six
|
||||||
|
upstream Common Ground modules (OpenZaak, Open Notificaties, Objecten, Objecttypen,
|
||||||
|
Keycloak, Flowable), their databases and workers, five .NET services, four portals, six
|
||||||
|
one-shot bootstrap containers, and an observability backplane (off by default here). Compose is the
|
||||||
|
CI-canonical stack: `make verify` and every `verify-*` script drive it.
|
||||||
|
|
||||||
|
We now also want the stack on Kubernetes — first target a **single-node Talos VM on a
|
||||||
|
laptop**. Four properties of this particular stack shape the answer:
|
||||||
|
|
||||||
|
- **The upstream images are used verbatim** and read their configuration from a mounted
|
||||||
|
directory (`setup_configuration/data.yaml`, Keycloak realm exports, BPMN/DMN). Compose
|
||||||
|
streams those files into external volumes (`infra/seed-config.sh`) because bind mounts
|
||||||
|
don't reach sibling containers on the CI runner. Kubernetes needs the same files as
|
||||||
|
ConfigMaps — from *somewhere*.
|
||||||
|
- **Django's `URLValidator` rejects single-label hosts.** Compose works around it by
|
||||||
|
handing the ACL and the seeds a container *IP* (ADR-0009, ADR-0020, ADR-0029, and the
|
||||||
|
`objecten.local` network alias). In Kubernetes a Service FQDN is already multi-label, so
|
||||||
|
the workaround has a natural replacement — but the hosts have to line up exactly, since
|
||||||
|
Objecten reflects the request Host into the URLs it publishes to NRC.
|
||||||
|
- **The OIDC issuer must be one string** for both the browser and the BFF (ADR-0010).
|
||||||
|
`infra/host-browser.yml` already solved this for a host browser: pin `KC_HOSTNAME`, keep
|
||||||
|
backchannel discovery in-cluster, and mount a `config.json` per portal.
|
||||||
|
- **Nothing here is highly available.** One replica of everything, on one node.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
**One chart — `infra/helm/big-reference` — whose `values.yaml` is a near-literal
|
||||||
|
transcription of the compose file, rendered by three generic templates (Deployment, Job,
|
||||||
|
Service) over a `workloads` map.** Adding a service is a values edit.
|
||||||
|
|
||||||
|
Consequences of that shape, each chosen deliberately:
|
||||||
|
|
||||||
|
- **Config files are not copied into the chart.** `infra/helm/seed-configmaps.sh` creates
|
||||||
|
the ConfigMaps from the files that already live in the repo — the Kubernetes sibling of
|
||||||
|
`infra/seed-config.sh`. The chart therefore needs `make k8s-seed` before `helm install`,
|
||||||
|
which is the same two-step dance compose already has.
|
||||||
|
- **Bootstrap one-shots become Jobs, with no ordering mechanism.** Every one is idempotent
|
||||||
|
(ADR-0020); each waits for the TCP ports it needs via a busybox init container and
|
||||||
|
Kubernetes retries the rest. `make k8s-reseed` re-runs them.
|
||||||
|
- **The four Django services apply their own `setup_configuration`** —
|
||||||
|
`args: [sh, -c, "/setup_configuration.sh && exec /start.sh"]` — instead of getting a
|
||||||
|
separate `*-init` Job like compose. Both of those image scripts run
|
||||||
|
`manage.py migrate`, and compose serialises them with
|
||||||
|
`depends_on: service_completed_successfully`; Kubernetes has no such edge, so a Job and
|
||||||
|
its web pod migrate the same database concurrently and Django dies with
|
||||||
|
*"relation zgw_consumers_service already exists"*. Running the two steps in order inside
|
||||||
|
the one container leaves exactly one migrator per database, and deletes four workloads.
|
||||||
|
- **`args`, never `command`.** Compose's `command:` replaces the image's CMD; Kubernetes'
|
||||||
|
`command:` replaces its ENTRYPOINT. Transcribing one to the other silently broke every
|
||||||
|
upstream image that relies on its entrypoint — postgres ran as root and refused to
|
||||||
|
start, Keycloak tried to exec `start-dev` as a binary. The chart now `fail`s at render
|
||||||
|
time if a workload sets `command`, because the symptom (a crashloop three layers down)
|
||||||
|
is nothing like the cause.
|
||||||
|
- **Published ports are NodePorts.** No ingress controller, no LoadBalancer, no TLS. The
|
||||||
|
four portals are the exception in *use*, not in wiring: PKCE needs `crypto.subtle`, which
|
||||||
|
browsers expose only in a secure context, so a portal has to be reached over `localhost`
|
||||||
|
(`make k8s-portals` forwards them) or eventually over HTTPS. `.Values.host` is therefore
|
||||||
|
"the address the browser uses", not "the node's address" — it pins Keycloak's issuer and
|
||||||
|
each portal's `config.json`, and both must agree with the URL bar (ADR-0010).
|
||||||
|
- **Databases are `emptyDir` by default**, so the stack comes up on a cluster with no CSI
|
||||||
|
driver; setting `persistence.storageClass` switches every database to a PVC.
|
||||||
|
- **Only two hosts become FQDNs** — OpenZaak (for the ACL and the zaaktype seed) and
|
||||||
|
Objecten (for the ACL's register writes), the two that Django validates as URLs.
|
||||||
|
Everything else keeps the short compose service name, because the upstream
|
||||||
|
`setup_configuration` files name those and Objecten matches an objecttype URL against the
|
||||||
|
one it was configured with. The portals used to be a third case — nginx's `resolver` never
|
||||||
|
appends search domains, so the bare `bff` upstream could not resolve on Kubernetes — which
|
||||||
|
ADR-0034 removed by serving them with Caddy, whose resolver honours `/etc/resolv.conf`.
|
||||||
|
- **Compose stays CI-canonical.** The chart is a second deployment target, not a
|
||||||
|
replacement; the acceptance, verify and e2e lanes are unchanged.
|
||||||
|
|
||||||
|
### Alternatives considered
|
||||||
|
|
||||||
|
- **A chart per service, or an umbrella of 30 subcharts.** The conventional layout, and
|
||||||
|
roughly 1,500 lines of near-identical YAML for a stack where 28 of 30 workloads are
|
||||||
|
"one pod, one image, some env". It buys independent versioning we don't want (the stack
|
||||||
|
is demoed as a whole) and costs the eye-diffability against the compose file that keeps
|
||||||
|
the two stacks honest.
|
||||||
|
- **`kompose convert`.** One-shot generation, no ongoing artefact to maintain — but it
|
||||||
|
drops exactly the parts that carry the design (init ordering, the config volumes, the
|
||||||
|
issuer pinning) and produces output nobody owns.
|
||||||
|
- **Bitnami PostgreSQL/Redis subcharts.** Six more dependencies (CLAUDE.md §13) and a
|
||||||
|
second way of expressing the same three-line database.
|
||||||
|
- **ingress-nginx with hostname routing.** Needs a controller, `/etc/hosts` entries and a
|
||||||
|
matching issuer host; NodePorts need none of it and reuse the mechanism
|
||||||
|
`infra/host-browser.yml` already proves.
|
||||||
|
- **A registry on the laptop** (the obvious home for images built there). Talos cannot
|
||||||
|
side-load an image, so a registry is required either way — but reaching one on the host
|
||||||
|
means opening an inbound port on firewalld's `libvirt` zone, which needs root, and
|
||||||
|
pushing to it over plain HTTP means an `insecure-registries` entry in the Docker daemon,
|
||||||
|
which needs root again. `infra/helm/registry.yaml` runs the registry *in* the cluster on
|
||||||
|
a NodePort instead: pushing laptop → node is outbound and unfiltered, the node pulls from
|
||||||
|
its own NodePort, and `docker save | crane push --insecure` needs no daemon
|
||||||
|
configuration. Cost: one more (throwaway, `emptyDir`) workload, and a re-push if its pod
|
||||||
|
is replaced.
|
||||||
|
- **Helm hooks (`pre-install`/`post-install`) for bootstrap ordering.** Hooks run after
|
||||||
|
`--wait`, which would deadlock: OpenZaak's readiness needs the migrations that the hook
|
||||||
|
is supposed to run. Idempotent Jobs plus retries need no such sequencing.
|
||||||
|
|
||||||
|
- ponytail ceiling: single-node assumptions are baked in — one replica per workload,
|
||||||
|
`Recreate` rollouts, ReadWriteOnce volumes, no PodDisruptionBudgets, no resource
|
||||||
|
requests or limits (a laptop VM schedules everything or nothing), plain HTTP.
|
||||||
|
Upgrade path for a real cluster: add requests/limits per workload (the field is already
|
||||||
|
passed through), swap NodePorts for an Ingress with TLS, and give the databases a real
|
||||||
|
StorageClass — none of which changes the workload graph.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
**Positive**
|
||||||
|
|
||||||
|
- One file to read to see what the cluster runs, and it lines up with the compose file
|
||||||
|
line for line.
|
||||||
|
- The compose IP workarounds disappear: cluster DNS supplies multi-label hosts.
|
||||||
|
- `make k8s-lint` renders and schema-checks the whole stack without a cluster.
|
||||||
|
- The config inputs have exactly one home (the repo) for both stacks — no fork to drift.
|
||||||
|
|
||||||
|
**Negative / costs**
|
||||||
|
|
||||||
|
- A second deployment description to keep in step with compose. Nothing enforces that
|
||||||
|
today; a drift check belongs in CI (follow-up).
|
||||||
|
- `helm install` alone is not enough — the ConfigMaps must be seeded first, and a missing
|
||||||
|
one surfaces as `ContainerCreating`, not as a clear error.
|
||||||
|
- Generic templates mean a values typo can render valid-but-wrong YAML; `k8s-lint` catches
|
||||||
|
schema errors, not intent.
|
||||||
|
- The verify/e2e lanes do not run against the chart, so the Kubernetes path is verified by
|
||||||
|
hand (docs/runbooks/kubernetes-talos.md §5) rather than by CI.
|
||||||
|
- The chart deviates from compose in four places now (args, self-configuring Django pods,
|
||||||
|
FQDN hosts, NodePorts). Each is forced by the platform and commented where it appears,
|
||||||
|
but it is four more things that can drift.
|
||||||
|
|
||||||
|
## Coupling rules touched (CLAUDE.md §8)
|
||||||
|
|
||||||
|
None. The chart deploys the same graph: portals reach only the BFF (§8.3), only the ACL
|
||||||
|
holds ZGW credentials (§8.1), only the Workflow Client talks to Flowable (§8.2), each
|
||||||
|
service keeps its own database (§8.5). No workload gained a peer it didn't have in compose.
|
||||||
|
|
||||||
|
## Verified
|
||||||
|
|
||||||
|
Brought up from scratch on a single-node Talos v1.14.0 VM (6 vCPU / 10 GB, virtio disk)
|
||||||
|
under virt-manager: 29 pods ready and four bootstrap Jobs complete in under three minutes,
|
||||||
|
with zero restarts, using ~4.4 GB of the VM's 10 GB. The smoke test in the runbook's §5
|
||||||
|
walks the whole path — portal proxy → BFF → domain → Flowable → ACL → OpenZaak + Objecten →
|
||||||
|
NRC → event-subscriber → projection → public register — plus a werkbak read with an
|
||||||
|
MFA'd medewerker token. The browser flow itself was driven with Playwright against
|
||||||
|
`http://localhost:30140`: secure context, PKCE, Keycloak form, login, no console errors.
|
||||||
|
|
||||||
|
## Process note
|
||||||
|
|
||||||
|
CLAUDE.md §14 wants the ADR proposal issue opened before the code, and §7 wants a slice
|
||||||
|
issue behind the work. This landed the other way round — chart first, on request. The
|
||||||
|
issue and the CI drift check are the outstanding follow-ups.
|
||||||
@@ -0,0 +1,370 @@
|
|||||||
|
# Deploying the stack to a single-node Talos cluster
|
||||||
|
|
||||||
|
The Helm chart in `infra/helm/big-reference` is a port of `infra/docker-compose.yml`
|
||||||
|
(ADR-0033). This runbook is the walkthrough that was actually used to bring the stack up
|
||||||
|
on a Talos VM under virt-manager on a laptop, including the parts that bite.
|
||||||
|
|
||||||
|
Compose remains the CI-canonical stack — `make verify`, the acceptance lane and the
|
||||||
|
Playwright e2e all still drive it. Kubernetes is a second deployment target.
|
||||||
|
|
||||||
|
## 0. What you need
|
||||||
|
|
||||||
|
On the laptop, four static binaries, all installable to `~/.local/bin` without root:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
curl -sSLo ~/.local/bin/talosctl https://github.com/siderolabs/talos/releases/download/v1.14.0/talosctl-linux-amd64
|
||||||
|
curl -sSLo ~/.local/bin/kubectl https://dl.k8s.io/release/v1.37.0/bin/linux/amd64/kubectl
|
||||||
|
curl -sSL https://get.helm.sh/helm-v3.16.4-linux-amd64.tar.gz | tar xz -O linux-amd64/helm > ~/.local/bin/helm
|
||||||
|
curl -sSL https://github.com/google/go-containerregistry/releases/download/v0.20.2/go-containerregistry_Linux_x86_64.tar.gz | tar xz -O crane > ~/.local/bin/crane
|
||||||
|
chmod +x ~/.local/bin/{talosctl,kubectl,helm,crane}
|
||||||
|
```
|
||||||
|
|
||||||
|
Match `talosctl` to the Talos ISO you booted (`talosctl version --insecure -n <ip>` reports
|
||||||
|
the server's tag). `crane` is what pushes images to a plain-HTTP registry without a
|
||||||
|
root-level Docker daemon change — see §2.
|
||||||
|
|
||||||
|
**VM sizing.** 6 vCPU / 10 GB RAM / 27 GB disk runs the whole stack with room to spare
|
||||||
|
(measured: ~4.4 GB used, 5.4 GB available with all 29 pods up). 4 GB is not enough. The
|
||||||
|
chart sets no resource requests or limits on purpose — on a single node the VM's RAM is the
|
||||||
|
only budget there is. Resize a stopped VM with:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
virsh -c qemu:///system destroy talos # it's in maintenance mode; nothing is lost
|
||||||
|
virsh -c qemu:///system setmaxmem talos 10G --config
|
||||||
|
virsh -c qemu:///system setmem talos 10G --config
|
||||||
|
virsh -c qemu:///system setvcpus talos 6 --config --maximum
|
||||||
|
virsh -c qemu:///system setvcpus talos 6 --config
|
||||||
|
```
|
||||||
|
|
||||||
|
Two addresses matter throughout:
|
||||||
|
|
||||||
|
| Name | Meaning | Example |
|
||||||
|
|---|---|---|
|
||||||
|
| `TALOS_HOST` | the VM's IP — used by the browser, `talosctl` and `kubectl` | `192.168.122.33` |
|
||||||
|
| `K8S_REGISTRY` | `TALOS_HOST:30500` — the in-cluster registry (§2) | `192.168.122.33:30500` |
|
||||||
|
|
||||||
|
Find the VM's address with `virsh -c qemu:///system net-dhcp-leases default`.
|
||||||
|
|
||||||
|
## 1. Install Talos onto the VM
|
||||||
|
|
||||||
|
### The virt-manager trap
|
||||||
|
|
||||||
|
virt-manager treats the install ISO as one-shot: on the VM's **first shutdown** it ejects
|
||||||
|
the CD and rewrites the boot order to `hd`. A Talos VM booted from `metal-amd64.iso` runs
|
||||||
|
entirely in RAM, so the disk is still empty — the next start lands on
|
||||||
|
`Boot failed: not a bootable disk`. Put the ISO back before installing:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
virsh -c qemu:///system change-media talos sda /path/to/metal-amd64.iso --config --insert
|
||||||
|
virt-xml -c qemu:///system talos --edit --boot cdrom,hd
|
||||||
|
virsh -c qemu:///system start talos
|
||||||
|
```
|
||||||
|
|
||||||
|
Wait for the maintenance-mode API, then confirm the install disk's device name — on virtio
|
||||||
|
it is `/dev/vda`, and Talos's default selector expects `/dev/sda`:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
talosctl get disks --insecure -n <TALOS_HOST> -e <TALOS_HOST>
|
||||||
|
```
|
||||||
|
|
||||||
|
### Generate the machine config
|
||||||
|
|
||||||
|
Talos 1.14 moved several v1alpha1 fields into their own config documents. In particular
|
||||||
|
`machine.install` is now `UnattendedInstallConfig`, and patching the old field is rejected
|
||||||
|
with *"UnattendedInstallConfig config is incompatible with v1alpha1 config"*. Write
|
||||||
|
`patch.yaml` as a multi-document patch:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
machine:
|
||||||
|
certSANs:
|
||||||
|
- 192.168.122.33
|
||||||
|
registries:
|
||||||
|
mirrors:
|
||||||
|
# The in-cluster registry (§2) speaks plain HTTP.
|
||||||
|
"192.168.122.33:30500":
|
||||||
|
endpoints:
|
||||||
|
- http://192.168.122.33:30500
|
||||||
|
---
|
||||||
|
apiVersion: v1alpha1
|
||||||
|
kind: UnattendedInstallConfig
|
||||||
|
provisioning:
|
||||||
|
diskSelector:
|
||||||
|
match: disk.dev_path == "/dev/vda"
|
||||||
|
```
|
||||||
|
|
||||||
|
```bash
|
||||||
|
talosctl gen config big https://<TALOS_HOST>:6443 --output-dir ~/.talos/big --config-patch @patch.yaml
|
||||||
|
talosctl apply-config --insecure -n <TALOS_HOST> -e <TALOS_HOST> --file ~/.talos/big/controlplane.yaml
|
||||||
|
```
|
||||||
|
|
||||||
|
Talos installs to the disk and **kexecs straight into the installed system**, so the CD
|
||||||
|
boot order doesn't get in the way here. Then point the client at the node and bootstrap:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
talosctl config merge ~/.talos/big/talosconfig
|
||||||
|
talosctl config endpoint <TALOS_HOST>
|
||||||
|
talosctl config node <TALOS_HOST>
|
||||||
|
talosctl bootstrap # wait for `talosctl version` to answer first
|
||||||
|
talosctl kubeconfig -f ~/.kube/config
|
||||||
|
```
|
||||||
|
|
||||||
|
A single-node cluster must run workloads on the control plane, or CoreDNS never schedules:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
kubectl taint node --all node-role.kubernetes.io/control-plane-
|
||||||
|
```
|
||||||
|
|
||||||
|
Finally, make a VM restart boot the installed system rather than the ISO (takes effect at
|
||||||
|
the next full power cycle):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
virsh -c qemu:///system change-media talos sda --eject --config
|
||||||
|
virt-xml -c qemu:///system talos --edit --boot hd
|
||||||
|
```
|
||||||
|
|
||||||
|
## 2. A registry the node can pull from
|
||||||
|
|
||||||
|
Talos has no Docker daemon and no way to side-load an image, so this repo's images have to
|
||||||
|
come from a registry. The registry runs **inside the cluster**, published on NodePort
|
||||||
|
30500 (`infra/helm/registry.yaml`):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
make k8s-registry
|
||||||
|
```
|
||||||
|
|
||||||
|
Why in-cluster rather than on the laptop: a laptop-side registry needs an inbound port
|
||||||
|
opened on firewalld's `libvirt` zone (`sudo firewall-cmd --zone=libvirt --add-port=5000/tcp`),
|
||||||
|
which needs root. Pushing from the laptop *to* the node is outbound and always allowed, and
|
||||||
|
the node pulls from its own NodePort. If you do open that port, put a registry on the
|
||||||
|
laptop instead and point `K8S_REGISTRY` at `<laptop-ip>:5000` — the mirror patch in §1 has
|
||||||
|
an entry ready for it.
|
||||||
|
|
||||||
|
Its storage is `emptyDir`, so if the registry pod is ever replaced, re-run `make k8s-images`.
|
||||||
|
|
||||||
|
## 3. Build and push the images
|
||||||
|
|
||||||
|
```bash
|
||||||
|
make k8s-images K8S_REGISTRY=<TALOS_HOST>:30500
|
||||||
|
```
|
||||||
|
|
||||||
|
This builds the nine images with `docker compose build` — same contexts and Dockerfiles as
|
||||||
|
compose, no second build definition — then `docker save | crane push --insecure` each one.
|
||||||
|
`docker push` is not used: the registry speaks plain HTTP, which the Docker daemon refuses
|
||||||
|
without a root-level `insecure-registries` entry, while crane just takes `--insecure`.
|
||||||
|
|
||||||
|
## 4. Deploy
|
||||||
|
|
||||||
|
```bash
|
||||||
|
make k8s-up TALOS_HOST=<TALOS_HOST> K8S_REGISTRY=<TALOS_HOST>:30500
|
||||||
|
```
|
||||||
|
|
||||||
|
That does two things:
|
||||||
|
|
||||||
|
1. `make k8s-seed` — creates the ConfigMaps the chart mounts, from the config files that
|
||||||
|
already live in this repo (`infra/helm/seed-configmaps.sh`): the four
|
||||||
|
`setup_configuration/data.yaml` files, the Keycloak realm exports, the BPMN + DMN, and
|
||||||
|
the two bootstrap scripts. Re-run it after editing any of them.
|
||||||
|
2. `helm upgrade --install` of the chart into namespace `big`.
|
||||||
|
|
||||||
|
First bring-up takes a few minutes: the four Django services migrate their databases and
|
||||||
|
apply their `setup_configuration`, Flowable creates its schema, and the bootstrap Jobs
|
||||||
|
deploy the BPMN/DMN, seed the zaaktype and register the NRC abonnement.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
kubectl -n big get pods -w
|
||||||
|
kubectl -n big get jobs # all four must reach COMPLETIONS 1/1
|
||||||
|
```
|
||||||
|
|
||||||
|
The Jobs are the stack's wiring; if one is not complete, the flow is broken somewhere
|
||||||
|
specific:
|
||||||
|
|
||||||
|
| Job | What breaks without it |
|
||||||
|
|---|---|
|
||||||
|
| `flowable-init` | no `registratie` process, no diploma DMN |
|
||||||
|
| `registerrecord-init` | the register has no RegisterRecord objecttype, so writes are refused |
|
||||||
|
| `seed-zaaktype` | the ACL can't resolve `BIG-REGISTRATIE`, so no zaak is created |
|
||||||
|
| `nrc-subscribe` | register writes never reach the projection — the public register stays empty |
|
||||||
|
|
||||||
|
## 5. Use it
|
||||||
|
|
||||||
|
### The portals must be reached over `localhost`
|
||||||
|
|
||||||
|
The portals' OIDC flow uses PKCE, which needs `crypto.subtle` — and browsers only expose
|
||||||
|
that in a **secure context**: HTTPS, or an origin on `localhost`/`127.0.0.1`. A NodePort on
|
||||||
|
the VM's IP is neither, so `http://<TALOS_HOST>:30140` fails before it can even build the
|
||||||
|
authorize URL:
|
||||||
|
|
||||||
|
```
|
||||||
|
ERROR TypeError: Cannot read properties of undefined (reading 'digest')
|
||||||
|
at t.calcHash → t.generateCodeChallenge → t.createUrlCodeFlowAuthorize
|
||||||
|
```
|
||||||
|
|
||||||
|
So deploy with `TALOS_HOST=localhost` — which pins Keycloak's issuer and the portals'
|
||||||
|
`config.json` authority to `http://localhost:30180` — and forward the browser-facing
|
||||||
|
services to those same ports:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
make k8s-up TALOS_HOST=localhost K8S_REGISTRY=<TALOS_HOST>:30500
|
||||||
|
make k8s-portals # stays in the foreground; Ctrl-C stops all five forwards
|
||||||
|
```
|
||||||
|
|
||||||
|
| URL (needs `make k8s-portals`) | What |
|
||||||
|
|---|---|
|
||||||
|
| `http://localhost:30140` | self-service portal (DigiD) |
|
||||||
|
| `http://localhost:30141` | openbaar register (anonymous) |
|
||||||
|
| `http://localhost:30142` | behandel portal (medewerker) |
|
||||||
|
| `http://localhost:30143` | beheer portal (medewerker) |
|
||||||
|
| `http://localhost:30180` | Keycloak (admin/admin) |
|
||||||
|
|
||||||
|
The port numbers are deliberately the NodePort numbers: Keycloak's issuer is one fixed
|
||||||
|
string, so the port the browser uses has to match the one baked into `config.json`.
|
||||||
|
|
||||||
|
This is the same mechanism `infra/host-browser.yml` uses for the compose stack (which pins
|
||||||
|
`localhost:8180`); only the addresses differ.
|
||||||
|
|
||||||
|
### The admin UIs work straight off the NodePorts
|
||||||
|
|
||||||
|
These are server-rendered and need no secure context, so they are reachable at the VM's
|
||||||
|
address with no forwarding:
|
||||||
|
|
||||||
|
| URL | What |
|
||||||
|
|---|---|
|
||||||
|
| `http://<TALOS_HOST>:30000` | OpenZaak admin (admin/admin) |
|
||||||
|
| `http://<TALOS_HOST>:30001` | Open Notificaties admin (admin/admin) |
|
||||||
|
| `http://<TALOS_HOST>:30020` / `:30021` | Objecttypen / Objecten admin |
|
||||||
|
| `http://<TALOS_HOST>:30080` | BFF (`/health`) |
|
||||||
|
| `http://<TALOS_HOST>:30090` | Flowable REST (rest-admin/test) |
|
||||||
|
|
||||||
|
### Credentials
|
||||||
|
|
||||||
|
Log in with the test users from `docs/synthetic-data.md` (all password `test123`, e.g.
|
||||||
|
`jan-burger` for self-service, `merel-behandelaar` for behandel). The `medewerker` realm
|
||||||
|
enforces MFA (ADR-0031) — print a current code with
|
||||||
|
`python3 infra/keycloak/check_realms.py otp`. Walk the flow in `docs/demo-script.md`.
|
||||||
|
|
||||||
|
`TALOS_HOST` is not cosmetic: it pins Keycloak's issuer (`KC_HOSTNAME`) and the portals'
|
||||||
|
OIDC authority to the same string, which is what makes a browser token pass the BFF's
|
||||||
|
validation (ADR-0010). Change it and you must re-run `make k8s-up` — the chart rolls the
|
||||||
|
portals for you, because their `config.json` is a subPath mount and would otherwise keep
|
||||||
|
serving the old authority.
|
||||||
|
|
||||||
|
### Smoke-test the whole chain without a browser
|
||||||
|
|
||||||
|
With the forwards running:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
TOK=$(curl -s -X POST http://localhost:30180/realms/digid/protocol/openid-connect/token \
|
||||||
|
-d grant_type=password -d client_id=big-portal \
|
||||||
|
-d username=jan-burger -d password=test123 -d scope=openid | jq -r .access_token)
|
||||||
|
|
||||||
|
# through the portal's Caddy, so this also proves the BFF reverse proxy
|
||||||
|
curl -s -X POST http://localhost:30140/self-service/registrations \
|
||||||
|
-H "Authorization: Bearer $TOK" -H 'Content-Length: 0'
|
||||||
|
# → {"registrationId":"…","status":"Ingediend"}
|
||||||
|
|
||||||
|
curl -s http://localhost:30141/openbaar/register
|
||||||
|
# → [{"id":"…","status":"INGEDIEND","reference":"<the registrationId>"}]
|
||||||
|
```
|
||||||
|
|
||||||
|
The second call proves the whole Common Ground path: portal → BFF → domain → Flowable →
|
||||||
|
ACL → OpenZaak + Objecten → NRC → event-subscriber → projection → openbaar register.
|
||||||
|
|
||||||
|
## 6. Keeping the databases (recommended if you iterate on the chart)
|
||||||
|
|
||||||
|
By default every database is an `emptyDir`: no CSI driver needed, and the data lives as
|
||||||
|
long as the pod. Note what that means in practice — **any** change to a database pod's
|
||||||
|
template (an image policy, an env value, a probe) recreates the pod and wipes it. The stack
|
||||||
|
then needs its bootstrap re-run:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
make k8s-reseed TALOS_HOST=... K8S_REGISTRY=...
|
||||||
|
```
|
||||||
|
|
||||||
|
which re-runs the four Jobs *and* restarts `event-subscriber` + `projection-api`, because
|
||||||
|
those two create the projection schema on start and otherwise keep writing to a
|
||||||
|
schema-less database (`relation "processed_notifications" does not exist`). For persistence, install Rancher's local-path-provisioner — on Talos it
|
||||||
|
must write under `/var` and its namespace needs the privileged Pod Security label:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
# kustomization.yaml
|
||||||
|
apiVersion: kustomize.config.k8s.io/v1beta1
|
||||||
|
kind: Kustomization
|
||||||
|
resources:
|
||||||
|
- github.com/rancher/local-path-provisioner/deploy?ref=v0.0.31
|
||||||
|
patches:
|
||||||
|
- patch: |-
|
||||||
|
kind: ConfigMap
|
||||||
|
apiVersion: v1
|
||||||
|
metadata:
|
||||||
|
name: local-path-config
|
||||||
|
namespace: local-path-storage
|
||||||
|
data:
|
||||||
|
config.json: |-
|
||||||
|
{ "nodePathMap":[ { "node":"DEFAULT_PATH_FOR_NON_LISTED_NODES", "paths":["/var/local-path-provisioner"] } ] }
|
||||||
|
- patch: |-
|
||||||
|
apiVersion: v1
|
||||||
|
kind: Namespace
|
||||||
|
metadata:
|
||||||
|
name: local-path-storage
|
||||||
|
labels:
|
||||||
|
pod-security.kubernetes.io/enforce: privileged
|
||||||
|
```
|
||||||
|
|
||||||
|
```bash
|
||||||
|
kubectl apply -k .
|
||||||
|
make k8s-up TALOS_HOST=... K8S_REGISTRY=... K8S_SET='--set persistence.storageClass=local-path'
|
||||||
|
```
|
||||||
|
|
||||||
|
The PVCs carry `helm.sh/resource-policy: keep`, so `make k8s-down` leaves the data behind;
|
||||||
|
`make k8s-purge` drops the namespace and with it the volumes.
|
||||||
|
|
||||||
|
## 7. Day-to-day
|
||||||
|
|
||||||
|
```bash
|
||||||
|
make k8s-lint # render + schema-check the chart, no cluster needed
|
||||||
|
make k8s-portals # forward the portals + Keycloak to localhost (browser access)
|
||||||
|
make k8s-images K8S_REGISTRY=... # after changing a service or a portal
|
||||||
|
make k8s-up TALOS_HOST=... K8S_REGISTRY=...
|
||||||
|
make k8s-seed # after editing a data.yaml, a realm export, or the BPMN
|
||||||
|
make k8s-reseed TALOS_HOST=... K8S_REGISTRY=... # re-run the bootstrap Jobs + reset the projection schema
|
||||||
|
make k8s-down # uninstall, keep the database PVCs
|
||||||
|
make k8s-purge # uninstall and drop the namespace
|
||||||
|
```
|
||||||
|
|
||||||
|
This repo's images are pulled with `imagePullPolicy: Always` (the `dev` tag is mutable), so
|
||||||
|
`kubectl -n big rollout restart deploy/<name>` after `make k8s-images` picks up a rebuild.
|
||||||
|
Upstream images stay `IfNotPresent`: their tags are pinned, and keeping them out of the pod
|
||||||
|
template avoids needless churn — a changed template makes a Job unpatchable.
|
||||||
|
|
||||||
|
`k8s-reseed` is also the path for *changing* a Job in the chart: a Job's pod template is
|
||||||
|
immutable, so `helm upgrade` is rejected with `cannot patch "…" with kind Job`.
|
||||||
|
|
||||||
|
## 8. When it doesn't work
|
||||||
|
|
||||||
|
| Symptom | Cause |
|
||||||
|
|---|---|
|
||||||
|
| `Boot failed: not a bootable disk` | virt-manager ejected the install ISO on first shutdown — see §1 |
|
||||||
|
| The VM comes back in maintenance mode after a restart | the ISO is still attached and boots first; eject it and set `--boot hd` (§1) |
|
||||||
|
| `apply-config` rejects the patch with *"incompatible with v1alpha1"* | Talos ≥1.14 owns that field in its own config document — patch the document, not `machine.*` (§1) |
|
||||||
|
| CoreDNS `Pending` forever | the control-plane taint is still on the only node (§1) |
|
||||||
|
| `ImagePullBackOff` … `pull QPS exceeded` | transient: the kubelet rate-limits pulls when ~30 pods start at once. It recovers on retry |
|
||||||
|
| `ImagePullBackOff` on a `register-referentie/*` image | the registry mirror patch is missing: `talosctl get registriesconfig` |
|
||||||
|
| Pod stuck in `ContainerCreating`, event names a ConfigMap | `make k8s-seed` |
|
||||||
|
| `seed-zaaktype` retrying | publishing a zaaktype validates the resultaattype against `selectielijst.openzaak.nl`, so this one Job needs outbound internet from the VM (ADR-0006) |
|
||||||
|
| `TypeError: Cannot read properties of undefined (reading 'digest')` on a portal | not a secure context: `crypto.subtle` is absent on `http://<ip>`. Use `localhost` + `make k8s-portals` (§5) |
|
||||||
|
| Login redirects but the portal stays logged out, or the BFF answers 401 | `TALOS_HOST` doesn't match the address in the browser's URL bar — issuer mismatch. Re-run `make k8s-up` with the right value |
|
||||||
|
| A portal returns 502 on `/self-service/…` | the BFF is unreachable from the portal pod: check `kubectl -n big get svc bff` and the BFF's own readiness |
|
||||||
|
| Public register empty after a submit | usually a wiped `emptyDir` database (§6): `make k8s-reseed`. Confirm with `kubectl -n big logs deploy/event-subscriber \| grep 42P01` |
|
||||||
|
| `helm upgrade` fails with `cannot patch … with kind Job` | see §7 — use `make k8s-reseed` |
|
||||||
|
| Pods `Evicted` / `OOMKilled` | the VM is too small (§0) |
|
||||||
|
| A Job shows `BackoffLimitExceeded` | read it: `kubectl -n big logs job/<name>` |
|
||||||
|
|
||||||
|
## What is not ported
|
||||||
|
|
||||||
|
- **Observability** (Tempo, Prometheus, Grafana) is defined but disabled — those are built
|
||||||
|
images too, so switching them on means pushing them as well:
|
||||||
|
`K8S_SET='--set workloads.tempo.enabled=true --set workloads.prometheus.enabled=true --set workloads.grafana.enabled=true'`.
|
||||||
|
The .NET services still export OTLP; the exporter fails harmlessly when Tempo is absent.
|
||||||
|
- **The verify/e2e lanes.** `make verify*` and the Playwright e2e drive compose, not the
|
||||||
|
chart. The Kubernetes path is verified with §5's smoke test.
|
||||||
|
- **Ingress, TLS, and resource requests.** See the ponytail ceiling in ADR-0033.
|
||||||
Reference in New Issue
Block a user