diff --git a/docs/architecture/adr-0033-kubernetes-via-one-helm-chart.md b/docs/architecture/adr-0033-kubernetes-via-one-helm-chart.md new file mode 100644 index 0000000..ef7b39e --- /dev/null +++ b/docs/architecture/adr-0033-kubernetes-via-one-helm-chart.md @@ -0,0 +1,161 @@ +# ADR-0033: Kubernetes deployment is one values-driven Helm chart, not a chart per service + +- **Status:** Accepted +- **Date:** 2026-09-04 +- **Deciders:** Respellion engineering +- **Slice:** _(none yet — raised directly as a deployment-target request; see + "Process note" at the end)_ + +## Context + +The stack is defined once, in `infra/docker-compose.yml`: 30-odd containers made of six +upstream Common Ground modules (OpenZaak, Open Notificaties, Objecten, Objecttypen, +Keycloak, Flowable), their databases and workers, five .NET services, four portals, six +one-shot bootstrap containers, and an observability backplane (off by default here). Compose is the +CI-canonical stack: `make verify` and every `verify-*` script drive it. + +We now also want the stack on Kubernetes — first target a **single-node Talos VM on a +laptop**. Four properties of this particular stack shape the answer: + +- **The upstream images are used verbatim** and read their configuration from a mounted + directory (`setup_configuration/data.yaml`, Keycloak realm exports, BPMN/DMN). Compose + streams those files into external volumes (`infra/seed-config.sh`) because bind mounts + don't reach sibling containers on the CI runner. Kubernetes needs the same files as + ConfigMaps — from *somewhere*. +- **Django's `URLValidator` rejects single-label hosts.** Compose works around it by + handing the ACL and the seeds a container *IP* (ADR-0009, ADR-0020, ADR-0029, and the + `objecten.local` network alias). In Kubernetes a Service FQDN is already multi-label, so + the workaround has a natural replacement — but the hosts have to line up exactly, since + Objecten reflects the request Host into the URLs it publishes to NRC. +- **The OIDC issuer must be one string** for both the browser and the BFF (ADR-0010). + `infra/host-browser.yml` already solved this for a host browser: pin `KC_HOSTNAME`, keep + backchannel discovery in-cluster, and mount a `config.json` per portal. +- **Nothing here is highly available.** One replica of everything, on one node. + +## Decision + +**One chart — `infra/helm/big-reference` — whose `values.yaml` is a near-literal +transcription of the compose file, rendered by three generic templates (Deployment, Job, +Service) over a `workloads` map.** Adding a service is a values edit. + +Consequences of that shape, each chosen deliberately: + +- **Config files are not copied into the chart.** `infra/helm/seed-configmaps.sh` creates + the ConfigMaps from the files that already live in the repo — the Kubernetes sibling of + `infra/seed-config.sh`. The chart therefore needs `make k8s-seed` before `helm install`, + which is the same two-step dance compose already has. +- **Bootstrap one-shots become Jobs, with no ordering mechanism.** Every one is idempotent + (ADR-0020); each waits for the TCP ports it needs via a busybox init container and + Kubernetes retries the rest. `make k8s-reseed` re-runs them. +- **The four Django services apply their own `setup_configuration`** — + `args: [sh, -c, "/setup_configuration.sh && exec /start.sh"]` — instead of getting a + separate `*-init` Job like compose. Both of those image scripts run + `manage.py migrate`, and compose serialises them with + `depends_on: service_completed_successfully`; Kubernetes has no such edge, so a Job and + its web pod migrate the same database concurrently and Django dies with + *"relation zgw_consumers_service already exists"*. Running the two steps in order inside + the one container leaves exactly one migrator per database, and deletes four workloads. +- **`args`, never `command`.** Compose's `command:` replaces the image's CMD; Kubernetes' + `command:` replaces its ENTRYPOINT. Transcribing one to the other silently broke every + upstream image that relies on its entrypoint — postgres ran as root and refused to + start, Keycloak tried to exec `start-dev` as a binary. The chart now `fail`s at render + time if a workload sets `command`, because the symptom (a crashloop three layers down) + is nothing like the cause. +- **Published ports are NodePorts.** No ingress controller, no LoadBalancer, no TLS. The + four portals are the exception in *use*, not in wiring: PKCE needs `crypto.subtle`, which + browsers expose only in a secure context, so a portal has to be reached over `localhost` + (`make k8s-portals` forwards them) or eventually over HTTPS. `.Values.host` is therefore + "the address the browser uses", not "the node's address" — it pins Keycloak's issuer and + each portal's `config.json`, and both must agree with the URL bar (ADR-0010). +- **Databases are `emptyDir` by default**, so the stack comes up on a cluster with no CSI + driver; setting `persistence.storageClass` switches every database to a PVC. +- **Only two hosts become FQDNs** — OpenZaak (for the ACL and the zaaktype seed) and + Objecten (for the ACL's register writes), the two that Django validates as URLs. + Everything else keeps the short compose service name, because the upstream + `setup_configuration` files name those and Objecten matches an objecttype URL against the + one it was configured with. The portals used to be a third case — nginx's `resolver` never + appends search domains, so the bare `bff` upstream could not resolve on Kubernetes — which + ADR-0034 removed by serving them with Caddy, whose resolver honours `/etc/resolv.conf`. +- **Compose stays CI-canonical.** The chart is a second deployment target, not a + replacement; the acceptance, verify and e2e lanes are unchanged. + +### Alternatives considered + +- **A chart per service, or an umbrella of 30 subcharts.** The conventional layout, and + roughly 1,500 lines of near-identical YAML for a stack where 28 of 30 workloads are + "one pod, one image, some env". It buys independent versioning we don't want (the stack + is demoed as a whole) and costs the eye-diffability against the compose file that keeps + the two stacks honest. +- **`kompose convert`.** One-shot generation, no ongoing artefact to maintain — but it + drops exactly the parts that carry the design (init ordering, the config volumes, the + issuer pinning) and produces output nobody owns. +- **Bitnami PostgreSQL/Redis subcharts.** Six more dependencies (CLAUDE.md §13) and a + second way of expressing the same three-line database. +- **ingress-nginx with hostname routing.** Needs a controller, `/etc/hosts` entries and a + matching issuer host; NodePorts need none of it and reuse the mechanism + `infra/host-browser.yml` already proves. +- **A registry on the laptop** (the obvious home for images built there). Talos cannot + side-load an image, so a registry is required either way — but reaching one on the host + means opening an inbound port on firewalld's `libvirt` zone, which needs root, and + pushing to it over plain HTTP means an `insecure-registries` entry in the Docker daemon, + which needs root again. `infra/helm/registry.yaml` runs the registry *in* the cluster on + a NodePort instead: pushing laptop → node is outbound and unfiltered, the node pulls from + its own NodePort, and `docker save | crane push --insecure` needs no daemon + configuration. Cost: one more (throwaway, `emptyDir`) workload, and a re-push if its pod + is replaced. +- **Helm hooks (`pre-install`/`post-install`) for bootstrap ordering.** Hooks run after + `--wait`, which would deadlock: OpenZaak's readiness needs the migrations that the hook + is supposed to run. Idempotent Jobs plus retries need no such sequencing. + +- ponytail ceiling: single-node assumptions are baked in — one replica per workload, + `Recreate` rollouts, ReadWriteOnce volumes, no PodDisruptionBudgets, no resource + requests or limits (a laptop VM schedules everything or nothing), plain HTTP. + Upgrade path for a real cluster: add requests/limits per workload (the field is already + passed through), swap NodePorts for an Ingress with TLS, and give the databases a real + StorageClass — none of which changes the workload graph. + +## Consequences + +**Positive** + +- One file to read to see what the cluster runs, and it lines up with the compose file + line for line. +- The compose IP workarounds disappear: cluster DNS supplies multi-label hosts. +- `make k8s-lint` renders and schema-checks the whole stack without a cluster. +- The config inputs have exactly one home (the repo) for both stacks — no fork to drift. + +**Negative / costs** + +- A second deployment description to keep in step with compose. Nothing enforces that + today; a drift check belongs in CI (follow-up). +- `helm install` alone is not enough — the ConfigMaps must be seeded first, and a missing + one surfaces as `ContainerCreating`, not as a clear error. +- Generic templates mean a values typo can render valid-but-wrong YAML; `k8s-lint` catches + schema errors, not intent. +- The verify/e2e lanes do not run against the chart, so the Kubernetes path is verified by + hand (docs/runbooks/kubernetes-talos.md §5) rather than by CI. +- The chart deviates from compose in four places now (args, self-configuring Django pods, + FQDN hosts, NodePorts). Each is forced by the platform and commented where it appears, + but it is four more things that can drift. + +## Coupling rules touched (CLAUDE.md §8) + +None. The chart deploys the same graph: portals reach only the BFF (§8.3), only the ACL +holds ZGW credentials (§8.1), only the Workflow Client talks to Flowable (§8.2), each +service keeps its own database (§8.5). No workload gained a peer it didn't have in compose. + +## Verified + +Brought up from scratch on a single-node Talos v1.14.0 VM (6 vCPU / 10 GB, virtio disk) +under virt-manager: 29 pods ready and four bootstrap Jobs complete in under three minutes, +with zero restarts, using ~4.4 GB of the VM's 10 GB. The smoke test in the runbook's §5 +walks the whole path — portal proxy → BFF → domain → Flowable → ACL → OpenZaak + Objecten → +NRC → event-subscriber → projection → public register — plus a werkbak read with an +MFA'd medewerker token. The browser flow itself was driven with Playwright against +`http://localhost:30140`: secure context, PKCE, Keycloak form, login, no console errors. + +## Process note + +CLAUDE.md §14 wants the ADR proposal issue opened before the code, and §7 wants a slice +issue behind the work. This landed the other way round — chart first, on request. The +issue and the CI drift check are the outstanding follow-ups. diff --git a/docs/runbooks/kubernetes-talos.md b/docs/runbooks/kubernetes-talos.md new file mode 100644 index 0000000..c4c1f54 --- /dev/null +++ b/docs/runbooks/kubernetes-talos.md @@ -0,0 +1,370 @@ +# Deploying the stack to a single-node Talos cluster + +The Helm chart in `infra/helm/big-reference` is a port of `infra/docker-compose.yml` +(ADR-0033). This runbook is the walkthrough that was actually used to bring the stack up +on a Talos VM under virt-manager on a laptop, including the parts that bite. + +Compose remains the CI-canonical stack — `make verify`, the acceptance lane and the +Playwright e2e all still drive it. Kubernetes is a second deployment target. + +## 0. What you need + +On the laptop, four static binaries, all installable to `~/.local/bin` without root: + +```bash +curl -sSLo ~/.local/bin/talosctl https://github.com/siderolabs/talos/releases/download/v1.14.0/talosctl-linux-amd64 +curl -sSLo ~/.local/bin/kubectl https://dl.k8s.io/release/v1.37.0/bin/linux/amd64/kubectl +curl -sSL https://get.helm.sh/helm-v3.16.4-linux-amd64.tar.gz | tar xz -O linux-amd64/helm > ~/.local/bin/helm +curl -sSL https://github.com/google/go-containerregistry/releases/download/v0.20.2/go-containerregistry_Linux_x86_64.tar.gz | tar xz -O crane > ~/.local/bin/crane +chmod +x ~/.local/bin/{talosctl,kubectl,helm,crane} +``` + +Match `talosctl` to the Talos ISO you booted (`talosctl version --insecure -n ` reports +the server's tag). `crane` is what pushes images to a plain-HTTP registry without a +root-level Docker daemon change — see §2. + +**VM sizing.** 6 vCPU / 10 GB RAM / 27 GB disk runs the whole stack with room to spare +(measured: ~4.4 GB used, 5.4 GB available with all 29 pods up). 4 GB is not enough. The +chart sets no resource requests or limits on purpose — on a single node the VM's RAM is the +only budget there is. Resize a stopped VM with: + +```bash +virsh -c qemu:///system destroy talos # it's in maintenance mode; nothing is lost +virsh -c qemu:///system setmaxmem talos 10G --config +virsh -c qemu:///system setmem talos 10G --config +virsh -c qemu:///system setvcpus talos 6 --config --maximum +virsh -c qemu:///system setvcpus talos 6 --config +``` + +Two addresses matter throughout: + +| Name | Meaning | Example | +|---|---|---| +| `TALOS_HOST` | the VM's IP — used by the browser, `talosctl` and `kubectl` | `192.168.122.33` | +| `K8S_REGISTRY` | `TALOS_HOST:30500` — the in-cluster registry (§2) | `192.168.122.33:30500` | + +Find the VM's address with `virsh -c qemu:///system net-dhcp-leases default`. + +## 1. Install Talos onto the VM + +### The virt-manager trap + +virt-manager treats the install ISO as one-shot: on the VM's **first shutdown** it ejects +the CD and rewrites the boot order to `hd`. A Talos VM booted from `metal-amd64.iso` runs +entirely in RAM, so the disk is still empty — the next start lands on +`Boot failed: not a bootable disk`. Put the ISO back before installing: + +```bash +virsh -c qemu:///system change-media talos sda /path/to/metal-amd64.iso --config --insert +virt-xml -c qemu:///system talos --edit --boot cdrom,hd +virsh -c qemu:///system start talos +``` + +Wait for the maintenance-mode API, then confirm the install disk's device name — on virtio +it is `/dev/vda`, and Talos's default selector expects `/dev/sda`: + +```bash +talosctl get disks --insecure -n -e +``` + +### Generate the machine config + +Talos 1.14 moved several v1alpha1 fields into their own config documents. In particular +`machine.install` is now `UnattendedInstallConfig`, and patching the old field is rejected +with *"UnattendedInstallConfig config is incompatible with v1alpha1 config"*. Write +`patch.yaml` as a multi-document patch: + +```yaml +machine: + certSANs: + - 192.168.122.33 + registries: + mirrors: + # The in-cluster registry (§2) speaks plain HTTP. + "192.168.122.33:30500": + endpoints: + - http://192.168.122.33:30500 +--- +apiVersion: v1alpha1 +kind: UnattendedInstallConfig +provisioning: + diskSelector: + match: disk.dev_path == "/dev/vda" +``` + +```bash +talosctl gen config big https://:6443 --output-dir ~/.talos/big --config-patch @patch.yaml +talosctl apply-config --insecure -n -e --file ~/.talos/big/controlplane.yaml +``` + +Talos installs to the disk and **kexecs straight into the installed system**, so the CD +boot order doesn't get in the way here. Then point the client at the node and bootstrap: + +```bash +talosctl config merge ~/.talos/big/talosconfig +talosctl config endpoint +talosctl config node +talosctl bootstrap # wait for `talosctl version` to answer first +talosctl kubeconfig -f ~/.kube/config +``` + +A single-node cluster must run workloads on the control plane, or CoreDNS never schedules: + +```bash +kubectl taint node --all node-role.kubernetes.io/control-plane- +``` + +Finally, make a VM restart boot the installed system rather than the ISO (takes effect at +the next full power cycle): + +```bash +virsh -c qemu:///system change-media talos sda --eject --config +virt-xml -c qemu:///system talos --edit --boot hd +``` + +## 2. A registry the node can pull from + +Talos has no Docker daemon and no way to side-load an image, so this repo's images have to +come from a registry. The registry runs **inside the cluster**, published on NodePort +30500 (`infra/helm/registry.yaml`): + +```bash +make k8s-registry +``` + +Why in-cluster rather than on the laptop: a laptop-side registry needs an inbound port +opened on firewalld's `libvirt` zone (`sudo firewall-cmd --zone=libvirt --add-port=5000/tcp`), +which needs root. Pushing from the laptop *to* the node is outbound and always allowed, and +the node pulls from its own NodePort. If you do open that port, put a registry on the +laptop instead and point `K8S_REGISTRY` at `:5000` — the mirror patch in §1 has +an entry ready for it. + +Its storage is `emptyDir`, so if the registry pod is ever replaced, re-run `make k8s-images`. + +## 3. Build and push the images + +```bash +make k8s-images K8S_REGISTRY=:30500 +``` + +This builds the nine images with `docker compose build` — same contexts and Dockerfiles as +compose, no second build definition — then `docker save | crane push --insecure` each one. +`docker push` is not used: the registry speaks plain HTTP, which the Docker daemon refuses +without a root-level `insecure-registries` entry, while crane just takes `--insecure`. + +## 4. Deploy + +```bash +make k8s-up TALOS_HOST= K8S_REGISTRY=:30500 +``` + +That does two things: + +1. `make k8s-seed` — creates the ConfigMaps the chart mounts, from the config files that + already live in this repo (`infra/helm/seed-configmaps.sh`): the four + `setup_configuration/data.yaml` files, the Keycloak realm exports, the BPMN + DMN, and + the two bootstrap scripts. Re-run it after editing any of them. +2. `helm upgrade --install` of the chart into namespace `big`. + +First bring-up takes a few minutes: the four Django services migrate their databases and +apply their `setup_configuration`, Flowable creates its schema, and the bootstrap Jobs +deploy the BPMN/DMN, seed the zaaktype and register the NRC abonnement. + +```bash +kubectl -n big get pods -w +kubectl -n big get jobs # all four must reach COMPLETIONS 1/1 +``` + +The Jobs are the stack's wiring; if one is not complete, the flow is broken somewhere +specific: + +| Job | What breaks without it | +|---|---| +| `flowable-init` | no `registratie` process, no diploma DMN | +| `registerrecord-init` | the register has no RegisterRecord objecttype, so writes are refused | +| `seed-zaaktype` | the ACL can't resolve `BIG-REGISTRATIE`, so no zaak is created | +| `nrc-subscribe` | register writes never reach the projection — the public register stays empty | + +## 5. Use it + +### The portals must be reached over `localhost` + +The portals' OIDC flow uses PKCE, which needs `crypto.subtle` — and browsers only expose +that in a **secure context**: HTTPS, or an origin on `localhost`/`127.0.0.1`. A NodePort on +the VM's IP is neither, so `http://:30140` fails before it can even build the +authorize URL: + +``` +ERROR TypeError: Cannot read properties of undefined (reading 'digest') + at t.calcHash → t.generateCodeChallenge → t.createUrlCodeFlowAuthorize +``` + +So deploy with `TALOS_HOST=localhost` — which pins Keycloak's issuer and the portals' +`config.json` authority to `http://localhost:30180` — and forward the browser-facing +services to those same ports: + +```bash +make k8s-up TALOS_HOST=localhost K8S_REGISTRY=:30500 +make k8s-portals # stays in the foreground; Ctrl-C stops all five forwards +``` + +| URL (needs `make k8s-portals`) | What | +|---|---| +| `http://localhost:30140` | self-service portal (DigiD) | +| `http://localhost:30141` | openbaar register (anonymous) | +| `http://localhost:30142` | behandel portal (medewerker) | +| `http://localhost:30143` | beheer portal (medewerker) | +| `http://localhost:30180` | Keycloak (admin/admin) | + +The port numbers are deliberately the NodePort numbers: Keycloak's issuer is one fixed +string, so the port the browser uses has to match the one baked into `config.json`. + +This is the same mechanism `infra/host-browser.yml` uses for the compose stack (which pins +`localhost:8180`); only the addresses differ. + +### The admin UIs work straight off the NodePorts + +These are server-rendered and need no secure context, so they are reachable at the VM's +address with no forwarding: + +| URL | What | +|---|---| +| `http://:30000` | OpenZaak admin (admin/admin) | +| `http://:30001` | Open Notificaties admin (admin/admin) | +| `http://:30020` / `:30021` | Objecttypen / Objecten admin | +| `http://:30080` | BFF (`/health`) | +| `http://:30090` | Flowable REST (rest-admin/test) | + +### Credentials + +Log in with the test users from `docs/synthetic-data.md` (all password `test123`, e.g. +`jan-burger` for self-service, `merel-behandelaar` for behandel). The `medewerker` realm +enforces MFA (ADR-0031) — print a current code with +`python3 infra/keycloak/check_realms.py otp`. Walk the flow in `docs/demo-script.md`. + +`TALOS_HOST` is not cosmetic: it pins Keycloak's issuer (`KC_HOSTNAME`) and the portals' +OIDC authority to the same string, which is what makes a browser token pass the BFF's +validation (ADR-0010). Change it and you must re-run `make k8s-up` — the chart rolls the +portals for you, because their `config.json` is a subPath mount and would otherwise keep +serving the old authority. + +### Smoke-test the whole chain without a browser + +With the forwards running: + +```bash +TOK=$(curl -s -X POST http://localhost:30180/realms/digid/protocol/openid-connect/token \ + -d grant_type=password -d client_id=big-portal \ + -d username=jan-burger -d password=test123 -d scope=openid | jq -r .access_token) + +# through the portal's Caddy, so this also proves the BFF reverse proxy +curl -s -X POST http://localhost:30140/self-service/registrations \ + -H "Authorization: Bearer $TOK" -H 'Content-Length: 0' +# → {"registrationId":"…","status":"Ingediend"} + +curl -s http://localhost:30141/openbaar/register +# → [{"id":"…","status":"INGEDIEND","reference":""}] +``` + +The second call proves the whole Common Ground path: portal → BFF → domain → Flowable → +ACL → OpenZaak + Objecten → NRC → event-subscriber → projection → openbaar register. + +## 6. Keeping the databases (recommended if you iterate on the chart) + +By default every database is an `emptyDir`: no CSI driver needed, and the data lives as +long as the pod. Note what that means in practice — **any** change to a database pod's +template (an image policy, an env value, a probe) recreates the pod and wipes it. The stack +then needs its bootstrap re-run: + +```bash +make k8s-reseed TALOS_HOST=... K8S_REGISTRY=... +``` + +which re-runs the four Jobs *and* restarts `event-subscriber` + `projection-api`, because +those two create the projection schema on start and otherwise keep writing to a +schema-less database (`relation "processed_notifications" does not exist`). For persistence, install Rancher's local-path-provisioner — on Talos it +must write under `/var` and its namespace needs the privileged Pod Security label: + +```yaml +# kustomization.yaml +apiVersion: kustomize.config.k8s.io/v1beta1 +kind: Kustomization +resources: + - github.com/rancher/local-path-provisioner/deploy?ref=v0.0.31 +patches: + - patch: |- + kind: ConfigMap + apiVersion: v1 + metadata: + name: local-path-config + namespace: local-path-storage + data: + config.json: |- + { "nodePathMap":[ { "node":"DEFAULT_PATH_FOR_NON_LISTED_NODES", "paths":["/var/local-path-provisioner"] } ] } + - patch: |- + apiVersion: v1 + kind: Namespace + metadata: + name: local-path-storage + labels: + pod-security.kubernetes.io/enforce: privileged +``` + +```bash +kubectl apply -k . +make k8s-up TALOS_HOST=... K8S_REGISTRY=... K8S_SET='--set persistence.storageClass=local-path' +``` + +The PVCs carry `helm.sh/resource-policy: keep`, so `make k8s-down` leaves the data behind; +`make k8s-purge` drops the namespace and with it the volumes. + +## 7. Day-to-day + +```bash +make k8s-lint # render + schema-check the chart, no cluster needed +make k8s-portals # forward the portals + Keycloak to localhost (browser access) +make k8s-images K8S_REGISTRY=... # after changing a service or a portal +make k8s-up TALOS_HOST=... K8S_REGISTRY=... +make k8s-seed # after editing a data.yaml, a realm export, or the BPMN +make k8s-reseed TALOS_HOST=... K8S_REGISTRY=... # re-run the bootstrap Jobs + reset the projection schema +make k8s-down # uninstall, keep the database PVCs +make k8s-purge # uninstall and drop the namespace +``` + +This repo's images are pulled with `imagePullPolicy: Always` (the `dev` tag is mutable), so +`kubectl -n big rollout restart deploy/` after `make k8s-images` picks up a rebuild. +Upstream images stay `IfNotPresent`: their tags are pinned, and keeping them out of the pod +template avoids needless churn — a changed template makes a Job unpatchable. + +`k8s-reseed` is also the path for *changing* a Job in the chart: a Job's pod template is +immutable, so `helm upgrade` is rejected with `cannot patch "…" with kind Job`. + +## 8. When it doesn't work + +| Symptom | Cause | +|---|---| +| `Boot failed: not a bootable disk` | virt-manager ejected the install ISO on first shutdown — see §1 | +| The VM comes back in maintenance mode after a restart | the ISO is still attached and boots first; eject it and set `--boot hd` (§1) | +| `apply-config` rejects the patch with *"incompatible with v1alpha1"* | Talos ≥1.14 owns that field in its own config document — patch the document, not `machine.*` (§1) | +| CoreDNS `Pending` forever | the control-plane taint is still on the only node (§1) | +| `ImagePullBackOff` … `pull QPS exceeded` | transient: the kubelet rate-limits pulls when ~30 pods start at once. It recovers on retry | +| `ImagePullBackOff` on a `register-referentie/*` image | the registry mirror patch is missing: `talosctl get registriesconfig` | +| Pod stuck in `ContainerCreating`, event names a ConfigMap | `make k8s-seed` | +| `seed-zaaktype` retrying | publishing a zaaktype validates the resultaattype against `selectielijst.openzaak.nl`, so this one Job needs outbound internet from the VM (ADR-0006) | +| `TypeError: Cannot read properties of undefined (reading 'digest')` on a portal | not a secure context: `crypto.subtle` is absent on `http://`. Use `localhost` + `make k8s-portals` (§5) | +| Login redirects but the portal stays logged out, or the BFF answers 401 | `TALOS_HOST` doesn't match the address in the browser's URL bar — issuer mismatch. Re-run `make k8s-up` with the right value | +| A portal returns 502 on `/self-service/…` | the BFF is unreachable from the portal pod: check `kubectl -n big get svc bff` and the BFF's own readiness | +| Public register empty after a submit | usually a wiped `emptyDir` database (§6): `make k8s-reseed`. Confirm with `kubectl -n big logs deploy/event-subscriber \| grep 42P01` | +| `helm upgrade` fails with `cannot patch … with kind Job` | see §7 — use `make k8s-reseed` | +| Pods `Evicted` / `OOMKilled` | the VM is too small (§0) | +| A Job shows `BackoffLimitExceeded` | read it: `kubectl -n big logs job/` | + +## What is not ported + +- **Observability** (Tempo, Prometheus, Grafana) is defined but disabled — those are built + images too, so switching them on means pushing them as well: + `K8S_SET='--set workloads.tempo.enabled=true --set workloads.prometheus.enabled=true --set workloads.grafana.enabled=true'`. + The .NET services still export OTLP; the exporter fails harmlessly when Tempo is absent. +- **The verify/e2e lanes.** `make verify*` and the Playwright e2e drive compose, not the + chart. The Kubernetes path is verified with §5's smoke test. +- **Ingress, TLS, and resource requests.** See the ponytail ceiling in ADR-0033.