Compare commits
3
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
c7f06b35fa | ||
|
|
804031eeb8 | ||
|
|
6cfcc4cf83 |
@@ -27,10 +27,10 @@ jobs:
|
||||
# `kubectl port-forward` — runbook §5. Override with repo variables.
|
||||
TALOS_VM_IP: ${{ vars.TALOS_VM_IP }}
|
||||
TALOS_HOST: ${{ vars.TALOS_HOST }}
|
||||
# Set it and the stack is published over TLS on <sub>.<domain> by the
|
||||
# in-cluster edge (ADR-0035, runbook §10). Empty = NodePorts, as before.
|
||||
PUBLIC_DOMAIN: ${{ vars.PUBLIC_DOMAIN }}
|
||||
PUBLIC_EMAIL: ${{ vars.PUBLIC_EMAIL }}
|
||||
# Set it when the labs Caddy publishes the portals: Keycloak's public https
|
||||
# origin, e.g. https://big-auth.labs.respellion.tech (runbook, "Publishing
|
||||
# through the labs Caddy").
|
||||
KEYCLOAK_URL: ${{ vars.KEYCLOAK_URL }}
|
||||
steps:
|
||||
- uses: https://github.com/actions/checkout@v4
|
||||
|
||||
@@ -94,12 +94,10 @@ jobs:
|
||||
# Job template from wedging the upgrade (`cannot patch … with kind Job`).
|
||||
- name: Deploy the chart
|
||||
run: |
|
||||
set -euo pipefail
|
||||
publish="${PUBLIC_DOMAIN:+--set public.domain=$PUBLIC_DOMAIN --set public.email=${PUBLIC_EMAIL:-}}"
|
||||
make k8s-reseed \
|
||||
TALOS_HOST=${TALOS_HOST:-localhost} \
|
||||
K8S_REGISTRY=${TALOS_VM_IP:-192.168.122.173}:30500 \
|
||||
K8S_SET="$publish"
|
||||
K8S_SET="${KEYCLOAK_URL:+--set keycloakUrl=$KEYCLOAK_URL}"
|
||||
|
||||
# `dev` is a mutable tag and helm sees an unchanged pod template, so the
|
||||
# new images only land on a restart (pullPolicy is already Always).
|
||||
@@ -115,12 +113,6 @@ jobs:
|
||||
- name: Smoke the public register
|
||||
run: curl -fsS --retry 10 --retry-delay 6 --retry-all-errors http://localhost:30141/openbaar/register
|
||||
|
||||
# Cluster-wide, not just `big`: the first thing that can fail is the registry
|
||||
# in its own namespace, and a scheduling problem shows up in the events, not
|
||||
# in `rollout status` — which only ever says "timed out waiting".
|
||||
- name: Pods and events on failure
|
||||
- name: Pods on failure
|
||||
if: failure()
|
||||
run: |
|
||||
kubectl get pods -A -o wide || true
|
||||
kubectl -n big get jobs || true
|
||||
kubectl get events -A --sort-by=.lastTimestamp | tail -30 || true
|
||||
run: kubectl -n big get pods,jobs || true
|
||||
|
||||
@@ -366,36 +366,6 @@ immutable, so `helm upgrade` is rejected with `cannot patch "…" with kind Job`
|
||||
every time a PR is squash-merged to `main` (and on demand via *Run workflow*). PR CI is the
|
||||
merge gate, so the workflow deploys without re-running the checks.
|
||||
|
||||
**Prerequisite: the VM must have been installed with the §1 patch.** A stock Talos config
|
||||
gives you a node that still carries the control-plane taint and knows nothing about the
|
||||
plain-HTTP registry, and the deploy hits those in that order: the `registry` pod sits
|
||||
`Pending` until `rollout status` times out, and once that is fixed every repo image fails to
|
||||
pull. Two separate fixes:
|
||||
|
||||
```bash
|
||||
# on the Fedora host — 1. let workloads onto the only node (§1)
|
||||
export KUBECONFIG=~/talos-kubeconfig-local
|
||||
kubectl taint node --all node-role.kubernetes.io/control-plane-
|
||||
|
||||
# 2. trust the in-cluster registry over plain HTTP (§2)
|
||||
cat > /tmp/registry-patch.yaml <<'YAML'
|
||||
machine:
|
||||
registries:
|
||||
mirrors:
|
||||
"<TALOS_VM_IP>:30500":
|
||||
endpoints:
|
||||
- http://<TALOS_VM_IP>:30500
|
||||
YAML
|
||||
talosctl -n <TALOS_VM_IP> -e <TALOS_VM_IP> patch mc --patch @/tmp/registry-patch.yaml
|
||||
```
|
||||
|
||||
Keep those two apart. On Talos 1.14 a patch that also sets
|
||||
`cluster.allowSchedulingOnControlPlanes` is rejected with *".cluster.allowSchedulingOnControlPlanes
|
||||
is already set in v1alpha1 config"* — the field moved out of the v1alpha1 schema, the same way
|
||||
`machine.install` did (§1) — and the rejection takes the whole patch with it, so the mirror
|
||||
silently doesn't land either. `kubectl taint` is the documented way (§1); it is undone if the
|
||||
node ever re-registers, which is a reboot, not a deploy.
|
||||
|
||||
The cluster's API and registry are not exposed publicly, so the job forwards them over the
|
||||
same SSH hop the Gitea-runner pipeline uses:
|
||||
|
||||
@@ -427,30 +397,46 @@ Settings, all on the repository in Gitea:
|
||||
The last step smokes `GET /openbaar/register` through the openbaar portal, which exercises
|
||||
portal → Caddy → BFF → projection. An empty register passes; a 502 does not.
|
||||
|
||||
### Reaching the portals from a laptop
|
||||
|
||||
The deployed portals are pinned to `http://localhost:30180` for Keycloak (§5), so a browser
|
||||
needs **all five** browser-facing ports on its own localhost — the portal alone is not
|
||||
enough, and a missing Keycloak shows up as `ERR_CONNECTION_REFUSED` on
|
||||
`/realms/*/.well-known/openid-configuration` followed by an opaque `ERROR Error: [object Object]`.
|
||||
`make k8s-portals` does this when kubectl can reach the cluster; through the lab server one
|
||||
SSH does it without a kubeconfig at all:
|
||||
|
||||
```bash
|
||||
ssh -N -p 6667 \
|
||||
-L 30140:<TALOS_VM_IP>:30140 -L 30141:<TALOS_VM_IP>:30141 \
|
||||
-L 30142:<TALOS_VM_IP>:30142 -L 30143:<TALOS_VM_IP>:30143 \
|
||||
-L 30180:<TALOS_VM_IP>:30180 \
|
||||
user@labs.respellion.tech
|
||||
```
|
||||
|
||||
Then the §5 table's URLs work as written. The admin UIs (OpenZaak, Flowable, …) need no
|
||||
forward — they are server-rendered, so the VM's address is fine.
|
||||
|
||||
Not covered: the portals still need `make k8s-portals` (or an SSH forward) to be usable in a
|
||||
browser, because PKCE needs a secure context (§5). Giving the server a hostname + TLS is the
|
||||
upgrade path.
|
||||
|
||||
## Publishing through the labs Caddy
|
||||
|
||||
The portals can be reached on real hostnames through the Caddy that already fronts
|
||||
`*.labs.respellion.tech` (repo `Infra`, `infra/development/`). The chain:
|
||||
|
||||
```
|
||||
browser → Caddy (labs server, TLS) → openssh-server:3014x/30180
|
||||
→ reverse SSH tunnel → Fedora host → <TALOS_VM_IP>:3014x/30180 (NodePorts)
|
||||
```
|
||||
|
||||
| URL | NodePort |
|
||||
|---|---|
|
||||
| `https://big-register.labs.respellion.tech` | 30141 openbaar |
|
||||
| `https://big-mijn.labs.respellion.tech` | 30140 self-service |
|
||||
| `https://big-behandel.labs.respellion.tech` | 30142 behandel |
|
||||
| `https://big-beheer.labs.respellion.tech` | 30143 beheer |
|
||||
| `https://big-auth.labs.respellion.tech` | 30180 Keycloak (`/admin` blocked) |
|
||||
|
||||
HTTPS makes the portals a secure context, so PKCE works without port-forwards — but
|
||||
Keycloak's issuer must be the public origin. Deploy with it:
|
||||
|
||||
```bash
|
||||
make k8s-up TALOS_HOST=localhost K8S_REGISTRY=<TALOS_HOST>:30500 \
|
||||
K8S_SET="--set keycloakUrl=https://big-auth.labs.respellion.tech"
|
||||
```
|
||||
|
||||
For deploy-on-merge, set the repository variable `KEYCLOAK_URL` to the same value.
|
||||
With it set, the `localhost` port-forwards (§5) no longer log in: the issuer is one string.
|
||||
|
||||
One-time setup:
|
||||
|
||||
1. Fedora host: install `infra/development/big-portals-tunnel.service` from the Infra repo
|
||||
(instructions in the file).
|
||||
2. Labs server: deploy the Infra `Caddyfile` + `compose.yml` (Caddy joins the
|
||||
`openssh_default` network to reach the tunnel ends).
|
||||
|
||||
## What is not ported
|
||||
|
||||
- **Observability** (Tempo, Prometheus, Grafana) is defined but disabled — those are built
|
||||
|
||||
@@ -57,6 +57,9 @@ services:
|
||||
# share this anchor and ignore it — they don't run uwsgi.
|
||||
UWSGI_PROCESSES: "1"
|
||||
UWSGI_THREADS: "2"
|
||||
# Same lever for oz-celery: unset, the worker forks one process per CPU (22 on the lab node,
|
||||
# ~225 MB each), which OOM-killed the shared runner mid-verify-stack. Only celery reads it.
|
||||
CELERY_WORKER_CONCURRENCY: "2"
|
||||
DJANGO_SETTINGS_MODULE: openzaak.conf.docker
|
||||
SECRET_KEY: ${OZ_SECRET_KEY:-dev-only-not-for-production}
|
||||
DB_HOST: oz-db
|
||||
@@ -144,6 +147,8 @@ services:
|
||||
# 1 uWSGI worker, not the image default of 4×4 (#147) — see the oz-env note above.
|
||||
UWSGI_PROCESSES: "1"
|
||||
UWSGI_THREADS: "2"
|
||||
# Two celery workers, not one per CPU — see the oz-env note above.
|
||||
CELERY_WORKER_CONCURRENCY: "2"
|
||||
DJANGO_SETTINGS_MODULE: nrc.conf.docker
|
||||
SECRET_KEY: ${NRC_SECRET_KEY:-dev-only-not-for-production}
|
||||
DB_HOST: nrc-db
|
||||
|
||||
@@ -135,6 +135,14 @@ cluster-internal hosts ({{ .Release.Namespace }}) and the node address
|
||||
{{- end }}
|
||||
{{- end -}}
|
||||
|
||||
{{/*
|
||||
The origin a browser reaches Keycloak on: the issuer Keycloak pins and the
|
||||
authority the portals use, from one place so they cannot drift (ADR-0010).
|
||||
*/}}
|
||||
{{- define "big.keycloakUrl" -}}
|
||||
{{- .Values.keycloakUrl | default (printf "http://%s:%v" .Values.host (index .Values.nodePorts "keycloak")) -}}
|
||||
{{- end -}}
|
||||
|
||||
{{- define "big.labels" -}}
|
||||
app.kubernetes.io/name: {{ .name }}
|
||||
app.kubernetes.io/instance: {{ .root.Release.Name }}
|
||||
|
||||
@@ -40,5 +40,5 @@ metadata:
|
||||
{{- include "big.labels" (dict "root" $ "name" (printf "portal-config-%s" $realm)) | nindent 4 }}
|
||||
data:
|
||||
config.json: |
|
||||
{ "authority": "{{ printf "http://%s:%v" $.Values.host (index $.Values.nodePorts "keycloak") }}/realms/{{ $realm }}" }
|
||||
{ "authority": "{{ include "big.keycloakUrl" $ }}/realms/{{ $realm }}" }
|
||||
{{- end }}
|
||||
|
||||
@@ -28,7 +28,7 @@ spec:
|
||||
{{- range $w.files }}
|
||||
{{- if hasPrefix "portal-config-" .configMap }}
|
||||
annotations:
|
||||
checksum/portal-config: {{ printf "%s|%v" $.Values.host (index $.Values.nodePorts "keycloak") | sha256sum }}
|
||||
checksum/portal-config: {{ include "big.keycloakUrl" $ | sha256sum }}
|
||||
{{- end }}
|
||||
{{- end }}
|
||||
labels:
|
||||
|
||||
@@ -25,6 +25,11 @@
|
||||
# string, so browser tokens and the BFF's discovered issuer agree.
|
||||
host: 192.168.122.100
|
||||
|
||||
# Set when a TLS proxy outside the cluster publishes Keycloak: the full origin, no
|
||||
# trailing slash. It replaces `host` + Keycloak's NodePort as the issuer and the
|
||||
# portals' authority (runbook, "Publishing through the labs Caddy").
|
||||
keycloakUrl: ""
|
||||
|
||||
# Set when pulling from a private registry (e.g. the Gitea Container Registry).
|
||||
imagePullSecrets: []
|
||||
|
||||
@@ -71,6 +76,7 @@ envGroups:
|
||||
oz:
|
||||
UWSGI_PROCESSES: "1"
|
||||
UWSGI_THREADS: "2"
|
||||
CELERY_WORKER_CONCURRENCY: "2"
|
||||
DJANGO_SETTINGS_MODULE: openzaak.conf.docker
|
||||
SECRET_KEY: dev-only-not-for-production
|
||||
DB_HOST: oz-db
|
||||
@@ -93,6 +99,7 @@ envGroups:
|
||||
nrc:
|
||||
UWSGI_PROCESSES: "1"
|
||||
UWSGI_THREADS: "2"
|
||||
CELERY_WORKER_CONCURRENCY: "2"
|
||||
DJANGO_SETTINGS_MODULE: nrc.conf.docker
|
||||
SECRET_KEY: dev-only-not-for-production
|
||||
DB_HOST: nrc-db
|
||||
@@ -268,7 +275,7 @@ workloads:
|
||||
# Pin the issuer to the address the browser uses, and let backchannel calls
|
||||
# keep using keycloak:8080 — the BFF discovers metadata in-cluster and gets
|
||||
# this issuer back, which is what browser tokens carry (infra/host-browser.yml).
|
||||
KC_HOSTNAME: "http://{{ .Values.host }}:{{ index .Values.nodePorts \"keycloak\" }}"
|
||||
KC_HOSTNAME: '{{ include "big.keycloakUrl" . }}'
|
||||
KC_HOSTNAME_BACKCHANNEL_DYNAMIC: "true"
|
||||
ports: [{ name: http, port: 8080 }]
|
||||
# TCP, not /health/ready on the management port: nothing here gates on realm
|
||||
|
||||
Reference in New Issue
Block a user