Compare commits

..
Author SHA1 Message Date
notandClaude Opus 5 dfa1a370ac ci(deploy): publish the stack over TLS when PUBLIC_DOMAIN is set (refs #175)
The in-cluster edge (#178) is off unless the chart is given a domain, so pass
one through from a repository variable. Unset, the deploy is exactly what it was
— NodePorts, and the portals reachable only over the SSH forwards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 16:31:47 +02:00
notandClaude Opus 5 e8cb1ec7e9 docs(k8s): how to reach the deployed portals from a laptop (refs #175)
The five forwards are not optional and not independent: the portals' OIDC
authority is pinned to localhost:30180, so forwarding the portal without
Keycloak gets ERR_CONNECTION_REFUSED on the discovery document and an opaque
"[object Object]" in the console. One ssh replaces `make k8s-portals` for the
lab server, and needs no kubeconfig.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 16:19:02 +02:00
notandClaude Opus 5 de6db7d35b docs(k8s): split the taint fix from the registry patch (refs #175)
Talos 1.14 rejects a `patch mc` that sets `cluster.allowSchedulingOnControlPlanes`
— the field left the v1alpha1 schema, the way `machine.install` did — and the
rejection discards the rest of the patch with it, so the registry mirror never
lands and the failure moves from Pending pods to ImagePullBackOff without ever
saying so. Document the two as separate steps, and correct the claim that the
taint has to be patched away: `kubectl taint` is what §1 prescribes and it holds
until the node re-registers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 16:06:24 +02:00
notandClaude Opus 5 d17e79959b ci(deploy): say why a deploy stalled, and document the cluster prerequisite (refs #175)
The first run against the lab server's VM timed out on `kubectl rollout status`
for the in-cluster registry with nothing but "timed out waiting for the
condition". The cause was three commands up the runbook: that VM was installed
from a stock Talos config, so the only node still carries the control-plane
taint and no pod can schedule — and the missing registry mirror would have
failed the image pulls right after.

`rollout status` can only ever report the symptom, so dump the whole cluster's
pods and the recent events on failure instead of `big`'s pods alone; the
scheduler's "untolerated taint" message is the answer and it lives in the
events. Runbook §9 now opens with the machine-config patch the deploy assumes,
as one applied-live patch rather than a `kubectl taint` that the controller
undoes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 16:01:47 +02:00
notandClaude Opus 5 fc036d53d5 ci(deploy): deploy the stack to Talos on merge to main (refs #175)
CI / k8s (pull_request) Successful in 5s
CI / lint (pull_request) Successful in 1m25s
CI / build (pull_request) Successful in 1m22s
CI / unit (pull_request) Successful in 1m24s
CI / frontend (pull_request) Successful in 1m45s
CI / mutation (pull_request) Successful in 3m1s
CI / verify-stack (pull_request) Successful in 6m22s
The chart has been deployable by hand since #25 and linted in CI since #168;
this makes a merged PR actually ship it to the lab server's Talos VM.

Neither the Kubernetes API nor the in-cluster registry is publicly reachable,
so the job forwards 6443, 30500 and 30141 over the same SSH hop into the Fedora
host that the Gitea-runner pipeline uses. That splits the registry into two
names for one store: images are pushed through the tunnel to localhost:30500,
and the node pulls them from its own NodePort — the address its registry-mirror
patch trusts over plain HTTP.

It deploys with `make k8s-reseed` rather than `make k8s-up`: the bootstrap Jobs
are idempotent, and deleting them first is what stops a changed Job template
from wedging `helm upgrade`. The nine deployments are then rolled explicitly,
because `dev` is a mutable tag and helm sees an unchanged pod template.

PR CI is the merge gate, so this workflow does not re-run the checks. Deploys
queue instead of cancelling: a `helm upgrade` killed half-way leaves the release
in `pending-upgrade` and needs unwedging by hand.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 15:40:42 +02:00
2 changed files with 69 additions and 3 deletions
+19 -3
View File
@@ -27,6 +27,10 @@ jobs:
# `kubectl port-forward` — runbook §5. Override with repo variables. # `kubectl port-forward` — runbook §5. Override with repo variables.
TALOS_VM_IP: ${{ vars.TALOS_VM_IP }} TALOS_VM_IP: ${{ vars.TALOS_VM_IP }}
TALOS_HOST: ${{ vars.TALOS_HOST }} TALOS_HOST: ${{ vars.TALOS_HOST }}
# Set it and the stack is published over TLS on <sub>.<domain> by the
# in-cluster edge (ADR-0035, runbook §10). Empty = NodePorts, as before.
PUBLIC_DOMAIN: ${{ vars.PUBLIC_DOMAIN }}
PUBLIC_EMAIL: ${{ vars.PUBLIC_EMAIL }}
steps: steps:
- uses: https://github.com/actions/checkout@v4 - uses: https://github.com/actions/checkout@v4
@@ -89,7 +93,13 @@ jobs:
# The jobs are idempotent, and deleting them first is what keeps a changed # The jobs are idempotent, and deleting them first is what keeps a changed
# Job template from wedging the upgrade (`cannot patch … with kind Job`). # Job template from wedging the upgrade (`cannot patch … with kind Job`).
- name: Deploy the chart - name: Deploy the chart
run: make k8s-reseed TALOS_HOST=${TALOS_HOST:-localhost} K8S_REGISTRY=${TALOS_VM_IP:-192.168.122.173}:30500 run: |
set -euo pipefail
publish="${PUBLIC_DOMAIN:+--set public.domain=$PUBLIC_DOMAIN --set public.email=${PUBLIC_EMAIL:-}}"
make k8s-reseed \
TALOS_HOST=${TALOS_HOST:-localhost} \
K8S_REGISTRY=${TALOS_VM_IP:-192.168.122.173}:30500 \
K8S_SET="$publish"
# `dev` is a mutable tag and helm sees an unchanged pod template, so the # `dev` is a mutable tag and helm sees an unchanged pod template, so the
# new images only land on a restart (pullPolicy is already Always). # new images only land on a restart (pullPolicy is already Always).
@@ -105,6 +115,12 @@ jobs:
- name: Smoke the public register - name: Smoke the public register
run: curl -fsS --retry 10 --retry-delay 6 --retry-all-errors http://localhost:30141/openbaar/register run: curl -fsS --retry 10 --retry-delay 6 --retry-all-errors http://localhost:30141/openbaar/register
- name: Pods on failure # Cluster-wide, not just `big`: the first thing that can fail is the registry
# in its own namespace, and a scheduling problem shows up in the events, not
# in `rollout status` — which only ever says "timed out waiting".
- name: Pods and events on failure
if: failure() if: failure()
run: kubectl -n big get pods,jobs || true run: |
kubectl get pods -A -o wide || true
kubectl -n big get jobs || true
kubectl get events -A --sort-by=.lastTimestamp | tail -30 || true
+50
View File
@@ -366,6 +366,36 @@ immutable, so `helm upgrade` is rejected with `cannot patch "…" with kind Job`
every time a PR is squash-merged to `main` (and on demand via *Run workflow*). PR CI is the every time a PR is squash-merged to `main` (and on demand via *Run workflow*). PR CI is the
merge gate, so the workflow deploys without re-running the checks. merge gate, so the workflow deploys without re-running the checks.
**Prerequisite: the VM must have been installed with the §1 patch.** A stock Talos config
gives you a node that still carries the control-plane taint and knows nothing about the
plain-HTTP registry, and the deploy hits those in that order: the `registry` pod sits
`Pending` until `rollout status` times out, and once that is fixed every repo image fails to
pull. Two separate fixes:
```bash
# on the Fedora host — 1. let workloads onto the only node (§1)
export KUBECONFIG=~/talos-kubeconfig-local
kubectl taint node --all node-role.kubernetes.io/control-plane-
# 2. trust the in-cluster registry over plain HTTP (§2)
cat > /tmp/registry-patch.yaml <<'YAML'
machine:
registries:
mirrors:
"<TALOS_VM_IP>:30500":
endpoints:
- http://<TALOS_VM_IP>:30500
YAML
talosctl -n <TALOS_VM_IP> -e <TALOS_VM_IP> patch mc --patch @/tmp/registry-patch.yaml
```
Keep those two apart. On Talos 1.14 a patch that also sets
`cluster.allowSchedulingOnControlPlanes` is rejected with *".cluster.allowSchedulingOnControlPlanes
is already set in v1alpha1 config"* — the field moved out of the v1alpha1 schema, the same way
`machine.install` did (§1) — and the rejection takes the whole patch with it, so the mirror
silently doesn't land either. `kubectl taint` is the documented way (§1); it is undone if the
node ever re-registers, which is a reboot, not a deploy.
The cluster's API and registry are not exposed publicly, so the job forwards them over the The cluster's API and registry are not exposed publicly, so the job forwards them over the
same SSH hop the Gitea-runner pipeline uses: same SSH hop the Gitea-runner pipeline uses:
@@ -397,6 +427,26 @@ Settings, all on the repository in Gitea:
The last step smokes `GET /openbaar/register` through the openbaar portal, which exercises The last step smokes `GET /openbaar/register` through the openbaar portal, which exercises
portal → Caddy → BFF → projection. An empty register passes; a 502 does not. portal → Caddy → BFF → projection. An empty register passes; a 502 does not.
### Reaching the portals from a laptop
The deployed portals are pinned to `http://localhost:30180` for Keycloak (§5), so a browser
needs **all five** browser-facing ports on its own localhost — the portal alone is not
enough, and a missing Keycloak shows up as `ERR_CONNECTION_REFUSED` on
`/realms/*/.well-known/openid-configuration` followed by an opaque `ERROR Error: [object Object]`.
`make k8s-portals` does this when kubectl can reach the cluster; through the lab server one
SSH does it without a kubeconfig at all:
```bash
ssh -N -p 6667 \
-L 30140:<TALOS_VM_IP>:30140 -L 30141:<TALOS_VM_IP>:30141 \
-L 30142:<TALOS_VM_IP>:30142 -L 30143:<TALOS_VM_IP>:30143 \
-L 30180:<TALOS_VM_IP>:30180 \
user@labs.respellion.tech
```
Then the §5 table's URLs work as written. The admin UIs (OpenZaak, Flowable, …) need no
forward — they are server-rendered, so the VM's address is fine.
Not covered: the portals still need `make k8s-portals` (or an SSH forward) to be usable in a Not covered: the portals still need `make k8s-portals` (or an SSH forward) to be usable in a
browser, because PKCE needs a secure context (§5). Giving the server a hostname + TLS is the browser, because PKCE needs a secure context (§5). Giving the server a hostname + TLS is the
upgrade path. upgrade path.