## What & why After the Gitea 1.27 + act_runner 2.0.0 upgrade, `verify-stack` never starts: the run sits in `waiting` forever with no logs for that job, while the other five jobs pass — so `main` stays pending/red (P0). See #134. Closes #134 ### Root cause `verify-stack` was the only job gated by a status-function `if` on top of `needs`: ```yaml verify-stack: needs: [mutation] if: ${{ !cancelled() }} ``` Gitea 1.27 reworked cancellation/aggregation so that `always()`/`cancelled()`-gated `needs` jobs route through a new transitional **`Cancelling`** state + server↔runner **capability negotiation** ("Requires Gitea Runner 2.0.0"). On this 1.27 + 2.0.0 pairing that handshake doesn't resolve, so the job is never dispatched and never leaves `waiting`. Plain jobs (no `if`/`needs`) are unaffected — exactly the observed pattern. It worked pre-upgrade (old runner). ### Fix Drop the `if: ${{ !cancelled() }}`; keep `needs: [mutation]`. Default `if: success()` dispatches normally and still serialises the two memory-heavy jobs (OOM avoidance, #126). **Trade-off:** the `!cancelled()` (added in #127) let verify-stack run even when the mutation ratchet fails. Now a failing mutation skips verify-stack; the fix-and-re-push re-run exercises it, so the signal isn't lost — just deferred to the green-mutation run. If we later want both signals on one run, serialise via a `concurrency` group rather than `needs` + `always()`. Documented as §7 in `docs/runbooks/gitea-actions-gotchas.md`. ## Note on the stuck run Run 582 (the #133 merge) will **not** clear itself and must be force-cancelled from the Actions UI (plain cancel can also stall on this version, gitea#35782). This PR's own run is the first real test of the fix — if `verify-stack` dispatches and runs here, the fix holds. ## Definition of Done - [x] Linked issue (#134). - [x] Conventional Commit referencing the issue. - [ ] CI green — this PR's run is the verification (verify-stack must dispatch). - [x] Runbook updated (gotchas §7). - [ ] Closed by the merging PR. 🤖 Generated with [Claude Code](https://claude.com/claude-code)Reviewed-on: #135
188 lines
8.0 KiB
YAML
188 lines
8.0 KiB
YAML
name: CI
|
|
|
|
on:
|
|
push:
|
|
branches: [main]
|
|
pull_request:
|
|
branches: [main]
|
|
|
|
permissions:
|
|
contents: read
|
|
|
|
# Supersede stale runs: a new push to the same branch/PR cancels the previous run, so the runner's
|
|
# concurrency slots aren't spent on commits nobody is waiting for (refs #127).
|
|
concurrency:
|
|
group: ci-${{ github.workflow }}-${{ github.ref }}
|
|
cancel-in-progress: true
|
|
|
|
# Self-hosted runner — see docs/runbooks/ci.md for the runner setup.
|
|
# `uses:` are absolute, tag-pinned URLs (CLAUDE.md §8.7 / §15).
|
|
|
|
# Each job calls a `make` target — the same one developers run locally
|
|
# (`make ci`). The Makefile is the single source of truth; see docs/runbooks/ci.md.
|
|
|
|
jobs:
|
|
lint:
|
|
runs-on: ubuntu-latest
|
|
steps:
|
|
- uses: https://github.com/actions/checkout@v4
|
|
- uses: https://github.com/actions/setup-dotnet@v4
|
|
with:
|
|
dotnet-version: '10.0.x'
|
|
# Cache the NuGet package store so each .NET job restores from disk, not the network. There are
|
|
# no lock files (so setup-dotnet's built-in cache doesn't apply); key on the project files. @v3
|
|
# avoids the GHES guard that breaks @v4 on Gitea (gitea-actions-gotchas.md); cache is best-effort
|
|
# — a miss just restores from the network. See issue #73.
|
|
- uses: https://github.com/actions/cache@v3
|
|
with:
|
|
path: ~/.nuget/packages
|
|
key: nuget-${{ runner.os }}-${{ hashFiles('**/*.csproj') }}
|
|
restore-keys: |
|
|
nuget-${{ runner.os }}-
|
|
- run: make lint
|
|
|
|
build:
|
|
runs-on: ubuntu-latest
|
|
steps:
|
|
- uses: https://github.com/actions/checkout@v4
|
|
- uses: https://github.com/actions/setup-dotnet@v4
|
|
with:
|
|
dotnet-version: '10.0.x'
|
|
- uses: https://github.com/actions/cache@v3
|
|
with:
|
|
path: ~/.nuget/packages
|
|
key: nuget-${{ runner.os }}-${{ hashFiles('**/*.csproj') }}
|
|
restore-keys: |
|
|
nuget-${{ runner.os }}-
|
|
- run: make build
|
|
|
|
unit:
|
|
runs-on: ubuntu-latest
|
|
steps:
|
|
- uses: https://github.com/actions/checkout@v4
|
|
- uses: https://github.com/actions/setup-dotnet@v4
|
|
with:
|
|
dotnet-version: '10.0.x'
|
|
- uses: https://github.com/actions/cache@v3
|
|
with:
|
|
path: ~/.nuget/packages
|
|
key: nuget-${{ runner.os }}-${{ hashFiles('**/*.csproj') }}
|
|
restore-keys: |
|
|
nuget-${{ runner.os }}-
|
|
- run: make unit
|
|
|
|
# Frontend (Nx/Angular) lane: install with pnpm, then Nx lint + test + build.
|
|
frontend:
|
|
runs-on: ubuntu-latest
|
|
steps:
|
|
- uses: https://github.com/actions/checkout@v4
|
|
- uses: https://github.com/pnpm/action-setup@v4
|
|
with:
|
|
version: 11
|
|
- uses: https://github.com/actions/setup-node@v4
|
|
with:
|
|
node-version: '24'
|
|
cache: 'pnpm'
|
|
- run: make frontend
|
|
|
|
mutation:
|
|
runs-on: ubuntu-latest
|
|
steps:
|
|
- uses: https://github.com/actions/checkout@v4
|
|
- uses: https://github.com/actions/setup-dotnet@v4
|
|
with:
|
|
dotnet-version: '10.0.x'
|
|
- uses: https://github.com/actions/cache@v3
|
|
with:
|
|
path: ~/.nuget/packages
|
|
key: nuget-${{ runner.os }}-${{ hashFiles('**/*.csproj') }}
|
|
restore-keys: |
|
|
nuget-${{ runner.os }}-
|
|
- run: make mutation
|
|
# Publish the Stryker HTML reports. `if: always()` uploads them even when the
|
|
# ratchet fails — that is exactly when you want to inspect the survivors.
|
|
# `continue-on-error` keeps the upload best-effort: the mutation *gate* is the
|
|
# ratchet (make mutation's exit code), not the report, so a Gitea artifact-backend
|
|
# 500 must not fail the job (gitea-actions-gotchas.md §4). Glob handles Stryker's
|
|
# non-deterministic StrykerOutput/<timestamp>/ dir. Pinned @v3: @v4's bundled
|
|
# @actions/artifact hard-aborts on non-github.com (GHES guard) — see the runbook.
|
|
- uses: https://github.com/actions/upload-artifact@v3
|
|
if: always()
|
|
continue-on-error: true
|
|
with:
|
|
name: acl-mutation-report
|
|
path: services/acl/StrykerOutput/**/reports/mutation-report.html
|
|
if-no-files-found: warn
|
|
- uses: https://github.com/actions/upload-artifact@v3
|
|
if: always()
|
|
continue-on-error: true
|
|
with:
|
|
name: event-subscriber-mutation-report
|
|
path: services/event-subscriber/StrykerOutput/**/reports/mutation-report.html
|
|
if-no-files-found: warn
|
|
- uses: https://github.com/actions/upload-artifact@v3
|
|
if: always()
|
|
continue-on-error: true
|
|
with:
|
|
name: domain-mutation-report
|
|
path: services/domain/StrykerOutput/**/reports/mutation-report.html
|
|
if-no-files-found: warn
|
|
- uses: https://github.com/actions/upload-artifact@v3
|
|
if: always()
|
|
continue-on-error: true
|
|
with:
|
|
name: bff-mutation-report
|
|
path: services/bff/StrykerOutput/**/reports/mutation-report.html
|
|
if-no-files-found: warn
|
|
|
|
# One stage for every check that needs the live stack. Booting OpenZaak once (instead
|
|
# of once per job) is the cheapest layout (issue #58). No setup-dotnet: the ACL test runs
|
|
# in a built image and everything reaches services by container IP. Needs Docker + egress
|
|
# (base images, nuget, selectielijst.openzaak.nl).
|
|
#
|
|
# `needs: [mutation]` is NOT a data dependency — it serialises the two memory-heavy jobs so
|
|
# they never co-schedule now the runner has capacity >1. A concurrent Stryker run + full-stack
|
|
# bring-up + Playwright browser on one host is what OOMs the e2e (commit d5e5fa2, #126). The
|
|
# light .NET/frontend jobs have no `needs`, so they still parallelise up to runner capacity.
|
|
#
|
|
# No `if: ${{ !cancelled() }}` here (removed in #134): on Gitea 1.27 + act_runner 2.0.0, a job
|
|
# gated by a status-function `if` (always()/cancelled()) on top of `needs` routes through the new
|
|
# transitional "Cancelling" state + capability negotiation and never leaves `waiting` — it's never
|
|
# dispatched (gitea-actions-gotchas.md §7). Default `if: success()` dispatches normally. Cost: a
|
|
# failing mutation ratchet now skips verify-stack instead of running it anyway; the fix-and-re-push
|
|
# re-run exercises verify-stack, so we still get the signal.
|
|
verify-stack:
|
|
needs: [mutation]
|
|
runs-on: ubuntu-latest
|
|
steps:
|
|
- uses: https://github.com/actions/checkout@v4
|
|
# Bring the full stack up + wait for health — this also is the DoD "compose up
|
|
# reaches green health" smoke (it replaces the old compose-smoke job).
|
|
- name: Bring up the full stack & wait for health
|
|
run: make verify-up
|
|
- name: Observability backplane (Grafana + Tempo + Prometheus datasources)
|
|
run: OBS_TIMEOUT=180 make verify-observability
|
|
- name: ACL ↔ OpenZaak integration tests
|
|
run: make verify-acl
|
|
- name: OpenZaak → NRC notification delivery
|
|
run: make verify-nrc
|
|
- name: OpenZaak → NRC → Event Subscriber → projection-api
|
|
run: make verify-projection
|
|
- name: Domain → Flowable → ACL → OpenZaak
|
|
run: make verify-domain
|
|
- name: BFF → Keycloak + domain + projection
|
|
run: make verify-bff
|
|
- name: Distributed traces reach Tempo (one connected trace across services)
|
|
run: TRACING_TIMEOUT=120 make verify-tracing
|
|
- name: Golden-signal metrics scraped by Prometheus (/metrics on every service)
|
|
run: METRICS_TIMEOUT=120 make verify-metrics
|
|
- name: Self-service e2e (Playwright, login → submit → success)
|
|
run: make verify-e2e
|
|
# Log dump must precede teardown (which removes the containers).
|
|
- name: Dump container logs on failure
|
|
if: failure()
|
|
run: docker compose -f infra/docker-compose.yml logs --no-color --tail=100 oz-init openzaak nrc-init nrc-web nrc-celery nrc-beat flowable-db flowable-rest flowable-init keycloak acl bff domain projection-db event-subscriber projection-api self-service openbaar behandel beheer tempo prometheus grafana 2>&1 || true
|
|
- name: Tear down
|
|
if: always()
|
|
run: make down
|