Compare commits
5
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
4ab2ef4285 | ||
|
|
8474b72bf4 | ||
|
|
271c54197e | ||
|
|
b32c352f20 | ||
|
|
4274fd30d1 |
@@ -142,6 +142,8 @@ jobs:
|
||||
# reaches green health" smoke (it replaces the old compose-smoke job).
|
||||
- name: Bring up the full stack & wait for health
|
||||
run: make verify-up
|
||||
- name: Observability backplane (Grafana + Tempo + Prometheus datasources)
|
||||
run: OBS_TIMEOUT=180 make verify-observability
|
||||
- name: ACL ↔ OpenZaak integration tests
|
||||
run: make verify-acl
|
||||
- name: OpenZaak → NRC notification delivery
|
||||
@@ -152,12 +154,14 @@ jobs:
|
||||
run: make verify-domain
|
||||
- name: BFF → Keycloak + domain + projection
|
||||
run: make verify-bff
|
||||
- name: Distributed traces reach Tempo (one connected trace across services)
|
||||
run: TRACING_TIMEOUT=120 make verify-tracing
|
||||
- name: Self-service e2e (Playwright, login → submit → success)
|
||||
run: make verify-e2e
|
||||
# Log dump must precede teardown (which removes the containers).
|
||||
- name: Dump container logs on failure
|
||||
if: failure()
|
||||
run: docker compose -f infra/docker-compose.yml logs --no-color --tail=100 oz-init openzaak nrc-init nrc-web nrc-celery nrc-beat flowable-db flowable-rest flowable-init keycloak acl bff domain projection-db event-subscriber projection-api self-service openbaar behandel 2>&1 || true
|
||||
run: docker compose -f infra/docker-compose.yml logs --no-color --tail=100 oz-init openzaak nrc-init nrc-web nrc-celery nrc-beat flowable-db flowable-rest flowable-init keycloak acl bff domain projection-db event-subscriber projection-api self-service openbaar behandel tempo prometheus grafana 2>&1 || true
|
||||
- name: Tear down
|
||||
if: always()
|
||||
run: make down
|
||||
|
||||
+7
-1
@@ -253,10 +253,16 @@ Split (issue #11 closed) into two independently-demoable slices per §13 — the
|
||||
|
||||
**Outcome:** Beheer portal lets an admin view ZTC catalogi (read-only first), and manage the ACL's default-fill configuration via a CRUD UI. MFA on the medewerker realm enforced.
|
||||
|
||||
### S-16 · OpenTelemetry traces + Grafana dashboard
|
||||
### S-16 · OpenTelemetry traces + Grafana dashboard *(split — #17 closed)*
|
||||
|
||||
**Outcome:** Traces span portal → BFF → Domain → ACL → OpenZaak and portal → BFF → Domain → Flowable. Grafana dashboards pre-built for golden signals.
|
||||
|
||||
Split into independently deployable sub-slices (CLAUDE.md §13):
|
||||
|
||||
- **S-16a** (#122) · Observability backplane — Grafana Tempo + Prometheus + Grafana in compose, datasources auto-provisioned (ADR-0023). No collector; config baked into built images.
|
||||
- **S-16b** (#123) · Distributed traces across the five .NET services (OTLP → Tempo; traceparent propagates via the typed HttpClients). Depends on S-16a. ✅
|
||||
- **S-16c** (#124) · Prometheus metrics + golden-signal Grafana dashboards. Depends on S-16a.
|
||||
|
||||
### S-17 · Quartz.NET scheduler — herregistratie reminder sweep ✅
|
||||
|
||||
**Outcome:** Daily Quartz.NET cron job finds inscriptions within 90 days of their herregistratie deadline and reminds each (flag on the aggregate + log). No outbound notification and no domain event in v1 — the reminder is the persisted flag, surfaced on the read model (ADR-0022, #120). Quartz fires time-triggered sweeps; the existing pumps stay as queue-drainers.
|
||||
|
||||
@@ -43,7 +43,7 @@ export DOCKER_HOST := unix://$(PODMAN_SOCK)
|
||||
endif
|
||||
endif
|
||||
|
||||
.PHONY: ci lint build unit mutation frontend integration verify verify-up verify-acl verify-nrc verify-projection verify-bff verify-domain verify-notifications smoke up down local verify-local local-down changelog openzaak-up openzaak-smoke openzaak-seed openzaak-down stack-up stack-smoke stack-down keycloak-up keycloak-smoke keycloak-down flowable-up flowable-smoke flowable-down help
|
||||
.PHONY: ci lint build unit mutation frontend integration verify verify-up verify-acl verify-nrc verify-projection verify-bff verify-domain verify-observability verify-tracing verify-notifications smoke up down local verify-local local-down changelog openzaak-up openzaak-smoke openzaak-seed openzaak-down stack-up stack-smoke stack-down keycloak-up keycloak-smoke keycloak-down flowable-up flowable-smoke flowable-down help
|
||||
|
||||
## ci: run the full pipeline — lint, build, unit, mutation, frontend, verify (mirrors Gitea Actions)
|
||||
## `verify` is the live-stack stage (full stack up once → ACL + notification checks).
|
||||
@@ -170,6 +170,16 @@ verify-bff:
|
||||
verify-e2e:
|
||||
bash infra/run-e2e-check.sh
|
||||
|
||||
## verify-observability: assert the observability backplane (Grafana + provisioned Tempo &
|
||||
## Prometheus datasources) is live, against the already-running stack (S-16a).
|
||||
verify-observability:
|
||||
bash infra/run-observability-check.sh
|
||||
|
||||
## verify-tracing: assert one connected distributed trace spans the .NET services in Tempo
|
||||
## (S-16b), against the already-running stack.
|
||||
verify-tracing:
|
||||
bash infra/run-tracing-check.sh
|
||||
|
||||
## verify: local mirror of the CI verify-stack job — full stack up once, all checks,
|
||||
## tear down (always). For fast single-concern local iteration use `integration`
|
||||
## (oz-only) or `verify-notifications` (oz+nrc) instead.
|
||||
|
||||
@@ -0,0 +1,74 @@
|
||||
# ADR-0023: Grafana-native observability stack (Tempo + Prometheus + Grafana)
|
||||
|
||||
- **Status:** Accepted
|
||||
- **Date:** 2026-07-23
|
||||
- **Deciders:** Respellion engineering
|
||||
- **Slice:** S-16a (#122), first of the S-16 (#17) split
|
||||
|
||||
## Context
|
||||
|
||||
The PRD calls for "OpenTelemetry traces, Prometheus metrics; a local Grafana with
|
||||
pre-built dashboards" (§80). S-16 was split (CLAUDE.md §13) into a backplane slice
|
||||
(this one), distributed tracing (#123), and metrics + dashboards (#124). The
|
||||
backplane must stand up first: a local, CI-friendly place for traces and metrics to
|
||||
land, viewable in one UI, reaching green health within the 3-minute compose budget.
|
||||
|
||||
Two shape decisions are non-obvious enough to record.
|
||||
|
||||
## Decision
|
||||
|
||||
**Run a Grafana-native stack — Grafana Tempo (traces) + Prometheus (metrics) +
|
||||
Grafana (UI) — with the services exporting OTLP straight to Tempo (no collector),
|
||||
and ship the config baked into small built images.**
|
||||
|
||||
### Trace backend: Tempo (not Jaeger)
|
||||
|
||||
Tempo keeps everything under one Grafana pane alongside metrics (and later logs),
|
||||
which is exactly the "local Grafana with dashboards" the PRD asks for. Jaeger would
|
||||
add a second UI and a second mental model for no benefit at this scale.
|
||||
|
||||
### No OTLP collector
|
||||
|
||||
Tempo ingests OTLP directly (gRPC 4317 / HTTP 4318) and Prometheus scrapes each
|
||||
service's `/metrics`, so a collector would be a hop that processes nothing. Skipped.
|
||||
If we later need fan-out, tail sampling, or log processing, a collector is an
|
||||
additive change — the services already speak OTLP.
|
||||
|
||||
### Config baked into built images, not config volumes
|
||||
|
||||
The upstream Common Ground modules (OpenZaak, NRC, Keycloak, Flowable) run as
|
||||
**verbatim** images and get their config streamed into external named volumes by
|
||||
`infra/seed-config.sh`, because bind mounts don't reach sibling containers on the
|
||||
CI runner (see `docs/runbooks/gitea-actions-gotchas.md`). The observability tools
|
||||
are **not** peer modules we must run verbatim, so we take the simpler path: a
|
||||
three-line `Dockerfile` per tool that `COPY`s its config in. This reaches sibling
|
||||
containers everywhere (docker, podman, CI) with no seed step, no `CFG_VOLS` entry,
|
||||
and no Makefile sprawl.
|
||||
|
||||
### Verified, not assumed
|
||||
|
||||
`infra/run-observability-check.sh` (the `verify-observability` step, run early in CI
|
||||
`verify-stack`) asks Grafana to reach both datasources — Prometheus via its health
|
||||
method, Tempo via the datasource proxy (Tempo's Grafana plugin implements no health
|
||||
method) — so the check proves the datasources are actually wired, not merely that
|
||||
containers started. The containers are not in `WAIT_SVCS`; the check polls Grafana
|
||||
itself, so no in-image healthcheck tool is required.
|
||||
|
||||
## Consequences
|
||||
|
||||
**Positive**
|
||||
|
||||
- One UI for traces + metrics + (future) logs. Config is versioned in
|
||||
`infra/observability/` and self-contained in the images.
|
||||
- Backplane is independent of app instrumentation — #123 and #124 build on it.
|
||||
|
||||
**Negative / costs**
|
||||
|
||||
- Three more images built each CI run (kept small; not on the health-gate list).
|
||||
- Storage is ephemeral container fs — a demo backplane, not a retention target.
|
||||
Object storage for Tempo / remote-write for Prometheus is a later concern.
|
||||
|
||||
## Coupling rules touched (CLAUDE.md §8)
|
||||
|
||||
None. The stack is passive infrastructure: services *push* OTLP and *expose*
|
||||
`/metrics`; nothing in the stack calls into a service or a peer module.
|
||||
@@ -5,6 +5,58 @@ copy-pasteable walkthrough against a local `make up` stack.
|
||||
|
||||
---
|
||||
|
||||
## S-16b — distributed traces across the .NET services (#123, ADR-0023)
|
||||
|
||||
**Outcome:** the five .NET services (BFF, Domain, ACL, projection-api, event-subscriber) now emit
|
||||
OpenTelemetry traces — ASP.NET Core + `HttpClient` auto-instrumentation, exported over OTLP to Tempo.
|
||||
Because every cross-service call goes through a typed `HttpClient`, the W3C `traceparent` propagates for
|
||||
free, so a request is **one connected trace** across the services (bff → domain → acl → openzaak;
|
||||
bff → projection-api). `/health` is filtered out. No browser-side instrumentation yet, so the trace
|
||||
begins at the BFF; the async Flowable-poll boundary is a separate trace (ADR-0023).
|
||||
|
||||
```bash
|
||||
# 1. Automated (a CI verify-stack step): generate BFF traffic and assert Tempo holds one trace
|
||||
# spanning multiple services.
|
||||
make verify-tracing # → OK — trace <id> spans ['bff', 'projection-api']
|
||||
|
||||
# 2. By hand: drive the stack, then explore traces in Grafana.
|
||||
make up
|
||||
curl -s localhost:8080/openbaar/register >/dev/null # BFF → projection-api
|
||||
open http://localhost:3000 # Grafana → Explore → Tempo → Search → service.name = bff → open a trace
|
||||
```
|
||||
|
||||
**The path:** each host wires `AddOpenTelemetry().WithTracing(AddAspNetCoreInstrumentation +
|
||||
AddHttpClientInstrumentation + AddOtlpExporter)`; `OTEL_SERVICE_NAME` / `OTEL_EXPORTER_OTLP_ENDPOINT`
|
||||
come from compose; spans export to **tempo:4317** and render in Grafana against the provisioned Tempo
|
||||
datasource.
|
||||
|
||||
---
|
||||
|
||||
## S-16a — observability backplane: Tempo + Prometheus + Grafana (#122, ADR-0023)
|
||||
|
||||
**Outcome:** the compose stack now includes a Grafana-native observability backplane — **Tempo** (OTLP
|
||||
trace ingest on 4317/4318), **Prometheus**, and **Grafana** with both datasources auto-provisioned.
|
||||
Nothing is instrumented yet (traces land in S-16b, metrics + dashboards in S-16c); this slice stands the
|
||||
backplane up and proves Grafana can reach both datasources. Config is baked into small built images
|
||||
(`infra/observability/`) — no collector, no config-volume seeding.
|
||||
|
||||
```bash
|
||||
# 1. Bring the stack up, then assert the backplane is live (Grafana healthy + Tempo/Prometheus
|
||||
# datasources reachable through Grafana). This is a CI verify-stack step.
|
||||
make up
|
||||
make verify-observability # → ✓ Grafana healthy ✓ Prometheus reachable ✓ Tempo reachable
|
||||
|
||||
# 2. Or just the backplane, no full stack needed (no external egress):
|
||||
docker compose -f infra/docker-compose.yml up -d --build tempo prometheus grafana
|
||||
open http://localhost:3000 # Grafana (admin/admin) → Connections → Data sources: Prometheus + Tempo
|
||||
open http://localhost:9090 # Prometheus
|
||||
```
|
||||
|
||||
**The path:** services will export OTLP → **Tempo:4317** and expose `/metrics` ← **Prometheus** scrapes;
|
||||
**Grafana** (:3000) reads both via provisioned datasources with fixed uids `tempo` / `prometheus`.
|
||||
|
||||
---
|
||||
|
||||
## S-17 — herregistratie reminder sweep on a Quartz cron (#18, ADR-0022)
|
||||
|
||||
**Outcome:** an inscription (INGESCHREVEN) now carries the moment it was entered in the register, from
|
||||
|
||||
@@ -296,6 +296,10 @@ services:
|
||||
dockerfile: Dockerfile
|
||||
image: register-referentie/acl:dev
|
||||
environment:
|
||||
# OpenTelemetry traces → Tempo (S-16b, ADR-0023).
|
||||
OTEL_EXPORTER_OTLP_ENDPOINT: http://tempo:4317
|
||||
OTEL_EXPORTER_OTLP_PROTOCOL: grpc
|
||||
OTEL_SERVICE_NAME: acl
|
||||
# Overridable so verify-domain can point the ACL at the same OpenZaak host that
|
||||
# owns the seeded zaaktype URL (host-consistent zaak creation, ADR-0009).
|
||||
Acl__OpenZaak__BaseUrl: ${ACL_OPENZAAK_BASEURL:-http://openzaak:8000/}
|
||||
@@ -334,6 +338,10 @@ services:
|
||||
dockerfile: Dockerfile
|
||||
image: register-referentie/domain:dev
|
||||
environment:
|
||||
# OpenTelemetry traces → Tempo (S-16b, ADR-0023).
|
||||
OTEL_EXPORTER_OTLP_ENDPOINT: http://tempo:4317
|
||||
OTEL_EXPORTER_OTLP_PROTOCOL: grpc
|
||||
OTEL_SERVICE_NAME: domain
|
||||
Flowable__BaseUrl: http://flowable-rest:8080/flowable-rest/
|
||||
Flowable__Username: rest-admin
|
||||
Flowable__Password: test
|
||||
@@ -360,6 +368,10 @@ services:
|
||||
dockerfile: Dockerfile
|
||||
image: register-referentie/bff:dev
|
||||
environment:
|
||||
# OpenTelemetry traces → Tempo (S-16b, ADR-0023).
|
||||
OTEL_EXPORTER_OTLP_ENDPOINT: http://tempo:4317
|
||||
OTEL_EXPORTER_OTLP_PROTOCOL: grpc
|
||||
OTEL_SERVICE_NAME: bff
|
||||
# The BFF is the portals' only backend; it validates digid tokens and fans out (ADR-0010).
|
||||
# Keycloak (start-dev) derives the issuer from the request host, so the BFF authority and the
|
||||
# verify token request both use keycloak:8080 to keep the issuer consistent.
|
||||
@@ -412,6 +424,10 @@ services:
|
||||
dockerfile: services/event-subscriber/Dockerfile
|
||||
image: register-referentie/event-subscriber:dev
|
||||
environment:
|
||||
# OpenTelemetry traces → Tempo (S-16b, ADR-0023).
|
||||
OTEL_EXPORTER_OTLP_ENDPOINT: http://tempo:4317
|
||||
OTEL_EXPORTER_OTLP_PROTOCOL: grpc
|
||||
OTEL_SERVICE_NAME: event-subscriber
|
||||
ConnectionStrings__Projection: Host=projection-db;Database=projection;Username=projection;Password=projection
|
||||
# The subscriber enriches the projection with each zaak's reference (identificatie) by asking
|
||||
# the ACL — the only code allowed to read ZGW (§8.1, #78).
|
||||
@@ -441,6 +457,10 @@ services:
|
||||
dockerfile: services/projection-api/Dockerfile
|
||||
image: register-referentie/projection-api:dev
|
||||
environment:
|
||||
# OpenTelemetry traces → Tempo (S-16b, ADR-0023).
|
||||
OTEL_EXPORTER_OTLP_ENDPOINT: http://tempo:4317
|
||||
OTEL_EXPORTER_OTLP_PROTOCOL: grpc
|
||||
OTEL_SERVICE_NAME: projection-api
|
||||
ConnectionStrings__Projection: Host=projection-db;Database=projection;Username=projection;Password=projection
|
||||
ports:
|
||||
- "8120:8080"
|
||||
@@ -524,6 +544,50 @@ services:
|
||||
condition: service_started
|
||||
networks: [cg]
|
||||
|
||||
# ── Observability backplane (S-16a, ADR-0023) ──────────────────────────────
|
||||
# Grafana-native stack: Tempo ingests OTLP traces (the .NET services export
|
||||
# straight to it — no collector hop, S-16b), Prometheus scrapes service
|
||||
# /metrics (S-16c), and Grafana reads both with datasources auto-provisioned.
|
||||
# Config is baked into small built images (COPY) rather than streamed into
|
||||
# external config volumes like the upstream CG modules — these aren't verbatim
|
||||
# peer images, so a built image is the simpler path that still reaches sibling
|
||||
# containers on the CI runner. Not in WAIT_SVCS: run-observability-check.sh
|
||||
# polls Grafana itself, so no in-image healthcheck tool is needed.
|
||||
tempo:
|
||||
build:
|
||||
context: ./observability/tempo
|
||||
image: register-referentie/tempo:dev
|
||||
command: ["-config.file=/etc/tempo.yaml"]
|
||||
# Cap the backplane's footprint so it can't starve the app stack + the Playwright browser on the
|
||||
# memory-tight CI runner (verify-e2e OOM history, commit d5e5fa2). Generous vs idle (~150M).
|
||||
mem_limit: 400m
|
||||
networks: [cg]
|
||||
|
||||
prometheus:
|
||||
build:
|
||||
context: ./observability/prometheus
|
||||
image: register-referentie/prometheus:dev
|
||||
mem_limit: 400m
|
||||
ports:
|
||||
- "9090:9090"
|
||||
networks: [cg]
|
||||
|
||||
grafana:
|
||||
build:
|
||||
context: ./observability/grafana
|
||||
image: register-referentie/grafana:dev
|
||||
mem_limit: 512m
|
||||
environment:
|
||||
GF_SECURITY_ADMIN_USER: admin
|
||||
GF_SECURITY_ADMIN_PASSWORD: admin
|
||||
GF_AUTH_ANONYMOUS_ENABLED: "true"
|
||||
ports:
|
||||
- "3000:3000"
|
||||
depends_on:
|
||||
- tempo
|
||||
- prometheus
|
||||
networks: [cg]
|
||||
|
||||
volumes:
|
||||
oz-db:
|
||||
nrc-db:
|
||||
|
||||
@@ -0,0 +1,4 @@
|
||||
# Grafana with datasources baked in via provisioning (S-16a, ADR-0023).
|
||||
# Dashboards (S-16c, #124) are added under provisioning/dashboards later.
|
||||
FROM grafana/grafana:11.3.0
|
||||
COPY provisioning/ /etc/grafana/provisioning/
|
||||
@@ -0,0 +1,17 @@
|
||||
# Auto-provisioned datasources (S-16a, ADR-0023). Fixed uids so dashboards (S-16c)
|
||||
# and the verify-observability check can reference them by a stable id.
|
||||
apiVersion: 1
|
||||
|
||||
datasources:
|
||||
- name: Prometheus
|
||||
uid: prometheus
|
||||
type: prometheus
|
||||
access: proxy
|
||||
url: http://prometheus:9090
|
||||
isDefault: true
|
||||
|
||||
- name: Tempo
|
||||
uid: tempo
|
||||
type: tempo
|
||||
access: proxy
|
||||
url: http://tempo:3200
|
||||
@@ -0,0 +1,2 @@
|
||||
FROM prom/prometheus:v2.55.1
|
||||
COPY prometheus.yml /etc/prometheus/prometheus.yml
|
||||
@@ -0,0 +1,10 @@
|
||||
# Prometheus scrape config (S-16a, ADR-0023). For the backplane slice it scrapes
|
||||
# only itself; the .NET services' /metrics scrape targets are added in S-16c
|
||||
# (#124) when the services expose metrics.
|
||||
global:
|
||||
scrape_interval: 15s
|
||||
|
||||
scrape_configs:
|
||||
- job_name: prometheus
|
||||
static_configs:
|
||||
- targets: ['localhost:9090']
|
||||
@@ -0,0 +1,4 @@
|
||||
# Tempo with our config baked in — so it reaches sibling containers on the CI
|
||||
# runner without the external-config-volume dance the upstream CG images need.
|
||||
FROM grafana/tempo:2.6.1
|
||||
COPY tempo.yaml /etc/tempo.yaml
|
||||
@@ -0,0 +1,27 @@
|
||||
# Grafana Tempo — single-binary, all-in-one, local storage (S-16a, ADR-0023).
|
||||
# Ingests OTLP directly (services export straight to Tempo; no collector hop).
|
||||
# Storage is ephemeral container fs — this is a local/CI demo backplane, not a
|
||||
# retention target. ponytail: local backend, swap for object storage if traces
|
||||
# must outlive the stack.
|
||||
server:
|
||||
http_listen_port: 3200
|
||||
|
||||
distributor:
|
||||
receivers:
|
||||
otlp:
|
||||
protocols:
|
||||
grpc:
|
||||
endpoint: 0.0.0.0:4317
|
||||
http:
|
||||
endpoint: 0.0.0.0:4318
|
||||
|
||||
ingester:
|
||||
max_block_duration: 5m
|
||||
|
||||
storage:
|
||||
trace:
|
||||
backend: local
|
||||
local:
|
||||
path: /var/tempo/blocks
|
||||
wal:
|
||||
path: /var/tempo/wal
|
||||
Executable
+45
@@ -0,0 +1,45 @@
|
||||
#!/usr/bin/env bash
|
||||
#
|
||||
# S-16a (#122): assert the observability backplane is live against an ALREADY-RUNNING
|
||||
# stack. Runs curl INSIDE the compose network (like the other verify checks) because
|
||||
# the stack's published ports aren't on the CI runner's localhost — the stack is a set
|
||||
# of sibling containers on the host daemon. It asks Grafana to reach its provisioned
|
||||
# datasources — Prometheus via its health method, Tempo via the datasource proxy (Tempo's
|
||||
# Grafana plugin implements no health method) — so it proves the datasources are wired,
|
||||
# not merely that the containers started. Polls, so it tolerates a cold Grafana.
|
||||
#
|
||||
# Does NOT manage the stack lifecycle (the caller owns bring-up + teardown).
|
||||
set -euo pipefail
|
||||
|
||||
TIMEOUT="${OBS_TIMEOUT:-60}"
|
||||
AUTH="${GRAFANA_AUTH:-admin:admin}"
|
||||
|
||||
gf="$(docker ps -q --filter 'name=[-_]grafana[-_]' | head -1)"
|
||||
[ -n "$gf" ] || { echo "ERROR: no running grafana container — bring the stack up first" >&2; exit 1; }
|
||||
net="$(docker inspect -f '{{range $k,$_ := .NetworkSettings.Networks}}{{$k}}{{"\n"}}{{end}}' "$gf" | head -1)"
|
||||
gf_ip="$(docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' "$gf")"
|
||||
base="http://$gf_ip:3000"
|
||||
echo ">> grafana=$gf_ip network=$net"
|
||||
|
||||
# Run curl inside a throwaway container on the stack network (reaches services by IP).
|
||||
net_curl() { docker run --rm --network "$net" curlimages/curl:latest "$@"; }
|
||||
|
||||
# poll <description> <grep -E pattern> <curl args...>
|
||||
poll() {
|
||||
local desc="$1" pat="$2"; shift 2
|
||||
local deadline=$(( $(date +%s) + TIMEOUT ))
|
||||
while :; do
|
||||
if net_curl -fsS "$@" 2>/dev/null | grep -Eq "$pat"; then echo " ✓ $desc"; return 0; fi
|
||||
if [ "$(date +%s)" -ge "$deadline" ]; then echo " ✗ $desc ($*)" >&2; return 1; fi
|
||||
sleep 3
|
||||
done
|
||||
}
|
||||
|
||||
echo "Checking observability backplane at $base ..."
|
||||
poll "Grafana is healthy" \
|
||||
'"database":[[:space:]]*"ok"' "$base/api/health"
|
||||
poll "Prometheus datasource reachable" \
|
||||
'"status":[[:space:]]*"OK"' -u "$AUTH" "$base/api/datasources/uid/prometheus/health"
|
||||
poll "Tempo datasource reachable (via Grafana proxy)" \
|
||||
'"version"' -u "$AUTH" "$base/api/datasources/proxy/uid/tempo/api/status/buildinfo"
|
||||
echo "Observability backplane OK."
|
||||
Executable
+27
@@ -0,0 +1,27 @@
|
||||
#!/usr/bin/env bash
|
||||
#
|
||||
# S-16b (#123): assert one connected distributed trace spans the .NET services in Tempo,
|
||||
# against an ALREADY-RUNNING full stack. Runs the driver in a python:3-slim container on the
|
||||
# stack network (services reached by container IP; the runner can't reach published ports —
|
||||
# gitea-actions-gotchas.md §5/§6). Does NOT manage the stack lifecycle.
|
||||
set -euo pipefail
|
||||
|
||||
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
ip() { docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' "$1"; }
|
||||
|
||||
bff="$(docker ps -q --filter 'name=[-_]bff[-_]' | head -1)"
|
||||
tempo="$(docker ps -q --filter 'name=[-_]tempo[-_]' | head -1)"
|
||||
[ -n "$bff" ] && [ -n "$tempo" ] || { echo "ERROR: bff and/or tempo not running — bring the stack up first" >&2; exit 1; }
|
||||
net="$(docker inspect -f '{{range $k,$_ := .NetworkSettings.Networks}}{{$k}}{{"\n"}}{{end}}' "$bff" | head -1)"
|
||||
bff_ip="$(ip "$bff")"; tempo_ip="$(ip "$tempo")"
|
||||
echo ">> network=$net bff=$bff_ip tempo=$tempo_ip"
|
||||
|
||||
cid="$(docker create --network "$net" \
|
||||
-e "BFF=http://$bff_ip:8080" -e "TEMPO=http://$tempo_ip:3200" \
|
||||
-e "TRACING_TIMEOUT=${TRACING_TIMEOUT:-90}" \
|
||||
python:3-slim python /tracing-check.py)"
|
||||
docker cp "$here/tracing-check.py" "$cid:/tracing-check.py" >/dev/null
|
||||
rc=0; docker start -a "$cid" || rc=$?
|
||||
docker rm -f "$cid" >/dev/null
|
||||
exit $rc
|
||||
Executable
+81
@@ -0,0 +1,81 @@
|
||||
#!/usr/bin/env python3
|
||||
"""S-16b (#123): prove distributed tracing works end to end.
|
||||
|
||||
Generate anonymous BFF traffic (GET /openbaar/register, which the BFF serves by
|
||||
calling projection-api — no auth, no OpenZaak egress), then query Tempo and assert
|
||||
that ONE trace contains spans from both `bff` and `projection-api`. That proves the
|
||||
services export OTLP to Tempo AND that the W3C traceparent propagates across the
|
||||
HttpClient hop, stitching the request into a single connected trace.
|
||||
|
||||
Stdlib only (urllib/json) so it runs in a bare python:3-slim container in-network.
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
import urllib.error
|
||||
import urllib.parse
|
||||
import urllib.request
|
||||
|
||||
BFF = os.environ["BFF"] # http://<bff-ip>:8080
|
||||
TEMPO = os.environ["TEMPO"] # http://<tempo-ip>:3200
|
||||
TIMEOUT = int(os.environ.get("TRACING_TIMEOUT", "90"))
|
||||
WANT = {"bff", "projection-api"} # the two services that must share one trace
|
||||
|
||||
|
||||
def _get(url):
|
||||
with urllib.request.urlopen(url, timeout=10) as r:
|
||||
return r.read()
|
||||
|
||||
|
||||
def generate_traffic():
|
||||
# A non-2xx still produces spans; only total unreachability of the BFF is fatal.
|
||||
for _ in range(3):
|
||||
try:
|
||||
_get(f"{BFF}/openbaar/register")
|
||||
except urllib.error.HTTPError:
|
||||
pass
|
||||
|
||||
|
||||
def search_trace_ids():
|
||||
q = urllib.parse.quote('{ resource.service.name = "bff" }')
|
||||
try:
|
||||
data = json.loads(_get(f"{TEMPO}/api/search?q={q}&limit=50"))
|
||||
except Exception:
|
||||
return []
|
||||
return [t["traceID"] for t in data.get("traces", [])]
|
||||
|
||||
|
||||
def services_in_trace(trace_id):
|
||||
try:
|
||||
data = json.loads(_get(f"{TEMPO}/api/traces/{trace_id}"))
|
||||
except Exception:
|
||||
return set()
|
||||
names = set()
|
||||
for batch in data.get("batches", []):
|
||||
for attr in batch.get("resource", {}).get("attributes", []):
|
||||
if attr.get("key") == "service.name":
|
||||
names.add(attr.get("value", {}).get("stringValue"))
|
||||
return names
|
||||
|
||||
|
||||
def main():
|
||||
deadline = time.time() + TIMEOUT
|
||||
generate_traffic()
|
||||
seen = set()
|
||||
while time.time() < deadline:
|
||||
for tid in search_trace_ids():
|
||||
names = services_in_trace(tid)
|
||||
seen |= names
|
||||
if WANT.issubset(names):
|
||||
print(f"OK — trace {tid} spans {sorted(names)}")
|
||||
return 0
|
||||
time.sleep(3)
|
||||
generate_traffic()
|
||||
print(f"FAIL — no single trace spanned {sorted(WANT)}; services seen: {sorted(seen)}",
|
||||
file=sys.stderr)
|
||||
return 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -5,6 +5,13 @@
|
||||
<ProjectReference Include="..\Acl.Infrastructure\Acl.Infrastructure.csproj" />
|
||||
</ItemGroup>
|
||||
|
||||
<ItemGroup>
|
||||
<PackageReference Include="OpenTelemetry.Exporter.OpenTelemetryProtocol" Version="1.17.0" />
|
||||
<PackageReference Include="OpenTelemetry.Extensions.Hosting" Version="1.17.0" />
|
||||
<PackageReference Include="OpenTelemetry.Instrumentation.AspNetCore" Version="1.17.0" />
|
||||
<PackageReference Include="OpenTelemetry.Instrumentation.Http" Version="1.17.0" />
|
||||
</ItemGroup>
|
||||
|
||||
<PropertyGroup>
|
||||
<TargetFramework>net10.0</TargetFramework>
|
||||
<Nullable>enable</Nullable>
|
||||
|
||||
@@ -1,8 +1,21 @@
|
||||
using Acl.Application;
|
||||
using Acl.Infrastructure;
|
||||
using OpenTelemetry.Resources;
|
||||
using OpenTelemetry.Trace;
|
||||
|
||||
var builder = WebApplication.CreateBuilder(args);
|
||||
|
||||
// OpenTelemetry tracing (S-16b, ADR-0023): auto-instrument incoming ASP.NET Core requests and
|
||||
// outgoing HttpClient calls (the ACL → OpenZaak hop), exported over OTLP to Tempo. Service name +
|
||||
// OTLP endpoint come from OTEL_* env (compose); the exporter no-ops when Tempo is unreachable.
|
||||
builder.Services.AddOpenTelemetry()
|
||||
.ConfigureResource(r => r.AddService(
|
||||
builder.Configuration["OTEL_SERVICE_NAME"] ?? builder.Environment.ApplicationName))
|
||||
.WithTracing(tracing => tracing
|
||||
.AddAspNetCoreInstrumentation(o => o.Filter = ctx => ctx.Request.Path != "/health")
|
||||
.AddHttpClientInstrumentation()
|
||||
.AddOtlpExporter());
|
||||
|
||||
builder.Services.AddSingleton<IClock, SystemClock>();
|
||||
builder.Services.AddSingleton(sp => sp.GetRequiredService<IConfiguration>()
|
||||
.GetSection("Acl:Defaults").Get<AclDefaults>()
|
||||
|
||||
@@ -10,6 +10,10 @@
|
||||
<!-- OIDC/JWT validation of Keycloak-issued tokens (ADR-0010) and OpenAPI generation. -->
|
||||
<PackageReference Include="Microsoft.AspNetCore.Authentication.JwtBearer" Version="10.0.8" />
|
||||
<PackageReference Include="Microsoft.AspNetCore.OpenApi" Version="10.0.8" />
|
||||
<PackageReference Include="OpenTelemetry.Exporter.OpenTelemetryProtocol" Version="1.17.0" />
|
||||
<PackageReference Include="OpenTelemetry.Extensions.Hosting" Version="1.17.0" />
|
||||
<PackageReference Include="OpenTelemetry.Instrumentation.AspNetCore" Version="1.17.0" />
|
||||
<PackageReference Include="OpenTelemetry.Instrumentation.Http" Version="1.17.0" />
|
||||
</ItemGroup>
|
||||
|
||||
</Project>
|
||||
|
||||
@@ -3,9 +3,23 @@ using System.Text.Json;
|
||||
using System.Text.Json.Serialization;
|
||||
using Bff.Api;
|
||||
using Microsoft.AspNetCore.Authentication.JwtBearer;
|
||||
using OpenTelemetry.Resources;
|
||||
using OpenTelemetry.Trace;
|
||||
|
||||
var builder = WebApplication.CreateBuilder(args);
|
||||
|
||||
// OpenTelemetry tracing (S-16b, ADR-0023): auto-instrument incoming ASP.NET Core requests and
|
||||
// outgoing HttpClient calls (BFF → Domain, BFF → projection-api), exported over OTLP to Tempo, so a
|
||||
// portal request is one connected trace across the services. Service name + OTLP endpoint come from
|
||||
// OTEL_* env (compose); the exporter no-ops when Tempo is unreachable. /health is filtered out.
|
||||
builder.Services.AddOpenTelemetry()
|
||||
.ConfigureResource(r => r.AddService(
|
||||
builder.Configuration["OTEL_SERVICE_NAME"] ?? builder.Environment.ApplicationName))
|
||||
.WithTracing(tracing => tracing
|
||||
.AddAspNetCoreInstrumentation(o => o.Filter = ctx => ctx.Request.Path != "/health")
|
||||
.AddHttpClientInstrumentation()
|
||||
.AddOtlpExporter());
|
||||
|
||||
var keycloakAuthority = builder.Configuration["Keycloak:Authority"]
|
||||
?? throw new InvalidOperationException("Missing configuration 'Keycloak:Authority'");
|
||||
// Behandelaars authenticate against a *different* Keycloak realm (medewerker) than citizens (digid),
|
||||
|
||||
@@ -6,6 +6,10 @@
|
||||
</ItemGroup>
|
||||
|
||||
<ItemGroup>
|
||||
<PackageReference Include="OpenTelemetry.Exporter.OpenTelemetryProtocol" Version="1.17.0" />
|
||||
<PackageReference Include="OpenTelemetry.Extensions.Hosting" Version="1.17.0" />
|
||||
<PackageReference Include="OpenTelemetry.Instrumentation.AspNetCore" Version="1.17.0" />
|
||||
<PackageReference Include="OpenTelemetry.Instrumentation.Http" Version="1.17.0" />
|
||||
<PackageReference Include="Quartz.Extensions.Hosting" Version="3.18.2" />
|
||||
</ItemGroup>
|
||||
|
||||
|
||||
@@ -1,10 +1,25 @@
|
||||
using Big.Application;
|
||||
using Big.Domain;
|
||||
using Big.Infrastructure;
|
||||
using OpenTelemetry.Resources;
|
||||
using OpenTelemetry.Trace;
|
||||
using Quartz;
|
||||
|
||||
var builder = WebApplication.CreateBuilder(args);
|
||||
|
||||
// OpenTelemetry tracing (S-16b, ADR-0023): auto-instrument incoming ASP.NET Core requests and
|
||||
// outgoing HttpClient calls, exported over OTLP to Tempo, so a request is one connected trace across
|
||||
// the services. Service name + OTLP endpoint come from OTEL_* env (compose); the exporter no-ops
|
||||
// harmlessly when Tempo is unreachable (e.g. a service run standalone). /health is filtered out so
|
||||
// liveness polls don't flood the traces.
|
||||
builder.Services.AddOpenTelemetry()
|
||||
.ConfigureResource(r => r.AddService(
|
||||
builder.Configuration["OTEL_SERVICE_NAME"] ?? builder.Environment.ApplicationName))
|
||||
.WithTracing(tracing => tracing
|
||||
.AddAspNetCoreInstrumentation(o => o.Filter = ctx => ctx.Request.Path != "/health")
|
||||
.AddHttpClientInstrumentation()
|
||||
.AddOtlpExporter());
|
||||
|
||||
// Options bound from configuration (compose sets Flowable__* and Acl__* env vars).
|
||||
builder.Services.AddSingleton(sp => sp.GetRequiredService<IConfiguration>()
|
||||
.GetSection("Flowable").Get<FlowableOptions>()
|
||||
|
||||
@@ -5,6 +5,13 @@
|
||||
<ProjectReference Include="..\..\projection-api\Projection.ReadModel\Projection.ReadModel.csproj" />
|
||||
</ItemGroup>
|
||||
|
||||
<ItemGroup>
|
||||
<PackageReference Include="OpenTelemetry.Exporter.OpenTelemetryProtocol" Version="1.17.0" />
|
||||
<PackageReference Include="OpenTelemetry.Extensions.Hosting" Version="1.17.0" />
|
||||
<PackageReference Include="OpenTelemetry.Instrumentation.AspNetCore" Version="1.17.0" />
|
||||
<PackageReference Include="OpenTelemetry.Instrumentation.Http" Version="1.17.0" />
|
||||
</ItemGroup>
|
||||
|
||||
<PropertyGroup>
|
||||
<TargetFramework>net10.0</TargetFramework>
|
||||
<Nullable>enable</Nullable>
|
||||
|
||||
@@ -1,9 +1,22 @@
|
||||
using System.Text.Json;
|
||||
using EventSubscriber.Application;
|
||||
using OpenTelemetry.Resources;
|
||||
using OpenTelemetry.Trace;
|
||||
using Projection.ReadModel;
|
||||
|
||||
var builder = WebApplication.CreateBuilder(args);
|
||||
|
||||
// OpenTelemetry tracing (S-16b, ADR-0023): auto-instrument the incoming NRC notification callback and
|
||||
// the outgoing ACL enrichment call, exported over OTLP to Tempo. Service name + OTLP endpoint come
|
||||
// from OTEL_* env (compose); the exporter no-ops when Tempo is unreachable.
|
||||
builder.Services.AddOpenTelemetry()
|
||||
.ConfigureResource(r => r.AddService(
|
||||
builder.Configuration["OTEL_SERVICE_NAME"] ?? builder.Environment.ApplicationName))
|
||||
.WithTracing(tracing => tracing
|
||||
.AddAspNetCoreInstrumentation(o => o.Filter = ctx => ctx.Request.Path != "/health")
|
||||
.AddHttpClientInstrumentation()
|
||||
.AddOtlpExporter());
|
||||
|
||||
var connectionString = builder.Configuration.GetConnectionString("Projection")
|
||||
?? throw new InvalidOperationException("Missing connection string 'ConnectionStrings:Projection'");
|
||||
// The exact Authorization header value Open Notificaties sends on each abonnement callback.
|
||||
|
||||
@@ -1,8 +1,21 @@
|
||||
using Microsoft.EntityFrameworkCore;
|
||||
using OpenTelemetry.Resources;
|
||||
using OpenTelemetry.Trace;
|
||||
using Projection.ReadModel;
|
||||
|
||||
var builder = WebApplication.CreateBuilder(args);
|
||||
|
||||
// OpenTelemetry tracing (S-16b, ADR-0023): auto-instrument incoming ASP.NET Core requests, exported
|
||||
// over OTLP to Tempo, so a BFF → projection-api read is one connected trace. Service name + OTLP
|
||||
// endpoint come from OTEL_* env (compose); the exporter no-ops when Tempo is unreachable.
|
||||
builder.Services.AddOpenTelemetry()
|
||||
.ConfigureResource(r => r.AddService(
|
||||
builder.Configuration["OTEL_SERVICE_NAME"] ?? builder.Environment.ApplicationName))
|
||||
.WithTracing(tracing => tracing
|
||||
.AddAspNetCoreInstrumentation(o => o.Filter = ctx => ctx.Request.Path != "/health")
|
||||
.AddHttpClientInstrumentation()
|
||||
.AddOtlpExporter());
|
||||
|
||||
var connectionString = builder.Configuration.GetConnectionString("Projection")
|
||||
?? throw new InvalidOperationException("Missing connection string 'ConnectionStrings:Projection'");
|
||||
|
||||
|
||||
@@ -4,6 +4,13 @@
|
||||
<ProjectReference Include="..\Projection.ReadModel\Projection.ReadModel.csproj" />
|
||||
</ItemGroup>
|
||||
|
||||
<ItemGroup>
|
||||
<PackageReference Include="OpenTelemetry.Exporter.OpenTelemetryProtocol" Version="1.17.0" />
|
||||
<PackageReference Include="OpenTelemetry.Extensions.Hosting" Version="1.17.0" />
|
||||
<PackageReference Include="OpenTelemetry.Instrumentation.AspNetCore" Version="1.17.0" />
|
||||
<PackageReference Include="OpenTelemetry.Instrumentation.Http" Version="1.17.0" />
|
||||
</ItemGroup>
|
||||
|
||||
<PropertyGroup>
|
||||
<TargetFramework>net10.0</TargetFramework>
|
||||
<Nullable>enable</Nullable>
|
||||
|
||||
Reference in New Issue
Block a user