## What & why S-16c, the last of the S-16 (#17) split, on top of the backplane (#122) and distributed tracing (#123). The five .NET services now expose OpenTelemetry **metrics** in Prometheus format at `/metrics`; Prometheus scrapes each (one job per service); and Grafana ships a pre-built **Request path — golden signals** dashboard (traffic / errors / latency / saturation), split by service. Closes #124 ### How - Each service adds `.WithMetrics(AddAspNetCoreInstrumentation + AddHttpClientInstrumentation + AddMeter("System.Runtime") + AddPrometheusExporter)` and maps `/metrics`. Same shape as the S-16b tracing wiring already in these `Program.cs` files. - `infra/observability/prometheus/prometheus.yml`: one scrape job per service (`acl`, `domain`, `bff`, `event-subscriber`, `projection-api`), reached by compose service name. - `infra/observability/grafana/provisioning/dashboards/`: dashboard provider + `golden-signals.json` (baked into the Grafana image by the existing `COPY provisioning/`). - `verify-metrics` (new CI verify-stack step + Makefile target): generates BFF traffic and asserts Prometheus scraped the golden-signal metric from every service. Mirrors `verify-tracing`. ### Dependency (CLAUDE.md §13/§14) Adds `OpenTelemetry.Exporter.Prometheus.AspNetCore` `1.17.0-beta.1` (matched to the `1.17.0` core already in use). It gives the OTel-native `/metrics` pull endpoint; replacing it would mean hand-rolling Prometheus exposition over a `MeterListener`; the risk is that it is a **prerelease** package (the whole OTel .NET Prometheus line is `-beta`) — pinned, wired only in `Program.cs`, and gated by `verify-metrics`. Recorded in **ADR-0024**. ## Definition of Done - [x] Linked Gitea issue (#124). - [x] Failing test committed before the implementation (`test(bff): /metrics exposes http-server request duration`). - [x] Implementation makes the test pass. - [ ] CI green — pending Gitea Actions run. - [x] `docker compose up` reaches green health within 3 min (backplane images unchanged in shape; not on the health gate, ADR-0023). - [x] Docs updated — demo-script S-16c entry. - [x] ADR added — ADR-0024. - [x] Demo note in `docs/demo-script.md`. ## Notes for reviewers - `/health` polls are counted as traffic (metrics aren't path-filtered, unlike traces). Fine for a demo dashboard and honest — real load stacks on top. - `projection-api` has no Stryker config (unchanged); the four mutated services carry the metrics wiring in `Program.cs`, same as the merged S-16b tracing code. - Metric names verified against a live service: `http_server_request_duration_seconds{,_bucket,_count}`, label `http_response_status_code`, `dotnet_process_cpu_time_seconds_total`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Reviewed-on: #129
This commit was merged in pull request #129.
This commit is contained in:
Executable
+75
@@ -0,0 +1,75 @@
|
||||
#!/usr/bin/env python3
|
||||
"""S-16c (#124): prove the golden-signal metrics pipeline works end to end.
|
||||
|
||||
Generate anonymous BFF traffic (GET /openbaar/register — no auth, no OpenZaak egress),
|
||||
then query Prometheus and assert (1) every .NET service's scrape target is UP, and (2)
|
||||
the http.server.request.duration histogram is actually being scraped — i.e. the services
|
||||
expose /metrics AND Prometheus collects it, which is exactly what the golden-signal
|
||||
dashboard reads.
|
||||
|
||||
Stdlib only (urllib/json) so it runs in a bare python:3-slim container in-network.
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
import urllib.error
|
||||
import urllib.parse
|
||||
import urllib.request
|
||||
|
||||
BFF = os.environ["BFF"] # http://<bff-ip>:8080
|
||||
PROM = os.environ["PROMETHEUS"] # http://<prometheus-ip>:9090
|
||||
TIMEOUT = int(os.environ.get("METRICS_TIMEOUT", "90"))
|
||||
SERVICES = {"acl", "domain", "bff", "event-subscriber", "projection-api"}
|
||||
|
||||
|
||||
def _get(url):
|
||||
with urllib.request.urlopen(url, timeout=10) as r:
|
||||
return r.read()
|
||||
|
||||
|
||||
def generate_traffic():
|
||||
for _ in range(3):
|
||||
try:
|
||||
_get(f"{BFF}/openbaar/register")
|
||||
except urllib.error.HTTPError:
|
||||
pass # a non-2xx still records an http.server metric
|
||||
|
||||
|
||||
def query(promql):
|
||||
q = urllib.parse.quote(promql)
|
||||
try:
|
||||
data = json.loads(_get(f"{PROM}/api/v1/query?query={q}"))
|
||||
except Exception:
|
||||
return []
|
||||
return data.get("data", {}).get("result", [])
|
||||
|
||||
|
||||
def jobs_up():
|
||||
return {r["metric"].get("job") for r in query("up == 1")}
|
||||
|
||||
|
||||
def jobs_with_request_metric():
|
||||
return {r["metric"].get("job")
|
||||
for r in query("http_server_request_duration_seconds_count")}
|
||||
|
||||
|
||||
def main():
|
||||
deadline = time.time() + TIMEOUT
|
||||
while time.time() < deadline:
|
||||
generate_traffic()
|
||||
up = jobs_up()
|
||||
scraped = jobs_with_request_metric()
|
||||
if SERVICES.issubset(up) and SERVICES.issubset(scraped):
|
||||
print(f"OK — targets up: {sorted(up & SERVICES)}; "
|
||||
f"request metric scraped from: {sorted(scraped & SERVICES)}")
|
||||
return 0
|
||||
time.sleep(3)
|
||||
print(f"FAIL — up: {sorted(jobs_up() & SERVICES)}; "
|
||||
f"request metric from: {sorted(jobs_with_request_metric() & SERVICES)}; "
|
||||
f"expected all of {sorted(SERVICES)}", file=sys.stderr)
|
||||
return 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -1,4 +1,4 @@
|
||||
# Grafana with datasources baked in via provisioning (S-16a, ADR-0023).
|
||||
# Dashboards (S-16c, #124) are added under provisioning/dashboards later.
|
||||
# Grafana with datasources + the golden-signals dashboard baked in via provisioning
|
||||
# (S-16a/S-16c, ADR-0023). Everything under provisioning/ is copied in below.
|
||||
FROM grafana/grafana:11.3.0
|
||||
COPY provisioning/ /etc/grafana/provisioning/
|
||||
|
||||
@@ -0,0 +1,13 @@
|
||||
# Dashboard provider (S-16c, ADR-0023): Grafana loads every *.json in this folder as a
|
||||
# read-only, code-owned dashboard. The golden-signals board is versioned here, not
|
||||
# clicked together in the UI.
|
||||
apiVersion: 1
|
||||
|
||||
providers:
|
||||
- name: register-referentie
|
||||
type: file
|
||||
disableDeletion: true
|
||||
allowUiUpdates: false
|
||||
options:
|
||||
path: /etc/grafana/provisioning/dashboards
|
||||
foldersFromFilesStructure: false
|
||||
@@ -0,0 +1,87 @@
|
||||
{
|
||||
"uid": "golden-signals",
|
||||
"title": "Request path — golden signals",
|
||||
"tags": ["s-16c", "golden-signals"],
|
||||
"timezone": "browser",
|
||||
"schemaVersion": 39,
|
||||
"version": 1,
|
||||
"editable": true,
|
||||
"refresh": "10s",
|
||||
"time": { "from": "now-15m", "to": "now" },
|
||||
"templating": {
|
||||
"list": [
|
||||
{
|
||||
"name": "job",
|
||||
"type": "query",
|
||||
"datasource": { "type": "prometheus", "uid": "prometheus" },
|
||||
"query": "label_values(http_server_request_duration_seconds_count, job)",
|
||||
"includeAll": true,
|
||||
"multi": true,
|
||||
"current": { "text": "All", "value": "$__all" },
|
||||
"refresh": 2
|
||||
}
|
||||
]
|
||||
},
|
||||
"panels": [
|
||||
{
|
||||
"id": 1,
|
||||
"title": "Traffic — requests/sec",
|
||||
"type": "timeseries",
|
||||
"datasource": { "type": "prometheus", "uid": "prometheus" },
|
||||
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 0 },
|
||||
"fieldConfig": { "defaults": { "unit": "reqps", "custom": { "drawStyle": "line", "fillOpacity": 10 } }, "overrides": [] },
|
||||
"targets": [
|
||||
{
|
||||
"refId": "A",
|
||||
"expr": "sum by (job) (rate(http_server_request_duration_seconds_count{job=~\"$job\"}[$__rate_interval]))",
|
||||
"legendFormat": "{{job}}"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"title": "Errors — 5xx responses/sec",
|
||||
"type": "timeseries",
|
||||
"datasource": { "type": "prometheus", "uid": "prometheus" },
|
||||
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 0 },
|
||||
"fieldConfig": { "defaults": { "unit": "reqps", "custom": { "drawStyle": "line", "fillOpacity": 10 }, "color": { "mode": "fixed", "fixedColor": "red" } }, "overrides": [] },
|
||||
"targets": [
|
||||
{
|
||||
"refId": "A",
|
||||
"expr": "sum by (job) (rate(http_server_request_duration_seconds_count{job=~\"$job\",http_response_status_code=~\"5..\"}[$__rate_interval]))",
|
||||
"legendFormat": "{{job}}"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"title": "Latency — p95 request duration",
|
||||
"type": "timeseries",
|
||||
"datasource": { "type": "prometheus", "uid": "prometheus" },
|
||||
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 8 },
|
||||
"fieldConfig": { "defaults": { "unit": "s", "custom": { "drawStyle": "line", "fillOpacity": 10 } }, "overrides": [] },
|
||||
"targets": [
|
||||
{
|
||||
"refId": "A",
|
||||
"expr": "histogram_quantile(0.95, sum by (job, le) (rate(http_server_request_duration_seconds_bucket{job=~\"$job\"}[$__rate_interval])))",
|
||||
"legendFormat": "{{job}} p95"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": 4,
|
||||
"title": "Saturation — CPU cores in use",
|
||||
"type": "timeseries",
|
||||
"datasource": { "type": "prometheus", "uid": "prometheus" },
|
||||
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 8 },
|
||||
"fieldConfig": { "defaults": { "unit": "none", "custom": { "drawStyle": "line", "fillOpacity": 10 } }, "overrides": [] },
|
||||
"targets": [
|
||||
{
|
||||
"refId": "A",
|
||||
"expr": "sum by (job) (rate(dotnet_process_cpu_time_seconds_total{job=~\"$job\"}[$__rate_interval]))",
|
||||
"legendFormat": "{{job}}"
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -1,6 +1,7 @@
|
||||
# Prometheus scrape config (S-16a, ADR-0023). For the backplane slice it scrapes
|
||||
# only itself; the .NET services' /metrics scrape targets are added in S-16c
|
||||
# (#124) when the services expose metrics.
|
||||
# Prometheus scrape config (S-16c, ADR-0023). Each .NET service exposes OTel metrics
|
||||
# at /metrics (Prometheus text format); one scrape job per service, so the service is
|
||||
# identified by the `job` label in the golden-signal dashboard. Targets are reached by
|
||||
# compose service name on the shared `cg` network (internal port 8080).
|
||||
global:
|
||||
scrape_interval: 15s
|
||||
|
||||
@@ -8,3 +9,19 @@ scrape_configs:
|
||||
- job_name: prometheus
|
||||
static_configs:
|
||||
- targets: ['localhost:9090']
|
||||
|
||||
- job_name: acl
|
||||
static_configs:
|
||||
- targets: ['acl:8080']
|
||||
- job_name: domain
|
||||
static_configs:
|
||||
- targets: ['domain:8080']
|
||||
- job_name: bff
|
||||
static_configs:
|
||||
- targets: ['bff:8080']
|
||||
- job_name: event-subscriber
|
||||
static_configs:
|
||||
- targets: ['event-subscriber:8080']
|
||||
- job_name: projection-api
|
||||
static_configs:
|
||||
- targets: ['projection-api:8080']
|
||||
|
||||
Executable
+28
@@ -0,0 +1,28 @@
|
||||
#!/usr/bin/env bash
|
||||
#
|
||||
# S-16c (#124): assert the golden-signal metrics pipeline works — the .NET services expose
|
||||
# /metrics and Prometheus scrapes them — against an ALREADY-RUNNING full stack. Runs the
|
||||
# driver in a python:3-slim container on the stack network (services reached by container IP;
|
||||
# the runner can't reach published ports — gitea-actions-gotchas.md §5/§6). Does NOT manage
|
||||
# the stack lifecycle.
|
||||
set -euo pipefail
|
||||
|
||||
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
ip() { docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' "$1"; }
|
||||
|
||||
bff="$(docker ps -q --filter 'name=[-_]bff[-_]' | head -1)"
|
||||
prom="$(docker ps -q --filter 'name=[-_]prometheus[-_]' | head -1)"
|
||||
[ -n "$bff" ] && [ -n "$prom" ] || { echo "ERROR: bff and/or prometheus not running — bring the stack up first" >&2; exit 1; }
|
||||
net="$(docker inspect -f '{{range $k,$_ := .NetworkSettings.Networks}}{{$k}}{{"\n"}}{{end}}' "$bff" | head -1)"
|
||||
bff_ip="$(ip "$bff")"; prom_ip="$(ip "$prom")"
|
||||
echo ">> network=$net bff=$bff_ip prometheus=$prom_ip"
|
||||
|
||||
cid="$(docker create --network "$net" \
|
||||
-e "BFF=http://$bff_ip:8080" -e "PROMETHEUS=http://$prom_ip:9090" \
|
||||
-e "METRICS_TIMEOUT=${METRICS_TIMEOUT:-90}" \
|
||||
python:3-slim python /metrics-check.py)"
|
||||
docker cp "$here/metrics-check.py" "$cid:/metrics-check.py" >/dev/null
|
||||
rc=0; docker start -a "$cid" || rc=$?
|
||||
docker rm -f "$cid" >/dev/null
|
||||
exit $rc
|
||||
Reference in New Issue
Block a user