## What & why S-16c, the last of the S-16 (#17) split, on top of the backplane (#122) and distributed tracing (#123). The five .NET services now expose OpenTelemetry **metrics** in Prometheus format at `/metrics`; Prometheus scrapes each (one job per service); and Grafana ships a pre-built **Request path — golden signals** dashboard (traffic / errors / latency / saturation), split by service. Closes #124 ### How - Each service adds `.WithMetrics(AddAspNetCoreInstrumentation + AddHttpClientInstrumentation + AddMeter("System.Runtime") + AddPrometheusExporter)` and maps `/metrics`. Same shape as the S-16b tracing wiring already in these `Program.cs` files. - `infra/observability/prometheus/prometheus.yml`: one scrape job per service (`acl`, `domain`, `bff`, `event-subscriber`, `projection-api`), reached by compose service name. - `infra/observability/grafana/provisioning/dashboards/`: dashboard provider + `golden-signals.json` (baked into the Grafana image by the existing `COPY provisioning/`). - `verify-metrics` (new CI verify-stack step + Makefile target): generates BFF traffic and asserts Prometheus scraped the golden-signal metric from every service. Mirrors `verify-tracing`. ### Dependency (CLAUDE.md §13/§14) Adds `OpenTelemetry.Exporter.Prometheus.AspNetCore` `1.17.0-beta.1` (matched to the `1.17.0` core already in use). It gives the OTel-native `/metrics` pull endpoint; replacing it would mean hand-rolling Prometheus exposition over a `MeterListener`; the risk is that it is a **prerelease** package (the whole OTel .NET Prometheus line is `-beta`) — pinned, wired only in `Program.cs`, and gated by `verify-metrics`. Recorded in **ADR-0024**. ## Definition of Done - [x] Linked Gitea issue (#124). - [x] Failing test committed before the implementation (`test(bff): /metrics exposes http-server request duration`). - [x] Implementation makes the test pass. - [ ] CI green — pending Gitea Actions run. - [x] `docker compose up` reaches green health within 3 min (backplane images unchanged in shape; not on the health gate, ADR-0023). - [x] Docs updated — demo-script S-16c entry. - [x] ADR added — ADR-0024. - [x] Demo note in `docs/demo-script.md`. ## Notes for reviewers - `/health` polls are counted as traffic (metrics aren't path-filtered, unlike traces). Fine for a demo dashboard and honest — real load stacks on top. - `projection-api` has no Stryker config (unchanged); the four mutated services carry the metrics wiring in `Program.cs`, same as the merged S-16b tracing code. - Metric names verified against a live service: `http_server_request_duration_seconds{,_bucket,_count}`, label `http_response_status_code`, `dotnet_process_cpu_time_seconds_total`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Reviewed-on: #129
This commit was merged in pull request #129.
This commit is contained in:
@@ -5,6 +5,34 @@ copy-pasteable walkthrough against a local `make up` stack.
|
||||
|
||||
---
|
||||
|
||||
## S-16c — Prometheus metrics + golden-signal Grafana dashboard (#124, ADR-0023)
|
||||
|
||||
**Outcome:** the five .NET services now expose OpenTelemetry metrics in Prometheus format at `/metrics`
|
||||
— ASP.NET Core + `HttpClient` instrumentation plus the built-in `System.Runtime` meter. Prometheus
|
||||
scrapes each service (one job per service), and a **pre-built Grafana dashboard** — *Request path —
|
||||
golden signals* — plots the four golden signals: **traffic** (req/s), **errors** (5xx/s), **latency**
|
||||
(p95 request duration), and **saturation** (CPU cores in use), split by service. It populates under load.
|
||||
|
||||
```bash
|
||||
# 1. Automated (a CI verify-stack step): generate BFF traffic and assert Prometheus scraped the
|
||||
# golden-signal metric from every service.
|
||||
make verify-metrics # → OK — targets up: [...]; request metric scraped from: [...]
|
||||
|
||||
# 2. By hand: drive the stack, generate some load, then open the dashboard.
|
||||
make up
|
||||
for i in $(seq 1 50); do curl -s localhost:8080/openbaar/register >/dev/null; done # BFF → projection-api
|
||||
open http://localhost:3000 # Grafana → Dashboards → "Request path — golden signals"
|
||||
open http://localhost:9090/targets # Prometheus → every service target UP
|
||||
```
|
||||
|
||||
**The path:** each host adds `.WithMetrics(AddAspNetCoreInstrumentation + AddHttpClientInstrumentation +
|
||||
AddMeter("System.Runtime") + AddPrometheusExporter)` and maps `/metrics`; Prometheus scrapes
|
||||
`<service>:8080/metrics` (config in `infra/observability/prometheus/prometheus.yml`); Grafana ships the
|
||||
dashboard via provisioning against the fixed `prometheus` datasource uid. No metrics are pushed over
|
||||
OTLP — Prometheus pulls, so there is no collector hop (ADR-0023).
|
||||
|
||||
---
|
||||
|
||||
## S-16b — distributed traces across the .NET services (#123, ADR-0023)
|
||||
|
||||
**Outcome:** the five .NET services (BFF, Domain, ACL, projection-api, event-subscriber) now emit
|
||||
|
||||
Reference in New Issue
Block a user