Compare commits

..
Author SHA1 Message Date
not 60d556b46e docs(architecture): record why single-binary Tempo has its ingester health check off (refs #156)
CI / lint (pull_request) Successful in 1m41s
CI / build (pull_request) Successful in 1m8s
CI / unit (pull_request) Successful in 1m23s
CI / frontend (pull_request) Successful in 3m21s
CI / mutation (pull_request) Successful in 6m32s
CI / verify-stack (pull_request) Successful in 9m31s
2026-09-01 09:45:28 +02:00
not e54dbe9d5d fix(observability): stop Tempo evicting its only ingester under load (refs #156)
`verify-tracing` flaked on run 722: `FAIL — no single trace spanned ['bff',
'projection-api']`, green on a plain re-run of the same commit. Not a broken trace
chain — Tempo could not ingest:

    removing distributor_pool failing healthcheck addr=127.0.0.1:9095
      reason="rpc error: code = DeadlineExceeded"
    pusher failed to consume trace data  err="context canceled"   (x18)

Root cause is the mechanism of the data loss, not whatever caused the stall. Tempo
runs single-binary, so distributor and ingester are the *same process* and the
distributor's ingester pool holds exactly one, in-process, member. dskit still
health-checks it over loopback gRPC with a 1s deadline; on the shared runner a
transient stall blows that, the only ingester is dropped from the pool, and every
push fails until the next 15s check interval — spans silently lost. With one
in-process ingester the check can never route around a failure, so it can only ever
discard data.

Fix: `ingester_client.pool_config.healthcheckenabled: false`. This addresses both
candidate triggers (GC pressure near `mem_limit`, CPU contention) at the point where
they turn into lost data, so `mem_limit: 400m` stays untouched — raising it would
risk reintroducing the verify-e2e OOM of #144 on a memory-tight runner.

Also print `tempo_distributor_ingester_clients` on the check's failure path: a
recurrence then names Tempo-dropped-spans instead of costing another container-log
dive, since from the check's side that is indistinguishable from missing
instrumentation.

Verified against the built image: config parses (`-config.verify`), the effective
`/status/config` reports `healthcheckenabled: false`, and the diagnostic reads the
metric off a live Tempo.
2026-09-01 09:45:07 +02:00
not 94742a261f feat: read projection sourced from the register in Objecten (closes #153) (#155)
CI / build (push) Successful in 1m7s
CI / lint (push) Successful in 1m26s
CI / unit (push) Successful in 1m37s
CI / frontend (push) Successful in 3m36s
CI / mutation (push) Successful in 6m42s
CI / verify-stack (push) Failing after 11m26s
## What & why

S-19b-2, closing out ADR-0028's stated direction: **the read projection is now derived from the
`RegisterRecord` in Objecten, not from ZGW zaak events.**

Until now the subscriber listened on `zaken` and *inferred* register state from case events — a
`zaak/create` meant INGEDIEND, and any `status/create` was assumed to be the approval (it may not
read OpenZaak, so it could not tell statustypen apart). The reference wasn't in the notification
at all, so every projection made a second hop to the ACL. The register — a fact about a person —
was being reconstructed by guessing at the lifecycle of the case that produced it.

- The subscriber's abonnement moves to the `objecten` kanaal (S-19b-1 made it publish).
- An Objecten notification carries **no record data**, only the object URL, so the record is read
  back through the ACL (`POST /register-records/read`) — §8.1 applies to Objecten exactly as
  ADR-0028 established.
- The record carries `id`, `status` and `reference`, so the row *is* the record: `IsZaakCreated`,
  `IsZaakStatusSet`, `ZaakUrl`, `ZaakId` and `ToEntry`'s `Resource == "status"` inference are all
  gone, and so is the ACL enrichment hop.
- **The ACL now writes an INGEDIEND record on submit.** Without it, re-sourcing would silently
  drop every submitted registration from the public register, since only approval wrote a record.
- `processed_notifications` holds the projected row (`register_id`, `status`, `reference`) instead
  of the ZGW event, so a rebuild is a replay with no mapping rules and no upstream reads at all.

**ADR-0030** records it. ADR-0028's open caveat — record written but not yet read, "the two must
agree" — is closed: there is one source now.

Closes #153

## Definition of Done

- [x] Linked Gitea issue (above).
- [x] Failing tests committed before the implementation — two red/green pairs, ACL side
      (06c0444 → 566ef7d) and subscriber side (142ed45 → 8af09b2).
- [x] Refactor commit follows (b496ac9).
- [x] Conventional Commits referencing the issue (`refs #153`).
- [x] CI green — all six jobs on b30fa66, `verify-stack` end to end including the e2e.
- [x] `docker compose up` from a fresh clone reaches green health checks within 3 minutes
      (`verify-stack`'s bring-up step — see the wait-healthy fix below).
- [x] Docs updated — ADR-0030 added, ADR-0028's consequence + caveat annotated, BACKLOG.md,
      e2e header comment.
- [x] ADR added in `docs/architecture/`.
- [x] Demo note in `docs/demo-script.md` — n/a: no user-visible change. The openbaar register
      shows the same two statuses for the same registrations; only where they come from changed.

## Notes for reviewers

**The decision I'd most like a second opinion on** is the one the issue didn't settle: what
happens to INGEDIEND. Objecten held only INGESCHREVEN records, so re-sourcing forced a choice
between (a) the ACL also writing on submit, (b) a public register that lists only actual
registrations, or (c) a hybrid keeping both kanalen. I took (a): visible behaviour is unchanged
and the register holds the whole lifecycle. (b) is arguably the better *semantics* for a public
register but narrows what the portal shows and reads against PRD §68 ("~50 register entries with
diverse statuses"); (c) leaves the projection half-derived from ZGW, which is the coupling
ADR-0028 set out to remove. All three are laid out in ADR-0030.

**The dedup key is the projected row**, `objecten:object:{url}:{status}:{reference}` — not the
object URL (the ACL upserts *one object per registration*, so submit and approval notify about
the same URL and the approval would be swallowed as a duplicate) and not URL+actie (a retried
approval is a second `update`). Redeliveries collapse, genuine state changes don't. §8.6.

**The migration drops columns rather than renaming them.** EF scaffolded renames — `resource` →
`register_id`, `zaak_id` → `status` — which would have carried ZGW values into columns meaning
something else, and a rebuild would then have projected that garbage. It also empties both
tables: a pre-slice row describes a zaak event the new projector can't reproject, and those
registrations have no RegisterRecord in Objecten either, so they're not re-derivable from the new
source. Stated as a ceiling in the ADR — fine while stacks are ephemeral, backfill from Objecten
if a long-lived environment ever needs it.

**`run-projection-check.sh` now opens its zaak through the ACL** instead of straight against
OpenZaak, because the ACL is what writes the record. A zaak created behind the ACL's back
produces no projection row — that's the re-source working, not a gap.

## Three fixes CI found, none of them in the projection logic

1. **`wait-healthy.sh` matched the wrong container** (744f91a). Bring-up timed out with
   `TIMEOUT: 'objecten' not healthy (status=none)` while the `docker ps` it dumps showed
   objecten `Up 9 minutes (healthy)`. `--filter name=` is a substring match, so `objecten` also
   matches `objecten-db`/`objecten-redis`/`objecten-celery`, and `head -1` took whichever docker
   listed first — the celery worker has no healthcheck, hence `status=none`. Latent since those
   services landed and decided purely by listing order; `objecttypen` matches `objecttypen-db`
   the same way. Anchored on the compose replica suffix, which the verify scripts already do.
2. **The ACL had to be repointed at OpenZaak's IP** (7e0897a). Opening the zaak through the ACL
   put this check in the same bind run-domain-check.sh already handles:
   `400 {"name":"zaaktype","code":"bad-url","reason":"Voer een geldige URL in."}`. OpenZaak
   reflects the request Host into the zaaktype URL and then rejects it on zaak-create when
   single-label — the mechanism compose already documents on `ACL_OPENZAAK_BASEURL`.
3. **Approval arrives as `partial_update`, not `update`** (0dd26a7 → b30fa66) — the one real bug
   in the slice. The ACL upserts with PATCH; DRF routes it through the notifying `update()` but
   names the action `partial_update`, so the projector dropped every approval. Only the e2e could
   catch it: `verify-projection` drives a submit, and per ADR-0028 the e2e is the only check that
   drives a *real* approval.

`verify-tracing` also failed once (run 722) on a path this PR doesn't touch, and passed on a
plain re-run of the same commit. Tempo logged `pusher failed to consume trace data` /
`distributor_pool failing healthcheck` — it dropped spans under runner load rather than the trace
chain being broken. Filed as **#156** rather than absorbed here.

**Correction to the #152 PR notes:** I wrote there that celery concurrency was "the next knob" if
verify-stack got tight. It isn't — `CELERY_WORKER_CONCURRENCY` already defaults to 1 in the Maykin
image, so `objecten-celery` is already a single-process worker. Noted in #156.

**Possible follow-up, deliberately not done here:** an `openzaak.local` network alias mirroring
`objecten.local` would remove the ACL-repoint dance from both run-domain-check.sh and
run-projection-check.sh. It changes the host in every zaak URL the system produces, which is too
broad a ripple to land inside an unrelated slice — worth its own issue.

**Known costs, all in the ADR:** submission is now two writes across two modules and eventually
consistent (same posture ADR-0028 accepted for approval); projecting now depends on the ACL being
reachable on the main path, not just for enrichment (NRC retries, so it converges); and OpenZaak
still publishes to `zaken` with nothing in the product listening — kept because `verify-nrc`
asserts that path.Reviewed-on: #155
2026-09-01 07:26:33 +00:00
8 changed files with 89 additions and 13 deletions
@@ -67,6 +67,14 @@ itself, so no in-image healthcheck tool is required.
- Three more images built each CI run (kept small; not on the health-gate list).
- Storage is ephemeral container fs — a demo backplane, not a retention target.
Object storage for Tempo / remote-write for Prometheus is a later concern.
- Tempo runs **single-binary**, so its distributor and ingester are one process and
some of its distributed-mode machinery is not just redundant but harmful. Its
ingester-pool health check is disabled (`ingester_client.pool_config`) because with
a single in-process ingester the check can never route around a failure — a 1s
loopback-gRPC deadline missed under CI load only evicted the one ingester and made
Tempo drop spans, which is how `verify-tracing` flaked (#156). Expect the same
shape from other distributed-mode knobs if we tune them; the fix is to switch to
real multi-ingester Tempo, not to re-enable them here.
## Coupling rules touched (CLAUDE.md §8)
@@ -34,6 +34,12 @@ notification points at. The projection is a cache of the register; ZGW is no lon
as a kenmerk — so the record is read back through the ACL (`POST /register-records/read`).
§8.1 applies to Objecten exactly as ADR-0028 established: the ACL is the only code that talks
to it.
- The accepted acties are `create`, `update` and `partial_update`. The last one is not
defensive breadth: the ACL upserts with PATCH, and DRF routes a PATCH through the notifying
`update()` while naming the action `partial_update` — which is what Objecten publishes. So
every approval arrives as `partial_update`, and accepting only `create`/`update` drops the
one state change this slice exists to project. `destroy` is deliberately not accepted:
removing a registration from the public register is its own decision.
- The record already carries `id`, `status` and `reference`, so the row is the record. The
zaak-shaped surface goes: `IsZaakCreated`, `IsZaakStatusSet`, `ZaakUrl`, `ZaakId`, and
`ToEntry`'s `Resource == "status"` inference are replaced by `IsRegisterRecordWritten` +
+12
View File
@@ -25,3 +25,15 @@ storage:
path: /var/tempo/blocks
wal:
path: /var/tempo/wal
# #156: don't let the distributor evict its own ingester. Tempo runs single-binary here, so the
# distributor and the ingester are the same process and the "pool" holds exactly one, in-process,
# member. dskit still health-checks it over loopback gRPC with a 1s deadline (checkinterval 15s);
# on the shared CI runner a transient stall blows that deadline, the only ingester is dropped from
# the pool ("removing distributor_pool failing healthcheck"), and every push then fails ("pusher
# failed to consume trace data", err="context canceled") until the next check — silently losing
# spans, which is how verify-tracing flaked. With one in-process ingester the check can never route
# around a failure, so it can only ever discard data. Turn it off.
ingester_client:
pool_config:
healthcheckenabled: false
+19 -2
View File
@@ -12,11 +12,15 @@
# which is the point of the re-source.
#
# All in-network, reaching services by container IP — single-label hosts aren't URL-valid and
# the runner can't reach published ports (gitea-actions-gotchas.md §5/§6). Does NOT manage the stack
# lifecycle (the caller owns bring-up + teardown). Plain docker primitives only. See ADR-0007/0008/0030.
# the runner can't reach published ports (gitea-actions-gotchas.md §5/§6). Does not own the stack
# lifecycle (the caller brings it up and tears it down), but does recreate the `acl` service to
# repoint it — see below, and run-domain-check.sh, which does the same. Plain docker primitives only.
# See ADR-0007/0008/0030.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
root="$(cd "$here/.." && pwd)"
compose="$root/infra/docker-compose.yml"
WEBHOOK_AUTH="${NOTIFICATION_WEBHOOK_TOKEN:-Bearer big-reference-notifications}"
cleanup() { docker rm -f rr-pverify rr-pquery >/dev/null 2>&1 || true; }
@@ -54,6 +58,19 @@ docker cp "$here/local/register-abonnement.py" "$drv:/subscribe.py" >/dev/null
docker start -a "$drv"
docker rm -f rr-pverify >/dev/null
# OpenZaak reflects the request Host into the zaaktype `url` it returns, and then rejects that same
# URL on zaak-create when the host is single-label ("Voer een geldige URL in."). The stack's ACL is
# configured with `http://openzaak:8000/`, so it must be repointed at OpenZaak's container IP before
# it can open a zaak — exactly what run-domain-check.sh does, and the same class of constraint as the
# `objecten.local` alias (ADR-0029). The ACL resolves the zaaktype itself (S-27, ADR-0021), so the
# base URL is the only thing to inject.
echo ">> recreating the acl service pointed at OpenZaak's IP"
ACL_OPENZAAK_BASEURL="http://$oz_ip:8000/" docker compose -f "$compose" up -d acl
WAIT_TIMEOUT="${WAIT_TIMEOUT:-120}" bash "$here/wait-healthy.sh" acl
# The container is replaced, so its IP may have changed.
acl="$(docker ps -q --filter 'name=[-_]acl[-_]' | head -1)"
acl_ip="$(ip "$acl")"
echo ">> opening a zaak through the ACL (which writes the INGEDIEND register record)"
reference="PROJ-$(date +%s)"
zaak_url="$(docker run --rm --network "$net" curlimages/curl:latest \
+15
View File
@@ -59,6 +59,20 @@ def services_in_trace(trace_id):
return names
def tempo_ingest_state():
"""#156: distinguish a broken trace chain from Tempo dropping spans. `ingester_clients` is 0
when the distributor has evicted its (single, in-process) ingester over a failed loopback
health check — pushes fail and spans are lost, which looks identical to missing instrumentation
from here. Diagnostics only; never fails the check."""
try:
for line in _get(f"{TEMPO}/metrics").decode().splitlines():
if line.startswith("tempo_distributor_ingester_clients "):
return f"tempo {line.strip()} (0 = no ingester in the pool — evicted, so pushes\n are failing and spans are being dropped; see #156)"
except Exception as e:
return f"tempo /metrics unreadable: {e}"
return "tempo_distributor_ingester_clients not reported"
def main():
deadline = time.time() + TIMEOUT
generate_traffic()
@@ -74,6 +88,7 @@ def main():
generate_traffic()
print(f"FAIL — no single trace spanned {sorted(WANT)}; services seen: {sorted(seen)}",
file=sys.stderr)
print(f" {tempo_ingest_state()}", file=sys.stderr)
return 1
+7 -3
View File
@@ -15,9 +15,13 @@ set -euo pipefail
timeout="${WAIT_TIMEOUT:-420}"
deadline=$(( $(date +%s) + timeout ))
# compose service name -> container id. The name filter matches both docker
# compose ("infra-openzaak-1") and podman-compose ("infra_openzaak_1") naming.
cid_for() { docker ps -aq --filter "name=$1" | head -1; }
# compose service name -> container id. `--filter name=` is a substring match, so it is anchored on
# the compose replica suffix — otherwise 'objecten' also matches objecten-db / objecten-redis /
# objecten-celery, and 'objecttypen' matches objecttypen-db. Whichever docker listed first won, so a
# service with a sibling that has no healthcheck timed out with status=none while it was in fact
# healthy. The pattern matches both docker compose ("infra-objecten-1") and podman-compose
# ("infra_objecten_1") naming; the same anchoring the verify check scripts use.
cid_for() { docker ps -aq --filter "name=$1[-_][0-9]+\$" | head -1; }
for svc in "$@"; do
echo "waiting for '$svc' to be healthy (timeout ${timeout}s)..."
@@ -18,10 +18,20 @@ public sealed record Notification(
string Actie,
Uri ResourceUrl)
{
/// <summary>A register record written to Objecten — <c>create</c> on submit, <c>update</c> on
/// approval, since the ACL upserts the same object for a registration (§8.6).</summary>
/// <summary>
/// A register record written to Objecten — <c>create</c> on submit and <c>partial_update</c> on
/// approval, since the ACL upserts the same object for a registration (§8.6).
/// </summary>
/// <remarks>
/// <c>partial_update</c> is what a PATCH actually reports: DRF routes it through the notifying
/// <c>update()</c> but names the action <c>partial_update</c>, and that is what Objecten puts in
/// the notification. <c>update</c> is accepted too, so a PUT-shaped write would project the same
/// way. <c>destroy</c> is deliberately not: removing a registration from the public register is
/// its own decision, not a side effect of this one.
/// </remarks>
public bool IsRegisterRecordWritten =>
Kanaal == "objecten" && Resource == "object" && Actie is "create" or "update";
Kanaal == "objecten" && Resource == "object"
&& Actie is "create" or "update" or "partial_update";
/// <summary>The object holding the register record. For a <c>resource: object</c> notification
/// Objecten sends the object as both <c>hoofdObject</c> and <c>resourceUrl</c> — the object is
@@ -38,13 +38,17 @@ public sealed class NotificationProjectorTests
Assert.Equal("REG-2026-0001", entry.Reference);
}
[Fact]
public async Task approval_updates_the_same_row_from_ingediend_to_ingeschreven()
// The ACL PATCHes the same object on approval. DRF routes a PATCH through `update()` but reports
// the action as `partial_update`, which is what Objecten puts in the notification — so accepting
// only `create`/`update` silently drops every approval.
[Theory]
[InlineData("partial_update")]
[InlineData("update")]
public async Task approval_updates_the_same_row_from_ingediend_to_ingeschreven(string actie)
{
var projector = Projector();
await projector.HandleAsync(RecordWritten());
// The ACL PATCHes the same object on approval, so Objecten publishes an `update`.
await projector.HandleAsync(RecordWritten("update", status: RegistrationStatus.Ingeschreven));
await projector.HandleAsync(RecordWritten(actie, status: RegistrationStatus.Ingeschreven));
var entry = Assert.Single(await _store.AllAsync());
Assert.Equal(ZaakId, entry.Id);
@@ -112,7 +116,7 @@ public sealed class NotificationProjectorTests
{
var projector = Projector();
await projector.HandleAsync(RecordWritten());
await projector.HandleAsync(RecordWritten("update", status: RegistrationStatus.Ingeschreven));
await projector.HandleAsync(RecordWritten("partial_update", status: RegistrationStatus.Ingeschreven));
var callsAfterProjection = _acl.CallCount;
await projector.RebuildAsync();