verify-tracing is flaky — Tempo intermittently drops spans under runner load #156
Closed
opened 2026-08-28 12:30:22 +00:00 by not
·
0 comments
No Branch/Tag Specified
main
ci/175-deploy-on-merge
feat/177-public-tls-edge
ci/168-helm-chart-ci-gate
docs/169-mkdocs-nav
feat/25-helm-kubernetes-caddy
fix/161-e2e-bounded-and-diagnosable
feat/162-werkbak-live-refresh
feat/132-medewerker-mfa
fix/156-tempo-ingester-healthcheck
feat/153-projection-sourced-from-objecten
feat/152-objecten-publishes-to-nrc
feat/149-acl-writes-registerrecord
feat/141-registerrecord-objecttype
perf/verify-stack-uwsgi-oz-nrc
fix/144-verify-stack-uwsgi
feat/140-objecten-up
feat/139-objecttypen-up
feat/131-default-fill-crud
chore/136-ci-job-summaries
fix/134-verify-stack-scheduling
feat/130-beheer-catalogi
feat/124-metrics-dashboards
ci/127-parallel-jobs
feat/123-distributed-traces
feat/111-self-service-resume
feat/113-acl-zaaktype-by-identificatie
fix/110-compose-local-flow
fix/115-e2e-single-worker
docs/111-backlog-s26
feat/106-close-zaak-on-timeout
feat/103-diploma-upload-documenten
feat/102-document-wait-timeout
feat/14-dmn-diploma-eligibility
feat/15-beoordeling-escalation
fix/portal-nginx-resolver
fix/local-eventsubscriber-acl
feat/12-withdrawal-portal
fix/91-local-compose-parity
feat/12-withdrawal-bff
feat/12-withdrawal-workflow
feat/12-withdrawal
feat/13-behandel-portal
feat/13-behandel-decide
feat/13-behandel-bff-auth-werkbak
feat/13-workflow-user-tasks
feat/13-behandel-decision-model
chore/release-2026.07.0
feat/78-reference-correlation
feat/75-approval-flow
feat/10-openbaar-portal
chore/73-ci-speedups
feat/68-e2e
feat/67-self-service-form
feat/66-api-client
feat/65-nx-workspace
feat/8-bff
feat/6-domain-service
feat/7-event-subscriber-projection
feat/56-nrc-notification-wiring
test/46-acl-openzaak-integration
feat/47-acl-mutation-baseline
ci/30-gitea-actions-ci
feat/5-acl-open-zaak
feat/4-flowable
feat/3-keycloak
feat/2-opennotificaties
feat/2-catalogus-seed
feat/10-openzaak-compose
feat/32-docs-scaffold
feat/31-contributor-workflow
feat/30-gitea-actions-ci
feat/29-bff-docker-compose
chore/remove-bootstrap-scripts
feat/28-bff-health
docs/split-s00
v2026.07.0
Milestone
No items
No Milestone
Projects
Clear projects
No projects
Assignees
eho (Edwin van den Houdt)
Clear assignees
No Assignees
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: eho/register-referentie#156
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Observed on
verify-stackrun 722 (PR #155).verify-tracingfailed:An identical run on the same branch 20 minutes earlier (721) passed, and the tracing path was
untouched by that PR. A plain re-run of the same job on the same commit went green, so it is
intermittent, not a regression.
It is not a broken trace chain — Tempo could not ingest. From the same run's container logs:
No OOM kill and no restart — Tempo stayed up but dropped spans during the window, so
projection-api's never landed and the check saw only
bff.Why this is worth fixing rather than tolerating
CLAUDE.md §15: flaky tests are fixed, not retried. The check gates merges, so every occurrence
costs a full ~20-minute
verify-stackcycle, and it trains people to re-run on red — which isexactly how a real tracing regression would get waved through.
Notes for whoever picks this up
temporuns withmem_limit: 400m(set in commitd5e5fa2to stop the backplane starving theapp stack + Playwright on the memory-tight runner). Worth checking whether that is now too
tight: the stack has gained containers since, most recently
objecten-celery(#152).CELERY_WORKER_CONCURRENCYis a dead end. The Maykin image already defaults it to 1, soobjecten-celeryis a single-process worker and there is nothing to turn down. (The #152 PRnotes name it as "the next knob" — that is wrong; this issue supersedes it.)
run-tracing-check.shpolls withTRACING_TIMEOUT(CI passes 120s). If the real cause isingestion lag under load rather than dropped spans, the fix is on Tempo's side — a longer
timeout would just paper over dropped data.
err="context canceled"suggests the exporter gave up on the push, so the .NET OTLPexporter timeout is also worth a look alongside Tempo's limits.