Every platform eventually has an observability conversation, and it almost always starts in the wrong place: which dashboard product to buy. The more useful starting point is a list of things that will actually go wrong, and an honest check of how long it would take to answer each one today. Framing the work around real failures, rather than a schedule of tooling, changed both what we found and what we fixed first.
01Judge it by questions, not dashboards
The goal we set was end-to-end traceability: any log entry, from any service we operate, can be traced back to the request or event that caused it, across every integration path. That is a long way off on a platform built from a dozen SaaS products. So instead of measuring progress by coverage, we wrote down four questions that come straight from the failure modes of a ticketed venue, and assessed every system against them:
Four questions is deliberately few. Each one names a person who will ask it (a visitor services lead, a finance officer, an on-sale manager, a content editor) and a moment when they will ask it, which is always the worst possible moment.
02An estate that speaks three dialects
No platform built over a few years has one telemetry story, and ours was no exception. The assessment found three dialects in use:
- OpenTelemetry in the containerised .NET services, exporting to more than one backend.
- The cloud's built-in APM integration in the serverless integration functions, with no OpenTelemetry, and the API gateway logging to its own workspace at a sampled rate.
- An error tracker on the customer journey: the mobile app, the identity broker and the booking shell. Some web projects had the project created but the SDK never wired, so they had received no events at all.
That mix is not a failure. Each choice was reasonable for its team and its runtime. The problem is what happens at the seams between them, and the best existing work was already a template for the fix: our most mature ingestion function reports data freshness, processing time, resource use and failures as structured logs, and the event publisher reports notification volumes, delays and publishing problems, with a read-only query tool and a manual dead-letter redrive. The pattern existed. It had simply not been made the standard.
03Where the chain breaks
Correlation stopped at nearly every boundary, each for a locally sensible reason:
- The gateway used W3C trace context on some APIs and a legacy correlation protocol on others, and no policy forwarded a request ID.
- One API generated a request ID on the way in, then did not send it on its own upstream calls.
- Events on the bus carried no correlation identifier at all. There was a to-do in the event base class to add one.
- Some services called a vendor system directly rather than through the integration layer, so their calls never appeared in its telemetry.
- Much of the integration telemetry landed in a central workspace that the team operating those services could not read.
The last one is the quiet killer. Telemetry that the operating team cannot query is, for practical purposes, telemetry that does not exist. Access to the logs is part of observability, not an afterthought for the security review.
04Nobody is told
This was the finding that mattered most, and it had nothing to do with tracing. When we looked, there was a respectable number of metric alerts, and almost none of them had an action group, so they fired into the portal and nowhere else. The few that did route somewhere routed to receivers nobody on the team could see. And the alerts that existed were platform-level; the application layer (the gateway, the functions, their errors and latency) was not covered at all.
The motivating incident was plain and slightly embarrassing: a development API had been returning server errors for an unknown length of time. It was development, so the cost was small. The same wiring in production would have meant visitors finding the outage before we did.
The audit question
Do not count your alerts. For each one, ask: who receives it, on what channel, and when did they last actually receive one? An alert with no answer to all three is decoration.
05The four questions, against what exists
Delayed ticket loads
The integration layer measures how fresh its copy of ticketing data is, which answers half the question. The other half is the app: it finds a new booking by polling a handful of times over several seconds, and nothing signals when the ticketing system has actually created the order. So "why was it slow" has an answer on one side of the boundary and a guess on the other.
Double payments
Web checkout detected payment completion by polling rather than by consuming the payment provider's notification. Polling a payment is not neutral: every extra check is a chance to act twice on one event, and it leaves no authoritative record of when the payment actually completed. The fix is structural, not a better dashboard. Obtain a payment confirmation signal and treat it as the source of truth.
Queue backups
The virtual waiting room's configuration was not in any repository, and the booking shell could only see that a visitor had arrived from the queue. The ticketing vendor confirmed its own logs cannot be exported to any external endpoint, so its view of an incident lives only in its admin panel. Some questions will always need a human with a login. The job is to know which ones, in advance.
Stale content
The integration API does not cache; it recomposes every response from its store, so its side of freshness is simply the ingest schedule. The caches live downstream: the content API holds responses at the edge for an hour with no purge wired, and one backend refreshes on a timer and never expires entries. Only one product on the estate captions its data with the time it was true. That caption is the cheapest observability feature there is.
06Join on business keys, because traces will not cross
Here is the uncomfortable truth about a platform built on SaaS: trace IDs will never cross into the ticketing vendor, the payment provider or the waiting room. No amount of OpenTelemetry changes that. So the join between our telemetry and theirs has to run on something both sides already record: business keys.
That makes one unglamorous discipline the whole foundation: log the same business keys, with the same field names, on our side of every boundary. A proposed common set of fields looks like this:
// one log line, any service, any runtime { "ts": "2026-09-29T09:41:07.412Z", "service": "booking-broker", "env": "prod", "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736", // inside our boundary "correlation_id": "c-7f2e…", // survives the event bus "customer_id": "auth|…", // the business keys: "order_ref": "…", // these cross boundaries "performance": "…", "venue": "…", "outcome": "ok | retry | fail", "latency_ms": 184 }
Illustrative field set, not a live record. Keys are pseudonymous identifiers, never names, emails or card data.
Those keys are exactly what a finance officer, an on-sale manager or a visitor services lead already has in front of them when they ask the question. Logging them is what turns "we will look into it" into a query.
07The order matters more than the tools
With the findings on the table, the plan was mostly about sequence. Monitoring existed across several components; what was missing was standardisation, correlation and one view. We chose to do those in a strict order:
FIRST An alert that reaches a person
Action groups with real receivers on every existing alert. Failure, server-error and latency alerts on the gateway and every function app. Read access for the team that operates the services.
cheap · fast · closes the worst risk before launch
SECOND Correlation
W3C trace context on every gateway API, a policy that forwards the trace ID, a correlation ID on every event, and the common log fields in every service, starting from the most mature function as the template.
standards work · one repository at a time
THIRD One panel
A shared view on consistent definitions, joining application telemetry with vendor evidence on business keys, surfaced to management through the reporting product the organisation already uses.
only worth building once the first two are true
A beautiful dashboard built before step one would have shown green while nobody was paged, which is worse than no dashboard, because it teaches people to trust it.
The tooling question, reframed
I had originally framed tooling as a choice between two dashboard products. The code had partly answered it already, and the answer was less interesting than the question behind it: whether OpenTelemetry becomes the common language across serverless functions, container apps and the edge platform, with trace context carried through the gateway and the identity broker, so that one trace can follow a request as far as our side of the boundary goes. The dashboard product matters less than that. And it matters less again than an alert that reaches a person.
An alert first, correlation second, the panel third. Every other order produces a very good-looking outage.
Still open
What is not done yet
- A scheduled smoke check with chat notification exists as a script, and is not yet running on a schedule.
- The edge cache on the content API needs a wired purge, so freshness becomes controllable, and then an alert on content age.
- The common log fields are a proposal. They need agreement, and every plan point needs a named owner.
- Some web projects need their error-tracking SDK actually wired before they can report anything.