Solution Architecture  /  Observability  /  Operations

Plenty of alerts.
Nobody was told.

The test of observability is not how many dashboards exist. It is how quickly you can answer a specific question when something goes wrong, and whether a person finds out at all. This is how we assessed a multi-channel visitor platform against four real failure modes rather than a tooling wishlist, what the assessment found, and the order we chose to fix it in: an alert that reaches a person first, correlation second, the pretty panel last.

ONE VISITOR JOURNEY · HOW FAR DOES THE TRACE ID GET? APP error tracker GATEWAY two trace formats OUR APIS id made, then dropped EVENT BUS no correlation TICKETING vendor, closed logs PAYMENTS vendor, own records ×××× ALERT FIRES into a portal A PERSON never finds out NO ROUTE

Every platform eventually has an observability conversation, and it almost always starts in the wrong place: which dashboard product to buy. The more useful starting point is a list of things that will actually go wrong, and an honest check of how long it would take to answer each one today. Framing the work around real failures, rather than a schedule of tooling, changed both what we found and what we fixed first.


01Judge it by questions, not dashboards

The goal we set was end-to-end traceability: any log entry, from any service we operate, can be traced back to the request or event that caused it, across every integration path. That is a long way off on a platform built from a dozen SaaS products. So instead of measuring progress by coverage, we wrote down four questions that come straight from the failure modes of a ticketed venue, and assessed every system against them:

Why did this visitor's tickets take so long to appear?A booking is made; the app shows it late, or not at all.
LATENCY
Was this customer charged twice?The most expensive question on the list, in money and in trust.
PAYMENT
Is the booking queue backing up, and why?High-demand on-sales sit behind a virtual waiting room we do not host.
QUEUE
Is what the visitor sees stale, and how stale?Several caches sit between the content source and the screen.
FRESHNESS

Four questions is deliberately few. Each one names a person who will ask it (a visitor services lead, a finance officer, an on-sale manager, a content editor) and a moment when they will ask it, which is always the worst possible moment.


02An estate that speaks three dialects

No platform built over a few years has one telemetry story, and ours was no exception. The assessment found three dialects in use:

That mix is not a failure. Each choice was reasonable for its team and its runtime. The problem is what happens at the seams between them, and the best existing work was already a template for the fix: our most mature ingestion function reports data freshness, processing time, resource use and failures as structured logs, and the event publisher reports notification volumes, delays and publishing problems, with a read-only query tool and a manual dead-letter redrive. The pattern existed. It had simply not been made the standard.


03Where the chain breaks

Correlation stopped at nearly every boundary, each for a locally sensible reason:

The last one is the quiet killer. Telemetry that the operating team cannot query is, for practical purposes, telemetry that does not exist. Access to the logs is part of observability, not an afterthought for the security review.


04Nobody is told

This was the finding that mattered most, and it had nothing to do with tracing. When we looked, there was a respectable number of metric alerts, and almost none of them had an action group, so they fired into the portal and nowhere else. The few that did route somewhere routed to receivers nobody on the team could see. And the alerts that existed were platform-level; the application layer (the gateway, the functions, their errors and latency) was not covered at all.

The motivating incident was plain and slightly embarrassing: a development API had been returning server errors for an unknown length of time. It was development, so the cost was small. The same wiring in production would have meant visitors finding the outage before we did.

The audit question

Do not count your alerts. For each one, ask: who receives it, on what channel, and when did they last actually receive one? An alert with no answer to all three is decoration.


05The four questions, against what exists

Delayed ticket loads

The integration layer measures how fresh its copy of ticketing data is, which answers half the question. The other half is the app: it finds a new booking by polling a handful of times over several seconds, and nothing signals when the ticketing system has actually created the order. So "why was it slow" has an answer on one side of the boundary and a guess on the other.

Double payments

Web checkout detected payment completion by polling rather than by consuming the payment provider's notification. Polling a payment is not neutral: every extra check is a chance to act twice on one event, and it leaves no authoritative record of when the payment actually completed. The fix is structural, not a better dashboard. Obtain a payment confirmation signal and treat it as the source of truth.

Queue backups

The virtual waiting room's configuration was not in any repository, and the booking shell could only see that a visitor had arrived from the queue. The ticketing vendor confirmed its own logs cannot be exported to any external endpoint, so its view of an incident lives only in its admin panel. Some questions will always need a human with a login. The job is to know which ones, in advance.

Stale content

The integration API does not cache; it recomposes every response from its store, so its side of freshness is simply the ingest schedule. The caches live downstream: the content API holds responses at the edge for an hour with no purge wired, and one backend refreshes on a timer and never expires entries. Only one product on the estate captions its data with the time it was true. That caption is the cheapest observability feature there is.


06Join on business keys, because traces will not cross

Here is the uncomfortable truth about a platform built on SaaS: trace IDs will never cross into the ticketing vendor, the payment provider or the waiting room. No amount of OpenTelemetry changes that. So the join between our telemetry and theirs has to run on something both sides already record: business keys.

THE JOIN RUNS ON KEYS BOTH SIDES ALREADY RECORD OUR SIDE app + gateway + APIs functions + event bus OpenTelemetry, one set of log fields BUSINESS KEYS customer id constituent id order number performance venue THEIR SIDE identity sign-in logs payment records ticketing audit trails no trace ids, ever ONE PANEL · CONSISTENT DEFINITIONS · JOINED ON KEYS
Figure 1. Traces carry the story inside our boundary. Business keys carry it across boundaries we will never control.

That makes one unglamorous discipline the whole foundation: log the same business keys, with the same field names, on our side of every boundary. A proposed common set of fields looks like this:

// one log line, any service, any runtime
{
  "ts":             "2026-09-29T09:41:07.412Z",
  "service":        "booking-broker",
  "env":            "prod",
  "trace_id":       "4bf92f3577b34da6a3ce929d0e0e4736",  // inside our boundary
  "correlation_id": "c-7f2e…",                         // survives the event bus
  "customer_id":    "auth|…",                           // the business keys:
  "order_ref":      "…",                                // these cross boundaries
  "performance":    "…",
  "venue":          "…",
  "outcome":        "ok | retry | fail",
  "latency_ms":     184
}

Illustrative field set, not a live record. Keys are pseudonymous identifiers, never names, emails or card data.

Those keys are exactly what a finance officer, an on-sale manager or a visitor services lead already has in front of them when they ask the question. Logging them is what turns "we will look into it" into a query.


07The order matters more than the tools

With the findings on the table, the plan was mostly about sequence. Monitoring existed across several components; what was missing was standardisation, correlation and one view. We chose to do those in a strict order:

FIRST An alert that reaches a person

Action groups with real receivers on every existing alert. Failure, server-error and latency alerts on the gateway and every function app. Read access for the team that operates the services.

cheap · fast · closes the worst risk before launch

SECOND Correlation

W3C trace context on every gateway API, a policy that forwards the trace ID, a correlation ID on every event, and the common log fields in every service, starting from the most mature function as the template.

standards work · one repository at a time

THIRD One panel

A shared view on consistent definitions, joining application telemetry with vendor evidence on business keys, surfaced to management through the reporting product the organisation already uses.

only worth building once the first two are true

A beautiful dashboard built before step one would have shown green while nobody was paged, which is worse than no dashboard, because it teaches people to trust it.

The tooling question, reframed

I had originally framed tooling as a choice between two dashboard products. The code had partly answered it already, and the answer was less interesting than the question behind it: whether OpenTelemetry becomes the common language across serverless functions, container apps and the edge platform, with trace context carried through the gateway and the identity broker, so that one trace can follow a request as far as our side of the boundary goes. The dashboard product matters less than that. And it matters less again than an alert that reaches a person.

An alert first, correlation second, the panel third. Every other order produces a very good-looking outage.


Still open

What is not done yet

← All writing rummanahmed.com Share on LinkedIn