A platform for a venue with half a dozen visitor and staff channels and a dozen SaaS products is not one system. It is a handful of shared layers that every channel reaches through, and the discipline of keeping them shared. Most of what follows is about that discipline. It starts with the smallest app on the estate, because small apps are where rules get skipped first.
01Write the rule you actually follow
I started with a clean rule: no external client talks to a source system directly. It is a good rule. It is also not quite true, and reading the code showed exactly where.
It holds perfectly for content. No visitor-facing client composes a page from the content platform, the venue scheduling system or the collection database. But a profile service calls the ticketing system for a visitor's bookings. The identity broker binds a session through the ticketing web API. The mobile app reads the shop's public storefront API. Each of those is a backend we own, or a public API designed for exactly that call, doing one specific transaction against its system of record. That is the right shape. It just was not the rule as I first wrote it.
So the rule was rewritten to match the practice, and it came out narrower and more useful:
No channel composes content from source systems, and no channel holds a private source-system credential.
The first half keeps content consistent, because there is one place where source data becomes a canonical shape. The second half keeps the blast radius small, because a stolen device, a leaked bundle or a curious browser tab gets nothing it can use. Both halves are checkable in a code review. The original rule was not, because every reviewer would have argued about what counted as "directly".
Practice
When the code and the rule disagree, find out which one is right before you fix either. Sometimes the code is wrong. Often the rule was written before anyone knew the edge cases, and the honest fix is a narrower rule that everyone can actually follow.
02A backend-for-frontend at the counter
The front-of-house dashboard is a Next.js application on Vercel, built for a wall-mounted tablet behind the visitor counter and installable as a progressive web app on staff phones. The browser never calls an upstream service. Every request goes to the app's own route handlers, which act as a backend-for-frontend: they hold every credential server-side, call the upstream, validate what comes back, and return a small, typed envelope.
Three choices in how that layer is built do most of the work:
- Every upstream is its own package. The repository is a monorepo, and each upstream (content, events, schedules, occupancy, weather) lives behind its own client package. A change to one supplier's API is a change to one package, reviewed on its own.
- Contracts are shared, not copied. The shapes the browser receives live in a contracts package that both sides import, so the frontend and its own backend cannot quietly drift apart.
- Everything from outside is validated. Upstream responses are parsed with Zod before they are trusted. A supplier that changes a field name produces a clear, logged validation error at the boundary, not a blank screen at the counter on a busy morning.
None of this is exotic. It is the pattern you would want in any app, applied to a small one on purpose, because a small app with credentials in the browser is the easiest breach on an estate to create and the hardest to notice.
03Use the operations team's tools without calling them
The operations team already runs its day on a set of work-management boards: a daily run sheet, a staff FAQ, and the day's school groups. The obvious build is to call that product's API from the dashboard. We deliberately did not.
Instead, the integration layer syncs those boards into its own store on a timer, and serves them through the same gateway as every other feed, on the same base URL with the same key. From the dashboard's point of view there is one supplier with one contract. The ops team keeps the tool they already know, and nobody has to learn a new admin screen.
Cache the good answer, never the failure
The dashboard caches board data for five minutes, and it never caches a failure. That one line matters more than the number. If the upstream has an outage, the app keeps serving the last good copy rather than replacing it with an error that then sticks for five minutes. A scheduled warm-up call runs during opening hours, so a staff member is rarely the first person to read an expired entry and wait for it.
Why this shape
Syncing into a store you own turns someone else's rate limits, outages and API changes into your ingestion problem, handled once, instead of every channel's runtime problem, handled badly and repeatedly.
04Offline is a feature, and so is forgetting
Venue networks drop, usually at the busiest moment. The dashboard treats that as a normal state rather than an error.
- A service worker serves pages network-first, with the last copy as the fallback.
- The query cache is persisted to IndexedDB for 24 hours, so a reload draws the last data immediately and refreshes behind it.
- Search runs on the device, over data it already has: one bar across months of program, spaces and what is on in them, and the collection. It keeps working with no network at all.
The part people skip is deciding what must not survive. Live occupancy and weather are deliberately excluded from the persisted cache. A four-hour-old occupancy count shown with confidence is worse than no count, because a staff member will act on it. Staleness is not a property of the app. It is a property of each field, and some fields are better forgotten.
An offline copy is a promise about how old the truth is allowed to be. Make that promise per field.
05Every project starts as a written decision
Our architecture governance is not a board that approves designs. It is a sequence of documents, each small enough to be read, that make the reasoning visible before money is spent and keep it visible after the design changes.
What a good decision record looks like
Anything expensive to reverse gets an architecture decision record in the repository it governs: context, options considered, the decision, and its consequences. The best one on our estate, on how collection images are delivered, is the model I point people to. It compares four options. It states the organisational drivers as plainly as the technical ones. It carries a cost model with the crossover point where the answer flips. And it lists explicit conditions for revisiting the decision. It was accepted, built as decided, and one of its revisit conditions has since been met, which is exactly what it was written to catch.
Designs change. Records make that visible.
Several of our early designs did not survive contact with reality. A hybrid ticketing interface became the vendor's native flow with our own template, because the vendor's constraints were confirmed in writing. A multi-tier event design with several topics became one topic with an outbox. A proof of concept was built exactly as designed, then replaced by something simpler that did the same job.
None of those is a failure of the process. Each is the process working. The gap that matters is the other one: a change that happens in code and never gets written back into the decision. So the rules for new work are short:
- A short written decision before code: systems of record, where sensitive data lives and does not, the integration pattern, and what we are deliberately not building.
- An ADR, in the implementing repository, for anything expensive to reverse.
- Decisions are superseded, never deleted. Dated updates are appended, not rewritten.
- Proofs of concept list their deferrals: what was out of scope for the proof but required for production.
- When the code departs from the decision, the decision is updated, not the other way round.
- Every write-up ends with what we actually stood up and what is honestly still open. A design is more trustworthy when it names its own edges.
06One path to production
Standards that name a repository go stale the day a new repository is created, so ours are written as rules that apply to every repository we own, whatever its branches are called.
FEATURE One concern, one approval
Small enough to review properly. This is where review actually happens.
RELEASE Grouped, one approval
The question here is whether the release is what was reviewed, not a first read of the work inside it.
INTEGRATION → PROD Tests pass, reviewed PR
Production is reached only by a reviewed pull request from the integration branch after a green test run. Nothing skips a stage.
The rules
What an approval means
An approval is a statement that you have read the change and are comfortable with it going to production. It carries your name. Reviewers check correctness first, then risk (authentication, data handling, anything touching customer information or payment flows), then maintainability. Style comes last and never blocks. Asking for a split is a legitimate review outcome, and the turnaround target is one business day.
When the standards landed, work already in flight was reviewed at a high level rather than retro-split. Rules that punish the work that was already in the queue do not get adopted; they get resented.
07If it is not in code, it is not in the landing zone
Infrastructure standards are only useful if a deployment fails when they are broken. Ours are enforced by the cloud platform itself rather than by a checklist:
- One naming pattern: company prefix, region, environment, workload and type, with dash-free forms where a resource type demands it.
- Seven mandatory tags on every resource group (environment, application, purpose, business owner, system owner, creator and criticality). A policy denies an untagged group and pushes the tags down to its resources, so a missing tag fails the deployment outright.
- Delete locks on production resources, and a monthly budget with alerts on every landing zone.
- Pipelines that lint, build and run a what-if preview before deploying through each environment in order.
Split ownership where drift has already bitten
Drift is not theoretical. Once, re-running an application's infrastructure template reset configuration the application team had set by hand. The fix was not more care. It was an ownership line: the infrastructure team owns the template that provisions the foundation, and the application team owns images and configuration through its own repository and pipeline. Two teams, two sources of truth, no overlap for a re-run to overwrite.
Lesson
Whenever two tools can write the same setting, one of them will eventually undo the other. Draw the ownership line before that happens, not after.
08A rule that is not checked is a preference
The next step is making the decision record do more than record. In the platform's own control application, an ADR is designed to be the input to a build pipeline. Each record has two parts: narrative fields, which carry the reasoning, and a typed block of machine-actionable facts, validated with a schema. The narrative is context for a language model. The typed block drives everything that follows.
The reason is stated plainly in the design: a model is bad at being a deterministic configuration source and good at narrative reasoning, so the facts that drive a deployment never come from its prose. From the typed block comes a provider-neutral plan (a message bus, storage tiers, serverless transforms, a gateway), then generated infrastructure code that must compile before it is shown as ready, then a mandatory what-if preview. The model never triggers a deployment. A person does, and the approval is bound to the plan version and the preview hash, so approving version two does not authorise deploying version three. Budget and region guardrails are enforced by code, not by the model.
To be clear about status: only the first stage, ingesting the ADR, has shipped. The rest is an accepted design. But the principle behind it is already how the repository itself runs: specifications written to be executed by an agent with no prior context, a scheduled queue that sweeps for stale decisions and out-of-date documentation, and architectural invariants enforced as CI checks.
An architectural rule that is not mechanically checked is just a preference.
Still open
What we have not finished
These practices are adopted, not complete. In the spirit of naming our own edges:
- Not every repository has automated checks yet. Where one does not, review carries the whole load, and that is a gap, not a design.
- Some decisions changed in code before they were written back into their records. Catching those up is ongoing work.
- Pull request templates are meant to live in one organisation-level repository that every repository inherits. That repository still needs an owner.
- Ticket tracking and source control are not yet linked automatically, so traceability from a request to a commit still relies on people.
- The ADR-to-deploy pipeline is a design beyond its first stage, and is described that way on purpose.