rumman.ahmed
All writing

ai / ai / myai

What's New in myAI: Universal Boot, a Smarter Brain, and Self-Hosted CI

Two months of changes to myAI since the last architecture post — a universal daily boot, a brain that surfaces prior art and grades its own health, a fleet-wide MCP bridge, self-hosted CI runners, and the doorway that lets me drive the whole thing from my phone.

I wrote a long post about myAI's architecture back in July — the brain, the agent fleet, the gateway. That post is still accurate, but it's a snapshot, and the system hasn't stood still. This one covers what actually changed since then: not a roadmap, not a wishlist — commits that landed, in most cases because something broke in production first.

If you haven't read the first post, the short version: myAI is a git-versioned memory layer plus an agent fleet plus a task queue, all sitting behind one gateway. The CLI is ai-management on npm, the dashboard runs at ai.rummanahmed.com, and there's a public mirror at github.com/knofler/myai.

Boot is universal now, not scheduled

The original setup split boot duty by machine and day — office Mac Monday through Thursday, home Mac Friday — mostly to avoid two Macs racing to rebuild the same shared state. That schedule was itself a workaround, and it meant a machine sitting idle on its "off" days had stale context if I picked it up.

The fix was simpler than the problem it replaced: every machine boots every day. Each host now writes its own boot report instead of one shared file, so two Macs never contend for the same write. The operator instruction behind the change, verbatim from the commit: "each separate machine, separate agent, single git repo, single brain repo, dropbox ignore to avoid conflict, always sync with git and brain and dropbox file, boot is universal for all device always." Once state lives in git and the brain rather than in a machine-pinned file, there was no longer a reason to gate boot by a calendar.

The brain surfaces prior art and grades itself

The brain was already git-versioned memory with delta-based boot — the July post covers that mechanism in depth. Two things changed since:

It stopped losing signal to noise. A stop-hook auto-generates a fallback memory atom when a session closes without a real handoff, so continuity doesn't just die because I forgot to write one. But early on, that fallback dumped raw command lists into the next session's boot — technically present, practically useless. It's now distilled to a few lines, skipped entirely if the session made no real changes, and ranked below atoms a session actually wrote on purpose.

It started telling the next session what it already knows. If the same fact shows up independently across three or more session atoms, the distiller promotes it to a standing memory fact instead of leaving it buried in session history. More usefully, boot-time delta now explicitly flags when a new session's topic matches something already in memory — a direct fix for a real incident where prior art existed in the brain but no session ever surfaced it, so work got redone. And the brain now reports a health score at every session start — not just "is it there" but a graded signal, including what fraction of recent sessions actually produced a usable memory atom versus quietly closing without one.

flowchart LR
    S["session closes"] --> H{"real handoff written?"}
    H -->|yes| A["session atom<br/>ranked by what it says"]
    H -->|no| F["auto-fallback atom<br/>distilled, BRONZE-ranked,<br/>skipped if nothing changed"]
    A --> M["brain main"]
    F --> M
    M --> D["distiller"]
    D -->|"fact repeats ≥3x"| P["promoted to<br/>standing memory"]
    M --> N["next boot: delta<br/>+ prior-art flag<br/>+ health score"]

A fleet-wide MCP bridge — and the outage that hardened it

Every repo the agents work in needs its .mcp.json correctly wired to the myAI MCP server, or the agent boots without brain or gateway tools. That used to depend on a manually maintained list of repos. It now also scans the whole project tree for any directory carrying MCP config and folds in whatever it finds — a new repo doesn't need to be told about; it gets picked up.

The more interesting story is a bug the bridge itself caused. The tool that merges each repo's config against a shared template was set up to union-merge every list field — sensible for most config lists, wrong for one specific field: a server's command-line arguments, which are positional, not a set. Two overlapping copies of the same command got merged into one array with the arguments interleaved, which meant the wrong script ran. The result was 27 of 30 repos silently losing their MCP server on the next boot rollout. The fix taught the merge tool the difference between a set and an argv list, and the fleet went from mostly-broken to fully working on the next run. It's a good reminder that a "just merge the configs" tool is a small parser away from corrupting every consumer at once.

Self-hosted CI, because metered minutes ran out

Nothing exotic here, and I'd rather say that plainly than dress it up: this repo is private, GitHub Actions minutes on a private repo are metered, and enough workflows fire per pull request that the billing allowance ran out. When that happens, checks fail closed and automated PR review goes quiet with them — not an outage you notice until you go looking.

The fix was a self-hosted runner: an official actions/runner binary (not a third-party wrapper) in a container I control, registered against the repo, no long-lived token sitting on disk — it mints its own registration token at startup and deregisters cleanly on shutdown. It has no Docker socket mounted, deliberately, so it can't reach outside its own container. It's scoped to one pilot workflow — the fully hermetic unit-test suite, no gateway or network dependency — and it isn't a required check, so if the runner itself is offline, a PR queues instead of blocking. Everything else still gets verified the way it always has: locally, in Docker, before anything gets pushed.

The task queue got harder to fool

The overnight runner works a task queue against a shared store, and a queue that runs unattended has exactly one job: never lose a task, and never do a finished one twice. Two fixes since July, both provoked by watching the queue misbehave rather than by reading a spec.

First, a genuine retry loop needed a floor: a task that keeps failing now dead-letters after a bounded number of attempts instead of looping forever or silently vanishing, and re-queuing it clears the retry count.

Second, and more subtle: the backlog cursor that tracks which planned tasks have already been turned into queue entries was confirming its own progress by re-querying the task list — capped at a fixed page size. As the fleet-wide task table grew past that cap, a task that was already confirmed on one pass could fall outside the window on a later pass, and the cursor would interpret that as "never confirmed" and rewind — re-dispatching work that was already done, a little at a time, indefinitely. The fix was to stop trusting a capped list query as the only source of truth and add a durable ledger of exactly what's been confirmed, so a later pass can check the ledger instead of re-deriving the answer from a query that might not see far enough back.

flowchart TB
    P["planner writes<br/>backlog line"] --> C["cursor confirms<br/>via task list query"]
    C -->|"table grows past<br/>the query's page cap"| L["line falls out of view"]
    L --> R["cursor reads that as<br/>'not yet confirmed'"]
    R --> D["re-dispatches<br/>already-done work"]
    C -->|"fix: durable ledger"| K["confirmed once,<br/>recorded permanently"]
    K --> N["later pass checks ledger,<br/>never re-derives from a capped query"]

A doorway, reachable from my phone

This is the part that's easiest to undersell: a "doorway" is just an idle, phone-drivable session running under its own profile, sitting there costing nothing until I actually send it something. remote status shows which repos currently have one open; starting one is duplicate-guarded so I don't end up with three idle sessions burning RAM for no reason.

I found the gap in the safety net the annoying way: the morning boot ran, but no doorway ever opened, and the completion ping that was supposed to tell me that also never arrived — because the notifier's Telegram configuration was empty and it exited quietly instead of complaining. Both are now first-class, checked steps of the boot sequence rather than best-effort side effects: the doorway's existence gets verified, and the notification's outcome gets recorded as pass or warn-with-a-reason, not just fired and forgotten. Telegram pairing itself got a matching fix — it now fails closed instead of silently leaving the chat ID unset.

None of this is glamorous. It's the unglamorous half of building something autonomous: not the agent doing the work, but making sure you find out when it silently didn't.

What I'd take from these two months

The theme across all of the above isn't new features so much as closing gaps that only showed up under real, unattended use — a rewind bug that only appears once a table grows past a page size, a merge tool that's correct for every field except one, a notifier that fails silently instead of loudly. None of these were visible in a demo. All of them were visible in the queue, the logs, or a missing phone notification a week later.

If you're building something similar — agentic delivery, a memory layer, MCP infrastructure — I'd like to compare notes. Get in touch, try the CLI (npm i -g ai-management), or poke at the public mirror. And if you want the fuller architecture picture this post assumes, start with building myAI.