2026-05-24

Six Machines, One Brain, Eleven Thousand Commits Behind

I run a fleet of AI nodes across six machines. A Mac, some Windows boxes, a couple of Linux servers across the ocean. Each one runs its own Claude, and they all share one brain: a private git repo full of memory files. Every node pulls the repo when it wakes up, auto-commits after it works, and there's a WIP file so any node can pick up where another left off mid-task.

It's a beautiful design. I want to be clear about that, because everything after this paragraph is about how it rotted from the inside while telling me it was healthy.

The corkboard

The part that gets people is the interoffice memo file. The machines leave each other notes in it. Actual action requests, machine to machine. "Please add these redirect URIs." "M5, please fire this payload — this node has no messenger." It reads like a break-room corkboard for software. One of my computers filed a ticket asking another one of my computers to do a thing, because it lacked the hands. I have built a small bureaucracy, and the employees are all named after laptops.

I spent four years in the Marine Corps. Passing word up and down a chain of command is a thing I understand in my bones. What I did not anticipate was becoming the mail clerk for six machines that CC each other.

The rot

Here's the design flaw, and I'll show you the actual code, because the code is the punchline:


git push origin main 2>/dev/null || true

Every sync hook, on every node, wrapped its push in that. And every session-start pull ran with --ff-only, meaning: fast-forward or nothing.

Translate that out of git and into English. 2>/dev/null means "throw away the error message." || true means "and if it failed, report success anyway." --ff-only means "if this branch has diverged from the shared truth, do nothing — quietly."

So the failure mode was preordained. The moment any node fell even one commit behind origin, every subsequent push got rejected as non-fast-forward. Forever. Silently. The pull that should have fixed it refused, because fast-forward was impossible. Also silently. The node kept committing its work locally — dutiful, disciplined, every session — and every commit went straight into a pile of undelivered mail.

The machine believed it was current. The hooks told it so. || true, after all.

The body count

It surfaced the way these things always surface: by accident. A routine git status on a cloud node showed it sitting on 136 unpushed commits. Nobody noticed. No log, no banner, no symptom. The machine had been journaling into the void for weeks like a lighthouse keeper who doesn't know the war ended.

Then it got worse. A multi-node divergence — several machines all pushing, all failing, all merging their own private realities — spiraled the shared repo to 1,900-some commits on one side versus 3,400-some on the other before I sat down and reconciled it by hand.

And then, at rollout time, the true champion revealed itself. One of the Windows nodes was ahead 29, behind 11,231. Read that again. Eleven thousand two hundred thirty-one commits behind the shared brain, holding 29 commits of its own that it had been politely trying to deliver the whole time. The server node was ahead 2, behind 10,691, with about twenty merge conflicts waiting inside.

These machines all believed they were current. Every one of them. If you asked, the hooks would have said everything was fine, because the hooks had been surgically prevented from saying anything else.

The root cause was not auth. Not the network. Not GitHub. It was silent failure plus a reconcile step that couldn't reconcile. I had built an organization where every field office confirmed receipt of orders it never received.

The fix

The fix wasn't clever. The fix was to stop lying.

The reconciliations themselves were grim work. Origin was declared source of truth, the stragglers' conflicts resolved in its favor, backup branches kept in case I'd just declared the wrong reality the winner. The lighthouse keepers' journals were merged into history. Their work was not lost. It was just very, very late.

And here's the detail I keep coming back to. When I deployed the new session-start hook on my primary Mac — the flagship, the node that runs the show — its first live run immediately found and healed a real divergence. On the boss's own machine. The old hooks sitting right there were the exact broken pattern, --ff-only and 2>/dev/null || true, same as everywhere else. The command node had the same disease as the field offices. It just hadn't been asked a hard question yet.

One brain, with receipts

Rollout finished across all five online nodes on June 9th. Now every node self-heals, every sync attempt leaves a log line, and a diverged branch is a Tuesday instead of an archaeology project. The fleet genuinely shares one brain, and — this is the part that matters — it can prove it, which is more than it could do before, when it merely claimed it.

So that's the state of the org. Six machines, one memory, a corkboard where the software leaves each other notes, and a boss who spent a weekend as a git therapist, coaxing a computer through the realization that the last eleven thousand things it was told never actually arrived.

In the Corps we had a saying about communication: the message isn't sent until it's acknowledged. Turns out || true is not an acknowledgment. It took me six machines and fifteen thousand commits of drift to re-learn something they taught me at eighteen, for free, with yelling.

The machines are fine now. They talk, they merge, they log. Whether this is the org chart of the future or a support group for computers depends on the hour, but either way — attendance is mandatory, and now we take minutes.