In January 2026 I wired eight layers of coordination infrastructure around a codebase so that several coding agents could work on it at once. By May I was running one agent, one pull request per session, and most of the eight were gone. This is what each layer was for, which parts I still use, and why the rest went.

Note

Rewritten in September 2026. The original was a 4,500-word setup guide for this stack, complete with a “measuring success” table whose thresholds had no measurement behind them. The layers below are the same. The verdicts are new, and they come from the workflow that replaced this one.

The problem I was solving

Several instances of Claude Code, in an IDE, in a terminal, on a second machine, all in one repository. The failure modes were easy to picture and I built for every one of them: two agents picking up the same issue, one agent’s uv sync corrupting another’s virtual environment, an agent crashing mid-task with its work abandoned in a shared working directory, and no way to tell from the outside what any of them was doing.

The instinct was to build coordination infrastructure rather than run fewer agents. That instinct was half right.

The eight layers, and the verdict on each

Ruler, one source of rules for every agent. Ruler generates CLAUDE.md, .cursor/rules, Copilot instructions and the rest from one .ruler/AGENTS.md. Verdict: dropped, because I stopped running tools that want different files. One hand-written AGENTS.md at the repository root, read by every agent I use, does the job with nothing to regenerate. The lesson underneath it survived: the rules file is the file in the repository that matters most, and it should be written by hand, not generated.

Beads, issues in git. Beads stores issues as one JSON line each in a tracked file, with dependencies, so an agent can run bd ready and get unblocked work without a network call. Verdict: dropped with the multi-agent setup, but of all eight this is the one I would reach for again. Agents are bad at web issue trackers and good at CLIs over local files. When there is one agent, a STATUS.md file naming the current task does the same job in fewer moving parts.

PostgreSQL advisory locks, atomic claims. Two agents can both see an issue as open and both claim it. A session-level pg_try_advisory_lock on a hash of the issue id closes the window, and it releases itself if the agent crashes. Verdict: dropped, because there is nothing to claim when one agent works at a time. The design was sound. It solved a problem I chose to stop having.

Git worktrees, one directory per task. Each task gets its own checkout on its own branch, sharing .git, so build artifacts, virtual environments and uncommitted work cannot collide. Verdict: kept, for a different reason. With one agent there is nothing to isolate from, but a checkout per task still means a half-done branch never sits in the directory the next session starts from. Worktrees cost disk and nothing else.

Skillz, validation skills over MCP. Read-only skills (allowed-tools: Read, Grep, Glob) that check API contracts, layer boundaries and security patterns on request. Verdict: replaced. The checks moved out of prompts and into a pre-push hook, where they cannot be skipped. A validation an agent can decline to run is a suggestion. A hook that exits non-zero is a gate.

Agent learning, a corpus of insights. The layer I was proudest of: a git hook extracted one- to three-sentence learnings from commits, stored them with the files and commits that backed them, and fed the relevant ones back to the next agent within a token budget, with confidence decaying as the underlying files changed. Verdict: replaced by a file. LEARNINGS.md, appended to by a /retro step at the end of every session, read by an /orient step at the start of the next. No database, no vector search, no decay function. The idea that hard-won context should be stored in the repository and fed back on demand was right. The machinery around it was more than the idea needed.

Git hooks. Beads installed its own, and pre-commit ran the formatter, the type checker and a rules regeneration. Verdict: kept and grown. The pre-push gate is now eight sequential checks, and it is the only layer from this list that got bigger. Everything that must happen lives in a hook. Everything that should happen lives in a prompt, and the difference between those two words is the whole design.

Session protocols. A written start sequence (sync, find ready work, claim, plan) and a close sequence ending in “work is not done until pushed”. Verdict: kept, in a different form. The start and end of a session are still the two moments I care most about, and they are now slash commands, /orient and /retro, rather than a checklist the agent might read.

Why the multi-agent part went

Three reasons, and only the third is about the tools.

The first is what running several agents did to me. The pull requests merged cleanly and I could not have rewritten the modules they touched. That failure mode has nothing to do with coordination and no lock fixes it; the fix was one agent, one pull request per session, and a rule that I can explain every merged hunk in thirty seconds. That is the May post. Once there is one agent, layers one, three and most of two have nothing to do.

The second is cost. Every parallel agent is another context loaded with the same repository, and the coordination layers add their own tokens on every turn. For a solo project the parallelism bought me speed on work I then had to slow down to understand.

The third is that the checks were in the wrong place. Skills and protocols live in prompts, and a prompt is advice. Moving the load-bearing checks into hooks made most of the coordination scaffolding unnecessary, because the thing I was coordinating against, an agent doing something wrong, is now caught at push time whether one agent or five is pushing.

What I would tell January me

Build the hooks first. Keep worktrees. Write the rules and the learnings by hand, in files the agent reads at the start of every session. Add coordination only when a second agent is actually contending for the same work, and measure that it is, because the failure modes are vivid enough to make you build for them before they happen.

And do not publish the setup guide before the verdicts. The layers were easy to describe and the verdicts took four months.

Further reading