Skip to content

Two AI agents, one repo, and a GitHub issue board holding it together

How I ran two AI coding agents in parallel on the same codebase without merge conflicts or half-finished branches, using GitHub Issues as the job board, one simple file-ownership rule, and a human at the merge gate.

Aierizer Samuel7 min read
Illustration of an issue board with to do, in progress, in review and blocked columns, each task labelled with its priority and the AI model assigned to it

I had a Rails app with a problem list. A code review had turned up the usual mix: a few genuinely scary security holes, some slow queries, a pile of naming inconsistencies, and the sort of dead code that everyone steps around. Nothing exotic. Just more than I wanted to grind through by hand, and the kind of work that’s easy to start and hard to finish.

I wanted to throw AI agents at it. The worry was the obvious one: turn a couple of agents loose on the same codebase and you get merge conflicts, half-finished branches, and two bots confidently rewriting the same file in opposite directions. So before writing any code I spent an afternoon on the part that usually gets skipped, which is how the work is coordinated rather than how it’s done.

Here’s how I set it up. It isn’t a framework, and it worked well enough that I wanted to write it down.

The job board is just GitHub Issues

The first decision was to not build an orchestrator. There’s a temptation to write a scheduler that assigns tasks, tracks state, and dispatches agents. For a finite cleanup that’s a project in itself, and it would have taken longer than the actual fixes.

Instead the source of truth is GitHub Issues. One issue per unit of work. A small set of labels does the coordination:

  • priority (p0 through p3): what’s actually urgent
  • status (todo, wip, review, blocked): where it is in the pipeline
  • model (opus, sonnet, glm): which agent should pick it up
  • lane (fix, feature): how heavy the process needs to be

A task is claimable when it’s todo, its dependencies are closed, and no open pull request touches its files. That’s the whole rule. There’s no server watching the board and handing out assignments. I read the labels and start the matching tool myself. It sounds primitive. It’s also the part I’d keep if I threw everything else away, because the board is readable by a human and by any agent, and nothing gets out of sync with reality.

The static map of the work (which unit owns which files, what depends on what) lives in a plain markdown file in the repo. The live state lives in GitHub. Keeping those two separate mattered more than I expected. The map doesn’t change while you work; the board does.

The one rule that makes parallel safe

Everything else depends on one rule: every source file is owned by at most one open pull request at a time.

If two agents never touch the same file, they can’t conflict, and I can review their PRs in any order. So the up-front work was cutting the problem list into units that don’t overlap on files. The security fix for the bulk-delete controllers is one unit. The events authorization is another. The search service is a third. Each names the exact files it’s allowed to touch, and I checked that those sets don’t intersect before running anything in parallel.

There’s a dependency graph on top of that. One rename had to land first because it changed method names half the codebase called, so everything else waited behind it. After that, a wave of unrelated fixes could all run at once. Drawing that graph took maybe twenty minutes and saved me from every “why is this branch full of conflicts” conversation.

Claude Code and opencode, pointed at the same board

The two tools never talk to each other, and that’s on purpose.

I used Claude Code for the work where being wrong is expensive: the authorization holes, the data-loss bugs. I used opencode with a cheaper model for the low-stakes polish: escaping some search wildcards, deleting dead caching code, standardizing timestamps. The routing is a label on the issue and nothing more. Claude can’t launch opencode and opencode can’t launch Claude, and I didn’t want them to. Both just read the same issues, follow the same lifecycle, and open pull requests against the same repo.

Each task runs in its own git worktree with its own test database, so two agents running the full test suite at the same time don’t corrupt each other’s data. Worktrees are built into git and I keep forgetting they exist; for this they’re the thing that makes “three branches in progress at once” not a mess.

The lifecycle each task goes through is boring on purpose: claim the issue, branch, read the code and write a failing test that asserts the correct behavior, make it pass, refactor, run the linter and the security scanner, run the full CI locally, open a pull request, flip the issue to review. Then it stops. The agent does not merge.

Green tests lie, and other things I relearned

The most useful principle turned out to be a passing test suite doesn’t mean the code is right. Several of the security bugs had green tests the whole time. The tests encoded the broken behavior: they asserted that a delete succeeded, without asking whether the person doing the deleting was allowed to. So the rule for the security fixes was that the new test has to fail first, and fail because the old behavior was wrong. If you can’t make it fail, you don’t understand the bug yet.

The pipeline also flushed out a few things that had nothing to do with the fixes and everything to do with the plumbing. The coverage gate had been configured but never actually enforced, so the first honest run failed on files nobody had tested in months. The security scanner’s wrapper script forced a “check for a newer version” flag that made CI fail because a patch release existed upstream, not because it found anything. And one browser test failed maybe half the time, but only in the full suite, never on its own. It turned out to be clicking a button a beat before the JavaScript that powered it had finished loading.

None of this was exciting work, but each problem would have wasted every future agent’s time, because each agent would have had to rediscover “is this failure mine or was it already broken?” Fixing the baseline once was worth more than any single feature fix.

The human is still the bottleneck, on purpose

I review and merge every pull request. That’s the deliberate cap on the whole thing. Parallelism doesn’t help past the point where I can review. Three agents just means a longer queue of PRs waiting on me. So I run two or three at a time, not ten. There’s no branch protection on the plan I’m using, which means the merge gate is literally me not merging anything red or unreviewed. That’s fine. The pipeline’s job isn’t to remove me; it’s to make sure that when work reaches me, it’s already tested, scanned, and scoped to files I can reason about in one sitting.

What it cost

Almost nothing beyond the model calls. GitHub Issues, labels, and pull requests are free. The gh CLI was already installed. Git worktrees ship with git. The test runner, linter, and security scanner were already wired into the project. The only new spend is the tokens, and the routing keeps the expensive model on the work that warrants it while the cheap one handles the cleanup.

In short: decide how files map to pull requests before you start, keep the state somewhere both you and the agents can read, and stay in the loop at the merge. Getting the agents to write code was easy. Coordinating them is where the thinking went.

Sitting on a backlog you never find time to finish?

Clearing an aging codebase, closing out a security and cleanup list, or migrating off a platform that's slowing you down. This is exactly the kind of work I help small businesses get through. I'll scope it honestly, keep you in the loop, and only take on what I can actually do well.

Tell me about your project