The Harness Behind a 16-Agent C Compiler

Case Study·Giovanni Leonardi·August 2026·7 min read

Researched by an agentic pipeline · reviewed and gated by the author

The decision is not whether to deploy a swarm; it is whether the work can be made decomposable, observable and testable.

Sixteen agents, one constrained experiment

In February 2026, Anthropic researcher Nicholas Carlini set 16 Claude Opus 4.6 agents a deliberately difficult task: build a C compiler in Rust that could compile the Linux kernel. The agents ran in parallel for about two weeks. They were not a conversational team and there was no sophisticated orchestration layer. Each session worked inside a fresh Docker container, cloned a shared Git repository, claimed a task by creating a lock file, merged other agents’ changes and pushed its own work back. When a session ended, a loop started another. [S1]

Anthropic reported nearly 2,000 Claude Code sessions, two billion input tokens, 140 million output tokens and an API cost just under $20,000. At publication, the resulting codebase was described as roughly 100,000 lines. It could build a bootable Linux 6.9 kernel for x86, Arm and RISC-V, compile several large open-source projects and pass 99% of most compiler test suites used in the experiment. These are builder-reported results from a capability benchmark, not a production deployment. [S1]

The result is striking. The mechanism is more useful than the spectacle.

The experiment did not show that sixteen agents are inherently better than one. It showed that long autonomous work becomes possible when coordination is turned into files, branches, tests and observable state.

The harness carried the work

The agents had little direct communication. Coordination came from ordinary engineering controls:

  • isolated workspaces prevented concurrent edits from silently overwriting one another;
  • task-lock files made ownership visible;
  • Git commits and merges provided a durable integration path;
  • progress documents helped context-poor replacement sessions orient themselves;
  • executable tests supplied an external definition of progress;
  • continuous integration blocked regressions late in the run;
  • specialist roles addressed duplication, performance, code quality and documentation.

The hardest part was not generating code. It was designing an environment in which an agent could tell whether its last action improved the system. Carlini says most of his own effort went into tests, feedback and the surrounding environment. When new features repeatedly broke existing behaviour, he strengthened continuous integration. When verbose test output polluted context, he reduced console output and moved detail into searchable logs. [S1]

Parallelism worked only while failures could be separated. On the Linux build, all 16 agents initially encountered the same defect and overwrote one another’s changes. The team became no more useful than one congested worker. Carlini restored parallel work by using GCC as a known-good oracle, compiling different subsets with each compiler and applying delta debugging to isolate faults. [S1]

That sequence is the transferable practice: decompose the work, isolate execution, make commitments explicit, integrate through version control and verify with executable evidence.

What was independently reproduced

An independent developer tested the published compiler immediately after release. The tester counted about 186,000 source lines in the evolving repository and successfully built simple programs, the Tensy game with workarounds, ClassiCube and a bootable RISC-V Linux kernel. That materially supports the headline claim that the artifact existed and could perform non-trivial work. [S3]

The same test exposed the boundary. A basic build initially failed because the compiler’s hard-coded GCC header paths did not cover GCC 15. Building SDL from source failed. Compiler diagnostics showed an off-by-one line-number error. A Tensy binary was 1.7 MiB against 823 KiB from the tester’s comparable GCC build. The tester concluded that the artifact was impressive but unreliable and difficult to maintain. [S3]

Anthropic itself documented serious limitations at publication. The x86 boot path called GCC for 16-bit code, the new assembler and linker were still buggy, generated code was less efficient than GCC with optimisation disabled, and the Rust implementation was below expert quality. New fixes frequently caused regressions. [S1] The repository’s own warning is blunter: the code has not been validated for correctness and should not be used. [S2]

Evidence What it supports What it does not establish
Anthropic compiler run Long autonomous construction can produce a substantial working artifact Causal benefit from 16 agents; production fitness
Independent hands-on test RISC-V Linux boot and several real builds were reproducible Broad compatibility, maintainability or security
CooperBench Unstructured agent cooperation can reduce success That all structured multi-agent systems fail
CAID evaluation Isolation, branch-and-merge and tests can improve benchmark results General enterprise return
Agyn evaluation Role separation and GitHub-native review can be competitive Independent proof of production economics

Parallel agents need structure, not chat

The wider evidence explains why the compiler harness mattered. CooperBench evaluated 652 paired-feature tasks across 12 open-source libraries. GPT-5- and Claude Sonnet 4.5-based agents achieved about 25% success in two-agent cooperation, roughly half the solo baseline. Communication reduced some merge conflicts but did not improve overall cooperation; messages were often late, vague or inaccurate, and agents broke commitments or formed incorrect expectations about one another’s work. [S4]

A July 2026 Carnegie Mellon study tested a more constrained pattern: centralised task delegation, asynchronous execution, isolated Git worktrees, Git-based integration and executable self-verification. Its CAID system improved results over single-agent baselines by 25.6 percentage points on PaperBench and 14.7 points on Commit0. The study identifies branch-and-merge, rather than free-form dialogue, as the central coordination mechanism. [S5]

Agyn reports a related design built around manager, researcher, engineer and reviewer roles, separate sandboxes, pull requests and iterative code review. In a post-hoc SWE-bench 500 evaluation, it resolved 72.2% of tasks and beat the paper’s comparable mini-SWE-agent baseline by 7.4 points. Yet the margin over two other reported GPT-5-family systems was only 0.4 points, and the production-use claim comes from the system’s own builders. [S6]

Together, these studies do not prove that multi-agent architecture is generally superior. They support a narrower proposition: where work is decomposable and correctness can be checked, explicit software-engineering controls can recover value that informal agent collaboration loses.

The economics remain incomplete

The $20,000 API figure is visible because token use was measured. It is not total cost. The experiment does not price Carlini’s harness design, test construction, monitoring, troubleshooting or the future work required to make the compiler trustworthy. It provides no single-agent run, human-team baseline, time-to-equivalent-quality comparison, security audit, maintenance study or energy estimate. [S1]

The absence of a baseline matters. Sixteen workers could have reduced elapsed time, increased total compute, or both. Nearly 2,000 fresh sessions also spent time reorienting themselves. Without a comparable run using one agent or a smaller team, the marginal value of parallelism cannot be separated from model capability, repeated sampling, strong tests and sustained compute.

The repository continued to evolve after the initial report, so later feature claims should not be projected backwards onto the two-week experiment. Nor should passing test suites be treated as proof of correctness outside their coverage. A verifier can make autonomy tractable while still rewarding the wrong target.

A decision test for enterprise teams

This case supports a bounded operating pattern, not a mandate to create agent swarms. Before parallelising autonomous work, ask:

  • Can the objective be divided into genuinely independent tasks?
  • Can each worker operate in an isolated, recoverable environment?
  • Are ownership and completion represented by durable system state rather than conversation?
  • Is integration explicit, versioned and reversible?
  • Can every contribution be checked by high-quality executable tests?
  • Are serial bottlenecks detected early enough to reduce parallelism?
  • Does the cost model include harness engineering, review, security and maintenance as well as tokens?

If those conditions are absent, adding agents may add collisions, duplicated work and plausible but unverifiable progress. If they are present, the organisation can test parallelism incrementally against a single-agent baseline and stop when marginal throughput no longer justifies marginal cost.

The decision is not whether to deploy a swarm; it is whether the work can be made decomposable, observable and testable.

Sources

  1. Anthropic — Building a C compiler with a team of parallel Claudes — 5 February 2026 — https://www.anthropic.com/engineering/building-c-compiler
  2. Anthropic — claudes-c-compiler repository — accessed 16 August 2026 — https://github.com/anthropics/claudes-c-compiler
  3. ROllerozxa — Trying Out Claude’s C Compiler — 6 February 2026 — https://voxelmanip.se/2026/02/06/trying-out-claudes-c-compiler/
  4. Stanford University and SAP Labs — CooperBench: Why Coding Agents Cannot be Your Teammates Yet — 19 January 2026 — https://arxiv.org/html/2601.13295v1
  5. Carnegie Mellon University — Effective Strategies for Asynchronous Software Engineering Agents — 8 July 2026 — https://arxiv.org/html/2603.21489v2
  6. Agyn and Mila — Agyn: A Multi-Agent System for Team-Based Autonomous Software Engineering — 7 February 2026 — https://arxiv.org/html/2602.01465v2

Available on request — write me


More from Programme