arXiv:2608.23740nvidia/nemotron-3.5-lightning-30b-a3bAugust 24, 2026

AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared WorkspaceExplained for Beginners

Seonglae Cho, Donghyun Lee

Artificial IntelligenceSoftware Engineering

Abstract

Concurrent multi-agent coding promises division of labor across modules, robustness through redundancy, and parallel exploration at the natural granularity of multi-file projects. Realtime collaborative editing protocols solve this coordination problem for human teams via Conflict-free Replicated Data Types (CRDTs), but the LLMs underneath generate one token at a time and existing multi-agent coding systems inherit this serial limit: they either sequence agents through phase handoffs or pool independent samples without coordination, and a single agent abandons up to half of hard tasks with a one-file stub-and-exit. AgentRoom is a realtime collaborative editing protocol for concurrent coding agents. Its runtime layer exposes file-level claim, status, and broadcast as MCP tools on a CRDT-merged shared filesystem. Five frontier coding-CLI models ran four backend coding tasks, with cross-language checks in Python DevBench and Rust+axum. For CLI-stable models, AgentRoom with 2 agents abandons fewer tasks than Solo and has less run-to-run variation. At matched-compute, one positive mean LLM-judge contrast puts AgentRoom over parallel-merge. The other contrast, a bundle probe, puts full AgentRoom above each partial case: an ordering rather than a percentage split. Coordination, not parallelism or CRDT-merge, bears the load.

AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace

1. The Problem: When Parallelism Breaks Down

The dream of multi-agent coding is simple: throw more agents at a complex programming task, and they’ll split the work, finish faster, and produce better code. The reality, however, is messier. Existing multi-agent coding systems fundamentally treat agents as sequential entities. They either queue agents through handoff phases ("design → implement → review") or simply pool independent samples without any coordination. When multiple agents work on the same file, they end up overwriting each other’s changes, leading to compilation breaks and test failures.

The paper identifies a specific, pervasive failure mode: the "stub-and-exit" behavior. When a single agent encounters a difficult module, it often emits just one file of skeleton code and bails out early, judging the task too hard. Across the experiments, a solo agent abandoned up to 34% of hard tasks with this one-file stub.

Why should you care? If you’re a software engineering manager, this means your investment in "AI-assisted development" is vulnerable to silent failures. If you’re a product builder, it means the promised productivity gains from multi-agent systems are fragile—relying on agents to "just figure out coordination" on their own. The paper argues that we’ve been treating multi-agent coding as a problem of parallelism (getting more agents to run at once) when it’s actually a problem of explicit coordination (making agents negotiate who works on what).

2. How It Works: The AgentRoom Protocol

AgentRoom introduces a deceptively simple but powerful idea: treat the shared codebase like a real-time collaborative document, but with structured ownership rules. The system consists of three layers:

  1. The CRDT-merged shared workspace: All agents write to a unified filesystem. The underlying technology (Y.js CRDT) handles the low-level merging of character-level edits in under 2 seconds, so if two agents type on the same line simultaneously, both changes survive the merge.
  2. The MCP coordination layer: On top of the shared workspace, AgentRoom exposes five Model Context Protocol (MCP) tools that agents call explicitly:
    • room_claim(path): Atomically claims ownership of a file. If another agent already holds it, the claim is rejected.
    • room_release(path): Releases ownership when done.
    • room_read(path): Reads the current state of a file.
    • room_broadcast(message): Sends a message to all other agents.
    • room_state(): Queries peer status.
  3. The advisory collaboration protocol: Agents follow a six-step workflow: read room state, claim files, write, poll for updates, and report completion. Critically, this protocol is advisory—if an agent ignores a claim conflict and writes anyway, the system surfaces the violation rather than hard-blocking it, allowing the "cross-agent bug-fix" pattern where one agent fixes another's mistake.

The analogy: Think of this like Google Docs, but for code. In a regular shared document, you might two people type the same paragraph, and you’d end up with a messy merge. AgentRoom is more like a shared Google Doc where, before anyone types, they "claim" a section. If two people try to claim the same section, the system flags the conflict, and one person yields. The actual typing can happen concurrently because the CRDT handles the merge, but the ownership prevents the chaotic overwriting that plagues current AI coding tools.

The paper uses a concrete software engineering analogy to explain the coordination cost: without coordination, two agents writing to the same file "silently overwrite the first agent’s file," a phenomenon they call "naive concurrent ensembling." This doesn't average the two attempts; it amplifies the failures of the first agent who stubs out and exits.

3. Key Results & Benchmarks

The experiments are extensive, but the headline findings are striking and translate into plain-language impact:

  • Dramatic reduction in task abandonment: This is the most significant result. Across all models and tasks (12 model × task strata), AgentRoom with 2 agents abandons tasks at a rate 13.7 times lower than Solo. The 95% confidence interval for this odds ratio is [3.9, 48], with a p-value < 0.00001. In plain terms: if a solo agent gives up on about 30% of hard tasks with a one-file stub, AgentRoom cuts that abandonment rate to roughly 2-3%.
  • Reduced run-to-run variation: For the three CLI-stable models (Sonnet 4.6, Haiku 4.5, Codex GPT-5.4), AgentRoom reduces the standard deviation of quality scores by 30–45%. The solo Sonnet on T4 had a quality standard deviation of 0.23; AgentRoom with 2 agents brought that down to 0.14. This means the system is more reliable—you’re less likely to get a wildly different result each time you run it.
  • Quality improvement at matched compute: When comparing against a "parallel-merge" baseline (two agents working in separate copies of the codebase and then stitching the results together), AgentRoom achieves a +0.213 mean quality advantage on the LLM-judge composite (p = 0.003). The parallel-merge actually underperforms a single agent because the second agent’s changes silently overwrite the first agent’s work. AgentRoom fixes this coordination problem.
  • The coordination layer is the key driver: The paper runs a "bundle probe" that strips away components to measure their individual contribution. The results form a clear ordering: shared-only (no coordination) < prompt-only (collaboration prompt without tools) < AgentRoom (full coordination). The single largest jump comes from adding the MCP coordination tools (+0.081 quality points), beating even the collaboration prompt alone (+0.013). In other words, the CRDT merge and the prompts matter, but the explicit coordination channel (claim/broadcast tools) bears the heaviest load.

4. Why It Matters: Key Takeaways

The paper concludes with several actionable insights:

  • Coordination trumps parallelism: Simply running more agents in parallel isn't the answer. The paper shows that 3 and 4 agents actually decline in quality compared to 2, largely because the single broadcast coordination channel becomes a bottleneck. The "sweet spot" is 2 agents.
  • Explicit coordination is the differentiator: The CRDT substrate (the merge technology) is necessary but not sufficient. The real gain comes from the advisory locking protocol (room_claim, etc.). This suggests that future AI coding tools should build in explicit "who's editing what" mechanisms rather than relying on agents to negotiate via chat.
  • Cost-effective pairing: An interesting finding: pairing two Haiku (a cheaper, smaller model) AgentRooms reaches a mean quality of 0.662, which is competitive with a single Sonnet (0.544) at roughly half the cost. This hints that for many tasks, you don't need the most expensive model—you need the right coordination protocol.
  • Beyond TypeScript: The results extend to Rust + axum and Python DevBench projects, showing the approach isn't locked into one language or framework.

The bottom line: AgentRoom solves the coordination problem that has long limited multi-agent AI coding. By giving agents explicit tools to claim files and broadcast intent, it tames the "stub-and-exit" failure mode, reduces quality variation, and produces better code without needing more compute. For anyone building or using AI-assisted development tools, this paper suggests that the next frontier isn't just smarter models, but better plumbing for agents to work together.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →