Microsoft Ran 1,024 Coding Agents With No Manager. 128 Times the Agents Bought 9.5 Points

September 29, 2026
9 min read

September 29, 2026
9 min read
The usual way to run many coding agents is to put one in charge: a planner splits the work and hands it out. That planner becomes the bottleneck as soon as there are more workers than it can track. A paper from Microsoft Research last week removes it entirely and scales the result to 1,024 agents. The design is worth studying. The headline numbers need reading carefully.
Everything here comes from "Agensh: Scaling Organizational Intelligence to 1,024 Agents" by Zhihao Zhan, Ting Song, Li Dong, Shaohan Huang, Jianxun Lian, Yan Xia and Furu Wei of Microsoft Research, posted to arXiv on 22 September 2026. It is a preprint and has not been peer-reviewed. The reading from section 3 onward is mine.
ProgramBench gives an agent a compiled program and asks it to rebuild the software from scratch — a complete codebase that compiles and behaves the same — within six hours, with internet access disabled. The score is the share of hidden tests the rebuilt program passes.
Agensh was tested on the five hardest of ProgramBench's 200 tasks, chosen by how poorly current models do on them. They cover multimedia processing, molecular simulation, document conversion, language interpretation and code indexing. Every worker used the same model, GPT-5.6-sol at high reasoning effort.
Mean test-pass rate across the five tasks:
| Agents | Mean test-pass rate | Change from previous row |
|---|---|---|
| 1 | 19.31% | — |
| 8 | 20.68% | +1.37 points |
| 32 | 26.52% | +5.84 points |
| 128 | 28.78% | +2.26 points |
From 1 agent to 128, that is about 9.5 points, which the authors describe as roughly a 49% relative improvement.
Only one task, the document converter pandoc, was run at the full 1,024 agents:
Eight times as many agents as the 128 run added about 4 points on the one task where it was tried.
The curve goes up. It goes up slowly, and each step costs many more agents than the one before.
The most practical number in the paper is not the final score. It is how quickly each configuration got somewhere useful. The authors report that 128 agents passed a 30% test-pass rate at the 30-minute mark. With 32 agents that took 60 minutes, and with 8 agents it took 90.
That is what parallelism reliably buys: getting to a given level sooner. The ceiling moved too, but modestly. The time to reach a working level moved a lot.
For a founder, that distinction matters. If what you are paying for is a faster turnaround on work you could otherwise wait for, more agents are a latency purchase, and you can price that. If you are hoping more agents will solve problems a single agent cannot, this paper suggests the effect exists but is small relative to the number of agents added.
The design is the part worth copying, and it is simpler than the scale suggests. There are three shared pieces:
A shared workspace. A git server. Each worker has its own checkout and branch and merges into a main branch. Git records who changed what and detects merge conflicts. If a merge is blocked, the worker pulls the latest work, resolves the conflict and merges again.
A message interface. Shared channels per task plus direct messages, delivered asynchronously with history kept, so workers can settle who owns what and untangle dependencies.
A shared context. An append-only log of typed entries — OBSERVED, FACT, FAIL, CLAIM, PATCH_SUMMARY — that any worker can search. Before starting, a worker posts a CLAIM describing the scope it is taking.
Each worker runs the same loop: read the shared context, claim a piece of work, do it, check it against that piece's acceptance criteria, merge, and publish what changed and why. There are no locks. Overlap is handled by claims and conversation.
One detail shows where the real cost of coordination sits. To reduce contention over who owns what, the authors did not start all the agents at once. They started one every 30 seconds for the first hour, then one every 3 seconds after that. Even with the whole design built around self-organization, arrival had to be managed.
The authors also describe behaviour that emerged as the group grew: at 8 agents, workers agreed an interface and built to it independently; at 128, they chose reviewers based on who had worked on related code; at 1,024, several workers took on integration, and workers picked up tasks others had abandoned.
The paper does not report what any of these runs cost: no token totals, no API spend, no compute per point gained. It gives the per-worker limits — up to 272,000 input tokens and 128,000 output tokens — and a six-hour budget, and that is all.
That matters because the scaling claim is really a price claim. "128 times the agents for 9.5 points" is only a good trade if you know what 128 times costs. Without that, the results show that more agents can help, not that they are worth it. I am not going to estimate the cost from the token limits. Workers do not necessarily use their full allowance, and a guess would look more precise than it is.
The question to ask of any agent-scaling result
Not "did the score go up?" but "how many points per dollar at each step, and how many minutes saved?" Those are the two numbers that decide whether the next doubling is worth buying.
You do not need 1,024 agents to use this design. Most of it is useful as soon as you run more than one.
Each agent writes down what it is taking before it starts. Our earlier look at fanning out five coding agents found that isolating agents is the easy part and landing their work is where they collide — one 2026 study found textual merge conflicts on 27.67% of agent pull requests. A visible claim, made before the work starts, is the cheapest way to reduce both collisions and duplicated effort.
A FAIL entry saves every later agent from repeating a dead end. It is the most valuable line in the log and the one most setups do not record.
Private branches, one main branch, conflicts resolved by the agent that caused them. You already have this infrastructure, and it already records who did what.
Agensh workers check their work against the sub-task's criteria before merging. Without them, "done" means whatever the agent decides it means.
If Microsoft needed to space out agent starts at this scale, a burst of five agents all reading the same empty task list will also collide. A short delay between starts is free.
Record how long it takes to reach "good enough" at 1, 2 and 4 agents on your own work. That, together with cost, tells you where to stop adding agents.
It is a preprint. It has not been peer-reviewed, and the results come from one team on one benchmark.
It is five tasks and one model. And only one task was run at 1,024 agents. The pattern may not hold for other tasks or other models.
There is no orchestrator baseline. The paper argues that a central orchestrator limits cooperation, but does not compare Agensh with an orchestrated system on the same tasks. We do not know from this paper whether removing the manager is what produced the gains.
No variance is reported that I could find. Whether the step from 32 to 128 agents is larger than run-to-run noise is not something I can check from what is published.
ProgramBench has an unusually clear answer key. The agents can run the reference program and compare outputs at any time. Most real software work has no such oracle, and verification is far harder. The design may scale less well where "correct" is a judgement.
The cost is unknown. Section 5 is a complaint about missing data, not a finding that the gains are too expensive.
Agensh shows that coding agents can coordinate at a scale where a central manager would not cope, using infrastructure most teams already have: git, a message channel and a shared log. That design is useful at five agents, and most of it costs nothing to adopt.
The scaling result is more modest than its headline. Going from 1 agent to 128 added about 9.5 points on the hardest tasks. Going from 128 to 1,024 added about 4 on the one task where it was tried. The clearest gain was speed. And the paper leaves out the one number that would say whether any of it is worth paying for.
Copy the claim log and the failure log this week. Before you add more agents, measure what each extra one gets you, in minutes and in money.
Source: Zhihao Zhan, Ting Song, Li Dong, Shaohan Huang, Jianxun Lian, Yan Xia and Furu Wei, "Agensh: Scaling Organizational Intelligence to 1,024 Agents", Microsoft Research, arXiv:2609.26781v1, 22 September 2026 — the ProgramBench setup, the five-task selection, the model and token limits, the 19.31%, 20.68%, 26.52% and 28.78% mean test-pass rates, the pandoc figures of 33.89%, 50.94% and 55.06%, the roughly 49% relative improvement, the 30-, 60- and 90-minute threshold times, the three infrastructure components and typed context entries, the staggered activation schedule, and the emergent behaviours are all as reported there. The point differences in section 2 are my arithmetic from those figures. The readings in sections 3 to 6 are mine. For the earlier question of how to split work across a handful of coding agents, see two ways to fan out five coding agents. For why an agent that reads everything costs more than it looks, see every file your agent reads stays on the bill.
IdeaToMVP Academy
4-week live cohort for founders. Learn to ship AI agents, scope MVPs, and automate your business — taught by the same team that writes these guides.