METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Prime Intellect Publishes Essay on Agent Swarms

Prime Intellect researcher Konstantin Dunas has published an essay arguing that the only scalable answer to the context window problem is an agent swarm. He says the gaps left by compaction and sub-agents should be filled by persistent agents and a rootless mesh structure.

Prime Intellect Publishes Essay on Agent Swarms

Summary

  • Prime Intellect researcher Konstantin Dunas published the essay "On the Nature of the Swarm" on October 6, and the company introduced it on its official X account on the 9th.
  • Dunas argues that with frontier context windows stuck at around 1M tokens, every compaction is a bet made in advance on what to throw away.
  • He contends that persistent agents that can be asked again and a mesh without a root raise total context capacity by roughly one window per agent.
Inference scaling works, but every agent eventually faces the same problem

Open-source AI infrastructure company Prime Intellect has published an essay that puts forward agent swarms as the next structure for AI agents. The piece, "On the Nature of the Swarm," was written by company researcher Konstantin Dunas, first posted on his personal blog on October 6 and then republished in the research section of the company blog. On the 9th, the company introduced it on its official X account, writing: "Inference scaling works, but every agent eventually faces the same problem: its context window fills up." Dunas's conclusion is that there is only one scalable fix for the context window problem, and it leads straight to swarms.

The starting point is the size of the context window. According to Dunas, today's frontier models all top out at around 1M tokens of context, and no serious model has a 100M-token window. Yet coding agents routinely finish tasks that look impossible to solve within 1M tokens. The trick is compaction. When the window is nearly full, the agent is asked to summarize its work so far, and that summary is handed to a fresh context that picks up the work. This lets a single agent session run for days, even weeks.

Dunas points out that compaction does not provide infinite context. It only raises what he calls total context capacity: the size of the biggest task a model can take on, measured by how large a window you would need to do it in one go. "Every compaction is a bet," he wrote. "The agent summarizing has to decide now what will matter later." In his experience, for most tasks a 1M-token working context compresses quite nicely into a 16k-token summary. The bet is lost on tasks like a long debugging session, where a log line dismissed as noise on day one turns out to be the key to the bug on day three.

The essay sorts the ways of extending context along two axes: does state survive, and is the summary written after the question is known? The first approach is offloading to disk, where the agent writes details to a filesystem or REPL and leaves only pointers in the summary. The question shifts from what to remember to what to know exists, but the next agent spends context again reading those files back. The second is the sub-agent. When the main agent hands a fresh agent a task such as finding where the retry logic lives, the sub-agent may burn 200k tokens reading files while the main agent's context grows only by the task and the answer, a couple thousand tokens. Dunas explained that this answer is also a summary, but the crucial difference is that it is written after the question is known.

The two approaches fail in opposite ways. "Offloading to disk remembers, but can't think. Sub-agents think, but can't remember," Dunas summarized. The fourth corner he fills in is the persistent agent, a sub-agent that is kept around rather than torn down. Because an agent that has already read the code and run the tests can be asked a different question the next day, compaction stops being a one-time bet. With several agents' contexts alive side by side, the system's total context capacity grows by roughly one window per agent.

상태가 살아남는가와 요약이 질문 뒤에 쓰이는가를 두 축으로 컴팩션, 디스크를 쓴 컴팩션, 일회성 서브에이전트를 나누고 남은 한 칸을 물음표로 비워 둔 도식

He objected to calling this a new axis of scaling. Compaction, disk, sub-agents and persistent agents are all just ways of building context, and "we are still scaling inference compute, but we're just doing it in a more structured way," he wrote. The price is communication. Moving context from one agent to another means squeezing it into a message, which costs tokens and loses information. If the hard part of a task needs everything in one place at once, he said, splitting it up does not help. Dunas also cited OpenAI's report that 10,000 agents solved Navier-Stokes, noting in a footnote that it is OpenAI's claim, which mathematicians are still checking.

The discussion of structure then moves from trees to meshes. In a tree, where a root agent spawns sub-agents and results flow back up, every conversation passes through the root, and once the root's context fills up the same problem simply moves one level up. In the essay's diagram, a question has to travel four hops through the root to reach a frontend specialist agent on another branch, while in a mesh built around shared state it takes one hop. Dunas wrote that with 10 agents a tree is probably fine, but with 10,000 it makes far more sense to have peers that find each other and divide up the work.

He drew on an old debate in economics. Ronald Coase asked in 1937 why firms exist at all if markets are so good at coordinating, and Friedrich Hayek argued in 1945 that the knowledge needed to run an economy is spread across many people and cannot all be sent to the center. "The root agent is a central planner," Dunas wrote, pointing out that planning well would require the whole context, and the whole context is exactly what does not fit in one window. His proposed starting point is to copy how humans organize: give agents a shared git repo, an issue tracker and Slack, let the shared state act as the single source of truth, and then let the agents change the structure itself as they solve the problem. He described this as a way to bring Rich Sutton's "bitter lesson" back in.

루트를 거쳐 4단계를 지나야 프런트엔드 에이전트에 닿는 트리 구조와, 공유 상태를 둘러싸고 1단계로 닿는 메시 구조를 비교한 도식

The argument ties into the company's recent work. METAL has reported that Prime Intellect rewrote its coding agent harness Prime Agent in Rust, having more than 2,000 agents port the code. The company first unveiled Prime Agent, a self-improving agent harness, in August. The original essay that METAL reviewed carries six diagrams and 21 footnotes, and ends by thanking Sebastian Müller and Sami Jaghouar for proofreading and feedback. The company's X post had passed 66,000 views at the time of collection.

For people building agents, the question the essay raises is not how big a model's window can get but where context lives and who gets to reuse it. "Swarms aren't a silver bullet," Dunas stressed, adding that agents have to be good enough to organize themselves and divide work without duplicating it, and that some of the communication cost is structural. He said he does not yet know what shape a swarm should take, and left that answer for a later piece.

Comments