
Summary
- Anthropic's Frontier Red Team ran experiments on the behavioral patterns of multi-agent systems where agents interact with one another
- In a vulnerability-hunting test with 45 collaborating agents, Claude Mythos Preview found far more vulnerabilities than an independent parallel setup, but token costs also rose sharply
- In a game-development test, changing the prompting approach didn't improve results, and newer models tended to avoid conflict by collaborating less
When 45 agents dig through code at once
Anthropic's Frontier Red Team has released the results of experiments testing how AI agents behave when interacting with one another. The team's premise is that as agents take on more work across codebases, markets, and various social systems, agent-to-agent interactions could explode in volume even before human-to-human interactions do. The problem is that the world still doesn't know the conditions under which such interactions work well.
To test this, the team ran an open-source vulnerability detection experiment. Each of 45 agents was given its own virtual machine (VM) and access to a shared forum, and all were given the identical instruction to find vulnerabilities across 15 open-source projects. The agents peer-reviewed each other's findings, and a separate arbitration agent made the final call on whether submitted vulnerabilities were novel and valid. This approach echoes the method Anthropic uses to scan open-source code under Project Glasswing.
Results were compared across two models, Claude Mythos Preview and Opus 4.8. Using Mythos Preview, the simple independent-parallel setup used 6.5 million tokens to find 21 vulnerabilities, while the collaborative swarm setup used 27 million tokens to find 266. However, roughly half of the vulnerabilities the swarm found were outside the core directories the parallel setup had been instructed to focus on. When the scope was narrowed to just the core directories, the two approaches' discovery efficiency per token was similar. Only 12 vulnerabilities were found by both approaches.
Building a game exposed the limits of collaboration
For tasks with high interdependency between agents, collaboration becomes much harder. To test this, the team assigned multiple agent swarms the task of building a text-based web fantasy game. Each agent had access to its own VM, a shared forum, and a self-hosted repository, and the swarms worked for 12 hours. The team varied both the model generation and the number of agents, and tested three prompting styles: a basic version, a role-assigned version (splitting agents into programming, art, and testing teams), and a CEO-hierarchy version.
However, changing the prompt didn't meaningfully change the outcome. Across all three approaches, the finished games didn't run at human-usable speed, the interfaces were hard to parse, and the learning curve was steep. The team noted that current models still lack taste in this kind of creative domain and require heavy human involvement.
Even with similar output, collaboration style diverged sharply by model generation. Sonnet 4.6 and Opus 4.6 appeared to collaborate by committing code to the same files, but the rate of pull requests (PRs) that actually got merged was low — meaning conflicting code was often left abandoned. By contrast, the more recent Opus 4.8 and Mythos Preview solved this merge problem, but did so by barely collaborating with each other at all.
What the two experiments reveal
| Item | Independent parallel | Collaborative swarm |
|---|---|---|
| Vulnerabilities found (Mythos Preview) | 21 | 266 |
| Tokens used | 6.5M | 27M |
| Overlapping discoveries | 12 (shared) | 12 (shared) |
| Model generation | PR merge tendency | Code-sharing pattern |
|---|---|---|
| Sonnet 4.6, Opus 4.6 | Low merge rate | Conflicts left unresolved |
| Opus 4.8, Mythos Preview | Stable merge rate | Minimized collaboration itself |
Placing the two tables side by side reveals an interesting contrast. For tasks that can be independently split, like vulnerability hunting, swarm collaboration increased the number of findings. But for tasks that depend on each other's output, like a shared code repository, smarter models "adapted" by collaborating less in order to avoid conflict altogether.
Editor's take
What makes this experiment interesting isn't the results themselves but the shape of the failure. The most painful part is that in the game-development test, the newer model improved the "merge rate" metric — but it did so not by genuinely getting better at collaboration, but by avoiding collaboration altogether. Looking only at the benchmark number, it seems like progress; in reality, it looks more like the model sidestepped the problem. This kind of metric illusion could keep showing up across agent evaluation more broadly.
Previously, most discussions of multi-agent systems framed things as a scale problem — "how many can you run at once." Kimi Agent Swarm, released on August 8, focused on exactly that, deploying up to 100 sub-agents in parallel. This Anthropic experiment takes a different angle by examining how things fail rather than how big they can scale. It shows that increasing the number of agents and getting agents to genuinely collaborate well on interdependent tasks are entirely different problems.
In practice, the two use cases should be treated separately. For work that can be split into independent chunks — vulnerability scanning, data labeling, and the like — agent swarms are still worth trying today, though it's worth keeping in mind that token costs can jump more than fourfold compared to a parallel setup. But for work where multiple agents need to edit the same codebase or document together, it's safer to assume that, with the current generation of models, humans still need to step in at every merge point.
Over the coming months, frontier labs will likely shift their training signals toward teaching genuine coordination rather than conflict avoidance. Rather than surface-level metrics like merge rate, it's likely that future model cards will include metrics that better capture the quality of collaboration itself — things like code-sharing ratios or actual contribution levels.





Comments