METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Prime Agent technical report shows ARC-AGI-3 score jump from 30% to 95.5%

Prime Intellect says the gain came from changing the execution shell around the model, not the model itself. The code is on GitHub.

Prime Agent technical report shows ARC-AGI-3 score jump from 30% to 95.5%

Image: METAL

Summary

  • Prime Intellect has released the full technical report for Prime Agent, its open-source agent harness. This is the August 24 revision of a document first published on August 5.
  • The report's abstract says Prime Agent lifted the ARC-AGI-3 RHAE Best@1 score from 30% to 95.5%, and matched or beat existing harnesses on long-context coding, GPU kernel generation, emulator building, and autonomous nanoGPT speedruns.
  • Its four pillars are a persistent IPython REPL, a Continual Harness that preserves history, memory, and skills, recursive sub-agents that talk to each other directly, and an Agents View that lets humans inspect sessions.

The score on the same benchmark jumped from 30% to 95.5%. According to Prime Intellect, that number didn't come from swapping the model — it came from changing the execution shell wrapped around it. The company announced on X that it had released the full technical report for Prime Agent, its open-source agent harness, and the report's abstract states that the RHAE Best@1 metric on ARC-AGI-3 rose from 30% to 95.5%.

On the left, a model is shown as a single circle, and its capability points along a dotted arrow toward a score, with a gate shaped like a broken ring sitting on that arrow. This gate represents the "harness" — previously, execution failures leaked through and dragged the score down, but now those failures are filtered out, and the score climbs in a staircase of points from 30% up to 95.5%.

Why a harness can triple a score

A harness is the execution layer that runs outside the language model when you hook it up to real-world tasks. Calling tools, reading and writing files, running code, recovering from errors, and deciding how much conversation history to retain — all of that happens here. The report's introduction defines a language model as a "bounded sequential processor": it decides its next move using only its weights and whatever information sits in its currently open context. The paper's starting premise is that the harness fills the gap between what a single computer can do and what a single model can do.

That means benchmark scores have always been a joint product of model and harness. If one tool call goes wrong and kills a session, the scoreboard just records it as a failure. The report states that Prime Agent "prevents harness failures from masquerading as model failures" (from the Prime Agent technical report's abstract). The goal is to push the measured score closer to a model's true ceiling.

Launched August 5, revised August 24

Prime Agent itself first launched on August 5. At the time, Prime Intellect described it as a harness for coding and long-running autonomous tasks, built around programmatic tool calls, context handled as variables, multi-agent messaging, and a self-repairing harness state — but no benchmark numbers came with that announcement. This new report fills that gap. The paper's cover lists an initial publication date of August 5 and a current version dated August 24.

There are 11 authors, with affiliations listed as Princeton University, Prime Intellect, and MIT. Two corresponding addresses appear: seth@primeintellect.ai and altzhang@mit.edu. The post describes the paper as an expansion of the company's blog post, putting the question of how a harness should be designed and evaluated at the center of the discussion.

Four components

According to the abstract, the architecture breaks down into four parts.

ComponentFunction
Persistent IPython REPLA Python execution environment that stays alive for the whole session. It follows the recursive language model (RLM) abstraction, treating context as a program and running test-time computation
Continual HarnessPreserves conversation history, memory, skills, prompts, and sub-agent specifications across multiple task trajectories
Recursive sub-agentsAgents collaborate by communicating directly with each other, with no intermediary
Agents ViewA screen that lets humans observe and manage sessions running as daemons

The key idea is treating context not as a block of text but as a variable handled by a program. Instead of stuffing an entire long log into the model, the REPL trims it, summarizes it, and pulls out only the pieces that are needed. The design principle is that execution, recovery, verification, and resource accounting are standardized by the harness, while strategy is left to the model.

The X post numbered its innovations into several categories — agentic context management and swarm/depth-n+ RLMs both come through fully, while a third item on Verifiers support cuts off mid-sentence. For the remaining items, the full report is the more reliable source.

What was measured — from games to kernels

The list of measured tasks is distinctive: long-context coding, GPU kernel generation, emulator building, and autonomous nanoGPT speedruns. That last one is a race to train a small GPT model to a target performance level as fast as possible, with the agent running the whole thing without human intervention. The abstract states that across these domains, Prime Agent matched or outperformed each model's native harness or other widely used harnesses.

Games were included too. In the factory-building game Factorio, the system was able to progress through the tech tree without interruption through repeated iteration, and attaching dedicated sub-agents let it split work in parallel. The choice reads as a deliberate one: multi-hour goals broken into parallel threads show up more clearly in a game environment than in a coding benchmark.

Trying it yourself

The code is open-sourced at the GitHub repository at github.com/PrimeIntellect-ai/prime-agent, the address the abstract cites as the code's location. Since a harness is a layer you can swap models into, running the same task on a coding agent you already use — changing only the harness — makes it easy to see where the difference actually comes from. A practical approach would be to clone the repo, spin up a REPL session, run a familiar task through it, and then check the Agents View to see exactly where the session stalls.

Editor's take

Whether or not you take the jump from 30% to 95.5% at face value, what this report really disturbs is the credibility of a year's worth of agent benchmark scores. When a model's performance on something like SWE-bench gets announced as a percentage, we've been reading that number as a measure of the model's raw capability. But if putting the same model inside a different shell can swing the score by that much, then what got announced wasn't a model score — it was a model-plus-harness score. That's presumably why the paper's title pins down the harness itself as the object of evaluation. Our own tracking has recently picked up a string of related efforts — attempts to quantify infrastructure noise in agentic coding evaluations, and discussions of scaling managed agents by separating "brain" from "hands." That's a signal that multiple teams are digging into the same problem at the same time.

The practical takeaway is that fixing your harness should come before swapping your model. Teams that have deployed coding agents tend to hit the same wall. It's rarely that the model doesn't know the answer — tasks stall because a tool call goes wrong, a long log eats up the context window, or there's no way to revive a session that died midway. Switch to a more expensive model in that state and you'll likely see token costs double or triple with little improvement in completion rate. The four things Prime Agent calls out for standardization — execution, recovery, verification, and resource accounting — target exactly that failure point. Resource accounting in particular, tracking how many tokens each sub-agent burns, doesn't show up in demos, but it's exactly what determines the monthly bill in production.

The part that's immediately useful for teams here is the Continual Harness concept. Right now, most agent pipelines are built so that once a task ends, the conversation history and whatever was learned along the way disappear with it. A design that carries history, memory, skills, prompts, and sub-agent specifications across tasks pays off a lot when it's applied to recurring internal work — if you're generating the same kind of report every week, your instructions get shorter starting from the second week. Swarms, on the other hand, still feel early. Running multiple sub-agents in parallel works well in environments like Factorio, where the goal is clear and verification is automatic, but in work where verification depends entirely on human judgment, parallelizing just shifts the burden onto review.

In the coming weeks, expect to see evaluations that hold the harness fixed while comparing models, alongside evaluations that hold the model fixed while comparing harnesses. And it's likely that model announcements will start routinely specifying which harness was used to produce a given score. Whether Prime Agent becomes the standard is a separate question — but any announcement that reports a score without naming the shell around it now reads as incomplete.

Comments