
Summary
- On October 9, Prime Intellect released Prime Agent rewritten from the ground up in Rust, a port that Prime Agent itself completed in two weeks by orchestrating more than 2,000 agents.
- The work used more than 10,000 sandboxes and over 200 billion GLM-5.3 tokens, while humans were responsible for setting up four kinds of parity checks.
- The Rust version reaches usable input in 51.9 milliseconds, 14.18 times faster than the TypeScript version, and uses 4.76 times less memory on large sessions.
Open-source AI research company Prime Intellect on October 9 released a version of its coding agent harness Prime Agent rewritten from scratch in Rust instead of TypeScript. The rewrite itself was handled by Prime Agent. According to the company, Prime Agent orchestrated more than 2,000 agents over two weeks to port its own code end to end, using more than 10,000 Prime Sandboxes and over 200 billion tokens from the GLM-5.3 endpoint of its inference service, Prime Inference. The company's X account said the agents exchanged 16,000 messages with one another.
The result shows up as speed. In benchmarks published by the company, the Rust version took 51.9 milliseconds from a fresh launch to accepting input, 14.18 times faster than the TypeScript version's 736.1 milliseconds. On a repeat launch it took 41.4 milliseconds, 13.34 times faster, and whole-process memory after opening a 10 MiB session fell 4.76 times, from 1,130.0 MB to 237.3 MB. Installed size dropped from 172.1 MB to 59.6 MB, and switching to the agents view went from 78.0 milliseconds to 12.8 milliseconds.
Prime Intellect launched Prime Agent in August and says it has since been downloaded more than 300,000 times and processed over 8 trillion tokens. METAL has previously reported that Prime Intellect unveiled Prime Agent, a self-improving agent harness. The company cited several reasons for moving to Rust: TypeScript's types are optional and disappear at runtime, CPU-heavy work like rendering and parsing competes with keyboard input on a single event loop, and every process pays for a JavaScript runtime and garbage collector.
The human work was building verification. The company set up four parity checks: a TUI check comparing the terminal frames both versions render against the same scripted model, a harness check comparing session transcripts and the requests sent to the model provider from the same inputs, a protocol check matching every message type in the daemon protocol, and a feature check in which agents audited features one by one and classified them as matching, partial or missing. The company said objective checks let agents measure their own progress and catch regressions before merging, which allowed it to cut back on human review.
The orchestration structure was simple. A single root agent divided the work into a dependency order; it wrote no product code and only monitored tasks and merged finished work. Each task moved through four agents: a planner writing the specification and verification criteria, an implementer writing Rust code in a separate worktree, a reviewer inspecting the change adversarially using a different model in a separate context from the implementer, and a verifier compiling and running the checks in a fresh sandbox. "An agent that writes code is biased when evaluating it," the company wrote in its blog, explaining why generation and verification were split across agents. The orchestrators and subagents ran on two 8-core on-demand CPU nodes, each handling more than 100 concurrent subagents.
After parity was reached, a second orchestrator ran a loop for three days to improve the benchmarks. The loop logged more than 144 experiment and audit records and merged more than 69 changes that improved performance. No numeric targets were given on purpose. "A fixed threshold tends to become a stopping point," the company wrote. Each change went in only after two reviewer agents, each on a different frontier model, confirmed its behavior still matched the TypeScript version. Most of the gains came from three kinds of change: moving work off the startup and render paths, replacing polling loops with event-driven waits, and releasing memory as soon as large sessions finished loading.

Comparisons with other harnesses were also published. In the company's measurements on a 4-core, 8 GB sandbox, the Rust version showed its first visible output in 23.6 milliseconds, faster than Claude Code v2.1.289 at 264.6 milliseconds, Codex CLI v0.160.0 at 296.8 milliseconds, Pi v1.0.3 at 306.3 milliseconds and Hermes Agent v0.21.5 at 1,715.3 milliseconds. Whole-process memory after startup was 106.0 MB, lower than Claude Code's 226.9 MB and Codex CLI's 344.4 MB. Measurements used a scripted model and exclude inference time, and the company added that without a common benchmark standard, comparisons with external harnesses should be interpreted with caution.
The code structure changed too. The code is split into nine crates with a one-way dependency graph, and the largest source file is now about 2,500 lines, down from about 15,000 in TypeScript. The TypeScript version had four files over 5,000 lines; the Rust version has none. Each session runs in its own worker process, so one session crashing leaves the others running. This release also adds native Windows support in beta and installation through Homebrew, and Prime Agent remains open source. The original blog post METAL reviewed states that bringing the autonomous work to release quality took several weeks of follow-up, and that in this phase humans found problems, directed changes and reviewed all results.

The weight of this announcement lies less in the speed figures than in the method. What made it possible for more than 2,000 agents to port a product to another language in two weeks was not only model capability but four objective checks that humans set up in advance. Prime Intellect said it will open the multi-agent workflows behind the rewrite to users. The case shows what development organizations need first if they want to hand large code migrations to agent swarms: checks the agents can use to grade themselves.





Comments