AI GlossaryRTechnical words in the news
RHAE Best@1
A scoring standard that grades an AI agent's single best attempt at a task while filtering out attempts that failed only because of execution-environment errors, not the agent's own reasoning.
In plain words
RHAE Best@1 is a scoring method that counts the score from a single attempt at a problem as the final result. It's a bit like diving judges scoring only a diver's actual performance — if the diving board breaks and the diver falls because of that, it isn't counted as their fault and they get to try again. This metric works the same way: it separates cases where the agent genuinely failed to solve a problem from cases where the attempt itself was voided because a tool call crashed the program.
This metric made news because a score on the same benchmark jumped from 30% to 95.5%. The model doing the actual reasoning stayed the same — only the execution environment wrapping the model, which handles tool calls and error recovery, was changed. Yet the score more than tripled. Under the old execution environment, many cases where the model actually reasoned correctly were recorded as failures simply because the program crashed midway. RHAE Best@1 is a scoring standard meant to filter out that kind of noise and show a number closer to the model's true capability.
So when looking at this number, two things need to be checked together: which benchmark was used, and which execution environment scored it. This case shows that changing the execution environment alone can produce completely different scores for the same model.
How it shows up in the news
An article reports that "ARC-AGI-3's RHAE Best@1 score rose from 30% to 95.5%." The easy mistake here is reading this as the model itself getting smarter. As Prime Intellect explains, what actually changed was the execution environment wrapping the model, and RHAE Best@1 is the number that reflects the reduction in attempts that were wrongly recorded as failures once that environment was replaced.
See also
Stories using this term
- Prime Agent technical report shows ARC-AGI-3 score jump from 30% to 95.5%AI · 2026.08.27
- Prime Intellect unveils self-improving agent harness 'Prime Agent'AI · 2026.08.09
- Kimi launches 100-agent parallel swarm systemAI · 2026.08.08
- AWS Adds Open-Source Agent Skills for Bedrock Automated Reasoning PoliciesAI · 2026.08.09
- AWS adds AI traffic rate limiting to AgentCore gatewayAI · 2026.08.09
- Job seeker sends ChatGPT to face an AI recruiter instead of himselfAI · 2026.09.03
