AI GlossaryㅇWords you meet while using AI
ExploitBench
OpenAI's in-house test that scores an AI model's ability to hack and exploit security vulnerabilities
In plain words
ExploitBench is OpenAI's test for checking how good an AI model is at hacking. Like a test paper that gives a student several problems to solve and then scores them, it presents the model with a series of tasks that involve finding hidden weaknesses in simulated computer systems and actually breaking in, then tallies up how many it succeeds at.
Getting a perfect score on this test means the model can find unknown weaknesses in computer systems and carry out actual attacks without human help. OpenAI has said its new model, Astra, scored perfectly on this test — but that same capability could also be misused for dangerous purposes, which is why the score is used as a basis for assigning a safety rating.
However, this test was created and scored by OpenAI itself. It's important to note in the article that this is not a result independently verified by an outside organization.
How it shows up in the news
The article states, "OpenAI said Astra achieved a perfect score on ExploitBench, which evaluates hacking capability." A common misunderstanding here is assuming this perfect score came from an independent verification body — but it should be read with the distinction that this was actually an internal evaluation designed and conducted by OpenAI itself.
See also
Stories using this term
- Tencent's Zhuque Lab Open-Sources AI Agent/MCP Security ScannerAI · 2026.08.21
- Databricks Finds $1.2M in Annual Losses From 7 Agent BugsAI · 2026.09.03
- Hermes Agent builds its own skills the more you use itAI · 2026.08.24
- Meta Unveils First 10 Tasks in WildArtifactBench, a Benchmark for AI AgentsAI · 2026.08.21
- Perplexity's local agent beats Hermes, Pi in benchmarksAI · 2026.08.26
- Shepherd, open-source runtime for rewinding agent execution unveiledAI · 2026.08.09
