AI GlossaryㅍWords you meet while using AI
Petri
An open-source testing tool in which one AI model holds multi-turn conversations with another AI model to uncover misaligned behavior like lying or sycophancy.
In plain words
Petri is a testing tool that runs multi-turn conversations with an AI to see how badly it might behave. Instead of a human tester, another AI plays the examiner, setting up situations that try to push the target AI into lying, breaking rules, or otherwise slipping up, then watching whether it takes the bait.
What sets it apart is that it doesn't stop after a single question—it keeps talking back and forth, pressing the AI over multiple turns. Weaknesses that a single trick question from a human might miss often surface once the conversation is repeated and pushed further. That makes Petri useful as a verification step when someone claims a new method has made an AI safer, checking whether that claimed safety actually holds up.
It's known that Anthropic used this tool as one of the verification methods in an alignment research experiment involving Claude, but based on this material alone, it's hard to say for certain who originally created Petri itself.
How it shows up in the news
The article explains that "the methods Claude found held up not only on a separate validation benchmark it had never seen during the research process, but also on Petri, an open-source tool that tests for misalignment using adversarial multi-turn scenarios." Here it's worth noting that Petri isn't the solution Claude came up with—it's the separate testing ground used to double-check whether that solution actually works.
See also
Stories using this term
- Anthropic has Claude tackle AI alignment research, and it outperforms humansAI · 2026.08.31
- Playco cuts manual game-fix work 50% with AstraAI · 2026.09.04
- Anthropic funds $5M research program for AI wellbeing evaluationsAI · 2026.08.26
- Tencent's Zhuque Lab Open-Sources AI Agent/MCP Security ScannerAI · 2026.08.21
- OpenAI's Astra Nears Launch Carrying 'Critical' Cyber RatingAI · 2026.09.02
- Anthropic releases second risk reportAI · 2026.08.15
