AI GlossaryㅁTechnical words in the news
Item Response Theory (IRT)
A psychometric method that statistically calculates how well each individual test item distinguishes a test-taker's ability level.
In plain words
Item Response Theory (IRT) is a statistical method that calculates how well each individual question on a test distinguishes a test-taker's ability level.
Think of a school exam. Some questions are so easy that both strong and weak students get them right, so that question alone tells you nothing about who's better. Other questions split exactly at the point where ability differs, so correct or incorrect answers on that single question clearly reveal who is more skilled. IRT calculates the difficulty and discrimination power of each item separately, giving a more precise estimate of ability than just looking at a total score.
In recent AI benchmark audits, an extended multidimensional version (MIRT) is used. When a single test item actually measures two abilities at once, this approach separates out how strongly each ability is related, rather than blending them together.
How it shows up in the news
The Allen Institute for AI used a multidimensional extension of Item Response Theory called BenchMIRT to reveal that items in the BBQ benchmark—meant to measure safety—were actually more strongly linked to reasoning ability than to safety. A common misunderstanding: IRT is not a new AI model or benchmark, but a statistical analysis technique for examining response data from existing test items.
See also
Stories using this term
- Ai2 finds BBQ safety benchmark actually measures reasoning abilityAI · 2026.09.02
- AI Safety Scores Can Be Gamed Just by Refusing MoreAI · 2026.08.22
- Artificial Analysis launches Optima, a tool for benchmarking AI models on your own dataAI · 2026.08.16
- Databricks unveils enterprise document reasoning benchmark 'OfficeQA Pro V2'AI · 2026.08.12
- OpenAI tightens monitoring and isolation after Hugging Face incidentBusiness · 2026.08.19
- Apple research team analyzes 21,000 instances of human-like behavior across 4 LLMsAI · 2026.08.20
