AI GlossaryㅅSafety and controversy
Psychometrics
psychometrics
A set of statistical techniques—like those behind IQ tests or aptitude exams—for measuring invisible traits with scores, now also used to validate AI safety benchmarks.
In plain words
Psychometrics is the field and set of techniques used to measure things you can't see directly, like intelligence, personality, or aptitude, through test questions and scores. When schools build exams, they need to statistically check whether a question actually distinguishes stronger students from weaker ones, and whether a set of questions measures one single ability or ends up mixing several different abilities together. Psychometrics is the toolbox for exactly this kind of verification.
For example, if every single student answers a question correctly, or every single student gets it wrong, that question is useless for telling who is better at the subject. Psychometrics helps identify and filter out such useless items, and it also reveals whether several tests that look similar on the surface are actually measuring different things underneath.
Recently, these techniques have started being applied not just to people but to evaluations of AI language models. Psychometric analysis has revealed that safety benchmarks—which score how well an AI refuses dangerous requests—were often not measuring one single, unified notion of "safety," but were instead mixing together several unrelated properties.
How it shows up in the news
The article reports that "the same psychometric techniques used to build human IQ tests and aptitude exams were applied directly to analyzing language model safety benchmarks." A common misunderstanding here is that psychometrics is not a technology that makes AI smarter or safer. Rather, it's a statistical methodology for checking whether a test itself was properly constructed and what it's actually measuring.
See also
Stories using this term
- Ai2 finds BBQ safety benchmark actually measures reasoning abilityAI · 2026.09.02
- AI Safety Scores Can Be Gamed Just by Refusing MoreAI · 2026.08.22
- Tencent's Zhuque Lab Open-Sources AI Agent/MCP Security ScannerAI · 2026.08.21
- OpenAI tightens monitoring and isolation after Hugging Face incidentBusiness · 2026.08.19
- OpenAI Reverses Course, Now Pushes to Strengthen California AI Safety Law It Once OpposedBusiness · 2026.08.24
- Hume AI Measures Benchmark Memorization in Speech Recognition ModelsAI · 2026.08.22
