METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅁWhere everyone starts

Randomized Controlled Trial

An experiment design that randomly splits participants into two groups, applying a new method to only one, so the results can be fairly compared.

In plain words

A Randomized Controlled Trial (RCT) splits people into two groups by something like a coin flip, then applies a new method to only one group, to check whether the results actually changed because of that method.

Why split people randomly? If only naturally healthier people happened to try a new method while everyone else didn't, you couldn't tell whether good results came from the method itself or just from those people being healthier to begin with. Random assignment spreads out other factors—age, health status, habits—evenly across both groups, so any remaining difference can be credited to the method itself.

The same logic applies to evaluating AI. Comparing an AI's score when it solves problems on its own against the score when real people talk to the AI to solve the same problems—using randomly split groups—reveals not just how good the AI is on its own, but what actually happens when people use it in practice.

How it shows up in the news

The article reports that in a randomized controlled trial involving 1,298 adults in the UK, performance dropped significantly when people judged their symptoms by talking with an AI, compared to when the AI solved problems on its own. It's easy to misread this as meaning the AI itself is bad at the task. But the RCT separated the AI's own capability from the outcome of people actually using that AI, and it was that gap the trial revealed.

Try it yourself

Next time you see the phrase "randomized controlled trial" in health or tech news, try checking:

  1. How many participants were there, and how many groups were they split into?
  2. What exactly was being compared (e.g., AI alone vs. a person working with AI)?
  3. Does the article explain whether the difference in results came from the method itself, or from the two groups being different to begin with?

See also

Stories using this term

Browse every entry