METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅂTechnical words in the news

Classifier

An AI program that automatically sorts input into one of several predefined categories.

In plain words

A classifier is like an automatic sorter that puts whatever comes in into one of a few boxes. Just as a post office worker glances at an envelope and quickly sorts it into 'domestic mail' or 'international mail,' a classifier looks at a piece of text or a command and decides things like 'this was written by a human' vs. 'this was written by AI,' or 'this command is safe' vs. 'this command is dangerous.'

The important thing is that a classifier doesn't know the answer for certain — it guesses based on probability. After learning patterns from a huge number of examples, it looks at a new input and decides, 'this looks similar to that pattern, so I'll put it in this box.' Because of this, it sometimes gets things wrong. That's why classifiers are usually paired with human review or other safeguards rather than being trusted on their own.

These days, classifiers show up in two opposite roles. One is as a gatekeeper that filters out sloppy AI-generated content. The other is as a brake that stops an AI program from executing a dangerous command on its own. Both rely on the same underlying principle: deciding 'pass' or 'stop.'

How it shows up in the news

In news coverage, classifiers show up in two contrasting uses. LinkedIn announced it had switched to a 'new and improved' classifier for filtering out AI-written posts, while Claude Code used a classifier to replace the step where a human had to approve every command, automatically screening out dangerous ones. A common misunderstanding is that a classifier judges as perfectly as a human would — in reality, it's making probabilistic guesses based on patterns. In the Claude Code case, the reported 89% detection rate for dangerous commands also means the remaining 11% could slip through.

Try it yourself

You can try having a chatbot act as a classifier yourself.

  1. Ask any conversational AI something like: 'Classify each of these five sentences as positive, negative, or neutral: (list the sentences).'
  2. Check the label the AI assigns to each sentence.
  3. Pick one ambiguous sentence and ask, 'Why did you give it this label?' This lets you guess at what clues the classifier used to pick that box.
  4. Ask again with the same sentence slightly reworded — the classification may change. This is a good way to feel firsthand that a classifier is making probabilistic judgments, not certain ones.

See also

Stories using this term

Browse every entry