METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅅSafety and controversy

Circuit Breaking

A safeguard that automatically blocks an AI's problematic behavior once it accumulates past a certain threshold, until a human confirms it's safe to resume

In plain words

Circuit breaking is a mechanism that automatically shuts down a specific behavior once an AI has repeated a problematic action a certain number of times — and keeps it shut down until a human confirms it's safe to resume.

Think of the circuit breaker in your home's electrical panel. If the current suddenly spikes, the breaker trips on its own and cuts the power, stopping a fire before it starts. Someone has to check what caused the problem and manually flip the switch back on before power returns. Circuit breaking in AI systems works the same way. When monitoring tools flag a risky signal multiple times, the system locks itself out of that type of behavior until a human steps in and decides it's okay to continue.

It doesn't trigger on a single flagged instance — it kicks in only after flagged cases build up past some threshold, similar to how a breaker trips based on an overload level rather than any tiny fluctuation. How much of this kind of safeguard AI companies actually have in place internally isn't very visible from the outside, but a recent nonprofit evaluation used this as one of the criteria for distinguishing safety practices across companies.

How it shows up in the news

In coverage of this evaluation, circuit breaking refers to a standard clause stating that "once the number of flagged instances accumulates past a certain point, the gated behavior will be blocked entirely until a human determines it's safe and lifts the restriction." A common misunderstanding: this isn't an alarm that reacts instantly to a single anomalous act, but a blocking mechanism that activates once problem cases have piled up. Anthropic and OpenAI scored 2-3 out of 5 on this criterion, xAI scored 2, Google scored 1, and Meta scored 0.

See also

Stories using this term

Browse every entry