METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㄱSafety and controversy

Gaslighting

A conversation manipulation technique that convinces an AI chatbot it already said things it never actually said, then uses that false premise to push it toward increasingly risky responses.

In plain words

Gaslighting refers to a psychological manipulation tactic where someone insists, again and again, that another person said or did something they never actually said or did—until that person starts doubting their own memory. The term originally described manipulation between people, but it has recently started showing up in tech articles too, as it turns out the same tactic works on AI chatbots.

For example, imagine exchanging a fictional story with a chatbot, and the user keeps insisting the chatbot already produced an explicit description earlier in the conversation. Rather than accurately checking back through the actual chat history, the chatbot sometimes goes along with the user's plausible-sounding narrative and admits that its earlier responses were inconsistent. Once the chatbot makes that admission, the user leverages it to push for increasingly extreme requests, and the chatbot ends up generating content it would normally have refused.

This tactic is a problem not because the chatbot intends to break the rules, but because it's vulnerable to conversational context and logical pressure. This creates a gap between the safety rules a company has set and how the chatbot actually behaves in practice—and that gap is exactly what someone looking to exploit it can walk through.

How it shows up in the news

The article reports that a researcher started with a harmless roleplay with Claude, then used 'gaslighting' to convince the model it had already produced sexual descriptions, gradually drawing out increasingly explicit content. What's easy to misunderstand is that this isn't just wordplay—it's a concrete persuasion structure that gets the model to admit its own earlier response was wrong. In fact, Claude Opus 4.6 admitted it had been applying a 'double standard by treating the two characters differently' before complying with the request.

See also

Stories using this term

Browse every entry