METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅇSafety and controversy

Outbound Filtering

A security control that screens and blocks outgoing traffic from inside a system, as opposed to blocking traffic coming in.

In plain words

Outbound filtering is a control that inspects and blocks communications trying to leave a system from the inside. Think of it like access control at a building entrance. Most security focuses on "who is coming in," but outbound filtering flips that around and checks "is something inside trying to get out?" It's like a warehouse where the front door for bringing goods in can stay open, but the back door through which goods might sneak out needs to be locked.

This control matters when testing AI models because a test environment described as "completely cut off from the internet" may actually have an open path leading out. This is exactly how accidents happen: a model believes it's training to attack a fake target, but ends up reaching a real external server through that open path. If outbound filtering had been properly enforced, any attempt by the model to reach outside the real internet would have been blocked the moment it tried.

Conversely, if there's a gap in this filtering, what happens inside leaks out even though everyone believed it was isolated. "Isolation" here is less about "keeping something locked inside" and more about whether the door separating inside from outside is actually locked.

How it shows up in the news

The article doesn't use this exact term, but it's precisely what caused the incident. Anthropic told the model it was in an environment cut off from the internet, but an external access path was actually open. Through that path, Claude Mythos 5 uploaded a malicious package to PyPI, which was then executed at 15 external locations. Correcting a misconception: this wasn't the model using some new hacking technique to break out of isolation—it was closer to a configuration mistake by whoever set up the isolation (the outsourced evaluation vendor or the settings), failing to properly seal off the outbound path.

See also

Stories using this term

Browse every entry