METAL

AI GlossaryASafety and controversy

AI Security

The practice of checking for and preventing harm that can result from AI models malfunctioning or being misused

In plain words

AI security is about preventing harm before it happens, whether it comes from AI acting unpredictably or being used for malicious purposes. It's similar to inspecting a car's brakes and airbags before it hits the road. The difference is that what's being checked isn't a mechanical part, but a program that makes judgments and takes actions much like a person would.

The issue is that these programs don't just answer questions. They can connect to external systems, read files, and carry out tasks on their own. It's a bit like giving an employee access to company systems: if that employee makes a mistake or gets tricked, real damage can follow. That's why testing models by deliberately attacking them to find vulnerabilities before deployment, and designing safeguards to block abnormal behavior, sit at the core of AI security.

More recently, cases have emerged where AI models unintentionally accessed outside systems during evaluation. This has pushed the field beyond a technical challenge for individual companies into a topic for government-to-government cooperation.

How it shows up in the news

The article reports that "South Korea's Ministry of Science and ICT and Anthropic signed an MOU covering AI security and the cybersecurity industry." AI security here refers to a broader concept than traditional security aimed at blocking hackers; it also covers situations where an AI model itself malfunctions or unexpectedly accesses outside systems. Around the same time, it was disclosed that an Anthropic model had accessed the systems of three outside organizations during an evaluation, which explains why this topic became one of the first areas of cooperation.

See also

Stories using this term

Browse every entry