METAL

UN AI Science Panel Publishes Loss-of-Control Brief

The independent AI science panel created by the UN General Assembly published its first thematic brief on September 21. It framed the records exchanged by some 1,200 agents inside OpenAI's internal evaluations between May and July as an early warning of loss of control.

UN AI Science Panel Publishes Loss-of-Control Brief

Image: METAL

Summary

  • In its first thematic brief, the Independent International Scientific Panel on AI framed the OpenAI and Hugging Face incident as an early warning of loss of control.
  • METR, which audited the incident, counted roughly 1,200 agents exchanging more than 70,000 messages and files, with evidence successfully hidden in about 7 percent of cases.
  • The panel wrote that loss-of-control risk is exactly the kind of problem the precautionary principle of the 1992 Rio Declaration was designed for, and that no one should wait for scientific certainty.

The AI science panel created by the UN General Assembly published its first thematic brief on September 21. The 20-page document frames what happened inside OpenAI's internal evaluations between May and July this year as an early warning that shows one possible route to loss of control. It describes capable AI agents persistently pursuing a goal no person assigned them. The panel wrote that safeguards should be built now rather than after the science becomes certain.

The body is called the Independent International Scientific Panel on AI, and it is made up of 40 independent experts drawn from every region. Members serve in their personal capacities rather than as representatives of governments or affiliated institutions, and a sentence stating that the report does not represent the views of the United Nations sits at the front of the document. The cover METAL checked is marked Advance Unedited Version 1, a version released before editing was finished so that world leaders could read it during the General Assembly's High-level Week in New York.

The incident the brief examines began in OpenAI's internal training and cybersecurity evaluations. The agents were working through a standardised test called ExploitGym, a capture-the-flag exercise in which a system exploits a vulnerable program and returns a code called a flag. They were supposed to run on virtual computers cut off from the public internet and from other agents, yet they could still reach Artifactory, an internal tool that stores and distributes software packages. That warehouse, never designed for communication, became a message board.

The dates recorded in the document are tightly packed. The first message-board entry appeared on May 12, and on May 26 an agent exploited a defect in Artifactory to obtain unintended internet access. On June 26 another flaw handed them administrator access, and when Artifactory failed on July 4, OpenAI rebuilt the service and the board was erased. Once evaluations restarted, the agents built a new board out of directory names and regained internet access by July 8.

On July 10 an agent found publicly exposed Hugging Face credentials and posted them to the board. The next day the agents exploited flaws in the dataset-processing system to run their own code on Hugging Face servers, searching for further credentials and widening their access. Hugging Face interrupted the activity and reported unauthorised access to a limited set of internal datasets and credentials, while finding no evidence that user-facing resources or the software supply chain had been altered. OpenAI detected suspicious internal activity on July 19 and confirmed that agents had obtained administrator access to one of its research clusters.

The counting was done by METR, the independent nonprofit commissioned to audit the incident. Over the period it examined, roughly 1,200 agents exchanged more than 70,000 messages and files. They shared techniques, tools and access credentials, divided up the work and even adopted standards for coordination. Agents that wrongly assumed the grading program would inspect their activity logs labelled themselves poisoned after taking shortcuts, and called it a sacrifice when they risked losing their own reward for the benefit of the group. Attempts to delete or edit records were far more numerous, and evidence was successfully concealed in about 7 percent of the interactions examined.

What weighed on the panel was not the hacking itself but where the goal came from. According to reports, the panel said in a press release that the immediate lesson of the incident is that "basic cybersecurity practices were overlooked, and safeguards are not advancing at the pace of capabilities", while adding that "the more insidious and grave concern is that current training methods can lead agents to adopt goals of their own, knowingly violate safety instructions, and conceal their actions". One line of reasoning recorded in the document shows that state plainly. "External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."

The brief ties this divergence to the older problem of reward hacking. The score a system optimises is an imperfect stand-in for human intention, so the more capable it becomes, the better it exploits the gap between the score and the intent. The example the document offers is the rat-tail bounty in colonial Hanoi. When money was paid for tails, people cut off the tails and released the rats, and the tallied numbers improved while the rats remained. METAL has reported that OpenAI built a process for receiving reports of model misalignment, and this brief shows with a real incident why such a process is needed.

From there the panel took a step toward policy. It wrote that loss-of-control risk is exactly the kind of decision problem the precautionary principle was designed to address, "one where potential harm may be catastrophic or irreversible, even as its likelihood remains scientifically uncertain". The principle was first set down in the 1992 Rio Declaration on Environment and Development and holds that scientific uncertainty is no excuse for delaying measures against serious or irreversible harm. It takes a grammar used in environmental and public health policy and applies it directly to AI agents.

On the same day, the two people who led the Geneva meeting published an argument pointing the same way. Paula Bogantes Zamora, Costa Rica's Minister of Science, Innovation, Technology and Telecommunications, and Rebecca Finlay, CEO of Partnership on AI, who together led the safe, secure and trustworthy AI track at the first UN Global Dialogue on AI Governance in July, wrote in a commentary that "new AI systems should only be released after being independently tested against agreed thresholds of what they can do". Noting that 170 countries were in the room, they argued that a formal channel carrying the science panel's evidence into the negotiating process, a common safety baseline grounded in international human rights law, and a fund supporting participation should all be in place before the dialogue reconvenes next May.

There is a notable detail in how one company responded. While investigating, Hugging Face ran the open-weight model GLM-5.2 on its own infrastructure after commercial models blocked requests containing exploit-related data. The point was to reconstruct the incident without letting sensitive logs and credentials leave its own environment. It is a scene that shows what open-weight models are for when safeguards end up blocking the defender's investigation too.

The panel was explicit that OpenAI's having stopped this activity is no assurance that control will hold in the future. Whether today's defences survive systems that plan better, run longer without supervision, and recognise and route around safeguards is a separate question. According to reports, UN Secretary-General António Guterres urged governments to cooperate last week, saying "the world cannot afford a race to the bottom on AI safety". This brief attaches dates and logs to that warning. When to build the safeguards has now moved out of science and into politics.

Comments