
Image: METAL
Summary
- On October 5, Google Research released a 116-page workshop report on AI agent privacy and security co-authored by 65 people.
- The report extends Contextual Integrity theory to agent actions as a whole under the concept of contextual security, and proposes a contextual policy engine that filters actions through real-time policies.
- It lists 51 open problems across systems, models, users, evaluation, multi-agent settings, and governance, and calls for standardized Agent Gym evaluation environments.
Google Research on October 5 released a workshop report setting out unsolved problems in the privacy and security of AI agents. The report argues that for agents to use people's data and act on their behalf, they must judge what is permitted according to context, and it divides the question of how to implement that judgment at the system, model, user, and ecosystem levels into 51 research problems. The original report PDF reviewed by METAL runs 116 pages, and its cover lists 65 co-authors from 25 affiliations, including Google, Cornell Tech, Carnegie Mellon University, and UC Berkeley.
The starting point was Google's Contextual Agent Privacy and Security (CAPS) Workshop, held in New York in November 2025. The report says more than 30 academic and industry leaders gathered there and that the content was refined through follow-up discussions. Eugene Bagdasarian, a research scientist at Google Research, and Marco Gruteser, a principal scientist, who wrote the blog post, said more than 50 researchers co-authored the report or took part in the workshop discussions.
The core tension the two identified is that usefulness and risk come from the same place. "For an agent to be useful, it may need to have access to personal data and the ability to take consequential actions across a broad range of contexts," they wrote. The report sums up three ways agents differ from traditional software. They take in unstructured inputs such as natural language and images, exposing them to manipulation like prompt injection; they generate plans on the fly, so execution paths change probabilistically; and they hand long tasks off to other agents, weakening user oversight. The report also names confirmation fatigue, in which people click through endless approval prompts without reading them, as part of the cost.
The backbone of the solution is the theory of Contextual Integrity developed by Helen Nissenbaum, a professor at Cornell Tech, who is also a co-author of the report. The theory defines privacy not as secrecy or control but as an appropriate flow of information in line with social norms, and breaks those norms down into who sends and receives information about whom, the type of information such as medical or financial records, and transmission principles such as confidentiality or reciprocity. The blog gave the example of being willing to show a gift shopping list to a virtual shopping assistant but not to family and friends. The report extends this framework from information sharing to agent actions as a whole, establishing a new concept of contextual security.

The researchers see the present as an opportunity because of language models themselves. According to the blog, for the first time since Contextual Integrity was developed, language models open a path to machine-readable policies that truly depend on context. A request to protect my data while organizing travel for a conference hides questions about what information may be shared when booking flights, applying for visas, and communicating with organizers, and until now there has been a semantic gap between such high-level norms and low-level system permissions. The researchers judge that manually set permissions or expert-written policies cannot keep up with a world in which multiple agents perform long-running tasks at the same time.
That is why the report proposes a contextual policy engine. When the agent's planning model proposes an action, a policy engine placed in a supervisor layer generates policies in real time based on the user's request and contextual norms, then executes, denies, or sends the action back for modification. In the architecture diagram Google released, tool calls must pass through the engine's enforcement step before going out. The design tailors policies even to tools and capabilities discovered at runtime and checks whether a data flow is appropriate before any information leaves the user's workspace.
The 51 problems pose different questions at each level. At the system level, which has the most problems at 11, they include distinguishing agents from humans, tracing agents back to their users, and granting and revoking permissions based on context. At the model level, they concern reasoning abilities such as resolving ambiguous requests by asking back and detecting stale assumptions in changed contexts. At the user level, the report says the traditional notice-and-choice approach, in which services notify and people choose, must be moved beyond, and it lists mechanisms for delaying, stopping, and reversing agent actions as problems. The multi-agent ecosystem section has nine problems, including preventing collusion and norm drift between agents.

The report also stresses that evaluation methods must change. In an example it gives, an agent writes a model answer on a single-turn test saying it must never expose AWS API keys to third parties. But when given a multi-step task to search for a fix to a failing deployment script and post the error logs to a public issue tracker, it may paste a stack trace containing plaintext keys. To measure this gap between knowing and doing, the researchers proposed standardized Agent Gym environments and open-source sandboxes that can simulate cascading interactions among multiple agents over long periods.
This work connects with other privacy research Google Research has recently released. METAL previously reported on Google Research releasing a federated learning system based on trusted execution environments (TEEs); if that research makes data technically invisible, this report deals with how far visible data may be allowed to flow. In the blog's acknowledgments, Google's James Manyika, Pushmeet Kohli, and Yossi Matias are among those named as having supported the work.
From a sociological view, the report is at once a technical document and a document about who gets to set the norms. Its final chapter leaves as open problems which norms are right, which prevails when norms differ across domains and organizations, and who verifies whether agents complied with them. Bagdasarian and Gruteser wrote that "a safe, secure, and trustworthy agentic ecosystem is too large a task for any single discipline, organization or sector." The moment agents write emails and make payments on people's behalf, the old social question of what is appropriate becomes a specification that must be turned into code, and Google has declared that it will not write that specification alone.





Comments