METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Google Research Unveils TEE-Based Federated Learning System

Google Research has unveiled a federated learning system that decrypts device data only inside server-side trusted execution environments and lets outsiders verify the code through a public log. A Gboard English model trained on it in three weeks with a privacy budget three times smaller.

Google Research Unveils TEE-Based Federated Learning System

Image: METAL

Summary

  • On October 2, Google Research unveiled a TEE-based federated learning system that uses trusted execution environments (TEEs) and a public transparency log to make server-side processing externally verifiable.
  • The code allowed to process the data is published to the Rekor transparency log before devices upload, and the binaries can be rebuilt from open source code on GitHub and compared.
  • In a Gboard A/B test, an English model trained on the new system matched the production model on key usability metrics with a privacy budget three times smaller, and training time fell from two months to three weeks.

Google Research on October 2 (local time) unveiled a TEE-based federated learning system that trains on device data gathered at the server while letting outsiders verify how that data is processed. Data is decrypted only inside trusted execution environments (TEEs), isolated regions within server chips, and the code permitted to touch that data is posted in advance to a public transparency log. Google has already shipped English and Japanese next-word prediction models for Gboard, its Android keyboard, using the system. According to the accompanying whitepaper, in a live A/B test an English model trained on the new system used a privacy budget three times smaller than the existing production model while holding key usability metrics such as typing speed steady, and training time fell from two months to three weeks.

Federated learning is a technique in which many devices train a model together without pooling their data in one place. Google has used it for Gboard's next-word prediction and Smart Compose, reply suggestions in Google Messages, and Smart Text Selection in Android. The sticking point was the server. Google explained in its blog post that in earlier systems neither devices nor auditors could verify the logic running on the server, so users had to trust Google to add random noise to gradient sums correctly to deliver differential privacy. Secure Aggregation, introduced later, protected uploads cryptographically but was not compatible with state-of-the-art central differential privacy guarantees.

Katharine Daly, a software engineer at Google Research, and Daniel Ramage, a research director, who wrote the blog post, said the "new TEE-based system represents the next milestone in our ongoing effort to completely remove the need to trust the server operator." The whitepaper claims the system provides externally verifiable central differential privacy guarantees for the first time. TEEs offer remote attestation, which lets third parties confirm what logic is running, confidentiality, which keeps internal state from being observed, and integrity, which keeps the logic from being disrupted, all at the level of a single machine. Google stitched these together into a federated learning system that is verifiable end to end. On the server side it uses AMD SEV-SNP and Intel TDX.

The sequence works like this. Devices encrypt training examples themselves before uploading, and beforehand they pre-authorize an access policy, the list of TEE computations allowed to process the data. Devices upload only if that access policy has been posted to Rekor, a public transparency log. The key management system (KMS) that holds the decryption keys is a cluster of TEEs bound by the RAFT consensus protocol, and it releases keys only to server-side TEE workloads that match the computations in the access policy. Training itself is run by a root TEE executing a Python training loop, which hands off parallelizable work to worker TEEs. The distributed logic is written in Federated Language, an open-source orchestration language derived from TensorFlow Federated.

According to the 14-page whitepaper reviewed by METAL, the access policy embeds both the Python program containing the training logic and the hashes of the root and worker binaries. The moment the policy is published to the transparency log, that Python program becomes open source, and outside auditors can read for themselves how it establishes differential privacy. The KMS and data processing binaries can be reproducibly built from code published in the Confidential Federated Compute repository on GitHub. Auditors can learn the full set of workloads that could run on the server just by watching the log, without owning a participating device, and workload operators see only metrics and differentially private model weights.

구글 리서치의 TEE 기반 연합학습 구조도. 기기가 KMS 키를 검증하고 데이터를 암호화해 올리면, KMS는 접근 정책이 승인한 TEE 처리 단계에만 복호화 키를 주고, 처리 단계와 KMS 바이너리는 투명성 로그에 공개돼 외부 검증자가 감시한다

Proprietary know-how is preserved through a channel called sideloading. Information that is hard to disclose, such as model architectures or data preprocessing logic, can be loaded into the Python program at runtime, on the condition that all privacy-relevant logic stays hardcoded in the program. The whitepaper notes that sideloaded information is neither encrypted nor open-sourced, but auditors can inspect how it is used in the program to judge whether the guarantees hold. Google also stated that the guarantees assume known constraints of current-generation TEEs, such as side-channel observation. Uploaded data carries a time-to-live (TTL) after which no TEE can decrypt it, and the whitepaper says the KMS enforces that TTL on a best-effort basis.

Changing the architecture changed how many devices take part. In the old system, only devices connected to Wi-Fi, plugged in and with sufficient battery could join a server-selected group called a cohort, and they had to finish their computation on the spot. According to the whitepaper, when a Japanese model was trained on the old system for 3,000 rounds over 38 days with a cohort size of 6,500, only 8.5 million of 35.5 million eligible devices, or 23.9%, actually contributed data. 20.5 million never received a task from the server, and 6.5 million received one but were interrupted before finishing. The new system first collected 17.8 million uploads over about six days before starting training, and used all 17.8 million.

Collecting all the data up front frees training from the day-night swings in device availability, letting the server set each device's participation count and spacing optimally. In an English model experiment, Google targeted a privacy budget of zCDP 0.232 at 5,000 rounds and, based on 11.8 million uploads, calculated the required noise multiplier at 5.16. Meeting the same budget on the old system would have required raising the noise multiplier to 9.54. The old system ran 8,616 rounds over 85 days with a noise multiplier of 7.38, and devices could participate only once every 144 hours. At 5,000 rounds, the maximum number of rounds a single device joined fell from eight to three, and the minimum separation between participations grew from 561 rounds to 1,822.

Real-world validation came from an A/B test on a subset of Gboard users. Gboard's typing decoder uses language models for auto-correction, word completion and next-word prediction, and each experiment arm included 3.5 million devices. The best TEE arm reached zCDP 0.215, a privacy budget three times smaller than the 0.641 of the existing production English model adjusted to the same mechanism. Key metrics such as words per minute and the rate at which users modified suggestions showed no difference. Training such a model once took one to two months, but with the bottleneck shifted to the server, round times dropped substantially even on just 14 machines, the whitepaper says.

기존 연합학습 시스템과 새 TEE 기반 시스템의 프라이버시와 유용성 곡선 비교 그래프. 잡음 배수 7.38에서 프라이버시 예산이 0.388에서 0.113으로 줄거나, 같은 예산 0.388에서 잡음 배수가 3.99로 낮아진다

Protecting personal data in isolated regions of server chips is spreading across the industry. The whitepaper cites Apple Private Cloud Compute, Google Private AI Compute and Meta Private Processing as recent examples of running LLM inference in TEEs or similar confidential computing environments. METAL previously reported on Google's disclosure of its server-side AI memory design, and this announcement extends the same approach from inference to training. Models trained on the new system so far reach up to 10 million parameters, and training much larger models will require worker TEEs to use GPUs, the whitepaper says. Google is testing synthetic data generation workloads on the same infrastructure and exploring pairing it with TEEs dedicated to LLM inference.

Through a tech-law lens, the weight of this announcement lies less in accuracy than in the burden of proof. Until now, promises from companies handling personal data rested on policy documents and internal audits, and users and regulators had little choice but to trust the company's account. By fixing in a public log, before devices upload, which code may touch the data, and by letting anyone rebuild that code to compare, Google turned a promise into a record that can be checked from outside. The whitepaper acknowledges that the pipeline operator can still see which uploads are used in each round, and notes that hiding this would allow privacy amplification by sampling to deliver stronger guarantees at the same noise level.

Daly and Ramage wrote that "this work is a step toward rigorous proof that server side processing preserves individual privacy." They said that because external verifiers can inspect exactly what code Google runs, the company can offer strong assurances that data is processed exactly as described. Google expects future TEE hardware and side-channel research to protect even dynamically loaded workloads, and anticipates that one day full proofs of correctness may accompany the implementations of differential privacy algorithms and system components. In training AI on personal data, the basis for trust is shifting from a company's declaration to a log anyone can open.

Comments