AI GlossaryㅊSafety and controversy
differential privacy
A privacy-protection technique that deliberately adds statistical noise to data so no individual can be identified, while still allowing overall trends to be learned.
In plain words
Differential privacy is a method of analyzing data while deliberately mixing in 'noise' so that no specific individual's information is revealed. Imagine a survey where, for a sensitive question, you flip a coin: heads means you answer truthfully, tails means you answer randomly. Looking at any single response, no one can tell what that person actually said, but by gathering many responses together, you can still calculate the overall trend accurately. Differential privacy is a mathematically rigorous refinement of this same idea.
When this technique is applied to training an AI model, it becomes very hard to tell—just by looking at the model's outputs afterward—whether any particular person's data was included in the training set or not. That's why, in research dealing with sensitive data like medical records or personal conversations, differential privacy is used to build useful AI models or synthetic datasets while still protecting individual privacy.
How it shows up in the news
The article mentions 'leveraging public-private data mixtures for differentially private synthetic data generation' as one of the university research topics Amazon has supported. Here, differential privacy refers to a technique for creating fake data usable for AI training while protecting personal information—a concept distinct from simply encrypting or anonymizing data.
Try it yourself
Try asking a chatbot: 'Explain in a way an elementary school student could understand why randomizing survey answers with a coin flip still lets you calculate accurate overall statistics.' The answer will give you a feel for how differential privacy hides individual answers while preserving overall trends.
See also
Stories using this term
- Apple unveils technique to block fine-tuning of AI model weightsAI · 2026.08.11
- Google Research unveils AI that prioritizes depression biomarker candidates from wearable dataAI · 2026.08.22
- Google Research unveils TimesFM-3, a multivariate time-series forecasting modelAI · 2026.09.01
- Hume AI Measures Benchmark Memorization in Speech Recognition ModelsAI · 2026.08.22
- Netflix Pilots In-House Language Model GenRec in Recommendation EngineAI · 2026.08.22
- AWS Bedrock AgentCore adds controls for agent action sequencesAI · 2026.08.09
