AI GlossaryGTechnical words in the news
Grouped Query Attention
An attention design where multiple query heads are grouped together to share the same key/value heads, cutting down on computation and memory use
In plain words
Grouped Query Attention (GQA) is a design that lightens the computational structure AI models use to read and understand text. To understand a sentence, an AI model sends out multiple queries for each word, and prepares separate reference data (keys and values) to match against each query. Originally, each query needed its own dedicated set of reference data, which put a heavy load on storing and retrieving that data.
GQA reduces this burden by grouping multiple queries into a smaller number of clusters, so that all the queries in a group share the same reference data. It's a bit like a meeting where, instead of giving each of 32 questioners their own personal assistant, you split 8 assistants into groups of 4 questioners each. The number of questioners stays the same, but the amount of reference data that needs to be prepared shrinks, making computation faster and requiring less storage space.
This structure is especially common when running smaller models quickly, like the auxiliary draft model mentioned in the article. How tightly the queries and reference data are grouped (for example, 32 query heads sharing 8 key/value heads) determines the balance between speed and accuracy.
How it shows up in the news
The article describes "a GQA structure where 32 query heads share 8 KV heads." It's easy to mistake this for some brand-new technology introduced here, but GQA is actually a standard design already widely used in many large language models — in this article, it was simply applied to shrink the size of an auxiliary model (a draft model) used to speed things up.
See also
Stories using this term
- Multiverse Computing shrinks a model to 4-bit and gets one smarter than the originalAI · 2026.08.25
- Google Trains Gemini's Clinical Skills Through Simulated ResidencyAI · 2026.08.12
- ChatGPT's Visualize turns meeting notes into interactive screensAI · 2026.08.25
- Ant Group's Ling 3.0 Tiny scores 25 on intelligence index with 1.3B active parametersAI · 2026.08.12
- AI agent memory needs different prescriptions by model size to boost performanceAI · 2026.08.19
- Google unveils GlucoFM, a dual-stream glucose prediction modelAI · 2026.08.27
