METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅎTechnical words in the news

Mixture of Experts Architecture

An AI design approach that, instead of using the entire model, activates only a subset of specialized "expert" components suited to the task at hand.

In plain words

Mixture of Experts architecture is easiest to understand through a hospital analogy. A general hospital has dozens of specialists, but when a patient arrives, not every doctor examines them at once. The front desk looks at the symptoms and calls in only the specialists who are actually needed. AI models work the same way: they contain many small "expert" components internally, and when a query comes in, only some of them are selected to take part in the computation.

The reason for this is to grow the model's size while keeping computation costs down. Even if the model's total size (total parameters) reaches into the trillions, the portion that actually moves to produce a single answer (active parameters) is only a fraction of that. This makes it possible to keep the knowledge and capability of a large model while cutting the speed or cost of generating answers down to something closer to a much smaller model.

However, a larger architecture doesn't always mean better performance or lower cost. The actual results depend heavily on how well the right experts are selected and how well those experts have been trained.

How it shows up in the news

The article describes Alibaba's Qwen3.8 Max as having "a Mixture of Experts (MoE) structure with 2.4T (2.4 trillion) total parameters." A common misunderstanding here is that a larger total parameter count doesn't automatically mean the model is smarter or cheaper. In fact, the same article notes that Qwen3.8 Max fell behind the open-source model Kimi K3 in both benchmark scores and cost per task. This illustrates that size (total parameters) and actual performance or efficiency are separate matters.

Try it yourself

Next time you come across news about a new LLM, look for two numbers together: total parameters and active parameters. If the two figures are listed separately, that means the model uses a Mixture of Experts architecture. The bigger the gap between them, the more you can understand it as a model designed to be large in size but low in actual computation cost.

See also

Stories using this term

Browse every entry