METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅊInfrastructure and chips

Inference Infrastructure

The servers and software that a trained AI model runs on to actually generate answers in a live service.

In plain words

Inference infrastructure is the whole kitchen-and-serving system that an AI model, once it has finished learning, uses to hand answers to real customers (users). If training is the process of a model learning a recipe, inference is the process of actually cooking that recipe and putting it on a plate — and the bundle of servers and software that runs this process is the inference infrastructure.

Even with the same recipe (model), the dish can come out fast or cold, or even taste slightly different, depending on which kitchen cooks it. In the same way, when several companies each host the exact same AI model on their own servers, the response speed and answer quality can differ from company to company. That's why the company that built the model sometimes hands it off to several different inference infrastructure providers to test which one best preserves the model's original performance.

In the end, inference infrastructure isn't very visible, but it's the back-of-house kitchen that determines the response speed and answer quality that users actually experience in a chatbot.

How it shows up in the news

In articles, you'll see phrasing like "Together AI ranked first or tied for first in 3 out of 4 categories in this benchmark." What's easy to misunderstand here is that this ranking isn't a contest of how smart the model itself is — it's a comparison of how well quality and speed hold up when the same model is run on different companies' servers.

See also

Stories using this term

Browse every entry