METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅇInfrastructure and chips

on-demand inference

A way of using an AI model where you don't sign a contract in advance — you just pay for what you use and get answers in real time.

In plain words

On-demand inference is a bit like calling a taxi. Instead of buying a car and leaving it parked, you call one when you need it, ride to your destination, and pay only for that trip. AI models work the same way: a company doesn't have to rent a dedicated server or sign a long-term contract. You just send a question, get an answer right away, and get charged based on the amount of text in that question and answer.

Pricing for this approach is usually calculated in small units based on how much text you send in and how much comes back out. That means when the company that built the model lowers its prices, everyone using that model automatically gets the cheaper rate on their next use, without having to change any settings. Unlike bulk, pre-reserved usage plans, you're billed only for what you actually use, which makes this especially well suited to services with unpredictable, up-and-down usage.

How it shows up in the news

An article might say, "Amazon Bedrock cut on-demand inference pricing for OpenAI's GPT-5.6 series models." This means the usage fee charged each time the model is used went down — not the model itself — and it's applied automatically to users without them needing to change any plan.

See also

Stories using this term

Browse every entry