AI GlossaryㅇInfrastructure and chips
on-demand inference
A way of using an AI model where you don't sign a contract in advance — you just pay for what you use and get answers in real time.
In plain words
On-demand inference is a bit like calling a taxi. Instead of buying a car and leaving it parked, you call one when you need it, ride to your destination, and pay only for that trip. AI models work the same way: a company doesn't have to rent a dedicated server or sign a long-term contract. You just send a question, get an answer right away, and get charged based on the amount of text in that question and answer.
Pricing for this approach is usually calculated in small units based on how much text you send in and how much comes back out. That means when the company that built the model lowers its prices, everyone using that model automatically gets the cheaper rate on their next use, without having to change any settings. Unlike bulk, pre-reserved usage plans, you're billed only for what you actually use, which makes this especially well suited to services with unpredictable, up-and-down usage.
How it shows up in the news
An article might say, "Amazon Bedrock cut on-demand inference pricing for OpenAI's GPT-5.6 series models." This means the usage fee charged each time the model is used went down — not the model itself — and it's applied automatically to users without them needing to change any plan.
See also
Stories using this term
- GPT-5.6 model pricing on Bedrock cut by up to 80%AI · 2026.08.09
- NVIDIA to Guarantee Up to 25% of Its Own Chips' Value, Raising $500 BillionBusiness · 2026.08.12
- NVIDIA halves guarantee for OpenAI data centerBusiness · 2026.08.16
- NVIDIA's rumored $12.9 billion Hugging Face deal, and the math behind an 86x revenue multipleBusiness · 2026.08.27
- Cognition adopts Fable 5.1 in Devin, cuts coding task costs 54%AI · 2026.09.02
- Starcloud raises additional $250M for orbital data centersBusiness · 2026.08.22
