METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅅInfrastructure and chips

Serverless Inference

A way to use AI models by calling them on demand and paying only for what you use, without buying or managing servers yourself

In plain words

Serverless inference lets you use an AI model by calling it whenever you need it and paying only for what you actually use, instead of buying or managing the computers needed to run the model.

Think of electricity. Nobody builds their own power plant just to turn on a light at home. You plug into an outlet, and the bill reflects exactly how much power you used. Serverless inference works the same way. Instead of a company buying computers (usually high-performance chips called GPUs) to run an AI model and keeping them on 24/7, it rents just enough computing power each time a request comes in, and pays only for that usage.

The alternative is reserving a fixed chunk of computing resources in advance. For a large company that already knows its traffic will be heavy, reserving capacity ahead of time can be cheaper. But for a team just launching a service, with no way to predict when or how much it will be used, reserving capacity upfront is a real risk. Serverless inference removes that risk: you pay only for what you use, and capacity automatically scales up as traffic grows.

How it shows up in the news

The article explains that when Microsoft launched Fireworks AI models on Azure Foundry, it said billing is based by default on pay-per-token serverless inference. A common misunderstanding here is thinking the name means there are no servers at all — in reality, servers are still running behind the scenes. It's called "serverless" because users don't buy or manage those servers themselves; they simply pay for what they call.

See also

Stories using this term

Browse every entry