AI GlossaryㅍInfrastructure and chips
Pay-per-token
A usage-based pricing model where you're charged only for the number of text chunks (tokens) an AI processes.
In plain words
Pay-per-token means you're billed for exactly how much you use an AI service — specifically, for the number of small text chunks exchanged with the AI. These chunks are called tokens, roughly equivalent to a word or part of one.
Think of it like an electricity or water bill: you're metered for what you actually consume. It's different from a subscription where you pay a flat monthly fee for unlimited use, and different from reserving a whole server capacity that stays on all the time. In months with heavy demand, the bill goes up; when usage is light, it barely costs anything.
Startups and small dev teams often choose this model when first integrating an AI service, since there's no need to buy or reserve servers in advance — you're only charged for the requests you actually send. But once traffic grows large and becomes predictable, reserved-capacity pricing can end up being cheaper.
How it shows up in the news
The article explains that "billing is based on pay-per-token serverless inference by default." One common misunderstanding is that this model isn't always the cheapest option. For teams with stable, predictable traffic, PTU (reserved capacity) pricing can be more cost-effective — and the article notes that startup credits don't apply to the reserved-capacity option.
Try it yourself
Open the pricing page for an AI service you use or are considering. If the price is listed as a rate per million tokens, that's pay-per-token; if it's a flat monthly fee or a reserved-capacity unit, that's a different pricing model. Whether this works better for you depends on whether your usage pattern is spiky or steady.
See also
Stories using this term
- Same AI Model Shows a 15x Speed Gap Across Inference ProvidersAI · 2026.08.11
- Microsoft Foundry Opens Fireworks AI to StartupsBusiness · 2026.08.07
- 2026 Comparison of the Top 4 AI Video Generation APIsAI · 2026.08.05
- Firebird launches CIS region's largest AI factory in ArmeniaBusiness · 2026.08.08
- Firefox Smart Window to Add Real-Time Source Links to AI AnswersAI · 2026.08.19
- AWS unveils build guide for automated web insight extraction with Bedrock AgentCoreAI · 2026.08.09
