AI GlossaryㅅInfrastructure and chips
Scale-In
Reducing the number of running servers (instances) as traffic drops — the opposite of scale-out
In plain words
Scale-in means shrinking back the number of servers you were running in parallel once you no longer need them all. It's the same idea as opening extra checkout counters when a store gets busy, then closing them again once things quiet down.
Cloud services are often set up to automatically adjust the number of servers based on how many people are using them. During busy periods, more servers are added — that's called scale-out. When traffic quiets down, servers are removed again through scale-in, saving money on capacity that's no longer needed.
This is different from changing the performance of a single server. Scale-in and scale-out always refer to the number of servers, while adjusting a server's own performance is a separate concept.
See also
Stories using this term
- Claude Managed Agents update memory, domain controls, and session viewerAI · 2026.08.20
- Reddit developer releases 'Unswarm' to auto-switch between multiple local LLMsAI · 2026.08.23
- Cloudflare Unveils Kitesurf, a Browser Built Exclusively for AI AgentsAI · 2026.08.09
- Databricks compresses agent infrastructure into a single databaseBusiness · 2026.08.11
- Same AI Model Shows a 15x Speed Gap Across Inference ProvidersAI · 2026.08.11
- NVIDIA Details How to Build Fully Isolated K8s Tenants on Shared GPUsAI · 2026.08.09
