METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AWS adds AI traffic rate limiting to AgentCore gateway

New feature controls per-user traffic by request, token, and connection units

AWS adds AI traffic rate limiting to AgentCore gateway

Summary

  • AWS has added rate limiting to the Amazon Bedrock AgentCore gateway
  • Traffic can be controlled using three metrics: requests (RPS/RPM), tokens (TPM), and connections (CPS)
  • Granular rules can be set per user or group using JWT claims or IAM identities

AWS announced it has added rate limiting to Amazon Bedrock AgentCore gateway, its fully managed, serverless AI gateway. The AgentCore gateway serves as a single entry point for AI traffic heading to managed web search, knowledge bases, MCP servers, LLM inference models, agents (including A2A), and HTTP endpoints.

Three rate-limiting metrics

With this update, administrators can now control traffic using three metrics: request count (RPS/RPM), token throughput (TPM), and concurrent connections (CPS). Request-based limits apply to all target types, with each request counted as a single unit regardless of processing time. Token-based limits apply only to inference targets and cover both input and output tokens — the gateway deducts a preliminary estimate using a general-purpose tokenizer, then reconciles it afterward against the actual model response. Connection-based limits target long-lived sessions such as streaming, where a 100-second request occupies a connection slot for the entire duration.

Granular control by user and group

As an example, AWS presented a configuration with three user groups — Basic, Advanced, and Beta — combining JWT authentication via Microsoft Entra ID with AgentCore Identity and policy-based role access control (RBAC). In this setup, Basic users are given stricter limits while Beta users receive relatively relaxed limits for specific models, allowing an organization to validate a new model's performance and fit before rolling it out broadly. Rate limiting rules consist of "dimension keys," which group requests, and "entries," which define the allowed throughput for each group.

Comments