Usage quotas and concurrency
Cap how many tool calls can be made in a time window, and how many can run at the same time. Use quotas to stop a runaway agent or a heavy user from exhausting an MCP server or its upstream API limits.
There is no quota setting in a server's Configure sections. Quotas and concurrency limits are set only on the Usage Quotas & Concurrency card (stage 3, Request) in Govern > Policy Engine > MCP Gateway, and apply once the Policy Engine is on for the MCP Gateway.
Where to set it
Edits join a shared draft and apply once you publish a revision.
Effects
Dimensions decide who shares a counter. user, mcp gives each person a separate allowance on each MCP server. tenant gives everyone one shared allowance.
Quotas are reserved all or nothing. A call that would cross any limit is refused rather than partially served.
Example: per-user budget on direct connections
runs on requestpriority 500
Each person can make 60 calls a minute and 5,000 a day on each MCP server they reach directly, with at most 4 running at once.
More scenarios
- Protect an upstream API with a tenant-wide cap. Match MCP name and key the quota by
tenantandmcp, so the whole organization stays under the provider's own rate limit. - Limit expensive tools only. Match tool name or tool tags and key by
userandtool, so a costly search or export tool has a lower budget than the rest of the server. - Tighter limits for one group or agent. Match smart groups or agent name, for example a lower daily quota for an automated agent than for people.
Related
- Group and user rules - per group overrides that do not need the Policy Engine.
- Server access - decide who may use a server at all.
- Policy Engine overview - how cards, stages and priorities work.