Skip to main content

Rate, Token & Timeout Limits

V2 console

This card lives in Policy Engine > LLM Gateway at web.quilr.ai/policy. Edits join the shared draft and take effect once you review and publish a revision.

Cap concurrency, request rate, token volume, tokens per request and provider timeout, for a scope or per model. Limits are enforced at the gateway before the provider call.

Rate, Token & Timeout Limits card collapsed with the Application limits and Per-model limits summary rows

How limits combine​

Limits do not resolve by priority. For each limit, the strictest matching value wins, so a narrower row can tighten a broader one but never loosen it. Window limits compare per second, so 1,000 per minute is stricter than 100,000 per day.

Rate, Token & Timeout Limits card expanded with an application limits row, an empty Per-model limits section and Add configuration

Sections​

SectionWhat it doesEmpty state
Application limitsConcurrency, request rate, token windows, tokens per request and provider timeout for a scope. Add limits.No limit.
Per-model limitsThe same limits keyed by model, or by provider credential and model. Add model limit (scope first, then the editor).No model carries its own limit.

Both apply on assistants, bedrock, chat, embeddings, realtime, rerank, responses, stt, text, tts and vertex. For limits keyed on an API key, request metadata or content size, use Add configuration and finish in the full editor.

Limit settings​

Switch on only the limits you want. A limit left off keeps the value from a broader configuration.

New application limits dialog with all six limits switched on: Concurrency, Requests per minute, Input tokens per hour, Output tokens per hour, Tokens per request and Timeout

SettingUnitWindow optionsNotes
Applies to--Everyone, People, Smart group, Application, App tag, Requested model, Provider, API surface, Environment, Source network, Except... Default Everyone.
ConcurrencyRequests in flight-
RequestsRequestsper minute (default), per hour, per day
Input tokensTokensper minute, per hour (default), per day
Output tokensTokensper minute, per hour (default), per day
Tokens per requestInput tokens-Ceiling for one request.
TimeoutSeconds-How long the gateway waits for the provider response.
Severity--Not set (default) to Very critical. Reported only.

Examples​

support_copilot_limitsrequest

runs on requestpriority 500

WhenApplicationisSupport Copilot
Then
Concurrency limit50
Rate limit1,000 per minute
Input token limit100,000 per day
Output token limit200,000 per day
Tokens per request80,000
Timeout30s
contractors_tighter_raterequest

runs on requestpriority 600

WhenApplicationisSupport Copilot
andSmart groupsincludes (ignoring case)Contractors
Then
Rate limit100 per minute

Contractors get 100 requests a minute; everyone else in the app keeps 1,000. A narrower row asking for 5,000 per minute would have no effect, because the stricter value always wins.

Scoping and precedence​

  • Strictest value wins per limit, regardless of priority.
  • Limits cannot be scoped by Prompt complexity, Prompt text or Tool.
  • For USD caps or allowances that reset on a calendar, use Budgets & Usage Limits.

Legacy app setting​

Rate and Token Limits in the app's settings.