Budgets & Usage Limits
This card lives in Policy Engine > LLM Gateway at web.quilr.ai/policy. Edits join the shared draft and take effect once you review and publish a revision.
Cap spend in USD, or cap requests and tokens, over a reset period. Count it shared across everyone in scope, or separately per person, application, key, model or credential.

How budgets work

- Every matching budget applies. Budgets do not resolve by priority. A request must fit inside all of them.
- Requests count at the start; tokens and spend count at completion. One in-flight request can carry a counter past its limit.
- Once a budget is reached, later matching requests get HTTP
429 quota_exceededuntil the period resets. - Spend needs prices. A spend budget prices provider-reported usage with the rates in Model pricing. With the Policy Engine on, this card is the only price source (prices set under Settings > Models are not used): a model with no rate under a spend budget is rejected with HTTP
503 quota_price_unavailable. - With no budget configured, spend and usage are recorded but never capped.
Sections
For a budget keyed on request metadata or an API key, use Add configuration at the bottom of the card and finish in the full editor.
Budget settings

Each budget gets a generated ID that its usage is tracked against. See The Budget ID matters before changing a published budget's measure, period or grouping.
For a Spend (USD) budget, Model pricing must cover every model the budget can reach, including fallback models from routing, for the whole budget scope. Coverage is checked when you add the budget and again before publishing.
Budget groups
A budget group is the set of requests that share one counter. Two settings shape it:
Because every matching budget applies, the tightest one stops the request first.
Model pricing settings
The Add pricing dialog first asks who the rates apply to (same scope chips as a budget), then opens the editor to set the rates.
Rates are scoped like any other setting. The simplest setup is one Everyone configuration listing every model you use.
Examples
runs on requestpriority 500
Each person gets 500 USD a month, and the whole app shares 10M tokens a rolling week. A person under their own budget is still stopped once the team cap is reached.
runs on requestpriority 600
A usage allowance needs no prices: each intern may send 200 requests a day.
Scoping and precedence
- Budgets add up; they never override each other.
- Every matching Model pricing configuration contributes its rates. Avoid pricing the same model differently in overlapping scopes.
- Budgets cannot be scoped by Prompt complexity, Prompt text or Tool.
Legacy app setting
Classic app settings have no USD budget. App-wide request and token limits are in Rate and Token Limits.