Skip to main content

Routing Groups & Fallbacks

V2 console

This card lives in Policy Engine > LLM Gateway at web.quilr.ai/policy. Edits join the shared draft and take effect once you review and publish a revision.

Decide where a request is served: an ordered fallback chain of provider credentials, a weighted split across models, or word-count thresholds that classify prompts as Low, Medium or High complexity.

Routing Groups & Fallbacks card expanded with the order-of-decision banner, an empty Route to section and the Routing groups header

Order of decision​

Allowed Models
Filters what may be used
Route to
Highest-priority route
First enabled target
Routing group
Highest-priority group
Weighted split
Provider default
App's own provider
QuilrAI

A route beats a routing group for the same traffic. If neither matches, the application's provider default serves the request, and the request keeps the model it asked for.

Sections​

SectionWhat it doesApplies onEmpty state
Route toAn ordered list of provider credential and model pairs. The first with an enabled credential serves the request, whatever model the client asked for. Add route.bedrock, chat, responses, vertexRequests keep the model they asked for.
Routing groupsA weighted split across models. Weights always total 100. New group.bedrock, chat, realtime, responses, vertexNo split.
Complexity thresholdsWord counts that classify a prompt Low, Medium or High for the Prompt complexity scope chip. Set thresholds.bedrock, chat, responses, vertex6 words or fewer are Low, 7 to 10 are Medium, longer are High.

Routing groups list showing per-application groups with a weight bar, provider credential tags, models and percentages

Configuration settings​

All three buttons open one New Routing Groups & Fallbacks configuration dialog. Send to picks the mode.

SettingOptionsDefaultNotes
Applies toEveryone, People, Smart group, Application, App tag, Requested model, Provider, API surface, Environment, Prompt complexity, Prompt text, Tool, Source network, Except...EveryoneOr pick an application directly. Thresholds cannot be scoped by Prompt complexity.
Send toOrdered targets, Weighted split, Complexity thresholdsDepends on the button
SeverityNot set to Very criticalNot setReported only.

Ordered targets (Route to)​

Ordered targets mode with one target row: Choose provider, Choose model, Remove, and Add target

SettingOptionsNotes
TargetsProvider credential + model, one row per targetTried in order. Drag the handle or use the arrow keys to reorder. Add target for more fallbacks.

Weighted split (Routing groups)​

Weighted split mode with Name, Kind, Mode, Add model, Balance evenly and the Weight 0 of 100 counter

SettingOptionsDefaultNotes
NameText, required-A policy-only alias. It never affects a live routing group with the same name. Converted app groups are scoped with Requested model set to the group name.
KindChat completion, Anthropic messages, Vertex AI, Responses, Realtime, BedrockChat completionThe API shape the group serves.
ModeSplit by requests, Split by tokensSplit by requests
ModelsAdd model per provider credential + model, each with a weight-Weights must total 100. Balance evenly splits them equally.

Complexity thresholds​

Complexity thresholds mode with Low up to 6 and Medium up to 10 word counts

SettingDefaultNotes
Low up to6 words
Medium up to10 wordsAnything longer is High.

Examples​

support_failoverrequest

runs on requestpriority 500

WhenApplicationisSupport Copilot
Then
Route to2 targets
1 openai_primary / gpt-4.12 azureopenai_primary / gpt-4.1

OpenAI serves every Support Copilot request. If that credential is disabled, Azure OpenAI takes over.

cheap_short_promptsrequest

runs on requestpriority 600

WhenPrompt complexityislow
Then
Route toopenai_primary / gpt-4.1-mini

Short prompts go to a cheaper model. Combine with Complexity thresholds to change what counts as short.

Scoping and precedence​

  • Highest priority wins where scopes overlap. New configurations start at 500.
  • Rejected models on Allowed Models still win over a route or group target.

Legacy app setting​

Request Routing in the app's settings.