Provider Support
Every provider type, the gateway endpoints it serves, and the credentials it needs. Provider support is the same in the V1 and V2 consoles.
You add a provider to an app in Create App or in the app's LLM Providers section. In the V2 console you can also add it once in Settings > AI Gateway > Models and link it to many apps (see Providers and Models). Your code always authenticates with a Quilr key; provider credentials never leave the gateway.
Capability matrix
One row per provider type. The type is what you select in the console and what you pass as X-Provider-Name.
- Chat on
bedrock,vertex_aiand theanthropic_messages*types is translated from OpenAI Chat Completions. See Unified Completions. bedrockalso serves native Bedrock Runtime calls from boto3.vertex_aialso serves the native Vertex AI routes.- Speech is text-to-speech and speech-to-text. Sarvam also serves translation, transliteration and language detection.
- Manual: enter Oracle model IDs yourself.
- QuilrAI-provided models back an app through a
generalprovider with base URLhttps://models.quilrai.dev/v1. Direct integration is coming soon. See QuilrAI-Provided Models.
Endpoints
Combine a regional base URL with a path below.
Credentials by provider
Every provider also has a Provider label and a Models list. Field names in code font are the API names used by the Management APIs. The auth option value goes in auth_type (Vertex AI, Oracle) or aws_auth_mode (AWS).
Rules that apply to every app:
- Every Oracle provider also needs the OCI region (
oci_region) and Generative AI project OCID (oci_project_id).oraclealso needs the compartment OCID (oci_compartment_id); it is optional fororacle_responses. Oracle does not acceptbase_url; the endpoint comes from the region. - Fields from a different auth option are rejected (for example, AWS access keys with
assume_role). Switching the auth option removes the old option's credentials. - The Base URL and Azure API version overrides on Responses, Assistants, Realtime and rerank types are in the V1 console form. In the V2 console, set them through the Management APIs.
- Every model-serving provider needs at least one model. For Azure, the model ID is the deployment name.
- Labels may use letters, numbers, spaces,
_,.and-, and must be unique within the app.primaryis reserved. - The consoles accept each provider type once per app. The API accepts several providers of the same type under different labels; select one with
provider_label, sinceprovideralone is then ambiguous. quilr_sdkandcopilot_studioapps cannot add other providers.- For a Bedrock IAM role, see AWS Bedrock - Assume Role Setup.
- Credentials not listed here (Azure Entra ID or managed identity, AWS default chain or web identity, OpenAI organization or project headers, custom upstream headers) are not supported.
What the forms look like
Chat Completions
/openai_compatible/v1/chat/completions works with OpenAI SDKs and OpenAI-compatible wrappers. It reaches providers that already support the OpenAI format (OpenAI, Azure OpenAI, Anthropic OpenAI-compatible, DeepSeek, Gemini, Oracle, Sarvam, custom endpoints) and translates for bedrock (Converse), vertex_ai (Gemini generateContent) and the anthropic_messages* types. Translation details: Unified Completions.
Sarvam serves only its chat models here. Its speech and text models use the Sarvam routes.
Anthropic Messages
/anthropic_messages/v1/messages takes the native Anthropic request shape. Use it with the Anthropic SDKs and Claude Code. Served by anthropic_messages, anthropic_messages_bedrock and anthropic_messages_azure.
AWS Bedrock Runtime (boto3)
Point a boto3 bedrock-runtime client at https://guardrails-usa-2.quilr.ai/bedrock-runtime (or your region) and sign with the Quilr key as both access key ID and secret. Paths: /model/{model_id}/converse, /converse-stream and /invoke (also under /bedrock-runtime/).
Only Bedrock Runtime is proxied, not the Bedrock control plane or Agent Runtime. See AWS Bedrock - boto3 Runtime.
Vertex AI
/vertex_ai/ is a native passthrough to generateContent, so any Gemini model on the app works, including multimodal and image-output models. Guardrails scan the text parts of the request. Image, audio and video parts, and non-text outputs, pass through unscanned.
Embeddings
/openai_compatible/v1/embeddings takes the OpenAI embeddings shape for openai, azureopenai, general and bedrock_embeddings (Titan and Cohere Embed on Bedrock).
Rerank
All three rerank paths take a Cohere-compatible body (model, query, documents, optional top_n, return_documents) and return a response in Cohere format. Guardrails scan query and documents. Responses are scores and indices, so they are not scanned. bedrock_rerank serves Cohere Rerank 3.5 and Amazon Rerank and uses the bedrock:InvokeModel permission.
TTS & STT
/openai_compatible/v1/audio/speech, /audio/transcriptions and /audio/translations work with openai, azureopenai and sarvam. Azure deployments use the /openai/deployments/{deployment}/ prefix.
On a Sarvam provider these routes adapt to Sarvam's APIs:
Transcriptions accept response_format of json, verbose_json or text, and timestamp_granularities[]=segment. model is required.
Sarvam Speech and Text
Native routes accept and return Sarvam's own formats. Auth: Authorization: Bearer, api-key or api-subscription-key, each with the Quilr key.
Chat models (sarvam-105b, sarvam-105b-conversations, and the beta glm5.2, gemma4, deepseekv4-flash) use the standard Chat Completions endpoint.
Responses API
/openai_responses/v1/responses is a native passthrough served by openai_responses, openai_responses_azure and oracle_responses. Create, retrieve, cancel, delete and list input items are supported. Azure-style paths work too: /openai_responses/openai/deployments/{deployment}/responses. The deployment goes in body.model.
Guardrails scan input_text parts and instructions on the request, and output_text on non-streaming responses. previous_response_id and built-in tools pass through.
An openai or azureopenai provider cannot serve this endpoint. Add a Responses provider type to the app.
Assistants API
/openai_assistants/ serves the OpenAI Assistants API (threads, runs, file search) through openai_assistants and openai_assistants_azure. Use it only for apps already built on Assistants.
Realtime API
wss://<base>/openai_realtime/v1/realtime is a websocket passthrough for OpenAI Realtime voice and text, served by openai_realtime and openai_realtime_azure. Aliases: /openai/v1/realtime, /openai/realtime, /openai_realtime/openai/v1/realtime, /openai_realtime/openai/realtime.
The Quilr key is accepted, in priority order, as an Authorization: Bearer header, an api-key header, an api-key or api_key query parameter, an authorization query parameter, or the openai-insecure-api-key.<key> subprotocol (stripped before forwarding).
Guardrails are not yet applied to live Realtime events. Session logging (handshake status, byte counts, usage) still runs.
Selecting a Provider on Multi-Provider Apps
When an app has several providers that can serve a request, choose one by provider type or label. Without a selector, the gateway uses the one enabled provider that has the requested model; if several have it, one is selected at random.
For apps linked to platform providers, the label is the platform provider's label.
SDK
/sdk/v1/check scans text, messages or JSON for guardrail findings without calling any LLM. Create an app with the quilr_sdk provider. Install with pip install quilrai or npm install quilrai. See SDK Mode, or try it in the LLM Gateway Playground.
Microsoft Copilot Studio
Create an app with the copilot_studio provider and register https://guardrails-usa-2.quilr.ai/copilot_studio/<Quilr key> (or your region) in Power Platform. Copilot Studio calls /validate and /analyze-tool-execution before tool execution. Block, redact and partial-redact outcomes return blockAction: true; scanning errors fail open. See Copilot Studio.