Skip to main content

Provider Support

Every provider type, the gateway endpoints it serves, and the credentials it needs. Provider support is the same in the V1 and V2 consoles.

You add a provider to an app in Create App or in the app's LLM Providers section. In the V2 console you can also add it once in Settings > AI Gateway > Models and link it to many apps (see Providers and Models). Your code always authenticates with a Quilr key; provider credentials never leave the gateway.

Capability matrix​

One row per provider type. The type is what you select in the console and what you pass as X-Provider-Name.

Provider typeProviderChatMessagesResponsesAssistantsRealtimeEmbeddingsRerankSpeechList models
openaiOpenAI✓----✓-✓✓
openai_responsesOpenAI--✓-----✓
openai_assistantsOpenAI---✓----✓
openai_realtimeOpenAI----✓---✓
azureopenaiAzure OpenAI✓----✓-✓✓
openai_responses_azureAzure OpenAI--✓-----✓
openai_assistants_azureAzure OpenAI---✓----✓
openai_realtime_azureAzure OpenAI----✓---✓
anthropicAnthropic (OpenAI-compatible)✓-------✓
anthropic_messagesAnthropic✓✓------✓
anthropic_messages_bedrockAnthropic on AWS Bedrock✓✓------✓
anthropic_messages_azureAnthropic on Azure AI Foundry✓✓------✓
bedrockAWS Bedrock (Converse, boto3 runtime)✓-------✓
bedrock_embeddingsAWS Bedrock-----✓--✓
bedrock_rerankAWS Bedrock------✓-✓
vertex_aiGoogle Vertex AI✓-------✓
gemini_chatcompletionsGoogle Gemini API (OpenAI-compatible)✓-------✓
deepseekDeepSeek✓-------✓
oracleOracle OCI Generative AI✓-------Manual
oracle_responsesOracle OCI Generative AI--✓-----Manual
sarvamSarvam✓------✓✓
cohere_rerankCohere------✓-✓
jina_rerankJina------✓-✓
voyage_rerankVoyage------✓-✓
generalCustom endpoint (vLLM, Ollama, LiteLLM, any OpenAI-compatible URL)✓----✓--✓
general_rerankCustom endpoint with a Cohere-shaped /rerank------✓-✓
quilr_sdkQuilrAI SDK (guardrails only, no upstream)---------
copilot_studioMicrosoft Copilot Studio (guardrails only)---------
  • Chat on bedrock, vertex_ai and the anthropic_messages* types is translated from OpenAI Chat Completions. See Unified Completions.
  • bedrock also serves native Bedrock Runtime calls from boto3. vertex_ai also serves the native Vertex AI routes.
  • Speech is text-to-speech and speech-to-text. Sarvam also serves translation, transliteration and language detection.
  • Manual: enter Oracle model IDs yourself.
  • QuilrAI-provided models back an app through a general provider with base URL https://models.quilrai.dev/v1. Direct integration is coming soon. See QuilrAI-Provided Models.

Endpoints​

Combine a regional base URL with a path below.

SurfacePathAuth
Chat Completions/openai_compatible/v1/chat/completionsAuthorization: Bearer <Quilr key>
Embeddings/openai_compatible/v1/embeddingsAuthorization: Bearer <Quilr key>
Speech/openai_compatible/v1/audio/speech, /audio/transcriptions, /audio/translationsAuthorization: Bearer <Quilr key>
Anthropic Messages/anthropic_messages/v1/messagesx-api-key: <Quilr key>
Responses/openai_responses/v1/responsesAuthorization: Bearer <Quilr key>
Assistants/openai_assistants/Authorization: Bearer <Quilr key>
Realtimewss://<base>/openai_realtime/v1/realtimeAuthorization: Bearer <Quilr key>
Bedrock Runtime (boto3)/bedrock-runtime/model/{model_id}/converse and related operationsAWS SigV4 with the Quilr key
Vertex AI/vertex_ai/Authorization: Bearer <Quilr key>
Rerank/rerank/v2/rerank, /rerank/v1/rerank, /rerank/rerankAuthorization: Bearer <Quilr key>
Sarvam native/sarvam/...Authorization: Bearer <Quilr key>
QuilrAI SDK/sdk/v1/checkAuthorization: Bearer <Quilr key>
Copilot Studio/copilot_studio/{Quilr key}Quilr key in the path

Credentials by provider​

Every provider also has a Provider label and a Models list. Field names in code font are the API names used by the Management APIs. The auth option value goes in auth_type (Vertex AI, Oracle) or aws_auth_mode (AWS).

Provider typeAuth optionRequiredOptionalWhere to set
openai, anthropic_messages, sarvamAPI keyAPI key (api_key)-Console, API
openai_responses, openai_assistants, openai_realtimeAPI keyAPI keyBase URL (base_url, default OpenAI)Console, API
anthropic, gemini_chatcompletions, deepseekAPI keyAPI key-Console, API
cohere_rerank, jina_rerank, voyage_rerankAPI keyAPI keyBase URL (default: the vendor's API)Console, API
azureopenaiAPI keyAPI key, Azure endpoint (azure_endpoint)Azure API version (azure_api_version, default 2024-10-21; the consoles require it)Console, API
openai_responses_azure, openai_assistants_azure, openai_realtime_azureAPI keyAPI key, Azure endpointAzure API version (defaults 2025-03-01-preview, 2024-05-01-preview, 2025-04-01-preview)Console, API
anthropic_messages_azureAPI keyAPI key, Base URL (the Azure AI Foundry base URL)Anthropic version (anthropic_version, default 2023-06-01)Console, API
general, general_rerankAPI keyAPI key, Base URL-Console, API
bedrock, anthropic_messages_bedrock, bedrock_embeddings, bedrock_rerankStatic credentials (static, default)AWS access key (aws_access_key), AWS secret key (aws_secret_key)AWS region (aws_region, default us-east-1), AWS session token (aws_session_token)Console, API
sameAssume role (assume_role)Role ARN (aws_role_arn), External ID (aws_external_id)AWS region, Role session name (aws_role_session_name), Session duration (aws_session_duration_seconds, 900 to 43200, default 3600)Console, API
vertex_aiAPI key (api_key, default)API key, GCP project ID (gcp_project_id)GCP region (gcp_region, default us-central1)Console, API
sameService account (service_account)Service account JSON (service_account_json)GCP project ID (the consoles require it; the API falls back to the JSON's project_id), GCP regionConsole, API
sameExpress (express)API key (Gemini API key; calls generativelanguage.googleapis.com, no project or region)-Management API only
sameApplication Default Credentials (adc)GCP project ID. Uses the gateway's own Google identity.GCP regionManagement API only
oracle, oracle_responsesAPI key (api_key, default)API key-Console, API
sameGateway user principal (gateway_user_principal)No secret. See Oracle OCI - Gateway Sign-In Setup.-Console, API
sameUser principal (user_principal)Tenancy OCID (oci_tenancy_id), User OCID (oci_user_id), Key fingerprint (oci_fingerprint), Private key (oci_private_key)Private key passphrase (oci_private_key_passphrase)Console, API
sameSession principal (session_principal)Session token (oci_session_token), Private keyPrivate key passphraseConsole, API
sameInstance principal (instance_principal), Resource principal (resource_principal)No secret-Console, API
quilr_sdk, copilot_studioNoneLabel only-Console, API

Rules that apply to every app:

  • Every Oracle provider also needs the OCI region (oci_region) and Generative AI project OCID (oci_project_id). oracle also needs the compartment OCID (oci_compartment_id); it is optional for oracle_responses. Oracle does not accept base_url; the endpoint comes from the region.
  • Fields from a different auth option are rejected (for example, AWS access keys with assume_role). Switching the auth option removes the old option's credentials.
  • The Base URL and Azure API version overrides on Responses, Assistants, Realtime and rerank types are in the V1 console form. In the V2 console, set them through the Management APIs.
  • Every model-serving provider needs at least one model. For Azure, the model ID is the deployment name.
  • Labels may use letters, numbers, spaces, _, . and -, and must be unique within the app. primary is reserved.
  • The consoles accept each provider type once per app. The API accepts several providers of the same type under different labels; select one with provider_label, since provider alone is then ambiguous.
  • quilr_sdk and copilot_studio apps cannot add other providers.
  • For a Bedrock IAM role, see AWS Bedrock - Assume Role Setup.
  • Credentials not listed here (Azure Entra ID or managed identity, AWS default chain or web identity, OpenAI organization or project headers, custom upstream headers) are not supported.

What the forms look like​

Chat Completions​

/openai_compatible/v1/chat/completions works with OpenAI SDKs and OpenAI-compatible wrappers. It reaches providers that already support the OpenAI format (OpenAI, Azure OpenAI, Anthropic OpenAI-compatible, DeepSeek, Gemini, Oracle, Sarvam, custom endpoints) and translates for bedrock (Converse), vertex_ai (Gemini generateContent) and the anthropic_messages* types. Translation details: Unified Completions.

Sarvam serves only its chat models here. Its speech and text models use the Sarvam routes.

Anthropic Messages​

/anthropic_messages/v1/messages takes the native Anthropic request shape. Use it with the Anthropic SDKs and Claude Code. Served by anthropic_messages, anthropic_messages_bedrock and anthropic_messages_azure.

AWS Bedrock Runtime (boto3)​

Point a boto3 bedrock-runtime client at https://guardrails-usa-2.quilr.ai/bedrock-runtime (or your region) and sign with the Quilr key as both access key ID and secret. Paths: /model/{model_id}/converse, /converse-stream and /invoke (also under /bedrock-runtime/).

OperationCoverage
converseAny selected model that supports Converse. Request and response guardrails.
converse_streamRequest guardrails; the event stream passes through unchanged.
invoke_modelAmazon Nova, Anthropic and OpenAI-style Bedrock models. Request and response guardrails.
invoke_model_with_response_streamReturns ValidationException.

Only Bedrock Runtime is proxied, not the Bedrock control plane or Agent Runtime. See AWS Bedrock - boto3 Runtime.

Vertex AI​

/vertex_ai/ is a native passthrough to generateContent, so any Gemini model on the app works, including multimodal and image-output models. Guardrails scan the text parts of the request. Image, audio and video parts, and non-text outputs, pass through unscanned.

Embeddings​

/openai_compatible/v1/embeddings takes the OpenAI embeddings shape for openai, azureopenai, general and bedrock_embeddings (Titan and Cohere Embed on Bedrock).

Rerank​

All three rerank paths take a Cohere-compatible body (model, query, documents, optional top_n, return_documents) and return a response in Cohere format. Guardrails scan query and documents. Responses are scores and indices, so they are not scanned. bedrock_rerank serves Cohere Rerank 3.5 and Amazon Rerank and uses the bedrock:InvokeModel permission.

TTS & STT​

/openai_compatible/v1/audio/speech, /audio/transcriptions and /audio/translations work with openai, azureopenai and sarvam. Azure deployments use the /openai/deployments/{deployment}/ prefix.

On a Sarvam provider these routes adapt to Sarvam's APIs:

OpenAI fieldSarvam field
inputtext
voicespeaker
speedpace
response_formatoutput_audio_codec (default MP3; pcm is raw linear16)

Transcriptions accept response_format of json, verbose_json or text, and timestamp_granularities[]=segment. model is required.

Sarvam Speech and Text​

Native routes accept and return Sarvam's own formats. Auth: Authorization: Bearer, api-key or api-subscription-key, each with the Quilr key.

EndpointPurposeModels
/sarvam/text-to-speechSpeech synthesisbulbul:v3 (default), bulbul:v2
/sarvam/speech-to-textTranscription (multipart/form-data, one file)saaras:v3 (default), saaras:v4
/sarvam/speech-to-text-translateSpeech translationsaaras:v3, saaras:v4 (both need mode=translate), saaras:v2.5
/sarvam/translateText translationmayura:v1 (default), sarvam-translate:v1
/sarvam/transliterateTransliterationsarvam-transliterate
/sarvam/text-lidLanguage detectionsarvam-text-lid

Chat models (sarvam-105b, sarvam-105b-conversations, and the beta glm5.2, gemma4, deepseekv4-flash) use the standard Chat Completions endpoint.

TopicLimit or behavior
ModelsEnable every model and alias you call, including sarvam-transliterate and sarvam-text-lid.
Synthesislanguage_code is required. bulbul:v3: 2500 characters, speaker shubh, 24000 Hz. bulbul:v2: 1500 characters, speaker anushka, 22050 Hz.
RecognitionModes transcribe, translate, verbatim, translit, codemix. Segment timestamps only. Keep recordings under 30 seconds; bodies are capped at 25 MB.
Translationmayura:v1: 11 languages, auto source, 1000 characters. sarvam-translate:v1: 23 languages, explicit source, 2000 characters.
GuardrailsScan synthesized text, transcripts, text hints and translated output. Uploaded audio is forwarded unchanged and never logged.
Not coveredEmbeddings, rerank, Responses, Realtime, batch and streaming speech.

Responses API​

/openai_responses/v1/responses is a native passthrough served by openai_responses, openai_responses_azure and oracle_responses. Create, retrieve, cancel, delete and list input items are supported. Azure-style paths work too: /openai_responses/openai/deployments/{deployment}/responses. The deployment goes in body.model.

Guardrails scan input_text parts and instructions on the request, and output_text on non-streaming responses. previous_response_id and built-in tools pass through.

An openai or azureopenai provider cannot serve this endpoint. Add a Responses provider type to the app.

Assistants API​

/openai_assistants/ serves the OpenAI Assistants API (threads, runs, file search) through openai_assistants and openai_assistants_azure. Use it only for apps already built on Assistants.

Realtime API​

wss://<base>/openai_realtime/v1/realtime is a websocket passthrough for OpenAI Realtime voice and text, served by openai_realtime and openai_realtime_azure. Aliases: /openai/v1/realtime, /openai/realtime, /openai_realtime/openai/v1/realtime, /openai_realtime/openai/realtime.

The Quilr key is accepted, in priority order, as an Authorization: Bearer header, an api-key header, an api-key or api_key query parameter, an authorization query parameter, or the openai-insecure-api-key.<key> subprotocol (stripped before forwarding).

Guardrails coverage

Guardrails are not yet applied to live Realtime events. Session logging (handshake status, byte counts, usage) still runs.

Selecting a Provider on Multi-Provider Apps​

When an app has several providers that can serve a request, choose one by provider type or label. Without a selector, the gateway uses the one enabled provider that has the requested model; if several have it, one is selected at random.

EndpointBody fieldHeaderQuery parameter
Chat Completions, Anthropic Messages, Vertex AI, Embeddings, Rerank, Responsesprovider or provider_labelX-Provider-Name / X-Provider-Label-
Sarvam native (multipart routes: as a form field)provider or provider_labelX-Provider-Name / X-Provider-Label-
Realtime-X-Provider-Name / X-Provider-Labelprovider or provider_label

For apps linked to platform providers, the label is the platform provider's label.

SDK​

/sdk/v1/check scans text, messages or JSON for guardrail findings without calling any LLM. Create an app with the quilr_sdk provider. Install with pip install quilrai or npm install quilrai. See SDK Mode, or try it in the LLM Gateway Playground.

Microsoft Copilot Studio​

Create an app with the copilot_studio provider and register https://guardrails-usa-2.quilr.ai/copilot_studio/<Quilr key> (or your region) in Power Platform. Copilot Studio calls /validate and /analyze-tool-execution before tool execution. Block, redact and partial-redact outcomes return blockAction: true; scanning errors fail open. See Copilot Studio.