Integration Guide
Connect to the QuilrAI gateway with your existing SDK. Change the base URL and the API key; keep everything else.
1. Choose Your Endpoint
Region
Use the regional endpoint closest to your application for production traffic. Use https://guardrails.quilr.ai only when you want global auto-routing. The examples below use US East.
API Format
Combine a region with a path, for example:
https://guardrails-usa-2.quilr.ai/openai_compatible/
Each path is served only by matching provider types on the app. For example, an openai provider cannot serve /openai_responses/; add an openai_responses provider. See the capability matrix.
sk-quilr-xxx stands for your Quilr key. Copy it from the app's API Integration section (see Applications and Keys). The model you send must be enabled on the app.
2. Code Examples
OpenAI-compatible chat
- Python
- JavaScript
- cURL
from openai import OpenAI
# Point the client to QuilrAI's gateway
client = OpenAI(
base_url='https://guardrails-usa-2.quilr.ai/openai_compatible/',
api_key='sk-openai-xxx'
api_key='sk-quilr-xxx'
)
# Everything below stays exactly the same
response = client.chat.completions.create(
model='gpt-4o-mini',
messages=[{'role': 'user', 'content': 'Hello!'}]
)
print(response.choices[0].message.content)
# Embeddings work too
embedding = client.embeddings.create(
model='text-embedding-3-small',
input='The quick brown fox'
)
print(embedding.data[0].embedding[:5])
import OpenAI from "openai";
// Point the client to QuilrAI's gateway
const client = new OpenAI({
baseURL: "https://guardrails-usa-2.quilr.ai/openai_compatible/",
apiKey: "sk-openai-xxx",
apiKey: "sk-quilr-xxx",
});
// Everything below stays exactly the same
const response = await client.chat.completions.create({
model: "gpt-4o-mini",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0].message.content);
# Point the request to QuilrAI's gateway
curl https://api.openai.com/v1/chat/completions \
curl https://guardrails-usa-2.quilr.ai/openai_compatible/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-openai-xxx" \
-H "Authorization: Bearer sk-quilr-xxx" \
-d '{
"model": "gpt-4o-mini",
"messages": [{"role": "user", "content": "Hello!"}]
}'
Bedrock, Vertex AI and Anthropic through OpenAI-compatible chat
Keep the OpenAI client and send a provider-native model name. The gateway translates the request to Bedrock Converse, Vertex AI generateContent or Anthropic Messages. This path is text-only; see Unified Completions for supported parameters, tools and streaming.
- AWS Bedrock
- Vertex AI
- Anthropic Messages
App provider: bedrock. Send any selected Bedrock model ID or inference profile ID that supports Converse.
from openai import OpenAI
client = OpenAI(
base_url='https://guardrails-usa-2.quilr.ai/openai_compatible/',
api_key='sk-quilr-xxx',
)
response = client.chat.completions.create(
model='amazon.nova-lite-v1:0',
messages=[{'role': 'user', 'content': 'Hello from an OpenAI client.'}],
max_tokens=256,
)
print(response.choices[0].message.content)
App provider: vertex_ai. Send a selected Gemini model name. Use the native /vertex_ai/ endpoint for multimodal calls.
from openai import OpenAI
client = OpenAI(
base_url='https://guardrails-usa-2.quilr.ai/openai_compatible/',
api_key='sk-quilr-xxx',
)
response = client.chat.completions.create(
model='gemini-2.5-flash',
messages=[{'role': 'user', 'content': 'Hello from an OpenAI client.'}],
max_tokens=256,
)
print(response.choices[0].message.content)
App provider: anthropic_messages, anthropic_messages_bedrock or anthropic_messages_azure. Send the Claude model name, or the Bedrock Claude model ID for anthropic_messages_bedrock.
from openai import OpenAI
client = OpenAI(
base_url='https://guardrails-usa-2.quilr.ai/openai_compatible/',
api_key='sk-quilr-xxx',
)
response = client.chat.completions.create(
model='claude-sonnet-4-5',
messages=[{'role': 'user', 'content': 'Hello from an OpenAI client.'}],
max_tokens=256,
)
print(response.choices[0].message.content)
Embeddings
Every embeddings provider (openai, azureopenai, bedrock_embeddings) takes the OpenAI embeddings shape. For Bedrock, the AWS credentials stay on the app's provider and the gateway makes the Bedrock call.
- Python
- cURL
from openai import OpenAI
client = OpenAI(
base_url='https://guardrails-usa-2.quilr.ai/openai_compatible/',
api_key='sk-quilr-xxx',
)
# Same call for OpenAI, Azure OpenAI, or AWS Bedrock embeddings providers.
# Use a model name enabled on your app
# (e.g. 'text-embedding-3-small', 'amazon.titan-embed-text-v2:0',
# 'cohere.embed-english-v3').
embedding = client.embeddings.create(
model='amazon.titan-embed-text-v2:0',
input='The quick brown fox',
)
print(embedding.data[0].embedding[:5])
curl https://guardrails-usa-2.quilr.ai/openai_compatible/v1/embeddings \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-quilr-xxx" \
-d '{
"model": "amazon.titan-embed-text-v2:0",
"input": "The quick brown fox"
}'
Rerank
Every rerank provider takes the Cohere-compatible shape. Point the Cohere SDK at https://guardrails-usa-2.quilr.ai/rerank; /rerank/rerank, /rerank/v1/rerank and /rerank/v2/rerank all work.
- Python
- cURL
import cohere
co = cohere.ClientV2(
base_url='https://guardrails-usa-2.quilr.ai/rerank',
api_key='co-xxx',
api_key='sk-quilr-xxx',
)
# Same call for any configured rerank provider.
result = co.rerank(
model='rerank-english-v3.0',
query='What is the capital of France?',
documents=[
'Paris is the capital of France.',
'Berlin is the capital of Germany.',
'The Eiffel Tower is in Paris.',
],
top_n=2,
)
for r in result.results:
print(r.index, r.relevance_score)
curl https://guardrails-usa-2.quilr.ai/rerank/v2/rerank \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-quilr-xxx" \
-d '{
"model": "rerank-english-v3.0",
"query": "What is the capital of France?",
"documents": [
"Paris is the capital of France.",
"Berlin is the capital of Germany.",
"The Eiffel Tower is in Paris."
],
"top_n": 2
}'
Anthropic
- Python
- JavaScript
- cURL
import anthropic
# Point the client to QuilrAI's gateway
client = anthropic.Anthropic(
# uses default base URL
base_url='https://guardrails-usa-2.quilr.ai/anthropic_messages/',
api_key='sk-ant-xxx'
api_key='sk-quilr-xxx'
)
# Everything below stays exactly the same
message = client.messages.create(
model='claude-sonnet-4-20250514',
max_tokens=1024,
messages=[{'role': 'user', 'content': 'Hello!'}]
)
print(message.content[0].text)
import Anthropic from "@anthropic-ai/sdk";
// Point the client to QuilrAI's gateway
const client = new Anthropic({
// uses default base URL
baseURL: "https://guardrails-usa-2.quilr.ai/anthropic_messages/",
apiKey: "sk-ant-xxx",
apiKey: "sk-quilr-xxx",
});
// Everything below stays exactly the same
const message = await client.messages.create({
model: "claude-sonnet-4-20250514",
max_tokens: 1024,
messages: [{ role: "user", content: "Hello!" }],
});
console.log(message.content[0].text);
# Point the request to QuilrAI's gateway
curl https://api.anthropic.com/v1/messages \
curl https://guardrails-usa-2.quilr.ai/anthropic_messages/v1/messages \
-H "Content-Type: application/json" \
-H "x-api-key: sk-ant-xxx" \
-H "x-api-key: sk-quilr-xxx" \
-H "anthropic-version: 2023-06-01" \
-d '{
"model": "claude-sonnet-4-20250514",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Hello!"}]
}'
Vertex AI
Pass the Quilr key as a Bearer token. project and location should match the GCP project ID and region on the app's vertex_ai provider.
- Google GenAI SDK
- LangChain
from google import genai
from google.genai.types import HttpOptions
from google.oauth2 import service_account
from google.auth import credentials as auth_credentials
class APIKeyCredentials(auth_credentials.Credentials):
"""Pass the QuilrAI API key as a Bearer token."""
def __init__(self, api_key):
super().__init__()
self.api_key = api_key
self.token = api_key
def refresh(self, request):
self.token = self.api_key
@property
def valid(self):
return True
credentials = service_account.Credentials.from_service_account_file(
'service.json',
scopes=['https://www.googleapis.com/auth/cloud-platform']
)
credentials = APIKeyCredentials('sk-quilr-xxx')
client = genai.Client(
vertexai=True,
project='your-gcp-project',
location='us-central1',
credentials=credentials,
# uses default Vertex AI endpoint
http_options=HttpOptions(base_url='https://guardrails-usa-2.quilr.ai/vertex_ai'),
)
# Everything below stays exactly the same
response = client.models.generate_content(
model='gemini-2.5-flash',
contents='Hello!'
)
print(response.text)
from google.oauth2 import service_account
from google.oauth2 import credentials as ga_credentials
from langchain_google_genai import ChatGoogleGenerativeAI
class _NoopCredentials(ga_credentials.Credentials):
"""Inject the QuilrAI API key as a Bearer token."""
def __init__(self, api_key):
super().__init__(token=api_key)
def refresh(self, request):
pass
@property
def valid(self):
return True
credentials = service_account.Credentials.from_service_account_file(
'service.json',
scopes=['https://www.googleapis.com/auth/cloud-platform']
)
credentials = _NoopCredentials('sk-quilr-xxx')
llm = ChatGoogleGenerativeAI(
model='gemini-2.5-flash',
credentials=credentials,
base_url='https://guardrails-usa-2.quilr.ai/vertex_ai',
project='your-gcp-project',
location='us-central1',
vertexai=True,
)
# Everything below stays exactly the same
response = llm.invoke('Hello!')
print(response.content)
AWS Bedrock Runtime - boto3
App provider: bedrock. Point the Bedrock Runtime client at the gateway and sign with the Quilr key.
import boto3
from botocore.config import Config
QUILR_KEY = "sk-quilr-xxx"
bedrock = boto3.client(
"bedrock-runtime",
region_name="us-east-1",
endpoint_url="https://guardrails-usa-2.quilr.ai/bedrock-runtime",
aws_access_key_id="AKIA...",
aws_access_key_id=QUILR_KEY,
aws_secret_access_key="aws-secret",
aws_secret_access_key=QUILR_KEY,
config=Config(read_timeout=300),
)
response = bedrock.converse(
modelId="amazon.nova-lite-v1:0",
messages=[
{
"role": "user",
"content": [{"text": "Hello!"}],
}
],
inferenceConfig={"maxTokens": 256},
)
print(response["output"]["message"]["content"][0]["text"])
converse, converse_stream and invoke_model are supported. See AWS Bedrock - boto3 Runtime for coverage and troubleshooting.
OpenAI Responses
App provider: openai_responses or openai_responses_azure. For Azure, send the deployment name as model.
- Python
- JavaScript
- cURL
from openai import OpenAI
# Point the client to QuilrAI's gateway
client = OpenAI(
base_url='https://guardrails-usa-2.quilr.ai/openai_responses/v1',
api_key='sk-openai-xxx'
api_key='sk-quilr-xxx'
)
# Everything below stays exactly the same
response = client.responses.create(
model='gpt-5',
input=[{'role': 'user', 'content': 'Hello!'}],
instructions='You are a helpful assistant.'
)
print(response.output_text)
import OpenAI from "openai";
// Point the client to QuilrAI's gateway
const client = new OpenAI({
baseURL: "https://guardrails-usa-2.quilr.ai/openai_responses/v1",
apiKey: "sk-openai-xxx",
apiKey: "sk-quilr-xxx",
});
// Everything below stays exactly the same
const response = await client.responses.create({
model: "gpt-5",
input: [{ role: "user", content: "Hello!" }],
instructions: "You are a helpful assistant.",
});
console.log(response.output_text);
# Point the request to QuilrAI's gateway
curl https://api.openai.com/v1/responses \
curl https://guardrails-usa-2.quilr.ai/openai_responses/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-openai-xxx" \
-H "Authorization: Bearer sk-quilr-xxx" \
-d '{
"model": "gpt-5",
"input": [{"role": "user", "content": "Hello!"}]
}'
OpenAI Realtime
App provider: openai_realtime or openai_realtime_azure. Sessions are a websocket passthrough; guardrails are not yet applied to live events (see Realtime API).
- Python
- JavaScript
import asyncio
from openai import AsyncOpenAI
async def main():
client = AsyncOpenAI(
base_url='https://guardrails-usa-2.quilr.ai/openai/v1',
api_key='sk-openai-xxx',
api_key='sk-quilr-xxx',
)
# Everything below stays exactly the same
async with client.realtime.connect(model='gpt-realtime') as conn:
await conn.session.update(session={'modalities': ['text']})
await conn.conversation.item.create(item={
'type': 'message',
'role': 'user',
'content': [{'type': 'input_text', 'text': 'Hello!'}],
})
await conn.response.create()
async for event in conn:
if event.type == 'response.output_text.delta':
print(event.delta, end='', flush=True)
elif event.type == 'response.done':
break
asyncio.run(main())
import { OpenAIRealtimeWebSocket } from "openai/realtime/websocket";
const rt = new OpenAIRealtimeWebSocket({
baseURL: "wss://guardrails-usa-2.quilr.ai/openai/v1",
apiKey: "sk-openai-xxx",
apiKey: "sk-quilr-xxx",
model: "gpt-realtime",
});
rt.on("response.output_text.delta", (e) => process.stdout.write(e.delta));
rt.send({
type: "conversation.item.create",
item: {
type: "message",
role: "user",
content: [{ type: "input_text", text: "Hello!" }],
},
});
rt.send({ type: "response.create" });
Sarvam speech and text
App provider: sarvam. Chat goes through the OpenAI-compatible client:
from openai import OpenAI
client = OpenAI(
base_url='https://guardrails-usa-2.quilr.ai/openai_compatible/',
api_key='sk-quilr-xxx'
)
resp = client.chat.completions.create(
model='sarvam-105b',
messages=[{'role': 'user', 'content': 'Hello!'}]
)
Speech uses the OpenAI audio methods. voice is a Sarvam speaker and language_code is required:
speech = client.audio.speech.create(
model='bulbul:v3',
input='Hello world',
voice='shubh',
response_format='wav',
extra_body={'language_code': 'en-IN'}
)
with open('hello.wav', 'wb') as out:
out.write(speech.content)
with open('audio.wav', 'rb') as audio:
transcript = client.audio.transcriptions.create(model='saaras:v4', file=audio)
The native /sarvam/ routes take Sarvam's own fields. Translation, transliteration and language detection are only here:
# Speech synthesis - returns {"request_id": "...", "audios": ["<base64>"]}
curl https://guardrails-usa-2.quilr.ai/sarvam/text-to-speech \
-H "Authorization: Bearer sk-quilr-xxx" \
-H "Content-Type: application/json" \
-d '{
"text": "Hello world",
"model": "bulbul:v3",
"speaker": "shubh",
"language_code": "en-IN",
"output_audio_codec": "wav"
}'
# Transcription - multipart, one file field
curl https://guardrails-usa-2.quilr.ai/sarvam/speech-to-text \
-H "Authorization: Bearer sk-quilr-xxx" \
-F file=@audio.wav \
-F model=saaras:v4 \
-F mode=transcribe \
-F language_code=hi-IN
# Text translation
curl https://guardrails-usa-2.quilr.ai/sarvam/translate \
-H "Authorization: Bearer sk-quilr-xxx" \
-H "Content-Type: application/json" \
-d '{
"model": "mayura:v1",
"input": "Hello",
"source_language_code": "en-IN",
"target_language_code": "hi-IN"
}'
See Sarvam Speech and Text for every endpoint, model and limit.
Microsoft Copilot Studio
App provider: copilot_studio. Register this endpoint base in Power Platform admin center:
https://guardrails-usa-2.quilr.ai/copilot_studio/sk-quilr-xxx
See Copilot Studio for setup.
TrueFoundry custom guardrails
Add QuilrAI as a TrueFoundry custom input/output guardrail: URL = your regional base plus /sdk/v1/check/truefoundry, mode Mutate, Custom Bearer Auth with a Quilr key from a quilr_sdk app. See TrueFoundry Integration.
3. Optional Headers
4. Using Routing Groups
Send a routing group name as model. The gateway load-balances and fails over across the group's models.
response = client.chat.completions.create(
model='Group1', # your routing group name
messages=[{'role': 'user', 'content': 'Hello!'}]
)
5. Selecting a Provider
On an app with several providers, pick one per request by provider type or label. The fields for each endpoint are in Selecting a Provider on Multi-Provider Apps.
# Responses: pick a specific additional provider
response = client.responses.create(
model='gpt-5',
input=[{'role': 'user', 'content': 'Hello!'}],
extra_body={'provider_label': 'azure-westus'},
)
# Realtime: select via query string (headers also work)
async with client.realtime.connect(
model='gpt-realtime',
extra_query={'provider_label': 'azure-westus'},
) as conn:
...