Skip to main content

Unified Completions

Use OpenAI Chat Completions clients with provider-native chat models. QuilrAI accepts an OpenAI-style /chat/completions request, translates it to the selected provider, and returns a chat completion response in OpenAI format.

Scope​

This page covers translated providers on:

/openai_compatible/v1/chat/completions
Provider typeUpstream callModel valueStreamingContent support
bedrockBedrock Converse / ConverseStreamSelected Bedrock model ID or inference profile IDYesText only
vertex_aiVertex AI Gemini generateContent / streamGenerateContentSelected Gemini model nameYesText only
anthropic_messagesAnthropic Messages messages.create / streamingSelected Claude model nameYesText only
anthropic_messages_bedrockAnthropic Messages on BedrockSelected Bedrock Claude model IDYesText only
anthropic_messages_azureAzure Anthropic MessagesSelected Claude model nameYesText only

This page does not cover:

  • OpenAI, Azure OpenAI, Anthropic OpenAI-compatible, DeepSeek, Gemini public API, Sarvam, or custom providers that already expose an OpenAI-compatible upstream API
  • Native Vertex AI /vertex_ai/ routes
  • Native Anthropic Messages /anthropic_messages/ routes
  • AWS Bedrock Runtime boto3 routes such as /bedrock-runtime/model/{model_id}/converse
  • Bedrock embeddings or rerank
  • OpenAI Responses API

Use Unified Completions when your application already uses OpenAI Chat Completions and you want to call selected Bedrock, Vertex AI Gemini, or Anthropic Messages models without switching SDKs.

Native multimodal routes

The translated OpenAI-compatible path is text-only today. Use native Vertex AI, Anthropic Messages, or Bedrock Runtime routes when you need provider-native image, audio, video, document, or other multimodal request shapes.

Request Flow​

  1. Add a provider of type bedrock, vertex_ai, anthropic_messages, anthropic_messages_bedrock, or anthropic_messages_azure to your app (or, in the V2 console, link a platform provider of that type).
  2. Enable the models the app is allowed to call.
  3. Point your OpenAI SDK or OpenAI-compatible wrapper at the closest regional endpoint, such as https://guardrails-usa-2.quilr.ai/openai_compatible/.
  4. Send the provider model name in the OpenAI SDK model parameter.
  5. QuilrAI translates the OpenAI-style request to the provider-native chat API and translates the provider response back to OpenAI Chat Completions.
from openai import OpenAI

client = OpenAI(
base_url="https://guardrails-usa-2.quilr.ai/openai_compatible/",
api_key="sk-quilr-xxx",
)

response = client.chat.completions.create(
model="amazon.nova-lite-v1:0",
messages=[{"role": "user", "content": "Summarize this in one sentence."}],
max_tokens=256,
)

print(response.choices[0].message.content)

For Vertex AI Gemini, use the same OpenAI client configuration and pass a selected Gemini model:

response = client.chat.completions.create(
model="gemini-2.5-flash",
messages=[{"role": "user", "content": "Write a concise release note."}],
max_tokens=256,
)

For Anthropic Messages, use the same OpenAI client configuration and pass a selected Claude model:

response = client.chat.completions.create(
model="claude-sonnet-4-5",
messages=[{"role": "user", "content": "Write a concise release note."}],
max_tokens=256,
)

The normal gateway behavior still applies: authentication, provider and model routing, Prompt Store substitution, request-side DLP, response-side DLP for non-streaming responses, logging, rate limits, token estimates, Guardian checks, and performance metrics.

Common Contract​

These translators use allowlists. Unknown OpenAI parameters are rejected instead of silently dropped.

OpenAI parameterBedrock supportVertex AI supportAnthropic Messages support
messagesSupportedSupportedSupported
modelSelected Bedrock model ID or inference profile IDSelected Gemini model nameSelected Claude model name or Bedrock Claude model ID
streamBedrock ConverseStreamVertex streamGenerateContentAnthropic Messages streaming
max_tokensinferenceConfig.maxTokensgenerationConfig.maxOutputTokensmax_tokens
max_completion_tokensinferenceConfig.maxTokensgenerationConfig.maxOutputTokensmax_tokens
temperatureinferenceConfig.temperaturegenerationConfig.temperaturetemperature
top_pinferenceConfig.topPgenerationConfig.topPtop_p
stopinferenceConfig.stopSequencesgenerationConfig.stopSequencesstop_sequences
toolsFunction toolsFunction declarationsAnthropic tools
functionsLegacy function toolsLegacy function declarationsLegacy function tools
tool_choiceauto, none, required, and function-specific choicesauto, none, required, and function-specific choicesauto, none, required, and function-specific choices
function_callSupportedSupportedSupported
response_formattext, json_schematext, json_object, json_schematext only
stream_optionsinclude_usageinclude_usageinclude_usage
parallel_tool_callsAccepted as a boolean; false is not enforcedAccepted as a boolean; false is not enforcedAccepted as a boolean; false is not enforced
nMust be absent or 1Must be absent or 1Must be absent or 1
metadataAccepted, not sent upstreamAccepted, not sent upstreamSent as Anthropic metadata
userAccepted, not sent upstreamAccepted, not sent upstreamSent as metadata.user_id when not already set
systemRejectedRejectedAccepted as a top-level convenience field and merged with system/developer messages
frequency_penaltyRejectedgenerationConfig.frequencyPenaltyRejected
presence_penaltyRejectedgenerationConfig.presencePenaltyRejected
seedRejectedgenerationConfig.seedRejected

max_tokens and max_completion_tokens can both be present only when they have the same value. Token limits must be positive integers. stop must be a string or a list of strings.

Anthropic Messages requires max_tokens; if both max_tokens and max_completion_tokens are omitted, QuilrAI sends max_tokens: 1024.

Common rejected OpenAI parameters include logit_bias, logprobs, top_logprobs, reasoning_effort, modalities, audio, store, service_tier, prediction, and provider-specific extra_body.

Message Support​

OpenAI roleBedrock translationVertex AI translationAnthropic Messages translation
systemTop-level Bedrock system blockTop-level systemInstruction.partsTop-level Anthropic system
developerTop-level Bedrock system blockTop-level systemInstruction.partsTop-level Anthropic system
userBedrock message with role userVertex content with role userAnthropic message with role user
assistantBedrock message with role assistantVertex content with role modelAnthropic message with role assistant
toolBedrock toolResult block when tools are activeVertex functionResponse partAnthropic tool_result block
functionBedrock toolResult block for legacy function historyVertex functionResponse partAnthropic tool_result block

Content is text-only on translated paths:

  • String content is supported.
  • Content arrays are supported only when every part is text-like.
  • Images, audio, files, Bedrock document blocks, Bedrock image blocks, Vertex inline media, Vertex file data, Anthropic image blocks, and Anthropic document blocks are not supported on this OpenAI-compatible path.

Tools and Function Calling​

OpenAI tools are supported when each tool is type: "function".

OpenAI fieldBedrock fieldVertex AI fieldAnthropic Messages field
tools[].function.nametoolSpec.namefunctionDeclarations[].nametools[].name
tools[].function.descriptiontoolSpec.descriptionfunctionDeclarations[].descriptiontools[].description
tools[].function.parameterstoolSpec.inputSchema.jsonfunctionDeclarations[].parameterstools[].input_schema
tools[].function.stricttoolSpec.strictValidated through schema handling; not emitted as a Vertex fieldNot translated

Legacy OpenAI functions and function_call are also supported.

Modern assistant.tool_calls entries should include id. Bedrock and Anthropic Messages reject tool calls without IDs because the later tool result would have no stable identifier. Vertex maps modern assistant.tool_calls and legacy assistant.function_call to functionCall parts, parsing JSON argument strings when possible.

Tool Choice​

OpenAI tool_choiceBedrock behaviorVertex AI behaviorAnthropic Messages behavior
absentNo explicit Bedrock toolChoiceNo Vertex toolConfigNo Anthropic tool_choice
autoNo explicit Bedrock toolChoiceNo Vertex toolConfig{"type": "auto"}
noneNo Bedrock toolConfig; tool history is serialized as textfunctionCallingConfig.mode = "NONE"No tools are sent
requiredBedrock toolChoice.anyfunctionCallingConfig.mode = "ANY"{"type": "any"}
function-specific choiceBedrock toolChoice.tool.namefunctionCallingConfig.mode = "ANY" with allowedFunctionNames{"type": "tool", "name": ...}

If a request requires a tool but no tools are present, QuilrAI rejects the request before the upstream call.

Parallel Tool Results​

OpenAI-compatible clients often send one role: "tool" message per tool call:

[
{"role": "assistant", "tool_calls": [{"id": "call_a"}, {"id": "call_b"}]},
{"role": "tool", "tool_call_id": "call_a", "content": "result A"},
{"role": "tool", "tool_call_id": "call_b", "content": "result B"}
]

Bedrock, Vertex AI, and Anthropic Messages expect matching tool results for a model turn to stay together in the next user-side content entry. QuilrAI groups consecutive OpenAI role: "tool" or role: "function" messages into one provider-native user message containing multiple tool-result blocks.

This grouping matters for parallel tool calls. Sending each tool result as a separate provider turn can produce upstream validation errors about missing tool results or function responses.

Structured Output​

response_formatBedrock supportVertex AI supportAnthropic Messages support
{"type": "text"}SupportedSupportedSupported
{"type": "json_object"}RejectedMaps to generationConfig.responseMimeType = "application/json"Rejected
{"type": "json_schema", "json_schema": {...}}Supported on Bedrock models that accept outputConfigMaps to generationConfig.responseMimeType = "application/json" plus generationConfig.responseSchemaRejected

For Bedrock json_schema, QuilrAI maps:

OpenAI fieldBedrock field
json_schema.nameoutputConfig.textFormat.jsonSchema.name
json_schema.descriptionoutputConfig.textFormat.jsonSchema.description
json_schema.schemaoutputConfig.textFormat.jsonSchema.schema

json_schema.strict is validated as a boolean, but it is not separately mapped. Bedrock structured output is schema-constrained through outputConfig.textFormat.

For Vertex AI json_schema, QuilrAI sanitizes the OpenAI schema into the subset accepted by Vertex Gemini. name, description, and strict are type-checked but are not emitted as separate Vertex fields.

Vertex schema normalization:

  • Local $ref values are inlined.
  • Cyclic or unresolvable local refs are rejected.
  • Nullable single-type unions collapse to type plus nullable: true.
  • Unsupported JSON Schema metadata keys are dropped.
  • oneOf and anyOf are accepted only for nullable single-type unions.
  • Other union arrays are rejected.

Dropped Vertex schema metadata keys include $defs, $id, $schema, additionalProperties, default, definitions, patternProperties, title, allOf, and unsupported oneOf or anyOf. Schema property names are preserved; for example, a property named default remains valid.

Anthropic Messages does not map OpenAI JSON mode or structured output on this translated path today. Only {"type": "text"} is accepted.

Streaming Responses​

Streaming returns OpenAI-compatible server-sent events.

ProviderUpstream streamOpenAI stream behavior
BedrockConverseStreamBedrock message, text, tool-use, stop, and usage events are converted to OpenAI deltas
Vertex AIstreamGenerateContent with alt=sseVertex candidate text, function calls, finish reasons, and usage metadata are converted to OpenAI deltas
Anthropic MessagesAnthropic Messages streamAnthropic message, text, tool-use, stop, and usage events are converted to OpenAI deltas

When stream_options.include_usage is true, Bedrock emits usage chunks from metadata.usage; Vertex emits one final usage chunk using the latest usageMetadata observed in the stream; Anthropic Messages emits one final usage chunk from tracked input and output token counts.

Streaming response-side DLP is not applied on this path. QuilrAI performs request-side scanning, forwards chunks, and accumulates text and tool-call data for logging.

Non-Streaming Responses​

Non-streaming provider responses are converted back to OpenAI chat completions:

  • Provider text parts are joined into choices[].message.content.
  • Provider tool-use or function-call parts become OpenAI choices[].message.tool_calls.
  • Tool-call-only responses return message.content: null.
  • Provider usage maps to OpenAI prompt_tokens, completion_tokens, and total_tokens.

Finish reason mapping:

OpenAI finish_reasonBedrockAnthropic MessagesVertex AI
stopend_turn, stop_sequenceend_turn, stop_sequence, pause_turnSTOP and any unmapped reason
lengthmax_tokensmax_tokensMAX_TOKENS
tool_callstool_usetool_useAny response containing function calls
content_filtercontent_filtered, guardrail_intervenedrefusalSAFETY, RECITATION, BLOCKLIST, PROHIBITED_CONTENT, SPII, IMAGE_SAFETY
null--FINISH_REASON_UNSPECIFIED

Unknown Bedrock stop reasons pass through unchanged.

Vertex usage details include:

Vertex usage fieldOpenAI usage field
promptTokenCountprompt_tokens
candidatesTokenCount plus thoughtsTokenCountcompletion_tokens
totalTokenCounttotal_tokens
thoughtsTokenCountcompletion_tokens_details.reasoning_tokens
cachedContentTokenCountprompt_tokens_details.cached_tokens

Anthropic Messages usage maps usage.input_tokens to prompt_tokens, usage.output_tokens to completion_tokens, and their sum to total_tokens.

Provider Setup​

Credentials for these provider types (API key, AWS static keys or assume role, Vertex API key or service account, Azure Foundry base URL) are listed once in Provider Support.

  • anthropic_messages_azure needs an API key and the Azure AI Foundry Base URL.
  • Direct and Azure Anthropic send anthropic_version: 2023-06-01 by default.
  • For bedrock, select chat models that support Converse, and send the Bedrock model ID or inference profile ID as model. The default AWS region is us-east-1.
  • Vertex AI model listing is best effort and falls back to a curated Gemini list if fetching fails.

Error Handling​

QuilrAI returns errors in OpenAI format for request translation validation failures and preserves upstream provider messages where possible.

ProblemBedrockAnthropic MessagesVertex AI
Parameter not translatedunsupported_bedrock_openai_parameterunsupported_anthropic_openai_parameterunsupported_vertex_openai_parameter
Invalid parameter value or typeinvalid_bedrock_openai_parameterinvalid_anthropic_openai_parameterinvalid_vertex_openai_parameter
Unsupported content (image, audio, file, document parts)unsupported_bedrock_openai_contentunsupported_anthropic_openai_contentunsupported_vertex_openai_content
Invalid message order or tool-result historyinvalid_bedrock_openai_messagesinvalid_anthropic_openai_messagesinvalid_vertex_openai_messages
Unsupported roleunsupported_bedrock_openai_roleunsupported_anthropic_openai_roleunsupported_vertex_openai_role
Malformed tool definitions or historyinvalid_bedrock_openai_toolsinvalid_anthropic_openai_toolsinvalid_vertex_openai_tools
Tool shape cannot be translatedunsupported_bedrock_openai_toolsunsupported_anthropic_openai_toolsunsupported_vertex_openai_tools
Error codeMeaning
bedrock_credentials_errorBedrock credentials could not be loaded or used.
bedrock_converse_error / bedrock_converse_stream_errorBedrock Converse / ConverseStream returned an error.
missing_provider_keyThe Anthropic provider API key is missing.
missing_azure_anthropic_base_urlThe Azure Anthropic provider is missing its Base URL.
anthropic_messages_bedrock_credentials_errorBedrock credentials could not be loaded or used for Anthropic Messages on Bedrock.
anthropic_messages_bedrock_package_missingThe Anthropic Bedrock package needed for the upstream call is missing.
anthropic_messages_error / anthropic_messages_stream_errorAnthropic Messages (or its stream) returned an error.
vertex_credentials_errorVertex credentials could not be loaded or used.
vertex_generate_content_error / _timeout / _parse_errorVertex generateContent failed, timed out, or returned an unparseable response.
vertex_stream_generate_content_errorVertex streamGenerateContent returned an error.
vertex_stream_timeout / vertex_stream_parse_errorThe Vertex stream timed out or sent an unparseable event.

Upstream Vertex HTTP errors map to OpenAI-style types: 401 to authentication_error, 429 to rate_limit_error, other 4xx to invalid_request_error, and 5xx to upstream_error.

Guardrail Behavior​

Request-side DLP scans user text before the upstream call. Non-streaming responses are scanned before they are returned to the client.

Streaming responses are different: request-side DLP still runs, but response-side DLP is skipped so chunks can pass through as they arrive.

Tool messages are carried through without changing tool IDs, function response names, or result ordering. Changing a tool_call_id, dropping a role: "tool" message, or reordering tool results can break provider tool-result validation.

Expected Failures​

Besides the unsupported parameters and content above, the following are rejected:

  • Tool result messages missing tool_call_id, or assistant tool_calls without id (Bedrock and Anthropic Messages)
  • A user message right after assistant tool calls, without the matching tool results
  • Parallel tool results that are not consecutive, so they cannot be grouped into one provider-native user turn
  • n > 1, log probabilities, token bias, audio modes, reasoning-effort controls and provider-specific extra_body
  • response_format other than text on Anthropic Messages; json_object on Bedrock; json_schema on Bedrock models without outputConfig
  • Vertex JSON Schema unions other than nullable single-type unions, and cyclic or unresolvable local refs