Skip to main content

Providers and models

V2 console only

Settings > AI Gateway > Models, platform providers and credential reuse across apps are available in the V2 console only. In V1, each app keeps its own provider credentials.

The Models page (Settings > AI Gateway > Models) is where you manage every model the gateway can reach. It has four tabs:

TabWhat it is for
ModelsYour models (your own provider connections and their costs) and QuilrAI provided models (the hosted catalog with retail prices).
API keysModel API keys for calling QuilrAI-provided models directly.
UsageCredit remaining, spend, requests, tokens and per-model throughput for QuilrAI-provided models.
PlaygroundChat with any active QuilrAI-provided model.

A platform provider is a provider connection you set up once for the whole tenant and then link to any number of apps. You manage them in Settings > AI Gateway > Models under Your models, or with Provider configuration on the Settings > AI Gateway > LLM Gateway page, which opens the same list.

Connect
Provider and API
Label and credentials
Models
List or add by ID
Pick what it serves
Pricing
USD per 1M tokens
Published prices prefilled
Review & save
Everything saves together
Link to apps
Create App
Platform providers
QuilrAI

Platform providers or app-only credentials​

Platform providerApp-only credentials
Set up inSettings > AI Gateway > Models, or Provider configuration on the LLM Gateway pageCreate App, or the app's LLM Providers section
CredentialsStored once, shared by every linked appStored on one app
Models added laterReach every linked app automaticallyAdded app by app
Credential change, disable or deleteAffects every linked appAffects one app
Available inV2 consoleV1 and V2 consoles
The choice is fixed when the app is created

An app uses either platform providers or app-only credentials, never both. You cannot convert an app later. To switch, create a new app.

Your models​

Go to Settings > AI Gateway > Models > Models and select Your models.

Each card is one provider connection. It shows the label, the provider and API (for example Anthropic · Messages), its status, and each model's input, cached input and output price in USD per 1M tokens. Price source shows Published when the price came from the provider's published list.

The QuilrAI provided models switch shows the models QuilrAI hosts. See QuilrAI-provided models.

Add a provider​

Click Add your own models. Nothing is saved until the last step. Use + Add provider at the top to connect several providers in one save (each provider type once per batch).

1. Connect​

  1. Select the provider tile.
  2. Under Which API does the gateway talk to?, select the API your apps call. This sets the gateway provider type.
  3. Set the Provider label. The console suggests one such as openai_02_oct_2026_1150. The label is permanent once saved.
  4. Fill in the credentials. See Provider Support for each provider's fields.
  5. Click Check key and list models.
TileAPI options (gateway provider type)
OpenAIChat completions (openai), Responses (openai_responses), Assistants (openai_assistants), Realtime (openai_realtime)
AnthropicMessages (anthropic_messages), Chat completions (anthropic)
AzureOpenAI chat (azureopenai), Responses (openai_responses_azure), Assistants (openai_assistants_azure), Realtime (openai_realtime_azure), Anthropic messages (anthropic_messages_azure)
Amazon BedrockConverse (bedrock), Anthropic messages (anthropic_messages_bedrock), Embeddings (bedrock_embeddings), Rerank (bedrock_rerank)
GoogleVertex AI (vertex_ai), Gemini OpenAI-compatible (gemini_chatcompletions)
DeepSeekChat completions (deepseek)
SarvamSpeech, text & chat (sarvam)
Cohere, Jina, VoyageRerank (cohere_rerank, jina_rerank, voyage_rerank)
Custom endpointChat completions (general), Rerank (general_rerank)

2. Models​

Only the models you select here are reachable through this provider.

  • Get available models asks the provider which models the key can reach.
  • Add by ID adds a model by name. Select Add without checking to skip the reachability test.
  • Validation set to Don't validate skips checks for the whole step.

3. Pricing​

Enter what each model costs you in USD per 1M tokens. Published prices are prefilled, so you only need to adjust them to match your contract. Input and output are required; leave cached input empty if the provider has no separate cached price. The console computes estimated spend from these numbers.

4. Review & save​

Click Review all providers, check the summary, and save.

Add models to an existing provider​

Click + Add models to this provider under a card, or Add models in its gear menu.

  • The stored key is reused. You never re-enter it.
  • Provider, API and label stay fixed.
  • Models already on the provider are marked Already added.
  • Every linked app can call the new models once they are saved.

Enable, disable and delete​

ActionEffect
Add modelsOpens the add-models flow above.
Disable providerStops traffic to this provider in every linked app. Credentials and models are kept.
Delete providerRemoves the provider. Its label can never be reused.
Remove model (model row gear)Removes one model. Linked apps can no longer call it.
Changes reach every linked app

Disabling or deleting a platform provider, removing a model, or changing its credentials (for example with the Management API) applies to every app linked to it. Changes can take a short time to appear.

In Create App, step 1 selects Platform providers by default.

  1. Select one or more providers. The first one is the Primary; the rest are fallbacks. To change the order, use the up and down arrows.
  2. The app inherits every model and credential of the linked providers.

If none exist, the wizard says No platform providers yet and offers Add a platform provider (adds one without leaving the form), Provider configuration and Use app-only credentials. See Quick Start.

To send a request to one specific linked provider, pass its label. See Selecting a provider.

QuilrAI-provided models​

QuilrAI hosts a catalog of chat models behind one OpenAI-compatible endpoint, https://models.quilrai.dev/v1. You do not need a provider account: create a model API key, select the models it may call, and send requests.

  • Billing: each request is charged at the catalog's listed prices (USD per 1M tokens) and drawn from your organization's model credit. Contact QuilrAI support to raise your credit limit.
  • Shared credit: the Playground, Workflow Agents and Red Teaming runs that use QuilrAI-provided models draw from the same credit.
You want toUse
Call a hosted model directly, billed from your QuilrAI creditA model API key against https://models.quilrai.dev/v1
Use QuilrAI-provided models with gateway guardrails, routing, limits and loggingAn LLM Gateway app with a General LLM provider (steps)
Use your own OpenAI, Anthropic, Azure, Bedrock or Vertex accountsAn LLM Gateway app with your own providers

A model API key can only call QuilrAI-provided models. Your own provider models cannot be added to it.

Browse the catalog​

Go to Settings > AI Gateway > Models > Models and select QuilrAI provided models.

ControlWhat it does
Search modelsFilters by model ID or provider prefix.
Sort by priceSorts by input, cached input or output price, low to high or high to low.
Chat (per row)Opens the Playground with that model selected.

Columns: Model (the exact ID to send as model), Capabilities, Input price, Cached input price ("Not available" when the model has none), Output price and API schema (informational; every model is called the same way). These docs do not keep a copy of the catalog, because models and prices change: Settings > AI Gateway > Models > Models is the reference for current model IDs and prices (USD per 1M tokens).

Model API keys​

  1. Go to Settings > AI Gateway > Models > API keys and click Create API key.
  2. Enter a Key name (required, 1 to 128 characters; "Console playground" is reserved) and pick Allowed models (at least one). The key can call only these models.
  3. Click Create API key. The full key (sk-quilrllm-...) and the inference base URL are shown once. The console keeps only the key prefix.

warning

If you lose the key, revoke it and create a new one. Revocation is permanent.

The list shows each key's name, key prefix, allowed models (View), status (ACTIVE / REVOKED), requests and spend this month, and Revoke. Keys have no expiry and no per-key spend limit; the organization credit is the only cap. You may also see keys you did not create: Console playground (used by the Playground) and Agent run <id> (a Workflow Agent run, limited to the run's model and revoked after it).

Permissions: viewing the catalog, keys and usage requires LLM Gateway read access. Creating a key requires LLM Gateway create access (under RBAC V2, llm.apps.create and secrets.reveal). Revoking requires LLM Gateway delete access (llm.apps.delete). Every key creation is recorded in the audit log.

Call a model​

Base URLhttps://models.quilrai.dev/v1
EndpointPOST /chat/completions (OpenAI Chat Completions format)
AuthAuthorization: Bearer $QUILR_MODEL_API_KEY
ModelThe exact catalog ID, including any prefix (for example deepseek/deepseek-v3.2). It must be one of the key's allowed models.
curl 'https://models.quilrai.dev/v1/chat/completions' \
-H "Authorization: Bearer $QUILR_MODEL_API_KEY" \
-H "Content-Type: application/json" \
--data '{
"model": "deepseek/deepseek-v3.2",
"messages": [{"role": "user", "content": "Explain zero trust in one sentence."}],
"max_tokens": 1024
}'

The Playground's Use this model in your app panel generates these snippets for the selected model, prompt and settings.

Use QuilrAI-provided models in a gateway app​

Route QuilrAI-provided models through an LLM Gateway app to apply the app's guardrails, routing, limits and logs.

  1. Create a model API key whose Allowed models include every model the app should use.

  2. Add a General LLM provider to the app (the Custom endpoint tile with the Chat completions API, general), as a platform provider or with app-only credentials:

    FieldValue
    base_urlhttps://models.quilrai.dev/v1
    api_keyYour model API key (sk-quilrllm-...)
    ModelsThe exact catalog IDs, for example deepseek/deepseek-v3.2
  3. Call the app with its QuilrAI gateway key (sk-quilr-...) as in the Quick Start, using the catalog ID as model.

Usage is billed from your organization credit either way and appears in the key's spend and the Usage tab. Selecting QuilrAI-provided models directly in Create App, without a General LLM provider, is coming soon.

Playground​

Go to Settings > AI Gateway > Models > Playground (or click Chat on a catalog row). No key is needed; the Playground uses its own server-held key. Requests count toward your organization's usage, and responses are not stored by the console or streamed.

SettingRangeDefault
ModelAny active catalog model
System prompt / User promptUp to 16,000 characters each
Temperature0 to 20.7
Top P0 to 11
Max output tokens1 to 4,0961,024

Conversations are limited to 20 messages and 64,000 characters.

Usage​

Settings > AI Gateway > Models > Usage shows credit and spend. Usage can take a few moments to appear; click Refresh to update.

TileShows
Credit remainingRemaining credit, and amount spent of your lifetime limit
Spend this monthSpend in the current UTC month
Requests this monthSuccessful and failed requests
Tokens this monthInput and output tokens
Total throughputTokens per second, input and output

Below the tiles, a per-model table lists requests, tokens, input/output/total TPS, average latency, error rate, month spend and lifetime spend. If a banner says some token totals are estimated, the provider did not report exact usage for those requests. Per-key spend is on the API keys tab.