Providers and models
Settings > AI Gateway > Models, platform providers and credential reuse across apps are available in the V2 console only. In V1, each app keeps its own provider credentials.
The Models page (Settings > AI Gateway > Models) is where you manage every model the gateway can reach. It has four tabs:
A platform provider is a provider connection you set up once for the whole tenant and then link to any number of apps. You manage them in Settings > AI Gateway > Models under Your models, or with Provider configuration on the Settings > AI Gateway > LLM Gateway page, which opens the same list.
Platform providers or app-only credentials
An app uses either platform providers or app-only credentials, never both. You cannot convert an app later. To switch, create a new app.
Your models
Go to Settings > AI Gateway > Models > Models and select Your models.
Each card is one provider connection. It shows the label, the provider and API (for example Anthropic · Messages), its status, and each model's input, cached input and output price in USD per 1M tokens. Price source shows Published when the price came from the provider's published list.
The QuilrAI provided models switch shows the models QuilrAI hosts. See QuilrAI-provided models.
Add a provider
Click Add your own models. Nothing is saved until the last step. Use + Add provider at the top to connect several providers in one save (each provider type once per batch).
1. Connect
- Select the provider tile.
- Under Which API does the gateway talk to?, select the API your apps call. This sets the gateway provider type.
- Set the Provider label. The console suggests one such as
openai_02_oct_2026_1150. The label is permanent once saved. - Fill in the credentials. See Provider Support for each provider's fields.
- Click Check key and list models.
2. Models
Only the models you select here are reachable through this provider.
- Get available models asks the provider which models the key can reach.
- Add by ID adds a model by name. Select Add without checking to skip the reachability test.
- Validation set to Don't validate skips checks for the whole step.
3. Pricing
Enter what each model costs you in USD per 1M tokens. Published prices are prefilled, so you only need to adjust them to match your contract. Input and output are required; leave cached input empty if the provider has no separate cached price. The console computes estimated spend from these numbers.
4. Review & save
Click Review all providers, check the summary, and save.
Add models to an existing provider
Click + Add models to this provider under a card, or Add models in its gear menu.
- The stored key is reused. You never re-enter it.
- Provider, API and label stay fixed.
- Models already on the provider are marked Already added.
- Every linked app can call the new models once they are saved.
Enable, disable and delete
Disabling or deleting a platform provider, removing a model, or changing its credentials (for example with the Management API) applies to every app linked to it. Changes can take a short time to appear.
Link providers to an app
In Create App, step 1 selects Platform providers by default.
- Select one or more providers. The first one is the Primary; the rest are fallbacks. To change the order, use the up and down arrows.
- The app inherits every model and credential of the linked providers.
If none exist, the wizard says No platform providers yet and offers Add a platform provider (adds one without leaving the form), Provider configuration and Use app-only credentials. See Quick Start.
To send a request to one specific linked provider, pass its label. See Selecting a provider.
QuilrAI-provided models
QuilrAI hosts a catalog of chat models behind one OpenAI-compatible endpoint, https://models.quilrai.dev/v1. You do not need a provider account: create a model API key, select the models it may call, and send requests.
- Billing: each request is charged at the catalog's listed prices (USD per 1M tokens) and drawn from your organization's model credit. Contact QuilrAI support to raise your credit limit.
- Shared credit: the Playground, Workflow Agents and Red Teaming runs that use QuilrAI-provided models draw from the same credit.
A model API key can only call QuilrAI-provided models. Your own provider models cannot be added to it.
Browse the catalog
Go to Settings > AI Gateway > Models > Models and select QuilrAI provided models.
Columns: Model (the exact ID to send as model), Capabilities, Input price, Cached input price ("Not available" when the model has none), Output price and API schema (informational; every model is called the same way). These docs do not keep a copy of the catalog, because models and prices change: Settings > AI Gateway > Models > Models is the reference for current model IDs and prices (USD per 1M tokens).
Model API keys
- Go to Settings > AI Gateway > Models > API keys and click Create API key.
- Enter a Key name (required, 1 to 128 characters; "Console playground" is reserved) and pick Allowed models (at least one). The key can call only these models.
- Click Create API key. The full key (
sk-quilrllm-...) and the inference base URL are shown once. The console keeps only the key prefix.
If you lose the key, revoke it and create a new one. Revocation is permanent.
The list shows each key's name, key prefix, allowed models (View), status (ACTIVE / REVOKED), requests and spend this month, and Revoke. Keys have no expiry and no per-key spend limit; the organization credit is the only cap. You may also see keys you did not create: Console playground (used by the Playground) and Agent run <id> (a Workflow Agent run, limited to the run's model and revoked after it).
Permissions: viewing the catalog, keys and usage requires LLM Gateway read access. Creating a key requires LLM Gateway create access (under RBAC V2, llm.apps.create and secrets.reveal). Revoking requires LLM Gateway delete access (llm.apps.delete). Every key creation is recorded in the audit log.
Call a model
- cURL
- Python
- Node
curl 'https://models.quilrai.dev/v1/chat/completions' \
-H "Authorization: Bearer $QUILR_MODEL_API_KEY" \
-H "Content-Type: application/json" \
--data '{
"model": "deepseek/deepseek-v3.2",
"messages": [{"role": "user", "content": "Explain zero trust in one sentence."}],
"max_tokens": 1024
}'
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["QUILR_MODEL_API_KEY"],
base_url="https://models.quilrai.dev/v1",
)
response = client.chat.completions.create(
model="deepseek/deepseek-v3.2",
messages=[{"role": "user", "content": "Explain zero trust in one sentence."}],
max_tokens=1024,
)
print(response.choices[0].message.content)
import OpenAI from 'openai';
const client = new OpenAI({
apiKey: process.env.QUILR_MODEL_API_KEY,
baseURL: "https://models.quilrai.dev/v1",
});
const response = await client.chat.completions.create({
model: "deepseek/deepseek-v3.2",
messages: [
{ role: "user", content: "Explain zero trust in one sentence." }
],
temperature: 0.7,
top_p: 1,
max_tokens: 1024,
});
console.log(response.choices[0].message.content);
The Playground's Use this model in your app panel generates these snippets for the selected model, prompt and settings.
Use QuilrAI-provided models in a gateway app
Route QuilrAI-provided models through an LLM Gateway app to apply the app's guardrails, routing, limits and logs.
-
Create a model API key whose Allowed models include every model the app should use.
-
Add a General LLM provider to the app (the Custom endpoint tile with the Chat completions API,
general), as a platform provider or with app-only credentials: -
Call the app with its QuilrAI gateway key (
sk-quilr-...) as in the Quick Start, using the catalog ID asmodel.
Usage is billed from your organization credit either way and appears in the key's spend and the Usage tab. Selecting QuilrAI-provided models directly in Create App, without a General LLM provider, is coming soon.
Playground
Go to Settings > AI Gateway > Models > Playground (or click Chat on a catalog row). No key is needed; the Playground uses its own server-held key. Requests count toward your organization's usage, and responses are not stored by the console or streamed.
Conversations are limited to 20 messages and 64,000 characters.
Usage
Settings > AI Gateway > Models > Usage shows credit and spend. Usage can take a few moments to appear; click Refresh to update.
Below the tiles, a per-model table lists requests, tokens, input/output/total TPS, average latency, error rate, month spend and lifetime spend. If a banner says some token totals are estimated, the provider did not report exact usage for those requests. Per-key spend is on the API keys tab.