Skip to main content

QuilrAI-Provided Models

V2 console only

Settings > Models is available in the V2 console only.

QuilrAI hosts a catalog of chat models behind one OpenAI-compatible endpoint. You do not need a provider account: create a model API key, pick the models it may call, and send requests to https://models.quilrai.dev/v1.

BillingEach request is charged at the catalog's listed prices (USD per 1M tokens) and drawn from your organization's model credit. Contact QuilrAI support to raise your credit limit.
Shared creditThe Playground, Workflow Agents, and Red Teaming runs that use QuilrAI-provided models all draw from the same credit.
WhereSettings > Models: Models (catalog), API keys, Usage, Playground tabs.

When to use which​

You want toUse
Call a hosted model directly, without a provider account, billed from your QuilrAI creditA model API key against https://models.quilrai.dev/v1 (steps 1 to 3)
Use QuilrAI-provided models with LLM Gateway guardrails, routing, rate and token limits, and loggingAn LLM Gateway app with a General LLM provider that points at QuilrAI-provided models (steps)
Use your own OpenAI, Anthropic, Azure, Bedrock, or Vertex accountsAn LLM Gateway app with your own providers (Quick Start, Providers and Models)

A model API key can only call QuilrAI-provided models. Your own provider models cannot be added to it.

1. Browse the catalog​

Go to Settings > Models > Models and select QuilrAI provided models.

QuilrAI-provided models catalog

ControlWhat it does
Search modelsFilters by model id or provider prefix.
Sort by priceSorts by input, cached input, or output price, low to high or high to low.
Chat (per row)Opens the Playground with that model selected.

Columns: Model (the exact id to send as model), Capabilities, Input price, Cached input price ("Not available" when the model has no cached-input price), Output price, and API schema. Prices are USD per 1M tokens. See the full catalog below.

2. Create a model API key​

  1. Go to Settings > Models > API keys and click Create API key.

    API keys tab with the Create API key button

  2. Fill in the dialog:

    FieldNotes
    Key nameRequired, 1 to 128 characters. Example: "Production inference". The name "Console playground" is reserved.
    Allowed modelsRequired, at least one. Searchable multi-select of the QuilrAI catalog. The key can call only these models.

    Create model API key dialog

    Allowed models picker open

    One allowed model selected

  3. Click Create API key. The full key (sk-quilrllm-...) and the inference base URL are shown once. Copy the key and store it securely; the console keeps only the key prefix.

warning

If you lose the key, revoke it and create a new one. Revocation is permanent.

The API keys list shows each key's name, key prefix, allowed models (click View to expand), status (ACTIVE / REVOKED), requests and spend this month, and a Revoke action. Keys have no expiry and no per-key spend limit; the organization credit is the only cap.

Models a key may call

You may also see keys you did not create:

Key nameCreated by
Console playgroundThe Playground. Older playground keys stay listed after they are replaced.
Agent run <id>A Workflow Agent run. Limited to the run's model and revoked after the run.

Permissions: viewing the catalog, keys, and usage needs LLM Gateway read access. Creating a key needs LLM Gateway create access (under RBAC V2, llm.apps.create and secrets.reveal). Revoking needs LLM Gateway delete access (llm.apps.delete). Every key creation is recorded in the audit log.

3. Call a model​

Base URLhttps://models.quilrai.dev/v1
EndpointPOST /chat/completions (OpenAI Chat Completions format)
AuthAuthorization: Bearer $QUILR_MODEL_API_KEY
ModelThe exact catalog id, including any prefix (for example deepseek/deepseek-v3.2, gpt-oss-120b). It must be one of the key's allowed models.
export QUILR_MODEL_API_KEY="sk-quilrllm-..."
curl 'https://models.quilrai.dev/v1/chat/completions' \
-H "Authorization: Bearer $QUILR_MODEL_API_KEY" \
-H "Content-Type: application/json" \
--data '{
"model": "deepseek/deepseek-v3.2",
"messages": [
{
"role": "user",
"content": "Explain zero trust in one sentence."
}
],
"temperature": 0.7,
"top_p": 1,
"max_tokens": 1024
}'
tip

The Playground's Use this model in your app panel generates these snippets for the selected model, prompt, and settings.

Use through an LLM Gateway app​

Put QuilrAI-provided models behind an LLM Gateway app to get the app's guardrails, routing, limits, and logs. The app reaches them through the General LLM provider (an OpenAI-compatible custom endpoint).

  1. Create a model API key whose Allowed models include every model the app should use.

  2. In the LLM Gateway, create an app (or edit an existing one) and add a General LLM provider, either as a global provider or with app-specific credentials:

    FieldValue
    base_urlhttps://models.quilrai.dev/v1
    api_keyYour model API key (sk-quilrllm-...)
    ModelsThe exact catalog ids, for example deepseek/deepseek-v3.2
  3. Call the app with its QuilrAI gateway key (sk-quilr-...) as in the Quick Start, using the catalog id as model.

See Provider Support for the General LLM provider's capabilities.

Usage is billed from your organization credit either way, and shows up on the model API key's spend and the Usage tab.

note

Native integration, selecting QuilrAI-provided models directly in Create App without a General LLM provider, is coming soon.

Try a model in the Playground​

Go to Settings > Models > Playground (or click Chat on a catalog row). Playground requests count toward your organization's usage, and responses are not stored by the console. You do not need to create a key; the Playground uses its own server-held key.

Models Playground

SettingRangeDefault
ModelAny active catalog model
System promptUp to 16,000 characters
Temperature0 to 20.7
Top P0 to 11
Max output tokens1 to 4,0961,024
User promptUp to 16,000 characters

Conversations are limited to 20 messages and 64,000 characters. Playground responses are not streamed.

Request settings and the Use this model in your app panel

Track usage​

Settings > Models > Usage shows your credit and spend. Usage can take a few moments to appear after a request completes; click Refresh to update.

TileShows
Credit remainingRemaining credit, and amount spent of your lifetime limit
Spend this monthSpend in the current UTC month
Requests this monthSuccessful and failed requests
Tokens this monthInput and output tokens
Total throughputTokens per second, input and output

Below the tiles, a per-model table lists requests, tokens, input/output/total TPS, average latency, error rate, month spend, and lifetime spend. Models that have only lifetime spend stay listed with zeros for the month. If a banner says some token totals are estimated, the provider did not report exact usage for those requests.

Per-key spend for the current month is on the API keys tab.

Model catalog​

Catalog as of 2026-10-02; the console shows the current list and prices. Prices are USD per 1M tokens. API schema is informational: every model is called the same way.

Model (use as model)CapabilitiesInputCached inputOutputAPI schema
anthracite-org/magnum-v4-72bChat, Streaming$5Not available$10v1
baidu/ernie-4.5-vl-424b-a47bChat, Streaming$0.84Not available$2.5v1
bytedance/ui-tars-1.5-7bChat, Streaming$0.2$0.2$0.4v1
cognitivecomputations/dolphin-mistral-24b-venice-editionChat, Streaming$0.4Not available$1.8v1
deepseek-r1-distill-qwen-32bChat, Streaming$0.994Not available$9.762v2
deepseek-v4-flash-0731Chat, Streaming$0.88$0.028$2.64v2
deepseek-v4-pro-0813Chat, Streaming$2.64$0.088$7.92v2
deepseek/deepseek-chatChat, Streaming$0.64Not available$1.78v1
deepseek/deepseek-chat-v3-0324Chat, Streaming$0.48$0.27$1.8v1
deepseek/deepseek-chat-v3.1Chat, Streaming$0.5$0.26$1.9v1
deepseek/deepseek-r1Chat, Streaming$1.4Not available$5v1
deepseek/deepseek-r1-0528Chat, Streaming$1$0.7$4.3v1
deepseek/deepseek-r1-distill-llama-70bChat, Streaming$1.6Not available$1.6v1
deepseek/deepseek-v3.1-terminusChat, Streaming$0.54$0.27$2v1
deepseek/deepseek-v3.2Chat, Streaming$0.52$0.26$0.76v1
deepseek/deepseek-v3.2-expChat, Streaming$0.54Not available$0.82v1
deepseek/deepseek-v4-flash-vision-expChat, Streaming$0.4312$0.1372$1.2936v1
gemma-4-26b-a4b-itChat, Streaming$0.2Not available$0.6v2
gemma-sea-lion-v4-27b-itChat, Streaming$0.702Not available$1.11v2
glm-4.7-flashChat, Streaming$0.121Not available$0.8v2
glm-5.2Chat, Streaming$2.8$0.52$8.8v2
glm-5.3Chat, Streaming$2.8$0.52$8.8v1
glm-5.3-flashChat, Streaming$0.3$0.06$1v1
google/gemma-2-27b-itChat, Streaming$1.3Not available$1.3v1
google/gemma-3-12b-itChat, Streaming$0.1Not available$0.3v1
google/gemma-3-27b-itChat, Streaming$0.16Not available$0.32v1
google/gemma-3-4b-itChat, Streaming$0.1Not available$0.2v1
google/gemma-4-31b-itChat, Streaming$0.18$0.1$0.68v1
gpt-oss-120bChat, Streaming$0.7Not available$1.5v2
gpt-oss-120b-ultrafastChat, Streaming$0.7Not available$1.5v1
gpt-oss-20bChat, Streaming$0.4Not available$0.6v2
granite-4.0-h-microChat, Streaming$0.034Not available$0.224v2
gryphe/mythomax-l2-13bChat, Streaming$0.12Not available$0.12v1
ibm-granite/granite-4.1-8bChat, Streaming$0.1$0.1$0.2v1
ibm-granite/granite-4.2-8bChat, Streaming$0.2$0.1$0.3v1
inclusionai/ling-3.0-flashChat, Streaming$0.042$0.0084$0.126v1
kimi-k2.6Chat, Streaming$1.9$0.32$8v2
kimi-k2.7-codeChat, Streaming$1.9$0.38$8v2
llama-3.1-8b-instruct-fp8Chat, Streaming$0.304Not available$0.574v2
llama-3.2-11b-vision-instructChat, Streaming$0.097Not available$1.352v2
llama-3.2-1b-instructChat, Streaming$0.054Not available$0.402v2
llama-3.2-3b-instructChat, Streaming$0.1018Not available$0.67v2
llama-3.3-70b-instruct-fp8-fastChat, Streaming$0.586Not available$4.506v2
llama-4-scout-17b-16e-instructChat, Streaming$0.54Not available$1.7v2
llama-guard-3-8bChat, Streaming$0.968Not available$0.06v2
meta-llama/llama-3.1-70b-instructChat, Streaming$0.8Not available$0.8v1
meta-llama/llama-4-maverickChat, Streaming$0.4Not available$1.392v1
meta-llama/llama-guard-4-12bChat, Streaming$0.36Not available$0.36v1
meta/muse-glimmer-30bChat, Streaming$0.6$0.08$2.2v1
microsoft/phi-4Chat, Streaming$0.14Not available$0.28v1
microsoft/wizardlm-2-8x22bChat, Streaming$1.24Not available$1.24v1
minimax/minimax-m2Chat, Streaming$0.51Not available$2.04v1
minimax/minimax-m2.1Chat, Streaming$0.6$0.06$2.4v1
minimax/minimax-m2.5Chat, Streaming$0.54$0.06$1.9v1
minimax/minimax-m2.7Chat, Streaming$0.48Not available$1.92v1
minimax/minimax-m3Chat, Streaming$0.46$0.1$1.92v1
mistral-small-3.1-24b-instructChat, Streaming$0.702Not available$1.11v2
mistralai/devstral-2512Chat, Streaming$0.8$0.08$4v1
mistralai/ministral-14b-2512Chat, Streaming$0.4$0.04$0.4v1
mistralai/ministral-3b-2512Chat, Streaming$0.2$0.02$0.2v1
mistralai/ministral-8b-2512Chat, Streaming$0.3$0.03$0.3v1
mistralai/mistral-nemoChat, Streaming$0.038Not available$0.06v1
mistralai/mistral-small-24b-instruct-2501Chat, Streaming$0.1Not available$0.16v1
mistralai/mistral-small-2603Chat, Streaming$0.3$0.03$1.2v1
mistralai/mistral-small-3.2-24b-instructChat, Streaming$0.15Not available$0.4v1
mistralai/mixtral-8x22b-instructChat, Streaming$4.4$0.44$13.2v1
mistralai/voxtral-small-24b-2507Chat, Streaming$0.2$0.02$0.6v1
moonshotai/kimi-k2Chat, Streaming$1.14Not available$4.6v1
moonshotai/kimi-k2-0905Chat, Streaming$1.2Not available$5v1
moonshotai/kimi-k2-thinkingChat, Streaming$1.2Not available$5v1
moonshotai/kimi-k2.5Chat, Streaming$0.9$0.14$4.5v1
moonshotai/kimi-k3Chat, Streaming$5.2$0.58$26v1
nemotron-3-120b-a12bChat, Streaming$1Not available$3v2
nex-agi/nex-n2-proChat, Streaming$1$0.5$5v1
nousresearch/hermes-3-llama-3.1-405bChat, Streaming$2Not available$2v1
nousresearch/hermes-3-llama-3.1-70bChat, Streaming$1.4Not available$1.4v1
nousresearch/hermes-4-405bChat, Streaming$2Not available$6v1
nousresearch/hermes-4-70bChat, Streaming$0.26Not available$0.8v1
nvidia/nemotron-3-nano-30b-a3bChat, Streaming$0.1$0.06$0.4v1
nvidia/nemotron-3-ultra-550b-a55bChat, Streaming$1$0.2$4.4v1
nvidia/nemotron-3.5-lightningChat, Streaming$0.16$0.08$0.4v1
openai/gpt-oss-safeguard-20bChat, Streaming$0.15$0.075$0.6v1
qwen-3.8-27b-ultrafastChat, Streaming$1.98Not available$2.98v1
qwen/qwen-2.5-72b-instructChat, Streaming$0.72Not available$0.8v1
qwen/qwen-2.5-7b-instructChat, Streaming$0.2Not available$0.4v1
qwen/qwen2.5-vl-72b-instructChat, Streaming$0.5Not available$1.5v1
qwen/qwen3-14bChat, Streaming$0.2Not available$0.44v1
qwen/qwen3-235b-a22b-2507Chat, Streaming$0.18Not available$1.1v1
qwen/qwen3-235b-a22b-thinking-2507Chat, Streaming$0.6Not available$6v1
qwen/qwen3-30b-a3b-instruct-2507Chat, Streaming$0.18Not available$0.6v1
qwen/qwen3-32bChat, Streaming$0.16Not available$0.56v1
qwen/qwen3-coderChat, Streaming$0.6$0.2$2v1
qwen/qwen3-coder-30b-a3b-instructChat, Streaming$0.14Not available$0.54v1
qwen/qwen3-coder-nextChat, Streaming$0.24$0.14$1.6v1
qwen/qwen3-next-80b-a3b-instructChat, Streaming$0.18Not available$2.2v1
qwen/qwen3-next-80b-a3b-thinkingChat, Streaming$0.3Not available$2.4v1
qwen/qwen3-vl-235b-a22b-instructChat, Streaming$0.4$0.22$1.76v1
qwen/qwen3-vl-235b-a22b-thinkingChat, Streaming$1.96Not available$7.9v1
qwen/qwen3-vl-30b-a3b-instructChat, Streaming$0.3Not available$1.2v1
qwen/qwen3-vl-30b-a3b-thinkingChat, Streaming$0.58Not available$2v1
qwen/qwen3-vl-8b-instructChat, Streaming$0.5$0.24$1.5v1
qwen/qwen3.5-122b-a10bChat, Streaming$0.52Not available$4.16v1
qwen/qwen3.5-27bChat, Streaming$0.5Not available$4v1
qwen/qwen3.5-35b-a3bChat, Streaming$0.28$0.1$2v1
qwen/qwen3.5-397b-a17bChat, Streaming$0.9$0.44$6v1
qwen/qwen3.5-9bChat, Streaming$0.2Not available$0.3v1
qwen/qwen3.6-27bChat, Streaming$0.64$0.3$5.4v1
qwen/qwen3.6-35b-a3bChat, Streaming$0.2$0.1$1.8v1
qwen/qwen3.8-2.4t-a95bChat, Streaming$4$0.4$12v1
qwen2.5-coder-32b-instructChat, Streaming$1.32Not available$2v2
qwen3-30b-a3b-fp8Chat, Streaming$0.1018Not available$0.67v2
qwen3.8-27bChat, Streaming$0.9$0.1$6.4v2
qwq-32bChat, Streaming$1.32Not available$2v2
rekaai/reka-edgeChat, Streaming$0.2Not available$0.2v1
rekaai/reka-flash-3Chat, Streaming$0.2Not available$0.4v1
sao10k/l3-lunaris-8bChat, Streaming$0.08Not available$0.1v1
sao10k/l3.1-euryale-70bChat, Streaming$1.7Not available$1.7v1
sao10k/l3.3-euryale-70bChat, Streaming$1.3Not available$1.5v1
stepfun/step-3.5-flashChat, Streaming$0.2Not available$0.6v1
stepfun/step-3.7-flashChat, Streaming$0.32$0.064$1.84v1
tencent/hunyuan-a13b-instructChat, Streaming$0.28Not available$1.14v1
tencent/hy-mt2-1.8bChat, Streaming$0.088Not available$0.354v1
tencent/hy-mt2-30b-a3bChat, Streaming$0.148Not available$0.59v1
tencent/hy-mt2-7bChat, Streaming$0.148Not available$0.59v1
tencent/hy3Chat, Streaming$0.28$0.07$1.16v1
thedrummer/cydonia-24b-v4.1Chat, Streaming$0.6$0.3$1v1
thedrummer/skyfall-36b-v2Chat, Streaming$1.1$0.5$1.6v1
thedrummer/unslopnemo-12bChat, Streaming$0.8Not available$0.8v1
thinkingmachines/inklingChat, Streaming$1.9$0.32$8.1v1
thinkingmachines/inkling-smallChat, Streaming$0.9$0.2$2.4v1
undi95/remm-slerp-l2-13bChat, Streaming$0.7Not available$1.3v1
xiaomi/mimo-v2.5Chat, Streaming$0.336$0.0068$0.672v1
xiaomi/mimo-v2.5-proChat, Streaming$0.96048$0.007912$1.92096v1
z-ai/glm-4.5Chat, Streaming$1.2$0.22$4.4v1
z-ai/glm-4.5-airChat, Streaming$0.26$0.05$1.7v1
z-ai/glm-4.5vChat, Streaming$1.2$0.22$3.6v1
z-ai/glm-4.6Chat, Streaming$0.86$0.16$3.5v1
z-ai/glm-4.6vChat, Streaming$0.6$0.11$1.8v1
z-ai/glm-4.7Chat, Streaming$0.8$0.16$3.5v1
z-ai/glm-5Chat, Streaming$1.2$0.24$4.16v1
z-ai/glm-5.1Chat, Streaming$2.1$0.41$7v1