Skip to main content

Token saving

Rewrite request content into fewer tokens before it reaches the provider. Responses are returned untouched, and your code does not change.

Turn it on for an app​

Go toQuilrAI consoleSettingsAI GatewayLLM Gatewayyour appConfigureToken Saving

Turn on each strategy that matches the traffic the app sends. Apps with any strategy on show a Token saving flag in the app list.

Strategies​

StrategyWhat it doesBeforeAfter
Smart JSON compressionCompacts eligible JSON objects and arrays, converting to TOON where that saves more. Best for tool results and structured data. Up to about 20% savings.{"name": "John", "age": 30}name:John|age:30
HTML to textStrips HTML markup down to readable text.<p><b>Hello</b> world</p>Hello world
Markdown to textStrips Markdown syntax that costs tokens without adding meaning.## Hello **world**Hello world
Text compressionRemoves low-value prose and separator noise from long text while keeping its meaning. Structured-looking lines are left alone.Please review the following statement and the context which was actually very repetitive.Review statement and context.

A transform only applies when it reduces token usage. Savings appear as Tokens saved in the app's Analytics tab and per request in Activity > Requests, and tenant-wide in Costs and savings.

Turn on only the methods that match your traffic. A retrieval app that sends JSON search results and Markdown snippets can enable Smart JSON compression and Markdown to text without enabling every transform.

Shared with the MCP Gateway​

The same four methods, with the same names in settings, logs and analytics, are available for MCP tool output. On the MCP Gateway they rewrite tools/call text results before they return to the client, and OneMCP smart tool search also reduces tool-schema tokens. See MCP Gateway token saving.

Setting keyUse it for
smart_json_compressionStructured JSON payloads, API responses and tool results.
html_to_textScraped pages, rich emails, dashboards and HTML reports.
markdown_to_textMarkdown documents, README files, issue bodies and generated notes.
text_compressionVerbose prose, repeated separator lines and low-signal plain text.

For the Management API, the app's switches are (all off by default):

{
"smart_json_compression": false,
"html_to_text": false,
"markdown_to_text": false,
"text_compression": false
}

Going further with the Policy Engine​

Token saving is also the Token Savings card in Policy Engine > LLM Gateway. When the engine is on for the LLM Gateway, the Token Saving tab freezes and the card applies instead. See What happens to classic settings.

Each strategy is a section with a Turn on button; all four apply on chat, responses and vertex. Each strategy is a three-way switch (Leave as is, On, Off), and the highest-priority configuration wins per strategy. Scenarios the card supports:

  • One configuration for several apps. Scope by App tag or several applications; it shows as a shared_* row on each section.
  • Exempt a narrow scope. Turn a strategy Off for one Requested model, Provider or Smart group while it stays on for Everyone.
  • Compress by content. Scope by Prompt complexity or Prompt text, for example only long prompts.
  • Structural strategies for a pipeline. A document-ingestion app turns on JSON, HTML and Markdown compression and leaves prose untouched:
compress_document_ingestionrequest

runs on requestpriority 500

WhenApplicationisDoc Ingestion Pipeline
Then
JSON compressiontrue
HTML to texttrue
Markdown to texttrue