Token saving
Rewrite request content into fewer tokens before it reaches the provider. Responses are returned untouched, and your code does not change.
Turn it on for an app
Turn on each strategy that matches the traffic the app sends. Apps with any strategy on show a Token saving flag in the app list.
Strategies
A transform only applies when it reduces token usage. Savings appear as Tokens saved in the app's Analytics tab and per request in Activity > Requests, and tenant-wide in Costs and savings.
Turn on only the methods that match your traffic. A retrieval app that sends JSON search results and Markdown snippets can enable Smart JSON compression and Markdown to text without enabling every transform.
Shared with the MCP Gateway
The same four methods, with the same names in settings, logs and analytics, are available for MCP tool output. On the MCP Gateway they rewrite tools/call text results before they return to the client, and OneMCP smart tool search also reduces tool-schema tokens. See MCP Gateway token saving.
For the Management API, the app's switches are (all off by default):
{
"smart_json_compression": false,
"html_to_text": false,
"markdown_to_text": false,
"text_compression": false
}
Going further with the Policy Engine
Token saving is also the Token Savings card in Policy Engine > LLM Gateway. When the engine is on for the LLM Gateway, the Token Saving tab freezes and the card applies instead. See What happens to classic settings.
Each strategy is a section with a Turn on button; all four apply on chat, responses and vertex. Each strategy is a three-way switch (Leave as is, On, Off), and the highest-priority configuration wins per strategy. Scenarios the card supports:
- One configuration for several apps. Scope by App tag or several applications; it shows as a
shared_*row on each section. - Exempt a narrow scope. Turn a strategy Off for one Requested model, Provider or Smart group while it stays on for Everyone.
- Compress by content. Scope by Prompt complexity or Prompt text, for example only long prompts.
- Structural strategies for a pipeline. A document-ingestion app turns on JSON, HTML and Markdown compression and leaves prose untouched:
runs on requestpriority 500