Skip to main content

Token saving

Compress tool results before they reach the model, so agents use fewer tokens on each MCP call.

Configure it on the server​

Go toQuilrAI consoleSettingsAI GatewayMCP Gatewayserver cardConfigureToken Saving

The Strategies card reads: "Rewrites tool results on the way back to the client to cut model token usage." New servers have every strategy off.

Strategies​

StrategyWhat it doesTurn it on when the server returns
Smart JSON compressionCompacts verbose JSON tool results before they reach the model.Large JSON objects or lists.
HTML to textStrips HTML markup from tool results and keeps the readable text.Web pages, HTML emails or reports.
Markdown to textFlattens Markdown formatting in tool results.Markdown documents or wiki pages.
Text compressionCompresses long text results while preserving meaning.Long plain text.

Turn on the strategies that match what the server returns, then click Save settings in the footer. The Active strategies bar and the section list (for example 2 of 4 strategies on) show how many are on.

Token saving runs after guardrails, so detections are made on the full tool result. The four strategies are the same methods the LLM Gateway applies to requests; see LLM Gateway token saving for a before-and-after example of each.

Where savings show​

PlaceWhat you see
Server cardToken saving chip (Yes / No) and tokens saved for the period.
Overall analytics > AnalyticsThe Tokens saved tile across servers.
Costs & SavingsSavings from the MCP channel.

OneMCP saves tokens too​

OneMCP reduces the tokens agents use to discover tools. With dynamic tool calling on, an agent receives a small set of tools instead of every server's full tool list:

ToolPurpose
list_mcp_connectionsLists the MCP servers the person can use and whether each is connected.
find_relevant_toolsFinds the tools that fit the current task.
call_toolCalls the chosen tool.

Set it with the OneMCP endpoint button in the MCP Gateway header, under Dynamic tool calling (User preference, Always on or Always off).

Going further with the Policy Engine​

When the Policy Engine is on for the MCP Gateway, this section turns read-only and the Token Savings card (stage 4, Response) in Govern > Policy Engine > MCP Gateway applies instead. Edit anyway changes the stored values, which are used only if the Policy Engine is disabled (see What happens to classic settings). The card has one effect per strategy: smart JSON compression, HTML to text, Markdown to text and text compression.

Scenarios the card supports that server settings cannot:

  • Compress by tool, not by server. Match tool name or tags, so only the tools that return large payloads are compressed.
  • Compress for some callers or agents. Match smart groups or agent name, for example compress results only for an agent with a small context window.
  • Compress on one route. Match route kind to treat OneMCP traffic differently from direct connections.
compress_web_searchresponse

runs on responsepriority 400

WhenMCP nameisWeb Search
Then
smart JSON compressiontrue
HTML to texttrue

The OneMCP side has its own card. OneMCP Features (stage 1, Session) sets OneMCP dynamic tools and memory per caller, for example turning dynamic tools off for one smart group.