Skip to main content

Token Savings

V2 console

This card lives in Policy Engine > LLM Gateway at web.quilr.ai/policy. Edits join the shared draft and take effect once you review and publish a revision.

Rewrite request content into fewer tokens before it reaches the provider. Responses come back untouched.

Token Savings card collapsed with Smart JSON compression, HTML to text, Markdown to text and Text compression summary rows

Sections​

Each strategy is its own section with a Turn on button. All four apply on chat, responses and vertex.

SectionWhat it doesOff by default
Smart JSON compressionCompresses JSON in supported request bodies.Yes
HTML to textReduces HTML in supported request bodies to its text.Yes
Markdown to textReduces Markdown in supported request bodies to its text.Yes
Text compressionCompresses prose in supported request bodies.Yes

HTML to text and Markdown to text sections, each with a Turn on button and a shared configuration row

With no configuration, request bodies are sent to the provider as received.

Configuration settings​

Any Turn on button opens one New Token Savings configuration dialog with all four controls, preset to the one you clicked.

Controls section with Leave as is, On and Off for each of the four strategies

SettingOptionsDefaultNotes
Applies toEveryone, People, Smart group, Application, App tag, Requested model, Provider, API surface, Environment, Prompt complexity, Prompt text, Tool, Source network, Except...EveryoneOr pick an application directly.
Each strategyLeave as is, On, OffLeave as is (On for the strategy you clicked)Leave as is keeps the broader configuration's value. Off switches a strategy off for a narrower scope.
SeverityNot set to Very criticalNot setReported only.

Example​

compress_document_ingestionrequest

runs on requestpriority 500

WhenApplicationisDoc Ingestion Pipeline
Then
JSON compressiontrue
HTML to texttrue
Markdown to texttrue

A pipeline that sends scraped pages and API payloads turns on the three structural strategies and leaves prose untouched.

Scoping and precedence​

  • Highest-priority configuration wins per strategy. New configurations start at 500.
  • One configuration can cover several apps; it shows as a shared_* row on each section.

Legacy app setting​

Token Saving in the app's settings. For how the strategies work, see Token Saving.