Skip to main content

Security guardrails

Scan what agents send to an MCP server and what the server sends back, and block, redact or flag sensitive data and attacks.

Configure it on the server​

Go toQuilrAI consoleSettingsAI GatewayMCP Gatewayserver cardConfigureGuardrails

Request and response​

DirectionWhat is scanned
Request (tool arguments)The inputs an agent sends when it calls a tool.
Response (tool results)What the tool returns before the agent sees it.

Actions​

ActionWhat happens
BlockThe tool call is stopped.
RedactThe detected data is removed and the call proceeds.
Partial (Partial Redact in the category lists)Part of each match is masked and the rest stays visible for context.
MonitorThe call goes ahead and the detection is recorded.

The Actions card sets the defaults:

SettingMeaning
DefaultThe action for any detection that has no more specific action. Starts as Monitor.
Request (tool arguments)The action for detections in tool inputs. Default uses the setting above.
Response (tool results)The action for detections in tool results. Default uses the setting above.

Categories​

Select a category to scan for it, then optionally select its own Request and Response action. Leave a category on Default to use the Actions card.

GroupCategories
Data risk categoriesPII, PHI, PFI, PCI, Insurance, Auth Secrets, Device Network Online Identifiers, Telecom Subscriber Data, Employee HR Data
Adversarial categoriesPrompt Injection Techniques, Jailbreak Techniques, Prompt Context Corruption, Semantic Adversarial Prompts, Social Engineering Prompts, Response Risks, Hateful Or Offensive Content, Violence And Harmful Content, Fraudulent Or Illegal Activity Content, System Guardrail And Security Disclosure, Security Exploit And Payload Enablement, Cybersecurity Frameworks And Standards Mention

Each group shows how many categories are on, for example 3 of 9 on. The Coverage card totals both groups and Request actions across categories shows how many categories use each action.

Click Save settings in the footer to apply your changes.

Where detections show​

  • The Stopped by guardrails tile on the MCP Gateway page counts guardrail flags.
  • The Guardrails chip on the server card opens this section.
  • Overall analytics > Analytics shows Blocked and Guardrail flags.
  • Findings & Interactions lists each detection, filterable by input and output DLP.

Going further with the Policy Engine​

When the Policy Engine is on for the MCP Gateway, this section turns read-only and the Data & Adversarial Risks card (stages 3 and 4, Request and Response) in Govern > Policy Engine > MCP Gateway applies instead. Edit anyway changes the stored values, which are used only if the Policy Engine is disabled (see What happens to classic settings). Edits join a shared draft and apply once you publish a revision.

The card's effects are Actions per sensitive data type, Default sensitive data action, Sensitive data detectors and a risk level. Scenarios it supports that server settings cannot:

  • A different action per data type in one rule. One map can block secrets, redact national IDs and monitor names on the same server. Each data type resolves independently, so a rule about one type never erases another rule's action on a different type. Default sensitive data action covers any enabled category the map does not name.
  • Scope by caller, agent or tool. Apply stricter actions for one smart group, one AI client, or only the tools tagged as writes, instead of one setting for the whole server.
  • React to what was found. A rule with a data found condition (detections by exact catalog name) can set the sensitive data action or raise the call's risk level, for example on a response from one tool.
crm_data_actionsrequest

runs on requestpriority 700

WhenMCP nameisCustomer CRM
Then
Actions per sensitive data type3 data types
Auth & Secrets blockAadhaar Number / VID redactName monitor
Default sensitive data actionmonitor
note

Actions per sensitive data type, Default sensitive data action and Sensitive data detectors are evaluated before content is scanned, so a rule carrying one of them cannot also carry a data found condition. Put the data condition in a separate rule.