Security guardrails
Detect sensitive data and adversarial input in prompts and responses, then monitor, redact or block it. Set the basics per app in the app's settings; use the Policy Engine's Data & Adversarial Risks card when you need rules that depend on who is calling, which model, or which part of the request.
Turn it on for an app
The Guardrails tab has six parts:
Defaults for a new app
Create App has no guardrail step. Every new app starts with these guardrails:
Because the default action is Monitor, a new app records detections without changing traffic. Review findings before you change any category to Redact or Block.
Actions
- Adversarial risks and the hallucination check support Block or Monitor only. There is no value to redact.
- Action resolution: the category's own action takes precedence, then the app's default action, then Monitor.
- Set everything to default clears per-category actions so every category follows the default action.
- The Action mix panel counts how many of the 20 categories (7 data + 13 adversarial) use each action or are off.
Data risks
Each category has these settings:
Risk level and sub-category sensitivity
Every sub-category has a sensitivity. The category's risk level determines which sensitivities trigger detection: a low risk level detects only clearly sensitive values, and a high risk level also detects weaker contextual signals.
Example for PII, where passport is a high-sensitivity sub-category and name, home address and email are low-sensitivity:
Change one sub-category's sensitivity to tune that value without changing the whole category.
The app's risk level wins over the organization-wide setting in Detection Models. When the Policy Engine is enabled, the policy wins.
Adversarial risks
Each category has an on/off switch and a Block | Monitor action. Its scope is fixed to the side of the conversation it targets.
Enable all turns every adversarial category on.
Precision detections
Exact-match patterns for structured identifiers. Use them when you need a specific identifier detected in addition to the contextual data risk categories. Each detection has its own switch and action (Redact, Partial redact, Block or Monitor; default Monitor), and all are off until you turn them on. A red dot marks a high-sensitivity identifier.
The built-in catalog has more than 80 identifiers:
To match your own identifiers with a regex, create a Custom Detection instead.
Hallucination check
Flags responses the gateway scores as likely fabricated, using a fixed threshold of 0.8.
It runs on non-streaming responses only, because a streamed response cannot be blocked once it has started. To set a different threshold per app or group, use the Policy Engine's Hallucination protection card.
Source IP restrictions
Accept gateway calls only from listed networks. Turn on Enabled and enter Allowed source IPs as comma-separated IPv4, IPv6 or CIDR values, for example 203.0.113.10, 10.0.0.0/8. Calls from other addresses are rejected.
Endpoint coverage
- Streaming: request-side checks run as normal. Response-side redaction and blocking apply only when the gateway holds the full response, so streamed chunks pass through unmodified.
- Copilot Studio cannot accept rewritten tool input, so Redact and Partial redact become Block. See Copilot Studio.
- Realtime sessions log the handshake, byte counters and usage, but do not run DLP. Use SDK Mode to scan Realtime transcripts out of band.
Going further with the Policy Engine
The same controls live on the Data & Adversarial Risks card in Policy Engine > LLM Gateway. When the engine is on for the LLM Gateway, the Guardrails tab freezes and the card decides what happens on live requests. See What happens to classic settings. Edits join a shared draft and apply once you publish a revision.
A detection rule picks data types (a whole category, single types or custom detections), a findings threshold (at least N), an action (Monitor, Partial redact, Redact, Block), a stage (Request, Response, Both) and an optional severity that is reported but never changes the action. The card applies on the assistants, bedrock, chat, copilot, embeddings, rerank, responses, sdk_check, stt, text, tts and vertex API surfaces.
Add more lines to one rule with + Add line. Each line keeps its own types, threshold, action and stage, and all lines share the configuration's scope and priority. Adding a rule copies the scope from the rule above it.
Scenarios the card supports that app settings cannot express:
-
Block secrets for everyone, monitor PII for one team. Scope rules to Everyone, People, a Smart group, an Application or App tag. Narrower scopes take precedence, and the highest priority wins per data type.
-
Different rules per model or provider. Scope by Requested model, Provider, API surface or Environment, for example redact PHI only on requests to one provider.
-
Thresholds. Act only when a request contains at least N findings of a type, such as 5 or more email addresses.
-
Tool-call arguments. Turn on Scan tool-call arguments to evaluate each tool call's arguments on their own (Monitor and Block only).
-
Language blocking. Monitor or block passages outside a list of allowed languages (streamed responses are skipped).
-
Several data types, several actions. One configuration for Support Copilot redacts Aadhaar numbers, records names and blocks secrets. A request carrying all three has the Aadhaar redacted, the name left alone but recorded, and the whole call refused because of the key.
runs on requestpriority 700
Redact and Partial redact rewrite only the findings their own rule selected, and a finding no rule selects is left unchanged. Where two rules select the same finding, the higher priority wins; at equal priority the more restrictive action wins.
Hallucination scoring and source IP lists move to their own cards: Hallucination protection and Identity and network trust.
Related
- Custom detections - your own regex and intent detections.
- Guardian Agent - dependency checks and task adherence.
- Policy Engine overview - how cards, scopes and priorities work.