Skip to main content

Security guardrails

Detect sensitive data and adversarial input in prompts and responses, then monitor, redact or block it. Set the basics per app in the app's settings; use the Policy Engine's Data & Adversarial Risks card when you need rules that depend on who is calling, which model, or which part of the request.

Turn it on for an app​

Go toQuilrAI consoleSettingsAI GatewayLLM Gatewayyour appConfigureGuardrails

The Guardrails tab has six parts:

PartWhat it does
Default actionThe action used where a category does not set its own.
Data risksSensitive data such as PII, PHI, card data and secrets.
Adversarial risksPrompt injection, jailbreaks and harmful content.
Precision detectionsExact-match identifiers such as SSN, Aadhaar or IBAN.
Hallucination checkFlags responses that score as likely fabricated.
Source IP restrictionsAccepts calls only from listed networks.

Defaults for a new app​

Create App has no guardrail step. Every new app starts with these guardrails:

SettingDefault
Default actionMonitor
Data risksAll 6 on: PII, PHI, PFI, PCI, Insurance data, Authentication secrets
Adversarial risks12 of 13 on. Malicious scripts is off (opt-in).
Data risk scopeRequest and response
Precision detections, hallucination check, source IP restrictionsOff
Guardian AgentOff. See Guardian Agent.

Because the default action is Monitor, a new app records detections without changing traffic. Review findings before you change any category to Redact or Block.

Actions​

ActionWhat happensAvailable on
MonitorThe request passes unchanged and the detection is logged.All categories
Partial redactMasks part of each detected value and forwards the rest.Data risks, precision detections
RedactReplaces each detected value and forwards the request.Data risks, precision detections
BlockRejects the whole request.All categories
  • Adversarial risks and the hallucination check support Block or Monitor only. There is no value to redact.
  • Action resolution: the category's own action takes precedence, then the app's default action, then Monitor.
  • Set everything to default clears per-category actions so every category follows the default action.
  • The Action mix panel counts how many of the 20 categories (7 data + 13 adversarial) use each action or are off.

Data risks​

CategorySub-categories
Personally Identifiable Information (PII)Date of birth, driver's license number, email address, employee ID, home address, name, national ID, passport number, phone number, social security number, other PII
Protected Health Information (PHI)Medical appointment, condition, facility, record number, treatment, prescription, other PHI
Payment and Financial Information (PFI)Bank account number, bank identification code, customer ID or account number, financial amount, invoice number, PAN card, payment processor detail, tax information, transaction ID, other PFI
Protected Card Information (PCI)Credit/debit card
Insurance DataHealth insurance, insurance policy
Auth & SecretsUsername, username or alias, access token, API key, AWS credentials, password, other secret
Code Scripts and Queries (off by default)Shell, database queries, YAML/config, Python, JavaScript/TypeScript, Java/Kotlin/Scala, Go, Rust, C/C++/C#, PHP/Ruby, Swift, other code

Each category has these settings:

SettingOptionsDefault
EnabledOn / offOn (Code Scripts and Queries off)
Risk levelNone, Low, Medium, HighHigh
Applies toRequest only, Request & response, Response onlyRequest & response
ActionRedact, Partial redact, Block, MonitorApp default action
Sub-category sensitivityLow, Medium, High per sub-categorySet per sub-category

Risk level and sub-category sensitivity​

Every sub-category has a sensitivity. The category's risk level determines which sensitivities trigger detection: a low risk level detects only clearly sensitive values, and a high risk level also detects weaker contextual signals.

Risk levelDetects sub-categories marked
LowHigh
MediumHigh, Medium
HighHigh, Medium, Low

Example for PII, where passport is a high-sensitivity sub-category and name, home address and email are low-sensitivity:

RequestRisk level LowRisk level High
My name is Jane Doe and I live in BengaluruAllowedDetected (NAME, HOME ADDRESS)
Reach me at jane@example.comAllowedDetected (EMAIL ADDRESS)
My passport number is M1234567DetectedDetected

Change one sub-category's sensitivity to tune that value without changing the whole category.

The app's risk level wins over the organization-wide setting in Detection Models. When the Policy Engine is enabled, the policy wins.

Adversarial risks​

Each category has an on/off switch and a Block | Monitor action. Its scope is fixed to the side of the conversation it targets.

CategoryChecked onDefault
Prompt Injection TechniquesRequestOn
Jailbreak TechniquesRequestOn
Prompt Context CorruptionRequestOn
Semantic Adversarial PromptsRequestOn
Social Engineering PromptsRequestOn
System, Guardrail & Security DisclosureRequestOn
Security Exploit & Payload EnablementRequestOn
Cybersecurity Frameworks & Standards MentionRequestOn
Response RisksResponseOn
Hateful or Offensive ContentBothOn
Violence & Harmful ContentBothOn
Fraudulent or Illegal Activity ContentBothOn
Malicious ScriptsRequestOff (opt-in)

Enable all turns every adversarial category on.

Precision detections​

Exact-match patterns for structured identifiers. Use them when you need a specific identifier detected in addition to the contextual data risk categories. Each detection has its own switch and action (Redact, Partial redact, Block or Monitor; default Monitor), and all are off until you turn them on. A red dot marks a high-sensitivity identifier.

The built-in catalog has more than 80 identifiers:

GroupIdentifiers
Personal identitySSN, Aadhaar, PAN card, driver's license, passport, national ID, phone, email, CA social insurance number, UK national insurance number, pensioner/student/senior citizen ID, vehicle number, VIN, vehicle registration certificate, customer signature, photographic image, age, nationality, language, gender orientation, marital status, father's, mother's and maiden names, mother's maiden name, anniversary date, current location, latitude/longitude, local reference
HealthMedical record number, medical registration number, patient ID, Medicare number, health insurance, insurance policy, CA health number, UK NHS number, blood group
FinancialBank account number, account holder's name and signature, routing number, bank identification code, SWIFT code, IBAN, UPI ID, customer ID or account number, credit score, tax information, UK unique taxpayer reference, monthly/annual income
CardCard number, expiry date, CVV, cardholder name
SecretsPIN
Device and networkIP address, MAC address, SIM number, IMEI, IMSI, handset make/model, ADID/IDFA, cookies
Telecom subscriberCall data records, SMS, roaming, voice and VAS pattern records, PUK code, unique portability code, SIM contacts, credit history and limit, last billed and unbilled amounts, previous payments, usage details, services subscribed, talk plan
Employee and HRSalary details and components, CTC, PF account number, investment details, education, designation, department, employee type, date of joining, work experience, biometric information

To match your own identifiers with a regex, create a Custom Detection instead.

Hallucination check​

Flags responses the gateway scores as likely fabricated, using a fixed threshold of 0.8.

SettingOptionsDefault
EnabledOn / offOff
ActionBlock, MonitorMonitor
Risk levelLow, Medium, HighMedium

It runs on non-streaming responses only, because a streamed response cannot be blocked once it has started. To set a different threshold per app or group, use the Policy Engine's Hallucination protection card.

Source IP restrictions​

Accept gateway calls only from listed networks. Turn on Enabled and enter Allowed source IPs as comma-separated IPv4, IPv6 or CIDR values, for example 203.0.113.10, 10.0.0.0/8. Calls from other addresses are rejected.

Endpoint coverage​

SurfaceRequestResponse
Chat completions, including Bedrock Converse, Vertex Gemini and Anthropic models reached through translationYesYes
Native Anthropic MessagesYesYes
OpenAI ResponsesYesYes
Bedrock Runtime (boto3)YesYes. converse_stream is request only.
Native Vertex / Gemini generateContentYesYes
Embeddings, text-to-speechInput-
Speech-to-text-Output
Sarvam speech synthesisInput-
Sarvam transcription-Output
Sarvam translation, transliteration, language detectionYesYes
Copilot StudioUser context and tool inputs-
OpenAI Realtime (websocket)Not scannedNot scanned
  • Streaming: request-side checks run as normal. Response-side redaction and blocking apply only when the gateway holds the full response, so streamed chunks pass through unmodified.
  • Copilot Studio cannot accept rewritten tool input, so Redact and Partial redact become Block. See Copilot Studio.
  • Realtime sessions log the handshake, byte counters and usage, but do not run DLP. Use SDK Mode to scan Realtime transcripts out of band.

Going further with the Policy Engine​

The same controls live on the Data & Adversarial Risks card in Policy Engine > LLM Gateway. When the engine is on for the LLM Gateway, the Guardrails tab freezes and the card decides what happens on live requests. See What happens to classic settings. Edits join a shared draft and apply once you publish a revision.

A detection rule picks data types (a whole category, single types or custom detections), a findings threshold (at least N), an action (Monitor, Partial redact, Redact, Block), a stage (Request, Response, Both) and an optional severity that is reported but never changes the action. The card applies on the assistants, bedrock, chat, copilot, embeddings, rerank, responses, sdk_check, stt, text, tts and vertex API surfaces.

Detection line settingDefault
Data types (several types on one line match any of them)None
Findings threshold (at least N)1
ActionMonitor
StageRequest
Scan tool-call argumentsOff

Add more lines to one rule with + Add line. Each line keeps its own types, threshold, action and stage, and all lines share the configuration's scope and priority. Adding a rule copies the scope from the rule above it.

Scenarios the card supports that app settings cannot express:

  • Block secrets for everyone, monitor PII for one team. Scope rules to Everyone, People, a Smart group, an Application or App tag. Narrower scopes take precedence, and the highest priority wins per data type.

  • Different rules per model or provider. Scope by Requested model, Provider, API surface or Environment, for example redact PHI only on requests to one provider.

  • Thresholds. Act only when a request contains at least N findings of a type, such as 5 or more email addresses.

  • Tool-call arguments. Turn on Scan tool-call arguments to evaluate each tool call's arguments on their own (Monitor and Block only).

  • Language blocking. Monitor or block passages outside a list of allowed languages (streamed responses are skipped).

  • Several data types, several actions. One configuration for Support Copilot redacts Aadhaar numbers, records names and blocks secrets. A request carrying all three has the Aadhaar redacted, the name left alone but recorded, and the whole call refused because of the key.

support_copilot_data_actionsrequest

runs on requestpriority 700

Data rule 1
WhenApplicationisSupport Copilot
anddata foundis any ofAadhaar Number / VID
Then
Sensitive data actionredact
Risk levelhigh
Data rule 2
WhenApplicationisSupport Copilot
anddata foundis any ofName
Then
Sensitive data actionmonitor
Data rule 3
WhenApplicationisSupport Copilot
anddata foundis any ofAuth & Secrets
Then
Sensitive data actionblock
Risk levelcritical

Redact and Partial redact rewrite only the findings their own rule selected, and a finding no rule selects is left unchanged. Where two rules select the same finding, the higher priority wins; at equal priority the more restrictive action wins.

Hallucination scoring and source IP lists move to their own cards: Hallucination protection and Identity and network trust.