Skip to main content

Detection Models

Detection Models define what counts as sensitive data or an adversarial prompt. Policies and controls refer to these detections by name, so this page is where you decide what Quilr looks for, and how risky each finding is by default.

Go toQuilrAI ConsoleGovernDetection Models

The page has two tabs: Data Risks and AI Adversarial Risks.

Data Risks​

Data Risks has two views:

ViewContains
Out-of-the-boxBuilt-in detections, split into Contextual and Non-contextual
CustomYour own detection models. See Custom detections and library.

Contextual​

Contextual detections use the surrounding text to decide whether something is sensitive. They are grouped into categories such as:

  • PII (personally identifiable information)
  • PHI (protected health information)
  • PFI (payment and financial information)
  • PCI (protected card information)
  • Insurance Data
  • Auth & Secrets
  • Code Scripts and Queries

Each category row shows how many subcategories it contains, its Risk level (None, Low, Medium or High), a View & test action and an enable toggle. Expanding a category shows each subcategory with a priority of Low, Medium or High.

Non-contextual​

Non-contextual detections match well-defined entities by format, for example SSN, Aadhaar, PAN, passport number, phone number, email address, IBAN, UPI ID, card number and CVV. Each entity has an Enabled toggle.

AI Adversarial Risks​

Adversarial detections look for attacks on the model rather than sensitive data. They are grouped by technique:

GroupTechniques
Prompt Injection6
Jailbreak6
Prompt Context Corruption1
Semantic Adversarial Prompts3
Social Engineering Prompts2
Response Risks8

Each group shows its risk level, Show techniques to list the individual techniques, and View & test.

Risk levels​

"Risk level" means different things in different places. Check which one you are looking at.

WhereUI label and valuesWhat it does
Data Risks > Contextual, per categoryRisk level: None, Low, Medium, HighA detection threshold, not a severity. It decides which subcategory priorities are detected: None detects nothing, Low only High-priority subcategories, Medium High and Medium, High all three.
Data Risks > Contextual, per subcategorypriority: Low, Medium, HighHow strong a signal the subcategory is. Combined with the category's risk level above.
AI Adversarial Risks, per techniqueRisk level badgeThe default severity of that technique. Shown for reference.
LLM Gateway app Guardrails tabRisk level per category, with sub-category sensitivityThe same threshold model as Contextual, set per app. See Risk level and sub-category sensitivity.
Policy Engine rule effectRisk level: very_low to very_criticalThe severity assigned to a matching call. It only climbs: a call takes the highest level any matching rule sets.

Example: PII has Risk level set to Low, and its name subcategory has priority Low. A prompt containing only a person's name is not detected as PII at all, so no Policy Engine rule on PII matches it. Raise the category to High and the name is detected; a rule with Risk level high on PII then marks that call as high risk.

When the same category's risk level is set in more than one place, an LLM Gateway app's own Guardrails setting wins over the organization-wide Detection Models setting. When the Policy Engine is enabled, the policy wins.

None of these levels choose the action. Monitor, redact or block always comes from the policy or control.

View & test​

View & test opens a category or group so you can see what it covers and try sample text against it before you rely on it in a policy. Test with synthetic data, never real customer or employee records.

How detections are used​

  • LLM Gateway and MCP Gateway: the Data & Adversarial Risks card in the Policy Engine picks detections by category or individual type and assigns an action per stage.
  • Findings in Findings & Interactions name the detection that fired.