Skip to main content

Hallucination protection

Score model responses for hallucinations and monitor or block them above a confidence threshold. Runs at the response stage.

Where it is configured​

Hallucination protection is a Policy Engine feature. There is no dedicated app setting or Configure tab for it. The only per-app option is the on/off Hallucination check in an app's Guardrails tab, which uses a fixed threshold of 0.8 and freezes once the Policy Engine is on (see Switching from classic settings).

For a tunable threshold, scoping and follow-up severities, use the Hallucination Protection card in Policy Engine > LLM Gateway. Edits join the shared draft and take effect once you review and publish a revision.

Sections​

SectionWhat it doesEmpty state
ProtectionSwitch, threshold, action and risk level. Add protection.Responses are not checked for hallucinations.
After a hallucination is detectedFollow-up rules that tag a severity on the detection result. Add follow-up.Nothing is tagged on the detection result.

Applies on bedrock, chat, responses and vertex, and only to non-streaming responses. Streaming responses cannot be blocked, so they are not checked.

Settings​

Add protection opens the Hallucination protection dialog.

SettingOptionsDefaultNotes
ProtectionOn, OffOnOff exempts the scope from a broader configuration that switched it on.
Threshold0 to 1, step 0.050.8Lower catches more, with more false positives.
When flagged: actionMonitor, BlockMonitorBlock replaces the response with a refusal.
When flagged: risk levelLow, Medium, HighMediumTags the hallucination itself.
SeverityNot set, Very low, Low, Medium, High, Critical, Very criticalNot setApplies to the whole request. Reported only.

Follow-up rules​

Add follow-up opens Hallucination follow-up. When is one of: a hallucination is detected, the score is at least (0 to 1), or the risk level is (Low, Medium, High). The rule then tags a severity. The highest severity among matching rules is reported.

Scoping and advanced scenarios​

  • Applies to: Everyone, People, Smart group, Application, App tag, Requested model, Provider, API surface, Environment, Tool, Source network, or Except....
  • Threshold: the lowest matching threshold applies, whatever its priority. A narrower scope cannot raise the threshold set for Everyone.
  • Action and risk level follow the highest-priority matching configuration.
  • Off on a narrower scope beats On for Everyone.
  • Need to key protection off attachment types, streaming or a metadata field? Use Add configuration.

For example, block likely hallucinations for one customer-facing app while the rest of the tenant only monitors, or set a lower (stricter) threshold for one Smart group. Start at Monitor, then move the action or the threshold:

monitor_high_confidence_hallucinationsresponse

runs on responsepriority 500

WhenApplication method typeis any ofchatresponsesbedrockvertex
Then
Hallucination checktrue
Hallucination threshold0.82
Hallucination risk levelhigh
Hallucination actionmonitor