Hallucination protection
Score model responses for hallucinations and monitor or block them above a confidence threshold. Runs at the response stage.
Where it is configured
Hallucination protection is a Policy Engine feature. There is no dedicated app setting or Configure tab for it. The only per-app option is the on/off Hallucination check in an app's Guardrails tab, which uses a fixed threshold of 0.8 and freezes once the Policy Engine is on (see Switching from classic settings).
For a tunable threshold, scoping and follow-up severities, use the Hallucination Protection card in Policy Engine > LLM Gateway. Edits join the shared draft and take effect once you review and publish a revision.
Sections
Applies on bedrock, chat, responses and vertex, and only to non-streaming responses. Streaming responses cannot be blocked, so they are not checked.
Settings
Add protection opens the Hallucination protection dialog.
Follow-up rules
Add follow-up opens Hallucination follow-up. When is one of: a hallucination is detected, the score is at least (0 to 1), or the risk level is (Low, Medium, High). The rule then tags a severity. The highest severity among matching rules is reported.
Scoping and advanced scenarios
- Applies to: Everyone, People, Smart group, Application, App tag, Requested model, Provider, API surface, Environment, Tool, Source network, or Except....
- Threshold: the lowest matching threshold applies, whatever its priority. A narrower scope cannot raise the threshold set for Everyone.
- Action and risk level follow the highest-priority matching configuration.
- Off on a narrower scope beats On for Everyone.
- Need to key protection off attachment types, streaming or a metadata field? Use Add configuration.
For example, block likely hallucinations for one customer-facing app while the rest of the tenant only monitors, or set a lower (stricter) threshold for one Smart group. Start at Monitor, then move the action or the threshold:
runs on responsepriority 500
Related
- Security guardrails - the basic per-app hallucination check.
- Policy Engine overview