Guardian agent
Guide model behavior with gateway checks for dependency safety and task adherence.
Guardian Agent runs inside the gateway's request and response flow. It is not a separate agent or model endpoint. When enabled, the gateway can add instructions before a request reaches the model, retry unsafe dependency output once with corrective guidance, append an advisory, or block off-task requests.
Turn it on for an app
Guardian Agent is off for new apps.
How it works
For dependency output, Guardian Agent adds a response-side review before the final answer is returned:
Feature groups
Guardian Agent currently has two feature groups.
Coding helpers
Coding helpers focus on dependency-related prompts and generated dependency output.
On the request side, Guardian Agent detects dependency intent in user messages, such as requirements.txt, pip install, pyproject.toml, dependency lists, and package version questions. When matched, it adds a system instruction directing the model to avoid vulnerable versions and prefer current stable patched versions, depending on configuration.
On the response side, Guardian Agent scans dependency-like output, including:
pip installcommandsrequirements.txtandpyproject.tomlpackage.jsonCargo.tomlGemfilego.mod.csprojPackageReferenceentriespom.xmlcomposer.json
Package extraction is best-effort across PyPI, npm, crates.io, RubyGems, NuGet, Go, Maven, and Packagist. Exact pinned versions can be checked against OSV for known vulnerabilities. Bare package installs resolve the latest registry version first, then check that version. Version ranges are not checked.
Latest-version suggestions are supported for exact pins on PyPI, npm, crates.io, RubyGems, NuGet, and Go. Latest-version checks are not available for Maven and Packagist.
Task adherence
Task adherence checks whether the latest user message stays within the agent's purpose. Set Agent purpose to one sentence describing what the agent is for, and use Guardian prompt to state the boundary. The request's system prompt also describes the purpose; if a request has no system prompt, the check is skipped and the request is allowed.
When the latest user message is classified as off-task, Guardian Agent records a guardian_task_adherence finding and applies the action:
Task adherence runs on the request side only.
Writing a guardian prompt
One or two direct sentences are usually enough: name what the agent may handle, then say what is outside its scope. Keep evaluation logic out of the prompt, and put persona, tone and formatting in the agent's own system prompt.
Streaming and retry behavior
Request-side Guardian Agent checks run before upstream calls for both streaming and non-streaming requests.
For non-streaming responses, dependency findings trigger one retry with Guardian dependency instructions. If the retry still contains dependency advisories, the gateway appends a Guardian note to the final response. Vulnerability advisories suppress latest-version advisories for the same response.
For streaming requests with dependency checks enabled, the gateway first sends a preliminary non-streaming provider request to inspect a full draft response. If no dependency findings are found, the gateway streams that draft back to the client as SSE in the provider's format. If Guardian Agent finds vulnerabilities or update advisories, the gateway adds corrective instructions and sends a second streaming upstream request, then streams the second response to the client.
Other response-side Guardian Agent checks are skipped for normal streaming passthrough.
Latency impact
Guardian Agent runs additional checks inside the request and response path, so it adds latency on top of the normal gateway overhead.
As a planning figure, expect Guardian Agent to add ~700 ms per request when it is enabled. The actual latency varies with the scenario and the complexity of the request:
- Which feature groups are enabled. Running coding helpers and task adherence together takes longer than running one of them.
- Request size and complexity. Longer conversations and larger dependency manifests take longer to evaluate.
- Dependency lookups. OSV vulnerability checks and registry latest-version lookups are network calls, and their latency grows with the number of packages extracted from the response.
- Retries. A dependency finding triggers one corrective retry, which adds a second upstream model call to the request.
- Streaming with dependency checks. The gateway first issues a preliminary non-streaming draft request, so time to first token reflects the full draft rather than the first upstream token.
~700 ms is a guideline, not a guarantee. Requests that need no retry and no registry lookups take substantially less time, and requests that trigger a retry or many package lookups can take longer.
Endpoint coverage
Guardian Agent supports these LLM Gateway APIs:
OpenAI-compatible chat includes provider-native chat models reached through gateway translations, including Bedrock Converse, Vertex AI Gemini generateContent, and Anthropic Messages.
Configuration keys
For the Management APIs, Guardian Agent is the guardian_agent object in the app configuration:
{
"guardian_agent": {
"enabled": true,
"coding_helpers": {
"enabled": true,
"dependency_security_check": true,
"latest_version_suggestions": true
},
"task_adherence": {
"enabled": true,
"sensitivity": "medium",
"action": "nudge"
}
}
}
Setting guardian_agent to null on update removes the Guardian Agent configuration.
Logging
Guardian Agent findings are logged in the same prediction format used by guardrails:
type:classifymatch_type:guardianid:guardian_task_adherence,guardian_coding_dependency_security_review, orguardian_coding_latest_version_review
Guardian categories are also written to metadata.extra_data.guardian_agent.request and metadata.extra_data.guardian_agent.response in exported logs. Nudge and monitor findings appear under actions_and_categories.request.monitored or actions_and_categories.response.monitored. Blocked task-adherence findings appear under actions_and_categories.request.blocked.
If Guardian Agent records a finding and no content was blocked or anonymized, the request outcome becomes monitor_detected. If task adherence is configured with block and the latest user message is classified as unrelated, the request outcome becomes blocked and the upstream model is not called.
Current limits
- Dependency extraction is best-effort, and version ranges are not checked.
- Vulnerable dependency output is not blocked: the gateway retries once and then appends an advisory.
- Task adherence checks only the latest user message, on the request side.
- Dependency and task-adherence network checks fail open on transient errors.
Going further with the Policy Engine
Guardian Agent is also a card in Policy Engine > LLM Gateway. When the engine is on for the LLM Gateway, the Guardian tab freezes and the card applies instead. See What happens to classic settings.
The card has four sections: Guardian (master switch), Coding helpers, Task adherence and After Guardian runs (follow-up rules). Each control is a three-way switch (Leave as is, On, Off), so a narrow configuration can change one setting and inherit the rest. The card applies on the assistants, bedrock, chat, copilot, responses, sdk_check and vertex API surfaces.
- In a new configuration, Guardian starts On when opened from the Guardian section; Coding helpers and Task adherence start at Leave as is. Turning Coding helpers On ticks both Dependency security check and Latest version suggestions. At least one control must be On or Off.
- Guardian results are request-stage fields, so follow-up rules cannot see them on the response leg.
Scenarios it supports:
- Coding models only. Turn on coding helpers and strict task adherence when the Requested model matches a pattern such as
*code*. - Exempt one app or group. Guardian Off on a narrower scope (People, Smart group, Application, App tag) beats On for Everyone. The highest-priority configuration wins per setting.
- Severity follow-ups. Tag a severity when Guardian blocked, when there are at least N findings, or for a specific finding type, for dashboards and alerts.
- Metadata conditions. Use Add configuration to key Guardian off languages, repos or request metadata.
runs on requestpriority 700
Related
- Security guardrails - data and adversarial risk detection.
- Policy Engine overview
- HA and SLA - gateway latency budget.