Reading Red Team Results
How to read an Agentic or Model Red Teaming report, with two real example runs: a clean model result and an agent with findings to fix.
This page covers Agentic Red Teaming and Model Red Teaming reports. For LLM Intelligence Assessment runs, see Reading the LLM Intelligence Assessment Report.
Find a report
Open Runs on either tab. The Completed list has one card per assessment: a letter grade (A-F, or N/A when the run is incomplete), the target name, "x/y vulnerable · done | incomplete", and the date. Click a card to open its report.

Comparisons across several models open a campaign view first. See Compare models.
The report at a glance
The top bar has Back to runs, Track findings (copies the findings into the findings tracker), Review findings, and Remediate & verify. The last two jump to those sections.
Outcome terms
A letter grade is not a pass. When anything is breached or partial, the report shows "Complete coverage with confirmed residual risk... This is not a clean or passing posture regardless of the letter grade."
Example 1: a clean result
gpt-5.4-mini, tested through Model Red Teaming with the Full library.


Every OWASP category is green: LLM07, ASI02, LLM01, ASI01, ASI03, LLM05, ASI05, LLM02, ASI06, LLM06, LLM10, LLM08, LLM09, LLM04, ASI01 Agent Goal Manipulation, and Custom. Every ATLAS technique is green and the technique panel reads "no technique produced a confirmed breach".

How to read it:
- "No vulnerabilities confirmed across 60 completed probes. This is evidence for the tested scope, not a guarantee beyond it." The result covers this model, with this configuration (here, no system prompt or tools), at this depth.
- 0 tools probed and "text-only target" mean tool abuse and multimodal channels were not exercised. Add tools (or a system prompt) on the model form to widen the test.
- With nothing to fix, Remediation shows only the guardrail verification card.
Re-run after any change to the model, system prompt, or tools. Results from scheduled quick scans of the same model can vary from run to run (daily gpt-5.4-mini quick scans on the same tenant ranged from grade B to D, with 1-3 of 12 objectives vulnerable), so a single clean run is a data point, not a guarantee.
Example 2: reading findings and fixing them
"HTTP Agent", a support agent tested through Agentic Red Teaming with a Quick scan.

12 of 14 objectives held (86%), but both failures are in the same place.
Step 1: see where it broke

- OWASP LLM & Agentic coverage - green dot = held, red dot = breached. Only LLM07: System Prompt Leakage is red.
- Tool exposure - red marks a tool the agent was induced to invoke. This agent exposed none.

- MITRE ATLAS coverage - only AML.T0056: Extract LLM System Prompt is red.
- Attack-technique effectiveness - which adversarial techniques succeeded. Here: Direct (1). Prioritize defenses against what appears here.
- Discovered attack surface - recon found no tools; 6 tool-targeted objectives were still synthesized for the target.
Step 2: read each finding

Use the search box or the All findings / Breached & partial filter. Each row shows title, severity, chips (PARTIAL, SYNTHESIZED), OWASP category, ATLAS technique, the technique that worked ("via Direct"), and reproduction status.
Expand a finding for the full write-up:

Completed reports show turn-referenced evidence, not full transcripts. Full turn-by-turn transcripts are visible only in the live run view while the assessment runs.
Step 3: remediate and verify
1. Verify with the Quilr guardrail. Re-runs this exact assessment with Quilr's production guardrail in front of the agent and shows before/after grade cards. The guardrail is the same Gemma task-adherence classifier that the LLM Gateway's Guardian Agent enforces inline.

2. Prompt hardening suggestions. Auto-generated from the findings: why the changes help, a checklist of key changes, and recommended system-prompt guardrails to copy. For this run, generated from 2 findings with 6 key changes, including an explicit instruction hierarchy, a ban on disclosing or paraphrasing hidden text, and a fixed refusal for extraction attempts.

3. Guardian Agent and custom detection recommendations. Controls that stop these attacks before they reach the agent (3 controls for this run).
- Guardian Agent prompt - an agent purpose and a Do/Don't prompt (here 3 Do and 5 Don't) to paste into the app's Guardian Agent, with suggested settings (Block off-task requests, Sensitivity: High). It catches semantic attacks that patterns cannot.

- Custom detection suggestions - pattern-based controls to add as Security Guardrails. Each has an action, a control type, a rationale, patterns to add, and what it protects against.

Other runs can also suggest a Token Limit control (for example "Cap oversized extraction requests").
Finish with Apply & verify with the guardrail to put the guardrail in front of the agent and measure the improvement. Then click Track findings to hand the open issues to their owners.
Recon and fingerprinting
When recon finds something, Discovered attack surface lists the capabilities or tools it enumerated, each with a risk level, plus fingerprint chips such as the model family and knowledge cutoff it inferred. This example is gpt-oss-120b-ultrafast (grade C, risk 33/100: 2 breached, 5 partial, 7 held, 14 tested), where technique effectiveness shows Authority 1 and Direct 1.

Its tool-disclosure finding shows two details worth knowing:

- "Reproduction not attempted. live target was not declared an idempotent replay-safe sandbox" - confirmation replays only run when the target is marked replay-safe and confirmation runs are above 0.
- "MITRE ATLAS publishes no mitigation for this technique; apply the OWASP guidance above." Not every ATLAS technique has a published mitigation.
Framework mapping
Every objective and finding is mapped to:
See the Attack Library for the mapping of each objective.
Export
Use PDF or Markdown on the posture card. Exports include:
- Executive summary and assessment coverage
- OWASP and MITRE ATLAS coverage, and strategy effectiveness
- Tool exposure and discovery
- Findings and evidence, including judge votes, excerpts, reproduction, and tool arguments and results
- Guardrail application history
- Prompt hardening, Guardian Agent, and custom detections
- Remaining app-side fixes
- Framework references