Agentic Red Teaming
Run adaptive, multi-turn attacks against your live AI agent, including the tools it can call, and turn what breaks into actionable findings.
The Red Teaming page
Open Assessments → Red Teaming. The page has four tabs:
Agentic and Model Red Teaming share one sub-navigation: New assessment, Runs, Findings, and Schedules. Runs, findings, and schedules are shared between both tabs.
How an assessment works
The new-assessment page is one form with three steps. A sticky right rail helps you track it:
- Assessment Plan - a Target / Objectives / Review checklist (for example "1/3 ready"), a one-line summary such as "HTTP agent · Standard · 8 selected objectives", the connection status, and a Go to target button.
- Smart Assist - an assistant with quick prompts such as "Help me connect my agent" and "Choose the right attack coverage". Keep keys and credentials out of the chat.
01 Connect your agent
Give the assessment a Name (for example "Support Bot"), then choose How do you reach this agent?:
HTTP endpoint
Click Test connection before you launch. It sends one benign message and validates the response mapping.
Example mapping
If your agent is called like this:
POST https://agent.example.com/chat
Content-Type: application/json
Authorization: Bearer <your-key>
{ "message": "What is my balance?", "session_id": "rt-abc123" }
and replies:
{
"reply": "Your balance is $420.00.",
"session_id": "rt-abc123",
"actions": [ { "name": "get_balance", "arguments": { "account": "self" } } ]
}
then the defaults already match: Reply path reply, Session-id path session_id, Tool-calls path actions, and the default body template. Add Authorization: Bearer <your-key> as a custom header.
Returning the session id keeps multi-turn attacks coherent. Exposing tool calls lets the engine detect when the agent was induced to invoke a tool.
To try this end to end, run the HTTP agent example. It uses exactly this request and reply shape, so the defaults work as is.
WebSocket endpoint
Choose WebSocket endpoint when your agent keeps a WebSocket open and answers each message with one or more frames, for example a typing indicator, streamed tokens, and a final frame.
Click Test connection before you launch. It connects, sends the opening message if you set one, sends one benign message, and checks the reply mapping.
Example: a streaming agent
If your agent greets a new connection, then streams each reply:
agent -> {"type": "ready", "session_id": "rt-abc123"}
you -> {"message": "What is my balance?"}
agent -> {"type": "typing"}
agent -> {"type": "delta", "text": "Your "}
agent -> {"type": "delta", "text": "alice "}
agent -> {"type": "delta", "text": "balance: "}
agent -> {"type": "delta", "text": "$420.00."}
agent -> {"type": "complete", "actions": [{"name": "get_balance", "arguments": {"customer": "alice"}}]}
then set When is a reply finished? to Done marker with Done path type and Done value complete, turn on Join the text of every frame in the reply, and set Reply path text, Session-id path session_id, Tool-calls path actions, and the Message template to {"message": "{{message}}"}.
The WebSocket agent example implements this protocol, including bearer auth on the handshake, an optional subprotocol, and an optional opening auth frame, with the matching settings in its README.
Voice agent
Pick a Voice transport:
For the audio HTTP endpoint, the engine sends each spoken turn as a WAV file in a multipart field named audio and reads the base64 reply audio from the Reply audio path in your JSON response.
All transports also have:
- Attacker voice (default
verse) - the voice the assessment speaks its probes with. - Replay-safe checkbox - leave it unchecked if the agent takes real actions; confirmation replays are then skipped.
- Test connection.
Voice assessments need speech-to-text and text-to-speech enabled for the assessment engine. If they are not, the form shows a banner asking an administrator to turn on speech access.
02 Define the attack coverage
Every mode also adds tool-targeted objectives synthesized live for the target after recon. In the example runs in Reading a report, recon synthesized 6, so a quick scan ran 14 objectives (8 library + 6 synthesized).
Expand Attack library to see every objective grouped by OWASP category, with its MITRE ATLAS technique and severity. In Quick scan mode, objectives outside the quick scan are dimmed and tagged "full only".
In Select attacks mode, click a row to toggle it, or use Select all and Clear.
The full list is in the Attack Library reference.
Custom objectives
Click + Add custom objective to test something specific to your agent. Give it a name, a severity (default high), and describe what the agent should be made to do or reveal that it should not. Custom objectives are reported under a Custom category.
03 Set the boundaries and launch
Click Show evaluation settings to tune the Shared evaluation policy:
Click Launch assessment. It runs against the authorized target only, and you can leave the page and come back.
Only test agents you own or are explicitly authorized to assess. Leave the replay-safe box unchecked for any agent that takes real actions (payments, record changes, messages).
While it runs
Open the run from Runs to watch it live. Select an objective to see its progress and live turn-by-turn transcript. Leaving the page does not cancel the run.
Full transcripts are visible only in this live view. The completed report shows turn-referenced evidence quotes instead of embedding transcripts.
Next steps
- Reading a report - grades, findings, framework mapping, remediation, and two worked example reports.
- Findings and Schedules - track remediation and run the same assessment on a recurring schedule.
- Model Red Teaming - test a model directly, or compare up to eight.