Skip to main content

Agentic Red Teaming

Run adaptive, multi-turn attacks against your live AI agent, including the tools it can call, and turn what breaks into actionable findings.

The Red Teaming page​

Open Assessments → Red Teaming. The page has four tabs:

TabWhat it testsDocs
LLM Intelligence AssessmentFixed adversarial and benchmark suites against a gateway app's modelLLM Intelligence Assessment
MCP Threat DetectionSupply-chain scan of an MCP server (repository, dependencies, live tool surface)MCP Threat Detection
Agentic Red TeamingAdaptive attacks against a live agent over HTTP, WebSocket, or voiceThis page
Model Red TeamingThe same adaptive engine pointed directly at one model, or 2-8 models side by sideModel Red Teaming

Agentic and Model Red Teaming share one sub-navigation: New assessment, Runs, Findings, and Schedules. Runs, findings, and schedules are shared between both tabs.

How an assessment works​

01 Connect
HTTP, WebSocket, or voice agent
Test connection
02 Challenge
Quick scan, full library, or selected attacks
Optional custom objectives
03 Launch
Authorized scope
Depth + evaluation policy
Engine runs
Benign recon
Tool-targeted objectives synthesized
Multi-turn attacks
Report
Grade + findings
Remediation + verify
QuilrAI

The new-assessment page is one form with three steps. A sticky right rail helps you track it:

  • Assessment Plan - a Target / Objectives / Review checklist (for example "1/3 ready"), a one-line summary such as "HTTP agent · Standard · 8 selected objectives", the connection status, and a Go to target button.
  • Smart Assist - an assistant with quick prompts such as "Help me connect my agent" and "Choose the right attack coverage". Keep keys and credentials out of the chat.

01 Connect your agent​

Give the assessment a Name (for example "Support Bot"), then choose How do you reach this agent?:

OptionUse it when
HTTP endpoint (default)The agent answers over an HTTPS chat API your application already calls.
WebSocket endpointThe agent answers over a secure WebSocket (wss://) your application keeps open.
Voice agentThe agent answers over a spoken channel: a realtime voice model, a hosted voice agent, or an audio endpoint.

HTTP endpoint​

FieldDefaultNotes
Endpoint URLplaceholder https://api.example.com/chatYour agent's HTTPS chat API. Put authentication in custom headers, not the URL.
Reply pathreplyDotted path to the assistant reply in the JSON response.
Session-id pathsession_idWhere the response returns a session id to reuse on the next turn.
Tool-calls pathactionsWhere the response lists the tool calls the agent made, so tool abuse can be analyzed.
Timeout (s)placeholder 240Per-request timeout.
Request body template{"message": "{{message}}", "session_id": "{{session_id}}"}{{message}} is the next attacker turn. {{session_id}} is a stable id per conversation.
Custom headers (optional)One Authorization rowUse + Add header for more.
Confirmation replay is safe for this HTTP endpointOffEnable only when repeating a successful request cannot cause a harmful or duplicate side effect. Confirmation runs stay off otherwise.

Click Test connection before you launch. It sends one benign message and validates the response mapping.

Example mapping​

If your agent is called like this:

POST https://agent.example.com/chat
Content-Type: application/json
Authorization: Bearer <your-key>

{ "message": "What is my balance?", "session_id": "rt-abc123" }

and replies:

{
"reply": "Your balance is $420.00.",
"session_id": "rt-abc123",
"actions": [ { "name": "get_balance", "arguments": { "account": "self" } } ]
}

then the defaults already match: Reply path reply, Session-id path session_id, Tool-calls path actions, and the default body template. Add Authorization: Bearer <your-key> as a custom header.

tip

Returning the session id keeps multi-turn attacks coherent. Exposing tool calls lets the engine detect when the agent was induced to invoke a tool.

To try this end to end, run the HTTP agent example. It uses exactly this request and reply shape, so the defaults work as is.

WebSocket endpoint​

Choose WebSocket endpoint when your agent keeps a WebSocket open and answers each message with one or more frames, for example a typing indicator, streamed tokens, and a final frame.

FieldDefaultNotes
WebSocket URLplaceholder wss://agent.example.com/chatYour agent's secure WebSocket. Only wss:// is accepted. Put authentication in custom headers, not the URL. Socket.IO endpoints are not supported.
SubprotocolsEmptyOptional. Offered during the handshake, for example graphql-transport-ws. Separate several with commas. Treated as a secret and never stored with results.
Opening messageEmptyOptional. Sent once after connecting, before the first message, for example an auth or session-start frame. Treated as a secret and never stored with results.
When is a reply finished?First reply messageFirst reply message: the first frame that carries reply text. Done marker: a frame where Done path is present and true, or equals Done value if you set one. Quiet period: the agent sends no frame for Quiet period (s) (default 2).
Join the text of every frame in the replyOffWith Done marker or Quiet period, joins the text of every frame. Use it for agents that stream a reply token by token.
ConnectionKeep the connection openKeep the connection open: one connection per attack, so the agent keeps its own conversation context. Reconnect for each message: a new connection for every message; include {{history}} in the template if the agent keeps no context.
Reply path, Session-id path, Tool-calls path, Timeout (s)As for HTTPApplied to each frame the agent sends. Without a reply path, only a plain-text frame or a common reply field counts as reply text; status frames such as {"type": "typing"} are ignored.
Message template{"message": "{{message}}", "session_id": "{{session_id}}"}The frame sent for each turn. Same placeholders as the HTTP request body template.
Custom headers (optional)One Authorization rowSent with the WebSocket handshake.
Confirmation replay is safe for this WebSocket endpointOffSame as for HTTP endpoints.

Click Test connection before you launch. It connects, sends the opening message if you set one, sends one benign message, and checks the reply mapping.

Example: a streaming agent​

If your agent greets a new connection, then streams each reply:

agent  -> {"type": "ready", "session_id": "rt-abc123"}
you -> {"message": "What is my balance?"}
agent -> {"type": "typing"}
agent -> {"type": "delta", "text": "Your "}
agent -> {"type": "delta", "text": "alice "}
agent -> {"type": "delta", "text": "balance: "}
agent -> {"type": "delta", "text": "$420.00."}
agent -> {"type": "complete", "actions": [{"name": "get_balance", "arguments": {"customer": "alice"}}]}

then set When is a reply finished? to Done marker with Done path type and Done value complete, turn on Join the text of every frame in the reply, and set Reply path text, Session-id path session_id, Tool-calls path actions, and the Message template to {"message": "{{message}}"}.

The WebSocket agent example implements this protocol, including bearer auth on the handshake, an optional subprotocol, and an optional opening auth frame, with the matching settings in its README.

Voice agent​

Pick a Voice transport:

TransportWhat it isFields
Realtime voice model (default)A speech-to-speech model over WebSocket (OpenAI Realtime or Azure). Runs your instructions and tools.Realtime model id (gpt-realtime), API key, Realtime endpoint URL (optional override, wss:// or https://), Agent voice (optional)
ElevenLabs agentAn ElevenLabs Conversational AI agent, addressed by its agent id.Agent id, API key (only private agents need one; it is used to obtain a signed URL)
Audio HTTP endpointYour own HTTPS endpoint that accepts spoken audio and returns spoken audio.Audio endpoint URL, Reply audio path (default audio_base64)

For the audio HTTP endpoint, the engine sends each spoken turn as a WAV file in a multipart field named audio and reads the base64 reply audio from the Reply audio path in your JSON response.

All transports also have:

  • Attacker voice (default verse) - the voice the assessment speaks its probes with.
  • Replay-safe checkbox - leave it unchecked if the agent takes real actions; confirmation replays are then skipped.
  • Test connection.
note

Voice assessments need speech-to-text and text-to-speech enabled for the assessment engine. If they are not, the form shows a banner asking an administrator to turn on speech access.

02 Define the attack coverage​

ModeRunsBest for
Quick scan (default)~8 core objectivesA fast, representative sweep across the top risk categories
Full libraryAll 64 objectivesA complete baseline before release
Select attacksThe objectives you tickFocusing on the risks that matter to this agent

Every mode also adds tool-targeted objectives synthesized live for the target after recon. In the example runs in Reading a report, recon synthesized 6, so a quick scan ran 14 objectives (8 library + 6 synthesized).

Expand Attack library to see every objective grouped by OWASP category, with its MITRE ATLAS technique and severity. In Quick scan mode, objectives outside the quick scan are dimmed and tagged "full only".

In Select attacks mode, click a row to toggle it, or use Select all and Clear.

The full list is in the Attack Library reference.

Custom objectives​

Click + Add custom objective to test something specific to your agent. Give it a name, a severity (default high), and describe what the agent should be made to do or reveal that it should not. Custom objectives are reported under a Custom category.

03 Set the boundaries and launch​

ControlDefaultNotes
Authorized assessment scopeEmpty (required)Record the approved target, ticket or reference, and testing boundary.
DepthStandardStandard is a focused assessment with standard attack effort. Deep explores more persistently and may take longer.
I acknowledge managed assessment analysisOff (required)Pattern-redacted transcripts may be processed by the deployment-managed evaluation engine. Use only approved targets.

Click Show evaluation settings to tune the Shared evaluation policy:

SettingDefaultWhat it does
Evaluation modeHybridHybrid combines deterministic evidence and model-judge consensus. Deterministic uses evidence detectors without model judges. Judge uses model judges and the consensus threshold.
Judge count1Number of model judges.
Consensus threshold0.67Agreement required among judges.
Confirmation runs0Replays of a successful attack to confirm it. Disabled until the endpoint is marked replay-safe.
Deterministic evidence overrideOnLets strong deterministic evidence override judge disagreement.

Click Launch assessment. It runs against the authorized target only, and you can leave the page and come back.

warning

Only test agents you own or are explicitly authorized to assess. Leave the replay-safe box unchecked for any agent that takes real actions (payments, record changes, messages).

While it runs​

Open the run from Runs to watch it live. Select an objective to see its progress and live turn-by-turn transcript. Leaving the page does not cancel the run.

note

Full transcripts are visible only in this live view. The completed report shows turn-referenced evidence quotes instead of embedding transcripts.

Next steps​