Model Red Teaming
Point the adaptive red-team engine directly at a model: probe its guardrails, its system prompt, and the tools it is allowed to call. Test one model, or compare two to eight on the same objectives.
Model Red Teaming is the Model Red Teaming tab on Assessments → Red Teaming. It uses the same three-step form, attack library, evaluation policy, reports, findings tracker, and schedules as Agentic Red Teaming. Only step 01 differs: you choose models instead of an agent endpoint.

01 Choose the model
-
Enter a Name (for example "Checkout assistant model").
-
Pick an Assessment mode:
-
For each model card, pick a Model source.
-
Optionally add a system prompt and tools.
Model sources

Enter manually

Point the endpoint URL at a gateway or proxy to assess the model exactly as your app reaches it.
Supported providers (17):

System prompt and tools
Both are optional, and both make the test closer to your real application.

Paste a JSON array of tool definitions (OpenAI function format is supported) into Tools JSON, or click Open Tool Studio.
Tool Studio
Tool Studio is a side-by-side editor for tool definitions:
- Paste or write definitions in Tools JSON. Use Format to tidy them, or Example to load a sample.
- Check the footer, for example "Valid JSON · 2 tools".
- Review Discovered tools on the right. Each tool card shows its description and parameters (type, required) and is classed READ or WRITE.
- Click Apply N tools.

The built-in example declares get_account_balance (READ) and transfer_funds (WRITE). WRITE tools are the ones to watch: the red agent will try to induce unsafe calls to them.
Compare models
Choose Compare models to run one campaign across two to eight models. Each model card has its own source and credentials; coverage, evaluation policy, and authorization apply to every model. By default the form shows two cards (OpenAI and Anthropic).

Steps 02 and 03 are the same as in Agentic Red Teaming. The plan rail reads, for example, "2 models · Standard · 8 selected objectives".
The campaign view
Open the campaign from Runs. The header shows progress, such as "Multi-model assessment · 2 of 2 targets finished".


How to read the matrix:
- Rows are objectives, aligned by objective fingerprint, then attack id.
- Each cell is outcome · severity · judge confidence, for example "Partial · Medium · 75% confidence".
- Not reported means that objective did not run against that model. Tool-targeted objectives are synthesized per target, so each model can get different ones. The page warns "Targets were evaluated with different coverage" when this happens.
- Comparable risk score stays N/A until every target result is comparable.
Example. Comparing two QuilrAI-provided models on a quick scan: gpt-oss-120b-ultrafast had 2 vulnerable (grade C, risk 33/100: 2 breached, 5 partial, 7 held, 14 tested), while gpt-oss-120b had 4 vulnerable (grade F). Click Report on either row to open its full report.
Next steps
- Reading Red Team Results - including a worked example:
gpt-5.4-minion the full library, grade A. - Attack Library - all 64 objectives.
- Findings and Schedules.