Skip to main content

LLM Gateway Log Export API

Use the Log Export API to read LLM Gateway request logs from your own data platform, SIEM, warehouse, or scheduled export job.

The API returns newline-delimited JSON. Each response line is one complete JSON object, so clients can stream, parse, and checkpoint logs incrementally.

GET https://guardrails.quilr.ai/llmgateway/logs/export

Response content type:

Content-Type: application/x-ndjson

The endpoint also serves an aggregated executive dashboard view. Add view=metrics to receive a single summary JSON document instead of the per-request NDJSON stream. See Metrics View.

Authentication​

Pass a log export key from the QuilrAI LLM Gateway UI:

X-Quilr-Log-Export-Key: sk-export-...

Do not use your QuilrAI gateway API key as the request credential for this endpoint. The log export key is separate from model-call authentication.

The UI exposes two export scopes:

Export keyScope
log_export_keyExports logs only for the underlying QuilrAI API key it belongs to.
all_apps_log_export_keyExports logs for all active, non-expired LLM Gateway apps in the tenant.

Both scopes use the same endpoint, header, query parameters, pagination model, and response format. In all-apps exports, each llmgateway.request event still includes the concrete app name in app.name.

Query Parameters​

All query parameters are optional.

ParameterDescription
start_timeISO 8601 lower bound for exported logs. Naive timestamps are treated as UTC.
end_timeISO 8601 upper bound for exported logs. Naive timestamps are treated as UTC.
cursorOpaque cursor from the previous checkpoint.next_cursor. When provided, it wins over start_time.
limitMaximum request rows to export in this response. Default 1000. Values above 5000 are silently clamped to 5000. Values below 1 or non-integer values return 400.
viewResponse shape. logs (default) streams per-request NDJSON events. metrics returns one aggregated JSON document. See Metrics View. Any other value returns 400.

Logs are available for a maximum of 15 days. Choose start_time within that retention window when backfilling. Requests with an effective start_time, end_time, or cursor timestamp before the retention window fail with 400.

If neither start_time nor cursor is provided, the API exports a default 24-hour window ending at the effective export end time.

Export Lag​

The API does not export logs newer than 15 minutes. Gateway logs and prediction payloads are written asynchronously, so this lag keeps exported rows stable.

If end_time is newer than now - 15 minutes, the server clamps it to the maximum exportable time. The request still succeeds. The export_started and checkpoint events include the effective export bounds.

Request Examples​

Start an export window, here the hour that ended two hours ago (logs are kept for 15 days, so fixed dates go stale):

# macOS (BSD date) first, GNU date as fallback
START=$(date -u -v-3H +%Y-%m-%dT%H:%M:%SZ 2>/dev/null || date -u -d '3 hours ago' +%Y-%m-%dT%H:%M:%SZ)
END=$(date -u -v-2H +%Y-%m-%dT%H:%M:%SZ 2>/dev/null || date -u -d '2 hours ago' +%Y-%m-%dT%H:%M:%SZ)

curl -N \
-H "X-Quilr-Log-Export-Key: sk-export-..." \
"https://guardrails.quilr.ai/llmgateway/logs/export?start_time=$START&end_time=$END&limit=1000"

Resume from the previous checkpoint:

curl -N \
-H "X-Quilr-Log-Export-Key: sk-export-..." \
"https://guardrails.quilr.ai/llmgateway/logs/export?cursor=<next_cursor>"

When resuming with cursor, you do not need to pass start_time or end_time.

Pagination​

Rows are ordered by:

timestamp ASC, request_id ASC

The cursor is opaque. Store it exactly as returned in checkpoint.next_cursor and send it back as the cursor query parameter on the next request.

If checkpoint.has_more is true, call the endpoint again immediately with cursor=<next_cursor>.

If checkpoint.has_more is false, there are no more rows in the current effective window. Store next_cursor and poll later with that cursor to continue incremental export.

When an initial request (no cursor supplied) returns zero rows, the API returns a checkpoint cursor pinned to the effective end time. This lets exporters store one cursor value even for empty windows.

When a request with a cursor returns zero rows, checkpoint.next_cursor echoes the inbound cursor unchanged and has_more is false. Re-poll later with the same cursor.

A reliable collector​

A minimal Python collector (standard library only) that you can run on a schedule or as a service. It:

  • stores rows and the cursor in one SQLite transaction, so the cursor only advances after the page is safely written;
  • writes idempotently (INSERT OR IGNORE on a unique id), so a replayed page never creates duplicates;
  • treats a mid-stream error event or a stream without a checkpoint as a failed page, and retries it with exponential backoff and jitter;
  • stops on 400, 401 and 403, which retrying cannot fix.
collector.py
"""Minimal LLM Gateway log collector: cursor checkpoint, retries, idempotent writes."""
import json, os, random, sqlite3, time, urllib.error, urllib.parse, urllib.request
from datetime import datetime, timedelta, timezone

URL = os.environ.get("EXPORT_URL", "https://guardrails.quilr.ai/llmgateway/logs/export")
KEY = os.environ["QUILR_LOG_EXPORT_KEY"]
EVENT = "llmgateway.request"
row_id = lambda e: e["request"]["id"] # unique per request

db = sqlite3.connect("quilr_logs.db")
db.execute("CREATE TABLE IF NOT EXISTS events (id TEXT PRIMARY KEY, ts TEXT, body TEXT)")
db.execute("CREATE TABLE IF NOT EXISTS state (k TEXT PRIMARY KEY, v TEXT)")

def saved_cursor():
row = db.execute("SELECT v FROM state WHERE k = 'cursor'").fetchone()
return row[0] if row else None

def fetch_page(params):
"""Return (events, checkpoint) for one complete page, or raise."""
req = urllib.request.Request(URL + "?" + urllib.parse.urlencode(params),
headers={"X-Quilr-Log-Export-Key": KEY})
events, checkpoint = [], None
with urllib.request.urlopen(req, timeout=300) as r: # raises HTTPError on 4xx/5xx
for line in r:
if not line.strip():
continue
msg = json.loads(line)
if msg["type"] == "error": # can arrive after HTTP 200
raise RuntimeError(f"export error: {msg['error']}")
if msg["type"] == EVENT:
events.append(msg)
elif msg["type"] == "checkpoint":
checkpoint = msg
if checkpoint is None:
raise RuntimeError("stream ended without a checkpoint")
return events, checkpoint

def with_retries(fn, *args, attempts=6):
for n in range(attempts):
try:
return fn(*args)
except urllib.error.HTTPError as e:
if e.code in (400, 401, 403) or n == attempts - 1:
raise # fix the key or the window; retrying will not help
except (OSError, RuntimeError, ValueError): # network, mid-stream error, truncation
if n == attempts - 1:
raise
time.sleep(min(60, 2 ** n) + random.random()) # backoff with jitter

def run_once():
cursor = saved_cursor()
if cursor:
params = {"cursor": cursor, "limit": 1000}
else: # first run: start one hour back, well inside the 15-day retention
start = datetime.now(timezone.utc) - timedelta(hours=1)
params = {"start_time": start.strftime("%Y-%m-%dT%H:%M:%SZ"), "limit": 1000}
while True:
events, checkpoint = with_retries(fetch_page, params)
with db: # one transaction: rows and cursor commit together
db.executemany(
"INSERT OR IGNORE INTO events VALUES (?, ?, ?)",
[(row_id(e), e["request"]["timestamp"], json.dumps(e)) for e in events])
db.execute("INSERT OR REPLACE INTO state VALUES ('cursor', ?)",
(checkpoint["next_cursor"],))
print(f"stored {checkpoint['rows']} rows up to {checkpoint['effective_end_time']}")
if not checkpoint["has_more"]:
return
params = {"cursor": checkpoint["next_cursor"], "limit": 1000}

if __name__ == "__main__":
while True:
run_once()
time.sleep(300) # new rows appear about 15 minutes after the request
export QUILR_LOG_EXPORT_KEY='sk-export-...'
python3 collector.py

To forward to a SIEM instead of SQLite, replace the with db: block with your sink's write, and persist next_cursor only after the sink acknowledges the batch. If the collector is down for longer than the 15-day retention, the saved cursor fails with 400; delete it to restart from a recent start_time.

Verify delivery​

  1. Send one request through an app covered by the export key and note the time.
  2. After at least 15 minutes, run the collector once.
  3. Check that the row arrived: sqlite3 quilr_logs.db "SELECT id, ts FROM events ORDER BY ts DESC LIMIT 5".
  4. For a closed window, compare the number of rows you stored with governance_metrics.total_requests from the metrics view for the same start_time and end_time (check coverage.complete is true). A gap means pages were skipped; re-export that window, and the idempotent writes absorb the overlap.

Coverage​

The export covers LLM Gateway traffic for the selected export scope, including:

Traffic typeExported
OpenAI-compatible chat completionsYes
Anthropic MessagesYes
OpenAI ResponsesYes
OpenAI Realtime session logsYes
OpenAI speech-to-textYes
OpenAI text-to-speechYes
Sarvam speech-to-text and text-to-speechYes
Sarvam translation, transliteration, and language detectionYes
EmbeddingsYes
RerankYes
AWS Bedrock Runtime boto3Yes
Vertex AI GeminiYes
Streaming requestsYes
SDK mode checksYes
Copilot Studio checksYes

Response Events​

Every successful response starts with export_started, contains zero or more llmgateway.request events, and ends with checkpoint.

export_started​

The first line describes the effective export window.

{
"type": "export_started",
"schema_version": "v1",
"scope": "app",
"app_name": "my-app",
"app_count": 1,
"effective_start_time": "2026-05-14T00:00:00.000Z",
"effective_end_time": "2026-05-14T01:00:00.000Z",
"max_exportable_time": "2026-05-14T10:45:00.000Z",
"end_time_clamped": false,
"limit": 1000
}
FieldTypeDescription
typestringAlways export_started.
schema_versionstringEvent schema version. Current value is v1.
scopestringapp for a single-key export, or all_apps for a tenant-wide all-apps export.
app_namestring or nullLLM Gateway app name for a single-key export, or "" if that app has no configured name. null for all-apps exports because the export spans multiple apps; the concrete per-request app name is on each llmgateway.request event at app.name.
app_countnumberNumber of active, non-expired apps included in the export scope.
effective_start_timestringISO 8601 timestamp where this export starts.
effective_end_timestringISO 8601 timestamp where this export ends.
max_exportable_timestringNewest timestamp eligible for export after the 15-minute lag.
end_time_clampedbooleantrue when the requested end_time was newer than max_exportable_time.
limitnumberMaximum request rows returned in this response.

For all-apps export, the first line uses this scope shape:

{
"type": "export_started",
"schema_version": "v1",
"scope": "all_apps",
"app_name": null,
"app_count": 3,
"effective_start_time": "2026-05-14T00:00:00.000Z",
"effective_end_time": "2026-05-14T01:00:00.000Z",
"max_exportable_time": "2026-05-14T10:45:00.000Z",
"end_time_clamped": false,
"limit": 1000
}

llmgateway.request​

Each request row is emitted as one llmgateway.request event.

{
"type": "llmgateway.request",
"schema_version": "v1",
"cursor": "<opaque-cursor>",
"app": {
"name": "my-app"
},
"request": {
"id": "request-id",
"timestamp": "2026-05-14T00:00:01.123Z",
"endpoint": "/openai_compatible/v1/chat/completions",
"model": "gpt-4.1",
"provider": "openai",
"stream": false,
"status_code": 200,
"error_type": null,
"error_message": null
},
"tokens": {
"request": 100,
"response": 200,
"cache_read": 0,
"cache_write": null,
"reasoning": null,
"max_requested": 1000
},
"latency_ms": {
"upstream": 800,
"quilr_processing": 120,
"guardrails": 90,
"first_response": 920,
"total": 950
},
"guardrails": {
"outcome": "normal",
"is_blocked": false,
"is_anonymized": false,
"actions_and_categories": {},
"request_predictions": [],
"response_predictions": []
},
"payload": {
"hydration_status": "complete",
"request_text": {},
"response_text": {}
},
"metadata": {
"user_email": null,
"conversation_id": null,
"client_ip": "203.0.113.10",
"extra_data": {},
"sdk": null
},
"routing": {
"group_id": null,
"mode": null
},
"telemetry": {
"processing_times": null,
"chunk_funnel": null
}
}

Top-Level Fields​

FieldTypeDescription
typestringAlways llmgateway.request.
schema_versionstringEvent schema version. Current value is v1.
cursorstringOpaque cursor for this request row.
appobjectApp metadata.
requestobjectGateway request metadata.
tokensobjectToken counts and token limits.
latency_msobjectLatency measurements in milliseconds.
guardrailsobjectGuardrail outcome and prediction metadata.
payloadobjectHydrated request and response payloads when available.
metadataobjectUser, client, SDK, and extra request metadata.
routingobjectRouting group metadata when routing is used.
telemetryobjectAdditional processing telemetry.

app​

FieldTypeDescription
namestringLLM Gateway app name.

request​

FieldTypeDescription
idstringUnique request ID.
timestampstringRequest timestamp in ISO 8601 format.
endpointstringGateway endpoint path used by the request.
modelstring or nullRequested model or routing group name.
providerstring or nullProvider selected for the request.
streambooleanWhether the request used a streaming response mode.
status_codenumber or nullHTTP status code returned to the client.
error_typestring or nullError category when the request failed.
error_messagestring or nullError message when the request failed. Credential-shaped substrings are redacted before export.

tokens​

FieldTypeDescription
requestnumber or nullInput token count.
responsenumber or nullOutput token count.
cache_readnumberTokens read from provider prompt cache. 0 when the provider did not report a cache read.
cache_writenumber or nullTokens written to provider prompt cache, when available.
reasoningnumber or nullReasoning token count, when reported by the provider.
max_requestednumber or nullMaximum output tokens requested by the client.

latency_ms​

FieldTypeDescription
upstreamnumber or nullTime spent waiting on the upstream provider.
quilr_processingnumber or nullTime spent in QuilrAI gateway processing.
guardrailsnumber or nullTime spent evaluating guardrails.
first_responsenumber or nullTime to first response token or first response byte, when available.
totalnumber or nullTotal gateway request duration.

guardrails​

FieldTypeDescription
outcomestring or nullFinal guardrail outcome, such as normal, blocked, or another configured outcome.
is_blockedbooleanWhether the request or response was blocked.
is_anonymizedbooleanWhether anonymization was applied.
actions_and_categoriesobjectGuardrail actions grouped by detected categories.
request_predictionsarrayRequest-side prediction results.
response_predictionsarrayResponse-side prediction results.

Guardian Agent findings are included in these same prediction arrays with match_type: "guardian". Guardian request and response categories are also available under metadata.extra_data.guardian_agent when present.

payload​

FieldTypeDescription
hydration_statusstringcomplete when payload data is available, or missing_prediction when the request log exists but payload hydration is unavailable.
request_textobject, array, string, or nullHydrated request payload. The field name matches the dashboard concept and is not limited to plain strings.
response_textobject, array, string, or nullHydrated response payload. The field name matches the dashboard concept and is not limited to plain strings.

When hydration is unavailable, the payload object uses this shape:

{
"hydration_status": "missing_prediction",
"request_text": null,
"response_text": null
}

metadata​

FieldTypeDescription
user_emailstring or nullUser email associated with the request, when identity-aware tracking is configured. Also present in extra_data when populated.
conversation_idstring or nullConversation ID from X-Conversation-Id, when provided. Also present in extra_data when populated.
client_ipstring or nullClient IP observed by the gateway. Also present in extra_data when populated.
extra_dataobjectAdditional request metadata. The hoisted fields above (user_email, conversation_id, client_ip) are not removed from this object. The jwt_claims field is always stripped.
sdkobject or nullSDK metadata when the request came from SDK mode or a tracked SDK client.

routing​

FieldTypeDescription
group_idstring or nullRouting group identifier when request routing is used.
modestring or nullRouting mode used for the request.

telemetry​

FieldTypeDescription
processing_timesobject or nullAdditional internal processing timings, when available.
chunk_funnelobject or nullStreaming chunk telemetry, when available.

checkpoint​

The final line on a successful response is a checkpoint.

{
"type": "checkpoint",
"schema_version": "v1",
"next_cursor": "<opaque-cursor>",
"rows": 1000,
"has_more": true,
"effective_end_time": "2026-05-14T01:00:00.000Z",
"max_exportable_time": "2026-05-14T10:45:00.000Z"
}
FieldTypeDescription
typestringAlways checkpoint.
schema_versionstringEvent schema version. Current value is v1.
next_cursorstringOpaque cursor to store and use on the next request.
rowsnumberNumber of llmgateway.request events emitted in this response.
has_morebooleantrue when another page is available for the same effective export window.
effective_end_timestringEffective upper bound used for this export response.
max_exportable_timestringNewest timestamp eligible for export after the 15-minute lag.

Metrics View​

Add view=metrics to aggregate the whole export window server-side and return one JSON document shaped for executive, security, and governance dashboards. Use this when you want headline numbers rather than a copy of every request row.

GET https://guardrails.quilr.ai/llmgateway/logs/export?view=metrics
Content-Type: application/json

This view uses the same endpoint, the same export key, and the same window rules as the default logs view: the 15-minute export lag, the 15-day retention limit, and start_time / end_time clamping all behave identically. Two differences:

  • The response is one JSON object, not an NDJSON stream. There are no export_started, llmgateway.request, or checkpoint events, and no pagination.
  • cursor and limit are accepted but ignored, because the whole window is aggregated in a single response.
curl -H "X-Quilr-Log-Export-Key: sk-export-..." \
"https://guardrails.quilr.ai/llmgateway/logs/export?view=metrics&start_time=$START&end_time=$END"

Counting Model​

Every count is request-level: one request contributes at most one to a given metric, no matter how many findings of that kind it carried. A single request that trips four separate PII rules counts once toward pii_phi_pci_violations.total.

Findings are classified by matching the category identifiers recorded on each request, so tenant-defined custom categories roll up into the standard buckets when their identifiers contain the relevant marker (for example a custom category whose id contains jailbreak counts toward jailbreak_attempts).

Example Response​

{
"type": "llmgateway.executive_metrics",
"schema_version": "v1",
"scope": "all_apps",
"window": {
"effective_start_time": "2026-05-07T00:00:00.000Z",
"effective_end_time": "2026-05-14T10:45:00.000Z",
"max_exportable_time": "2026-05-14T10:45:00.000Z",
"end_time_clamped": false
},
"security_metrics": {
"prompt_injection_attempts": 40,
"jailbreak_attempts": 18,
"secrets_exposure_events": 9,
"pii_phi_pci_violations": {
"total": 88,
"pii": 28,
"phi": 0,
"pci": 60,
"pfi": 0
},
"blocked_requests": 67,
"high_severity_incidents": 67
},
"risk_metrics": {
"top_applications_by_violations": [
{
"name": "internal-assistant",
"count": 87
},
{
"name": "support-bot",
"count": 68
}
],
"users_with_repeated_violations": [
{
"name": "a@corp.com",
"count": 100
},
{
"name": "b@corp.com",
"count": 46
},
{
"name": "c@corp.com",
"count": 9
}
],
"top_violation_categories": [
{
"name": "Protected Card Information (PCI)",
"count": 60
},
{
"name": "prompt_injection_direct",
"count": 40
},
{
"name": "Personally Identifiable Information (PII)",
"count": 28
},
{
"name": "jailbreak_roleplay",
"count": 18
},
{
"name": "auth_secrets_api_key",
"count": 9
}
],
"client_data_exposure_attempts": 88
},
"governance_metrics": {
"total_requests": 1245,
"traffic_share_by_app": [
{
"name": "support-bot",
"requests": 768,
"percent": 61.69
},
{
"name": "internal-assistant",
"requests": 387,
"percent": 31.08
},
{
"name": "batch-summarizer",
"requests": 90,
"percent": 7.23
}
],
"integrated_applications": {
"count": 4,
"names": [
"batch-summarizer",
"internal-assistant",
"legacy-classifier",
"support-bot"
],
"active_in_window": [
"batch-summarizer",
"internal-assistant",
"support-bot"
]
},
"applications_without_traffic_in_window": [
"legacy-classifier"
],
"guardrail_effectiveness": {
"detections_total": 155,
"blocked": 67,
"anonymized": 28,
"monitored_only": 60,
"prevented_percent": 61.29
}
},
"executive_kpis": {
"critical_events_prevented": 67,
"secrets_blocked": 9,
"prompt_injection_attempts_blocked": 40,
"jailbreak_attempts_blocked": 18,
"sensitive_records_protected": 28
},
"coverage": {
"scanned_requests": 1245,
"complete": true,
"max_scan_rows": 250000
}
}

security_metrics​

FieldTypeDescription
prompt_injection_attemptsnumberRequests where a prompt-injection category was detected.
jailbreak_attemptsnumberRequests where a jailbreak category was detected.
secrets_exposure_eventsnumberRequests where a secret or credential category was detected.
pii_phi_pci_violations.totalnumberRequests with any data-risk detection. Less than or equal to the sum of the breakdown, since one request can carry several data types.
pii_phi_pci_violations.piinumberRequests with a PII detection.
pii_phi_pci_violations.phinumberRequests with a PHI detection.
pii_phi_pci_violations.pcinumberRequests with a PCI detection.
pii_phi_pci_violations.pfinumberRequests with a PFI detection.
blocked_requestsnumberRequests the gateway blocked outright.
high_severity_incidentsnumberRequests carrying prompt-injection, jailbreak, security-exploit, or secrets detections.

risk_metrics​

FieldTypeDescription
top_applications_by_violationsarrayUp to 10 {name, count} entries, highest first, for apps with at least one violation.
users_with_repeated_violationsarrayUp to 10 {name, count} entries for users with 2 or more violations, taken from extra_data.user_email. Empty when identity headers are not in use.
top_violation_categoriesarrayUp to 10 {name, count} entries by category display name.
client_data_exposure_attemptsnumberRequests with any data-risk detection.

governance_metrics​

FieldTypeDescription
total_requestsnumberEvery request that reached the gateway in the window.
traffic_share_by_apparray{name, requests, percent} per application, highest first, ties broken by name.
integrated_applications.countnumberNumber of apps configured under this export scope.
integrated_applications.namesarrayAll configured app names.
integrated_applications.active_in_windowarrayConfigured apps that served at least one request in the window.
applications_without_traffic_in_windowarrayConfigured apps that served no traffic in the window.
guardrail_effectiveness.detections_totalnumberRequests that were blocked, anonymized, or flagged by a monitoring rule.
guardrail_effectiveness.blockednumberRequests blocked.
guardrail_effectiveness.anonymizednumberRequests anonymized.
guardrail_effectiveness.monitored_onlynumberRequests flagged but allowed through.
guardrail_effectiveness.prevented_percentnumber | nullBlocked plus anonymized, as a percentage of detections_total. null when there were no detections.

percent in traffic_share_by_app is each app's share of total_requests, so the values describe how gateway traffic is distributed across your applications and sum to approximately 100 (individual values are rounded to two decimals). Applications with no traffic are omitted from this array and listed in applications_without_traffic_in_window instead.

This is deliberately a share of traffic that reached the gateway, not a share of all AI usage in your organisation. The gateway can only observe requests routed through it, so it cannot measure traffic that bypasses it.

executive_kpis​

FieldTypeDescription
critical_events_preventednumberBlocked requests carrying prompt-injection, jailbreak, security-exploit, or secrets detections.
secrets_blockednumberBlocked requests carrying a secrets detection.
prompt_injection_attempts_blockednumberBlocked requests carrying a prompt-injection detection.
jailbreak_attempts_blockednumberBlocked requests carrying a jailbreak detection.
sensitive_records_protectednumberRequests where data-risk content was blocked or anonymized.

coverage​

FieldTypeDescription
scanned_requestsnumberRequests aggregated for this response.
completebooleanfalse when the window exceeded the scan cap and the numbers are therefore partial. Narrow the window and retry.
max_scan_rowsnumberMaximum rows a single metrics response will scan.

Always check coverage.complete. When it is false, the figures cover only the first max_scan_rows requests in the window and should not be reported as totals.

Metrics View Errors​

Unlike the logs view, metrics errors are returned as a plain JSON object rather than NDJSON, because no stream has started.

StatusCodeCause
400invalid_viewview was neither logs nor metrics.
500metrics_failedAggregation failed. Retry, and narrow the window if it persists.

All authentication and time-window errors behave exactly as they do for the logs view.

Redaction​

The export endpoint applies a best-effort credential scrub before emitting any row. Expect the following to be missing or rewritten in exported events:

  • extra_data.jwt_claims is removed from every row.
  • Any object key named headers, request_headers, response_headers, or http_headers is replaced with the string [REDACTED_HEADERS].
  • Object keys that name a credential (such as authorization, api_key, x_api_key, quilr_api_key, access_token, refresh_token, client_secret, password, private_key, AWS credential field names) and any key suffixed with _api_key, _apikey, _access_token, _refresh_token, _client_secret, or _private_key are replaced with [REDACTED].
  • String values are scanned for common credential patterns. Matches are rewritten to placeholders such as [REDACTED_API_KEY], [REDACTED_QUILR_API_KEY], [REDACTED_LOG_EXPORT_KEY], or Bearer [REDACTED].

Redaction is applied recursively to payload.request_text, payload.response_text, guardrails.actions_and_categories, guardrails.request_predictions, guardrails.response_predictions, metadata.extra_data, metadata.sdk, telemetry.processing_times, telemetry.chunk_funnel, and the top-level error_message field. This is a safety layer, not a formal DLP pass over exported payloads.

Errors​

Errors are returned as NDJSON too.

{"type":"error","error":{"message":"<message>","code":"<code>"}}
FieldTypeDescription
typestringAlways error.
error.messagestringHuman-readable error message.
error.codestringMachine-readable error code.

Errors before streaming starts return an HTTP error status with a single NDJSON error line as the response body. Errors after streaming has started return HTTP 200 and emit an error event line in the body because the HTTP response has already been committed.