LLM Gateway Log Export API
Use the Log Export API to read LLM Gateway request logs from your own data platform, SIEM, warehouse, or scheduled export job.
The API returns newline-delimited JSON. Each response line is one complete JSON object, so clients can stream, parse, and checkpoint logs incrementally.
GET https://guardrails.quilr.ai/llmgateway/logs/export
Response content type:
Content-Type: application/x-ndjson
The endpoint also serves an aggregated executive dashboard view. Add view=metrics
to receive a single summary JSON document instead of the per-request NDJSON stream.
See Metrics View.
Authentication
Pass a log export key from the QuilrAI LLM Gateway UI:
X-Quilr-Log-Export-Key: sk-export-...
Do not use your QuilrAI gateway API key as the request credential for this endpoint. The log export key is separate from model-call authentication.
The UI exposes two export scopes:
Both scopes use the same endpoint, header, query parameters, pagination model, and response format. In all-apps exports, each llmgateway.request event still includes the concrete app name in app.name.
Query Parameters
All query parameters are optional.
Logs are available for a maximum of 15 days. Choose start_time within that retention window when backfilling. Requests with an effective start_time, end_time, or cursor timestamp before the retention window fail with 400.
If neither start_time nor cursor is provided, the API exports a default 24-hour window ending at the effective export end time.
Export Lag
The API does not export logs newer than 15 minutes. Gateway logs and prediction payloads are written asynchronously, so this lag keeps exported rows stable.
If end_time is newer than now - 15 minutes, the server clamps it to the maximum exportable time. The request still succeeds. The export_started and checkpoint events include the effective export bounds.
Request Examples
Start an export window, here the hour that ended two hours ago (logs are kept for 15 days, so fixed dates go stale):
# macOS (BSD date) first, GNU date as fallback
START=$(date -u -v-3H +%Y-%m-%dT%H:%M:%SZ 2>/dev/null || date -u -d '3 hours ago' +%Y-%m-%dT%H:%M:%SZ)
END=$(date -u -v-2H +%Y-%m-%dT%H:%M:%SZ 2>/dev/null || date -u -d '2 hours ago' +%Y-%m-%dT%H:%M:%SZ)
curl -N \
-H "X-Quilr-Log-Export-Key: sk-export-..." \
"https://guardrails.quilr.ai/llmgateway/logs/export?start_time=$START&end_time=$END&limit=1000"
Resume from the previous checkpoint:
curl -N \
-H "X-Quilr-Log-Export-Key: sk-export-..." \
"https://guardrails.quilr.ai/llmgateway/logs/export?cursor=<next_cursor>"
When resuming with cursor, you do not need to pass start_time or end_time.
Pagination
Rows are ordered by:
timestamp ASC, request_id ASC
The cursor is opaque. Store it exactly as returned in checkpoint.next_cursor and send it back as the cursor query parameter on the next request.
If checkpoint.has_more is true, call the endpoint again immediately with cursor=<next_cursor>.
If checkpoint.has_more is false, there are no more rows in the current effective window. Store next_cursor and poll later with that cursor to continue incremental export.
When an initial request (no cursor supplied) returns zero rows, the API returns a checkpoint cursor pinned to the effective end time. This lets exporters store one cursor value even for empty windows.
When a request with a cursor returns zero rows, checkpoint.next_cursor echoes the inbound cursor unchanged and has_more is false. Re-poll later with the same cursor.
A reliable collector
A minimal Python collector (standard library only) that you can run on a schedule or as a service. It:
- stores rows and the cursor in one SQLite transaction, so the cursor only advances after the page is safely written;
- writes idempotently (
INSERT OR IGNOREon a unique id), so a replayed page never creates duplicates; - treats a mid-stream
errorevent or a stream without acheckpointas a failed page, and retries it with exponential backoff and jitter; - stops on
400,401and403, which retrying cannot fix.
"""Minimal LLM Gateway log collector: cursor checkpoint, retries, idempotent writes."""
import json, os, random, sqlite3, time, urllib.error, urllib.parse, urllib.request
from datetime import datetime, timedelta, timezone
URL = os.environ.get("EXPORT_URL", "https://guardrails.quilr.ai/llmgateway/logs/export")
KEY = os.environ["QUILR_LOG_EXPORT_KEY"]
EVENT = "llmgateway.request"
row_id = lambda e: e["request"]["id"] # unique per request
db = sqlite3.connect("quilr_logs.db")
db.execute("CREATE TABLE IF NOT EXISTS events (id TEXT PRIMARY KEY, ts TEXT, body TEXT)")
db.execute("CREATE TABLE IF NOT EXISTS state (k TEXT PRIMARY KEY, v TEXT)")
def saved_cursor():
row = db.execute("SELECT v FROM state WHERE k = 'cursor'").fetchone()
return row[0] if row else None
def fetch_page(params):
"""Return (events, checkpoint) for one complete page, or raise."""
req = urllib.request.Request(URL + "?" + urllib.parse.urlencode(params),
headers={"X-Quilr-Log-Export-Key": KEY})
events, checkpoint = [], None
with urllib.request.urlopen(req, timeout=300) as r: # raises HTTPError on 4xx/5xx
for line in r:
if not line.strip():
continue
msg = json.loads(line)
if msg["type"] == "error": # can arrive after HTTP 200
raise RuntimeError(f"export error: {msg['error']}")
if msg["type"] == EVENT:
events.append(msg)
elif msg["type"] == "checkpoint":
checkpoint = msg
if checkpoint is None:
raise RuntimeError("stream ended without a checkpoint")
return events, checkpoint
def with_retries(fn, *args, attempts=6):
for n in range(attempts):
try:
return fn(*args)
except urllib.error.HTTPError as e:
if e.code in (400, 401, 403) or n == attempts - 1:
raise # fix the key or the window; retrying will not help
except (OSError, RuntimeError, ValueError): # network, mid-stream error, truncation
if n == attempts - 1:
raise
time.sleep(min(60, 2 ** n) + random.random()) # backoff with jitter
def run_once():
cursor = saved_cursor()
if cursor:
params = {"cursor": cursor, "limit": 1000}
else: # first run: start one hour back, well inside the 15-day retention
start = datetime.now(timezone.utc) - timedelta(hours=1)
params = {"start_time": start.strftime("%Y-%m-%dT%H:%M:%SZ"), "limit": 1000}
while True:
events, checkpoint = with_retries(fetch_page, params)
with db: # one transaction: rows and cursor commit together
db.executemany(
"INSERT OR IGNORE INTO events VALUES (?, ?, ?)",
[(row_id(e), e["request"]["timestamp"], json.dumps(e)) for e in events])
db.execute("INSERT OR REPLACE INTO state VALUES ('cursor', ?)",
(checkpoint["next_cursor"],))
print(f"stored {checkpoint['rows']} rows up to {checkpoint['effective_end_time']}")
if not checkpoint["has_more"]:
return
params = {"cursor": checkpoint["next_cursor"], "limit": 1000}
if __name__ == "__main__":
while True:
run_once()
time.sleep(300) # new rows appear about 15 minutes after the request
export QUILR_LOG_EXPORT_KEY='sk-export-...'
python3 collector.py
To forward to a SIEM instead of SQLite, replace the with db: block with your sink's write, and persist next_cursor only after the sink acknowledges the batch. If the collector is down for longer than the 15-day retention, the saved cursor fails with 400; delete it to restart from a recent start_time.
Verify delivery
- Send one request through an app covered by the export key and note the time.
- After at least 15 minutes, run the collector once.
- Check that the row arrived:
sqlite3 quilr_logs.db "SELECT id, ts FROM events ORDER BY ts DESC LIMIT 5". - For a closed window, compare the number of rows you stored with
governance_metrics.total_requestsfrom the metrics view for the samestart_timeandend_time(checkcoverage.completeistrue). A gap means pages were skipped; re-export that window, and the idempotent writes absorb the overlap.
Coverage
The export covers LLM Gateway traffic for the selected export scope, including:
Response Events
Every successful response starts with export_started, contains zero or more llmgateway.request events, and ends with checkpoint.
export_started
The first line describes the effective export window.
{
"type": "export_started",
"schema_version": "v1",
"scope": "app",
"app_name": "my-app",
"app_count": 1,
"effective_start_time": "2026-05-14T00:00:00.000Z",
"effective_end_time": "2026-05-14T01:00:00.000Z",
"max_exportable_time": "2026-05-14T10:45:00.000Z",
"end_time_clamped": false,
"limit": 1000
}
For all-apps export, the first line uses this scope shape:
{
"type": "export_started",
"schema_version": "v1",
"scope": "all_apps",
"app_name": null,
"app_count": 3,
"effective_start_time": "2026-05-14T00:00:00.000Z",
"effective_end_time": "2026-05-14T01:00:00.000Z",
"max_exportable_time": "2026-05-14T10:45:00.000Z",
"end_time_clamped": false,
"limit": 1000
}
llmgateway.request
Each request row is emitted as one llmgateway.request event.
{
"type": "llmgateway.request",
"schema_version": "v1",
"cursor": "<opaque-cursor>",
"app": {
"name": "my-app"
},
"request": {
"id": "request-id",
"timestamp": "2026-05-14T00:00:01.123Z",
"endpoint": "/openai_compatible/v1/chat/completions",
"model": "gpt-4.1",
"provider": "openai",
"stream": false,
"status_code": 200,
"error_type": null,
"error_message": null
},
"tokens": {
"request": 100,
"response": 200,
"cache_read": 0,
"cache_write": null,
"reasoning": null,
"max_requested": 1000
},
"latency_ms": {
"upstream": 800,
"quilr_processing": 120,
"guardrails": 90,
"first_response": 920,
"total": 950
},
"guardrails": {
"outcome": "normal",
"is_blocked": false,
"is_anonymized": false,
"actions_and_categories": {},
"request_predictions": [],
"response_predictions": []
},
"payload": {
"hydration_status": "complete",
"request_text": {},
"response_text": {}
},
"metadata": {
"user_email": null,
"conversation_id": null,
"client_ip": "203.0.113.10",
"extra_data": {},
"sdk": null
},
"routing": {
"group_id": null,
"mode": null
},
"telemetry": {
"processing_times": null,
"chunk_funnel": null
}
}
Top-Level Fields
app
request
tokens
latency_ms
guardrails
Guardian Agent findings are included in these same prediction arrays with match_type: "guardian". Guardian request and response categories are also available under metadata.extra_data.guardian_agent when present.
payload
When hydration is unavailable, the payload object uses this shape:
{
"hydration_status": "missing_prediction",
"request_text": null,
"response_text": null
}
metadata
routing
telemetry
checkpoint
The final line on a successful response is a checkpoint.
{
"type": "checkpoint",
"schema_version": "v1",
"next_cursor": "<opaque-cursor>",
"rows": 1000,
"has_more": true,
"effective_end_time": "2026-05-14T01:00:00.000Z",
"max_exportable_time": "2026-05-14T10:45:00.000Z"
}
Metrics View
Add view=metrics to aggregate the whole export window server-side and return one
JSON document shaped for executive, security, and governance dashboards. Use this
when you want headline numbers rather than a copy of every request row.
GET https://guardrails.quilr.ai/llmgateway/logs/export?view=metrics
Content-Type: application/json
This view uses the same endpoint, the same export key, and the same window rules as
the default logs view: the 15-minute export lag, the 15-day retention limit, and
start_time / end_time clamping all behave identically. Two differences:
- The response is one JSON object, not an NDJSON stream. There are no
export_started,llmgateway.request, orcheckpointevents, and no pagination. cursorandlimitare accepted but ignored, because the whole window is aggregated in a single response.
curl -H "X-Quilr-Log-Export-Key: sk-export-..." \
"https://guardrails.quilr.ai/llmgateway/logs/export?view=metrics&start_time=$START&end_time=$END"
Counting Model
Every count is request-level: one request contributes at most one to a given
metric, no matter how many findings of that kind it carried. A single request that
trips four separate PII rules counts once toward pii_phi_pci_violations.total.
Findings are classified by matching the category identifiers recorded on each
request, so tenant-defined custom categories roll up into the standard buckets when
their identifiers contain the relevant marker (for example a custom category whose
id contains jailbreak counts toward jailbreak_attempts).
Example Response
{
"type": "llmgateway.executive_metrics",
"schema_version": "v1",
"scope": "all_apps",
"window": {
"effective_start_time": "2026-05-07T00:00:00.000Z",
"effective_end_time": "2026-05-14T10:45:00.000Z",
"max_exportable_time": "2026-05-14T10:45:00.000Z",
"end_time_clamped": false
},
"security_metrics": {
"prompt_injection_attempts": 40,
"jailbreak_attempts": 18,
"secrets_exposure_events": 9,
"pii_phi_pci_violations": {
"total": 88,
"pii": 28,
"phi": 0,
"pci": 60,
"pfi": 0
},
"blocked_requests": 67,
"high_severity_incidents": 67
},
"risk_metrics": {
"top_applications_by_violations": [
{
"name": "internal-assistant",
"count": 87
},
{
"name": "support-bot",
"count": 68
}
],
"users_with_repeated_violations": [
{
"name": "a@corp.com",
"count": 100
},
{
"name": "b@corp.com",
"count": 46
},
{
"name": "c@corp.com",
"count": 9
}
],
"top_violation_categories": [
{
"name": "Protected Card Information (PCI)",
"count": 60
},
{
"name": "prompt_injection_direct",
"count": 40
},
{
"name": "Personally Identifiable Information (PII)",
"count": 28
},
{
"name": "jailbreak_roleplay",
"count": 18
},
{
"name": "auth_secrets_api_key",
"count": 9
}
],
"client_data_exposure_attempts": 88
},
"governance_metrics": {
"total_requests": 1245,
"traffic_share_by_app": [
{
"name": "support-bot",
"requests": 768,
"percent": 61.69
},
{
"name": "internal-assistant",
"requests": 387,
"percent": 31.08
},
{
"name": "batch-summarizer",
"requests": 90,
"percent": 7.23
}
],
"integrated_applications": {
"count": 4,
"names": [
"batch-summarizer",
"internal-assistant",
"legacy-classifier",
"support-bot"
],
"active_in_window": [
"batch-summarizer",
"internal-assistant",
"support-bot"
]
},
"applications_without_traffic_in_window": [
"legacy-classifier"
],
"guardrail_effectiveness": {
"detections_total": 155,
"blocked": 67,
"anonymized": 28,
"monitored_only": 60,
"prevented_percent": 61.29
}
},
"executive_kpis": {
"critical_events_prevented": 67,
"secrets_blocked": 9,
"prompt_injection_attempts_blocked": 40,
"jailbreak_attempts_blocked": 18,
"sensitive_records_protected": 28
},
"coverage": {
"scanned_requests": 1245,
"complete": true,
"max_scan_rows": 250000
}
}
security_metrics
risk_metrics
governance_metrics
percent in traffic_share_by_app is each app's share of total_requests, so the
values describe how gateway traffic is distributed across your applications and sum
to approximately 100 (individual values are rounded to two decimals). Applications
with no traffic are omitted from this array and listed in
applications_without_traffic_in_window instead.
This is deliberately a share of traffic that reached the gateway, not a share of all AI usage in your organisation. The gateway can only observe requests routed through it, so it cannot measure traffic that bypasses it.
executive_kpis
coverage
Always check coverage.complete. When it is false, the figures cover only the
first max_scan_rows requests in the window and should not be reported as totals.
Metrics View Errors
Unlike the logs view, metrics errors are returned as a plain JSON object rather than NDJSON, because no stream has started.
All authentication and time-window errors behave exactly as they do for the logs view.
Redaction
The export endpoint applies a best-effort credential scrub before emitting any row. Expect the following to be missing or rewritten in exported events:
extra_data.jwt_claimsis removed from every row.- Any object key named
headers,request_headers,response_headers, orhttp_headersis replaced with the string[REDACTED_HEADERS]. - Object keys that name a credential (such as
authorization,api_key,x_api_key,quilr_api_key,access_token,refresh_token,client_secret,password,private_key, AWS credential field names) and any key suffixed with_api_key,_apikey,_access_token,_refresh_token,_client_secret, or_private_keyare replaced with[REDACTED]. - String values are scanned for common credential patterns. Matches are rewritten to placeholders such as
[REDACTED_API_KEY],[REDACTED_QUILR_API_KEY],[REDACTED_LOG_EXPORT_KEY], orBearer [REDACTED].
Redaction is applied recursively to payload.request_text, payload.response_text, guardrails.actions_and_categories, guardrails.request_predictions, guardrails.response_predictions, metadata.extra_data, metadata.sdk, telemetry.processing_times, telemetry.chunk_funnel, and the top-level error_message field. This is a safety layer, not a formal DLP pass over exported payloads.
Errors
Errors are returned as NDJSON too.
{"type":"error","error":{"message":"<message>","code":"<code>"}}
Errors before streaming starts return an HTTP error status with a single NDJSON error line as the response body. Errors after streaming has started return HTTP 200 and emit an error event line in the body because the HTTP response has already been committed.