Gain end-to-end visibility into your AI workloads. Inspect individual inference events, monitor token expenditure, and diagnose production issues with millisecond-resolution telemetry.
Live Request Logs
The Logs Section in Console captures all inbound API requests in real time. Each log entry includes:
Timestamp: UTC timestamp with millisecond precision.
Status Code: HTTP 200 (Success), 400 (Bad Request), 401 (Auth Failure), 429 (Rate Limit), or 500 (Internal Server Error).
Model: e.g. openai/gpt-oss-120b.
Tokens: Total prompt and completion token counts.
Duration: Time-to-first-token (TTFT) and total generation time.
Tracing with Request IDs
Every response from api.cortiqa.co returns a unique tracing header: x-request-id:
HTTP Response Header
x-request-id: req_01JHE8B1VNMZ8PQ5
When contacting support@cortiqa.co regarding an unexpected error or slow response, always provide the x-request-id so our engineers can instantly pull server-side diagnostic traces.
Usage & Latency Metrics
The console visualizes key SLIs (Service Level Indicators):
Tokens Per Second (TPS): Model generation speed on Cortiqa LPU hardware.
P50 / P95 / P99 Latency: Response time percentiles to measure tail latency.
Cache Hit Ratio: Percentage of requests served from prompt cache.
Automated Alerts & Webhooks
Configure proactive alerts to receive Slack, Discord, or Email notifications when:
Error rate exceeds 1% of total requests over a 5-minute window.
Monthly token spend crosses 80% or 100% of your configured budget.
Unusual request spikes occur from unexpected IP ranges.