Platform

Logs & Observability

Gain end-to-end visibility into your AI workloads. Inspect individual inference events, monitor token expenditure, and diagnose production issues with millisecond-resolution telemetry.


Live Request Logs

The Logs Section in Console captures all inbound API requests in real time. Each log entry includes:

  • Timestamp: UTC timestamp with millisecond precision.
  • Status Code: HTTP 200 (Success), 400 (Bad Request), 401 (Auth Failure), 429 (Rate Limit), or 500 (Internal Server Error).
  • Model: e.g. openai/gpt-oss-120b.
  • Tokens: Total prompt and completion token counts.
  • Duration: Time-to-first-token (TTFT) and total generation time.

Tracing with Request IDs

Every response from api.cortiqa.co returns a unique tracing header: x-request-id:

HTTP Response Header
x-request-id: req_01JHE8B1VNMZ8PQ5

When contacting support@cortiqa.co regarding an unexpected error or slow response, always provide the x-request-id so our engineers can instantly pull server-side diagnostic traces.

Usage & Latency Metrics

The console visualizes key SLIs (Service Level Indicators):

  • Tokens Per Second (TPS): Model generation speed on Cortiqa LPU hardware.
  • P50 / P95 / P99 Latency: Response time percentiles to measure tail latency.
  • Cache Hit Ratio: Percentage of requests served from prompt cache.

Automated Alerts & Webhooks

Configure proactive alerts to receive Slack, Discord, or Email notifications when:

  • Error rate exceeds 1% of total requests over a 5-minute window.
  • Monthly token spend crosses 80% or 100% of your configured budget.
  • Unusual request spikes occur from unexpected IP ranges.
Was this page helpful?