Core Concepts
Rate limits ensure infrastructure reliability and fair access across developer workloads. Limits are evaluated based on Requests Per Minute (RPM) and Tokens Per Minute (TPM).
| Tier | Requests/min (RPM) | Tokens/min (TPM) | Daily Limit |
|---|---|---|---|
| Starter / Free | 60 RPM | 100,000 TPM | 1,000,000 tokens/day |
| Pro / Scale | 500 RPM | 1,000,000 TPM | 50,000,000 tokens/day |
| Enterprise | Custom SLA | Uncapped / Dedicated LPU | Custom Volume |
Every response returns live quota headers so your applications can dynamically monitor available bandwidth:
x-ratelimit-limit-requests: 500
x-ratelimit-limit-tokens: 1000000
x-ratelimit-remaining-requests: 489
x-ratelimit-remaining-tokens: 978400
x-ratelimit-reset-requests: 2026-09-22T14:30:00Z
x-ratelimit-reset-tokens: 2026-09-22T14:30:00ZIf you implement custom HTTP calls, catch status 429 and retry with exponential backoff:
import time
import random
from cortiqa import Cortiqa
from cortiqa.exceptions import RateLimitError
client = Cortiqa()
def execute_with_backoff(prompt, max_retries=3):
for attempt in range(max_retries):
try:
return client.chat.completions.create(
model="openai/gpt-oss-120b",
messages=[{"role": "user", "content": prompt}]
)
except RateLimitError as e:
if attempt == max_retries - 1:
raise e
sleep_duration = (2 ** attempt) + random.uniform(0.1, 0.5)
time.sleep(sleep_duration)You can request quota increases directly from platform.cortiqa.co/billing, or reach out to team@cortiqa.co for dedicated enterprise throughput agreements.