Core Concepts

Rate Limits

Rate limits ensure infrastructure reliability and fair access across developer workloads. Limits are evaluated based on Requests Per Minute (RPM) and Tokens Per Minute (TPM).


Tier Rate Limits

TierRequests/min (RPM)Tokens/min (TPM)Daily Limit
Starter / Free60 RPM100,000 TPM1,000,000 tokens/day
Pro / Scale500 RPM1,000,000 TPM50,000,000 tokens/day
EnterpriseCustom SLAUncapped / Dedicated LPUCustom Volume

Rate Limit HTTP Headers

Every response returns live quota headers so your applications can dynamically monitor available bandwidth:

HTTP Response Headers
x-ratelimit-limit-requests: 500
x-ratelimit-limit-tokens: 1000000
x-ratelimit-remaining-requests: 489
x-ratelimit-remaining-tokens: 978400
x-ratelimit-reset-requests: 2026-09-22T14:30:00Z
x-ratelimit-reset-tokens: 2026-09-22T14:30:00Z

Handling 429 Status Codes

Automatic SDK Retries
All official Cortiqa SDKs (Python, TypeScript, Go, Java, C#, Swift) include automatic retry logic with exponential backoff and randomized jitter by default.

If you implement custom HTTP calls, catch status 429 and retry with exponential backoff:

backoff.py
import time
import random
from cortiqa import Cortiqa
from cortiqa.exceptions import RateLimitError

client = Cortiqa()

def execute_with_backoff(prompt, max_retries=3):
    for attempt in range(max_retries):
        try:
            return client.chat.completions.create(
                model="openai/gpt-oss-120b",
                messages=[{"role": "user", "content": prompt}]
            )
        except RateLimitError as e:
            if attempt == max_retries - 1:
                raise e
            sleep_duration = (2 ** attempt) + random.uniform(0.1, 0.5)
            time.sleep(sleep_duration)

Increasing Rate Limits

You can request quota increases directly from platform.cortiqa.co/billing, or reach out to team@cortiqa.co for dedicated enterprise throughput agreements.

Was this page helpful?