Guides

Prompt & Context Caching

Prompt caching stores the computed key-value (KV) attention states of long prompts directly in high-speed LPU memory. Repeated requests sharing common prefixes bypass redundant prefill computation, delivering instant time-to-first-token.


How It Works

When requests start with the exact same prefix (e.g., a 20,000 token legal document or enterprise knowledge base followed by changing user questions), Cortiqa automatically detects the match and serves the precomputed KV cache.

Optimizing for Cache Hits

No special headers are required—caching is built into the Cortiqa inference engine. To maximize cache hits, structure your prompts with static content first:

cache_structure.py
# Prompt Structure for Optimal Cache Re-use:
messages = [
    # 1. STATIC: Long documentation or reference manual (Cached!)
    {"role": "system", "content": LARGE_KNOWLEDGE_BASE_TEXT},
    # 2. STATIC: Few-shot examples or schema definitions (Cached!)
    {"role": "system", "content": "Always output in JSON conforming to Schema X."},
    # 3. DYNAMIC: Changing user question (Appended at end)
    {"role": "user", "content": "How do I configure OAuth2?"}
]

Cache Eviction & Lifetime

Cached prefixes remain hot in LPU memory using an LRU (Least Recently Used) policy. Active prefixes accessed every few minutes remain resident indefinitely.

Pricing & Cost Savings

Cached input tokens are billed at a 50% discount compared to uncached inputs ($0.075 / 1M tokens vs $0.15 / 1M tokens for openai/gpt-oss-120b), making high-volume document retrieval highly economical.

Was this page helpful?