Guides
Prompt caching stores the computed key-value (KV) attention states of long prompts directly in high-speed LPU memory. Repeated requests sharing common prefixes bypass redundant prefill computation, delivering instant time-to-first-token.
When requests start with the exact same prefix (e.g., a 20,000 token legal document or enterprise knowledge base followed by changing user questions), Cortiqa automatically detects the match and serves the precomputed KV cache.
No special headers are required—caching is built into the Cortiqa inference engine. To maximize cache hits, structure your prompts with static content first:
# Prompt Structure for Optimal Cache Re-use:
messages = [
# 1. STATIC: Long documentation or reference manual (Cached!)
{"role": "system", "content": LARGE_KNOWLEDGE_BASE_TEXT},
# 2. STATIC: Few-shot examples or schema definitions (Cached!)
{"role": "system", "content": "Always output in JSON conforming to Schema X."},
# 3. DYNAMIC: Changing user question (Appended at end)
{"role": "user", "content": "How do I configure OAuth2?"}
]Cached prefixes remain hot in LPU memory using an LRU (Least Recently Used) policy. Active prefixes accessed every few minutes remain resident indefinitely.
Cached input tokens are billed at a 50% discount compared to uncached inputs ($0.075 / 1M tokens vs $0.15 / 1M tokens for openai/gpt-oss-120b), making high-volume document retrieval highly economical.