Core Concepts

Context Windows & Token Limits

A context window defines the maximum amount of text (in tokens) that a model can process in a single inference session. This includes both the input prompt history and the model's output tokens.


Overview

Cortiqa's active flagship model, GPT-OSS 120B (openai/gpt-oss-120b), features an expansive 131,072 token (131K) context window running with KV-cache optimization on Cortiqa LPU hardware. This allows you to ingest entire code repositories, research papers, or dozens of chat turns in a single prompt.

Token Limits by Model

ModelContext WindowMax OutputStatus
GPT-OSS 120B (openai/gpt-oss-120b)131,072 tokens (131K)4,096 tokensActive / Production
Falin 01 (falin-01)128,000 tokens (128K)4,096 tokensComing Soon
Falin Pro (falin-pro)200,000 tokens (200K)8,192 tokensComing Soon
Falin Vision (falin-vision)128,000 tokens (128K)4,096 tokensComing Soon
Falin Ultra (falin-ultra)256,000 tokens (256K)8,192 tokensComing Soon

Inspecting Token Usage

Every response returns precise token consumption metrics in the usage payload:

token_usage.py
from cortiqa import Cortiqa

client = Cortiqa()

response = client.chat.completions.create(
    model="openai/gpt-oss-120b",
    messages=[{"role": "user", "content": "Analyze this query."}]
)

print(f"Prompt tokens: {response.usage.prompt_tokens}")
print(f"Completion tokens: {response.usage.completion_tokens}")
print(f"Total tokens: {response.usage.total_tokens}")

Optimization Strategies

  • Sliding Window: For multi-turn chats, prune old messages when total tokens approach the limit.
  • Summarization: Condense older historical conversation turns into a short recap system message.
  • Chunking: Divide multi-megabyte documents into semantically coherent sections before prompting.
Was this page helpful?