Core Concepts

Token Counting & Estimation

Tokens are the fundamental units of text processed by Cortiqa reasoning engines. Understanding how words translate to tokens helps optimize request size and predict inference costs.


What Are Tokens

Tokens are subword fragments. In English, 1 token is approximately 4 characters or 0.75 words. For code, punctuation and whitespace characters are also tokenized individually.

Tracking Token Usage

The Cortiqa API reports exact token accounting in every completion response:

count.py
from cortiqa import Cortiqa

client = Cortiqa()

response = client.chat.completions.create(
    model="openai/gpt-oss-120b",
    messages=[{"role": "user", "content": "Explain binary search trees."}]
)

usage = response.usage
print(f"Prompt Tokens: {usage.prompt_tokens}")
print(f"Completion Tokens: {usage.completion_tokens}")
print(f"Total Tokens: {usage.total_tokens}")

Rule of Thumb Estimation

Text UnitApproximate Tokens
1 word~1.3 tokens
1 sentence~20 - 30 tokens
1 paragraph~75 - 120 tokens
1 page of prose (500 words)~650 tokens
1,000 lines of Python code~4,000 tokens
Was this page helpful?