Guides

Multimodal & Vision

Cortiqa models can process both text and images in a unified request, enabling visual reasoning, document OCR, chart analysis, and user interface inspection.


Overview

You can provide images as base64-encoded strings or public URLs inside the message content array.

Dedicated Vision Architecture
The upcoming dedicated Falin Vision model (falin-vision, 128K context) will provide state-of-the-art native document perception, fine-grained diagram OCR, and spatial UI navigation.

Sending Images

vision_demo.py
1import base64
2from cortiqa import Cortiqa
3
4client = Cortiqa()
5
6with open("architecture_diagram.png", "rb") as f:
7 base64_image = base64.b64encode(f.read()).decode("utf-8")
8
9response = client.chat.completions.create(
10 model="openai/gpt-oss-120b",
11 messages=[
12 {
13 "role": "user",
14 "content": [
15 {"type": "text", "text": "Analyze this architecture diagram and highlight any single points of failure."},
16 {
17 "type": "image_url",
18 "image_url": {
19 "url": f"data:image/png;base64,{base64_image}"
20 }
21 }
22 ]
23 }
24 ]
25)
26
27print(response.choices[0].message.content)

Supported Formats

FormatMIME TypeMax Upload Size
PNGimage/png20 MB
JPEGimage/jpeg20 MB
WebPimage/webp20 MB
GIFimage/gif20 MB

Best Practices

  • Resolution: High resolution images (e.g. 1080p - 4K) yield much better text extraction for diagrams and spreadsheets.
  • Clarity: Crop irrelevant borders or margins before encoding to minimize token usage and latency.
  • Prompts: Formulate specific visual questions (e.g., “What is the value in row 3 column 2?” rather than “What is this?”).
Was this page helpful?