Guides

Vector Embeddings

Cortiqa embeddings convert natural language text into dense, high-dimensional floating-point vectors that capture semantic relationships, concepts, and intent.


Overview

Embeddings enable computers to understand semantic similarity. Texts with similar meanings (such as “How do I reset my password?” and “Forgot credentials recovery”) map close together in vector space.

Generating Embeddings

You can generate embeddings via direct REST HTTP or using any OpenAI-compatible client library:

Terminal
curl https://api.cortiqa.co/api/v1/embeddings \
  -H "Authorization: Bearer sk-cortiqa-YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "text-embedding-3-small",
    "input": "Cortiqa is an ultra-fast AI inference platform."
  }'

Example response containing the 1536-dimensional float vector:

Response
{
  "object": "list",
  "data": [
    {
      "object": "embedding",
      "index": 0,
      "embedding": [-0.0069, -0.0053, 0.0125, -0.0241, 0.0089, ...]
    }
  ],
  "model": "text-embedding-3-small",
  "usage": {
    "prompt_tokens": 10,
    "total_tokens": 10
  }
}

Common Use Cases

  • Retrieval-Augmented Generation (RAG): Embed documentation chunks and retrieve the top-K relevant passages before querying Cortiqa models.
  • Semantic Search: Match user queries to product catalogs and knowledge bases based on intent rather than exact keyword matches.
  • Clustering & Topic Analysis: Group large volumes of customer feedback or support tickets into automated clusters.
  • Anomaly & Duplicate Detection: Flag redundant database records or unusual textual inputs.

Calculating Cosine Similarity

Use standard cosine similarity or inner product (dot product on normalized vectors) to compare two vectors:

similarity.py
import numpy as np

def cosine_similarity(v1, v2):
    return np.dot(v1, v2) / (np.linalg.norm(v1) * np.linalg.norm(v2))

# Scores range from -1.0 (opposite) to 1.0 (identical meaning)

Best Practices

  • Batch Requests: Pass arrays of strings in input: ["text 1", "text 2"] to maximize throughput.
  • Vector Databases: Store generated embeddings in vector indexes like pgvector, Pinecone, Qdrant, or Chroma for sub-millisecond similarity search.
  • Chunk Length: Keep text passages under 512-800 tokens so vectors remain focused on specific concepts.
Was this page helpful?