Safety & Policies

Content Filtering & Moderation

Cortiqa incorporates real-time safety classifiers to safeguard against the generation of illegal, abusive, or dangerous outputs while preserving high utility for developers.


How Filtering Works

Classification checks operate synchronously during token prefill and generation. If a prompt or generated completion crosses severity thresholds for critical harm categories, the request is refused gracefully with a standardized code.

Harm Categories

CategoryDescriptionAction
Malicious Cyber ActionsExploit generation, automated malware creation, and unauthorized network penetration instructions.Immediate Block
Hate & HarassmentTargeted abuse, derogatory slurs, and promotion of violence against protected groups.Immediate Block
Self-Harm & ViolenceInstructional guidance or encouragement of self-harm, suicide, or physical bodily injury.Immediate Block
CBRN ThreatsChemical, biological, radiological, or nuclear weapon synthesis.Immediate Block & Alert

Detecting & Handling Filter Triggers

When a completion is halted by the safety layer, the response choice sets finish_reason: "content_filter":

filter_check.py
response = client.chat.completions.create(
    model="openai/gpt-oss-120b",
    messages=[{"role": "user", "content": "..."}]
)

choice = response.choices[0]
if choice.finish_reason == "content_filter":
    print("Content generation was stopped by the safety filter.")

Enterprise Customization

For regulated industries, cybersecurity research firms, and clinical compliance workflows requiring custom sensitivity thresholds, contact our team at team@cortiqa.co.

Was this page helpful?