Safety & Policies
Cortiqa incorporates real-time safety classifiers to safeguard against the generation of illegal, abusive, or dangerous outputs while preserving high utility for developers.
Classification checks operate synchronously during token prefill and generation. If a prompt or generated completion crosses severity thresholds for critical harm categories, the request is refused gracefully with a standardized code.
| Category | Description | Action |
|---|---|---|
| Malicious Cyber Actions | Exploit generation, automated malware creation, and unauthorized network penetration instructions. | Immediate Block |
| Hate & Harassment | Targeted abuse, derogatory slurs, and promotion of violence against protected groups. | Immediate Block |
| Self-Harm & Violence | Instructional guidance or encouragement of self-harm, suicide, or physical bodily injury. | Immediate Block |
| CBRN Threats | Chemical, biological, radiological, or nuclear weapon synthesis. | Immediate Block & Alert |
When a completion is halted by the safety layer, the response choice sets finish_reason: "content_filter":
response = client.chat.completions.create(
model="openai/gpt-oss-120b",
messages=[{"role": "user", "content": "..."}]
)
choice = response.choices[0]
if choice.finish_reason == "content_filter":
print("Content generation was stopped by the safety filter.")For regulated industries, cybersecurity research firms, and clinical compliance workflows requiring custom sensitivity thresholds, contact our team at team@cortiqa.co.