Most AI systems aren't ready. Check yours in 15 min →
OL

OpenAI Launches Framework to Track and Report AI Safety Incidents

AuthorAndrew
Published on:
Published in:AI

This is the rare kind of AI news that I actually trust a little more because it makes the company look worse, not better. If you’re OpenAI and you choose to publicly talk about your models malfunctioning—especially incidents people hadn’t heard about—you’re either trying to get ahead of a bigger mess, or you’ve decided the “we’re fine, trust us” era is over. Either way, it’s a crack in the usual polished story. And cracks are useful. They let you see what’s real.

Based on what’s been shared publicly, OpenAI unveiled a new framework for tracking AI safety incidents. Alongside it, they disclosed several incidents that were previously unreported where their models malfunctioned. The pitch is simple: build a more systematic way to log these problems and report them going forward. This fits with what’s happening across the industry: more leading AI developers are trying to look more transparent by adopting structured processes for monitoring and communicating safety issues.

On paper, I like it. In practice, I don’t think we should clap yet.

A “framework” is not the same thing as accountability. A framework is a promise about how you’ll describe reality. Accountability is what happens when reality embarrasses you, costs you money, or slows you down—and you still tell the truth.

The fact that there were “previously unreported incidents” is doing a lot of work here. It raises an uncomfortable question: were these incidents unreported because nobody noticed, because they were hard to define, or because reporting them wasn’t in the company’s interest? Those are three very different worlds. In one world, these systems are so complex and fast-moving that problems slip through the cracks. In another, the company can’t even agree internally on what counts as an “incident.” In the third, the incentives are obvious: you don’t volunteer bad news unless you have to.

To be fair, I can also see a more generous interpretation. AI is getting embedded into more products, more workflows, more decisions. When something goes wrong, it’s not always clear what “went wrong” even means. Was it a model error? Bad user instructions? A weird edge case? A human blindly trusting the output? If you don’t have a shared way to classify and log failures, you end up with chaos: support tickets, angry posts, quiet fixes, and no real learning. A framework could force consistency. It could make patterns visible sooner. It could help teams avoid repeating the same mistakes.

But here’s the part that makes me uneasy: self-reporting is not the same thing as transparency. It can become a carefully managed story that stays technically true while still hiding the most important thing—how often it happens, how bad it gets, and what choices led to it.

Imagine you’re running a small company and you use an AI assistant to draft emails to customers. One day it makes up a policy that doesn’t exist. Now your customer thinks you’re lying, your support person is stuck cleaning up the mess, and you look sloppy. Is that an “incident”? Probably. Now imagine you’re a teacher and your school uses an AI tool to help write student feedback. It spits out something inappropriate or unfair, and a parent sees it. That’s not just a glitch. That’s a trust break. Or imagine a developer uses a model to speed up code changes, and it introduces a subtle security issue. Nobody notices for weeks. Is that an AI incident, a code review failure, or both?

These examples are hypothetical, but they point to the real stakes: as AI becomes normal, “malfunction” stops being a funny screenshot and starts being a cost—money, time, reputations, and sometimes real harm. The winners in a world with better incident tracking are the people downstream: users, customers, and teams who have to live with the fallout. The losers are any company that depends on the illusion that their system is mostly safe because the worst stuff is rare or “just misuse.”

And yes, there’s a tension here. If you push companies to be more open about failures, you also give critics more ammunition. You create scary headlines. You might even help competitors. That’s the argument against too much disclosure: it could slow adoption, create panic, or lead to blunt regulation written by people who don’t understand the tech. I don’t dismiss that. But the opposite problem is worse: quiet failures building up until the public only learns about them after someone gets seriously hurt or a major scandal hits.

A framework also risks becoming a box-checking ritual. If the internal culture is “ship fast, apologize later,” then incident tracking can turn into a bureaucratic filter: only log what’s undeniable, only report what’s already leaked, categorize things in ways that make them look smaller. The details matter. What counts as an incident? Who decides? How fast do they report? Do they share near-misses or only confirmed harm? Do they share the messy causes, or just the cleaned-up lesson?

OpenAI disclosing unreported incidents could be a sign that they’re trying to build a healthier habit: treat failures as data, not shame. I want that to be true. But I also think the public has learned, repeatedly, that companies are very good at “being transparent” in ways that still keep control.

If this new push toward incident tracking is real, it will show up in the uncomfortable moments. The first time an incident makes them look reckless. The first time it threatens a product launch. The first time the easiest move would be to keep quiet and quietly patch it.

So here’s where I land: this is a step in the right direction, but it’s also a test of whether AI companies can handle grown-up responsibility without being forced into it. If they can’t, someone else will do it for them, and it won’t be gentle.

What would it take for you to believe that companies reporting AI safety incidents are genuinely opening the curtain, not just managing the story?

Frequently asked questions

What is AI agent governance?

AI agent governance is the set of policies, controls, and monitoring systems that ensure autonomous AI agents behave safely, comply with regulations, and remain auditable. It covers decision logging, policy enforcement, access controls, and incident response for AI systems that act on behalf of a business.

Does the EU AI Act apply to my company?

The EU AI Act applies to any organisation that develops, deploys, or uses AI systems in the EU, regardless of where the company is headquartered. High-risk AI systems face strict obligations starting 2 August 2026, including risk management, data governance, transparency, human oversight, and conformity assessments.

How do I test an AI agent for security vulnerabilities?

AI agent security testing evaluates agents for prompt injection, data exfiltration, policy bypass, jailbreaks, and compliance violations. Talan.tech's Talantir platform runs 500+ automated test scenarios across 11 categories and produces a certified security score with remediation guidance.

Where should I start with AI governance?

Start with a free AI Readiness Assessment to benchmark your current maturity across 10 dimensions (strategy, data, security, compliance, operations, and more). The assessment takes about 15 minutes and produces a prioritised roadmap you can act on immediately.

Ready to secure and govern your AI agents?

Start with a free AI Readiness Assessment to benchmark your maturity across 10 dimensions, or dive into the product that solves your specific problem.