The Standards Already Said This

OpenAI announced this week that it built a framework for tracking, investigating, and disclosing "misalignment" in its models — cases where a model acted without authorization, coordinated with other models, or evaded oversight — and published six incident reports to kick it off. Any employee can flag a case. Cases get triaged. They disclose before they fully understand the cause. An incident doesn't have to cause harm to count.

The story is everywhere, so I won't rehash the six incidents. What I want to talk about is that none of this is new. It's what the frameworks have been saying for almost three years.

Where it comes from

NIST's AI Risk Management Framework came out in January 2023. ISO/IEC 42001 was published in December 2023. Between them, they already describe what OpenAI just built:

NIST AI RMF GOVERN 4.3 says organizations should have practices in place to enable AI testing, identification of incidents, and information sharing. MANAGE 4.3 says incidents and errors get communicated to the relevant people, with a documented process for tracking, responding, and recovering. MEASURE 3.3 says there should be a way for users and affected people to report problems.

ISO 42001 asks every organization running an AI management system to build the same thing: A.3.3, a process for reporting concerns; A.8.3, external reporting; A.8.4, communication of incidents.

An employee flags a problem. It gets investigated. It gets communicated. The standards spelled it out. OpenAI is now doing it—in public —which goes further than either framework requires. That part deserves credit.

Why this matters for everyone else

I hear a lot of "we're not building models, so that stuff doesn't apply to us." It does. Whether you're a producer of AI or a consumer of it — building on someone's API, rolling out Copilot, running a chatbot on your website — you're going to have little things go sideways. A model that hides a mistake in a summary. An output that's confidently wrong. A coworker using a tool in a way the acceptable use policy never allowed. The question isn't whether those things happen. It's whether anyone in your organization knows how to raise their hand when they see one.

That's the part of a management system that actually does the work day to day. Not the certificate — the plumbing. User training that tells people what a red flag looks like. An acceptable use policy that doesn't just list prohibitions but tells people how to report what they're seeing: in their own daily interactions with AI, in the outputs they get back, or in how someone else is using it outside what the policy and the paperwork they signed allow. A channel that actually goes somewhere. People with enough knowledge to triage what comes in.

Defense in depth — twice

Whether it was the CISSP manual or the AAISM manual, both land on the same understanding: when you build AI systems, they need to sit inside layered defenses. Not just in the sandbox — and some of these models have escaped their sandboxes this year — but everywhere your crown jewels and sensitive information live. "Defense in depth" gets thrown around so much it goes in one ear and out the other. But the idea is still right: no single control is going to hold, so you stack them, and you assume the model will eventually get past one of them.

Here's the part I think people miss. The same principle applies to reporting. One channel isn't enough. Layers of reporting means an employee can raise a flag through their manager, through a security team, through an anonymous channel, through the AUP process — whatever gets it out of their head and into the system. It means management keeps an open line and says out loud that these things will happen, so nobody's surprised when they do. And it means people are allowed to know what to do when they happen: who to tell, what to write down, what not to touch.

Two things kill a reporting layer. Fear — people who think speaking up will land on them. And fatigue — people who spoke up before and watched nothing happen. Neither should be an issue in a functioning program. OpenAI's framework says an incident doesn't have to cause harm to be worth reporting; that's the right bar, and it only works if the people reporting believe they'll be heard.

The takeaway

OpenAI just showed that the little details and the little implementations work — at the scale of a frontier lab. It's refreshing to see one of the biggest producers of AI land on the same answer the frameworks have had all along. If it's worth doing there, it's worth doing in a 200-person company rolling out its first chatbot. Use these frameworks as best you can, and put people who understand them in the room.

Defense in depth for the data. Defense in depth for the people who notice when something's wrong with it. You need both — and the second one is cheaper.

Sources

  • OpenAI — Our framework for reporting model misalignment: openai.com/index/model-misalignment-reporting-framework/

  • NPR/AP — OpenAI flags new concerning AI behavior, to track model misalignment regularly: npr.org/2026/09/17/g-s1-143774/openai-concerning-ai-behavior

  • NIST — AI Risk Management Framework 1.0 (January 2023): nist.gov/itl/ai-risk-management-framework

  • ISO/IEC 42001:2023 — Artificial intelligence management system