A claim comes in after a policyholder's car is totaled in a multi-vehicle wreck. Within seconds, an AI model scores it as high severity, routes it to a specialist adjuster, and sets the reserve, all before a human reviews the file. Months later, a state market conduct examiner asks the carrier's compliance team to reconstruct what happened: Which model version scored the claim, who was accountable for the decision, and why it made the decision. The room goes quiet. Nothing went wrong. The claim was handled correctly. But nobody can produce that answer on the spot.
This moment has recurred over the past two decades of compliance and audit work, most recently as AI moved from a pilot project into daily
That inability is no longer a hypothetical risk. As of August 6, 2026,
Examiners in those states are asking carriers for bias-testing records, a named executive accountable for
The gap between how quickly carriers have adopted AI and how ready they are to prove it works safely is an organizational problem: Staffing, ownership, and, most critically, whether governance was built into the AI system from day one or added after the fact. The frameworks already exist and are mature, including the NIST AI Risk Management Framework, ISO 42001, and the NAIC's own bulletin structure. Carriers that can already produce evidence, rather than intent, when an examiner asks to see it are the ones positioned to handle the next 12 to 18 months of enforcement. The rest have real work to do, and not much time to do it.
Where the gap widens first
Oversight lags hardest at the exact moment a system moves from flagging something for a human to reviewing it, to deciding it outright. Claims severity scoring, FNOL routing, and underwriting pre-screening were funded and deployed first precisely because they are high-volume and repetitive, which is also what allows the exposure to compound quickly once governance is not in place.
This pattern recurs across compliance functions: A model gets deployed to speed up throughput, and governance only enters the conversation later, usually when exam preparation forces someone to reconstruct what the model actually did. The risk is adopting without a plan for the moment someone asks why the model decided what it decided.
Audit trails are the first thing examiners ask for
Regulators are now probing four components of AI governance: Audit trails, model validation, human-in-the-loop review, and vendor accountability. Audit trails are the least mature of the four. Most carriers can say what a model decided. Far fewer can say why, which version made the call, or who had the authority to override it.
The NAIC's 12-state pilot is built around exactly this gap, pulling bias-testing records and a named accountable owner in live exams rather than hypothetical ones. What most carriers can produce today is a policy document describing intent, not evidence that the policy was followed in practice. That distinction, between having a governance policy and proving it was followed, is exactly where exams are starting to bite.
Build the inventory before anything else
The most common mistake is treating AI governance as a documentation exercise that happens after deployment, rather than a requirement built into the system from the start. Retrofitting audit trails and explainability onto a production system already in use is harder and reads far less credibly to an examiner than designing for it from the outset.
The single concrete step compliance leaders should take now is building a real model inventory: Every AI or machine learning model touching underwriting or claims, who owns each one, and whether a documented audit trail exists for it. Most carriers cannot produce this in under a week. Everything that follows, bias testing, human-in-the-loop design, exam readiness, depends on having that inventory first.
Adoption pressure is not slowing down, and regulatory tooling is still catching up to it. That gap will likely widen before it narrows over the next 12 to 18 months. The carriers that come out ahead will be the ones who can already answer "show me" with evidence instead of scrambling to reconstruct it after the fact. If a regulator asked for your model inventory today, could you produce it, or would it take a week?










