A decision made by an employee leaves a trail almost without anyone planning it. There are emails, an approval under somebody's name and a job title that tells you what they were allowed to sign off. If that decision later leads to a claim, the insurer has a clear trail to follow.
An AI agent works differently. It can read a prompt, query an internal database, call an outside tool, hand a task to another AI agent and then produce a payment or customer message. Reconstructing how the AI agent got there can mean pulling records from several systems that were never designed to tell the same story.
Once an AI-generated answer causes a loss, the argument moves beyond whether the answer was wrong. Two legal cases offered an early warning, even though neither involved the kind of multi-step agent now entering business workflows.
In
An AI agent that gathers data, changes systems and passes work on creates a much longer chain for an insurer to untangle.
A claims trail in pieces
A general explanation of the model will not settle a disputed claim. The insurer needs the version running at the time, the instructions given to the AI agent and the information it pulled in.
Companies usually retain pieces of this. Prompts and outputs sit in the model platform, the security team has access records, and an approval lives in workflow software with the transaction recorded somewhere else again. Each system has done the job it was bought to do, and none has produced a claims file.
By the time a claim is investigated, the AI agent under review can be several versions removed from the one involved in the loss. A new model has gone live, the system prompt has been edited, permissions have expanded and routine retention rules have cleared older logs.
Regulators have started pressing insurers on the same weakness in their own AI use.
Pricing what nobody can see
Everything above concerns what happens after a loss. The underwriter pricing that same deployment before anything goes wrong is just as blind.
The information that would help is being generated constantly and thrown away. An employee rewrites a wrong answer before it reaches a customer. A permission check refuses an action, and correction rates climb after a model update and settle again later. The business fixes each thing and carries on, nobody outside the team ever hears about it, and the insurer prices the risk from a questionnaire and a conversation.
Nearly half of the underwriters in the
A company that found a weakness and fixed it can still expect more questions at renewal, tighter terms or a higher premium for having said so.
Insurance has dealt with this before by paying for the information. Nuclear insurers offer premium differentials of up to 40% on the strength of engineering safety reports. Cyber insurers discount by as much as 25% where a policyholder can evidence its security posture.
An AI agent deployment able to produce its operating record should be worth the same, and that is what would make the record worth keeping.
Where the control gives way
A record only helps if somebody can work through it. A claims file tends to contain the opening prompt and the final answer. I would look first at everything in between. Extracts get chosen by somebody who does not yet know what went wrong. Every method that summarizes or indexes a record before review carries the same flaw.
An AI agent that refuses a request at the first prompt will often agree to it by the tenth, according to the Underwriting the Agent Economy report's own testing.
An insurer will want more than 'the model did it'. Whether the business can show what actually happened was decided the day the AI agent went live.










