- AI drift causes claims and underwriting answers to vary
- Lack of direct connection to information worsens AI drift
- AI drift tests can be conducted
An AI revenue boom could cause a lack of accountability for AI model drift, according to projections by Gartner.
The
Gartner senior director analyst Mario Capellari defined AI behavioral drift as AI outputs diverging from the behavior originally tested, validated and expected.
To stop AI drift from
For insurance underwriting and claims workflows using agentic AI, drift happens when the same queries end up yielding different answers, said Grace Apea, AVP, product marketing and market insights at Equisoft, an insurtech services provider specializing in life insurance functions.
"What if every single time you ask the agent to do something, or what if every time you're expecting it to do something, it changes? Even if the change is very small, a fraction of a percent of a change, what is that impact going to be?"
AI drift can also be caused by AI operating based on different sources or third parties, according to James Hannay, chief revenue officer at Sapiens, an AI software platform serving insurers. Hannay described AI drift using the human nervous system as an illustration.
"If the brain says to pick up the right hand, make a circle with it, and put it back down again, it will do it the same every single time," Hannay said. "If the hand wasn't connected to the brain, every time it was asked to carry out that action, you would get a variation. One minute, it's a square. Next minute, it's a circle. Or it might be an oblong or a rectangle." AI needs to be connected to the same source of instructions and information to avoid results drifting, he added.
Insurers studying how to use AI agents are not connecting them directly to sources of knowledge, Hannay said. This contributes to the AI drift that Capellari and the Gartner report describe, he added. Using A/B testing of different sets of data in an AI platform can determine if it is navigating directly to data sources or needs information patched in, Hannay said.
Similarly, a "golden set," a group of underwriting test cases already reviewed by a person, then run through a large language model or AI, can show if the AI is drifting, said Apea of Equisoft. "A good golden set has to cover edge cases, not just the usual run of the mill underwriting cases," she said. "You want to make sure that it's really measuring the drift correctly. You need to make sure that it's covering different types of risks, different types of demographics, different product types as well. Most importantly, Most importantly, review versions of output to know if there is a change in the data inputs and be able to track that."
For insurance claims, Apea said, a rule of thumb is if a particular claims scenario typically leads to a 98% approval rate, and the rejection rate climbs to 5%, the use of AI should be reviewed.









