InsureThink

Why neural networks and embeddings improve insurance models

Visualization created with AI assistance.

Every day, people interact with systems that recognize patterns without relying on explicit rules. Streaming platforms group related shows, online retailers recommend products that align with customer behavior and photo libraries retrieve images based on context rather than tags. 

Processing Content

These systems are not driven by lists of handcrafted rules. Instead, they rely on learned representations, called embeddings, that capture patterns, context and relationships across large volumes of data. 

The insurance industry is beginning to apply similar techniques. As data sources expand and risks become more dynamic, embeddings are influencing how insurers assess risks, segment portfolios and support underwriting and pricing decisions. These approaches enable insurers to rely less on fixed categories and more on continuous, adaptive understanding of risks.

In my previous article, How insurance data is evolving from raw reports to embeddings, we traced the evolution of insurance data practices: from underwriters using raw data reports, to engineered attributes and predictive models. We also highlighted the limits of traditional approaches as data volumes grow and relationships become harder to identify. Finally, we introduced neural networks and embeddings as a more flexible foundation for modern risk modeling. This article builds on that foundation by taking a closer look at embeddings and why they matter now, how they are learned and what they enable for modern insurance risk modeling.

Traditional insurance models are built on attributes explicitly defined by humans: counts, flags, averages, ratios and categorical indicators derived from raw data. These variables can be highly predictive, but they reflect assumptions about which signals matter and how they should be combined. As a result, model performance often depends on how thoroughly those assumptions anticipate real-world risk behavior. 

As insurance data expands to include time‑based behavior, geospatial patterns, imagery, text and event sequences, those assumptions become harder to maintain. Feature sets grow larger and more fragile, segmentation schemes multiply and models require frequent re‑engineering simply to stay current. Over time, this creates maintenance challenges and limits how quickly insurers can adapt to changing risk conditions. 

Neural networks approach this challenge differently. Rather than relying on predefined combinations of inputs, they learn patterns directly from data. During training, the model identifies which signals tend to occur together, how those interactions relate to outcomes and which relationships are most informative for prediction.

This ability makes neural networks well suited to insurance risk scenarios with nonlinear relationships, interaction across data sources and dependencies that unfold over time. Just as importantly, neural networks can update what they learn as data and behavior change. When risk patterns shift due to economic cycles, climate trends or evolving customer behavior, models can adjust without requiring a full redesign of the attribute structure.

A neural network's primary advantage lies in its capacity to learn embeddings. An embedding is a dense numeric representation of an entity, event or behavior that summarizes complex information in a compact form. Instead of describing a property, customer or location through hundreds of loosely related attributes, an embedding captures the characteristics that are most relevant to modeling risk in a single numerical vector.

The vector reflects how the observation compares to other observations in the data set. For example, properties exposed to similar environmental conditions or customers exhibiting comparable behavioral patterns naturally cluster together in the embedding space when their surface‑level attributes differ. This allows models to assess risk on learned relationships rather than rigid, predefined categories.

Embeddings are not created manually. Neural networks learn them as part of the training process. As a neural network attempts to predict outcomes such as loss likelihood or claim severity, it continuously adjusts its internal layers to better reflect the structure of the data. These layers serve as representation learners, transforming raw inputs into progressively more informative abstractions.

Embeddings can be learned in different ways, depending on the modeling objective and each with practical implications for insurance:

  • Supervised learning optimizes embeddings against known outcomes such as claims or losses, making them highly predictive for specific objectives.
  • Self‑supervised learning allows models to learn structure by predicting missing, future or related information within the data.
  • Multitask learning trains a shared embedding across multiple objectives, such as underwriting decisions, pricing accuracy and fraud detection, resulting in richer and more reusable representations.

By capturing underlying patterns rather than surface‑level attributes, embedding-based models can support more accurate risk assessment than traditional models alone. 

High‑dimensional, heterogeneous data is difficult to manage and costly to maintain. Embeddings help reduce that complexity by concentrating information into compact representations, which can reduce noise, improve model stability and simplify downstream modeling.

Once learned, embeddings are not limited to a single model or use case. They can serve as shared inputs across underwriting, pricing, segmentation, forecasting and portfolio management. Rather than rebuilding feature sets for each new modeling problem, insurers can reuse embeddings as a common analytical foundation, supporting greater consistency and scalability across teams and decisions. 

Property risk offers a clear example of where this approach is valuable. Risk emerges from the interaction of environment, infrastructure, exposure and human activity, often in ways that are difficult to capture with predefined categories. Two properties may look different on paper yet behave similarly under certain weather patterns or loss scenarios.

Neural networks and embeddings are well suited to capturing these relationships. By learning directly from geospatial data, weather history, property characteristics and prior losses, embeddings provide a more faithful representation of how property risk accumulates and manifests over time.

The next article in the series will focus on applying embeddings to property risk modeling, exploring how property-level embeddings can support sharper underwriting decisions, more precise risk segmentation and pricing that better reflects conditions on the ground.

For insurers seeking to move beyond rigid assumptions toward models that better reflect how risk behaves in practice, embeddings represent an important step in the evolution of insurance risk modeling. 


For reprint and licensing requests for this article, click here.
Artificial Intelligence Data modeling
MORE FROM DIGITAL INSURANCE
Load More