EconBase
← Back to paper

From Rating Factors to Crash Mechanisms: A Multiscale Causal DAG Framework Linking Motor Insurance and Road Safety

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

81,664 characters

From Rating Factors to Crash Mechanisms: A Multiscale Causal DAG Framework Linking Motor Insurance and Road Safety



\title{From Rating Factors to Crash Mechanisms: A Multiscale Causal DAG Framework Linking Motor Insurance and Road Safety}
\author{Arthur Charpentier\\
\small Université du Québec à Montréal, Montréal, Canada\\
\small Kyoto University, Kyoto, Japan}
\date{}

\maketitle

\begin{abstract}
Road safety mechanisms operate within seconds, minutes and trips, whereas motor insurance observes liability claims aggregated over policy years. An annual rating coefficient can therefore predict claims accurately while leaving the crash-generating process unresolved. We propose a multiscale causal DAG framework with three parts: a proposed crash-occurrence graph constructed from a structured, non-exhaustive map of 72 study--edge records; a separate observation layer linking conventional rating variables to latent exposure, context and behaviour; and a downstream crash-to-claim process that includes reporting, responsibility attribution and claim administration. The formal contribution is set-valued: it characterizes which annual mechanism laws and claim-observation mappings are compatible with an observed insurance contrast and retained external evidence, rather than estimating a causal effect of a rating factor. Diagnostic examples show the limits of that interpretation. A sublinear mileage relation constrains aggregate exposure without identifying its composition. In the French \texttt{freMTPL2freq} portfolio, the 18--20 versus 40--49 claim-frequency relativity is 3.388 after vehicle/geographic adjustment and 1.235 after conditioning on medium-resolution bonus--malus categories; the latter is a different conditional predictive contrast because bonus--malus summarizes endogenous prior insurance history. A Spanish age-mediation estimate narrows only one coarse bookkeeping block under explicit transport-sensitivity assumptions, and the resulting region remains wide. The practical implication is a data requirement: stronger mechanistic claims need trip-level intermediate states and linked crash--claim observations.
\end{abstract}

\noindent\textbf{Keywords:} motor insurance, road safety, causal inference, directed acyclic graph, telematics, risk classification, evidence synthesis, partial identification
\medskip


\section{Introduction}
\label{sec:intro}

Accurate prediction of annual claims does not identify the mechanisms that generate crashes. Motor insurance observes policy characteristics and claims over months or years; road safety research studies road, traffic, vehicle and behavioural states that change within trips. A mechanistic interpretation of an annual claim contrast therefore requires two bridges: one across time scales and another from crash occurrence to the recorded insurance outcome.

Let $N_i$ denote the annual number of recorded motor-liability claims for policyholder $i$ and let $X_i$ contain relatively stable rating information such as age, vehicle characteristics, territory, declared use, mileage and prior insurance history. The reduced-form insurance target is
\begin{equation}
  \lambda_{\mathrm{claim}}(X_i)=\mathbb{E}(N_i\mid X_i).
  \label{eq:glm}
\end{equation}
A Poisson GLM is one possible estimator of this mean, but the estimand itself is annual predictive claim frequency. It does not describe the sequence of circumstances and actions preceding a collision. Nor is a recorded claim identical to a crash: reporting, coverage, attribution of responsibility and claim administration intervene after crash occurrence. Section~\ref{sec:claimobservation} makes that downstream observation process explicit.

Road safety studies provide evidence on the shorter-horizon process. Speed, traffic variation, adverse weather, sleepiness, distraction and passenger configuration are analysed with naturalistic, quasi-experimental, cohort and synthesis designs \cite{elvik2019speed,rosandel2015traffic,qiu2008weather,moradi2019sleepiness,simmons2016distraction,ouimet2015passengers}. Causal diagrams have also been used in transportation safety \cite{karwa2011causal,davis2011conflicts,dufournet2016selection}, while ESC-DAG offers a structured method for building graphs from heterogeneous evidence \cite{ferguson2020escdag}. The contribution here goes beyond the familiar distinction between prediction and causation \cite{greenland1999causal,pearl2009causality}: it specifies the assumptions needed to connect annual insurance information to mechanisms defined within driving opportunities.

The paper makes three connected contributions. First, it proposes a time-ordered crash-occurrence DAG whose retained arrows are traceable to a structured evidence map of 72 study--edge records. Second, it separates that structural graph from an actuarial observation layer and from the downstream process that turns a crash into a recorded liability claim. Third, it formulates mechanistic interpretation as a compatibility problem: which latent annual mechanism laws and claim-observation mappings can reproduce an observed insurance contrast while satisfying the retained graph and external road safety restrictions? The mileage and age calculations are diagnostic examples of this third contribution, not validation of the full DAG.

Let $H_i$ denote persistent driver heterogeneity, $R_i$ accumulated driving experience, $C_{it}$ the context of driving opportunity $t$, $S_{it}$ a transient driver state, $B_{it}$ immediate behaviour, $K_{it}$ a critical conflict/avoidance state, and $Y_{it}$ the indicator of a crash during that opportunity. Write $V_i=(V_i^{\mathrm{perf}},V_i^{\mathrm{safe}})$ for vehicle-performance and safety/avoidance mechanisms, and collect the opportunity-level states in
\[
 \Omega_{it}=(H_i,R_i,C_{it},S_{it},B_{it},K_{it},V_i^{\mathrm{safe}}),
 \qquad
 p_{it}=\Pr(Y_{it}=1\mid\Omega_{it},X_i).
\]
The core graph $G_0$ imposes the local Markov restriction $Y_{it}\perp(\Omega_{it}\setminus K_{it},X_i)\mid K_{it}$, so that $p_{it}=\pi_Y(K_{it})$ under the maintained graph. A driving opportunity is taken to be a trip or a pre-specified trip segment over which the short-term variables can be represented by one ordered snapshot. If $M_i$ is the number of such opportunities in the policy period and $D_{it}$ their travelled distances, observed annual mileage may be written $m_i=\sum_{t=1}^{M_i}D_{it}$ when the segment distances are available. Thus $M_i$ is not mileage and a mile is not treated as a crash opportunity. Annual crash-occurrence frequency is
\begin{equation}
\lambda_{\mathrm{crash}}(X_i)
=
\mathbb{E}\!\left[
\sum_{t=1}^{M_i}p_{it}
\,\middle|\,X_i
\right].
\label{eq:central}
\end{equation}
The annual latent history $U_i$ is defined in Section~\ref{sec:integrated}; $Q_x=\mathcal L(M_i,U_i\mid X_i=x)$ denotes its full conditional law. Equation~\eqref{eq:central} is a random-sum expectation under $Q_x$ and does not require independent or exchangeable driving opportunities.

The empirical examples are diagnostic. The mileage analysis isolates what an aggregate exposure relation can and cannot imply. The age analysis first shows how conditioning on an endogenous history variable changes the insurance contrast, then uses one external road safety estimate to exclude part of a coarse compatibility set. Neither example estimates a causal effect of age or mileage or validates the full DAG; both identify information that annual claims do not contain.

\section{Two time scales of motor risk}
\label{sec:timescales}

\subsection{Annual actuarial risk as a reduced-form target}

Claim-frequency models aggregate heterogeneous driving over a policy year. Separating mileage $m_i$ from the other rating variables, write $X_i=(m_i,Z_i)$. A specification such as
\begin{equation}
  \log \lambda_{\mathrm{claim},i} = \log m_i + \beta^\top Z_i
\end{equation}
treats mileage as exposure and assigns the remaining risk per unit distance to variables measured at policy inception or updated infrequently. That representation is useful for pricing and classification, but it leaves the composition of exposure implicit.

Telematics records time of day, speed, braking, acceleration and route-related signals that conventional rating variables leave latent \cite{verbelen2018telematics,baecke2017telematics,henckaerts2022dynamic}. Hidden-state models likewise summarize trip-level behaviour and relate those states to insurance losses \cite{jiang2024hmm}. This evidence supports treating $X_i$ as information about a distribution of driving states, not as a list of direct crash causes.

\subsection{Short-term crash generation}

Evidence closer to collision time covers speed \cite{kloeden1997speed,elvik2019speed}, traffic state \cite{golob2004traffic,xu2012trafficstate,rosandel2015traffic}, weather \cite{qiu2008weather}, sleepiness \cite{martiniuk2013sleep,moradi2019sleepiness}, distraction \cite{simmons2016distraction}, young-driver passenger exposure and behaviour \cite{simonsmorton2005passengers,goodwin2012passengers,ouimet2015passengers}, and novice nighttime or passenger restrictions \cite{fell2011gdl}. Conflict models provide an intermediate scale between ordinary driving and rare crashes \cite{davis2011conflicts}. These studies define the mechanisms represented in the short-term graph; annual rating labels remain in the observation layer.

\begin{figure}[t]
\centering
\begin{tikzpicture}[x=1.2cm,y=1cm,>=Latex]
  \draw[-{Latex[length=2mm]},thick] (0,0) -- (10.5,0);
  \foreach \x/\lab in {0/second,2/minute,4/trip,7/month,10/year}{
    \draw (\x,0.12)--(\x,-0.12);
    \node[below=5pt,font=\small] at (\x,0) {\lab};
  }
  \node[align=center,above=10pt,font=\small] at (1,0) {reaction\\speed\\conflict};
  \node[align=center,above=10pt,font=\small] at (3.2,0) {traffic\\fatigue\\distraction};
  \node[align=center,above=10pt,font=\small] at (5.2,0) {road\\weather\\passengers};
  \node[align=center,above=10pt,font=\small] at (8.7,0) {rating factors\\claim count};
  \draw[decorate,decoration={brace,amplitude=5pt},yshift=34pt] (0,.75)--(6,.75)
     node[midway,yshift=13pt,font=\small]{road safety research};
  \draw[decorate,decoration={brace,amplitude=5pt},yshift=34pt] (7,0.25)--(10.5,0.25)
     node[midway,yshift=13pt,font=\small]{insurance};
\end{tikzpicture}
\caption{The time-scale mismatch. Road safety mechanisms are predominantly studied over seconds, minutes and trips, whereas traditional motor-insurance risk classification aggregates experience over months or years.}
\label{fig:timescale}
\end{figure}

\subsection{From crashes to recorded liability claims}
\label{sec:claimobservation}

The insurance outcome adds an observation process after crash occurrence. Let $J_{it}\in\mathbb N_0$ denote the number of recorded liability-claim contributions attributed, under the portfolio's accounting convention, to opportunity $t$, with $J_{it}=0$ when $Y_{it}=0$. The count-valued definition permits zero, one or more recorded claim contributions from an underlying crash event. Define
\begin{equation}
 \rho_{it}
 =\mathbb E(J_{it}\mid Y_{it}=1,\Omega_{it},X_i).
 \label{eq:rho}
\end{equation}
Thus $\rho_{it}$ is a conditional mean claim contribution per crash, not a probability unless the accounting rule makes $J_{it}$ binary. It integrates over post-crash features not represented in the occurrence DAG, including severity, third-party involvement, coverage, reporting, responsibility attribution and claim administration. With $N_i=\sum_{t=1}^{M_i}J_{it}$ under the portfolio's accounting convention, iterated expectation gives
\begin{equation}
 \lambda_{\mathrm{claim}}(X_i)
 =\mathbb E\!\left[
   \sum_{t=1}^{M_i}p_{it}\rho_{it}
   \,\middle|\,X_i
 \right].
 \label{eq:claimobservation}
\end{equation}
Equality of $\lambda_{\mathrm{claim}}(X_i)$ and $\lambda_{\mathrm{crash}}(X_i)$ requires the additional restriction $\rho_{it}=1$ almost surely, which we do not impose. If reporting, coverage or responsibility assignment varies with driver, context or crash characteristics, a claim-frequency relativity need not match the corresponding crash-frequency relativity. External road safety evidence therefore reaches an insurance contrast through two mappings: from the study estimand to the target crash mechanism, and from crash occurrence to the recorded claim. The second mapping belongs to the observation process, not to the crash-mechanism DAG.

\begin{table}[H]
\centering
\small
\begin{tabularx}{\textwidth}{@{}p{2.0cm}p{2.5cm}X@{}}
\toprule
Symbol & Level & Role \\
\midrule
$X_i$ & policy / annual & observed rating information used by the insurer \\
$m_i$ & policy / annual & observed annual mileage \\
$M_i$ & policy / annual & number of trips or pre-specified trip segments in the policy period \\
$W_i$ & policy / slowly varying & latent experience, heterogeneity, activity, mobility, exposure and vehicle states \\
$\Omega_{it}$ & opportunity & state vector relevant to crash occurrence at opportunity $t$ \\
$U_i$ & annual latent history & $W_i$ plus the random-length sequence of opportunity states up to $M_i$ \\
$Y_{it}$ & opportunity & crash-occurrence indicator \\
$J_{it}$ & opportunity / observation & recorded liability-claim contribution generated after a crash \\
$\rho_{it}$ & opportunity / observation & conditional mean of $J_{it}$ given a crash and observed/latent state \\
$Q_x$ & annual law & conditional law $\mathcal L(M_i,U_i\mid X_i=x)$ \\
\bottomrule
\end{tabularx}
\caption{Core notation and the scale at which each object is defined.}
\label{tab:notation}
\end{table}


\section{Constructing the crash-mechanism DAG}
\label{sec:escdag}

\subsection{Target outcome and graph scale}

The structural target is \emph{crash occurrence} during a driving opportunity, not culpability or severity conditional on a crash. Crash occurrence, responsibility attribution, injury severity and recorded insurance claims are distinct outcomes. Responsibility is deliberately handled downstream in the crash-to-claim observation process rather than folded into $Y_{it}$. Conditioning on severe crashes can create selection or collider bias when a determinant of the exposure also affects entry into the analysed sample \cite{dufournet2016selection}. Studies of culpability, injury or fatal crashes can still inform the graph, but their estimands are not substituted for crash occurrence without an explicit mapping.

The graph is a time-ordered snapshot of one driving opportunity. Fatigue, attention, speed and distraction can change within a trip and can affect one another at later instants; those lagged relations are compressed into an acyclic ordering at the chosen scale. Interventions with within-trip or cross-trip feedback require a dynamic DAG or structural time-series model.

\subsection{Evidence-map protocol and search boundary}

We adapt the ESC-DAG logic of Ferguson et al.~\cite{ferguson2020escdag}. The evidence database contains 72 estimand-bearing study--edge records. A publication can contribute more than one record when it reports different exposures, outcomes or designs. The map is used for graph construction and criticism, not as a PRISMA review of one common estimand and not as a graph-wide meta-analysis. This distinction matters because a naturalistic near-crash odds ratio, a culpability estimate and a prospective crash rate can support the same broad mechanism while remaining numerically non-commensurate.

The targeted search was last updated on 8 August 2026 and covered four families: proximal crash mechanisms; driver experience and exposure; the actuarial bridge; and activity or mobility environment. Search terms combined mechanism terms with crash occurrence, responsibility, conflict/near-crash and safety-critical outcomes, together with design terms for cohort, case-control, case-crossover, naturalistic, quasi-experimental and synthesis studies. Severity studies were screened separately so that the larger injury literature did not determine an occurrence graph by volume alone. The search was iterative and was not prospectively registered as a systematic-review protocol; accordingly, we do not claim a complete PRISMA search history or exhaustive coverage of all databases.

The search is structured but not exhaustive, so absence from the map is not evidence that a mechanism is absent. Its evidential role is local: every retained solid arrow has a traceable rationale, while uncertain and proxy-only relations remain visible without entering the structural core. The resulting $G_0$ is therefore a proposed evidence-informed core rather than a validated or definitive road-safety DAG. Appendix~\ref{app:extraction} gives the extraction schema, and Appendix~\ref{app:evidencetables} reports the detailed study-level tables.

\subsection{Adjudication and quantitative eligibility}
\label{sec:adjudication}

Each candidate relation is documented through the same sequence of questions. First, the source and target are checked against the paper's target estimand and temporal scale. Second, temporal ordering and a plausible mechanism are required before an arrow can enter the structural core. Third, design-specific threats---confounding, self-report, outcome-dependent sampling, severity selection and measurement error---are recorded. Fourth, evidence across designs is compared for directional consistency and estimand compatibility. Only then is the relation assigned a graph status and a separate quantitative-eligibility class. This sequence makes the decision rule auditable, but it does not turn judgement into an automated procedure.

All extraction and adjudication were conducted by one author, so no inter-rater agreement statistic is available. This is a material limitation because graph construction itself is part of the contribution. The fixed adjudication sequence reduces undocumented discretion but cannot substitute for independent coding. For that reason, $G_0$ is treated throughout as a proposed evidence-informed structure whose arrows can be re-adjudicated or replaced, not as a graph established by reviewer agreement. The study-level tables report the study, outcome, rationale and status behind each displayed relation so that those choices remain inspectable.

Causal admissibility and numerical reusability are assessed separately. A study may support the direction or existence of a mechanism while remaining unsuitable for calibration because its outcome, population or time scale differs from the target.

\begin{table}[H]
\centering
\small
\begin{tabularx}{\textwidth}{@{}p{1.5cm}X@{}}
\toprule
Class & Interpretation \\
\midrule
\textbf{Q1} & A numerical effect can plausibly anchor an edge-level constraint or calibration, subject to transportability sensitivity. \\
\textbf{Q2} & Numerically informative, but the estimand is a policy effect, surrogate outcome, population-specific contrast or otherwise requires a mapping before entering the DAG. \\
\textbf{Q3} & Structural evidence only: useful for retaining or orienting an edge, but not for direct parameterization. \\
\textbf{P} & Predictive/measurement evidence for the actuarial observation layer, not a causal edge parameter. \\
\bottomrule
\end{tabularx}
\caption{Quantitative-eligibility classes used in the study-level extraction.}
\label{tab:qclasses}
\end{table}


Each candidate relation also receives one graph status. The label ``causal'' is reserved for a relation retained as a solid arrow under the maintained assumptions; it is not a claim that every source study identifies an unconfounded intervention effect.

\begin{table}[H]
\centering
\small
\begin{tabularx}{\textwidth}{@{}p{2.3cm}X@{}}
\toprule
Status & Operational rule \\
\midrule
\textbf{causal} & Retained as a solid arrow when temporal ordering, mechanism plausibility and the combined evidence support a structural relation for the target process. This status is conditional on the stated causal assumptions; it is not a generic label for every association reported by the source studies. \\
\textbf{proxy-only} & Retained only in the actuarial observation layer: the source predicts or measures the target but is not interpreted as causing it. \\
\textbf{uncertain} & Mechanistically plausible, but current evidence is too heterogeneous, selected, self-reported or confounded to justify a solid arrow. \\
\textbf{excluded} & Not drawn in the integrated graph because it is a reduced-form shortcut, reverses the measurement relation, conditions on a problematic downstream variable, or lacks adequate support. \\
\bottomrule
\end{tabularx}
\caption{Edge-adjudication categories used for the integrated DAG.}
\label{tab:adjudication}
\end{table}



Risk of bias, edge status and quantitative eligibility are recorded separately. For prospective non-randomized exposure studies, the domain vocabulary is informed by ROBINS-E \cite{higgins2024robinse}, but we do not report a summary ROBINS-E score; design-specific concerns are used directly in adjudication. Other designs are assessed with the concerns appropriate to their sampling and outcome definitions rather than forced through one instrument. Culpability studies illustrate why these dimensions must remain separate: they are close to the target mechanism but condition on crash involvement, whereas a prospective cohort avoids that selection while often measuring a persistent exposure over a longer horizon than the transient state represented in the DAG.

\subsection{What the evidence map contributes}

The 72 records have two different uses. Most determine graph structure or the observation layer; only a subset are eligible for numerical calibration. The full structural-core, actuarial-bridge, primary-study and excluded-shortcut tables are reported in \ref{app:evidencetables}. Moving them out of the main text does not change their role in adjudication.

The structural core is supported most directly for short-horizon speed, traffic, weather, fatigue, distraction, activity/exposure and collision-avoidance mechanisms. Passenger, errors/violations, vehicle-performance and several broad vehicle relations remain uncertain because their estimates mix selection, severity, self-report or heterogeneous outcomes. The observation layer is supported by exposure and telematics studies showing that age, sex, territory, declared use, mileage and vehicle information carry information about route, timing, speed and behavioural distributions without thereby becoming proximal causes.

\subsection{Study-level quantitative anchors}
\label{sec:quantanchors}

Table~\ref{tab:quantanchors} collects numerical anchors that can constrain local parts of the model. They are heterogeneous by construction: a pooled odds ratio for drowsy driving, a mileage elasticity and a collision-avoidance risk ratio refer to different estimands. They constrain different components of the multiscale model and are not coefficients of a common regression.

\begin{table}[H]
\centering
\scriptsize
\setlength{\tabcolsep}{2.4pt}
\begin{tabularx}{\textwidth}{@{}p{1.65cm}p{0.75cm}p{2.10cm}p{2.20cm}X@{}}
\toprule
Evidence & Class & Contrast & Numerical anchor & Use in the model \\
\midrule
Elvik \cite{elvik2023mileage} & Q1 & annual distance $\to$ annual accidents & accident count approximately $\propto$ distance$^{1/2}$ & Benchmark for the exposure-intensity aggregation; not a per-km causal coefficient. \\
Connor et al. \cite{connor2002sleepiness} & Q1 & sleepy vs alert & OR $8.2$ (95\% CI $3.4$--$19.7$) & Acute-state anchor; transported cautiously from injury crashes to the crash-occurrence branch. \\
Connor et al. \cite{connor2002sleepiness} & Q1 & $\leq5$ h sleep vs more & OR $2.7$ (95\% CI $1.4$--$5.4$) & Supports the upstream sleep-deficit $\to$ fatigue pathway. \\
Moradi et al. \cite{moradi2019sleepiness} & Q1 & drowsy vs non-drowsy driving & random-effects OR $1.34$ (95\% CI $1.25$--$1.43$) & Broader pooled constraint; heterogeneity is retained rather than collapsing it with the acute estimate. \\
Fell et al. \cite{fell2011gdl} & Q2 & nighttime restriction & about 10\% reduction in fatal nighttime crash involvement & Policy constraint on nighttime exposure for novice drivers. \\
Fell et al. \cite{fell2011gdl} & Q2 & teen-passenger restriction & about 9\% reduction in fatal crashes with teen passengers & Policy constraint; not a universal passenger coefficient. \\
Gomes-Franco et al. \cite{gomesfranco2020age} & Q2 & age 18--24 vs 35--44, environment- and vehicle-mediated path & indirect OR $1.09$ (95\% CI $1.08$--$1.10$) & Worked transport-sensitive anchor in Section~\ref{sec:workedintersection}; Spanish culpability outcome and different age bands require explicit discrepancy. \\
Qiu and Nixon \cite{qiu2008weather} & Q1 & snow vs reference weather & crash-rate ratio about $1.84$; injury-rate ratio about $1.75$ & Context constraint with weather/road-type heterogeneity. \\
Cicchino \cite{cicchino2017aeb} & Q1 & FCW+AEB vs same models without system & rear-end striking crash-rate ratio about $0.50$ & Direct technology parameter for the relevant rear-end branch only. \\
Cicchino \cite{cicchino2023aeb} & Q1 & AEB-equipped vs non-equipped pickups & rear-end striking risk ratio about $0.66$ & Replication/sensitivity anchor for vehicle-class transportability. \\
Winlaw et al. \cite{winlaw2019telematics} & Q2 & 75th vs 50th percentile speeding penalty & $>50\%$ greater estimated crash chance & Validation of the speed-behaviour branch; small crash count prevents treating it as a universal coefficient. \\
\bottomrule
\end{tabularx}
\caption{Numerical anchors already available for an evidence-calibrated DAG. The contrasts live on different estimand scales and are mapped to different parts of the multiscale model rather than pooled indiscriminately.}
\label{tab:quantanchors}
\end{table}



Table~\ref{tab:quantanchors} is broader than the worked examples because most estimates inform local graph structure without being transportable to the age decomposition. The small numerical subset reflects estimand mismatch, not lack of relevance. \ref{app:evidencetables} retains the heterogeneous primary estimates on their native scales instead of forcing them into a common coefficient.

\section{An integrated multiscale model}
\label{sec:integrated}
\label{sec:model}

\subsection{Persistent, contextual and transient components}

Five slowly varying objects separate signals that conventional rating variables tend to compress. $R_i$ denotes accumulated driving experience and $H_i$ residual persistent driver heterogeneity. $A_i$ is a relatively stable activity/use profile, covering commuting, work, leisure and recurring timing or route patterns. $G_i$ represents the mobility environment associated with residence and habitual activity space, including road network, urban/rural structure and recurrent traffic or environmental conditions. $E_i$ is latent exposure intensity. Annual mileage $m_i$ measures one aspect of $E_i$ but is not identified with the number of trips or driving opportunities $M_i$. Let
\[
 W_i=(R_i,H_i,A_i,G_i,E_i,V_i^{\mathrm{perf}},V_i^{\mathrm{safe}})
\]
collect these slowly varying states.

The actuarial observation layer specifies their conditional predictive law,
\begin{equation}
  W_i\mid X_i\sim q_W(\cdot\mid X_i).
  \label{eq:measurementlayer}
\end{equation}
The kernel $q_W(\cdot\mid x)$ is the marginal law of the slowly varying coordinates of the annual latent history $U_i$. Together with the opportunity-count and within-opportunity kernels below, it induces the full law $Q_x=\mathcal L(M_i,U_i\mid X_i=x)$. It is a predictive measurement model, not a causal factorization. Age and licence tenure inform $R_i$; claims history updates information about $H_i$; declared use informs $A_i$; territory informs $G_i$; mileage informs $E_i$; and vehicle rating information informs the physical vehicle mechanisms. Because persistent heterogeneity contributes to earlier crashes and claims, the dashed history $\to H_i$ link points in the direction of statistical information, not data generation. The other dashed links in Figure~\ref{fig:dag} have the same interpretation.

Conditional on the slowly varying objects, let $G_0$ denote the \emph{core} graph formed by the solid arrows in Figure~\ref{fig:dag}. At the block level its factorization is
\begin{align}
M_i\mid(E_i,A_i,G_i) &\sim \kappa_M(\cdot\mid E_i,A_i,G_i), \\
C_{it} &\sim \kappa_C(c\mid A_i,G_i), \\
S_{it} &\sim \kappa_S(s\mid R_i,C_{it}), \\
B_{it} &\sim \kappa_B(b\mid R_i,H_i,C_{it},S_{it}), \\
K_{it} &\sim \kappa_K(k\mid C_{it},B_{it},V_i^{\mathrm{safe}}), \\
Y_{it} &\sim \mathrm{Bernoulli}\{\pi_Y(K_{it})\}.
\label{eq:factorization}
\end{align}
Equation~\eqref{eq:factorization} is the formal block-level graph specification. Together with the local Markov property it implies, for example, $C_{it}\perp(R_i,H_i,E_i,V_i)\mid(A_i,G_i)$ and $S_{it}\perp(H_i,A_i,G_i,E_i,V_i)\mid(R_i,C_{it})$ at the displayed block resolution, as well as the $Y_{it}$ restriction stated in Section~\ref{sec:intro}. Appendix~\ref{app:graphspec} lists the parent sets and the principal omitted-parent sensitivities used in the analysis; no independence across opportunities is imposed.
For notational clarity, imagine an infinite sequence of potential opportunity-level states and observe only the first $M_i$ during the policy period. The annual latent history is the resulting random-length sequence
\[
 U_i=
 \left(
 W_i,\{C_{it},S_{it},B_{it},K_{it}\}_{t=1}^{M_i}
 \right).
\]
Thus $Q_x$ is induced by the slowly varying kernel $q_W$, the opportunity-count kernel $\kappa_M$ and the within-opportunity structural kernels in Eq.~\eqref{eq:factorization}; it is not an additional model specified independently of them.

These parent sets are restrictions of the block-level core graph. $C_{it}$ contains trip context (road, traffic, weather, passengers, night), $S_{it}$ transient state (fatigue, distraction, attention/reaction), and $B_{it}$ immediate behaviour (speed, violations, errors). The block factorization is coarser than the component-level adjacency file because components within a block need not share identical parent sets. The dotted $V_i^{\mathrm{perf}}\to\mathrm{Speed}$ relation, for example, is excluded from $G_0$ and considered only in graph sensitivity. Vehicle safety acts on the conflict/avoidance state, and $K_{it}$ summarizes the proximal configuration through which the core graph reaches $Y_{it}$. The final line makes explicit the Markov reduction already used in Section~\ref{sec:intro}: under $G_0$, $p_{it}=\pi_Y(K_{it})$. Experience remains separate from residual heterogeneity, and $M_i$ is the realised opportunity count generated from the slower exposure process $E_i$. The variable $E_i$ indexes that conditional count distribution; it is not assumed to be a Poisson rate or to equal observed mileage.

Two block-level exclusions are worth making explicit. The core graph omits $H_i\to S_{it}$ and $R_i\to C_{it}$. These omissions restrict direct effects of persistent heterogeneity on transient state and of accumulated experience on context selection beyond $A_i$ and $G_i$; they are not claims of scientific impossibility. Both relations enter the sensitivity family in Section~\ref{sec:graphsensitivity} and \ref{app:graphspec}.

\subsection{The DAG and the actuarial observation layer}

Figure~\ref{fig:dag} separates three kinds of relation. Solid arrows define the structural core, densely dotted arrows are plausible but uncertain, and dashed arrows belong only to the observation layer. Absence of a solid arrow is a maintained structural restriction. D-separation claims are therefore conditional on the variables shown and on the assumed absence of omitted common causes \cite{greenland1999causal,pearl2009causality}. Dashed links are not used to derive adjustment sets or $do$-operator effects. Declared use measures $A_i$, territory measures $G_i$, and claims history measures $H_i$; none is inserted into the crash mechanism simply because it predicts claims.

\begin{figure}[H]
\centering
\resizebox{0.99\textwidth}{!}{
\begin{tikzpicture}[
  x=1cm,y=1cm,>=Latex,
  causal/.style={draw,rounded corners,align=center,minimum height=7mm,minimum width=18mm,fill=white,font=\scriptsize},
  latent/.style={draw,dashed,rounded corners,align=center,minimum height=7mm,minimum width=22mm,fill=gray!8,font=\scriptsize},
  rating/.style={draw,dotted,rounded corners,align=center,minimum height=7mm,minimum width=18mm,fill=gray!4,font=\scriptsize},
  info/.style={-{Latex},dashed,line width=0.65pt,gray!55},
  carrow/.style={-{Latex},line width=0.85pt},
  uarrow/.style={-{Latex},densely dotted,line width=1.05pt,black},
  layer/.style={draw=none,font=\scriptsize\itshape,gray}
]

\node[rating] (sex) at (-0.7,10.2) {Sex};
\node[rating] (age) at (1.1,10.2) {Age};
\node[rating] (tenure) at (3.0,10.2) {Licence\\tenure};
\node[rating] (mileage) at (5.1,10.2) {Annual\\mileage};
\node[rating] (use) at (7.2,10.2) {Declared\\use};
\node[rating] (territory) at (9.5,10.2) {Territory};
\node[rating] (history) at (11.9,10.2) {Claims /\\offences};
\node[rating] (vrating) at (14.7,10.2) {Vehicle\X_i info};

\node[latent] (experience) at (2.4,8.2) {Accumulated\R_i $R_i$};
\node[latent] (H) at (5.6,8.2) {Persistent\S_{it} $H_i$};
\node[latent] (A) at (8.5,8.2) {Activity / use\\profile $A_i$};
\node[latent] (G) at (11.4,8.2) {Mobility\\environment $G_i$};
\node[causal] (vperf) at (14.3,8.2) {Vehicle\\performance};
\node[causal] (adas) at (16.9,8.2) {Vehicle safety\\ / ADAS};

\node[latent] (E) at (1.0,6.0) {Exposure\\intensity $E_i$};
\node[causal] (pass) at (3.5,6.0) {Passengers};
\node[causal] (night) at (5.7,6.0) {Night};
\node[causal] (road) at (8.3,6.0) {Road type};
\node[causal] (traffic) at (10.8,6.0) {Traffic};
\node[causal] (weather) at (13.2,6.0) {Weather};

\node[causal] (distract) at (3.5,3.8) {Distraction};
\node[causal] (fatigue) at (5.7,3.8) {Fatigue};
\node[causal] (attention) at (8.3,3.8) {Attention /\\reaction};
\node[causal] (viol) at (10.8,3.8) {Violations};
\node[causal] (speed) at (13.2,3.8) {Speed};

\node[causal] (errors) at (8.3,1.6) {Errors};
\node[causal] (conflict) at (13.8,1.6) {Critical\\conflict};
\node[causal] (crash) at (16.8,1.6) {Crash\\occurrence};

\draw[carrow] (experience) to[bend right=12] (attention);
\draw[carrow] (experience) to[bend right=22] (errors);
\draw[carrow] (A) -- (pass);
\draw[carrow] (A) -- (night);
\draw[carrow] (A) to[bend left=10] (road);
\draw[carrow] (G) -- (road);
\draw[carrow] (G) -- (traffic);
\draw[carrow] (G) -- (weather);

\draw[uarrow] (pass) -- (distract);
\draw[carrow] (night) -- (fatigue);
\draw[carrow] (distract) -- (attention);
\draw[carrow] (fatigue) -- (attention);
\draw[carrow] (attention) -- (errors);
\draw[carrow] (H) to[bend right=14] (viol);
\draw[carrow] (H) to[bend left=13] (speed);
\draw[carrow] (H) to[bend right=24] (errors);
\draw[carrow] (road) to[bend right=8] (speed);
\draw[carrow] (traffic) -- (speed);
\draw[uarrow] (vperf) -- (speed);

\draw[uarrow] (errors) to[bend right=8] (conflict);
\draw[uarrow] (viol) -- (conflict);
\draw[carrow] (speed) -- (conflict);
\draw[carrow] (weather) to[bend left=8] (conflict);
\draw[carrow] (adas) to[bend left=10] (conflict);
\draw[carrow] (conflict) -- (crash);

\draw[info] (age) -- (experience);
\draw[info] (tenure) -- (experience);
\draw[info] (sex) -- (H);
\draw[info] (sex) to[bend right=12] (A);
\draw[info] (mileage) -- (E);
\draw[info] (use) -- (A);
\draw[info] (territory) -- (G);
\draw[info] (history) -- (H);
\draw[info] (vrating) -- (vperf);
\draw[info] (vrating) -- (adas);

\node[layer,anchor=west] at (18.3,10.2) {actuarial observations};
\node[layer,anchor=west] at (18.3,8.2) {persistent / assignment};
\node[layer,anchor=west] at (18.3,6.0) {trip context};
\node[layer,anchor=west] at (18.3,3.8) {state / behaviour};
\node[layer,anchor=west] at (18.3,1.0) {proximal outcome};

\end{tikzpicture}
}
\caption{The integrated DAG with explicit observation and assignment layers. Solid black arrows represent retained structural relations, heavier dotted arrows represent relations adjudicated as uncertain, and lighter dashed arrows represent prediction or measurement links from actuarial variables. Age and licence tenure inform accumulated experience $R_i$; declared use measures an activity/use profile $A_i$; territory measures a mobility-environment profile $G_i$; mileage measures exposure intensity $E_i$; and claims/offence history measures persistent heterogeneity $H_i$. The per-opportunity crash graph is a time-ordered snapshot aggregated over the number and composition of driving opportunities; dashed links are not part of the causal DAG.}
\label{fig:dag}
\end{figure}

\subsection{Graph uncertainty, coarse-graining and structural sensitivity}
\label{sec:graphsensitivity}

Let $G_0$ contain the solid arrows in Figure~\ref{fig:dag}, and let $G_+$ add the densely dotted relations. Dashed actuarial links belong to neither graph. The larger family $\mathbb G$ also includes appendix-listed alternatives that are not drawn in the main figure, such as a persistent-state $\to$ transient-state link or an experience $\to$ context-selection link. Write $\mathcal A_G$ for the annual exposure--mechanism laws compatible with graph $G$ under the maintained temporal ordering and no-unmodelled-common-cause assumptions, and let $\mathcal F_G(r)$ denote the elements that additionally reproduce annual claim contrast $r$; Section~\ref{sec:partial} gives the full definition including the crash-to-claim observation process.

Within the nested nonparametric structural model considered here, adding an arrow removes a conditional-independence restriction. If $G_1$ contains every structural arrow of $G_0$ and possibly more, then
\begin{equation}
 \mathcal A_{G_0}\subseteq\mathcal A_{G_1}
 \quad\Longrightarrow\quad
 \mathcal F_{G_0}(r)\subseteq\mathcal F_{G_1}(r),
 \label{eq:graphmonotonicity}
\end{equation}
for the same claim functional and rating contrast. Thus admitting an uncertain edge can preserve or enlarge the compatible set but cannot sharpen identification within this nested model class. The main danger runs in the opposite direction: if a true edge is omitted, the smaller graph can create spurious precision. For a finite family $\mathbb G$ of plausible graphs, define
\begin{equation}
 \mathcal F_{\mathbb G}(r)
 =\bigcup_{G\in\mathbb G}\mathcal F_G(r),
 \qquad
 \mathcal F_{\mathbb G,\mathrm{ext}}(r)
 =\bigcup_{G\in\mathbb G}
 \bigl(\mathcal F_G(r)\cap\mathcal F_{\mathrm{RS},G}\bigr),
 \label{eq:graphunion}
\end{equation}
where $\mathcal F_{\mathrm{RS},G}$ denotes the road safety restrictions that are meaningful under graph $G$. This union keeps graph uncertainty separate from uncertainty in effect magnitude and transportability.

The low-dimensional age illustration does not recover these graph-specific sets. Let $\Gamma$ map a fine-scale mechanism description to the three displayed bookkeeping blocks $(c_E,c_C,c_B)$. The simplex and transport-restricted polygons in Section~\ref{sec:workedintersection} impose only the annual log contrast, non-negativity and one external bound on $c_C$. For every graph to which those common restrictions apply,
\begin{equation}
 \Gamma\{\mathcal F_{G,\mathrm{ext}}(r)\}
 \subseteq
 \mathcal C_{\mathrm{ext}}(\log r;\delta).
 \label{eq:coarseouter}
\end{equation}
The displayed polytope is a \emph{coarse outer compatibility set}, not the exact graph-specific projection of $G_0$. Candidate edges such as passengers $\to$ distraction, vehicle performance $\to$ speed, and errors/violations $\to$ conflict alter fine-scale pathways, while a residual age $\to$ crash shortcut changes what the residual block absorbs. The three-block projection contains no constraint that distinguishes these alternatives, so a second simplex labelled $G_+$ would add no information. Failure to see graph uncertainty after projection reflects coarse-graining, not equivalence of $G_0$ and $G_+$.

\ref{app:graphspec} lists the principal candidate additions and identifies which structural restriction each one relaxes. A graph-sensitive empirical implementation would need measurements or external constraints below the present block resolution---for example distraction conditional on passenger configuration, speed conditional on vehicle performance, or transient-state measurements stratified by persistent driver state.

\subsection{From trip-level causation to annual risk}

No independence assumption is needed to separate exposure quantity from average opportunity risk. To keep the zero-opportunity case explicit, define for $M_i>0$
\[
 \bar p_i = \frac{1}{M_i}\sum_{t=1}^{M_i}p_{it},
 \qquad
 \pi_+(X_i)=\Pr(M_i>0\mid X_i).
\]
Because the random sum in Eq.~\eqref{eq:central} is zero when $M_i=0$,
\begin{equation}
\begin{aligned}
 \lambda_{\mathrm{crash}}(X_i)
 =\pi_+(X_i)\Bigl[
 &\mathbb E(M_i\mid X_i,M_i>0)\,
  \mathbb E(\bar p_i\mid X_i,M_i>0)\\
 &+\operatorname{Cov}(M_i,\bar p_i\mid X_i,M_i>0)
 \Bigr].
\end{aligned}
\label{eq:quantityquality}
\end{equation}
If $\Pr(M_i>0\mid X_i)=1$, the leading factor is one and this reduces to the familiar covariance decomposition. The covariance term is an exact accounting term, not an estimated mechanism: it records dependence between how much a driver is exposed and the risk composition of that exposure.

Observed mileage need not behave as a linear offset in an annual frequency model. Drivers with different mileage can differ in route, time, traffic and behavioural composition, so exposure quantity may be associated with average opportunity risk. Elvik \cite{elvik2023mileage} documents a substantially sublinear aggregate relation between annual distance and accident involvement. That relation constrains quantity and composition jointly without selecting a mechanism.

\section{Rating information and latent mechanisms}
\label{sec:rating}

Conventional rating variables enter primarily as information about latent mechanisms. A more proximal structural role is retained only where the evidence supports one. \ref{app:evidencetables} reports the detailed bridge table.

Age illustrates the distinction. A direct $\mathrm{Age}\to\mathrm{Crash}$ arrow would collapse accumulated experience, exposure composition and behavioural pathways into one reduced-form edge. For two age groups $a$ and $a_0$, the structural crash relativity is
\begin{equation}
 RR^{\mathrm{crash}}_{\mathrm{age}}(a;a_0)
 =
 \frac{\mathbb{E}[\sum_{t=1}^{M_i} p_{it}\mid \mathrm{Age}_i=a]}
 {\mathbb{E}[\sum_{t=1}^{M_i} p_{it}\mid \mathrm{Age}_i=a_0]}.
 \label{eq:rrage}
\end{equation}
The observed insurance analogue replaces the crash sum by $\sum_t p_{it}\rho_{it}$ from Eq.~\eqref{eq:claimobservation}. The two quantities coincide only under additional assumptions on the claim-observation process. Licence tenure is closer to accumulated driving experience $R_i$ than age alone, and novice-driver studies document substantial changes with early independent-driving experience \cite{mccartt2003experience,gulliver2013learner,curry2015experience,ehsani2020learner}. Age and tenure can therefore inform $q_W(R_i\mid X_i)$ jointly, but they do not by themselves separate maturation, cohort and time-since-licensure. The \texttt{freMTPL2freq} illustration contains age but not licence tenure.

Sex and prior claims are also observation-layer variables in the core model. Exposure studies and telematics show that sex can carry information about mileage, time of day and speed distributions \cite{massie1997gender,regev2018agegender,ayuso2016gender,guillen2021speeding}; assigning a single sex-to-crash mechanism would discard that heterogeneity. Claims history has an even clearer temporal interpretation:
\begin{equation}
\begin{aligned}
 H_i,R_i,\ldots &\longrightarrow Y_{i,<t}
 \longrightarrow J_{i,<t}
 \longrightarrow \mathrm{PastClaims}_i,\\
 H_i &\longrightarrow B_{it}\longrightarrow K_{it}\longrightarrow Y_{it}.
\end{aligned}
 \label{eq:claimhistory}
\end{equation}
The dashed $\mathrm{PastClaims}_i\to H_i$ relation reverses the generative direction on purpose: it is a statistical update about persistent risk, not a physical input to the next collision. Bonus--malus inherits this endogeneity because it is constructed from responsible-claim history and insurance duration.

Declared use and territory are routed through stable latent bridge objects rather than directly to crash. Declared use informs an activity/use profile $A_i$ that affects opportunity count and context composition; territory informs a mobility environment $G_i$ associated with recurrent road, traffic and environmental conditions. Activity-based, work-driving, spatial and telematics studies support these intermediate objects \cite{elias2010activity,robb2008work,newnam2022systems,chipman1993exposure,blatt1998residence,lee2014residence,ma2018context,guillen2024weekly,shi2017territorial}. The evidence does not identify which individual road safety mechanism explains a given declared-use or territorial relativity.

Vehicle information requires a different separation. Collision-avoidance technology can act directly on crash occurrence through the conflict/avoidance branch. Broad vehicle age, weight, power or model categories mix physical mechanisms, severity and driver--vehicle selection. Høye's analysis of registration year, age and weight concerns killed-or-seriously-injured outcomes \cite{hoye2019vehicle}, so it cannot by itself establish an occurrence effect. Power and performance are associated with operating speed or crash involvement \cite{mccartt2017power,keall2013performance}, but selection into vehicle type remains a competing explanation. The graph distinguishes safety technology and performance mechanisms from the coarse vehicle label.

\section{Compatibility analysis, sensitivity and partial identification}
\label{sec:partial}

\subsection{Scope of the quantitative illustrations}

The set-valued perspective follows the partial-identification tradition in distinguishing point identification from the information contained in a maintained model and its assumptions \cite{manski2003partial,tamer2010partial}. Here that distinction is used diagnostically. The graph-specific set $\mathcal F_G(r)$ below is the collection of structural and observation laws compatible with an annual contrast under the maintained graph. By contrast, the $\delta$-indexed transport envelopes are \emph{sensitivity sets conditional on user-specified discrepancy bounds}; $\delta$ is not identified from the French portfolio.

The calculations below are lower-dimensional than the DAG. They do not estimate causal effects of mileage or age, recover the full mechanism distribution, validate the graph, or estimate transport discrepancy. They show where annual aggregate information constrains a mechanistic explanation and where it does not. The log-additive age representation is only a bookkeeping device for one coarse projection; the nonparametric DAG allows interactions.


\subsection{Edge-by-edge evidence synthesis}

Numerical evidence is retained on the scale on which it was estimated. The extraction table stores the reported study-level effect $\widehat\eta_{ej}$ separately from the target structural quantity $\theta_e$. For a ratio estimate with confidence limits $(L_{ej},U_{ej})$, define
\[
 z_{ej}=\log(\widehat\eta_{ej}),
 \qquad
 s_{ej}\simeq\frac{\log(U_{ej})-\log(L_{ej})}{2\times1.96}.
\]
For non-ratio estimands an appropriate design-specific transformation replaces the logarithm. The pairs $(z_{ej},s_{ej})$ are not assumed exchangeable.

Let $T_{d,o,\mathsf{pop}}$ map the target edge-level structural parameter or response function $\theta_e$ to the estimand generated by design $d$, outcome $o$ and population $\mathsf{pop}$. For study $j$ on edge $e$,
\begin{equation}
 z_{ej}=T_{d_j,o_j,\mathsf{pop}_j}(\theta_e)
       +\Delta_{ej}+\varepsilon_{ej},
 \qquad
 \varepsilon_{ej}\sim N(0,s_{ej}^2),
 \label{eq:transport}
\end{equation}
where $\Delta_{ej}$ is residual estimand or population discrepancy. The mapping $T$ and the discrepancy bound play different roles: $T$ encodes a substantive relation between target and study estimand; $\Delta$ absorbs remaining mismatch. If neither can be defended, the study remains structural evidence and is not used numerically. The worked Gomes-Franco example below makes this explicit by using $T(\theta_C)=\theta_C$ on the log-odds scale before a separate discrepancy bound is introduced; no comparable mapping is asserted for the other heterogeneous anchors.

For an edge-specific vector $\delta_e=(\delta_{e1},\ldots,\delta_{eJ_e})$ with $|\Delta_{ej}|\le\delta_{ej}$, define
\begin{equation}
 \Pi_e(\delta_e)=\left\{\theta_e:
 |z_{ej}-T_{d_j,o_j,\mathsf{pop}_j}(\theta_e)|
 \le 1.96s_{ej}+\delta_{ej}
 \quad\text{for every }j=1,\ldots,J_e
 \right\}.
 \label{eq:edgeenvelope}
\end{equation}
This is a sensitivity-feasible set, not a joint 95\% confidence region. With several heterogeneous studies, $\Pi_e(\delta_e)$ need not be an interval and can be disconnected or empty if the transported estimands are mutually incompatible. Emptiness is informative: it signals that the chosen mappings, discrepancy bounds and retained evidence cannot all hold simultaneously.

Suppose a collection of edges $\mathcal E^\star$ enters a block contribution $c_k=b_k(\vartheta)$, where $\vartheta=(\theta_e:e\in\mathcal E^\star)$ is the joint structural parameter. Rather than taking a Cartesian product and thereby suggesting statistical independence, define the joint feasible set directly as
\begin{equation}
\begin{aligned}
 \Pi_{\mathcal E^\star}(\boldsymbol\delta)
 =\bigl\{\vartheta:\;&
 \theta_e\in\Pi_e(\delta_e)
 \quad\text{for every }e\in\mathcal E^\star,\\
 &\vartheta\text{ satisfies the stated cross-edge structural restrictions}
 \bigr\}.
\end{aligned}
 \label{eq:jointenvelope}
\end{equation}
No independence between study estimates or edge parameters is implied by this notation. When dependence information is available it belongs in the joint restriction; when it is unavailable, the construction is a deterministic compatibility analysis. Block-level bounds are then projections,
\begin{equation}
 \ell_k(\boldsymbol\delta)=
 \inf_{\vartheta\in\Pi_{\mathcal E^\star}(\boldsymbol\delta)} b_k(\vartheta),
 \qquad
 u_k(\boldsymbol\delta)=
 \sup_{\vartheta\in\Pi_{\mathcal E^\star}(\boldsymbol\delta)} b_k(\vartheta).
 \label{eq:blockprojection}
\end{equation}
Larger discrepancy bounds weakly enlarge the feasible set by construction; that monotonicity is a logical property of the sensitivity analysis, not an empirical finding.

The classes in Table~\ref{tab:qclasses} enter at different stages. Q1 estimates can constrain $\theta_e$ after limited transport sensitivity. Q2 estimates require a substantive mapping or a wider discrepancy set. Q3 records affect graph structure only, while class P records inform the observation layer. Culpability odds ratios, policy effects on fatal crashes and naturalistic near-crash estimates are therefore not treated as repeated measurements of one coefficient.

Quantitative synthesis remains local to an edge or mechanism. The worked example below has $J_e=1$, so its feasible set is deliberately simple; the more general notation above is needed because several evidence-map edges contain non-commensurate studies for which intersection, not pooling, is the appropriate operation.

\subsection{Rating relativities do not identify their explanations}

Suppose an actuarial model estimates a \emph{claim-frequency} relativity $r_{\mathrm{claim}}(x)$ for $X=x$ relative to a baseline $x_0$. The full annual law introduced in Section~\ref{sec:intro} is
\[
 Q_x=\mathcal L(M_i,U_i\mid X_i=x),
\]
with the slowly varying marginal $q_W(\cdot\mid x)$ specified in Eq.~\eqref{eq:measurementlayer}. For a given annual history $(M,U)$, define the conditional claim sum
\begin{equation}
 \mu_{\mathrm{claim}}(M,U;x)
 =\sum_{t=1}^{M} p_t(U)\,\rho_t(U,x),
 \label{eq:annualclaimfunctional}
\end{equation}
where $p_t(U)$ is the opportunity-level crash probability and $\rho_t$ is the crash-to-claim conditional mean from Eq.~\eqref{eq:rho}. The observed relativity satisfies
\begin{equation}
 r_{\mathrm{claim}}(x)
 =
 \frac{\mathbb{E}_{Q_x}\{\mu_{\mathrm{claim}}(M,U;x)\}}
      {\mathbb{E}_{Q_{x_0}}\{\mu_{\mathrm{claim}}(M,U;x_0)\}}.
 \label{eq:inverse}
\end{equation}
Setting $\rho_t\equiv1$ recovers the corresponding crash-frequency functional, but claim data do not justify that restriction.

For graph $G$, let $\mathcal Q_G(x)$ be the set of annual laws $Q_x$ induced by structural models compatible with $G$, its temporal ordering and the maintained no-unmodelled-common-cause assumptions. Let $\rho_x(t,U)$ be a measurable crash-to-claim conditional-mean mapping and let $\mathcal R_x$ denote its admissible class. Define the claim functional
\begin{equation}
 \mu(Q_x,\rho_x)
 =
 \mathbb E_{Q_x}\!\left[
   \sum_{t=1}^{M}p_t(U)\rho_x(t,U)
 \right].
 \label{eq:claimfunctionalQ}
\end{equation}
The graph-specific compatibility set is then
\begin{equation}
\begin{aligned}
 \mathcal F_G(r)=\biggl\{(Q_x,Q_{x_0},\rho_x,\rho_{x_0}):\;&
 Q_x\in\mathcal Q_G(x),\quad Q_{x_0}\in\mathcal Q_G(x_0),\\
 &\rho_x\in\mathcal R_x,\quad \rho_{x_0}\in\mathcal R_{x_0},\\
 &\frac{\mu(Q_x,\rho_x)}
 {\mu(Q_{x_0},\rho_{x_0})}=r
 \biggr\}.
\end{aligned}
 \label{eq:Fset}
\end{equation}
This definition supplies the $\mathcal F_G(r)$ used in Section~\ref{sec:graphsensitivity}. Road safety evidence restricts crash mechanisms and the study-to-target mappings that are meaningful under graph $G$; it does not by itself identify the observation functions $\rho$. Denote those graph-specific external restrictions by $\mathcal F_{\mathrm{RS},G}$. Then
\begin{equation}
 \mathcal F_{G,\mathrm{ext}}(r)
 =\mathcal F_G(r)\cap\mathcal F_{\mathrm{RS},G}.
 \label{eq:idset}
\end{equation}
The baseline analysis uses $G_0$, while Eq.~\eqref{eq:graphunion} describes graph-robust inference over $\mathbb G$. These are compatibility sets, not claims that the available insurance data point-identify causal mechanisms. A tariff can be predicted precisely while both the crash pathways and the crash-to-claim bridge behind its rating contrast remain weakly constrained.

The next construction is only a low-dimensional geometric illustration. Neither the DAG nor Eq.~\eqref{eq:inverse} implies log additivity, and arbitrary interactions are allowed in the underlying nonparametric model. For visualization, let a user-chosen coarse map $\Gamma$ summarize a fine-scale compatible mechanism law into $K$ log-scale bookkeeping blocks $c=(c_1,\ldots,c_K)$ whose sum reproduces $L=\log r$. Writing $\mathbf 1$ for the $K$-vector of ones, the unrestricted display set is
\begin{equation}
 \mathcal{C}(L)=\{c\in\mathbb{R}^K:\mathbf{1}^{\top}c=L\}.
 \label{eq:hyperplane}
\end{equation}
is an unbounded affine hyperplane for $K\ge2$. Imposing $c_k\ge0$ gives the bounded simplex
\begin{equation}
 \mathcal{C}_{+}(L)=\mathcal{C}(L)\cap\mathbb{R}_{+}^{K},
 \label{eq:simplexgeneral}
\end{equation}
while external bounds $\ell_k(\boldsymbol\delta)\le c_k\le u_k(\boldsymbol\delta)$ obtained from Eq.~\eqref{eq:blockprojection} produce the convex polytope
\begin{equation}
 \mathcal{C}_{\mathrm{ext}}(L;\boldsymbol\delta)=
 \mathcal{C}(L)\cap\{c:\ell_k(\boldsymbol\delta)\le c_k\le u_k(\boldsymbol\delta),\ k=1,\ldots,K\}.
 \label{eq:polytope}
\end{equation}
The simplex is not a causal decomposition theorem. With interactions, the compatible set is defined by a nonlinear constraint $g(c)=L$ and need not be convex; with an explicit claim-observation model, additional dimensions may be required as well.

\subsection{Two quantitative submodels}
\label{sec:quantitative}

Two small submodels make the identification problem concrete without fitting the full graph. The mileage example supplies an aggregate constraint; the age example defines an annual insurance contrast and studies the mechanism sets compatible with it.

\subsubsection{Mileage: an aggregate constraint on risk per unit distance}

Use observed annual mileage $m$ as the exposure unit, without equating a mile to a trip or to latent intensity $E_i$. The object in this subsection is the \emph{marginal aggregate} accident relation reported in the mileage literature, not the conditional insurance mean $\lambda_{\mathrm{claim}}(m,Z)$ from Section~\ref{sec:timescales}. Denote it by $\lambda_{\mathrm{agg}}(m)$ and define aggregate accident risk per unit distance as
\begin{equation}
  \bar p_{\mathrm{agg}}(m)=\frac{\lambda_{\mathrm{agg}}(m)}{m},
  \qquad\text{so that}\qquad
  \lambda_{\mathrm{agg}}(m)=m\,\bar p_{\mathrm{agg}}(m).
  \label{eq:mileagefactor}
\end{equation}
Taking log derivatives of this identity gives
\begin{equation}
  \eta_{\lambda,m}=1+\eta_{\bar p,m},
  \qquad
  \eta_{\lambda,m}=\frac{d\log\lambda_{\mathrm{agg}}(m)}{d\log m}.
  \label{eq:mileageelasticity}
\end{equation}
More generally, an aggregate power approximation $\lambda_{\mathrm{agg}}(m)\propto m^{\alpha}$ implies
\begin{equation}
  \eta_{\bar p,m}=\alpha-1.
  \label{eq:compositionelasticity}
\end{equation}
Elvik's synthesis uses $\alpha\approx1/2$ as a useful approximation to the sublinear mileage--accident relation in the studies considered \cite{elvik2023mileage}. It is not reported as a universal exponent with a sampling interval. At $\alpha=1/2$, four times the mileage corresponds to about twice as many annual accidents and half the average risk per mile. A non-inferential sensitivity check with $\alpha\in\{0.4,0.5,0.6\}$ gives fourfold accident multipliers of 1.74, 2.00 and 2.30 and risk-per-mile multipliers of 0.44, 0.50 and 0.57.

The decline in $\bar p_{\mathrm{agg}}(m)$ follows arithmetically from sublinear $\lambda_{\mathrm{agg}}(m)$ and does not imply that additional mileage protects a driver. Elvik cautions against a causal reading of risk-per-distance ratios \cite{elvik2023mileage}. Route and time composition, driver selection, accumulated experience and unobserved heterogeneity may all contribute. Naturalistic analyses of the low-mileage bias support roles for composition and heterogeneity without separating their shares \cite{janke1991mileage,langford2006lowmileage,antin2017lowmileage}.

If average risk per mile is factorized multiplicatively into mechanism blocks for bookkeeping, the corresponding elasticities add:
\begin{equation}
 \eta_{\bar p,m}=\eta_{\mathrm{road}}+\eta_{\mathrm{time}}+\eta_{\mathrm{traffic}}+\eta_{\mathrm{experience/selection}}+\eta_{\mathrm{other}}.
 \label{eq:mileageblocks}
\end{equation}
The aggregate relation identifies only the sum in Eq.~\eqref{eq:mileageblocks}, approximately $-1/2$ under the square-root approximation. The decomposition is not implied by the DAG and fails under non-additive interactions.

\begin{figure}[H]
\centering
\includegraphics[width=0.78\textwidth]{fig_mileage_elasticity_revised.pdf}
\caption{Aggregate elasticity identity under the working approximation $\lambda_{\mathrm{agg}}(m)\propto m^{1/2}$. Observed mileage grows linearly by definition of the horizontal exposure scale, while the associated accident count grows as $m^{1/2}$ and the ratio $\lambda_{\mathrm{agg}}(m)/m$ declines as $m^{-1/2}$. The declining ratio is an aggregate constraint, not an identified causal composition effect.}
\label{fig:mileageelasticity}
\end{figure}

\subsubsection{Age: the rating relativity depends on the information set}

The age illustration uses \texttt{freMTPL2freq} from \texttt{CASdatasets}, an open French motor third-party-liability portfolio with 677,991 policy records \cite{dutang2026casdatasets,noll2020fremtpl}. All counts reported here refer to the \texttt{CASdatasets} version loaded by the reproducibility script rather than to external mirrors. Restricting to $0<\mathrm{Exposure}\leq1$ leaves 676,767 records: fractional-year exposures are retained, while records exceeding one policy-year are excluded to keep the annual exposure convention. Claim counts are fitted by Poisson pseudo-maximum likelihood with log exposure as offset; sandwich inference does not rely on equidispersion. Driver age is grouped into 18--20, 21--24, 25--29, 30--39, 40--49, 50--59, 60--69, 70--79 and 80+, with 40--49 as reference. Vehicle age is grouped as 0, 1--4, 5--9, 10--14 and 15+ years; these are pragmatic descriptive bands, not estimated cut points. Density enters through empirical type-7 deciles. The vehicle/geographic model also includes vehicle power, fuel, area and region.

A \emph{rating cell} is an observed combination of the fine age, vehicle and geographic partition, augmented by bonus--malus when that variable enters the mean model. Counts and exposures are aggregated within cells, and uncertainty is computed with a cell-level Huber--White sandwich. The age-only model is evaluated on the same fine partition, avoiding a saturated nine-cell sandwich calculation while preserving its point estimates. The computational appendix records the full construction.

The vehicle/geographic specification gives an annual claim-frequency relativity of
\begin{equation}
 \widehat{RR}_{18--20:40--49}=3.388\quad (95\%\ \mathrm{CI}:\ 3.114,\ 3.686).
 \label{eq:agerelativity}
\end{equation}
Vehicle and geographic adjustment does not attenuate the young-driver contrast: the 18--20 versus 40--49 relativity is 3.388 (95\% CI 3.114--3.686), compared with a raw value of 3.182 (95\% CI 2.923--3.464). There is no reason for adjustment to move a predictive coefficient monotonically toward one; vehicle and geographic covariates are associated with both age and claims, so the conditional contrast can increase or decrease. Adding the medium bonus--malus discretisation reduces the conditional relativity to 1.235 (95\% CI 1.129--1.352). Table~\ref{tab:bmsensitivity} varies only the resolution of the observed \texttt{BonusMalus} score. Fixed-width bins of 25, 10 and 5 points, anchored at 50, yield 8, 19 and 34 realised categories and relativities from 1.381 to 1.209. These bins are a sensitivity device for the dataset's recorded score, not an attempt to reconstruct the statutory bonus--malus step grid. The finest confidence interval remains above one. This attenuation reflects conditioning on an endogenous summary of prior insurance history, not removal of confounding or recovery of a causal age effect.

\begin{table}[H]
\centering
\begingroup
\scriptsize
\setlength{\tabcolsep}{3pt}
\begin{tabular}{lrrrr}
\toprule
Bonus--malus specification & BM levels & Rating cells & $\widehat{RR}_{18--20:40--49}$ & 95\% CI \\
\midrule
Coarse & 8 & 94,840 & 1.381 & [1.264, 1.510] \\
Medium & 19 & 125,251 & 1.235 & [1.129, 1.352] \\
Fine & 34 & 153,155 & 1.209 & [1.104, 1.324] \\
\bottomrule
\end{tabular}

\endgroup
\caption{Sensitivity of the incremental 18--20 versus 40--49 age relativity to alternative fixed-width categorical resolutions of bonus--malus. The coarse, medium and fine specifications use bin widths 25, 10 and 5 respectively, anchored at 50. All specifications retain the same vehicle and geographic controls; only the resolution of bonus--malus changes.}
\label{tab:bmsensitivity}
\end{table}

More generally, let $\mathcal Z$ denote the \emph{set of rating variables included in the fitted claim model}, rather than a random covariate value. For a log-link model without age interactions, write the fitted age relativity as
\begin{equation}
 r^{\mathrm{claim}}_{\mathrm{age}}(a;\mathcal Z)
 =\exp\{\beta_a^{(\mathcal Z)}-\beta_{a_0}^{(\mathcal Z)}\}.
 \label{eq:conditionalrating}
\end{equation}
Changing $\mathcal Z$ changes the predictive contrast before any mechanistic interpretation. Bonus--malus requires particular caution because it is not a baseline confounder. In the French system the coefficient evolves mechanically with claim-free insurance periods and responsible claims \cite{france2025bonusmalus}. Figure~\ref{fig:bmdag} summarizes the relevant history structure. Persistent heterogeneity and experience contribute to earlier crashes, earlier crash-to-claim realizations feed the recorded bonus--malus score, and insurance-history duration is another cause of that score. The same persistent states also contribute to future crash risk.

\begin{figure}[H]
\centering
\resizebox{0.96\textwidth}{!}{
\begin{tikzpicture}[>=Latex, node distance=12mm and 14mm,
  n/.style={draw,rounded corners,align=center,font=\scriptsize,inner sep=3pt},
  cond/.style={draw,double,rounded corners,align=center,font=\scriptsize,inner sep=3pt}]
\node[n] (age) {Age /\ tenure};
\node[n, right=of age] (latentbm) {$R_i,H_i$};
\node[n, right=of latentbm] (pastcrash) {past\ crash};
\node[n, right=of pastcrash] (pastclaim) {past liability\ claim};
\node[cond, right=of pastclaim] (bm) {bonus--malus\ conditioned on};
\node[n, below=of bm] (duration) {insurance-history\ duration};
\node[n, below=of pastcrash] (futurecrash) {future\ crash};
\node[n, right=of futurecrash] (futureclaim) {future liability\ claim};
\draw[->] (age) -- (latentbm);
\draw[->] (latentbm) -- (pastcrash);
\draw[->] (pastcrash) -- (pastclaim);
\draw[->] (pastclaim) -- (bm);
\draw[->] (duration) -- (bm);
\draw[->] (latentbm) -- (futurecrash);
\draw[->] (futurecrash) -- (futureclaim);
\end{tikzpicture}
}
\caption{Schematic history sub-DAG for interpreting bonus--malus conditioning. The double border marks the variable conditioned on in the predictive model. The diagram does not assert that bonus--malus causes future crashes; it shows why conditioning on a downstream summary of past claims changes the target and can, under fuller graphs, induce associations among causes of the conditioned variable.}
\label{fig:bmdag}
\end{figure}

The bonus--malus-adjusted value 1.235 is therefore a \emph{conditional predictive contrast}: under the fitted no-age-interaction log-link model, it compares fitted claim rates for the two age groups at the same included vehicle, geographic and bonus--malus covariate values. It is not a controlled direct effect of age, a deconfounded crash-risk contrast, or an effect ``among otherwise identical drivers.'' Conditioning removes variation associated with histories summarized by bonus--malus and, depending on the fuller graph, may also open non-causal associations among its multiple causes; these are history-conditioning and possible collider concerns rather than ordinary baseline-confounding control \cite{greenland1999causal,pearl2009causality}. Bonus--malus may proxy persistent heterogeneity and insurance duration, but it is not accumulated driving experience $R_i$, and the portfolio does not identify those channels separately.

\begin{figure}[H]
\centering
\includegraphics[width=0.82\textwidth]{fig_age_relativity_fremtpl2_v14.pdf}
\caption{Annual claim-frequency relativities by age in \texttt{freMTPL2freq}, with ages 40--49 as reference. Raw relativities are compared with the vehicle/geographic model and with the same model augmented by the medium bonus--malus discretisation (10-point bins; 19 realised levels). Intervals are cell-level sandwich 95\% confidence intervals. Bonus--malus is an endogenous history variable, so the comparison describes alternative predictive conditioning sets; it is not a causal adjustment or an intervention on age.}
\label{fig:agerelativity}
\end{figure}

\subsubsection{Coarse compatibility sets for the age explanation}

The information-set comparison changes the constraint before external road safety evidence is introduced. Let
\[
 L_{\mathrm{base}}=\log(3.387706)=1.220153,
 \qquad
 L_{\mathrm{BM}}=\log(1.235315)=0.211326,
\]
where the first target uses vehicle/geographic information and the second adds the displayed medium bonus--malus specification. Both are claim-scale contrasts. For the worked geometry, the coarse map $\Gamma$ groups the fine DAG into three log-additive bookkeeping blocks. The experience/human block $c_E$ collects contrast routed through $R_i$, $H_i$ and the transient/behavioural states they influence; the environment/vehicle/context block $c_C$ collects contrast routed through $A_i$, $G_i$, $C_{it}$ and vehicle mechanisms; and the balance term $c_B$ absorbs unresolved interactions, omitted mechanisms and crash-to-claim differences not represented in the first two blocks. The symbol $c_B$ is a balance term and should not be confused with the behavioural state $B_{it}$. This grouping is a display map, not a unique path decomposition or a claim that the underlying nonparametric DAG is log additive. These blocks are not asserted to be causal shares. Under the visualization restriction $c_E,c_C,c_B\ge0$, the compatible set for either target is
\begin{equation}
 \mathcal{C}_{+}(L_z)=
 \{(c_E,c_C,c_B)\in\mathbb{R}_+^3:c_E+c_C+c_B=L_z\},
 \qquad z\in\{\mathrm{base},\mathrm{BM}\}.
 \label{eq:agesimplex}
\end{equation}
Every point in either simplex reproduces the corresponding annual claim relativity under this bookkeeping restriction. Adding bonus--malus changes the claim target itself: the log contrast falls from 1.220 to 0.211 before external safety evidence enters. We do not interpret the smaller value as a causally adjusted target. Because bonus--malus is downstream of prior crash and claim histories, the road safety intersection below uses $L_{\mathrm{base}}$, not $L_{\mathrm{BM}}$. Experience, context and behaviour are plausible age-related pathways \cite{curry2015experience,simonsmorton2005passengers,goodwin2012passengers,regev2018agegender,guillen2021speeding}, but the insurance contrast does not determine their shares. Non-negativity is imposed only to obtain a bounded visualization.

\subsubsection{A worked intersection with external road safety evidence}
\label{sec:workedintersection}

One study is carried through Eqs.~\eqref{eq:transport}--\eqref{eq:polytope} to illustrate an external restriction. Gomes-Franco et al.~\cite{gomesfranco2020age} analyse Spanish police-recorded crashes and compare drivers aged 18--24 with ages 35--44. Their mediation model reports a total culpability OR of 2.15, a direct-path OR of 1.97 and an indirect OR of 1.09 (95\% CI 1.08--1.10) through environmental and vehicle circumstances. The reported components are multiplicatively compatible on the odds-ratio scale because $\log(1.97)+\log(1.09)\simeq\log(2.15)$. They are not commensurate with the French claim contrast: population, age bands, culpability definition and crash-to-claim observation all differ.

In this worked example, $c_C$ is the environment/vehicle/context block defined above; $c_E+c_B$ contains the remaining human and unresolved pathways. The study's indirect-path estimate is used as a Q2 anchor with identity map $T(\theta_C)=\theta_C$ before transport discrepancy is added. On the log scale,
\begin{equation}
 z_C=\log(1.09)=0.0862,
 \qquad
 s_C\simeq\frac{\log(1.10)-\log(1.08)}{2\times1.96}=0.00468.
 \label{eq:gomesanchor}
\end{equation}
We do not assume identity transport from a Spanish culpability OR to a French liability-claim relativity. The discrepancy $\Delta_C$ collects at least three gaps: population and age-band differences, culpability versus the target crash-occurrence construct, and the crash-to-recorded-claim mapping in Eq.~\eqref{eq:claimobservation}. Bounding their combined effect by $|\Delta_C|\le\delta$ gives
\begin{equation}
 \ell_C(\delta)=\max\{0,z_C-1.96s_C-\delta\},
 \qquad
 u_C(\delta)=z_C+1.96s_C+\delta.
 \label{eq:gomesbound}
\end{equation}
For the vehicle/geographic actuarial target $L_{\mathrm{base}}=1.220153$, the externally restricted set is
\begin{equation}
\begin{aligned}
 \mathcal C_{\mathrm{ext}}(L_{\mathrm{base}};\delta)
 =\bigl\{(c_E,c_C,c_B)\in\mathbb R_+^3:\;&
 c_E+c_C+c_B=L_{\mathrm{base}},\\
 &\ell_C(\delta)\le c_C\le u_C(\delta)\bigr\}.
\end{aligned}
 \label{eq:workedpolytope}
\end{equation}
Table~\ref{tab:transportworked} uses three discrepancy levels. The first, $\delta=0$, is a consistency check: since $s_C$ is reconstructed from the published confidence interval, this row simply reproduces the study uncertainty on the log scale. The next two permit multiplicative discrepancies of 1.05 and 1.10 in either direction, using $\delta=\log(1.05)$ and $\delta=\log(1.10)$. These are sensitivity scenarios, not empirically calibrated bounds on cross-country or crash-to-claim transport. Empirical calibration of $\delta$ would require comparable estimates across populations and outcome definitions.

\begin{table}[H]
\centering
\small
\begin{tabular}{lccc}
\toprule
Transport sensitivity & $\delta$ & bound on $c_C$ & equivalent ratio range \\
\midrule
Study uncertainty only & 0 & $[0.077,\,0.095]$ & $[1.080,\,1.100]$ \\
Allow factor 1.05 & $\log(1.05)$ & $[0.028,\,0.144]$ & $[1.029,\,1.155]$ \\
Allow factor 1.10 & $\log(1.10)$ & $[0.000,\,0.191]$ & $[1.000,\,1.210]$ \\
\bottomrule
\end{tabular}

\caption{Worked transport-sensitivity bounds obtained from the environmental/vehicle indirect-path estimate of Gomes-Franco et al.~\cite{gomesfranco2020age}. The last two rows enlarge the study uncertainty by a user-specified log-scale discrepancy. They are sensitivity scenarios, not confidence intervals for transportability.}
\label{tab:transportworked}
\end{table}

Figure~\ref{fig:externalcalibration} holds $L_{\mathrm{base}}$ fixed. The polygons narrow because the external study restricts $c_C$, not because the insurance target changes. Substantial width remains: the Gomes-Franco direct path does not separate accumulated experience, risky behaviour and other driver mechanisms, so $c_E$ and $c_B$ remain unresolved even under the tightest scenario. The external estimate excludes some decompositions but does not select one.

The figure is also coarse with respect to graph uncertainty. Under Eq.~\eqref{eq:coarseouter}, its polytope is an outer compatibility region after many component-level pathways have been collapsed into three blocks. Candidate arrows in $G_+$ change the fine-scale admissible laws, but no sub-block measurement in the illustration can reveal those changes. Separate $G_0$ and $G_+$ polygons would therefore suggest unsupported precision.

\begin{figure}[H]
\centering
\includegraphics[width=0.78\textwidth]{fig_age_external_calibration_pass2.pdf}
\caption{Worked coarse compatibility region for the vehicle/geographic age relativity after adding external road safety evidence. The outer triangle is $\mathcal C_+(L_{\mathrm{base}})$ with $L_{\mathrm{base}}=\log(3.387706)=1.220153$; the nested polygons impose the environmental/vehicle indirect-path anchor from Gomes-Franco et al.~\cite{gomesfranco2020age} under increasingly permissive transport discrepancies. These polygons are outer bookkeeping sets at the three-block resolution, not exact graph-specific projections, posterior probability regions, or evidence that the Spanish mediation estimand transports without discrepancy to the French portfolio.}
\label{fig:externalcalibration}
\end{figure}

\subsection{What the quantitative examples identify}

The three numerical objects play different roles. Mileage supplies an aggregate crash-involvement elasticity constraint. The French portfolio supplies claim-frequency contrasts under different predictive conditioning sets. The Gomes-Franco estimate restricts one bookkeeping block of the vehicle/geographic claim contrast only after a stated mapping that includes crash-to-claim discrepancy. The remaining width is a diagnostic of unresolved mechanism and observation-process information; the analysis does not convert it into a causal percentage or a graph-wide posterior.

\section{Implications for road safety research and validation}
\label{sec:telematics}

Annual liability claims are an aggregated, administratively selected outcome. For road safety research they can provide predictive contrasts and constraints on exposure or context, but they are not direct observations of crash occurrence unless the crash-to-claim process is modelled or observed in linked data. The converse caution applies to insurance interpretation: an age, territory or claims-history coefficient predicts recorded losses but does not establish the rating label as a proximal crash cause.

The immediate implication is a data requirement, not a new estimator. Sharper mechanistic inference needs observations between the policy-year predictor and the recorded claim: trip context, exposure composition, transient driver states, conflicts or near-crashes, and preferably links between crashes and subsequent claims. Telematics can supply speed, braking, timing, road type and route context; fatigue, attention and persistent behavioural propensity still require measurement models. Existing insurance and road safety studies show that these intermediate signals add predictive resolution beyond conventional rating factors \cite{verbelen2018telematics,baecke2017telematics,henckaerts2022dynamic,guillen2019rates,ma2018context,guillen2024weekly,boylan2024review,guillen2020nearmiss,jiang2024hmm,jackson2015hmm}.

The observation layer also defines a direct validation target. Among policy records with the same conventional $X$, measured telematics-state distributions can be compared with those implied by $q_W$ and the opportunity-level kernels. Disagreement would challenge the bridge even if the annual claim model remained predictive. Linked crash--claim data would test the downstream mapping by informing $\rho_{it}$ in Eq.~\eqref{eq:claimobservation}. Neither validation is carried out here.

\section{Discussion}
\label{sec:discussion}

Annual insurance data make it easy to conflate two questions. A fitted tariff describes how recorded claim frequency varies with rating information. A mechanistic explanation asks which exposure, context, behaviour and vehicle pathways generate a crash contrast. The two coincide only under assumptions linking time scale, mechanism and claim observation.

The age application makes the distinction concrete. Vehicle and geographic adjustment leaves the 18--20 versus 40--49 claim contrast slightly larger than the raw contrast, whereas medium-resolution bonus--malus conditioning reduces it from 3.388 to 1.235. That change does not estimate an effect of claims history or experience and does not represent deconfounding. Bonus--malus is downstream of prior responsible claims and insurance duration; conditioning on it changes the predictive target, blocks parts of the historical pathway and may induce collider-type associations among its causes. For that reason the external calibration starts from the vehicle/geographic contrast, not from the smaller bonus--malus-adjusted value.

Several limitations are structural. The DAG compresses within-trip dynamics into a time-ordered snapshot and assumes no unmodelled common causes for the d-separation claims it uses. Missed common causes or a poorly chosen opportunity unit can invalidate those restrictions; graph sensitivity to added arrows addresses only part of that problem. The target mechanism is crash occurrence, whereas supporting studies also analyse involvement, culpability, injury or fatal crashes. Those outcomes require transportability assumptions, and responsibility analyses can be selected by crash severity \cite{dufournet2016selection}. The French portfolio adds another mismatch because it observes policy-level liability claims rather than verified unique drivers or crash events. A policy can cover exposure generated by more than one driver, and the data do not identify the reporting, coverage and responsibility process represented by $\rho_{it}$.

The evidence map has a separate limitation. Extraction and adjudication were conducted by one author without an independent second coding pass. The fixed decision sequence and public study-level rationale make the choices inspectable but do not remove judgement. The search is targeted, not exhaustive. The 72 records provide a traceable basis for the proposed graph; they are not a complete sample of the road safety literature or a source of pooled graph-wide uncertainty.

The set-valued analysis is useful as a diagnostic. The age example excludes some coarse decompositions under tight transport assumptions, yet the remaining region is wide and expands mechanically as the discrepancy bound is relaxed. The width is not presented as a substantive discovery. It marks the missing information: mechanism-specific states, linked crash--claim outcomes and external estimates closer to the target population and outcome.

\section{Conclusion}
\label{sec:conclusion}

Annual insurance claims and within-trip crash mechanisms observe different parts of the same risk process. The proposed evidence-informed DAG represents crash occurrence, the observation layer links conventional rating information to latent states, and the claim-observation layer separates crash occurrence from the recorded insurance outcome.

An annual rating relativity consequently supports only limited mechanistic inference. The mileage example constrains aggregate exposure without identifying its composition. The age example shows that an endogenous history variable can change the predictive contrast substantially before causal interpretation begins, while one transported road safety estimate narrows only a coarse compatibility region. The constructive implication is specific. A sharper analysis would need trip- or segment-level exposure and context, measurements of transient state or conflict close to the crash, and linkage from the crash event to responsibility and the subsequent insurance claim. Those measurements would make graph-specific restrictions testable, allow the crash-to-claim mapping to be estimated rather than absorbed into sensitivity, and provide a basis for validating or revising the proposed DAG.