The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
87,962 characters
Axiomatizing Local Asymptotic Minimax Risk
\maketitle
\begin{abstract}
Local asymptotic minimax (LAM) risk is a foundational efficiency criterion in statistics and econometrics. The literature uses two definitions of LAM risk: one which appears in classical lower bounds and another which appears in arguments establishing attainment of those bounds. Conventional efficiency arguments are consistent with \emph{any} estimator-dependent weighted average of the two, and consequently do not reveal which of these \emph{generalized $\alpha$-LAM} risk indices describes researchers' actual preferences. We take a decision-theoretic approach to systematically resolve this ambiguity. We axiomatically characterize the set of preferences with a generalized $\alpha$-LAM representation, and we document that such preferences may violate basic rationality requirements such as Monotonicity. Motivated by this, we axiomatically characterize the subset with constant-weight representations. Within this class, a \emph{Sample Uncertainty Aversion} axiom uniquely selects attainment LAM. We argue that Sample Uncertainty Aversion is normatively appealing, and we therefore recommend attainment LAM as the functional form for LAM risk.
\end{abstract}
\clearpage
\section{Introduction}
\label{sec:intro}
The \emph{local asymptotic minimax} (LAM) risk criterion is a foundational notion of efficiency in parametric and semiparametric statistics \citep{hajek1972local, bickel1993efficient, van2000asymptotic}. Econometricians use it to evaluate efficient estimation in structural models \citep{hirano2003asymptotic}, efficient estimation of non-smooth functionals \citep{song2014local,fang2014optimal}, efficient treatment assignment rules \citep{hirano2009asymptotics}, and efficient policies in sequential decision problems \citep{hirano2023asymptotic,adusumilli2025risk}. Despite this central role, existing motivations for adopting LAM are heuristic. A standard motivation is that localizing the parameter at the rate at which sampling uncertainty vanishes ensures that the statistical difficulty of the estimation problem remains non-negligible in large samples, thereby yielding more useful comparisons than fixed-parameter asymptotics.\footnote{See Sections 7.3 and 8.3 of \cite{van2000asymptotic} and Section 2.5.2 of \cite{hirano2020asymptotic}. LAM also reveals superefficient procedures' poor performance near points of superefficiency (e.g., Hodges' estimator in \citealt{lehmann1998theory} pg. 442 and the ``oracle property'' in \citealt{leeb2008sparse}).} Although this intuitively motivates a \emph{local asymptotic} perspective towards optimality, it does not uniquely recommend the LAM criterion, nor does it specify a functional form for LAM risk. This note aims to fill these gaps using a decision-theoretic approach.
A key starting point for our analysis is that, while finite-sample optimality criteria evaluate estimators, asymptotic optimality criteria evaluate \emph{estimator sequences} indexed by sample size. Consequently, such criteria necessarily take a stance on how to trade off risk across sample sizes, even though the researcher's actual estimation problem features a fixed sample size. This is a convenient fiction for asymptotic statistics, but it poses a challenge for understanding the LAM criterion, since existing motivations do not fully specify how to aggregate risk across sample sizes.\footnote{They also do not explain taking the \emph{worst-case} risk over local parameters for a fixed sample size. A naive motivation is to view LAM as a limiting analog of finite-sample minimaxity and point to previous decision-theoretic work studying the latter \citep{stoye2012new,gilboa1989maxmin}. However, ``[LAM] optimality$\ldots$ is conceptually different from finite sample minmaxity, and neither$\ldots$ implies the other'' (\citealt{hirano2020asymptotic} pg. 311). Our main axiomatic results make this ``conceptual difference'' precise.} Indeed, the literature uses two functional forms for LAM risk: classical lower bounds \citep{hajek1972local,le1979theorem} apply to the \emph{limiting best-case} local risk over sample size subsequences ($\liminf$), while subsequent arguments which establish attainment of these bounds \citep{van2002semiparametric,song2014local} use the corresponding \emph{limiting worst-case} ($\limsup$). Observation \ref{obs:gen-alpha-LAM-iff-obv-cons} shows that standard efficiency arguments are \emph{transparently consistent} with any estimator-dependent weighted average of these two endpoints. It is therefore unclear which functional form for LAM risk from this class of \emph{generalized $\alpha$-LAM} risk indices (if any) describes researchers' actual preferences.
Proposition~\ref{prop:stable-local-risks} and Corollary~\ref{cor:regular-stable-losses} provide one reason why this issue may not have been formally addressed: all generalized $\alpha$-LAM risk indices coincide for estimator sequences with limiting local estimation error laws and asymptotically uniformly integrable local losses, which are standard regularity conditions. Consequently, observing the researcher's rankings of such estimator sequences cannot distinguish among this class of functional forms. Nevertheless, we exhibit a simple estimator sequence (Example \ref{ex:gaussian}) which does not satisfy these conditions, such that the difference among generalized $\alpha$-LAM risk indices is maximally stark. Given that a commonly stated appeal of the LAM criterion is that it evaluates \emph{any} estimator sequence,\footnote{See, e.g., \cite{van2000asymptotic} Section 8.7 and \cite{van2002semiparametric} pg. 348.} our research question remains relevant: which index describes researchers' preferences?
We use the toolkit of microeconomic decision theory to systematically answer this question. For each class of LAM risk indices of interest, we offer a set of conditions (called \emph{axioms}) that are necessary and sufficient for the researcher's rankings to be consistent with some index in the class. Our axioms specify how the researcher aggregates risk across sample sizes and local parameters. Since the researcher ranks \emph{sequences} indexed by sample size, we interpret the decision environment as one in which the researcher forms their rankings \emph{before} learning the sample size of their dataset, and thus faces \emph{ex-ante uncertainty} about the sample size (as well as the local parameter). This is the decision-theoretic interpretation of the convenient fiction described above that is inherent in any asymptotic optimality criterion.
With this interpretation in mind, we briefly summarize and interpret our main results. As an initial benchmark, Theorem~\ref{thm-main-result} axiomatically characterizes the class of generalized $\alpha$-LAM preferences, thereby providing a complete understanding of the behavioral content of merely being (transparently) consistent with standard efficiency arguments. With this result in hand, we document that generalized $\alpha$-LAM preferences are sufficiently broad to allow for violations of basic rationality requirements like \emph{Monotonicity}. Our examples suggest that these violations are due to unrestricted dependence of the weights on the procedure being evaluated.
To address this limitation, we focus on the subset of \emph{$\alpha$-LAM} preferences represented by a \emph{fixed constant}-weighted average of classical and attainment LAM. Theorem~\ref{thm:fixed-alpha-LAM} restores the basic rationality restrictions, in addition to other axioms, to axiomatically characterize $\alpha$-LAM. Within this subclass, axioms governing the researcher's attitudes towards \emph{sample size uncertainty} play a key role: \emph{Sample Uncertainty Aversion} (which requires a \emph{preference for hedging}) uniquely selects attainment LAM, while \emph{Sample Uncertainty Affinity} (which requires an \emph{aversion to hedging}) uniquely selects classical LAM (Corollary \ref{cor:CLAM-and-ALAM}). This lays bare the difference between risk indices in terms of a concrete and intuitive consequence for researcher behavior.
Finally, we argue that Sample Uncertainty Aversion is normatively appealing. Consequently, out of the candidate LAM risk representations we study, we recommend attainment LAM. Our axiomatic characterization of attainment LAM therefore complements and refines the heuristic motivations for the LAM criterion described above.
\paragraph{Related literature and key contributions.}
Providing decision-theoretic foundations for notions of statistical optimality has a long tradition \citep{wald1950statistical,hurwicz1951optimality,savage1954foundations}. More recent work has studied finite-sample statistical optimality criteria from an axiomatic perspective. \cite{stoye2011statistical} draws upon previous microeconomic decision theory work to characterize Bayes, minimaxity, minimax regret, and admissibility, as well as the \emph{Hurwicz} or \emph{$\alpha$-maxmin} criterion which mixes between the best- and worst-case according to a fixed weight, while \cite{stoye2012new} studies these and other finite-sample criteria in a decision environment consisting of risk functions. \cite{andrews2026misspecification} provide axioms which characterize various classes of misspecification-averse criteria used in the literature; we adapt some of their technical tools to elicit and aggregate conditional preferences. In contrast to the finite-sample axiomatic analysis of these papers, our focus is on \emph{local asymptotic} optimality criteria. Rather than studying an established class of representation functional forms, we begin with the largest class that is in a precise sense transparently consistent with standard efficiency arguments in econometrics, and we use our axiomatic analysis to make a principled recommendation for a functional form for LAM risk.
Many of our axioms and results are adapted from or build upon related axioms and representation results in the microeconomic decision theory literature. Key building blocks include the Consistency and Caution axioms \citep{gilboa2010objective,cerreia2024making}, maxmin expected utility \citep{gilboa1989maxmin}, and the complete patience representation studied in \cite{marinacci1998axiomatic} Theorem 7. The axiomatic results in this note contribute to the literature which studies $\alpha$-maxmin and generalized $\alpha$-maxmin representations \citep{arrow1972optimality,ghirardato2004differentiating,cerreia2011rational}. In particular, \cite{chateauneuf2007choice} study a class of representations which mixes between expected utility, the best-case, and the worst-case. The special case of their subclass which puts zero weight on expected utility (their \emph{Hurwicz capacity}) is analogous to our $\alpha$-LAM preferences restricted to a particular subdomain of our choice environment, although they do not axiomatize this case. \cite{drugeon2023alpha} study a class of asymptotic $\alpha$-maxmin criteria in a choice environment of intertemporal utility streams. Their analysis takes as primitive two preference relations, which they interpret as the preferences of society for the near and far future.
Relative to these papers, we view our key contributions as follows. First, we define and axiomatically characterize a collection of asymptotic Hurwicz-type criteria whose functional forms are tightly connected to standard efficiency arguments in statistics and econometrics, and we use our axioms to provide a unique recommendation. Given that asymptotic efficiency is the benchmark notion of optimality among practitioners in these fields, we view this as a compelling use case of the microeconomic decision theory toolkit. Second, our axioms make transparent the role that attitudes towards sample size risk aggregation play in adjudicating among different candidate LAM risk indices. On a technical level, we streamline the axiomatic analysis by taking only one preference as primitive---the researcher's actual ranking of procedures---and elicit other preferences of interest from the primitive preference. The multidimensional nature of uncertainty in our setting (in which outcomes jointly depend on the sample size and local parameter) requires specialized techniques which we hope are useful for future work in decision-making under multidimensional uncertainty.
\paragraph{Roadmap.}
The rest of this note proceeds as follows. Section~\ref{sec:setup} sets up the standard parametric point estimation environment and shows that conventional efficiency arguments do not pin down a functional form for LAM risk. Section~\ref{sec:axioms} states and interprets our axioms, characterizes generalized $\alpha$-LAM and $\alpha$-LAM preferences, and then characterizes classical and attainment LAM. Appendix~\ref{app:LAN} provides a brief exposition of local asymptotic normality for non-familiar readers. Appendix~\ref{app-sec-2-proofs} proves the results in Section~\ref{sec:setup}. Appendix~\ref{app:AA} develops axiomatic foundations for the main text's decision environment in a standard Anscombe--Aumann setup. Appendix~\ref{app:util-act-lemmas} proves the axiomatic representation results in Section~\ref{sec:axioms}.
\section{Functional forms for LAM risk}
\label{sec:setup}
\paragraph{Statistical decision environment.} We adapt the standard parametric point estimation setup of \cite{van2000asymptotic} Chapter 8. Let $\Theta\subseteq \mathbb{R}^k$ be an open set of parameters, and let $\psi: \Theta \rightarrow \mathbb{R}^m$ be the estimand of interest. For each sample indexed $n\geq 1$, the researcher observes data $X_n \in \mathcal{X}_n$ drawn from a distribution in the model $\mathcal{P}_n=\{P_{\theta,n} : \theta \in \Theta\} \subseteq \Delta(\mathcal{X}_n)$,\footnote{Assume that $\Theta$ and each $\mathcal{X}_n$ are Polish spaces endowed with their Borel $\sigma$-algebras. Endow $\mathbb{N}$ with the power set $\sigma$-algebra, and endow product spaces with their product $\sigma$-algebras.} chooses an estimate $a \in \mathbb{R}^m$, and incurs loss $\ell_n(a-\psi(\theta))$. An \emph{estimator} for the sample $n$ is a measurable function $\delta_n : \mathcal{X}_n \times [0,1] \to \mathbb{R}^m$.\footnote{To allow for estimators which randomize conditional on the sample, we let $\delta(X_n,\cdot)$ additionally depend on the realization of an auxillary independent $\text{Unif}[0,1]$ random variable.} An \emph{estimator sequence} is a sequence of estimators $\delta = (\delta_n)_{n \geq 1}$. Let $\mathcal{D}$ be the set of estimator sequences. We begin by stating three standard assumptions on the statistical primitives which collect regularity conditions on the sequence of loss functions, the sequence of models, and the estimand of interest.
\begin{assumption}\label{as:1-loss}
$\ell_n(e)=\ell(\sqrt{n}e)$ for some Borel measurable loss function $\ell: \mathbb{R}^m \rightarrow \mathbb{R}_+$ whose sublevel sets are convex and symmetric about the origin.
\end{assumption}
Assumption \ref{as:1-loss} ensures nontrivial comparisons between $\sqrt{n}$-consistent estimator sequences, which is a standard property in parametric settings. Throughout the main text, unless otherwise stated, fix a centering value $\theta_0 \in \Theta$. Let $\overset{\theta_0}{\rightsquigarrow}$ denote convergence in distribution under the sequence of laws $P_{\theta_0,n}$. Let $\overset{\theta_0,n}{\sim}$ denote distribution under $P_{\theta_0,n}$. Say that a sequence of random variables is $o_{P_{\theta_0,n}}(1)$ if it converges in probability to $0$ under the sequence of laws $P_{\theta_0,n}$.
\begin{assumption}[\citealt{van2000asymptotic} Definition 7.14]\label{as:1}
$(\mathcal{P}_n)_n$ is \emph{locally asymptotically normal (LAN)} at $\theta_0$: there exists an invertible matrix $I_{\theta_0} \in \mathbb{R}^{k\times k}$ and random vectors $\Delta_{\theta_0,n} \overset{\theta_0}{\rightsquigarrow} N(0,I_{\theta_0})$ such that, for each converging sequence $h_n \to h$,
\begin{equation}\label{eq:LAN}
\log \frac{dP_{\theta_0+h_n/\sqrt{n},n}}{dP_{\theta_0,n}}=h'\Delta_{\theta_0,n}-\frac{1}{2}h'I_{\theta_0}h+o_{P_{\theta_0,n}}(1) \tag{LAN}
\end{equation}
\end{assumption}
To interpret Assumption \ref{as:1}, we consider the following sequence of \emph{local informational environments} at the fixed centering value $\theta_0 \in \Theta$. Suppose that the researcher conditions on the value $\theta_0$ and instead faces uncertainty about the value of the \emph{local parameter} $h\in H=\mathbb{R}^k$,\footnote{This interpretation is often motivated by the observation that ``asymptotically, [$\theta_0$] can be known with unlimited precision," whereas the ``true statistical difficulty is$\ldots$ determined by the nature of the measures'' local to $P_{\theta_0}$ (\citealt{van2000asymptotic} Section 7.3).} where for each $n\geq 1$, $h \in H$ parametrizes the $\sqrt{n}$-local value $\theta_0+h/\sqrt{n} \in \Theta$.\footnote{Here and throughout, whenever $\theta_0+h/\sqrt{n} \notin \Theta$, set it equal to an arbitrary value in $\Theta$.} We reinterpret the data $X_n \in \mathcal{X}_n$ as arising from the \emph{local experiment} $\mathcal{P}_{\theta_0,n,H}=\{P_{\theta_0+h/\sqrt{n},n}: h\in H\}$ which is informative about the unknown local parameter $h$.\footnote{Since we have fixed the centering value $\theta_0$, $P_{\theta_0+h/\sqrt{n},n}$ is fully determined up to the local parameter $h$.} Note that as $n$ increases, two forces are at play: the informativeness of the experiment $\{P_{\theta,n}: \theta \in \Theta\}$ changes, and each parameter $\theta_0+h/\sqrt{n}$ shrinks towards $\theta_0$. \eqref{eq:LAN} requires that these two forces balance each other out, such that the sequence of local experiments $\mathcal{P}_{\theta_0,n,H}$ converges in an informational sense to the \emph{limit experiment} $\mathcal{G}_{\theta_0,H}=\{N(h,I_{\theta_0}^{-1}): h\in H\}$. For readers who are less familiar with local asymptotics, Appendix \ref{app:LAN} provides a more detailed exposition of LAN.
A leading setting satisfying Assumption \ref{as:1} is i.i.d. sampling from a smooth parametric model (\citealt{van2000asymptotic} Chapter 7).\footnote{In this case, the \emph{Fisher information} of the model $\mathcal{P}_1$ at $\theta_0$ provides the matrix $I_{\theta_0}$.} With this case in mind, we henceforth refer to the index $n\geq 1$ as the \emph{sample size}, with the understanding that our results nevertheless apply to more general statistical environments which satisfy Assumption \ref{as:1}. We conclude this subsection with a smoothness condition on the estimand of interest.
\begin{assumption}\label{as:diff-estimand}
$\psi$ is differentiable at $\theta_0$ with derivative $\dot{\psi}_{\theta_0}: \mathbb{R}^k \rightarrow \mathbb{R}^m$.
\end{assumption}
\paragraph{Local risk function sequences.} To interpret Assumption \ref{as:diff-estimand} and future results, we introduce a unifying expositional framework for the note. We consider a sequence of \emph{local decision problems} at the fixed centering value $\theta_0$ in which the researcher conditions on the value $\theta_0$ and, for each sample size $n\geq 1$, observes data $X_n$ from the local experiment $\mathcal{P}_{\theta_0,n,H}$ about the unknown local parameter $h$, chooses an estimate $a_{\text{local}} \in \mathbb{R}^m$, and incurs loss $\ell(a_{\text{local}}-\dot{\psi}_{\theta_0,n}(h))$, where $\dot{\psi}_{\theta_0,n}: H \rightarrow \mathbb{R}^m$ is the \emph{local estimand of interest} at $\theta_0$ defined as $\dot{\psi}_{\theta_0,n}(h)=\sqrt{n}\left(\psi\left(\theta_0+h/\sqrt{n}\right)-\psi(\theta_0) \right)$. Note that $\dot{\psi}_{\theta_0,n}$ is a finite-step size approximation to the derivative $\dot{\psi}_{\theta_0}$. In this sense, Assumption \ref{as:diff-estimand} ensures the existence of a \emph{limiting} local estimand of interest.
It is convenient to summarize the performance of an estimator sequence $\delta$ in the sequence of local decision problems described above. To this end, for an estimator $\delta_n$, define its \emph{local version at $\theta_0$} to be the function $\delta_{\theta_0,n}: \mathcal{X}_n \rightarrow \Delta(\mathbb{R}^m)$ defined as: $\delta_{\theta_0,n}(X_n)=\sqrt{n}(\delta_n(X_n)-\psi(\theta_0))$. We allow $\delta_{\theta_0,n}$ to depend on $\theta_0$ since the researcher has already conditioned on the value of $\theta_0$. Loosely speaking, $\delta_{\theta_0,n}$ translates $\delta_n=\psi(\theta_0)+\delta_{\theta_0,n}/\sqrt{n}$ ``into local coordinates.'' Let $P_{\theta_0,n,h}=P_{\theta_0+h/\sqrt{n},n}$. Let $E_{\theta_0,n,h}$ denote expectation under $P_{\theta_0,n,h}$,\footnote{More precisely for randomized estimators, under $P_{\theta_0,n,h} \times \mathcal{L}(\text{Unif}[0,1])$, where $\mathcal{L}(\text{Unif}[0,1])$ is the law of an auxillary $\text{Unif}[0,1]$ random variable.} and let $\overset{\theta_0,n,h}{\sim}$ denote distribution under $P_{\theta_0,n,h}$. Let $\overline{\mathbb{R}}_+=[0,+\infty]$.
\begin{definition}\label{defn:lrfs}
The \emph{local risk function sequence of $\delta$ at $\theta_0$} is the function $R_{\theta_0,\delta}: \mathbb{N} \times H \rightarrow \overline{\mathbb{R}}_+$ defined as:
\[
R_{\theta_0,\delta}(n,h)=E_{\theta_0,n,h}\big[\ell(\delta_{\theta_0,n}(X_n)-\dot{\psi}_{\theta_0,n}(h)) \big]
\]
\end{definition}
In words, $R_{\theta_0,\delta}(n,h)$ is the expected loss (\emph{risk}) of using the local estimator $\delta_{\theta_0,n}$ to estimate the local estimand $\dot{\psi}_{\theta_0,n}(h)$ when the data is drawn according to $P_{\theta_0,n,h}$.\footnote{By substituting definitions, we may write $R_{\theta_0,\delta}(n,h)=E_{\theta_0,n,h}[\ell(\sqrt{n}(\delta_n(X_n)-\psi(\theta_0+h/\sqrt{n})))]$. This version of the expression appears in standard references (\citealt{van2000asymptotic} Ch. 8).} In this sense, $R_{\theta_0,\delta}$ encodes the performance of $\delta$ in the sequence of local decision problems at $\theta_0$, for each sample size $n$ and local parameter $h$. Say that a local risk function sequence \emph{$R$ is induced at $\theta_0$ by $\delta$} if $R=R_{\theta_0,\delta}$.
\begin{example}
\label{ex:gaussian}
\normalfont We consider as a running example the problem of estimating the mean of a one-dimensional Gaussian shift model under squared error loss with i.i.d. data. In our notation, let $\Theta=\mathbb{R}$, $\psi(\theta)=\theta$, $\mathcal{X}_n=\mathbb{R}$ for each $n\geq 1$, $P_{\theta,n}=N(\theta,1)^{\otimes n}$, and $\ell_n(x) = nx^2$. Assumption \ref{as:1-loss} follows since squared-error loss $\ell(x)=x^2$ satisfies the stated conditions. Fix any centering value $\theta_0 \in \mathbb{R}$, and note that the Fisher information at $\theta_0$ is $I_{\theta_0}=1$. A straightforward computation yields:
\[
\log \frac{dP_{\theta_0+h_n/\sqrt{n},n}}{dP_{\theta_0,n}}(X_n)=h_n \sqrt{n}(\Bar{X}_n-\theta_0)-\frac{h_n^2}{2}
\]
where $\Bar{X}_n$ is the sample mean and $\sqrt{n}(\Bar{X}_n-\theta_0) \overset{\theta_0,n}{\sim} N(0,1)$. Since the sequence of random variables $(h_n-h)\sqrt{n}(\Bar X_n-\theta_0)+(h^2-h_n^2)/2$ is $o_{P_{\theta_0,n}}(1)$, Assumption \ref{as:1} follows. Assumption \ref{as:diff-estimand} holds with $\dot{\psi}_{\theta_0}(h)=h$.
The sample mean induces the constant function $R_{\theta_0,\Bar{X}}=1$.
\end{example}
\paragraph{Classical LAM risk.} The literature contains two key functional forms of LAM risk, both of which evaluate estimator sequences by their local risk function sequences. We refer to the first one as \emph{classical} LAM risk since it appears in the classical LAM lower bounds developed by \cite{hajek1972local} and \cite{le1979theorem}. Let $\mathcal{I}$ be the set of nonempty finite subsets of $H$.
\begin{definition}\label{def:LAM-estimator}
The \emph{classical LAM risk} of estimator sequence $\delta$ at $\theta_0$ is:
\begin{equation}\label{eq:LAM-risk}
\underline{V}_{\mathrm{LAM},\theta_0}(\delta) := \sup_{I \in \mathcal{I}} \liminf_{n \to \infty} \sup_{h \in I} R_{\theta_0,\delta}(n,h) \tag{CLAM}
\end{equation}
\end{definition}
In words: for each \emph{local neighborhood} $I$ and sample size $n$, $\delta_n$ is evaluated by the \emph{worst-case} risk in the $n$-th local decision problem at $\theta_0$ over local parameters $h\in I$, the resulting sequence of worst-case risks is aggregated by its \emph{limiting best-case} across sample size subsequences, and finally the \emph{worst-case} is taken over all local neighborhoods $I$. The \emph{local asymptotic minimax theorem} (\citealt{hajek1972local,le1979theorem}; \citealt{van2000asymptotic} Theorem~8.11) lower bounds the classical LAM risk of \emph{every} estimator sequence $\delta \in \mathcal{D}$ at $\theta_0$:
\begin{equation}
\label{eq:LAM thm}
\underline{V}_{\mathrm{LAM},\theta_0}(\delta) \geq B_{\mathrm{LAM},\theta_0}:=\int \ell \, dN(0, \dot{\psi}_{\theta_0}I_{\theta_0}^{-1}\dot{\psi}_{\theta_0}') \tag{LAM bound}
\end{equation}
Note that $B_{\mathrm{LAM},\theta_0}$ (or \emph{LAM bound}) is the optimal minimax risk for estimating $\dot{\psi}_{\theta_0}(h)$ with data from the limit experiment $\mathcal{G}_{\theta_0,H}$. In this sense, \eqref{eq:LAM thm} establishes that the sequence of local decision problems at $\theta_0$ is \emph{weakly more difficult than} this limit decision problem at $\theta_0$.
\paragraph{Attainment LAM risk.}
Standard arguments in statistics and econometrics show that a proposed estimator sequence attains the LAM bound and conclude that it is efficient \citep{chamberlain1992efficiency,hirano2003asymptotic,graham2011efficiency,adusumilli2025risk}. Many of these arguments (e.g., pg. 348 of \citealt{van2002semiparametric} and Theorem~4 of \citealt{song2014local}) use an alternative risk index.
\begin{definition}
The \emph{attainment LAM risk} of estimator sequence $\delta$ at $\theta_0$ is:
\begin{equation}
\label{eq:LAM-limsup-estimator}
\overline{V}_{\mathrm{LAM},\theta_0}(\delta):=\sup_{I \in \mathcal{I}} \limsup_{n \to \infty} \sup_{h \in I} R_{\theta_0,\delta}(n,h) \tag{ALAM}
\end{equation}
\end{definition}
The interpretation of attainment LAM risk is exactly analogous with that of classical LAM risk, except that each sequence of worst-case risks is aggregated by its \emph{limiting worst-case} across sample size subsequences. Note that we may also write
\[
\overline{V}_{\mathrm{LAM},\theta_0}(\delta)=\sup_{h\in H} \limsup_{n\to\infty} R_{\theta_0,\delta}(n,h)
\]
so attainment LAM risk is also the \emph{worst-case limiting} risk taken over all local parameters $h\in H$ and all sample size subsequences. Following the above references, say that an estimator sequence $\delta$ \emph{attains the LAM bound at $\theta_0$} if $\overline{V}_{\mathrm{LAM},\theta_0}(\delta)\leq B_{\mathrm{LAM},\theta_0}$.
\addtocounter{example}{-1}
\begin{example}[Continued]
\label{ex:CLAM-ALAM-bounds}
\normalfont Fix any centering value $\theta_0\in \mathbb{R}$. By our previous computations, the LAM bound is $B_{\mathrm{LAM},\theta_0}=\int x^2 \, dN(0,1) = 1$. Since the local risk function sequence $R_{\theta_0,\Bar{X}}=1$ is constant, both risk indices coincide: $\overline{V}_{\mathrm{LAM},\theta_0}(\Bar{X})=\underline{V}_{\mathrm{LAM},\theta_0}(\Bar{X})=1$. Hence, the sample mean attains the LAM bound at $\theta_0$.\footnote{More generally, under mild conditions the \emph{maximum likelihood estimator} attains the LAM bound at each $\theta_0\in \Theta$. See, e.g., \cite{hirano2020asymptotic} page 311.}
\end{example}
\paragraph{Researcher preferences.} The appearance of classical and attainment LAM risk in the literature raises a natural question: when a researcher professes to apply the LAM criterion, what risk index describes their actual preferences? Although the motivation described in Section \ref{sec:intro} of normalizing the difficulty of the estimation problem to be the same order as sampling-based uncertainty suggests a local asymptotic approach in a broad sense, it is too coarse to distinguish between these two functional forms for LAM risk. Indeed, in this section we demonstrate that conventional efficiency arguments are consistent with a \emph{much larger class} of LAM-based risk indices.
To make this claim precise, recall that standard efficiency arguments conclude that an estimator sequence $\delta^*$ is optimal whenever it attains the LAM bound. Hence, say that a risk index $V_{\theta_0}: \mathcal{D} \rightarrow \overline{\mathbb{R}}_+$ is \emph{consistent} with such arguments at $\theta_0$ if, for each $\delta^* \in \mathcal{D}$,
\begin{equation}
\label{eq:cons}
\delta^* \text{ attains the LAM bound at } \theta_0 \implies V_{\theta_0}(\delta)\geq V_{\theta_0}(\delta^*) \quad \forall \delta\in\mathcal{D} \tag{Cons.}
\end{equation}
The set of consistent risk indices is the largest set of functional forms for risk which does not contradict the property that attaining the LAM bound is a sufficient condition for global optimality. However, \eqref{eq:cons} by itself is not an appealing desideratum, since it places no restrictions on how $V_{\theta_0}$ ranks estimator sequences which do not attain the LAM bound. This means that the revealed preference exercise embodied by \eqref{eq:cons} has little bite, allowing for, e.g., the risk index which assigns $1$ to any $\delta^*$ which attains the bound and $+\infty$ to anything else. We therefore focus on the following subset of \emph{transparently consistent} risk indices.
\begin{definition}
\label{defn:obv-cons}
A risk index $V_{\theta_0}: \mathcal{D} \rightarrow \overline{\mathbb{R}}_+$ is \emph{transparently consistent} with standard efficiency arguments at $\theta_0$ if
\begin{equation}
\label{eq:obv-cons}
\underline{V}_{\mathrm{LAM},\theta_0}(\delta)\leq V_{\theta_0}(\delta)\leq \overline{V}_{\mathrm{LAM},\theta_0}(\delta) \quad \forall \delta \in \mathcal{D} \tag{Trans. Cons.}
\end{equation}
\end{definition}
To justify this terminology, note that if $V_{\theta_0}$ satisfies \eqref{eq:obv-cons}, then verifying consistency is immediate, since $\delta^*$ attaining the LAM bound at $\theta_0$ implies
\[
\underline{V}_{\mathrm{LAM},\theta_0}(\delta) \geq \overline{V}_{\mathrm{LAM},\theta_0}(\delta^*) \quad\forall \delta\in \mathcal{D}
\]
by \eqref{eq:LAM thm}, which when combined with \eqref{eq:obv-cons} immediately yields \eqref{eq:cons}. Let $\mathcal{V}_{\theta_0}$ be the set of transparently consistent risk indices at $\theta_0$. In sum, observing that the researcher weakly prefers an estimator sequence which attains the LAM bound to every alternative does not actually specify a functional form for the LAM criterion, since such efficiency arguments are \emph{transparently consistent} with every risk index $V \in \mathcal{V}_{\theta_0}$. Our first results explain why this ambiguity may not have been formally addressed: on the choice domain of estimator sequences whose local risk function sequences converge pointwise, every transparently consistent risk index coincides.
\begin{proposition}
\label{prop:stable-local-risks}
Fix any estimator sequence $\delta$. Suppose there exists a \emph{pointwise limiting local risk function} $r_{\theta_0,\delta}: H\rightarrow \overline{\mathbb{R}}_+$ such that
\begin{equation}
\label{eq:stable-local-risks}
R_{\theta_0,\delta}(n,h)
\to r_{\theta_0,\delta}(h)
\qquad\text{for each }h\in H
\end{equation}
Then $V_{\theta_0}(\delta)=V_{\theta_0}'(\delta)$ for all $V_{\theta_0},V_{\theta_0}'\in \mathcal{V}_{\theta_0}$.
\end{proposition}
Two transparently consistent risk indices can differ only if there exists some local parameter where local risk fails to converge across sample sizes. Our next result shows that estimator sequences whose local estimation errors possess limit laws and whose local losses are asymptotically uniformly integrable satisfy the pointwise convergence of local risk functions in \eqref{eq:stable-local-risks}. Define the \emph{local estimation error} of estimator sequence $\delta$ at sample size $n\geq 1$ and local parameter $h\in H$ to be the random vector:
\[
Z_{\theta_0,\delta,n}(h):=\delta_{\theta_0,n}(X_n)-\dot{\psi}_{\theta_0,n}(h)
\]
Recall that this is the random estimation error which arises from using the local version of $\delta_n$ at $\theta_0$ to estimate the local estimand at $\theta_0$.
\begin{corollary}
\label{cor:regular-stable-losses}
Fix any estimator sequence $\delta$. Suppose that, for each $h\in H$, there exists a random vector $Z_{\theta_0,\delta}(h)$ such that $Z_{\theta_0,\delta,n}(h) \overset{\theta_0,n,h}{\rightsquigarrow} Z_{\theta_0,\delta}(h)$, $\ell$ is a.s. continuous under the law of $Z_{\theta_0,\delta}(h)$, and the sequence of local losses $\{\ell(Z_{\theta_0,\delta,n}(h))\}_{n\geq 1}$ under the sequence of laws $P_{\theta_0,n,h}$ is \emph{asymptotically uniformly integrable}:
\begin{equation}
\label{eq:local-loss-AUI}
\lim_{M\to\infty}\limsup_{n\to\infty}
E_{\theta_0,n,h}\!\big[
\ell\!\left(Z_{\theta_0,\delta,n}(h)\right)
\mathbf 1\!\left\{
\ell\!\left(Z_{\theta_0,\delta,n}(h)\right)>M
\right\}
\big]
=0
\end{equation}
Then $V_{\theta_0}(\delta)=V_{\theta_0}'(\delta)$ for all $V_{\theta_0},V_{\theta_0}'\in \mathcal{V}_{\theta_0}$.
\end{corollary}
Definition \ref{defn:obv-cons}, Proposition~\ref{prop:stable-local-risks} and Corollary~\ref{cor:regular-stable-losses} make precise the issue at the heart of this note: conventional efficiency arguments in econometrics (and more generally, researcher rankings on a class of procedures satisfying standard regularity conditions) do not distinguish among a class of risk indices bounded by classical and attainment LAM risk. This does not imply, however, that the distinction is not meaningful. While the regularity conditions in Corollary~\ref{cor:regular-stable-losses} are standard, simple examples violate them and make the choice of $V_{\theta_0}\in \mathcal{V}_{\theta_0}$ consequential. Furthermore, a stated appeal of the LAM criterion (as opposed to, e.g., convolution arguments which restrict attention to regular estimator sequences\footnote{See, e.g., \cite{van2000asymptotic} Theorem 8.8.}) is that it provides a \emph{universal} theory of statistical optimality which does not rely on restrictions to special classes of estimators, such as the ones studied in Proposition~\ref{prop:stable-local-risks} and Corollary~\ref{cor:regular-stable-losses}. It is therefore important for the LAM criterion to evaluate such simple examples. Indeed, the following example shows that the distinction can be maximally stark.
\addtocounter{example}{-1}
\begin{example}[Continued]
\label{ex:powers-of-10}
\normalfont For ease of exposition, fix the centering value $\theta_0=0$. For any arbitrarily large but finite constant $C>0$, consider the estimator sequence
\[
\delta_n(X_n)=\begin{cases}
\Bar{X}_n & n=10^k \text{ for some } k\in \mathbb{N} \\
\Bar{X}_n+C/\sqrt{n} & \text{ else}
\end{cases}
\]
which outputs the sample mean when the sample size is a power of 10 and otherwise outputs the sample mean with constant local bias $C$. By previous computations,
\[
R_{\theta_0,\delta}(n,h)=\begin{cases}
1 & n=10^k \text{ for some } k\in \mathbb{N} \\
1+C^2 & \text{ else}
\end{cases}
\]
Since $\delta$ achieves the same performance as the sample mean along the subsequence $n_k=10^k$ and incurs risk $1+C^2$ otherwise,
\[
1+C^2=\overline{V}_{\mathrm{LAM},\theta_0}(\delta)>\underline{V}_{\mathrm{LAM},\theta_0}(\delta)=B_{\mathrm{LAM},\theta_0}=1
\]
which implies that $\delta$ is a \emph{best} estimator sequence under classical LAM risk for any value of $C>0$, but an arbitrarily bad estimator sequence under attainment LAM risk for sufficiently large values of $C$. More generally, a transparently consistent risk index may assign $\delta$ any risk in $[1,1+C^2]$.
\end{example}
\paragraph{Generalized $\boldsymbol{\alpha}$-LAM risk.} We conclude this section with a useful reparametrization of the set of transparently consistent risk indices, in terms of a set of \emph{estimator sequence-dependent weights} on classical and attainment LAM risk.
\begin{definition}
\label{defn-gen-alpha-LAM-estimators}
A risk index $V_{\theta_0}: \mathcal{D} \rightarrow \overline{\mathbb{R}}_+$ is \emph{generalized $\alpha$-LAM} at $\theta_0$ if there exists a function $\alpha_{\theta_0}: \mathcal{D} \rightarrow [0,1]$ such that\footnote{We use the convention $0\cdot(+\infty)=0$, such that $\alpha_{\theta_0}=1$ selects classical LAM risk and $\alpha_{\theta_0}=0$ selects
attainment LAM risk, even when an endpoint is infinite.}
\begin{equation}
V_{\theta_0}(\delta)=V_{\mathrm{gen}\text{-}\alpha\text{-}\mathrm{LAM},\theta_0}(\delta):=\alpha_{\theta_0}(\delta) \underline{V}_{\mathrm{LAM},\theta_0}(\delta)+(1-\alpha_{\theta_0}(\delta))\overline{V}_{\mathrm{LAM},\theta_0}(\delta) \tag{gen $\alpha$-LAM}
\end{equation}
\end{definition}
The following observation verifies that Definitions \ref{defn:obv-cons} and \ref{defn-gen-alpha-LAM-estimators} are essentially equivalent.
\begin{observation}
\label{obs:gen-alpha-LAM-iff-obv-cons}
\begin{itemize}
\item[(i)] Every generalized $\alpha$-LAM risk index at $\theta_0$ is transparently consistent at $\theta_0$.
\item[(ii)] For every risk index $V_{\theta_0}$ which is transparently consistent at $\theta_0$, there exists $\alpha_{\theta_0}: \mathcal{D} \rightarrow [0,1]$ such that:
\[
V_{\theta_0}(\delta)=V_{\mathrm{gen}\text{-}\alpha\text{-}\mathrm{LAM},\theta_0}(\delta) \quad \forall \delta \text{ s.t. } \underline{V}_{\mathrm{LAM},\theta_0}(\delta)=+\infty \text{ or } \overline{V}_{\mathrm{LAM},\theta_0}(\delta)<+\infty
\]
\end{itemize}
\end{observation}
Hence, excluding estimator sequences which have finite classical LAM risk but infinite attainment LAM risk, a risk index is transparently consistent if and only if it is generalized $\alpha$-LAM. In particular, the equivalence holds over the set of estimator sequences which induce bounded local risk function sequences at $\theta_0$. Restating transparent consistency in the equivalent language of weighted averages of classical and attainment LAM risk is useful for our forthcoming decision theoretic-approach.
\section{Axiomatic results}
\label{sec:axioms}
\subsection{Roadmap of Section \ref{sec:axioms}}
In the context of Example~\ref{ex:powers-of-10}, it seems intuitive to consider $\delta$ unappealing, since it behaves well only when the sample size happens to be a power of 10. Consequently, it seems like classical LAM risk may be a poor fit for researcher preferences. To make this intuition precise, and to comprehensively understand which transparently consistent risk indices (if any) are a better fit, we take a decision-theoretic approach.
With the equivalence from Observation \ref{obs:gen-alpha-LAM-iff-obv-cons} in mind, we begin by seeking an axiomatic characterization of the set of researcher preferences with a generalized $\alpha$-LAM risk representation (Theorem \ref{thm-main-result}). We view this exercise as an initial benchmark for understanding which conditions on the researcher's rankings are necessary and sufficient to \emph{merely} be (transparently) consistent with standard efficiency arguments. Given the flexibility of the procedure-dependent weights $\alpha$, we find that this is not a particularly demanding restriction (although it does still carry behavioral content). In particular, we document that generalized $\alpha$-LAM preferences may violate basic rationality requirements, such as \emph{Monotonicity} and \emph{Mixture Continuity}. This suggests that requiring that $\alpha$ be a \emph{fixed} weight may be a more appealing fit, a class of risk representations we denote as $\alpha$-LAM.
Therefore, we next seek an axiomatic characterization of the set of $\alpha$-LAM preferences (Theorem \ref{thm:fixed-alpha-LAM}). Our axioms, which now include the basic rationality requirements mentioned above, determine how the researcher aggregates risk across sample sizes and local parameters. Within this class, two axioms specifying contrasting attitudes towards aggregating risk across sample sizes uniquely select classical LAM and attainment LAM. We argue that the latter axiom, which requires a \emph{preference for hedging}, is normatively appealing. We therefore recommend attainment LAM as the functional form for the LAM criterion.
\subsection{Shared axioms}
We begin by stating and interpreting several variants of axioms which fulfill a shared purpose for each of the axiomatic characterizations we pursue. The first axiom clarifies our focus on optimality criteria which evaluate estimator sequences by their performance \emph{local to $\theta_0$}. For ease of exposition, we state this axiom informally in the main text and defer the formal statement to Axiom \ref{ax:cond-risk-relevance} in Appendix \ref{app:util-act-lemmas}, where we additionally take as primitive the researcher's preference over estimator sequences conditional on $\theta_0$.
\begin{axiom}[Local Risk Sufficiency, informally stated]\label{as:2}
Each estimator sequence $\delta$ is evaluated by its induced local risk function sequence $R_{\theta_0,\delta}$.
\end{axiom}
We view (the formal statement of) Axiom \ref{as:2} as a precise statement of the heuristic motivation for local asymptotics given in Section~\ref{sec:intro}. Recall that $R_{\theta_0,\delta}$ encodes $\delta$'s performance in the sequence of local decision problems at $\theta_0$ defined in Section~\ref{sec:setup}, which localize the parameter of interest around $\theta_0$ at the same rate as sampling uncertainty vanishes. Axiom \ref{as:2} therefore ensures that performance in this sequence of decision problems is \emph{all that matters} for ranking estimator sequences. Recall that classical and attainment LAM risk both satisfy Axiom \ref{as:2}. More generally, this holds for every generalized $\alpha$-LAM risk index where $\alpha_{\theta_0}(\delta)$ depends on $\delta$ only through $R_{\theta_0,\delta}$. Given that such heuristic motivations for local asymptotics are standard, we view Axiom \ref{as:2} as a natural restriction for studying local asymptotic optimality.\footnote{At this stage, the distinction between the local risk function sequence induced by $\delta$ at $\theta_0$ and the risk function sequence induced by $\delta$ at $\theta_0$ is vacuous: one may be obtained from the other by the usual reparametrization defined in Section \ref{sec:setup}. In this sense, Local Risk Sufficiency is equivalent to Risk Sufficiency. However, the distinction has bite when paired with our forthcoming axioms.} However, it rules out criteria which evaluate estimator sequences by other desiderata beyond their performance local to $\theta_0$, such as simplicity or computational complexity.
\addtocounter{example}{-1}
\begin{example}[Continued]
\normalfont Axiom \ref{as:2} requires that $(\Bar{X}_n)_n$ is preferred to $\delta$ if and only if $R_{\theta_0,\Bar{X}}$ is preferred to $R_{\theta_0,\delta}$.
\end{example}
\paragraph{Preferences over local risk function sequences.} With Axiom \ref{as:2} in hand, it is without loss of generality to directly study the researcher's preferences at $\theta_0$ over \emph{local risk function sequences}. A \emph{local risk function} is a bounded, measurable function $r: H \rightarrow \mathbb{R}_+$. Let $R_H$ be the set of local risk functions. A \emph{local risk function sequence} is a bounded, measurable function $R: \mathbb{N} \times H \rightarrow \mathbb{R}_+$. Let $\mathcal{R}_H$ be the set of local risk function sequences. We take as primitive a binary relation $\succsim_{\theta_0}$ on $\mathcal{R}_H$, which we interpret as the preferences of a researcher facing the sequence of local informational environments defined in Section \ref{sec:setup}, in which the researcher has conditioned on the fixed centering value $\theta_0$.
Recall from Section \ref{sec:intro} that, although in practice the researcher’s problem features a fixed sample size, the criteria we study operate under a convenient fiction in which the researcher ranks local risk function \emph{sequences}
indexed by sample size. We may therefore interpret the decision environment as one in
which $\succsim_{\theta_0}$ is formed \emph{before} the researcher learns the sample size of their dataset. Under this interpretation, at the time of their decision, the researcher faces \emph{uncertainty} about the sample size $n$ (in addition to the local parameter $h$),\footnote{Since $\succsim_{\theta_0}$ is conditional on $\theta_0$, the representation results developed in the main text are also stated conditional on $\theta_0$. Appendix \ref{app:util-act-lemmas} studies a more primitive setup which does not condition on a fixed $\theta_0$, and adds an \emph{unanimity} axiom across centering values (Axiom \ref{ax:Theta-unanimity}) to ensure that the criterion considers performance conditional on $\theta_0$ for each $\theta_0\in \Theta$.} and their stance towards how to trade off risk across $n$ and $h$ is determined by their \emph{attitudes towards uncertainty} about $n$ and $h$. Following standard conventions in the microeconomic decision theory literature, we draw on this uncertainty-based interpretation to organize and discuss our axioms.
We briefly discuss the decision environment. While the boundedness restriction on the choice domain $\mathcal{R}_H$ is standard in the literature on decision-making under uncertainty, it rules out some local risk function sequences induced by estimator sequences where risk tends to infinity as the sample gets large (or yields infinite risk in finite sample). The LAM risk representations we obtain from our axioms therefore do not apply to such local risk function sequences. As is standard in the literature studying optimality criteria from a decision-theoretic perspective \citep{stoye2012new,andrews2026misspecification}, we study the researcher's preferences over \emph{all} local risk function sequences, even though the set of (bounded and measurable) local risk function sequences induced by some estimator sequence is a strict subset. To illustrate our axioms, we sometimes ask the researcher to consider their preferences over infeasible local risk function sequences, including those induced by \emph{oracle estimator sequences} which may condition their estimates on $h$. We have endeavored to make the sequences we draw on simple and interpretable, regardless of their feasibility.
We introduce some useful notation. For each $k\geq 0$, let $\overrightarrow{k} \in R_H$ denote the \emph{constant local risk function} which yields risk $k$ for each $h \in H$.\footnote{When the context is clear, we abuse notation and also use $\overrightarrow{k}$ to refer to the \emph{constant sequence} of constant local risk functions $(\overrightarrow{k},\overrightarrow{k},\ldots) \in \mathcal{R}_H$.} Estimators which are \emph{equivariant-in-law}\footnote{An estimator $\delta_n$ is \emph{equivariant-in-law} if the law of $\sqrt{n}(\delta_n(X_n)-\psi(\theta))$ under $X_n \sim P_{\theta,n}$ does not depend on $\theta$.} (such as the sample mean in Example \ref{ex:gaussian}) yield constant local risk functions. For local risk function sequences $R,R' \in \mathcal{R}_H$ and $\alpha \in [0,1]$, define the mixture $\alpha R+(1-\alpha)R' \in \mathcal{R}_H$ as: $(\alpha R+(1-\alpha)R')(n,h)=\alpha R(n,h)+(1-\alpha)R'(n,h)$. Note that this mixture occurs pointwise in $(n,h)$. To interpret this object, suppose $R$ is induced at $\theta_0$ by $\delta$ and $R'$ is induced at $\theta_0$ by $\delta'$. Then, $\alpha R+(1-\alpha)R'$ is induced at $\theta_0$ by the estimator sequence which flips a coin (independently of the data) whose probability of heads is $\alpha$, and uses $\delta$ if heads and $\delta'$ if tails.\footnote{Since the local risk function sequence induced by an estimator sequence only depends on the set of \emph{marginal distributions} of estimation errors for each $(n,h)$, it does not matter whether the coin is flipped before or after $(n,h)$ is realized (as long as the true value of $(n,h)$ does not depend on the outcome of the coin flip).}
We emphasize that comparing two local risk function sequences $R,R'\in \mathcal{R}_H$ is a nontrivial task, since $R$ may yield lower risk than $R'$ at some values of $(n,h)$ and higher risk at other values. Any complete comparison must therefore take a stance on how to trade off risk across sample sizes $n\geq 1$ and local parameters $h\in H$. A \emph{risk index} $V: \mathcal{R}_H \rightarrow \mathbb{R}$ provides a complete ranking of local risk function sequences in $\mathcal{R}_H$, where $V(R)$ aggregates the performance of $R$ across each $(n,h)$ into a single index of risk. A risk index $V: \mathcal{R}_H \rightarrow \mathbb{R}$ is a \emph{risk representation of $\succsim_{\theta_0}$ on $\mathcal{R}_H$} if: for each $R,R' \in \mathcal{R}_H$, $R\succsim_{\theta_0} R'$ if and only if $V(R)\leq V(R')$. The following definition collects the key LAM-based risk representations of interest.
\begin{definition}\label{defn:LAM-risk-main-text}
\begin{itemize}
\item[(i)] $\succsim_{\theta_0}$ has a \emph{generalized $\alpha$-LAM risk representation on $\mathcal R_H$} if there exists a function $\alpha: \mathcal{R}_H \rightarrow [0,1]$ such that the function
\begin{equation}\label{eq:gen-LAM-risk-act}
V_{\mathrm{gen}\text{-}\alpha\text{-}\mathrm{LAM}}(R)
:=\alpha(R)\underline V_{\mathrm{LAM}}(R)+(1-\alpha(R)) \overline V_{\mathrm{LAM}}(R) \tag{gen-$\alpha$-LAM risk}
\end{equation}
represents it, where
\begin{equation}
\label{eq:clam-risk}
\underline V_{\mathrm{LAM}}(R):=\sup_{I\in\mathcal I}\liminf_{n\to\infty}\max_{h\in I}R(n,h) \tag{CLAM risk}
\end{equation}
and
\begin{equation}
\label{eq:alam-risk}
\overline V_{\mathrm{LAM}}(R):=\sup_{I\in\mathcal I}\limsup_{n\to\infty}\max_{h\in I}R(n,h) \tag{ALAM risk}
\end{equation}
\item[(ii)] $\succsim_{\theta_0}$ has an \emph{$\alpha$-LAM risk representation on $\mathcal R_H$} if it has a generalized $\alpha$-LAM risk representation on $\mathcal R_H$ where $\alpha \in [0,1]$ is a constant.
\end{itemize}
\end{definition}
Recall that the set of generalized $\alpha$-LAM researcher preferences on $\mathcal{R}_H$ is precisely the set of researcher preferences on $\mathcal{R}_H$ that have a risk representation that is \emph{transparently consistent} with standard efficiency arguments.\footnote{Since local risk function sequences in $\mathcal{R}_H$ are bounded, the equivalence holds exactly: $V$ is generalized $\alpha$-LAM on $\mathcal{R}_H$ if and only if $\underline{V}_{\mathrm{LAM}}\leq V\leq \overline{V}_{\mathrm{LAM}}$ on $\mathcal{R}_H$.} As an initial benchmark, we axiomatically characterize this set. Our axioms reveal a lack of basic rationality requirements such as Monotonicity and Mixture Continuity. Amending our axioms to account for these will lead us to the (fixed) $\alpha$-LAM class, and eventually to the classical and attainment LAM risk representations \eqref{eq:clam-risk} and \eqref{eq:alam-risk}.
Our axiomatic exercise is tightly connected to the motivation and decision environment of Section \ref{sec:setup}. In particular, each risk index $V: \mathcal{R}_H \rightarrow \mathbb{R}_+$ studied in Definition \ref{defn:LAM-risk-main-text}(i) and (ii) induces a corresponding optimality criterion for estimator sequences conditional on $\theta_0$, by defining the risk index of $\delta$ as $V_{\theta_0}(\delta):=V(R_{\theta_0,\delta})$. We therefore seek axiomatic characterizations of the set of binary relations $\succsim_{\theta_0}$ on $\mathcal{R}_H$ which have the risk representations in Definition \ref{defn:LAM-risk-main-text}(i) and (ii) with the understanding that, under Axiom~\ref{as:2}, obtaining such a representation for local risk function sequences produces the desired representation for estimator sequences.\footnote{More precisely, for estimator sequences whose induced local risk function sequences at $\theta_0$ are bounded, measurable
functions $\mathbb{N} \times H\rightarrow \mathbb{R}_+$.}
\paragraph{Basic axioms.} We begin with two variants of standard regularity conditions on preferences. We use the conditions in Axiom \ref{ax: basics-main-text} as part of our characterization of $\alpha$-LAM preferences. However, in this section we document that generalized $\alpha$-LAM need not satisfy even these standard conditions, including Monotonicity and Mixture Continuity. Consequently, Axiom \ref{ax:gen-basics-main-text} provides a weaker variant to use for our characterization of generalized $\alpha$-LAM.
\begin{axiom}[Basic axioms]\label{ax: basics-main-text}
\begin{itemize}
\item[(i)] \textnormal{Weak Order:} $\succsim_{\theta_0}$ is complete and transitive on $\mathcal{R}_H$.
\item[(ii)] \textnormal{Strong Monotonicity:} For each $R,R' \in \mathcal{R}_H$,
\[
R'(n,h)\geq R(n,h) \quad \forall (n,h)\in \mathbb{N} \times H \implies R \succsim_{\theta_0} R'
\]
For any $j>k$, $\overrightarrow{k} \succ_{\theta_0} \overrightarrow{j}$.
\item[(iii)] \textnormal{Mixture Continuity:} For each $R,R',R'' \in \mathcal{R}_H$, the following sets are closed:
\[
\{\alpha \in [0,1]: \alpha R+(1-\alpha)R' \succsim_{\theta_0} R''\} \quad \text{and} \quad \{\alpha \in [0,1]: R'' \succsim_{\theta_0} \alpha R+(1-\alpha)R'\}
\]
\end{itemize}
\end{axiom}
Axiom \ref{ax: basics-main-text}(i) requires that the researcher can rank any two local risk function sequences, and that their rankings are transitive.\footnote{Although completeness is a standard axiom in the microeconomic decision theory literature, it may reasonably be viewed as strong in the context of comparing local risk function sequences. Nevertheless, completeness is required by any risk representation.} Axiom \ref{ax: basics-main-text}(ii) requires that a local risk function sequence which yields weakly lower risk at every sample $n$ and local parameter value $h$ is weakly preferred, and strictly lower values of constant risk are strictly preferred. Axiom \ref{ax: basics-main-text}(iii) requires that small perturbations in mixtures of local risk function sequences do not alter their ranking.
The following example documents that, due to the arbitrary dependence of $\alpha(R)$ on $R$, generalized $\alpha$-LAM preferences may violate Monotonicity and Mixture Continuity.
\addtocounter{example}{-1}
\begin{example}[Continued]
\label{ex:violations}
\normalfont Consider the statistical environment of our running example. For a bounded sequence $(a_n)_n$, consider the estimator sequence $\gamma_n=\Bar{X}_n+a_n/\sqrt{n}$, which shifts the sample mean by local bias $a_n$. Since each $\gamma_n$ is equivariant-in-law, analogous computations as before yield that $R_{\theta_0,\gamma}=(\overrightarrow{1+a_n^2})_{n}$. This implies that every bounded sequence of constant local risk functions $K=(\overrightarrow{k_n})_n$ with each $k_n\geq 1$ is induced by some $\gamma$ and $(a_n)_n$. We therefore work directly with such local risk function sequences. For ease of notation, we drop the overhead arrow and refer to $K=(k_n)_n$ as a \emph{local risk sequence}.
\begin{itemize}
\item[(i)] Consider the local risk sequences $K=(1,5,1,5,\ldots)$ and $K'=(2,10,2,10,\ldots)$, respectively. Set $\alpha(K)=0$ and $\alpha(K')=1$, and extend $\alpha$ arbitrarily to $\mathcal{R}_H$. Then $K'\geq K$ pointwise, but $V_{\mathrm{gen}\text{-}\alpha\text{-}\mathrm{LAM}}(K)=5>2=V_{\mathrm{gen}\text{-}\alpha\text{-}\mathrm{LAM}}(K')$, which implies $K'\succ_{\theta_0}K$: the sequence with strictly higher risk at every sample size is strictly preferred. This violates Monotonicity.
\item[(ii)] For each $\beta \in [0,1]$, define $R_\beta:=\beta K+(1-\beta)K'$. Set $\widetilde\alpha(R_\beta)=1$ for rational $\beta$ and $0$ otherwise, and extend $\widetilde\alpha$ to $\mathcal{R}_H$ arbitrarily. Then
\[
\{\beta \in [0,1]:R_\beta\succsim_{\theta_0}\overrightarrow 3\}=\mathbb{Q} \cap [0,1]
\]
which is not closed. This violates Mixture Continuity.
\end{itemize}
\end{example}
We must therefore weaken Axiom \ref{ax: basics-main-text} to characterize generalized $\alpha$-LAM.
\begin{axiom}[Basic Axioms II]\label{ax:gen-basics-main-text}
\begin{itemize}
\item[(i)] \textnormal{Weak Order:} $\succsim_{\theta_0}$ is complete and transitive on $\mathcal R_H$.
\item[(ii)] \textnormal{Constant Calibration:} For each $k,j\in\mathbb R_+$, $\overrightarrow{k}\succsim_{\theta_0}\overrightarrow{j}
\iff k\leq j$.
\item[(iii)] \textnormal{Constant Mixture Continuity:} For each $R\in\mathcal R_H$ and $k,j\in\mathbb R_+$, the following sets are closed:
\[
\left\{\beta \in [0,1]: \beta \overrightarrow{k}+(1-\beta)\overrightarrow{j} \succsim_{\theta_0}R\right\}
\quad \text{and} \quad
\left\{\beta \in [0,1]:R\succsim_{\theta_0}\beta \overrightarrow{k}+(1-\beta)\overrightarrow{j} \right\}
\]
\end{itemize}
\end{axiom}
Axiom \ref{ax:gen-basics-main-text}(i) is identical to Axiom \ref{ax: basics-main-text}(i). Axiom \ref{ax:gen-basics-main-text}(ii) merely requires that (strictly) smaller risks are (strictly) better. Axiom \ref{ax:gen-basics-main-text}(iii) requires that small perturbations in mixtures of \emph{constant risks} do not alter their ranking. Note that each of Axiom \ref{ax:gen-basics-main-text}(i)-(iii) is implied by Axiom \ref{ax: basics-main-text}(i)-(iii), respectively.
\addtocounter{example}{-1}
\begin{example}[Continued]
\label{ex:gamma}
\normalfont The estimator sequence $\gamma_n=\Bar{X}_n+1/\sqrt{n}$ induces the constant local risk function sequence $R_{\theta_0,\gamma}=\overrightarrow{2}$. Both Axioms \ref{ax: basics-main-text}(ii) and \ref{ax:gen-basics-main-text}(ii) require that $R_{\theta_0,\Bar{X}} \succ_{\theta_0} R_{\theta_0,\gamma}$.
\end{example}
\paragraph{Axioms for uncertainty towards $\boldsymbol{I}$.} Observe that each of the risk representations we consider may be written as $V(R)=\sup_{I\in \mathcal{I}} V_I(R)$ for a collection of risk indices $\{V_I\}_{I\in \mathcal{I}}$, which takes the \emph{worst-case risk} over all local neighborhoods $I\in \mathcal{I}$. In this section, we state axioms which ensure that the researcher's attitudes towards uncertainty about $I$ are described by this worst-case evaluation. Representations related to this form are axiomatically studied in \cite{gilboa2010objective}, and our axioms are partially adapted from their analysis.\footnote{\cite{gilboa2010objective} study the \emph{maxmin expected utility} model of \cite{gilboa1989maxmin}, which takes the worst case over \emph{beliefs}. In their case, preferences conditional on a belief are \emph{subjective expected utility}. By contrast, in our setting the researcher takes the worst case over \emph{local neighborhoods} $I\in \mathcal{I}$, and conditional on $I$, preferences have representations which require their own axiomatic analysis. \cite{gilboa2010objective} also make the additional assumption that the unanimity rule induced by conditional preferences is separately observed, whereas we derive the $I$-conditional preferences directly from the researcher's actual preference. For these reasons, our approach requires different techniques (such as Lemma \ref{lem:I-rep} in Appendix \ref{app:util-act-lemmas}).} However, since we study generalized $\alpha$-LAM risk and $\alpha$-LAM risk, we again require two variants of these axioms specialized to our setting and functional forms of interest.
To define these axioms, we first derive the researcher's preferences over $\mathcal{R}_H$ \emph{conditional} on each local neighborhood $I \in \mathcal{I}$, denoted $\succsim_{\theta_0,I}$, from the researcher's actual preferences $\succsim_{\theta_0}$. Our axioms are then jointly imposed on $\succsim_{\theta_0}$ and $\{\succsim_{\theta_0,I}\}_{I\in\mathcal{I}}$. We begin by defining a useful \emph{splicing} operation on the set of local risk function sequences.
\begin{definition}
For each $R,R' \in \mathcal{R}_H$ and $I \in \mathcal{I}$, define $R_I R' \in \mathcal{R}_H$ as:
\[
(R_IR')(n,h)=\begin{cases}
R(n,h) & \text{if } h\in I \\
R'(n,h) & \text{else}
\end{cases}
\]
\end{definition}
To interpret the splicing operation, suppose $R$ and $R'$ are induced at $\theta_0$ by $\delta$ and $\delta'$. Then $R_I R'$ is induced by an oracle estimator sequence which uses $\delta$ when $h\in I$ and $\delta'$ otherwise. $I$-conditional preferences are derived by comparing spliced local risk function sequences after fixing $R'$ to be the best constant risk $\overrightarrow{0}$.
\begin{definition}\label{def:I-cond}
For each $I\in\mathcal I$, define the binary relation $\succsim_{\theta_0,I}$ on $\mathcal R_H$ as:
\[
R\succsim_{\theta_0,I}R'
\iff
R_I\overrightarrow 0\succsim_{\theta_0}R'_I\overrightarrow 0
\]
\end{definition}
We refer to $\succsim_{\theta_0,I}$ as the researcher's \emph{$I$-conditional preference}. To interpret it, let $R$ and $R'$ be induced at $\theta_0$ by $\delta$ and $\delta'$, and suppose that for each $h\notin I$, the oracle perfectly reveals the local estimand at $h$ for every $n\geq 1$, such that the researcher incurs \emph{zero} risk. Then, the researcher $I$-conditionally prefers $R$ to $R'$ when they nevertheless prefer to use $\delta$ over $\delta'$ for local parameters $h\in I$ where the local estimand is not revealed.
As discussed, each of our axioms in this section relate the preferences $\succsim_{\theta_0}$ and $\{\succsim_{\theta_0,I}\}_{I\in \mathcal{I}}$ to specify a \emph{maximally pessimistic} attitude towards uncertainty about $I\in \mathcal{I}$.
\begin{axiom}[Consistency and Caution]\label{ax:cons-and-caution-main-text}
\begin{itemize}
\item[(i)] \emph{Consistency:} For each $R,R' \in \mathcal{R}_H$,
\[
R \succsim_{\theta_0,I} R' \quad \forall I \in \mathcal{I} \implies R \succsim_{\theta_0} R'
\]
\item[(ii)] \emph{Caution:} For each $R \in \mathcal{R}_H$ and $k\in \mathbb{R}_+$,
\[
\overrightarrow{k}\succ_{\theta_0,I} R \quad \text{ for some } I\in\mathcal{I} \implies \overrightarrow{k}\succsim_{\theta_0} R
\]
\end{itemize}
\end{axiom}
Axiom \ref{ax:cons-and-caution-main-text}(i) requires that if a local risk function sequence is $I$-conditionally preferred \emph{for each local neighborhood $I\in \mathcal{I}$}, then it is unconditionally preferred. Since Axiom \ref{ax:cons-and-caution-main-text}(i) only requires that $\succsim_{\theta_0}$ respects $\{\succsim_{\theta_0,I}\}_{I\in\mathcal{I}}$ \emph{when the latter unanimously agree} (and is silent otherwise), it ensures that $\succsim_{\theta_0}$ is \emph{minimally consistent} with $\{\succsim_{\theta_0,I}\}_{I\in\mathcal{I}}$.\footnote{In this sense, we may also interpret Axiom \ref{ax:cons-and-caution-main-text}(i) as a Pareto principle for aggregating $\{\succsim_{\theta_0,I}\}_{I\in \mathcal{I}}$.} Axiom \ref{ax:cons-and-caution-main-text}(ii) ensures that as long as there exists \emph{some} local neighborhood $I\in \mathcal{I}$ for which the researcher would $I$-conditionally prefer to incur a certain risk over a local risk function sequence, they would unconditionally prefer to do so as well. Put another way, say that a preference exhibits \emph{caution} if it ranks a certain risk over a local risk function sequence (whose risk may depend on the uncertain values of $n$ and $h$). Axiom \ref{ax:cons-and-caution-main-text}(ii) requires that as soon as some $\succsim_{\theta_0,I}$ (strictly) exhibits caution, $\succsim_{\theta_0}$ does so (weakly) as well. In this sense, $\succsim_{\theta_0}$ is \emph{maximally cautious} with respect to $\{\succsim_{\theta_0,I}\}_{I\in\mathcal{I}}$.
We use the conditions in Axiom \ref{ax:cons-and-caution-main-text} as part of our characterization of $\alpha$-LAM. However, as in the previous section, unrestricted dependence of $\alpha(R)$ on $R$ means that generalized $\alpha$-LAM need not satisfy Consistency and Caution.
\addtocounter{example}{-1}
\begin{example}[Continued]
\normalfont Consider the local risk sequences $K=(1,10,1,10,\ldots)$ and $K'=(2,5,2,5,\ldots)$. As previously discussed, each of these local risk sequences is feasible in the statistical environment of our running example. Suppose $\succsim_{\theta_0}$ has a generalized $\alpha$-LAM representation $V_{\mathrm{gen}\text{-}\alpha\text{-}\mathrm{LAM}}$.
\begin{itemize}
\item[(i)] For each $I\in \mathcal{I}$, set $\alpha(K_I\overrightarrow{0})=\alpha(K'_I\overrightarrow{0})=1$ and $\alpha(K)=\alpha(K')=0$. Then,
\begin{align*}
V_{\mathrm{gen}\text{-}\alpha\text{-}\mathrm{LAM}}(K'_I\overrightarrow{0})\geq V_{\mathrm{gen}\text{-}\alpha\text{-}\mathrm{LAM}}(K_I\overrightarrow{0}) \quad \forall I\in \mathcal{I} \\
\iff 2=\liminf_{n\to\infty} K'_n\geq \liminf_{n\to\infty} K_n=1
\end{align*}
but
\begin{align*}
10=\limsup_{n\to\infty} K_n=V_{\mathrm{gen}\text{-}\alpha\text{-}\mathrm{LAM}}(K)>V_{\mathrm{gen}\text{-}\alpha\text{-}\mathrm{LAM}}(K')=\limsup_{n\to\infty} K'_n=5
\end{align*}
\item[(ii)] Fix any $I\in \mathcal{I}$ and set $\alpha(K_I\overrightarrow{0})=0$, $\alpha(K)=1$, and $k=5$. Then, $k<\limsup_{n\to\infty} K_n=10$ but $k>\liminf_{n\to\infty} K_n=1$.
\end{itemize}
\end{example}
For generalized $\alpha$-LAM, we therefore instead require \emph{limiting versions} of Consistency and Caution.
\begin{axiom}[Asymptotic Consistency and Caution]
\label{ax:asymptotic-benchmark-dominance}
For each $R\in\mathcal R_H$ and $k\in\mathbb R_+$,
\begin{enumerate}[(i)]
\item[(i)] \emph{Asymptotic Consistency:}
\[
R_n \succsim_{\theta_0,I} \overrightarrow{k} \text{ e.v.} \quad \forall I\in \mathcal{I} \implies R\succsim_{\theta_0}\overrightarrow{k}
\]
\item[(ii)] \emph{Asymptotic Caution:}
\[
\overrightarrow{k}\succsim_{\theta_0,I} R_n \text{ e.v.} \quad \text{ for some } I \in \mathcal{I} \implies \overrightarrow{k}\succsim_{\theta_0}R
\]
\end{enumerate}
\end{axiom}
Axiom \ref{ax:asymptotic-benchmark-dominance}(i)-(ii) are asymptotic analogs of Axioms \ref{ax:cons-and-caution-main-text}(i)-(ii). Axiom \ref{ax:asymptotic-benchmark-dominance}(i) requires that if each local risk function along a sequence is \emph{eventually} $I$-conditionally preferred to a constant risk for \emph{every} local neighborhood $I\in \mathcal{I}$, then the entire sequence is unconditionally preferred. Axiom \ref{ax:asymptotic-benchmark-dominance}(ii) requires that as soon as each local risk function along a sequence is \emph{eventually} $I$-conditionally dominated by a certain risk for \emph{some} local neighborhood, then the entire sequence is unconditionally dominated.
Since Axioms \ref{ax:cons-and-caution-main-text} and \ref{ax:asymptotic-benchmark-dominance} characterize the aggregator $\sup_{I\in \mathcal{I}}$ in the functional forms of $\alpha$-LAM risk and generalized $\alpha$-LAM risk, our remaining axioms are applied to $\succsim_{\theta_0,I}$ for each $I\in \mathcal{I}$. We make a further simplifying observation: $R$'s ranking according to $\succsim_{\theta_0,I}$ depends on $R$ only through its restriction to $I$, denoted $R_I:\mathbb N\times I\to\mathbb R_+$. The interpretation of this observation is straightforward: if two estimator sequences yield the same risk at every $(n,h)$ with $h\in I$, an oracle who is constrained to provide zero risk outside $I$ has no scope to alter the risk profile. Hence, the researcher is $I$-conditionally indifferent between them.
Let $\mathcal{R}_I$ be the set of bounded, measurable functions $R_I: \mathbb{N} \times I \rightarrow \mathbb{R}_+$. An \emph{$I$-local risk function sequence} $R_I \in \mathcal{R}_I$ may be interpreted analogously as before: if $R$ is induced at $\theta_0$ by $\delta$, $R_I$ records $\delta$'s performance in the sequence of local decision problems at $\theta_0$ for local parameters $h\in I$. The above observation ensures that we may impose the remaining axioms on the induced $I$-conditional preference $\succsim_{\theta_0,I}$ on the simpler domain $\mathcal{R}_I$.\footnote{This follows from the axioms we have imposed on $\succsim_{\theta_0}$ and by definition of $\succsim_{\theta_0,I}$. Lemma \ref{lem:I-relevance} in Appendix \ref{app:util-act-lemmas} formalizes this observation.}
\paragraph{Axioms for uncertainty towards $\boldsymbol{h}$.} Observe that, for a fixed local neighborhood $I\in \mathcal{I}$ and sample size $n\geq 1$, each of the risk representations we study depend on $R(n,\cdot)$ only through its \emph{worst-case risk} over $h\in I$: $\max_{h\in I} R(n,h)$. In this section, we offer axioms which ensure that the researcher's $I$-conditional attitudes towards uncertainty about $h\in I$ are described by this worst-case evaluation, and which apply to each of the risk representations we consider. Representations related to the form $R(n,\cdot) \mapsto \max_{h\in I} R(n,h)$ were axiomatized by \cite{gilboa1989maxmin} (henceforth GS89), and our axioms adapt their analysis to our setting.
To isolate the researcher's attitudes towards uncertainty about $h\in I$, we focus on the researcher's ranking of elements of $\mathcal{R}_I$ which are \emph{constant in $n$}. Formally, an \emph{$I$-local risk function} is a function $r_I: I \rightarrow \mathbb{R}_+$. Let $\mathbb{R}_+^I \subseteq \mathcal{R}_I$ denote the set of $I$-local risk functions.\footnote{By identifying each $r_I \in \mathbb{R}_+^I$ with the $I$-local risk function sequence $(r_I,r_I,\ldots) \in \mathcal{R}_I$, we may view $\mathbb{R}_+^I$ as a subset of $\mathcal{R}_I$. Hence, any binary relation on $\mathcal{R}_I$ induces a binary relation on $\mathbb{R}_+^I$.} To interpret this object, suppose $R_I$ is induced at $\theta_0$ by $\delta$. For each fixed $n\geq 1$, $(R_I)_n$ is the $I$-local risk function induced by $\delta_n$ at $\theta_0$, which describes $\delta_n$'s performance in the $n$-th local decision problem at $\theta_0$ for local parameters $h\in H$.
\begin{axiom}[GS89 Axioms]\label{ax:gs-main-text}
$\succsim_{\theta_0,I}$ on $\mathbb{R}_+^I$ satisfies:
\begin{itemize}
\item[(i)] \textnormal{Nontrivial Weak Order:} $\succsim_{\theta_0,I}$ is nontrivial, complete and transitive on $\mathbb{R}_+^I$.
\item[(ii)] \textnormal{Monotonicity:} For each $r_I,r_I' \in \mathbb{R}_+^I$, $r_I'(h)\geq r_I(h)$ for all $h\in I$ implies $r_I\succsim_{\theta_0,I} r_I'$.
\item[(iii)] \textnormal{Mixture Continuity:} For each $r_I,r_I',r_I'' \in \mathbb{R}_+^I$, the following sets are closed:
\[
\{\alpha \in [0,1]: \alpha r_I+(1-\alpha)r_I' \succsim_{\theta_0,I} r_I''\} \quad \text{and} \quad \{\alpha \in [0,1]: r_I'' \succsim_{\theta_0,I} \alpha r_I+(1-\alpha)r_I'\}
\]
\item[(iv)] \textnormal{Certainty Independence:} For each $r_I,r_I' \in \mathbb{R}_+^I$, $k \in \mathbb{R}_+$, and $\alpha \in (0,1)$,
\[
r_I \succsim_{\theta_0,I} r_I' \iff
\alpha r_I + (1-\alpha)\overrightarrow{k} \succsim_{\theta_0,I} \alpha r_I' + (1-\alpha)\overrightarrow{k}
\]
\item[(v)] \textnormal{Uncertainty Aversion}: For each $r_I,r_I' \in \mathbb{R}_+^I$ and $\alpha \in (0,1)$,
\[
r_I \sim_{\theta_0,I} r_I' \implies \alpha r_I + (1-\alpha)r_I' \succsim_{\theta_0,I} r_I
\]
\end{itemize}
\end{axiom}
Axiom \ref{ax:gs-main-text}(i)-(iii) have analogous interpretations to Axiom \ref{ax: basics-main-text}(i)-(iii), adapted to rankings of $I$-local risk functions. Axiom \ref{ax:gs-main-text}(iv) requires that rankings of $I$-local risk functions are preserved under mixing with (identical) constant local risks. To understand this, suppose $r_I$ and $r_I'$ are induced by estimators $\delta_n$ and $\delta_n'$ at $\theta_0$, and let $\gamma_n$ be any equivariant-in-law estimator. Axiom \ref{ax:gs-main-text}(iv) requires that if the researcher $I$-conditionally prefers $\delta_n$ to $\delta_n'$, then they $I$-conditionally prefer the randomized estimator which flips a coin and uses $\delta_n$ if heads and $\gamma_n$ if tails over the randomized estimator flips a coin and uses $\delta_n'$ if heads and $\gamma_n$, regardless of the distribution of the coin. Axiom \ref{ax:gs-main-text}(v) ensures that the researcher exhibits a \emph{preference for hedging} against uncertainty towards $h$: if the researcher is $I$-conditionally indifferent between $\delta_n$ and $\delta_n'$, then they $I$-conditionally prefer the randomized estimator which flips a coin between them to either $\delta_n$ or $\delta_n'$, regardless of the distribution of the coin. Intuitively, the coin acts as an external randomization device which flattens the risk profile across $h\in I$.\footnote{In statistics and econometrics, we often have the intuition (motivated by Jensen’s inequality and convex loss functions) that adding extra randomness is undesirable. Axioms such as Axiom \ref{ax:gs-main-text}(v) do not contradict this intuition: they do not require randomized estimators to be \emph{undominated}, merely that some randomized estimators dominate other estimators of a particular form.}
The GS89 axioms ensure that $\succsim_{\theta_0,I}$ on $\mathbb{R}_+^I$ has a representation of the form $r_I \mapsto \max_{p\in C} \int r_I \ dp$ for some set of priors $C\subseteq \Delta(I)$. To ensure that the researcher takes the worst-case over \emph{all} $h\in I$, we impose the following axiom, adapted from \cite{ghirardato2002ambiguity}.
\begin{axiom}[$I$-Maximal Caution]\label{ax:I-Caution-Main-Text}
For each $r_I \in \mathbb R_+^I$ and $k\in \mathbb{R}_+$,
\[
r_I \succsim_{\theta_0,I} \overrightarrow{k} \implies \overrightarrow{r_I(h)} \succsim_{\theta_0,I} \overrightarrow{k} \quad \forall h\in I
\]
\end{axiom}
Axiom \ref{ax:I-Caution-Main-Text} requires that whenever the researcher $I$-conditionally prefers an estimator $\delta_n$ to an equivariant-in-law estimator $\gamma_n$, $\delta_n$ must incur less risk than $\gamma_n$ for every $h\in I$.
\paragraph{Characterizing generalized $\boldsymbol{\alpha}$-LAM.} The following result provides an axiomatic characterization of generalized $\alpha$-LAM.
\begin{theorem}[Generalized $\alpha$-LAM]\label{thm-main-result}
The following are equivalent:
\begin{itemize}
\item[(i)] $\succsim_{\theta_0}$ satisfies Axiom~\ref{ax:gen-basics-main-text}; $\succsim_{\theta_0}$ and $\{\succsim_{\theta_0,I}\}_{I\in \mathcal{I}}$ satisfy Axiom \ref{ax:asymptotic-benchmark-dominance}; and each $\succsim_{\theta_0,I}$ satisfies Axioms \ref{ax:gs-main-text}-\ref{ax:I-Caution-Main-Text}.
\item[(ii)] $\succsim_{\theta_0}$ has a generalized $\alpha$-LAM risk representation on $\mathcal R_H$.
\end{itemize}
\end{theorem}
The axioms in Theorem \ref{thm-main-result} constitute a necessary and sufficient set of conditions on researcher behavior such that they are merely \emph{transparently consistent} with conventional efficiency arguments. In this sense, Theorem \ref{thm-main-result} serves as a useful benchmark towards eventually selecting a functional form for LAM risk. However, as we have previously discussed, transparent consistency by itself is not a sufficiently demanding desideratum, since it allows for preferences which violate some basic tenets of rational decision-making. As Example \ref{ex:violations} suggests, these violations are due to the arbitrary dependence of $\alpha(R)$ on $R\in \mathcal{R}_H$. This motivates the focus of the rest of this note on $\alpha$-LAM, which restores the basic axioms we want and brings us one step closer to pinning down a functional form for LAM risk.
\subsection{Remaining axioms for $\boldsymbol{\alpha}$-LAM}
Taking stock, recall that any risk representation of $\succsim_{\theta_0}$ must take a stance on how to aggregate risk across $n\geq 1$ and $h\in H$, and axioms allow us to translate this stance into concrete implications for researcher behavior. In particular, the functional form for $\alpha$-LAM risk summarizes risk via three aggregators---$\sup_{I\in\mathcal{I}}$, $\alpha \liminf_{n\to\infty}+(1-\alpha)\limsup_{n\to\infty}$, and $\max_{h\in I}$---that determine the criterion's attitudes towards uncertainty about $I\in\mathcal{I}$, $n\geq 1$, and $h\in I$, respectively. We have already imposed axioms relating $\succsim_{\theta_0}$ and $\{\succsim_{\theta_0,I}\}_{I\in\mathcal{I}}$ which ensure that the researcher takes the \emph{worst-case} over $I\in \mathcal{I}$ (Axiom \ref{ax:cons-and-caution-main-text}), as well as axioms on each $\succsim_{\theta_0,I}$ which ensure that the researcher takes the \emph{worst-case} over $h\in I$ (Axioms \ref{ax:gs-main-text}-\ref{ax:I-Caution-Main-Text}). It remains to provide axioms on each $\succsim_{\theta_0,I}$ which ensure that the researcher's $I$-conditional attitudes towards uncertainty about $n\geq1$ are described by the aggregator $\alpha \liminf_{n\to\infty}+(1-\alpha)\limsup_{n\to\infty}$ for some $\alpha \in [0,1]$. Recall that we may impose these axioms on the induced $I$-conditional preference $\succsim_{\theta_0,I}$, which ranks \emph{$I$-local risk function sequences} $\mathcal{R}_I$.
\paragraph{Axioms for uncertainty towards $\boldsymbol{n}$.}
We begin with a basic rationality condition which requires \emph{monotonicity} in sample size conditional on $I$.
\begin{axiom}[$I$-Sample Monotonicity]\label{ax:-sample-size-mono-main-text}
For each $R_I,R_I' \in \mathcal{R}_I$,
\[
(R_I)_n \succsim_{\theta_0,I} (R_I')_n \quad \forall n\geq 1 \implies R_I \succsim_{\theta_0,I} R_I'
\]
\end{axiom}
Suppose $R_I$ and $R_I'$ are induced at $\theta_0$ by $\delta$ and $\delta'$, respectively. Axiom \ref{ax:-sample-size-mono-main-text} requires that if the researcher $I$-conditionally prefers the estimator $\delta_n$ to the estimator $\delta_n'$ at every sample size $n\geq1$, then they $I$-conditionally prefer the entire estimator sequence $\delta$ to the entire estimator sequence $\delta'$.
To isolate the researcher's attitudes towards uncertainty about $n\geq 1$, we focus on the researcher's ranking of elements of elements of $\mathcal{R}_I$ which are \emph{constant in $h$}. Formally (as defined in Example \ref{ex:violations}), a \emph{local risk sequence} is a bounded sequence $K=(k_n)_{n\geq 1}$, where each $k_n\geq 0$. Let $B_+^{\mathbb{N}} \subseteq \mathcal{R}_H$ denote the set of local risk sequences.\footnote{By identifying each $K \in B_+^{\mathbb{N}}$ with the local risk function sequence $(\overrightarrow{k}_n)_{n\geq 1} \in \mathcal{R}_H$, we may view $B_+^{\mathbb{N}}$ as a subset of $\mathcal{R}_H$. Hence, any binary relation on $\mathcal{R}_H$ induces a binary relation on $B_+^{\mathbb{N}}$.} Since constant local risk functions may be generated by equivariant-in-law estimators, a local risk sequence may be generated by a sequence of equivariant-in-law estimators whose (constant in $h$) risk may depend on the sample index $n$. For example, the estimator sequences $\gamma$ defined in Example \ref{ex:violations} for bounded sequences $(a_n)_n$ induce local risk sequences.
We therefore impose our axioms on $\succsim_{\theta_0,I}$ restricted to the domain $B_+^{\mathbb{N}}$. Among other representations, \cite{marinacci1998axiomatic} studies representations of the form $K\mapsto \liminf_{n\to\infty} K_n$. We build upon the analysis in \cite{marinacci1998axiomatic} Theorem~7 to offer axioms for representations of the more general form $K\mapsto \alpha \liminf_{n\to\infty} K_n+(1-\alpha) \limsup_{n\to\infty} K_n$ for some $\alpha \in [0,1]$.
Stating these axioms requires several more definitions. Say that $K,K'\in B_+^{\mathbb{N}}$ are \emph{comonotonic} if $(K_n-K_m)(K'_n-K'_m)\geq 0$ for all $m,n\in\mathbb{N}$, and say that a set of local risk sequences is \emph{pairwise comonotonic} if every pair from it is comonotonic. In words, comonotonic sequences exhibit the same ordering of risks across $n$, which implies that any mixture retains this ordering. A useful special case is when $K$ and $K'$ are each monotone decreasing in $n$. In Example \ref{ex:gaussian}, the estimator sequence $\gamma_n$ defined above yields a local risk sequence which is monotone decreasing in $n$ when $|a_n|\downarrow 0$. Next, for each $K \in B_+^{\mathbb{N}}$ and $m \in \mathbb{N}$, define $K^m = (K_{m+1},K_{m+2},\ldots)$ to be the local risk sequence which truncates the first $m$ risks in $K$. In Example \ref{ex:gaussian}, if $K$ is induced at $\theta_0$ by $\gamma$, $K^m$ is induced at $\theta_0$ by the estimator sequence $(\gamma_n^m)_{n\geq 1}$, where each $\gamma_n^m=\Bar{X}_n+a_{m+n}/\sqrt{n}$. Finally, for each $K \in B_+^{\mathbb{N}}$ and each bijection $\pi:\mathbb{N}\rightarrow \mathbb{N}$, define $K^\pi$ to be the local risk sequence satisfying $K^\pi_n = K_{\pi(n)}$ for all $n\in\mathbb{N}$. If $K$ is induced at $\theta_0$ by an equivariant-in-law estimator sequence $\delta$, $K^\pi$ is induced at $\theta_0$ by an equivariant-in-law estimator sequence which, for each sample $n$, reproduces the constant local risk incurred by $\delta_{\pi(n)}$. In Example \ref{ex:gaussian}, if $K$ is induced by
$\gamma_n=\Bar X_n+a_n/\sqrt n$, then $K^\pi$ is induced by $\gamma_n^\pi
=\Bar X_n+a_{\pi(n)}/\sqrt n$.
\begin{axiom}[\citealt{marinacci1998axiomatic} Axioms]\label{ax:mar-98-main-text}
$\succsim_{\theta_0,I}$ on $B_+^{\mathbb{N}}$ satisfies:
\begin{itemize}
\item[(i)] \textnormal{Nontrivial Weak Order:} $\succsim_{\theta_0,I}$ is nontrivial, complete, and transitive on $B_+^{\mathbb{N}}$.
\item[(ii)] \textnormal{Archimedean Continuity:} For each $K,K',K'' \in B_+^{\mathbb{N}}$, if $K \succ_{\theta_0,I} K'$ and $K' \succ_{\theta_0,I} K''$, then there exist $\beta,\gamma \in (0,1)$ such that
\[
\beta K + (1-\beta)K'' \succ_{\theta_0,I} K'
\quad\text{and}\quad
K' \succ_{\theta_0,I} \gamma K + (1-\gamma)K''
\]
\item[(iii)] \textnormal{Comonotonic Independence:} For all pairwise comonotonic $K,K',K'' \in B_+^{\mathbb{N}}$ and all $\beta \in (0,1)$,
\[
K \succsim_{\theta_0,I} K' \implies \beta K + (1-\beta)K''
\succsim_{\theta_0,I}
\beta K' + (1-\beta)K''
\]
\item[(iv)] \textnormal{Truncation Invariance:} For each $K \in B_+^{\mathbb{N}}$ and $m \in \mathbb{N}$, $K \sim_{\theta_0,I} K^m$.
\item[(v)] \textnormal{Permutation Invariance:} For each $K \in B_+^{\mathbb{N}}$ and each bijection $\pi:\mathbb{N}\rightarrow \mathbb{N}$, $K \sim_{\theta_0,I} K^\pi$.
\end{itemize}
\end{axiom}
Axiom \ref{ax:mar-98-main-text}(i)--(ii) are standard regularity conditions. Axiom \ref{ax:mar-98-main-text}(iii) requires that rankings are preserved under randomization with a common alternative when the local risk sequences being compared and the alternative order risk across sample sizes in the same way. For example, suppose $K$, $K'$, and $K''$ are induced at $\theta_0$ by $\gamma$, $\gamma'$, and $\gamma''$ defined as in Example \ref{ex:violations}, where each sequence of local biases is monotone decreasing in $n$. Axiom \ref{ax:mar-98-main-text}(iii) requires that if the researcher prefers $\delta$ to $\delta'$, then they prefer randomizing between $\delta$ and $\delta''$ to randomizing between $\delta'$ and $\delta''$ (where the probability of $\delta''$ is the same). Axioms \ref{ax:mar-98-main-text}(iv)-(v) enforce the asymptotic nature of the criterion, requiring that the researcher is indifferent to truncating and permuting risk sequences.
\paragraph{Characterizing $\boldsymbol{\alpha}$-LAM.} We may now use the axioms we have discussed to characterize $\alpha$-LAM.
\begin{theorem}[$\alpha$-LAM]\label{thm:fixed-alpha-LAM}
The following are equivalent:
\begin{itemize}
\item[(i)] $\succsim_{\theta_0}$ satisfies Axiom~\ref{ax: basics-main-text}; $(\succsim_{\theta_0},\{\succsim_{\theta_0,I}\}_{I\in\mathcal I})$ satisfy Axiom~\ref{ax:cons-and-caution-main-text}; and each $\succsim_{\theta_0,I}$ satisfies Axioms~\ref{ax:gs-main-text}-\ref{ax:mar-98-main-text}.
\item[(ii)] $\succsim_{\theta_0}$ has an $\alpha$-LAM risk representation on $\mathcal R_H$ for some $\alpha\in[0,1]$.
\end{itemize}
\end{theorem}
\paragraph{Characterizing classical and attainment LAM.} To understand what additional axioms select classical and attainment LAM from this class, we consider a motivating example.
\addtocounter{example}{-1}
\begin{example}[Continued]
\label{ex:uncertainty-affinity}
\normalfont Fix any centering value $\theta_0$. Let $A=\{10^k: k\in \mathbb{N}\}$, and define the estimator sequences $\delta,\delta'$ as:
\[
\delta_n(X_n)=
\begin{cases}\Bar{X}_n,& n\in A,\\\Bar{X}_n+1/\sqrt{n},& \text{ else},\end{cases}
\qquad
\delta_n'(X_n)=
\begin{cases}\Bar{X}_n+1/\sqrt{n},&n\in A,\\\Bar{X}_n,& \text{ else}\end{cases}
\]
In words, $\delta$ outputs the sample mean on sample sizes which are powers of 10 and otherwise outputs the sample mean with constant local bias $1$, while $\delta'$ does the opposite. The induced local risk sequences are:
\[
K_n=
\begin{cases}1,&n\in A,\\2,& \text{ else},\end{cases}
\qquad
K_n'=
\begin{cases}2,&n\in A,\\1,&\text{ else}\end{cases}
\]
The mixture $(K+K')/2$ yields the constant local risk sequence $(3/2,3/2,\ldots)$. Assume $\succsim_{\theta_0}$ has an $\alpha$-LAM representation. Then,
\[
V_{\alpha\text{-}\mathrm{LAM}}(K)
=V_{\alpha\text{-}\mathrm{LAM}}(K')=\alpha+(1-\alpha)2
\quad \text{and} \quad
V_{\alpha\text{-}\mathrm{LAM}}((K+K')/2)=3/2
\]
and hence
\[
\frac{1}{2}K+\frac{1}{2}K' \succsim_{\theta_0} K \sim_{\theta_0} K' \iff \alpha \leq 1/2
\]
Hence, comparing mixtures of two local risk sequences to each puts restrictions on the value of $\alpha$. In particular, $(K+K')/2 \succ_{\theta_0} K$ for $\alpha=0$ (attainment LAM) and $K \succ_{\theta_0} (K+K')/2$ for $\alpha=1$ (classical LAM).
\end{example}
Example \ref{ex:uncertainty-affinity} suggests that the key property of preferences which distinguishes classical LAM risk from attainment LAM risk is the researcher's attitudes towards \emph{hedging against uncertainty about $n$}. We may therefore state the sample-size-analog of Axiom \ref{ax:gs-main-text}(v), as well as its opposite.
\begin{axiom}[Sample Uncertainty Aversion]\label{ax:sample-size-uncertainty-aversion}
For each $K,K'\in B_+^{\mathbb N}$,
\[
K\sim_{\theta_0}K'
\implies
\beta K+(1-\beta)K'\succsim_{\theta_0}K
\quad\forall\beta\in(0,1).
\]
\end{axiom}
\begin{axiom}[Sample Uncertainty Affinity]\label{ax:sample-size-uncertainty-affinity}
For each $K,K'\in B_+^{\mathbb N}$,
\[
K\sim_{\theta_0}K'
\implies
K\succsim_{\theta_0}\beta K+(1-\beta)K'
\quad\forall\beta\in(0,1).
\]
\end{axiom}
Sample Uncertainty Aversion requires the researcher to prefer hedging between indifferent local risk sequences and selects $\alpha=0$, while Sample Uncertainty Affinity requires the researcher to dislike hedging between indifferent local risk sequences and selects $\alpha=1$.
\begin{corollary}[Classical and attainment LAM]\label{cor:CLAM-and-ALAM}
Suppose $\succsim_{\theta_0}$ has an $\alpha$-LAM risk representation on $\mathcal{R}_H$.
\begin{itemize}
\item[(i)] $\succsim_{\theta_0}$ satisfies Axiom \ref{ax:sample-size-uncertainty-aversion} if and only if $\succsim_{\theta_0}$ is represented by $\overline{V}_{\mathrm{LAM}}$ on $\mathcal{R}_H$.
\item[(ii)] $\succsim_{\theta_0}$ satisfies Axiom \ref{ax:sample-size-uncertainty-affinity} if and only if $\succsim_{\theta_0}$ is represented by $\underline{V}_{\mathrm{LAM}}$ on $\mathcal{R}_H$.
\end{itemize}
\end{corollary}
As the running example demonstrates, Sample Uncertainty Aversion seems normatively appealing: when two procedures are indifferent, randomizing between them hedges their sample-size-dependent risks. We therefore uniquely recommend attainment LAM as the functional form for LAM risk. Corollary~\ref{cor:CLAM-and-ALAM} provides a precise axiomatic foundation for this recommendation.
\setstretch{0.8}
\small
\setlength{\bibsep}{0pt}
\bibliography{WorksCited}
\newpage
\setstretch{1.2}
\normalsize