EconBase
← Back to paper

Characterization of Efficient Influence Function for Off-Policy Evaluation Under Optimal Policies

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

74,465 characters · 22 sections · 39 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness

\def\spacingset#1{ {#1}} \spacingset{1}

\def\spacingset#1{ {#1}} \spacingset{1}

\if11 \fi

\if01 \fi

abstractOff-policy evaluation (OPE) constructs confidence intervals for the value of a target policy using data generated under a different behavior policy. Most existing inference methods focus on fixed target policies and may fail when the target policy is estimated as optimal, particularly when the optimal policy is non-unique or nearly deterministic. We study inference for the value of optimal policies in Markov decision processes. We characterize the existence of the efficient influence function and show that non-regularity arises under policy non-uniqueness. Motivated by this analysis, we propose a novel Nonparametric SequentiAl Value Evaluation (NSAVE) method, which achieves semiparametric efficiency and retains the double robustness property when the optimal policy is unique, and remains stable in degenerate regimes beyond the scope of existing asymptotic theory. We further develop a smoothing-based approach for valid inference under non-unique optimal policies, and a post-selection procedure with uniform coverage for data-selected optimal policies. Simulation studies support the theoretical results. An application to the OhioT1DM mobile health dataset provides patient-specific confidence intervals for optimal policy values and their improvement over observed treatment policies.

{\it Keywords:} Efficient estimator; non-regular inference; optimal policy; off-policy evaluation.

\spacingset{1.9}

Introduction

Reinforcement learning (RL) is concerned with learning optimal decision rules for sequential decision problems in order to maximize long-term cumulative rewards sutton2018reinforcement. A fundamental statistical task within RL is off-policy evaluation (OPE), which seeks to estimate the value of a target policy using data generated under a potentially distinct behavior policy. OPE plays a pivotal role in offline RL, where new data collection is either costly or ethically constrained, necessitating that inference rely entirely on historical trajectories luedtke2016statistical,agarwal2019reinforcement,uehara2022review.

The majority of existing statistical analyses of OPE concentrate on the classical setting in which the evaluation policy is fixed and known a priori. In this regime, an extensive body of literature has established doubly robust and semiparametrically efficient estimators under various modeling assumptions jiang2016doubly,kallus2020double,shi2021deeply. However, in many empirical applications, the policy of interest is not pre-specified but is itself estimated from the data as an optimal policy. This setting introduces a qualitatively different statistical structure: the target functional involves a maximization over policies, and the resulting value function can be non-smooth and non-regular, particularly when the optimal policy is not unique or is nearly deterministic.

Analogous issues have been extensively studied in the causal inference literature regarding optimal treatment regimes laber2014dynamic,kosorok2019precision,athey2021policy, where it is now well-established that the non-uniqueness of optimal rules leads to non-regularity and renders standard asymptotic theory invalid. Extending such insights to the sequential decision-making framework of Markov decision processes (MDPs) is substantially more challenging due to temporal dependence, the Bellman fixed-point structure, and the complex interaction between policy optimization and value estimation.

Recently, shi2022statistical proposed the SAVE estimator, which establishes semiparametric efficiency for the value of an optimal policy under a linear $Q$-function model and a set of non-degeneracy conditions. While SAVE represents a significant step toward principled inference for optimal policy values, its theory relies on stringent structural and regularity assumptions. In particular, it requires (i) a low-dimensional linear approximation of the $Q$-function, and (ii) well-conditioned Bellman estimating equations under the target policy. When the optimal policy is unique and deterministic, or nearly so, the latter condition often fails: the feature covariance induced by the target policy becomes ill-conditioned, leading to numerical instability and the breakdown of the associated inference. Moreover, in such regimes, SAVE no longer admits a doubly robust representation and loses its efficiency guarantees; furthermore, no alternative valid confidence sets are provided once these non-degeneracy conditions are violated.

This paper develops a unified inferential framework for the value of optimal policies in MDPs that explicitly addresses such non-regular phenomena. Our contributions are threefold.

itemize• First, we characterize the existence of the efficient influence function (EIF) for the optimal policy value and derive its explicit form under the regime in which the optimal policy is unique and deterministic, and demonstrate that the classical EIF does not exist when the optimal policy is not unique. • Second, building on this characterization, we propose a novel Nonparametric SequentiAl Value Evaluation (NSAVE). NSAVE achieves semiparametric efficiency in the regular regime of a unique optimal policy and, in this case, also retains a doubly robust representation, while remaining well-defined and yielding valid inference in degenerate or near-degenerate regimes where existing methods become unstable. • Third, we develop a complementary smoothing-based approach that regularizes the policy optimization map via softmax approximation, thereby restoring differentiability and enabling first-order inference through a smoothed value functional. This construction provides an alternative route to valid uncertainty quantification under policy non-uniqueness and bridges optimal policy evaluation in MDPs with recent advances in post-selection and non-regular inference. Finally, beyond pointwise inference for the optimal value, we also consider a post-selection inference formulation. When the optimal policy is not unique, rather than targeting a single value functional, we construct confidence sets that uniformly cover the collection of values associated with the set of data-dependent estimated optimal policies. This provides a complementary form of uncertainty quantification that remains valid under policy non-uniqueness.

Through theoretical analysis, simulations, and an application to the OhioT1DM dataset, we demonstrate that the proposed NSAVE and smoothing procedures yield stable and valid confidence intervals across both regular and non-regular regimes, significantly outperforming existing methods in settings where the optimal policy is deterministic or nearly deterministic.

The remainder of the paper is organized as follows. In Sections (ref) and (ref), we first characterize the efficient influence function for the optimal policy value and establish the non-regularity that arises under policy non-uniqueness. In Sections (ref) and (ref), we then develop the proposed NSAVE estimator together with its efficiency and stability properties. Section (ref) introduces the smoothing-based approach and the associated post-selection confidence sets for handling non-unique optimal policies. The finite-sample performance of NSAVE, the smoothing approach, and existing methods is investigated through extensive simulation studies in Section (ref). In Section (ref), we apply the proposed methods to the OhioT1DM mobile health dataset and conduct patient-specific off-policy inference. Finally, the last section concludes with a discussion of the implications, limitations, and directions for future research.

Problem Formulation

Data Generating Process and Parameter of Interest

We consider observational data generated from a canonical Markov decision process (MDP). At any given time $t$, let $(S_t, A_t, R_t)$ denote the state-action-reward triplet. Let $O$ be shorthand for the data tuple $(S, A, R, S^{\prime})$. We observe an offline dataset $\big\{ O_{it}: 1 \leq i \leq N, 0 \leq t \leq T\big\}$ with $O_{it} = (S_{it}, A_{it}, R_{it})$, generated by a behavior policy $b(\cdot \mid S)$, where $i$ indexes the episode and $t$ indexes the time point. For any fixed target policy $\pi(a \mid s)$, OPE generally aims to evaluate the mean return $ \eta(\pi) = \mathrm{E}^{\sim \pi} \left[ \sum_{t = 0}^{+\infty} \gamma^t R_t \right] $ and construct a valid confidence interval, where $\mathrm{E}^{\sim \pi}$ denotes the expectation when the system follows policy $\pi$. Distinct from existing semiparametric studies on OPE that consider an arbitrary $\pi$, we focus on a specific target policy: the optimal policy $\pi^*$, which maximizes $\eta(\pi)$ over the set of all possible policies $\Pi$. Specifically, the parameter of interest is \[

aligned\eta^* := \eta(\pi^*) = \mathrm{E}^{\sim \pi^*} \bigg[ \sum_{t = 0}^{+\infty} \gamma^t R_t \bigg]\quad such that \quad \pi^* = \arg\max_{\pi \in \Pi} \eta(\pi).

\]

To ensure the value function is identifiable, we adopt standard assumptions in the OPE literature. For simplicity, we use $f(x \mid y)$ to represent the conditional density of $X$ given $Y = y$.

assumption[Data Structure & Observations] The observations are i.i.d. copies of the trajectory $\{ (S_t, A_t, R_t, S_{t + 1})\}_{t \geq 0}$, following the data-generating process (DGP): $S_{t + 1} \sim f(s_{t + 1} \mid A_t, S_t)$, $ R_t \sim f(r_t \mid A_t, S_t)$, and $A_t \sim b(a_t \mid S_t)$.
assumption[Markov, Conditional Independence, & Time-Homogeneity] $f(a_t, s_t \mid a_{t-1}, s_{t-1}, \penalty 0 a_{t-2}, s_{t-2}, \ldots) = f(a_t, s_t \mid a_{t-1}, s_{t-1})$ for any $t \geq 1$; $f(a_t \mid a_{t-1}, s_t) = b(a_t \mid s_t)$. The reward $R_t$ depends only on $A_t$ and $S_t$; All conditional density functions $b(a \mid s)$, $f(r \mid a, s)$, and $f(s' \mid a, s)$ remain fixed over time.

Assumptions (ref) and (ref) are sufficient for identifying $\eta(\pi)$ for any given policy $\pi \in \mathcal{P}$. We briefly review standard estimation methods. The first method involves analyzing the aggregate mean return via the $Q$-function, defined as

equation[equation omitted — 152 chars of source]

The value function can be expressed as \[

aligned& \eta(\pi) = \mathrm{E}^{\sim \pi} \Bigg[ \mathrm{E}^{\sim \pi} \bigg[ \sum_{t = 0}^{+\infty} \gamma^t R_t \mid A_0, S_0 \bigg] \Bigg] = \int Q(a_0, s_0 ; \pi) \pi(a_0 \mid s_0) f(s_0) \mathrm{d} a_0 \, \mathrm{d} s_0, \\

\] where $f(s_0)$ is the initial state density. We also define the value function $V(s; \pi) = \int Q(a, s; \pi) \pi(a \mid s) \, \mathrm{d} a$.

The second method is the marginal importance sampling (MIS) estimator, which addresses the curse of horizon. The marginal ratio is defined as

equation[equation omitted — 286 chars of source]

where $f_{\sim \pi, t}$ and $f_{+ \infty}$ denote the time-dependent and stationary densities, respectively. Under stationarity, we have the identity \[ \eta(\pi) = \frac{1}{1 - \gamma} \mathrm{E} \big[ \omega(A, S; \pi) \mathrm{E}^{\sim \pi} [R \mid S]\big]. \] Crucially, both the $Q$-function (ref) and MIS ratio (ref) are required for the semiparametrically efficient estimation of $\eta(\pi)$.

Characterization and Issues

Let the trajectory $O = O_{1:T} \sim P_0 \in \mathcal{M}$. We define the functional $\Psi^*: \mathcal{M} \rightarrow \mathbb{R}$ as \[ \Psi^*(P) := \mathrm{E}_{P} \big[ Q(P) \big(A_0, S_0; \pi^*(P) \big)\pi^*(P)(A_0\mid S_0)\big], \] where $ \pi^*(P)(\cdot \mid s) = \arg\max_{\pi \in \mathcal{P}} Q(P)(a, s; \pi)$ is the optimal policy under the law $P$. Thus, $\Psi^*(P_0) = \eta(\pi^*)$ is well-defined. We focus on the value function $\Psi^*(P_0)$; discussions regarding $\pi^*(P_0)$ itself can be found in kosorok2019precision,athey2021policy,luo2024policy. Define the auxiliary functional \[ \Psi(P; \pi) = \mathrm{E}_{P} \big[ Q(P) \big(A_0, S_0; \pi \big)\pi(A_0\mid S_0)\big]. \] While $\Psi\big(P; \pi^*(P)\big) = \Psi^*(P)$, in general $\Psi\big(P_1; \pi^*(P_2)\big) \neq \Psi^*(P_1)$ if $P_1 \neq P_2$.

Assume $\{ S_t\}_{t \geq 0}$ is stationary. As shown in uehara2020minimax,shi2024off,shi2021deeply, for any fixed $\pi$, the efficient influence function (EIF) for $\Psi(P; \pi)$ at $P_0$, evaluated at $O$ in the full nonparametric space $\mathcal{M}_{\text{nonpar}}$, is

equation[equation omitted — 523 chars of source]

assuming $O \sim P_0$. We denote the estimating functions for a single point $O$ and the trajectory $O_{0:T}$ as

equation[equation omitted — 461 chars of source]

The EIF for $\eta(\pi)$ satisfies \[ S^{\text{eff, nonpar}}_{\eta(\pi)}(O; Q, \omega, V, \pi) = {\psi}^{\text{point}}_{\eta(\pi)}(O; Q, \omega, V, b, \pi) - \eta(\pi), \] and under stationarity, \[ \mathrm{E} \big[ S^{\text{eff, nonpar}}_{\eta(\pi)}(O; Q, \omega, V, \pi) \big] = \mathrm{E} \big[ {\psi}^{\text{traj}}_{\eta(\pi)}(O_{0:T}; Q, \omega, V, \pi) \big] - \eta(\pi). \]

For any regular asymptotically linear (RAL) estimator $\widehat{\Psi}(\widehat{P}; \pi)$, there exists a unique influence function tsiatis2006semiparametric such that \[ \widehat{\Psi}(\widehat{P}; \pi) - \Psi(P_0; \pi) = \frac{1}{\sqrt{N}}\sum_{i=1}^{N}\operatorname{IF}(P_0; \pi)(O_i) + o_{P_0} \big( (N - \ell_N )^{-1 / 2}\big). \] The EIF $S^{\text{eff, nonpar}} \{ \Psi(P; \pi)\}\big|_{P = P_0}(O)$ is the unique influence function that minimizes the variance $\operatorname{var}_{P_0} \big( \operatorname{IF}(P_0; \pi)(O) \big)$.

Geometrically, the EIF is characterized via the tangent space $\mathcal{T}$. For any differentiable path $\{ P_{\epsilon}: \epsilon \in \mathbb{R}\} \subsetneq \mathcal{M}$ passing through $P_0$ at $\epsilon = 0$, the pathwise differentiability of $\Psi(P; \pi)$ implies $ \frac{\mathrm{d}}{\mathrm{d} \epsilon} \Psi(P_{\epsilon}; \pi) \Big|_{\epsilon = 0} = \mathrm{E}_{P_0} \Big[ \widetilde{\psi}(O) \dot{\ell} (O) \Big] $ by the Riesz representation theorem, where $\dot{\ell}(O)$ is the score function. When $\widetilde{\psi}(O) \in \mathcal{T}$, then $\widetilde{\psi}(O) = S^{\text{eff, nonpar}} \{ \Psi(P; \pi)\}\big|_{P = P_0}(O)$, and

equation[equation omitted — 242 chars of source]

However, for the optimal value functional $\Psi^*(P) = \Psi \big(P; \pi^*(P)\big)$, (ref) may fail. The optimal policy $\pi^*(P_{\epsilon})$ along the path may contribute to the derivative if (i) $\Psi \big(P; \pi \big)$ is sensitive to $\pi$, and (ii) $\pi^*(P)$ is sensitive to $P$. By the chain rule of Gateaux differentials shapiro1990concepts: \[

aligned& \mathrm{E}_{P_0} \Big[ S^{eff, nonpar} \{ \Psi^*(P)\}\big|_{P = P_0}(O) \dot{\ell}_{eff} (O)\Big] = \frac{\mathrm{d}}{\mathrm{d} \epsilon} \Psi\big(P_{\epsilon}; \pi^*(P_{\epsilon})\big) \Big|_{\epsilon = 0} \\ = & \mathrm{E}_{P_0} \Big[ S^{eff, nonpar} \{ \Psi(P; \pi)\}\big|_{P = P_0, \pi = \pi^*(P_0)}(O) \dot{\ell}_{eff} (O) \Big] + \mathds{D}\Big|_{\frac{\mathrm{d}}{\mathrm{d} \epsilon} \pi^*(P_{\epsilon}) |_{\epsilon = 0}} \Psi\big(P_{0}; \pi^*(P_{\epsilon})\big) \Big|_{\epsilon = 0}.

\] If the second term is non-zero, then $ S^{\text{eff, nonpar}} \{ \Psi^*(P)\}\big|_{P = P_0} \neq S^{\text{eff, nonpar}} \{ \Psi(P; \pi)\}\big|_{P = P_0, \pi = \pi^*(P_0)}. $ We establish a concise expression for $ S^{\text{eff, nonpar}} \{ \Psi^*(P)\}$ in the next section.

Existence of EIF and Its Expression

We present regularity conditions on the distributions of $S$, $A$, and $R$ to ensure the EIF exists. Let $\mathcal{S}$ and $\mathcal{A}$ denote the supports of the states and actions, respectively.

assumption[Regularity] (States) The state space $\mathcal{S}$ is compact, and $S_0$ is not a point mass; (Actions) The action space $\mathcal{A}$ is compact; (Policies): There exist positive constants $\underline{c}_{\pi}$ and $\overline{c}_{\pi}$ such that $\underline{c}_{\pi} \leq \inf_{\pi \in \Pi} \inf_{(a, s) \in \mathcal{A} \times \mathcal{S}} \pi(a \mid s) \leq \sup_{\pi \in \Pi} \|\pi\|_{\infty} \leq \overline{c}_{\pi}$. Furthermore, if either $\mathcal{A}$ or $\mathcal{S}$ is not finite, then for any policy $\pi \in \mathcal{P}$, $\pi(a \mid s)$ is lower semicontinuous in the argument corresponding to the non-finite space(s); (Rewards): $R$ is bounded.

Assumption (ref) contains standard conditions levine2020offline,uehara2022review,shi2022statistical, with the exception of the policy bounds, which are nonetheless mild and standard in OPE xu2021doubly,shi2024off,bian2024off. Our first result shows that when the optimal policy is unique and deterministic, the EIF $S^{\text{eff, nonpar}} \{ \Psi^*(P)\}$ exists and equals $S^{\text{eff, nonpar}} \{ \Psi(P; \pi)\}$ at $\pi = \pi^*$.

assumption[Unique Deterministic Optimal Policy] The optimal policy $\pi^* \in \Pi$ is unique and deterministic, satisfying $ \pi^*(P) (a \mid s) = \mathds{1} \big\{ a = \arg \max_{a' \in \mathcal{A}} Q(P) (a', s; \pi^*) \big\} $ for $s \in \mathcal{S}$.
theoremSuppose that Assumptions (ref), (ref), (ref), and (ref) hold. Then the efficient influence function of $\Psi^*$ exists and satisfies $ S^{\text{eff, nonpar}} \{ \Psi^*(P)\}\big|_{P = P_0} = S^{\text{eff, nonpar}} \{ \Psi(P; \penalty 0 \pi)\}\big|_{P = P_0, \pi = \pi^*(P_0)}. $

Conversely, if the optimal policy is not unique—specifically, if a significant set of states exists where multiple optimal actions are indifferent—the influence function does not exist.

assumption[Unrestricted Optimal Rules] There exists at least one $\Pi \ni \pi^{\star} \neq \pi^* $ such that $ Q(P) \big(a, s; \pi^{\star}(P) \big) = Q(P) \big(a, s; \pi^*(P) \big) = \max_{\pi \in \mathcal{P}} Q(P) \big(a, s; \pi \big) $ and $ \mu \big\{ s \in \mathcal{S}: \mu \{ a \in \mathcal{A}: \pi^{\star}(a \mid s) \neq \pi^{*}(a \mid s) \} > 0 \big\} > 0. $

Assumption (ref) describes Unrestricted Optimal Rules robins2014discussion, leading to non-regularity where standard margin conditions shi2022statistical fail.

theoremSuppose that Assumptions (ref), (ref), (ref), and (ref) hold. Then $\Psi^*(P)$ does not have any influence function at $P = P_0$.

Estimation Under Possible Non-Uniqueness

The Challenge and Current Gap

When $\pi$ is fixed, as studied thoroughly in the literature, the one-step estimator is obtained by solving the estimating equation $\mathbb{P}_{NT} \big\{ S^{\text{eff, nonpar}}_{\eta(\pi)}(O; \widehat{Q}, \widehat{\omega}, \widehat{V}, \widehat{b}, \pi) \big\} = 0$, which yields

equation*[equation* omitted — 383 chars of source]

with estimated nuisance functions $\widehat{Q}, \widehat{\omega}, \widehat{V}, \widehat{b}$. Here, $\widehat{\eta}_{\text{DR, 1}}(\pi)$ and $\widehat{\eta}_{\text{DR, 2}}(\pi)$ are asymptotically equivalent and share desirable statistical properties:

itemize• Double-Robustness: Assuming either $\widehat{Q}$ (and thus $\widehat{V}$) or $\widehat{\omega}$ is consistent, both $\widehat{\eta}_{\text{DR, 1}}(\pi)$ and $\widehat{\eta}_{\text{DR, 2}}(\pi)$ are consistent estimators for $\eta(\pi)$. • Semiparametric Efficiency: If both $\widehat{Q}$ and $\widehat{\omega}$ are $o(N^{1/4})$-consistent, $\widehat{\eta}_{\text{DR, 1}}(\pi)$ and $\widehat{\eta}_{\text{DR, 2}}(\pi)$ achieve semiparametric efficiency, satisfying $ \sqrt{N} \big( \widehat{\eta}_{\text{DR, j}}(\pi) - \eta(\pi) \big) \rightsquigarrow \mathcal{N} \big( 0, \penalty 0 \mathrm{E} \big[ S^{\text{eff, nonpar}}_{\eta(\pi)}(O; Q, \omega, V, b, \pi) \big]^2 \big) $ for $j = 1, 2$.

When the optimal policy (or policies) $\pi^*$ is unknown, and we have an estimated optimal policy $\widehat{\pi}$ computed by a $Q$-learning type algorithm such that \[ \widehat{\pi}(a \mid s) := \mathds{1} \Big\{ a = \arg \max_{a' \in \mathcal{A}} \widehat{Q}_{\text{opt}} (a, s ) \Big\}, \] where $\widehat{Q}_{\text{opt}} (a, s ) $ denotes a consistent estimator for the optimal $Q$-function, i.e., $Q(a, s; \pi^* )$, an intuitive “plug-in" estimator for the parameter of interest $\eta(\pi^*)$ is $\widehat{\eta}_{\text{DR, j}}(\widehat{\pi})$. The above double-robustness and semiparametric efficiency property still holds as a standard result established in the literature uehara2022review. However, as we pointed out in Theorem (ref), when the optimal policy is not unique, there is no basis for discussing semiparametric efficiency, as RAL estimators do not exist. More importantly, in such cases, the argmax of the optimal $Q$-function may not be uniquely defined; consequently, $\widehat{\pi}(\cdot \mid s)$ might not converge to a fixed quantity for some $s \in \mathcal{S}$. As a result, the plug-in estimator $\widehat{\eta}_{\text{DR, j}}(\widehat{\pi})$ will fluctuate randomly and fail to maintain a stable limiting distribution shi2022statistical.

To overcome this issue, shi2022statistical proposed a new estimator called SequentiAl Value Evaluation (SAVE), denoted as $\widehat{\eta}_{\text{SAVE}}$, assuming that the $Q$-function follows a linear sieve model such that $Q (a, s; \pi) \approx \Phi^{\top}(s) \bm{\beta}_{\pi, a}$, where $\Phi(s)$ is a vector of sieve basis functions. This novel estimator enjoys bidirectional asymptotic normality: \[

aligned\sqrt{NT (K - 1) / K} \widehat{\sigma}_{SAVE}^{-1}\big( \widehat{\eta}_{SAVE} - \eta (\widehat{\pi})\big) \, \rightsquigarrow \, \mathcal{N}(0, 1); \\ \sqrt{NT (K - 1) / K} \widehat{\sigma}_{SAVE}^{-1}\big( \widehat{\eta}_{SAVE} - \eta ({\pi}^*)\big) \, \rightsquigarrow \, \mathcal{N}(0, 1),

\] as either $N \rightarrow \infty$ or $T \rightarrow \infty$, where $K$ is the number of data partitions and $\widehat{\sigma}_{\text{SAVE}}^2$ is a “plug-in" type variance estimator. Thus, the readily applicable estimator $\widehat{\eta}_{\text{SAVE}}$ can be used for statistical inference.

Nonetheless, despite its appealing theoretical guarantees, the SAVE estimator $\widehat{\eta}_{\text{SAVE}}$ suffers from several significant limitations. First, SAVE relies critically on a linear structural assumption for the $Q$-function, namely that $Q(s,a;\pi)$ can be well-approximated by a low-dimensional linear form $Q(s,a;\pi)\approx \Phi^\top(s)\boldsymbol{\beta}_{\pi,a}$. This assumption may be violated in many realistic sequential decision problems. Second, when the optimal policy is unique and deterministic, $\widehat{\eta}_{\text{SAVE}}$ no longer admits a doubly robust representation and consequently loses both the double robustness property and the associated semiparametric efficiency guarantees. Third, and most critically, SAVE requires strong non-degeneracy conditions on the target policy. In particular, its inference theory implicitly relies on well-conditioned Bellman estimating equations, which may fail when the target policy is deterministic or nearly deterministic. In such cases, the feature covariance induced by the target policy becomes nearly singular, leading to unstable estimation and invalid uncertainty quantification. Consequently, SAVE does not provide a principled fallback inference procedure once these marginal conditions are violated. In Section (ref), we demonstrate through simulation that this issue is not merely theoretical: under deterministic or highly concentrated target policies, SAVE can exhibit severely distorted coverage, whereas our proposed method remains stable and valid.

To address these three challenges, we adopt the conceptual framework of shi2022statistical while introducing a revised sequential value evaluation procedure and a corresponding estimator, which we term Nonparametric SequentiAl Value Evaluation (NSAVE).

Nonparametric SequentiAl Value Evaluation Approach

Assume $\{ O_{\tau(i)}\}_{i = 1}^N$ is a random permutation of the original i.i.d. trajectory observations $\{ O_i\}_{i = 1}^N$. Let $\{\ell_N\}$ be a sequence of non-negative integers representing the size of the initial sample used to estimate the initial optimal policy from the estimated $Q$-function, denoted by $\widehat{\pi}_{\tau(\ell_N - 1)}^{(0)}$. We initialize such that $\widehat{\pi}_{\tau(\ell_N - 1)}^{(Q)} = \widehat{\pi}_{\tau(\ell_N - 1)}^{(\omega)} = \widehat{\pi}_{\tau(\ell_N - 1)}^{(0)}$.

For $j = \ell_N + 1, \ldots, N$, we perform the following steps:

itemize• Optimizing: Using $\widehat{Q}_{\tau(j - 2)} (\cdot, \cdot ; \cdot)$, we obtain the $Q$-based estimated optimal policy $\widehat{\pi}_{\tau(j - 1)}^{(Q)}(a \mid s)$ as \[ \widehat{\pi}_{\tau(j - 1)}^{(Q)}(a \mid s) := \mathds{1} \Big\{ a = \arg \max_{a' \in \mathcal{A}} \widehat{Q}_{\tau(j - 1)} \big(a, s; \widehat{\pi}_{\tau(j - 2)}^{(Q)} \big) \Big\}, \] and compute the estimated marginal ratio under the optimal $\omega$-based estimated optimal policy, denoted as $\widehat{\omega}_{\tau(j - 1)} (\cdot, \cdot ; \widehat{\pi}_{\tau(j - 1)}^{(\omega)})$. • Training: Using the data up to the previous step, i.e., $\{ O_{\tau(i)}\}_{i \leq j - 1}$, we estimate the $Q$ nuisance functions in (ref), i.e., $\widehat{Q}_{\tau(j - 1)} (\cdot, \cdot ; \cdot)$. The value function $\widehat{V}_{\tau(j - 1)}(\cdot; \cdot)$ can then be obtained directly. • Evaluating: Using the above nuisance functions, the estimated trajectory estimating functional in the EIF is calculated as \[ \begin{aligned} \widehat{\psi}^{\text{traj-step}}_{\tau(j)} &:= \widehat{\psi}^{\text{traj-step}}_{\tau(j)} \{ \Psi^* \}(O_{\tau(j), t}) \\ & = \sum_{t = 0}^T \gamma^t \widehat{\omega}_{\tau(j - 1)}\big(A_{\tau(j), t}, S_{\tau(j), t}; \widehat{\pi}_{\tau(j - 1)}^{(\omega)}\big) \big[ R_{\tau(j), t} + \gamma \widehat{V}_{\tau(j - 1)} \big(S_{\tau(j), t + 1}; \widehat{\pi}_{\tau(j - 1)}^{(Q)} \big)\\ &~~~~~~~~~~~~~~ - \widehat{Q}_{\tau(j - 1)} \big(A_{\tau(j), t}, S_{\tau(j), t}; \widehat{\pi}_{\tau(j - 1)}^{(Q)} \big)\big] + \widehat{V}_{\tau(j - 1)} \big(S_{\tau(j), 0}; \widehat{\pi}_{\tau(j - 1)}^{(Q)} \big). \end{aligned} \]

It is worth noting that there is no need to explicitly use or estimate the $\omega$-based estimated optimal policy, defined as $ \widehat{\pi}_{\tau(j - 1)}^{(\omega)} (a \mid s) := \arg\max_{a \in \mathcal{A}} \widehat{\omega}_{\tau(j - 1)} (a, s ; \widehat{\pi}_{\tau(j - 1)}^{(\omega)}). $ This $\omega$-based estimated optimal policy is primarily introduced for notational convenience, emphasizing that the sequential marginal ratio is derived from the Optimizing Step.

Define the online one-step variance as \[

aligned\widetilde{\sigma}_{\tau(j - 1)}^2 : = \operatorname{var} \Big( S^{eff, nonpar} \{ \Psi \}\big(O; \widehat{Q}_{\tau(j - 1)}, \widehat{\omega}_{\tau(j - 1)}, \widehat{\pi}_{\tau(j - 1)}^{(Q)}, \widehat{\pi}_{\tau(j - 1)}^{(\omega)}) \big) \mid \{ O_{\tau(i)}\}_{i \leq j - 1} \Big).

\] Here, the function $S^{\text{eff, nonpar}} \{ \Psi \}$ is regarded purely as a functional of the observation $O$ and the nuisance functions $\{Q, \omega, \pi \}$ (independent of the distribution $P$), regardless of whether the optimal policies are unique. Let $\widehat{\sigma}_{\tau(j - 1)}^2$ denote its corresponding consistent estimator. In practice, $\widehat{\sigma}_{\tau(j - 1)}^2$ can be estimated using the sample variance over a specific sliding window, such as $\big\{ \widehat{\psi}^{\text{traj-step}}_{\tau(j - m)}, \widehat{\psi}^{\text{traj-step}}_{\tau(j - m + 1)} , \ldots, \widehat{\psi}^{\text{traj-step}}_{\tau(j - 1)} \big\}$, for a sufficiently large $m$. Then, our final estimator is similarly defined as the weighted average: \[

aligned\widehat{\eta}_{NSAVE} := \bigg\{ \sum_{j = \ell_N + 1}^N \frac{1}{\widehat{\sigma}_{\tau(j - 1)}}\bigg\}^{-1} \sum_{j = \ell_N + 1}^N \frac{\widehat{\psi}^{traj-step}_{\tau(j)}}{\widehat{\sigma}_{\tau(j - 1)}}.

\] Intuitively, our novel estimator $\widehat{\eta}_{\text{NSAVE}}$ approximates, but is distinct from, the average weighted empirical historical value $\overline{\eta}_w(\widehat{\pi}^{(Q)})$, defined as \[ \overline{\eta}_{w}(\widehat{\pi}^{(Q)}) := \bigg\{ \sum_{j = \ell_N + 1}^N \frac{1}{\widehat{\sigma}_{\tau(j - 1)}}\bigg\}^{-1} \sum_{j = \ell_N + 1}^N \frac{\eta ({\widehat{\pi}_{\tau(j - 1)}^{(Q)})}}{\widehat{\sigma}_{\tau(j - 1)}}. \]

In the following, we analyze the theoretical properties of our novel estimator $\widehat{\eta}_{\text{NSAVE}}$ to demonstrate its advantages. We assume that $\mathcal{S}$ and $\mathcal{A}$ are finite. Furthermore, we assume that the function classes for both $Q$ and $\omega$ are uniformly bounded Donsker classes.

Nuisance Estimation Approaches

There are various approaches for estimating the nuisance components. Here, we mainly focus on the estimation of $\widehat{\omega}_{\tau(j - 1)} (\cdot, \cdot ; \widehat{\pi}_{\tau(j - 1)}^{(\omega)})$, while the estimation of other nuisance components, such as obtaining $\widehat{Q}_{\tau(j - 1)} (\cdot, \cdot ; \cdot)$ and $\widehat{b}(\cdot \mid \cdot)$, has been discussed in detail in the literature shi2025statistical.

The most common strategy is using dual linear programming. Specifically, the true nuisance $\omega(\cdot, \cdot ; \pi^*)$ is the solution to the following maximization problem: \[ \omega(a, s ; \pi^*) = \arg\max_{\omega \in \Omega_{\text{flow}}} \mathrm{E}_{P_0} \big[\omega(A; S) R\big], \] where $\Omega_{\text{flow}}$ is the polytope of valid density ratios satisfying the Bellman flow constraints, i.e., \[

aligned\Omega_{flow} := \bigg\{ \omega(a, s) \in P(a, s) : & \sum_{a \in \mathcal{A}} \omega(a, s') b(a \mid s') f_0(s') = (1 - \gamma) f_0(s') \\ & + \gamma \sum_{(a, s) \in \mathcal{A} \times \mathcal{S}} f(s' \mid a, s) \omega(a, s) b(a \mid s) f_0(s) \quad for any s \in \mathcal{S}\bigg\}.

\] Let $\widehat{\Omega}_{\text{flow}}$ be the corresponding estimated $\Omega_{\text{flow}}$ obtained by replacing the unknown nuisance functions $b(a \mid s)$, $f_0(s)$, and $f(s' \mid a, s)$ with their estimated counterparts. Then, a practical estimator $\omega(a, s ; \pi^*)$ at Step $j$ can be defined as \[

aligned\widehat{\omega}_{\tau(j - 1)} (\cdot, \cdot ; \widehat{\pi}_{\tau(j - 1)}^{(\omega)}) := \arg\max_{\omega \in \widehat\Omega_{flow}} \frac{1}{T (j - 1)} \sum_{t = 1}^T \sum_{\iota = 1}^{j - 1} \omega(A_{\tau(\iota), t}, S_{\tau(\iota), t}) R_{\tau(\iota), t}.

\]

Another important approach is Minimax Weight Learning (MWL), which constructs two nuisance estimators for $\big( Q (\cdot, \cdot; \pi^*), \, \omega(\cdot, \cdot; \pi^*)\big)$ simultaneously nachum2019dualdice,duan2020minimax,uehara2020minimax. We adapt their idea and extend this minimax framework to the estimation of the optimal policy value, where the target policy itself is data-dependent and potentially non-unique. To do this, we define the Lagrangian function \[

aligned\mathcal{L} (Q^{\pi}, \omega^{\pi}; P_0) & : = (1 - \gamma) \mathrm{E}_{P_0} \big[ Q(A, S; \pi)\big] \\ & + \mathrm{E}_{P_0} \Big[ \omega(A, S ; \pi) \Big\{ R + \gamma \mathrm{E}_{P_0} \big[ Q (A, S' ; \pi)\big] - Q(A, S ; \pi) \Big\} \Big].

\] Let $\mathbb{P}_{\tau(j - 1) T} := \frac{1}{T (j - 1)} \sum_{t = 1}^T \sum_{\iota = 1}^{j - 1} [\boldsymbol{\cdot}]_{\tau(\iota), t}$ be the empirical distribution measure at Step $j$. We can then construct the estimated $Q$-function and $\omega$-function under the (estimated) policies at Step $j$ by solving the following minimax problem:

equation[equation omitted — 330 chars of source]

We then set $\widehat{\omega}_{\tau(j - 1)} (\cdot, \cdot ; \widehat{\pi}_{\tau(j - 1)}^{(\omega)}) = \widehat{\omega}_{\text{opt}, \tau(j - 1)} (\cdot, \cdot)$, while one may choose whether to use $\widehat{Q}_{\text{opt}, \tau(j - 1)}$ as $\widehat{Q}_{\tau(j - 1)} (\cdot, \cdot ; \widehat{\pi}_{\tau(j - 1)}^{(Q)})$.

Inference

Although it is intuitively plausible that our novel estimator $\widehat{\eta}_{\text{NSAVE}}$ will be consistent as long as the nuisance components are consistently estimated (we will also formally demonstrate its consistency), such intuition is insufficient for inference. The latter typically requires stronger conditions. To characterize the specific requirements, consider that the remainder term can be decomposed into two distinct components corresponding to two different inferential strategies:

itemize• For the Conservative Lower Bound: \[ \begin{aligned} \widehat{\eta}_{\text{NSAVE}} - \eta(\pi^*) = \underbrace{\widehat{\eta}_{\text{NSAVE}} - \overline{\eta}_w(\widehat{\pi}^{(Q)})}_{=: R_{\text{CLB}, 1N}} + \underbrace{\overline{\eta}_w(\widehat{\pi}^{(Q)}) - \eta(\pi^*)}_{=: R_{\text{CLB}, 2N}}. \end{aligned} \] This is a relatively coarse decomposition, as our primary goal here is to establish a lower bound. It is straightforward to see that $R_{\text{CLB}, 1N}$ represents the empirical error for the average value functions under the estimated policies, while $R_{\text{CLB}, 2N}$ represents the cumulative regret arising from the estimated policies. • For the Two-Sided Confidence Interval: \[ \begin{aligned} \widehat{\eta}_{\text{NSAVE}} - \eta(\pi^*) = \underbrace{\widehat{\eta}_{\text{NSAVE}} - \eta({\widehat{\pi}_{\tau(N)}^{(Q)}})}_{=: R_{\text{TCI}, 1N}} + \underbrace{\eta({\widehat{\pi}_{\tau(N)}^{(Q)}}) - \eta(\pi^*)}_{=: R_{\text{TCI}, 2N}}. \end{aligned} \] Here, we use a more refined decomposition consistent with standard analyses: $R_{\text{TCI}, 1N}$ represents the statistical error, and $R_{\text{TCI}, 2N}$ captures the policy-value error.

Conservative Lower Bound

In both decompositions, the second terms, $R_{\text{CLB}, 2N}$ and $R_{\text{TCI}, 2N}$, are non-positive by the definition of $\pi^*$. Consequently, we have \[

aligned\eta(\pi^*) = \left\{ \begin{array}{cc} \widehat{\eta}_{NSAVE} - \big( R_{CLB, 1N} + R_{CLB, 2N} \big) & \geq \widehat{\eta}_{NSAVE} - R_{CLB, 1N} \\ \widehat{\eta}_{NSAVE} - \big( R_{\text{TCI}, 1N} + R_{\text{TCI}, 2N} \big) & \geq \widehat{\eta}_{\text{NSAVE}} - R_{\text{TCI}, 1N} \end{array}\right. .

\] If we can construct a valid $(1 - \alpha)$ upper bound $\operatorname{UB}(R_{1N}; \alpha)$ for either $R_{\text{CLB}, 1N}$ or $R_{\text{TCI}, 1N} $ such that \[ \liminf_{N \rightarrow \infty} \mathrm{P} \big(R_{\text{CLB}, 1N} \leq \operatorname{UB}(R_{1N}; \alpha) \big) \geq 1 - \alpha ~~\text{ or } ~~\liminf_{N \rightarrow \infty} \mathrm{P} \big( R_{\text{TCI}, 1N} \leq \operatorname{UB}(R_{1N}; \alpha) \big) \geq 1 - \alpha, \] then \[

aligned\liminf_{N \rightarrow \infty} \mathrm{P} \big( \eta(\pi^*) \geq \widehat{\eta}_{NSAVE} - \operatorname{UB}(R_{1N}; \alpha) \big) \geq 1 - \alpha.

\] This implies that $\widehat{\eta}_{\text{NSAVE}} - \operatorname{UB}(R_{1N}; \alpha)$ serves as a valid lower bound for the optimal value. The following theorem formally states how to construct a valid $\operatorname{UB}(R_{1N}; \alpha)$ and its corresponding estimator $\widehat{\operatorname{UB}}(R_{1N}; \alpha)$. Let $ \sigma_{R_{1N}} := \frac{1}{\sqrt{N - \ell_N}} \left\{ \sum_{j = \ell_N + 1}^N \frac{1}{\widetilde{\sigma}_{\tau(j - 1)}}\right\} $ and \[ {\psi}^{\text{traj, *}}_{\eta(\pi), \tau(j)} ( \cdot, \cdot, \cdot, \cdot ) := {\psi}^{\text{traj}}_{\eta(\pi)} (O_{0:\tau(j)} \, ; \, \cdot, \cdot, \cdot, \cdot ) - \mathrm{E} \big[ {\psi}^{\text{traj}}_{\eta(\pi)} (O_{0:\tau(j)} \, ; \, \cdot, \cdot, \cdot, \cdot ) \mid \sigma \langle O_{0:\tau(j - 1)} \rangle \big]. \] The upcoming theorem, which establishes the asymptotic normality for the first terms in the two types of decompositions, relies on the following assumptions:

assumption[Convergence Rates for Nuisance Parameters] $\widehat{Q}_{\tau(j - 1)} (\cdot, \cdot ; \widehat{\pi}_{\tau(j - 1)}^{(Q)})$ and $\widehat{\omega}_{\tau(j - 1)}(\cdot, \cdot ;\widehat{\pi}_{\tau(j - 1)}^{(\omega)})$ are $j^{\kappa_Q}$-consistent estimators of $Q(\cdot, \cdot ; \widehat{\pi}_{\tau(j - 1)}^{(Q)})$ and $j^{\kappa_\omega}$-consistent estimators of {$\omega(\cdot, \cdot ; \widehat{\pi}_{\tau(j - 1)}^{(Q)})$} such that $\kappa_Q + \kappa_\omega \geq 1 / 2$.
assumption[Non-Zero Variances] $\inf_{j > \ell_N} \widehat{\sigma}_{\tau(j - 1)} > \sigma_0$ and $\inf_{j > \ell_N} \widetilde{\sigma}_{\tau(j - 1)} > \sigma_0$ for some $\sigma_0 > 0$.
assumption[Conditions for Estimated Variances] $\frac{1}{(N - \ell_N)} \sum_{j = \ell_N + 1}^N \mathrm{E} \left[\left( \frac{\widehat{\sigma}_{\tau(j - 1)}}{\widetilde{\sigma}_{\tau(j - 1)}} - 1 \right)^2 \right] = o_{P_0}(1)$.
assumption[Lindeberg Condition] For any $\epsilon > 0$, \[ \begin{aligned} \sum_{j = \ell_N + 1}^N \mathrm{E} \Bigg[ \bigg( \frac{{\psi}^{\text{traj, *}}_{\eta(\pi), \tau(j)}}{\sqrt{N - \ell_N} \widetilde{\sigma}_{ \tau(j - 1)}} \bigg)^2 \mathds{1} \bigg\{ \bigg| \frac{{\psi}^{\text{traj, *}}_{\eta(\pi), \tau(j)}}{\sqrt{N - \ell_N} \widetilde{\sigma}_{ \tau(j - 1)}} \bigg| > \epsilon \bigg\} \Bigg] = o_{P_0} (1). \end{aligned} \]
theoremSuppose that Assumptions (ref), (ref), and (ref) hold. In addition, assume that Assumptions (ref)--(ref) hold. Then \[ \begin{aligned} \sigma_{R_{1N}}^{-1} R_{\text{CLB}, 1N} \rightsquigarrow \mathcal{N} \big( 0, 1\big) \end{aligned} \] as $N \to \infty$.

Let ${\operatorname{UB}}(R_{1N}; \alpha) := \frac{z_{\alpha} \widetilde\sigma_{R_{1N}}^{-1}}{\sqrt{N - \ell_N}}$ and $\widehat{\operatorname{UB}}(R_{1N}; \alpha) := \frac{z_{\alpha} \widehat\sigma_{R_{1N}}^{-1}}{\sqrt{N - \ell_N}}$ with $ \widehat\sigma_{R_{1N}} := \frac{1}{\sqrt{N - \ell_N}} \left\{ \sum_{j = \ell_N + 1}^N \frac{1}{\widehat{\sigma}_{\tau(j - 1)}}\right\}$. Then Theorem (ref) implies that $\widehat{\eta}_{\text{NSAVE}} - \widehat{\operatorname{UB}}(R_{1N}; \alpha)$ provides a readily applicable conservative lower bound for $\eta(\pi^*)$.

corollaryUnder the conditions in Theorem (ref), we have that \[ \lim_{N \rightarrow \infty} \mathrm{P}_{P_0} \Big( \eta(\pi^*) \geq \widehat{\eta}_{\text{NSAVE}} - \widehat{\operatorname{UB}}(R_{1N}; \alpha) \Big) \geq 1 - \alpha. \]

Here, we only establish asymptotic normality for the first term in the coarse decomposition. The reason is that ensuring asymptotic normality for $R_{\text{TCI}, 1N}$ requires regularity conditions for the Estimated Optimal Policies. In contrast, as can be seen from Theorem (ref) and Corollary (ref), we do not impose any conditions on the estimated policy sequence $\{ \widehat{\pi}_{\tau(j - 1)} \}_{j > \ell_N}$. Therefore, compared with the conditions in shi2022statistical, which require regularity and so-called margin conditions for both the estimated and true optimal policies, our novel estimator $\widehat{\eta}_{\text{NSAVE}}$ admits valid inference procedures without such restrictions.

Two-Sided Confidence Interval

Under stronger conditions, more accurate inference via a Two-Sided Confidence Interval for $\eta^*$ is possible. To achieve this, we first need to establish the asymptotic properties of the two terms $R_{\text{TCI}, 1N}$ and $R_{\text{TCI}, 2N}$. Here, $R_{\text{TCI}, 1N}$ represents the fluctuation of our novel estimator around the true value function evaluated at the estimated optimal policy $\widehat{\pi}_{\tau(N)}^{(Q)}$. Consequently, as the construction of $\widehat{\eta}_{\text{NSAVE}}$ is conditional on past observations, we can apply the martingale CLT to show that $R_{\text{TCI}, 1N}$ is $\sqrt{N - \ell_N}$-consistent and converges to a Gaussian distribution under conditional Lindeberg conditions and regularity conditions for the final estimated policy $\widehat{\pi}_{\tau(N)}^{(Q)}$. This two-sided technique is also applied in shi2022statistical. On the other hand, $R_{\text{TCI}, 2N}$ represents the systematic error arising from replacing the unknown optimal policy sequence with $\{ \widehat{\pi}_{\tau(j - 1)} \}_{j > \ell_N}$, which may exhibit a slower convergence rate.

assumption[{$Q$-based} Estimated Optimal Policies] $\big\| Q (\cdot, \cdot ; \widehat{\pi}_{\tau(j - 1)}^{(Q)}) - Q (\cdot, \cdot ; \pi^*) \big\|_{P_0, 2} = O_{P_0} \big((j - \ell_N)^{- \kappa_\pi} \big)$ for some $\kappa_\pi > 1 / 2$.

These additional conditions will help guarantee the CLT for $R_{\text{TCI}, 1N}$. {An important note here is that $\kappa_\pi > 1 / 2$ should NOT be regarded as a super-consistent convergence rate, as we have another dimension of sampling: the time horizon dimension $T$.}

theoremUnder the conditions in Theorem (ref) and Assumption (ref), we have \[ \begin{aligned} \sigma_{R_{1N}}^{-1} R_{\text{TCI}, 1N} \rightsquigarrow \mathcal{N} \big( 0, 1\big) \end{aligned} \] as $N \to \infty$.

Define the sub-optimality gaps as $\mathbf{\Delta} (a, s; Q, \pi) := V(s; \pi^*) - Q(a, s; \pi)$. In addition to Assumption (ref), to guarantee favorable limiting behavior for $R_{\text{TCI}, 2N}$ as a functional of the estimated policy sequence, we introduce a margin-type condition below. {Let $\mathcal{A}_{\text{sub-opt}}(s) = \mathcal{A} \, \setminus \, \arg\max_{a \in \mathcal{A}} Q(a, s; \pi^*)$.}

assumption[Margin-Type Condition] There exists some constant $\alpha > 0$ such that $ \mathrm{P}_{P_0} \big( 0 \leq {\min_{a \in \mathcal{A}_{\text{sub-opt}}(S)}} \mathbf{\Delta} (a, S; Q, \pi^*) \leq \delta \big) \lesssim \delta^{\alpha}$.
theoremUnder the conditions in Theorem (ref), Assumption (ref), and Assumption (ref), we have $ R_{\text{TCI}, 2N} = o_{P_0}\big( (N- \ell_N)^{-1 / 2} \big) $ if $\kappa_Q \geq 1 / 2$.
corollaryUnder the conditions in Theorem (ref), we have \[ \sigma_{R_{1N}}^{-1} \big( \widehat{\eta}_{\text{NSAVE}} - \eta(\pi^*) \big) \, \rightsquigarrow \, \mathcal{N} (0, 1). \] Furthermore, if Assumption (ref) holds, $\widehat{\eta}_{\text{NSAVE}}$ achieves semiparametric efficiency.

As partially shown in Corollary (ref), compared with the results in shi2022statistical, our estimator $\widehat{\eta}_{\text{NSAVE}}$ demonstrates several advantages. Specifically:

itemize• Semiparametric efficiency: Our estimator does not lose any efficiency as long as the optimal policy is uniquely defined as in Assumption (ref), whereas there is no discussion of efficiency in shi2022statistical. • Double-Robustness: Under Assumption (ref), our estimator also retains the typical double-robustness property shared by standard EIF-based estimators. Again, such a robustness property is absent in the SAVE estimator. We will formally state this advantage in the next section. • Weaker restrictions on the estimated optimal policies: The convergence rate requirement for the estimated optimal policies is the same as that in shi2022statistical, yet we do not require the associated effective sample size to be larger than a specific number inversely proportional to the convergence rate. • Weaker constraints for the margin conditions: We only require that the probability, rather than the Lebesgue measure, satisfies the margin condition.

Double-Robustness and Efficiency under Uniqueness

To explicitly state the first two advantages of our new estimator compared to the estimator in shi2022statistical when the optimal policy $\pi^*$ is deterministic and unique, we use the following theorem to characterize the theoretical asymptotic properties of $\widehat{\eta}_{\text{NSAVE}}$.

assumption[Flow Constraint] $\lim_{N \to \infty } \sup_{j > \ell_N }\mathrm{P}_{P_0} \Big\{ \widehat{\omega}_{\tau(j - 1)}\big(\cdot, \cdot \, ; \, {\widehat{\pi}_{\tau(j - 1)}^{(\omega)}} \big) \in \widehat{\Omega}_{\text{flow}}\Big\} = 1$.
assumption[Saddle Points] For $j = \ell_N + 1, \ldots, N$, there exists a constant $\kappa_{\mathcal{L}} > 0$ such that \[ \begin{aligned} & \mathbb{D}_Q \mathcal{L} \big( \widehat{Q}_{\tau(j - 1)} (\cdot, \cdot ; \widehat{\pi}_{\tau(j - 1)}^{(Q)}) \, , \widehat{\omega}_{\tau(j - 1)} (\cdot, \cdot ; \widehat{\pi}_{\tau(j - 1)}^{(\omega)}) \, ; \, \mathbb{P}_{\tau(j - 1) T} \big) = O_{P_0}(j^{-\kappa_{\mathcal{L}}}) \\ \text{and} \qquad & \mathbb{D}_{\omega} \mathcal{L} \big( \widehat{Q}_{\tau(j - 1)} (\cdot, \cdot ; \widehat{\pi}_{\tau(j - 1)}^{(Q)}) \, , \widehat{\omega}_{\tau(j - 1)} (\cdot, \cdot ; \widehat{\pi}_{\tau(j - 1)}^{(\omega)}) \, ; \, \mathbb{P}_{\tau(j - 1) T} \big) = O_{P_0}(j^{-\kappa_{\mathcal{L}}}). \end{aligned} \]
theoremAssume that the conditions in Theorem (ref) hold. In addition, assume that Assumptions (ref)--(ref), and Assumptions (ref)--(ref) hold. \begin{itemize} • (Double Robustness) Assume either of the following conditions holds: (i) $\widehat{Q}_{\tau(j - 1)} \big(\cdot, \cdot ; \widehat{\pi}_{\tau(j - 1)}^{(Q)} \big)$ is a consistent estimator of $Q(\cdot, \cdot ; \pi^*)$; (ii) $\widehat{\omega}_{\tau(j - 1)}\big(\cdot, \cdot ; {\widehat{\pi}_{\tau(j - 1)}^{(\omega)}} \big)$ is a consistent estimator of $\omega(\cdot, \cdot ; \pi^*)$. Then our proposed estimator $\widehat{\eta}_{\text{NSAVE}}$ is consistent for $\eta^* = \eta(\pi^*)$. • (Semiparametric Efficiency) Assume that for some $\kappa > 1 / 4$, both of the following conditions hold: (i) $\widehat{Q}_{\tau(j - 1)} (\cdot, \cdot ; \widehat{\pi}_{\tau(j - 1)} )$ is a $j^{\kappa}$-consistent estimator of $Q(\cdot, \cdot ; \pi^*)$; (ii) $\widehat{\omega}_{\tau(j - 1)}(\cdot, \cdot ;\widehat{\pi}_{\tau(j - 1)})$ is a $j^{\kappa}$-consistent estimator of $\omega(\cdot, \cdot ; \pi^*)$. Then our proposed estimator $\widehat{\eta}_{\text{NSAVE}}$ satisfies $\sqrt{N} \big(\widehat{\eta}_{\text{NSAVE}} - \eta^* \big) \rightsquigarrow \mathcal{N} \big( 0, \mathrm{E} \big[ S^{\text{eff, nonpar}} \{ \Psi \} (P_0)\big]^2 \big)$ given Assumption (ref). \end{itemize}

{Assumptions (ref) and (ref) are quite mild. In particular, the MWL estimator from (ref) would naturally satisfy these two assumptions, as the two Gateaux differentials are exactly zero, and the probability for the flow constraint would also be exactly one.} Theorem (ref) explicitly states the advantages of our novel estimator $\widehat{\eta}_{\text{NSAVE}}$: Compared with the naive “plug-in" estimators $\widehat{\eta}_{\text{DR, j}}(\widehat{\pi})$, our estimator can adapt to potentially non-unique optimal policies; Compared with $\widehat{\eta}_{\text{SAVE}}$, $\widehat{\eta}_{\text{NSAVE}}$ retains both double-robustness and efficiency when the optimal policy is unique and deterministic.

Alternative Inference Approaches

Smoothing

As recently proposed by whitehouse2025inference, the non-differentiability inherent in $\eta(\pi^*) = \max_{\pi \in \Pi}$ can be overcome by carefully combining softmax smoothing with first-order de-biasing in the single-period setting. Here, we adopt their concept and extend it to the dynamic setting of MDPs. To the best of our knowledge, this work is the first to consider such a smoothing technique in the context of multiple time periods.

For any real-valued function $h: \mathbb{R}^p \rightarrow \mathbb{R}$ and vector $\mathbf{v} \in \mathbb{R}^d$, define the softmax smoothing approximation $\varphi_{\beta} \{ \cdot \}$ and the multiple softmax operator $\mathbf{sm}_{\beta} \{ \cdot \}$ such that

equation[equation omitted — 298 chars of source]

where $\beta > 0$ denotes the degree of smoothing.

In the static case, the $Q$-function under a policy $\pi$ reduces to $Q (a, s ;\pi) = Q(a, s)$, as $\pi$ simply selects an action $a$ given $x$ to maximize the $Q$-function. As pointed out in whitehouse2025inference, by smoothing the value function $\eta^*(P) = \mathrm{E}_P [\max_{a} Q(a, S)]$ with $ \eta_{\beta}^* = \mathrm{E}_P \big[\mathbf{sm}_{\beta} \big\{ Q(\mathbf{a}, S) \big\} \big], $ one can differentiate $\eta_{\beta}^*(P_{\epsilon})$ with respect to $\epsilon$ and then use the first-order condition to construct a Neyman orthogonal score (or the estimating equation for the pathwise derivative at $\epsilon = 0$). However, challenges arise when extending their framework to the dynamic setting inherent in MDPs: for any fixed $\pi$, the (dynamic) $Q$-function is derived from the fixed point of the Bellman equation, such that $ \eta(\pi) = \mathrm{E} \big[ R + \gamma K_{\pi} \eta(\pi) \big], $ where $K_{\pi}$ is the transition kernel defined in (ref). Consequently, smoothing the value function directly would disrupt the above contraction structure (the basis of fixed-point theory). Fortunately, we leverage the policy optimization perspective found in Entropy-Regularized MDPs neu2017unified: instead of smoothing the value function, we choose to smooth the policy $\pi$. Specifically, we define the smoothing-greedy policy $\pi_{\beta}(P)$ as

equation[equation omitted — 184 chars of source]

It is straightforward to see that $\lim_{\beta \rightarrow +\infty}\pi_{\beta}(P) = \pi^*(P)$.

Our procedure for smoothed nuisance estimation proceeds sequentially:

itemize• First, we estimate the optimal Q-function $Q^*$ using any off-policy algorithm (e.g., Fitted Q-Iteration), yielding $\widehat{Q}_{\text{opt}}(\cdot, \cdot)$; • Second, using this estimate and under a chosen smoothing sequence $\beta_N$, we construct the plug-in policy $\widehat{\pi}_{\beta_N}$ using $\widehat{Q}$ as $ \widehat{\pi}_{\beta_N}(a \mid s) := \frac{\exp\{\beta_N \widehat{Q}_{\text{opt}}(a, s)\}}{\sum_{a' \in \mathcal{A}} \exp\{\beta_N \widehat{Q}_{\text{opt}}(a', s)\}}; $ • Finally, we estimate the density ratio $\widehat{\omega}_{\text{opt}}(\cdot, \cdot)$ corresponding specifically to this fixed policy $\widehat{\pi}_{\beta_N}$ using a method such as Minimax Weight Learning.

These nuisance estimates are then plugged into the one-step estimator, leading to

equation[equation omitted — 229 chars of source]

The smoothed estimator (ref) can be regarded as a modified version of our NSAVE estimator, where we replace the sequential evaluation with the smoothing technique. Specifically, we decompose the difference between $\widehat{\eta}_{\beta_N}$ and the true value $\eta(\pi^*)$ as follows: \[

aligned\widehat{\eta}_{\beta_N} - \eta(\pi^*) = \underbrace{\widehat{\eta}_{\beta_N} - \eta(\widehat\pi_{\beta_N})}_{:= R_{SM, 1N}} + \underbrace{\eta(\widehat\pi_{\beta_N}) - \eta(\pi^*)}_{:= R_{SM, 2N}}.

\] As shown in the proof of Theorem (ref), provided consistent nuisance estimates with appropriate convergence rates are selected, the statistical error $R_{\text{SM}, 1N}$ will converge to a normal distribution. Meanwhile, the policy-value error $R_{\text{SM}, 2N}$, controlled by the smoothing parameter $\beta_N$, becomes $o_{P_0}(N^{-1 / 2})$ via the smoothing mechanism rather than sequential evaluation. Thus, intuitively, $\widehat{\eta}_{\beta_N}$ remains a RAL estimator and can achieve semiparametric efficiency. We formalize these results in the following theorem.

theoremSuppose the conditions hold in addition to the conditions in Theorem (ref) as well as Assumption (ref). Furthermore, assume the smoothing parameter $\beta_N$ satisfies: \[ \beta_N \to \infty, \quad {\beta_N = o \big(N^{\omega_Q - {1 / 2}} \big), \quad \text{and} \quad \beta_N^{-1} = o\Big(N^{- \max \big\{ \frac{1}{2(1 + \alpha)}, \, \frac{2 + \alpha}{2 \alpha (3 + \alpha)} \big\}} \Big)}, \] and the estimated nuisances $ \{ \widehat{Q}_{\text{opt}}, \widehat\omega_{\text{opt}} \} $ satisfy the convergence rates in Assumption (ref). In addition, suppose $ {\omega_Q > \frac{1}{2} + \max \left\{ \frac{1}{2(1 + \alpha)}, \, \frac{2 + \alpha}{2 \alpha (3 + \alpha)} \right\}} $ for compatibility. Then, the smoothed one-step estimator $\widehat{\eta}_{\beta_N}$ defined in (ref) satisfies: \[ \sigma_{R_{1N}}^{-1} \big( \widehat{\eta}_{\beta_N} - \eta(\pi^*) \big) \, \rightsquigarrow \, \mathcal{N} (0, 1). \] Thus, $\widehat{\eta}_{\beta_N}$ also achieves semiparametric efficiency if Assumption (ref) holds.

From a computational perspective, our smoothed estimator $\widehat{\eta}_{\beta_N}$ is more straightforward to calculate: it only requires estimating two nuisance functionals once, followed by a direct plug-in procedure. Theorem (ref) reveals the trade-off: we require stronger convergence rates for the nuisance parameters. Again, the condition $\omega_Q > 1 / 2$ should not be regarded as a super-consistent convergence rate, given the additional time horizon dimension $T$. Nonetheless, our estimator still achieves semiparametric efficiency under Assumption (ref).

Post-Selection Inference

When $\mathcal{A} \times \mathcal{S}$ is finite and small, or more generally when the candidate policy class $\Pi = \{ \pi_1, \dots, \pi_K \}$ is finite, we can employ Post-Selection Inference (PSI) techniques to address the non-regularity. Unlike the smoothing approach, which modifies the target parameter to a smooth approximation $\eta(\pi_{\beta})$, PSI aims to construct a valid confidence interval for the value of the empirically selected policy itself, denoted as $\eta(\widehat{\pi}_N)$, where $\widehat{\pi}_N = \arg\max_{\pi \in \Pi} \widehat{\eta}(\pi)$.

Standard inference that treats $\widehat{\pi}_N$ as a fixed policy fails to account for the winner's curse: the selection process systematically favors policies with positive estimation noise, leading to an upward bias in the naive estimator.

To rigorously correct for this bias while accounting for the potential non-uniqueness of optimal policies (ties) and the high correlation between OPE estimates, we adopt the Two-Step Inference on Multiple Winners framework proposed by petrou2024inference. Our PSI procedure proceeds sequentially as follows:

itemize• Step 0: OPE estimation and selection. We estimate the values for all candidate policies. Let $\widehat{\bm{\eta}}=(\widehat{\eta}(\pi_1),\dots,\widehat{\eta}(\pi_K))^\top$ be the OPE estimates (e.g., doubly robust for each $\pi_k$). We also estimate the asymptotic covariance matrix $\widehat{\boldsymbol{\Sigma}}$ (e.g., via EIFs), such that $ \sqrt{N}(\widehat{\bm{\eta}} - \bm{\eta}) \rightsquigarrow \mathcal{N}(\bm{0}, \boldsymbol{\Sigma}) \quad \text{with} \quad \widehat{\boldsymbol{\Sigma}} \overset{\mathrm{P}}{\to} \boldsymbol{\Sigma} \quad \text{and} \quad \widehat{\mathcal{A}}_{\text{opt}} = \arg\max_{k \in [K]} \widehat{\bm{\eta}}$; • Step 1: A $(1-\delta_1)$ confidence region for the nuisance governing selection. Construct a confidence region $\mathcal{C}_\eta(\widehat{\bm{\eta}};\delta_1)\subseteq \mathbb{R}^K$ such that $\liminf_{N\to\infty} \mathrm{P}\Big(\bm{\eta}\in \mathcal{C}_\eta(\widehat{\bm{\eta}};\delta_1)\Big)\ \ge\ 1-\delta_1$. Using $\mathcal{C}_\eta$, define the plausible optimal set $\widehat{\mathcal{A}}^{+} := \bigcup_{\bm{\eta} \in \mathcal{C}_\eta(\widehat{\bm{\eta}};\delta_1)} \arg\max_{k\in[K]} \penalty 0 \eta_k$. By construction, on ${\bm{\eta}\in \mathcal{C}_\eta(\widehat{\bm{\eta}};\delta_1)}$, the true optimal set $\widehat{\mathcal{A}}_{\text{opt}} $ is contained in $\widehat{\mathcal{A}}^{+}$. • Step 2: Calibrate a simultaneous critical value over $\widehat{\mathcal{A}}^{+}$. Let the standardized errors be $Z_k := \widehat{\boldsymbol{\Sigma}}_{kk}^{-1 / 2} \big(\widehat{\eta}(\pi_k) - \eta(\pi_k) \big)$. We choose a (data-dependent) critical value $q_{1-(\delta_2 - \delta_1)}(\widehat{\mathcal{A}}^{+})$ satisfying the asymptotic guarantee $ \lim_{N\to\infty} \mathrm{P}\big( \max_{k\in \widehat{\mathcal{A}}^{+}} |Z_k| \le q_{1-(\delta_2 - \delta_1)}(\widehat{\mathcal{A}}^{+}) \big) \penalty 0 \geq 1-(\delta_2 - \delta_1). $ • Final PSI confidence set (reported on the selected set). Define $\mathcal{C}_{\text{PSI}} := \bigtimes_{k\in \widehat{\mathcal{A}}_{\text{opt}}} \Bigl[ \widehat{\eta}(\pi_k)\ \pm\ q_{1-(\delta_2 - \delta_1)}(\widehat{\mathcal{A}}^{+})\cdot\sqrt{\widehat{\boldsymbol{\Sigma}}_{kk}/N}\Bigr]$.

Various approaches instantiate the template required in the above steps. Here, we utilize the worst-case construction (the primary construction in petrou2024inference), postponing other constructions to Appendix (ref).

To complete Step 1, we first identify indices that are NOT “significantly” worse than any winner $\widehat{k} \in \widehat{\mathcal{A}}_{\text{opt}}$: $\widehat{\mathcal{A}}^{+}: = \left\{ j \in [K] : \big| \widehat\eta(\pi_{\widehat{k}}) - \eta(\pi_{j}) \big| \le z_{1 - \delta_1 / 2} \widehat{\boldsymbol{\Sigma}}_{\widehat{k}, j} / \sqrt{N} ~~~ \text{ for any } ~~~ \widehat{k} \in \widehat{\mathcal{A}}_{\text{opt}} \right\}$, where $z_{1 - \delta_1 / 2} $ denotes the upper $\delta_1/2$-th quantile of a standard normal distribution. The worst-case PSI confidence set is correspondingly given by: \[ \mathcal{C}_{\text{PSI}}^{(\text{WC})} := \bigtimes_{k \in \widehat{\mathcal{A}}_{\text{opt}}} \Bigl[\, \widehat{\eta}(\pi_k) \pm q_{1-(\delta_2 - \delta_1)}(\widehat{\mathcal{A}}^{+}) \,\widehat{\boldsymbol{\Sigma}}_{kk}^{1 / 2}\Bigr], \] where $q_{1-(\delta_2 - \delta_1)}(\widehat{\mathcal{A}}^{+}) $ is the quantile functional defined in (ref) in Appendix (ref).

corollary$ \liminf_{N \to \infty} \mathrm{P}_{P_0} \Big\{ \big\{ \eta(\pi_k) :k \in \widehat{\mathcal{A}}_{\text{opt}} \big\} \in \mathcal{C}_{\text{PSI}} \Big\} \geq 1 - \delta_2. $ Specifically, the worst-case PSI confidence set $\mathcal{C}_{\text{PSI}}^{(\text{WC})}$ satisfies the above inequality.

A straightforward consequence of Corollary (ref) is that if the optimal policy is unique, the length of $\mathcal{C}_N^{\text{PSI}}$ converges to that of the standard oracle confidence interval (oracle efficiency). If there are multiple optimal policies (ties), the interval remains valid by adapting to the worst-case distribution over the set of winners.

Simulations

We investigate the finite-sample performance of the proposed NSAVE and smoothing-based inference procedures, comparing them with the SAVE estimator shi2022statistical. We consider infinite-horizon Markov decision processes with discrete state and action spaces, a uniform initial state distribution, and a behavior policy satisfying the overlap condition. Three representative regimes are examined: (i) a regular, well-specified setting (Scenario A); (ii) a setting with heavy-tailed reward contamination (Scenario B); and (iii) a structurally misspecified setting with severe state aliasing (Scenario C). The full data-generating mechanisms, tuning parameters, and implementation details are provided in Appendix (ref).

We evaluate both a fixed oracle-optimal policy and data-driven policies learned via double Fitted Q-Iteration. All methods are implemented using cross-fitting. We report the mean squared error (MSE) and empirical coverage probability (ECP) of 95% confidence intervals over 100 Monte Carlo replications. Additional experimental results for Scenarios A and B are deferred to Appendix (ref).

\paragraph{Main findings.} In the regular regimes (Scenarios A and B), both NSAVE and SAVE are asymptotically consistent; however, their finite-sample behaviors differ substantially. As detailed in Appendix (ref), NSAVE attains nominal coverage at markedly smaller sample sizes, reflecting the stability of trajectory-level efficient influence function–based inference combined with studentized batch means. In contrast, SAVE requires significantly larger $N$ and $T$ for its blockwise variance approximation to stabilize. Furthermore, under heavy-tailed reward contamination (Scenario B), NSAVE maintains well-calibrated coverage, whereas SAVE exhibits noticeable distortion.

Most notably, in the structurally misspecified setting with state aliasing (Scenario C), purely model-based methods fail. As shown in Table (ref), SAVE and the smoothing-based plug-in estimator suffer from persistent bias and near-zero coverage. In contrast, NSAVE remains accurate and achieves orders-of-magnitude smaller MSE by leveraging its double-robust correction via importance weighting. Overall, these results demonstrate that NSAVE provides both sharper finite-sample calibration and greater robustness across regular and non-regular regimes.

table[table omitted — 1,423 chars of source]

Application to the OhioT1DM Dataset

We apply NSAVE and the smoothing-based method to the OhioT1DM mobile health dataset previously analyzed by shi2022statistical. The data consist of continuous glucose monitoring records, insulin delivery logs, and self-reported events for six patients with type 1 diabetes over an eight-week period. Following the construction in shi2022statistical, we discretize time into non-overlapping 3-hour intervals and define patient-specific state, action, and reward trajectories; full preprocessing details are provided in Appendix (ref). We set the discount factor to $\gamma=0.5$.

For each patient, we estimate an optimal policy and construct confidence intervals for its value using NSAVE and the smoothing approach. We also estimate the value of the observed clinician behavior policy via a plug-in model-based estimator and conduct inference on the value difference $D_i = V(S_{i0};\pi^*) - V(S_{i0}; b_i)$, which quantifies the potential improvement of the learned optimal policy over the observed treatment rule.

Inference is performed separately for two clinically relevant starting times (8:00 am and 2:00 pm on Day 1). Figure (ref) reports the 95% confidence intervals for $D_i$ with $\gamma=0.5$. Across patients and starting times, the estimated value differences are consistently positive, with several confidence intervals excluding zero, indicating statistically significant improvements. Compared with SAVE (as reported in prior analyses), NSAVE yields stable confidence intervals without relying on stringent non-degeneracy conditions. The smoothing-based approach provides a complementary regularized alternative, particularly effective when the optimal policy is nearly deterministic or non-unique. Sensitivity analyses for $\gamma \in \{0.4, 0.7\}$ are provided in Appendix (ref).

figure[figure omitted — 300 chars of source]

Final Remarks

Finally, we emphasize that the existence of an efficient influence function for optimal policy values hinges critically on the regularity of the policy optimization map. When the optimal policy is unique, the problem reduces locally to inference for a fixed policy, and classical semiparametric theory applies uehara2022review,shi2025statistical. When optimal policies are non-unique, the value functional becomes non-smooth and non-regular, and standard root-$N$ inference can fail, a phenomenon closely related to non-regular parameters in optimal treatment regimes and post-selection inference laber2014dynamic,whitehouse2025inference.

Our NSAVE procedure provides a stable, efficient solution in the regular regime, while the smoothing approach connects optimal policy evaluation in MDPs to entropy-regularized control neu2017unified and recent smoothing-based inference for max-type functionals whitehouse2025inference. The post-selection confidence sets further complement these methods by offering valid worst-case coverage for sets of optimal policies, extending ideas from selective and multiple-inference frameworks chernozhukov2015valid.

These results clarify both the scope and the limitations of existing approaches such as SAVE shi2022statistical, and suggest that non-regularity is an intrinsic feature of optimal policy inference rather than a technical artifact.

Data Availability Statement

The OhioT1DM dataset analyzed in this study is publicly available at \url{https://webpages.charlotte.edu/rbunescu/data/ohiot1dm/OhioT1DM-dataset.html}.

Disclosure Statement

The authors report there are no competing interests to declare.

Supplementary Material

The online Supplementary Material contains detailed configurations for the simulations and real data application (Appendix A), alternative confidence set constructions for post-selection inference (Appendix B), proofs of the theoretical results (Appendices C--E), and auxiliary lemmas (Appendix F).