Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
74,465 characters · 22 sections · 39 citation commands
Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness
\def\spacingset#1{ {#1}} \spacingset{1}
\def\spacingset#1{ {#1}} \spacingset{1}
\if11 \fi
\if01 \fi
{\it Keywords:} Efficient estimator; non-regular inference; optimal policy; off-policy evaluation.
\spacingset{1.9}
Reinforcement learning (RL) is concerned with learning optimal decision rules for sequential decision problems in order to maximize long-term cumulative rewards sutton2018reinforcement. A fundamental statistical task within RL is off-policy evaluation (OPE), which seeks to estimate the value of a target policy using data generated under a potentially distinct behavior policy. OPE plays a pivotal role in offline RL, where new data collection is either costly or ethically constrained, necessitating that inference rely entirely on historical trajectories luedtke2016statistical,agarwal2019reinforcement,uehara2022review.
The majority of existing statistical analyses of OPE concentrate on the classical setting in which the evaluation policy is fixed and known a priori. In this regime, an extensive body of literature has established doubly robust and semiparametrically efficient estimators under various modeling assumptions jiang2016doubly,kallus2020double,shi2021deeply. However, in many empirical applications, the policy of interest is not pre-specified but is itself estimated from the data as an optimal policy. This setting introduces a qualitatively different statistical structure: the target functional involves a maximization over policies, and the resulting value function can be non-smooth and non-regular, particularly when the optimal policy is not unique or is nearly deterministic.
Analogous issues have been extensively studied in the causal inference literature regarding optimal treatment regimes laber2014dynamic,kosorok2019precision,athey2021policy, where it is now well-established that the non-uniqueness of optimal rules leads to non-regularity and renders standard asymptotic theory invalid. Extending such insights to the sequential decision-making framework of Markov decision processes (MDPs) is substantially more challenging due to temporal dependence, the Bellman fixed-point structure, and the complex interaction between policy optimization and value estimation.
Recently, shi2022statistical proposed the SAVE estimator, which establishes semiparametric efficiency for the value of an optimal policy under a linear $Q$-function model and a set of non-degeneracy conditions. While SAVE represents a significant step toward principled inference for optimal policy values, its theory relies on stringent structural and regularity assumptions. In particular, it requires (i) a low-dimensional linear approximation of the $Q$-function, and (ii) well-conditioned Bellman estimating equations under the target policy. When the optimal policy is unique and deterministic, or nearly so, the latter condition often fails: the feature covariance induced by the target policy becomes ill-conditioned, leading to numerical instability and the breakdown of the associated inference. Moreover, in such regimes, SAVE no longer admits a doubly robust representation and loses its efficiency guarantees; furthermore, no alternative valid confidence sets are provided once these non-degeneracy conditions are violated.
This paper develops a unified inferential framework for the value of optimal policies in MDPs that explicitly addresses such non-regular phenomena. Our contributions are threefold.
Through theoretical analysis, simulations, and an application to the OhioT1DM dataset, we demonstrate that the proposed NSAVE and smoothing procedures yield stable and valid confidence intervals across both regular and non-regular regimes, significantly outperforming existing methods in settings where the optimal policy is deterministic or nearly deterministic.
The remainder of the paper is organized as follows. In Sections (ref) and (ref), we first characterize the efficient influence function for the optimal policy value and establish the non-regularity that arises under policy non-uniqueness. In Sections (ref) and (ref), we then develop the proposed NSAVE estimator together with its efficiency and stability properties. Section (ref) introduces the smoothing-based approach and the associated post-selection confidence sets for handling non-unique optimal policies. The finite-sample performance of NSAVE, the smoothing approach, and existing methods is investigated through extensive simulation studies in Section (ref). In Section (ref), we apply the proposed methods to the OhioT1DM mobile health dataset and conduct patient-specific off-policy inference. Finally, the last section concludes with a discussion of the implications, limitations, and directions for future research.
We consider observational data generated from a canonical Markov decision process (MDP). At any given time $t$, let $(S_t, A_t, R_t)$ denote the state-action-reward triplet. Let $O$ be shorthand for the data tuple $(S, A, R, S^{\prime})$. We observe an offline dataset $\big\{ O_{it}: 1 \leq i \leq N, 0 \leq t \leq T\big\}$ with $O_{it} = (S_{it}, A_{it}, R_{it})$, generated by a behavior policy $b(\cdot \mid S)$, where $i$ indexes the episode and $t$ indexes the time point. For any fixed target policy $\pi(a \mid s)$, OPE generally aims to evaluate the mean return $ \eta(\pi) = \mathrm{E}^{\sim \pi} \left[ \sum_{t = 0}^{+\infty} \gamma^t R_t \right] $ and construct a valid confidence interval, where $\mathrm{E}^{\sim \pi}$ denotes the expectation when the system follows policy $\pi$. Distinct from existing semiparametric studies on OPE that consider an arbitrary $\pi$, we focus on a specific target policy: the optimal policy $\pi^*$, which maximizes $\eta(\pi)$ over the set of all possible policies $\Pi$. Specifically, the parameter of interest is \[
\]
To ensure the value function is identifiable, we adopt standard assumptions in the OPE literature. For simplicity, we use $f(x \mid y)$ to represent the conditional density of $X$ given $Y = y$.
Assumptions (ref) and (ref) are sufficient for identifying $\eta(\pi)$ for any given policy $\pi \in \mathcal{P}$. We briefly review standard estimation methods. The first method involves analyzing the aggregate mean return via the $Q$-function, defined as
The value function can be expressed as \[
\] where $f(s_0)$ is the initial state density. We also define the value function $V(s; \pi) = \int Q(a, s; \pi) \pi(a \mid s) \, \mathrm{d} a$.
The second method is the marginal importance sampling (MIS) estimator, which addresses the curse of horizon. The marginal ratio is defined as
where $f_{\sim \pi, t}$ and $f_{+ \infty}$ denote the time-dependent and stationary densities, respectively. Under stationarity, we have the identity \[ \eta(\pi) = \frac{1}{1 - \gamma} \mathrm{E} \big[ \omega(A, S; \pi) \mathrm{E}^{\sim \pi} [R \mid S]\big]. \] Crucially, both the $Q$-function (ref) and MIS ratio (ref) are required for the semiparametrically efficient estimation of $\eta(\pi)$.
Let the trajectory $O = O_{1:T} \sim P_0 \in \mathcal{M}$. We define the functional $\Psi^*: \mathcal{M} \rightarrow \mathbb{R}$ as \[ \Psi^*(P) := \mathrm{E}_{P} \big[ Q(P) \big(A_0, S_0; \pi^*(P) \big)\pi^*(P)(A_0\mid S_0)\big], \] where $ \pi^*(P)(\cdot \mid s) = \arg\max_{\pi \in \mathcal{P}} Q(P)(a, s; \pi)$ is the optimal policy under the law $P$. Thus, $\Psi^*(P_0) = \eta(\pi^*)$ is well-defined. We focus on the value function $\Psi^*(P_0)$; discussions regarding $\pi^*(P_0)$ itself can be found in kosorok2019precision,athey2021policy,luo2024policy. Define the auxiliary functional \[ \Psi(P; \pi) = \mathrm{E}_{P} \big[ Q(P) \big(A_0, S_0; \pi \big)\pi(A_0\mid S_0)\big]. \] While $\Psi\big(P; \pi^*(P)\big) = \Psi^*(P)$, in general $\Psi\big(P_1; \pi^*(P_2)\big) \neq \Psi^*(P_1)$ if $P_1 \neq P_2$.
Assume $\{ S_t\}_{t \geq 0}$ is stationary. As shown in uehara2020minimax,shi2024off,shi2021deeply, for any fixed $\pi$, the efficient influence function (EIF) for $\Psi(P; \pi)$ at $P_0$, evaluated at $O$ in the full nonparametric space $\mathcal{M}_{\text{nonpar}}$, is
assuming $O \sim P_0$. We denote the estimating functions for a single point $O$ and the trajectory $O_{0:T}$ as
The EIF for $\eta(\pi)$ satisfies \[ S^{\text{eff, nonpar}}_{\eta(\pi)}(O; Q, \omega, V, \pi) = {\psi}^{\text{point}}_{\eta(\pi)}(O; Q, \omega, V, b, \pi) - \eta(\pi), \] and under stationarity, \[ \mathrm{E} \big[ S^{\text{eff, nonpar}}_{\eta(\pi)}(O; Q, \omega, V, \pi) \big] = \mathrm{E} \big[ {\psi}^{\text{traj}}_{\eta(\pi)}(O_{0:T}; Q, \omega, V, \pi) \big] - \eta(\pi). \]
For any regular asymptotically linear (RAL) estimator $\widehat{\Psi}(\widehat{P}; \pi)$, there exists a unique influence function tsiatis2006semiparametric such that \[ \widehat{\Psi}(\widehat{P}; \pi) - \Psi(P_0; \pi) = \frac{1}{\sqrt{N}}\sum_{i=1}^{N}\operatorname{IF}(P_0; \pi)(O_i) + o_{P_0} \big( (N - \ell_N )^{-1 / 2}\big). \] The EIF $S^{\text{eff, nonpar}} \{ \Psi(P; \pi)\}\big|_{P = P_0}(O)$ is the unique influence function that minimizes the variance $\operatorname{var}_{P_0} \big( \operatorname{IF}(P_0; \pi)(O) \big)$.
Geometrically, the EIF is characterized via the tangent space $\mathcal{T}$. For any differentiable path $\{ P_{\epsilon}: \epsilon \in \mathbb{R}\} \subsetneq \mathcal{M}$ passing through $P_0$ at $\epsilon = 0$, the pathwise differentiability of $\Psi(P; \pi)$ implies $ \frac{\mathrm{d}}{\mathrm{d} \epsilon} \Psi(P_{\epsilon}; \pi) \Big|_{\epsilon = 0} = \mathrm{E}_{P_0} \Big[ \widetilde{\psi}(O) \dot{\ell} (O) \Big] $ by the Riesz representation theorem, where $\dot{\ell}(O)$ is the score function. When $\widetilde{\psi}(O) \in \mathcal{T}$, then $\widetilde{\psi}(O) = S^{\text{eff, nonpar}} \{ \Psi(P; \pi)\}\big|_{P = P_0}(O)$, and
However, for the optimal value functional $\Psi^*(P) = \Psi \big(P; \pi^*(P)\big)$, (ref) may fail. The optimal policy $\pi^*(P_{\epsilon})$ along the path may contribute to the derivative if (i) $\Psi \big(P; \pi \big)$ is sensitive to $\pi$, and (ii) $\pi^*(P)$ is sensitive to $P$. By the chain rule of Gateaux differentials shapiro1990concepts: \[
\] If the second term is non-zero, then $ S^{\text{eff, nonpar}} \{ \Psi^*(P)\}\big|_{P = P_0} \neq S^{\text{eff, nonpar}} \{ \Psi(P; \pi)\}\big|_{P = P_0, \pi = \pi^*(P_0)}. $ We establish a concise expression for $ S^{\text{eff, nonpar}} \{ \Psi^*(P)\}$ in the next section.
We present regularity conditions on the distributions of $S$, $A$, and $R$ to ensure the EIF exists. Let $\mathcal{S}$ and $\mathcal{A}$ denote the supports of the states and actions, respectively.
Assumption (ref) contains standard conditions levine2020offline,uehara2022review,shi2022statistical, with the exception of the policy bounds, which are nonetheless mild and standard in OPE xu2021doubly,shi2024off,bian2024off. Our first result shows that when the optimal policy is unique and deterministic, the EIF $S^{\text{eff, nonpar}} \{ \Psi^*(P)\}$ exists and equals $S^{\text{eff, nonpar}} \{ \Psi(P; \pi)\}$ at $\pi = \pi^*$.
Conversely, if the optimal policy is not unique—specifically, if a significant set of states exists where multiple optimal actions are indifferent—the influence function does not exist.
Assumption (ref) describes Unrestricted Optimal Rules robins2014discussion, leading to non-regularity where standard margin conditions shi2022statistical fail.
When $\pi$ is fixed, as studied thoroughly in the literature, the one-step estimator is obtained by solving the estimating equation $\mathbb{P}_{NT} \big\{ S^{\text{eff, nonpar}}_{\eta(\pi)}(O; \widehat{Q}, \widehat{\omega}, \widehat{V}, \widehat{b}, \pi) \big\} = 0$, which yields
with estimated nuisance functions $\widehat{Q}, \widehat{\omega}, \widehat{V}, \widehat{b}$. Here, $\widehat{\eta}_{\text{DR, 1}}(\pi)$ and $\widehat{\eta}_{\text{DR, 2}}(\pi)$ are asymptotically equivalent and share desirable statistical properties:
When the optimal policy (or policies) $\pi^*$ is unknown, and we have an estimated optimal policy $\widehat{\pi}$ computed by a $Q$-learning type algorithm such that \[ \widehat{\pi}(a \mid s) := \mathds{1} \Big\{ a = \arg \max_{a' \in \mathcal{A}} \widehat{Q}_{\text{opt}} (a, s ) \Big\}, \] where $\widehat{Q}_{\text{opt}} (a, s ) $ denotes a consistent estimator for the optimal $Q$-function, i.e., $Q(a, s; \pi^* )$, an intuitive “plug-in" estimator for the parameter of interest $\eta(\pi^*)$ is $\widehat{\eta}_{\text{DR, j}}(\widehat{\pi})$. The above double-robustness and semiparametric efficiency property still holds as a standard result established in the literature uehara2022review. However, as we pointed out in Theorem (ref), when the optimal policy is not unique, there is no basis for discussing semiparametric efficiency, as RAL estimators do not exist. More importantly, in such cases, the argmax of the optimal $Q$-function may not be uniquely defined; consequently, $\widehat{\pi}(\cdot \mid s)$ might not converge to a fixed quantity for some $s \in \mathcal{S}$. As a result, the plug-in estimator $\widehat{\eta}_{\text{DR, j}}(\widehat{\pi})$ will fluctuate randomly and fail to maintain a stable limiting distribution shi2022statistical.
To overcome this issue, shi2022statistical proposed a new estimator called SequentiAl Value Evaluation (SAVE), denoted as $\widehat{\eta}_{\text{SAVE}}$, assuming that the $Q$-function follows a linear sieve model such that $Q (a, s; \pi) \approx \Phi^{\top}(s) \bm{\beta}_{\pi, a}$, where $\Phi(s)$ is a vector of sieve basis functions. This novel estimator enjoys bidirectional asymptotic normality: \[
\] as either $N \rightarrow \infty$ or $T \rightarrow \infty$, where $K$ is the number of data partitions and $\widehat{\sigma}_{\text{SAVE}}^2$ is a “plug-in" type variance estimator. Thus, the readily applicable estimator $\widehat{\eta}_{\text{SAVE}}$ can be used for statistical inference.
Nonetheless, despite its appealing theoretical guarantees, the SAVE estimator $\widehat{\eta}_{\text{SAVE}}$ suffers from several significant limitations. First, SAVE relies critically on a linear structural assumption for the $Q$-function, namely that $Q(s,a;\pi)$ can be well-approximated by a low-dimensional linear form $Q(s,a;\pi)\approx \Phi^\top(s)\boldsymbol{\beta}_{\pi,a}$. This assumption may be violated in many realistic sequential decision problems. Second, when the optimal policy is unique and deterministic, $\widehat{\eta}_{\text{SAVE}}$ no longer admits a doubly robust representation and consequently loses both the double robustness property and the associated semiparametric efficiency guarantees. Third, and most critically, SAVE requires strong non-degeneracy conditions on the target policy. In particular, its inference theory implicitly relies on well-conditioned Bellman estimating equations, which may fail when the target policy is deterministic or nearly deterministic. In such cases, the feature covariance induced by the target policy becomes nearly singular, leading to unstable estimation and invalid uncertainty quantification. Consequently, SAVE does not provide a principled fallback inference procedure once these marginal conditions are violated. In Section (ref), we demonstrate through simulation that this issue is not merely theoretical: under deterministic or highly concentrated target policies, SAVE can exhibit severely distorted coverage, whereas our proposed method remains stable and valid.
To address these three challenges, we adopt the conceptual framework of shi2022statistical while introducing a revised sequential value evaluation procedure and a corresponding estimator, which we term Nonparametric SequentiAl Value Evaluation (NSAVE).
Assume $\{ O_{\tau(i)}\}_{i = 1}^N$ is a random permutation of the original i.i.d. trajectory observations $\{ O_i\}_{i = 1}^N$. Let $\{\ell_N\}$ be a sequence of non-negative integers representing the size of the initial sample used to estimate the initial optimal policy from the estimated $Q$-function, denoted by $\widehat{\pi}_{\tau(\ell_N - 1)}^{(0)}$. We initialize such that $\widehat{\pi}_{\tau(\ell_N - 1)}^{(Q)} = \widehat{\pi}_{\tau(\ell_N - 1)}^{(\omega)} = \widehat{\pi}_{\tau(\ell_N - 1)}^{(0)}$.
For $j = \ell_N + 1, \ldots, N$, we perform the following steps:
It is worth noting that there is no need to explicitly use or estimate the $\omega$-based estimated optimal policy, defined as $ \widehat{\pi}_{\tau(j - 1)}^{(\omega)} (a \mid s) := \arg\max_{a \in \mathcal{A}} \widehat{\omega}_{\tau(j - 1)} (a, s ; \widehat{\pi}_{\tau(j - 1)}^{(\omega)}). $ This $\omega$-based estimated optimal policy is primarily introduced for notational convenience, emphasizing that the sequential marginal ratio is derived from the Optimizing Step.
Define the online one-step variance as \[
\] Here, the function $S^{\text{eff, nonpar}} \{ \Psi \}$ is regarded purely as a functional of the observation $O$ and the nuisance functions $\{Q, \omega, \pi \}$ (independent of the distribution $P$), regardless of whether the optimal policies are unique. Let $\widehat{\sigma}_{\tau(j - 1)}^2$ denote its corresponding consistent estimator. In practice, $\widehat{\sigma}_{\tau(j - 1)}^2$ can be estimated using the sample variance over a specific sliding window, such as $\big\{ \widehat{\psi}^{\text{traj-step}}_{\tau(j - m)}, \widehat{\psi}^{\text{traj-step}}_{\tau(j - m + 1)} , \ldots, \widehat{\psi}^{\text{traj-step}}_{\tau(j - 1)} \big\}$, for a sufficiently large $m$. Then, our final estimator is similarly defined as the weighted average: \[
\] Intuitively, our novel estimator $\widehat{\eta}_{\text{NSAVE}}$ approximates, but is distinct from, the average weighted empirical historical value $\overline{\eta}_w(\widehat{\pi}^{(Q)})$, defined as \[ \overline{\eta}_{w}(\widehat{\pi}^{(Q)}) := \bigg\{ \sum_{j = \ell_N + 1}^N \frac{1}{\widehat{\sigma}_{\tau(j - 1)}}\bigg\}^{-1} \sum_{j = \ell_N + 1}^N \frac{\eta ({\widehat{\pi}_{\tau(j - 1)}^{(Q)})}}{\widehat{\sigma}_{\tau(j - 1)}}. \]
In the following, we analyze the theoretical properties of our novel estimator $\widehat{\eta}_{\text{NSAVE}}$ to demonstrate its advantages. We assume that $\mathcal{S}$ and $\mathcal{A}$ are finite. Furthermore, we assume that the function classes for both $Q$ and $\omega$ are uniformly bounded Donsker classes.
There are various approaches for estimating the nuisance components. Here, we mainly focus on the estimation of $\widehat{\omega}_{\tau(j - 1)} (\cdot, \cdot ; \widehat{\pi}_{\tau(j - 1)}^{(\omega)})$, while the estimation of other nuisance components, such as obtaining $\widehat{Q}_{\tau(j - 1)} (\cdot, \cdot ; \cdot)$ and $\widehat{b}(\cdot \mid \cdot)$, has been discussed in detail in the literature shi2025statistical.
The most common strategy is using dual linear programming. Specifically, the true nuisance $\omega(\cdot, \cdot ; \pi^*)$ is the solution to the following maximization problem: \[ \omega(a, s ; \pi^*) = \arg\max_{\omega \in \Omega_{\text{flow}}} \mathrm{E}_{P_0} \big[\omega(A; S) R\big], \] where $\Omega_{\text{flow}}$ is the polytope of valid density ratios satisfying the Bellman flow constraints, i.e., \[
\] Let $\widehat{\Omega}_{\text{flow}}$ be the corresponding estimated $\Omega_{\text{flow}}$ obtained by replacing the unknown nuisance functions $b(a \mid s)$, $f_0(s)$, and $f(s' \mid a, s)$ with their estimated counterparts. Then, a practical estimator $\omega(a, s ; \pi^*)$ at Step $j$ can be defined as \[
\]
Another important approach is Minimax Weight Learning (MWL), which constructs two nuisance estimators for $\big( Q (\cdot, \cdot; \pi^*), \, \omega(\cdot, \cdot; \pi^*)\big)$ simultaneously nachum2019dualdice,duan2020minimax,uehara2020minimax. We adapt their idea and extend this minimax framework to the estimation of the optimal policy value, where the target policy itself is data-dependent and potentially non-unique. To do this, we define the Lagrangian function \[
\] Let $\mathbb{P}_{\tau(j - 1) T} := \frac{1}{T (j - 1)} \sum_{t = 1}^T \sum_{\iota = 1}^{j - 1} [\boldsymbol{\cdot}]_{\tau(\iota), t}$ be the empirical distribution measure at Step $j$. We can then construct the estimated $Q$-function and $\omega$-function under the (estimated) policies at Step $j$ by solving the following minimax problem:
We then set $\widehat{\omega}_{\tau(j - 1)} (\cdot, \cdot ; \widehat{\pi}_{\tau(j - 1)}^{(\omega)}) = \widehat{\omega}_{\text{opt}, \tau(j - 1)} (\cdot, \cdot)$, while one may choose whether to use $\widehat{Q}_{\text{opt}, \tau(j - 1)}$ as $\widehat{Q}_{\tau(j - 1)} (\cdot, \cdot ; \widehat{\pi}_{\tau(j - 1)}^{(Q)})$.
Although it is intuitively plausible that our novel estimator $\widehat{\eta}_{\text{NSAVE}}$ will be consistent as long as the nuisance components are consistently estimated (we will also formally demonstrate its consistency), such intuition is insufficient for inference. The latter typically requires stronger conditions. To characterize the specific requirements, consider that the remainder term can be decomposed into two distinct components corresponding to two different inferential strategies:
In both decompositions, the second terms, $R_{\text{CLB}, 2N}$ and $R_{\text{TCI}, 2N}$, are non-positive by the definition of $\pi^*$. Consequently, we have \[
\] If we can construct a valid $(1 - \alpha)$ upper bound $\operatorname{UB}(R_{1N}; \alpha)$ for either $R_{\text{CLB}, 1N}$ or $R_{\text{TCI}, 1N} $ such that \[ \liminf_{N \rightarrow \infty} \mathrm{P} \big(R_{\text{CLB}, 1N} \leq \operatorname{UB}(R_{1N}; \alpha) \big) \geq 1 - \alpha ~~\text{ or } ~~\liminf_{N \rightarrow \infty} \mathrm{P} \big( R_{\text{TCI}, 1N} \leq \operatorname{UB}(R_{1N}; \alpha) \big) \geq 1 - \alpha, \] then \[
\] This implies that $\widehat{\eta}_{\text{NSAVE}} - \operatorname{UB}(R_{1N}; \alpha)$ serves as a valid lower bound for the optimal value. The following theorem formally states how to construct a valid $\operatorname{UB}(R_{1N}; \alpha)$ and its corresponding estimator $\widehat{\operatorname{UB}}(R_{1N}; \alpha)$. Let $ \sigma_{R_{1N}} := \frac{1}{\sqrt{N - \ell_N}} \left\{ \sum_{j = \ell_N + 1}^N \frac{1}{\widetilde{\sigma}_{\tau(j - 1)}}\right\} $ and \[ {\psi}^{\text{traj, *}}_{\eta(\pi), \tau(j)} ( \cdot, \cdot, \cdot, \cdot ) := {\psi}^{\text{traj}}_{\eta(\pi)} (O_{0:\tau(j)} \, ; \, \cdot, \cdot, \cdot, \cdot ) - \mathrm{E} \big[ {\psi}^{\text{traj}}_{\eta(\pi)} (O_{0:\tau(j)} \, ; \, \cdot, \cdot, \cdot, \cdot ) \mid \sigma \langle O_{0:\tau(j - 1)} \rangle \big]. \] The upcoming theorem, which establishes the asymptotic normality for the first terms in the two types of decompositions, relies on the following assumptions:
Let ${\operatorname{UB}}(R_{1N}; \alpha) := \frac{z_{\alpha} \widetilde\sigma_{R_{1N}}^{-1}}{\sqrt{N - \ell_N}}$ and $\widehat{\operatorname{UB}}(R_{1N}; \alpha) := \frac{z_{\alpha} \widehat\sigma_{R_{1N}}^{-1}}{\sqrt{N - \ell_N}}$ with $ \widehat\sigma_{R_{1N}} := \frac{1}{\sqrt{N - \ell_N}} \left\{ \sum_{j = \ell_N + 1}^N \frac{1}{\widehat{\sigma}_{\tau(j - 1)}}\right\}$. Then Theorem (ref) implies that $\widehat{\eta}_{\text{NSAVE}} - \widehat{\operatorname{UB}}(R_{1N}; \alpha)$ provides a readily applicable conservative lower bound for $\eta(\pi^*)$.
Here, we only establish asymptotic normality for the first term in the coarse decomposition. The reason is that ensuring asymptotic normality for $R_{\text{TCI}, 1N}$ requires regularity conditions for the Estimated Optimal Policies. In contrast, as can be seen from Theorem (ref) and Corollary (ref), we do not impose any conditions on the estimated policy sequence $\{ \widehat{\pi}_{\tau(j - 1)} \}_{j > \ell_N}$. Therefore, compared with the conditions in shi2022statistical, which require regularity and so-called margin conditions for both the estimated and true optimal policies, our novel estimator $\widehat{\eta}_{\text{NSAVE}}$ admits valid inference procedures without such restrictions.
Under stronger conditions, more accurate inference via a Two-Sided Confidence Interval for $\eta^*$ is possible. To achieve this, we first need to establish the asymptotic properties of the two terms $R_{\text{TCI}, 1N}$ and $R_{\text{TCI}, 2N}$. Here, $R_{\text{TCI}, 1N}$ represents the fluctuation of our novel estimator around the true value function evaluated at the estimated optimal policy $\widehat{\pi}_{\tau(N)}^{(Q)}$. Consequently, as the construction of $\widehat{\eta}_{\text{NSAVE}}$ is conditional on past observations, we can apply the martingale CLT to show that $R_{\text{TCI}, 1N}$ is $\sqrt{N - \ell_N}$-consistent and converges to a Gaussian distribution under conditional Lindeberg conditions and regularity conditions for the final estimated policy $\widehat{\pi}_{\tau(N)}^{(Q)}$. This two-sided technique is also applied in shi2022statistical. On the other hand, $R_{\text{TCI}, 2N}$ represents the systematic error arising from replacing the unknown optimal policy sequence with $\{ \widehat{\pi}_{\tau(j - 1)} \}_{j > \ell_N}$, which may exhibit a slower convergence rate.
These additional conditions will help guarantee the CLT for $R_{\text{TCI}, 1N}$. {An important note here is that $\kappa_\pi > 1 / 2$ should NOT be regarded as a super-consistent convergence rate, as we have another dimension of sampling: the time horizon dimension $T$.}
Define the sub-optimality gaps as $\mathbf{\Delta} (a, s; Q, \pi) := V(s; \pi^*) - Q(a, s; \pi)$. In addition to Assumption (ref), to guarantee favorable limiting behavior for $R_{\text{TCI}, 2N}$ as a functional of the estimated policy sequence, we introduce a margin-type condition below. {Let $\mathcal{A}_{\text{sub-opt}}(s) = \mathcal{A} \, \setminus \, \arg\max_{a \in \mathcal{A}} Q(a, s; \pi^*)$.}
As partially shown in Corollary (ref), compared with the results in shi2022statistical, our estimator $\widehat{\eta}_{\text{NSAVE}}$ demonstrates several advantages. Specifically:
To explicitly state the first two advantages of our new estimator compared to the estimator in shi2022statistical when the optimal policy $\pi^*$ is deterministic and unique, we use the following theorem to characterize the theoretical asymptotic properties of $\widehat{\eta}_{\text{NSAVE}}$.
{Assumptions (ref) and (ref) are quite mild. In particular, the MWL estimator from (ref) would naturally satisfy these two assumptions, as the two Gateaux differentials are exactly zero, and the probability for the flow constraint would also be exactly one.} Theorem (ref) explicitly states the advantages of our novel estimator $\widehat{\eta}_{\text{NSAVE}}$: Compared with the naive “plug-in" estimators $\widehat{\eta}_{\text{DR, j}}(\widehat{\pi})$, our estimator can adapt to potentially non-unique optimal policies; Compared with $\widehat{\eta}_{\text{SAVE}}$, $\widehat{\eta}_{\text{NSAVE}}$ retains both double-robustness and efficiency when the optimal policy is unique and deterministic.
As recently proposed by whitehouse2025inference, the non-differentiability inherent in $\eta(\pi^*) = \max_{\pi \in \Pi}$ can be overcome by carefully combining softmax smoothing with first-order de-biasing in the single-period setting. Here, we adopt their concept and extend it to the dynamic setting of MDPs. To the best of our knowledge, this work is the first to consider such a smoothing technique in the context of multiple time periods.
For any real-valued function $h: \mathbb{R}^p \rightarrow \mathbb{R}$ and vector $\mathbf{v} \in \mathbb{R}^d$, define the softmax smoothing approximation $\varphi_{\beta} \{ \cdot \}$ and the multiple softmax operator $\mathbf{sm}_{\beta} \{ \cdot \}$ such that
where $\beta > 0$ denotes the degree of smoothing.
In the static case, the $Q$-function under a policy $\pi$ reduces to $Q (a, s ;\pi) = Q(a, s)$, as $\pi$ simply selects an action $a$ given $x$ to maximize the $Q$-function. As pointed out in whitehouse2025inference, by smoothing the value function $\eta^*(P) = \mathrm{E}_P [\max_{a} Q(a, S)]$ with $ \eta_{\beta}^* = \mathrm{E}_P \big[\mathbf{sm}_{\beta} \big\{ Q(\mathbf{a}, S) \big\} \big], $ one can differentiate $\eta_{\beta}^*(P_{\epsilon})$ with respect to $\epsilon$ and then use the first-order condition to construct a Neyman orthogonal score (or the estimating equation for the pathwise derivative at $\epsilon = 0$). However, challenges arise when extending their framework to the dynamic setting inherent in MDPs: for any fixed $\pi$, the (dynamic) $Q$-function is derived from the fixed point of the Bellman equation, such that $ \eta(\pi) = \mathrm{E} \big[ R + \gamma K_{\pi} \eta(\pi) \big], $ where $K_{\pi}$ is the transition kernel defined in (ref). Consequently, smoothing the value function directly would disrupt the above contraction structure (the basis of fixed-point theory). Fortunately, we leverage the policy optimization perspective found in Entropy-Regularized MDPs neu2017unified: instead of smoothing the value function, we choose to smooth the policy $\pi$. Specifically, we define the smoothing-greedy policy $\pi_{\beta}(P)$ as
It is straightforward to see that $\lim_{\beta \rightarrow +\infty}\pi_{\beta}(P) = \pi^*(P)$.
Our procedure for smoothed nuisance estimation proceeds sequentially:
These nuisance estimates are then plugged into the one-step estimator, leading to
The smoothed estimator (ref) can be regarded as a modified version of our NSAVE estimator, where we replace the sequential evaluation with the smoothing technique. Specifically, we decompose the difference between $\widehat{\eta}_{\beta_N}$ and the true value $\eta(\pi^*)$ as follows: \[
\] As shown in the proof of Theorem (ref), provided consistent nuisance estimates with appropriate convergence rates are selected, the statistical error $R_{\text{SM}, 1N}$ will converge to a normal distribution. Meanwhile, the policy-value error $R_{\text{SM}, 2N}$, controlled by the smoothing parameter $\beta_N$, becomes $o_{P_0}(N^{-1 / 2})$ via the smoothing mechanism rather than sequential evaluation. Thus, intuitively, $\widehat{\eta}_{\beta_N}$ remains a RAL estimator and can achieve semiparametric efficiency. We formalize these results in the following theorem.
From a computational perspective, our smoothed estimator $\widehat{\eta}_{\beta_N}$ is more straightforward to calculate: it only requires estimating two nuisance functionals once, followed by a direct plug-in procedure. Theorem (ref) reveals the trade-off: we require stronger convergence rates for the nuisance parameters. Again, the condition $\omega_Q > 1 / 2$ should not be regarded as a super-consistent convergence rate, given the additional time horizon dimension $T$. Nonetheless, our estimator still achieves semiparametric efficiency under Assumption (ref).
When $\mathcal{A} \times \mathcal{S}$ is finite and small, or more generally when the candidate policy class $\Pi = \{ \pi_1, \dots, \pi_K \}$ is finite, we can employ Post-Selection Inference (PSI) techniques to address the non-regularity. Unlike the smoothing approach, which modifies the target parameter to a smooth approximation $\eta(\pi_{\beta})$, PSI aims to construct a valid confidence interval for the value of the empirically selected policy itself, denoted as $\eta(\widehat{\pi}_N)$, where $\widehat{\pi}_N = \arg\max_{\pi \in \Pi} \widehat{\eta}(\pi)$.
Standard inference that treats $\widehat{\pi}_N$ as a fixed policy fails to account for the winner's curse: the selection process systematically favors policies with positive estimation noise, leading to an upward bias in the naive estimator.
To rigorously correct for this bias while accounting for the potential non-uniqueness of optimal policies (ties) and the high correlation between OPE estimates, we adopt the Two-Step Inference on Multiple Winners framework proposed by petrou2024inference. Our PSI procedure proceeds sequentially as follows:
Various approaches instantiate the template required in the above steps. Here, we utilize the worst-case construction (the primary construction in petrou2024inference), postponing other constructions to Appendix (ref).
To complete Step 1, we first identify indices that are NOT “significantly” worse than any winner $\widehat{k} \in \widehat{\mathcal{A}}_{\text{opt}}$: $\widehat{\mathcal{A}}^{+}: = \left\{ j \in [K] : \big| \widehat\eta(\pi_{\widehat{k}}) - \eta(\pi_{j}) \big| \le z_{1 - \delta_1 / 2} \widehat{\boldsymbol{\Sigma}}_{\widehat{k}, j} / \sqrt{N} ~~~ \text{ for any } ~~~ \widehat{k} \in \widehat{\mathcal{A}}_{\text{opt}} \right\}$, where $z_{1 - \delta_1 / 2} $ denotes the upper $\delta_1/2$-th quantile of a standard normal distribution. The worst-case PSI confidence set is correspondingly given by: \[ \mathcal{C}_{\text{PSI}}^{(\text{WC})} := \bigtimes_{k \in \widehat{\mathcal{A}}_{\text{opt}}} \Bigl[\, \widehat{\eta}(\pi_k) \pm q_{1-(\delta_2 - \delta_1)}(\widehat{\mathcal{A}}^{+}) \,\widehat{\boldsymbol{\Sigma}}_{kk}^{1 / 2}\Bigr], \] where $q_{1-(\delta_2 - \delta_1)}(\widehat{\mathcal{A}}^{+}) $ is the quantile functional defined in (ref) in Appendix (ref).
A straightforward consequence of Corollary (ref) is that if the optimal policy is unique, the length of $\mathcal{C}_N^{\text{PSI}}$ converges to that of the standard oracle confidence interval (oracle efficiency). If there are multiple optimal policies (ties), the interval remains valid by adapting to the worst-case distribution over the set of winners.
We investigate the finite-sample performance of the proposed NSAVE and smoothing-based inference procedures, comparing them with the SAVE estimator shi2022statistical. We consider infinite-horizon Markov decision processes with discrete state and action spaces, a uniform initial state distribution, and a behavior policy satisfying the overlap condition. Three representative regimes are examined: (i) a regular, well-specified setting (Scenario A); (ii) a setting with heavy-tailed reward contamination (Scenario B); and (iii) a structurally misspecified setting with severe state aliasing (Scenario C). The full data-generating mechanisms, tuning parameters, and implementation details are provided in Appendix (ref).
We evaluate both a fixed oracle-optimal policy and data-driven policies learned via double Fitted Q-Iteration. All methods are implemented using cross-fitting. We report the mean squared error (MSE) and empirical coverage probability (ECP) of 95% confidence intervals over 100 Monte Carlo replications. Additional experimental results for Scenarios A and B are deferred to Appendix (ref).
\paragraph{Main findings.} In the regular regimes (Scenarios A and B), both NSAVE and SAVE are asymptotically consistent; however, their finite-sample behaviors differ substantially. As detailed in Appendix (ref), NSAVE attains nominal coverage at markedly smaller sample sizes, reflecting the stability of trajectory-level efficient influence function–based inference combined with studentized batch means. In contrast, SAVE requires significantly larger $N$ and $T$ for its blockwise variance approximation to stabilize. Furthermore, under heavy-tailed reward contamination (Scenario B), NSAVE maintains well-calibrated coverage, whereas SAVE exhibits noticeable distortion.
Most notably, in the structurally misspecified setting with state aliasing (Scenario C), purely model-based methods fail. As shown in Table (ref), SAVE and the smoothing-based plug-in estimator suffer from persistent bias and near-zero coverage. In contrast, NSAVE remains accurate and achieves orders-of-magnitude smaller MSE by leveraging its double-robust correction via importance weighting. Overall, these results demonstrate that NSAVE provides both sharper finite-sample calibration and greater robustness across regular and non-regular regimes.
We apply NSAVE and the smoothing-based method to the OhioT1DM mobile health dataset previously analyzed by shi2022statistical. The data consist of continuous glucose monitoring records, insulin delivery logs, and self-reported events for six patients with type 1 diabetes over an eight-week period. Following the construction in shi2022statistical, we discretize time into non-overlapping 3-hour intervals and define patient-specific state, action, and reward trajectories; full preprocessing details are provided in Appendix (ref). We set the discount factor to $\gamma=0.5$.
For each patient, we estimate an optimal policy and construct confidence intervals for its value using NSAVE and the smoothing approach. We also estimate the value of the observed clinician behavior policy via a plug-in model-based estimator and conduct inference on the value difference $D_i = V(S_{i0};\pi^*) - V(S_{i0}; b_i)$, which quantifies the potential improvement of the learned optimal policy over the observed treatment rule.
Inference is performed separately for two clinically relevant starting times (8:00 am and 2:00 pm on Day 1). Figure (ref) reports the 95% confidence intervals for $D_i$ with $\gamma=0.5$. Across patients and starting times, the estimated value differences are consistently positive, with several confidence intervals excluding zero, indicating statistically significant improvements. Compared with SAVE (as reported in prior analyses), NSAVE yields stable confidence intervals without relying on stringent non-degeneracy conditions. The smoothing-based approach provides a complementary regularized alternative, particularly effective when the optimal policy is nearly deterministic or non-unique. Sensitivity analyses for $\gamma \in \{0.4, 0.7\}$ are provided in Appendix (ref).
Finally, we emphasize that the existence of an efficient influence function for optimal policy values hinges critically on the regularity of the policy optimization map. When the optimal policy is unique, the problem reduces locally to inference for a fixed policy, and classical semiparametric theory applies uehara2022review,shi2025statistical. When optimal policies are non-unique, the value functional becomes non-smooth and non-regular, and standard root-$N$ inference can fail, a phenomenon closely related to non-regular parameters in optimal treatment regimes and post-selection inference laber2014dynamic,whitehouse2025inference.
Our NSAVE procedure provides a stable, efficient solution in the regular regime, while the smoothing approach connects optimal policy evaluation in MDPs to entropy-regularized control neu2017unified and recent smoothing-based inference for max-type functionals whitehouse2025inference. The post-selection confidence sets further complement these methods by offering valid worst-case coverage for sets of optimal policies, extending ideas from selective and multiple-inference frameworks chernozhukov2015valid.
These results clarify both the scope and the limitations of existing approaches such as SAVE shi2022statistical, and suggest that non-regularity is an intrinsic feature of optimal policy inference rather than a technical artifact.
The OhioT1DM dataset analyzed in this study is publicly available at \url{https://webpages.charlotte.edu/rbunescu/data/ohiot1dm/OhioT1DM-dataset.html}.
The authors report there are no competing interests to declare.
The online Supplementary Material contains detailed configurations for the simulations and real data application (Appendix A), alternative confidence set constructions for post-selection inference (Appendix B), proofs of the theoretical results (Appendices C--E), and auxiliary lemmas (Appendix F).