The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
46,994 characters
The Adaptive Doubly Robust Estimator for Policy Evaluation in Adaptive Experiments and a Paradox Concerning Logging Policy
\title{The Adaptive Doubly Robust Estimator\\
for Policy Evaluation in Adaptive Experiments\\
and a Paradox Concerning Logging Policy}
\author[1]{Masahiro Kato\thanks{Corresponding author: masahiro\[email removed]}\ \ }
\author[1]{Shota Yasui}
\author[2]{Kenichiro McAlinn}
\affil[1]{CyberAgent Inc.}
\affil[2]{Temple University}
\maketitle
\begin{abstract}
The \emph{doubly robust} (DR) estimator, which consists of two nuisance parameters, the conditional mean outcome and the logging policy (the probability of choosing an action), is crucial in causal inference. This paper proposes a DR estimator for \emph{dependent samples} obtained from adaptive experiments. To obtain an asymptotically normal semiparametric estimator from dependent samples with non-Donsker nuisance estimators, we propose \emph{adaptive-fitting} as a variant of sample-splitting. We also report an empirical \emph{paradox} that our proposed DR estimator tends to show better performances compared to other estimators utilizing the true logging policy. While a similar phenomenon is known for estimators with i.i.d. samples, traditional explanations based on asymptotic efficiency cannot elucidate our case with dependent samples. We confirm this hypothesis through simulation studies\footnote{A code of the conducted experiments is available at
\url{https://github.com/MasaKat0/adr}.}.
\end{abstract}
\section{Introduction}
\emph{Adaptive experiments}, including efficient treatment effect estimation \citep{Laan2008TheCA,Hahn2011}, treatment regimes \citep{Zhang2012,zhao2012est_ind}, and multi-armed bandit (MAB) problems \citep{Gittins89,lattimore2020bandit}, are widely accepted and used in real-world applications, for example, in social experiments \citep{Hahn2011}, online advertisement \citep{zhang2012jointoptimization}, clinical trials \citep{ChowChang201112,Villar2018}, website optimization \citep{white2012bandit}, and recommendation systems \citep{Lihong2010contextual}. In those various applications, there is a significant interest in off-line counterfactual inference using samples obtained from past trials generated by a logging policy \citep[the probability of choosing an action:][]{ Lihong2010contextual,Li2011}. Among several off-line evaluation criteria, this paper focuses on \emph{off-policy value estimation} \citep[OPVE:][]{Li2015}. The goal of OPVE is to estimate the expected value of the weighted outcome, which includes average treatment effect (ATE) estimation as a special case \citep{hirano2003,bang2005drestimation}.
In OPVE, when the true logging policy is known, there are mainly three types of estimators: inverse probability weighting \citep[IPW:][]{Horvitz1952}, direct method (DM), and augmented IPW \citep[AIPW][]{bang2005drestimation} estimators. IPW-type estimators refer to a sample average of the outcomes weighted by the true logging policy. DM-type estimators refer to a sample average of an estimated conditional outcome. AIPW-type estimators consist of two components: the IPW part and the DM part. More details are given in Section~\ref{sec:prelim}. In addition, when the true logging policy is unknown, we call IPW-type and AIPW-type estimators that use the estimated logging policy EIPW-type and DR-type estimators, respectively. In particular, and the focus of this paper, DR-type estimators are crucial in OPVE, since in most applications of interest the true logging policy is either unknown, difficult, or costly to obtain, making IPW and AIPW-type estimators often infeasible.
While existing OPVE methods typically assume that the samples are \emph{independent and identically distributed} \citep[i.i.d.:][]{hirano2003,dudik2011doubly}, adaptive experiments usually allow logging policies to be updated based on past observations, under which the generated samples are non-i.i.d. In this case, theoretical results shown under i.i.d. assumptions, such as consistency and asymptotic normality, are not guaranteed. In particular, the asymptotic normality is critical, as we can obtain $\sqrt{T}$-rate confidence intervals, where $T$ is the sample size. For this problem, \citet{Laan2008TheCA} proposed an asymptotically normal AIPW estimator under dependent samples, with several follow-up studies, including \citet{hadad2019}, \citet{Kato2020adaptive}, and \citet{kelly2020batched}.
However, a DR-type estimator for dependent samples has never been formally proposed, despite its importance.
This paper proposes an asymptotically normal DR estimator for dependent samples. We call this DR estimator the \emph{adaptive DR} (ADR) estimator. There are mainly two difficulties concerning this estimator. First, we cannot use the CLT for i.i.d. or martingale difference sequences, as done in \citet{Laan2008TheCA}, \citet{hadad2019}, and \citet{Kato2020adaptive}. Second, when nuisance parameters are estimated with complicated models, the Donsker condition does not hold in general, which is required to show the asymptotic normality of semiparametric estimators, such as our proposed ADR estimator. In this paper, to solve these problems, we propose adaptive-fitting, inspired by the double/debiased machine learning for i.i.d. samples \citep{ChernozhukovVictor2018Dmlf}. We also find that by using the ADR estimator, not only can OPVE be done without knowing the true logging policy, the ADR estimator paradoxically often outperforms the performance of AIPW estimators that use the true logging policy. We list our contributions as follows.
{\bf (I) Asymptotic normality of the ADR estimator.}
We show the asymptotic normality of our proposed ADR estimator. To the best of our knowledge, a DR-type estimator for dependent samples obtained in adaptive experiments has not been formally proposed.
{\bf (II) Adaptive fitting.}
While the asymptotic normalities of IPW and AIPW estimators are shown by the martingale CLT, we cannot directly apply this technique to our proposed ADR estimator. This is because the martingale condition does not hold when the true logging policy is unknown. To solve this problem, we utilize a novel sample-splitting method called adaptive-fitting to show the asymptotic normality of the ADR estimator. We generalize this technique as a variant of sample-splitting \citep{ChernozhukovVictor2018Dmlf} for semiparametric inference with dependent samples without the Donsker condition. Unlike existing sample-splitting methods, which assume that samples are i.i.d., the proposed adaptive-fitting is applicable to non-i.i.d. samples. Note that this adaptive-fitting is different from martingale based estimators, such as \citet{Laan2008TheCA} and \citet{hadad2019}, which requires the true logging policy.
{\bf (III) Empirical paradox on using estimated logging policy.}
Through experimental studies, we investigate the empirical performance of our proposed ADR estimator. We find that the ADR estimator often exhibits improved performances over other estimators using the true logging policy, such as the AIPW estimator. Estimators requiring the true logging policy tend to empirically be unstable due to the instability of the nuisance parameter; the instability being caused by the logging policy being near zero before convergence. Our finding implies that the ADR estimator mitigates this instability. We call this phenomenon a \emph{paradox concerning logging policy} because using more information (the true logging policy) does not improve empirical performance. A similar paradox is known for IPW-type estimators for i.i.d. samples \citep{hahn1998role,Henmi2004paradox}. However, we cannot explain our paradox from a conventional semiparametric efficiency perspective \citep{hahn1998role,hirano2003}, as the asymptotic variances are the same between the ADR estimator (with an estimated logging policy) and the AIPW estimator (with the true logging policy), unlike IPW-type estimators.
{\bf Organization of this paper.} In Section~\ref{sec:prb}, we formulate our problem. In Section~\ref{sec:prelim}, we introduce existing OPVE estimators for dependent samples. In Section~\ref{sec:main}, we propose the ADR estimator, which is our main contribution. We call the method used for the ADR estimator adaptive-fitting, and generalize it in Section~\ref{sec:relation_ddm}. In Section~\ref{sec:paradox}, we report a paradox through simulation studies and explain the cause of this phenomenon. In Section~\ref{sec:exp}, we conduct experiments using benchmark datasets.
\vspace{-0.1cm}
\section{Problem setting}
\label{sec:prb}
Suppose that there is a time series $1,2,\dots, T$, and denote the set as $[T] = \{1,\dots,T\}$. For $t\in[T]$, let $A_t$ be an \emph{action} in $\mathcal{A}=\{1,2,\dots,K\}$, $X_t$ be a \emph{covariate} observed by a decision maker when choosing an action, and $\mathcal{X}$ be its space. Following the Neyman–Rubin causal model \citep{rubin1974}, let a reward at period $t$ be $Y_t=\sum^K_{a=1}\mathbbm{1}[A_t = a]Y_t(a)$, where $Y_t(a):\mathcal{A}\to\mathbb{R}$ is a potential (random) outcome. We have a dataset $\big\{(X_t, A_t, Y_t)\big\}^{T}_{t=1}$ with the following data-generating process (DGP):
\begin{align}
\label{eq:DGP}
(X_t, A_t, Y_t(A_t))\sim p(x)p_t(a|x)p(y_{a}|x),
\end{align}
where $p(x)$ denotes the density of the covariate $X_t$, $p_t(a| x)$ denotes the probability of choosing an action $a$ conditional on a covariate $x$ at period $t$, and $p(y_a| x)$ denotes the density of a reward $Y_t(a)$ conditional on a covariate $x$. We assume that $p(x)$ and $p(y_a| x)$ are invariant across periods; that is, $\{(X_t, Y_t(1),\dots, Y_t(K))\}^T_{t=1}$ is i.i.d., but $p_t(a| x)$ can take different values across periods based on past observations. In this case, the samples $\big\{(X_t, A_t, Y_t)\big\}^{T}_{t=1}$ are correlated over time, that is, the samples are not i.i.d. Let $\Omega_{t-1}=\{X_{t-1}, A_{t-1}, Y_{t-1}, \dots, X_{1}, A_1, Y_{1}\}$ be the history and $\mathcal{M}_{t-1}$ be a set of possible histories until the $t$-th period. The probability $p_t(a| x)$ is determined by a \emph{logging policy} $\pi_t:\mathcal{A}\times\mathcal{X}\times\mathcal{M}_{t-1}\to(0,1)$, such that $\sum^K_{a=1}\pi_t(a| x, \Omega_{t-1}) = 1$, which is a function of a covariate $X_t$, an action $A_t$, and history $\Omega_{t-1}$. We also assume that $\pi_t$ is conditionally independent of $Y_t(a)$ to satisfy unconfoundedness (Remark~\ref{rem:unconfoundedness}).
\begin{remark}[Stable unit treatment value assumption (SUTVA)]
\label{rem:sutva}
The DGP implies SUTVA; that is, $p(y_a| x)$ is invariant for any $p(a| x) $\citep{Rubi:86}.
\end{remark}
\begin{remark}[Unconfoundedness]
\label{rem:unconfoundedness}
In this paper, unconfoundedness refers to independence between $(Y_t(1),\dots, Y_t(K))$ and $A_t$, conditional on $X_t$ and $\Omega_{t-1}$, which is required for identification.
\end{remark}
\subsection{OPVE}
\label{sec:OPVE}
The goal of OPVE is to estimate the expected value of the sum of the outcomes $Y_t(a)$ weighted by an evaluation function $\pi^\mathrm{e}:\mathcal{A}\times \mathcal{X} \to\mathbb{R}$; that is,
\begin{align}
R(\pi^\mathrm{e}) &:=\mathbb{E}_{(X_t, Y_t(1),\dots,Y_t(K))}\left[\sum^K_{a=1}\pi^\mathrm{e}(a | X_t)Y_t(a)\right],
\end{align}
where the expectation $\mathbb{E}_{(X_t, Y_t(1),\dots, Y_t(K))}$ is taken over $(X_t, Y_t(1), \dots, Y_t(K))$. We denote it as $\mathbb{E}$, when there is no ambiguity. As with \citet{dudik2011doubly}, the evaluation function is usually referred to as an evaluation policy, where $\pi^\mathrm{e}(a| x)\in [0,1]$ and the sum is $1$. However, to not restrict it so we can include other forms, such as the ATE, we refer to it differently. The ATE is also a special case of OPVE for $\mathcal{A}=\{1,2\}$, where $\pi^\mathrm{e}(1| x) = 1$ and $\pi^\mathrm{e}(2| x) = -1$. To identify $R(\pi^\mathrm{e})$, we assume the overlap in the distributions of policies, convergence of $\pi_{t-1}$, and the boundedness of reward.
\begin{assumption}\label{asm:overlap_pol}
For all $t\in[T]$, $a\in\mathcal{A}$, $x\in\mathcal{X}$, and $\Omega_{t-1}\in\mathcal{M}_{t-1}$, there exists a constant $C_\pi > 0$, such that $\left| \frac{\pi^\mathrm{e}(a| x)}{\pi_t(a| x, \Omega_{t-1})}\right| \leq C_\pi$.
\end{assumption}
Assumption~\ref{asm:overlap_pol} equivalently means that $\pi_{t}(a| x) > 0$ for all $a\in\mathcal{A}$ and $x\in\mathcal{X}$.
\begin{assumption}
\label{asm:stationarity}
For all $x\in\mathcal{X}$ and $a\in\mathcal{A}$, $\pi_t(a| x, \Omega_{t-1})-\tilde{\pi}(a| x)\xrightarrow{\mathrm{p}}0$, where $\tilde{\pi}: \mathcal{A}\times\mathcal{X}\to(0,1)$.
\end{assumption}
\begin{assumption}\label{asm:bounded_reward}
For all $t\in[T]$ and $a\in\mathcal{A}$, there exists a constant $C_Y > 0$, such that $|Y_t(a)| \leq C_Y$.
\end{assumption}
Although the reader may feel that Assumption~\ref{asm:overlap_pol}--\ref{asm:stationarity} and the SUTVA \ref{rem:sutva} are strong, we adopt it as a simple and basic case for the application, in order to introduce adaptive-fitting and the ADR estimator. We can extend the proposed method for different cases, such as when the data has the structure of batches or when the average of logging policy converges \citep{Kato2021nonstationary}. Note that the convergence assumption (Assumption~\ref{asm:stationarity}) is also explicitly or implicitly required in other studies, such as \citet{Laan2008TheCA} and \citet{hadad2019}.
\subsection{Notations}
We denote $\mathbb{E}[Y_t(a)| x]$ and $\mathrm{Var}(Y_t(a)| x)$ as $f^*(a, x)$ and $v^*(a, x)$, respectively. Let $\hat{f}_{t}(a, x)$ be an estimator of $f^*(a, x)$ constructed from $\Omega_{t}$. Let $\mathcal{N}(\mu, \mathrm{var})$ be the normal distribution with the mean $\mu$ and the variance $\mathrm{var}$. For a random variable $Z$ with density $p(z)$ and function $\mu$, let $\|\mu(Z)\|_2=\int |\mu(z)|^2 p(z) dz $ be the $L^{2}$-norm.
\section{Preliminaries of OPVE}
\label{sec:prelim}
\subsection{Existing estimators}
For estimating $R(\pi^\mathrm{e})$ from dependent samples, existing studies propose various estimators. An adaptive version of the IPW estimator is defined as $R^{\mathrm{AdaIPW}}_T(\pi^\mathrm{e})=\frac{1}{T}\sum^T_{t=1}\sum^K_{a = 1}\frac{\pi^\mathrm{e}(A_t| X_t)\mathbbm{1}[A_t=a]Y_t }{\pi_{t-1}(A_t| X_t, \Omega_{t-1})}$ \citep{Laan2008TheCA}. If the model specification is correct, the direct method (DM) estimator $\frac{1}{T}\sum^T_{t=1}\sum^K_{a=1}\pi^\mathrm{e}(a | x)\hat{f}_{T}(a| X_t)$ is known to be consistent to $R(\pi^\mathrm{e})$. As an adaptive version of the AIPW estimator, \citet{Laan2008TheCA} proposed an estimator $\widehat{R}^{\mathrm{AIPW}}_T(\pi^\mathrm{e})$ defined as
\begin{align*}
\frac{1}{T}\sum^T_{t=1}\sum^K_{a=1}\Bigg\{\frac{\pi^\mathrm{e}(a| X_t)\mathbbm{1}[A_t=a]\left(Y_t - \hat{f}_{t-1}(a, X_t)\right) }{\pi_{t-1}(a| X_t, \Omega_{t-1})}+ \pi^\mathrm{e}(a| X_t)\hat{f}_{t-1}(a, X_t)\Bigg\}.
\end{align*}
Using the martingale property, \citet{Laan2008TheCA} showed asymptotic normality under Assumption~\ref{asm:stationarity}. \citet{hadad2019} and \citet{Kato2020adaptive} organized the results ( Proposition~\ref{prp:asymp_dist_a2ipw}).
\subsection{Asymptotic efficiency}
\label{sec:asymp_eff}
We are often interested in the asymptotic efficiency of estimators. The lower bound of the asymptotic variance is defined for an estimator under some posited models of the DGP \eqref{eq:DGP}. As with the Cram\'{e}r-Rao lower bound for the parametric model, we can also define the lower bound for the non- or semiparametric model \citep{bickel98}.
The semiparametric lower bound of the DGP~\eqref{eq:DGP} under $p_1(a| x)=\cdots=p_T(a| x)=p(a| x)$ is given as follows \citep{hahn1998role,narita2019counterfactual}:
\begin{align*}
\Psi(\pi^\mathrm{e}, p) = \mathbb{E}\Bigg[\sum^{K}_{a=1}\frac{\big(\pi^\mathrm{e}(a| X_t)\big)^2v^*(a, X_t)}{p(a| X_t)}+ \left(\sum^{K}_{a=1}\pi^\mathrm{e}(a| X_t)f^*(a, X_t) - R(\pi^\mathrm{e})\right)^2\Bigg].
\end{align*}
The asymptotic variance of the asymptotic distribution is also known as the asymptotic mean squared error (MSE); that is, an OPVE estimator achieving the semiparametric lower bound also minimizes the MSE to the true value $R(\pi^\mathrm{e})$, not just obtaining a tight confidence interval.
\subsection{Related work}
There are various studies related to OPVE, including ATE estimation, under the assumption that samples are i.i.d. \citep{hahn1998role,hirano2003,dudik2011doubly,wang2017optimal,narita2019counterfactual,Bibaut2019moreffficient,Oberst2019}. There are also several studies extending these methods to OPVE from dependent samples \citep{Laan2008TheCA,Laan2016onlinetml,Luedtke2016,hadad2019,Kato2020adaptive}.
The AIPW estimator for dependent samples are proposed by \citet{Laan2008TheCA}. \citet{hadad2019} proposed an evaluation weight to stabilize the estimator, which shares a similar motivation with weight clipping, or shrinkage, when i.i.d. samples are given \citep{Bembom2008,Bottou2013,Wang2017,Su2019,Su2020}. \citet{Kato2020adaptive} showed a non-asymptotic confidence interval of the AIPW estimator. \citet{Laan2016onlinetml} and \citet{Luedtke2016} proposed an OPVE method without the convergence of the logging policies by using batches and standardization, respectively. Note that \citet{hadad2019} cannot weaken the assumptions regarding the logging policy, unlike \citet{Luedtke2016}. \citet{kelly2020batched} proposed an estimator similar to \citet{Laan2016onlinetml} and applied it to linear regression. Estimators proposed by \citet{Laan2008TheCA}, \citet{Luedtke2016}, \citet{hadad2019}, \citet{Kato2020adaptive}, and \citet{kelly2020batched} require the true logging policy, unlike our ADR estimator.
A semiparametric estimator usually requires the Donsker condition for its asymptotic normality \citep{bickel98}. For semiparametric inference without the Donsker condition, sample-splitting is a typical approach \citep{klaassen1987,ZhengWenjing2011CTME,ChernozhukovVictor2018Dmlf}. \citet{ChernozhukovVictor2018Dmlf} referred to sample-splitting as \emph{cross-fitting} and the semiparametric inference using cross-fitting as \emph{double-debiased machine learning} (DML). For off-policy evaluation of reinforcement learning from dependent samples, \citet{KallusNathan2019EBtC} proposed a mixingale-based sample-splitting.
\section{The ADR estimator}
\label{sec:main}
For OPVE with dependent samples, this paper proposes the ADR estimator $\widehat{R}^{\mathrm{ADR}}_T(\pi^\mathrm{e})$ defined as
\begin{align*}
\frac{1}{T}\sum^T_{t=1}\sum^K_{a=1}\Bigg\{\frac{\pi^\mathrm{e}(a| X_t)\mathbbm{1}[A_t=a]\left(Y_t - \hat{f}_{t-1}(a, X_t)\right) }{\hat{g}_{t-1}(a| X_t)}+ \pi^\mathrm{e}(a| X_t)\hat{f}_{t-1}(a, X_t)\Bigg\},
\end{align*}
where $\hat{g}_{t-1}$ is an estimator of $\pi_{t-1}$, constructed only from $\Omega_{t-1}$. We can use standard regression methods for constructing $\hat{f}_{t-1}$ and $\hat{g}_{t-1}$ if they satisfy the following assumptions.
\begin{assumption}
\label{asm:conv_rate1}
For $p, q > 0$ such that $p+q = 1/2$, $\|\hat{g}_{t-1}(a| X_t) - \pi_{t-1}(a| X_t, \Omega_{t-1})\|_{2}=\mathrm{o}_{p}(t^{-p})$, and $\|\hat{f}_{t-1}(a,X_t)-f^*(a,X_t)\|_2=\mathrm{o}_{p}(t^{-q})$, where the expectation of the norm is taken over $X_t$.
\end{assumption}
\begin{assumption}
\label{asm:bound_nuisance1}
There exists a constant $C_f$ such that $|\hat{f}_{t-1}(a, x)| \leq C_f$ $\forall a\in\mathcal{A},x\in\mathcal{X}$.
\end{assumption}
\begin{assumption}
\label{asm:bound_nuisance2}
There exist a constant $C_g$ such that $0 < \left|\frac{\pi^\mathrm{e}(a| x)}{\hat{g}_{t-1}(a| x)}\right| \leq C_g$ $\forall a\in\mathcal{A},x\in\mathcal{X}$.
\end{assumption}
Assumption~\ref{asm:conv_rate1} requires convergence rates standard in regression estimators. For instance, we can apply nonparametric estimators proposed in MAB problems \citep{yang2002,qian2016kernel}. Under the assumptions, we show the asymptotic normality of the ADR estimator.
\begin{theorem}[Asymptotic distribution of an ADR estimator]
\label{thm:asymp_dist_adr}
Under Assumptions~\ref{asm:overlap_pol}--\ref{asm:bound_nuisance2},
\begin{align*}
\sqrt{T}\left(\widehat{R}^{\mathrm{ADR}}_T(\pi^\mathrm{e})-R(\pi^\mathrm{e})\right)\xrightarrow{d}\mathcal{N}\left(0, \Psi(\pi^\mathrm{e}, \tilde{\pi})\right).
\end{align*}
\end{theorem}
Let us put forward the following assumption.
\begin{assumption}
\label{asm:pointwise_f}
For all $x\in\mathcal{X}$ and $a\in\mathcal{A}$, $\hat{f}_{t-1}(a, x)-f^*(a, x)\xrightarrow{\mathrm{p}}0$.
\end{assumption}
The proof is shown in Appendix~\ref{appdx:proof_main}, which uses the following proposition from \citet{Kato2020adaptive}.
\begin{proposition}[Asymptotic distribution of an AIPW estimator (Corollary~1, \citet{Kato2020adaptive}).]
\label{prp:asymp_dist_a2ipw}
Under Assumptions~\ref{asm:overlap_pol}--\ref{asm:bounded_reward}, \ref{asm:bound_nuisance1}, and \ref{asm:pointwise_f}, $\sqrt{T}\left(\widehat{R}^{\mathrm{AIPW}}_T(\pi^\mathrm{e})-R(\pi^\mathrm{e})\right)\xrightarrow{d}\mathcal{N}\left(0, \Psi(\pi^\mathrm{e}, \tilde{\pi})\right)$.
\end{proposition}
\begin{remark}[Consistency and double robustness]
The ADR estimator has double robustness, as with standard DR estimators; that is, if either $\hat{f}$ or $\hat{g}$ is consistent, the ADR estimator is also consistent.
\end{remark}
\begin{theorem}[Consistency]
Under Assumptions~\ref{asm:overlap_pol}--\ref{asm:bounded_reward} and \ref{asm:bound_nuisance1}--\ref{asm:bound_nuisance2}, if either $\hat{f}_{t-1}$ or $\hat{g}_{t-1}$ is consistent, $\widehat{R}^{\mathrm{ADR}}_T\xrightarrow{\mathrm{p}} R(\pi^\mathrm{e})$.
\end{theorem}
We can prove this by the law of large numbers for martingales (Proposition~\ref{prp:mrtgl_WLLN} in Appendix~\ref{appdx:prelim}).
\vspace{-0.2cm}
\paragraph{Donsker condition.} The main reason for using step-wise estimators $\{\hat{f}_{t-1}\}^T_{t=1}$ and $\{\hat{g}_{t-1}\}^T_{t=1}$ is to regard them as constants in the expectation conditioned on $\Omega_{t-1}$. The motivation is shared with DML. For asymptotic normality shown in Theorem~\ref{thm:asymp_dist_adr}, we do not impose the Donsker condition on the nuisance estimators, $\hat{f}_{t-1}$ and $\hat{g}_{t-1}$, but only require the convergence rate conditions. We call this sample-splitting method, using $\hat{f}_{t-1}$ and $\hat{g}_{t-1}$, \emph{adaptive-fitting} and discuss again in Section~\ref{sec:relation_ddm}. In the MAB problem, convergence rate conditions in nonparametric regression, such as a nearest-neighbor regression \citep{yang2002}, Nadaraya-Watson regression \citep{qian2016kernel}, random forest \citep{Feraud2016}, kernelized linear models \citep{Chowdhury2017ker}, and neural networks \citep{Zhou2020}, have been shown. Note that we need to slightly modify the results for each situation because they are influenced by time-series behavior of logging probabilities.
\vspace{-0.2cm}
\paragraph{Convergence rate of the logging policy.} In the main theorem, we do not explicitly describe the convergence rate of the logging policy. However, from Assumption~\ref{asm:conv_rate1}, which requires $\|\hat{g}_{t-1}(a| X_t) - \pi_{t-1}(a| X_t, \Omega_{t-1})\|_{2}=\mathrm{o}_{p}(t^{-p})$, $\|\tilde{\pi}(a| x) - \pi_{t-1}(a| X_t, \Omega_{t-1})\|_{2}=\mathrm{o}_{p}(t^{-p})$ is also required.
\vspace{-0.2cm}
\paragraph{Theoretical comparison between AIPW and ADR estimators.} There are two major differences between AIPW and ADR estimators. First, the AIPW estimator requires {\it a priori} knowledge of the true logging policy, but the ADR estimator does not. Another main difference between the two is the convergence rate of nuisance estimators $\hat{g}_{t-1}$ and $\hat{f}_{t-1}$. For the AIPW estimator, only the uniform convergence in probability is required, but the ADR estimator requires specific convergence rates on $\hat{g}_{t-1}$ and $\hat{f}_{t-1}$. This difference comes from unbiasedness. The AIPW estimator is unbiased; therefore, the convergence of the asymptotic variance is essential, where it converges with $\mathrm{o}_{p}(1)$ if $\hat{f}$ and $\pi_t$ is $\mathrm{o}_{p}(1)$. Thus, the AIPW estimator does not require specific convergence rates. On the other hand, the ADR estimator requires the asymptotic bias term to vanish in a specific order. For this purpose, it imposes specific convergence rates on the nuisance estimators. From another perspective, a standard DML for i.i.d. samples and Theorem~\ref{thm:asymp_dist_adr} require $\|\hat{g}_{t-1}(a| X_t) - \pi_{t-1}(a| X_t, \Omega_{t-1})\|_{2}=\mathrm{o}_{p}(t^{-p})$, $\|\hat{f}_{t-1}(a,X_t)-f^*(a,X_t)\|_2=\mathrm{o}_{p}(t^{-q})$, and $p+q=1/2$.
\vspace{-0.3cm}
\paragraph{Asymptotic efficiencies.}
As shown in Theorem~\ref{thm:asymp_dist_adr} and Proposition~\ref{prp:asymp_dist_a2ipw}, ADR and AIPW estimators achieve the semiparametric lower bound. On the other hand, the asymptotic variance of the IPW estimator using the true logging policy $\pi_t$ is larger than the lower bound \citep{hirano2003,Laan2008TheCA,Kato2020adaptive}. Although it is known that the IPW estimator using the estimated logging policy can achieve the lower bound under some conditions with i.i.d. samples \citep{hirano2003}, the asymptotic property under dependent samples is still unknown.
\section{Adaptive-fitting: DML for dependent samples}
\label{sec:relation_ddm}
We generalize the method used to derive the asymptotic normality of the ADR estimator as a variant of DML. Let us define the parameter of interest $\theta_0$ that satisfies $\mathbb{E}[\psi(W_t; \theta_0, \eta_0)] = 0$, where $\{W_t\}^T_{t=1}$ are observations, $\eta_0$ is a nuisance parameter, and $\psi$ is a score function. We consider obtaining an asymptotic normal estimator of $\theta_0$ when using complex and data-adaptive regression methods, such as random forests, neural networks, and Lasso, to estimate the nuisance parameter $\eta_0$. Under such a situation, the Donsker condition does not hold, in general. Sample-splitting is a typical approach to control the complexities of semiparametric inference without the Donsker condition \citep{klaassen1987,ZhengWenjing2011CTME,ChernozhukovVictor2018Dmlf}. In DML of \citet{ChernozhukovVictor2018Dmlf}, the dataset with i.i.d. samples $\{W_t\}^T_{t=1}$ is separated into several subgroups. Then, a semiparametric estimator for each subgroup is constructed, but the nuisance estimators are constructed from the other subgroups. Here, the nuisance estimators are independent in the expectation conditioned on the other subgroups. Thus, the complexities are controlled without the Donsker condition, and standard nonparametric convergence rate conditions on nuisance estimators suffice to show the asymptotic normality of the semiparametric estimator. \citet{ChernozhukovVictor2018Dmlf} called the method cross-fitting.
\begin{figure}[t]
\vspace{-0.5cm}
\begin{center}
\includegraphics[width=100mm]{concept2-eps-converted-to.pdf}
\end{center}
\vspace{-0.3cm}
\caption{The difference between cross-fitting (left) and adaptive-fitting (right). }
\label{fig:concept}
\vspace{-0.15cm}
\end{figure}
However, cross-fitting of \citet{ChernozhukovVictor2018Dmlf} cannot be applied when samples $\{W_t\}^T_{t=1}$ are dependent. Therefore, the ADR estimator uses step-wise nuisance estimators based on past observations $\Omega_{t-1}$. This construction is inspired by \citet{Laan2016onlinetml}, which proposed a sample-splitting for batched dependent samples. The nuisance estimators are independent in the expectation over the $t$-th samples conditioned on $\Omega_{t-1}$. Thus, we can regard our sample-splitting as another approach for DML. We call this step-wise construction \emph{adaptive-fitting}, in contrast to cross-fitting. We briefly summarize the procedure as follows: (i) at each step $t=1,2,\dots,T$, we estimate $\eta_0$ only using $\{W_s\}^{t-1}_{s=1}$ and denote the estimaters as $\{\eta_{t-1}\}^T_{t=1}$; (ii) then, we substitute $W_{t}$ and $\eta_{t-1}$ into $\psi$ and obtain an estimator $\hat{\theta}_T$ of $\theta_0$ by solving $\frac{1}{T}\sum^T_{t=1}\psi(W_t, \hat{\theta}_T, \eta_{t-1}) = 0$. Under this construction, if (a) an estimator $\check{\theta}_T$ obtained by solving $\frac{1}{T}\sum^T_{t=1}\psi(W_t, \check{\theta}_T, \eta_{0}) = 0$ has the asymptotically normal distribution and (b) the asymptotic bias decays with the rate faster than $\mathrm{o}_{p}(-1/\sqrt{t})$, we can obtain the asymptotically normal estimator of $\theta_0$. In Figure~\ref{fig:concept}, we illustrate the difference between cross-fitting (left) and adaptive-fitting (right). The y-axis represents samples used for constructing the OPVE estimator with nuisance estimators using samples represented by the x-axis. In the left graph, for $T/K$ samples, we calculate the sample average of the component including nuisance estimators based on the other $(K-1)T/K$ samples. In the right graph, at period $t$, we use nuisance estimators constructed from samples at $t=1,2,\dots,t-1$. To satisfy (a), we assumed that the convergence of the logging policy (Assumption~\ref{asm:stationarity}). There are other ways to satisfy this condition. For example, we can assume the existence of batches as \citet{Laan2016onlinetml} and \citet{kelly2020batched}. We briefly introduce the ADR estimator for this case in Appendix~\ref{appdx:batch}.
\section{Paradox concerning logging policy}
\label{sec:paradox}
The advantage of the ADR estimator is that, unlike the AIPW estimator, it does not require the true logging policy. However, we find that the ADR estimator often achieves a smaller MSE than the AIPW estimator, which requires the true logging policy. We refer to this phenomenon as a paradox concerning logging policy.
\subsection{Instability of the AIPW estimator}
\label{sec:a2ipw}
\citet{hadad2019} pointed out that the AIPW estimator tends to be unstable when the nuisance parameter $\pi_{t-1}$ can take a value close to zero before converging to $\tilde{\pi}$. To prevent this instability, \citet{hadad2019} proposed the Adaptively Weighted AIPW (AW-AIPW) estimator by adding evaluation weights, which converge to a constant almost surely, to the AIPW estimator proposed in other studies, such as \citet{Laan2008TheCA}. Note that the AW-AIPW estimator cannot mitigate the assumptions on the logging policy $\pi_{t-1}$ required in the AIPW \citep{Laan2008TheCA} and our ADR estimator, namely $\pi_{t-1}$ is non-zero and converges. The aim of adaptive weighting is not in theoretical improvement of the AIPW estimator, but in its empirical stabilization.
\subsection{Sensitivity of the ADR estimator to the logging policy.}
The ADR estimator replaces the true logging policy $\pi_{t-1}$ with its estimator. We find that by replacing the true logging policy with its estimator, we can control the instability caused from $\pi_{t-1}$ by constructing well-formed $\hat{g}_{t-1}$. For example, if we know that the range of $\tilde{\pi}$ is $(\varepsilon, 1-\varepsilon)$ for $0< \varepsilon < 1-\varepsilon < 1$, we can add the clipping technique when estimating $\hat{g}_{t-1}$. For example, in early periods, we can set $\hat{g}_{t-1}$ as $0.5$. Such techniques is greatly beneficial.
Note that if the range of $\tilde{\pi}$ is truly $(\varepsilon, 1-\varepsilon)$, the clipping of $\hat{g}_{t-1}$ does not cause clipping bias; that is, the estimator correctly converges to the asymptotic distribution shown in Theorem~\ref{thm:asymp_dist_adr}.
The main difference between the AW-AIPW estimator and the ADR estimator is that the former requires information of the true logging policy $\pi_{t-1}$, while the latter does not by replacing it with its estimator. Note that the AW-AIPW estimator requires the same assumptions on $\pi_{t-1}$ as the AIPW and ADR estimator and loses unbiasedness due to the evaluation weight. In addition, the construction of the adaptive weight is not obvious when there are covariates $X_t$ because all but one method for weight construction in \citet{hadad2019} does not consider covariates. Our finding also shares motivation with weight clipping for i.i.d. samples, such as \citet{Bottou2013}. We can regard the AW-AIPW estimator as a variant of this method, where the weight clipping can decay as $\pi_t$ converges, and be ignored asymptotically. However, our interest is in asymptotic normality from dependent samples without the true logging policy, which is not discussed in the literature.
\begin{table*}[t]
\vspace{-0.1cm}
\caption{The results of Section~\ref{sec:a2ipw} with sample sizes $T=250, 500,750$. We show the RMSE, SD, and coverage ratio of the confidence interval (CR). We highlight in red bold two estimators with the lowest RMSE.
Estimators with asymptotic normality are marked with $\dagger$, and estimators that do not require the true logging policy are marked with $*$.}
\vspace{-0.3cm}
\label{tbl:exp_table1}
\begin{center}
\scalebox{0.61}[0.61]{
\begin{tabular}{|l||rrr||rrr|rrr|rrr||rrr|rrr|}
\hline
\multicolumn{19}{|c|}{LinUCB policy}\\
\hline
{$T$} & \multicolumn{3}{c||}{ADR $\dagger*$} & \multicolumn{3}{c|}{IPW $\dagger$} & \multicolumn{3}{c|}{AIPW $\dagger$} &
\multicolumn{3}{c||}{AW-AIPW $\dagger$} & \multicolumn{3}{c|}{DM $*$} & \multicolumn{3}{c|}{EIPW $*$}\\
\hline
& RMSE & SD & CR & RMSE & SD & CR & RMSE & SD & CR & RMSE & SD & CR & RMSE & SD & CR & RMSE & SD & CR\\
\hline
250 & \textcolor{red}{\textbf{0.054}} & 0.003 & 0.97 & 0.117 & 0.034 & 0.87 & 0.106 & 0.036 & 0.90 & 0.152 & 0.022 & 0.09 & 0.073 & 0.006 & 0.12 & 0.103 & 0.017 & 0.82 \\
500 & \textcolor{red}{\textbf{0.040}} & 0.002 & 0.94 & 0.078 & 0.014 & 0.93 & 0.059 & 0.007 & 0.98 & 0.179 & 0.020 & 0.00 & 0.046 & 0.002 & 0.16 & 0.121 & 0.013 & 0.40 \\
750 & \textcolor{red}{\textbf{0.033}} & 0.001 & 0.97 & 0.064 & 0.008 & 0.90 & 0.059 & 0.006 & 0.97 & 0.176 & 0.015 & 0.00 & 0.040 & 0.002 & 0.15 & 0.122 & 0.012 & 0.22 \\
\hline
\end{tabular}
}
\end{center}
\vspace{-0.3cm}
\begin{center}
\scalebox{0.61}[0.61]{
\begin{tabular}{|l||rrr||rrr|rrr|rrr||rrr|rrr|}
\hline
\multicolumn{19}{|c|}{LinTS policy}\\
\hline
{$T$} & \multicolumn{3}{c||}{ADR $\dagger*$} & \multicolumn{3}{c|}{IPW $\dagger$} & \multicolumn{3}{c|}{AIPW $\dagger$} &
\multicolumn{3}{c||}{AW-AIPW $\dagger$} & \multicolumn{3}{c|}{DM $*$} & \multicolumn{3}{c|}{EIPW $*$}\\
\hline
& RMSE & SD & CR & RMSE & SD & CR & RMSE & SD & CR & RMSE & SD & CR & RMSE & SD & CR & RMSE & SD & CR\\
\hline
250 & \textcolor{red}{\textbf{0.049}} & 0.002 & 0.96 & 0.128 & 0.037 & 0.82 & 0.107 & 0.031 & 0.93 & 0.121 & 0.016 & 0.20 & 0.069 & 0.005 & 0.15 & 0.077 & 0.011 & 0.89 \\
500 & \textcolor{red}{\textbf{0.035}} & 0.001 & 0.95 & 0.151 & 0.139 & 0.91 & 0.122 & 0.086 & 0.91 & 0.136 & 0.014 & 0.06 & 0.048 & 0.002 & 0.13 & 0.090 & 0.007 & 0.57 \\
750 & \textcolor{red}{\textbf{0.028}} & 0.001 & 0.97 & 0.069 & 0.010 & 0.90 & 0.053 & 0.005 & 0.91 & 0.154 & 0.013 & 0.00 & 0.036 & 0.001 & 0.14 & 0.113 & 0.009 & 0.23 \\
\hline
\end{tabular}
}
\end{center}
\vspace{-0.5cm}
\end{table*}
\begin{figure}[t]
\begin{center}
\includegraphics[width=130mm]{fig_compare-eps-converted-to.pdf}
\end{center}
\vspace{-0.5cm}
\caption{This figure illustrates the error distributions of estimators for OPVE from dependent samples generated with the LinUcB(left) and LinTS(right) with the sample size $750$. We smoothed the error distributions using kernel density estimation. Estimators with asymptotic normality are marked with $\dagger$, and estimators that do not require a true logging policy are marked with $*$.}
\label{fig:bias}
\vspace{-0.5cm}
\end{figure}
\subsection{Simulation studies for OPVE estimators}
\label{sec:num_exp}
For reasons given in the previous section, the ADR estimator may empirically perform better than the AIPW estimator because the estimator $\hat{g}_{t-1}$ stabilizes the ADR estimator by absorbing the instability of $\pi_{t-1}$.
We empirically show this paradox using a synthetic dataset. We generate an artificial pair of covariate and potential outcome, $(X_t, Y_t(1), Y_t(2), Y_t(3))$. The covariate $X_t$ is a $10$ dimensional vector generated from the standard normal distribution. For $a\in\{1,2,3\}$, $Y_t(a)=1$ if $a$ is chosen with probability $q(a| x) = \frac{\exp(g(a, x))}{\sum^3_{a'=1}\exp(g(a', x))}$, where $g(1, x) = \sum^{10}_{d=1} X_{t,d}$, $g(2, x) = \sum^{10}_{d=1} W_dX^2_{t,d}$, $g(3, x) = \sum^{10}_{d=1} W_d|X_{t,d}|$, and $W_d$ is uniform randomly chosen from $\{-1, 1\}$. We generate three datasets, $\mathcal{S}^{(1)}_{T_{(1)}}$, $\mathcal{S}^{(2)}_{T_{(2)}}$, and $\mathcal{S}^{(3)}_{T_{(3)}}$, where $\mathcal{S}^{(m)}_{T_{(m)}} = \{(X^{(m)}_t, Y^{(m)}_t(1), Y^{(m)}_t(2), Y^{(m)}_t(3))\}^{T_{(m)}}_{t=1}$. First, we train an evaluation function $\pi^\mathrm{e}$ by solving a prediction problem between $X^{(1)}_t$ and $Y^{(1)}_t(1), Y^{(1)}_t(2), Y^{(1)}_t(3)$ using $\mathcal{S}^{(1)}_{T_{(1)}}$. Then, we regard the evaluation function $\pi^\mathrm{e}$ as a logging policy and apply it on the independent dataset $\mathcal{S}^{(2)}_{T_{(2)}}$ to obtain $\{(X'_t, A'_t, Y'_t)\}^{T_{(2)}}_{t=1}$, where $A'_t$ is an action chosen from the evaluation function and $Y^{(m)}_t = \sum^3_{a=1}\mathbbm{1}[A^{(m)}_t = a]Y^{(m)}_t(a)$. Then, we estimate the true value $R(\pi^\mathrm{e})$ as $\frac{1}{T_{(2)}}\sum^{T_{(2)}}_{t=1}Y^{(m)}_t$. Next, using the datasets $\mathcal{S}^{(3)}_{T_{(3)}}$ and an MAB algorithm, we generate a bandit dataset $\mathcal{S}=\{(X_t, A_t, Y_t)\}^{T_{(3)}}_{t=1}$. For $\mathcal{S}$, we apply the ADR estimator, IPW estimator with the true logging policy (IPW), AIPW estimator (AIPW), AW-AIPW estimator (AW-AIPW) IPW estimator with estimated logging policy (EIPW), and DM estimator (DM). For the AW-AIPW estimator, we use the evaluation weight $\sqrt{\pi_{a,x}(t)}/T$, proposed in \citet{hadad2019}. The other proposed weights of \citet{hadad2019} seem to be specific to situations without covariate $X_t$. For estimating $f^*$ and the logging policy, we use the kernelized Ridge least squares and logistic regression, respectively. We use a Gaussian kernel, and both hyper-parameters of the Ridge and kernel are chosen from $\{0.01, 0.1, 1\}$. We define the estimation error as $R(\pi^\mathrm{e}) - \widehat{R}(\pi^\mathrm{e})$. We conduct ten experiments by changing sample sizes and MAB algorithms. For the sample size $T_{(3)}$, we use $100$, $250$, $500$, $750$, and $10,000$ with the LinUCB and LinTS algorithms. For the sample sizes $T_{(1)}$ and $T_{(2)}$, we use $1,000$ and $100,000$ respectively. We conduct $100$ trials to obtain the root MSEs (RMSEs), the standard deviations of MSEs (SDs), and the coverage ratios (CRs) of the $95\%$ confidence interval (percentage that the confidence interval covers the true value). The results with $T=250$, $500$, $750$ are shown in Table~\ref{tbl:exp_table1} and those with $T=750$ are shown in Figure~\ref{fig:bias}. The other results are shown in Appendix~\ref{appdx:det_exp}. In all cases, the ADR estimator performs well compared to other methods. Although the AIPW estimator has asymptotic normality, the performance is not comparable with the ADR estimator. The reason why the AW-AIPW estimator underperforms is that the performance depends on the choice of evaluation weights, and the weight proposed by \citet{hadad2019} is not suitable for our case with covariates. The EIPW estimator suffers from the dependency problem. We postulate that this is owing to the instability of $\pi_t$.
\begin{remark}[Paradox for i.i.d. samples]
This paradox is similar to the well-known property that the IPW estimator using an estimated propensity score shows a smaller asymptotic variance than the IPW using the true one \citep{hirano2003,Henmi2004paradox,Henmi2007imp}. However, as discussed above, we consider this to be a different phenomenon. In these studies, the paradox is mainly explained by differences in the asymptotic variance between IPW estimators with the true and estimated propensity score. On the other hand, for our case, AIPW and ADR estimators have the same asymptotic variance, unlike IPW-type estimators. Therefore, we cannot resolve the paradox by the aforementioned explanations. The instability of the AIPW estimator is consistent with the findings of \citet{hadad2019}, but, in our paper, we suggest using the ADR estimator as another solution.
\end{remark}
\section{Experiments}
\label{sec:exp}
Following \citet{dudik2011doubly}, we evaluate the estimators using classification datasets by transforming them into contextual bandit data. From the LIBSVM repository \citep{CC01a}, we use the {\tt mnist}, {\tt satimage}, {\tt sensorless}, and {\tt connect-4} datasets \footnote{\url{https://www.csie.ntu.edu.tw/~cjlin/libsvmtools/datasets/.}}. The dataset description is in Table~\ref{Dataset} of Appendix~\ref{appdx:det_exp}. We construct a logging policy as $\pi_t(a| x) = \alpha\pi^m_t(a| x)+(1-\alpha)\frac{0.1}{K}$, where $\alpha\in (0,1)$ is a constant, and $\pi^m_t(a| x)$ is a policy such that $\pi^m_t(a| x) = 1$ for an action $a\in\mathcal{A}$ and $\pi^m_t(a'| x) = 0$ for the other actions ($a'\neq a$). The policy $\pi^m_t(a| x)$ is determined by logistic regression and MAB algorithms. For MAB algorithms, we use upper confidence bound and Thompson sampling with a linear model, which are denoted as LinUCB \citep{Wei2011} and LinTS \citep{Agrawal2013}. The evaluation function is fixed at $\pi^\mathrm{e}(a| x) = 0.9\pi^d(a| x)+\frac{0.1}{K}$, where $\pi_d$ is a prediction of a logistic regression.
We focus on the ADR estimator and compare it to the IPW, EIPW, AIPW, and DM estimators, as done in Section~\ref{sec:num_exp}. We omit the AW-AIPW estimator, as we found that the performance depends on the choice of the evaluation weight (see Section~\ref{sec:paradox}). Additionally, the paper does not consider or discuss situations where there exist covariates. We compare those estimators using the benchmark datasets generated through the process explained above. For $\alpha\in\{0.7, 0.4, 0.1\}$ and the sample sizes $800$, $1,000$, and $1,200$, we calculate the RMSEs and the SDs over $10$ trials. The results of {\tt mnist} and {\tt satimage} with the LinUCB policy and $T=1,000$ are shown in Tables~\ref{tbl:exp_man_table1}. The full results, including the results with the LinTS policy and i.i.d. samples, are shown in Appendix~\ref{appdx:det_exp}. As discussed in Section~\ref{sec:paradox}, the ADR estimator performs well. Although the asymptotic distributions of AIPW and ADR estimators are the same, the AIPW estimator shows poorer performance. As Section~\ref{sec:num_exp}, we consider this to be because the estimator of $\pi_t$ is absorbing the instability of $\pi_t$.
\begin{table*}[t]
\vspace{-0.2cm}
\caption{The results of benchmark datasets with the LinUCB policy. We highlight in red bold the estimator with the lowest RMSE and highlight in under line the estimator with the lowest RMSE among estimators that do not use the true logging policy. Estimators with asymptotic normality are marked with $\dagger$, and estimators that do not require the true logging policy are marked with $*$.}
\label{tbl:exp_man_table1}
\begin{center}
\scalebox{0.70}[0.70]{
\begin{tabular}{|l||rr||rr|rr||rr|rr|}
\hline
{\tt mnist}$\ \ \ \ \ \ \ \ $ & \multicolumn{2}{|c||}{ADR\ $\dagger*$} & \multicolumn{2}{c|}{IPW $\dagger$} & \multicolumn{2}{c||}{AIPW $\dagger$} & \multicolumn{2}{c|}{DM $*$} & \multicolumn{2}{c|}{EIPW $*$} \\
\hline
$\alpha$ & RMSE & SD & RMSE & SD & RMSE & SD & RMSE & SD & RMSE & SD \\
\hline
0.7 & \underline{\textcolor{red}{\textbf{0.046}}} & 0.002 & 0.100 & 0.011 & 0.162 & 0.027 & 0.232 & 0.013 & 0.148 & 0.014 \\
0.4 & \underline{\textcolor{red}{\textbf{0.028}}} & 0.001 & 0.068 & 0.005 & 0.112 & 0.009 & 0.249 & 0.010 & 0.080 & 0.004 \\
0.1 & \underline{0.086} & 0.006 & \textcolor{red}{\textbf{0.078}} & 0.006 & 0.085 & 0.008 & 0.299 & 0.024 & 0.091 & 0.008 \\
\hline
\end{tabular}
}
\end{center}
\vspace{-0.3cm}
\begin{center}
\scalebox{0.70}[0.70]{
\begin{tabular}{|l||rr||rr|rr||rr|rr|}
\hline
{\tt satimage}$\ \ $ & \multicolumn{2}{|c||}{ADR\ $\dagger*$} & \multicolumn{2}{c|}{IPW $\dagger$} & \multicolumn{2}{c||}{AIPW $\dagger$} & \multicolumn{2}{c|}{DM $*$} & \multicolumn{2}{c|}{EIPW $*$} \\
\hline
$\alpha$ & RMSE & SD & RMSE & SD & RMSE & SD & RMSE & SD & RMSE & SD \\
\hline
0.7 & \underline{\textcolor{red}{\textbf{0.013}}} & 0.000 & 0.098 & 0.011 & 0.060 & 0.004 & 0.037 & 0.001 & 0.056 & 0.002 \\
0.4 & \underline{\textcolor{red}{\textbf{0.022}}} & 0.000 & 0.078 & 0.008 & 0.019 & 0.000 & 0.043 & 0.001 & 0.060 & 0.002 \\
0.1 & \underline{\textcolor{red}{\textbf{0.029}}} & 0.001 & 0.078 & 0.005 & 0.061 & 0.008 & 0.041 & 0.002 & 0.041 & 0.002 \\
\hline
\end{tabular}
}
\end{center}
\vspace{-0.65cm}
\end{table*}
Note that the ADR, IPW, and AIPW estimators are asymptotically normal, but the IPW and APIW estimators are not feasible when the true logging policy is not given. To the best of our knowledge, asymptotic normality of the EIPW and DM estimators has not been shown when samples are dependent, and the proof is non-trivial owing to the dependency and the Donsker condition.
\section{Conclusion}
DR-type estimators are crucial in causal inference because they do not assume {\it a priori} knowledge of the true logging policy and they asymptotically follow a normal distribution under standard convergence rate conditions of nuisance estimators. However, existing studies have rarely discussed DR-type estimators when samples are dependent. We derived the ADR estimator and proposed adaptive-fitting as a variant of DML to obtain asymptotic normality. In experiments, we found a paradox that the ADR estimator tends to be more stable than the AIPW estimator and conjectured that this is because the ADR estimator absorbs the instability of the true logging policy $\pi_t$.
In OPVE with dependent samples, we need to put some assumptions on the behavior of $\pi_t$ to obtain asymptotic normality (Assumption~\ref{asm:stationarity}). These assumptions can be broken if the time-series is very complicated. However, we note that this is a limitation of the entire field. The asymptotic normality and double robustness of OPVE estimators are necessary theoretical properties to avoid deriving false causality. Since causal inference is often used in applications related closely to public policy, we consider understanding these limitations critical.
\bibliographystyle{icml2021}
\bibliography{arXiv.bbl}
\clearpage
\onecolumn