EconBase
← Back to paper

Policy Choice in Time Series by Empirical Welfare Maximization

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

125,168 characters

Policy Choice in Time Series by Empirical Welfare Maximization



\maketitle

\begin{abstract}
This paper develops a novel method for policy choice in a dynamic setting where the available data is a multivariate time series. Overcoming challenges unique to time-series setting such as time-varying environments, history-dependent welfare, dynamic causal effects, and statistical dependence, we propose Time-series Empirical Welfare Maximization (T-EWM) methods. We characterize conditions for T-EWM to consistently learn optimal policies conditional or unconditinal on the time-series history, and derive
nonasymptotic upper bounds for the  welfare regrets. We illustrate a use of T-EWM for optimal restriction rules against Covid-19.
\end{abstract}
\textit{Keywords}:  Causal inference, potential outcome time series, treatment choice, regret bounds, concentration inequalities.
\thispagestyle{empty}
\pagebreak

\section{Introduction}
\setcounter{page}{1}
\onehalfspacing

A central topic in economics is the nature of the causal relationships between economic outcomes and government policies, both within and across time periods. To investigate this, empirical research makes use of time-series data, with the aim of finding  desirable policy rules.
For instance, a monetary policy authority may wish to use past and current macroeconomic data  to learn an interest rate policy that is optimal in terms of a social welfare criterion.
Building on the recent development of potential outcome time series (\cite{white2010granger}, \cite{angrist2018semiparametric}, \cite{bojinov2019time}, and \cite{rambachan2021when}), this paper proposes a novel method to inform policy choice when the available data is a multivariate time series.

In contrast to the structural and semi-structural approaches that are common in macroeconomic policy analysis, such as dynamic stochastic general equilibrium (DSGE) models and structural vector autoregressions (SVAR), we set up the policy choice problem from the perspective of the statistical treatment choice proposed by \cite{manski2004statistical}. The existing statistical treatment choice literature typically focuses on microeconomic applications in a static setting, and the applicability of these methods to a time-series setting has yet to be explored. In this paper, we propose a novel statistical treatment choice framework for time-series data and study how to learn an optimal policy rule. Specifically, we consider extending the conditional empirical success (CES) rule of \cite{manski2004statistical} and the empirical welfare maximization rule of \cite{kitagawa2018should} to time-series policy choice, and characterize the conditions under which these approaches can inform welfare-optimal policy. These conditions do not require functional form specifications for structural equations or the exact temporal dependence of the time-series observations, but can be
connected to the structural approach under certain conditions.

In the standard microeconometric setting considered in the treatment choice literature, the planner has access to a random sample of cross-sectional units, and it is often assumed that the populations from which the sample was drawn and to which the policy will be applied are the same. These assumptions are not feasible or credible in the time-series context, which leads to  several non-trivial challenges. First, the economic environment and the economy's causal response to it may be time-varying. Assumptions are required to make it possible to learn an optimal policy rule for future periods based on available past data.
In addition,
the outcomes and policies observed in the available data can be statistically and causally dependent in a complex manner,
and accordingly, the identifiability of social welfare under counterfactual policies becomes non-trivial and requires some conditions on how past policies were assigned over time.
Second, in defining an optimal policy in a time-series context, it is natural to consider both unconditional welfare and conditional  welfare given the history of observables available at the time the policy is chosen. The conditional perspective differs from the unconditional one in that it evaluates welfare given the realized history, whereas the unconditional criterion averages these conditional measures across all possible histories. These welfare criteria are unique in the time-series setting, contrasting with the standard settings with cross-sectional microdata. Whether the unconditional or the conditional welfare criterion is more relevant depends on applications and the data environment. Third, when past data are used to inform policy, we have only a single realization of a time series in which the observations are dependent across the periods and possibly nonstationary. Such statistical dependence complicates the characterization of the statistical convergence of the welfare performance of an estimated policy. Fourth, if the planner wishes to learn a dynamic assignment policy, which prescribes a policy in each period over multiple periods on the basis of observable information available at the beginning of every period, the policy learning problem becomes substantially more involved. This is because a policy choice in the current period may affect subsequent policy choices through the current policy assignment and a realized outcome under the assigned treatment.

Taking into account these challenges, we propose time-series empirical welfare maximization (T-EWM) methods that construct an empirical welfare criterion based on a historical average of the outcomes and obtain a policy rule by maximizing the empirical welfare criterion over a class of policy rules. We then clarify the conditions on the causal structure and data-generating process under which T-EWM methods consistently estimate a policy rule that is optimal in terms of unconditional or conditional welfare. Extending the regret bound analysis of \cite{manski2004statistical} and \cite{kitagawa2018should} to time-series dependent observations, we obtain a finite-sample uniform bound for welfare regret. Moreover, we characterize the convergence of welfare regret and establish the minimax rate optimality of the T-EWM rule.

Our development of T-EWM builds on the recent potential outcome time-series literature including \cite{white2010granger}, \cite{angrist2018semiparametric}, \cite{bojinov2019time}, and \cite{rambachan2021when}. In particular, to identify the counterfactual welfare criterion, we employ the sequential exogeneity restriction considered in \cite{bojinov2019time}. These studies focus on retrospective evaluation of the causal impact of policies observed in historical data, and do not analyze how to perform future policy choices based on the historical evidence.

Since the seminal works of \cite{manski2004statistical} and \cite{Dehejia2005}, statistical treatment choice and empirical welfare maximization have been active topics of research, e.g., \cite{stoye2009minimax, stoye2012minimax}, \cite{QianMurphy2011}, \cite{tetenov2012statistical}, \cite{BhattacharyaDupas2012}, \cite{Zhao2012JASA}, \cite{kitagawa2018should, kitagawa2021equality}, \cite{kallus2021more}, \cite{athey2021policy}, \cite{MT17}, \cite{KST21}, among others. These works focus on a setting where the available data is a cross-sectional random sample obtained from an experimental or observational study with randomized treatment, possibly conditional on observable characteristics. \cite{Viviano21} and \cite{Abhishek2020OptimalTreatment} consider EWM approaches for treatment allocations where the training data features cross-sectional dependence due to network spillovers, while to our knowledge, this paper is the first to consider policy choice with time-series data. As a related but distinct problem, there is a large literature on the estimation of dynamic treatment regimes, \cite{murphy2003optimal}, \cite{zhao2015new}, \cite{han2021optimal},  \cite{Sakaguchi21}, and \cite{ida2024dynamic}, among others. The problem of dynamic treatment regimes assumes that training data is a short panel in which treatments have been randomized both among cross-sectional units and across time periods. Recently, \cite{adusumilli2019dynamically} consider an optimal policy in a dynamic treatment assignment problem with a budget constraint where the planner allocates treatments to subjects arriving sequentially.
The T-EWM framework, in contrast, assumes observations are drawn from a single time series as is common in empirical macroeconomics.

A large literature on multi-arm bandit algorithms analyzes learning and dynamic allocations of treatments when there is a trade-off between exploration and exploitation. See \cite{lattimore2020bandit} and references therein,  and \cite{Dimakopoulou_et_al_2017}, \cite{Kock2020functional}, \cite{kasy2019adaptive}, \cite{adusumilli2021risk}, and \cite{Kitagawa_Rowley_2024} for recent works in econometrics. The setting in this paper differs from the standard multi-arm bandit setting in the following three respects. First, our framework treats the available past data as a training sample and focuses on optimizing short-run welfare. We are hence concerned with the performance of the method in terms of short-term regret rather than cumulative regret over a long horizon.
Second, in the standard multi-arm bandit problem, subjects to be treated are assumed to differ across rounds, which implies that the outcome-generating process is independent over time. This is not the case in our setting, and we include the realization of outcomes and policies in the past periods as contextual information for the current decision. {Third, suppose that bandit algorithms can be adjusted to take into account the dependence of observations, our method is then analogous to the ``pure exploration'' class,
involving a long exploration phase followed by a one-period exploitation at the very end. However, a major difference is that the bandit algorithm concerns data in a random experiment while our method is aimed at data in quasi-random experiments. }


The analysis of welfare regret bounds is similar to the derivation of risk bounds in empirical risk minimization, as reviewed by \cite{Vapnik1998Book} and \cite{Lugosi2002}. Risk bounds studied in the empirical risk minimization literature typically assume independent and identically distributed (i.i.d.) training data. A few exceptions, \cite{jiang2010risk}, \cite{brownlees2021performance}, and \cite{brownlees2021empirical}, obtain risk bounds for empirical risk minimizing predictions with time-series data, but they do not consider welfare regret bounds for causal policy learning.

The rest of the paper is organized as follows. Section \ref{model_example}  describes the setting using a simple illustrative model with a single discrete covariate. Section \ref{continuous} discusses the general model with continuous covariates and presents the main theorems. In Section \ref{ext}, we discuss extensions to our proposed framework, including  a case of multi-period welfare functions, how T-EWM is related to the Lucas critique, as well as T-EWM's links with structural vector autoregressive models, Markov Decision Processes, and reinforcement learning. In Section \ref{applica}, we present  an empirical application. Technical proofs, simulation studies, and other details are presented in Appendices.

\section{Model and illustrative example} \label{model_example}

In this section, we introduce the basic setting, notation, the welfare criterion we aim to maximize, and conditions on the data-generating process that are important for the learnability of an optimal policy. Then, we illustrate the main analytical tools used to bound welfare regret through a heuristic model with a simple dynamic structure.


\subsection{Notation, timing, and welfare} \label{sec:notation_timing_welfare}
We suppose that the social planner is at the beginning of time $T$. Let $W_t \in \{0, 1 \}$ denote a treatment or policy (e.g., nominal interest rate) implemented at time $t = 0,1,2,\dots$ To simplify the analysis, we assume that $W_t$ is binary (e.g., a high or low interest rate). The planner sets $W_T \in \{0, 1\}$, $T \geq 1$, making use of the history of observable information up to period $T-1$ to inform her decision. This observable information consists of an economic outcome (e.g., GDP, unemployment rate, etc.), $Y_{0:T-1}=(Y_0,Y_1,Y_2, \cdots, Y_{T-1})$, the history of implemented policies, $W_{0:T-1}=(W_0,W_1,W_2, \cdots, W_{T-1})$, contextual covariates (e.g., inflation), $Z_{0:T-1}=(Z_0, Z_1,Z_2, \cdots, Z_{T-1})$, and potential mediators between treatment and outcome, $M_{0:T-1}=(M_0, M_1,M_2, \cdots, M_{T-1})$. $Z_t$ and $M_t$ can be  multidimensional vectors, but $Y_t$ and $W_t$ are assumed to be univariate.


Following \cite{bojinov2019time}, we refer to a sequence of policies $w_{0:t} = (w_0, w_1,$ $ \dots, w_t) \in \{0,1\}^{t+1}$, $t \geq 0$, as a treatment path. A realized treatment path observed in the data $0 \leq t \leq T-1$ is a stochastic process $W_{0:T-1}$ drawn from the data-generating process.
The timing of realizations is therefore
\begin{figure}[H]
	\centering
	\includegraphics[scale=0.5]{pictures/time-label.png}
\end{figure}
  \noindent i.e., the transition between periods happens between the $Z_{t-1}$ is realised and before $W_t$ is realised.
Furthermore, let
\begin{equation} \label{covariate}
X_{t}=\{W_{t},M_{t}^{\prime}, Y_{t}, Z_{t}^{\prime}\}^{\prime}
\end{equation}
collect the observable variables for period $t$. For $t = 0,1,2,\dots$,. let $\mathcal{X}$ denote the sample space of $X_t$. Furthermore,  we define the filtration
$$\mathcal{F}_{t-1}=\sigma (X_{0:t-1}),$$
where $\sigma(\cdot)$ denotes the Borel $\sigma$-algebra generated by the variables specified in the argument. The filtration $\mathcal{F}_{T-1}$ corresponds to the planner's information set at the time of making her decision in period $T$.

Following the framework of \cite{bojinov2019time}, we introduce potential outcome time series. At each $t = 0, 1, 2, \dots,$ and for every treatment path $w_{0:t} \in \{0 ,1 \}^{t+1}$, let $Y_t(w_{0:t}) \in \mathbb{R}$ be the realized period $t$ outcome if the treatment path from $0$ to period $t$ were $w_{0:t}$. Hence, we have a collection of potential outcome paths indexed by treatment path,
$$
\{ Y_t(w_{0:t}) : w_{0:t} \in \{0, 1 \}^{t+1}, t=0,1,2,\dots \},
$$
which defines $2^{t+1}$ potential outcomes in each period $t$. This is an extension of the Neyman-Rubin causal model originally developed for cross-sectional causal inference. As maintained in \cite{bojinov2019time}, the potential outcomes for each $t$ are indexed by the current and past treatments $w_{0:t}$ only. This imposes the restriction that any future treatment $w_{t+p}$, $p \geq 1$, does not causally affect the current outcome, i.e., an exclusion restriction for future treatments.

For a realized treatment path $W_{0:t}$, the observed outcome $Y_t$ and the potential outcomes satisfy
$$
Y_t = \sum_{w_{0:t} \in \{ 0, 1\}^{t} } 1 \{W_{0:t} = w_{0:t} \} Y_t(w_{0:t})
$$
for all $t \geq  0$.

The baseline setting of the current paper considers the choice of policy $W_T$ for a single period $T$.\footnote{Section \ref{sec_multi} discusses how to extend the single-period policy choice problem to multi-period settings.} We denote the policy choice based on observations up to period $T-1$ by
\begin{equation} \label{eq:policy_f}
g: \mathcal{X}^{T} \to \{0,1\},
\end{equation}
where $\mathcal{X}^T$ denotes the sample space of $X_{0:T-1}$. The period-$T$ treatment is $W_T = g(X_{0:T-1})$, and we refer to $g(\cdot)$ as a \emph{decision rule}.

We assume that the planner's preferences for policies in period-$T$ are embodied in a social welfare criterion. In particular, we define one-period \emph{welfare conditional on $\mathcal{F}_{T-1}$} (conditional welfare, for short) to be\footnote{Throughout the paper, we acknowledge that the expectation, $\mathop{\mbox{\sf E}}$, and the probability, $\Pr$, are corresponding to the outer measure whenever a measurability issue is encountered.}
\begin{eqnarray} \label{wel:condi}
\mathcal{W}_T(g|\mathcal{F}_{T-1})
= \mathop{\mbox{\sf E}} \left[ Y_{T}(W_{0:T-1},1) g(X_{0:T-1})+ Y_{T}(W_{0:T-1},0)(1- g(X_{0:T-1}))|\mathcal{F}_{T-1} \right].
\end{eqnarray}


As an alternative welfare criterion, we can define the \emph{unconditional welfare} by averaging out the conditioning information in \eqref{wel:condi},
\begin{eqnarray} \label{eq:uncon_g}
    \mathcal{W}_{T}(g) = \mathop{\mbox{\sf E}} \left[ Y_{T}(W_{0:T-1},1) g(X_{0:T-1})+ Y_{T}(W_{0:T-1},0)(1- g(X_{0:T-1}))\right].
\end{eqnarray}
Whether the planner's objective function should be conditional or unconditional welfare depends on applications and the goal of analysis. If the planner's goal is to perform policy choice in a potentially nonstationary environment given the realized path of scenarios, it is natural to set the conditional welfare as the criterion to maximize. This case contrasts with the setting of the cross-sectional treatment choice with micro data where the unconditional welfare that averages out the observable characteristics of a unit is common.\footnote{\cite{manski2004statistical} also considers a conditional welfare criterion in the cross-sectional setting.}

 On the other hand, there are cases in which studying optimal policy from the perspective of unconditional welfare is suitable. First, an optimal policy in terms of the unconditional welfare can inform an optimal policy action that the planner would choose under a counterfactual history. Such analysis could be of interest in the practice of \textit{ex post} counterfactual policy analysis.
 Second, in settings where multiple units exist, learning an unconditionally optimal policy from the time-series data of a particular unit informs the planner of optimal policies for other similar units but with different histories.
 Third, an unconditionally optimal policy can also inform an optimal policy in the setting where the planner implements the same Markovian policy over multiple time periods, assuming the data-generating process is stationary.

Note that the T-EWM based on unconditional welfare yields the same empirical average as the EWM in cross-sectional data. The two, however, differ markedly in substance. In time-series settings, observed outcomes and policies are subject to statistical and causal dependence across periods, whereas cross-sectional data are typically assumed to be i.i.d. across the units. Due to this difference, external-validity type assumptions become more debatable in the time-series setting than in the cross-sectional setting.

 Learning conditionally optimal policies is generally more challenging than learning unconditionally optimal ones since estimation of the unconditional welfare can use all the observations in the sample, whereas estimation of the conditional welfare may have to rely on a subset of observations.
In Section \ref{sec:uncon_con}, however, we show that under certain assumptions, the unconditional welfare regret can bound the conditional welfare regret so that one can learn the conditionally optimal policy as fast as the unconditionally optimal one.






For the welfare targets \eqref{wel:condi} and \eqref{eq:uncon_g}, define the planner's optimal  conditional and unconditional policies that maximize her one-period welfare as
\begin{eqnarray*}
g^{\ast}_{X_{0:T-1}} &\in & \arg \max_g \mathcal{W}_T(g| \mathcal{F}_{T-1}),\label{p:opti_p}\\
g^{\ast} &\in & \arg \max_g \mathcal{W}_T(g),\nonumber
\end{eqnarray*}
respectively. The former, $g^{\ast}_{X_{0:T-1}}$, can be regarded as the optimal policy $g^{\ast}$ in the latter evaluated at the conditioning value of $X_{0:T-1}$. The planner does not know $g^{\ast}$, so she instead seeks a statistical treatment choice rule (\citeauthor{manski2004statistical}, \citeyear{manski2004statistical}) $\hat{g}$, which is a decision rule selected on the basis of the available data $X_{0:T-1}$.

Our goal is to develop a way of obtaining $\hat{g}$ that performs well in terms of the conditional welfare criteria (\ref{wel:condi}) and \eqref{eq:uncon_g}. Specifically, we assess the statistical performance of an estimated policy rule $\hat{g}$ in terms of the convergence of conditional and unconditional welfare regrets,
\begin{eqnarray*}
\mathcal{R}(\hat{g}|x_{0:T-1})&=&\mathcal{W}_T(g^{\ast}|X_{0:T-1} = x_{0:T-1}) - \mathcal{W}_T(\hat{g}|X_{0:T-1} = x_{0:T-1}), \\
\mathcal{R}(\hat{g})&=&\mathcal{W}_T(g^{\ast}) -\mathcal{W}_T(\hat{g}),\label{cond_regret_general}
\end{eqnarray*}
as well as their convergence rate with respect to the sample size $T$. When examining convergence, we accommodate statistical uncertainty over $\hat{g}$ by focusing on convergence with probability approaching one uniformly over a class of sampling distributions for $X_{0:T-1}$.

\subsection{An illustrative model with a discrete covariate} \label{one-p}
We begin our analysis with a simple illustrative  model, which provides a heuristic exposition of the main idea of T-EWM and its statistical properties. We cover more general settings and extensions in Sections \ref{continuous} and \ref{ext}.

Suppose that the data consist of a bivariate time series $X_{0:T-1} = ((Y_t,W_t) \in \mathbb{R} \times \{0, 1 \} : t=0,1, \dots, T-1)$ with no other covariates. To simplify exposition for the illustrative model, we impose the following restrictions on the dynamic causal structure and dependence of the observations.

\begin{assumption}  [Markov properties] \label{ass:toy_example}The time series of potential outcomes and observable variables satisfy the following conditions:

(i) \emph{Markovian exclusion}: for $t=2, \dots, T$ and for arbitrary treatment paths $(w_{0:t-2},w_{t-1},w_{t})$ and $(w_{0:t-2}^{\prime},w_{t-1},w_{t})$, where $ w_{0:t-2} \neq w_{0:t-2}^{\prime}$,
\begin{eqnarray} \label{exclu}
Y_t(w_{0:t-2}, w_{t-1},w_{t} ) = Y_t(w_{0:t-2}^{\prime}, w_{t-1},w_{t}) := Y_t(w_{t-1},w_t)
\end{eqnarray}
holds with probability one.

(ii) \emph{Markovian exogeneity}: for $t=1, \dots, T$ and any treatment path $w_{0:t}$,
\begin{eqnarray}
Y_t(w_{0:t}) \perp X_{0:t-1} | W_{t-1},
\end{eqnarray}
and for $t=1, \dots, T-1$,
\begin{eqnarray}
W_t \perp X_{0:t-1} | W_{t-1}.
\end{eqnarray}
\end{assumption}

These assumptions significantly simplify the dynamic structure of the problem. Markovian exclusion, Assumption \ref{ass:toy_example}(i), says that only the current treatment $W_{t}$ and treatment in the previous period $W_{t-1}$ can have a causal impact on the current outcome. This allows the indices of the potential outcomes to be compressed to the latest two treatments $(w_{t-1},w_t)$, as in (\ref{exclu}). Markovian exogeneity, Assumption \ref{ass:toy_example}(ii), states that once one conditions on the policy implemented in the previous period $W_{t-1}$, the potential outcomes and treatment for the current period are statistically independent of any other past variables.

It is important to note that these assumptions do not impose stationarity: we allow the distribution of potential outcomes to vary across time periods. In addition, under Assumption \ref{ass:toy_example},
we can reduce the class of policy rules to those that map from $W_{T-1} \in \{0,1 \} $ to $W_T \in \{0,1 \}$, i.e., $$g : \{0 , 1\} \to \{0, 1\}.$$
Consequently, welfare conditional on $\mathcal{F}_{T-1}$ can be simplified to
\begin{eqnarray}
\mathcal{W}_T(g|W_{T-1})
&=& \mathop{\mbox{\sf E}}  [ Y_{T}(W_{T-1},1) g(X_{0:T-1})+ Y_{T}(W_{T-1},0)(1- g(X_{0:T-1}))|X_{0:T-1} ]. \nonumber \\
&=& \mathop{\mbox{\sf E}} [ Y_{T}(W_{T-1},1) g(W_{T-1})+ Y_{T}(W_{T-1},0)(1- g(W_{T-1}))|W_{T-1} ],\label{obj_simple}
\end{eqnarray}
where the first equality follows from the definition of $\mathcal{F}_{T-1}$ and Assumption \ref{ass:toy_example}(i); the second equality follows from Assumption \ref{ass:toy_example}(ii).

To make sense of Assumption \ref{ass:toy_example} and illustrate the relationship between the potential outcome time series and the standard structural equation modeling, we provide a toy example.


\begin{exmp} \label{example_q1_Markov}
Suppose the planner (monetary policy authority) is interested in setting a low or high interest rate at period $T$. Let $W_t$ denote the indicator for whether the interest rate in period $t$ is high ($W_t = 1$) or low ($W_t = 0$). $Y_t$ denotes a measure of social welfare, which can be a function of aggregate output, inflation, and other macroeconomic variables.
Let $\varepsilon_t$ be an i.i.d.\ shock that is statistically independent of $X_{0:t-1}$, and we assume the following structural equation for the causal relationship of $Y_t$ on $W_t$ (and its lag) and the regression dependence of $W_t$ on its lag.\footnote{The distribution of $W_{t}$ follows \cite{hamilton1989new}. However, the Markov switching model of \cite{hamilton1989new} has unobserved $W_t$, which differs from this example. }
\begin{align}
Y_{t} & =\beta_0 + \beta_{1}W_{t}+\beta_{2}W_{t-1}+\varepsilon_{t},\label{eq:Yt}\\
W_{t} & =(1-q)+\lambda W_{t-1}+V_{t},\quad \lambda  =p+q-1,\label{eq:Wt}\\
 \varepsilon_t & \perp (W_t, X_{0:t-1}) \qquad \forall 1 \leq t \leq T-1  \qquad \text{and} \qquad  \varepsilon_T  \perp X_{0:T-1}.& \label{eq:et}
 \end{align}
 \begin{align}
\text{If }W_{t-1} & =1,\text{ }\begin{cases}
V_{t}=1-p & \text{with probability }p\\
V_{t}=-p & \text{with probability }1-p,
\end{cases} \label{eq:Vt_1} \\
\text{if }W_{t-1} & =0,\text{ }\begin{cases}
V_{t}=-(1-q) & \text{with probability }q\\
V_{t}=q & \text{with probability }1-q \label{eq:Vt_2}.
\end{cases}
\end{align}
\end{exmp}
Compatibility with Assumption \ref{ass:toy_example} can be seen as follows. Assumption \ref{ass:toy_example}(i) is implied by \eqref{eq:Yt}, where the structural equation of $Y_t$ involves only $(W_t,W_{t-1})$ as the factors of direct cause. Assumption \ref{ass:toy_example}(ii) is implied by \eqref{eq:Yt}, \eqref{eq:et}, and the fact that the distribution of $V_{t}$ depends solely on $W_{t-1}$, i.e., under \eqref{eq:Wt},\eqref{eq:Vt_1}, and \eqref{eq:Vt_2}, we have $\Pr(W_t|\mathcal{F}_{t-1})=\Pr(W_t|W_{t-1}).$\\


To examine the learnability of the optimal policy rule, we further restrict the data-generating process. First, we impose a strict overlap condition on the propensity score.
\begin{assumption}[Strict overlap]  \label{bound1} Let $e_t(w) := \Pr(W_t = 1|W_{t-1} = w)$ be the period-$t$ propensity score. There exists a constant $\kappa \in (0,1/2)$, such that for any $t = 1, 2, \dots, T-1$ and $w\in \{0,1\}$,
        \begin{equation*}
                \kappa \leq e_t(w) \leq 1-\kappa.
        \end{equation*}
\end{assumption}
The next assumption imposes an unconfoundedness condition on observed policy assignment.
\begin{assumption} [Sequential unconfoundedness]\label{unconf}
        For any $t = 1,2,\dots, T-1$ and $w\in \{0,1\}$,
        \[
        Y_{t}(W_{t-1},w)\perp W_t|X_{0:t-1}.
        \]
\end{assumption}
This assumption states that the treatments observed in the data are sequentially randomized conditional on lagged observable variables. This is a key assumption to make unbiased estimation of the welfare feasible in each period in the sample, as employed in \cite{bojinov2019time} and others. It is worth noting that the above assumption together with Assumption \ref{ass:toy_example}(ii) implies $Y_{t}(W_{t-1},w)\perp W_t|W_{t-1}.$

Under Assumption \ref{ass:toy_example}, we have that, for any measurable function $f$ of the potential outcome $Y_{t}(W_{0:t-1})$ and treatment $W_t$, it holds
\begin{eqnarray} \label{simple_cond_exp}
\mathop{\mbox{\sf E}} (f(Y_{t}(W_{0:t}),W_t)|\mathcal{F}_{t-1})=\mathop{\mbox{\sf E}} (f(Y_{t}(W_{t-1},W_t),W_t)|W_{t-1}) = \mathop{\mbox{\sf E}} (f(Y_{t},W_t)|W_{t-1}).
\end{eqnarray}

\noindent {\bf Example \ref{example_q1_Markov} continued.} Assumption \ref{bound1} is satisfied if $0<p<1$ and $0<q<1$; Assumption \ref{unconf} is implied by \eqref{eq:Yt} and \eqref{eq:et}.\\

Imposing Assumption \ref{bound1} and \ref{unconf} and assuming propensity scores are known, we consider constructing a sample analogue of (\ref{obj_simple}) conditional on $W_{T-1}=w$ based on the historical average of the inverse propensity score weighted outcomes,
\begin{eqnarray} \label{sample_simple}
\widehat{\mathcal{W}}(g|W_{T-1}=w)
= \frac{1}{T(w)}\sum_{1 \leq t \leq T-1: W_{t-1} = w}\left [\frac{Y_{t}W_{t}g(W_{t-1})}{e_t(W_{t-1})}+ \frac{Y_{t}(1-W_{t})\{1- g(W_{t-1})\}}{1- e_t(W_{t-1})} \right],
\end{eqnarray}
where $T(w) = \#\{1\leq t \leq T-1: W_{t-1} = w\} $ is the number of observations where the policy in the previous period took value $w$, i.e., the subsample corresponding to $W_{t-1} = w$. Unlike the microeconometric setting considered in, e.g., \cite{kitagawa2018should}, we do not necessarily have $\widehat{\mathcal{W}}(g|W_{T-1}=w)$ as a direct sample analogue for the planner's social welfare objective, since we allow a non-stationary environment in which the historical average of the conditional welfare criterion can diverge from the conditional welfare in the current period. Nevertheless, we refer to $\widehat{\mathcal{W}}(g|W_{T-1}=w)$ as the empirical welfare of the policy rule $g$.

Denoting $(\cdot|W_{T-1}=w)$ by $(\cdot|w)$, we define the true optimal conditional policy and its empirical analogue to be,
\begin{eqnarray}
g^*(w) &\in& \mbox{argmax}_{g: \{w\}\to  \{0,1\}} \mathcal{W}_T(g|w),\label{best_g} \\
\hat{g}(w) &\in& \mbox{argmax}_{g: \{w\}\to  \{0,1\}} \widehat{\mathcal{W}}(g|w),\label{best_g_hat}
\end{eqnarray}
where $\hat{g}$ is constructed by maximizing empirical welfare over a class of policy rules (four policy rules in total). We call a policy rule constructed in this way the \textit{Time-series Empirical Welfare Maximization} (T-EWM) rule. The construction of the T-EWM rule $\hat{g}$ is analogous to the conditional empirical success rule with known propensity scores considered by \cite{manski2004statistical} in the i.i.d. cross-sectional setting. In the time-series setting, however, the assumptions imposed so far do not guarantee that $\widehat{\mathcal{W}}(g|w)$ is an unbiased estimator of the true conditional welfare $\mathcal{W}_{T}(g|w)$.

\subsection{Bounding the conditional welfare regret of the T-EWM rule} \label{discrete_bound}
A major contribution of this paper is characterizing conditions that justify the T-EWM rule $\hat{g}$ in terms of the convergence of conditional welfare. This section clarifies these points in the context of our illustrative example.

To bound conditional welfare regret, our strategy is to decompose empirical welfare $\widehat{\mathcal{W}}(g|w)$ into a conditional mean component and a deviation from it. The deviation is the sum of a martingale difference sequence (MDS), and this allows us to apply concentration inequalities for the sum of MDS.  Define an intermediate welfare function,
\begin{eqnarray} \label{E_sum_c}
\bar{\mathcal{W}}(g|w)
=T(w)^{-1}\sum_{1 \leq t \leq T-1 : W_{t-1} = w}\mathop{\mbox{\sf E}}\left[ Y_{t}(W_{t-1},1)g(W_{t-1})+Y_{t}(W_{t-1},0)\left[1-g(W_{t-1})\right]|W_{t-1}\right].
\end{eqnarray}
Under the strict overlap and unconfoundedness assumptions (i.e., Assumptions \ref{bound1} and \ref{unconf}), the difference between empirical welfare and  \eqref{E_sum_c} is a sum of MDS. A concentration inequality for average MDS then implies that the empirical welfare concentrates around $\bar{\mathcal{W}}(g|w)$. Since $\bar{\mathcal{W}}(g|w)$ is not guaranteed to inform an optimal policy in terms of $\mathcal{W}_T(g|w)$, we impose the assumption:

 \begin{assumption}  [Invariance of the welfare ordering]  \label{equiv_W}
      Given $w\in \{0,1\}$, let $g^*=g^*(w)$ as defined in \eqref{best_g}.   There exist a positive constant $c>0$, such that for any  $g\in\{0,1\}$,
                \begin{eqnarray} \label{eq:invariance_toy}
        \mathcal{W}_T(g^*|w)- \mathcal{W}_T(g|w) \leq c\big[\bar{\mathcal{W}}(g^*|w)-\bar{\mathcal{W}}(g|w)\big],
        \end{eqnarray}
        with probability one, i.e., $P_T \left( \text{inequality (\ref{eq:invariance_toy}) holds} \right) = 1$,
        where $P_T$ is the probability distribution for $X_{0:T-1}$.
\end{assumption}

Noting that the left-hand side of (\ref{eq:invariance_toy}) is nonnegative for any $g$ by construction and $c>0$, this assumption implies that $\bar{\mathcal{W}}(g^*|w)-\bar{\mathcal{W}}(g|w) > 0$ must hold whenever $\mathcal{W}_T(g^*|w)- \mathcal{W}_T(g|w) > 0$. That is, the optimality of $g^{\ast}$ in terms of the conditional value of welfare at $T$ is maintained in the historical average of the conditional values of welfare. Under this assumption, having an estimated policy $\hat{g}$ that attains a convergence of $\bar{\mathcal{W}}(g^*|w)-\bar{\mathcal{W}}(\hat{g}|w)$ to zero guarantees that $\hat{g}$ is also consistent for the optimal policy $g^{\ast}$ in terms of the conditional welfare at $T$.

\begin{rem} \label{rem:invar_order}
Assumption \ref{equiv_W} can be restrictive in a situation where the dynamic causal structure of the current period is believed to be different from the past, but is weaker than stationarity. In the current example, Assumption \ref{equiv_W} is implied by the following condition.
        \textit{\begin{description}
                        \item [{A\ref{equiv_W}' }] The stochastic process
        $$S_t(w) \equiv Y_{t}(W_{t-1},1)g(W_{t-1})+Y_{t}(W_{t-1},0)\left[1-g(W_{t-1})\right]|_{W_{t-1}=w}$$ is weakly stationary.
        \end{description}}
   Under A\ref{equiv_W}',  $\mathop{\mbox{\sf E}}\left[ Y_{t}(W_{t-1},1)g(W_{t-1})+Y_{t}(W_{t-1},0)\left[1-g(W_{t-1})\right]|W_{t-1}=w\right] $
    is invariant for $2\leq t\leq T$. Then, Assumption \ref{equiv_W} will hold naturally.

  Furthermore,  Assumption \ref{equiv_W} can be satisfied by many classic non-stationary processes in linear time-series models, including series with deterministic or stochastic time trends.
  \end{rem}

    \noindent{\bf  Example \ref{example_q1_Markov} continued.}\\
    $\varepsilon_{t}$ remains an i.i.d. noise in the following settings.\\
(i) By Remark \ref{rem:invar_order}, Assumption \ref{equiv_W} holds
for Example 1 since
$S_{t}(w)=\beta_{0}+\beta_{1}\cdot g+\beta_{2}\cdot w+\varepsilon_{t}$
is weakly stationary.



\noindent(ii) If we replace \eqref{eq:Yt} by
\[
Y_{t}=\delta_{t}+\beta_{1}W_{t}+\beta_{2}W_{t-1}+\varepsilon_{t},
\]
where $\delta_{t}$ is an arbitrary deterministic time trend. The
process $Y_{t}$ is trend stationary (non-stationary), but Assumption
\ref{equiv_W} still holds with $c=1$ since those
deterministic trends are canceled out by differences, i.e.,
\[
\mathcal{W}_{T}(g^{*}|w)-\mathcal{W}_{T}(g|w)=\beta_{1}\left(g^{*}-g\right)=\bar{\mathcal{W}}(g^{*}|w)-\bar{\mathcal{W}}(g|w).
\]

\noindent(iii) If we replace \eqref{eq:Yt} by
\[
Y_{t}=\beta_0 + \beta_{1}W_{t}+\beta_{2}W_{t-1}+\sum_{i=0}^{t}\varepsilon_{i},
\]
the process $Y_{t}$ is non-stationary with stochastic trends, but
Assumption \ref{equiv_W} still holds with $c=1$
since the stochastic trends are canceled out by differences.

\noindent(iv) If we replace \eqref{eq:Yt} by
\[
Y_{t}=\delta_t + \beta_{1,t}W_{t}+\beta_{2,t}W_{t-1}+\varepsilon_{t}
\]
to allow heterogeneous treatment effect, then
\begin{align*}
\mathcal{W}_{T}(g^{*}|w)-\mathcal{W}_{T}(g|w) & =\beta_{1,T}\left(g^{*}-g\right),\\
\bar{\mathcal{W}}(g^{*}|w)-\bar{\mathcal{W}}(g|w) & =\bar{\beta}_{w}\left(g^{*}-g\right),
\end{align*}
where $\bar{\beta}_{w}:=\frac{1}{T(w)}\sum_{1\leq t\leq T-1:W_{t-1}=w}\beta_{1,t}$.
Since $\mathcal{W}_{T}(g^{*}|w)-\mathcal{W}_{T}(g|w)$ is non-negative
by the definition of $g^{*}$ in  \eqref{best_g}, Assumption \ref{equiv_W}
holds if $\beta_{1,T}$ and $\bar{\beta}_{w}$ have the same
sign and $\frac{|\beta_{1,T}|}{|\bar{\beta}_{w}|}\leq c.$.

Without loss of generality, we can assume that both $\beta_{1,T}$
and $\bar{\beta}_{w}$ are positive, then a sufficient condition for Assumption \ref{equiv_W} is: There
are positive numbers $l$ and $u$, such that $0<l\leq\beta_{1,t}\leq u$
holds for all $t$. In this case, $c=\frac{u}{l}$.



Assumption \ref{equiv_W} implies that $g^*$ also maximizes $\bar{\mathcal{W}}$. Hence, if empirical welfare $\widehat{\mathcal{W}}(\cdot|w)$ can approximate $\bar{\mathcal{W}}(\cdot|w)$ well, intuitively the T-EWM rule $\hat{g}$ should converge to $g^{\ast}$.  The motivation for Assumption \ref{equiv_W} is to create a bridge between $\mathcal{W}_{T}(g^*|w)-\mathcal{W}_{T}(\hat{g}|w)$, the population regret,  and $\widehat{\mathcal{W}}(\cdot|w) - \bar{\mathcal{W}}(\cdot|w) $, which is a sum of MDS  with respect to filtration $\{ \mathcal{F}_{t-1} : t = 1,2, \dots, T \}$.\footnote{Note that by \eqref{simple_cond_exp},  $\mathop{\mbox{\sf E}}(\cdot|\mathcal{F}_{t-1})=\mathop{\mbox{\sf E}}(\cdot|W_{t-1} )$.}
 Specifically,
  \begin{align} \label{ineq_s}
        & \mathcal{W}_{T}(g^*|w)-\mathcal{W}_{T}(\hat{g}|w)\nonumber
        \leq  c\left[\bar{\mathcal{W}}(g^*|w)-\bar{\mathcal{W}}(\hat{g}|w)\right]\nonumber\\
        = & c\left[\bar{\mathcal{W}}(g^*|w)-\widehat{\mathcal{W}}(\hat{g}|w)+\widehat{\mathcal{W}}(\hat{g}|w)-\bar{\mathcal{W}}(\hat{g}|w)\right]\nonumber
        \leq  c\left[\bar{\mathcal{W}}(g^*|w)-\widehat{\mathcal{W}}(g^*|w)+\widehat{\mathcal{W}}(\hat{g}|w)-\bar{\mathcal{W}}(\hat{g}|w)\right]\nonumber\\
        \leq & 2c\sup_{g:\{w\}\to  \{0,1\}}|\bar{\mathcal{W}}(g|w)-\widehat{\mathcal{W}}(g|w)|,
 \end{align}
 where the first inequality follows from Assumption \ref{equiv_W}. The second inequality follows from the definition of T-EWM rule $\hat{g}$ in \eqref{best_g_hat}.

 To bound the right-hand side of \eqref{ineq_s}, define
 \begin{align*}
 \widehat{\mathcal{W}}_t(g|w)=&1(W_{t-1}=w)\left[\frac{Y_{t}W_{t}g(W_{t-1})}{e_t(W_{t-1})}+ \frac{Y_{t}(1-W_{t})\{1- g(W_{t-1})\}}{1- e_t(W_{t-1})}\right],\\
\bar{\mathcal{W}}_t(g|w)=&1(W_{t-1}=w)\mathop{\mbox{\sf E}} \{ Y_{t}(W_{t-1},1) g(W_{t-1})+ Y_{t}(W_{t-1},0)(1- g(W_{t-1}))|W_{t-1}=w\}.
\end{align*}
 Then, we can express \eqref{sample_simple} and \eqref{E_sum_c} as
 \begin{eqnarray*}
 \widehat{\mathcal{W}}(g|w)
 =&\frac{T-1}{T(w)}\cdot \frac{1}{T-1}\sum_{t  = 1}^{T-1}\widehat{\mathcal{W}}_t(g|w),\\
\bar{\mathcal{W}}(g|w)
 =&\frac{T-1}{T(w)}\cdot \frac{1}{T-1}\sum_{t  = 1}^{T-1}\bar{\mathcal{W}}_t(g|w),
 \end{eqnarray*}
 and
 \begin{eqnarray}
 \widehat{\mathcal{W}}(g|w) - \bar{\mathcal{W}}(g|w) = \frac{T-1}{T(w)} \cdot \frac{1}{T-1} \sum_{t=1}^{T-1} \left[ \widehat{\mathcal{W}}_t(g|w) - \bar{\mathcal{W}}_t(g|w) \right] \notag
 \end{eqnarray}
 follows. Next, we impose
 \begin{assumption} [Bounded outcomes]\label{s_bounded_y}
		There exists $M<\infty$ such that the support of outcome variable
	$Y_t$ is contained in $[-M/2,M/2]$.
\end{assumption}
 \begin{prop} \label{prop:simple_mds}
 	Under Assumptions \ref{ass:toy_example}-\ref{unconf} and \ref{s_bounded_y}, the sequence $\{\widehat{\mathcal{W}}_t(g|w)-\bar{\mathcal{W}}_t(g|w)\}_{t=1}^{T-1}$ is an MDS.
 	\end{prop}
  The proof of this proposition can be found in Appendix \ref{app:proof_simple_mds}. With Proposition \ref{prop:simple_mds}, we can apply a concentration inequality for sums of MDS to obtain a high-probability bound for $\widehat{\mathcal{W}}(g|w) - \bar{\mathcal{W}}(g|w)$ that is uniform in $g$.
 \begin{theorem}\label{thm:discrete_bound}
Under Assumptions \ref{ass:toy_example} to   \ref{s_bounded_y}, it holds that
 \begin{equation} \label{simp_rate}
        \sup_{g: \{0,1\}\to \{0,1\}}\mathop{\mbox{\sf E}}|\widehat{\mathcal{W}}(g|w) - \bar{\mathcal{W}}(g|w)| \leq \frac{C}{\sqrt{T-1}},
 \end{equation}
where $C$ is a constant defined in Appendix \ref{simple_mds_proof}.\end{theorem}
The proof of Theorem \ref{thm:discrete_bound} can be found in Appendix \ref{simple_mds_proof}.  Combining (\ref{ineq_s}) (an implication of Assumption \ref{equiv_W}) and (\ref{simp_rate}), we can conclude that the convergence rate of expected regret $\mathop{\mbox{\sf E}}[\mathcal{W}_{T}(g^*|w)-\mathcal{W}_{T}(\hat{g}|w)]$ is upper-bounded by $2c\cdot\frac{C}{\sqrt{T-1}}$ uniformly in $w\in \{0,1\}$.
  \begin{rem} {[}Higher Markov orders{]} \label{rem:higher/inf Markov}
	Theorem \ref{thm:discrete_bound}
	presents our main results of welfare regret upper bounds under the simple
	first-order Markovian structure outlined in Assumption \ref{ass:toy_example}.
	We can extend our analysis to a higher or infinite order Markovian structure as summarized below. We defer the detailed discussion and formal proofs to Appendix \ref{app_multi_markov}.

	If current observations can depend causally or statistically on the realized treatment over the preceding $q$ periods (for some $1<q<T$), we can define $\mathcal{W}_{T}(g|w_{T-q:T-1})$, $\widehat{\mathcal{W}}(\hat{g}|w_{T-q:T-1})$, $\bar{\mathcal{W}}(g|w_{T-q:T-1})$, $g^*$, and $\hat{g}$,  similar to the definitions in \eqref{obj_simple}, \eqref{sample_simple}, \eqref{E_sum_c}, \eqref{best_g}, and \eqref{best_g_hat}, respectively. Let $w_{T-q:T-1}\in \{0,1\}^q$ be a realization of the treatment path spanning from time $T-q$ to $T-1$. Given the modified Assumptions detailed in Appendix \ref{app_finite_markov}, we can apply similar reasoning as in Theorem \ref{thm:discrete_bound} to establish a convergence rate of $\frac{1}{\sqrt{T-q}}$.

	\end{rem}
 \begin{rem}{[}Comparison with \cite{bojinov2019time}{]}\label{Com_B&S}
	The major distinction
	between our work and \cite{bojinov2019time} is that we focus on policy decisions and future welfare, while \cite{bojinov2019time} study estimation and inference on the
	retrospective causal effects. An estimand of interest in \cite{bojinov2019time} is the temporal (zero-lag)
	average treatment effect (ATE) defined as
	\begin{equation}
		\bar{\tau}_{0} :=\frac{1}{T-1}\sum_{t=1}^{T-1}\left[Y_{t}\left(W_{0:t-1},1\right)-Y_{t}\left(W_{0:t-1},0\right)\right].\label{eq:BS_ATE_maintext}
	\end{equation}
	In contrast,  our paper focuses on maximizing
	\begin{equation}
		\mathcal{W}_{T}(G|\mathcal{F}_{T-1}) =\mathsf{E}\left[ \tau_{T}(\mathcal{F}_{T-1}) g(X_{0:T-1})+Y_{T}(W_{0:T-1},0) |\mathcal{F}_{T-1}\right]
	\end{equation}
	where $\tau_{T}(\mathcal{F}_{T-1})$ is the conditional ATE (CATE) at time $T$,
	\begin{equation}
		\tau_{T}(\mathcal{F}_{T-1})=\mathsf{E}\left[Y_{T}(W_{0:T-1},1)-Y_{T}(W_{0:T-1},0)|\mathcal{F}_{T-1}\right]. \label{eq:T-EWM-ATE_maintext}
	\end{equation}

	$\tau_{T}(\mathcal{F}_{T-1})$ is the conditional ATE at an upcoming time
	period of $T$, while $\bar{\tau}_0$ sums up the causal
	effects of the past realized time periods from $0$ to $T-1$. This distinction necessitates different sets of assumptions between our work and \cite{bojinov2019time}. Specifically, the sequential
	unconfoundedness assumption (Assumption \ref{unconf}) is sufficient for unbiased estimation for $\bar{\tau}_0$, whereas it
	falls short for $\tau_{T}(\mathcal{F}_{T-1})$. Consequently,  Assumptions \ref{ass:toy_example}
	(Markov properties) and \ref{equiv_W} (Invariance of welfare ordering) are not imposed in \cite{bojinov2019time}.

	Analogous to the construction of our empirical welfare criterion, we can consider the following estimator for $\tau_{T}(\mathcal{F}_{T-1})$:
	\begin{equation}
		\hat{\tau}_{T}(w)=T(w)^{-1}\sum_{1\leq t\leq T-1:W_{t-1}=w}\left[\frac{Y_{t}W_{t}}{e_{t}(W_{t-1})}-\frac{Y_{t}(1-W_{t})}{1-e_{t}(W_{t-1})}\right].\label{eq:e_ATE_maintext}
	\end{equation}
	To validate $\hat{\tau}_T(w)$ as an estimator for $\tau_{T}(\mathcal{F}_{T-1})$, the crucial
	step is to link the past CATEs (or the welfare in our context) from past periods to the future ones. This
	motivates our introduction of Assumption \ref{equiv_W}, the invariance of welfare ordering.


	The construction of $\hat{\tau}_T(w)$ is model-free and
	selects a subset of past periods that
	share the same conditioning with time $T$. To guarantee the availability of observations sharing the conditioning states, we impose Assumption \ref{ass:toy_example}, limiting the persistence of carryover effects of treatment.
\end{rem}


  \begin{rem} We construct the empirical welfare $\widehat{\mathcal{W}}(g|W_{T-1} = w)$ in (\ref{sample_simple}) by the average with uniform weights. To generalize, we can specify $\widehat{\mathcal{W}}(g|W_{T-1} = w)$ as a weighted average:
\begin{eqnarray} \label{sample_weighted}
\widehat{\mathcal{W}}(g|W_{T-1}=w)
= \sum_{1 \leq t \leq T-1: W_{t-1} = w} a_t \left [\frac{Y_{t}W_{t}g(W_{t-1})}{e_t(W_{t-1})}+ \frac{Y_{t}(1-W_{t})\{1- g(W_{t-1})\}}{1- e_t(W_{t-1})} \right],
\end{eqnarray}
where $a_t \geq 0$ is a prespecified weight assigned to period $1 \leq t \leq T-1$. Modifying intermediate welfare $\bar{\mathcal{W}}(g|w)$ accordingly by
\begin{eqnarray} \label{E_weighted_sum_c}
\bar{\mathcal{W}}(g|w)
=\sum_{1 \leq t \leq T-1 : W_{t-1} = w}a_t\mathop{\mbox{\sf E}}\left[ Y_{t}(W_{t-1},1)g(W_{t-1})+Y_{t}(W_{t-1},0)\left[1-g(W_{t-1})\right]|W_{t-1}\right]
\end{eqnarray}
and imposing Assumption \ref{equiv_W} with the modified intermediate welfare, we can study a condition for the weights that generalizes the welfare regret convergence of Theorem \ref{thm:discrete_bound}. The specification of nonuniform weights can reflect the planner's belief or knowledge on how the period $T$ potential outcome distribution differs from those of the past periods or which past period observations are more informative for the current period decision making. See, e.g., \cite{ishihara2024} for an optimal weighting of pieces of evidence for policy choice when the population in which the policy is implemented  differs from the populations in which pieces of evidence were collected. We leave a formal investigation for the current time-series setting for future research.
\end{rem}

\subsection{Infinite Markov order}
\label{sec:infi_markov}
Time-series models commonly used in macroeconomic policy analysis imply an infinite order Markovian structure. For example, we consider the following modification to  Example \ref{example_q1_Markov}.
\begin{exmp} \label{Example_inf_Markov}
Replace \eqref{eq:Yt} in  Example \ref{example_q1_Markov} with a structural MA($\infty$) model:
\begin{equation}
Y_{t} =\alpha+\sum_{i=0}^{\infty}\beta_{i}W_{t-i}+\sum_{i=0}^{\infty}\gamma_{i}\varepsilon_{t-i},\label{eq:Yt_inf}
\end{equation}
while keeping the conditions \eqref{eq:Wt} to \eqref{eq:Vt_2} unchanged.
This type of $\text{MA}(\infty)$ process underlies the causal impulse response analysis of structural vector autoregressions; see, e.g., \cite{kilian2017structural}.  For the effect of shocks to diminish, both $\gamma_i$ and $\beta_i$ must decay in absolute value with $i\to \infty$. In particular, if $Y_t$ is an AR(1) process, then $\beta_i$ and $\gamma_i$ are polynomials of the AR coefficient.
\end{exmp}

Without requiring stationarity or functional form restrictions, we can extend our framework to infinite Markovian order and obtain convergence of welfare regret conditional on a treatment path of infinite length, $\mathcal{W}_{T}(g^{*}|w_{-\infty:T-1})-\mathcal{W}_{T}(\hat{g}|w_{-\infty:T-1})$.
Let $1< m < T$ and $\hat{g}$ be a policy that maximizes the empirical welfare $\widehat{\mathcal{W}}(g|w_{T-m:T-1})$ with conditioning policy path truncated to $w_{T-m:T-1}$. Appendix \ref{rem: Inf_order} shows the following upper bound for the welfare regret of infinite order Markov models:
\begin{align*}
	\mathcal{W}_{T}(g^{*}|w_{-\infty:T-1})-\mathcal{W}_{T}(\hat{g}|w_{-\infty:T-1})& \leq 2c\sup_{g:\{w_{T-m:T-1}\}\to\{0,1\}}|\bar{\mathcal{W}}(g|\mathcal{F}_{t-1})-\widehat{\mathcal{W}}(g|w_{T-m:T-1})|\\
	& +2\cdot \widetilde{\text{w-bias}}_{\infty}\left(m\right),
\end{align*}
where $\bar{\mathcal{W}}(g|\mathcal{F}_{t-1})$ is the intermediate welfare defined by \eqref{eq:wel_bar_F_m} in Appendix \ref{rem: Inf_order}, around which the empirical welfare is expected to concentrate, and $\widetilde{\text{w-bias}}_{\infty}\left(m\right)$ is the welfare bias due to truncation of the empirical welfare defined by \eqref{eq:inf_markov_bias} in Appendix \ref{rem: Inf_order}.
The first term on the right-hand side represents the average of an MDS. Under the regularity conditions presented in Appendix \ref{rem: Inf_order}, we can show that this term converges at a rate of $\frac{1}{\sqrt{T-m}}$. For the second term on the right-hand side, we shall have $\text{plim}_{m\to\infty}\widetilde{\text{w-bias}}_{\infty}\left(m\right)=0$ under additional conditions that ensure the decay temporal dependence. See Appendix \ref{rem: Inf_order} for more details. Furthermore, in Appendix \ref{App:exmple_2}, we show that the infinite order T-EWM proposed in Appendix \ref{rem: Inf_order} can be applied to  the policy choice problem specified in Example \ref{Example_inf_Markov}.
\section{Continuous covariates} \label{continuous}
This section extends the illustrative example of Section \ref{model_example} by allowing $X_t$ to contain continuous variables. We first introduce the continuous setting in Section \ref{sec:conti_setting}. Then,  Section \ref{sec:conti_kernel} presents a natural approach for handling the continuous policy variable, which employs a kernel function to construct an analogue of the  conditional empirical welfare function \eqref{sample_simple} presented in Section \ref{model_example}.
Finally, having introduced the unconditional welfare in the time-series setting in Section \ref{sec:notation_timing_welfare}, we formally present the \emph{unconditional} T-EWM method in Section \ref{sec:uncon}, developed in parallel with the kernel-based approach.






\subsection{Setting} \label{sec:conti_setting}

In addition to $(Y_t,W_t)$, we incorporate general covariates $Z_t$ into $X_t \in \mathcal{X}$, which can be continuous. Now, $X_t=(Y_t,W_t,Z_t)$.  For simplicity of exposition, we maintain the first-order Markovian structure similarly to the illustrative example, while modifying Assumptions \ref{ass:toy_example}, \ref{bound1}, and \ref{unconf} as follows. It is straightforward to incorporate a higher-order Markovian structure.

\begin{assumption}  [Markov properties] \label{ass:continuous_Markov}The time series of potential outcomes and observable variables satisfy the following conditions:

(i) \emph{Markovian exclusion}: the same as Assumption \ref{ass:toy_example} (i).

(ii) \emph{Markovian exogeneity}: for $t=1, \dots, T$ and any treatment path $w_{0:t}$,
\begin{eqnarray}
Y_t(w_{0:t}) \perp X_{0:t-1} | X_{t-1},
\end{eqnarray}
and for $t=1, \dots, T-1$,
\begin{eqnarray}
W_t \perp X_{0:t-1} | X_{t-1}.
\end{eqnarray}
\end{assumption}
Similarly to (\ref{obj_simple}) in the illustrative example, Assumption \ref{ass:continuous_Markov} implies that we can reduce the conditioning information of $\mathcal{F}_{t-1}$ to only $X_{t-1}$ and reduce the policy to a binary map of $X_{t-1}$ without any loss of conditional welfare.
Following these reductions and considering the planner's focus on the policy choice in period $T$, we can formulate the planner's objective function as follows:
\begin{eqnarray} \label{conti_obj_kernel}
\mathcal{W}_{T}(g|X_{T-1})
=\mathop{\mbox{\sf E}}\left[ Y_{T}(W_{T-1},1)g(X_{T-1})+Y_{T}(W_{T-1},0)[1-g(X_{T-1})]|X_{T-1}\right].
\end{eqnarray}
where the policy function $g$ is defined in \eqref{eq:policy_f}.
We assume the strict overlap and unconfoundedness restrictions under the general covariates as follows.
\begin{assumption}[Strict overlap]\label{bound_c}
         Let $e_t(x)=\Pr(W_{t}=1|X_{t-1}=x)$ be the propensity score at time $t$. There exists $\kappa \in (0,1/2)$, such that
        \begin{equation*}
                \kappa \leq e_t(x) \leq 1-\kappa
        \end{equation*}
        holds for every $t = 1, \dots, T-1$ and each $x\in\mathcal{X}$.
        \end{assumption}
{
\begin{assumption}[Unconfoundedness] \label{unconf_c}
        For any $t=1,2,\dots,T-1$ and $w\in \{0,1\}$
\[
Y_{t}(W_{t-1},w)\perp W_t|X_{0:t-1}.
\]
\end{assumption}

Under Assumptions \ref{ass:continuous_Markov} and \ref{unconf_c}, we can generalize (\ref{simple_cond_exp}) by including the set of covariates in the conditioning variables: for any measurable function $f$,
\begin{eqnarray} \label{equiv_markov}
\mathop{\mbox{\sf E}}(f(Y_t(W_{0:t}),W_t)| \mathcal{F}_{t-1})=\mathop{\mbox{\sf E}}(f(Y_t(W_{t-1},W_t),W_t)|X_{t-1})= \mathop{\mbox{\sf E}} (f(Y_t,W_t)|X_{t-1}).
\end{eqnarray}

\subsection{Kernel approach for conditional welfare} \label{sec:conti_kernel}
For continuous conditioning covariates $X_{T-1}$, a simple sample analogue of the objective function is not available due to the lack of multiple observations at any single conditioning value of $X_{T-1}$. One approach is to use nonparametric smoothing to construct an estimate for conditional welfare.

For simplicity, we let $X_{T-1}\in \mathbb{R}$. Then, with a kernel function $K(\cdot)$ and a bandwidth $h$, the empirical conditional welfare can be rewritten as\footnote{If the set of conditioning variables $X_{T-1}$ contains both continuous and discrete components, we can adapt a hyper method to construct a valid sample analogue combing kernel-smoothing (for continuous variables) and subsamples (for discrete variables). In this section, we focus on the case where the target welfare function is conditional on a univariate continuous variable}
{\small \begin{eqnarray} \label{eq:sample_kernel}
        \widehat{\mathcal{W}}(g|x)
        =\frac{\sum_{t=1}^{T-1}K_h(X_{t-1},x)\widehat{\mathcal{W}}_t(g)}{\sum_{t=1}^{T-1}K_h(X_{t-1},x)},
        \end{eqnarray}}
where $(\cdot|x)$ denotes $(\cdot|X_{T-1}=x)$, $\widehat{\mathcal{W}}_t(g) := \frac{Y_{t}W_{t}}{e_t(X_{t-1})} g(X_{T-1})+ \frac{Y_{t}(1-W_{t})}{1- e_t(X_{t-1})}[1-g(X_{T-1})]$, and $K_h(a,b):=\frac{1}{h}K(\frac{a-b}{h})$.

Using the same notation, we can define $\mathcal{W}_{t}(g|x)
=\mathop{\mbox{\sf E}}[ Y_{t}(W_{t-1},1)g(X_{t-1})+Y_{t}(W_{t-1},0)[1-g(X_{t-1})]|X_{t-1}=x]$.
Then, similar to \eqref{best_g} and \eqref{best_g_hat}, we define
\begin{eqnarray*}
g_x^* &\in& \mbox{argmax}_{g}\mathcal{W}_T(g|x),\\
\hat{g}_x &\in& \mbox{argmax}_{g}\widehat{\mathcal{W}}(g|x),
\end{eqnarray*}
to be the maximizers of $\mathcal{W}_T(g|x)$  and $\widehat{\mathcal{W}}(g|x)$. Moreover,
define an intermediate welfare function:
\begin{eqnarray} \label{eq:kernel_bar}
\bar{\mathcal{W}}_{h}(g|x)
&=&\frac{\sum_{t=1}^{T-1}K_h(X_{t-1},x)\mathcal{W}_t(g|x)}{\sum_{t=1}^{T-1}K_h(X_{t-1},x)}.
\end{eqnarray}

The invariance of welfare ordering assumption is modified to:
\begin{assumption}[Invariant of ordering for the conditional welfare] \label{equiv_W_c}
        For any  $g:\mathcal{X}\to \{0,1\}$ and $x\in \mathcal{X}$, there exists some positive constant $c$
        \begin{equation} \label{uncon_con_kernel}
        \mathcal{W}_{T}(g_x^*|x)- \mathcal{W}_{T}(g|x) \leq c\big[\bar{\mathcal{W}}_h(g_x^*|x)-\bar{\mathcal{W}}_h(g|x)\big].
        \end{equation}
\end{assumption}
Similar to Assumption  \ref{equiv_W}, Assumption  \ref{equiv_W_c} holds if the stochastic process $S_t(x):=Y_{T}(W_{T-1},1)$ $g(X_{T-1})+Y_{t}(W_{T-1},0)[1-g(X_{T-1})]|_{X_{T-1}=x}$ is weakly stationary.

The following theorem shows an upper bound for conditional regret in the one-dimensional covariate case (i.e., $X_t\in \mathbb{R}^1$). This result can be readily extended to the multiple-covariate case.
\begin{theorem}  \label{thm: kernel}
Under Assumptions \ref{ass:continuous_Markov} to \ref{equiv_W_c} and \ref{kernel}  specified in Appendix \ref{max_kernel},
        \begin{equation*}
        \underset{P_T\in\mathcal{P}_T(M,\kappa)}{\sup}\sup_{g:\mathcal{X} \to \{0,1\}}\sup_{x\in \mathcal{X}}\mathop{\mbox{\sf E}}{_{P_T}}[\mathcal{W}_{T}(g |x)-\mathcal{W}_{T}(\hat {g}_x|x)]  \leq c_1\left[\left(\sqrt{(T-1)h}\right)^{-1} +h^{2}\right].
\end{equation*}
\end{theorem}
Setting $h = O(T^{-1/5})$, the right-hand side bound becomes $O(T^{-2/5})$. A proof is presented in Appendix \ref{max_kernel}.

\subsection{Unconditional T-EWM}\label{sec:uncon}
In this section, we shift the focus to maximizing unconditional welfare. For ease of exposition with the unconditional T-EWM, given a policy function $g:\mathcal{X}^{T} \to \{0,1\}$, we define the corresponding region in the space of the covariate vector for which the decision rule chooses $W_T =1$ to be
 \begin{equation} \label{gG}
G = \{ X_{0:T-1} : g(X_{0:T-1}) = 1 \} \subset \mathcal{X}^{T}.
\end{equation}
We refer to $G$ as a \emph{decision set}. Under Assumption \ref{ass:continuous_Markov}, the unconditional welfare under policy $G$ is
\begin{equation}\label{wel:uncondi}
        \mathcal{W}_{T}(G) =\mathsf{E}\left[ Y_{T}(W_{T-1},1)1(X_{T-1}\in G)+Y_{T}(W_{T-1},0)1(X_{T-1}\notin G)
        \right].
\end{equation}
The sample analogue of $\mathcal{W}_T(G)$ can be expressed as
\begin{eqnarray} \label{eq:unconditional_sample}
\widehat{\mathcal{W}}(G)
=\frac{1}{T-1}\sum_{t=1}^{T-1} \left[\frac{Y_{t}W_{t}}{e_t(X_{t-1})}1(X_{t-1}\in G)+ \frac{Y_{t}(1-W_{t})}{1- e_t(X_{t-1})}1(X_{t-1}\notin G) \right].
\end{eqnarray}
For $\mathcal{G}$ being a class of decision sets, we define
\begin{eqnarray}
G_*&\in& \mbox{argmax}_{G\in\mathcal{G}}\mathcal{W}_{T}(G),\label{maxG}\\
\hat{G}  &\in&  \mbox{argmax}_{G\in\mathcal{G}}\widehat{\mathcal{W}}(G).\label{sample:maxG}
\end{eqnarray}
 In addition, define two intermediate welfare functions,
\begin{align} \label{2intermediate_welfare}
        \bar{\mathcal{W}}(G) &=\frac{1}{T-1}\sum_{t=1}^{T-1}\mathsf{E}\left[ Y_{t}(W_{t-1},1)1(X_{t-1}\in G)+Y_{t}(W_{t-1},0)1(X_{t-1}\notin G)
        |\mathcal{F}_{t-1}\right], \nonumber \\
        \widetilde{\mathcal{W}}(G) &=\frac{1}{T-1}\sum_{t=1}^{T-1}\mathsf{E}\left[ Y_{t}(W_{t-1},1)1(X_{t-1}\in G)+Y_{t}(W_{t-1},0)1(X_{t-1}\notin G)\right].
\end{align}
The need for an additional intermediate welfare function, $\widetilde{\mathcal{W}}(G)$, arises because we aim to bound the unconditional welfare. After centering the empirical welfare around its conditional mean, $\bar{\mathcal{W}}(G) $, we are left with another difference, $\bar{\mathcal{W}}(G)-\widetilde{\mathcal{W}}(G)$, which we will control below.

To obtain a regret bound for unconditional welfare, Assumption \ref{equiv_W} is modified to:
\begin{assumption} [Invariant of ordering for the unconditional welfare]\label{equiv_W_uc}
For any $G\in \mathcal{G}$, there exists some constant $c$, such that
\begin{equation} \label{uncon_con2}
\mathcal{W}_{T}(G_*)- \mathcal{W}_{T}(G) \leq c\big[\widetilde{\mathcal{W}}(G_*)-\widetilde{\mathcal{W}}(G)\big]
\end{equation}
holds with probability one, i.e., $P_T \left( \text{inequality (\ref{uncon_con2}) holds} \right) = 1$,
where $P_T$ is the probability distribution for $X_{0:T-1}$.
\end{assumption}
Note that by Assumption \ref{lower_p_dis_con}, we have
\begin{equation} \label{welfare_inequality}
        \mathcal{W}_{T}(G_{*})-\mathcal{W}_{T}(\hat{G})
         \leq c\left[\mathcal{\widetilde{W}}(G_{*})-\mathcal{\widetilde{W}}(\hat{G})\right]
         \leq 2c\sup_{G\in \mathcal{G}}\left|\widehat{\mathcal{W}}(G)-\mathcal{\widetilde{W}}(G)\right|,
\end{equation}
where the first inequality follows from \eqref{uncon_con2}, and the second inequality follows from an argument similar to the one for \eqref{ineq_s}.

Note that $\widehat{\mathcal{W}}(G)-\widetilde{\mathcal{W}}(G)$
is \emph{not} a sum of MDS. Instead, it can be decomposed as
\begin{eqnarray} \label{decompose_uncon}
\widehat{\mathcal{W}}(G)-\widetilde{\mathcal{W}}(G) & =& \bar{\mathcal{W}}(G)-\widetilde{\mathcal{W}}(G) +( \widehat{\mathcal{W}}(G)-\bar{\mathcal{W}}(G))
= I+II,
\end{eqnarray}
where $I:=\bar{\mathcal{W}}(G)-\widetilde{\mathcal{W}}(G)$ and $II:=\widehat{\mathcal{W}}(G)-\bar{\mathcal{W}}(G)$. Subject to assumptions specified later, Theorem \ref{thm:mds_bound} below shows that $II$, which is a sum of MDS,  converges at $\frac{1}{\sqrt{T-1}}$-rate, and Theorem \ref{mean_bound} below shows that $I$ converges at the same rate.

\eqref{decompose_uncon} reveals that our proof strategy is considerably more complicated than the proof for the EWM model with i.i.d. observations of \cite{kitagawa2018should}, although the rates are similar. Specifically, we need to derive a  bound for the tail probability of the sum of martingale difference
sequences. In addition, we need to handle complex functional classes induced by non-stationary processes.
For the EWM model, the main task is to show the convergence rate of a sample analogue of $II$, which can be achieved with standard empirical process theory for i.i.d.\  samples. In comparison, we not only have to treat our $II$ more carefully due to time-series dependence, but we also have to deal with $I$.

Similar to our approach with the simple model described in Section \ref{model_example},  we assume that $Y_t$, which is reintroduced in Section \ref{sec:conti_setting}, has a bounded support.
\begin{assumption}[Bounded outcome] \label{ass bounded y}

	There exists $M<\infty$ such that the support of outcome variable
	$Y_t$ is contained in $[-M/2,M/2]$.

\end{assumption}

\subsubsection{Bounding II}

Define empirical welfare at time $t$ and its population conditional expectation as follows,
\begin{align*}
\widehat{\mathcal{W}}_t(G)&=\frac{Y_{t}W_{t}}{e_t(X_{t-1})}1(X_{t-1}\in G)+ \frac{Y_{t}(1-W_{t})}{1- e_t(X_{t-1})}1(X_{t-1}\notin G),\\
\bar{\mathcal{W}}_t(G)&=\mathsf{E}\left[ Y_{t}(W_{t-1},1)1(X_{t-1}\in G)+Y_{t}(W_{t-1},0)1(X_{t-1}\notin G)|\mathcal{F}_{t-1}\right].
\end{align*}
Then, we examine two summations:
\begin{eqnarray*}
	\widehat{\mathcal{W}}(G)=\frac{1}{T-1}\sum_{t=1}^{T-1}\widehat{\mathcal{W}}_t(G), \quad
	\bar{\mathcal{W}}(G) =\frac{1}{T-1}\sum_{t=1}^{T-1} \bar{\mathcal{W}}_t(G).
	\end{eqnarray*}

For each $t=1, \dots, T-1$, define a function class indexed by $G \in \mathcal{G}$,
\begin{equation} \label{H_class}
	\mathcal{H}_t=\{h_t(\cdot;G)=\widehat{\mathcal{W}}_t(G)-\bar{\mathcal{W}}_t(G):G\in\mathcal{G}\},
\end{equation}
where the arguments of the function $h_t(\cdot;G)$ are $Y_{t}$, $W_t$, and $X_{t-1}$. In the following, we use $n$ to represent the number of summands since the endpoints of samples may vary across different settings in this section, subsequent sections, and appendices. For example, in the case of multi-period welfare functions, the endpoint of a sample is no longer fixed at $T-1$. Given the class of functions $\mathcal{H}_t$, we consider a martingale difference array $\{h_t(Y_{t},W_t,X_{t-1};G)\}_{t=1}^{n}$,
and denote its average by
$${\mathop{\mbox{\sf E}}}_{n} h\stackrel{\mathrm{def}}{=} \frac{1}{n}\sum_{t=1}^{n} h_t(Y_{t},W_t,X_{t-1}; G),$$
where $h \stackrel{\mathrm{def}}{=} \left\{ h_{1}(\cdot; G),h_{2}(\cdot;,G),\dots,h_{n}(\cdot;G)\right\} $, and we suppress $n$ and $G$ if there is no confusion in the context.

 Since we do not restrict $X_t$ to be stationary, we shall handle a vector of function classes that possibly vary over $t$. To this end, we define the following set of notations. Let $H_t$ denote the envelope for the function class $\mathcal{H}_t$, and $\overline{H}_{n}=(H_{1},H_{2},\cdots,H_{n})^{\prime}$, and $\mathbf{H_{n}}=\mathcal{\mathcal{H}}_{1}\times \mathcal{\mathcal{H}}_{2}\times \dots \times \mathcal{H}_{n}.$
 For a function $f$ supported on $\mathcal{X}$, define $\|f\|_{Q,r} \stackrel{\mathrm{def}}{=} (\int_{x\in\mathcal{X}} |f(x)|^r dQ(x))^{1/r} $, and for an $n$-dimensional vector $v=\left\{ v_{1},\dots,v_{n}\right\} $,
its $l_{2}$ norm is denoted by $|v|_{2}\overset{\text{def}}{=}\left(\sum_{i=1}^{n}v_{i}^{2}\right)^{1/2}$.
The covering number of a function class $\mathcal{H}$ w.r.t. a metric $\rho$ is denoted by $\mathcal{N}(\varepsilon, \mathcal{H}, \rho(.))$. For two series of functions $f=\{f_t
\}_{i=1}^n$ and $g=\{g_t
\}_{i=1}^n$, define the metrics $\rho_{2,n}(f,g) = (n^{-1}\sum_t|f_t-g_t|^2)^{1/2}$ and   $\sigma_n(f,g) = \left( n^{-1} \sum_t \mathop{\mbox{\sf E}}[(f_t-g_t)^2| \mathcal{F}_{t-1}]\right)^{1/2}$. Let $\alpha_{n}$ denote an $n$-dimensional vector in $\mathbb{\mathbb{R}}^{n}$ and $\circ$ denote the element-wise product.
In the next assumption, we want to bound the covering number of $\mathcal{N}(\delta|\overline{H}{}_{n}|_{2}, \mathbf{H_{n}}, \rho_{2,n})$ by the covering number of all its one-dimensional projection.
 Now, let $A\lesssim B$ denote that there exists some constant $c_0$ such that $A \leq c_0\cdot B$.
\begin{assumption}[Function classes]\label{cover1}
Let $n=T-1$. For  any discrete measures $Q$, any  $\alpha_{n}\in\mathbb{R}_{+}^{n}$, and all $\delta>0$, we have
\begin{align}\label{cover}
 	\mathcal{N}(\delta|\tilde{\alpha}_{n}\circ\overline{H}{}_{n}|_{2},\mathbf{\tilde{\alpha}_{n}\circ H_{n}},|.|_{2})
\leq \max_t\sup_Q \mathcal{N} ( \delta \|H_t\|_{Q,2}, \mathcal{H}_t, \|.\|_{Q,2}) \lesssim K (v+1) (4e)^{v +1}(\frac{2}{\delta})^{crv},
\end{align}
where $K$, $v$, $c$, and $e$ are positive constants; $r$ is a positive integer and $\tilde{\alpha}_{n,t}=\frac{\sqrt{\alpha_{n,t}}}{\sqrt{\sum_{t}\alpha_{n,t}}}$.
\end{assumption}

Assumption \ref{cover1} restricts the complexity of the function class to be of polynomial discrimination, and the complexity index $v$ appears in the derived regret bounds. See Appendix \ref{justify_entropy} for a justification for this assumption.

\begin{assumption}[Empirical sum]\label{norm}
{
There exists a constant $L>0$ such that  $\Pr (\sigma_n(f,g)/$ $ \rho_{2,n}(f,g) >L) \to 0$ as $n \to \infty$.}
Also, $\Pr (\left( n^{-1} \sum_t \mathop{\mbox{\sf E}}[(f_t-g_t)^2| \mathcal{F}_{t-2}]\right)^{1/2}/ \rho_{2,n}(f,g) >L) \to 0$ as $n \to \infty$.
\end{assumption}

$\rho_{2,n}(f,g)^2$ is the quadratic variation difference and $\sigma_n(f,g)^2$ is its conditional equivalent. It is evident that $\rho_{2,n}(f,g)^2- \sigma_n(f,g)^2$ involves  martingale difference sequences. In the special case of i.i.d. observations, $\sigma_n(f,g)^2$ is equivalent to the sample average of unconditional expectations. Assumption \ref{norm} can thus be viewed as specifying that $\rho_{2,n}(f,g)^2$ and $\sigma_n(f,g)^2$ are asymptotically equivalent in a probability sense. A similar condition can be seen, for example, in Theorem 2.23 of \cite{hall2014martingale}.

Now, let $A\lesssim_p B$ denote $A=O_p(B)$. Then, we have for $II$:

\begin{theorem}\label{thm:mds_bound}
Under Assumptions \ref{ass:continuous_Markov} to \ref{unconf_c}, and \ref{ass bounded y} to \ref{norm},
\begin{equation*}
\sup_{G\in \mathcal{G}}|\widehat{\mathcal{W}}(G)-\bar{\mathcal{W}}(G)|
\lesssim_p C\sqrt{\frac{v}{T-1}},
\end{equation*}
where $C$ is a constant that depends only on $M$ and $\kappa$.
\end{theorem}
The proof of Theorem \ref{thm:mds_bound} is presented in Appendix \ref{proof_MDS_t}.

\subsubsection{Bounding I}

Here, we complete the process of bounding unconditional regret. Let us define
 \begin{equation*}
S_{t}(G) = \begin{array}{c}
	Y_{t}(W_{t-1},1)1(X_{t-1}\in G)
		+Y_{t}(W_{t-1},0)1(X_{t-1}\notin G),
	\end{array}
\end{equation*}
and
\begin{equation} \label{eq:bar_St}
\overline{S}_{t}(G)= \mathop{\mbox{\sf E}}(S_{t}(G)| \mathcal{F}_{t-1}) - \mathop{\mbox{\sf E}}(S_{t}(G)| \mathcal{F}_{t-2}),
\end{equation}
\begin{equation} \label{eq:tilde_St}
\tilde{S}_{t}(G) =  \mathop{\mbox{\sf E}}(S_{t}(G)| \mathcal{F}_{t-2}) -  \mathop{\mbox{\sf E}}(S_{t}(G)).
\end{equation}
We can apply a similar technique in Theorem \ref{thm:mds_bound} (See Lemma \ref{bound} in Appendix \ref{proof_MDS_t}) to bound the sum of $\overline{S}_{t}(G)$. The second term $\tilde{S}_{t}(G)$ is handled below.
Recalling \eqref{2intermediate_welfare} and \eqref{decompose_uncon}, we have $I=\frac{1}{T-1}\sum_{t=1}^{T-1}\tilde{S}_{t}(G)+ \frac{1}{T-1}\sum_{t=1}^{T-1}\overline{S}_{t}(G)$. Define the function classes,
\begin{eqnarray*}
\overline{\mathcal{S}}_{t} &=&\{ f_{t} = \mathop{\mbox{\sf E}}(S_{t}(G)| \mathcal{F}_{t-1}) - \mathop{\mbox{\sf E}}(S_{t}(G)|\mathcal{F}_{t-2}): G \in \mathcal{G}  \},\\
\tilde{\mathcal{S}}_{t} &=&\{ f_{t} = \mathop{\mbox{\sf E}}(S_{t}(G)| \mathcal{F}_{t-2}) - \mathop{\mbox{\sf E}}(S_{t}(G)): G \in \mathcal{G}  \}.
\end{eqnarray*}
Note that by Assumption \ref{ass:continuous_Markov}, we have $\mathop{\mbox{\sf E}}(S_{t}(G)| \mathcal{F}_{t-1}) = \mathop{\mbox{\sf E}}(S_{t}(G)| X_{t-1})$, and $\mathop{\mbox{\sf E}}(S_{t}(G)| \mathcal{F}_{t-2})=\mathop{\mbox{\sf E}}(S_{t}(G)| X_{t-2})$.
\begin{definition}  Let $\{{\varepsilon}_t\}_{t=-\infty}^{\infty}$ be a sequence of i.i.d.\ random variables, and $\{g_t\}_{t=-\infty}^{\infty}$ is a sequence of measurable functions of ${\varepsilon}$'s, which might vary with time $t$. For a process $\xi_{.}\stackrel{\mathrm{def}}{=}\{\xi_t\}_{t=-\infty}^{\infty}$ with $\xi_t \stackrel{\mathrm{def}}{=} g_t({\varepsilon}_t, {\varepsilon}_{t-1},\cdots)$ and integers $l,q\geq 0$,
we define the dependence adjusted norm for an arbitrary process $\xi_t $ as
\begin{equation}
\theta_{\xi,q} =  \sum_{l = 0}^{\infty} \max_t  \|\xi_t - \xi_{t,l}^*\|_{q},
\end{equation}
where $\|\cdot\|_q$ denotes $(\mathop{\mbox{\sf E}}|\cdot|^q)^{1/q}$, and $\xi_{t,\ell}^*=g_t({\varepsilon}_t,\cdots, {\varepsilon}^{\prime}_{t-l},\cdots)$ is the random variable $\xi_{t}$ with its $l$-th lag replaced by $\varepsilon^{\prime}_{t-l}$,  an independent copy of $\varepsilon_{t-l}$. The subexponential/Gaussian dependence adjusted norm is given by:
\begin{equation}
    \Phi_{\phi_{\tilde{v}}} (\xi_{.})= \sup_{q\geq 2}(\theta_{\xi,q}/q^{\tilde{v}}),\label{eq:DA_norm}
\end{equation}
where  $\tilde{v} = 1/2$ (resp. 1) corresponds to the case that the process $\xi_{i}$ is sub-Gaussian (resp. sub-exponential). \end{definition}
\begin{assumption}
[Data generating process] \label{D.1} $X_t={g}{}_{t}({\varepsilon}_t, {\varepsilon}_{t-1}, \cdots),$ where ${\varepsilon}_t$ is a sequence of i.i.d. random variables, and $g_t$  is a measurable function of ${\varepsilon}$'s.
\end{assumption}

{It shall be noted that Assumption \ref{D.1} implies  $\tilde{S}_{t}(G) = \tilde{g}_{t}({\varepsilon}_t,{\varepsilon}_{t-1}, \cdots )$, where $\tilde{g}$ is another measurable function of ${\varepsilon}$'s.}

\begin{assumption}
\label{D.2}
(i) (Markov exogeneity for $Z$) $Z_t \perp X_{0:t-1} | X_{t-1}$.
(ii) (Envelope functions)
$ \tilde{S}_{t}(G) $ has an envelope $\tilde{F}_{t}(\cdot)$,
i.e., $\sup_{G\in \mathcal{G}}  |\mathop{\mbox{\sf E}}(S_{t}(G)| X_{t-2}=x) -  \mathop{\mbox{\sf E}}(S_{t}(G))|\leq |\tilde{F}_t(x)|$ for every $t$  and $x\in\mathcal{X}$; for every $t$ and any discrete measures $Q$, it holds that $\mbox{max}_t\mbox{sup}_Q     \|\tilde{F}_{t}\|_{Q,2}\leq \infty$.
\end{assumption}


The next assumption pertains to the tail of the underlying innovations of time series.
   \begin{assumption}[Tail assumption] \label{D.3}For $\tilde{v} =1/2$ or $1$, it holds that $\sup_{G\in \mathcal{G}}\Phi_{\phi_{\tilde{v}}}(\tilde{S}_{.}(G))   < \infty.$
   \end{assumption}


}
        \begin{assumption}[Further function classes]
        \label{D.4}  Suppose that $F_{t}$ (resp. $\tilde{F}_{t}$) is the envelope of the function class $ \overline{\mathcal{S}_t}$ (resp. $\tilde{\mathcal{S}_t}$).
        Define $\overline{F}_n = (F_{1}, F_2, \cdots,F_{n})$ (resp. $\overline{\tilde{F}}_n = (\tilde{F}_{1}, \tilde{F}_2, \cdots,\tilde{F}_{n})$), and $\mathbf{F}_{n} = \{\overline{\mathcal{S}}_{1},\overline{\mathcal{S}}_{2}, \cdots, \overline{\mathcal{S}}_{n}\}$ (resp. $\tilde{\mathbf{F}}_{n} = \{\tilde{\mathcal{S}}_{1},\tilde{\mathcal{S}}_{2}, \cdots, \tilde{\mathcal{S}}_{n}\}$).
 Let $Q$ denote a discrete measure over a finite number of $n$ points. For $n=T-1$  and all $\delta>0$,
{
there exist positive constants $v$, $V$, and $c$, such that, }
\begin{eqnarray*}
\mathcal{N}(\delta|\tilde{\alpha}_{n}\circ\overline{F}_{n}|_{2},\mathbf{\tilde{\alpha}_{n}\circ F_{n},|.|_{2}})&\leq&\max_t \sup_Q \mathcal{N}(\delta \|\tilde{F}_{t}\|_{Q,2}, \overline{\mathcal{S}}_t, \|.\|_{Q,2}) \lesssim (1/\delta)^{cV},\\
\mathcal{N}(\delta|\tilde{\alpha}_{n}\circ\overline{\tilde{F}}{}_{n}|_{2},\mathbf{\tilde{\alpha}_{n}\circ \tilde{F}_{n},|.|_{2}})&\leq&\max_t \sup_Q \mathcal{N}(\delta \|\tilde{F}_{t}\|_{Q,2}, \tilde{\mathcal{S}}_t, \|.\|_{Q,2}) \lesssim (1/\delta)^{cv}.
\end{eqnarray*}
\end{assumption}
Assumption \ref{D.1} imposes that the time series $\tilde{S}_{t}(G)$ and $X_t$ can be expressed as measurable functions of i.i.d.\ innovations ${\varepsilon}_t$. Assumption \ref{D.2}(i) completes the Markov exogeneity (Assumption \ref{ass:continuous_Markov}) in the presence of the covariate $Z_t$, which is introduced at the beginning of Section \ref{continuous}.  Assumption \ref{D.2}(ii) is a standard envelope assumption on the function class, and it states that the function  of interest is enveloped by a function of $X_{t-2}$.
Assumption \ref{D.3} implies that $\Phi_{\phi_{\tilde{v}}}(\tilde{S}_.(G))< \infty$  for $\tilde{v} = 1/2$ or 1. Assumption \ref{D.4} restricts the complexity of the function classes. Based on these assumptions, we have the following rate:


\begin{theorem} \label{mean_bound}

Under Assumptions \ref{ass:continuous_Markov} to \ref{unconf_c}, \ref{ass bounded y}, and  \ref{norm} to \ref{D.4},
\begin{equation*}
\underset{G\in \mathcal{G}}{\sup} \left|\bar{\mathcal{W}}(G)-\widetilde{\mathcal{W}}(G)\right|\lesssim_p \frac{c_T [2V(\log T)e\gamma ]^{1/\gamma}  \sup_{G\in \mathcal{G}}\Phi_{\phi_{\tilde{v}}}(\tilde{S}_{.}(G))}{\sqrt{T-1}}+ C\sqrt{\frac{v}{T-1}},
\end{equation*}
where $c_T$ is a large enough constant; $\tilde{v} = 1/2$ or 1, and $\gamma = 1/(1+2\tilde{v})$; $V$ and $v$ are the constants defined in Assumption \ref{D.4}; and $C$ is a similar constant in Theorem \ref{thm:mds_bound}, which depends only on $M$ and $\kappa$.
\end{theorem}
A proof is presented in Appendix \ref{Proof_of_mean_bound}.
The bound depends on the complexity of the function class, $V$ and $v$, and the time-series dependency, $\mbox{sup}_{G}\Phi_{\phi_{\tilde{v}}}(\tilde{S}_{.}(G))$. As we consider the sample analogue of unconditional welfare, all the observations are utilized, resulting in a $\frac{1}{\sqrt{T-1}}-$rate of convergence.

\subsubsection{Regret bound}

Now, we can obtain the overall bound for unconditional welfare, using \eqref{welfare_inequality}. Let $P_T$ be a joint probability distribution of a sample path of length $(T-1)$, $\mathcal{P}_T(M,\kappa)$ be the class of $P_T$, which satisfies Assumptions \ref{ass:continuous_Markov} to \ref{unconf_c}, and \ref{equiv_W_uc} to  \ref{D.4}; let ${\mathop{\mbox{\sf E}}}_{P_{T}}$ be the expectation taken over different realizations of random samples. Recall that $x_{obs}$ is defined to be the observed value of $X_{T-1}$.
\begin{theorem} \label{thm:mds_mean_bound_unconditional}
Under Assumptions \ref{ass:continuous_Markov} to \ref{unconf_c}, and \ref{equiv_W_uc} to  \ref{D.4},
{\small
\begin{equation} \label{eq:main_Theorem}
\underset{P_{T}\in\mathcal{P}_T(M,\kappa)}{\sup} {\mathop{\mbox{\sf E}}}_{P_T}[\mathcal{W}_{T}(G_*)-\mathcal{W}_{T}(\hat{G})]\lesssim C \sqrt{\frac{v}{T-1}} + \frac{c_T [2V(\log T)e\gamma]^{1/\gamma} \sup_{G\in \mathcal{G}} \Phi_{\phi_{\tilde{v}}}(\tilde{S}_{.}(G))}{\sqrt{T-1}},
\end{equation}}
where $G_*$ is the optimal policy defined in \eqref{maxG}.
\end{theorem}
This theorem follows from \eqref{welfare_inequality}, Theorems \ref{thm:mds_bound} and \ref{mean_bound}, and analogous reasoning presented in Appendix \ref{simple_mds_proof}.


\subsubsection{Using unconditional T-EWM to bound conditional regret}\label{sec:uncon_con}

The kernel method for the conditional regret discussed in Section \ref{sec:conti_kernel} is a direct way to estimate an optimal policy with the conditional welfare criterion. However, the localization by bandwidth slows down the speed of learning; the regret of conditional welfare can only achieve a $\frac{1}{\sqrt{(T-1)h}}$-rate of convergence rather than a $\frac{1}{\sqrt{T-1}}$-rate.
This slow convergence rate may limit its practicality when the covariates $X_{t-1}$ are multi-dimensional or when we extend the order of Markovian dependence to multiple periods.  Therefore, in what follows, we instead pursue a novel approach that estimates a conditional optimal policy rule by maximizing an empirical analogue of unconditional welfare over a specified class of decision sets, $\mathcal{G}$. We show that under additional assumptions, this approach can lead to the convergence rate of the conditional welfare that is free from the curse of dimensionality. To this end, we rewrite the conditional welfare \eqref{conti_obj_kernel} with the notation for the decision set $G$:
\begin{eqnarray} \label{conti_obj}
\mathcal{W}_{T}(G|x)
=\mathop{\mbox{\sf E}}\left[ Y_{T}(W_{T-1},1)1(X_{T-1}\in G)+Y_{t}(W_{T-1},0)1(X_{T-1}\notin G)|X_{T-1}=x\right].
\end{eqnarray}
The relationship between $g$ and $G$ is shown in \eqref{gG}. We first clarify how a maximizer of conditional welfare  \eqref{conti_obj} can be linked to a maximizer of unconditional welfare. With this result in hand, we can focus on estimating unconditional welfare and choosing a policy by maximizing it.  In our setup, the complexity of the functional class needs to be specified by the user.

 We will show that this approach can attain a $\frac{1}{\sqrt{T-1}}$ rate of convergence. Faster convergence relative to the kernel approach comes at the cost of imposing an additional restriction on the data-generating process, as we spell out in the next assumption.



\begin{assumption}[Correct specification] \label{ass:correct_specify}


For the conditional welfare as defined in \eqref{conti_obj} and
the unconditional welfare under policy $G$ defined in \eqref{wel:uncondi}. At every $x\in \mathcal{X}$, it holds that
\[ \mbox{argmax}_{G\in\mathcal{G}} \mathcal{W}_{T}(G)\subset \mbox{argmax}_{G\in\mathcal{G}} \mathcal{W}_{T}(G|x). \]
\end{assumption}
This assumption ensures that maximizing unconditional welfare corresponds to maximizing conditional welfare over $\mathcal{G}$. A sufficient condition for the equivalence of maximizing conditional and unconditional welfare is that the specified class of policy rules, $\mathcal{G}$, includes the first best policy $G^*_{FB}$ for the unconditional problem, where
\begin{equation} \label{eq:first_best}
G^*_{FB}:=\left\{x\in \mathcal{X}:\mathsf{E}[Y_{T}(W_{T-1},1)-Y_{T}(W_{T-1},0)|X_{T-1}=x]\geq 0\right\}.
\end{equation}
This sufficient condition states that the class of policy rules over which unconditional empirical welfare is maximized contains the set of points in $\mathcal{X}$ where the conditional average treatment effect $\mathsf{E}[Y_{T}(W_{T-1},1)-Y_{T}(W_{T-1},0)|X_{T-1}=x]$ is positive. This assumption thus restricts the distribution of potential outcomes at $T$ and its dependence on $X_{T-1}$. We refer to Assumption \ref{ass:correct_specify} as `correct specification'.\footnote{\cite{kitagawa2018should}, \cite{KST21}, and \cite{Sakaguchi21} consider correct specification assumptions exclusively for unconditional welfare criteria. These assumptions correspond to $G_{FB}^{\ast} \in \mathcal{G}$.} The planner can be confident about Assumption \ref{ass:correct_specify} if, for instance, $\mathsf{E}[Y_{T}(W_{T-1},1)-Y_{T}(W_{T-1},0)|X_{T-1}=x]$ is believed to be monotonic in $x$ (element-wise) and the class $\mathcal{G}$ consists of decision sets with monotonic boundaries \citep{MT17,KST21}.





With Assumption \ref{ass:correct_specify}, we can shift the focus to maximizing unconditional welfare, even when the planner's ultimate objective function is conditional welfare.
We also note several side-benefits of considering optimal policy in terms of unconditional welfare.
The following proposition directly results from Assumptions \ref{ass:continuous_Markov} (i), Assumption \ref{ass:correct_specify}, and the definition of $G^*_{FB}$.

\begin{prop} \label{first_b_equal}
        Under Assumptions \ref{ass:continuous_Markov} (i) and \ref{ass:correct_specify}, an optimal policy rule $G_* \in \mathcal{G}$ defined by \eqref{maxG} maximizes the conditional welfare function,
    $G_*  \in \mbox{argmax}_{G\in\mathcal{G}}\mathcal{W}_{T}(G|X_{T-1})$.
        Furthermore, if the first best solution belongs to the class of feasible policy rules,  $G^*_{FB}\in\mathcal{G}$, then we have
\begin{eqnarray*}
G^*_{FB} \in \mbox{argmax}_{G\in\mathcal{G}}\mathcal{W}_{T}(G|X_{T-1}).
\end{eqnarray*}
\end{prop}


Having assumed the relationship of optimal policies between the two welfare criteria, we now show how the unconditional welfare function can bound the conditional function. For $G \in \mathcal{G}$ and $X_{T-1}=x$, define conditional regret as
\[
R_{T}(G|x)=\mathcal{W}_{T}(G_{*}|x)-\mathcal{W}_{T}(G|x).
\]
 Note that the unconditional regret can be expressed as an integral of the conditional regret,
\begin{align*}
\mathcal{W}_{T}(G)&=\int\mathcal{W}_{T}(G|x)dF_{X_{T-1}}(x),\\
R_{T}(G)&=\mathcal{W}_{T}(G_*)-\mathcal{W}_{T}(G)=\int R_{T}(G|x)dF_{X_{T-1}}(x).
\end{align*}
For $x' \in \mathcal{X}$, define
\begin{align}
A(x^{\prime},G)&=\{x \in\mathcal{X} : R_{T}(G|x)\geq R_{T}(G|x^{\prime})\}, \notag \\
p_{T-1}(x^{\prime},G)&=\Pr\left(X_{T-1}\in A(x^{\prime},G)\right)=\int_{x\in A(x^{\prime},G)}d F_{X_{T-1}}(x), \notag
\end{align}
and let $x^{obs}$ denote  the observed value of $X_{T-1}$. We assume the following:
\begin{assumption} (Lower bound of conditional density)\label{lower_p_dis_con}
        For  $x^{obs}\in \mathcal{X} $ and any $G\in \mathcal{G}$, there exists a positive constant $\underline{p}$ such that
        \begin{equation}
                p_{T-1}(x^{obs},G)\geq\underline{p}>0.
        \end{equation}
\end{assumption}
\begin{rem}
This assumption is satisfied if $X_{T-1}$ is a discrete random variable taking a finite number of different values. In this case, $p_{T-1}(x^{obs},G)\geq \min_{x\in \mathcal{X}}\Pr\left(X_{T-1}= x\right)>0$, so we can set $\underline{p}=\min_{x\in \mathcal{X}}\Pr\left(X_{T-1}= x\right)$.  If $X_t$ is continuous, then we need to exclude a set of points around the maximum of the function $R_{T}(G|x)$ for the assumption to hold.
Namely, we can assume that we focus on $x$ belonging to a compact subset $\tilde{\mathcal{X}}\subset \mathcal{X}$ such that $\mbox{arg}\max_{x\in\mathcal{X} }R_{T}(G|x)\notin \tilde{\mathcal{X}}$.
If we would like to include the whole support of $X_t$, we can modify the proof by imposing an additional uniform continuity condition on $R_{T}(G|\cdot)$.
\end{rem}
The following lemma provides a bound for conditional regret $ R_{T}(G|x^{obs})$ using  unconditional regret $ R_{T}(G)$.

\begin{lemma} \label{lem:lower_RT(G)} Under Assumptions \ref{lower_p_dis_con},
\begin{align}
    R_{T}(G|x^{obs})  \leq  \frac{1}{\underline{p}} R_{T}(G).
\end{align}
\end{lemma}
The proof of this lemma can be found in Appendix \ref{app:proof_uncon_con_bound}. Using Assumption \ref{lower_p_dis_con} and Lemma \ref{lem:lower_RT(G)}, we proceed to bound the regret for conditional welfare $ \mathcal{W}_{T}(G_{*}|x^{obs})-\mathcal{W}_{T}(\hat{G}|x^{obs})$ by regret for unconditional welfare $\left[\mathcal{W}_{T}(G_{*})-\mathcal{W}_{T}(\hat{G})\right]$, and further by $\left[\mathcal{\widetilde{W}}(G)-\widehat{\mathcal{W}}{}(G)\right]$ (up to constant factors):
\begin{align} \label{welfare_inequality2}
        \mathcal{W}_{T}(G_{*}|x^{obs})-\mathcal{W}_{T}(\hat{G}|x^{obs})
        & \leq\frac{1}{\underline{p}}\left[\mathcal{W}_{T}(G_{*})-\mathcal{W}_{T}(\hat{G})\right]\nonumber\\
        & \leq\frac{c}{\underline{p}}\left[\mathcal{\widetilde{W}}(G_{*})-\mathcal{\widetilde{W}}(\hat{G})\right]
         \leq\frac{2c}{\underline{p}}\sup_{G\in \mathcal{G}}\left|\widehat{\mathcal{W}}(G)-\mathcal{\widetilde{W}}(G)\right|.
\end{align}
The first inequality
follows from Lemma \ref{lem:lower_RT(G)} and Assumption \ref{lower_p_dis_con}.
The second inequality follows from Assumption \ref{equiv_W_uc}.
The last inequality follows from an argument similar to the one below \eqref{ineq_s}. As a result, the bound of the conditional regret can be further established using Theorems \ref{thm:mds_bound} and \ref{mean_bound}.












\begin{theorem} \label{thm:mds_mean_bound}
Under Assumptions \ref{ass:continuous_Markov} to \ref{lower_p_dis_con},
{\small
\begin{equation} \label{eq:main_Theorem}
\underset{P_{T}\in\mathcal{P}_T(M,\kappa)}{\sup} {\mathop{\mbox{\sf E}}}_{P_T}[\mathcal{W}_{T}(G_*|x_{obs})-\mathcal{W}_{T}(\hat{G}|x_{obs})]\lesssim \frac{1}{\underline{p}}\left(C \sqrt{\frac{v}{T-1}} + \frac{c_T [2V(\log T)e\gamma]^{1/\gamma} \sup_{G\in \mathcal{G}} \Phi_{\phi_{\tilde{v}}}(\tilde{S}_{.}(G))}{\sqrt{T-1}}\right).
\end{equation}}
\end{theorem}

\begin{rem}
 [on Assumptions \ref{ass:correct_specify} and \ref{lower_p_dis_con}]
 In this theorem, the conditional regret is compared with the conditional welfare achieved under $G_*$, which is the optimal policy within the class of feasible unconditional decision sets, $\mathcal{G}$. Under Assumption \ref{ass:correct_specify}, $G_*$ can be substituted with $G_*^{FB}$ without altering the regret bound. However, if Assumption \ref{ass:correct_specify} is violated, the regret bound relative to $G_*^{FB}$ becomes
 {\small
 \begin{eqnarray*}&&\underset{P_{T}\in\mathcal{P}_T(M,\kappa)}{\sup} {\mathop{\mbox{\sf E}}}_{P_T}[\mathcal{W}_{T}(G_*^{FB}|x_{obs})-\mathcal{W}_{T}(\hat{G}|x_{obs})]
\\&&\lesssim \frac{1}{\underline{p}}\left(C \sqrt{\frac{v}{T-1}} + \frac{c_T [2V(\log T)e\gamma]^{1/\gamma} \sup_{G\in \mathcal{G}} \Phi_{\phi_{\tilde{v}}}(\tilde{S}_{.}(G))}{\sqrt{T-1}}\right)
+\left[\mathcal{W}_{T}(G_*^{FB}|x_{obs})-\mathcal{W}_{T}(G_*|x_{obs})\right],
\end{eqnarray*}}
 where the last term represents the cost to social welfare incurred by adhering to policy restrictions (such as ethical, moral, or legislative considerations) that are implied by $\mathcal{G}$.

 The factor $\frac{1}{\underline{p}}$ arises from Assumption \ref{lower_p_dis_con}.  A violation of this assumption, i.e., setting $\underline{p}=0$, could result in unbounded conditional regret.  Under Assumption \ref{lower_p_dis_con}, Theorem \ref{thm:mds_mean_bound}  shows that the regret bound of our proposed policy choice converges at a rate of $\frac{1}{\sqrt{T-1}}$.
\end{rem}










\section{Extensions and Discussion}\label{ext}
In this section, we discuss a few possible extensions.
Section \ref{sec_multi} introduces a multi-period policy-making framework.
Section \ref{sec:lucas critique} concerns the ability of the current methods to handle the Lucas critiques. Section \ref{sec:connection} discusses a few connections of T-EWM to the literature on optimal policy choice. More extensions are discussed in Appendix \ref{App_2}.



\subsection{Multi-period welfare} \label{sec_multi}
So far, we have considered only the case of a one-period welfare function.
This subsection discusses how to extend  the current setting to a multiple-period policy framework. In the interest of space, we focus on cases with discrete covariates
and extend the simple model in Section \ref{one-p} to a two-period welfare function. Extending this model beyond two periods is straightforward using similar reasoning.
Recall Assumption  \ref{ass:toy_example}, which imposed Markov properties on the data-generating processes,
\begin{eqnarray*}
Y_{t}(w_{0:t}) &=& Y_t(w_{t-1},w_t),\nonumber\\
Y_{t}(w_{0:t}) &\perp& X_{0:t-1}|W_{t-1},\nonumber\\
W_{t} &\perp& X_{0:t-1}|W_{t-1}.
\end{eqnarray*}

The planner chooses policy rules for two periods, $g_1(\cdot)$ and $g_2(\cdot): \{0,1\}\to  \{0,1\}$, to maximise aggregate welfare over periods $T$ and $T+1$. The decision in the second period, $g_2(.)$, is not contingent on the functional form of $g_1(.)$. By Assumption \ref{ass:toy_example}, the period $T+1$ welfare is determined only by the treatment choice in the period $T$, i.e., $W_T=g_1(W_{T-1})$.
The  two-period welfare function can be written as
\begin{align} \label{two-p-two-g}
	\mathcal{W}_{T:T+1}(g_1(\cdot), g_2(\cdot)|\mathcal{F}_{T-1}) & =\mathcal{W}_{T:T+1}(g_1(\cdot),g_2(\cdot)|W_{T-1}) \nonumber \\
	& =\mathcal{W}_{T}(g_1(\cdot)|W_{T-1})+\mathcal{W}_{T+1}\left(g_2(\cdot)|W_{T}=g_1\left(W_{T-1}\right)\right),
\end{align}
where the last equality follows from Assumption \ref{ass:toy_example} and \eqref{simple_cond_exp}.
To be more specific with the analytical format, we suppress the $(\cdot)$ in $g_i(\cdot)$, when the meaning is clear from the context.
\begin{eqnarray} \label{eq:multi_welfare}
&& \mathcal{W}_{T:T+1}(g_1, g_2|W_{T-1}=w)\nonumber\\
& &=\mathop{\mbox{\sf E}}\left[ Y_{T}(W_{T-1},1)g_{1}(W_{T-1}) +Y_{T}(W_{T-1},0)\left( 1-g_{1}(W_{T-1}) \right)|W_{T-1}=w\right] \nonumber \\
& &+\mathop{\mbox{\sf E}}\left[ Y_{T+1}(W_T,1)g_{2}(W_{T}) +Y_{T+1}(W_T,0)(1-g_{2}(W_{T}) )|W_T=g_{1}(w) \right].
\end{eqnarray}
The second term of \eqref{eq:multi_welfare} follows from
\begin{eqnarray*}
&&\mathop{\mbox{\sf E}}\left[ Y_{T+1}(g_{1}(W_{T-1}) ,1)g_{2}(W_{T}) +Y_{T+1}(g_{1}(W_{T-1}) ,0)(1-g_{2}(W_{T}) )|W_{T-1}=w\right] \\
&&=\mathop{\mbox{\sf E}}\left[ Y_{T+1}(W_T,1)g_{2}(W_{T}) +Y_{T+1}(W_T,0)(1-g_{2}(W_{T}) )|W_T=g_{1}(W_{T-1}) , W_{T-1}=w\right]\\
&&=\mathop{\mbox{\sf E}}\left[ Y_{T+1}(W_T,1)g_{2}(W_{T}) +Y_{T+1}(W_T,0)(1-g_{2}(W_{T}) )|W_T=g_{1}(w) \right],
\end{eqnarray*}
where the last equality follows from Assumption \ref{ass:toy_example}. To estimate the above welfare function, we recall the definition of $T(w)=\#\{1\leq t\leq T-1:W_{t}=w\}$, and we define $T(g_{1}(w) )$ similarly. Then
the empirical analogue of \eqref{two-p-two-g} can be written as,
\begin{align} \label{sample_multi}
	\widehat{\mathcal{W}}_{T:T+1}(g_{1},g_{2}|w)
	= & \frac{1}{T(w)}\sum_{t:W_{t-1}=w}\left\{ \frac{Y_{t}W_{t}g_{1}(W_{t-1}) }{e_{t}(W_{t-1})}+\frac{Y_{t}\left(1-W_{t}\right)\left(1-g_{1}(W_{t-1}) \right)}{1-e_{t}(W_{t-1})}\right\} \nonumber \\
	+ & \frac{1}{T(g_{1}(w))}\sum_{t:W_{t-1}=g_{1}(w)}\left\{ \frac{Y_{t}W_{t}g_{2}(W_{t-1}) }{e_{t}(W_{t-1})}+\frac{Y_{t}\left(1-W_{t}\right)\left(1-g_{2}(W_{t-1}) \right)}{1-e_{t}(W_{t-1})}\right\}.
\end{align}

The maximizer of \eqref{sample_multi}, $\hat{g}_1$, $\hat{g}_2$,  can be obtained by backward induction, a technique widely applied in the Markov decision process (MDP) and dynamic treatment regime literature. See Section \ref{MDP} and Appendix \ref{MDP_A} for more discussion on the relationship between T-EWM and MDP.
To derive the theoretical property of the estimator, we also define
\begin{align*}
	\bar{\mathcal{W}}_{T:T+1}(g_1,g_2|w)
	& =\frac{1}{T(w)}\sum_{t:W_{t-1}=w}\mathop{\mbox{\sf E}}\left[ Y_{t}(1)g_{1}(W_{t-1}) +Y_{t}(0)\left[1-g_{1}(W_{t-1}) \right]|W_{t-1}=w\right] \\
	& +\frac{1}{T(g_{1}(w))}\sum_{t:W_{t-1}=g_{1}(w)}\mathop{\mbox{\sf E}}\left[ Y_{t}(1)g_{2}(W_{t-1}) +Y_{t}(0)\left[1-g_{2}(W_{t-1}) \right]|W_{t-1}=g_{1}(w) \right].
\end{align*}
Similarly to the derivation in the previous sections,   $\widehat{\mathcal{W}}_{T:T+1}(g_1,g_2|w)- \bar{\mathcal{W}}_{T:T+1}(g_1,g_2|w)$ is a (weighted) sum of MDS. Its upper bound can be shown by the method of Section \ref{discrete_bound}.
We show in Appendix \ref{appendmulti} the extension to multi-period welfare with continuous conditioning covariates.

\subsection{Accounting for  Lucas critique} \label{sec:lucas critique}

How can the framework of T-EWM handle the Lucas critique? In this section, we clarify a link between T-EWM and the  SVAR approach based on a three-equation New Keynesian model where a choice of a policy regime can take into account the economy's policy response.


We start with a three-equation New Keynesian model. (See, e.g., Chapter 8 of \cite{walshmonetary}.) At time $t$, let $\pi_{t}$ denote inflation, $x_{t}$ the output gap, and $i_{t}$ the interest rate.
\begin{eqnarray} \label{eq:three_equation}
&\text{Phillips curve: }&\pi_{t}=\beta {\mathop{\mbox{\sf E}}}_{t}\pi_{t+1}+\kappa x_{t}+\varepsilon_{t},\nonumber \\
&\text{IS curve: }&x_{t}={\mathop{\mbox{\sf E}}}_{t}x_{t+1}-\sigma^{-1}(i_{t}-{\mathop{\mbox{\sf E}}}_{t}\pi_{t+1}), \nonumber\\
&\text{Taylor rule: }&i_{t}=\delta\pi_{t}+v_{t},
\end{eqnarray}
Here $v_{t}$ is the monetary policy shock (baseline target rate), which can be viewed as a policy variable that the planner can manipulate. In the process of generating the sample, it is often assumed to follow an AR(1) process $ v_{t}=\rho v_{t-1}+e_{t}$. The AR coefficient of $v_t$, $\rho$, represents a monetary policy regime.
We also assume that the shocks in the Phillips curve follow an AR(1), $\varepsilon_t=\gamma \varepsilon_{t-1}+\delta_t$. Define $
d_{t}=\left( \begin{array}{c}
	v_{t}\\
	\varepsilon_{t}
\end{array}\right) \text{, } F=\left( \begin{array}{cc}
	\rho & 0\\
	0 & \gamma
\end{array}\right) $, and a vector of noises $\eta_t=\left( \begin{array}{c}
e_{t}\\
\delta_{t}
\end{array}\right) $. Then, the process for $d_t$ can be written as $d_t=Fd_{t-1}+\eta_t$.

Define the outcome variables of the system \eqref{eq:three_equation} to be $\tilde{Y}_{t}=\left( \begin{array}{c}
	x_t\\
	\pi_t
\end{array}\right)$. At the end of time $T-1$, the goal of the planner is to minimize (or maximize) the expectation of some function of $\tilde{Y}_{T}$. For example,  an objective function is the welfare cost that penalizes the time $T$ output gap and inflation, $Y_{T}= |x_{T}|^2+|\pi_{T}-\pi_0|^2$, where $\pi_0$ is the inflation target.

Appendix \ref{MDP_B} shows that the
VAR-reduced form of the system \eqref{eq:three_equation} can be expressed as:
\begin{equation}
	\tilde{Y}_{t}=M(\rho)d_{t},\label{eq:Yt_dt2}
\end{equation}
where  $M(\rho)$ is a non-random matrix defined in Appendix \ref{MDP_B}. If the model \eqref{eq:three_equation} is correctly specified, the solution to \eqref{eq:Yt_dt2} takes the Lucas critique into account since it solves for a \emph{deep} parameter $\rho$. A change in $\rho$ incorporates both the direct effect of the policy regime (through $\rho$ in $d_t$ equation) and private agents' anticipation of the policy change ($\rho$ in $M(\rho)$).

Now we show how this is related to the T-EWM framework. The treatment $v_t$ in \eqref{eq:three_equation} corresponds to $W_t$ in the previous sections.
We can write $M(\rho)=\left( \begin{array}{cc}
	m_{11}(\rho) & m_{12}(\rho)\\
	m_{21}(\rho) & m_{22}(\rho)
\end{array}\right) $. When $v_t$ is a binary variable (e.g., high target rate and low target rate), we can define the potential outcomes by setting $v_t = 1$ or $0$ in (\ref{eq:Yt_dt2}): $\tilde{Y}_t(1)=\left( \begin{array}{c} m_{11}(\rho)+m_{12}(\rho) \varepsilon_{t}\\m_{21}(\rho)+m_{22}(\rho) \varepsilon_{t}  \end{array}\right) $ and $\tilde{Y}_t(0)=\left( \begin{array}{c}m_{12}(\rho) \varepsilon_{t}\\m_{22}(\rho) \varepsilon_{t}  \end{array}\right) $. Both $\tilde{Y}_t(1)$ and $\tilde{Y}_t(0)$ depend on $\rho$, so we can write them as $\tilde{Y}_t(1;\rho)$ and $\tilde{Y}_t(0;\rho)$.
Transforming $\tilde{Y}_t=(x_t,\pi_t)^{\prime}$ into $Y_t=|x_t|^2 + |\pi_t - \pi_0|^2$, we can define the potential outcomes for the welfare $(Y_t(1; \rho), Y_t(0; \rho))$, which also depend on $\rho$. The policy choice problem for $v_T \in \{0, 1\}$ can be set to minimize the expected welfare cost:
\begin{equation}
	 \mathop{\mbox{\sf E}} \left[ Y_{T}(1;\rho) g(v_{T-1}) + Y_T(0 ; \rho) (1-g(v_{T-1})) | \mathcal{F}_{T-1} \right]\label{eq:vT_lucas}
\end{equation}
in $g(v_{T-1}) \in \{0, 1\}$, assuming that the one-time policy choice of $v_T$ does not change the value of $\rho$ governing the outcome-generating process (\ref{eq:Yt_dt2}). Thus, one can view the framework of T-EWM in the previous sections concerns a choice of binary policy shock $v_T$, assuming the fixed deep parameter of $\rho$.

In contrast, consider the case where the policy choice of interest concerns regime $\rho$ instead of $v_T$.
We can modify the T-EWM framework as follows in order to handle this case.
Assume that the monetary policy shock $v_t$ is continuously distributed and let $f_{v_{t}|\mathcal{F}_{t-1}}( v;\rho)$ be the conditional density of $v_t$ given the filtration $\mathcal{F}_{t-1}$ evaluated at $v$ and $\rho$. In the outcome process of $t=T-1$ and earlier, the distribution $ f_{v_{t}|\mathcal{F}_{t-1}}( v;\rho)$ at the value of $\rho$ generates the outcomes up to $T-1$. In the counterfactual policy scenario of period $t=T$ and later, the planner manipulates the distribution of $v_T$ by changing the value of $\rho$ instead of directly setting a particular value of $v_T$.
For such an intervention, we specify the planner's problem as a choice of $\rho$ to minimize
\begin{equation}
	\mathop{\mbox{\sf E}} \left[Y_{T}(v_{T};\rho)|\mathcal{F}_{T-1}\right]=\int \mathop{\mbox{\sf E}}\left[Y_{T}(v;\rho)|\mathcal{F}_{T-1}\right]f_{v_{T}|\mathcal{F}_{T-1}}(v;\rho)dv\label{eq:continuous_lucas}.
\end{equation}
This objective function differs from the welfare objective function of  \eqref{eq:vT_lucas} in the following two aspects.
First, in (\ref{eq:vT_lucas}), the policy rule for $v_T$ is deterministic, whereas in (\ref{eq:continuous_lucas}), the planner chooses $\rho$ to change the conditional probability density function of the treatment $v_T$, i.e., a randomized policy for $v_T$.
Second, we see that in (\ref{eq:vT_lucas}), the policy $v_T$ affects the outcome $Y_T$ only through $v_{T}$ with $\rho$ fixed, whereas in (\ref{eq:continuous_lucas}),
the policy $\rho$ affects the outcome through both the distribution of $v_T$
and the deep policy parameter $\rho$ underlying the potential outcomes.

Despite these differences, we can pursue a T-EWM approach for a statistical policy choice of $\rho$ as far as one can construct a sample analogue of the welfare criterion of (\ref{eq:continuous_lucas}).
For instance,  let $\rho_t$ be the policy regime in the sampling period $t=1, \dots, T-1$, and assume $\rho_t$ is observable or estimable.\footnote{$\rho_t$ can be estimated by the method proposed in \cite{schorfheide2005learning}, assuming that monetary policy follows a nominal interest rate rule that is subject to regime shifts.}
For the sake of illustration, we will use a simplified model by  assuming that $\rho_t$ takes values in a finite set $\Omega$ and $v_t\in\{0,1\}$.
 For a given $\rho\in\Omega$, we assume that $\rho_t$ is independent of $Y_{t}(v_t;\rho)$  conditional on $\mathcal{F}_{t-1}$, and we let
 Assumption \ref{ass:toy_example} hold with $W_{t} = v_{t}$, such that $\mathop{\mbox{\sf E}}(\cdot|\mathcal{F}_{t-1})=\mathop{\mbox{\sf E}}(\cdot|v_{t-1})$.
Then, for a policy function $g_{\rho}:\{0,1\}\to \Omega$ and $v\in\{0,1\}$, we can define the empirical welfare estimating\eqref{eq:continuous_lucas} as
 \begin{equation}
\widehat{\mathcal{W}}\left(g_{\rho}|v\right)
		=  T(v)^{-1}\sum_{t:v_{t-1}=v}\left[\sum_{\rho\in\Omega}\frac{Y_{t}\boldsymbol{1}(\rho_t=\rho)}{\Pr[\rho_t=\rho|v_{t-1}]}\boldsymbol{1}[g_{\rho}(v_{t-1})=\rho]\right], \label{eq:emp_w_lucas}
 \end{equation}
where $T(v):=\# \{ v_{t-1}=v\}$. We then maximize this empirical criterion with respect to $v=v_{T-1}$ to estimate an optimal policy regime.


It is worth noting that
as long as $\rho_t$ is observable (or estimable),
the empirical analogue of the welfare criterion is free from functional form restrictions of $Y_t(v; \rho)$ or a distributional restriction of $f_{v_t|\mathcal{F}_{t-1}}(v; \rho)$.
On the other hand, whether the empirical analogue can estimate the welfare criterion (\ref{eq:continuous_lucas}) well or not relies on whether the agents in the economy have correct knowledge of $\rho_t$ and behave in response to the shift of $\rho_t$.
We leave for future research a rigorous characterization of the learnability of an optimal policy regime along this T-EWM approach and its implementability in practice.









\subsection{Connections to other policy choice models in the literature}\label{sec:connection}
In this section, we discuss T-EWM's relation to the literature on treatment and optimal policy analysis.

\subsubsection{Connection to MDP and Reinforcement learning} \label{MDP}

Markov Decision Process (MDP) is the standard framework for reinforcement learning algorithms commonly applied to decision-making problems in dynamic environments. See, e.g.,  \cite{kallenberg2016markov} for a comprehensive introduction.
We can relate the current T-EWM model to a special case of MDP with a finite horizon.
The conditioning variables $X_{t-1}$ correspond to the Markov state at time $t$, and the welfare outcome $Y_{t}$ corresponds to the reward at $t$.
Before the planner intervenes at time $T$, the Markov state transitions follow the data-generating process in which the transitions of the policies are prescribed by the propensity score.
After the planner intervenes, the transition of policies is governed by a deterministic rule described by the (estimated) optimal policy function \eqref{eq:policy_f}.
As we show in Appendix \ref{MDP_A}, the population conditional welfare of T-EMW with a finite-horizon welfare target can be viewed as the value function of a finite-horizon MDP with a non-stationary solution. See Chapter 2 of  \cite{kallenberg2016markov} for more details about finite-horizon MDPs.


The reinforcement learning (RL) literature is vast and provides a rich toolbox to solve MDPs. There are, however, several main differences between T-EWM and RL approaches. First, T-EWM builds on the framework of potential outcome time series which allows for more flexibility in dynamic causal dependence than the reward-generating processes considered in the standard MDPs. Second, the framework of T-EWM can easily accommodate a nonstationary environment in both causal effects and data-generating processes without requiring modeling them explicitly.
The third difference lies in the estimation algorithms. T-EWM obtains a policy by optimizing the empirical welfare criterion, whereas RL typically employs iteration algorithms to maximize the value function.
Fourth, T-EWM performs an optimal policy choice with available data exogenously given to the planner. The RL literature in contrast studies the decision maker's joint strategies of sampling data and learning policies.
That is, the T-EWM method can be regarded as an estimation approach for the optimal policy in an offline and off-policy RL problem with a finite horizon, building on the potential outcome time-series framework.


\subsubsection{Comparison with impulse response functions}

In empirical macroeconomics, a common practice is to measure the causal effects of policies using the impulse response functions (IRF) to the structural policy shocks; see \cite{sims1980macroeconomics}, \cite{ramey2016macroeconomic}, \cite{plagborg2016essays}, \cite{stock2017twenty}, among others. This approach is based on the representation that the endogenous outcome variables are expressed as a weighted sum of the structural shocks, in which the coefficients of the structural shocks correspond to IRFs. With our notation, it can be expressed as
\begin{equation} \label{Structural MA}
Y_t = \sum_{h=0}^{\infty} ( \theta_{h} W_{t-h} + \phi_h \epsilon_{t-h} ),
\end{equation}
where $W_t$ is the policy shock and $\epsilon_t$ is a non-policy structural shock that the planner cannot control. Setting the value of structural shocks exogenously to $w_{-\infty:t}$ without intervening in the joint distribution of the other shocks $\epsilon_{-\infty:t}$ defines the potential outcome $Y_t(w_{-\infty:t})$. The coefficients $(\theta_h, \phi_h)$, $h = 0, 1, 2, \dots,$ are IRFs of the $h$-period ahead outcome variable to the unit changes in policy and non-policy shocks, respectively.

Our approach of T-EWM based on the potential outcome time series differs from the approach of causal IRF analysis in several aspects. First, we assume that the policy shocks of interest $\{W_t: t=0, \dots, T-1 \}$ are directly observable in the available time-series data, while the common framework of IRF analysis such as SVAR considers settings where the policy shocks are not observable and need to be identified through identification of a linear simultaneous equation system. Second, our framework imposes little restrictions on the functional form of how the current and past shocks affect the outcome, while the IRF analysis based on (\ref{Structural MA}) assumes that the structural equation for $Y_t$ is linear in the structural shocks and additive over the horizons. Third, our framework allows unrestricted heterogeneity (i.e., nonstationarity) of the causal effects, whereas the shock-based causal models (\ref{Structural MA}) assume either stationary IRFs that depend only on horizon $h$ or nonstationary IRFs with explicit modeling of how they evolve over time.\footnote{See \cite{bojinov2019time}
and \cite{rambachan2021when} for comparative discussions between the IRFs and the average treatment effects with potential outcome time series.}  Fourth, the existing IRF analysis focuses mainly on estimating and inferring the IRFs, whereas the literature has not formulated how to perform statistical policy choice based on the estimated IRFs. Our T-EWM approach, in contrast, explicitly formulates the policy choice problem using potential outcome time series and analyzes the welfare performance of the statistical policy choice.














\subsection{\label{sec: unknown propensity-score}Estimation with unknown propensity score}
To this point, we have treated the propensity score function as known, but this is infeasible in many applications. Here we consider the case where the propensity score at each time $t$,
$e_{t}(\cdot)$, is unknown. Estimation can be
either parametric or non-parametric. Let $\hat{e}_{t}(\cdot)$ denote the estimator of the propensity score function, and $\hat{G}_{e}$ denote the optimal policy obtained using $\hat{e}_t$.

\subsubsection{The convergence rate with estimated propensity scores}

In this subsection, we adapt Theorem 2.5 of \citet{kitagawa2018should}
to our setting and obtain a new regret bound with estimated propensity scores. We show that, with estimated propensity scores,
the convergence rate is determined by the slower one of the rate of convergence of
$\hat{e}_{t}(\cdot)$ and the rate of convergence of the estimated welfare loss (given
in Theorem \ref{thm:mds_mean_bound}).
\begin{theorem} \label{thm:e_propensity}
	 Let $\hat{e}_{t}(\cdot)$ be an estimated
	propensity score of $e_{t}(\cdot)$, and
		$\hat{\tau}_{t}=\frac{Y_{t}W_{t}}{\hat{e}_t(W_{t-1})}- \frac{Y_{t}(1-W_{t})}{1- \hat{e}_t(W_{t-1})}$
	be a feasible estimator for
$\tau_{t}=\frac{Y_{t}W_{t}}{e_t(W_{t-1})}- \frac{Y_{t}(1-W_{t})}{1- e_t(W_{t-1})}$.
	Given a class of data-generating processes $\mathcal{P}_{T}(M,\kappa)$ defined above equation \eqref{eq:main_Theorem}, we assume
	that there exists a sequence $\phi_{T}\to\infty$ such that the series
	of estimators $\hat{\tau}_{t}$ satisfy
		\begin{equation}
		\text{lim}\underset{T\to\infty}{\text{sup}}\underset{P_T\in\mathcal{P}_{T}(M,\kappa)}{\text{sup}}\phi_{T}{\mathop{\mbox{\sf E}}}_{P_{T}}\left[(T-1)^{-1}\sum_{t=1}^{T-1}|\hat{\tau}_{t}-\tau_{t}|\right]<\infty.
		\label{eq:ass phi}
	\end{equation}

	Then under Assumptions \ref{ass:continuous_Markov} to \ref{D.4}, we have
	\[
	\underset{P_T\in\mathcal{P}_{T}(M,\kappa)}{\text{sup}}{\mathop{\mbox{\sf E}}}_{\mathcal{P}_{T}}[\mathcal{W}(G_*)-\mathcal{W}(\hat{G}_{e})]\lesssim
	(\phi_{T}^{-1}\vee \frac{1}{\sqrt{T-1}}).
	\]

\end{theorem}
To ensure that $\hat{e}_t(W_{t-1})$ is a valid estimator of ${e}_t(W_{t-1})$ in the sense of satisfying (\ref{eq:ass phi}), we need to restrict the dynamics of $W_t$. It is, however, important to note that the condition of (\ref{eq:ass phi}) does not require stationarity of the outcome variable $Y_t(.)$.

		A proof is presented in Appendix \ref{sec:proof_e_propensity}. This theorem shows that if the propensity score is estimated with sufficient accuracy (a rate of $\phi_{T}^{-1} \lesssim \sqrt{T}^{-1}$),  we obtain a similar regret bound to the previous sections. It is not surprising to see that the rate is affected by the
		estimation accuracy of the propensity score, and it is the maximum of $\frac{1}{\sqrt{T-1}}$ and $\phi_{T}^{-1}$.
\begin{rem}
In the cross-sectional setting,  \cite{athey2021policy} show an improved rate of welfare convergence when propensity scores are unknown and estimated. It is possible to extend their analysis to our time-series setting and assess whether or not the rate shown in Theorem \ref{thm:e_propensity} can be improved. However, this is not a trivial extension, and we leave it for further research.  \end{rem}

\subsubsection{Estimation of propensity scores} \label{sec:estimated_prop}
In this subsection, we briefly review various methods of estimating propensity score functions. The propensity score function $e_{t}(\cdot)$ can be estimated parametrically
or nonparametrically.
An example of a parametric estimator is the (ordered) probit model,
which is employed by \cite{hamilton2002model}, \cite{scotti2018bivariate},
and \cite{angrist2018semiparametric}. Under Assumption \ref{ass:continuous_Markov}, $e_{t}$ can be expressed as a function of
$X_{t-1}$. Then, the structure of the propensity score can be given by a probit model,
\begin{align*}
	e_{t}(X_{t-1}) & \equiv P(W_{t}=1|X_{t-1})=\Phi(\beta^{\prime}X_{t-1}),\\
	1-e_{t}(X_{t-1}) & \equiv P(W_{t}=0|X_{t-1})=1-\Phi(\beta^{\prime}X_{t-1}).
\end{align*}

A more complicated structure, such as the dynamic
probit model (\cite{eichengreen1985bank}; \cite{davutyan1995operations}) can also be employed.

We can also use a nonparametric estimator to estimate $e_{t}(\cdot)$.
For example, \cite{frolich2006non} and \cite{park2017nonparametric} extend the
local polynomial regression of \cite{fan1995data} to a dynamic setting. Their methods can be employed here. For simplicity, we assume that the propensity score function is invariant across times, i.e., $e_t(\cdot)=e(\cdot)$ for any $t$, and $e(\cdot)$ is continuous, and we set $\mathcal{X}\subset \mathbb{R}^1$.
For a local polynomial of order $p=1$ and any $x\in\mathcal{X}$, a local likelihood logit model can be specified as
$$\text{log}\left[\frac{e(x)}{1-e(x)}\right]=\alpha_x,$$
for a local parameter $\alpha_x$.
By the continuity of the propensity score function $e(\cdot)$, for
some $X_{t-1}$ close to $x$, we can find a local parameter $\beta_{x}$, such that
$\text{log}\left[\frac{e(X_{t-1})}{1-e(X_{t-1})}\right]\approx\alpha_{x}+\beta_{x}(X_{t-1}-x)$.
The estimated propensity score evaluated at $x$, $\hat{e}(x)$, can be obtained by solving
\begin{align*}
(\hat{\alpha}_{x},\hat{\beta}_{x}) & =\text{argmax}_{\alpha,\beta}\frac{1}{T-1}\sum_{i=1}^{T-1}\bigg\{\bigg[W_{i}\log\left(\frac{\exp\left(\alpha+\beta\left(X_{t-1}-x\right)\right)}{1+\exp\left(\alpha+\beta\left(X_{t-1}-x\right)\right)}\right)\\
 & +(1-W_{i})\log\left(\frac{1}{1+\exp\left(\alpha+\beta\left(X_{t-1}-x\right)\right)}\right)\bigg]K\left(\frac{X_{t-1}-x}{h}\right)\bigg\}.
\end{align*}


where $K(\cdot)$ is a kernel function, and $h$ is the bandwidth. Then, we have
$\hat{e}(x)=\frac{\exp(\hat{\alpha}_{x})}{1+\exp(\hat{\alpha}_{x})}.$




\section{Application} \label{applica}

In the previous pandemic of COVID-19, policymakers
 around the world faced the problem of making effective policies to respond to the Covid-19 pandemic. In this section, we illustrate the usage of T-EWM with an application to choosing the stringency of restrictions imposed in the United States during the pandemic. We let  the treatment $W_{t}$ be a binary indicator for whether the government relaxes restrictions at week $t$. $W_{t}=1$ means the stringency of restrictions is maintained or increased from time $t-1$. $W_{t}=0$ means restrictions are relaxed. The stringency of restrictions is measured by the Oxford Stringency Index,
	which is described in Section \ref{subsec:Data-description}. We assume that a change in the stringency of restrictions at time $t$ will have a lagged effect on deaths, and set
	the outcome variable $Y_{t}$ to be  \emph{$-1\times$two-week-ahead deaths}. The time $t$ information set is,
	\begin{align}
		\ensuremath{}\ensuremath{X_{t}=} & (\text{cases}_{t},\text{deaths}_{t},\text{change in cases}_{t},\text{change in deaths}_{t},\nonumber \\
		& \text{restriction stringency}_{t},\text{vaccine coverage}_{t},\text{economic conditions}_{t}).\label{eq:app_X}
	\end{align}
	The variables in $X_{t}$ are chosen to include the most important factors considered by  policy
	makers when deciding the stringency of restriction.
	The inclusion of the economic conditions reflects policymakers'
    concerns over the economic effect of restrictions. Finally, the propensity score is estimated as a logit model with a linear
index,
\begin{equation}
	\text{log}\left(\frac{\Pr(\text{keeping or increasing restrictions at week }t)}{\Pr(\text{decreasing restriction at week }t)}\right)=\alpha+\beta X_{t-1}.\label{eq:p_score_app}
\end{equation}


\subsection{Data\label{subsec:Data-description}}

The dataset consists of weekly data for the United States. It runs from April 2020 to January 2022 and contains 92 observations. Data on cases, deaths, and vaccine coverage are downloaded from the website
of the Centers for Disease Control and Prevention (CDC). These series are
plotted in Figures \ref{fig:Cases-and-deaths} and \ref{fig:Vaccine} in Appendix \ref{app_app}. We use the Oxford Stringency Index as a measure
of the stringency of restrictions. This index is taken from the Oxford COVID-19 Government
Response Tracker (OxCGRT), which tracks policy interventions and constructs a suite of composite indices that measure
governments' responses.\footnote{The index is calculated based on a dataset
that is updated by a professional team of over one hundred students,
alumni, staff, and project partners. It is a composite index covering
different types of restriction, such as school and workplace
closures, restrictions on the size of gatherings, internal and international
movements, facial covering policies etc. See \cite{hale2020variation} for
more details.} The left panel of Figure \ref{fig:Restriction-econ} in Appendix \ref{app_app} plots the time series of the stringency index.

The economic impact of restrictions is a crucial factor that every policymaker has
to consider. To measure economic conditions, we use the Lewis-Mertens-Stock weekly economic index (WEI). \footnote{The WEI is an index of ten indicators of real economic activity. It represents the common component of series covering consumer behavior, the labor market, and production. See \cite{lewis2021measuring} for more details. }The right panel of Figure \ref{fig:Restriction-econ} in Appendix \ref{app_app} plots the WEI.
In general, none
of these time series seems stationary over the sample period. They exhibit both seasonal fluctuations and changes with respect to different stages of the pandemic.


\subsection{On T-EWM assumptions}\label{checkas}
Before we proceed with the analysis, we discuss the validity of some of the T-EWM assumptions in this context. The Markov properties (Assumption \ref{ass:continuous_Markov}) require that (i) the potential
outcome at time $t$ depends only on the time $t$ and time $t-1$ treatments; (ii) conditional on $X_{t-1}$, the potential outcomes
and treatment at time $t$ are not affected by the path $X_{0:t-2}$.
(i) can be justified here as the treatment is the
\emph{direction of change} in the stringency of restrictions. While the \emph{level}
before $t-1$ may affect the outcome at $t$, its directional
change may not have such a long-existing impact. (ii) requires that conditional on current cases and deaths, cases and deaths
one week ago (as $X_{t}$ includes the first difference of
cases and deaths, it includes the level of cases and deaths from one week ago), current economic
conditions, the stringency of restrictions, vaccine coverage, and deaths in two weeks
will be independent of lags of these variables if the lag is greater than one. For this assumption to hold, it must be that current infections are unrelated to
deaths in more than three weeks. In general, it is unclear
whether this is true, although there is some support
in the literature. For example, based on data collected during the
early stage of the pandemic in China, \cite{verity2020estimates} calculate
that the posterior mean time from infection to death is 17.8 days, with
a 95\%-confidence interval of {[}16.9, 19.2{]}.



Strict overlap (Assumption \ref{bound_c}) requires that, for all $x\in\mathcal{X}$,
the propensity score is strictly larger than 0 and smaller than
1. Figure \ref{fig:PS_app} in Appendix \ref{app_app} shows the histogram of estimated propensity
scores. The strict overlap assumption can be verified by visual inspection: the smallest (estimated)
propensity score is larger than 0.3, and the largest is smaller than
0.9. Sequential unconfoundedness (Assumption  \ref{unconf_c}) requires that conditional
on $X_{t-1}$, treatment assignment at time $t$ is quasi-random. This assumption cannot be tested in general. However, in the past two years, policymakers have had very limited knowledge of Covid-19 beyond the observable data. After controlling for the observables
in (\ref{eq:app_X}), it seems reasonable that the remaining
random factors in two-week-ahead deaths and changes in
the stringency of restrictions are independent.





\subsection{Estimation results and policy recommendations}
In this subsection, we summarize the estimation results and discuss the policy recommendations. We show that T-EWM leads to sensible and robust policy decisions.
After observing $X_{T-1}$, we aim to maximize expected welfare,
$\mathop{\mbox{\sf E}}(-1\cdot\text{deaths}_{T+1})$, over a set of quadrant policies. A class of quadrant treatment rules with $k$ policy variables, $x=(x_{1}\dots x_k)^{\prime}$, is defined as
\begin{equation*}
	\mathcal{G}\equiv\left\{ \begin{array}{c}
		\left\{x:s_{1}(x_{1}-b_{1})>0\;\&\;\dots\;\&\;  s_{k}(x_{k}-b_{k})>0\right\};\\
		s_{1},\dots,s_{k}\in\{-1,1\},b_{1},\dots,b_{k}\in \mathbb{R}
	\end{array}\right\}.
\end{equation*}
Let $X_{T-1}^{P}$
denote a vector of variables for policy choice. This can be any
subvector of $X_{T-1}$.
For our first set of results, we use $X_{T-1}^{P}=\left(\text{change in deaths}_{T-1},\text{restriction stringency}_{T-1}\right).$
Figure \ref{fig:Policy_k=00003D2} presents the estimated treatment region.
The estimated optimal decision rule states that restrictions should not be relaxed (i.e., $W_{T}=1$), if the weekly fall in deaths is below $745$, and the current level of restrictions is lower
than 62.6.

\begin{figure}[h]
	\caption{\label{fig:Policy_k=00003D2} Optimal policy based on $X_{T-1}^{P}=\left(\text{change in cases}_{T-1},\text{restriction stringency}_{T-1}\right)$}

	\centering

	\includegraphics[scale=0.35]{pictures/covid-19_app/policy}
		\noindent\begin{minipage}[t]{1\columnwidth}
		\centering{\footnotesize{}The $x$-axis is the change in deaths at
			$T-1$, and the $y$-axis is the stringency of restrictions  at $T-1$.}
	\end{minipage}
\end{figure}


We can examine the robustness of this result by expanding the set of variables for policy choice. We first add vaccine coverage: $X_{T-1}^{P}=(\text{change in deaths}_{T-1},$ $\text{restriction stringency}_{T-1},$ $\text{vaccine coverage}_{T-1})$.
The optimal estimated quadrant policy is then a 3d-quadrant. Figure
\ref{fig:Policy_k=00003D3}  in Appendix \ref{app_app} shows the projection of this 3d-quadrant
onto the 2d-planes of $(\text{change in deaths}_{T-1},$ $\text{restriction stringency}_{T-1})$
and $(\text{restriction stringency}_{T-1}$,\\  $\text{vaccine coverage}_{T-1}$).
We further expand the set of variables for policy choice by adding the change in cases, so that $X_{T-1}^{P}=(\text{change of deaths}_{T-1},$
$\text{restriction stringency}_{T-1},$ \\$\text{vaccine coverage}_{T-1},\text{ change in cases}_{T-1})$.
The optimal estimated quadrant policy in this case is a 4d-quadrant. Figure
\ref{fig:Policy_k=00003D4} in Appendix \ref{app_app} shows projections of this 4d-quadrant
onto the 2d-planes of $(\text{change in deaths}_{T-1}$, $\text{restriction stringency}_{T-1})$
and $(\text{vaccine coverage}_{T-1},$ $\text{change in cases}_{T-1}$).

Table \ref{tab: policy choice} summarises the estimation results of Figures \ref{fig:Policy_k=00003D2}, \ref{fig:Policy_k=00003D3},
and \ref{fig:Policy_k=00003D4}. It shows that the thresholds of the T-EWM policies are insensitive to the additions of the variables that the quadrant policies depend on. As the number of variables increases, the threshold of the change in weekly deaths
decreases slightly from $-745$ to $-1007.5$, while the other thresholds
remain stable. During the sample period, the mean of weekly cases is 594344.3. Therefore, the threshold in
the last column and the last row corresponds to a $10\%$ fall relative to the
average number of cases during the sample period. In sum, T-EWM suggests that the policymaker should not relax restrictions
($W_{T}=1$) if there are no significant drops in deaths and cases, current stringency is comparatively low, and vaccine coverage is comparatively low.

\begin{table}[H]
	\caption{\label{tab: policy choice}Estimated optimal quadrant rules}

	\centering{}{\small{}}
	\begin{tabular}{cccccc}
		\toprule
		& variables  & change in deaths & restriction & vaccine ($\%$) & change in cases\tabularnewline
		\midrule
		Figure \ref{fig:Policy_k=00003D2} & 2  & $>-745$ & $<62.6$ & -- & --\tabularnewline
		Figure \ref{fig:Policy_k=00003D3} & 3 & $>-1007.5$ & $<62.6$ & $<115.2$ & --\tabularnewline
		Figure \ref{fig:Policy_k=00003D4} & 4 & $>-1007.5$ & $<62.6$ & $<115.2$ & $>-50842.5$\tabularnewline
		\bottomrule
	\end{tabular}{\small\par}
\end{table}

Including more than four policy variables presents a significant computational challenge for the grid-search method we have used to optimize quadrant treatment rules. \footnote{When five variables are included in $X_{T-1}^{P}$, the estimated running time is 12 hours; with six variables, the estimated running time skyrockets to 4459 hours. (These estimates were made on a computer with a 12th Gen Intel(R) Core(TM) i7-1265U 2.7 GHz processor and 32GB of RAM.)} To overcome this issue, we employ decision trees to search for an optimal policy based on time-series empirical welfare.

The decision tree approach (\cite{breiman1984classification}) sequentially searches for policy variables and their threshold-based splits to maximize welfare. \cite{athey2021policy} and \cite{zhou2023offline} study the properties and implementation of this approach in maximizing the doubly robust empirical welfare criterion with brute force searches for tree partitions. \cite{ida2022choosing} implement a decision tree with a heuristic two-step optimization in the context of estimating a rebate assignment policy in an electricity market. We implement a classification tree algorithm based on \cite{hastie2009elements}. See Appendix \ref{app:emp_tree_algorithm} for the details of our implementation.




Figure  \ref{T-EWM tree} illustrates a policy rule obtained by a T-EWM decision tree, which considers all the seven policy variables within $X_{T-1}$, i.e., $ X_{T-1}^{P}=X_{T-1}=(\text{cases}_{T-1},\,\text{deaths}_{T-1},$  $\text{change in cases}_{T-1},\,\text{change in deaths}_{T-1},\,\text{restriction stringency}_{T-1},\,\text{vaccine coverage}_{T-1},$\\  $\text{economic conditions}_{T-1}).$ Due to the relatively small sample size of 92, we only generate trees with a depth of three.
\begin{figure}[H]
	\caption{\label{T-EWM tree} T-EWM decision tree with seven policy variables}

	\centering

	\includegraphics[width=85mm]{pictures/covid-19_app/tree_q1_p7}
	\noindent\begin{minipage}[t]{1\columnwidth}
	\end{minipage}
\end{figure}

The result from the  decision tree indicates that if the economy performs relatively well, policymakers should adopt a more aggressive approach in maintaining or tightening Covid restrictions.
(Note that the decision given by the left node of the second level aligns with the last one in the last row of Table \ref{tab: policy choice}, which considers four policy variables.) Conversely, if the economy performs relatively poorly, policymakers should refrain from further tightening Covid restrictions unless the increase in cases is significant.

 To wrap up, we illustrate the use and implementation of the proposed T-EWM approach in making policy decisions. Furthermore, utilizing decision trees with the T-EWM approach allows for the inclusion of more policy variables while keeping computation time manageable.  More results from the T-EWM decision trees including a specification with a higher order of the Markovian structure of $q=2$ can be found in Appendix \ref{app_app}.

\section{Conclusion}
This article proposes T-EWM, a framework and method for choosing optimal policies based on time-series data. We characterise assumptions under which this method can learn an optimal policy. We evaluate its statistical properties by deriving non-asymptotic upper and lower bounds of the unconditional and conditional welfare. We discuss its connections to the existing literature, including the Markov decision process and impulse response analysis. We present simulation results and empirical applications to illustrate the computational feasibility and applicability of T-EWM. As a benchmark formulation, this paper mainly focuses on a one-period social welfare function as a planner's objective. Extensions of the analysis to policy choices for the middle-run and long-run cumulative social welfare are left for future research.

\begin{spacing}{1.0}
\bibliographystyle{abbrvnat}
\bibliography{biball}
\end{spacing}
\newpage