EconBase
← Back to paper

A Semiparametric Instrumented Difference-in-Differences Approach to Policy Learning

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

48,426 characters · 10 sections · 67 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Semiparametric Instrumented Difference-in-Differences Approach to Policy Learning

abstractRecently, there has been a surge in methodological development for the difference-in-differences (DiD) approach to evaluate causal effects. Standard methods in the literature rely on the parallel trends assumption to identify the average treatment effect on the treated. However, the parallel trends assumption may be violated in the presence of unmeasured confounding, and the average treatment effect on the treated may not be useful in learning a treatment assignment policy for the entire population. In this article, we propose a general instrumented DiD approach for learning the optimal treatment policy. Specifically, we establish identification results using a binary instrumental variable (IV) when the parallel trends assumption fails to hold. Additionally, we construct a Wald estimator, novel inverse probability weighting (IPW) estimators, and a class of semiparametric efficient and multiply robust estimators, with theoretical guarantees on consistency and asymptotic normality, even when relying on flexible machine learning algorithms for nuisance parameters estimation. Furthermore, we extend the instrumented DiD to the panel data setting. We evaluate our methods in extensive simulations and a real data application.

{\bf Keywords:} individualized treatment rule, instrumental variable, multiple robustness, semiparametric efficiency, unmeasured confounding

Introduction

Data-driven individualized decision making has received increasing interests in many fields, such as precision medicine luedtke2016statistical,tsiatis2019dynamic, econometrics and quantitative social sciences imai2004causal,athey2021policy, computer science and operations research shi2022off,kallus2022doubly. The common goal is to learn optimal treatment assignment policies (also known as regimes, rules or plans) which map individual characteristics to treatment assignments so as to optimize some functional of the counterfactual outcome distributions, leveraging observational data where causal effects can be identified under various strategies and assumptions.

Popular existing methods in the statistical and machine learning literature include model-based approaches such as Q-learning watkins1992q,murphy2003optimal,linn2017interactive, A-learning robins2000marginal,shi2018high, and direct model-free policy search approaches zhang2012estimating,zhao2012estimating. Recent advances of policy learning have also considered a variety of data structures, optimization objectives, criteria or constraints, such as survival and longitudinal data goldberg2012q,ertefaie2018constructing,zhao2023efficient, networks viviano2019policy,sherman2020general, distributional robustness mo2021learning,sahoo2022learning, budget, fairness, or interpretability constraints luedtke2016optimal,fang2022fairness, among others luedtke2020performance,hadad2021confidence,nie2021learning,hu2022fast,jin2023sensitivity.

With few exceptions, most methods in prior work rely on the pivotal assumption that there is no unmeasured confounding. This is a key threat to credible causal inference in observational studies, and may lead to suboptimal policies, because this assumption is impossible to verify or test in practice. An ad hoc work-around commonly adopted by practitioners is to collect and appropriately adjust for a large number of covariates, which still lacks theoretical guarantee and seems likely to be error-prone. To address this limitation, there has been recent progress made in several directions. kallus2018confounding propose to minimize the worst-case regret of a policy under a marginal sensitivity model for the unmeasured confounding. zhang2021selecting utilize a randomization test to rank by a partial order and select treatment rules within a given finite collection. While partial identification results provide certain improvement, the performance of such a learned policy may still be suboptimal. qi2023proximal build on the semiparametric proximal causal inference framework introduced by cui2023semiparametric to establish point identification results on different policy classes and accordingly propose several classification-based approaches; but this framework requires the analyst to correctly classify the measured covariates into three types of proxies, and it may be difficult to estimate the confounding bridge functions.

Instrumental variable methods are widely used to handle unmeasured confounding in observational studies or randomized trials with non-compliance. The core requirements for a pretreatment variable to be a valid IV are: (i) it is associated with the treatment; (ii) it is independent of all unmeasured confounders; (iii) it does not have a direct causal effect on the outcome other than through the treatment. Along with the seminal work of imbens1994identification,angrist1996identification, extensive development has been made in using the IV to estimate the local average treatment effect tan2006regression,ogburn2015doubly, defined as the average treatment effect for the complier subgroup who would always comply with their treatment assignments. Since the complier subgroup is unknown and may have systematically different characteristics from the population, the population (conditional) average treatment effect is arguably the causal parameter of primary interest in most studies hernan2006instruments,aronow2013beyond, especially for policy learning. More recently, pu2021estimating consider a partial identification approach to optimal treatment rule estimation; and wang2018bounded formally establish point identification of the population average treatment effect under alternative no-interaction assumptions, upon which cui2021semiparametric propose various IV methods for estimating optimal treatment regimes. It is notable that all of these IV methods in the literature only consider the setting with a single time point, with the only exception of pmlr-v202-xu23x, where the authors propose an IV approach to off-policy evaluation in confounded Markov decision processes with infinite horizons.

There has always been interest in exploiting the longitudinal structure common in datasets such as electronic health records and medical claims in epidemiology and biomedicine robins2000marginal, as well as cross-sectional or panel data in program evaluations, economic censuses, and surveys athey2017state. DiD methods have been an important tool widely used by empirical researchers card1994minimum. The key identification assumption of DiD is that the trend in outcome of the control group over time is informative about what the trend would have been for the treatment group in the absence of the treatment. Specifically, under the standard (conditional) parallel trends assumption, which states that the (conditional) expected trends in the potential outcomes of the two groups in the absence of the treatment are identical, the average treatment effect on the treated can be identified abadie2005semiparametric,sant2020doubly; we refer interested readers to lechner2011estimation and roth2023s for detailed reviews. However, concerns often arise that the parallel trends assumption may be violated due to unmeasured confounding. athey2006identification develop a new changes-in-changes model that relates outcomes to an individual's group, time, and unobservable characteristics; and various recent extensions for DiD include partial identification ye2020negative, sensitivity analysis keele2019patterns and negative control sofer2016negative, among others dukes2022semiparametric,park2023universal. Moreover, DiD methods focus on the identification and estimation of the average treatment effect on the treated, which limits its application in policy learning since the treated cannot represent the population. To the best of our knowledge, this is the first work to systematically study policy learning under the DiD setting.

In this article, we combine the two natural experiments and propose an instrumented DiD approach to policy learning when the parallel trends assumption fails to hold in the presence of unmeasured confounding. Specifically, we adapt and extend the recent progress in ye2022instrumented and vo2022structural, relaxing some key assumptions of the conventional IV and DiD methods. We allow for the violation of the parallel trends assumption by leveraging an IV which has no direct effect on the the trend in outcome, and does not modify the average treatment effect. Notably, this exogenous variable is not necessarily a valid instrument for the conventional treatment-outcome association, since we allow it to have a direct effect on the outcome not just through the treatment at each time point.

The contributions of this article are summarized as follows. First, we propose the direct policy search approach to learn optimal treatment assignment policies, based on the conditional average treatment effect estimators using instrumented DiD. This approach essentially allows us to learn the optimal policy that maximizes the estimated value within a restricted policy class. Second, we establish novel identification results of optimal policies for the instrumented DiD design subject to unmeasured confounding. The new results give rise to new inverse probability weighting estimators of optimal policies without necessarily identifying the value function for a given policy. Another interesting progress is also made towards identifying optimal policies without necessarily using the subjects' realized treatment values. In summary, we construct a Wald estimator and novel inverse probability weighting estimators. A class of semiparametric efficient and multiply robust estimators is also proposed, which is consistent provided that a subset of several posited models indexing the observed data distribution is correctly specified. Third, we prove theoretical guarantees for the proposed multiply robust policy learning approaches. Specifically, we consider both parametric models and flexible data-adaptive machine learning algorithms with the cross-fitting procedure to estimate the nuisance parameters, to draw valid inferences under mild regularity conditions and certain rate of convergence conditions. In particular, we consider a restricted policy class indexed by an Euclidean parameter $\eta$ and establish the $n^{-1/3}$ convergence rate of $\hat{\eta}$, even though its resultant limiting distribution is not standard. Fourth, we extend our proposed methods to the panel data setup. We establish identification of the conditional average treatment effect under alternative assumptions and provide the direct policy search approaches for panel data. The theoretical results for panel data can be similarly derived.

The rest of this article is organized as follows. In Section (ref), we introduce the statistical framework of instrumental variable, DiD and policy learning. Section (ref) develops our main methodology of learning the optimal policy using the instrumented DiD. Semiparametric efficiency results and multiply robust estimators are presented in Section (ref). Section (ref) establishes the asymptotic properties of the proposed estimators. Extensive simulations are reported in Section (ref) to demonstrate the proposed methods, followed by a real data application in Section (ref). Next, we consider the extension of our methods to panel data in Section (ref). The article concludes in Section (ref) with a discussion of some remarks and future work. All proofs and additional results are provided in the Supplementary Material.

Statistical framework

We first introduce some notation. Let $X$ denote the $p$-dimensional vector of covariates that belongs to a covariate space $\mathcal{X} \subset \mathbb{R}^p$, $A \in \mathcal{A} = \{0, 1\}$ denote the binary treatment, $Y \in \mathbb{R}$ denote the outcome of interest, and $T \in \mathcal{T} = \{0, 1\}$ denote the time period. Suppose that $U = (U_{0}, U_{1})$ is an unmeasured confounder of the effect of $A$ on $Y$, and $Z \in \{0, 1\}$ is a binary instrumental variable; the observed data are $O = (X, A, Y, T, Z)$. We assume that the random samples $(O_{1}, \ldots, O_{n})$ collected at the two time periods are independent and identically distributed (i.i.d.) observations of $O \sim P_{0}$, and there is no overlap between individuals in these two time periods. This setup is commonly known as the repeated cross-sectional data. Extension to panel data setting is studied in Section (ref).

We use the potential outcomes framework neyman1923applications,rubin1974estimating to define causal effects. Let $A_{t}(z)$ denote the potential exposure at time $t$ if the instrument were set to level $z$, $Y_{t}(a)$ denote the potential outcome at time $t$ if the exposure were set to level $a$ and the instrument would take the same value it actually had, and $Y_{t}(z, a)$ denote the potential outcome at time $t$ had the instrument and exposure been set to $z, a$ respectively.

Without loss of generality, we assume that larger values of $Y$ are more desirable. Our aim is to identify and estimate an policy $d: \mathcal{X} \to \mathcal{A}$, that maximizes the expected potential outcome in a counterfactual world had this policy been implemented on the population. The optimal policy at time $t$ is given by $d_{{\rm opt},t} (x) = I\{\tau_{t} (x) > 0\}$, where $\tau_{t} (x) = E[Y_{t}(1) - Y_{t}(0) \mid X = x]$ is the conditional average treatment effect (CATE) at time $t$.

Let $Y_{t}(d) = d(X) Y_{t}(1) + (1 - d(X)) Y_{t}(0)$ denote the potential outcome under a hypothetical intervention that assigns treatment according to policy $d$. The value function of a policy $d$ at time $t$ is defined as $V_{t}(d) = E[Y_{t}(d)]$. Let $\mathcal{D}$ be the class of candidate policies of primary interest. The optimal policy can be obtained by directly maximizing the value function:

equation[equation omitted — 145 chars of source]

Throughout this article, we assume that the stable treatment effect over time assumption holds, which says that the CATE does not vary over time, and thus ensures that the optimal policy remains the same between the two time periods. The subscript $t$ is omitted when it is clear from the context.

remarkOur proposed instrumented DiD methodology can also be readily formulated in the weighted classification perspective. Pioneered by zhang2012estimating, this perspective has been widely used in the biostatistics and precision medicine literature, and enjoys certain robustness empirically. Specifically, the above maximization problem (ref) can be transformed into the following equivalent weighted classification problem: \begin{equation} d_{\rm{opt}} (x) = \arg\max_{d \in \mathcal{D}} E[W I\{A = d(X)\}], \end{equation} where $W$ is regarded as a weight that is motivated by standard outcome regression, inverse probability weighting and doubly robust methods. Many robust classification methods and off-the-shelf implementations can be utilized.

Instrumented difference-in-differences

In this section, we introduce a general instrumented DiD framework for policy learning under endogeneity, and provide novel identification results. Let $\pi (t, z, x) = Pr(T = t, Z = z \mid X = x)$, and for any random variable $C \in \{A, Y\}$, we define $\mu_{C}(t, z, x) = E[C \mid T = t, Z = z, X = x]$, $\delta_{C}(x) = \mu_{C}(1, 1, x) - \mu_{C}(0, 1, x) - \mu_{C}(1, 0, x) + \mu_{C}(0, 0, x)$. We make the following identification assumptions.

assumption[Consistency] $A = A_{T}(Z)$ and $Y = Y_{T}(A)$.
assumption[Positivity] $c_1 < \pi (t, z, x) < 1 - c_1$ for some $0< c_1 < 1/2$.
assumption[Random sampling] $T \perp \{A_{t}(z), Y_{t}(a) : t = 0, 1, z = 0, 1, a = 0, 1\} \,|\, X, Z$.
assumption[Stable treatment effect over time] $E[Y_{0}(1) - Y_{0}(0) \mid X] = E[Y_{1}(1) - Y_{1}(0) \mid X]$.

Assumption (ref) is also known as the stable unit treatment value assumption, which states that there is no interference between subjects and no multiple versions of the instrument and treatment. Assumption (ref) ensures the same support of $X$ for each $(T, Z)$ level. Assumption (ref) is commonly assumed for repeated cross-sectional data abadie2005semiparametric. Assumption (ref) requires that the CATE $\tau (x)$ does not vary over time, and thus ensures that the optimal policy remains the same between the two time periods.

assumption[Trend relevance] $E[A_{1}(1) - A_{0}(1) \mid Z = 1, X] \neq E[A_{1}(0) - A_{0}(0) \mid Z = 0, X]$.
assumption[Independence & exclusion restriction] $Z \perp \{A_{t}(1), A_{t}(0), Y_{t}(1) - Y_{t}(0), Y_{1}(0) - Y_{0}(0) : t = 0, 1\} \,|\, X$.
assumption[No unmeasured common effect modifier] $Cov\{A_{t}(1) - A_{t}(0), Y_{t}(1) - Y_{t}(0) \mid X\} = 0$ for $t = 0, 1$.

Assumption (ref) and (ref) are parallel to the core assumptions in the standard IV literature. Directed acyclic graphs illustrating the causal structure are provided in Section (ref) of the Supplementary Material. Assumption (ref) states that the IV affects the trend in treatment. Assumption (ref) requires that the IV is unconfounded, has no direct effect on the trend in outcome, and does not modify the treatment effect. This exogenous variable is not necessarily a valid instrument for the conventional treatment-outcome association, since we allow it to have a direct effect on the outcome not just through the treatment at each time point. Assumption (ref) essentially states that there is no common effect modifier by an unmeasured confounder, of the additive effect of treatment on the outcome, and the additive effect of the IV on treatment. It has been studied in cui2021semiparametric, and relax certain no additive interaction assumptions in wang2018bounded. We refer interested readers to ye2022instrumented for detailed discussion and concrete examples of an IV for DiD. Now we present our first identification result under the above assumptions.

theoremUnder Assumptions (ref)-(ref), the optimal policy is nonparametrically identified by \begin{equation} \arg\max_{d \in \mathcal{D}} E\left[\frac{\delta_{Y}(X)}{\delta_{A}(X)} d(X)\right]. \end{equation}

Theorem (ref) combines the Wald estimator for CATE and the direct policy search approach in Equation (ref). Similarly, the IPW estimator proposed by ye2022instrumented can also be used to learn the optimal policy. Semiparametric efficient and multiply robust estimators are presented in Section (ref). Next we propose our novel identification results, which also serves as basis for the estimators proposed in Section (ref).

theoremUnder Assumptions (ref)-(ref), the optimal policy is nonparametrically identified by \begin{equation} \arg\max_{d \in \mathcal{D}} E\left[\frac{(2Z - 1)(2T - 1)(2A - 1) Y I\{A = d(X)\}}{\pi(T, Z, X) \delta_{A}(X)}\right]. \end{equation}

Theorem (ref) extends prior identification of CATE, and proposes a novel IPW estimator of the optimal policy without necessarily identifying the value function. Semiparametric efficiency results based on (ref) are given in Section (ref) and (ref) of the Supplementary Material.

theoremUnder Assumptions (ref)-(ref), the optimal policy is nonparametrically identified by \begin{equation} \arg\max_{d \in \mathcal{D}} \, E\left[\frac{(2T - 1) Y I\{Z = d(X)\}}{\pi(T, Z, X) \delta_{A}(X)}\right]. \end{equation}

Theorem (ref) essentially proves that we can identify the optimal policy without necessarily using the subjects' realized treatment values, for instance when $\delta_{A}(X)$ is known a priori, or when a separate sample with data on $(A, X, T, Z)$ is available to estimate $\delta_{A}(X)$. To conclude this section, we propose the following estimators for optimal policies:

align*[align* omitted — 508 chars of source]

where $\hat{\delta}_{Y}$, $\hat{\delta}_{A}$ and $\hat{\pi}$ are estimated by parametric models or machine learning algorithms. Our simulation studies in Section (ref) empirically shows comparable performance of the IPW estimators (ref) and (ref).

remarkSimilarly, classification-based estimators based on Theorem (ref) and (ref) can be proposed: \begin{equation} \arg\max_{d \in \mathcal{D}} E[\Tilde{W}_{1} I\{A = d(X)\}], \quad \arg\max_{d \in \mathcal{D}} E[\Tilde{W}_{2} I\{Z = d(X)\}], \end{equation} respectively, where the weights are given by \begin{equation*} \Tilde{W}_{1} = \frac{(2Z - 1)(2T - 1)(2A - 1) Y}{\pi(T, Z, X) \delta_{A}(X)}, \quad \Tilde{W}_{2} = \frac{(2T - 1) Y}{\pi(T, Z, X) \delta_{A}(X)}. \end{equation*} The Fisher consistency, excess risk bound and universal consistency of the estimated policy can also be established zhao2012estimating.

Semiparametric efficiency and multiply robust estimators

In this section, we use semiparametric theory and propose multiply robust estimators. The Wald and the IPW approaches require the corresponding models to be correctly specified. Hence, methods that are robust against model misspecification are highly desired, where consistency is guaranteed when a subset of several posited models indexing the observed data distribution is correctly specified.

We consider the (uncentered) efficient influence function:

equation*[equation* omitted — 215 chars of source]

which has been proposed in ye2022instrumented. Therefore, the optimal policy is identified by $\arg\max_{\mathcal{D}} E\left[\Delta (X) d(X)\right]$. Moreover, in light of the optimization tasks formulated in (ref), we propose the following two choices of statistic:

equation*[equation* omitted — 226 chars of source]

and

equation*[equation* omitted — 206 chars of source]

which also enjoy the multiply robustness property.

First, we consider positing parametric models. Let $\mu_{A}(t,z,x;\alpha)$, $\mu_{Y}(t,z,x;\beta)$ and $\pi(t, z, x;\theta)$ denote the posited models. $\hat{\alpha}$, $\hat{\beta}$ and $\hat{\theta}$ can be estimated by maximum likelihood estimation. In Theorem (ref), we show the multiple robustness in the sense of maximizing the objective function (or minimizing the weighted classification error) in the union model of the following models:\\ $\mathcal{M}_1$: models for $\pi(t, z, x)$ and $\delta_{A}(x)$ are correct; \\ $\mathcal{M}_2$: models for $\pi(t, z, x)$ and $\delta_{Y}(x) / \delta_{A}(x)$ are correct; \\ $\mathcal{M}_3$: models for $\delta_{Y}(x) / \delta_{A}(x)$ and $\mu_{C}(0, 0, x)$, $\mu_{C}(1, 0, x)$, $\mu_{C}(0, 1, x)$ for $C \in \{A, Y\}$ are correct.

theoremUnder Assumptions (ref)-(ref), the optimal policy is identified by \begin{equation} \arg\max_{\mathcal{D}} \, E\left[W_{1} I\{A = d(X)\}\right] = \arg\max_{\mathcal{D}} \, E\left[W_{2} I\{Z = d(X)\}\right] = \arg\max_{\mathcal{D}} \, E\left[\Delta (X) d(X)\right], \end{equation} under the union model $\mathcal{M}_1 \cup \mathcal{M}_2 \cup \mathcal{M}_3$.

We also consider using modern machine learning methods to estimate these nuisance parameters. In practice, we apply the cross-fitting technique schick1986asymptotically,zheng2010asymptotic,chernozhukov2018double, which is easy to implement. The cross-fitting procedure goes as follows. We randomly split data into $K$ folds; the cross-fitted estimator is given by

equation*[equation* omitted — 139 chars of source]

where $P_{n,k}$ denote empirical averages only over the $k$-th fold, and $\hat{\mu}_{A,-k}$, $\hat{\mu}_{Y,-k}$ and $\hat{\pi}_{-k}$ denote the nuisance estimators constructed excluding the $k$-th fold. Similar cross-fitted estimators for $E\left[W_{1} I\{A = d(X)\}\right]$ and $E\left[W_{2} I\{Z = d(X)\}\right]$ can also be constructed in the same way.

Asymptotic analysis of policy learning

In this section, we study theoretical guarantees for our proposed policy learning approaches. While researchers have suggested applying machine learning algorithms to estimate the optimal policies from large classes which cannot be described by a finite dimensional parameter luedtke2016statistical,kunzel2019metalearners, it is also important to consider certain classes of policies for better interpretability and transparency, especially in clinical medicine and policy research zhang2015using,athey2021policy. Specifically, here we focus on a class of feasible policies $\mathcal{D} = \left\{I\{\eta^\top X > 0\} : \eta \in \mathbb{H}\right\}$, where $\eta$ indexes different policies and $\mathbb{H}$ is a compact subset of $\mathbb{R}^{p}$. That is, we analyze the following estimator:

equation*[equation* omitted — 166 chars of source]

where $\hat{M}(\eta)$ is estimated by posited parametric models, or the cross-fitted estimator. Let $\eta^\ast = \arg\max_{\eta \in \mathbb{H}} E[\Delta (X) d(X;\eta)]$ denote the Euclidean parameter that indexes the optimal policy. We detail the main large sample property of our proposed estimator, that $\hat{\eta}$ converges to $\eta^\ast$ at $n^{1/3}$ rate, and that $\hat{M}(\hat{\eta})$ is $n^{1/2}$-consistent and asymptotically normal under weak conditions (mostly requiring standard regularity conditions white1982maximum, or only that the nuisance parameters are estimated at faster than $n^{1/4}$ rates).

remarkIn order to obtain certain rates of convergence or regret bounds, it is necessary to require some control over the complexity of the class $\mathcal{D}$; see athey2021policy for examples of the VC-dimension of classes of linear rules, decision trees and monotone rules. Here we apply the empirical process techniques to establish theoretical guarantees for linear rules, which also hold on any other $\mathcal{D}$ indexed by finite-dimensional parameters. Also note that all identification and semiparametric efficiency results hold for any class of policies, and other optimization methods can be readily utilized.

We assume the following regularity conditions.

condition{\rm (i)} The supports of $X$ and $Y$ are bounded. {\rm (ii)} The functions $\mu_{Y}(t,z,x)$, $\mu_{A}(t,z,x)$ and $\pi(t, z, x)$ are smooth and bounded for all $(t,z,x)$. {\rm (iii)} The function $M(\eta)$ is twice continuously differentiable in a neighborhood of $\eta^\ast$; {\rm (iv)} For all $\delta > 0$, we have that $Pr(|X^{T} \eta^\ast| \leq \delta) \leq c_2 \delta$, for some constant $c_2 > 0$ such that $c_2 \delta \leq 1$.
condition{\rm (i)} $\sqrt{n} (\hat{\alpha} - \alpha^\ast) = O_{p}(1)$; {\rm (ii)} $\sqrt{n} (\hat{\beta} - \beta^\ast) = O_{p}(1)$; {\rm (iii)} $\sqrt{n} (\hat{\theta} - \theta^\ast) = O_{p}(1)$.
theoremUnder Assumptions (ref)-(ref), if Conditions (ref) and (ref) hold, we have {\rm (i)} $\|\hat{\eta} - \eta^\ast\|_{2} = O_{p}(n^{-1/3})$; {\rm (ii)} $\sqrt{n} \{M(\hat{\eta}) - M(\eta^\ast)\} = o_p(1)$; {\rm (iii)} $\sqrt{n} \{\hat{M}(\hat{\eta}) - M(\eta^\ast)\} \to \mathcal{N}(0, \sigma_1^2)$, where $\sigma_1^2$ is given in the Supplementary Material.

Condition (ref) (i), (ii) and (iii) are standard regularity conditions to establish uniform convergence. Condition (ref) (iv), also known as the margin condition, is often assumed in the literature of classification tsybakov2004optimal, reinforcement learning hu2022fast and treatment assignment policies luedtke2020performance, to guarantee fast convergence rates. Condition (ref) requires $\sqrt{n}$ convergence rates of parameter estimates of the posited models, which holds under mild conditions.

We assume the following conditions for the machine learning algorithms used to construct cross-fitted estimators.

condition$\|\hat{\mu}_{A}(t,z,X) - \mu_{A}(t,z,X)\|_{L_2} = o_{p}(n^{-1/4})$, $\|\hat{\mu}_{Y}(t,z,X) - \mu_{Y}(t,z,X)\|_{L_2} = o_{p}(n^{-1/4})$ and $\|\hat{\pi}(t, z, X) - \pi(t, z, X)\|_{L_2} = o_{p}(n^{-1/4})$, for $t,z = 0,1$.
theoremUnder Assumptions (ref)-(ref), if Conditions (ref) and (ref) hold, we have {\rm (i)} $\|\hat{\eta} - \eta^\ast\|_{2} = O_{p}(n^{-1/3})$; {\rm (ii)} $\sqrt{n} \{M(\hat{\eta}) - M(\eta^\ast)\} = o_p(1)$; {\rm (iii)} $\sqrt{n} \{\hat{M}(\hat{\eta}) - M(\eta^\ast)\} \to \mathcal{N}(0, \sigma_2^2)$, where $\sigma_2^2$ is given in the Supplementary Material.

Condition (ref) says the nuisance estimators must be consistent and converge at a fast enough rate (essentially $n^{1/4}$ in $L_2$ norm). This is quite general and can be achieved by many existing algorithms under nonparametric smoothness, sparsity, or other structural constraints. According to Theorems (ref) and (ref) (ii), the regret of our estimated regime vanishes as the sample size increases. Theorems (ref) and (ref) (iii) imply that $\hat M(\hat \eta)$ is a regular and asymptotic normal estimator of $M(\eta^\ast)$.

Simulations

In this section, we conduct extensive simulations to evaluate the finite-sample performance of the proposed estimators. Specifically, we compare them to the instrumental variable approach proposed by cui2021semiparametric, which is in principle valid only for a single time point. Replication code is available at \href{https://github.com/panzhaooo/policy-learning-instrumented-DiD}{GitHub}.

We first describe the complete data generation process as follows. Baseline covariates $X = (X_{1}, X_{2})^\top$ are generated from independent standard normal distributions. The time period indicator $T$ is generated from a Bernoulli distribution with probability $0.5$. The unmeasured confounders $U = (U_{0}, U_{1})^\top$ are generated from independent bridge distributions with parameter $0.5$ \footnote{The bridge density function is $p(u) = 1 / (2\pi\cosh{(u/2)})$. We use the bridge distribution because by wang2003matching, the data generation process ensures that upon marginalizing over $U$, the model for $Pr(A_{t} = 1 \mid Z, X)$ remains a logistic regression.}. The instrumental variable $Z$ is generated from a Bernoulli distribution with probability $0.5$. The potential treatments and outcomes at time points $t = 0, 1$ are generated from the models:

align*[align* omitted — 301 chars of source]

where $\mu_{0} = 200 + 10 (A_{0} (1.5 X_{1} + 2 X_{2} - 0.5) + 0.5 U_{0} + 2 Z + 1.5 X_{1} + 2 X_{2})$, and $\mu_{1} = 240 + 10 (A_{1} (1.5 X_{1} + 2 X_{2} - 0.5) + 0.5 U_{1} + 2 Z + 2 X_{1} + 1.5 X_{2})$. Therefore, the optimal policy is $d_{\rm{opt}} (x) = I\{3 x_{1} + 4 x_{2} - 1 > 0\}$. Let $A = T A_{1} + (1 - T) A_{0}, Y = T Y_{1} + (1 - T) Y_{0}$; thus the observed cross-sectional data are $(X, A, Y, T, Z)$.

A large test dataset of size $N = 1 \times 10^{6}$ is generated independently to evaluate the performance of different estimators. The percentage of correct decisions (PCD) of an estimated policy $\hat{d}(x)$ is computed by $1 - N^{-1} \sum_{i=1}^{N} |\hat{d}(X_i) - d_{\rm{opt}}(X_i)|$.

We compare $7$ estimators in our study: the two IPW estimators, the Wald estimator, and the two multiply robust estimators, along with the below IV estimators proposed by cui2021semiparametric:

align*[align* omitted — 404 chars of source]

where $n_{t0}$ and $n_{t1}$ are the sample sizes at time point $0, 1$, respectively; $\delta_{t0}(x) = \mu_{A}(0,1,x) - \mu_{A}(0,0,x)$, $\delta_{t1}(x) = \mu_{A}(1,1,x) - \mu_{A}(1,0,x)$, $\pi_{t0}(z,x) = Pr(Z = z \mid X = x, T = 0)$, $\pi_{t1}(z,x) = Pr(Z = z \mid X = x, T = 1)$ are the nuisance parameters, and $\hat{\delta}_{t0}, \hat{\delta}_{t1}, \hat{\pi}_{t0}, \hat{\pi}_{t1}$ can be estimated using parametric models or machine learning algorithms. We utilize the genetic algorithm implemented in the R package rgenoud mebane2011genetic to solve the optimization tasks.

First, we posit parametric models for the nuisance parameters. The linear/logistic regression models for $\mu_{A}(t,z,x;\alpha)$, $\mu_{Y}(t,z,x;\beta)$ and $\pi(t,z,x;\theta)$ are correctly specified. The sample size is $n = 5000$.

We also consider flexible machine learning algorithms for nuisance parameter estimation. Specifically, we apply the generalized random forests athey2019generalized implemented in the R package grf with default tuning parameters. For the cross-fitting procedure, we use $K = 4$ folds. The sample size is $n = 10^4$.

figure[figure omitted — 262 chars of source]

Figure (ref) reports the main simulation results from $500$ Monte Carlo replications. In both scenarios, the two standard IV estimators fail to learn the optimal policy, due to the direct effects of the treatment $A$ on the outcomes $Y_0, Y_1$. The two IPW estimators perform much better, but the variability can be large due to possibly extreme weights. The Wald and multiply robust estimators generally lead to lower variability, and attain superior performance. Additional simulation results are reported in Section (ref) of the Supplementary Material to illustrate how different sample sizes and the strength of the IV affect the performance of the estimated policies. We observe that a stronger strength of IV generally leads to lower variability and better accuracy, and also as sample size increases, our proposed methods have better performance.

Data application

In this section, we illustrate the use of the instrumented DiD approach for policy learning with a analysis of the Australian Longitudinal Survey (ALS) data. Researchers in labor economics have a longstanding interest in investigating the causal effect of education on earnings in the labor market. card2001estimating suggests that the endogeneity of education might partially explain the continuing interest “in this very difficult task of uncovering the causal effect of education in labor market outcomes", and argues that the effects of education are heterogeneous since the economic benefits are individual-specific. Besides the well acknowledged benefits of personal growth and social good from education, we aim to provide a personalized recommendation on whether an individual should pursuit more education or not, in order to gain higher earnings.

The Australian Longitudinal Survey was conducted annually since 1984. Specifically, we include the 1984 and 1985 waves as cross sectional data in our analysis. The 1984 wave surveyed a sample of $3000$ people aged $15 - 24$, and the 1985 wave consisted of $9000$ interviews with people aged $16 - 25$. The surveys aim mainly at providing data on the dynamics of the youth labour market, and include basic demographic variables, labour market variables, background variables and topics related to the main labour market theme. We follow the guidelines from su2013local,cai2006functional and vella1994gender, who was among the first researchers extensively working with the ALS data. Finally, our data include $2401$ subjects from the 1984 wave, and $8997$ subjects from the 1985 wave. We consider the following baseline covariates: whether a person is born in Australia, marital status, union membership, government employment, age and work experience. The treatment is the education level, and the outcome is the hourly wage. We use an index of labor market attitudes as the instrumental variable su2013local. The details of our analysis are provided in Section (ref) of the Supplementary Material.

The nuisance parameters are estimated by posited linear/logistic regression models, and we apply our proposed methods with the same configurations as Section (ref). The policy coefficient estimates of all covariates are reported in Table (ref).

table[table omitted — 1,341 chars of source]

The coefficients should be interpreted cautiously. We also find that there exists some discrepancies among the treatment recommendations by our proposed estimators. The Wald and multiply robust estimators usually agree, but the variability of the IPW estimators are a bit large. Due to the potentially different recommendations by different estimated policies, one may conservatively suggest a recommendation by the majority rule, and accordingly obtain an ensemble policy. It is also interesting to construct a decision tree to further explore which covariates indicate which treatment level qi2023proximal.

Extension to panel data

In this section, we consider extending the instrumented DiD approach to the panel data setup where a random sample from the population is followed up over two time points abadie2005semiparametric. The observed data are $O = (X, Z, A_{0}, Y_{0}, A_{1}, Y_{1})$. Let $\delta_{Y,z}(x) = E[Y_1 - Y_0 \mid X = x, Z = z]$, $\delta_{A,z}(x) = E[A_1 - A_0 \mid X = x, Z = z]$, and $\pi_{Z}(x) = Pr(Z = 1 \mid X = x)$. We make the following identification assumptions.

assumptionSuppose the following assumptions hold: {\rm (consistency)} $A_{t} = A_{t}(Z)$ and $Y_{t} = Y_{t}(A_{t})$ for $t = 0, 1$; {\rm (positivity)} $c_3 < \pi_{Z}(x) < 1 - c_3$ for some $0 < c_3 < 1/2$; {\rm (trend relevance)} $E[A_{1}(1) - A_{0}(1) \mid Z = 1, X] \neq E[A_{1}(0) - A_{0}(0) \mid Z = 0, X]$; {\rm (stable treatment effect over time)} $E[Y_{0}(1) - Y_{0}(0) \mid X] = E[Y_{1}(1) - Y_{1}(0) \mid X]$; {\rm (independence & exclusion restriction)} $Z \perp \{A_{t}(1), A_{t}(0), Y_{t}(1) - Y_{t}(0), Y_{1}(0) - Y_{0}(0) : t = 0, 1\} \mid X$; {\rm (no unmeasured common effect modifier)} $Cov\{A_{t}(1) - A_{t}(0), Y_{t}(1) - Y_{t}(0) \mid X\} = 0$ for $t = 0, 1$.

Assumption (ref) is the counterpart of Assumptions (ref)-(ref) for the panel/longitudinal structure. vo2022structural use a structural mean model and consider alternative assumptions to the no unmeasured common effect modifier assumption above. In Section (ref) of the Supplementary Material, we also prove the identification results under the following assumptions that replaces the no unmeasured common effect modifier assumption: (sequential ignorability) $Y_{t}(a) \perp A_{t} \mid U, X, Z$ for $t, a = 0, 1$, and there is no additive interaction of either {\rm (i)} $E[A_{1} - A_{0} \mid X, U, Z = 1] - E[A_{1} - A_{0} \mid X, U, Z = 0] = E[A_{1} - A_{0} \mid X, Z = 1] - E[A_{1} - A_{0} \mid X, Z = 0]$ or {\rm (ii)} $E[Y_{t}(1) - Y_{t}(0) \mid U, X] = E[Y_{t}(1) - Y_{t}(0) \mid X]$ for $t = 0, 1$. The sequential ignorability is intuitive, and commonly assumed in panel/longitudinal data analysis. We note that the no additive interaction assumption implies the no unmeasured common effect modifier assumption.

theoremUnder Assumption (ref), the CATE is nonparametrically identified by \begin{equation} \tau (x) = \frac{E[Y_{1} - Y_{0} \mid X = x, Z = 1] - E[Y_{1} - Y_{0} \mid X = x, Z = 0]}{E[A_{1} - A_{0} \mid X = x, Z = 1] - E[A_{1} - A_{0} \mid X = x, Z = 0]}, \end{equation} and the efficient influence function is \begin{align*} \phi_{\rm{panel}} = & \frac{\delta_{Y,1}(x) - \delta_{Y,0}(x)}{\delta_{A,1}(x) - \delta_{A,0}(x)} - \frac{z - \pi_{Z}(x)}{\pi_{Z}(x) (1 - \pi_{Z}(x)) (\delta_{A,1}(x) - \delta_{A,0}(x))^2} \left\{(y_{1} - y_{0})(\delta_{A,1}(x) - \delta_{A,0}(x)) \right. \\ & \left. - (a_{1} - a_{0})(\delta_{Y,1}(x) - \delta_{Y,0}(x)) + \delta_{Y,1}(x) \delta_{A,0}(x) - \delta_{Y,0}(x) \delta_{A,1}(x)\right\} - \tau (x). \end{align*}
theoremUnder Assumption (ref), the optimal policy is nonparametrically identified by \begin{equation} \arg\max_{\mathcal{D}} E\left[\frac{\delta_{Y,1}(X) - \delta_{Y,0}(X)}{\delta_{A,1}(X) - \delta_{A,0}(X)} d(X)\right] = \arg\max_{\mathcal{D}} E\left[\Delta_{\rm{panel}} (X) d(X)\right], \end{equation} where the uncentered efficient influence function $\Delta_{\rm{panel}}$ is \begin{align*} \Delta_{\rm{panel}} = & \frac{\delta_{Y,1}(X) - \delta_{Y,0}(X)}{\delta_{A,1}(X) - \delta_{A,0}(X)} - \frac{Z - \pi_{Z}(X)}{\pi_{Z}(X) (1 - \pi_{Z}(X)) (\delta_{A,1}(X) - \delta_{A,0}(X))^2} \left\{(Y_{1} - Y_{0})(\delta_{A,1}(X) - \delta_{A,0}(X)) \right. \\ & \left. - (A_{1} - A_{0})(\delta_{Y,1}(X) - \delta_{Y,0}(X)) + \delta_{Y,1}(X) \delta_{A,0}(X) - \delta_{Y,0}(X) \delta_{A,1}(X)\right\}. \end{align*}

Estimators of optimal policies can be constructed by the empirical versions of equations in Theorem (ref), and the cross-fitting procedure can also be applied when using the efficient influence function. Similarly, asymptotic analysis of policy learning as Theorems (ref) and (ref) can be established for panel data.

Discussion

Similar approaches as the instrumented difference-in-differences design has long been employed by econometricians duflo2001schooling and has also been formally considered as fuzzy differences-in-differences by de2018fuzzy, where the individuals can switch treatment in only one direction within each treatment group. We refer interested readers to ye2022instrumented and its rejoinder for discussions on the differences, and applications in biomedicine and epidemiology.

There are several interesting directions for future research and application. Our approach is the first work to systematically study policy learning under the DiD setting. It may be possible to consider alternative assumptions or structures in DiD design to learn the optimal policy. Our instrumented DiD may also be generalized to multiple time points, continuous time, or continuous IV.

Note that Assumption (ref) can be replaced by the monotonicity assumption, i.e. $A_{t}(1) \geq A_{t}(0)$ for $t = 0, 1$ with probability $1$, which identifies the complier treatment effects. Then we can also target complier optimal policies that would optimize the potential outcome among compliers.

Acknowledgements

The authors are grateful for helpful comments and feedback from Julie Josse and Antoine Chambaz.

Pan Zhao is supported in part by the French National Research Agency ANR-16-IDEX-0006. Yifan Cui is supported in part by the National Natural Science Foundation of China and the Open Research Fund Key Laboratory of Advanced Theory and Application in Statistics and Data Science (East China Normal University), Ministry of Education of the People’s Republic of China. The authors are grateful to the OPAL infrastructure from Université Côte d'Azur for providing resources and support.