Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
111,483 characters · 25 sections · 83 citation commands
Evaluating Program Sequences with Double Machine Learning: An Application to Labor Market Policies
\begingroup \let\relax \endgroup
\thispagestyle{empty}
\setcounter{page}{1}
\setcounter{page}{1}
Program evaluations based on observational data are used by economists to assess the impact of policy measures such as training programs, public transfer schemes, and healthcare interventions abadie2018econometric. The standard approach in the literature involves comparing outcomes between treatment groups of program participants and a control group of non-participants. If all confounding variables jointly influencing the outcome and program assignment are observed, the causal effect of the intervention can be identified. By collapsing program assignment, participation, and completion into a single treatment state, the described approach considerably simplifies the complexities prevalent in many non-experimental settings, limiting the ability to address many important questions. For example, program duration may vary between participants, which could have a strong impact on program efficiency. Furthermore, individuals may participate not only in one but in several successive measures, which makes it difficult to match them to one particular program type. Importantly, a specific program sequence may result from reassignment following the initial placement. For instance, a program may be shortened if an individual is no longer eligible, or the program may be switched if the initial program is deemed unsuccessful. After all, a particular program sequence may be effective for one person but not for another, which cannot be distinguished when only analyzing population-level effects.
This paper addresses the previously mentioned challenges by adopting a framework to evaluate program sequences instead of single time-point interventions. The framework is based on ideas from biostatistics, originally formulated by robins1986new, robins1987addendum and further developed in subsequent research, as reviewed in richardson2014causal. While these methods have frequently been applied to the analysis of sequences of medical treatments hernan2002estimating, taubman2009intervening, young2011comparative, they have so far seen limited adoption in the econometric program evaluation literature. Wider adoption has mainly been hindered by the stricter identification and estimation requirements in the sequential setting compared to one-time interventions. This paper demonstrates that these concerns can be addressed by (i) extending sequential estimands to dynamic policies that facilitate the credible identification of counterfactual treatment sequences using observational data and (ii) leveraging recently proposed machine learning-based estimation strategies to flexibly estimate these quantities. These innovations enhance the applicability of sequential analysis for assessing policy measures commonly studied in economics, ensuring a better reflection of their sequential nature.
The key challenge in analyzing program sequences from observational data is that the complete trajectory of treatments is typically not predetermined prior to the start of the sequence. This must be considered both when accounting for the program assignment mechanism underlying the observed data (dynamic confounding) and when designing counterfactual scenarios that depend on time-varying covariates (dynamic policies). While the former has been addressed in prior econometric research lechner2009sequential, the latter has thus far received little attention in evaluation studies. Instead, dynamic policies have primarily been used in the context of optimal dynamic treatment regimes murphy2003optimal, zhang2013robust, sakaguchi2024robust. This body of literature develops procedures to select an optimal dynamic policy from a class of feasible policies, with the objective of maximizing the mean response in the population. While such approaches allow to provide individualized treatment recommendations, identifying the full set of policies in a given class requires strong assumptions that are often implausible in non-experimental contexts. In contrast, the present work adopts an alternative objective. Rather than optimizing over an extensive policy class, it is proposed to selectively isolate specific policies that closely align with the assumed underlying assignment process and are therefore convincingly identifiable from observable data. Accordingly, instead of focusing on individualized treatment recommendations, dynamic policies are leveraged as a framework to uncover aggregate and group-level effects that enable robust estimation and inference.
\defcitealias{chernozhukov2018double}{Chernozhukov, Chetverikov, Demirer, Duflo, Hansen, Newey, & Robins, 2018;}
A major challenge for statistical inference in sequential settings is that the number of possible program sequences expands exponentially with the number of periods, leading to data sparsity for individual sequences. For example, a setting with five treatments and five time periods already yields $5^{5}=3,125$ possible treatment sequences, and ensuring sufficient observations for each becomes increasingly difficult. Hence, while nonparametric identification of dynamic policies is achievable, conventional estimation approaches, such as structural nested mean models robins1989analysis, robins1994correcting or marginal structural models robins2000marginal, rely on structural assumptions to extrapolate into regions with limited data support.\footnote{For an overview of these methods, see Appendix (ref)} However, if a fair number of observations is available for each program sequence of interest, more flexible, machine learning-based estimators can be employed. Aggregating information across periods and/or treatments can enhance feasibility, especially when credible information about the functional form of the underlying data-generating process is lacking. Building on this approach, the present paper demonstrates the potential of recently proposed double machine learning (DML) methods \citetext{\citetalias{chernozhukov2018double} bodory2022evaluating, bradic2021high} for analyzing dynamic policies. The key advantages of DML methods are that they do not require parametric assumptions and that they can handle a potentially large covariate space while maintaining desirable statistical properties.
In the empirical application, the analysis focuses on identifying which implementation aspects are most beneficial for individuals selected to participate in active labor market policies (ALMP). ALMP are interventions designed to improve employment outcomes for unemployed individuals and are widely used by governments as policy tools.\footnote{For an overview, refer to reviews and meta-studies by card2010active,card2018works,crepon2016active,vooren2019effectiveness} In Switzerland, which is the focus of this paper, two-thirds of unemployed are assigned to these measures within the first twelve months of their unemployment period, with approximately 60% of them participating in more than one program. The application is the first to jointly address dynamic confounding and dynamic policies in evaluating ALMP sequences, whereas prior work lechner2013does, laffers2024locking addressed only the former. Specifically, it is shown that introducing a dynamic policy based on intermediate outcomes allows to identify a more practically relevant estimand with only minor modifications to the identifying assumptions. Furthermore, the application innovates by applying flexible dynamic DML estimation in the context of ALMP, offering more reliable effect estimates compared to prior sequential studies that relied on parametric estimators. Leveraging DML in combination with a large dataset containing extensive intermediate details about the unemployed also allows to assess group-level effect heterogeneity in a sequential ALMP evaluation. The analysis shows that a particular temporary wage subsidy is most effective on average, even after aligning program duration across programs. Moreover, individuals with limited language skills profit more from this program in comparison to extended training courses. The findings emphasize the practical value of employing DML-based estimation for the empirical evaluation of programs sequences, allowing policymakers to develop more nuanced and targeted policies.
The remainder of the paper is structured as follows: The following Section (ref) introduces the notation, the conceptual framework, the estimands of interest and their identification. Section (ref), presents different estimators for DML under static and dynamic confounding and discusses their properties. Section (ref) applies these estimators to the empirical application, followed by the conclusion in Section (ref). Further derivations and additional results are provided in the Appendix.
This section introduces the framework for assessing sequential policies, drawing primarily on hernan2020causal, while employing a notation more commonly used in econometrics. The framework is presented in a setting with $T=2$ time periods, as used in the application in Section (ref). For clarity, the presentation avoids using a general $T$, although the framework naturally extends to cases with $T>2$. When $T=1$, it coincides with the standard setting of a single time-point intervention. Let $\mathcal{D}_t = \left\{0,1,...,M_t\right\}$ denote the set of existing programs in period $t\in \left\{1,2\right\}$ and $D_{i,t}$ a discrete random variable indicating the observed program of individual $i \in \left\{1,...,N\right\}$ in period $t$. Prior to each program, a vector of covariates $X_{i,t-1} \in \mathcal{X}_{t-1}$ is observed, which may contain lagged outcomes $Y_{i,t-1} \in \mathcal{Y}_t$. The main outcome of interest is observed after the final treatment $Y_{i}:=Y_{i,T}=Y_{i,2}$, where the time subscript is dropped for the sake of readability. This variable may include outcome information from any period following the start of the final treatment. In general, capital letters (except for $T$, $M$ and $N$) denote random variables, lower case letters denote their realizations and boldface letters denote vectors of variable histories up to $t$, i.e. $\mathbf{X}_{i,t} = (X_{i,0}, ..., X_{i,t})$ and $\mathbf{d}_t = (d_1, ..., d_t)$ such that for example $\mathbf{d}_1=d_1$ and $\mathbf{d}_2=(d_1, d_2)$. To simplify the notation further, the individual identifier $i$ is omitted when not explicitly needed and all random variables for which realizations can be observed are collected in the vector $\mathbf{W}_2=(Y,\,\mathbf{D}_2,\,\mathbf{X}_1)$.
The potential outcome framework of rubin1974estimating is adopted to study causal effects, where $Y^{\mathbf{d}_2}$ denotes the hypothetical outcome under a particular treatment sequence $\mathbf{d}_2 \in \mathcal{D}_1 \times \mathcal{D}_2$. Similarly, potential values of the time-varying covariates are defined as $X_1^{d_1}$. The relationship between potential and observed variables follows the standard observation rule, where an observed variable is assumed to correspond to the potential variable associated with the assigned treatments.
This assumption implicitly requires that there are no unrepresented programs in the population of interest (everyone is assigned to a particular program sequence $\mathbf{d}_2$) and that there are no relevant interactions between individuals, meaning that the program of one individual does not affect the final outcome and intermediate covariates of another individual.
The key challenge in analyzing program sequences from observational data is that the program assignment mechanism underlying the observed data might be affected by time-varying feedback between treatments, covariates and outcomes. Hence, to identify the causal effect, it is essential to consider that sequential treatment assignment is a decision process potentially influenced by dynamic confounding.
According to this definition, confounding is static if the expected outcome of those treated with $\mathbf{D}_2=\mathbf{d}_2$ equals the expected outcome if everyone were treated with $\mathbf{d}_2$, conditional on pre-treatment information. This implies that the entire treatment sequence is pre-determined given $X_0$. Conversely, if there is feedback between treatments, covariates, and outcomes across periods, this equality will generally not hold, and the confounding is considered dynamic. Under dynamic confounding, controlling only for pre-treatment information often proves inadequate when assessing the impacts of treatment sequences. Therefore, existing research has developed alternative identification strategies that rely on modified identification assumptions and also imply new challenges for effect estimation robins2009estimation.
The causal diagram pearl1995causal in Figure (ref) illustrates confounding the two-period setup. Prior to the first period, the pre-treatment covariates $X_0$ are observed. These variables may have a causal effect (arrow) on any future treatment, intermediate covariate, or outcome. In period $t=1$, individuals are treated with $D_1$, which may impact their time-varying characteristics in the current and subsequent periods as well as any future treatment assignments. The same happens in period $t=2$. Under dynamic confounding, the initial treatment assignment $D_1$ induces changes in the covariates $X_1$, which subsequently influence the second treatment $D_2$ and the outcome $Y$, as highlighted by the blue bold arrows. In contrast, if any one of the bold blue arrows is absent, confounding is considered static, meaning that program assignment in all periods is predetermined conditional on pre-treatment information.\footnote{In absence of a causal pathway from $D_1$ to $X_1$, the variable $X_1$ can also considered to be a pre-treatment variable.}
Besides addressing dynamic confounding, dynamics also need to be considered in the design of counterfactual scenarios. While confounding is determined by the observational setting, counterfactuals can be freely specified to meet the research objectives. However, as discussed below, aligning counterfactuals with the assumed confounding structure can substantially facilitate identification and estimation. In the sequential setting, static counterfactuals that do not depend on time-varying covariates, such as “two consecutive periods of a training course,” are often of limited relevance because they fail to account for the possibility of dynamic decision-making wager2024causal. For instance, the mean potential outcome for such a sequence reflects a scenario in which all individuals remain in the program for the entire sequence. This implies continued participation even in cases where individuals would no longer be eligible. Dynamic decision-making can be formalized through the concept of dynamic policies\footnote{Policies are also called treatment rules or regimes. Following wager2024causal, here they are denoted as policies as this is the standard terminology in the econometrics literature.} murphy2003optimal.
The policy $\mathbf{g}_2$ is defined as a mapping from the (potential) decision variables to a subset of the possible treatment states. In the most flexible setting, decision variables can include all covariates, i.e. $\mathbf{V}^{d_1}_1=\mathbf{X}^{d_1}_1$, and the policy can return any possible program, i.e. $\mathcal{R}_t = \mathcal{D}_t$. Often, however, it is useful to align $\mathbf{V}^{d_1}_1$ and $\mathcal{R}_t$ with the requirements of the counterfactual scenario under consideration. For instance, the subsequent application will define $\mathbf{V}^{d_1}_1$ in a way to ensure that policies can depend on the evolution of the outcome variable. Importantly, while dynamic policies may be restricted to depend on selected decision variables, such restrictions do not limit the nature of confounding present in the data.
A key property of dynamic policies is that two individuals following the same policy $\mathbf{g}_2$ may follow different program sequences $\mathbf{d}_2$ depending on their covariates. For instance, if $Y^d_1$ in Example (ref) represents a binary employment indicator, then (i) individuals who remain unemployed after receiving treatment $d$ in the first period under the sequence $(d,d)$, and (ii) individuals who become employed and follow sequence $(d,d')$, would both comply with the policy. A sequence $\mathbf{d}_2$ constitutes a specific instance of a policy, where $\mathbf{g}_2$ is expressed as a constant function. This type of policy is referred to as a static policy. In what follows, $\mathbf{d}_2$ is used to represent static policies exclusively, whereas $\mathbf{g}_2$ is used in contexts that encompass both static and dynamic policies. An overview of possible scenarios involving static and dynamic confounding and policies, along with accompanying examples, is given in Table (ref).
The analysis aims at evaluating effects of sequential policies at different levels of granularity. At the broadest level of the entire population, the primary target parameter is the average potential outcome (APO) under a particular policy, defined as
This parameter represents the mean outcome if all individuals were assigned according to the policy $\mathbf{g}_2$ (or according to the sequence $\mathbf{d}_2$ if the policy is static). To assess heterogeneity, the focus can be switched to a specific subgroup of interest, for example to individuals that participated in a labor market program during a previous unemployment spell. This can be studied by examining the group average potential outcome (GAPO)
where $Z_0$ is a column or deterministic function of $X_0$ with low cardinality.\footnote{Subgroups are defined based on pre-treatment covariates, since estimands involving time-varying covariates are not identifiable under the assumptions stated below.} The following discussion of identification and estimation will focus on the parameters $\theta^{\mathbf{g}_2}$ and $\theta^{\mathbf{g}_2}(z_0)$, respectively. The average treatment effect (ATE) between implementing policy $\mathbf{g}_2$ and alternative policy $\mathbf{g}'_2$ is obtained subsequently by taking the difference between two average potential outcomes, i.e.
which directly follows from the linearity of the expectation operator. Subgroup-specific average treatment effects (GATE) are similarly obtained as
Heterogeneous treatment effects of this type are also referred to as conditional average treatment effects (CATE) in the literature. Following knaus2021machine, they are denoted as GATE to emphasize the focus on large, discrete subgroups of the population rather than granular individualized effects.\footnote{More granular individualized effects are not considered because DML-based estimators of such effects lack statistical guarantees and perform poorly in finite samples under confounding, as shown for example by lechner2024comprehensive for single time-point interventions. The development of robust estimation methods for individualized effects in the sequential setting remains an open question for future research beyond the scope of this study.}
The prior section defined the estimands of interest using potential outcomes. However, since each individual is observed in only one particular program sequence, the remaining potential outcomes are unobservable, necessitating additional assumptions for their identification. The assumptions required vary based on the type of confounding and the nature of the policy of interest. Before introducing dynamic policies, the identification assumptions for static policies are revisited.
In the presence of static confounding, the APO of a static policy $\mathbf{g}_2(\mathbf{V}^{g_1}_1) = \mathbf{d}_2$ can be identified if the following conditional independence assumption (CIA) and overlap assumption are satisfied.
Assumption (ref)(ref) is called full conditional independence assumption in lechner2001potential. It requires that potential outcomes are independent of program assignment in the first period for given values of the pre-treatment covariates. Thus, $X_0$ must include all variables that influence both the first period program assignment and the outcomes simultaneously. Additionally, the assumption asserts that allocation to a program in the second period is random, conditional on pre-treatment covariates and prior program participation. This implies that intermediate characteristics $X_1$ are either unaffected by previous programs or do not simultaneously influence both program assignment in the second period and potential outcomes. Furthermore, Assumption (ref)(ref) requires that for a given history of previous program participation and any realization of pre-treatment characteristics, it must be possible to observe individuals with a program $D_t = d_t$. Otherwise it would be impossible to construct appropriate counterfactuals. lechner2001potential show that Assumption (ref) can be re-written using basic probability theory:
This formulation shows that the sequential framework with static confounding is equivalent to the static single-period setting for multiple treatments Imbens:2000,lechner2001identification, where sequence $\mathbf{d}_2$ is considered a single treatment state, and identification is achieved by conditioning on pre-treatment covariates only. It is simple to show that APO and GAPO of a static policy $\mathbf{g}_2(\mathbf{V}_1) = \mathbf{d}_2$ can be identified under this assumption. The proof follows the standard identification argument for single time-point interventions as seen e.g. in rosenbaum1983central.\footnote{While the focus is on constant static policies defined as a fixed program sequence $\mathbf{d}_2$, the results can be generalized to static policies $h_t: \mathcal{V}_0 \rightarrow \mathcal{D}_t$ that depend on pre-treatment decision variables $V_0$ but do not depend on time-varying information. Specifically, $E[Y^{\mathbf{h}_2}]$ is identified under the condition that Assumption (ref) or (ref) holds for all $\mathbf{d}_2$ within the range of policies $\mathbf{h}_2(V_0)$. Similar identification results are commonly applied in optimal policy learning in the single-period setting Athey:2021.}
In observational settings, static confounding may be implausible if the assignment process underlying the data depends on time-varying information. Causal effects under dynamic confounding have first been studied in epidemiology and biostatistics starting with the seminal work by robins1986new, robins1987addendum. He demonstrated that under feedback effects between treatments and covariates, standard approaches do not allow for a causal comparison of treatment sequences, even when all pre-treatment confounding factors are controlled for. Instead, using a graphical model, robins1986new came up with the idea of what is called today a sequential randomized experiment and the sequential randomization assumption richardson2014causal. For static policies, these ideas have been introduced to econometrics by lechner2009sequential and lechner2010identification who termed the requirement for identification as weak dynamic conditional independence:
For $t=1$, Assumption (ref) is equivalent to Assumption (ref). However, in the second period, conditioning on pre-treatment characteristics $X_0$ and the treatment history no longer suffices to establish independence between potential outcomes and treatment state $D_2$. Instead, Assumption (ref)(ref) requires conditioning on the whole covariate history $\mathbf{X}_{1}$, which may be influenced by previous treatments. Hence, additional information about time-varying covariates that are expected to determine dynamic program selection is necessary. Importantly, these covariates may not be influenced by programs in future periods in a way that is related to the outcome variable.\footnote{This exogeneity requirement can be explicitly stated as $\forall \mathbf{d}_2, \mathbf{d}'_2 \in \mathcal{D}_1 \times \mathcal{D}_2: X_0 ^{\mathbf{d}_2} = X_0^{\mathbf{d}'_2}$ and $X_1^{{d}_2} = X_1^{{d}'_2}$, where $X_t^{\mathbf{d}_2}$ denotes the potential covariates in period $t$ under treatment sequence $\mathbf{d}_2$.} The overlap assumption (ref)(ref) also requires additional conditioning on time-varying covariates for the treatment probability in the second period. Under Assumption (ref), the APO and GAPO are identified, as stated in the following theorem:
Hence, the identification result for dynamic confounding is analogous to the case of static confounding, with $\mu_{\mathbf{d}_2}(X_0)$ replaced by $\nu_{\mathbf{d}_2}(X_0)$.\footnote{Assumption (ref) allows to identify additional parameters such as the average potential outcome within a particular treatment group in the first period $\mathbb{E}[Y^{\mathbf{d}_2}|D_1=d'_1] \; \forall\; d'_1 \in \mathcal{D}_1$. However, identification for subgroups defined by later periods, i.e. $\mathbb{E}[Y^{\mathbf{d}_2}|\mathbf{D}_2=\mathbf{d}'_2]$, is not possible under Assumption (ref) if $d_1 \neq d_1'$, as noted by lechner2010identification.} The proof of Theorem (ref) first establishes that $\mu_{\mathbf{d}_2}(\mathbf{X}_1) = \mathbb{E}[Y^{\mathbf{d}_2}|\mathbf{X}_1, D_1=d_1]$. However, due to potential feedback between the treatments through $X_1$, the conditional outcome $\mu_{\mathbf{d}_2}(\mathbf{X}_1)$ cannot simply be averaged over the population to determine the APO. Instead, an extra averaging step over the conditional distribution of covariates $X_1$ is required. Identification through $\nu_{\mathbf{d}_2}(X_0)$ is known in the literature as the $g$-formula or iterated conditional expectations robins1986new, robins1987addendum, hernan2020causal.
As previously discussed, prior econometric applications have mainly focused on static policies while program assignment in practice often involves dynamic decision-making. When considering dynamic policies, it is necessary to adjust identification assumptions, as they are based on intermediate potential decision variables. As a result, the counterfactual scenarios become more complex, since the policy can deterministically depend on these intermediate variables. For instance, the policy defined in Example (ref) assigns an individual to program $d_2$ or $d'_2$ depending on whether the potential decision variable in the first period equals one or not. These potential intermediate values need to be considered in the conditional independence assumption.
Assumption (ref) extends the notion of `full' conditional exchangeability in robins2009estimation by employing $\mathbf{V}_1^{d_1}$ instead of $\mathbf{X}_1^{d_1}$ and $\mathcal{R}_t$ instead of $\mathcal{D}_t$. Compared to Assumption (ref) for static policies, Assumption (ref) introduces two key differences. First, conditional independence and overlap must hold for all possible treatment sequences that the dynamic policy could generate, rather than just for one prespecified sequence. Second, given the pre-treatment characteristics, the assignment in the first period must be independent not only of any potential final outcomes but also of the potential intermediate variables that govern the second-period treatment. This additional requirement ensures that no unmeasured confounding affects the relationship between the first-period treatment and these intermediate variables. For example, consider an unmeasured confounder that influences both $D_1$ and $V^{{g}_1}_1$ but does not affect $Y^{\mathbf{g}_2}$. In this case, $\mathbb{E}[Y^{\mathbf{g}_2}]$ is not identifiable as $\mathbf{g}_2$ is a function of $V^{{g}_1}_1$ robins1986new, robins2009estimation.
Importantly, in contrast to the existing literature that defines policies as functions of the entire covariate vector, the present framework introduces a distinction between confounders and decision variables. Consequently, even though all covariates $\mathbf{X}_1$ might influence the underlying dynamic selection, conditional independence of the treatment in $t=1$ needs to hold only with respect to the potential $V_1^{d_1}$ that are part of the dynamic policy (and $Y^{\mathbf{d}_2}$) but not with respect to all $X_1^{d_1}$. Therefore, the choice of the structure of the dynamic policy critically influences the increased restrictiveness of Assumption (ref). For instance, in Example (ref), the dynamic policy depends only on potential intermediate outcomes, making the additional requirements less demanding. This is because, in many practical contexts, the CIA between the first-period treatment and the final outcome naturally extends to intermediate outcomes as well. Furthermore, if the chosen dynamic policy $\mathbf{g}_2$ more closely aligns with the assignment process underlying the observed data than a static policy $\mathbf{d}_2$, the overlap condition in Assumption (ref)(ref) becomes more credible than its static counterpart (ref)(ref). An example of such a setting is provided below in the empirical application.
Under Assumption (ref), identification is achieved in a similar way as for static policies, as shown in Theorem (ref). Hence, despite the need for additional identification conditions under dynamic policies, the logic of identification mirrors that of static policies once these are met. As shown, the average outcome associated with following policy $\mathbf{g}_2(\mathbf{V}_1^{g_1})$ can be inferred from individuals whose observed treatments and covariates align with following strategy $\mathbf{g}_2(\mathbf{V}_1)$.\footnote{Thus far, the case of dynamic policies under static confounding has not been considered. When relying on dynamic policies, it is often desirable to construct counterfactuals that closely reflect observed practice. Hence, in observational studies, dynamic policies usually should be accompanied by dynamic confounding in the underlying data. Dynamic policies under static confounding might become relevant in experimental settings, for example when using stratified-on-$X_0$ randomization. For such cases, Assumption (ref) can be weakened by conditioning only on $X_0=x_0$ instead of $\mathbf{X}_1 = \mathbf{x}_1$ in the second period. An explicit statement of this assumption is omitted for the sake of brevity.}
In addition to the identification results using the $g$-formula discussed in the previous subsections, alternative identification approaches based on Assumptions (ref)-(ref) offer favorable robustness properties through additional reweighting by propensity scores. Let
denote so-called score functions for the settings under $st$atic and {$dy$}namic confounding, respectively. It can be shown that the estimands of interest are identified as expectations of these scores:
These identification results go back to robins1994estimation for static confounding and robins2000robust for dynamic confounding. The former possesses the well-known double robustness property in the sense that it even identifies the APO if either $\mu_{\mathbf{d}_2}(\cdot)$ or $p_{\mathbf{d}_2}(\cdot)$ is replaced by any alternative function. Under dynamic confounding, this result extends to multiple robustness, which ensures identification if at least one component from each of the two periods (i.e. $\nu_{\mathbf{g}_2}(\cdot)$ or $p_{g_1}(\cdot)$ and $\mu_{\mathbf{d}_2}(\cdot)$ or $p_{g_2}(\cdot)$) is correctly specified. The score functions ((ref)) and ((ref)) will become important for the estimation procedures discussed in the following section.
The previous section introduced several estimands of interest and showed that, under identification assumptions, aggregates of their unobserved components can be expressed in terms of random variables, from which observations can be sampled. Given a set of such random samples, different techniques proposed in the literature can be used to estimate the parameters of interest and to conduct inference. This paper considers flexible machine learning estimators of $\theta^{\mathbf{g}_2}$ and $\theta^{\mathbf{g}_2}(z_0)$ that do not require functional form assumptions about the underlying data generating process and allow to use high-dimensional covariates, enhancing the credibility of the identification arguments in observational settings.\footnote{The focus here is on machine learning-based estimators; for an overview of conventional parametric methods in the sequential setting see Appendix (ref) or the textbook introduction in hernan2020causal.} However, machine learning estimators trade-off bias against variance to obtain a low mean squared error. Hence, naïve plug-in estimators that solely rely on machine learning estimates of $\mu_{\mathbf{d}_2}(X_0)$ or $\nu_{\mathbf{g}_2}(X_0)$ are biased due to regularization. To address this issue, it has been shown that plug-in estimators can be de-biased using an orthogonal moment condition, which adds an influence function adjustment to the target parameter chernozhukov2018double. Estimators based on the influence function have desirable properties, such as $\sqrt{N}$-convergence and asymptotic normality. Furthermore, an estimator needs to solve the influence function in order to be asymptotically efficient robins1994estimation.
The orthogonal moment condition of the the average potential outcome $\theta^{\mathbf{g}_2}$ is given by
where $\eta^{st} = (\mu_{\mathbf{d}_2}(X_0), p_{\mathbf{d}_2}(X_0))$ and $\eta^{dy} = (\nu_{\mathbf{g}_2}(X_0), \mu_{\mathbf{g}_2}(\mathbf{X}_1), p_{d_2}(\mathbf{X}_1, g_1), p_{g_1}(X_0))$ denote the vectors of nuisance functions under static and dynamic confounding, respectively. For true $\theta^{\mathbf{g}_2}$ and $\eta^j$ the condition is satisfied, as shown by the identification result in the previous section. This suggests to construct an estimator solving an empirical equivalent of it,
where $\hat\eta^j$ refers to the estimated nuisance functions. In addition, for the group average potential outcome, $\theta^{\mathbf{g}_2}(z_0)$, the moment condition
suggests to use the estimator
Hence, estimation of both parameters of interest can be based on the same scores chernozhukov2024applied.
Under static confounding ($j=st$), estimation is implemented by separately predicting $\hat\mu_{\mathbf{d}_2}(X_0)$ and $\hat p_{\mathbf{d}_2}(X_0)$ by machine learning methods, plugging them into Equation ((ref)), i.e.
and finally averaging this score over the population of interest. For statistical inference, the variance of the scores can be used to construct a standard $t$-test statistic. The structure of the score ${\Theta}^{st}_{\mathbf{d}_2}(\mathbf{W}_{2}, \hat\eta^{st})$ demonstrates that the regularization bias in the estimation of $\hat{\mu}_{\mathbf{d}_2}(X_{0})$ is corrected by adding an adjustment term, consisting of the conditional outcome residuals, re-weighted by the inverse treatment probability. Hence, the adjustment increases as the prediction deviates further from the observed outcome and as the conditional treatment probability decreases.
DML in the single-period setting has been proposed by chernozhukov2018double. In their seminal paper, the authors demonstrate the importance of addressing regularization bias and overfitting in the estimates of the nuisance functions when employing the machine learning-based plug-in approach. This is achieved by (i) using an orthogonal moment function, and (ii) by using a cross-fitting procedure. The orthogonal moment condition (i) allows to obtain $\sqrt{N}$-consistency even when using machine learning estimators that typically converge at relatively slow rates. This is because it satisfies a certain `Neyman'-orthogonality property that makes it locally insensitive to small biases in the nuisance estimates. In analogy to the concept of identification double robustness discussed in Section (ref), this property is also referred to as rate double robustness in the literature knaus2022double. The cross-fitting procedure (ii) avoids overfitting by ensuring that observations are not used to predict their own nuisance functions. Therefore, the set of observations $\mathcal{W} = \left\{1, ..., N\right\}$ is split into $K$ equally sized subsamples $\mathcal{W}_k$. For each $k=1,...,K$ the nuisance parameters are first trained on the complementing subset $\mathcal{W}_{-k}$, then predicted in the subsample $\mathcal{W}_k$ and plugged into the score function. Finally, the DML estimate is obtained by taking the mean of the cross-fitted scores. In line with earlier reasoning, this procedure designed for the single-period setting, can be directly applied to sequential estimation under static confounding by treating each program sequence $\mathbf{d}_2$ as a distinct treatment state.
Several recent contributions bodory2022evaluating, bradic2021high, singh2021finite, chernozhukov2022automaticdyn propose extensions of single-period DML to the sequential setting under dynamic confounding. All these extensions are introduced within a framework of static policies. However, once identification is established, DML-based estimation can proceed similarly for both static and dynamic policies, as detailed below. Accordingly, the procedures are presented directly using the more general notation for dynamic policies.
All mentioned contributions are based on the orthogonal moment condition ((ref)) with score
which now consists of two re-weighted outcome residuals, one for each period. This corresponds to the efficient influence function for longitudinal treatment effects robins2000robust, bang2005doubly. The authors demonstrate that this score satisfies the Neyman orthogonality condition, indicating its suitability for the DML approach. However, the extension to dynamic confounding is complicated by the fact that $\nu_{\mathbf{g}_2}(X_0)$ cannot be directly represented by observable variables as it nests the conditional mean outcome of the second period $\mu_{\mathbf{g}_2}(\mathbf{X}_{1})$. Hence, estimation of $\nu_{\mathbf{g}_2}(X_0)$ requires estimates $\hat{\mu}_{\mathbf{g}_2}(\mathbf{X}_{1})$ as inputs, which can be implemented in various ways.
An initial approach proposed by bodory2022evaluating is to estimate $\hat\nu_{\mathbf{g}_2}(X_0)$ by regression of $\hat{\mu}_{\mathbf{g}_2}(\mathbf{X}_{1})$ on $X_0$ for observations following $D_1 = g_1(V_0)$, i.e.
where $\hat{\mathbb{E}}[A|B, C=c]$ denotes predictions from a regression of $A$ on $B$ for observations with $C=c$. This requires an additional split of the subsamples $\mathcal{W}_{-k}$ to avoid overfitting, such that $\hat\nu_{\mathbf{g}_2}(X_0)$ and $\hat{\mu}_{\mathbf{g}_2}(\mathbf{X}_{1})$ are not learned from the same sample. In a setting with more than two time periods, the number of splits would increase even further.
bradic2021high propose another method for implementing DML under dynamic confounding, introducing an additional bias correction term for estimating the nested conditional outcome. In particular, they estimate $\hat\nu_{\mathbf{g}_2}(X_0)$ as
where the pseudo-outcome that is regressed on $X_0$ in the subsample $D_1=g_1(V_0)$ corresponds to the doubly robust score of the treatment effect of the second period. Again, this requires second-order sample splitting, but the authors propose to regain full sample size efficiency by cross-fitting. In a first step, $\hat{\mu}_{\mathbf{g}_2}(\mathbf{X}_{1})$ is estimated using the first subsample of $\mathcal{W}_{-k}$, and predictions are made on the second subsample. These predictions are then used to estimate $\hat{\nu}^{\text{BJZ24}}_{\mathbf{g}_2}(X_0)$. The process is repeated by reversing the roles of the subsamples, resulting in two estimates of $\hat{\nu}^{\text{BJZ24}}_{\mathbf{g}_2}(X_0)$. Subsequently, the observations from fold $\mathcal{W}_{k}$ are applied to both estimated functions, and the predictions are averaged for each observation. While the original paper is only formulated for a binary treatment setting, it is extended here to the case of multiple treatments. The exact algorithm and a comparison to the method of bodory2022evaluating is provided in Appendix (ref).
For both approaches, the properties of $\sqrt{N}$-consistency and asymptotic normality extend from single-period DML to the sequential setting. Besides standard regularity conditions, the theoretical guarantees rely on four (instead of two in the single-period framework) consistent nuisance parameter predictions $\hat{\mu}_{\mathbf{g}_2}(\mathbf{X}_{1})$, $\hat{\nu}_{\mathbf{g}_2}(X_0)$, $\hat{p}_{g_1}(X_0)$ and $\hat{p}_{g_2}(\mathbf{X}_1, g_1)$. In addition, bodory2022evaluating require three product rate conditions: The two within-period products of the rates of convergence between $\hat{\nu}_{\mathbf{g}_2}(X_0)$ and $\hat{p}_{g_1}(X_0)$, and between $\hat{\mu}_{\mathbf{g}_2}(\mathbf{X}_{1})$ and $\hat{p}_{g_2}(\mathbf{X}_1, g_1)$, as well as the cross-period product rate between $\hat{\mu}_{\mathbf{g}_2}(\mathbf{X}_{1})$ and $\hat{p}_{g_1}(X_0)$ need to be at least as fast as $N^{1/2}$ (see Assumption 4(d) therein). In bradic2021high, the additional doubly robust step reduces the number of required product rate conditions to just two. Specifically, only the within-period products need to be considered, making the cross-period product between $\hat{\mu}_{\mathbf{g}_2}(\mathbf{X}_{1})$ and $\hat{p}_{g_1}(X_0)$ no longer necessary. As will be demonstrated later, the practical significance of the weakened assumption appears minimal within the context of the following empirical application.
Both methods were originally developed for static policies under dynamic confounding and are extended here to accommodate dynamic policies. While the identification assumptions differ between static and dynamic policies as discussed in Section (ref), both identification results have the same functional form, with $\mathbf{d}_2$ replaced by $\mathbf{g}_2(\mathbf{V}_1)$. DML is an estimator plugging nuisance estimates directly into the empirical representation of the orthogonal moment condition. These nuisance functions can be estimated in the same way for both static and dynamic policies. Moreover, dynamic policies are deterministic functions of the same covariates already appearing in the estimator for static policies under dynamic confounding. Hence, the estimators naturally extend to dynamic policies, allowing estimation to proceed in the same manner regardless of whether the interest lies in $\theta^{\mathbf{d}_2}$ or $\theta^{\mathbf{g}_2}$. For completeness, Appendix (ref) provides a proof of Neyman orthogonality for the setting with dynamic policies. A recent exposition of sequential DML in the context of dynamic policies is provided by wager2024causal.
The presented estimators offer ready-to-implement algorithms based on robust theoretical foundations under standard assumptions. In addition, several alternative machine learning estimators based on orthogonal scores have been proposed in the literature, including singh2021finite, chernozhukov2022automaticdyn, van2012targeted, lewis2020double. However, they rely on stronger assumptions or lack practical implementability in the present setting. A detailed overview of these methods is provided in Appendix (ref).
Active labor market policies (ALMP) are government programs provided to unemployed individuals with the objective of facilitating their return to the labor market. While much of the existing literature has focused on the effects of the first or the longest program, this study investigates two practical aspects among participants: the duration and the order of the programs.
A limited number of prior studies have analyzed ALMP sequences under dynamic confounding. lechner2013does offers the first comprehensive evaluation using Austrian data and inverse probability weighting. The authors find that jobs search assistance is more effective after a qualification program compared to the reverse order, and that two consecutive qualification programs outperform a single one. Adopting the same approach, dengler2015effectiveness,dengler2019effectiveness find similar results for public employment programs and classroom training in Germany. More recently, laffers2024locking analyze ALMP sequences for young adults in Slovakia using the DML estimator of bodory2022evaluating, also reporting positive effects from combining programs. However, all mentioned studies are restricted to static policies, which makes the resulting treatment effects hard to identify and difficult to interpret. For example, laffers2024locking define unemployed non-participants and employed non-participants as two distinct control groups, which imposes strong assumptions: the former requires fixing unemployment duration a priori, while the latter assumes that exiting unemployment is conditionally independent of potential outcomes. These limitations are addressed in the present paper by introducing dynamic policies.
Complementary work by vikstrom2017dynamic examines ALMP sequences in Sweden using a survival time framework, where the outcome of interest is the probability of remaining unemployed up to a specific period, which differs from this paper's emphasis on medium-term labor market outcomes. In contrast to earlier findings, this study suggests that participating in two consecutive programs offers no substantial advantage over a single program. Finally, another strand of literature considers dynamic assignment to ALMP, accounting for potential non-randomness of program start dates sianesi2004evaluation, crepon2009active, van2019long, kastoryano2022dynamic. While these papers adapt their identification and estimation procedures to dynamic confounding, they remain within a framework that compares single ALMP but does not allow the evaluation of program sequences.
In Switzerland, ALMP are nationally regulated but implemented by regional employment offices (REOs). Unemployed individuals who have worked at least 12 months in the previous two years can register at a REO to receive income maintenance based on their past salary for up to 24 months. To receive these benefits, individuals must actively search for a job and participate in assigned ALMP. The programs can be categorized into five groups:
Two thirds of individuals take part in one of these programs within the first twelve months of their unemployment spell, with approximately 60% of them participating in more than one program. Programs are assigned by caseworkers using information from the registration process and consultation sessions. They have the authority to sanction individuals who decline to participate.
The analysis is based on Swiss administrative records for the period 2004 to 2018. The population under study is defined as individuals aged 25-55\footnote{Younger and older individuals are excluded to avoid dealing with educational and (early) retirement choices.} who became eligible for programs and received unemployment benefits between April 2011 and January 2015,\footnote{The time frame is restricted by a major revision of Swiss unemployment insurance in early 2011 and the need for a follow-up period of at least three years after program start to measure outcomes.} following a minimum of three months of prior unsubsidized employment. Among these individuals, all individuals with a program lasting at least five business days and starting within 12 months of the beginning of their unemployment spell are selected.\footnote{These restrictions ensure that assessments conducted before the allocation to a program are excluded and that there is sufficient time left to participate in a program within the entitlement period.} For the final sample of 191,619 individuals, monthly information on employment status, program participation and covariates is observed. For details on the institutional background and data, see mascolo2024heterogeneous, who base their study on the same public records.
The information is aggregated to a panel of two three-month periods, starting from the month of the first program. This structure has been chosen to meet two key requirements: Firstly, it should capture as good as possible the true assignment process. Secondly, it should ensure a sufficient number of observations per program sequence to enable a meaningful econometric analysis.\footnote{Since aggregation inherently introduces inaccuracies, maintaining the dataset at the highest possible level of granularity would be ideal. When opting against aggregation, an alternative approach involves using estimators relying on parametric assumptions, such as marginal structural models robins2000marginal, which enable extrapolation into regions beyond observed data support. However, given limited knowledge about the true data-generating process in this application, an aggregation approach was selected instead, prioritizing flexible estimation.} Defining the panel’s reference point as the start month of the first program, rather than the start month of the unemployment spell, ensures that all individuals receive a treatment in the first period. This approach substantially increases the number of individuals with identical program sequences.\footnote{For example, consider two individuals who are unemployed for ten months. Individual 1 is assigned to a six-week training course in the first month of the unemployment spell. Individual 2 is assigned to the same program in the fifth month. Despite the different start times of the program, they both exhibit the same program sequence, “training course - no program”.} The lack of sequences beginning with non-participation is of minor concern, as the analysis targets implementation details among participants, while addressing dynamic confounding.\footnote{Alternative designs might be preferable if the focus of the evaluation is when to start programs nie2021learning or program duration without considering dynamics Imbens:2000.} The time between the start of unemployment and the first program is controlled for as a pre-treatment covariate.
Period length is determined in a data-driven way. The sample reveals a median interval of 3 months between program starts and a median program duration of 45 days. Opting for three-month periods maximizes the number of program starts in the second period, compared to two-month or four-month intervals. With a three-month period length, individuals have on average $2.1$ appointments with their caseworker in the first period, which seems reasonable since the assignment decision is expected to be reconsidered less frequently than at every single meeting. In addition, the proportion of individuals with changes in covariates $X_1$ relative to $X_0$ increases considerably when moving from 2-month to 3-month intervals, but shows a much smaller increase when comparing 3-month to 4-month intervals. This suggests that a three-month period is sufficient for time-varying selection to occur.
Finally, the number of periods is another key factor influencing the range of possible program sequences. With each additional period, the number of sequences for a given set of programs increases exponentially. This makes it increasingly difficult to find individuals with the same sequence as the number of periods rises. In the available data, sample sizes for most sequences proved insufficient when considering $T > 2$ periods. Therefore, the analysis is restricted to two periods, consistent with the formal notation introduced earlier.
To evaluate the effectiveness of different program sequences, the outcome is defined as the sum of months employed over 30 months starting from the second period. This measure captures the medium-term impact of program sequences on employment trajectories, which aligns with the Swiss government’s policy objective of promoting sustained employment through ALMP.
In each of the two periods of the panel, each individual is assigned to a treatment state corresponding to one of the four programs JA, TC, EP, or WS introduced above. For a program to be considered a treatment state, it must last at least five business days within the period, regardless of its start date. For instance, if an individual participates in an employment program for five months, the treatment state in the second period is EP, even if the program started in the first period. Individuals participating in multiple different programs within the same period are assigned to the longest program.\footnote{Hence, within-period dynamics in program assignment are disregarded, i.e. it is assumed that assignment to a second program within a period does not depend on the first program within the same period. This seems reasonable given the limited duration of the periods.} In the second period, two additional treatment states are introduced. Individuals who remain without program participation throughout the period are categorized as No program (NP). This group includes both unemployed individuals who are no longer assigned to a program and individuals who are not assigned because they have exited unemployment. This additional treatment state allows analyzing sequences with a program in the first period but no program in the second period. Individuals who are assigned to \textit{OP} in the second period remain in the sample for modeling treatment assignment in the first period. However, they are not considered in the analysis due to the small number of participants and special admission requirements for these programs, which prevent credible identification.
Figure (ref) illustrates the distribution of treatment states across the two periods, revealing insights about program size and duration. In the first period, temporary wage subsidies comprise nearly half of the beneficiaries, followed by job-search assistance. Most recipients of temporary wage subsidies in the first period continue in the second period, while recipients of job-search assistance often transition to other program states. Overall, more than half of the individuals remain in a program in the second period. The plot highlights the large variety of transitions, emphasizing the importance of sequential analysis.
When evaluating ALMP over time, program assignments critically depend on the evolution of employment status. For instance, if an individual exits unemployment during the first period, this directly influences program allocation decisions in the second period. Consequently, evaluating counterfactuals that impose a fixed sequence of two program periods $(d_1, d_2)$ for all individuals, without accounting for changes in employment status, may lack meaningful insights. The present analysis solves this issue using dynamic policies. Specifically, as previewed earlier in Example (ref), the decision variable $V^{d_1}_1$, which determines program allocation in the second period, is defined as the potential intermediate outcome $Y_1^{d_1}$. Here, $Y_1^{d_1}$ is a binary indicator equal to one if an individual treated with program $d_1$ exits unemployment during the first period, and zero otherwise. Using this decision variable, 20 dynamic policies of interest are defined as
with $d_1, d_2 \in \left\{\text{JA}, \text{TC}, \text{EP}, \text{WS}\right\} \times \left\{\text{JA}, \text{TC}, \text{EP}, \text{WS}, \text{NP}\right\}$. The policies are dynamic only in the second period, while they assign a fixed program $d_1$ in the first period. This allows to construct counterfactuals in which every individual is initially assigned to program $d_1$ during the first period, with program $d_2$ assigned in the second period only to those who remain unemployed throughout the first period. The treatment $d_2$ may also be defined as NP, in which case the policy is fully static. Policy ((ref)) is defined in terms of potential intermediate outcomes $Y_1^{d_1}$ as the assignment to the second program should depend on the employment status after the first counterfactual program $d_1$ and not on the employment status following the observed program $D_1$.
For single time-point interventions, the literature widely acknowledges that the effects of ALMP can be plausibly identified using observational data, provided a comprehensive set of control variables is available caliendo:2017, lechner2013sensitivity. Since individuals are randomly assigned to caseworkers, who then determine program assignments, the primary challenge lies in accounting for all factors considered by caseworkers during this process. In the current setting, the relevant covariates identified in prior research are available. These include, among others, socio-demographic characteristics, employment and earnings histories spanning the past seven years, details about the last job and prior unemployment spells, regional labor market conditions, and caseworkers' assessments of individuals' job search efforts prior to program start. Appendix (ref) provides an overview of all control variables used.
For the analysis of sequential treatments, the identification assumptions must be extended to address static or dynamic confounding, as well as static or dynamic policies, as discussed in Section (ref). Under static confounding, Assumption (ref) requires the entire program path to be predetermined, given the available information before the start of the sequence, which seems unlikely, as caseworkers may adjust their assignments over time. This concern is supported by an analysis of covariate means by assigned program type, which reveals substantial differences across program sequences for both baseline and time-varying covariates (see Tables (ref) and (ref) in the Appendix). Therefore, this analysis focuses on the setting under dynamic confounding, which permits a more flexible program assignment process in the underlying data.
For static policies under dynamic confounding, Assumption (ref)(ref) consists of two parts, one for every period. In the first period, the identification argument aligns with that of a standard non-sequential evaluation and should therefore be reasonable given the available information. The additional difficulty of dynamic confounding is addressed by the second part of the assumption, which requires controlling for all factors that influence both potential outcomes and program assignment in the second period, conditional on program participation in the first period. To address this assumption, five types of covariates are considered, that may change during the first period and are expected to determine dynamic selection into the second program. Firstly, it is essential to control for the intermediate outcome, employment status, as only unemployed individuals are eligible for assignment to a program in the second period. Secondly, financial aspects are taken into account by controlling for the amount of unemployment benefits and other state subsidies, as well as intermediate earnings during the first period. Thirdly, the selection likely hinges on an individual's job prospects. These are assessed through variables such as the number of applications written and caseworkers' assessments of individual's job search efforts, employability, and qualification needs. Fourthly, caseworkers' decisions might be influenced by the cooperativeness of individuals. This is measured through the number of scheduled, postponed, and canceled appointments at the REO, as well as incidents resulting in sanction days. Lastly, changes in the personal situation such as relocation, pregnancy, and the number of sickness days, are also accounted for.
A key difficulty in identifying static program sequences is that program participation depends on remaining unemployed. Consequently, both unemployment status and program assignment must be accounted for when addressing selection bias. This essentially requires that, conditional on covariates and prior program participation, not only program assignment but also the chance of remaining unemployed are random, which makes credible identification significantly more demanding. This issue can be addressed by redefining the estimand in terms of dynamic policies, as defined in equation ((ref)). These policies account for the possibility of leaving unemployment in the counterfactual scenario, such that assignment to the treatment group is no longer conditional on employment status. However, this advantage comes at the cost of the stricter identification requirements stated in Assumption (ref). In particular, conditional on pre-treatment information, assignment to the first program needs to be independent of final and intermediate potential outcomes. Given the context of this application, the additional restrictiveness of the assumption appears to be unproblematic. For static policies, the conditional independence assumption applies to the final potential outcome, i.e. the joint distribution of potential employment status over the 30 months starting from the second period. The additional complexity introduced by the dynamic strategy necessitates independence with respect to employment status during the three months of the first period as well. This restriction aligns with the assumption made in any single-period program evaluation that measures outcomes from the start of the first treatment.
Finally, the overlap assumptions (ref)(ref) and (ref)(ref) require that, in both periods, there is a non-zero probability of following the (static or dynamic) policy of interest, conditional on pre-period information. This assumption is expected to hold among unemployed individuals, as all the programs considered are accessible to them, irrespective of their covariate history. Thus, although treatment probabilities may vary based on previous characteristics, any unemployed individual should, in principle, be eligible for admission to any program. Again, the fact that individuals might exit unemployment during the first period poses a challenge to the credibility of the assumption for static policies. For example, individuals who leave unemployment in the first period are only eligible for treatment in the second period if they re-enter unemployment within that three-month window, resulting in a very low treatment probability. This issue does not arise for dynamic strategies, as such situations are excluded under any treatment strategy of interest. Specifically, for any realization of the function $g_2(Y_1^{d_1})$, individuals employed in the first period are not assigned to a program in the second period. Hence, carefully designing the structure of the dynamic policy allows to obtain counterfactuals that are more credibly identified.
To estimate the effects of interest, the two dynamic estimation procedures by bodory2022evaluating\footnote{Note that the implementation of the procedure by bodory2022evaluating used in this paper differs from that in the original causalweight R-package. In particular, bodory2022evaluating fit $\hat\mu$ stratified by programs but include $D_1$ as a covariate when estimating $\hat p_{d_2}$. In contrast, the approach adopted here stratifies both $\hat\mu$ and $\hat p_{d_2}$. This adjustment is expected to reduce bias in the estimates of $\hat p_{d_2}$, but may increase variance if there are few individuals following a particular program $d_1$.} and bradic2021high described in Section (ref) are applied with 5-fold cross-fitting. The nuisance functions are estimated by random forests using the Python package scikit-learn pedregosa2011scikit. In line with standard practice, predictors with low variance or high collinearity are removed prior to the analysis. Following bach2024hyperparameter, the hyperparameters of the random forest are tuned on the full sample using FLAML wang2021flaml with a maximum time budget of 10 minutes per nuisance estimation. The minimum number of trees used in a forest is set to 500.\footnote{Python code implementing the procedures is available at \href{https://github.com/fmuny/dynamicDML}{https://github.com/fmuny/dynamicDML}}
Estimators based on propensity score reweighting, such as DML, are known to produce unstable estimates when estimated propensity scores assign large weights to certain observations khan2010irregular. Hence, even if overlap holds in the population, it may not be satisfied in the sample if the treatment probabilities for certain programs are very low for specific individuals. To address this issue, observations with extreme propensity scores are dropped from the sample, extending the minmax trimming procedure proposed in lechner2019practical to the sequential setting. In particular, for the first period propensity scores $\hat p_{d_1}(X_0)$, first the minimum value among individuals within the subgroups $D_1 = d_1$ and $D_1 \neq d_1$ are identified, respectively. Then, the largest of these two values is selected, and all individuals with a $\hat p_{d_1}(X_0)$ smaller than this threshold are removed from the sample. Analogously, all units with $\hat p_{d_1}(X_0)$ larger than the smallest maximum are also removed. For the second period propensity scores $\hat p_{g_2}(\mathbf{X}_1, d_1)$ the same procedure is applied for individuals with $(D_1 = d_1$ and $D_2 = g_2(Y_1))$ vs. $(D_1 \neq d_1$ or $D_2 \neq g_2(Y_1))$, separately for the subgroups with $g_2(Y_1) = d_2$ and $g_2(Y_1) = \text{NP}$, respectively. After repeating the procedure for all 20 dynamic policies of interest, 7% of the observations are dropped, resulting in a final sample size of 177,856 observations. A comparison of pre-treatment covariate means shows that trimmed units are more likely to come from the German-speaking region of Switzerland, have longer prior unemployment durations, and have received greater previous government support. Nonetheless, overall differences between trimmed and retained observations appear moderate, as shown in Appendix Table (ref). Plots of the propensity score densities after trimming for dynamic policies are presented in Figures (ref) and (ref).
Table (ref) presents the main results of the analysis. It illustrates average treatment effects for different comparisons of policies and estimation techniques. Each column refers to a comparison of programs in the first period while in the panels below, several scenarios of second period policies are compared.\footnote{While Table (ref) only reports ATEs and standard errors, Table (ref) in the Appendix provides an extended version including the trimmed number of observations for each policy and results for both dynamic DML estimators discussed in Section (ref).} Panel A provides a setting where the second period programs are unrestricted. These are the results one would obtain from a standard single-period evaluation analyzing the first program only. The discussion first focuses on these baseline results before highlighting the added value of the dynamic evaluation methods. The results indicate that WS is the most effective program on average, leading to significantly more months in employment in the medium term compared to all other programs. The effect sizes range from 2.65 months more employment compared to JA to 1.94 months more employment compared to TC in the 30 months starting from the second period. TC is identified as the second most effective program, yielding 0.71 and 0.59 months of increased employment compared to JA and EP, respectively. No statistically significant difference is observed between the programs \textit{JA} and \textit{EP}.
The median duration of the four programs, measured by the number of business days, are 23 for JA, 39 for TC, 80 for EP, and 65 for WS. A natural question arising from these differences is whether program duration affects effectiveness, i.e., whether some programs are superior to others due to their varying lengths. One approach to address this question is to use the sequential framework, comparing sequences as if individuals only participated during the first period.\footnote{An alternative approach to account for program duration would be to estimate a continuous treatment effect Imbens:2000. This framework, however, requires program duration to be determined before treatment start, while here it is allowed that program duration can be updated once in-between. Repetitions of the same program type are also considered.} Policies of this type are not dynamic as NP can occur in the second period regardless of whether an individual is unemployed or not. The results are presented in Panel B of Table (ref). The first row of the panel presents the ATE estimated using the dynamic DML method of bodory2022evaluating, while the second row shows estimates ignoring dynamic confounding.
The results show that when restricting program duration to a maximum of three months, the ranking of the programs slightly changes. While WS remains the most beneficial and JA the least beneficial program, there is no longer a significant difference between TC and EP, when considering dynamic confounding. Furthermore, aligning program duration enhances the advantage of WS relative to the other programs, for instance, by 1.5 months compared to JA and 0.8 months compared to \textit{TC}. When comparing dynamic to static confounding, a systematic over-estimation of ATEs is observed for programs with a larger average duration. For example, when comparing the shortest program (\textit{JA}) to the longest program (\textit{EP}), the relative effectiveness of the latter almost doubles when dynamic confounding is ignored (1.47 vs. 2.74). Consequently, disregarding the feedback between treatments and intermediate outcomes results in an overly optimistic assessment of longer programs. Furthermore, Appendix Table (ref) shows that results from the methods of bodory2022evaluating and bradic2021high yield very similar estimates, indicating no clear advantage of either method in this application.
Instead of restricting programs to the first period, another approach to align program duration is to require at least two consecutive periods of the same program. Panel C of Table (ref) presents estimates for the static scenario in which all individuals are assigned to the same program in both periods. Instead, Panel D presents results for dynamic policies in which individuals are reassigned to the same program only if they remain unemployed throughout the first period. Overall, the results show patterns similar to those already observed for shorter or unrestricted sequences. Compared to programs lasting only one period, longer programs tend to slightly narrow the differences in effectiveness. However, these reductions are modest, suggesting that program duration is unlikely to be the primary driver of the observed differences between programs. The comparison between static and dynamic policies under dynamic confounding reveals that effect sizes increase for all comparison when considering dynamic policies. For example, under static policies, two periods of WS result in 2.86 months more employment than JA, whereas under dynamic policies, the difference increases to 3.51 months. This underscores that the choice of counterfactual can have economically significant implications.
Besides duration or multiple participation, the sequential treatment effect framework can be exploited to obtain insights on the effective ordering of programs. As seen in Figure (ref), a considerable number of individuals is assigned to different programs across the two periods. In such cases, it could be interesting to assess whether reversing the order of two programs leads to better outcomes. Table (ref) presents the results of such an analysis. Panel A and B present the average treatment effects implementing a specific pair of programs, compared to implementing the same pair in reverse order, assuming static and dynamic policies, respectively.
The results for static policies indicate that when WS is combined with TC or EP, it should be implemented as the second program rather than the first. This suggests that quitting a subsidized job to join an alternative program is less effective than first participating in the alternative program and then transitioning to the subsidized job. However, under dynamic policies, all effects become smaller and statistically insignificant. Hence, the previous conclusion holds only if individuals remain unemployed for at least two periods, which is unknown a priori. Instead, when allowing for the possibility of reemployment during the first period in the counterfactual, no clear superiority of any combination is observed. This suggests that conclusions drawn from ALMP evaluations based solely on static policies may not be reliable.\footnote{A challenge with analyzing the order of programs is that the second program may extend into subsequent periods, resulting again in comparisons of programs with differing durations. While a three-period setup with a final NP period, as in lechner2013does, could address this issue, it is not used here due to limited sample sizes.}
A key advantage of causal machine learning methods is that they allow to flexibly analyze effect heterogeneities. In the DML setting, heterogeneous effects can be obtained by aggregating the estimated scores $\hat{\Theta}^{dy}_{\mathbf{g}_2}$ for specific sub-groups of interest. This analysis focuses on two types of heterogeneities: local language skill level and prior program participation in previous unemployment spells. Language skills, a proxy for migration background, are included since significant effect heterogeneity has been observed for this variable in multiple previous studies Cockx:2020. Previous program participation is included to enable an even more detailed analysis of program combinations. Of course, the procedure could be applied to any other discrete pre-treatment characteristic, provided there is a sufficiently large sample size.
The first row of panels in Figure (ref) presents the results for heterogeneities based on local language knowledge. Specifically, estimates of the difference GATE-ATE are shown for dynamic policies ((ref)) with $d_2=d_1$, estimated by the method of bodory2022evaluating. The effects are presented as the difference between GATE and ATE, where a significant deviation from zero indicates significant heterogeneity relative to the average.\footnote{As derived in Appendix (ref), standard errors for this difference can be computed as the square root of $\operatorname{Var}(\hat \theta (z_0)-\hat \theta) = \operatorname{Var}(\hat \theta (z_0))+\operatorname{Var}(\hat \theta) - \frac{2N_{z_0}}{N} \operatorname{Var}(\hat \theta (z_0))$, where $N_{z_0}$ denotes the number of observations with $Z_0=z_0$.} The results show that for this type of long programs, TC is particularly effective for individuals fluent in the local language, whereas those with limited language skills benefit significantly less than the average. This finding holds true when comparing TC to each of the other programs. The result is surprising, given that TC includes language courses alongside various other training programs. In addition, individuals with none to intermediate proficiency in the local language benefit significantly more than average from WS compared to TC, suggesting that extended programs incorporating work experience may be more effective than extended training courses for those with limited language skills.
The second row of Figure (ref) presents the same analysis using information about unemployment spells in the five years prior to program start. The results are mostly close to zero and insignificant with a few exceptions. In particular, individuals who participated in WS during a previous unemployment spell benefit above-average from programs other than WS. This suggests that repeated participation in WS across multiple unemployment spells leads to reduced program effectiveness. In addition, individuals who have not been unemployed in Switzerland before, profit less from TC compared to WS (significant at 10% level). This aligns with the findings for local language knowledge, as recent immigrants are overrepresented among the first-time unemployed.
This paper reviewed, explained, and applied methods for the evaluation of program sequences and introduced the concept of dynamic policies to the econometric program evaluation literature. By summarizing the identification process under static and dynamic confounding, it was demonstrated that assessing dynamic policies allows for the construction of more realistic counterfactuals, requiring only minor adjustments to identification assumptions. In addition, the analysis illustrated how dynamic DML can be employed to flexibly estimate the effects of dynamic policies. The presented methods provide a foundation for more effective policy design in settings where program assignments depend on time-varying characteristics.
The empirical application evaluated sequences of Swiss ALMP across two consecutive periods, starting at the beginning of the first program. An initial descriptive analysis revealed a significant variety of program sequences, highlighting the necessity of a sequential analysis. Using standard results from single time point interventions as a benchmark, it was demonstrated how an assessment of sequential policies can provide additional insights into implementation details of ALMP. This revealed that Temporary Wage Subsidies are most effective on average, even after adjusting program duration across different program types. In particular, individuals with limited proficiency of the local language profit more from programs related to obtaining work experience in comparison to extended training courses. In the application, disregarding dynamic confounding resulted in an overly optimistic assessment of longer programs. Moreover, the choice between static and dynamic policies was shown to result in economically and statistically significant differences in estimated effect sizes.
Overall, DML-based estimation of effects of dynamic policies turns out to be a valuable addition to the standard program evaluation toolkit. A key limitation of any sequential analysis is the need for sufficient sample sizes for all sequences of interest to ensure robust estimation. This challenge is particularly pronounced for flexible machine learning methods, which avoid parametric structural assumptions but require large datasets to capture non-linear relationships. In the application, this constraint necessitated the aggregation of program categories into four major groups. This aggregation potentially masks distinct effects of individual sub-programs, limiting the detail of the findings. For the same reason, only a few relatively long periods could be used in the panel, which might not perfectly capture the true dynamic selection process. Although sequential methods are data-intensive, their potential is expected to increase as larger datasets become available.