EconBase
← Back to paper

Residualised Treatment Intensity and the Estimation of Average Partial Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

87,243 characters · 18 sections · 33 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Residualised Treatment Intensity and the Estimation of Average Partial Effects

titlepage\begin{abstract} This paper introduces R-OLS, an estimator for the average partial effect (APE) of a continuous treatment variable on an outcome variable in the presence of non-linear and non-additively separable confounding of unknown form. Identification of the APE is achieved by generalising Stein's Lemma stein1981estimation, leveraging an exogenous error component in the treatment along with a flexible functional relationship between the treatment and the confounders. The identification results for R-OLS are used to characterize the properties of Double/Debiased Machine Learning Chernozhukov_et_al_2018_DebiasedML, specifying the conditions under which the APE is estimated consistently. A novel decomposition of the ordinary least squares estimand provides intuition for these results. Monte Carlo simulations demonstrate that the proposed estimator outperforms existing methods, delivering accurate estimates of the true APE and exhibiting robustness to moderate violations of its underlying assumptions. The methodology is further illustrated through an empirical application to fetzer2019_austerity_brexit. \\ \\ \noindentKeywords: average partial effect, distributional moments, residualisation, identification, ordinary least squares, debiased machine learning \\ \end{abstract} \setcounter{page}{0} \thispagestyle{empty}

Introduction

This paper is concerned with identification and estimation of the average partial effect $E[\partial Y/\partial X]$ for the model $Y = g(X,Z) + \varepsilon$, where $X$ is a continuous treatment variable of interest and $Z$ is a set of confounders. A novel estimator, R-OLS, is presented, which exploits an exogenous error component of the treatment to estimate the average partial effect (APE). Notably, R-OLS is suitable to estimate treatment effects in models characterized by non-linear and non-additively separable confounding of unknown form.

The estimator adds to two strands of the literature. The first one aims to identify and estimate the average partial effect within the aforementioned model by imposing restrictions on $g(X, Z)$. For instance, the canonical linear regression model assumes linear and additively separable confounding and accounts for confounders by adding them as controls. Building upon this, Robinson_1988 extends the framework to the partially linear model $Y = \beta X + \theta(Z) + \varepsilon$, accommodating additively separable but potentially non-linear confounding. Taking a further step, Graham_Pinto_2022 incorporate heterogeneous treatment effects by allowing for interactions between $X$ and $Z$ in the Y-DGP, but maintain the linearity of $X$. Without linearity in $X$, their estimator recovers a weighted average of derivatives of Y wrt. X with unknown weights.

A second strand of the literature achieves identification of the average partial effect by placing restrictions on the joint distribution of $X$ and $Z$. In a setting without confounding, Stein's Lemma stein1981estimation, Ross_Steins_Method shows that a linear regression of $Y$ on $X$ identifies the average partial effect if and only if $X$ is normally distributed. Under confounding, joint normality of $X$ and $Z$ enables the identification of the APE Stoker_1986,Powell_et_al_1989. LANDSMAN2008 extend Stein's Lemma to elliptical distributions, but only estimate the APE for a transformed variable $X^*$, resulting in a reweighting of the original APE. Considering endogeneity and unobservables, Cuesta_Steins_Lemma show that joint normality of the treatment and an instrumental variable enable the IV-Estimator to estimate the APE under additive separability of unobservables.

The R-OLS estimator combines elements of the previous two approaches and leverages the trade-off between the complexity of the $Y$-DGP and distributional assumptions on the covariates. Through the lens of the first strand of literature, restricting $g(.)$, it enables a more flexible functional form for $g(.)$, at the cost of making assumptions on the exogenous variation in $X$. Conversely, viewed within the context of the second strand of literature R-OLS relaxes the assumption of joint normality of covariates considerably, but restricts $g(.)$. Therefore, the R-OLS estimator can be seen as a combination of both previous approaches, allowing more flexibility in certain assumptions while imposing constraints on others. Furthermore, due to the structure of the assumptions made, R-OLS allows for customisation of the trade-off to the empirical application, since the flexibility in the $Y$-DGP is directly linked to the extend of the distributional assumptions imposed.

Concretely, the average partial effect of treatment $X$ on $Y$, $E\left[ \partial_{X_i} Y_i \right]$, is identified by a bivariate regression of $Y_i = g(X_i,Z_i) + \varepsilon_i$ on the exogenous variation $\nu_{i}$ in $X_i = r(Z_i) + \nu_{i}$. Here, $r(Z_i)$ can be any function describing the confounding between the treatment $X_i$ and covariates $Z_i$. The Y-DGP is characterised by $g(X_i, Z_i)$ and assumed to be an interaction of polynomials of the treatment and arbitrary functions of the confounders.

Under these assumptions, conditions on the distributional moments of $\nu$ establish an equivalence between R-OLS and the average partial effect of the treatment. The conditions on the distributional moments are directly related to the order of polynomials in the outcome DGP and hence enable a flexible trade-off between both assumptions. On one side of the trade-off, the distributional assumptions on $\nu$ only impose a mean of zero, but the $Y$-DGP is required to be linear in $X$. On the other end, the exogenous error $\nu$ unconditionally follows a normal distribution with mean zero, but the average partial effect is identified without placing assumptions on the DGP of $Y$. Specifically, if one is willing to assume a $Y$-DGP quadratic in $X$, then the exogeneous error just needs to have a first and third moment of zero to guarantee that R-OLS estimates the APE under arbitrarily complex confounding.

Similar to Robins_et_al_1992's E-estimator and MJ_Lee_2018's propensity score residual estimator, R-OLS is particularly useful under low-complexity treatment DGPs where the prediction of the error component is feasible or the DGP known. However, compared to earlier work, this paper goes one step further and shows that the average partial effect of a continuous treatment is identified by R-OLS, even under a more flexible semi-parametric model for the outcome, if certain distributional assumptions hold.

Following these identification results, it is shown how the average partial effect can be estimated under knowledge of the treatment DGP using the Frisch-Waugh-Lovell theorem. The asymptotic distribution of the estimator is derived and an estimator of the variance provided. Afterwards, Double/Debiased Machine Learning (DML) Belloni_et_al_Lasso_2014, Chernozhukov_et_al_2018_DebiasedML, chernozhukov2022autodml is leveraged to enable statistical inference for the APE in settings without knowledge of the treatment DGP by showing that, under the assumptions of R-OLS, the Neyman-orthogonal moment for the partially linear model estimates the APE.

Intuition for R-OLS is provided via a novel decomposition of OLS and IV estimands, showing that both estimands approximate the true outcome DGP with a Taylor expansion, take the derivative of said polynomials and apply weights to each polynomial. However, these weights do not equal unity and hence distort estimates away from the APE. The assumptions made in this paper guarantee that these weights equal unity and hence enable the estimation of the APE.

Simulation experiments demonstrate that R-OLS accurately estimates the APE even if the samples are small and/or the $Y$-DGP highly complex, improving on existing semi-parametric approaches. Estimates remain close to the true APE even under moderate violations of assumptions.

A real-world application re-estimating the effect of austerity on the UK Brexit referendum in fetzer2019_austerity_brexit is provided. The average partial effect of austerity on UKIP voteshare in local elections confirms the original finding, however, no statistically significant effect on European elections is detected.

The remainder of the paper is structered as follows: Section (ref) formalizes the set-up and states the main result. Section (ref) discusses the feasibility and asymptotic distribution of the estimator. Section (ref) decomposes the weights underlying R-OLS estimation. Simulation results for a variety of data generating processes are presented in Section (ref), and real-world applications are provided in Section (ref). Finally, Section (ref) concludes.

Model Assumptions and Estimand

The general estimation procedure underlying R-OLS is as follows: Assume $X$ is determined by a function of covariates with arbitrary dependencies and some additive exogeneous variation that is related to $Y$ only through $X$. Since changes in the exogenous variation only influence $X$, leaving everything else constant, regressing $Y$ on this exogenous variation identifies the a causal effect of $X$ on $Y$. This section shows that the average partial effect can be estimated this way.

Assumptions

We begin by describing the data generating process for which we want to estimate the average partial effect of $X$ on $Y$.

assumption[Data Generating Process] For independently and identically distributed random vectors $(Y_i, X_i, Z_i)$ it holds that \begin{align} &Y_i = \sum_{m=0}^M X_{i}^mg_m(Z_{i}) + \varepsilon_i \quad for \quad M \in \mathbb{N} \\ &X_{i} = r\left(Z_{i}\right)+\nu_{i} \end{align} with $X_i \in \mathbb{R}$, $Z_i \in \mathbb{R}^K$, $E[\varepsilon_i \mid X_i, Z_i] = 0$, $\nu_{i} \perp\!\!\!\perp Z_{i}$, $E[|\nu_i^2|] < \infty$ and $E[|\nu_iY_i|] < \infty$. Furthermore, $g_m: \mathbb{R}^{K} \to \mathbb{R}$ and $r: \mathbb{R}^{K} \to \mathbb{R}$.

Assumption (ref) defines the data generating process of $Y_i$ to be a polynomial of the independent variable of interest, $X_i$, multiplied by an arbitrarily complex function $g_m(Z_i)$ of all the other independent variables $Z_i$, where $Z_i$ can differ between equations ((ref)) and ((ref)). Although seemingly restrictive, the Y-DGP specified in Assumption (ref), is quite general. Firstly, $g_m(Z_{i})$ can be any function and differ with $m$. Secondly, by the Weierstrass Approximation Theorem every continuous function defined on a closed interval $[a, b]$ can be uniformly approximated as closely as desired by a polynomial function. Therefore, the functional form of the Y-DGP in Assumption (ref) allows for arbitrary flexibility in the outcome data generating process, if $M\to \infty$.

The variable of interest, $X_i$, is assumed to be determined by an unrestricted function $r(Z_i)$ of all other independent variables plus an exogenous error $\nu_i$.\footnote{Throughout the paper $\nu_i$ is referred to as exogenous error, error or exogenous variation. The estimate $\hat{\nu_i}$ will be referred to by residual variation or residual. In the machine learning literature $\nu_i$ is referred to as the irreducible error and $\hat{\nu_i}$ is called the prediction error.} The estimand presented in this paper does not require the functional form of $r(Z_i)$ to be known, but rather assumes that $\nu_i$ is known. This has implications for settings in which $X_i$ can be manipulated or its exogenous variation is known. When describing our estimation procedure, we will leverage orthogonalised moment conditions and estimate $\nu_i$ with machine learning.

The exogenous error $\nu_i$ is the main ingredient in the estimation procedure outlined in Theorem (ref) and enables the estimation of the average partial effect via a univariate regression under the following assumption:

assumption[Distributional Moments of the Error] The exogenous error $\nu_i$ satisfies the distributional moments \begin{align*} \frac{1}{(p+1) E\left[\nu_{i}^2\right]} E\left[\nu_{i}^{p+2}\right] &= E\left[\nu_{i}^{p}\right] \quad for \quad p \in \mathbb{N} \quad and \quad 0 \leq p \leq M-1 \\ E\left[v_i\right]&=0, \end{align*} where $M$ is determined by the highest order polynomial of $X$ in equation ((ref)).

Assumption (ref) implies that the distribution of the exogenous variation influencing the variable $X$ has odd moments that are equal to zero, while even moments conform to a specific series.

Following Lemma (ref) in the Appendix, Assumption (ref) is fulfilled for normally distributed errors up to $M=\infty$, regardless of the variance, as long as the expected value is zero. However, it is important to note that it is not necessary for the errors to satisfy these moments up to an infinite order; it suffices to adhere to them up to order $M+1$, where $M$ represents the maximum order of the polynomial terms of $X$ in equation ((ref)). For example, in the partial linear model of Robinson_1988 with $M=1$, Assumption (ref) requires that the errors have a mean of zero.

Estimand

Using Assumptions (ref) and (ref), we can define the R-OLS estimand for the average partial effect:

theorem[R-OLS] Under Assumptions (ref) and (ref), the R-OLS estimand: \begin{align} \beta = \frac{E[\nu_{i}Y_i]}{E[\nu_{i}^2]} \end{align} is equivalent to the average partial effect of $X_i$ on $Y_i$: \begin{align} E_{X,Z, \varepsilon}\left( \partial_{X_i} Y_i\right) \end{align}

For a proof of Theorem (ref) see Appendix A. The proof relies heavily on combining the functional form for the Y-DGP in Assumption (ref) with the moments of the exogenous error in Assumption (ref). Together, the assumptions allow matching R-OLS with the average partial effect after careful expansion of the polynomial structure.

Since the Y-DGP in Assumption (ref) is highly flexible and can approximate any functional form, if $M\to \infty$, we can conclude that the R-OLS estimand identifies the average partial effect without assumptions on the Y-DGP, as long as $\nu_i$ follows a normal distribution with expected value of zero. For finite orders of $M$, the distributional moments of $\nu_i$ need to match the respective moments of the normal distribution up to order $M+1$.

Theorem (ref) generalises Stein's Lemma stein1981estimation by leveraging a trade-off between functional form and distributional assumptions. It shows that normality of $X_i$ can be relaxed to the structure of moments in Assumption (ref) depending on the flexibility of the outcome model.

remark[Independence of $\nu_i$] Full independence of $\nu_i$ is not necessary. Only the moments of $\nu_i$ up to $M+1$ need to be independent of $Z_i$. However, even this assumption can be relaxed to mean independence, under an alternative version of Assumption (ref): \begin{align*} \frac{1}{(p+1) E\left[\nu_{i}^2\right]} E\left[\nu_{i}^{p+2} \mid Z_i \right] &= E\left[\nu_{i}^{p} \mid Z_i \right] \quad for \quad p \in \mathbb{N} \quad and \quad 0 \leq p \leq M-1 \\ E\left[v_i \mid Z_i \right]&=0. \end{align*} This means $\nu_i$ needs to satisfy the moments of the normal distribution up to order $M+1$ conditional on $Z_i$, if we also assume $E\left[\nu_{i}^2\right] = E\left[\nu_{i}^2 \mid Z_i \right]$. Given the difficulty of imagining a $\nu$-DGP that satisfies these conditions, but not independence, we assume independence.

Example: $M=2$

To better understand the implications of Theorem (ref), consider the following example. For $M=2$, the DGPs in Assumption (ref) can be written as

equation[equation omitted — 212 chars of source]

Theorem (ref) implies that the R-OLS estimand:

align*[align* omitted — 60 chars of source]

is equivalent to the average partial effect of $X_i$ on $Y_i$ if the distribution of $\nu_i$ satisfies the following two moment conditions:

align*[align* omitted — 76 chars of source]

These conditions hold for the normal distribution or any other distribution that is symmetric around zero. If a cubic term of $X$ was to be added to the DGPs in equation ((ref)), the errors would need to satisfy $Kurtosis(\nu_i)=\frac{E\left[ \nu_{i}^{4} \right]}{E^2\left[ \nu_{i}^2 \right]} = 3$, a property defining the class of mesokurtotic distributions\footnote{Examples of mesokurtotic distributions can be found in Appendix C.}. With a fourth degree polynomial of $X_i$, the errors would also need to fulfill $E\left[ \nu_{i}^{5} \right] = 0$. In general, the higher the order of the $X_i$-polynomial in the $Y$-DGP, the more moments of the error need to satisfy Assumption (ref) and the more the error distribution tends to a normal distribution.

Estimation

The previous section showed how the average partial effect can be estimated with R-OLS under knowledge of $\nu_i$. In this section we will relax this assumption and replace $\nu_i$ with an estimate. We will begin by discussing estimation using OLS, assuming the functional form linking the treatment variable $X_i$ and the confounders $Z_i$ is known. Afterwards, we will consider estimation via machine learning without knowledge of the functional form.

Regardless of which method is used to estimate $\nu_i$, the R-OLS estimator is given by:

align*[align* omitted — 122 chars of source]

Here, the population moments have been replaced with sample moments and $\nu_i$ with an estimate $\hat{\nu}_i$. Assuming the existence of a consistent estimator $\nu_i$, the weak law of large numbers can be applied to the numerator and denominator of the R-OLS estimator.\footnote{If $E[|\nu_i^2|] < \infty$ and $E[|\nu_iY_i|] < \infty$.} Then, by the continuous mapping theorem, $\hat{\beta} \xrightarrow{p} \beta$ as $n \to \infty$, proving the consistency of the R-OLS estimator for the APE following Theorem (ref). The R-OLS estimator itself is agnostic to the method used to predict $r(Z_i)$. The only assumption needed for the estimator to be consistent for the APE is that $\hat{\nu}_i \xrightarrow{p} \nu_i$ as $n \to \infty$.

However, when estimating $\nu_i$ via machine learning, without knowledge of the functional form of $r(Z_i)$, even though the estimator might be consistent, the rate of convergence can be slower than $\sqrt{n}$, making inference difficult.\footnote{Under knowledge of the functional form in OLS, the convergence rate is $\sqrt{n}$ hansen2022econometrics.} To enable inference with machine learning estimators the problem will be recast within the Double/Debiased Machine Learning framework Chernozhukov_et_al_2018_DebiasedML in Section (ref).

Estimation with OLS

We begin by assuming that the functional form of $r(Z_i)$ is known and allows for estimation via OLS. Under this assumption, $X_i$ can be modeled as $X_i \sim r(Z_i)$ and the resulting residuals can be utilised to estimate the APE with R-OLS. By the Frisch-Waugh-Lovell Theorem, this approach is equivalent to directly regressing $Y_i$ on $X_i$ and $r(Z_i)$, where in an abuse of notation $r(Z_i)$ is supposed to be the functional form of the X-DGP. This leads to:

align*[align* omitted — 280 chars of source]

where the R-OLS estimator was rewritten in matrix form using the FWL-Theorem. $[.]_X$ symbolises the element of the vector/matrix corresponding to $X$. $\Omega$ is defined as $\{X, r(Z)\}$ and $A$ is the coefficient vector in the population level regression of $Y$ on $\Omega$. $[A]_X=E\left[ \partial_{X_i} Y_i\right]$ by the FWL-Theorem and Theorem (ref). This estimator has the asymptotic distribution

align*[align* omitted — 254 chars of source]

under the assumptions that $E[Y^4] < \infty$, $E[|\Omega|^4] < \infty$ and $\Omega'\Omega$ positive definite. $\bm{V}$ can be estimated by replacing the population moments with their sample counterparts and $A$ with $\hat{A}$, the estimated coefficient vector from the regression of $Y$ on $\Omega$.

Simplifications of the variance under homoscedasticity are not recommended since they implicitly impose strong assumptions on the $Y$-DGP. Similarly to a FWL-type regression, heteroscedasticity robust standard errors for the coefficients have to be estimated. This can easily be seen by plugging Assumption (ref) into $E[\Omega\Omega'U^2]$:

align*[align* omitted — 153 chars of source]

Here, the expectation can not be split into $E[\Omega\Omega']$ and $E[(\sum_{m=0}^M X^mg_m(Z) + \varepsilon - \Omega A)^2]$ since $\sum_{m=0}^M X^mg_m(Z)$ and $\Omega A$ do not cancel out, creating dependence between both parts of the expectation and invalidating the use of homoscedastic standard errors.

Naïve Estimation with Machine Learning

If the functional form of $r(Z_i)$ is unknown, a machine learning estimator can be used to estimate it. For consistency of R-OLS, the estimation procedure must satisfy $r_n(Z_{i}) \xrightarrow{p} r(Z_{i})$ as $n \to \infty$. One algorithm satisfying these conditions is given by neural networks. Following HORNIK1989, artifical neural networks are universal approximators that can approximate any Borel measurable function from one finite dimensional space to another to any desired degree of accuracy, provided the network is sufficiently large, the sample size goes to infinity and the function is sufficiently smooth. Following this, we outline a simple two-step procedure to estimate the average partial effect under the R-OLS framework when the data-generating process (DGP) for $X_i$ is unknown.

First, $r(Z_i)$ is estimated using a neutral network. With this estimate, the residuals $\hat{\nu}_i = X_i - \hat{r}(Z_i)$ are computed for each observation. In the second step, a standard OLS regression of $Y_i$ on $\hat{\nu}_i$ is performed. The coefficient on $\hat{\nu}_i$ from this regression provides the R-OLS estimate of the average partial effect. To construct valid confidence intervals for the R-OLS estimate, a bootstrap procedure is employed. This involves repeatedly resampling the data, re-estimating the APE using described procedure, and calculating confidence intervals of the R-OLS estimator based on the set of estimates.

It is however important to note that any method that consistently estimates $r(Z_i)$ can be used to estimate the average partial effect. The only requirement is that the method's predictions converge to $r(Z_i)$ with increasing sample size, residualising $X_i$ correctly.\footnote{Incorrect residualisation of $X_i$ is the origin of the biases described in goldsmithpinkham_et_al_2022 and winkelmann_2023. Both papers show that by not modelling $X_{i}$'s DGP correctly, a correlation between the predicted residual, regressor and heterogeneous treatment effect can occur which results in a biased estimate.}

Although the estimator might be consistent, it does not necessarily attain $\sqrt{n}$-consistency. As shown in equation ((ref)), if the rate of convergence of $\hat{r}(Z_i)$ is slower than $\sqrt{n}$, the estimator will still converge to the true parameter, but at a rate slower than $\sqrt{n}$. To avoid asymptotic bias in this case, it is essential that the estimation errors are uncorrelated with $Y_i$. This issue can be addressed by combining the R-OLS identification result with the Double/Debiased Machine Learning framework of Chernozhukov_et_al_2018_DebiasedML.

Estimation with Double/Debiased Machine Learning

The Double/Debiased Machine Learning (DML) framework Chernozhukov_et_al_2018_DebiasedML introduces the concept of Neyman-orthogonality (NO) in influence functions, enabling the estimation of the averate treatment effect (ATE) under non-linear but additively separable confounding. In this section, we demonstrate that estimates based on the NO moment condition for the partially linear model can be reinterpreted as the average partial effect (APE), provided the conditions of Theorem (ref) are satisfied. This holds even though the partially linear model is misspecification under Assumption (ref). Crucially, the DML framework facilitates the estimation of the APE without requiring explicit knowledge of the treatment data generating process, allowing for the use of machine learning methods to approximate $r(Z_i)=E[X_i \mid Z_i]$. Additionally, it offers a rigorous foundation for conducting inference on the APE by leveraging the theoretical results of Chernozhukov_et_al_2018_DebiasedML. Compared to the two-step procedure described in Section (ref), the DML-based approach yields more robust estimation and inference procedures for the APE.

The particular DML-approach we use is based on the partially linear model of Robinson_1988:

align*[align* omitted — 59 chars of source]

In this model, $\theta$ is of interest to the econometrician and can be interpreted as the ATE. We will demonstrate that, provided the conditions of Theorem (ref) are satisfied, $\hat{\theta}$ estimated via the NO moment for the partially linear model can be interpreted as the APE, even though the partially linear model is misspecified under our assumptions.

The partialling-out based NO moment for the partially linear model is given by:

equation[equation omitted — 149 chars of source]

where $l(Z_i)=E[Y_i \mid Z_i]$ and $\eta_{i}=\{l(Z_i), r(Z_i)\}$ are the nuisance functions that need to estimated from the data, in this case with machine learning. This moment is first-order invariant to deviations in nuisance functions by satisfying the following conditions:

align*[align* omitted — 115 chars of source]

The solution to the NO sample moment is given by $$ \hat{\theta} = \frac{E_n[(X_i - \hat{r}(Z_i)) (Y_i - \hat{l}(Z_i))]}{E_n[(X_i - \hat{r}(Z_i))^2]} $$ where $E_n[X_i]= \frac{1}{n} \sum_{i=1}^n X_i$ and $\hat{\eta}_{i}=\{\hat{l}(Z_i), \hat{r}(Z_i)\}$ are the estimated nuisance functions.

Comparing this estimator to the R-OLS estimator, we observe that the estimator residualises the treatment variable similarly to R-OLS. However, unlike R-OLS, it also residualises the outcome variable. Under the required assumptions, Theorem (ref) demonstrates that the average partial effect (APE) is identified solely due to the residualization of the treatment variable. This suggests that the NO moment condition identifies the APE while accommodating slower convergence rates of the nuisance function estimators due to double residualization. This intuition is formalized in Lemmas (ref) and (ref), showing that $\hat{\theta}$ consistently estimates the APE under the conditions of Theorem (ref), cross-fitting, and sufficiently fast convergence of the nuisance function estimators.

lemma[Convergence of partially linear DML] Let $X_i = r(Z_i) + \nu_i$ with $E[\nu_i l(Z_i)] = 0$, $E[|\nu_i^2|] < \infty$ and $E[|\nu_iY_i|] < \infty$. Assume that the estimators of nuisance parameters $\eta_{i} = \{l(Z_i), r(Z_i)\}$ are fit using cross-fitting, and their convergence rates satisfy $\sqrt{n}n^{-(\varphi_r + \varphi_l)} \to 0$ and $n^{-2\varphi_r} \to 0$. Then, the method-of-moments estimator based on moment condition ((ref)) satisfies: \begin{equation*} \hat{\theta} = \frac{E_n[(X_i - \hat{r}(Z_i)) (Y_i - \hat{l}(Z_i))]}{E_n[(X_i - \hat{r}(Z_i))^2]} \xrightarrow{p} \frac{E[\nu_i Y_i]}{E[\nu_i^2]} \end{equation*}
lemma[DML estimation and inference for the APE] Under the assumptions of Lemma (ref) and Theorem (ref), the method-of-moments estimator based on equation ((ref)) satisfies: $$ \sqrt{n} \left( \hat{\theta} - E\left[ \partial_{X_i} Y_i \right] \right) \xrightarrow{p} 0. $$ where $E\left[ \partial_{X_i} Y_i\right]$ is the average partial effect of $X_i$ on $Y_i$. Under the additonal regularity conditions of Theorem 4.1 in Chernozhukov_et_al_2018_DebiasedML, the estimates are asymptotically normal.

We conclude that $\hat{\theta}$, based on the Neyman-orthogonal moment condition for the partially linear model, provides a consistent and asymptotically normal estimate of the average partial effect, under the conditions of Lemma (ref).

Applying the assumptions of Lemma (ref) to a cross-fit R-OLS estimator gives

align*[align* omitted — 352 chars of source]

which goes to $0$ if $\sqrt{n}n^{-\varphi_r} \xrightarrow{p} 0$, requiring a convergence rate of $\varphi_r > 1/2$ compared to the joint convergence rate of $\varphi_r + \varphi_l > 1/2$ in Lemma (ref). Alternatively, the estimation error $r(Z_i) - \hat{r}(Z_i)$ needs to be unrelated to $Y_i$, for example via conditional mean independence. Unfortunately, cross-fitting can't guarantee this, since $r(Z_i) - \hat{r}(Z_i)$ can have "remnants" correlated with $Z_i$, creating a correlation with $Y_i$ and inducing bias.

The advantage of using a NO moment to estimate the APE is in its robustness to estimation errors in $\hat{r}(X_i)$ compared to ROLS. Let

align*[align* omitted — 94 chars of source]

with $\varepsilon_i, \nu_i \sim N(0,1)$ and $Z_i \sim Unif(0,2)$. Figure (ref) shows the robustness of the NO moment condition when estimating the APE. The figure compares the estimates of the APE of R-OLS and DML over 200 simulations in which the nuisance functions are estimated with a neural network. The simulation design aims to reflect realistic challenges in nuisance function estimation, such as imperfect model training and limited sample size, to show the different robustness of R-OLS and DML. Variation between simulations is not only induced by different realisations of the errors $\nu_i$ and $\varepsilon_i$, but also by randomly chosing the number of training epochs for the neural networks predicting $l(Z_i)=E[Y_i \mid Z_i]$ and $r(Z_i)=E[X_i \mid Z_i]$. This creates artificial variation in the quality of the predictions of $l(Z_i)$ and $r(Z_i)$, showcasing the increased robustness of the NO moment condition compared to R-OLS. The quality of the predictions is measured by the correlation between $\hat{\nu_i}$ and $Z_i$, providing a simple measure of how well the neural network estimates $r(Z_i)$ and how much potential for bias exists. A measure of the quality with which $\hat{l}(Z_i)$ predicts $l(Z_i)$ is not shown, since the NO moment condition is utilising this prediction for debiasing and not for identification.

Figure (ref) shows the NO moment condition is robust to correlation between $\hat{\nu}_i$ and $Z_i$. The APE is estimated consistently even for imperfect estimation of $r(Z_i)$. The R-OLS estimator on the other hand is biased when the correlation between $\hat{\nu}_i$ and $Z_i$ is high, which is exactly the bias shown in equation ((ref)).

This confirms the theoretical results above and shows that the estimates given by the NO moment condition for the partially linear model ((ref)) can be interpreted as the APE, under the Assumptions of Lemma (ref).

figure[figure omitted — 1,032 chars of source]

Weights and Diagonistics

The OLS and IV estimand can be rewritten as weighted expectations of the partial derivative Yitzhaki1996, Angrist_1998, ANGRIST_KRUEGER_1999, Graham_Pinto_2022, kolesar2024dynamiccausaleffectsnonlinearworld. Using a novel decomposition, Theorem (ref) implicit uses this fact and make assumptions such that the weights ensure the APE is estimated. The following Lemma shows this for the R-OLS estimand, but the result applies to OLS immediately.\footnote{The weights applied to the OLS and IV estimand can be see in Appendix D.}

lemma[R-OLS as weighted sum of partial derivatives] Under Assumption (ref), the R-OLS estimand $$ \beta = \frac{Cov(\nu_i, Y_i)}{Var(\nu_i)} $$ can be decomposed as a weighted sum of partial derivatives $$ \beta = \sum_{m=1}^M \sum_{p=0}^{m-1} {m-1 \choose p} m E\left[ r\left(Z_{i}\right)^{m-1-p} \nu_{i}^{p} g_m\left(Z_{i}\right) \right] \frac{E\left[\nu_{i}^{p+2}\right] - E[\nu_{i}] E\left[\nu_{i}^{p+1}\right]}{ (p+1) Var(\nu_i) E\left[\nu_{i}^{p}\right]} $$

These weights are novel and show that one can interpret R-OLS estimates as a sum of partial derivatives of each element of the polynomial structure in the Y-DGP $$ m E\left[ r\left(Z_{i}\right)^{m-1-p} \nu_{i}^{p} g_m\left(Z_{i}\right) \right] $$ with weights $$ \frac{E\left[\nu_{i}^{p+2}\right] - E[\nu_{i}] E\left[\nu_{i}^{p+1}\right]}{ (p+1) Var(\nu_i) E\left[\nu_{i}^{p}\right]} $$ depending on the distribution of the treatment.

In contrast to existing literature, which typically emphasizes weights applied to the partial derivative evaluated at specific values of the regressor, this decomposition enables a trade-off between the complexity of the outcome DGP and the distributional assumptions on the treatment variable. This trade-off is particularly evident in the R-OLS framework.

Applying the assumptions of R-OLS to these weights, we can see that Assumption (ref) guarantees that the R-OLS regression of $Y_i$ on $\nu_i$ weights the average partial effect of each polynomial in equation ((ref)) by one.

Building on this intuition, practitioners can identify whether the R-OLS estimator under- or overestimates the APE and determine the extent to which the underlying distributional assumptions hold, by estimating the sample analog of the above weights and examining their behavior across values of p. Specifically, the order of p up to which the weights approximate unity provide guidance on the appropriateness of R-OLS for estimating the APE under a given data-generating processes.

Simulations

There are three variables in the simulation. A treatment variable $X$ for which the APE is estimated and two confounding variables, namely $Z = \{Z_1, Z_2\}$. The confounders $Z_1$ and $Z_2$ are observed and always $N(1,1)$ distributed.

The assumption of normally distributed confounders is made for two reasons. Firstly, the complexity in the simulations does not originate from the distribution of the confounders, but from the DGPs used to specify $Y$ and $X$. Secondly, assuming normally distributed confounders facilitates the generation of data with higher-order interactions, without the concern of infinite variance.

One of the aspects highlighted in the simulations is the importance of the moments of the errors in the X-DPG, $\nu$, since Theorem (ref) shows that they are crucial to estimate the APE with R-OLS. Hence, the simulations are split into two parts. In section (ref) the errors satisfy Assumption (ref) by following $N(0,1)$. Then in section (ref) this assumption is violated and $\nu$ is drawn from a Gaussian mixture satisfying only a limited number of moment conditions. In contrast to the $X$-DGP errors, the errors in the $Y$-DGP, denoted as $\varepsilon$, only need to satisfy an unconfoundedness condition and are therefore always generated from a standard normal distribution.

table[table omitted — 914 chars of source]
table[table omitted — 1,434 chars of source]

Table (ref) shows the DGPs used to generate the data in each specification of the simulations. For some of the DGPs, sine and cosine are used to achieve highly non-linear functional forms while keeping the variables bounded and non-normally distributed, apart from the error component.

All the estimated models compared in the simulations can be seen in Table (ref). PL-GAM is a semi-parametric model in which the non-parametric part is estimated with generalised additive models using third-degree BSplines. The R-OLS estimator predicts the residuals by first training a neural network with three hidden layers, each with 64 nodes, to predict $X$ for each individual. Then $\nu$ is approximated via $\hat{\nu}=X-\hat{X}$ on the same sample the algorithm was trained on.\footnote{I used the "tensorflow" library in Python to train the neural networks on an Nvidia Tesla 4 GPU. During training, no cross-validated grid search for the optimal hyperparameters of the neural network is done due to the high computational cost associated. Hence, the simulations show the lower bound of the potential performance of R-OLS.}

Throughout the next sections, the results of some simulations are highlighted. The appendix contains all remaining specifications.

Standard Normal Errors

Simple OLS Assumptions Satisfied

We begin with simulations showing the results for the additive $Y$-DGP and a variation of $X$-DGPs with standard normal errors. In these simulations, the assumptions of simple OLS are satisfied by the $Y$-DGP and hence any OLS estimator is unbiased.

The results can be seen in Table (ref). The table shows 10000 simulations for all five estimated models under different sizes of the dataset, $N$, and $X$-DGPs. Each field reports the average of the estimated APEs, specified in Table (ref), with its standard deviation in round brackets and mean squared error in square brackets.

The results show that all estimators perform well in this case due to the simplicity of the simulation. The simple OLS, PL-GAM and interacted OLS estimators are unbiased and have virtually no variance. The R-OLS estimator is unbiased, but has higher variance.

These simulations highlight that if the assumptions of simple OLS are satisfied, R-OLS obviously does not improve on the simple OLS estimator. However, it is important to note that the assumptions of simple OLS are rarely satisfied in practice and R-OLS offers increased robustness at the price of increased variance. The next simulations show the results for more complex $Y$-DGPs in which the assumptions of simple OLS are not satisfied anymore, leading to biased estimates.

sidewaystable\caption{APE estimation with additive $Y$-DGP and a variation of $X$-DGPs with standard normal errors} \resizebox{\textwidth}{!}{ \begin{tabular}{llcccclcccclcccc} & & \multicolumn{4}{c}{additive $X$-DGP with APE=1} & & \multicolumn{4}{c}{simple $X$-DGP with APE=1} & & \multicolumn{4}{c}{complex $X$-DGP with APE=1} \\ \cline{3-6} \cline{8-11} \cline{13-16} & & \multicolumn{1}{c|}{N=100} & \multicolumn{1}{c|}{N=500} & \multicolumn{1}{c|}{N=1000} & N=5000 & & \multicolumn{1}{c|}{N=100} & \multicolumn{1}{c|}{N=500} & \multicolumn{1}{c|}{N=1000} & N=5000 & & \multicolumn{1}{c|}{N=100} & \multicolumn{1}{c|}{N=500} & \multicolumn{1}{c|}{N=1000} & N=5000 \\ \cline{3-6} \cline{8-11} \cline{13-16} simple OLS & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular} \\ \cline{1-1} \cline{3-6} \cline{8-11} \cline{13-16} interacted OLS & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular} \\ \cline{1-1} \cline{3-6} \cline{8-11} \cline{13-16} PL-GAM & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular} \\ \cline{1-1} \cline{3-6} \cline{8-11} \cline{13-16} \textbf{R-OLS} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.05\\ (0.14)\\ {[}0.02{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.02\\ (0.05)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.01)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.01)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.01)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.01)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.0\\ (0.01)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}1.0\\ (0.0)\\ {[}0.0{]}\end{tabular} \end{tabular} } \newline \newline \justify \textbf{Note: }{Mean APE estimate after 10000 simulations. Round brackets show the standard deviation and square brackets the MSE of the estimates. $N$ is the sample size. Within each simulation, the same data is used for all estimators. Details of the DGP and estimated models are shown in Table (ref) and (ref), respectively. Errors in the $X$-DGP follow a standard normal distribution.}

Simple OLS Assumptions Violated

Results of the simulation with the simple $X$-DGP and the complex $Y$-DGP and standard normal errors can be seen in Table (ref). Simple OLS and PL-GAM severely underestimate the APE in this case and for some specifications even estimate the opposite sign. This comes from a negative correlation between the residual and the APE. Interacted OLS consistently estimates the APE since from a Frisch-Waugh-Lovell perspective the estimated equation residualises $X$ correctly, essentially replicating R-OLS, just with a linear regression.\footnote{Table (ref) in the appendix compares simulations in which all estimators replicate R-OLS from a FWL perspective and shows unbiasedness and similar variance for all estimators.} R-OLS, in which a neural network estimates the errors, does well in larger samples and/or low $M$.

sidewaystable\caption{APE estimation with complex $Y$ and simple $X$-DGPs and standard normal errors} \resizebox{\textwidth}{!}{ \begin{tabular}{llcccclcccclcccc} & & \multicolumn{4}{c}{M=1 with APE=0.17} & & \multicolumn{4}{c}{M=2 with APE=-0.14} & & \multicolumn{4}{c}{M=3 with APE=-2.06} \\ \cline{3-6} \cline{8-11} \cline{13-16} & & \multicolumn{1}{c|}{N=100} & \multicolumn{1}{c|}{N=500} & \multicolumn{1}{c|}{N=1000} & N=5000 & & \multicolumn{1}{c|}{N=100} & \multicolumn{1}{c|}{N=500} & \multicolumn{1}{c|}{N=1000} & N=5000 & & \multicolumn{1}{c|}{N=100} & \multicolumn{1}{c|}{N=500} & \multicolumn{1}{c|}{N=1000} & N=5000 \\ \cline{3-6} \cline{8-11} \cline{13-16} simple OLS & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-0.08\\ (0.14)\\ {[}0.08{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-0.08\\ (0.06)\\ {[}0.07{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-0.08\\ (0.04)\\ {[}0.06{]}\end{tabular}} & \begin{tabular}[c]{@c@}-0.08\\ (0.02)\\ {[}0.06{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-1.05\\ (0.96)\\ {[}1.76{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-1.07\\ (0.46)\\ {[}1.08{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-1.07\\ (0.34)\\ {[}0.98{]}\end{tabular}} & \begin{tabular}[c]{@c@}-1.07\\ (0.15)\\ {[}0.89{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-5.08\\ (10.08)\\ {[}110.72{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-5.01\\ (4.99)\\ {[}33.64{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-4.94\\ (3.8)\\ {[}22.76{]}\end{tabular}} & \begin{tabular}[c]{@c@}-4.98\\ (1.67)\\ {[}11.31{]}\end{tabular} \\ \cline{1-1} \cline{3-6} \cline{8-11} \cline{13-16} interacted OLS & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.22\\ (0.1)\\ {[}0.01{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.18\\ (0.04)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.18\\ (0.03)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.17\\ (0.01)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-0.01\\ (0.48)\\ {[}0.25{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-0.09\\ (0.22)\\ {[}0.05{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-0.11\\ (0.15)\\ {[}0.02{]}\end{tabular}} & \begin{tabular}[c]{@c@}-0.13\\ (0.07)\\ {[}0.01{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-1.49\\ (3.58)\\ {[}13.15{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-1.78\\ (1.69)\\ {[}2.93{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-1.87\\ (1.22)\\ {[}1.53{]}\end{tabular}} & \begin{tabular}[c]{@c@}-2.01\\ (0.58)\\ {[}0.34{]}\end{tabular} \\ \cline{1-1} \cline{3-6} \cline{8-11} \cline{13-16} PL-GAM & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-0.07\\ (0.13)\\ {[}0.07{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-0.09\\ (0.05)\\ {[}0.07{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-0.09\\ (0.04)\\ {[}0.07{]}\end{tabular}} & \begin{tabular}[c]{@c@}-0.08\\ (0.02)\\ {[}0.06{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-1.01\\ (0.7)\\ {[}1.26{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-1.13\\ (0.36)\\ {[}1.11{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-1.12\\ (0.27)\\ {[}1.04{]}\end{tabular}} & \begin{tabular}[c]{@c@}-1.09\\ (0.14)\\ {[}0.93{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-5.62\\ (5.08)\\ {[}38.51{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-5.82\\ (3.15)\\ {[}24.1{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-5.57\\ (2.58)\\ {[}19.04{]}\end{tabular}} & \begin{tabular}[c]{@c@}-5.18\\ (1.4)\\ {[}11.69{]}\end{tabular} \\ \cline{1-1} \cline{3-6} \cline{8-11} \cline{13-16} \textbf{R-OLS} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.18\\ (0.13)\\ {[}0.02{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.17\\ (0.04)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.17\\ (0.03)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.17\\ (0.02)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-0.05\\ (0.69)\\ {[}0.49{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-0.06\\ (0.26)\\ {[}0.07{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-0.12\\ (0.19)\\ {[}0.03{]}\end{tabular}} & \begin{tabular}[c]{@c@}-0.13\\ (0.08)\\ {[}0.01{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-1.17\\ (7.41)\\ {[}55.66{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-1.93\\ (2.13)\\ {[}4.55{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}-1.77\\ (1.89)\\ {[}3.67{]}\end{tabular}} & \begin{tabular}[c]{@c@}-1.98\\ (0.74)\\ {[}0.55{]}\end{tabular} \end{tabular} } \newline \newline \justify \textbf{Note: }{Mean APE estimate after 10000 simulations. Round brackets show the standard deviation and square brackets the MSE of the estimates. $M$ is the order of polynomials in the $Y$-DGP and $N$ the sample size. Within each simulation, the same data is used for all estimators. Details of the DGP and estimated models are shown in Table (ref) and (ref), respectively. Errors in the $X$-DGP follow a standard normal distribution.}

Moving to a more complicated DGP for $X$ in Table (ref), we can see that neither simple OLS nor PL-GAM can provide a consistent estimate of the APE since they can neither fit the $X$ nor the $Y$-DGP well. Furthermore, due to the increased complexity of the $X$-DGP, interacted OLS is not able to replicate the idea behind R-OLS anymore and its estimates are off, but not by much, hinting that the interacted model is surprisingly good at approximating the DGPs. In this setting, R-OLS with neural networks works very well if $M=2$. Surprisingly, for $M=1$ and $M=3$, the estimation gets worse with an increased sample size. However, it still improves on non-R-OLS alternatives.

sidewaystable[] \caption{APE estimation with complex $Y$ and complex $X$-DGPs and standard normal errors} \resizebox{\textwidth}{!}{ \begin{tabular}{llcccclcccclcccc} & & \multicolumn{4}{c}{M=1 with APE=0.17} & & \multicolumn{4}{c}{M=2 with APE=0.21} & & \multicolumn{4}{c}{M=3 with APE=1.52} \\ \cline{3-6} \cline{8-11} \cline{13-16} & & \multicolumn{1}{c|}{N=100} & \multicolumn{1}{c|}{N=500} & \multicolumn{1}{c|}{N=1000} & N=5000 & & \multicolumn{1}{c|}{N=100} & \multicolumn{1}{c|}{N=500} & \multicolumn{1}{c|}{N=1000} & N=5000 & & \multicolumn{1}{c|}{N=100} & \multicolumn{1}{c|}{N=500} & \multicolumn{1}{c|}{N=1000} & N=5000 \\ \cline{3-6} \cline{8-11} \cline{13-16} simple OLS & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.13\\ (0.05)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.13\\ (0.02)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.13\\ (0.02)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.13\\ (0.01)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.25\\ (0.17)\\ {[}0.03{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.24\\ (0.07)\\ {[}0.01{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.24\\ (0.05)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.24\\ (0.02)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.06\\ (0.77)\\ {[}0.8{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.04\\ (0.33)\\ {[}0.34{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.03\\ (0.24)\\ {[}0.29{]}\end{tabular}} & \begin{tabular}[c]{@c@}1.02\\ (0.1)\\ {[}0.26{]}\end{tabular} \\ \cline{1-1} \cline{3-6} \cline{8-11} \cline{13-16} interacted OLS & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.17\\ (0.08)\\ {[}0.01{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.17\\ (0.05)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.18\\ (0.04)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.18\\ (0.03)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.29\\ (0.33)\\ {[}0.12{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.31\\ (0.17)\\ {[}0.04{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.33\\ (0.13)\\ {[}0.03{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.33\\ (0.07)\\ {[}0.02{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.63\\ (1.45)\\ {[}2.11{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.55\\ (0.67)\\ {[}0.45{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.58\\ (0.49)\\ {[}0.25{]}\end{tabular}} & \begin{tabular}[c]{@c@}1.63\\ (0.25)\\ {[}0.07{]}\end{tabular} \\ \cline{1-1} \cline{3-6} \cline{8-11} \cline{13-16} PL-GAM & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.14\\ (0.07)\\ {[}0.01{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.12\\ (0.02)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.11\\ (0.02)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.11\\ (0.01)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.19\\ (0.21)\\ {[}0.04{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.15\\ (0.08)\\ {[}0.01{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.15\\ (0.05)\\ {[}0.01{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.14\\ (0.02)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.09\\ (1.04)\\ {[}1.27{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.99\\ (0.39)\\ {[}0.44{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.95\\ (0.27)\\ {[}0.4{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.92\\ (0.12)\\ {[}0.37{]}\end{tabular} \\ \cline{1-1} \cline{3-6} \cline{8-11} \cline{13-16} \textbf{R-OLS} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.16\\ (0.11)\\ {[}0.03{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.17\\ (0.04)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.15\\ (0.03)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.13\\ (0.01)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.25\\ (0.36)\\ {[}0.13{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.21\\ (0.14)\\ {[}0.02{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.21\\ (0.1)\\ {[}0.01{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.2\\ (0.05)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.67\\ (1.57)\\ {[}2.5{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.59\\ (0.67)\\ {[}0.46{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.35\\ (0.5)\\ {[}0.28{]}\end{tabular}} & \begin{tabular}[c]{@c@}1.37\\ (0.22)\\ {[}0.07{]}\end{tabular} \end{tabular} } \newline \newline \justify \textbf{Note: }{Mean APE estimate after 10000 simulations. Round brackets show the standard deviation and square brackets the MSE of the estimates. $M$ is the order of polynomials in the $Y$-DGP and $N$ the sample size. Within each simulation, the same data is used for all estimators. Details of the DGP and estimated models are shown in Table (ref) and (ref), respectively. Errors in the $X$-DGP follow a standard normal distribution.}

Impact of Non Normal Errors

The theoretical conditions on the errors given in Assumption (ref) are satisfied by normally distributed errors with mean zero up to any polynomial degree. To show that results also hold for non-normal errors satisfying the conditions up to a certain order, the following Gaussian mixture is used to generate the errors:

align*[align* omitted — 116 chars of source]
equation*[equation* omitted — 296 chars of source]

These errors do not satisfy the fourth-moment condition in Assumption (ref), which becomes relevant for a $Y$-DGP with polynomial degree $M\geq3$ in Assumption (ref).\footnote{A uniform distribution centered around zero would have the same implications.}

In Table (ref), the results for a simulation with non-normal errors are displayed. In the first two blocks with $M\leq2$ all assumptions of Theorem (ref) are satisfied and R-OLS is consistent for the APE. Increasing the polynomial degree in the DGP of $Y$ from $M=2$ to $M=3$ creates a scenario in which Assumption (ref) is not fulfilled anymore. As suggested by Theorem (ref), R-OLS does not provide an unbiased estimate of the average partial effect anymore. Luckily, the bias is not extreme and the estimates are still less biased than normal OLS or partially linear GAMs.

sidewaystable[] \caption{APE estimation with complex $Y$ and complex $X$-DGPs and non normal errors} \resizebox{\textwidth}{!}{ \begin{tabular}{llcccclcccclcccc} & & \multicolumn{4}{c}{M=1 with APE=0.17} & & \multicolumn{4}{c}{M=2 with APE=0.2} & & \multicolumn{4}{c}{M=3 with APE=1.52} \\ \cline{3-6} \cline{8-11} \cline{13-16} & & \multicolumn{1}{c|}{N=100} & \multicolumn{1}{c|}{N=500} & \multicolumn{1}{c|}{N=1000} & N=5000 & & \multicolumn{1}{c|}{N=100} & \multicolumn{1}{c|}{N=500} & \multicolumn{1}{c|}{N=1000} & N=5000 & & \multicolumn{1}{c|}{N=100} & \multicolumn{1}{c|}{N=500} & \multicolumn{1}{c|}{N=1000} & N=5000 \\ \cline{3-6} \cline{8-11} \cline{13-16} simple OLS & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.13\\ (0.05)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.13\\ (0.02)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.13\\ (0.02)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.13\\ (0.01)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.25\\ (0.16)\\ {[}0.03{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.24\\ (0.07)\\ {[}0.01{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.24\\ (0.05)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.24\\ (0.02)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.01\\ (0.64)\\ {[}0.67{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.98\\ (0.29)\\ {[}0.37{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.98\\ (0.2)\\ {[}0.33{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.98\\ (0.09)\\ {[}0.3{]}\end{tabular} \\ \cline{1-1} \cline{3-6} \cline{8-11} \cline{13-16} interacted OLS & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.16\\ (0.08)\\ {[}0.01{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.17\\ (0.05)\\ {[}0.01{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.17\\ (0.04)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.18\\ (0.03)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.32\\ (0.3)\\ {[}0.1{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.33\\ (0.15)\\ {[}0.04{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.33\\ (0.12)\\ {[}0.03{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.34\\ (0.07)\\ {[}0.02{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.59\\ (1.2)\\ {[}1.44{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.47\\ (0.55)\\ {[}0.3{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.5\\ (0.42)\\ {[}0.18{]}\end{tabular}} & \begin{tabular}[c]{@c@}1.55\\ (0.22)\\ {[}0.05{]}\end{tabular} \\ \cline{1-1} \cline{3-6} \cline{8-11} \cline{13-16} PL-GAM & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.14\\ (0.07)\\ {[}0.01{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.12\\ (0.02)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.11\\ (0.02)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.11\\ (0.01)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.19\\ (0.18)\\ {[}0.03{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.15\\ (0.07)\\ {[}0.01{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.15\\ (0.05\\ {[}0.01{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.14\\ (0.02)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.01\\ (0.87)\\ {[}1.02{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.91\\ (0.33)\\ {[}0.48{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.88\\ (0.23)\\ {[}0.46{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.86\\ (0.1)\\ {[}0.45{]}\end{tabular} \\ \cline{1-1} \cline{3-6} \cline{8-11} \cline{13-16} \textbf{R-OLS} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.15\\ (0.09)\\ {[}0.01{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.17\\ (0.03)\\ {[}0.0{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.15\\ (0.03)\\ {[}0.0{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.13\\ (0.01)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.24\\ (0.27)\\ {[}0.08{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.21\\ (0.1)\\ {[}0.01{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}0.21\\ (0.07)\\ {[}0.01{]}\end{tabular}} & \begin{tabular}[c]{@c@}0.21\\ (0.03)\\ {[}0.0{]}\end{tabular} & & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.53\\ (1.14)\\ {[}1.3{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.32\\ (0.45)\\ {[}0.24{]}\end{tabular}} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}1.27\\ (0.32)\\ {[}0.16{]}\end{tabular}} & \begin{tabular}[c]{@c@}1.15\\ (0.14)\\ {[}0.16{]}\end{tabular} \end{tabular} } \newline \newline \justify \textbf{Note: }{Mean APE estimate after 10000 simulations. Round brackets show the standard deviation and square brackets the MSE of the estimates. $M$ is the order of polynomials in the $Y$-DGP and $N$ the sample size. Within each simulation, the same data is used for all estimators. Details of the DGP and estimated models are shown in Table (ref) and (ref), respectively. Errors in the $X$-DGP follow a Gaussian mixture with distribution $0.5 \cdot N(\mu, 1-\mu^2) + 0.5 \cdot N(-\mu, 1-\mu^2)$ and $\mu = 0.9$.}

Empirical Illustration

Austerity and the Rise of UKIP (Fetzer, 2019)

This section demonstrates the application of the R-OLS estimator to evaluate the causal impact of austerity measures on political outcomes, leveraging the empirical framework and data from fetzer2019_austerity_brexit. The analysis examines the relationship between austerity-induced fiscal losses and the rise in support for the UK Independence Party (UKIP), with broader implications for understanding the Brexit referendum results.

The austerity measures implemented by the UK government after 2010 introduced significant reductions in welfare spending, disproportionately affecting rural and economically disadvantaged regions. These policy changes generated substantial heterogeneity in fiscal shocks, which provides a quasi-experimental setting for causal inference. Austerity exposure is measured as the average financial loss per county, encompassing reductions in tax credits, housing benefits, and disability allowances. This variation allows for an exploration of the political consequences of austerity across districts with differing socioeconomic profiles.

Previous findings by fetzer2019_austerity_brexit highlight a robust relationship between austerity exposure and increased support for UKIP. Specifically, districts with higher financial losses experienced greater gains in UKIP vote shares across local, European, and parliamentary elections. These effects were more pronounced in areas with lower educational attainment and higher shares of employment in routine or manufacturing jobs. Using the R-OLS and DML estimator, we refine these estimates by addressing potential non-linear interactions between austerity exposure and observable confounders, providing a more precise estimation of the average partial effect (APE) of austerity on electoral outcomes.

The R-OLS estimation involves two steps. In the first step, we use a Gradient Boosting Machine (GBM) to predict the treatment variable, defined as the interaction term $\mathds{1}(\text{Year} > 2010) \times \text{Austerity}$, where Austerity represents the total financial losses resulting from policy changes. Hyperparameters of the GBM, including the number of trees, interaction depth, minimum observations per node, and shrinkage (learning rate), are tuned using 5-fold cross-validation over a predefined grid. Once the optimal hyperparameters are determined, the data is bootstrapped 250 times. For each bootstrap sample, 5-fold cross-fitting is applied to train the GBM on 4 folds and predict the treatment variable for the left-out fold. The residuals $\hat{\nu}_i$, calculated as the difference between the observed and predicted treatment variable, are then used as inputs in the R-OLS estimator.

Figure (ref) displays the distribution of residuals after applying the above described residualisation to the entire sample with cross-fitting. It is overlaid with a normal density curve of the same mean and variance. The approximate normality of the residuals provides suggestive evidence that the assumptions under which R-OLS and DML estimate the APE are satisfied.\footnote{The original regression specification in fetzer2019_austerity_brexit includes region-by-year fixed effects. Our added controls are: TurnoutPct, populationukmillion, protectiontotalbillion, GVA_Agriculturepc, GVA_Manufacturinpc, GVA_Informationpc, GRetailUKShareWithin, LargeemphigherManAll_sh, IntermOccAll_sh, LLTunemployedAll_sh, LStudentAll_sh, CMiningAll_sh, DManufAll_sh, EUtilityAll_sh, FConstrAll_sh, HHotelsAll_sh, ITransportICTAll_sh, JFinancialAll_sh, LPublicAll_sh, NHealthAll_sh, OTHERAll_sh, instrument_for_shock, AgeAbove60UK, AC12MigrantShare. See fetzer2019_austerity_brexit for details. Without these added controls, the residuals show significant skew as shown in Figure (ref) in the Appendix and making it unlikely the assumptions of Theorem (ref) are satisfied.}

figure[figure omitted — 1,420 chars of source]

The DML estimator repeats the above outlined residualisation for the two outcome variables, support for UKIP in local and European elections. Residualisation is performed with a GBM, using the same hyper-parameters as before. The estimator is easily implemented with DoubleML package in R DoubleML2022Python, DoubleML2022R using the double machine learning estimator for the partially linear model with the 'partialling out'-score. Due to the results outlined in Section (ref), we interpret the estimate as an estimate for the APE and use the provided standard errors for inference.\footnote{The DML estimator was included in the bootstrapping framework used for inference for the R-OLS estimator. Estimates of standard error of the APE provided by the bootstrap procedure and the DoubleML package are nearly identical.}

Both estimators are repeated across all bootstrap samples, yielding 250 estimates of the APE of austerity exposure on UKIP vote shares in local and European elections. The results, summarized in Table (ref), report the mean and standard deviation of these estimates.

The R-OLS and DML estimates reveal a pronounced sensitivity of UKIP support to austerity-induced financial losses in local elections, exceeding the estimates obtained from conventional fixed-effects regressions (FEOLS) as in fetzer2019_austerity_brexit. In contrast, the effect of austerity on European elections is not deemed statistically significant, suggesting that the political consequences of austerity may vary across electoral contexts. These findings underscore the importance of accounting for non-linear treatment effects in empirical research.

table[table omitted — 1,227 chars of source]

Conclusion

This paper introduces a novel estimation method, R-OLS, for identifying and estimating the average partial effect of a continuous treatment in the presence of non-linear and non-additively separable confounding. The method leverages an exogenous error component of the treatment and imposes specific assumptions on the data-generating process to ensure equivalence to the APE. Double/Debiased Machine Learning is employed to facilitate valid inference for R-OLS estimates. Simulation studies demonstrate the strong performance of R-OLS across a range of complex data-generating processes, and applications to empirical data underscore its practical relevance. Overall, R-OLS offers a flexible and robust approach for estimating treatment effects in diverse empirical settings.

\singlespacing \setlength\bibsep{0pt}