EconBase
← Back to paper

Bounds for Bias-Adjusted Treatment Effect in Linear Econometric Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

84,356 characters · 30 sections · 58 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Bounds for Bias-Adjusted Treatment Effect in Linear Econometric Models

abstractIn linear econometric models with proportional selection on unobservables, omitted variable bias in estimated treatment effects are real roots of a cubic equation involving estimated parameters from a short and intermediate regression. The roots of the cubic are functions of $\delta$, the degree of selection on unobservables, and $R_{max}$, the R-squared in a hypothetical long regression that includes the unobservable confounder and all observable controls. In this paper I propose and implement a novel algorithm to compute roots of the cubic equation over relevant regions of the $\delta$-$R_{max}$ plane and use the roots to construct bounding sets for the true treatment effect. The algorithm is based on two well-known mathematical results: (a) the discriminant of the cubic equation can be used to demarcate regions of unique real roots from regions of three real roots, and (b) a small change in the coefficients of a polynomial equation will lead to small change in its roots because the latter are continuous functions of the former. I illustrate my method by applying it to the analysis of maternal behavior on child outcomes. \\ {\bf Keywords}: treatment effect, omitted variable bias.\\ {\bf JEL Codes}: C21.

Introduction

Researchers are often interested in estimating treatment effects in models where there are clear problems of unobserved or unobservable confounders. Such hypothetical `long' regressions cannot be estimated because of unobservability of the confounding regressor. Faced with this problem, researchers often compare the ordinary least square (OLS) estimate of the treatment effect between a `short' and an `intermediate' regression, both of which can be estimated. In the short regression both the observable and unobservable controls are left out; in the intermediate regression only the unobservable control is missing from the model. If the numerical magnitudes of the treatment effect are roughly similar in both the short and intermediate regressions, i.e. the estimate of the treatment effect is `stable', then researchers conclude that the bias from the omitted, unobservable confounder is small.

In a recent, innovative contribution, oster-2019 has demonstrated that such `coefficient stability' arguments to deal with possible omitted variable bias is misleading.\footnote{oster-2019 extends previous work on this issue by aet-2000,aet-2005.} In fact, what is needed to draw conclusions about the magnitude of possible bias due to the unobservable confounder is not the raw change in the estimate of the treatment effect, but an R-squared scaled change in the estimate of the treatment effect between the short and intermediate regressions. This becomes clear when we write the expression for the omitted variable bias of the treatment effect in the intermediate regression in terms of the R-squared in the short, intermediate and hypothetical long regressions, and relevant coefficients in the long regression. A little algebraic manipulation generates a cubic equation in the bias (of the treatment effect in the intermediate regression). The coefficients of this cubic equation are functions of estimated, and therefore known, parameters and, in addition, two unknown parameters: $\delta$, the relative degree of selection on unobservables, and $R_{max}$, the R-squared in the hypothetical long regression.

A cubic equation with real (or complex) coefficients will have either one or three real roots. When the cubic equation has a unique real root, the researcher is able to identify the bias, and use it to compute the bias-adjusted treatment effect, without any ambiguity. When the cubic equation has three real roots, the researcher is confronted with the problem of non-uniqueness. Confronted with this problem, previous researchers like oster-2019, aet-2005 and aet-2000 have proposed approaches that avoids computing roots of the cubic. oster-2019 proposes two approaches to deal with the problem of non-uniqueness that does not require the researcher to compute roots of the cubic equation. Unfortunately, these methods suffer from some serious shortcoming (which I point out later in this introductory section and then discuss in greater detail in section (ref)).

I propose an alternative methodology to quantify the bias in treatment effect. At the center stage of my method is an algorithm to compute and select the correct real root of the cubic equation. To understand my proposed algorithm let us return to the cubic equation that is at the heart of the bias calculations. The coefficients of the cubic equation are functions of two unknown parameters: $\delta$, and $R_{max}$. In any given econometric analysis, details of the question under investigation will allow a researcher to choose a plausible range (an interval of the real line) over which $\delta$ and $R_{max}$ can vary.

To understand how a plausible range can be chosen, note that a value of $\delta$ that is higher than $1$ means that the unobservable confounder is relatively more (less) important than the observed controls in explaining variation in the treatment variable. On the other hand, a relatively high (low) value of $R_{max}$ means that the unobservable confounder is relatively more (less) important than the observed controls in explaining variation in the outcome variable. Based on the details of the outcome, treatment, included control and omitted variables, a researcher will therefore be able to choose the plausible range for $\delta$ and $R_{max}$. Once this is done, we are able to define a bounded box in the ($\delta$-$R_{max}$) plane as the Cartesian product of the bounded intervals over which $\delta$ and $R_{max}$ span.

Using the magnitude of the discriminant of the cubic equation we can then divide the bounded box into two parts, the first corresponding to a region where the cubic equation will have unique real roots (this is the region where the discriminant is strictly positive) and the second corresponding to a region where the cubic equation will have three real roots (this is the region where the discriminant is nonpositive). Let us call the first part as the $URR$ (unique real root) area and the second part as the $NURR$ (nonunique real root) area.

As the first, and simplest case, I consider the situation where the bounded box is completely contained in $URR$. I impose a $N \times N$ grid on the bounded box, compute the unique real roots on all grid points of the box and collect the real roots in a vector, $B_U$, (whose length is equal to $N^2$, the number of points of the grid). We can now compute a bounding set, $S_{URR}$, for the `true' treatment effect as the interval formed by the $2.5$-th and the $97.5$-th quantile of the empirical distribution of $\beta^*=\tilde{\beta}-\nu$, where $\nu$ are elements of the vector $B_U$. This is because the unique real root is the bias of the treatment effect (estimated from the intermediate regression). Hence, the `true' treatment effect is the difference between the treatment effect estimated from the intermediate regression, $\tilde{\beta}$, and the root of the cubic, $\nu$. Thus, the difference between $\tilde{\beta}$ and the vector, $B_U$, gives us a $N^2 \times 1$ vector of `true' treatment effects. And since we have covered the whole bounded box, as long as the assumptions generating the bounded box are correct, the empirical distribution of $\beta^*$, the bias-adjusted treatment effect (BATE), will give us a good approximation of the 95% confidence interval of $\beta$.

As the next, and computationally more difficult, case, I consider the situation where the bounded box is partly contained in the URR and partly in the NURR. In this situation, I first compute all the unique real roots on $M_1^2$ grid points of the box that lies in URR. Then, I extend the analysis to the NURR by taking recourse to a result from complex analysis alexanderian-2013: the roots of a polynomial are continuous functions of the coefficients of the polynomial; moreover, real roots of multiplicity one will remain real when the coefficients are perturbed slightly.

After computing the $M_1^2$ unique real roots on URR, I impose a grid of $M_2^2$ points on NURR. I identify border points of the grid that belong to NURR by choosing all points of NURR that are within a pre-specified `small' distance of any point in URR. In each of these border points of NURR, I compute the three real roots and select the one that is closest in absolute value to the unique real root computed at its closest grid point in URR. This selection is theoretically justified by the fact that roots of the cubic equation are continuous functions of $\delta$ and $R_{max}$. Hence, if a grid point in NURR is within a small open ball of a grid point in URR, the unique real root in the latter will be `close' to one of the three real roots in the former. This `closeness' allows me to select this root as the correct estimate of bias on the grid point in NURR.

Once the border grid points in NURR are covered, I iterate the process deeper, layer by layer and in `small' steps at a time, into the NURR area until all grid points in NURR are exhausted. At the end of the process, I have therefore generated a $(M_1^2+M_2^2)$-vector of real roots. I use these to compute the bias-adjusted treatment effect, $\beta^*$, exactly as I did in the first case. Using the quantiles of the empirical distribution of $\beta^*$, I can now generate approximate confidence intervals for the `true' $\beta$.

The final case to consider is one when the bounded box is wholly contained in NURR. In this case, the algorithm extends dimensions of the bounded box in the $\delta$ direction in small steps so as to generate a non-empty intersection with the URR. As soon as the algorithm finds a non-empty area of intersection with URR, it then applies the logic of the second case to complete the analysis.

The choice of bounds for $\delta$ and $R_{max}$ define the meaningful area over which roots of the relevant cubic equation is solved. Hence, the results of the analysis that I propose depend crucially on these bounds. Since $\delta$ captures the relative importance of the unobservable confounder compared to the observed controls in explaining the variation in the treatment variable, we can define a low $\delta$ regime as one where $0 < \delta <1$ and a high $\delta$ regime as one where $1 < \delta <\delta_{high}$.\footnote{The choice of $\delta_{high}$ will be decided by the researcher. What we need is a large positive number, significantly higher than unity, which ensures that the box has non-empty intersection with URR.} On the other hand, $R_{max}/\tilde{R}$ captures the relative importance of the unobserved confounder compared to the observed controls in explaining the variation of the outcome variable (recall that $R_{max}$ and $\tilde{R}$ are the R-squared of the intermediate and long regressions). Since the hypothetical regressions has more regressors than the intermediate regression, $\tilde{R} \leq R_{max} \leq 1$. While the lower bound of $R_{max}$ is thereby fixed, the upper bound will need to be chosen by the researcher. A value of $1$ is too restrictive because even in the hypothetical long regression, the regressors cannot be expected to explain all the variation in the outcome variable (perhaps due to measurement error). Hence, we must use some upper bound for $R_{max}$ that is lower than unity. While oster-2019 proposes an upper bound of $1.3\tilde{R}$, researchers can also try other numbers that are theoretically or empirically justified.

My proposed method has clear advantages over the two ways to address nonuniqueness that was proposed in oster-2019. The first method proposed by oster-2019 involves computing the bias-adjusted treatment effect under the twin assumptions of $\delta=1$ (equal selection on observables and unobservables) and a sign restriction (which is stated as Assumption $ 3 $ in the paper). In section (ref), I show that even on theoretical grounds this does not solve the nonuniqueness problem. I also demonstrate the problem using real data in section (ref). The second method proposed by oster-2019 relies on choosing some value of $R_{max}$, and calculating the magnitude of $\delta$, i.e. degree of selection due to unobservables, that would be consistent with $\beta=0$ (no treatment effect). While this method solves the nonuniqueness issue, it suffers from serious problems of robustness and interpretation, as I discuss in section (ref).

In comparison to the first method of oster-2019, my method is free of the theoretical problems that I identify in her method. My method offers a transparent, theoretically grounded method of computing bounds for the `true' treatment effect. In comparison to the second method proposed by oster-2019, my method is more robust. Instead of choosing a specific value of $R_{max}$, as oster-2019 does, I compute and then use the treatment bias for all possible values of $R_{max}$ and $\delta$ over a meaningful area. While the methodology of oster-2019 can be extremely sensitive to the correct choice of $R_{max}$, my method is less so because it relies on computing bias over a whole region.

After presenting the theoretical results, I use my method on a data set that comes from the Children and Young Adults sample of the NLSY and is used to study the impact of maternal behaviour on child outcomes.\footnote{I would like to thank Emily Oster for making her data set available. I have downloaded the data set from her webpage: \url{https://drive.google.com/file/d/0B1U4uS7GkkxbV0VkZmd0ZVlDVDA/view?usp=sharing}} Using this data set, I highlight, at various points in the paper, the differences in my methodology from the results reported and discussed in oster-2019. The algorithm proposed in this paper has been implemented in an R package called bate (bias adjusted treatment effect) and can be downloaded from \url{https://github.com/dbasu-umass/bate}.

The rest of the paper is organized as follows. In section (ref), I discuss the basic set-up; in section (ref), I present my method of analyzing bias and discuss details of the proposed algorithm; in section (ref), I illustrate my method, using a data set (NLSY) to study the impact of maternal behaviour on child outcomes; in section (ref), I compare my approach with Oster's and point out some problems in the latter; in section (ref), I conclude with a summary of my proposed methodology.

Basic Set-Up

Four Regression Models

Consider a hypothetical `long' regression,

equation[equation omitted — 66 chars of source]

where $Y$ is the scalar outcome variable, $X$ is the scalar treatment variable, $\omega^0$ is a $k$-vector of observable control variables, $W_2$ is the unobserved control variable (understood as an index of any number of unobserved control variables), $\varepsilon$ is a stochastic error term, $\alpha$ is a scalar parameter, $\beta$ is the scalar parameter of interest to the researcher (which captures the treatment effect) and $\Psi$ is a $1 \times k$ vector of parameters. Let us denote the R-squared from this hypothetical long regression as $R_{max}$ and note that $R_{max}$ is unknown because this regression cannot be estimated (because $W_2$ is unobserved).

The researcher, instead, estimates an `intermediate' regression, by regressing $Y$ on $X$ and $\omega^o$ (where $W_2$ is left out of the regression). Let us denote the estimated coefficient on $X$ and the R-squared in the intermediate regression as $\tilde{\beta}$ as $\tilde{R}$, respectively. For analytical purposed, let us also consider two more regression models: (a) a `short' regression, where $Y$ is regressed on $X$; and (b) an auxiliary regression, where $X$ (treatment variable) is regressed on $\omega^o$ (all the observable control variables). Let us denote the estimated coefficient on $X$ and the R-squared from the short regression as $\mathring{\beta}$ and $\mathring{R}$, respectively; let us denote by $\tau_X$, the variance of the residual from the auxiliary regression. Finally, let $\sigma_X^2$ denote the variance of $X$ (treatment variable), and let $\sigma_Y^2$ denote the variance of $Y$ (outcome variable).

Proportion of Selection

Following oster-2019, let us define the measure of proportional selection on unobservables as,

equation[equation omitted — 80 chars of source]

where $\sigma_{1X}=\mathrm{Cov}\left( W_1, X\right) $, $\sigma_{2X}=\mathrm{Cov}\left( W_2, X\right) $, $\sigma^2_1=\mathrm{Var}\left( W_1\right) $, and $\sigma^2_2=\mathrm{Var}(W_2)$, and $W_1=\Psi \omega^0$ (an index of the observable controls). Let us try to understand the meaning of this parameter, $\delta$?

Consider a linear projection wooldridge of the treatment variables on the index of the observables, i.e. \[ X = \alpha_0 + \alpha_1 W_1 + u_1. \] Since $u_1$ is orthogonal to $W_1$ by definition of linear projections, we have

equation[equation omitted — 60 chars of source]

Now consider another linear projection of the treatment variables on the index of the unobservables, i.e. \[ X = \delta_0 + \delta_1 W_2 + u_2 \] and note, once again using the property of linear projections, that

equation[equation omitted — 60 chars of source]

Now we see clearly that the measure of proportional selection is just the ratio of the two coefficients from the two linear projections, i.e.

equation[equation omitted — 69 chars of source]

We will return to this expression when we try to look critically at the use of $\delta=1$ as a lower bound.

Cubic Equation in Bias

Let $\nu$ denote the asymptotic bias in the treatment effect estimated from the intermediate regression, i.e.

equation[equation omitted — 62 chars of source]

where $\mathrm{plim} \tilde{\beta}$ denotes the probability limit of $\tilde{\beta}$ as the sample size increases without bound.\footnote{Note that all estimators in this paper are functions of the sample size, $N$. We suppress this dependence for notational simplicity.} For $j = 1,2, \ldots, J$, let $\omega_{jj}=\mathrm{Var}(\omega^o_j)$ denote the variance of the $j$-th observed control variable, for $j \neq k$, let $\omega_{jk}=\mathrm{Cov}(\omega^o_j,\omega^o_k)$ denote the covariance between the $j$-th and $k$-th observed control variables, and for $j = 1,2, \ldots, J$, let $\sigma_{2j}=\mathrm{Cov}(W_2,\omega^o_j)$ denote the covariance between $W_2$ (the index of unobserved confounders) and the $j$-th observed control variable.

propositionIf, for $j,k = 1,2, \ldots, J$, $j \neq k$, \begin{equation} \omega_{jk}=0, \end{equation} and for $j = 1,2, \ldots, J$, \begin{equation} \sigma_{2j}=0, \end{equation} then we have, \begin{equation} \left( \mathring{\beta} - \tilde{\beta}\right) \overset{p}{\to} \frac{\sigma_{1X}}{\sigma_X^2} - \nu \left( \frac{\sigma_X^2 - \tau_X}{\sigma_X^2}\right), \end{equation} \begin{equation} \left( \tilde{R} - \mathring{R}\right) \sigma_y^2 \overset{p}{\to} \sigma_1^2 + \tau_X \nu^2 - \frac{1}{\sigma_X^2}\left( \sigma_{1X} + \nu \tau_X\right)^2, \end{equation} \begin{equation} \left(R_{max} - \tilde{R}\right) \sigma_y^2 \overset{p}{\to} \nu \left( \frac{\sigma_1^2 \tau_X}{\delta \sigma_{1X}} - \nu \tau_X\right) \end{equation} where $\overset{p}{\to}$ refers to convergence in probability as the sample size increases without bound.

The details of the proof can be found in Appendix (ref). The usefulness of the above result is that it leads to a cubic equation in the bias. To see this, note that the equations in ((ref)), ((ref)) and ((ref)) constitute a system of $ 3 $ equations in $ 3 $ unknowns: $\sigma^2_1$, the variance of $ W_1 $; $\sigma_{1X}$, the covariance of $ W_1 $ and $ X $ (treatment variable); and $\nu$ (the bias of the treatment effect in the intermediate regression). Algebraic manipulation can reduce the three equations into a single cubic equation in $\nu$ given by,

equation[equation omitted — 70 chars of source]

where,

align[align omitted — 633 chars of source]

This gives us the crucial result about the roots of the cubic equation in ((ref)).

propositionConsider the cubic equation given in ((ref)) with coefficients given in ((ref)), ((ref)), ((ref)), and ((ref)). \begin{enumerate} • When the cubic equation has one (unique) real root, denote it by $\nu_{U}$ and let $\beta^*_U=\tilde{\beta}-\nu_U$. Then \[ \beta^*_U \xrightarrow{p} \beta. \] • When the cubic equation has three (non-unique) real roots, denote them by $\nu_{NU,1},\nu_{NU,2},\nu_{NU,3}$ and for $i=1, 2, 3$, let $\beta^*_{NU,i}=\tilde{\beta}-\nu_{NU,i}$. Then, for $i=1$ or $i=2$ or $i=3$, \[ \beta^*_{NU,i} \xrightarrow{p} \beta. \] \end{enumerate}

The proof follows immediately from the fact that $\nu$ is defined to be the asymptotic bias in the treatment effect estimated from the intermediate regression.\footnote{This result is given in oster-2019.} For us, it is more important to pay attention to the two maintained assumptions of the whole analysis stated explicitly in Proposition (ref): (a) the observed controls are pairwise uncorrelated, i.e. for $j \neq k$, $\omega_{jk}=0$, and (b) the unobserved confounder is uncorrelated with each of the observed controls, i.e. $\sigma_{2j}=0$. Both these are strong assumptions and in working out the proof in Appendix (ref), I point out exactly where they are needed. While oster-2019 claims that these assumptions do not imply any loss of generality, the derivation in Appendix (ref) shows that that is not true. One way to see this is to note that the estimator for the treatment effect, $\beta^*$, is a function of the root of cubic equation in ((ref)). The cubic equation arises from ((ref)), ((ref)) and ((ref)), and these three equations, in turn, cannot be derived without the two orthogonality assumptions stated in Proposition (ref). Hence, the estimator relies crucially on the two orthogonality assumptions, contrary to the argument in oster-2019. Having noted these caveats, let me now turn to the main part of this paper, which is a discussion of a novel algorithm to compute omitted variable bias and bias-adjusted treatment effects (BATE).

Bounds for the Treatment Effect

Real Root as Bias and Overall Strategy

Finding the real roots of the cubic equation in ((ref)) is the key to constructing proper bounds for the `true' treatment effect. This follows from Proposition (ref). To do so we note that the coefficients of the cubic equation are composed of all known quantities other than the following two: $R_{max}$ (the R-squared in the hypothetical long regression), and $\delta$ (the measure of proportional selection on unobservables). My strategy consists of the following steps.

First, I select a bounded box of the $\left( \delta, R_{max}\right) $ plane that is relevant for the econometric analysis in question by choosing lower and upper bounds for $\delta$ and $R_{max}$, i.e. we choose $\delta_{low},\delta_{high}$, and $R_{low},R_{high}$, such that $\delta_{low} < \delta < \delta_{high}$ and $R_{low} < R_{max} < R_{high}$. This defines a bounded box, $B$. Second, I demarcate the portion of the bounded box where the cubic ((ref)) is guaranteed to have a unique real root (URR) from the portion where it has three real roots (NURR). Third, I impose a sufficiently granular grid of $N^2$ points on B and compute all real roots on the grid points in URR. Fourth, I use continuity of the roots of a polynomial equation with respect to its coefficients to select roots from the grid points in NURR by starting with the roots on the boundary of URR and NURR, and then covering all the grid points on NURR in pre-specified small steps. After I have covered all the $N^2$ points of the grid, I will have an empirical distribution of the omitted variable bias (the selected roots of the cubic equation), and an empirical distribution of the bias-adjusted treatment effect (treatment effect from intermediate regression minus the selected root of the cubic equation).

I will now discuss the details of an algorithm that will implement these ideas.

Algorithm

The algorithm relies on two results, the first relating to the roots of a cubic equation and the second regarding continuity of the roots of any polynomial equation with respect to its coefficients.

propositionConsider the cubic equation in ((ref)) and let $p=(3ac-b^2)/3a^2$ and $q=(27a^2d + 2b^3 - 9abc)/27a^3$. \begin{enumerate} • If $ 27q^2 + 4p^3>0$, then the cubic equation has a unique real root. Let us call the region of the $(\delta, R_{max})$ plane over which this condition is satisfied as URR, the unique real root region. • If $ 27q^2 + 4p^3 \leq 0$, then the cubic equation has three real roots. Let us call the region of the $(\delta, R_{max})$ plane over which this condition is satisfied as NURR, the nonunique real root region. \end{enumerate}
proofThis is a well-known result. See for instance, hellesland-etal-2013 and Appendix (ref) for details.

The basic idea driving the algorithm is to see how the bounded box, $B$, formed by the choice of the range of values for $\delta$ and $R_{max}$, intersects with the URR and NURR regions, and then to solve the cubic equation on relevant grids covering $B$, starting with the part where $B$ intersects with URR and then extending to NURR using continuity. The last and crucial step of the algorithm, therefore, relies on the following well-known result from complex analysis alexanderian-2013: the roots of any polynomial equation are continuous functions of the coefficients of the polynomial, and real roots, upon small changes in the coefficients, remain real.

propositionConsider the cubic equation in ((ref)), and assume that $b^2 \neq 3ac$. The roots of the cubic equation are continuous functions of $\delta$ and $R_{max}$. Moreover, if $\delta$ and $R_{max}$ are changed by sufficiently small magnitudes, the real roots of multiplicity one remain real.
proofNote that the coefficients of cubic equation in ((ref)) are polynomials in $\delta$ and $R_{max}$. Hence, the coefficients are continuous functions of $\delta$ and $R_{max}$ (because polynomials are everywhere continuous functions). Now, we use the results that the roots of any polynomial are continuous functions of the coefficients of the polynomial alexanderian-2013. This implies that the roots are compositions of continuous functions. Hence, they are continuous functions of $\delta$ and $R_{max}$. The condition, $b^2 \neq 3ac$, ensures that the real roots have multiplicity of one. To see this, note that the critical points of the cubic polynomial are given by the values of $\nu$ where the first derivative, $3ax^2+2bx+c$, is zero. These are given by \[ \nu_c = \frac{-b \pm \sqrt{b^2 - 3ac}}{3a}. \] The point of inflection of the cubic is given by the values of $\nu$ where the second derivative, $6ax+2b$, is zero. Hence, it is given by \[ \nu_i = -\frac{b}{3a}. \] The real root of the cubic equation has multiplicity of $3$ or $1$. A real root has multiplicity of $3$ if and only if $\nu_c=\nu_i$, if and only if $b^2=3ac$. Hence, a necessary and sufficient condition for real roots to have multiplicity of $1$ is that $b^2 \neq 3ac$. This implies, using Theorem 3.5 in alexanderian-2013, that small perturbations of $\delta$ and $R_{max}$ will produce real roots that are close to the original real roots.

The continuity result is crucial because it allows me to sequentially solve the cubic over grids imposed on $ B $. We start by solving for the cubic equation over grid points in URR (where each point gives a unique real root) and then incrementally move to cover the grid points on the NURR (by selecting among the three real roots that one whose difference in absolute value is least with respect to the unique real root computed on the closest grid point in URR). Since the roots of the cubic equation are continuous in $\delta$ and $R_{max}$, if we are at a point in the NURR that is within a “small” distance from a grid point in URR, then the real root at the former will be “close” to the unique real root at the latter. This is what allows me to choose one of the three real roots at the former point.

Case 1

In the first case, the bounded box, $B$, is wholly contained in URR. I solve the cubic on a sufficiently granular grid that covers the bounded box. For each point on the grid, using proposition (ref), we can find a unique real root, $\nu$, and use it to construct $\beta^* = \tilde{\beta}-\nu$. If the total number of points on the grid is $M$, we thereby generate a $M$-vector of $\beta^*$ values. The empirical distribution of this $\beta^*$ gives an approximation of the distribution of $\beta$ (the true treatment effect) for the case when the unknown parameters, $\delta$ and $R_{max}$, can range over the chosen bounded box, $B$. Hence, the empirical distribution of $\beta^*$ allows us to construct approximate confidence intervals for $\beta$. The accuracy of the approximation will increase with the number of points in the grid covering the bounded box.

Case 2

In the second case, the bounded box, $B$, is partly contained in URR and partly contained in NURR. This situation is depicted in Figure (ref), where $B$ is by the (blue) box contained in the Cartesian product of $[\delta_{low},\delta_{high}]$ and $[\tilde{R}, R_{high}]$. The (red) curve demarcates the plane into URR and NURR: the region above the curve is URR and the region below is NURR.

center[center omitted — 51 chars of source]

Let $S_1 = B \cap URR$, and $S_2 = B \cap NURR$. For $S_1$, we use the same method as in case 1. This generates, for instance, a $M_1^2$-vector of $\beta^*$ values. The real challenge is to select the `correct' real root for grid points in $S_2$ because each point in $S_2$ generates three real roots. To accomplish this task, we do the following:

itemize• We impose a granular grid on $S_2$ of $M_2^2$ points. • We identify points of $S_2$ that are within a pre-specified `small' distance of points in $S_1$, $ e $. What is `small' is determined by the choice of $e$ (which, in turn, determines $M_2$). As $e$ decreases, the distance separating the points on the two sides of the URR/NURR boundary falls. As the distance falls, the computational burden of the method increases, primarily because the number of points in the grid in NURR that has to be compared with points in URR increases and, secondarily because, the cubic equation has to be solved at a larger number of points. This is a trade-off that is inherent to this grid search methodology. • We call the points of $S_2$ identified in the previous step as the set of `closest' points to $S_1$ and denote this set as $S_{21}$. For every point in $S_{21}$, we compute the three real roots of the cubic equation in ((ref)). We select the real root that is closest, in absolute value, to the unique real root found for the corresponding closest point of $S_1$. This is the `correct' root and is justified by proposition (ref). Figure (ref) depicts the selection of the `correct' root at a grid point $N_1$, a point in the $NURR$ region, using one of the closest points in the URR region, $U_1$. • Next, we identify points of $S_2$ that are within a pre-specified small distance of the set $S_{21}$. We call these the `closest' points to $S_{21}$ and denote this set as $S_{22}$, For every point in $S_{22}$, we compute the three real roots of the cubic equation in ((ref)). We select the real root that is closest, in absolute value, to the real root that was selected (in the previous step) for the corresponding closest point of $S_{21}$. Once again, this is justified by proposition (ref). • We continue this process until we have exhausted all the points in $S_2$. This generates, for instance, a $M_2$-vector of $\beta^*$ values. • We combine the $M_1$ and $M_2$ vectors of $\beta^*$ values. We thereby generate a $M$-vector of $\beta^*$ values, where $M=M_1+M_2$, Now, following the same logic as in case 1, we can generate an empirical distribution of $\beta^*$ and use it to construct approximate confidence intervals for $\beta$.

Case 3

In the third, and final, case, the bounded box, $B$, is wholly contained in NURR. In this case, we extend the dimension of the bounded box, $B$, to the extent that is necessary to generate a non-empty intersection with URR. Once we have generated such a bounded box, we are back to case 2. Hence, we now use the method outlined for case 2 to compute the empirical distribution of the bias and $\beta^*$.

An Application

In this section, I report results of applying my method to the analysis of maternal behavior on child outcomes that was discussed in oster-2019. The substantive issue under investigation in this application is the impact of maternal behaviour on child outcomes. In particular, two child outcomes are studied: a child's standardized IQ score and a child's birth weight. In the study of child IQ, three treatment variables are used in turn: months of breastfeeding, any drinking of alcohol in pregnancy, and an indicator for being low birthweight and preterm. In studying child birthweight, two treatment variables are used, one by one: maternal smoking during pregnancy, and maternal drinking during pregnancy. The following control variables are used for both studies: child race, maternal age, maternal education, maternal income, maternal marital status. The question of interest is whether the treatment variables, each on their own, have any causal impact on the outcome variables.

Analysis of Bias and Bounding Sets

In Table (ref), I present the estimates of the treatment effect from the short and intermediate regressions. These results replicate the corresponding results in Table 3 in oster-2019. For instance, if we look at the first row of Table (ref), we see that the effect of (months of) breastfeeding on child IQ is $ 0.044 $ (column 1) in the short regression and $ 0.017 $ (column 4) in the intermediate regression. Thus, breastfeeding has a positive impact on a child's IQ score. Moving from the short to the intermediate regression, the R-squared increases from $ 0.045 $ (column 3) to $ 0.256 $ (column 6). We can read all the other numbers in columns 1 through 6 in a similar manner. Since these models are likely to have omitted variables, we would like to quantify the effect of the omitted variable bias.

center[center omitted — 62 chars of source]

I begin the analysis of bias by constructing two bounded boxes on the ($\delta, R_{max}$) plane. Box 1 is given by the Cartesian product of ($0.01<\delta<0.99$) and ($\tilde{R}<R_{max}<0.61$), and Box 2 is defined by the Cartesian product of ($1.01<\delta<3.99$) and ($\tilde{R}<R_{max}<0.61$). I use two bounded boxes so that I can compare results between a case when the relative selection on unobservables, $\delta$, is lower than $1$ to a case when it is larger than $1$. The lower limit of $R_{max}$ comes from the result that $\tilde{R} \leq R_{max}$ because the hypothetical long regression has at least one more regressor than the intermediate regression. The upper limit of $R_{max}=0.61$ (first three regression models) and $R_{max}=0.53$ (last two regression models) is taken to tally with the same assumption in oster-2019.

center[center omitted — 60 chars of source]

On each bounded box, I use proposition (ref) to identify the $ URR $ and $NURR$ areas. To construct the grid, I use a step size of $0.01$ and then use the algorithm of section (ref) to construct empirical distributions of the omitted variable bias and the bias-adjusted treatment effect (BATE). I summarize the results of this bounding analysis in Table (ref). Region plots showing the demarcation of the bounded boxes into URR and NURR, and contour plots of the estimated bias on the bounded boxes are presented in Figure (ref) through Figure (ref) in the appendix.

Effect of Breastfeeding on Child IQ

The first four rows of Table (ref) refer to the first row in Table (ref) and also to row 1 in oster-2019. In this case, the outcome variable is a child's IQ score and the treatment variable is months of breastfeeding. The first four rows of Table (ref) report quantiles of the empirical distribution of bias and BATE computed over the two bounded boxes displayed in Figure (ref) and (ref) in the appendix. From the second row of Table (ref), we can see that an approximate 95% confidence interval for the treatment effect would be $(-0.022, 0.017)$ if we chose to use the first bounded box (Box 1) as the relevant region over which to conduct the bounding exercise. On the other hand, if we chose to use the second bounded box (Box 2) as the relevant region over which $\delta$ and $R_{max}$ can vary, then the approximate 95% confidence interval for the treatment effect is given by $(-0.116, 0.017)$. In both cases, the substantive conclusion would be that the treatment effect can be completely nullified once the effect of omitted variables are taken into account.

Effect of Drinking during Pregnancy on Child IQ

The second block of four rows of Table (ref) refer to the second row in Table (ref) and the same row in oster-2019. In this case, the outcome variable is the same as in the previous analysis: a child's IQ score. The treatment variable is whether the mother reported drinking alcohol during pregnancy. From the sixth row of Table (ref), we can see that an approximate 95% confidence interval for the treatment effect would be $(-0.104, 0.050)$ if we chose to use the first bounded box (Box 1) as the relevant region over which to conduct the bounding exercise. On the other hand, if we chose to use the second bounded box (Box 2) as the relevant region over which $\delta$ and $R_{max}$ can vary, then the approximate 95% confidence interval for the treatment effect is given by $(-1.107, 0.050)$. In both cases, once again, the substantive conclusion would be that the treatment effect is likely to be completely nullified once the effect of omitted variables are taken into account.

Effect of LBW + Preterm on Child IQ

The third block of four rows of Table (ref) refer to the third row in Table (ref) and the same row in oster-2019. In this case, the outcome variable is the same as in the previous analysis: a child's IQ score. The treatment variable is whether the mother had low birth weight and was prematurely born. From the tenth row of Table (ref), we can see that an approximate 95% confidence interval for the treatment effect would be $(-0.125, -0.054)$ if we chose to use the first bounded box (Box 1) as the relevant region over which to conduct the bounding exercise. On the other hand, if we chose to use the second bounded box (Box 2) as the relevant region over which $\delta$ and $R_{max}$ can vary, then the approximate 95% confidence interval for the treatment effect is given by $(-0.125, 0.175)$. The substantive conclusion now changes depending on which box the researcher uses. If Box 1 is the relevant region to conduct the bounding exercise, then the treatment effect will remain negative and significantly different from zero even after we have accounted for omitted variable bias. On the other hand, if Box 2 is the relevant region to use, we cannot rule out the fact that the treatment effect might be completely nullified once the effect of omitted variables are taken into account.

Effect of Smoking during Pregnancy on Child Birth Weight

The fourth block of four rows of Table (ref) refer to the fourth row in Table (ref) and the same row in oster-2019. In this case, the outcome variable is a child's birth weight. The treatment variable is whether the mother reported smoking during pregnancy. From the fourteenth row of Table (ref), we can see that an approximate 95% confidence interval for the treatment effect would be $(-172.261, -21.121)$ if we chose to use the first bounded box (Box 1) as the relevant region over which to conduct the bounding exercise. On the other hand, if we chose to use the second bounded box (Box 2) as the relevant region over which $\delta$ and $R_{max}$ can vary, then the approximate 95% confidence interval for the treatment effect is given by $(-2070.193, 117.522)$. In first case, when $\delta$ is restricted to lie between $0$ and $1$, the treatment effect is likely to remain intact, i.e. different from zero, even when the effect of omitted variables are taken into account. On the other hand, if $\delta$ is allowed to be larger than $1$, then the approximate 95% confidence interval of the empirical distribution of the treatment effect contains zero. Hence, we cannot rule out the fact that allowing for the effect of the omitted variables will wipe out the treatment effect.

Effect of Drinking during Pregnancy on Child Birth Weight

The fifth block of four rows of Table (ref) refer to the fifth row in oster-2019. In this case, the outcome variable is a child's birth weight. The treatment variable is whether the mother reported smoking during pregnancy. From the eighteenth row of Table (ref), we can see that an approximate 95% confidence interval for the treatment effect would be $(-14.149, 2.895)$ if we chose to use the first bounded box (Box 1) as the relevant region over which to conduct the bounding exercise. On the other hand, if we chose to use the second bounded box (Box 2) as the relevant region over which $\delta$ and $R_{max}$ can vary, then the approximate 95% confidence interval for the treatment effect is given by $(-13.149, 138.120)$ in the twentieth row of Table (ref). Here we have another example where the substantive conclusion does not depend on which box is chosen as the correct one. Irrespective of whether we choose Box 1 or Box 2, the conclusion would be that the treatment effect is likely to be completely nullified once the effect of omitted variables are taken into account.

Step size of grid?

The choice of the step size of the grid over which the bias is computed needs to balance an important trade off. On the one hand, the smaller the size of the steps used in constructing the grid, the better the approximation of the bounding set to the the true confidence interval. On the other hand, the smaller the size of the steps, the larger than number of grid points. Hence, the more computationally intensive the implementation of the algorithm. To assess the step size on this trade off, in Table (ref), I report the confidence interval for the treatment effect for different step sizes. For this exercise, I use the model where the outcome variable is the child's IQ score and the treatment variable is the months of breast feeding (this case is reported in row 1, Table (ref)).

center[center omitted — 55 chars of source]

In Table (ref), I report the quantiles of the empirical distribution of $\beta^*$ (the bias-adjusted treatment effect) for step sizes of $e=1/25$, $e=1/50$, $e=1/100$, $e=1/250$, and $e=1/500$. From the results in the table, we see that the quantiles of the empirical distribution of $\beta^*$ remains largely unchanged for step sizes lower than $e=1/50$. On the other hand, the computing time increases rapidly as the step size is reduced. In terms of the computation-accuracy trade off, a choice of $e=1/100$ seems to good because it gives us a fairly accurate confidence interval without consuming too much computing power. This is the rationale behind by choice of $e=1/100$ for the analyses reported in Table (ref).

Comparison with Oster's Methodology

oster-2019 proposed two methods for dealing with the problem of non-uniqueness. The first involves constructing identified sets under the assumption of equal selection, i.e. $\delta=1$, and the second involves computing the value of $\delta$ that is necessary to force the treatment effect to become zero. I would now like to point towards some problems in both these methods.

Identified Sets Under Equal Selection

Bias Under Equal Selection

The first method proposed by oster-2019 is to compute `identified' sets under the assumption of equal selection, i.e. $\delta=1$. To compute these `identified' sets, one has to first solve for the bias under equal selection. If we impose the restriction that $\delta=1$ on the coefficients of the cubic in ((ref)) we get,

align*[align* omitted — 437 chars of source]

which converts the cubic in ((ref)) to a quadratic equation in $\nu$,

equation[equation omitted — 63 chars of source]

where the coefficients of this quadratic are given by,

align[align omitted — 478 chars of source]

The solutions of the quadratic in ((ref)) are given by \[ \nu = \frac{-c_1 \pm \sqrt{c_1^2 - 4d_1b_1}}{2b_1}, \] which are noted in Corollary 1 in oster-2019.

propositionThe quadratic equation in ((ref)) either has a unique real root or two distinct real roots. It does not have any complex roots.
proofThe proof follows by noting that the discriminant of this quadratic equation is non-negative, i.e. $c_1^2-4d_1b_1 \geq 0$, because $c_1^2 \geq 0$, and \begin{align*} -4d_1b_1 & = -4 \left\lbrace \left( R_{max} - \tilde{R} \right) \sigma^2_Y \left( \mathring{\beta} - \tilde{\beta}\right) \sigma^2_X \right\rbrace \left\lbrace -\tau_X \left( \mathring{\beta} - \tilde{\beta}\right) \sigma^2_X \right\rbrace \\ & = 4 \left( R_{max} - \tilde{R} \right) \sigma^4_X \sigma^2_Y \tau_X \left( \mathring{\beta} - \tilde{\beta}\right)^2 \\ & \geq 0 \end{align*} where the last inequality follows because $R_{max} \geq \tilde{R}$.

Identified Sets are not Unique

The bounding sets for the `true' treatment effect, for instance reported in column 5 in Table 3 in oster-2019are defined as $[ \tilde{\beta}, \beta^*(R_{max},\delta=1) ] $, where \[ \beta^*(R_{max},\delta=1) = \tilde{\beta} - \text{root of the quadratic equation in (\ref{eq:quad})}. \]

The implication of proposition (ref) is that, in general, there will be two real roots of the quadratic equation in ((ref)). Hence, in general, there will be two values of the bias in the treatment effect, and hence two values of $\beta^*$. Without the extremely restrictive assumption that the discriminant of the quadratic equation in ((ref)) is identically equal to zero, it is not possible to arrive at a unique `identified' set when $\delta=1$. Since oster-2019 uses the bias-adjusted treatment effect computed under the assumption of $\delta=1$ in constructing her `identified sets', it is not clear how one of these two sets are chosen.

To be more specific, I follow oster-2019 and and construct identified sets by choosing $R_{max}=0.61$ for the regressions corresponding to the first three rows of Table (ref) and $R_{max}=0.53$ for the regressions corresponding to the last two rows of Table (ref). I report the results from this exercise in Table (ref) and follow the same row numbering as Table 3 in oster-2019. Let us begin by noting, in column 1, the magnitude of the discriminant of the quadratic equation in ((ref)). We can see that the discriminant is always positive. Hence, in each case, there are two real roots. I use the first real root to define $\beta^*_1$ (column 2) and the second real root to define $\beta^*_2$ (column 3). The important conclusion to draw is that one cannot get a uniquely identified set.\footnote{The root selection algorithm outlined in section (ref) requires at least one set of coefficients to produce a unique real root. This will not work for the quadratic case because proposition (ref) shows that we do not have a unique real root for any set of coefficients.}

center[center omitted — 59 chars of source]

For instance, for the first row (where child IQ is the outcome and months of breastfeeding is the treatment), the first identified set is $[0.017, 0.375]$ and the second identified set is $[-0.034, 0.017]$. In row 1, Table 3, oster-2019 reports the second of these as the identified set. But there is no reason offered for this choice. What is basis on which one can choose between the two different identified sets? The same problem affects all the five rows in Table 3, oster-2019. No reason is given for choosing one over the other identified set. This is especially important because in three cases out of five, the conclusion of the bounding analysis would change if one set was chosen rather than the other. For instance, in the case of the first row (where child IQ is the outcome and months of breastfeeding is the treatment), the first identified does not include zero; the second identified set does include zero. The same is true for the analysis represented by row 2 and row 5.

If the quadratic in ((ref)) does not have a unique root for $\delta=1$ and $R_{max}=0.61$ (or $R_{max}=0.53$), then how can oster-2019 report one identified set? There seem to be two possibilities. First, it seems that she has taken recourse to Assumption 3 in her paper to generate a unique root of the quadratic equation. Assumption 3, in oster-2019, states that the sign of the covariance between the treatment variable and the actual index of observables is the same as the sign of the covariance between the treatment variable and the predicted index of observables. The meaning and import of this assumption is explained thus.

quoteEffectively, this assumes that the bias from the unobservables is not so large that it biases the direction of the covariance between the observable index and the treatment. Under Assumption 3, if $\delta=1$, there is a unique solution oster-2019.

It is not clear how the sign restriction on the covariance between the treatment variable and the index of observables can generate a unique root of the quadratic equation in ((ref)). oster-2019 does not provide a proof of this important claim in the paper or in the appendix.

The second possibility is that among the two real roots of the quadratic equation, oster-2019 choose the one that implies a lower absolute magnitude of omitted variable bias. In each of the five cases reported in oster-2019, one can see this by matching the identified sets that emerge from Table (ref). If this is the implicit reasoning behind the choice among the two real roots of the quadratic equation in ((ref)), then it needs to be justified on theoretical grounds. As it stands, it is unjustified and comes across as an ad hoc assumption. Why should we assume that the omitted variables are such that they will produce the lower of the two possible magnitudes of bias?

Does $\delta^*$ Provide Useful Information?

Let us consider the second strategy proposed in oster-2019 to deal with non-uniqueness, i.e. computing the value of $\delta$ (the relative selection on unobservables) that is consistent with a zero treatment effect, denoting this as $\delta^*$, and drawing conclusions about the problem of omitted variable bias by comparing $\delta^*$ with $1$. This strategy has at least three problems. First, it does not provide us with any identified set of values of the true treatment effect, $\beta$. It only gives us one number, $\delta^*$. Second, in many cases, as I demonstrate below, $\delta^*$ can be extremely sensitive to the choice of $R_{max}$. Even a small error in choosing $R_{max}$ can greatly magnify $\delta^*$ and thereby lead to incorrect conclusions. Third, in some cases, the conclusions drawn from $\delta^*$ does not accord with the conclusions drawn on the basis of the bounding set that I have proposed above. When there is such a conflict, it seems better to use the bounding set than to rely on one number, $\delta^*$, because the former strategy is more robust.

What is $\delta^*$?

Recall that $\delta^*$ is the degree of selection on unobservables that is consistent with zero treatment effect. If treatment effect is zero, i.e. $\beta=0$, then $\tilde{\beta}-\nu=0$. Hence, $\nu=\tilde{\beta}$. By plugging $\nu=\tilde{\beta}$ in ((ref)), we get a relationship between $\delta$ and $R_{max}$ that can be expressed with the following function,

equation[equation omitted — 140 chars of source]

where

align*[align* omitted — 514 chars of source]

are known constants that can be computed once we have estimated the short, intermediate and auxiliary regressions. When we plug in a value of $R_{max}$ in ((ref)), we get the corresponding $\delta^*$ by evaluating the function at that value of $R_{max}$.

For the function in ((ref)) to be meaningful and useful, we need to impose some restrictions. First, if $\tilde{\beta}=0$, then $C=0$, and hence $\delta=0$ for all values of $R_{max}$. This is not interesting. So, we assume $\tilde{\beta} \neq 0$. Second, given that $\tilde{\beta} \neq 0$, if $A=0$, then $\delta$ is the constant function. It does not vary with $R_{max}$. Once again, this is not interesting for us. Hence, we impose the condition that $A \neq 0$. Third, the function is not defined at the point $R_{max}=R^*$, where $R^*=\tilde{R}-(B/A)$. Hence, we need to exclude this point from the domain of definition of the function. There are two cases to consider.

Case 1

In this case, $R^* \notin [ \tilde{R}, 1 ] $. Hence, the function $f$ is defined on every point in $[ \tilde{R}, 1 ]$. Moreover, it is differentiable on $(\tilde{R}, 1)$ because it is a rational function and is defined on each point on this closed interval. The derivative of the function in ((ref)), on the open interval, is given by

equation[equation omitted — 121 chars of source]
propositionThe function in ((ref)) is monotone.
proofIf $C=0$, then $f'=0$ and $f$ is a constant function, i.e. both increasing and decreasing. If $C \neq 0$, we have either that $A$ and $C$ have the same sign or that they have opposite signs. If $A$ and $C$ are of the same sign, then using ((ref)), we see that $f'<0$. Hence, $f$ is strictly decreasing. If $A$ and $C$ are of opposite signs, then ((ref)) shows that $f'>0$ and $f$ is strictly increasing.

Two examples of the $f$ function are given in Figure (ref), one where its graph is upward sloping and another where its graph is downward sloping.

center[center omitted — 55 chars of source]

I give sufficient conditions for these two types of the $f$ function for the case when $R^* \notin [ \tilde{R}, 1 ] $ and provide some intuition for these conditions.

propositionIf $0 < \tilde{\beta} <\mathring{\beta}$ or $ \mathring{\beta} < \tilde{\beta} < 0$, then the function in ((ref)) has a downward sloping graph.
proofIf $0 < \tilde{\beta} <\mathring{\beta}$ or $ \mathring{\beta} < \tilde{\beta} < 0$, then it can be easily verified that the parameters $A$ and $C$ are of the same sign. Using ((ref)), we get the result.

Intuitively, what this sufficient condition says is this: if in moving from the short to the intermediate regression, the treatment effect moves towards zero without changing sign, then $\delta$ and $R_{max}$ have a negative relationship among themselves.

propositionIf either of the following two conditions are satisfied, \begin{enumerate} • $\tilde{\beta} > (\sigma_X^2/\tau_X)\mathring{\beta}$, and \begin{enumerate} • $\tilde{\beta}^2 < (\sigma_X^2/\tau_X)\mathring{\beta}^2 + (\sigma_Y^2/\tau_X)(\tilde{R}-\mathring{R})$, if $\tilde{\beta}>0$, or • $\tilde{\beta}^2 > (\sigma_X^2/\tau_X)\mathring{\beta}^2 + (\sigma_Y^2/\tau_X)(\tilde{R}-\mathring{R})$, if $\tilde{\beta}<0$, \end{enumerate} • $\tilde{\beta} < (\sigma_X^2/\tau_X)\mathring{\beta}$, and \begin{enumerate} • $\tilde{\beta}^2 > (\sigma_X^2/\tau_X)\mathring{\beta}^2 + (\sigma_Y^2/\tau_X)(\tilde{R}-\mathring{R})$, if $\tilde{\beta}>0$, or • $\tilde{\beta}^2 < (\sigma_X^2/\tau_X)\mathring{\beta}^2 + (\sigma_Y^2/\tau_X)(\tilde{R}-\mathring{R})$, if $\tilde{\beta}<0$, \end{enumerate} \end{enumerate} then the function in ((ref)) has a upward sloping graph.
proofNote that $A<0$ iff $\tilde{\beta} > (\sigma_X^2/\tau_X)\mathring{\beta}$. Similarly, note that $C>0$ iff $ \tilde{\beta}\left[ \sigma_X^2 \tau_X\mathring{\beta}^2 + \sigma_Y^2 \tau_X(\tilde{R}-\mathring{R}) - \tau_X^2 \tilde{\beta}^2\right] <0 $. Now, using ((ref)), we get the result.

One way in which this sufficient condition will be satisfied is this: if in moving from the short to the intermediate regression, the treatment effect changes sign and the difference of their absolute values is sufficiently large, then $\delta$ and $R_{max}$ has a positive relationship among themselves. As an example, consider $\tilde{\beta}=1$ and $\mathring{\beta}=-1.5$. Since $\sigma_X^2>\tau_X>0$, $\sigma_Y^2>0$ and $\tilde{R}>\mathring{R}$, this choice satisfies condition 1 (a) in proposition (ref). As another example, consider a scenario where $\tilde{\beta}=-1$ and $\mathring{\beta}=1.5$. Here, we can see that condition 2 (b) is satisfied.

Interpretation of Slope. When the graph of the function in ((ref)) is downward sloping, then this implies that for the treatment effect to be zero ($\beta=0$), a relatively high value of $R_{max}$ will be associated with a relatively low value of $\delta$. This can be interpreted in two different ways. On the one hand, this means that if the omitted variable is relatively more important in explaining the variation in the outcome variable than the included controls (high $R_{max}$), then only a small degree of selection on unobservables (low $\delta$) will ensure that the treatment effect is zero. If the degree of selection on unobservables is high, then the treatment effect is unlikely to be reduced to zero. On the other hand, it also means that if the degree of selection on unobservables is low, then only a relatively high importance of the omitted variable in explaining variation in the outcome variable compared to the included controls (high $R_{max}$) can ensure that the treatment effect is reduced to zero. If the omitted variable is relatively less important in explaining the variation in the outcome variable compared to the included controls, then the treatment effect cannot be washed out due to omitted variable bias.

On the other hand, when graph of the function in ((ref)) is upward sloping, exactly the opposite interpretation is valid. For the treatment effect to be zero ($ \beta=0 $), a relative high value of $R_{max}$ will be associated with a relatively high value of $\delta$. Once again, this can be interpreted in two different ways. On the one hand, this means that if the omitted variable is relatively more important in explaining the variation in the outcome variable (high $R_{max}$), then only a high degree of selection on unobservables than the included controls (high $\delta$) can ensure that the treatment effect is zero. A low degree of selection on unobservables will not reduce the treatment effect to zero. On the other, if the degree of selection on unobservable is low (low $\delta$) then only if the omitted variable is also relatively unimportant in explaining the variation of the outcome variable (low $R_{max}$), will the treatment effect be reduced to zero. If the omitted variable explains a relatively large part of the variation in the outcome variable, then the treatment effect cannot be reduced to zero due to omitted variable bias.

Whatever interpretation we accord to $\delta^*$ in one case (negative slope) will have to be completely reversed in the other case (positive slope). Since we cannot a priori rule out one or the other sign of the derivative of the function in ((ref)), if we use $\delta^*$ to draw conclusions about the severity or otherwise of the problem of omitted variable bias, our conclusion remains open to the need for a complete reversal of interpretation.

Case 2

In this case, $R^* \in [ \tilde{R}, 1 ] $. Hence, the function is only defined on the intersection of two half-open intervals, \[ \left\lbrace R_{max} | \tilde{R} \leq R_{max} < R^*\right\rbrace \cup \left\lbrace R_{max} | R^* < R_{max} \leq 1 \right\rbrace. \] The analysis of case 1 can now be applied to the two half-open intervals individually because the function is monotone on each of the half-open intervals. But there is an important implication of this case for using $\delta^*$ to draw conclusions about omitted variable bias. If a researcher computes the value of $\delta^*$ using ((ref)), compares it to $1$ and then draws conclusions about the problem of omitted variable bias, then, if this case holds, the researcher is likely to get very non-robust results. This is because the function is discontinuous at $R^*$. Hence, around $R^*$, the magnitude of $\delta^*$ is extremely sensitive to the choice of $R_{max}$. Even a slight error in choosing $R_{max}$ will greatly magnify the error in the magnitude of $\delta^*$.

$\delta^*$ Does not Match up with the Bounding Sets

In column 4 in Table (ref), I have reported the values of $\delta^*$ that was computed with ((ref)). In row 3, Table (ref), the value of $\delta^*$ is $1.36$. If we followed Oster's methodology, we would conclude that the reported estimate of the treatment effect is reliable, i.e. even after we take account of omitted variable bias, the true treatment effect is likely to remain different from zero. If we turn to the bounding sets reported in the third block of Table (ref), we see that this conclusion is not wholly warranted. This is because, if $\delta>1$, the bounding set, computed according to my methodology, will include zero (row 12 in Table (ref)). For instance, if the omitted variables are relatively more important that the observed control variables in explaining the variation in the treatment variable (breastfeeding), then the relative degree of selection on unobservables would be larger than unity. In that case, the conclusion drawn on the basis of $\delta^*=1.36$, that omitted variable bias is not a problem, would be incorrect.

If we look at row 4 in Table (ref) and compare it with the penultimate block of results in Table (ref), we will see the same problem. In row 4 in Table (ref), the value of $\delta^*$ is $1.08$. If we follow Oster's methodology, then we should conclude that omitted variable bias is not a serious problem. Now turn to row 16 in Table (ref). Using the numbers in that row, we can see that, if $\delta>1$, the bounding set for the treatment effect is $[-2070, 118]$. Hence, the bounding set includes zero. Hence, if the omitted variables are relatively more important than the observed controls in explaining the variation in the treatment variable (drinking during pregnancy), then the relative degree of selection on unobservables might very well be larger than unity. In that case, the conclusion drawn on the basis of $\delta^*=1.08$, that omitted variable bias is a not problem, would be incorrect.

Modified Procedure to Use $\delta^*$

The problem of discontinuity and non-correspondence with bounding sets suggests that the use of $\delta^*$ is fraught with problems. But, if a researcher must use it, then I would suggest a modified procedure. First, the researcher must ascertain whether $R^* \in [ \tilde{R}, 1 ] $, i.e. whether the point of discontinuity lies between $\tilde{R}$ and $1$. If the answer is yes, then the use of $\delta^*$ should be avoided. This is because the point of discontinuity lies in the interval of interest, $[ \tilde{R}, 1 ]$ and makes the analysis very unstable. Second, if $R^* \notin [ \tilde{R}, 1 ] $, then the researcher should ascertain the slope of the graph of the function in ((ref)). Since the interpretation is diametrically opposite depending on whether the sign is positive or negative, the researcher should note and report the sign of the derivative at any one point in the interval (monotonicity ensures that the sign of the derivative does not change). Third, the researcher can now report the value of $\delta^*$ and draw appropriate conclusion about omitted variable bias.

In the last three columns of Table (ref), I have reported these three things for the five regression models I have studied in this paper. In each case, we can see that $R^* \notin [ \tilde{R}, 1 ] $. Hence, it is valid to carry out the $\delta^*$ analysis. We also see that in each case, the slope of the graph of the function in ((ref)) is negative. Hence, with all the caveats noted above, this perhaps allows us to interpret $\delta^*$ as done in oster-2019.

Concluding Comments

Omitted variable bias is an ubiquitous problem in applied econometric work. Quantifying the magnitude of bias and computing bias-adjusted treatment effects is an important area of research. Building on earlier work by aet-2005, in a recent contribution, oster-2019 has proposed a novel methodology to compute bias-adjusted treatment effect when there is proportional selection on observables and unobservables. In this paper, I have argued that while oster-2019 posed the problem correctly, her proposed solutions are problematic. I have instead proposed an alternative methodology to compute bounding sets for the true treatment effect.

My proposed methodology relies on two mathematical results. First, for a cubic equation, it is possible to use the discriminant to demarcate regions of the parameter space where a unique real root is guaranteed. Second, the roots of any polynomial are continuous functions of the coefficients. Using these two ideas, I have proposed an algorithm to compute real roots of the cubic equation and use them to construct an empirical distribution of the bias-adjusted treatment effect (BATE). Using this empirical distribution, one can construct approximate confidence intervals for the true treatment effect.

To conclude the discussion, let me give a quick summary of the proposed methodology for the benefit of applied researchers.

itemize• Estimate the short regression and store the coefficient on the treatment variable as $\mathring{\beta}$ and the R-squared as $\mathring{R}$. • Estimate the intermediate regression and store the coefficient on the treatment variable as $\tilde{\beta}$ and the R-squared as $\tilde{R}$. • Estimate an auxiliary regression by regressing the treatment variable on all the controls that were excluded from the short regression. Store the variance of the residual from this regression as $\tau_X$. • Store the variance of the outcome variable as $\sigma^2_y$ and the variance of the control variable as $\sigma^2_X$. • Form the cubic equation in ((ref)). • Choose a bounded box in the ($\delta, R_{max}$) plane over which $\delta$ and $R_{max}$ can vary. Demarcate the URR and NURR regions in the box. • If the box is completely contained in the URR: Choose a $N \times N$ grid to cover the $URR$ area and solve the cubic at each point on the grid. Collect the $N^2 \times 1$ vector of real roots, $\nu$, of the cubic equation. This gives the empirical distribution of the treatment bias. Define $\beta^* = \tilde{\beta}-\nu$ and note that this is the bias-adjusted treatment effect. Use the empirical distribution of $\beta^*$ to define bounding sets for the `true' treatment effect. • If the box is partly contained in the NURR: Compute roots on the URR area as outlined above. Using grid points in URR that reside on the boundary of URR and NURR, select the boundary points of the grid that lie in NURR, which are within a `small' distance of the former points. Compute the three real roots on the boundary points of the grid that lie in NURR and select the root that is closest in absolute value to the unique real root in the closest grid point in URR. Iterate this process using the previously selected roots until all grid points in NURR are exhausted. Define $\beta^* = \tilde{\beta}-\nu$ using all the selected real roots and note that this is the bias-adjusted treatment effect. Use the empirical distribution of $\beta^*$ to define bounding sets for the `true' treatment effect. • If the box is wholly contained in NURR: Extend the dimension of the box until there is non-empty intersection with URR, and then repeat the steps outlined in the previous case.

The method outlined above will give bounding sets for the choice of the bounded box chosen by the researcher. It is important that a researcher draw on knowledge of the institutional details of the substantive issue under investigation in identifying the correct range for $\delta$ and $R_{max}$. For instance, if a particular research question has an omitted variable that is understood to be very important in explaining variation in the treatment variable, then it might be justified to use $\delta>1$. If, on the other hand, the researcher is sure that all important variables have been included in the model, and hence, that the omitted variable is relatively less important in explaining variation in the treatment variable, then a range of $\delta<1$ might be justified. Similar considerations should be used to infer a correct upper bound for $R_{max}$. The bounds generated for the BATE will only be as good as the choice of the bounded box chosen by the researcher.

table[table omitted — 1,395 chars of source]
sidewaystable[hbt!] \begin{center} \begin{threeparttable} \caption{Empirical Distributions of Omitted Variable Bias and the Bias-Adjusted Treatment Effect (BATE) related to the Effect of Maternal Behavior on Child Outcomes} \begin{tabular}{lccccc} \toprule & (1) & (2) & (3) & (4) & (5) \\ \cline{2-6} \\[-1.8ex] & 2.5% & 5% & 50% & 95% & 97.5% \\ \hline \\[-1.8ex] IQ, Breastfeed: Bias (Region 1) & $0$ & $0$ & $0.009$ & $0.034$ & $0.039$ \\ IQ, Breastfeed: BATE (Region 1) & $$-$0.021$ & $$-$0.017$ & $0.009$ & $0.017$ & $0.017$ \\ IQ, Breastfeed: Bias (Region 2) & $0$ & $0.004$ & $0.059$ & $0.119$ & $0.133$ \\ IQ, Breastfeed: BATE (Region 2) & $$-$0.116$ & $$-$0.102$ & $$-$0.041$ & $0.013$ & $0.017$ \\ & & & & &\\ IQ, Drink Preg: Bias (Region 1) & $0$ & $0.001$ & $0.036$ & $0.138$ & $0.155$ \\ IQ, Drink Preg: BATE (Region 1) & $$-$0.104$ & $$-$0.087$ & $0.014$ & $0.049$ & $0.050$ \\ IQ, Drink Preg: Bias (Region 2) & $0$ & $0.016$ & $0.221$ & $0.889$ & $1.158$ \\ IQ, Drink Preg: BATE (Region 2) & $$-$1.107$ & $$-$0.839$ & $$-$0.171$ & $0.034$ & $0.050$ \\ & & & & &\\ IQ, LBW + Preterm: Bias (Region 1) & $$-$0.071$ & $$-$0.063$ & $$-$0.017$ & $$-$0.001$ & $0$ \\ IQ, LBW + Preterm: BATE (Region 1) & $$-$0.125$ & $$-$0.124$ & $$-$0.108$ & $$-$0.062$ & $$-$0.054$ \\ IQ, LBW + Preterm: Bias (Region 2) & $$-$0.300$ & $$-$0.272$ & $$-$0.097$ & $$-$0.007$ & $0$ \\ IQ, LBW + Preterm: BATE (Region 2) & $$-$0.125$ & $$-$0.118$ & $$-$0.028$ & $0.147$ & $0.175$ \\ & & & & &\\ BW, Smoke Preg: Bias (Region 1) & $$-$151.390$ & $$-$120.190$ & $$-$18.521$ & $$-$0.421$ & $0$ \\ BW, Smoke Preg: BATE (Region 1) & $$-$172.510$ & $$-$172.090$ & $$-$153.990$ & $$-$52.321$ & $$-$21.121$ \\ BW, Smoke Preg: Bias (Region 2) & $$-$290.033$ & $$-$241.279$ & $217.838$ & $1,478.282$ & $1,897.682$ \\ BW, Smoke Preg: BATE (Region 2) & $$-$2,070.193$ & $$-$1,650.792$ & $$-$390.348$ & $68.769$ & $117.522$ \\ & & & & &\\ BW, Drink Preg: Bias (Region 1) & $$-$17.044$ & $$-$15.014$ & $$-$3.587$ & $$-$0.098$ & $0$ \\ BW, Drink Preg: BATE (Region 1) & $$-$14.149$ & $$-$14.051$ & $$-$10.562$ & $0.866$ & $2.895$ \\ BW, Drink Preg: Bias (Region 2) & $$-$152.269$ & $$-$142.776$ & $$-$24.933$ & $$-$1.510$ & $0$ \\ BW, Drink Preg: BATE (Region 2) & $$-$14.149$ & $$-$12.638$ & $10.784$ & $128.627$ & $138.120$ \\ \bottomrule \end{tabular} \begin{tablenotes} • Notes: This table reports the empirical distributions of bias and the bias-adjusted treatment effect (BATE) for the five models resported in Table (ref). Results about these models are discussed in oster-2019. Quantiles have been computed using the algorithm discussed in section (ref). \end{tablenotes} \end{threeparttable} \end{center}
table[table omitted — 1,186 chars of source]
sidewaystable[hbt!] \begin{center} \begin{threeparttable} \caption{Identified Set and $\delta^*$ Computed According to oster-2019} \begin{tabular}{lcccccc} \toprule & (1) & (2) & (3) & (4) & (5) & (6) \\ \cline{2-7} \\[-1.8ex] & $D$ & ID Set 1 & ID Set 2 & $\delta^*$ & Discont & Slope \\ \hline \\[-1.8ex] $R_{max}=0.61$ & & & & & & \\ IQ, Breastfeed & 20.32 & [0.017,0.375] & [-0.034,0.017] & 0.36 & FALSE & Negative \\ IQ, Drink Preg & 0.002 & [0.050,8.410] & [-0.147,0.050] & 0.26 & FALSE & Negative \\ IQ, LBW + Preterm & 0.0001 & [-79.874,-0.125] & [-0.125,-0.033] & 1.36 & FALSE & Negative \\ & & & & & & \\ $R_{max}=0.53$ & & & & & & \\ BW, Smoke Preg & 1905974 & [-3403.376,-172.511] & [-172.511,-49.713] & 1.16 & FALSE & Negative \\ BW, Drink Preg & 267359453 & [-3615.821,-14.149] & [-14.149,0.944] & 0.94 & FALSE & Negative \\ \bottomrule \end{tabular} \begin{tablenotes} • Notes: This table reports the identified set and $\delta^*$ computed according to the methodology in oster-2019. $D$ stands for the discriminant of the quadratic equation in ((ref)). ID Set 1 is $[\tilde{\beta}, \tilde{\beta}-\nu_1]$ and ID Set 2 is $[\tilde{\beta}, \tilde{\beta}-\nu_2]$, where $\nu_1$ and $\nu_2$ are the two roots of ((ref)), respectively. Proposition (ref) ensures that both these roots are real. $\delta^*$ is the value of $\delta$ that corresponds to $\beta=0$ and the relevant $R_{max}$ (which is $0.61$ for the first three rows and $0.53$ for the last two rows). $\delta^*$ has been computed with ((ref)). `Discont' is TRUE if $R^* \in [\tilde{R},1]$, and FALSE otherwise. `Slope' gives the slope the graph of the function $\delta=F(R_max)$ in ((ref)) on the domain $[\tilde{R},1]$; `Negative' denotes negative slope. For a discussion, see section (ref). \end{tablenotes} \end{threeparttable} \end{center}
figure[figure omitted — 595 chars of source]
figure[figure omitted — 437 chars of source]