EconBase
← Back to paper

Testing Identifying Assumptions in Parametric Separable Models: A Conditional Moment Inequality Approach

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

81,203 characters · 19 sections · 71 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Testing Identifying Assumptions in Parametric Separable Models: A Conditional Moment Inequality Approach

\onehalfspacing

abstractIn this paper, we propose a simple method for testing identifying assumptions in parametric separable models, namely treatment exogeneity, instrument validity, and/or homoskedasticity. We show that the testable implications can be written in the intersection bounds framework, which is easy to implement using the inference method proposed in chernozhukov2013intersection, and the Stata package of chernozhukov2015implementing. Monte Carlo simulations confirm that our test is consistent and controls size. We use our proposed method to test the validity of some commonly used instrumental variables, such as the average price in other markets in nevo2012identification, the Bartik instrument in card2009immigration, and the test rejects both instrumental variable models. When the identifying assumptions are rejected, we discuss solutions that allow researchers to identify some causal parameters of interest after relaxing functional form assumptions. We show that the IV model is nontestable if no functional form assumption is made on the outcome equation, when there exists a one-to-one mapping between the continuous treatment variable, the instrument, and the first-stage unobserved heterogeneity.

{ Keywords: Separable models, functional form, OLS, IV, testable implications, relaxed assumptions.

JEL subject classification: C14, C31, C35, C36.}

Introduction

Instrumental variable (IV) models are widely used in economics and related fields. In applied work, researchers often specify functional forms for the relationship between the outcome variable and the regressor of interest. The instrumental variable assumptions (exogeneity, exclusion restriction, and relevance) combined with functional form restrictions impose testable implications on the observed joint distribution of the outcome, the regressor, and the instrument.

Testing identifying assumptions is familiar in applied work. For example in the linear IV model, researchers usually test the relevance condition, which states that the covariance between the regressor and the instrument differs from zero. Accordingly, many papers in the literature propose solutions for how to deal with weak IVs (see for example staiger1997instrumental, wang1998inference, stock2002testing, hansen2014instrumental, andrews2019weak, etc.). Although there has been some attention devoted to testing the exogeneity and exclusion restriction assumptions, applied researchers do not often report a test result for these assumptions. A possible explanation could be that the existing approaches are difficult to implement in practice.

In this paper, we propose an easy-to-implement testing procedure for the IV assumptions in a broad class of parametric separable models. To do so, we transform the testable conditional moment equality implication of the model into two conditional moment inequalities, as this is more conducive to inference given recent developments in the moment inequality literature AS2013, chernozhukov2013intersection. The test can be implemented using existing Stata packages developed by chernozhukov2015implementing. Our approach relies on a simple two-step procedure. We first identify the model parameters using the standard IV methods, and we then plug the identified coefficients in the outcome equation to back out the error term. Afterwards, we check the exogeneity condition by testing nonparametrically the joint conditions that i) the supremum of the conditional expectation of the error term given the instrument values is less than or equal to zero; and ii) its infimum is bigger than or equal to zero.

The test is asymptotically consistent as the IV estimator converges at a parametric rate, and the approximation error that comes from replacing the true coefficients with the estimated ones vanishes when the sample size goes to infinity. We show through simulations how the test controls the size asymptotically (sample size $\gtrapprox$ 3000) and the power converges to one for large samples (1000 or bigger). We illustrate the proposed test on two real-world empirical examples: the average price in other markets IV nevo2012identification, and the Bartik IV card2009immigration. The test rejects the validity of both IVs. Finally, we discuss alternative solutions that are available to researchers when a parametric IV model is rejected. A rejection of the IV model could be due to a violation of the IV assumptions or functional form misspecification. nevo2012identification, Conley2012, Masten2021, among others have proposed alternative identification results when the IV is invalid. In this paper, we propose nonparametric identification when the IV is valid, but functional forms may be misspecified. Our approach is similar to imbens2009identification. We relax all functional form assumptions for the outcome equation but assume that the relationship between the regressor and first-stage unobserved heterogeneity is continuous and strictly monotonic for all values of the instrument. We then point-identify the marginal response function for each value of the regressor and thus various marginal treatment effects.

The testability of the instrumental variable model assumptions has been questioned over the years. Some researchers think the exogeneity and exclusion restriction assumptions are not testable because they relate the IV to an unobserved variable (potential outcomes or the error term in a linear model). However, a lot of progress has been made in recent years to elucidate when these IV assumptions can be testable or untestable. pearl1995testability appears to be the first paper that derives a testable implication for the IV independence (exogeneity) and exclusion restriction assumptions when the endogenous regressor is discrete. kedagni2020generalized added a new set of testable implications to those in pearl1995testability and showed that their testable implications are sharp when the outcome and treatment are binary but the instrument is unrestricted. gunsilius2021nontestability proved that the instrumental variable independence assumption is not testable when the endogenous treatment is continuously distributed and there are no structural assumptions. However, most of the models that researchers consider in applied work impose some additional structures that help identify the parameters of interest. For example, in the local average treatment effect model, the IV independence assumption is often coupled with the monotonicity assumption, which together yield testable implications as discussed in kitagawa2015test, mourifie2017testing, and Huber2015TestingConstraints. See also acerenza2023testing in the context of bivariate probit models, Arai2022TestingDesigns, and Hsu2023TestingDesigns in the regression discontinuity design framework. All the above mentioned papers focus on nonseparable models. In this paper, we focus on testing the IV assumptions in parametric separable models.

In the existing literature on testing identifying assumptions in parametric IV regression models, a conventional approach is to transform conditional moment restrictions to unconditional ones using weighting functions of the instrumental variable and then testing unconditional moment conditions (see the discussions in bierens1982consistent, bierens1990consistent, newey1985maximum, and related literature). Researchers should be cautious about choosing instrumental functions with this method since the tests could be inconsistent with most alternatives if improper instrumental functions are selected. Another standard method is to estimate the conditional moments with smoothing techniques, such as the kernel smoothed method (for example, hardle1993comparing, horowitz2001adaptive and related literature) and the smoothed empirical likelihood ratio method (see tripathi2003testing and kitamura2004empirical).

Despite available methods, most applied papers do not implement them to test identifying assumptions in regression models. When using instrumental variable methods, most applied papers focus on justifying the validity of their instrument based on some intuition. This motivates us to propose an easy-to-implement method that allows researchers to check whether or not their model assumptions are compatible with the data, and propose a relaxed IV model that still offers identification if a test rejection is thought to arise from functional form misspecification rather than a failure of instrument validity.

The remainder of the paper is organized as follows. Section (ref) presents the testable implications and the testing procedure in linear and nonlinear separable IV models. In Section (ref), we illustrate the power and the size of the proposed test through some Monte Carlo simulations. Section (ref) discusses results that relax the functional form assumption for the outcome equation. Section (ref) presents the empirical illustrations and Section (ref) concludes.

Testable Implications in Parametric Separable Models

In this section, we describe our method of testing the identifying assumptions, including exogeneity, instrument validity, and homoskedasticity assumptions, in some commonly used structural models. Our method is mainly based on the intersection bounds framework. We first identify the model parameters under the imposed identifying assumptions and then write the unobserved error term as a known function of observed variables and identified parameters. We derive the testable implications in a set of conditional moment inequalities, so that we can use the conditional moment inequality method developed in chernozhukov2013intersection.

Simple Linear IV Model

To give the general idea of our test, we begin with a simple example. Consider the basic linear model

equation[equation omitted — 68 chars of source]

where $Y \in \mathcal{Y}$ is an observed outcome variable with $\mathbb{E}|Y| < \infty$, $X \in \mathcal{X}$ is a scalar observed potentially endogeneous covariate, and $U$ is an unobserved error term. Equation (ref) could, for example, represent a treatment effect model in which the effect $\beta_1$ of increasing $X$ by one unit is common to all individuals and constant across levels of $X$.

Suppose that researchers propose an instrument variable $Z \in \mathcal{Z}$ to identify the parameters $(\beta_0, \beta_1)$, with $Z$ satisfying the following assumptions:

assumption[Exogeneity] $\mathbb{E}[U \mid Z]=0$ almost surely.
assumption[Relevance] $\operatorname{Cov}(X, Z) \neq 0$.

Note that a third assumption implicit in interpreting Equation (ref) causally is the so-called exclusion restriction that $Z$ does not have a direct effect on $Y$, only affecting it through $X$.\\

Under the model structure Equation ((ref)) and Assumptions (ref) - (ref), one can point identify $(\beta_0,\beta_1)$ as

equation*[equation* omitted — 138 chars of source]

Importantly, we can then express $U$ as a function of the data by substituting in the identified parameters, i.e., $U=Y - \mathbb{E}[Y] - (X - \mathbb{E}[X]) \operatorname{Cov}(Y, Z) / \operatorname{Cov}(X, Z)$. Combining this with Assumption (ref), we obtain the testable implication that

equation[equation omitted — 193 chars of source]

for all $z \in \mathcal{Z}$.

The question we may ask at this point is whether the implication (ref) can be rejected by the data. If $Z$ is binary, then implication (ref) holds automatically and cannot be rejected.\footnote{This can be shown directly by noting that when $Z$ is binary $\frac{\operatorname{Cov}(Y, Z)}{\operatorname{Cov}(X, Z)} = \frac{\mathbb{E}[Y|Z=1]-\mathbb{E}[Y|Z=0]}{\mathbb{E}[X|Z=1]-\mathbb{E}[X|Z=0]}$, and applying the law of iterated expectations over $Z$ for $\mathbb{E}[Y]$ and $\mathbb{E}[X]$.}

By contrast, if there are at least three elements in the support of $Z$, the condition (ref) may fail to hold. To see this, suppose that the support of $Z$ is equal to $\{0,1,2\}$. We can use two conditions (e.g., $\mathbb E[U\vert Z=0]=0$ and $\mathbb E[U\vert Z=1]=0$) to identify $\beta_0$ and $\beta_1$, and use a third condition (e.g., $\mathbb E[U\vert Z=2]=0$) to check whether the model is rejected or not. Intuitively, testability arises from the linearity assumption inherent in (ref), which only provides over-identification given information provided by more than two moments.

The above discussion leads to the following proposition.

propositionConsider the model specification (ref) together with Assumptions (ref) - (ref). Then, the testable implication (ref) holds and is sharp. On the other hand, whenever condition (ref) holds, there exist a random vector $(X^*,Y^*,U^*,Z)$ and a parameter vector $(\beta_0^*,\beta^*_1)$ s.t. $Y^*=\beta^*_0+\beta^*_1 X^* + U^*$, $(X^*,Y^*,Z)$ has the same distribution as $(X,Y,Z)$, and $E[U^*\vert Z]\neq 0$.
proofSee Appendix (ref).
remarkProposition (ref) shows that the standard linear IV model can generally be tested using the implication in Equation (ref). However, this is not the same as saying that the model is confirmed when the implication holds.

As more assumptions are added to the model, it may be possible to extend its testable impications. For example, one could add more structure to the above IV model by assuming homoskedasticity.

assumption[Homoskedasticity] $\mathbb{E}\left[U^2 \mid Z\right]=\sigma_0^2 < \infty$ almost surely.

Under Assumption (ref), we identify $\sigma^2_0$ as $\mathbb{E}\left[\left((Y - \mathbb{E}[Y]) - (X - \mathbb{E}[X]) \operatorname{Cov}(Y, Z) / \operatorname{Cov}(X, Z) \right)^2\right]$. Then, a testable implication for model Equation ((ref)) and Assumptions (ref) - (ref) is

equation[equation omitted — 556 chars of source]

for all $z \in \mathcal{Z}$.\\

remarkAll the results derived in this section also hold if $Z$ is replaced by $X$, which implies that the simple linear regression model (ref) combined with the exogeneity $(\mathbb E[U\vert X]=0)$, relevance $(Var(X)>0)$ and/or homoskedasticity $(\mathbb E[U^2\vert X]=\sigma^2_0)$ are also testable.

Before extending the basic idea to more general classes of models, we in the next section propose how to implement a practical test of the moment equalities (ref) and (ref) by re-expressing them as a set of moment inequalities.

Testing Procedure

We propose an easy-to-implement approach based on conditional moment inequalities using the intersection bound framework, which does not rely on a choice of instrumental functions to transform them into unconditional moments. Indeed, the implication (ref) is equivalent to the following.

equation[equation omitted — 430 chars of source]

The representation of the testable implication in Equation (ref) is now easier to test thanks to the recent advances on testing conditional moment inequalities. More precisely, chernozhukov2013intersection propose a method that allows us to test a set of conditional moment inequalities in the form of Equation (ref), which are often called intersection bounds.

To test inequalities in condition (ref), we will first plug in the $\sqrt{n}$-consistent estimators $(\hat{\beta}_0,\hat{\beta}_1)$ for $(\beta_0,\beta_1)$. To simplify notation, we define $W_1 = Y - \hat{\beta}_0-\hat{\beta}_1 X$ and $W_2 = - Y + \hat{\beta}_0+\hat{\beta}_1 X$, and a set $\mathcal{V}=\{(z, j): z \in \mathcal{Z}, j \in\{1,2\}\}$. Then, we denote $\theta(v) \equiv \mathbb{E}\left[W_j \mid Z=z\right]$ for $v=(z, j)$. We focus on testing the hypothesis

equation[equation omitted — 160 chars of source]

Second, to test the null hypothesis $H_0$ against the alternative $H_1$ in Equation (ref), we need to estimate the supremum statistic $\theta_0$, which requires estimating the conditional moment $\theta(v)$ for each $v$. To eliminate the first step plug-in estimation bias asymptotically, we recommend estimating the conditional moment $\theta(v)$ nonparametrically. Letting $\hat{\theta}(v)$ denote an estimator for $\theta(v)$, we show in Appendix (ref) that all asymptotic properties are preserved if we then plug in a $\sqrt{n}$-consistent estimator $(\hat{\beta}_0, \hat{\beta}_1)$ for $(\beta_0, \beta_1)$ in the first step. If we simply take supremum over $\hat{\theta}(v)$ and use it as an estimator for $\theta(v)$, we would have a finite-sample bias because of estimation errors. To correct this bias, chernozhukov2013intersection propose a precision-corrected estimator\footnote{See Appendix (ref) for how to choose a precision-corrected estimator.} for $\theta_0$,

equation[equation omitted — 153 chars of source]

The precision correction term is $k_{1-\alpha} \hat{s}(v)$. Here, $\hat{s}(v)$ is the standard error of $\hat{\theta}(v)$, $k_{1-\alpha}$ is $(1-\alpha)$-quantile of an approximated distribution of $\sup _{v \in \mathcal{V}} \{(\hat{\theta}(v)-\theta(v))/ \sigma(v)\}$, $\sigma(v)$ is the standard deviation of $\hat{\theta}(v)$. Heuristically speaking, the precision correction term adjusts the estimator with estimation error uniformly on $\mathcal{V}$.

The estimator can be used as a test statistic for testing the hypothesis ((ref)). We reject the null hypothesis at the significance level $\alpha$ if $\hat{\theta}_{1-\alpha} > 0$. We use the clrtest/clrbound Stata commands from chernozhukov2015implementing to test the hypothesis in ((ref)).

Extension to General Parametric Separable Models

Linear models with multiple regressors

Consider a linear IV model with multiple (potentially endogenous) regressors where $X$ is a vector instead of a scalar:

equation[equation omitted — 50 chars of source]

Let $Z=(Z_1, Z_2, \ldots, Z_k)'$ be a candidate IV for $X$. To identify and estimate the parameters $\beta$ in Equation ((ref)), researchers usually impose the following assumptions.

assumption(IV conditions) \begin{enumerate} • (Exogeneity assumption) $\mathbb{E}[U \mid Z]=0$ almost surely. • (Rank condition) The matrix $\mathbb E[ZX']$ is invertible. \end{enumerate}

Under the IV conditions in Assumption (ref),

equation*[equation* omitted — 61 chars of source]

As a result, $U=Y-X'\beta=Y-X'\mathbb E[ZX']^{-1}\mathbb E[ZY]$ where each term in $U$ is observable or identified from the data. Under the exogeneity condition in Assumption (ref),

equation[equation omitted — 157 chars of source]

which is equivalent to a set of joint conditional moment inequalities in Equation ((ref)) and consistent with the framework of intersection bounds.

proposition[Testable implications of linear IV models] Consider the linear model in Eq. (ref). Then, testable implications of Assumption (ref) are given by \begin{equation} \begin{aligned} & \sup_{z\in\mathcal{Z}} \mathbb E \left[Y- X' \beta^* \mid Z=z\right] \leq 0,\\ & \sup_{z\in\mathcal{Z}} \mathbb E \left[-Y+X' \beta^* \mid Z=z\right] \leq 0, \end{aligned} \end{equation} where $\beta^*=\mathbb E [ZX']^{-1} \mathbb E [ZY]$ is identified under Assumption (ref).

The testable implications in Equation (ref) are sharp, and the proof is similar to that of Proposition (ref) and is therefore omitted.

We can apply the same methods to test other identifying assumptions that are commonly used in the literature on linear IV models, for example, the homoskedasticity assumption $\mathbb E[U^2 \mid Z]=\sigma^2$. The testable implications for the linear IV model in Equation (ref) and Assumption (ref) coupled with the homoskedasticity assumption are

equation[equation omitted — 505 chars of source]

Nonlinear parametric separable models

Consider a more general model where $Y = m(X, \theta_0) + U$, where $\mathbb{E}\left(U \mid Z\right) = 0$, $m(\cdot, \theta)$ defined on $\mathbb{R}^{k} \times \Theta$ is a known function, and $\Theta$ is a parameter space. This model can generally be tested without assuming that there is a unique parameter $\theta_0$ such that $\mathbb E[Y-m(X,\theta_0) \mid Z]=0$. Instead, we can characterize the identified set for $\theta_0$ as:

eqnarray*[eqnarray* omitted — 212 chars of source]

The above identified set can be computed through a grid search using the intersection bounds framework by chernozhukov2013intersection. As a byproduct, if the identified set $\Theta_I$ is empty, then the model is rejected.

One can add other commonly used assumptions such as monotonicity of $m(\cdot,\theta)$ in $\theta$ for all values of $X$ (to ensure uniqueness of $\theta_0$), homoskedasticity, etc. to tighten the identified set.

example[Box-Cox regression model] For any $x > 0$, define \begin{equation*} x^{(\lambda)} \equiv \begin{cases}\frac{x^\lambda-1}{\lambda}, & if \lambda \neq 0 \\ \log (x), & if \lambda=0\end{cases} \end{equation*} and the Box-Cox model is specified as \begin{equation} Y=\beta_0+\beta_1 X^{(\lambda)}+U. \end{equation} In the Box-Cox model, the nonlinear function $m(\cdot)$ takes the form $m(X, \theta)=\beta_0+\beta_1 X^{(\lambda)}$, and the parameter $\mathbb{\theta} = \left(\beta_0, \beta_1, \lambda\right)$.
example[CES Function] Another typical example discussed in hansen2022econometrics is the constant elasticity of substitution (CES) function, which is a generalization of the Cobb-Douglas production function. Researchers often assume the CES structure to estimate production functions. The CES function with two inputs, $X_1, X_2$, is \begin{equation*} Y=\left\{\begin{array}{cc} A\left(\alpha X_1^\lambda+(1-\alpha) X_2^\lambda\right)^{\nu / \lambda}, & if \lambda \neq 0 \\ A\left(X_1^\alpha X_2^{(1-\alpha)}\right)^\nu, & if \lambda=0 \end{array}\right. \end{equation*} We can take the logarithm of Y and set $\log A = \beta_0 + U$, where $U$ represents the unobserved productivity. Then the CES function implies the following nonlinear model, \begin{equation} \log Y=\left\{\begin{array}{cl} \beta_0+\frac{\nu}{\lambda} \log \left(\alpha X_1^\lambda+(1-\alpha) X_2^\lambda\right)+U, & if \lambda \neq 0 \\ \beta_0 + \nu \left(\alpha \log X_1 + (1-\alpha) \log X_2\right) + U, & if \lambda=0, \end{array}\right. \end{equation} where the parameter $\mathbb{\theta} = (\lambda, \nu, \alpha, \beta_0)$, and the function $m(X, \mathbb{\theta})$ is nonlinear in $\mathbb{\theta}$.

Nonlinear parametric nonseparable models

Although we focus on parametric separable models in this paper, we explain in Example (ref) below how our proposed approach could be extended to some nonseparable models such as probit models.

example[Probit with endogeneity] Consider the following parametric model \begin{eqnarray*} \left\{\begin{array}{clc} Y &=& \mathbbm{1}\{X\beta -U \geq 0\}, \\ X &=& Z \delta + V, \end{array}\right. \end{eqnarray*} where $(U,V)'$ is jointly normally distributed with coefficient of correlation $\rho$, and $Z \perp \!\!\! \perp (U,V)$. We can write $U=\rho V + \varepsilon$, where $\varepsilon \perp \!\!\! \perp (V,Z),$ and $\varepsilon \sim N(0,1-\rho^2)$. The coefficients $(\beta, \delta, \rho)$ are identified and can be estimated through the maximum likelihood method. We can write \begin{eqnarray*} \mathbb E[Y|X=x, Z=z] &=& \mathbb P(U \leq X \beta | X=x, Z=z),\\ &=& \mathbb P(\rho V +\varepsilon \leq X \beta | X=x, Z=z),\\ &=& \mathbb P(\varepsilon \leq x \beta - \rho(x-z \delta) | X=x, Z=z),\\ &=& \Phi\left(x\frac{\beta-\rho}{\sqrt{1-\rho^2}}+z \frac{\rho \delta}{\sqrt{1-\rho^2}}\right), \end{eqnarray*} which implies $ \mathbb E\left[Y-\Phi\left(X(\beta-\rho) /\sqrt{1-\rho^2 }+Z \rho \delta / \sqrt{1-\rho^2}\right) | X=x, Z=z\right]=0$ for all $(x,z)$. Since the parameters $(\beta, \delta, \rho)$ are identified, this latter implication is testable.

This example above can be generalized in the following way. Suppose the observed vector $(Y,X,Z)$ satisfies

eqnarray*[eqnarray* omitted — 127 chars of source]

where $g$ and $h$ are known functions, $U=\rho V+\varepsilon,$ $\varepsilon \perp \!\!\! \perp V$, and $\varepsilon | X,Z \sim F_{\varepsilon}$ (known). This model is parametric and is separable in the “first-stage” equation for $X$, even though it is nonseparable in the outcome equation.

Suppose also that $(\beta, \delta, \rho)$ are identified. Then, we have a generalized version of the testable implication $$\mathbb E\left[Y-\int g(X,\beta, e+\rho(X-h(Z,\delta))d F_{\varepsilon}(e) | X=x,Z=z\right]=0,$$ which can be converted into conditional moment inequalities

eqnarray*[eqnarray* omitted — 243 chars of source]

Semiparametric linear single-index models

Our proposed approach can also be extended to binary response models that make no distributional assumptions on the error term, but restrict the dependence on covariates to take a parametric form such as a linear single-index model:

example[Binary response with linear index] Consider the model $$Y = \mathbbm{1}\{X'\beta -U > 0\},$$

where $X \perp \!\!\! \perp U$ but the distribution $G(u)$ of $U$ is left unrestricted. As pointed out by ichimura1993, identification of $\beta$ in this model is possible up to a normalization provided some regularity conditions hold, for example, that $G$ is differentiable and at least one regressor is continuously distributed.

Note that this model implies that $$\mathbb{E}[Y|X=x] = \mathbb{P}(U < x'\beta) = G(x'\beta),$$ which carries the testable restriction that for any $x_1$ and $x_2$ such that $x_1'\beta=x_2'\beta$, $\mathbb{E}[Y|X=x_1]=\mathbb{E}[Y|X=x_2]$. If the support of $X$ is a known subset $\mathcal{X} \in \mathbb{R}^k$, we can therefore test this model through the equality restriction that

equation[equation omitted — 175 chars of source]

We know that when $b=\beta^*$ the inner supremum is equal to zero, where $\beta^*$ is the true value of $\beta$. To bring the above into the intersection bounds framework, make use of the fact that $\beta^*$ is identified, and rewrite (ref) as $$\sup_{(x_1,x_2) \in \mathcal{V}(\beta^*)} \left\{\mathbb{E}[Y|X=x_1]-\mathbb{E}[Y|X=x_2]\right\} \le 0,$$ $$\sup_{(x_1,x_2) \in \mathcal{V}(\beta^*)} \left\{\mathbb{E}[Y|X=x_2]-\mathbb{E}[Y|X=x_1]\right\} \le 0,$$ i.e., that $\mathbb{E}[Y|X=x_1]-\mathbb{E}[Y|X=x_2]=0$ for all $(x_1,x_2) \in \mathcal{V}(b)$, where we define $\mathcal{V}(b) := \{(x_1,x_2) \in \mathcal{X} \times \mathcal{X}: x_1'b = x_2'b\}$. The results of chernozhukov2013intersection require the set over which the supremum or infimum is taken across be compact, so we must maintain this assumption here. A sufficient condition is that $\mathcal{X}$ be compact in $\mathbb{R}^k$.

We illustrate how sample splitting can help transform the inequalities above into conditional moment inequalities. Suppose we have an i.i.d. sample $\{Y_i,X_i\}_{i=1}^{n}$, and we randomly split the sample into two sub-samples $\{Y_i^{(1)},X_i^{(1)}\}_{i=1}^{n_1}$ and $\{Y_i^{(2)},X_i^{(2)}\}_{i=n_1+1}^{n}$. Then the inequalities above are equivalent to the following: $$\sup_{(x_1,x_2) \in \mathcal{V}(\beta^*)} \left\{\mathbb{E}[Y^{(1)}-Y^{(2)}|X^{(1)}=x_1, X^{(2)}=x_2]\right\} \le 0,$$ $$\sup_{(x_1,x_2) \in \mathcal{V}(\beta^*)} \left\{\mathbb{E}[Y^{(2)}-Y^{(1)}|X^{(1)}=x_1, X^{(2)}=x_2]\right\} \le 0.$$

Monte Carlo Simulations

Size of the test

First, we generate a linear model with an instrument variable that satisfies Assumption (ref). The model takes the form

equation[equation omitted — 186 chars of source]

where $\beta_0 = 0$, $\beta_1 = 2$, $\gamma_0 = 0$ and $\gamma_1 = 3$. We have done $500$ Monte Carlo replications. For each replication, we set the sample size to be $n$. We randomly draw $X_i \stackrel{\text { i.i.d. }}{\sim} \mathcal{U}[-3, 3]$, $(U_i, V_i)$ i.i.d. from $\mathcal{N}(0, \Sigma)$ with $\Sigma=(1, 0.5; 0.5, 2)$, and draw $Z_i \stackrel{\text { i.i.d. }}{\sim} \mathcal{U}[-3, 3]$ independent with $(U_i, V_i)$, $i = 1, \dots, n$. Table (ref) presents the simulation results of testing the above DGP under Assumption (ref).

table[table omitted — 675 chars of source]

To implement the test, we first estimate $\hat{\beta}_0$ and $\hat{\beta}_1$ under Assumptions (ref) and (ref) and obtain $\hat{U}_i = Y_i - \hat{\beta}_0 - {\beta}_1 X_i$. Then we use the Stata package of chernozhukov2015implementing to compute the rejection rate of the testable implications in Equation (ref). To estimate the supremum test statistic, we need to select grid points over the support of $Z_i$. We take $100$ grid points uniformly over the 1 centile to 99 centiles of the support of $Z_i$. Also, we select the series estimation to estimate conditional means $\mathbb{E}[\hat{U}_i \mid Z_i]$.

From the results in Table (ref), we can observe that the rejection rates are higher than nominal sizes when the sample sizes are small. The reason is that we use a nonparametric method to estimate the conditional expectations, which have large estimation errors when the sample sizes are small. As the sample size increases, the rejection rates drop at each significance level and finally are close to nominal sizes, which is consistent with our expectations. When the sample size increases to 3000, rejection rates fall below the nominal sizes at significance levels of $10\%$ and $5\%$. The low rejection rates suggest that our tests might be conservative, which is caused by the precision correction term when constructing the test statistics.

We also come up with the null hypothesis specifications under the exogenous assumption of the regressor $X$ and the homoskedasticity assumption. The results are presented in Appendix (ref), and we can observe a similar pattern for rejection rates as in Table (ref).

Next, we consider a nonlinear model that satisfies the IV assumption. We generate a DGP in the form of the Box-Cox model described in Equation (ref).

equation[equation omitted — 211 chars of source]

where $\beta_0 = 0$, $\beta_1 = 2$, $\gamma_0 = 0$, $\gamma_1 = 2$, and

equation[equation omitted — 166 chars of source]

for $\lambda = 0, -1, 1$. We randomly draw $(U_i, V_i)$ i.i.d. from $\mathcal{N}(0, \Sigma)$ with $\Sigma=(1, 0.5; 0.5, 2)$, and $V_{+i}$ denotes the positive part of $V_i$. We also draw $Z_i \stackrel{\text { i.i.d. }}{\sim} \mathcal{U}(0, 10]$ independent with $(U_i, V_i)$. The simulation results are presented in Table (ref).

table[table omitted — 1,270 chars of source]

When the sample size is $200$, the rejection rates at each significance level are larger than the nominal size for each value of $\lambda$. However, when the sample size increases, the rejection rates at each significance level decrease and are finally below the nominal size. Again, this shows that our tests are conservative, and researchers need to pay attention to this in practice. Note that the rejection rates are higher for $\lambda=-1$ compared to $\lambda=0,1$, and they do not fall below the nominal size for $\lambda=-1$ when $n=3000$ as they do for $\lambda=0,1$.

Power of the test

To investigate the power properties of our testing procedure, we consider the similar alternatives proposed in horowitz2001adaptive and tripathi2003testing. We first consider a specification of a linear model that violates the IV exogeneity assumption (Assumption (ref)). In the model specification (ref), we generate $Z_i$, $i = 1, \ldots, 1000$, from $Z_i \stackrel{\text { i.i.d. }}{\sim} \mathcal U{[-3,3]}$. Then, we generate $U_i = L / \sigma \cdot \phi(Z_i / \sigma) + \tilde{U}_i$, where $\tilde{U}_i = \min\{\max\{-3, \tilde{V}_i\}, 3\}$, $(\tilde{V}_i, V_i)$ is drawn independently of $Z_i$ and i.i.d. from $\mathcal{N}(0, \Sigma)$ with $\Sigma=(1, 0.5; 0.5, 1)$.\footnote{Note that $\tilde{U}_i$ has mean zero, as $\tilde{U}_i$ is a mean zero normal distribution truncated on $[-3,3]$ so that its distribution is symmetric around 0.} Under the constructed DGP, Assumption (ref) fails. The conditional mean function $\mathbb{E}[U_i \mid Z_i = z] = L / \sigma \cdot \phi(z / \sigma)$ is smooth in $z$ and has a unique maximizer. Under this constructed DGP, $L$ and $\sigma$ are two constants that determine the shape of $\mathbb{E}[U_i \mid Z_i = z]$. $L$ measures the average deviation of the conditional mean from zero, where a larger $L$ indicates a greater average deviation over the support. Also, $\sigma$ determines the shape of this function: a smaller $\sigma$ implies that the function $\mathbb{E}[U_i \mid Z_i =z]$ is more peaked around the maximizer. We conduct $500$ replications for each value of $L$ and $\sigma$. Table (ref) presents the rejection rates.

table[table omitted — 1,278 chars of source]

In Table (ref), we can observe that when $L$ increases, which suggests that the deviation of the conditional mean $\mathbb E[U_i \mid Z_i=z]$ from zero is more significant, the rejection rates of our test increase at each significance level. This is consistent with our conjectures that our test is more powerful when deviations of the null are more significant. In addition, we can observe that given the value of $L$, the rejection rates increase when $\sigma$ becomes smaller. This relates to the precision-correction term when constructing the supremum test statistics. As we mentioned in Section (ref), if the conditional moment function is more peaked, the method of chernozhukov2013intersection can correct errors of estimating supremum test statistic more accurately. Therefore, when $\sigma$ becomes smaller, the conditional mean function is more peaked, and our test with the supremum test statistic can achieve greater power. We also get the rejection rates for specifications that violate the exogenous assumption on $X$. The simulation results are in Appendix (ref) and display a similar pattern as in Table (ref).

To compare our testing method with the common overidentification test, we consider the same DGP with $L = 0.5, \sigma = 0.25$, and we convert the conditional moment restriction in Assumption (ref) to an unconditional moment condition $\mathbb{E}[h(Z_i) U_i] = 0$. We choose $h(Z) = (Z, Z^2, Z^3)'$ and estimate the parameters $(\beta_0, \beta_1)$ with two-stage least square method. Then, we apply Sargan's test to check the instrument mean independence validity of $Z$. We conduct $500$ simulations with different sample sizes and plot the rejection rates of those two tests in Figure (ref).

figure[figure omitted — 223 chars of source]

This Figure shows that if the instrument mean independence assumption is violated in the linear model, the rejection rates of both tests increase with the sample size and approach 1 as the sample size grows, which suggests that both tests are consistent. Also, our proposed conditional moment inequality test gives a higher rejection rate than Sargan's test for any sample size. Therefore, our test achieves better performance in power, and we suggest that researchers try our testing procedure to test the instrument validity even if they use the overidentification test and cannot reject the IV assumption. We also compare the power of our test with Sargan's test for other values of $L$ and $\sigma$, and the results are presented in the Appendix (ref) and show that our test can achieve higher powers than Sargan's test in most cases.

Solutions when the parametric IV model is rejected

Several reasons may lead to the rejection of the identifying assumptions in the IV model. The rejection could be due to the violation of the functional form assumptions (misspecification, heterogeneous effects, etc.), the invalidity of the instrumental variable (exogeneity and/or exclusion restriction), etc. The literature has offered many solutions. nevo2012identification develop an identification result in the parametric model when the correlation between the instrument and the error term has the same sign as the correlation between the regressor and the error term. They relax the IV exogeneity assumption while maintaining the exclusion restriction. Conley2012 present methods for performing inference while allowing for violations of the exclusion restriction but maintain exogeneity of the IV. In this framework, Masten2021 recommend reporting the set of parameters that are consistent with minimally nonfalsified models, which they refer to as the falsification adaptive set. So, when the IV model is rejected, we recommend that the researcher explores the possible relaxations discussed in the above mentioned papers if s/he believes in the functional form assumption. Below, we propose a way to relax the functional form while maintaining the exogeneity and exclusion restriction assumptions.

Relaxing Parametric Assumptions

Suppose we observe a random outcome variable $Y$, and a continuous random variable $X$ on $\mathbb{R}$, which is a potentially endogenous treatment variable. We also observe a random vector $Z$ on $\mathbb{R}^k$, $k \geq 1$, which is considered as an instrumental variable.

Instead of parameterizing the equation of $Y$ on $X$, we consider the following model

equation[equation omitted — 138 chars of source]

where $(U, V)$ represent the unobserved heterogeneity. The first equation in Equation (ref) is an outcome equation. We do not impose any functional form assumptions on function $g(\cdot, \cdot)$, and the unobservable $U$ can be multi-dimensional. We use $Y_x$ to denote the potential outcome when $X = x$, i.e. $Y_x = g(x, U)$. We use calligraphic characters to denote the supports of random variables. For example, $\mathcal{X}$ indicates the support of $X$. The second equation is a treatment selection equation for $X$, which takes as arguments the instrument $Z$ and a scalar unobservable $V$. The function $h$ may have a structural interpretation, with potential treatments $X_z = h(z, V)$. However, we can also consider $h$ as a reduced form representation of the conditional distribution of $X$, i.e. define $V=F_{X|Z}(X)$ to be the rank of a unit with treatment value $X$ in the conditional distribution among units sharing their value of $Z$. Then the equation $X=h(Z,V)$ holds with probability one if we define $h$ to yield the conditional quantile function of $X$ given the instrument: $h(z,v)=Q_{X|Z=z}(v)$.\footnote{We follow the standard definition of the quantile funciton, where for any random variable $A$ and $u \in [0,1]$ we define $Q_{A}(u) = \inf \{a: F_{A}(a) \ge u\}$.}

Given $p \in \mathcal{V}$, and $x, x' \in \mathcal{X}$, we define the marginal treatment effect from $x'$ to $x$ at a given value $V=p$ as $\operatorname{MTE}\left(p ; x, x^{\prime}\right) \equiv \mathbb{E}\left[Y_x-Y_{x^{\prime}} \mid V=p\right]$.\footnote{This parameter is defined similarly to the marginal treatment effect with binary treatment, for example, see heckman2005structural.} Also, given $x \in \mathcal{X}$, we define the average structural function as $\operatorname{ASF}(x) \equiv \mathbb{E}[Y_x]$. We can use these two parameters to identify other common treatment effects, such as average treatment effects or policy-relevant treatment effects. To identify the parameters of interest, we impose Assumptions (ref) - (ref).

assumption[Independence] $Z \perp \!\!\! \perp (U,V)$.
assumption[Strict monotonicity in V] For any $z \in \mathcal{Z}$, where $\mathcal{Z}$ is the support of $Z$, the function $h(z, v)$ is continuous and strictly monotonic in $v$.
assumptionThe CDF of $V$ is absolutely continuous and strictly increasing.

Assumption (ref) requires that the instrument variable $Z$ is independent of the unobservables $U$ and $V$. Assumption (ref) implies that $X$ is one-to-one mapped to $V$ given $z \in \mathcal{Z}$. Thus, given $Z$, we can invert the function $X = h(Z, V)$ with respect to $V$ and have $V = h^{-1}_{Z}(X)$ almost surely. Assumption (ref) is common in the existing literature. For example, when researchers assume that function $h(Z, V)$ is additively separable in $V$, say $X = f(Z) + V$, Assumption (ref) then holds straightforwardly if $f$ is continuous. Assumption (ref) assures that we can normalize $V \sim \mathcal{U}[0, 1]$, where $\mathcal{U}$ denotes the uniform distribution.\footnote{Note that when $h$ is simply defined to be the conditional quantile function of $h(z,v)=Q_{X|Z=z}(v)$, Assumption (ref) amounts to supposing that the conditional distribution of $X$ has no mass points and has a positive density everywhere on its support. This guarantees that $F_{X|Z}(x)$ is the inverse function of $Q_{X|Z}(v)$ and that $V|Z \sim V \sim \mathcal{U}[0, 1]$. Note that this implies $Z \perp \!\!\! \perp V$ by definition; however Assumption (ref) remains a substantive assumption because it takes $U$ and $V$ to be jointly independent of $Z$.}

For any $x \in \mathcal{X}$ and $z \in \mathcal{Z}$, similar to the setting with a binary treatment, we define propensity score function as

equation[equation omitted — 86 chars of source]

Under Assumptions (ref) - (ref), we show that the error term $V$ can be identified as the propensity score function on its support.

lemma[Identification of $h^{-1}_{Z}$] Consider the model defined in Equation (ref). If Assumptions (ref)-(ref) hold for $x \in \mathcal{X}$ and $z \in \mathcal{Z}$, the following equality holds: \begin{equation} P(z, x) = h^{-1}_{z}(x). \end{equation}
proofSee Appendix (ref).

Note that as a consequence of (ref) and Assumptions (ref) and (ref), we have that

equation[equation omitted — 145 chars of source]

We can identify the function $P(Z, X)$ from the population since $X$ and $Z$ are observed random variables, so that $h_{Z}^{-1}(X)$ is identifiable. Also, Assumption (ref) implies that $V = h^{-1}_{Z}(X)$ almost surely. Therefore, an important implication of Lemma (ref) is that we can identify the unobservable $V$ as $V=P(Z, X)$ on the support of $P(Z, X)$. We let $\mathcal{P}$ denote the support of propensity score function $P(Z, X)$ and $\mathcal{P}_x$ denote the support of $P(Z, X)$ given $X = x$ hereafter.

The next lemma proves that $P(Z, X)$ can serve as a control function.

lemma[Control function] Consider the model defined in Equation (ref). Under Assumptions (ref)-(ref), we can identify the conditional probability of potential outcome $\mathbb{P}(Y_x \in A \mid V = p)$ as \begin{equation} \mathbb{P}(Y \in A \mid X = x, P(Z, X) = p) = \mathbb{P}(Y_x \in A \mid V = p) \end{equation} for $x \in \mathcal{X}$, $p \in \mathcal{P}$, and any set $A \in \mathcal{F}_Y$, where $\mathcal{F}_Y$ denotes the Borel $\sigma$-field generated by $Y$. Also, we can identify the conditional expectation of $Y_x$ given $V=p$ as \begin{equation} \mathbb{E}\left[Y \mid X = x, P(Z, X) = p\right] = \mathbb{E} \left[Y_x \mid V = p\right] \end{equation} for $x \in \mathcal{X}$ and $p \in \mathcal{P}$.
proofSee Appendix (ref).
remarkLemma (ref) and an analog of Lemma (ref) have previously been established by imbens2009identification (henceforth IN). IN consider an augmented setup where $X$ can be a vector $X=(X_1,Z_1)$, where $Z_1$ represent exogenous variables that can enter the outcome equation directly. The second equation of (ref) is then replaced by $X_1=h(Z_1,Z_2,V)$, i.e. $Z_2$ represent excluded instruments that do not enter the outcome equation.
remarkA generalization that nests both the model described above and that of IN is $Y=g(X_1,Z_1,U)$ and $X_1=h(Z_1,Z_2,V)$ (as IN do) but where Assumptions (ref) and (ref) all hold conditional on $Z_1$. That is: $(Z_2 \perp \!\!\! \perp U)|Z_1,V$ and $(Z_2 \perp \!\!\! \perp V) |Z_1$, while the CDF of $V$ conditional on $Z_1$ is absolutely continuous and strictly increasing, with probability one (for all $Z_1$). This allows the observed variables $Z_1$ to serve as genuine control variables, rather than as “included” exogenous regressors as in IN. In this case all of our identification results carry through conditional on $Z_1$, but we omit this generalization for brevity.

We now show that in the case with no control variables $Z_1$, there are no testable implications of Assumptions (ref)-(ref), beyond some regularity and support conditions that can be directly verified. This stands in contrast to parametric models like those studied in Section (ref), which we have shown have testable implications that can be rejected by the data.

The following condition will be useful in establishing Proposition (ref):

conditionFor any fixed $v \in \mathcal{P}$, let $h^*_v(z) = Q_{X|Z=z}(v)$: \begin{enumerate} • $h^*_v$ is surjective ("onto"). That is, for any $x \in \mathcal{X}$, there exists $ z \in \mathcal{Z}$ such that $h^*_v(z)=x$. • $h^*_v$ is injective ("one-to-one"). That is, for any $x \in \mathcal{X}$, there exists at most one $z \in \mathcal{Z}$ such that $h^*_v(z)=x$. \end{enumerate}

Given Assumption (ref), part one of Condition (ref) is equivalent to what IN call “common support”.\footnote{To see this, note that the support of $V$ is $\mathcal{P}$, and the common support condition of IN says that $supp\{P(x,Z)\}=\mathcal{P}$ for all $x \in \mathcal{X}$, in other words for all $v \in \mathcal{P}$ and $x \in \mathcal{X}$, there exists a $z \in \mathcal{Z}$ such that $P(x,z)=v$. Meanwhile, the first part of Condition (ref) says that $v \in \mathcal{P}$ and $x \in \mathcal{X}$, there exists a $z \in \mathcal{Z}$ such that $Q_{X|Z=z}(v)=x$. The two are the same given Assumption (ref).} IN show that all of the quantiles of $g(x,U)$ are identified if Condition (ref).1 holds along with Assumptions (ref)-(ref). Condition (ref) overall holds naturally in parametric models in which $h^*(x,v)$ is additively separable and strictly increasing in $v$, for example, for example if $h_v(z) = \pi \cdot z + \phi(v)$ for some function $\phi$ and $\pi \ne 0$. More generally, both parts of Condition (ref) can be directly tested in the data, since they are features of the observed joint distribution of $X$ and $Z$.

Together, parts 1 and 2 of Condition (ref) say that the function $h^*_v$ is a bijection, for each $x \in \mathcal{X}$ there exists exactly one $z \in \mathcal{Z}$ such that $h^*_v(z)=x$ and vice-versa. A useful implication of Condition (ref) for the result that follows is that it enables us to rewrite events written in terms of $Z$ as events written in terms of $X$, for a fixed value of $P(Z,X)$. Specifically, for any Borel set $A$, the event $Z \in A$ is equivalent to the event $X \in h^*_{P(Z,X)}(A)$ conditional on $P(Z,X)$, where we define $h^*_v(A):=\{h^*_v(A): z \in A\}$.\footnote{To see this note that given Condition (ref) and conditional on the value of $P(Z,X)$, we have that $X=h^*_{P(Z,X)}(z)$ if and only if $Z=z$, since $h^*_{P(Z,X)}(Z)=Q_{X|Z}(F_{X|Z}(X))=X$ and for any $z' \ne z$ we have by Condition (ref) that $h^*_{P(Z,X)}(z') \ne h^*_{P(Z,X)}(z)$.}

Now we are ready to state the main result on the testability of this relaxed IV model:

propositionSuppose that the data satisfies Condition (ref). Then there are no further testable implications of Assumptions (ref)-(ref) together with the model structure (ref), beyond Equation (ref).
proofSee Appendix (ref).

Proposition (ref) contrasts with results that hold in settings without functional form restrictions but where the treatment is discrete kitagawa2015test, Huber2015TestingConstraints, mourifie2017testing, for example the LATE model. In this model, under monotonicity (no-defiers) and random assignment of the instrument, the joint probability of being a complier and having a potential outcome in a given Borel set is point-identified. Testability comes from the non-negativity of this joint probability. This approach relies on the ability to vary the instrument with the treatment value fixed, while maintaining overlap between the first-stage types. Given Condition (ref), this is not possible in the model of this section because there are one-to-one mappings between $X$ and $Z$ and between $X$ and $V$.

Proposition (ref) can also be seen as an extension of the result of gunsilius2021nontestability, who shows that a still weaker nonparametric IV model in which Assumption (ref) is dropped and $V$ is allowed to be multivariate lacks any testable implications whatsoever.

remarkIf Condition (ref) does not hold, which may be expected to occur especially when $Z$ is multidimensional, we have testable implications for Assumptions (ref) - (ref). If those assumptions hold, and there exist $z, z' \in \mathcal{Z}, z \neq z'$ such that $P(z, x) = P(z', x) = p$ given $x$, based on conclusions in Lemma (ref), \begin{equation*} \begin{aligned} & \mathbb{P}(Y \in A \mid X = x, Z = z) \\ =& \mathbb{P}(Y \in A \mid X = x, P(z, X)=p) \\ =& \mathbb{P}(Y_x \in A \mid V = p) \\ =& \mathbb{P}(Y \in A \mid X = x, P(z', X)=p) \\ =& \mathbb{P}(Y \in A \mid X = x, Z = z'). \end{aligned} \end{equation*} Therefore, a testable implication for Assumptions (ref) - (ref) is that $Y \mid \{X = x, Z = z\}$ has the same distribution as $Y \mid \{X = x, Z = z'\}$ if two different instrument values $z, z'$ give the same propensity score evaluated at a fixed $x$.

In our setting, $X$ is potentially endogenous, so that the equality $\mathbb{P}(Y_x \in A) = \mathbb{P}(Y \in A \mid X = x)$ may not hold. Lemma (ref) proves that $P(Z, X)$ serves as a control function. Once conditioning on $P(Z, X)$, $X$ is independent of the potential outcome $Y_x$. Also, since $\mathbb{P}(Y \in A \mid X = x, P(Z, X) = p)$ can be identified in the population level, $\mathbb{P}(Y_x \in A \mid V = p)$ is identifiable. Similarly, we can identify the conditional mean of potential outcome $\mathbb{E}\left[Y_x \mid V = p\right]$, which enables us to identify two parameters of interest, $\operatorname{MTE}\left(p ; x, x^{\prime}\right)$ and $\operatorname{ASF}\left(x\right)$.

theorem(Identifications of MTE and ASF) Consider the model defined in Equation (ref). Suppose that Assumptions (ref)-(ref) hold. Then, we can identify the $\operatorname{MTE} (p; x, x')$ as \begin{equation} \operatorname{MTE} (p; x, x') = \mathbb{E}\left[Y \mid X = x, P(Z, X) = p\right] - \mathbb{E}\left[Y \mid X = x', P(Z, X) = p\right] \end{equation} for any $x \in \mathcal{X}$ and $p \in \mathcal{P}$. If $\mathcal{P}_x = [0, 1]$ for any $x$, we can point identify the average structure function as \begin{equation} \operatorname{ASF} (x) = \int_0^1 \mathbb{E}\left[Y \mid X = x, P(Z, X) = p\right] d p. \end{equation} If $\mathcal{P}_x$ does not have full support but $\mathcal{P}_x \equiv [\underline{p}_x, \bar{p}_x]$, where $\underline{p}_x \geq 0$ and $\bar{p}_x \leq 1$.\footnote{The results can be generalized to cases where the support of the propensity score is not a single interval under some regularity conditions.} Then we can partially identify the average structure function $\operatorname{ASF} (x)$ as \begin{equation} \begin{bmatrix} \int_{[p_x, \bar{p}_x]} \mathbb{E}\left[Y \mid X = x, P(Z, X) = p\right] d p + Y_l (1 - \bar{p}_x + p_x), \\ \int_{[p_x, \bar{p}_x]} \mathbb{E}\left[Y \mid X = x, P(Z, X) = p\right] d p + Y_u (1 - \bar{p}_x + p_x) \end{bmatrix}. \end{equation} provided that $Y_x$ has bounded support $[Y_l, Y_u]$ for any $x \in \mathcal{X}$.
proofTheorem (ref) follows from Lemma (ref) and definitions of $\operatorname{MTE} (p; x, x')$ and $\operatorname{ASF} (x)$.

From Lemma (ref) and Theorem (ref), we can (partially) identify other parameters of interest, such as average treatment effects $\operatorname{ATE}(x, x') \equiv \mathbb{E}[Y_x - Y_x']$ and distributions of potential outcomes $\mathbb{P}\left(Y_x \in A\right)$.

Empirical Illustrations

In this section, we apply our testing procedure to check the instrument validity in two empirical studies. The results show that our testing method rejects the validity of instrumental variables in some settings. Therefore, we recommend researchers test the instrument validity using our test when working with parametric separable models.

Testing the Validity of the Bartik Instrument in card2009immigration

In this section, we revisit card2009immigration and examine the validity of the instrumental variable used in the paper. card2009immigration studies the impact of immigration on wage imbalances in the United States. Particularly, we revisit his analysis on substitution between immigrants and natives within the same skill group (Table 6 in card2009immigration). card2009immigration uses the 1980-2000 census data and the combined 2005 and 2006 American Community Surveys to construct panel data with city-level labor market variables in 1980, 1990, 2000, 2005, and 2006. He focuses on the 124 largest MSAs or PMSAs in the US as of 2000.

To estimate the elasticity of substitution, the paper focuses on a linear regression

equation[equation omitted — 163 chars of source]

where $r_{M j k}$ represents the mean wage residual for immigrant men in the city $j$ and skill group $k \in \{\textit{hs}, \textit{coll}\}$ ($k = \textit{hs}$ stands for high school or equivalent level, and $k = \textit{coll}$ for college or equivalent), $r_{N j k}$ represents the mean wage residual for native men in the city $j$ and skill group $k$, $S_{M j k} / S_{N j k}$ denotes the ratio of immigrant to native working hours in the city $j$ and skill group $k$, including both men and women workers, and $\mathbf{X}_{j}$ is a vector of city-level controls including log city size, college share, mfg. share in 1980 and 1990, and mean wage residuals for all natives and all immigrants in 1980. By the model structure in Equation (ref), the parameter $\beta_1$ measures the negative inverse elasticity of substitution between immigrants and natives in skill group $k$ of city $j$.

However, labor supply may respond to unobserved demand shocks (e.g., new constructions), which could also affect relative earnings. This situation may lead to endogeneity of the variable $S_{M j k} / S_{N j k}$ in this model. To address this potential endogeneity problem, the paper proposes an instrument for the ratio of immigrant to native working hours, $S_{M j k} / S_{N j k}$. The instrument is constructed as

equation[equation omitted — 122 chars of source]

where $\lambda_{m j} \equiv N_{m j} / N_m$ is the fraction of previous immigrants from country $m$ that live in city $j$, $N_m$ denotes the number of previous immigrants from country $m$ in the United States, while $N_{m j}$ denotes the number of previous immigrants from country $m$ living in city $j$. $M_m$ stands for the current number of immigrants from country $m$ to the United States, $P_j$ represents the number of the whole population in city $j$, and $\delta_{m k}$ represents the fraction of immigrants from country $m$ that are in skill group $k$. Specifically, to calculate the instrument $B_{j k}$ for immigration labor supply in city $j$ in 2000, card2009immigration uses national inflows of immigrants from 38 source countries/country groups over the period from 1990 to 2000 to construct $\lambda_{m j}$, and the shares of each group observed in each city in 1980 to construct $\delta_{m k}$. The intuition of choosing the instrument $B_{j k}$ is that since several studies showed that new immigrants tend to move to the same cities as earlier immigrants, $\lambda_{m j} M_m$ can serve as a prediction for the current immigrants from country $m$ to the city $j$ and could also be independent to the current unobserved demand shocks. Also, the instrument does not correlate with local shocks by such construction. The instrument $B_{j k}$ shares the same idea as the instrument in bartik1991benefits and is widely applied across labor economics and international trade.

To identify parameters $(\beta_0, \beta_1, \beta_2')'$, we need to impose the IV assumptions. In this specific case, we need $\mathbb{E} [u_{j k} \mid B_{j k}] = 0$ and $\operatorname{Cov}(S_{M j k} / S_{N j k}, B_{j k}) \neq 0$ to hold. When applying our testing method, the results show that the instrument validity is rejected in the linear structure of Eq.(ref) at all significance levels for both high-school- and college-equivalent groups. For the group with high school and equivalent levels, the test statistic $\hat{\theta}_{1-\alpha}$ in Equation (ref) is equal to $0.081 >0$ for all three significance levels $\alpha = 10\%, 5\%, 1\%$. For the group with college and equivalent level, the test statistic $\hat{\theta}_{1-\alpha}$ is $0.044>0$ for all three levels $\alpha = 10\%, 5\%, 1\%$. A potential reason could be that the unobserved demand shocks for immigrants are persistent over time, which leads to the correlation between the instrument $B_{j k}$ and the current demand shock. Hence, we have seen an example in which Bartik-type instruments may not satisfy the IV assumptions, and our test method provides a credible criterion for researchers to check validity of their instrument in other settings.

Testing the Validity of the Price IV in nevo2012identification

In industrial organization, researchers are often interested in estimating the demand for differentiated products. However, the price variable in the demand function is likely to be endogenous, and researchers commonly use instrument variables to address this issue. In this section, we revisit the application in nevo2012identification. They pointed out that the commonly used price instrument variable may not satisfy the exogeneity assumption. We apply our method to check their conjecture.

Researchers often assume that the difference in the log market shares is linear in the price and observed characteristics,

equation[equation omitted — 142 chars of source]

where $p_{j k}$, $w_{j k}$ represent the price and observed characteristics of product $j$ in market $k$, and $\xi_{j k}$ denotes the unobserved characteristics of product $j$ in the market $k$. The parameters of interest are $(\beta, \Gamma')'$. In practice, the price $p_{j k}$ is potentially correlated with the unobserved demand shock $\xi_{j k}$ (e.g., product quality), which leads to the endogeneity problem. The widely used instruments for $p_{j k}$ are the prices for the same product $j$ in other markets. Ideally, the prices for the same product in markets other than $k$ are related to the price $p_{j k}$ through the common supply shocks (e.g., common production costs), and the demand shocks across different markets are independent. However, the independence of demand shocks across different markets could be violated in some situations. For example, the advertisement within a region could affect the product preference in all markets in that region nevo2012identification. Therefore, nevo2012identification proposed to use prices in other markets as imperfect instrument variables (IIVs), and they study the ready-to-eat cereal industry at the brand-quarter-MSA (metropolitan statistical area) level.

We apply our testing procedure to check whether the prices in other markets satisfy the identifying assumptions in the data studied by nevo2012identification. When we use the average price in the other city as the instrument variable for price $p_{j k}$ and estimate $(\beta, \Gamma')'$ using the two-stage least squares method, the test results show that the instrument variable validity is rejected with imposed parametric assumptions at all three significance levels, as the supremum statistic $\hat{\theta}_{1-\alpha}$ in Equation (ref) is equal to $0.17>0$ for $\alpha = 1\%$, $0.19>0$ for $\alpha=5\%$, and $0.20>0$ for $\alpha = 10\%$. Therefore, we confirm the results in nevo2012identification that the average price in other markets is not a valid instrument for the price in the ready-to-eat cereal industry.

Applying Relaxed Assumptions: card2009immigration

In Section (ref), our conditional moment inequality test rejects the instrument constructed in card2009immigration. If one believes that the instrument $B_{jk}$ is independent of the unobserved demand shocks, then the rejection may caused by the misspecification of the model. In Eq.(ref), card2009immigration assumes that the gap between mean wage residuals, $(r_{M j k}-r_{N j k})$, is linear in the log of immigrant-to-native ratio $\log \left[S_{M j k} / S_{N j k}\right]$, and the observable $u_{j k}$ is additively separable. However, those specifications could be incorrect if the log of the immigrant-to-native ratio has nonlinear effects on the gap between mean wage residuals or if the unobservable enters into the equation in a more complicated way. In this case, we can consider the model (ref) instead to relax those parametric assumptions. card2009immigration also includes a vector of controls in his analysis. We assume that the controlled covariates affect the outcome, the treatment, and the log instrument linearly. To exclude the linear covariates' effect, we regress $(r_{M j k}-r_{N j k})$, $\log \left[S_{M j k} / S_{N j k}\right]$, and $\log \left(B_{jk}\right)$ linearly on controlled covariates and take residuals as the outcome $Y_k$, the treatment $X_k$, and the instrument $Z_k$ that we are interested. Table (ref) presents a summary of statistics of $Y_k$, $X_k$, and $Z_k$ for $k \in \{hs, coll\}$.

table[table omitted — 1,161 chars of source]

Since the instrument $Z_{k}$ is believed to be independent of demand shocks, and the first stage equation in Eq.(ref) can be defined as a conditional quantile function, Assumptions (ref) - (ref) hold under this setting. Then, applying Lemma (ref), we can point identify $\mathbb{P}\left(Y_{k x} \leq y \mid V_k = p\right)$, the conditional distribution of potential wage gaps given the first-stage unobservable, where $Y_{k x}$ stands for the potential wage gaps with the log of immigrant-to-native ratio set to be $x$ in group $k$. For estimation, we first use the local linear regression to estimate the propensity score in Equation (ref), and then plug in the estimated propensity score in Equation (ref) and use local linear regression to estimate the conditional distribution of the potential outcome. We select $x$ to be the median and the 75th percentile of the treatment for both groups\footnote{As shown in Table (ref), the median of treatment $X$ is -0.05 for the high school group and -0.02 for the college group, and the 75th percentile of the treatment $X$ is 0.75 for the high school group and 0.55 for the college group.} and set $p$ equal to 0.5 to plot the estimated conditional distributions of potential wage gaps in Figure (ref). For example, the upper left panel of Figure (ref) shows the conditional distributions of the potential wage gap if the log of immigrant-to-native ratio would equal the median given the first stage error set at 0.5. Notice that in the upper right panel of Figure (ref), the conditional probabilities are not always between 0 and 1. This is due to estimation error. We leave inference to future work.

figure[figure omitted — 256 chars of source]

Theorem (ref) allows us to identify the marginal treatment effect on potential wage gaps if changing the log of the immigrant-to-native ratio from $x'$ to $x$. To estimate the marginal treatment effects, we plug the nonparametrically estimated propensity scores into Equation (ref) and use local linear regression to estimate the conditional means in Equation (ref). We study the marginal treatment effects on potential outcome gaps if changing the treatment value from the 25th percentile to the 75th percentile of the sample\footnote{The 25th percentile of the treatment is -0.71 for the high school group and -0.55 for the college group, and the 75th percentile is 0.75 for the high school group and 0.55 for the college group.}, and the marginal treatment effect if changing the treatment value from the 45th percentile to the 55th percentile of the sample for both groups\footnote{The 45th percentile of the treatment is -0.28 for the high school group and -0.17 for the college group, and the 75th percentile is 0.13 for the high school group and 0.09 for the college group.}. Figure (ref) presents our estimated marginal treatment effects. For example, the upper left panel of Figure (ref) suggests the marginal treatment effect of log immigrant-to-native ratio on the potential wage gap if changing the log immigrant-to-native ratio from its 25th percentile to the 75th percentile, for the first stage error ranging from 0 to 1. We can observe that the marginal treatment effects are heterogeneous in the first stage error $p$.

figure[figure omitted — 255 chars of source]

Summary and Discussion

We propose an easy-to-implement testing procedure for commonly used identifying assumptions in parametric separable models, including the exogeneity, homoskedasticity, and instrument validity assumptions in linear IV models. We derive a testable implication for these assumptions. We transform this testable implication into a set of conditional moment inequalities and apply the method in chernozhukov2013intersection to implement the test. Doing this allows us to use their corresponding Stata packages to perform inference, so that researchers can easily implement our testing procedure in Stata.

We conduct Monte Carlo simulations, and the results suggest that our test performs well in large samples. We also apply our method to test several commonly used instrumental variables in empirical studies. Our proposed test rejects the validity of the Bartik IV used in card2009immigration, and the average price in other markets IV questioned in nevo2012identification.

Then, we provide solutions to researchers if our test rejects their identifying assumptions. We first relax the parametric assumptions. Instead of imposing structural assumptions on the outcome equation, we consider triangular system models with fully nonparametric outcome equations. We show that if researchers maintain the instrument independence assumption under this setting, they can point identify the distributions of potential outcomes and other standard parameters of interest. We further discuss how to “minimally” relax the functional form assumptions to the extent where the IV model becomes untestable. Essentially, we find that the IV model is nontestable if no functional form assumption is made on the outcome equation, when there exists a one-to-one mapping between the continuous treatment variable, the instrument, and the first-stage unobserved heterogeneity.

\onehalfspacing