Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
99,475 characters · 22 sections · 75 citation commands
Partially Linear Models under Data Combination
In this paper, we derive partial identification and inference results for a partially linear model, in a context where the outcome of interest and some of the covariates are observed in two different datasets that cannot be merged. Relevant situations include cases where the researcher is interested in the effect of a particular variable that is not observed jointly with the outcome variable, as well as cases where the outcome and covariates of interest are jointly observed but some of the potential confounders are observed in a different dataset.
Our analysis focuses on a partially linear model of the following form:
in a data combination environment where $F_{Y,X_c}$ and $F_{X_{nc},X_c}$ are supposed to be identified, but the joint distribution $F_{Y,X}$ is not. The variable $X_c$ is thus common to both datasets, whereas the variable $X_{nc}\in\mathbb{R}^p$ is only observed in one of the two datasets. In this setup, $\beta_0=(\beta_{01},...,\beta_{0p})'$ is generally not point-identified, and as a result we focus on the identified set of either $\beta_0$ or $\beta_{0k}$ for some $k\in\{1,...,p\}$; the identified set of $f$ can then be deduced from that of $\beta_0$.
We first derive a tractable characterization of the identified set of $\beta_0$. Unlike many other models considered in the partial identification literature, our setup does not deliver a tractable characterization of the identified set through the support function BontempsMagnacARE17,molinari2019econometrics. However, using Strassen's theorem strassen1965existence, a recent result in optimal transport by backhoff2019existence, and a convenient characterization of second-order stochastic dominance, we show that this set is convex, compact, includes the origin and can be simply constructed from its radial function.\footnote{The radial function $S$ of a closed, compact convex set $\mathcal{C}$ including the origin is defined, for any $q$ on the unit sphere, by $S(q)=\max_{\lambda q \in\mathcal{C}} \lambda$.} The identified set of $\beta_{0k}$, then, can also be computed at low computational cost by solving an unconstrained convex minimization problem.
The characterization of the identified set also implies that point identification may be achieved if $\beta_0=0$, or under a restriction on the unobserved term $Y- f(X_c) - X_{nc}'\beta_0$. While the latter condition is not directly testable, we show how to assess its plausibility when one has access to a validation sample in which the outcome and covariates are jointly observed.
In the partially identified case, the identification region may be reduced by adding restrictions on $f(\cdot)$. The two-sample two-stage least squares estimator (TSTSLS) relies on the assumption $f(X_c)= X_{c,i}'\gamma_0$ for some $\gamma_0$ and $X_c=(X_{c,e}', X_{c,i}')'$. In this context, $X_{c,e}$ (resp. $X_{c,i}$) corresponds to the excluded (resp. included) instruments. This is a leading example that results in point identification. But the exclusion restriction that $E(Y|X)$ does not depend on $X_{c,e}$ may not be credible. We show that alternative restrictions, such as imposing a lower bound on the $R^2$ of the “long regression” of $Y$ on $X_{nc}$ and $X_c$ (in a similar spirit as Oster19, Oster19) or shape restrictions such as monotonicity or convexity of $f$, may in practice dramatically reduce the identified set, and allow to, e.g., identify the sign of $\beta_{0k}$.
Our identification result is constructive, and readily leads to a simple, plug-in estimator of the identified sets for $\beta_0$ or $\beta_{0k}$. A difficulty arises, however, as the estimator of the radial function is generally not asymptotically normal. To construct asymptotically valid confidence regions on $\beta_0$ or confidence intervals on $\beta_{0k}$, we propose to use subsampling politis1999subsampling.
Our method is based on a specific characterization of the identified set, and one may wonder whether alternative characterizations would be more convenient. In particular, the identified set can also be expressed through an infinite collection of moment inequalities. Therefore, general approaches for such problems such as that developed by andrews2017inference could in principle be used instead. We show through simulations the key computational advantage of relying on the method we propose. With a univariate $X_{nc}$, confidence regions are typically computed in seconds, whereas they take up to 30 seconds with a bivariate $X_{nc}$. Compared to the method of andrews2017inference, this corresponds to a dramatic reduction by a factor of more than 1,000 in computational time.
We apply our method to study intergenerational income mobility over the period 1850 to 1930 in the United States, revisiting the analysis of olivetti2015name. In this context where the main variable and outcome of interest are observed in two different datasets that cannot be linked, we show that the confidence sets obtained using our method are quite informative in practice, while allowing us to relax the exclusion restrictions underlying the TSTSLS approach used in olivetti2015name. In the appendix, we consider another application where a key control variable is observed in a separate database. When incorporating sign constraints, our bounds are again very informative.
The method we develop in this paper can be used in a broad set of data combination environments. Two such contexts have attracted much attention in the empirical literature.
One can use our method to conduct inference on the relationship between a particular covariate and an outcome variable, in situations where both variables are not jointly observed. A large literature on intergenerational income mobility often faces the unavailability of linked income data across generations and relies on exclusion restrictions, as in the application we revisit SantavirtaStuhler22. Data combination issues are also common in consumption research, where income (or wealth) and consumption are often measured in two different datasets CLP22. More generally, this type of data combination environment frequently arises in various subfields of empirical microeconomics, including in education and returns to skill estimation RothsteinWozny13,PP16,garcia2016life,hanushek2020culture, health Manski18,RBB22 and labor ACI20. A leading example that has attracted much interest in the literature is one where the researcher seeks to combine experimental data with another observational dataset, in particular situations where data on long-term outcomes is not available in the experimental data.
Our approach can also be used to conduct inference on the causal effect of a variable of interest, in a setup where some of the confounders are observed in an auxiliary dataset. As such, our paper expands the range of data environments in which unconfoundedness is a credible assumption, complementing a literature that focuses on evaluating its reasonableness in the absence of data combination (see, e.g., AET05, AET05; Oster19, Oster19; DMP22, DMP22).
From a methodological standpoint, our paper is connected to the seminal article of CM02 and subsequent work by MP06. They consider the issue of identifying the “long regression”, in our context $E(Y|X_c,X_{nc})$, in the same data combination set-up as here. Importantly though, these two papers focus on deriving the identification region for $E(Y|X_c,X_{nc})$, but do not address the issue of inference. They also consider a setup where the covariates $X_{nc}$ have a discrete distribution with finite support, while we allow $X_{nc}$ to be continuously distributed. On the other hand their setup is entirely nonparametric, whereas we focus on a model that is linear in the covariates $X_{nc}$ and without interaction terms with $X_c$. The linearity assumption plays an important role in our ability to derive a tractable inference method. The absence of interaction further implies that in our set-up, and in contrast with these two papers, the identified set shrinks as one considers different values of $X_c$.
Our paper is also related to pacini2019two and hwang2022bounding. Both papers construct bounds on the best linear predictor of $Y$ on $X$ in a similar data combination framework as here. We show that if one is ready to impose the usual assumption that the model is partially linear, large identification gains may be achieved, possibly up to point identification. hwang2022bounding also considers a set-up where some of the $X$'s are only observed with $Y$ but not with $X_{nc}$, a case we do not study in this paper.
More generally speaking, our paper relates to the broader literature on data combination problems in econometrics and statistics. We refer the reader to RM2007 for a survey of this literature and to fan2014identifying, FSS16, BLL16, and ACI20 for recent contributions. Contrary to ours, most of these papers impose restrictions that entail point identification.
Within the data combination literature, our paper is technically closest to DGM21. Though that paper considered the entirely different context of rational expectation testing, we also relied therein on Strassen's theorem to obtain a characterization of the null hypothesis of rational expectations. Importantly, we extend here our previous main result in a highly non-trivial way, by relying in particular on backhoff2019existence to handle multivariate $X_{nc}$. Also, we previously based our inference on andrews2017inference. In contrast, a key contribution of our paper lies in the novel and tractable inference method that we derive.
Finally, by developing in this data combination context a feasible inference method that can be implemented at a very limited computational cost, our paper also adds to the growing set of papers that propose tractable computational methods for partially identified models (see BontempsMagnacARE17, BontempsMagnacARE17 and molinari2019econometrics, molinari2019econometrics for recent surveys). In particular, our paper fits into the strand of the literature that uses tools from optimal transport to devise computationally tractable identification and inference methods for partially identified models (GalichonHenry11, GalichonHenry11; galichon2016optimal, galichon2016optimal). By characterizing the sharp identified set based on the radial function, a novel approach in the partial identification literature, we show that it is possible to achieve very substantial tractability gains in this context, relative to a more standard characterization in terms of many moment inequalities.
The remainder of the paper is organized as follows. In Section (ref) we present our main identification results for the two-sample partially linear model described above. Section (ref) studies estimation and inference for this model. In Section (ref), we apply our method to intergenerational income mobility in the United States. Section (ref) concludes. The Appendix of the paper gathers additional results on robustness to measurement errors, identification in models with heterogeneous effects of $X_{nc}$ on $Y$, and a test for point-identification. It also presents our second application to the black-white wage gap in the United States. Monte Carlo simulation results, additional material on the application, and the proofs are collected in the online Appendix. Some complements of the proofs appear in supplementary material available in our working paper version DGM22. Finally, our inference method can be implemented using our companion R package, RegCombin, available at \href{https://CRAN.R-project.org/package=RegCombin}{CRAN.R-project.org/package=RegCombin}.
Before presenting our main identification results, we introduce some notation that will be used throughout the paper. We let $\|\cdot\|$, $0_p$ and $\mathcal{S}_p$ denote respectively the usual Euclidean norm in $\mathbb{R}^p$, the vector $0$ and the unit sphere in $\mathbb{R}^p$; we may omit the index $p$ in the absence of ambiguity. For any cumulative distribution function (cdf) $F$ defined on $\mathbb{R}$, we let $F^{-1}(t)=\inf\{x: F(x)\geq t\}$ denote its generalized inverse and $\overline{F}=1-F$ be the corresponding survival function. For any random variable $A$, we let $\text{Supp}(A)$ be its support, $F_A$ denote its cdf. and $V(A)$ its variance, if defined. We also let $\succ_{\!\text{cv}}$ denote the convex ordering, namely, for two random variables $A$ and $B$ with $E[|A|]<\infty$ and $E[|B|]<\infty$, $A\succ_{\!\text{cv}} B$ if $E[\phi(A)]\geq E[\phi(B)]$ for all convex functions $\phi$.\footnote{Even though we may have $E[|\phi(A)|]=\infty$, $E[\phi(A)]$ is always well-defined because $E[\max(0,-\phi(A))] <\infty$, since there exists $a,b$ such that for all $x$, $\phi(x)\ge a+bx$.} We write $A\not\succ_{\!\text{cv}} B$ when $A\succ_{\!\text{cv}} B$ does not hold. Finally, for any sets $C$ and $C'$, we denote by $\partial C$ the boundary of $C$ and by $d_H(C,C')$ the Hausdorff distance between $C$ and $C'$, defined by $$d_H(C,C')=\max\left(\sup_{c'\in C'} \inf_{c\in C} ||c- c'||, \; \sup_{c\in C} \inf_{c'\in C'} ||c- c'||\right).$$
We first consider a linear model and derive the sharp identified set of $\beta_0$ in the absence of common regressors observed in both datasets. We suppose that we observe from two samples that can not be merged the distributions of the outcome, $F_Y$, and covariates, $F_X$. We maintain the following assumption:
We focus hereafter on the identified set $\mathcal{B}$ of $\beta_0$. Since $\mathcal{B}$ is the set of all vectors in $\mathbb{R}^p$ that are compatible with the model and the marginal distributions of $Y$ and $X$, we have
where, for any random variable $A$ with $E[|A|]<\infty$, we let $A_0=A-E(A)$ and we have used that $E(Y|X)=\alpha_0+X'\beta_0$ for some $\alpha_0$ is equivalent to $E(Y_0|X_0)=X_0'\beta_0$. Now, our goal is to express $\mathcal{B}$ to make it amenable to (simple) estimation. To this end, we define, for any $\alpha\in (0,1)$, $F$ and $G$ cdfs with expectation 0, the following functions:
These two functions play an important role in our analysis. Remark that, since $F$ and $G$ are cdfs of mean zero distributions, $\int_{\alpha}^1 F^{-1}(t)dt$ and $\int_{\alpha}^1 G^{-1}(t)dt$ are both positive, so that the ratio of superquantiles $R(\alpha, F,G) $ is well-defined, with $R(\alpha, F,G)>0$ and $S(F,G)\geq 0$. Theorem (ref) is our main identification result.
We now give a sketch of the proof of (ref). Let $\mathcal{B}'$ denote the set on the right-hand side of (ref). First, one can show that by definition of $S(F_{Y_0}, F_{X_0'q})$, $$\mathcal{B}' = \left\{\beta\in\mathbb{R}^p: \forall \alpha \in (0,1), \, \int_\alpha^1 F^{-1}_{X_0'\beta}(t)dt \leq \int_\alpha^1 F^{-1}_{Y_0}(t)dt \right\}.$$ This, in turn, is equivalent to $F_{X_0'\beta}$ dominating $F_{Y_0}$ at the second order de2006stochastic, implying that $$\mathcal{B}'=\left\{\beta\in\mathbb{R}^p: Y_0\succ_{\!\text{cv}} X_0'\beta\right\}.$$ The inclusion $\mathcal{B} \subset \mathcal{B}'$ then follows essentially from Jensen's inequality. As a side remark, note that we can also express $\mathcal{B}'$ through infinitely many moment inequality restrictions:
This equality directly follows from Fubini-Tonelli, applied to the standard characterization of the second-order stochastic dominance condition, namely $\int_{-\infty}^y F_{Y_0}(t)dt \ge \int_{-\infty}^y F_{X_0'\beta}(t)dt$ $\forall y\in \mathbb{R}$. We return to this alternative characterization of the identified set in Subsections (ref) and (ref) of the online appendix, where we document the computational advantages of using our characterization instead.
The inclusion $\mathcal{B}'\subset \mathcal{B}$ is more intricate to prove. Assume $\beta \in \mathcal{B}'$. By what precedes, $Y_0\succ_{\!\text{cv}} X_0'\beta$. Then, by Strassen's theorem strassen1965existence,
This result was already used in DGM21 to characterize the restrictions on $F_Y$ and $F_\psi$ entailed by the rational expectation hypothesis $E(Y|\psi)=\psi$, where $\psi$ denotes the subjective expectations on an outcome $Y$. Importantly though, when $X$ is multivariate, (ref) is not sufficient to conclude that $\mathcal{B}'\subset \mathcal{B}$, as the $\sigma$-algebras generated by $X$ and $X'\beta$ are not equal in general. Nonetheless, we prove, using in particular Theorem 1.3 in backhoff2019existence, that for $\beta \in \mathcal{B}'$,\footnote{We thank Nathael Gozlan for his help in obtaining (ref).}
Together, (ref), (ref), and the existence of a minimizer on the left-hand side of (ref) backhoff2019existence, imply that we can find random variables $\widetilde{Y}$ and $\widetilde{X}$ such that $E[\widetilde{Y}_0 | \widetilde{X}_0]=\widetilde{X}_0'\beta$, $\widetilde{Y}\stackrel{d}{=} Y$ and $\widetilde{X}\stackrel{d}{=} X$. Thus, $\beta\in \mathcal{B}$.
Turning to the second part of the theorem, $0_p\in\mathcal{B}$ follows by noting that one can always rationalize, from the sole knowledge of their marginal distributions, that $X$ and $Y$ are independent. That $\mathcal{B} \subset \mathcal{B}^V$ comes from the inclusion $\mathcal{B} \subset \mathcal{B}'$, combined with the fact that $Y_0\succ_{\!\text{cv}} X_0'\beta$ implies $V(Y)\ge V(X'\beta)$. Hence, $\mathcal{B}$ is included in a bounded ellipsoid. The equality $\mathcal{B} =\mathcal{B}^V$ occurs for instance when $Y$ and $X$ are normally distributed. Otherwise, $\mathcal{B}$ may be substantially smaller than $\mathcal{B}^V$, as we illustrate below. In such cases, $\mathcal{B}^V$ remains a natural benchmark as it is very simple to characterize using $V(Y)$ and $V(X)$ only, and straightforward to estimate.
\paragraph{Radial vs. support function characterization of the identified set.} A key takeaway from Equation (ref) is that the identified set admits a very simple expression as a function of $S$, which is the inverse of the Minkowski gauge function of $\mathcal{B}$ hiriart2012fundamentals, also known as the radial function of $\mathcal{B}$. This function differs from the support function $\sigma$ of $\mathcal{B}$, defined by $\sigma(q,F_{Y_0}, F_{X_0})=\sup_{ b\in \mathcal{B}} \ q'b$. The difference between these two functions is illustrated in Figure (ref).
The partial identification literature has largely relied on support functions, as these are powerful tools that uniquely characterize their convex sets. But the radial function also uniquely characterizes convex sets if, as is the case here, these sets include the origin.\footnote{More generally, star-shaped sets are fully characterized by the radial function (and a given point, $0_p$ in our setup). See Molchanov17, p.156, for more details on this point.} Importantly, this approach allows us to characterize the sharp identified set by minimizing a simple function over the interval $(0,1)$. In contrast, the support function approach will generally be significantly less tractable in our context as it would require solving a high-dimensional constrained optimization problem. Namely, using the characterization of the identified set given in Equation ((ref)) above, the support function can be obtained by solving the following program:
where the constraint itself involves an optimization problem. Simulation results indicate that using the radial function rather than the support function approach does result in very large computational gains, see Online Appendix (ref) for details on this.
\paragraph{Partial identification of subcomponents of $\beta_0$.} The support function still plays a key role in our context when one is interested in a component of $\beta_0=(\beta_{0,1},...,\beta_{0,p})'$, say $\beta_{0,k}$. The following result shows that we can actually recover this function at a low computational cost once $S$ is known. Hereafter, we let $e_k$ denotes the $k$-th element of the canonical basis in $\mathbb{R}^p$ and use the convention $1/0=\infty$ and $1/\infty=0$.
We use the expression (ref) of the support function, rather than the simpler expression $\sigma(e_k, F_{Y_0}, F_{X_0}) = \sup_{q\in \mathbb{R}^p: q_k=1} S(F_{Y_0},F_{X_0'q})$, because $q\mapsto 1/S(F_{Y_0},F_{X_0'q})$ is convex (see the proof of Proposition (ref), which also applies when $\varepsilon=0$), whereas $q\mapsto S(F_{Y_0},F_{X_0'q})$ may not be concave. It follows that one can recover the support function $\sigma$, and in turn the sharp bounds on $\beta_{0,k}$, by simply minimizing a convex function over $\mathbb{R}^{p-1}$.
In some cases, our approach yields point identification of the parameters of interest, or subcomponents of it. Proposition (ref) below presents two such cases under which the identified sets $\mathcal{B}$ and $\mathcal{B}_1$, respectively, boil down to a singleton.
Recall from our main identification result above that the identified set $\mathcal{B}$ always includes the origin. The first point of Proposition (ref) further establishes point identification of $\beta_0=0_p$ when, basically, $Y$ has lighter tails than any linear index of $X$. The second point is similar but focuses on a subcomponent instead: if $Y$ and $X_{-1}'\beta_{-1}$ have lighter tails than $X_1$, then $\beta_{0,1}=0$ is point identified. As an example of function $\phi$ for which Proposition (ref) holds, one might consider for instance $\phi(x)=|x|^a$ for some $a>2$ (in which case $X'\beta$ or $X_1$ have heavy tails), or $\phi(x)=\exp(a |x|^b)$ for some $a,b>0$ (in which case $X'\beta$ or $X_1$ have exponential tails).
To illustrate Point 1 of Proposition (ref), suppose that $p=1$, $X$ follows a Laplace distribution (with density $\exp(-|x|)/2$ on $\mathbb{R}$) and $Y\sim\mathcal{N}(0,1)$. Then, by using $\phi(x)=\exp(|x|^{3/2})$, it follows from Point 1 of Proposition (ref) that $\beta_0=0$ is point identified in this case. On the other hand, the variance restrictions only set identify $\beta_0$, with an identified set given by $\mathcal{B}^V=[-1/\sqrt{2},1/\sqrt{2}]\simeq [-0.707,\,0.707]$. This example illustrates the (in this case point-) identifying power of higher-order moments of the distributions of $X$ and $Y$.
We now turn to the frequent situation where some regressors are observed in both datasets. Namely, suppose we observe regressors $X_c$ that are common to both datasets, and assume that the partially linear model (ref) holds:
The key here is to note, following Robinson88, that this case is equivalent to the previous setup without common regressors once we compute the following residuals, for all $x$ in the support of $X_c$:
It directly follows that $\beta_0$ satisfies $E(Y^x|X^x)=X^x{}'\beta_0$, which allows us to use the characterization of the identified set without common regressors obtained in Section (ref). \\
Let $\mathcal{B}^c$ and $\mathcal{F}$ denote the identified sets of $\beta_0$ and $f$, respectively. We have the following characterization of $\mathcal{B}^c$ and $\mathcal{F}$:
It is possible to extend (ref) by including interaction terms. Notably, such specification allows for heterogeneous effects of $X_{nc}$ on $Y$, which can be important in practice hausman2016fiscal. We consider this extension in Appendix (ref). Another interesting extension corresponds to cases where $E(Y|X)=f(X_c)+X_{nc}'\beta_0+X_a'\delta_0$ and we observe in a first dataset $(Y, X_a, X_c)$ and in a second dataset, $(X_c, X_{nc})$. This setup leads to qualitatively different results. For instance, if there is no common regressors and $(Y, X_a)$ and $X_{nc}$ are Gaussian, one can show that the sharp identified set of $(\beta_0,\delta_0)$ is not convex and does not include $0_{p+r}$ (with $r$ the dimension of $X_a$). We refer the reader to hwang2022bounding for outer bounds on the best linear predictor in this setup and leave its study for future research.
We now consider additional restrictions that may reduce the identified set, in some cases resulting in point identification of the parameters of interest.
A first way to reduce the identified set is to use a lower bound on the predictive power of $X_{nc}$ and $X_c$ with respect to $Y$. To formalize this idea, we assume that $R^2_\ell$, the coefficient of determination of the “long” regression of $Y$ on $X_{nc}$ and $X_{c}$ is higher than a certain threshold. This threshold may be absolute (e.g., 0.1) or relative to $R^2_s:=V(E(Y|X_c))/V(Y)$, the $R^2$ of the “short” regression of $Y$ on $X_c$, which is directly identified from the data. This is in the same spirit as Oster19, who suggests fixing $R^2_\ell/R^2_s$ to 1.3. Note that $$f(X_c)+X_{nc}'\beta= E(Y|X_c) + (X_{nc}- E(X_{nc}|X_c))'\beta,$$ and the two components on the right-hand side are uncorrelated. Thus, $$R^2_\ell = \frac{V(E(Y|X_c)) + \beta'E(V(X_{nc}|X_c))\beta}{V(Y)} = R^2_s + \frac{\beta'E(V(X_{nc}|X_c))\beta}{V(Y)}.$$ Then, if one imposes a lower bound $\underline{R}^2$ on $R^2_\ell$ such that $\underline{R}^2 \geq R^2_s$, the identified set on $\beta$ becomes $$\left\{\lambda q: q\in \mathcal{S}, \, \left(\frac{(\underline{R}^2 - R^2_s) V(Y)}{q'E(V(X_{nc}|X_c))q}\right)^{1/2} \le \lambda \le \overline{S}(F_{Y,X_c}, F_{X_{nc}'q,X_c}) \right\},$$ provided that $E(V(X_{nc}|X_c))$ is nonsingular. This restriction has three key attractive features. First, one can in practice motivate this restriction based on a “validation sample”, namely a subset of the population or another population (e.g., a different country than that under investigation), for which we identify the joint distribution of the outcome and covariates, and thus the $R^2$ of the “long” regression. Second, imposing a lower bound such that $\underline{R}^2>R^2_s$ allows one to exclude $0_p$ from the identified set. Third, the identified set still admits a very simple expression.
Another way to narrow the identified set $\mathcal{B}^c$ with common regressors is to impose some constraints on $f(\cdot)$. Shape restrictions such as monotonicity or convexity often follow from economic theory; see Matzkin94 and chetverikov2018econometrics for econometric reviews, and Tripathi00 and AbrevayaJiang05 for their use and testability with partially linear models. We characterize here the identified set when we impose such restrictions on $f$.
We model these restrictions by $[Rf](r)\ge \underline{c}(r)$ for all $r\in\mathcal{R}$, with $R$ a known linear operator, $\underline{c}$ a known, real function and $\mathcal{R}$ the domain of $[Rf]$ and $\underline{c}$. For instance, if $X_c$ is discrete such that $\text{Supp}(X_c)=\{x_{c,1},...,x_{c,K}\}\subset \mathbb{R}$, with $K>1$ and $x_{c,1}<...<x_{c,K}$, considering $[Rf](r)=f(x_{c,r+1})-f(x_{c,r})$ for $r\in\mathcal{R}=\{1,...,K-1\}$ (resp. $[Rf](r)=(f(x_{c,r+2})-f(x_{c,r+1}))/(x_{c,r+2} - x_{c,r+1}) - (f(x_{c,r+1})-f(x_{c,r}))/(x_{c,r+1} - x_{c,r})$ for $r\in\mathcal{R}=\{1,...,K-2\}$ with $K>2$) and $\underline{c}(r)=0$ corresponds to imposing that $f$ is non-decreasing (resp. convex). When $X_c$ is continuous, the same two constraints can be imposed by considering $[Rf](r)=f'(r)$ and $[Rf](r)=f''(r)$, with $\mathcal{R}=\text{Supp}(X_c)$.
This framework also accommodates restrictions on the magnitude of the effect of $X_c$ on $Y$. Namely, suppose for simplicity that $X_c$ is binary and consider $[Rf](1)=-[Rf](2)=f(x_{c,2})-f(x_{c,1})$ with $\mathcal{R}=\{1,2\}$ and $\underline{c}(1)=\underline{c}(2)=\underline{c}\ge 0$. The extreme case $\underline{c}=0$ corresponds to $X_c$ having no effect on $Y$, as in the two-sample two-stage least squares strategy (see the next subsection for a related, more general point identification result in this context). More generally, this corresponds to the constraint that the magnitude of the effect of $X_c$ is bounded by the cutoff $\underline{c}$, $|f(x_{c,2})-f(x_{c,1})| \leq \underline{c}$.\footnote{If $X_c$ has $K>2$ points of support, the same idea can be generalized by imposing restrictions on $|f(x_{c,k})-f(x_{c,j})|$ for specific pairs $(j,k)\in\{1,...,K\}^2$, $j\ne k$.} By increasing $\underline{c}$, one can therefore study how the identified set varies when relaxing the exclusion restriction, in a similar spirit to, e.g., masten2018identification.
Hereafter, we denote by $m_Y(\cdot)=E[Y|X_c=\cdot]$, $m_{X_{nc}}(\cdot)=E[X_{nc}|X_c=\cdot]$ and
where we let $\sup\emptyset=-\inf\emptyset=-\infty$ and we note that the two functions above may be infinite. We introduce limits to deal with the cases where $[Rm'_{X_{nc}}q](r)=0$. Proposition (ref) characterizes the identified sets of $\beta_0$ and $f$ under such shape restrictions.
In contrast to our baseline identification results in the absence of additional restrictions, the resulting identified set may exclude the origin. This illustrates the practical importance of imposing these types of shape restrictions in contexts where these are likely to hold. Suppose for instance that $p=1$, $X_c$ is binary ($\text{Supp}(X_c)=\{0,1\}$), $\mathcal{R}=\{1\}$ and $[Rf](1)=f(1)-f(0)$, namely we impose that $f$ is non-decreasing. If $f(1)-f(0)< (m_{X_{nc}}(0)-m_{X_{nc}}(1))\beta_0$, then $m_Y(1)<m_Y(0)$. As a result, $0\not\in\mathcal{B}^{\text{\text{con}}}$. The condition $f(1)-f(0)< (m_{X_{nc}}(0)-m_{X_{nc}}(1))\beta_0$ holds for instance if $m_{X_{nc}}$ is decreasing and $\beta_0$ is positive and large enough.
One may alternatively be willing to impose functional form restrictions on $f$. The following proposition shows that this may yield point identification.
This proposition encompasses several popular restrictions. We consider in particular three such restrictions, for which the key point-identifying condition $m_{X_{nc}}'\gamma\not\in\mathcal{G}$ has a simple interpretation:
Finally, if one is ready to impose a relative tail condition between the error term $U:=Y_0-X_0'\beta_0$ and $X_0'\beta_0$, the identified set is considerably reduced. For simplicity, we assume here that there are no common regressors but Proposition (ref) readily extends to accomodate such regressors.
With $X\in \mathbb{R}$, the condition $E[\phi(U\lambda)]<E[\phi(X\lambda)]=\infty$ for all $\lambda>0$ holds for instance if $E[|U|^a]<E[|X|^a]=\infty$ for some $a>2$. More generally, the condition $E[\phi(U\lambda)]<E[\phi(X_0'\beta_0 \lambda)]=\infty$ basically imposes that $X_0'\beta_0$ has fatter tails than $U$. In this sense, this condition is similar to those in Proposition (ref) above.
\paragraph{Testability.} Note that we cannot test the condition $E[\phi(U\lambda)]<E[\phi(X_0'\beta_0 \lambda)]=\infty$ for some convex function $\phi$ and all $\lambda>1$, simply because $U$ is not identified. On the other hand, we can assess the plausibility of $\beta_0 \in \partial \mathcal{B}$ using a validation sample, as defined above. Denoting by $(Y_v, X_v)$ the variables corresponding to this validation sample, it becomes possible to test whether the corresponding parameter $\beta_v=V(X_v)^{-1}\text{cov}(X_v,Y_v)$ is at the boundary of the identified set one would get from the sole knowledge of $F_{Y_v}$ and $F_{X_v}$. Provided that $\beta_v\ne 0$, this condition is indeed equivalent to $\left\|\beta_v \right\|= S(F_{Y_{v 0}}, F_{X'_{v0} \beta_v/\left\|\beta_v \right\|})$ or, in simpler terms, $$S(F_{Y_{v0}}, F_{X'_{v0} \beta_v})=1.$$ We consider a statistical test of this condition in Appendix (ref), and apply it in Section (ref) below.
We illustrate the previous results by considering the following model: $$Y = \gamma_{0,0} + X_c^{1.3} \gamma_{0,1} + X_{nc,1}\beta_{nc,1} + X_{nc,2}\beta_{nc,2} + U, \; U|X\sim\mathcal{N}(0,9).$$ We set the coefficients as follows: $\gamma_{0,0}=-0.1$, $\gamma_{0,1} = 0.3$, $\beta_{nc,1} = 1$ and $\beta_{nc,2}=1$. The variables $X$ are transformations of $(N_1,N_2,N_3)'$, which is supposed to follow a multivariate normal distribution with mean 0 and covariance matrix $$ \Sigma = \left(
\right).$$ Specifically, the common regressor is given by $ X_c = \sum_{k=1}^K (k-1) 1\{ c_{k-1} \leq N_1 \leq c_k \}$, $K=4$, $c_0 = -\infty $, $c_1,\dots, c_{K-1}$, are respectively the quantiles of order 0.1, 0.37, 0.67 and 0.9 of the standard normal, and $c_K=\infty$. We consider two cases for the regressors that are observed in one of the datasets only, $X_{nc}$. In the first case, $(X_{nc,1}, X_{nc,2})=(N_2, \exp(N_3))$ and in the second, $(X_{nc,1}, X_{nc,2})=(\exp(N_2), \exp(N_3))$.
Figure (ref) displays several identified sets for each of the two data-generating processes (DGPs) described above, each of them being associated with particular restrictions. Namely, the set in red, denoted by $\mathcal{B}^V$, is obtained from the variance restrictions only: $$\mathcal{B}^V=\left\{\beta: \beta'V(X^0)\beta\leq V(Y^0)\right\}\cap \left\{\beta: \beta'V(X^1)\beta\leq V(Y^1)\right\},$$ where $X^x$ and $Y^x$ are defined as in Section (ref). Hence, $\mathcal{B}^V$ is the intersection of two ellipses. The set in green, $\mathcal{B}^c$, is obtained as in Proposition (ref) and relies on the restrictions $E(Y^x|X_{nc},X_c=x)=X^{x}{}'\beta_0$ for $x\in\{0,1\}$. Finally, the set in blue, $\mathcal{B}^{\text{\text{con}}}$, is a subset of $\mathcal{B}^c$ that imposes both convexity and monotonicity constraints on $X_c$.
A couple of comments are in order. In case (a) the restrictions implied by the model are much more informative than the variance restrictions, because of the non-normality of $X_{nc,2}$, and in particular the fact that it has fatter tails than the residuals $U$. The true point is at the boundary of $\mathcal{B}^c$, illustrating Proposition (ref) applied conditional on $X_c=0$ and $X_c=1$. In this case, the shape restrictions are sufficient to imply that $0_2\not\in \mathcal{B}^{\text{\text{con}}}$ but also to rule out that $\beta_{nc,1}=0$ as well as $\beta_{nc,2}=0$. The identified set $\mathcal{B}^c$ is reduced further in case (b), as a result of the fatter tails of both $X_{nc,1}$ and $X_{nc,2}$. Like in case (a), the shape constraints on $f(X_c)$ allow to reduce dramatically the identified set.
Figure (ref) presents convexity constraints and different constraints on the $R^2$ on the first DGP. While unlike case (a) of Figure (ref), convexity constraints alone fail to reject $0_2\not\in \mathcal{B}^{\text{\text{con}}}$, imposing a constraint of the form $\underline{R}^2 \geq r R_s^2$ with $r>1$ rejects it by definition. In the latter case, the identified set is no longer convex, allowing to exclude some directions from the identified set and providing an informative lower bound on $|\beta_{nc,1}|$. Overall, that the sharp identified sets $\mathcal{B}^c$ are much more informative than the identified set $\mathcal{B}^V$ based on the variance restrictions highlights the importance of using all of the restrictions implied by the model. Another takeaway from these numerical illustrations is that sign constraints can be very informative in practice, resulting in significant shrinkage of the identified set.
An issue for estimation and inference on $\mathcal{B}$ is that when $\alpha\to 0$ or $\alpha\to 1$, $R(\alpha,F,G)$ is a ratio of two terms tending to 0. It follows that its plug-in estimator may become very unstable. To regularize the problem, we consider an outer set of $\mathcal{B}$ based on the removal of extreme values of $\alpha$. We will focus on this outer set when we turn to estimation and inference in Section (ref). Specifically, we define, for any $\varepsilon\in (0,1/2)$,
Note that for all $F,G$, $\alpha \mapsto R(\alpha, F,G)$ is continuous on $[\varepsilon,1-\varepsilon]$. Thus, the minimum in (ref) is well-defined. Proposition (ref) below describes some properties of $\mathcal{B}_\varepsilon$ and relates it to the sharp identified set $\mathcal{B}$.
The first part of Proposition (ref) states that the regularized set $\mathcal{B}_\varepsilon$, for all $\varepsilon \in (0,1/2)$, preserves the compactness and convexity of the sharp identified set $\mathcal{B}$. The second part states that $\mathcal{B}_\varepsilon$ is always a superset of $\mathcal{B}$, which is arbitrarily close to $\mathcal{B}$ as $\varepsilon\downarrow 0$. The third part states that if, basically, the tails of $\|X_0\|$ are thinner than those of $U$ (Condition (11)), the set $\mathcal{B}_\varepsilon$ coincides with the sharp set $\mathcal{B}$ for $\varepsilon$ small enough. Condition (ref) holds in particular if $X$ has a bounded support and $\text{Supp}(U)=\mathbb{R}$, or if $U$ is symmetric and has a tail index larger than that of $\|X_0\|$. Note that $\mathcal{B}=\mathcal{B}_{\varepsilon}$ may hold even without (ref). For instance, if both $U$ and $X$ are normally distributed, it is easy to check that $\mathcal{B}_\varepsilon=\mathcal{B}$ for all $\varepsilon\in (0,1/2)$.
On the other hand, when $U$ has thinner tails than $X'\beta_0$, $\mathcal{B}_\varepsilon$ will be a strict superset of $\mathcal{B}$ for $\varepsilon$ large enough. In such cases, and under additional restrictions, we provide upper bounds on the Hausdorff distance between $\mathcal{B}$ and $\mathcal{B}_\varepsilon$ in Proposition (ref). Intuitively, these bounds inform us about the maximal possible loss, in terms of identification, that is due to regularization.
The tail conditions imposed in Proposition (ref) are basically the opposite as in Point 3 of Proposition (ref), as they imply that $\|X\|$ has fatter tails than $U$. The assumption that $X$ has an elliptical distribution in Point 1 allows us to relate $S_\varepsilon(F_{Y_0}, F_{X_0'q})-S(F_{Y_0}, F_{X_0'q})$, for any $q\in\mathcal{S}$, with $S_\varepsilon(F_{Y_0},F_{X_0'\beta_0})-S(F_{Y_0},F_{X_0'\beta_0})$, but it is not necessary to obtain an upper bound on $S_\varepsilon(F_{Y_0},F_{X_0'\beta_0})-S(F_{Y_0},F_{X_0'\beta_0})$.
In the two cases of Proposition (ref), we produce upper bounds on the Hausdorff distance between $\mathcal{B}$ and $\mathcal{B}_\varepsilon$ that are, up to some constants, power of the regularization parameter $\varepsilon$. The upper bounds are close to 0 when $\varepsilon$ is small, in line with Point 2 of Proposition (ref). They are also closer to 0 the smaller $c$ is, i.e. the fatter the tails of $X'\beta_0$ (or $X'q$) are, or the larger $d$ is, i.e. the thinner the tails of $U$ are.
We now consider the estimation of the identified set, and how to conduct inference on the parameters of interest $\beta_0$. As in the previous section, we first consider the case without common regressors before showing how to incorporate such regressors and combine them with additional constraints. We conclude this section by discussing some computational aspects of our procedure. We illustrate the finite sample performances of our inference method in Online Appendix (ref).
We rely on random samples from the distributions of $Y$ and $X$.
For any $q\in\mathcal{S}$, let $\widehat{F}_Y$ and $\widehat{F}_{X'q}$ denote the empirical cdf of $Y$ and $X'q$ and let $\widehat{F}_{Y_0}(t)=\widehat{F}_Y(t+\overline{Y})$ and $\widehat{F}_{X'_0q}(t)=\widehat{F}_{X'q}(t+\overline{X}'q)$. We simply estimate $R(\alpha, F_{Y_0},F_{X_0'q})$ and $S_\varepsilon(F_{Y_0},F_{X_0'q})$ by their empirical counterpart $R(\alpha, \widehat{F}_{Y_0}, \widehat{F}_{X'_0q})$ and $S_\varepsilon(\widehat{F}_{Y_0}, \widehat{F}_{X'_0q})$. It turns out that these functions can be computed quickly, as detailed in Section (ref) below. We then also simply estimate the identified set $\mathcal{B}_\varepsilon$ by plug-in: $$ \widehat{\mathcal{B}}_\varepsilon := \left\{\lambda q: q\in \mathcal{S}, \ 0 \leq \lambda \leq S_{\varepsilon}\left(\widehat{F}_{Y_0}, \widehat{F}_{X_0'q}\right)\right\}.$$
Next, we build confidence regions on $\beta_0$. The asymptotic distribution of $S_{\varepsilon}\left(\widehat{F}_{Y_0}, \widehat{F}_{X_0'q}\right)$ is not Gaussian in general, so we rely on subsampling politis1999subsampling. One could alternatively use the numerical bootstrap, see the discussion pp. 18-19 in the first version of DGM22.
Let $n=(n_Xn_Y)/(n_X+n_Y)$ and let $b_n$ denote the size of the subsample. For any estimator $\widehat{\theta}$, let $\widehat{\theta}^*$ denotes its subsampling counterpart. For a nominal coverage of $1-\alpha$, the confidence region on $\beta_0$ we consider is given by $$\text{CR}_{1-\alpha}(\beta_0)= \left\{\lambda q: \ q\in\mathcal{S}, \ 0\leq \lambda \leq S_{\varepsilon}\left(\widehat{F}_{Y_0}, \widehat{F}_{X_0'q}\right) - \widehat{c}_{\alpha,\varepsilon}(q) n^{-1/2}\right\},$$ where $\widehat{c}_{\alpha,\varepsilon}(q) $ is the quantile of order $\alpha$ of the distribution of $b_n^{1/2}[ S_\varepsilon(\widehat{F}^*_{Y_0},\widehat{F}^*_{X_0'q}) - S_{\varepsilon}(\widehat{F}_{Y_0}, $ $\widehat{F}_{X_0'q})]$, conditional on the data.
\paragraph{Inference on subcomponents of $\beta_0$.} In practice, one is often interested in conducting inference on subcomponents of $\beta_0$. In view of (ref), the identified (outer) set $\mathcal{B}_{k,\varepsilon}$ of $\beta_{0,k}$ corresponding to $\mathcal{B}_\varepsilon$ satisfies
where $\sigma_\varepsilon(\cdot, F_{Y_0}, F_{X_0})$ denotes the support function associated to $q\mapsto S_\varepsilon(F_{Y_0},F_{X_0'q})$ and $e_k$ is the $k$-th element of the canonical basis of $\mathbb{R}^p$. To construct confidence intervals on $\beta_{0k}$, we first estimate $\sigma_\varepsilon(\cdot, F_{Y_0}, F_{X_0})$ by
see Corollary (ref). Then, denoting by $\widetilde{c}_{\beta,\varepsilon}(e)$ the quantile of order $\beta\in (0,1)$ of the distribution of $b_n^{1/2}( \sigma_{\varepsilon}(e, \widehat{F}^*_{Y_0}, \widehat{F}^*_{X_0}) - \sigma_\varepsilon(e, \widehat{F}_{Y_0}, \widehat{F}_{X_0}))$, conditional on the data, the confidence interval we consider for $\beta_{0,k}$ is $$\text{CI}_{1-\alpha}(\beta_{0,k})= \left[\left(-\sigma_\varepsilon(-e_k, \widehat{F}_{Y_0}, \widehat{F}_{X_0}) + \frac{\widetilde{c}_{\alpha,\varepsilon}(-e_k)}{n^{1/2}}\right)^-, \left(\sigma_\varepsilon(e_k, \widehat{F}_{Y_0}, \widehat{F}_{X_0}) - \frac{\widetilde{c}_{\alpha, \varepsilon}(e_k)}{n^{1/2}}\right)^+ \right],$$ where $x^-=\min(0,x)$ and $x^+=\max(0,x)$. The rationale for using $(\cdot)^-$ and $(\cdot)^+$ is to ensure that $0\in \text{CI}_{1-\alpha}(\beta_{0,k})$: recall that without constraints, $0\in \mathcal{B}_{k,\varepsilon}$. The advantage, then, is that we can still use the quantiles of order $\alpha$ while maintaining coverage even under point identification, as formally shown in Theorem (ref) below.
\paragraph{Choice of the regularization parameter $\varepsilon$.} Because $S_\varepsilon(F_{Y_0},F_{X_0'q}) \geq S(F_{Y_0},F_{X_0'q})$, the confidence regions and intervals above are conservative in general. To gain in efficiency, we suggest using several $\varepsilon$, and, basically, keep the one leading to the smallest confidence regions or intervals. We distinguish the cases $p=1$, where we can adapt the choice to the direction $q\in\mathcal{S}$ while preserving the convexity of $\widehat{\mathcal{B}}_\varepsilon$, from the case $p>1$. When $p=1$, let us define, for $q\in\mathcal{S}=\{-1,1\}$,
where $\mathcal{E}$ is a finite grid in $(0,1/2]$. Hence, $\varepsilon(q)$ simply minimizes the boundary value of the confidence region in the direction $q\in\mathcal{S}$. This idea is similar to that of chernozhukov2013intersection in the context of intersection bounds.
Now consider the case $p>1$. If one focuses on confidence intervals on $\beta_{0k}$, we need to choose the parameter $\varepsilon$ that appears in $\sigma_\varepsilon(\pm e_k, F_{Y_0}, F_{X_0})$. To this end, we simply use $\varepsilon(q)$ as given above, with $q=\pm e_k$. If we are interested instead in the set $\mathcal{B}$ itself, we recommend using $\underline{\varepsilon} = \min_{q \in \mathcal{Q}} \varepsilon(q)$, where $\mathcal{Q}$ is a finite subset of $\mathcal{S}$.
The following theorem shows that $\widehat{\mathcal{B}}_\varepsilon$ is consistent for $\mathcal{B}_\varepsilon$, in the sense of the Hausdorff distance, under mild regularity conditions.
Next, we establish the asymptotic validity of $\text{CR}_{1-\alpha}(\beta_0)$ and $\text{CI}_{1-\alpha}(\beta_{0,k})$, under Assumptions (ref) and (ref) respectively. Assumption (ref) (resp. (ref)) is used to establish the asymptotic validity of $\text{CR}_{1-\alpha}(\beta_0)$ (resp. $\text{CI}_{1-\alpha}(\beta_{0,k})$) using $\varepsilon(q)$ or $\underline{\varepsilon}$ (resp. $\varepsilon(\pm e_k)$), as defined above, instead of a fixed $\varepsilon$.
The second part of Assumption (ref) holds if for all $q\in\mathcal{S}$, the distributions of $X'q$ and $Y$ are continuous with respect to the Lebesgue distribution and their support is a (possibly unbounded) interval. The first part of Assumption (ref) is basically a reinforcement of Assumption (ref) to ensure that some of our results hold uniformly over $q$. This is needed when we consider the support function, as this function implies an optimization over $q$. A sufficient condition for (ref) is that, for all $q\in\mathcal{S}$, $X'q$ admits a density $f_{X'q}$ with respect to the Lebesgue measure and $\inf_{(q,\alpha)\in\mathcal{S}\times[\varepsilon,1-\varepsilon]} f_{X'q}(F_{X'q}^{-1}(\alpha))>0$. The conditions (ii) and (iii) in Assumption (ref) are sufficient conditions for the continuity of the asymptotic distribution of $n^{1/2}\left(\sigma_\varepsilon(e, \widehat{F}_{Y_0}, \widehat{F}_{X_0})- \sigma_\varepsilon(e,F_{Y_0},F_{X_0})\right)$, which is necessary for the validity of subsampling.
Assumption (ref) can accomodate DGPs where the tails of $\|X_0\|$ are thinner than those of $U$ (which may correspond to $a\mapsto R(a,F_{Y_0},F_{X_0'q})$ admitting a unique minimum) but also DGPs for which the opposite holds (since in this case we can have $S_{\varepsilon}(F_{Y_0},F_{X_0'q})>S(F_{Y_0},F_{X_0'q})$ for all $\varepsilon\in (0,1/2)$ and $q$). For instance, one can check that it holds if $Y=c+X+U$ with $c\in\mathbb{R}$, $X\perp \!\!\! \perp U$, and either $X$ follows a Laplace distribution while $U$ is uniform, or the other way around. But it fails to hold when both $X$ and $Y$ are Gaussian, since then $a\mapsto R(a,F_{Y_0},F_{X_0'q})$ is actually constant. Assumption (ref) is basically similar to Assumption (ref) but somewhat more complicated, as we consider therein the support function instead of the radial function.
To prove (ref)-(ref), we first show the weak convergence of $$\sqrt{n}\left(R(\alpha, \widehat{F}_{Y_0},\widehat{F}_{X_0'q}) - R(\alpha, F_{Y_0},F_{X_0'q})\right),$$ seen as a process indexed by either $\alpha$ or $(\alpha,q)$. The convergence in distribution of $S_\varepsilon(\widehat{F}_{Y_0}, \widehat{F}_{X_0'q})$ and $\sigma_\varepsilon(e, \widehat{F}_{Y_0}, \widehat{F}_{X_0})$, and in turn (ref)-(ref), then essentially follows by the Hadamard directional differentiability of the minimum and maximin maps, shown respectively by carcamo2019directional and firpo2021uniform.
Our results for a fixed $\varepsilon>0$ extend to the data-dependent $\varepsilon(q)$ and $\underline{\varepsilon}$, under the additional conditions provided above. Note that one could avoid these conditions by using sample splitting, with one subsample used to choose $\varepsilon(q)$ or $\underline{\varepsilon}$ and the other to construct the confidence regions/intervals. One drawback of this alternative solution, though, is that it increases the size of confidence regions/intervals, to a point that we may lose the benefits of using a data-dependent rather than a fixed $\varepsilon$.
We now turn to inference on $\beta_0$ with common regressors $X_c$. Recall from Proposition (ref) that the identified set on $\beta_0$ is $$\mathcal{B}^c = \left\{\lambda q: q\in \mathcal{S}, \; 0\leq \lambda\leq \overline{S}(F_{Y,X_c}, F_{X_{nc}'q,X_c})\right\},$$ with $\overline{S}(F_{Y,X_c}, F_{X_{nc}'q,X_c}) = \inf_{x\in \text{Supp}(X_c)} S(F_{Y^x|X_c=x}, F_{X^x{}'q|X_c=x})$.
Let us first assume that $X_c$ has a finite support. Let $\widehat{F}_{Y^x|X_c=x}$ and $\widehat{F}_{X^x{}'q|X_c=x}$ denote the empirical estimators of $F_{Y^x|X_c=x}$ and $F_{X^x{}'q|X_c=x}$, respectively. Following the same logic as above, we estimate $\overline{S}(F_{Y,X_c}, F_{X_{nc}'q,X_c})$ by $$\widehat{\overline{S}}(q, F_{Y,X_c},F_{X_{nc}'q,X_c}) = \min_{x \in \text{Supp}(X_c)} S_\varepsilon(\widehat{F}_{Y^x|X_c=x},\widehat{F}_{X^x{}'q |X_c=x}).$$ Let $\widehat{c}^c_{\alpha,\varepsilon}(q)$ be the quantile of order $\alpha\in(0,1)$ of the distribution of $b_n^{1/2}(\widehat{\overline{S}}{}^*(q, F_{Y,X_c},F_{X_{nc},X_c})$ $- \widehat{\overline{S}}(q, F_{Y,X_c}, F_{X_{nc},X_c}))$, conditional on the data. For a nominal coverage of $1-\alpha$, the confidence region on $\beta_0$ we consider is $$\text{CR}_{1-\alpha}^c(\beta_0)= \left\{\lambda q: q\in\mathcal{S}, \ 0 \le \lambda \le \widehat{\overline{S}}(q, F_{Y,X_c},F_{X_{nc},X_c}) - \widehat{c}^c_{\alpha,\varepsilon}(q)n^{-1/2} \right\}.$$
With continuous common regressors, one can adapt the earlier arguments using sieve estimation. Specifically, suppose that Model (ref) holds and consider a linear sieve approximation of $f(\cdot)$ by a step function $x_c\mapsto \sum_{k=1}^{K_n} \mathds{1}\left\{x_c\in I_{n,k}\right\}\gamma_k$ for some partition $(I_{n,k})_{k=1...K_n}$ of the support of $X_c$ and with $K_n$ tending to infinity at an appropriate rate. Then, one can construct a confidence region on $\beta_0$ by following a similar logic as above.\footnote{Establishing the asymptotic validity of such a confidence region would require to handle both the bias stemming from the approximation of $f(\cdot)$ and the increasing complexity of the approximation. We leave this analysis for future research.}
We now discuss how to conduct inference under constraints on the $R^2$ or shape restrictions, as considered in Subsections (ref) and (ref) respectively. The main difference with above is that for a given direction $q\in\mathcal{S}$, both the lower and upper bounds on the identified set need to be estimated. As before, we can estimate them with plug-in estimators. The only substantive difference is that in the confidence regions, we need to account for the variability of both bounds. For instance, with shape restrictions, we can consider the following confidence region:
where $\widehat{\underline{c}}^{con}_{\delta,\varepsilon}(q)$ is the quantile of order $\delta$ of $b_n^{1/2}(\widehat{\underline{S}}{}^{con*}(q, F_{Y,X_c},F_{X_{nc},X_c})$ $- \widehat{\underline{S}}{}^{con}(q, F_{Y,X_c},$ $F_{X_{nc},X_c}))$, conditional on the data and similarly for $\widehat{\overline{c}}{}^{con}_{\delta,\varepsilon}$. We conjecture that with a finite number of constraints, $X_c$ finitely supported and if $[Rm_{X_{nc}}'q](r)\ne 0$ for all $r\in\mathcal{R}$, $\text{CR}_{1-\alpha}^{con}(\beta_0)$ is pointwise asymptotically conservative. Alternatively, one could use the formulation of our problem with shape constraints as a set of infinitely many moments inequalities. While generally far less tractable that our baseline approach, confidence intervals based on the inversion of the test of these many moment inequalities have uniformly correct asymptotic size andrews2017inference.
We first discuss how to efficiently compute $S_\varepsilon(\widehat{F}_{Y_0},\widehat{F}_{X_0'q})$. Let $Y_{(1)}<...<Y_{(m_y)}$ represent the $m_y\le n_y$ distinct, ordered values of the $(Y_i)_{i=1,...,n_y}$ and let $W_{(j)}^Y=\#\{i:Y_{i}=Y_{(j)}\}/n_y$. Let us also define $I^Y=\{\sum_{j=1}^i W_{(j)}^Y: i=1,...,m_y-1\}$. We define similarly $W_{(j)}^{X'q}$ and $I^{X'q}$. By construction, the numerator $\widehat{f}^Y(\alpha):=\int_\alpha^1 \widehat{F}^{-1}_{Y_0}(t)dt$ of $R(\alpha, \widehat{F}_{Y_0}, \widehat{F}_{X'_0q})$ is linear on all intervals $[\sum_{j=1}^i W_{(j)}^Y, \sum_{j=1}^{i+1} W_{(j)}^Y]$ ($i=0,...,m_y-1$). Moreover, for any $\alpha=\sum_{j=1}^i W_{(j)}^Y\in I^Y$,
The same holds for the denominator $\widehat{f}^{X'q}(\alpha) $ of $R(\alpha,\widehat{F}_{Y_0}, \widehat{F}_{X'_0q})$. As a result, $R(\alpha, \widehat{F}_{Y_0},$ $\widehat{F}_{X'_0q})$ is of the form $(a \alpha + b)/(c\alpha + d)$ on intervals between two consecutive values of $I^Y \cup I^{X'q}$. Now, observe that the minimum of such a function is reached at one of the endpoints of the interval. As a result, we can compute $S_\varepsilon(\widehat{F}_{Y_0},\widehat{F}_{X_0'q})$ using the following algorithm:
To compute $\sigma_\varepsilon(\pm e_k, F_{Y_0}, F_{X_0})$, we solve (ref), in which $q \mapsto 1/S_\varepsilon(\widehat{F}_{Y_0},\widehat{F}_{X_0'q})$ is also convex. In practice, we use the BFGS quasi-Newton method implemented in the R package optim, using as a starting point the considered direction $e$.
Finally, the exact computation of $\widehat{\mathcal{B}}_\varepsilon$ and $\text{CR}_{1-\alpha}(\beta_0)$ requires the computation of $S_\varepsilon(\widehat{F}_{Y_0}, \widehat{F}_{X_0'q})$ and $\widehat{c}_{\alpha,\varepsilon}(q)$ for all $q\in\mathcal{S}$, which is in practice infeasible if $p>1$ as $\mathcal{S}$ is infinite. Instead, we suggest to (i) fix a grid $\widetilde{\mathcal{S}}\subset \mathcal{S}$; (ii) compute $S_\varepsilon(\widehat{F}_{Y_0}, \widehat{F}_{X_0'q})$ and $\widehat{c}_{\alpha,\varepsilon}(q)$ for each $q\in\widetilde{\mathcal{S}}$; (iii) construct an approximation of $\widehat{\mathcal{B}}_\varepsilon$ and $\text{CI}_{1-\alpha}$ by computing the convex hulls of $\{S_\varepsilon(\widehat{F}_{Y_0}, \widehat{F}_{X_0'q})q: \ q\in\widetilde{\mathcal{S}}\}$ and $\left\{\left(S_\varepsilon(\widehat{F}_{Y_0}, \widehat{F}_{X_0'q}) - \widehat{c}_{\alpha,\varepsilon}(q) n^{-1/2}\right)q:\right.$ $ \left. q\in\widetilde{\mathcal{S}}\right\}$, respectively.\footnote{The convex hull of $n$ points in $\mathbb{R}^p$ can be computed efficiently by the quickhull algorithm barber1996quickhull, which requires around $n^{p/2}$ operations.} The resulting sets, $\widetilde{\mathcal{B}}_\varepsilon$ and $\widetilde{\text{CR}}_{1-\alpha}(\beta_0)$ say, are convex, inner approximations of $\widehat{\mathcal{B}}_\varepsilon$ and $\text{CR}_{1-\alpha}(\beta_0)$, and satisfy, as $d_H(\mathcal{S},\widetilde{\mathcal{S}})\to 0$, $d_H(\widetilde{\mathcal{B}}_\varepsilon, \widehat{\mathcal{B}}_\varepsilon) \to 0$ and $d_H(\widetilde{\text{CR}}_{1-\alpha}(\beta_0), \text{CR}_{1-\alpha}(\beta_0)) \to 0$.
The computation of the estimated set, the confidence regions on $\beta_0$ and $\gamma_0$ in the specification $f(X_c) =X_c'\gamma_0$ (where $X_c$ is the vector of all dummy variables associated with a finitely supported variable) and the confidence intervals on the corresponding subcomponents are implemented in our companion R package RegCombin. The package also handles shape restrictions and lower bound on the $R^2$ of the long regression, as well as combinations of these. The RegCombin vignette, available through the description of the package on CRAN, provides additional details about the implementation, including the choice of the tuning parameters $\mathcal{E}$ and $b_n$.
We now apply our method to conduct inference on the intergenerational income mobility over the period 1850 to 1930 in the United States, revisiting the influential analysis of olivetti2015name on this question. We follow their paper and focus on the father-son and father-son-in-law intergenerational income elasticities. We conduct our analysis using 1 percent extracts from the decennial censuses of the United States, over the period 1850 to 1930 (1850-1930 IPUMS).\footnote{We refer the reader to Section 2 of olivetti2015name for a detailed discussion of the data used in the analysis. Note that they estimate the evolution of the intergenerational income mobility over a longer time window (1850 to 1940) than we do. We confine our analysis to the period 1850-1930 as the 1940 portion of the data (1% extract of the IPUMS Restricted Complete Count Data) is not publicly available.}
An important feature of the historical Census data used in this analysis is that father's and son's (as well as son-in-law's) incomes are not jointly observed. olivetti2015name address this measurement issue by predicting, for any given child (John, say) observed in one of the Census datasets, their father's log earnings using the mean log earnings of fathers whose children have the same first name (namely, John). olivetti2015name then estimate in a second step the intergenerational elasticity by regressing son's log earnings on the predicted father's log earnings computed from the previous step. This procedure boils down to a two-sample two-stage least squares estimator (TSTSLS).\footnote{Another limitation of the data used in olivetti2015name and in this application is that it does not allow us to directly calculate the intergenerational elasticity in income. Instead, we follow the baseline specification of olivetti2015name and proxy income using an index of occupational standing available from IPUMS (OCCSCORE), which is constructed as the median total income of the persons in each occupation in 1950.} The corresponding exclusion restriction that the son's first name does not predict his log earnings, once we control for his father's log earnings, may nonetheless be problematic; see SantavirtaStuhler22 for a critical review of the empirical literature using TSTSLS in this context of intergenerational mobility. For the periods 1860-1880 and 1880-1900 only, the IPUMS Linked Representative Samples link fathers and sons using information on first and last names, which allows us to estimate more directly the father-son elasticity using OLS.
Using our notation and consistent with olivetti2015name, the population parameter of interest here is given by $$\theta_0:=\frac{\text{Cov}(Y,X_{nc})}{V(X_{nc})} = \beta_0 + \left(\frac{\text{Cov}(X_c,X_{nc})}{V(X_{nc})}\right)'\gamma_0,$$ where $Y$ denotes the son's (or son-in-law's) log-income, $X_{nc}$ the father's log-income and $X_c$ the vector of indicators corresponding to the son's (or son-in-law's) first names observed in both datasets. The second equality follows from (ref), since $X_c$ is discrete and thus $f(X_c)=X_c'\gamma_0$ for some $\gamma_0$. In what follows, we report the upper bound of the estimated identified set and confidence interval on $\theta_0$.
Even though the sample sizes as well as the number of common regressors $X_c$ are quite large, our method can still be implemented at a very reasonable computational cost. For instance, for the sample of sons over the first period (1850-1870), the computation of the confidence intervals only takes less than 4 minutes with our R package. As expected, computational time is highest for the period 1910-1930 associated with the largest number of observations, with $n > 100,000$ for both samples of $Y$ and $X_{nc}$. Nonetheless, our inference procedure remains tractable in this case too, with a computational time of about 11 minutes.\footnote{These CPU times are obtained using our companion R package, parallelized on 20 CPUs on an Intel Xeon Gold 6130 CPU 2.10GHz with 382Gb of RAM.} Overall, this illustrates the applicability of our method, which can be easily implemented even in this type of rich and high-dimensional data environment.
Figures (ref)-(ref) and Table (ref) below display the results, for the father-son as well as father-son-in-law elasticities, obtained using our approach, the TSTSLS and, for the sample of sons over the years 1860-1880 and 1880-1900, the OLS.\footnote{In practice we need to restrict the set of first names included in $X_c$ to avoid very uncommon occurrences that are perfect predictors of the outcome variable $Y$. In our baseline specification, we implement this by restricting $X_c$ to the set of first names that account for at least 0.01% of the observations in the pooled sample, and appear at least 10 times in either of the samples. We discuss in the following the robustness of our results to alternative cutoffs.} Specifically, we report in Figures (ref)-(ref) the estimated upper bounds of the identified sets (in solid red) and the confidence intervals (dashed red) obtained with our method, the TSTSLS estimates and confidence intervals (solid and dashed blue, resp.) as well as, for 1860-1880 and 1880-1900 and the sample of sons only, the OLS estimates and confidence intervals (solid and dashed green, resp.).
A first conclusion from these results is that the upper bounds of the confidence intervals associated with our method range, depending on the periods, between 0.48 and 0.61 (0.51 and 0.6) for the sample of sons (sons-in-law). These values of the intergenerational coefficient are all well below the natural upper bound of 1. Also, even though the estimates vary depending on the data and econometric specification being used, most of the existing point estimates of the father-son income elasticity range between 0.40 and 0.50 olivetti2015name. Overall, this clearly indicates that our method leads to informative inference on the parameter of interest.
Second, consider the two cases where the linked data is available (1860-1880 and 1880-1900 for the sample of sons). Results in Table (ref) indicate that the corresponding OLS estimates of the intergenerational income elasticities are quantitatively very close to the estimated upper bound of our identified set. Recall that, from Proposition (ref) in Section (ref), the upper bound of our identified set ($\overline{\theta}_0$, say) plays a special role: under an additional restriction on the distributions of $X_{nc}$ and the error term, $\theta_0$ is actually point identified and equal to $\overline{\theta}_0$.\footnote{Proposition (ref) is obtained without $X_c$. Yet, it can be combined with Proposition (ref) to show that $\beta_0$, and in turn $\gamma_0$ (and thus $\theta_0$ here) are point identified with such $X_c$.} In other words, the results from these two periods support the hypothesis that the restriction on the distributions of $X_{nc}$ and the error term guaranteeing point identification of $\theta_0$ by $\overline{\theta}_0$ hold.
Besides, the fact that we do not reject at standard levels the null hypothesis of point identification with our formal test described in Section (ref) for the period 1860-1880 (p-value of 0.147) provides suggestive evidence in this direction.\footnote{Simulation results available from the authors upon request indicate that our choice of $\varepsilon$ tends to be conservative for the test of point identification. One would not reject either at the 1% level the null hypothesis for the period 1880-1900 with a less conservative choice of $\varepsilon$ (e.g. we obtain a p-value of 0.04 using $\varepsilon/2$).} Under this assumption, our results are informative not only on the maximal father-son elasticity coefficient for a given period of time, but also on its evolution. It follows in particular that our estimates point to a mild decrease in this elasticity coefficient for sons between 1850 and 1930.
Third, the results from the equality test reported in Table (ref) indicate that the TSTSLS estimates are in several cases statistically distinguishable from the estimated upper bounds of our identified sets. This includes, for the sample of sons, all periods except 1900-1920, and the periods 1850-1870 and 1880-1900 for the sample of sons-in-law. Besides, for the sample of sons in particular, the TSTSLS estimates exhibit a sharp increase, while our estimated upper bound decreases between the periods 1880-1900 and 1900-1920. In that sense, our results offer suggestive evidence that the intergenerational income correlation might have been more stable at the beginning of the 20th century than what one would infer from the TSTSLS estimates.
Fourth, we also report in Table (ref) the estimated identified set and confidence intervals associated with our method when we impose a lower bound on the $R^2$ of the long regression, namely $\underline{R}^2 \geq 1.3 R_s^2$ or $\underline{R}^2 \geq 2 R_s^2$ . Imposing any of these restrictions, which are satisfied for the periods 1860-1880 and 1880-1900 for which the linked data is available, results in substantially tighter confidence intervals. In particular, for the sample of sons, the confidence intervals obtained under the restriction $\underline{R}^2 \geq 2 R_s^2$ allow us to reject values of the intergenerational income elasticity coefficient smaller than 0.13 and larger than 0.48 for the years 1910-1930.
We consider in Tables (ref) and (ref), and Figure (ref) in online Appendix (ref) several robustness checks. They relate to the set of first names that we include as controls in our estimation procedure (Panel A), the choice of $\varepsilon$ (Panel B and Figure (ref)), and restrictions of the sample to the set of individuals whose first name is included in the set of controls $X_c$ (Panel C). Throughout the tables, we focus on the upper bound of the estimated identified set (“DGM, set”) and of the confidence interval (“DGM, CI.”).
The main takeaway from Table (ref) and Figure (ref) is that, for the sample of sons, the results from our inference procedure are qualitatively, and in most cases quantitatively, robust to these different sensitivity analyses. The one case that exhibits more sensitivity is the specification where we control for the first names that account for at least 0.02% of the sample, instead of 0.01% in our baseline specification. The upper bound of our confidence interval for the period 1900-1920 increases in this case from 0.48 to 0.58, the results remaining, however, stable for the other periods. The results for the sample of sons-in-law (Table (ref)) are also, for most periods at the exception of the same limit for 1900-1920, qualitatively, and in some cases quantitatively similar across specifications. The main difference with the sample of sons is that the choice of $\varepsilon$ does appear to matter more for the sons-in-law, a limitation that one should keep in mind when interpreting the findings for this subgroup. Nonetheless, to the extent that our baseline choice of $\varepsilon$ (see Section (ref)) is motivated by the theory and is found to perform well in our Monte Carlo simulation exercises, we do not view this as particularly worrisome.
\FloatBarrier
\linespread{1}\selectfont \linespread{1.3}\selectfont