Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
91,578 characters · 9 sections · 20 citation commands
Partial Identification of Expectations with Interval Data
\doublespacing
\nocite{Bratberg2017}
The value of a conditional expectation function (CEF) can at best be partially identified when the conditioning variable is interval censored Manski2002. When the observed intervals are coarse or the CEF slope is large in magnitude, existing methods may yield bounds that are minimally informative. In this paper, we develop three innovations that can yield narrower bounds on parameters of interest, and we develop analytical and numerical methods to calculate these bounds. We apply the methods in two policy-relevant settings: the estimation of mortality as a function of education Meara2008,Case2015, and the estimation of intergenerational mobility Solon1999,Guell2013,Chetty2014b.
First, we show that using information on the distribution of the conditioning variable leads to tighter bounds on the CEF. We prove sharp analytical bounds on the value of the CEF when the latent conditioning variable has a known distribution but is interval-censored. This approach is broadly applicable, because distributions are known or commonly assumed for many economic variables. For some conditioning variables (such as ranks), no additional assumptions are required; for example, ranks are uniform by construction. For others (such as income), distributional assumptions on the variable of interest are common and reasonable, and results under alternative assumptions can be tested.
Second, we derive a class of measures that describe the CEF mean across a fixed interval of the conditioning variable. Such interval means can in practice be bounded tightly in many cases and point estimated for some intervals. This makes meaningful inference possible for policy-relevant parameters even when the bounds on the CEF itself are very wide. For example, when the bottom bin is large, the mean value of the CEF in the bottom quintile of the conditioning variable may be bounded more tightly than the value of the CEF at any point in the bottom quintile.
Third, we show that a curvature constraint on the CEF can be implemented in a nonparametric setup using numerical constrained optimization, further narrowing bounds on the CEF and functions of the CEF. The assumption of limited curvature can also substitute for the monotonicity assumption that earlier approaches to this problem relied upon. In our applications, conservative curvature limits generate identified sets of similar size to those under assumptions of monotonicity. This result provides a tractable framework for nonparametric inference with interval-censored data even in contexts without monotonicity.
In practice, we find that to obtain informative bounds, the first and second innovations (known distribution and interval means) are required. Further, even a weak curvature constraint can tighten the bounds significantly. Although this paper focuses on conditional expectation functions, the method can be directly applied to any function or moment of the variable of interest. For example, the method can bound any percentile of the conditional distribution of $y$ given an interval-censored variable $x$.
This paper contributes to a growing literature focused on partial identification of solutions to problems where point identification is difficult without excessively restrictive assumptions Manski2003,Tamer2010,Ho2015a. This paper is most closely related to \citeasnoun{Manski2002}, who calculate analytical bounds on a CEF with an interval-censored conditioning variable from an unknown distribution. The \citeasnoun{Manski2002} bounds are sharp---we can only improve upon them by making additional assumptions, but the assumptions we make are weak and reasonable in many contexts, and tighten bounds in some cases by an order of magnitude or more. We also provide a tractable numerical framework for calculating nonparametric bounds under more complex constraints. Our analysis is limited to conditional expectation functions with a single parameter. As we show below, this setup nevertheless describes a broad class of problems and the innovations are more broadly applicable; extensions to more complex models are a subject for future research.\footnote{Other work on partial identification in contexts with interval data include \citeasnoun{Magnac2008}, who focus on cases with binary dependent variables. For this case, they show that bounds are tighter under known distributions and reduce to points under the uniform distribution. \citeasnoun{Bontemps2012} focus primarily on cases where the $y$ variable is interval-censored. Our focus is on continuous dependent variables with interval-censored conditioning variables.}
Here, we briefly describe the two applications that we will focus on below.
Application 1: Mortality as a Function of Education
In the first application, we resolve a long-standing problem in the estimation of mortality as a function of education. Researchers have noted recent increases in the mortality of less-educated individuals in the U.S. Meara2008,Cutler2010,Cutler2011,Olshansky2012,Case2015,Case2017. For example, mortality among women aged 50--54 with high school education or less (LEHS) has risen from 459 deaths per 100,000 people in 1992 to 587 deaths in 2015. A known concern with these estimates is that rising education levels over time, particularly among women, make these numbers difficult to interpret.
Figure (ref) shows mortality for 50--54 year old U.S. women as a function of the median education rank in each of three educational categories, illustrating the simultaneous changes in mortality and in the distribution of education. Women with a high school degree or less represented 64% of women in 1992 and only 39% of women in 2015. If mortality is a decreasing function of the latent education rank, then the increasing negative selection of LEHS women could explain some or all of the mortality change for this group, even if the underlying mortality-education rank relationship is unchanged. Whether and how to adjust for these compositional changes is an important debate in the mortality literature. Some studies have argued that the bias is close to zero, while others have suggested that it may explain all of the recent mortality increases.\footnote{Recent high profile work by Case and Deaton (2015, 2017) focuses on unadjusted estimates for non-Hispanic whites with high school education or less (rather than dropouts), arguing that their average school completion has not substantially changed over the sample period they study. For our sample period (1992 to the present, all races) LEHS men have gone from 54% to 44% of the population in 2015 and LEHS women have gone from 64% of the population to 39%. Changes are even larger for other age groups over other time periods. \citeasnoun{Dowd2014} and \citeasnoun{Currie2018} argue that the bias may be so large that estimates of mortality change among LEHS people are effectively uninformative.} Estimates of mortality within fixed education quantiles would solve the problem, but there is no established method to generate quantiles when intervals in the data do not correspond to quantile boundaries.\footnote{Mortality data typically report education in a small number of coarse categories. Most studies on mortality and education use only two or three categories of education.}\ensuremath{^{,}}\footnote{\citeasnoun{Bound2015} generate quantile point estimates, but only under the implicit assumption that the latent mortality-education gradient has zero slope within each education bin. This is a strong assumption given the important gradient across education bins. The partial identification approach that we propose lets us avoid making strong assumptions about censored data.}
To make progress on this problem, we make two assumptions. First, we assume that the observed education rank represents a latent, continuous rank that is observed only in coarse intervals, a common assumption in this literature Goldring2016.\footnote{The latent variable can be interpreted as the total net benefit of pursuing a given quantity of education. Those individuals at the high end of a latent education rank bin are the ones who would move to a higher education bin if their net benefit of education marginally increased.} Second, we assume that mortality is decreasing in latent educational rank. This assumption holds across bins for every year between 1992 and 2015 for both men and women, and across every income ventile Chetty2016b. Under just these assumptions, our method can generate bounds on the expectation of mortality at any rank or in any rank interval or quantile.
We focus on women age 50-54, because their increasing education over time means the selection bias for this group may be large. We bound the mortality rate of the bottom 64% of the education distribution -- a fixed share of the population representing LEHS in 1992. The bounds are tight and informative. The mortality increase for women from 1992--2015 in this part of the education distribution is between 29 and 38 additional deaths per 100,000. The unadjusted estimate (which compares the bottom 64% in 1992 to the bottom 39% in 2015) suggests an increase of 128 deaths, more than three times higher than the upper bound from our calculation. For some population groups, the mortality increases noted in the recent literature are sustained when we use our method to study constant rank groups, while for others, the unadjusted estimates are substantially biased or have the incorrect sign.
This application focuses on mortality, but similar compositional issues arise in any context where the researcher is interested in changes in the relationship between education and some outcome variable over time. For example, our method could resolve bias due to changing composition in studies on education gradients in birth outcomes, marriage patterns or disability Cutler2010a,Aizer2014,Bertrand2016.\footnote{In a context where education is strictly considered as an input to the production function (such as estimating the returns to education), unadjusted estimates may be preferred. But if education and the outcome are correlated with any omitted variable, then adjusting for population share and education rank will be a useful exercise. We take no stand on the health production function or whether the mortality-education relationship should be treated causally.}
Application 2: Intergenerational Educational Mobility
The methods presented here can also resolve several challenges in the estimation of intergenerational mobility. The object of interest in many studies of intergenerational mobility is the CEF of child education given parent education, in part because data on educational attainment are widely available and may be less subject to measurement error than parent income data Black2003,Guell2013.
The coarse binning of education data poses a key problem in this context. Many mobility measures require observation of the child CEF at a specific point in the parent rank distribution. Absolute upward mobility, for instance, is defined by \citeasnoun{Chetty2014b} as the expected outcome of a child who is born to a family at the 25\ensuremath{^{th}} percentile of the parent rank distribution. Binned education data make this challenging to estimate. In older generations in India, for example, over 50% of parents report having less than two years of education, the lowest recorded category in many datasets. In such a context, the 25\ensuremath{^{th}} percentile parent is not directly observed, so absolute upward mobility can at best be partially identified.\footnote{We assume that absolute upward mobility is measured in terms of the continuous latent educational rank rather than the directly observed rank bin. Treating the bin mean as the true value in the rank bin (the approach of \citeasnoun{Bound2015} to mortality) has the undesirable property that more granular measures of education will lead to lower measures of absolute upward mobility.} Expected child outcomes at constant parent ranks are also required for meaningful cross-group mobility comparisons Hertz2005.\footnote{The rank-rank gradient and other linear estimators of the parent-child outcome function are not informative about subgroup mobility, because they compare children of low-ranked parents with children of high-ranked parents from the same subgroup, which can be misleading Aaronson2008.} With education rank boundaries that change over time, subgroup educational mobility estimates are difficult to compare over time.
The CEF bounds proposed above are a direct solution to these problems, as they bound the expected outcome of a child born at arbitrary points or intervals in the parent rank distribution. We propose a new measure of mobility, upward interval mobility, which is the mean value of the child CEF in the bottom half of the parent rank distribution. Applying our method to Indian data, we show that conventional mobility measures are biased or uninformative about the mobility of older cohorts once we account for interval censoring, but upward interval mobility can be tightly bounded.\footnote{Upward interval mobility has very similar policy relevance to absolute upward mobility. Absolute upward mobility measures the expected outcome of the median child in the bottom half of the parent distribution, whereas upward interval mobility measures the mean.}
The interval problem for educational mobility is most severe in developing countries, but is important in other contexts as well. In wealthier countries, it is common for a large share of the population to be in a topcoded education bin.\footnote{In one mobility study from Sweden, for example, 40% of adoptive parents were topcoded with 15 or more years of education Bjorklund2006. Studies on the persistence of occupation across generations also frequently use a small number of categories and face a similar challenge when the occupational structure changes significantly over time, as it has with farm work in the United States. See, for example, \citeasnoun{Long2013}, \citeasnoun{Xie2013} and \citeasnoun{Guest1989}.} Internationally comparable censuses also frequently report education in as few as four categories; our method is thus particularly relevant for cross-country comparison.
In the next section, we describe the setup, prove the new bounds and present the numerical solution framework. In Section (ref), we explore properties of the bounds in a simulation. Sections (ref) and (ref) present the applications to the measurement of mortality and of intergenerational mobility in more detail. Section (ref) concludes. Stata and Matlab code to implement all methods in the paper are available on the corresponding author's web site.\footnote{Code can be downloaded at https://github.com/paulnov/anr-bounds.}
This section describes the main contribution of the paper. We calculate analytical and numerical bounds on a CEF where the conditioning variable is interval censored but has a known distribution. The bounds are sharp and depend either on the assumption of a weakly monotonic CEF or on the assumption that the CEF has limited curvature. The method can also bound any statistic that can be derived from the CEF, such as the mean over an arbitrary interval, or the best linear approximator to the CEF.
We describe the method by working through an example motivated by Figure (ref), which plots total mortality against education, where education is only observed in one of three education bins: (i) less than or equal to high school; (ii) some college; or (iii) bachelor's degree or higher.\footnote{Points are plotted at the midpoint of the education rank bins.} We focus on women aged 50--54, because (i) this age group has been highlighted in other recent research, and (ii) the change in education for this group has been large over the sample period. We wish to estimate some statistic that describes mortality in 1992 and in 2015 for a group of people occupying the same set of education ranks in the population. This is challenging because the rank bin boundaries change between 1992 and 2015. In 1992, 64% of women had less than or equal to a high school education, while in 2015, this number was 39%.
Our approach is to estimate the conditional expectation function of mortality given education in each year, which would allow us to partially identify mortality at any rank in any year. We implicitly assume that there exists a latent, continuous education rank that we only observe in discrete intervals. This section focuses on the problem of identifying a CEF given interval data, and Section (ref) explores the findings on mortality in more detail.
Figure (ref) depicts the setup for 2015. The points show mortality at the midpoints of three education bins and the vertical lines show the rank bin boundaries. The lines plot two (of many) possible nonparametric CEFs, each of which fit the sample means with zero error. These two functions have the same mean in each bin, even if they do not cross the mean at the bin midpoint.\footnote{A naive polynomial fit to the midpoints in the graph would be a biased fit to the data because of Jensen's Inequality.} These are the functions we aim to bound. We begin with the assumption that mortality is weakly decreasing in latent rank, and then show how a curvature constraint can supplement or substitute for this assumption.
Define the outcome as $y$ and the conditioning variable as $x$; the conditional expectation function is $Y(x)=E(y|x)$. Let the function $Y(x)$ be defined on $x \in [0,100]$, and assume $Y(x)$ is integrable. We also assume throughout that $\underline{Y} \leq Y(x) \leq \overline{Y}$, that is, the function is bounded absolutely.\footnote{In most applications, parameters of interest are likely to have upper and lower bounds either in theory or in practice. Loosening the absolute upper and lower bound restriction would result in wider bounds for the CEF in the bottom or top intervals, but informative inference is still possible even in these outer bins. In the case of mortality, we will impose that the upper bound is a mortality rate of 100%.}
With interval data, we do not observe $x$ directly, but only that it lies in one of $K$ bins. Let $f_k(x)$ be the probability density function of $x$ in bin $k$. Define the expected outcome in the $k^{th}$ bin as $$r_k = E \left(y|x \in [x_k,x_{k+1}] \right) = \int_{x_k}^{x_{k+1}} Y(x)f_k(x)dx,$$ where $x_k$ and $x_{k+1}$ define the bin boundaries of bin $k$. This expression holds due to the law of iterated expectations. The limits of the conditioning variable are assumed to be known, and are denoted by $x_1$ and $x_{K+1}$. Further define the expected outcomes in the intervals directly above and below the intervals of interest as $r_{k+1}=E \left(y|x \in [x_{k+1},x_{k+2}] \right)$ and $r_{k-1}=E\left(y|x \in [x_{k-1},x_{k}] \right)$, if they exist. Define $r_0 = \underline{Y}$ and $r_{K+1} = \overline{Y}$. The sample analog to $r_k$ is the observed mean outcome in bin $k$, which we denote $\overline{r}_k$.
Sharp bounds on $E(y|x)$ given interval measurement of $x$ are derived by \citeasnoun{Manski2002}, when the distribution of $x$ is unknown. The essential structural assumption that constrains the CEF is Monotonicity (M):
Note that we apply this assumption to the survival rate, which is one minus the mortality rate; however, our graphs show the mortality rate which is the parameter of interest. The CEF in the monotonic graphs is thus monotonically decreasing.\footnote{Mortality is decreasing in educational attainment for every group and time period in the CDC data; it is also a monotonically decreasing function of income Chetty2016b.} \citeasnoun{Manski2002} also introduce the following Interval (I) and Mean Independence (MI) assumptions. For $x$ which appears in the data as lying in bin $k$,
Assumption $I$ states that the rank of all people who report education ranks in category $k$ are actually in bin $k$. Assumption $MI$ states that censored observations are not different from uncensored observations. These always hold in our context because all of the data are interval censored.
If all observations of $x$ are interval censored, the \citeasnoun{Manski2002} bounds are:
The value of the CEF in each bin is bounded by the means in the previous and next bins.
We can improve upon these bounds if the distribution of $x$ is known. In some cases, as with ranks, the distribution is given by the definition of the variable. In other cases, conventional distributions are frequently assumed (such as lognormal or Pareto for income data). Alternatively, data could be transformed into a known distribution, for example, by transforming the conditioning variable into ranks. We first show bounds under the assumption that $x$ has a uniform distribution because the analytical results are particularly parsimonious, but we derive all of our results under a general known distribution. We therefore consider the following assumption (U):
where $U$ is the uniform distribution.
If $x$ is uniformly distributed, we know that:
We derive the following proposition.
The proposition is obtained from the insight that the value of $E(y|x=i)$ at a point $i$ in bin $k$ (below the midpoint) will only be minimized if all points in bin $k$ to the left of $i$ have the same value. Since all points to the right of $i$ are constrained by the outcome value in the subsequent bin $k+1$, $E(y|x=i)$ will need to rise above the \citeasnoun{Manski2002} lower bound as $i$ increases, in order to meet the bin mean. Intuitively, consider the point $E(y|x=x_{k+1}-\varepsilon)$. In order for this point to take on a value below the bin mean $r_k$, it needs to be the case that virtually all of the density in bin $k$ lies between $x_{k+1}-\varepsilon$ and $x_{k+1}$. This is ruled out by the uniform distribution, and indeed by most distributions; for many distributions, therefore, the \citeasnoun{Manski2002} bounds are too conservative. We prove the proposition and provide additional intuition in Appendix (ref).
We generalize the proposition to obtain the following result for an arbitrary known distribution of $x$:
A proof of the proposition is in Appendix B.
Figure (ref) compares \citeasnoun{Manski2002} bounds to those obtained under the additional assumption of uniformity, using the mortality data. The new bounds are a significant improvement, especially where the data are particularly coarse and near the bin boundaries. For example, without using information on the distribution type, one could not reject that mortality for people in the first bin is 100,000 per 100,000 until just before the first bin boundary. The improvements in the other bins are less extreme but still substantial.
In addition to bounding the value of $Y=E(y|x)$ at any given point, we can also bound many functions of the CEF, which we represent in the form $M(Y)$. One function of interest is the slope of the best linear approximation to the CEF; this is difficult to bound analytically, but we bound this numerically in Section (ref).
Here, we highlight a function that describes the average value of the CEF over an arbitrary interval of the conditioning space, or $\mu_a^b = E(y|x \in [a, b])$. This function has several desirable properties. First, it can be bounded analytically. Second, it is frequently bounded more tightly than $E(y|x)$. Third, it has a similar interpretation to $E(y|x)$ and is thus likely to be policy-relevant. We show in Sections (ref) and (ref) that for our applications, $\mu_a^b$ can be bounded considerably more tightly than $E(y|x)$.
Let $f(x)$ represent the probability density function of $x$. Define $\mu_a^b$ as
We now state analytical bounds on $\mu_a^b$ given uniformity. Let $Y_x^{max}$ be the analytical upper bound on $E(y \vert x)$, given by Proposition (ref). Let $Y_x^{min}$ be the analytical lower bound on $E(y \vert x)$. The following proposition defines sharp bounds on $\mu_a^b$ under the assumption that $x$ is uniformly distributed:
We prove this proposition under uniformity and under an arbitrary known conditioning distribution in Appendix (ref).
We note two special cases. First, if $a=b$, then $\mu_a^b = E(y|x=a)$. Second, if $a$ and $b$ correspond exactly to bin boundaries, then the bounds on $\mu_a^b$ collapse to a point: in this case, $\mu_a^b$ is just a weighted average of the bin means between $a$ and $b$.
In fact, $\mu_a^b$ can be very tightly bounded whenever $a$ and $b$ are close to bin boundaries. For intuition, consider the following examples. If $\delta \in [a,b]$, $\mu_a^b$ can be written as a weighted mean of the two subintervals $\frac{\delta - a}{b-a} \mu_a^{\delta} + \frac{b - \delta}{b-a} \mu_\delta^b$.\footnote{The weights on each subcomponent here assume that $x$ is uniformly distributed. A different distribution would use different weights.} If $\mu_a^{\delta}$ is known (because there are bin boundaries at $a$ and $\delta$), then any uncertainty about the value of the CEF in the range $[a, \delta]$ is not consequential for the bounds on $\mu_a^b$. If $b$ is close to $\delta$, the weight on the unknown value $\mu_\delta^b$ is very small, and $\mu_a^b$ can be tightly bounded. Similarly, if instead $\mu_a^b$ is known, and $b$ is again close to $\delta$, then $\mu_a^\delta$ can be tightly estimated even if $\mu_\delta^b$ has wide bounds.
Bounds on other functions of the CEF may be difficult to calculate analytically, but can be defined as the set of solutions to a pair of minimization and maximization problems that take the following structure. We write the conditional expectation function in the form $Y(x) = s(x,\gamma)$, where $\gamma$ is a finite-dimensional vector that lies in parameter space $G$ and serves to parameterize the CEF through the function $s$. For example, we could estimate the parameters of a linear approximation to the CEF by defining $s(x,\gamma)=\gamma_0+\gamma_1*x$. We can approximate an arbitrary nonparametric CEF by defining $\gamma$ as a vector of discrete values that give the value of the CEF in each of $N$ partitions; we take this approach in our numerical optimizations, setting $N$ to 100.\footnote{For example, $s(x,\gamma_{50})$ would represent $E(y|x \in [49,50])$.} Any statistic $m$ that is a single-valued function of the CEF, such as the average value of the CEF in an interval $(\mu_a^b)$, or the slope of the best fit line to the CEF, can be defined as $m(\gamma)=M(s(x,\gamma))$.
Let $f(x)$ again represent the probability distribution of $x$. Define $\Gamma$ as the set of parameterizations of the CEF that obey monotonicity and minimize mean squared error with respect to the observed interval data:
Decomposing this expression, $\frac{1}{\int_{x_k}^{x_{k+1}} f(x)dx } \int_{x_k}^{x_{k+1}} s(x,g) f(x) dx$ is the mean value of $s(x,g)$ in bin $k$, and $\int_{x_k}^{x_{k+1}} f(x) dx$ is the width of bin $k$. The minimand is thus a bin-weighted MSE.\footnote{While we choose to use a weighted mean squared error penalty, in principle $\Gamma$ could use other penalties.} Recall that for the rank distribution, $x_1=0$ and $x_{K+1}=100$.
The bounds on $m(\gamma)$ are therefore:
For example, bounds on the best linear approximation to the CEF can be defined by the following process. First, consider the set of all CEFs that satisfy monotonicity and minimize mean-squared error with respect to the observed bin means.\footnote{In many cases, and in all of our applications, there will exist many such CEFs that exactly match the observed data and the minimum mean-squared error will be zero.} Next, compute the slope of the best linear approximation to each CEF. The largest and smallest slope constitute $m^{min}$ and $m^{max}$. Stata code to generate bounds on the CEF and on $\mu_a^b$, and Matlab code to run these numerical optimizations for more complex functions (as well as with the curvature constraints described below) are posted on the corresponding author's web site.
CEF Bounds Under Constrained Curvature
The candidate CEFs that underlie the bounds in Proposition (ref) are step functions with substantial discontinuities. If such functions are implausible descriptions of the data, then the researcher may wish to impose an additional constraint on the curvature of the CEF, which will generate tighter bounds. For example, examination of the mortality-income relationship (which can be estimated at each of 100 income ranks, displayed in Figure (ref)) suggests no such discontinuities.\footnote{More complex structural restrictions can also be imposed. For example, the CEF might be continuous within education bins, but there could be large discontinuities due to sheepskin effects at the education bin boundaries Hungerford1987.} Alternately, in a context where continuity has a strong theoretical underpinning but monotonicity does not, a curvature constraint can substitute for a monotonicity constraint and in many cases deliver useful bounds.
We consider a curvature restriction with the following structure:
This is analogous to imposing that the first derivative is Lipshitz.\footnote{Let $X, Y$ be metric spaces with metrics $d_X, d_Y$ respectively. The function $f:X \to Y$ is Lipschitz continuous if there exists $K \geq 0$ such that for all $x_1,x_2 \in X$, $$d_Y(f(x_1),f(x_2)) \leq K d_X(x_1,x_2).$$} Depending on the value of $\overline{C}$, this constraint may or may not bind.
The most restrictive curvature constraint, $\overline{C}=0$, is analogous to the assumption that the CEF is linear. Note that the default practice in many studies of mortality is to estimate the best linear approximation to the CEF of mortality given education (e.g., \citeasnoun{Cutler2011} and \citeasnoun{Goldring2016}). In the study of intergenerational mobility (Section (ref)), the best linear approximation to the parent-child CEF is the canonical estimator. A moderate curvature constraint is therefore a less restrictive assumption than the approach in many studies. We discuss the choice of curvature restriction below.
In the rest of this section, we show results under a range of curvature restrictions to shed light on how these additional assumptions affect bounds in an empirical application. In our applications in Sections (ref) and (ref), we show all results under the most conservative approach of $\overline{C}=\infty$.
This section describes a method to numerically solve the constrained optimization problem suggested by Equations (ref) and (ref). We take a nonparametric approach for generality: explicitly parameterizing an unknown CEF with limited data is unsatisfying and could yield inaccurate results if the interval censoring conceals a non-linear within-bin CEF. In the context of mortality (and mobility, Section (ref)), many CEFs of interest do not appear to obey a familiar parametric form (see Figures (ref) and (ref)).
To make the problem numerically tractable, we solve the discrete problem of identifying the feasible mean value taken by $E(y|x)$ in each of $N$ discrete partitions of $x$. We thus assume $E(y|x) = s(x,\gamma)$, where $\gamma$ is a vector that defines the mean value of the CEF in each of the $N$ partitions. We use $N=100$ in our analysis, corresponding to integer rank bins, but other values may be useful depending on the application. Given continuity in the latent function, the discretized CEF will be a very close approximation of the continuous CEF; in our applications, increasing the value of $N$ increases computation time but does not change any of our results.
We solve the problem through a two-step process. Define a $N$-valued vector $\hat{\gamma}$ as a candidate CEF. First, we calculate the minimum MSE from the constrained optimization problem given by Equation (ref). We then run a second pair of constrained optimization problems that respectively minimize and maximize the value of $m(\hat{\gamma})$, with the additional constraint that the MSE is equal to the value obtained in the first step, denoted $\underbar{MSE}$. Equation (ref) shows the second stage setup to calculate the lower bound on $m(\hat{\gamma})$. Note that this particular setup is specific to the uniform rank distribution, but setups with other distributions would be similar.
$X_k$ is the set of discrete values of $x$ between $x_{k}$ and $x_{k+1}$ and $\Vert X_k \Vert$ is the width of bin $k$. The complementary maximization problem obtains the upper bound on $m(\hat{\gamma})$.
Note that setting $m(\gamma) = \gamma_x$ (the x\ensuremath{^{th}} element of $\gamma$) obtains bounds on the value of the CEF at point $x$. Calculating this for all ranks $x$ from 1 to 100 generates analogous bounds to those derived in proposition (ref), but satisfying the additional curvature constraint. Similarly $m(\gamma) = \frac{1}{b-a} \sum_{x=a}^{b}\gamma_x$ obtains bounds on $\mu_a^b$.
In this section, we demonstrate the bounding method using data from mortality in the United States, continuing with the mortality of 50--54 year-old women in 2015. We focus here on the properties of the bounds under different assumptions. We explore mortality change in more detail in Section (ref).
Panel A of Figure (ref) graphs the analytical upper and lower bounds on $E(y|x)$ at each value of $x$ under just the assumption of monotonicity. These bounds do not reflect statistical uncertainty but uncertainty about the CEF in the unobserved parts of the latent rank distribution.\footnote{We do not present standard errors because we are working with the universe of deaths in a large country and statistical imprecision is very small in this context. We discuss and present bootstrap confidence sets in Section (ref) where statistical imprecision is more important.}
We next consider a curvature-constrained CEF.\footnote{With neither the monotonicity nor the curvature constraint, the CEF cannot be bounded except by the maximum possible value of the variable of interest.} The mortality-education data are not in themselves informative regarding which curvature restriction to choose. To identify a conservative curvature constraint, we examine the curvature of a closely related conditional expectation function that is not interval censored: the CEF of mortality given income rank. We show this CEF in Figure (ref), using data from \citeasnoun{Chetty2016b}. Using a spline approximation to income rank data for 52-year-old women in 2015, we calculate a maximum $\overline{C}$ of 1.6; we use a constraint approximately twice as high as a conservative starting point. Panel B of Figure (ref) shows the bounds obtained under curvature constraints of 2, 3 and 5, but without the assumption of monotonicity. Relative to those under monotonicity, the curvature-constrained bounds are less informative at the tails of the distribution, and more informative close to the bin midpoints.
In Panel C, we impose the monotonicity and curvature constraints simultaneously. Panel D shows the limit case with $\overline{C}=0$; the CEF in this figure is identical to the predicted values from a regression of mortality on median education rank. Note that while stricter curvature restrictions can tighten the bounds, this may come at the expense of ruling out a plausible CEF, even if the MSE remains zero. In Figure (ref), only Panel D has a non-zero MSE.
Table (ref) presents estimates of $p_x=E(y|x)$ and $\mu_a^b=E(y|x \in (a, b))$ for women ages 50--54 in 2015, for various values of $x$, $a$ and $b$, under different constraints. We first highlight the statistics $p_{32}$ and $\mu_0^{64}$. In 1992, 64% of women had high school education or less, and thus occupied the bottom rank bin in the education distribution. $p_{32}$ and $\mu_0^{64}$ respectively describe the median and mean mortality of the comparably ranked group of women in 2015. These statistics give us mortality estimates for constant ranks in the education distribution, even though the distribution of education levels is changing over time.
We draw attention to two features of the table. First, the interval mean estimates ($\mu_a^b$) are in most cases considerably more tightly bounded than estimates of the CEF value at the midpoint of the interval ($p_x$). $\mu_0^{64}$ is nearly point identified in 1992 because $0$ and $64$ are very close to bin boundaries in 1992, and it is tightly bounded in 2015 as well, regardless of the constraint set.\footnote{We have used integer approximations to these parameters for convenience; if we used the average mortality for the precise proportion of women with less than or equal to a high school degree ($\mu_0^{63.658}$), then the parameter would be precisely point identified.} These two statistics are both useful summaries of mortality among the less educated, but $\mu_0^{64}$ is estimated with at least 22 times more precision than $p_{32}$. Similarly, $\mu_0^{39}$ is effectively point estimated in 2015, where 39% of women had attained high school or less. The advantage of $\mu_a^b$ over $p_x$ (where $x=\frac{a+b}{2}$) is greatest when $a$ and $b$ are close to boundaries in the data.
Second, $\mu_0^{39}$ and $\mu_0^{64}$ are very robust to different bounding assumptions. Inference is more difficult on a parameter like $\mu_0^{20}$ with boundaries far from any in the data, and the width of the bounds depends strongly on the assumptions being made. Mortality in the bottom 20% of the education distribution may be of policy interest, but our method shows that it cannot be precisely estimated with these data.
Because it is a frequently estimated parameter, in Column 5 we show the predicted values from the best linear approximation to the mortality-education CEF. This parameter is point estimated, but implicitly assumes away large increases in mortality at the bottom of the distribution, increases that are consistent with the data and in fact suggested by Figure (ref). In contrast, our method allows researchers to generate consistent bounds on mortality across the education distribution under considerably less restrictive assumptions.
In this section, we validate our method in a simulation by taking data from the fully supported U.S. mortality-income CEF Chetty2016b, interval censoring that data, and then recovering bounds on the true CEF from the interval censored data. The exercise shows that our approach works in practice. It also illustrates that studying partially identified bounds permits the researcher to recover important features of the CEF that she might miss if she attempted simply to fit a parametric form to the observed bin means. We use data on mortality by income percentile, gender, age and year Chetty2016b. We focus on women aged 52 in 2014, the group most comparable what we have examined so far.
First, we estimate the true CEF from the mortality-income data by fitting a cubic spline with four knots to the data, the same spline used to obtain an estimate of $\overline{C}$. We plot this in Appendix Figure (ref).\footnote{We use a spline approximation rather than the raw data because the variation across neighboring rank bins is most likely idiosyncratic given the small number of deaths in an age bin defined by a single year. By using information from neighboring points, the spline is a better estimate of mortality risk than the individual rank bin means.}
Next, we simulate interval censoring by obtaining the mean of the true CEF within income rank bins that cover the same ranks as the education bins observed in our 2015 mortality-education data. In this simulation, there are 39% of people in the bottom bin, 29% of people in the middle bin, and 33% of people in the top bin.\footnote{We round to the nearest integer, since we only observe integer percentiles in the data from \citeasnoun{Chetty2016b}.} After interval censoring, we have a dataset with average mortality in each of three bins, comparable to the data from Section (ref). We compute bounds on the CEF using only the binned data.
Panels A--D of Figure (ref) present CEF bounds generated from the binned data, under monotonicity and curvature limits that vary from $\infty$ (unconstrained) to $1$. The dashed lines show the underlying data. The solid circles show the constructed bin means of the censored data; these are the only data that we use for the optimization. The solid lines show the upper and lower envelopes that we calculate for the nonparametric CEF.
The suggested curvature constraint ($\overline{C}=3$) yields bounds that contain the true CEF at every point; but when we impose $\overline{C}=1$, the constraint is excessive and the bounds do not contain the true CEF. The true CEF is not always centered within the bounds; from ranks 25 to 40, the true CEF is near the bottom bound, and from ranks 90 to 100, it is nearer the upper bound.
The exercise also illustrates that assuming a parametric form for the underlying CEF can yield misleading results. A quadratic or linear fit to the data would fail to identify the convexity at the bottom of the distribution. The strength of our method is that it makes transparent how the structural assumptions affect the CEF bounds.
Table (ref) shows bounds on a range of statistics of interest under different curvatures, as well as the true estimate. We highlight three results. First, the interval mean measures ($\mu_a^b$) generate tighter bounds than the CEF values $p_x$, with no greater propensity for error. Second, $\overline{C}$ is consequential for $p_x$, but considerably less important for $\mu_a^b$. Third, the linear estimates generated with $\overline{C} = 0$ are biased by as much as 25% relative to the true estimates, and sometimes produce estimates outside the bounds even of CEFs with unconstrained curvature.
In this section, we apply our methodology to study changes in U.S. mortality for individuals at constant ranks in the education distribution. Many researchers have noted that mortality is rising for individuals in less educated groups; however, the changing composition of these groups over time has made this finding difficult to interpret. For example, women with a high school education or less (LEHS) represented the least educated 64% of the population in 1992, and the least educated 39% of the population in 2015. Those with LEHS are thus more negatively selected in 2015 than they were in 1992; the changing size and composition of this group may account for at least some of the mortality increase. This bias has been frequently noted in the literature, but different authors have reached widely different conclusions regarding its size and importance.\footnote{\citeasnoun{Cutler2011} adjust for compositional shifts by predicting propensity to attend college using region, marital status and income, and then using this propensity as a conditioning variable. They argue that compositional shifts are not important for mortality changes from the 1970s to the 1990s. This approach is limited by the extent to which these variables can predict education, and in many cases (e.g. with vital statistics data), these additional variables are unavailable. \citeasnoun{Case2015} and \citeasnoun{Case2017} argue that changes in the proportion of middle-aged whites with LEHS from the 1990s to the present are too small to influence mortality rates. In contrast, \citeasnoun{Dowd2014} and \citeasnoun{Bound2015} perform analytical exercises that suggest that compositional shifts can explain most or all of recent mortality changes. \citeasnoun{Bound2015} estimate mortality for the bottom quartile of the education distribution, implicitly assuming that mortality is constant within each interval-censored mortality rank bin. \citeasnoun{Currie2018} suggests that studying mortality for the least educated is entirely misleading because of the shrinking size of this group. \citeasnoun{Goldring2016} derive a one-tailed test for changes in the mortality-education gradient, but they do do not calculate the bias in existing mortality estimates or estimate mortality in constant rank bins. Our method requires no additional covariates, and bounds mortality at an arbitrary education rank under only the assumption of monotonicity.} In this section, we use the methods above to bound the value of the mortality CEF at constant education ranks, even if these ranks are not directly observed in the data. We can then study a group with constant size and education rank over time, and thus study any subset of the education rank distribution without bias from changing education levels over time.
Mortality by education records come from the U.S. Center for Disease Control's WONDER database and total population by age, gender and education come from the Current Population Survey, as in \citeasnoun{Case2017}. Additional details on data construction are available in Appendix (ref).
As above, we assume that the observed mortality data describe a monotonic relationship between mortality and latent education rank, the latter of which is observed only in coarse bins. Results are virtually identical if we constrain curvature using the parameter suggested in Section (ref) and forgo the monotonicity constraint. We focus in this section on women aged 50--54, because this is a group whose education composition has shifted substantially over time.\footnote{We use 5-year bins for ages rather than larger bins to ensure that the average age in the bin does not change over time Gelman2016.}
Panel A of Figure (ref) plots mean total mortality for women age 50-54 in each education group in 1992 and in 2015, along with analytical bounds on CEFs with unconstrained curvature. The bounds are largely overlapping across the entire education distribution, and too wide to infer very much about changes in mortality. In Panel B, we restrict curvature to approximately twice the maximum curvature from the income-mortality data, as discussed in Section (ref). Panel B identifies a clear decline in mortality at the top of the education distribution, but the bounds remain minimally informative at the bottom.
The interval mean measures ($\mu_a^b$) are more informative. We focus on mean mortality in the bottom 64%, denoted by $\mu_0^{64}$.\footnote{Specifically, we calculate $\mu_0^{63.658}$ for analytical monotonic bounds and $\mu_0^{64}$ for numerical curvature constrained bounds.} This measure describes mortality for the set of women who occupied positions in the rank distribution that would give them high school education or less in 1992. In 2015, the bottom 64% includes all women with high school or less, and some women with some college education, but none with bachelor's degrees or higher. For men, we focus on the bottom 54%, which is the population share with high school or less in 1992; by 2015, 44% of men have LEHS, so the bottom 54% again includes some men with two-year college degrees. As in \citeasnoun{Chetty2016b}, we rank men and women against members of their own gender, estimating mortality for a given percentile group of men or women; that is, the least educated group can be interpreted as the “the 64% of least educated women,” rather than “women in the bottom 64% of the population education distribution.” We chose own-gender reference points because women's and men's labor market opportunities and choices are often different and because women and men often share households and incomes, making population ranks misleading. However, alternate choices could be considered and estimated with the same method.
Panel A of Figure (ref) shows bounds on total mortality for women aged 50-54 in the bottom 64%. Mortality in 1992 can be point estimated, because the 0-64 rank bin interval is exactly observed in the data. As education levels diverge from those in 1992, the bounds progressively widen. The “x” markers in the figure plot the unadjusted estimates of mortality among women with less than or equal to high school education; these mortality estimates describe a group occupying a shrinking and more negatively selected share of the population over time. The unadjusted estimates, which are the object of study in most earlier work on the mortality-education relationship, significantly overstate mortality increases relative to the constant rank group. The upper bound on mortality gain for the bottom 64% is 8.5%, compared to the unadjusted estimate of 28%. Panel B shows the same figure for men. The unadjusted estimates are closer to the bounds here because men have gained less education than women over this period. We can bound the mortality change for men in the interval $[-7.1\%, +0.3\%]$, compared with the unadjusted estimate of $+1.2\%$.\footnote{This result is not directly comparable to Case and Deaton (2015, 2017), who focus on white men and women, whose unadjusted mortality is rising more substantially among the less educated. Estimating separate bounds for different racial groups requires additional assumptions about the relative positions of these groups in the unobserved part of the latent education distribution, and we leave this exercise for future work. The increases in mortality for less educated women that we identify here are still a cause for concern even if they are lower than previous estimates.} Panels C and D present analogous results for combined deaths from suicide, poisoning and liver disease, described by \citeasnoun{Case2017} as “deaths of despair.” The unadjusted mortality estimates continue to overstate the constant rank mortality changes, but the difference is small here because (i) deaths of despair have increased substantially among all groups; and (ii) the education gradient in deaths of despair was small in 1992. Appendix Figure (ref) shows the same plots, but removes the monotonicity assumption and instead imposes the curvature restriction of $\overline{C}=3$ suggested above. The plots are highly similar; as discussed in Section (ref), when $a$ and $b$ are close to bin boundaries, the bounds on $\mu_a^b$ are very robust to alternate bounding assumptions.
Table (ref) shows unadjusted and constant-rank estimates of women's mortality changes from 1992-2015 for age groups from 20 to 69, for all education categories. We fix education rank bins based on the 1992 rank bin divisions; results are very similar if we fix estimates at the 2015 boundaries. The unadjusted estimates systematically overstate mortality increases for all groups, because the mean rank in each group has declined over this period.
The extent of the bias on the naive estimates is increasing in the magnitude of the mortality-education gradient, and in the magnitude of the shift in bin boundaries. Given the significant variation across age groups and genders, blanket assumptions about the existence or lack of bias in unadjusted mortality differences are therefore unlikely to be useful. Unadjusted estimates of men's mortality changes from 1992-2015 are close to the constant rank bounds, as are unadjusted estimates of deaths of despair for both men and women. For women's total mortality, however, the naive estimates overstate mortality increases in many cases by a factor of three or more, and in some cases they have the wrong sign.
The study of intergenerational mobility is another research context where the conditional expectation function of interest in many cases has an interval-censored conditioning variable.\footnote{For a review of intergenerational mobility, see \citeasnoun{Solon1999}, \citeasnoun{Hertz2008}, \citeasnoun{Corak2013a}, \citeasnoun{Black2011}, and \citeasnoun{Roemer2016}.} Studies of intergenerational mobility typically rely upon some measure of rank in the social hierarchy which can be observed for both parents and children Chetty2014c,Chetty2017. In many contexts, the only measure of social rank available for parents is their level of education. In richer countries, this arises for studies of mobility in eras that predate the availability of administrative income data.\footnote{See, for example, \citeasnoun{Black2003} and \citeasnoun{Guell2013}.} In developing countries, matched parent-child data are considerably more rare, and educational mobility is often the only feasible object of study.\footnote{See, for example, \citeasnoun{Wantchekon2015}, \citeasnoun{Hnatkovska2013} or \citeasnoun{Emran2015}.} Interval-censored parent education data is ubiquitous in studies of intergenerational educational mobility. Table (ref) reports the number of parent education bins used in a set of recent studies of intergenerational mobility from several rich and poor countries. Several of the studies observe education in fewer than ten bins, the population share in the bottom bin is often above 20%, and sometimes it is above 50%.\footnote{We specifically selected a set of studies where coarse data is likely to be an important factor. Note that internationally comparable censuses often report education in as few as four or five categories.}
Studies on educational mobility typically focus on linear estimators of the parent-child outcome relationship, such as the slope of the best linear approximator to the CEF of child education rank given parent education rank, i.e., the rank-rank gradient. This is a useful mobility statistic but it has two important limitations. First, it is not useful for cross-group comparison. The within-group rank-rank gradient measures children's outcomes against better off members of their own group; a subgroup can therefore have a lower gradient (suggesting more mobility) in spite of having worse outcomes than other groups at every point in the parent distribution.\footnote{An extreme example makes this clear. Suppose children in some population subgroup A all end up at the 10th percentile of the outcome distribution with certainty. The rank-rank gradient for this group would be zero (assuming some variation in parent outcomes), implying perfect mobility. But in fact the group would have virtually no upward mobility.} Second, the rank-rank gradient aggregates information about mobility at the top and at the bottom of the parent distribution; it is not directly informative about upward mobility in the bottom half of the distribution.
Because of these limitations, recent studies have focused on measures based on the value of the parent-child CEF at a point in the parent distribution, termed absolute mobility at percentile $x$ by \citeasnoun{Chetty2014c} and denoted $p_x$. For example, \citeasnoun{Chetty2014c} focus on $p_{25}$, which describes the expected outcome of the child born to the median family in the bottom half of the rank distribution. Unlike the rank-rank gradient, these measures are both informative about child outcomes at arbitrary points in the parent rank distribution and can be meaningfully compared across population subgroups. These measures are central to current research on mobility, but there is no established method for calculating such measures with education data, where any given percentile in the parent distribution lies within some larger bin. The problem is most stark when the bins are very large, so we focus our application on measuring intergenerational mobility in India, where over 50% of older generation parents are in the bottom education bin ($<2$ years of education). Appendix Table (ref) shows the complete education transition matrices for decadal birth cohorts from 1950 to 1989. A tempting but misleading approach would be to simply assume that the expected child outcome is exactly the same at all ranks within a given rank bin. In this case, the value of $p_{25}$ will change when education is measured with a different degree of granularity. In contrast, our bounds will widen when the granularity of the measure decreases, but they will contain the bounds generated from more granular data.
We take the following approach. We assume that the latent parent-child rank CEF can be described by an increasing monotonic function; this relationship is monotonic in virtually every country Dardanoni2012, as well across every rank bin in every year of our data on India. We use the CEF bounding method derived in Section (ref) to obtain bounds on (i) the value of the parent-child CEF at arbitrary parent percentiles ($p_x$); and (ii) the average value of the parent-child CEF across arbitrary percentile ranges of the parent distribution.\footnote{Our method is loosely related to \citeasnoun{Chetty2016}, who use a numerical procedure with similar constraints to bound absolute mobility at the 25th percentile, given just the marginal distributions of children's and parents' incomes and no information on the joint distribution. However, the substantive problem they solve is very different from ours.} We call the latter statistic, which corresponds to $\mu_a^b$ from Section (ref), interval mobility. We show below that, (i) the rank-rank gradient may be biased or uninformative when estimated from interval data; and (ii) interval mobility ($\mu_a^b$) can be bounded considerably more tightly than the other measures that we consider. We combine data from two sources, including administrative data on the education of every person in India in 2012, to obtain a representative sample of every father-son pair in India.\footnote{We are restricted to the study of fathers and sons because the data do not match daughters to parents or children to mothers when they do not live in the same household.} The details of data construction are described in Appendix (ref).
We observe education for both fathers and sons in seven categories.\footnote{The categories are (i) less than two years of education; (ii) at least two years but no primary; (iii) primary; (iv) middle school; (v) secondary; (vi) senior secondary; and (vii) post-secondary or higher.} Because sons' education levels are also reported categorically, we do not directly observe the expected child outcome in each parent education bin. In this section, we instead assign to children the midpoint of their rank bin. We show in Appendix (ref) that data on son wages (for which the rank distribution is uncensored) suggests that the midpoint is a very close approximation to the true expected rank, because the residual correlation of father education and son wages is very small once son's education is controlled for. Note that such an exercise is impossible for interval censoring of parent data, because no additional data on parents is available, as is typical in studies of intergenerational mobility.\footnote{Because parental education is often obtained by asking children, it is common to have data on many child outcomes, but only the education level of parents, as we do here.} Appendix (ref) also provides a method that generates bounds under joint censoring, which can be used in contexts where additional data on sons is not available. An alternate approach would be to estimate child rank directly using a socioeconomic measure for sons that can be observed continuously.
Panel A of Figure (ref) shows the raw data for cohorts born in the 1950s and in the 1980s. Each point plots the midpoint of a father education rank bin against the expected child rank in that bin. The vertical lines plot the boundary for the lowest education bin for each cohort, which corresponds to fathers with less than two years of education. In the 1950s birth cohort (solid line), this group represents 60% of the population; it represents 38% for the 1980s cohort (dashed line). The points in the figure suggest that the rank-rank CEF has not changed in the bottom half of the parent distribution over this period: the bottom point in the 1950s lies almost directly between the bottom two points in the 1980s. However, when we estimate the rank-rank gradient directly on these bin means, we find small but unambiguous mobility gains over this 30-year period. The graph makes clear that the decrease in the gradient is driven by changes in mobility in the top half of the distribution. Alternately, if we treat the data as uncensored, such that the expected child outcome is the same at all latent ranks within each parent bin, we would conclude that absolute upward mobility ($p_{25}$, or the expected child outcome at the 25\ensuremath{^{th}} parent percentile) has unambiguously fallen from the 1950s to the 1980s. Neither of these conclusions appears to represent the true change in mobility.\footnote{Note also that the CEF is evidently non-linear, so a naive nonlinear parametric fit to the bin midpoints would be biased due to Jensen's Inequality. It also assumes away concavity at the bottom of the distribution, which is observed in many other countries (see Appendix Figure (ref)).} We therefore turn to estimating bounds on the CEF in each period.
Panel B of Figure (ref) shows the bounds on the parent-child CEFs for these birth cohorts; we select a curvature constraint of 0.1, which is approximately 1.5 times the maximum curvature observed in uncensored parent-child income data from the United States, Denmark, Sweden and Norway.\footnote{We selected these countries because we were able to obtain precise uncensored parent-child income rank data for them from \citeasnoun{Chetty2014b}, \citeasnoun{Boserup2014} and \citeasnoun{Bratberg2017}. Graphs for the spline estimations used to calculate the curvature constraints are displayed in Appendix Figure (ref). Results are substantively similar under different curvature constraints.} The bounds on the CEF are widest at the bottom of the distribution where interval censoring is most severe, and are worse for the older generation with the larger bottom rank bin. The bounds in the bottom half of the distribution are consistent with both large positive and large negative changes in mobility, and thus uninformative. Absolute upward mobility can evidently not be bounded informatively.\footnote{Note that a more restrictive curvature constraint would narrow the bounds, but at the expense of imposing excessive structure that would rule out plausible CEFs, especially given the evident nonlinearity in the data.}
We can make meaningful progress by focusing on an interval-based measure such as $\mu_a^{b}=E(y|x \in [a, b])$. We call this measure interval mobility, and focus in particular on $\mu_0^{50}=E(y|x \in [0, 50])$, which we call upward interval mobility. This statistic is closely related to absolute upward mobility ($p_{25}$). The latter describes the outcome of the median child born to a parent in the bottom half of the parent distribution, whereas upward interval mobility describes the mean child outcome in the bottom half of the parent distribution. These measures are of similar economic importance, but we show here that upward interval mobility can be bounded tightly in contexts with severe interval censoring, while absolute upward mobility cannot.
Figure (ref) shows bounds on the three mobility statistics discussed for each decadal cohort: the rank-rank gradient, absolute upward mobility ($p_{25}$), and upward interval mobility ($\mu_0^{50}$).\footnote{The bounds on the rank-rank gradient describe the slopes of the set of best linear approximators to feasible CEFs.} For reference, we plot recent estimates of similar educational mobility measures from USA and Denmark.\footnote{Rank-rank correlations of education are from \citeasnoun{Hertz2008}, which are equal to the slope of the rank-rank regression coefficient if estimated on uncensored rank data. For absolute mobility, we calculate $p_{25}$ for the U.S. and Denmark from the distributions shown in Figure (ref), with data from \citeasnoun{Chetty2014c}.}\ensuremath{^{,}}\footnote{We calculate and show bootstrap confidence sets using 1,000 bootstrap samples from the underlying datasets, following methods described in \citeasnoun{m_Imbens2004} and \citeasnoun{Tamer2010}.} Once we allow expected child outcomes to vary within the bottom parent education bin, both the rank-rank gradient and absolute upward mobility have wide and minimally informative bounds. In contrast, upward interval mobility is estimated with tight bounds in all periods. According to this measure, upward mobility has changed very little over the four decades studied; there is a small gain from the 1950s to the 1960s, followed by a small decline from the 1960s to the 1980s. On average, Indian mobility is as far below that in the United States as mobility in the United States is below Denmark. Table (ref) reports the bounds for each measure and cohort with bootstrap confidence sets under a range of curvature restrictions. Moderate curvature restrictions generate substantial improvements on the estimation of the value of the CEF (e.g., $p_{25}$), but are considerably less important for the interval mean measures (e.g., $\mu_{0}^{50}$), which are tightly bounded even with unconstrained curvature.
In conclusion, the most widely used mobility estimator, the rank-rank gradient, presents an incomplete and potentially biased picture of intergenerational educational mobility. Upward interval mobility, in contrast, yields informative estimates even without a curvature constraint, making it feasible to study upward mobility in the lower-ranked parts of the distribution, even in a context with extreme interval censoring. The advantage of upward interval mobility is likely to be replicated in mobility studies in the many other countries where older generations are clustered in less educated bins.
We propose a method that generates useful bounds on a conditional expectation function when the conditioning variable is interval-censored. Tight bounds on parameters of interest are possible because of three innovations. First, we show that CEF bounds are substantially improved when the distribution of the conditioning variable is known, and many economic contexts have distributions that are either known with certainty or are assumed by convention. Second, we show that there are many intervals in which the conditional mean can be bounded tightly, even when the bounds on the CEF itself are wide across its domain. Third, bounds can be improved by imposing a constraint on the curvature of the CEF, which is justified in many empirical contexts. A curvature constraint can further substitute for the assumption of monotonicity, making it possible to conduct inference in interval data contexts where there is not a strong theoretical basis for monotonicity.
We also propose a traactable numerical framework for bounding CEFs and functions of the CEF with arbitrary structural restrictions. Simulations of interval censoring indicate that the methods perform well in common empirical scenarios. In our applications, the first two innovations prove sufficient to bound parameters of interest informatively, but any of these three alone is insufficient.
A useful thought experiment when working with interval data is to explore how estimates are affected as intervals become more or less granular. The bounds presented in this paper become wider when the data become more coarse, as should be expected given that information has been removed. In contrast, with conventional point estimation approaches, the use of coarser intervals can lead to different point estimates, thus obscuring the loss of information to the researcher. Our method is transparent about what is known and what is not known.
We have shown that our method can be used to solve known problems in the study of mortality and of intergenerational mobility. Generating bounds on outcome variables by education quantile is an application with many other potential uses, given the large number of contexts where education is of interest as a dependent variable but available only in a small number of bins. Other useful applications may be found where the conditioning variable takes the form of interval-censored income data, or Likert scale responses, among others.