Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
68,807 characters · 22 sections · 23 citation commands
Partial Identification with Auxiliary Moment Restrictions
\noindentCredit authorship contribution statement:
Arie Beresteanu: Writing - review & editing, Writing - original draft, Methodology, Investigation, Conceptualization.
Behrooz Moosavi Ramezanzadeh: Writing - review & editing, Writing - original draft, Methodology, Investigation, Conceptualization.
\noindentConflict of Interest: The authors declare that they have no conflict of interest.
The literature on partial identification offers a principled way to learn from data under weak and credible assumptions. Yet it is adopted less often than its intellectual standing would warrant. A recurring complaint among practitioners is that the identified sets these methods deliver are large and, as a result, uninformative for policy. Confronted with wide bounds, researchers frequently impose strong behavioral or distributional assumptions that restore point identification but erode credibility---the very property that partial identification was designed to protect.
This paper develops a remedy that adds information rather than assumptions, and it does so in a setting where the information arises naturally. In a growing number of applications the outcome of interest is reported only as an interval, not by accident, but by design: to protect respondent confidentiality, the data custodian deliberately coarsens the true value $y^*$ into a published range $[y_L,y_U]$. Income is bracketed or top-coded in public-use survey files; sensitive variables are released only within bins; and modern formal-privacy systems perturb record-level data before release. At the same time, the very same custodian continues to publish accurate population-level summaries of the same latent outcome---official means, or the exact aggregates that a formal-privacy system releases as invariants---precisely because such aggregates pass disclosure review without exposing any individual record. The econometrician therefore confronts interval-valued microdata alongside known population aggregates of the same unobserved outcome. We show that this auxiliary information---a byproduct of the privacy regime rather than an external dataset that must be located and merged---can substantially shrink the identified set for a best linear predictor (BLP), without imposing any restriction on the data-generating process beyond the interval-containment condition $\mathbf{P}(y_L\le y^*\le y_U)=1$.
We study the predictor associated with squared loss: the best linear predictor that solves the population linear projection problem for the conditional expectation of the latent outcome. Our analysis builds directly on BM2008, who characterize the sharp identified set for the conditional-expectation BLP with interval-valued outcomes. We ask how much identifying power is contained in knowledge of (i) the unconditional mean of the latent outcome, (ii) a known moment of a transformation of the latent outcome, and (iii) a conditional mean given a subset of the covariates or given an external variable that is excluded from the model.
\paragraph{Overview of the analysis and contributions.} The central device in the paper is a restricted selection set. We leave the observed random interval $Y=[y_L,y_U]$ unchanged and restrict only the set of admissible selections from that interval. Each candidate selection represents a possible realization of the latent outcome $y^*$, and auxiliary information removes those selections whose moments do not agree with the known aggregate. The resulting identified region is therefore the image, under the best-linear-predictor map, of a selection set restricted by the available population information. Nonemptiness of the relevant restricted selection sets is established in the companion note beresteanu2025noterestrictedselectionset.
We use this framework to study several combinations of target parameters and auxiliary restrictions. The baseline target is the coefficient vector of the best linear predictor of $y^*$. When the unconditional mean $E[y^*]=\kappa$ is known, the identified region is the intersection of the unrestricted region with an affine hyperplane. Under a nondegeneracy condition, the restriction therefore lowers the affine dimension of the region by one. We then ask how the result changes when the target or the auxiliary information involves a transformation $f(y^*)$. If the target is the best linear predictor of $f(y^*)$, knowledge of $E[y^*]=\kappa$ restricts the admissible selections but does not impose a direct affine restriction on the transformed-outcome coefficients. If instead $E[f(y^*)]=\kappa_f$ is known, the transformed-outcome coefficients satisfy an affine hyperplane restriction. By contrast, when the target remains the best linear predictor of $y^*$ and $E[f(y^*)]=\kappa_f$ is used only as auxiliary information, the restriction generally narrows the identified region without reducing its dimension.
We next consider conditional auxiliary information. If the researcher knows $$ E[y^* \mid x_1]=\kappa(x_1) $$ for a subvector $x_1$ of the regressors, a population Frisch--Waugh--Lovell decomposition shows that the corresponding coefficient block is an affine function of the coefficients on the remaining regressors. The restriction therefore imposes $d_1+1$ exact affine restrictions and reduces the full identified set to an affine image of a lower-dimensional coefficient region. If instead the researcher knows $$ E[y^* \mid v]=\kappa(v) $$ for an external variable $v$ excluded from the predictor, no coefficient is generally pinned down by an affine equation. The restriction tightens identification by fixing how the available interval width must be allocated across values of $v$, and it improves on the pooled-mean restriction whenever that allocation differs from the one selected by the pooled optimization problem.
Beyond characterizing these geometric effects, we quantify the identifying value of the mean and conditional-mean restrictions. Directional bounds can be written as allocation problems in which the interval width $y_U-y_L$ is assigned across observations according to the score associated with a direction in coefficient space. A known mean fixes the total amount of width that may be allocated, while a known conditional mean fixes how that amount is divided across covariate cells or values of an external variable. This representation yields formulas for the loss in each directional support value and shows that a restriction matters only when it forces the allocation away from the one an unrestricted optimizer would choose.
Finally, we extend the restricted-selection framework to nonlinear and discontinuous transformations. For continuous transformations, we establish convexity and derive a Lagrangian support representation under suitable regularity conditions. For indicator transformations, which encode information such as a known distributional probability at a threshold, we provide a finite-sample sorting characterization. This characterization connects this paper to the quantile analysis of beresteanu-2021, but the role of the quantile information is different here: it is used as auxiliary information to sharpen identification of a mean-based best linear predictor rather than as the target of the analysis itself.
\paragraph{Related literature.} At its core, partial identification asks what can be learned about a parameter that is only set-identified, emphasizing credible assumptions; see Manski2003 for a monograph treatment and Tamer2010 for a survey. The foundational treatment of interval-valued regression itself is manskitamer, who derive nonparametric bounds on a regression function when an outcome or regressor is observed only within an interval; we work throughout in their interval-containment setting but target the BLP coefficient rather than the conditional expectation function directly, following BM2008. A central methodological theme is to translate credible restrictions on the data-generating process or on agents' behavior into inequality restrictions on the parameters, and to show that these inequalities are both necessary and sufficient, so that the identified set is sharp. We obtain sharpness through the random-set apparatus of Molchanov2005,BMM2011 and MolchanovMolinari2014, and we study how auxiliary information on the mean of the latent outcome, or of a transformation of it, tightens the corresponding selection sets and hence the identified regions.
The motivating environment connects our analysis to the literature on confidentiality protection. Agencies have long limited disclosure by bracketing or top-coding sensitive variables, which is precisely what makes the outcome interval-valued; more recently, national statistical offices have adopted formal privacy, as in the U.S. Census Bureau's 2020 Disclosure Avoidance System, which protects record-level data through differential privacy while releasing selected aggregates exactly as invariants Abowd2022. We take such released aggregates as given and ask what they identify, so that the auxiliary information in our results is a feature of the disclosure regime rather than an additional assumption.
Our use of external information is also related to, but distinct from, the data-combination literature. CrossManski2002 consider a setting in which $(y,x,z)$ are never jointly observed and the researcher knows the marginals $P(y\mid x)$ and $P(z\mid x)$ from two separate sources, which are then combined to bound $E(y\mid x,z)$. In our setting a single sample is observed in which the outcome is reported only as an interval, and the additional information takes the form of known moments of the unobserved outcome. Unlike in CrossManski2002, the conditional expectation of $y$ given the covariates is already partially identified without the auxiliary information; that information is used here to tighten the bounds. Our approach is also distinct from methods that treat the bounds on the outcome as themselves estimated or as an unknown object to be restricted directly: ChandrasekharEtAl2019 construct BLP regions from an estimated band around the outcome and develop inference for the resulting support function, while MagnacMaurin2008 and beresteanu-2021 restrict an unknown bounded function or transformation of the outcome. We instead hold the observed interval $Y=[y_L,y_U]$ fixed throughout and restrict only the set of selections admissible within it, so that auxiliary information enters through which completions of the observed interval are permitted rather than through the interval itself. Finally, shape restrictions such as monotonicity or concavity are typically imposed on selections by redefining the random set from which they are drawn; MolchanovMolinari2014 discuss restricting the values of selections through intersection with another (possibly random) set. We instead leave the observed random set intact and restrict the set of selections to those whose moments match externally known values. In spirit, this is close to the use of monotone-instrument and other auxiliary restrictions to sharpen partially identified objects, as in ManskiPepper2000; the present paper extends BM2008 by showing that auxiliary moments---of the outcome itself, of a subvector of covariates, of an external variable, and of transformations of the outcome, including quantiles---carry identifying power, and by quantifying it.
\paragraph{Roadmap.} (ref) develops the results summarized above: the unconstrained region and its basic properties, the hyperplane characterization under a known unconditional mean, directional and coordinate bounds, the value of a known mean, identification with transformations (with and without retargeting), a known conditional mean given a subvector of covariates, and a known conditional mean given an external variable. (ref) illustrates every result numerically using interval-valued wages constructed from the 2020 March Current Population Survey (CPS) via a stylized coarsening mechanism that plays the role of a privacy device applied to the observed wage, in the spirit of the data used by CHT2007 and BM2008. (ref) reports supplementary numerical detail. (ref) collects the background results from random-set theory and convex analysis used in the proofs, including the general treatment of vector-valued transformations and quantile information; (ref) contains all proofs.
In this section, we characterize the identification region for the parameter vector $\theta$ of a best linear predictor under squared loss when the outcome is interval-valued and the researcher has access to auxiliary population information about the latent outcome.
Let $(\Omega,\mathcal F,\mathbf{P})$ be a probability space. Let $y^*\in\mathbb{R}$ be a latent scalar outcome, and let \[ Y=[y_L,y_U] \] be an observed random interval satisfying \[ \mathbf{P}(y_L\leq y^*\leq y_U)=1. \] Let $x\in\mathbb{R}^d$ be an observed vector of covariates, and define the augmented covariate vector \[ \tilde x=(1,x_1,\dots,x_d)' \in\mathbb{R}^{d+1}. \]
Throughout this section, we impose the following regularity conditions.
Because $Q$ is a symmetric second-moment matrix, nonsingularity is equivalent to positive definiteness.
Define the set of integrable measurable selections of $Y$ by \[ \operatorname{\mathbf{Sel}}^1(Y) = \left\{ y\in L^1(\Omega,\mathcal F,\mathbf{P}): y_L\leq y\leq y_U \quad\mathbf{P}\text{-a.s.} \right\}. \]
Absent additional information, every $y\in\operatorname{\mathbf{Sel}}^1(Y)$ is a candidate for the latent outcome $y^*$. For a given selection $y$, the population best linear predictor coefficient is defined by the normal equations \[ \mathbb{E}\!\left[ \tilde x\bigl(y-\tilde x'\theta\bigr) \right] = \mathbf 0. \] Because $Q$ is nonsingular, the coefficient generated by $y$ is uniquely given by \[ \theta(y) = Q^{-1}\mathbb{E}[\tilde x y]. \]
The unconstrained sharp identification region is therefore
Equivalently, \[ \Theta^I = \left\{ \theta\in\mathbb{R}^{d+1}: \mathbb{E}\!\left[ \tilde x\bigl(y-\tilde x'\theta\bigr) \right] = \mathbf 0 \text{ for some }y\in\operatorname{\mathbf{Sel}}^1(Y) \right\}. \]
It is useful to define the attainable cross-moment set \[ \mathcal M = \left\{ \mathbb{E}[\tilde x y]: y\in\operatorname{\mathbf{Sel}}^1(Y) \right\}. \] Then \[ \Theta^I=Q^{-1}\mathcal M. \]
The proof represents $\mathcal M$ as the Aumann integral of an integrably bounded, measurable, compact-valued, and convex-valued correspondence. The relevant background results are collected in (ref) and all proofs for this section are in (ref).
We first consider the environment in which the researcher knows the unconditional population expectation of the latent outcome.
Under Assumption (ref), candidate outcomes are restricted to \[ \operatorname{\mathbf{Sel}}^1(Y\mid\kappa) = \left\{ y\in\operatorname{\mathbf{Sel}}^1(Y): \mathbb{E}[y]=\kappa \right\}. \]
The compatibility condition \[ \mathbb{E}[y_L]\leq\kappa\leq\mathbb{E}[y_U] \] guarantees that $\operatorname{\mathbf{Sel}}^1(Y\mid\kappa)$ is nonempty. To see this, assume that $\mathbb{E}[y_U-y_L]>0$, and define \[ \lambda_\kappa = \frac{\kappa-\mathbb{E}[y_L]}{\mathbb{E}[y_U-y_L]} \in[0,1]. \] Then \[ y_\kappa = y_L+\lambda_\kappa(y_U-y_L) \] belongs to $\operatorname{\mathbf{Sel}}^1(Y\mid\kappa)$ and $\mathbb{E}[y_{\kappa}]=\kappa$. If $\mathbb{E}[y_U-y_L]=0$, then $P(y_L<y_U)=0$ implies that $y_L=y_U$ almost surely, and Assumption (ref) implies $\kappa=\mathbb{E}[y_L]=\mathbb{E}[y_U]$.
Given Assumption (ref), we use the constrained selection set $\operatorname{\mathbf{Sel}}^1(Y|\kappa)$ to define the identification region as
The identification region in equation ((ref)) is not easy to compute directly by going over all selections in $\operatorname{\mathbf{Sel}}^1(Y|\kappa)$. The following Proposition shows how to characterize the constrained identification set as an intersection of two sets that can be computed.
The set $\Theta^I_\kappa$ is obtained by intersecting $\Theta^I$ with the affine hyperplane, $\Theta_{\kappa}$, induced by the known mean. Thus, knowing the mean of $y^*$ can only reduce the identified set, and under nondegeneracy it lowers the affine dimension by one. Before we can measure how much the known mean tightens the identification region, we first characterize the boundary of the unconstrained set $\Theta^I$ itself. Researchers are often interested in the projection of an identification region onto a specific linear combination of coordinates: given $\Theta^I$, what is the maximal or minimal value that $\theta_j$ can admit, for $j=0,1,\dots,d$? More generally, for $r\in\mathbb{R}^{d+1}$ representing a linear combination of the parameters in $\theta$, we would like to find the maximal or minimal value that $r'\theta$ can admit over $\Theta^I$. This apparatus is reused in (ref) below to quantify the impact of the auxiliary information $\mathbb{E}[y^*]=\kappa$.
To characterize the boundary of the identified set, fix a direction $r\in\mathbb{R}^{d+1}$ and define \[ s_r = r'Q^{-1}\tilde x. \] For a coordinate direction $r=e_j$, write \[ s_j=e_j'Q^{-1}\tilde x. \]
For every selection $y\in\operatorname{\mathbf{Sel}}^1(Y)$, \[ r'\theta(y) = r'Q^{-1}\mathbb{E}[\tilde x y] = \mathbb{E}[s_r y]. \]
Let \[ \Delta:=y_U-y_L\geq0. \] Every $y\in\operatorname{\mathbf{Sel}}^1(Y)$ can be represented as \[ y=y_L+\tau\Delta, \] where $\tau:\Omega\to[0,1]$ is measurable.
In the unconstrained problem, where $\mathbb{E}[y^*]$ is not given, \[ \sup_{y\in\operatorname{\mathbf{Sel}}^1(Y)}\mathbb{E}[s_r y] = \mathbb{E}[s_r y_L] + \sup_{0\leq\tau\leq1}\mathbb{E}[s_r\tau\Delta]. \] Because there is no aggregate restriction on $\tau$, an unconstrained maximizing allocation of $\tau$ satisfies $\tau^{r}(\omega)=1 \quad\text{on }\{\omega : s_r>0\}$, $ \tau^{r}(\omega)=0 \quad\text{on }\{\omega : s_r<0\}$, and with arbitrary values on $\{\omega : s_r=0\}$.
Therefore,
Similarly,
For bounds on $\theta_j$ we set $r=e_j$. The maximal and minimal values are \[ \theta_j^{\max} = \mathbb{E}[s_jy_L]+\mathbb{E}[s_j^+\Delta] \] and \[ \theta_j^{\min} = \mathbb{E}[s_jy_L]-\mathbb{E}[s_j^-\Delta]. \]
Equations ((ref)) and ((ref)) require the calculation of two expectations that include stochastic weights $s_r$. We next give a useful representation of the resulting directional and coordinate breadths.
The proof for Proposition (ref) is in the Appendix. The coordinate formula in Proposition (ref) is a Frisch--Waugh--Lovell type representation. In a point-identified linear projection, the coefficient on $\tilde x_j$ can be computed using the residualized regressor $\tilde x_j^*$. Here the same residualization determines the width of the identified interval for $\theta_j$. The only additional object is the interval width $\Delta=y_U-y_L$, which measures how much freedom the analyst has in choosing a selection from the observed interval. Expression (ref) shows that partial identification of $\theta_j$ is not determined by the average interval width alone. It depends on the joint distribution of $\Delta$ and the magnitude of the partialled-out regressor $|\tilde x_j^*|$. The quantity in (ref) can be consistently estimated using sample analogs of the expectations.
We now use the apparatus of (ref) to quantify the impact of knowing $\mathbb{E}[y^*]=\kappa$ on the identification region, by comparing the directional bounds of $\Theta^I_\kappa$ with those of $\Theta^I$.
Under Assumption (ref), the allocation $\tau$ must satisfy \[ \mathbb{E}[y_L+\tau\Delta]=\kappa, \] or equivalently \[ \mathbb{E}[\tau\Delta] = \kappa-\mathbb{E}[y_L] =: \alpha. \] Since $\kappa \in \left[\mathbb{E}[y_L],\mathbb{E}[y_U]\right]$, $0\leq\alpha\leq\mathbb{E}[\Delta]$.
The sharp directional upper bound is therefore obtained from
The problem in ((ref)) has a useful allocation interpretation. The variable $\tau(\omega)$ determines what fraction of the available interval width $\Delta(\omega)$ is assigned to the upper endpoint in state $\omega$. The mean restriction fixes the total amount of width that must be allocated: $$ \mathbb{E}[\tau\Delta]=\alpha. $$ The objective assigns value $s_r(\omega)$ to one unit of allocated width in state $\omega$. Thus, states with larger $s_r$ are more valuable for maximizing the directional bound $r'\theta$.
It is therefore natural to measure the size of a state not by its probability alone, but by how much interval width it contributes. Define the finite nonnegative measure $$ \mu(A)=\mathbb{E}[\Delta 1_A], \qquad A\in\mathcal F. $$ Under $\mu$, a set of states receives a weight equal to the expected interval width available on that set. With this notation, the constraint $\mathbb{E}[\tau\Delta]=\alpha$ becomes $$ \int_\Omega \tau\,d\mu=\alpha, $$ and the objective becomes $$ \mathbb{E}[s_r\tau\Delta]=\int_\Omega s_r\tau\,d\mu. $$ Hence (ref) can be written as $$ \sup_{\tau}\int_\Omega s_r\tau\,d\mu \quad\text{subject to}\quad 0\leq \tau\leq 1,\qquad \int_\Omega \tau\,d\mu=\alpha. $$ This is a continuous fractional-knapsack problem: allocate exactly $\alpha$ units of width-weighted mass to the states with the largest values of $s_r$, allowing fractional allocation at the cutoff if necessary.
The proposition expresses the tightening from the known mean as an area under the width-weighted quantile function of the score $s_r$. The unrestricted optimizer corresponds to the quantile level $1-u_r$, i.e. the zero cutoff $s_r=0$. The point $1-p_\kappa$ is the cutoff required by the known mean. The contraction is the value lost when moving from the unconstrained cutoff to the mean-constrained cutoff. Hence the restriction has no first-order effect when the two cutoffs are close and the distribution of $s_r$ is smooth around zero.
The formula has a direct plug-in analogue. In a sample, one computes $\hat s_{ri}=r'\hat Q^{-1}\tilde x_i$, weights observations by $\Delta_i/\sum_j\Delta_j$, sorts the scores, and evaluates the weighted area under the empirical quantile function between $1-\hat p_\kappa$ and $1-\hat u_r$. The full sample formula is given in Appendix (ref).
We now show that small deviations in the width budget have only second order effect on the directional support value.
Up to this point we have studied the best linear predictor of the latent outcome $y^*$. In many applications, however, the parameter of interest is instead the best linear predictor of a transformation $f(y)$, or a known population moment of $f(y^*)$ is used as auxiliary information about $y^*$ itself. Examples include logarithms of income, indicator functions, or nonlinear utility transformations. We consider both uses of a transformation $f$ in this subsection, since they share the same regularity conditions and rely on the same splicing argument, but lead to identification regions with markedly different geometry.
Let \[ f:\mathbb{R}\to\mathbb{R} \] be Borel measurable.
Since Assumption (ref) already gives $E[|\tilde x|^2]<\infty$, Assumption (ref) also implies $E[|\tilde x|F]<\infty$ by Cauchy-Schwarz.
We first retarget the object of interest to the best linear predictor of $f(y^*)$ itself, \[ \theta_f=Q^{-1}\mathbb{E}[\tilde x f(y^*)], \] maintaining the restriction \[ \mathbb{E}[y^*]=\kappa. \] The key difference from the previous section is that the objective now depends on $f(y^*)$, while the auxiliary information continues to constrain the untransformed variable through $E[y^*]=\kappa$. Unless $f$ is affine, the objective is nonlinear in $\tau$, so the fractional-knapsack representation of Section (ref) no longer applies. Proposition (ref) below shows that convexity nevertheless survives under nonatomlessness.
Define,
and
Although the restriction on admissible selections is unchanged, the optimization problem is generally no longer linear because the objective depends nonlinearily on the allocation variable through \[ f(y_L+\tau\Delta). \] Unless $f$ is affine, the optimization is no longer linear in $\tau$. Consequently, the fractional-knapsack representation of (ref). The following Proposition shows that despite this loss of linearity, convexity of the identified region is preserved.
The proposition shows that the geometric properties established earlier survive the introduction of arbitrary measurable transformations. Although the optimization problem becomes nonlinear, the attainable moment set remains convex because convexity is generated by the atomless probability space through Lyapunov's theorem rather than by linearity of $f$.
For every $r\in\mathbb{R}^{d+1}$, let
Convexity of the transformation itself yields an additional implication. Since every admissible selection satisfies $E[y]=\kappa$, Jensen's inequality implies
If $f$ is concave, the inequality is reversed. Thus, even though the restriction is imposed only on the mean of the latent outcome, convexity (concavity) of $f$ automatically generates a lower (upper) bound on the mean of every admissible transformed outcome.
The preceding discussion assumes that only $E[y^*]$ is known. If the researcher additionally knows the population mean of the transformed outcome itself, \[ \mathbb{E}[f(y^*)]=\kappa_f, \] then the transformed problem reduces to the same hyperplane characterization developed in (ref), namely \[ \mathbb{E}[\tilde x]'\theta_f=\kappa_f. \] Thus, auxiliary information about the transformed outcome can be incorporated exactly as before.
We now consider the complementary case: the target remains the original $\theta=Q^{-1}\mathbb{E}[\tilde x y^*]$, and a known moment of the transformation, $\mathbb{E}[f(y^*)]=\kappa_f$, is used purely as auxiliary information narrowing the selections admissible for $y^*$. No restriction on $\mathbb{E}[y^*]$ itself is imposed.
Define the achievable moment set \[ \mathcal K_f=\{\mathbb{E}[f(y)]:y\in\operatorname{\mathbf{Sel}}^1(Y)\}, \] and, for $\kappa_f\in\mathcal K_f$, \[ \operatorname{\mathbf{Sel}}^1(Y\mid f,\kappa_f) = \left\{ y\in\operatorname{\mathbf{Sel}}^1(Y):\mathbb{E}[f(y)]=\kappa_f \right\}. \] The compatibility condition, $\kappa_f\in\mathcal K_f$, guarantees this set is nonempty; unlike the linear case of Assumption (ref), $\mathcal K_f$ need not be an interval with known endpoints in closed form, so compatibility is stated directly as membership in $\mathcal K_f$ rather than via an explicit interval.
The resulting sharp identification region for the original target is
The proof applies Lyapunov's theorem for vector measures to the pair $(f(y_1)-f(y_0),\,\tilde x(y_1-y_0))$ for any two selections $y_0,y_1\in\operatorname{\mathbf{Sel}}^1(Y\mid f,\kappa_f)$: since both satisfy $\mathbb{E}[f(y_0)]=\mathbb{E}[f(y_1)]=\kappa_f$, the first component of this pair has mean zero, so splicing $y_0$ and $y_1$ along the sets furnished by Lyapunov's theorem preserves the moment restriction exactly while tracing out the line segment between $Q^{-1}\mathbb{E}[\tilde x y_0]$ and $Q^{-1}\mathbb{E}[\tilde x y_1]$ in $\theta$-space; see (ref) for the background result and (ref) for the full argument.
The geometric distinction is important. Unlike Proposition (ref), $\Theta^I_{\kappa_f}$ need not collapse to a lower-dimensional set. The restriction $\mathbb{E}[y^*]=\kappa$ is linear in $\theta$ because $\tilde x$'s first coordinate is $1$, so $\mathbb{E}[\tilde x]'\theta=\mathbb{E}[y]$ directly, pinning $\theta$ to an affine hyperplane. A nonlinear moment restriction $\mathbb{E}[f(y^*)]=\kappa_f$ has no such direct algebraic counterpart in $\theta$-space: it removes some selections from $\operatorname{\mathbf{Sel}}^1(Y)$ without confining the resulting $\theta$ values to any fixed hyperplane, so $\Theta^I_{\kappa_f}$ is generically full-dimensional -- a narrowed region rather than a segment. This is the same regularity condition and the same splicing argument as Proposition (ref), applied to a different target; the two propositions differ only in which functional of $y$ the objective retains and which the constraint restricts, and it is exactly this difference that separates a lower-dimensional segment from a full-dimensional band.
We now consider a stronger form of auxiliary information in which the researcher knows the conditional mean of the latent outcome given a subset of the covariates, and we show how this restriction reduces the full identification problem to the coefficient block associated with the remaining covariates.
Partition \[ x=(x_1',x_2')', \] where \[ x_1\in\mathbb{R}^{d_1}, \qquad x_2\in\mathbb{R}^{d_2}, \qquad d_1+d_2=d. \] Define \[ w =
. \] Then \[ \tilde x =
. \]
Define \[ \operatorname{\mathbf{Sel}}^1(Y\mid\kappa(x_1)) = \left\{ y\in\operatorname{\mathbf{Sel}}^1(Y): \mathbb{E}[y\mid x_1]=\kappa(x_1) \quad\mathbf{P}\text{-a.s.} \right\}. \]
This set is nonempty. Let \[ D(x_1)=\mathbb{E}[\Delta\mid x_1] \] and define \[ \lambda(x_1) =
\] On the event $D(x_1)=0$, nonnegativity of $\Delta$ implies $\Delta=0$ conditionally almost surely, and compatability therefore gives $\kappa(x_1)=E[y_L|x_1]=E[y_U|x_1]$. Then from the definition above, $0\leq\lambda(x_1)\leq1$ almost surely and \[ y_\kappa = y_L+\lambda(x_1)\Delta \] satisfies \[ \mathbb{E}[y_\kappa\mid x_1]=\kappa(x_1). \]
The sharp identified set is
Partition \[ \theta =
, \qquad \theta_1\in\mathbb{R}^{d_1+1}, \quad \theta_2\in\mathbb{R}^{d_2}, \] and define \[ \Sigma_{11}=\mathbb{E}[ww'], \qquad \Sigma_{12}=\mathbb{E}[wx_2'], \] \[ \Sigma_{21}=\Sigma_{12}', \qquad \Sigma_{22}=\mathbb{E}[x_2x_2']. \] Then \[ Q =
. \]
Define \[ g_1 = \mathbb{E}[w\kappa(x_1)] \] and the linear-projection residual \[ x_2^* = x_2-\Sigma_{21}\Sigma_{11}^{-1}w. \] Then \[ \mathbb{E}[x_2^*w']=0 \] and \[ \Sigma_{2\cdot1} = \mathbb{E}[x_2^*x_2^{*\prime}] = \Sigma_{22} - \Sigma_{21}\Sigma_{11}^{-1}\Sigma_{12} \] is positive definite.
Define
Consequently, $\Theta^I_{\kappa(x_1)}$ is the image of $\Theta_2^I$ under an injective affine map and therefore has affine dimension at most $d_2$. In particular, the conditional mean restriction imposes $d_1 +1$ exact linear restrictions on the full coefficient vector.
If $x_1$ has finite support, the conditional restriction is a finite collection of cell-specific moment restrictions. If $x_1$ is continuously distributed, it is an infinite-dimensional conditional restriction. The block normal-equation decomposition is the same in both cases.
An alternative residual is \[ \bar x_2=x_2-\mathbb{E}[x_2\mid x_1]. \] It satisfies \[ \mathbb{E}[\bar x_2\mid x_1]=0 \] and therefore \[ \mathbb{E}[\bar x_2y] = \mathbb{E}[\bar x_2\{y-\kappa(x_1)\}] \] for every conditionally admissible selection, provided that these expectations exist. This conditional-mean residualization corresponds to partialling out an unrestricted function of $x_1$ and should be distinguished from the finite-dimensional linear residual $x_2^*$.
We next consider auxiliary conditional-mean information indexed by an external variable v that does not enter the best linear predictor. Unlike conditioning on a subvector of the regressors, this restriction does not generally impose a direct linear restriction on the coefficient vector. Instead, it sharpens identification by fixing how the total interval-width allocation must be distributed across values of $v$.
Let $v$ be an observed external variable that need not enter the BLP specification.
Define \[ \operatorname{\mathbf{Sel}}^1(Y\mid\kappa(v)) = \left\{ y\in\operatorname{\mathbf{Sel}}^1(Y): \mathbb{E}[y\mid v]=\kappa(v) \quad\mathbf{P}\text{-a.s.} \right\}. \]
As before, conditional compatibility implies nonemptiness by taking \[ y_v = y_L+\lambda(v)\Delta, \] where \[ \lambda(v) =
\]
On the event $E[\Delta|v]=0$, nonnegativity of $\Delta$ implies that $\Delta=0$ conditionally almost surely. Compatibility, therefore, gives $\kappa(v)=E[y_L|v]=E[y_U|v]$ a common conditional expectation for all selections. As a result, $E[y_v|v]=\kappa(v)$
The sharp identified set is
Unlike the restriction studied in (ref), the external conditional mean generally has no direct finite-dimensional representation in coefficient space because v is excluded from $\Tilde{x}$. Its identifying content is instead revealed through the support-function comparison below.
Fix $r\in\mathbb{R}^{d+1}$. Then
Writing $y=y_L+\tau\Delta$, the conditional restriction becomes \[ \mathbb{E}[\tau\Delta\mid v] = \kappa(v)-\mathbb{E}[y_L\mid v] =: \alpha(v). \]
Suppose \[ v\in\{v_1,\dots,v_M\}, \qquad p_m=\mathbf{P}(v=v_m)>0. \] Let $\tau_m$ denote the allocation rule within event $\{v=v_m\}$. Then
Define the cell measures \[ \mu_m(A) = \mathbb{E}[\Delta\mathbf{1}_{A\cap\{v=v_m\}}], \] and the cell budgets \[ A_m = \mathbb{E}[ \{\kappa(v_m)-\mathbb{E}[y_L\mid v=v_m]\} \mathbf{1}\{v=v_m\} ]. \] Let \[ \kappa:=\mathbb{E}[\kappa(v)]=\mathbb{E}[y^*], \qquad \alpha=\kappa-\mathbb{E}[y_L]. \] Then \[ \sum_{m=1}^M A_m=\alpha. \]
For a finite measure $\nu$, define the upper-tail functional \[ \mathcal T_a^\nu(Z) = \sup\left\{ \int\tau Z\,d\nu: 0\leq\tau\leq1,\ \int\tau\,d\nu=a \right\}, \qquad 0\leq a\leq\nu(\Omega). \]
Let \[ \mu(A)=\mathbb{E}[\Delta\mathbf{1}_A]. \] Then \[ \mu(\Omega)=\mathbb{E}[\Delta]=\sum_{m=1}^M\mu_m(\Omega). \]
Thus, locally, the additional contraction is governed by dispersion of the cell-specific optimal cutoffs around the pooled cutoff. Cells contribute more to the additional contraction when their imposed budgets force their optimal cutoffs farther from the pooled cutoff, with the local cost weighted by the amount of interval width in the cell and by the width-weighted density of the score at the pooled cutoff. This is a statement about cutoff dispersion, not a general identity with between-group variance of the score.
When v is continuously distributed, the same logic applies pointwise in $v$: the conditional restriction fixes a separate width budget at almost every value of $v$, and the support value is obtained by integrating the corresponding conditional knapsack values.
Suppose $v$ takes values in a standard Borel space, a regular conditional distribution of $(y_L,y_U,x)$ given $v$ exists, and the conditional optimization admits a measurable selection of optimizers. Then
For almost every $t$, define \[ K_r(t) = \sup_{\substack{0\leq\tau\leq1\\ \mathbb{E}[\tau\Delta\mid v=t]=\alpha(t)}} \mathbb{E}[s_r\tau\Delta\mid v=t]. \] If $t\mapsto K_r(t)$ is measurable, then \[ h_{\Theta^I_{\kappa(v)}}(r) = \mathbb{E}\!\left[ \mathbb{E}[s_ry_L\mid v]+K_r(v) \right]. \] If $v$ has density $f_v$, this becomes \[ h_{\Theta^I_{\kappa(v)}}(r) = \int \left[ \mathbb{E}[s_ry_L\mid v=t] + K_r(t) \right] f_v(t)\,dt. \]
This section provides a numerical illustration of the identification results in (ref). The exercise is not intended as a substantive analysis of the returns to education. Instead, we use Current Population Survey (CPS) Annual Social and Economic Supplement (ASEC) observed income to construct interval-valued outcomes and then compare the identified regions obtained under different forms of auxiliary information. Throughout this section, expectations and identified regions are implemented using their empirical analogs. To avoid excessive notation, we retain the population notation from (ref). For simplicity, the numerical illustration treats the retained sample as an equally weighted empirical distribution. The exercise is intended to illustrate the geometry of the identification results rather than to provide population-representative estimates.
We restrict the sample to respondents with positive wage and salary income (WSAL_VAL) who report being employed, rescale income to units of \$1{,}000, and drop the top percentile of the income distribution as a cosmetic trim against extreme outliers. The variable measuring years of completed education (educ_numeric) enters the baseline covariate vector $\tilde x=(1,\mathrm{educ})$. Two further variables are retained for the extensions taken up later in this section: a categorical race indicator constructed from PRDTRACE (equal to $1$ for white respondents and $2$ for non-white), used in Section (ref) as $x_1$; and age (A_AGE), used in section (ref) as the external variable $v$. After these restrictions the working sample contains $n=22{,}397$ observations, with mean income of \$63{,}990 and mean educational attainment of $14.2$ years.
Because $y^*$ (income) is observed as a point value in the underlying CPS extract, we generate the interval $Y=[y_L,y_U]$ used throughout this section using a stylized interval-privacy mechanism based on DingDing2022, who introduce interval privacy as a privacy criterion distinct from differential privacy: rather than perturbing a respondent's value with additive noise, the mechanism narrows it to a random range that provably contains the truth, constructed so that the range's conditional distribution given the true value is uninformative beyond the range itself. For each observation $i$, we draw two quantile indices independently of $y^*_i$, convert them into income anchors using the empirical income distribution, and order the resulting values. Together with the lower and upper endpoints of the empirical support, these anchors partition the income domain into three intervals. The released interval $[y_{L,i},y_{U,i}]$ is the unique partition cell containing $y^*_i$. Because the anchors are constructed without reference to the respondent's own value, the resulting interval is not centered on $y_i^*$, which is what makes it a nontrivial input for the identification results below. A construction that released an interval centered exactly at $y^*_i$ would reveal the latent outcome through the interval midpoint, defeating the purpose of treating the outcome as interval-valued at all.
Figure (ref) plots $\Theta^I$ for $\tilde x=(1,\mathrm{educ})$, computed via the closed-form directional bounds (ref)--(ref) implied by Proposition (ref). The region is an elongated and negatively sloped polygon, reflecting a tradeoff between the intercept and education coefficient induced by the joint distribution of education and the interval endpoints: selections that generate higher fitted levels at low education values tend to require lower education slopes, and conversely.
Imposing $\mathbb{E}[y^*]=\kappa$ at the empirical mean of the underlying uncoarsened and unweighted income collapses $\Theta^I_\kappa$ to a one-dimensional segment (Figure (ref)), exactly as Proposition (ref) predicts: $\Theta_\kappa$ is a single linear equation in a two-dimensional $\theta$, so its intersection with $\Theta^I$ has area exactly zero, not merely small. Because Proposition (ref) already establishes that this intersection is exactly one-dimensional, we compute its endpoints directly from the support function evaluated at the single direction spanning $\Theta_\kappa$ and its reverse, rather than by sampling many directions and intersecting the resulting halfspaces as in Figure (ref). This generic method presumes a two-dimensional interior point to anchor the construction, which a segment does not have, and can return a spuriously nonzero area from sampling noise alone.
We retarget the object of interest to $\theta_f=Q^{-1}\mathbb{E}[\tilde x f(y^*)]$, $f(y)=y^2$, and distinguish two cases according to which population moment is taken as known, since they have markedly different consequences for the geometry of the identified region.
If only $\mathbb{E}[y^*]=\kappa$ is known, Proposition (ref) guarantees that the corresponding population region is convex under atomlessness. The restriction does not impose a hyperplane on $\theta_f$ because there is no direct algebraic link between $\mathbb{E}[y^*]$ and $\mathbb{E}[\tilde x f(y^*)]$ once $f$ is nonlinear, so the known mean prunes admissible selections without confining $\theta_f$ to a lower-dimensional set. The resulting region remains a genuine two-dimensional band, narrower than the unconstrained $\Theta^I_f$. We do not plot the finite-sample analogue of this case separately.
If instead $\mathbb{E}[f(y^*)]=\kappa_f$ is known, $\Theta^I_{f\mid\kappa_f}$ collapses to a one-dimensional segment (Figure (ref)), by exactly the hyperplane argument of Proposition (ref), now applied to $g=f(y)$ in place of $y$: since $\theta_f$ is itself built from $\mathbb{E}[\tilde x f(y)]$, knowing $\mathbb{E}[f(y^*)]=\kappa_f$ pins $\mathbb{E}[\tilde x]'\theta_f=\kappa_f$ directly. This is the informative case, and it holds regardless of whether $\mathbb{E}[y^*]=\kappa$ is imposed in addition.
Figure (ref) imposes only $\mathbb{E}[y^{*2}]=\kappa_f$, with no restriction on $\mathbb{E}[y^*]$ and the target left at the original $\theta$. Unlike the mean restriction, the $\mathbb{E}[y^{*2}]=\kappa_f$ restriction does not collapse the dimension of the identification region (see Proposition (ref)).\footnote{Areas are computed by sampling the support function over a fine grid of directions and constructing the convex hull of the resulting supporting halfplanes. Proposition (ref) establishes convexity of the population object $\Theta^I_{\kappa_f}$ via Lyapunov's theorem, which requires an atomless probability space; the empirical distribution used here is finite and atomic, so this convexity is not automatically inherited by the finite-sample analog. The reported region should therefore be read as an outer bound on the finite-sample identified set obtained by directional optimization, rather than as a verified reconstruction of that set; each additional sampled direction can only tighten this bound toward the true region.} The restricted region remains two-dimensional but retains only $16.4\%$ of the area of the unrestricted identification region. The restricted identification region is not a line segment since a nonlinear moment has no direct algebraic counterpart in $\theta$-space. This is the same mechanism as the mean-only case of (ref) above, with the roles of the original and transformed outcome reversed.
Extending the specification to $\tilde x=(1,\mathrm{race},\mathrm{educ})$ and imposing $\mathbb{E}[y^*\mid\mathrm{race}]=\kappa(\mathrm{race})$ (see Proposition (ref)) reduces the identified set to a one-dimensional affine segment parametrized by the return to education. For each admissible value of $\theta_{educ}$, Proposition (ref) uniquely determines the corresponding intercept and race coefficient through the affine FWL relation. The unconstrained interval for $\theta_{\mathrm{educ}}$, $[-9.79,\ 32.89]$, narrows to $[-8.52,\ 22.90]$ once race is conditioned on --- $73.6\%$ of the original width (Figure (ref)). The full three-dimensional region, and the exact projection confirming this width, are reported in Appendix (ref). Note that the unconstrained bounds for $\theta_{educ}$ reported here differ slightly from those in Sections (ref) and (ref) because the specification here additionally includes race as a regressor.
Age, binned into five cells, serves as the external variable $v$; it does not enter $\tilde x$. Figure (ref) compares three nested intervals for $\theta_{\mathrm{educ}}$: unconstrained, $[-9.96,\ 32.82]$; known pooled mean, $[-8.91,\ 23.02]$; and conditional on age, $[-8.53,\ 23.02]$ (see Proposition (ref)), a further $1.2\%$ narrower than the pooled interval. The conditional interval is contained in the pooled one, with the contraction concentrated entirely at the lower endpoint (from $-8.91$ to $-8.53$); the upper endpoint is identical. This is not a numerical coincidence: Proposition (ref)'s equality condition holds exactly at this endpoint's direction, since every age cell shares the same optimal cutoff as the pooled problem ((ref), Table (ref)), so the two support values coincide by construction rather than by approximation.
This paper studies how auxiliary population moments sharpen identification of best linear predictors when the outcome is observed only through an interval. Using restricted selection sets, we show that the identifying content of auxiliary information varies sharply with the form of the restriction. A known unconditional mean intersects the original identified region with an affine hyperplane. A known conditional mean given a subvector of the regressors imposes $d_1+1$ exact linear restrictions and reduces the full region to an affine image of the remaining coefficient block. A conditional mean given an external variable further tightens the pooled-mean region by restricting how the available interval width may be allocated across values of that variable. Transformation moments have different effects depending on the target: a known mean of the transformed outcome imposes a hyperplane restriction on its own BLP, whereas the same moment used as auxiliary information about the original outcome generally narrows the region without reducing its dimension.
Across these cases, the value of auxiliary information is determined not merely by how many moments are known, but by how those moments restrict the allocation of latent outcomes within the observed intervals. For the mean and conditional-mean restrictions, we quantify this value through constrained allocation problems for interval width; for transformation restrictions, we provide geometric, support-function, and finite-sample characterizations. The CPS illustration shows that the resulting gains can be quantitatively important: conditioning the latent mean on race removes roughly one quarter of the identified width for the return to education, although the magnitude varies substantially with the information supplied.
Relative to BM2008, whose sharp identification region for the interval-outcome BLP is the starting point here, the contribution is to show that a specific and commonly available kind of information --- population aggregates that already accompany coarsened data for confidentiality reasons --- has exploitable identifying content, and to give closed-form expressions for exactly how much. Coarsening is what creates the identification problem in the first place, but the same disclosure regime that requires coarsening typically also requires publishing exact aggregate invariants, and it is from those invariants, not from any external dataset, that the identifying power studied here is recovered. That said, this is a statement about the specific restrictions we study, not a general theory of how privacy regimes interact with partial identification; the results say what a mean, a conditional mean, or a transformation moment buys, not what an arbitrary disclosed statistic would buy.
The framework also has real limits, and they are not incidental to the results. Every closed form in (ref) relies on $\tilde x$ being exactly observed, so that the identification problem reduces to a fixed Aumann integral rather than a $\theta$-dependent containment check; if the covariates are themselves interval-valued, as in BMM2011, the machinery here does not directly apply. The quantification in (ref) is tied to squared loss --- it is the knapsack representation, not the restricted-selection idea itself, that depends on this --- so the value of a restriction under quantile or other asymmetric losses is not covered by the formulas derived here. We also treat $\Theta^I$ and its restricted analogs purely as population objects; this paper does not establish how their sample analogs behave, so the CPS numbers should be read as an illustration of the geometry rather than as a fitted or standard-error-attached estimate. Restrictions are studied one at a time, and it is not obvious from anything shown here whether the contraction from two simultaneously imposed restrictions is additive, larger, or smaller, since the restrictions can overlap in which selections they each rule out. Finally, throughout the paper the disclosed aggregate is taken as exogenously given; nothing here says whether it was chosen well, or what a custodian trying to balance disclosure risk against downstream identifying power should release instead. Addressing any of these would change the analysis rather than extend it as a footnote, and we have deliberately left them open rather than sketch results we have not derived.
\noindentData Availability: The data used in this paper are drawn from the Current Population Survey Annual Social and Economic Supplement (CPS ASEC), a public-use survey administered by the U.S. Census Bureau and the Bureau of Labor Statistics.
\noindentCode: Code implementing every result below, together with the data extract and every figure (main text and appendix), is available at \url{https://github.com/BehroozMoosavi/Codes/tree/main/PI_with_
\noindentFunding: This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.