EconBase
← Back to paper

Point-identification in multivariate nonseparable triangular models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

109,108 characters · 10 sections · 185 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Point-identification in multivariate nonseparable triangular models

abstractIn this article we introduce a general nonparametric point-identification result for nonseparable triangular models with a multivariate first- and second stage. Based on this we prove point-identification of Hedonic models with multivariate heterogeneity and endogenous observable characteristics, extending and complementing identification results from the literature which all require exogeneity. As an additional application of our theoretical result, we show that the BLP model berry1995automobile can also be identified without index restrictions.

Introduction

Over the last two decades several approaches towards identification of nonseparable triangular models of the form

align*[align* omitted — 45 chars of source]

have been developed, where $Y$ is the outcome, $X$ is an endogenous regressor, $U$ and $\varepsilon$ are latent error terms, and $m$ and $h$ are unknown production functions. Some results focus on point-identification of the second stage (d2015identification d2015identification, torgovitsky2015identification torgovitsky2015identification) while others identify average or marginal effects (blundell2003endogeneity blundell2003endogeneity, chesher2003identification chesher2003identification, imbens2009identification imbens2009identification, schennach2012local schennach2012local, matzkin2016independence matzkin2016independence).

All of the results aiming for point-identification of the second stage require a univariate and strictly increasing first- or second stage (in particular imbens2009identification imbens2009identification, d2015identification d2015identification, torgovitsky2015identification torgovitsky2015identification), which limits their practical applicability in settings with general multivariate heterogeneity like the Hedonic model or the BLP model (matzkin2007heterogeneous matzkin2007heterogeneous and berry2014identification berry2014identification, p. 1754). A generalization of identification results to a multivariate setting without strong artificial functional form assumptions is hence important, in particular for bridging the gap between economic and econometric theory.

We therefore provide a new framework for identification in nonseparable (triangular) models in this article, generalizing the seminal result in matzkin2003nonparametric. This allows us to prove point-identification in nonseparable triangular models while allowing for both first- and second stage to be multivariate, generalizing the point-identification results in torgovitsky2015identification and d2015identification. As the main application of our theoretical result we provide assumptions for the identification of multi-market Hedonic models with endogenous characteristics, complementing and building on the existing results in ekeland2004identification, heckman2010nonparametric, and chernozhukov2014single who all consider Hedonic models with exogenous characteristics. In particular, we provide an answer to the open question in the latter article, asking under which conditions one can nonparametrically identify Hedonic models when observable characteristics are endogenous. In a second application, we also show that the BLP model berry1995automobile can be nonparametrically identified without the need to assume that individual heterogeneity can be captured by an index, complementing the seminal result from berry2014identification.

The article is structured as follows. In section (ref) we give a brief overview of the current state of the literature and contrast our approach to other existing approaches. Section (ref) introduces the main theoretical framework and the general identification result for nonseparable triangular models: we introduce the theoretical framework in section (ref), the main assumptions in section (ref), and the main result in section (ref). Section (ref) contains the two applications of the main result: point-identification of the BLP model without index restrictions as well as the main application concerning the point-identification of Hedonic models with endogenous characteristics and multivariate heterogeneity. Section (ref) concludes. The appendix contains all proofs for the results in the main paper.

The literature on (point-) identification in nonseparable models

In this section we link our result to the literature on nonseparable models in general and nonseparable triangular models in particular while giving an intuitive overview of the standard assumptions in the literature.

Nonseparable triangular models are an extension of nonseparable models of the form $Y=m(X,\varepsilon)$ with exogenous $X$. The literature on identification in these models is large, with several authors seeking to point-identify the production function $m$ (matzkin2003nonparametric matzkin2003nonparametric, matzkin2007nonparametric matzkin2007nonparametric, imbens2007nonadditive imbens2007nonadditive and references therein) while others predominantly aim for identification of average or marginal effects (heckman2005structural heckman2005structural, altonji2005cross altonji2005cross, hoderlein2007identification hoderlein2007identification, chernozhukov2007instrumental chernozhukov2007instrumental, hoderlein2009identification hoderlein2009identification). The literature on identification of such models has been growing ever since and has been fruitfully applied to and extended in different scenarios like single-market Hedonic models (ekeland2004identification ekeland2004identification, heckman2010nonparametric heckman2010nonparametric, chernozhukov2014single chernozhukov2014single) or nonlinear Difference-in-Difference models (athey2006identification athey2006identification, d2013nonlinear d2013nonlinear). These results in turn also give rise to many identification results for simultaneous equation models (matzkin2008identification matzkin2008identification, berry2013identification berry2013identification, blundell2014control blundell2014control, matzkin2015estimation matzkin2015estimation and references therein).

All of these results in one way or another require injectivity of the production function $m$; the standard approach in the majority of approaches is the assumption of strictly increasing and continuous $m$ which in turn also requires that $Y$ and $\varepsilon$ are univariate. In this article we argue that while monotonicity is often a reasonable assumption to make, there is a more general assumption which allows for point-identification of the production function in more general settings: measure-preservation of $m$. In fact, we show that the combination of a unique nonparametric structure in combination with measure preservation of $m$ leads to identification results which are the natural extensions of the of the identification results which rely on strictly increasing and continuous production functions.

Allowing for endogenous $X$ is an important extension of the model, especially for practical purposes as many models of interest take this form. For instance, the BLP model berry1995automobile and Hedonic models with multivariate heterogeneity and endogenous characteristics, which are the two applications of our main result; a multivariate point-identification result is therefore important as the point-identification results in torgovitsky2015identification and d2015identification cannot be applied in these settings as they require a univariate first- and second stage. Our identification result has other possible applications, for example screening models with multidimensional consumer heterogeneity as in aryal2017identifying; the latter introduces an identification result for screening models based on d2015identification and requires strong high level assumptions on the observable distributions as well as strong functional form assumptions on the second stage. In particular, aryal2017identifying assumes the existence of a fixed point for identification, whereas we provide low-level and testable assumptions for a fixed set to exist.

Our general identification result covers four important general aspects which have not or only partially been dealt with in the literature. First, it is the only result for multivariate nonseparable triangular models, and can be applied to a wide variety of settings. Second, we can allow for the most general functional forms on the production functions $m$ and $h$, without being forced to require monotonicity as torgovitsky2015identification, d2015identification, or aryal2017identifying. Third, we prove a new mathematical result which leads to easy-to-check sufficient conditions for identification, generalizing the existing sequencing arguments in torgovitsky2015identification and d2015identification. Fourth, in our result we can allow for a continuous or discrete (or even binary) instrument $Z$, which can be of lower dimension than $X$, analogous to existing approaches; in our setting, we can allow for lower-dimensional $Z$ even when $m$ and $h$ are truly multivariate functions, and not element-wise monotonic.

\paragraph{Notation} The standard measurable space is defined by $(\mathbb{R}^d,\mathscr{B}_{\mathbb{R}^d})$ where $\mathscr{B}_{\mathbb{R}^d}$ is the Borel $\sigma$-algebra. For a random variable $X:\Omega\to \mathbb{R}^d$ in some measure space $(\Omega,\mathscr{A},P)$, the measure $P_X$ is the pushforward measure of $P$ via $X$, i.e. $P_X(E) = P(X^{-1}(E))$ for every Borel set $E\in\mathscr{B}_{\mathbb{R}^d}$. The corresponding distribution function $F_X:\mathbb{R}^d\to[0,1]$ is defined by $F_X(x) = P_X(X\leq x)$. The support of $X$ is $\mathcal{X}$. Conditional distributions are defined by $F_{X|Z=z,W=w}$ with support $\mathcal{X}_{zw}$. The standard partial order on $\mathbb{R}^d$ induced by the positive cone for $x,x'\in\mathbb{R}^d$ is defined as $x\leq x'$ if and only if $x_i\leq x_i'$ for $i=1,\ldots,d$. Based on this the distribution function $F_X$ is strictly increasing if $F_{X|Z=z_i}(x)< F_{X|Z=z_i}(x')$ whenever $x<x'$ in the standard partial order. Throughout, whenever we require the functions $X = h(Z,U)$ and $Y=m(X,\varepsilon)$ to be invertible, we always mean invertibility with respect to the second argument, i.e. between $X$ and $U$ as well as $Y$ and $\varepsilon$, never with respect to $Z$.

Theoretical section: point-identification of multivariate nonseparable triangular models

This is the main theoretical section where we introduce the general framework for the identification of multivariate nonseparable models as well as the main theorem, which generalizes the seminal results of torgovitsky2015identification and d2015identification to multivariate nonseparable triangular models.

The general framework for point-identification

Let us start this section with a general outline of the underlying idea for the identification result. We work within the following model throughout:

align[align omitted — 94 chars of source]

where the covariate $X$ is endogenous, i.e. depends on the unobservable error term $\varepsilon$. $U$ is the unobservable and independent error term in the first stage. $m$ and $h$ are unobservable. The variable $Z$ is an instrument in this model. Throughout, we have to assume that $\varepsilon$ is of the same dimension as $Y$ and $U$ is of the same dimension as $X$. The intuitive reason for this is that we need the production functions $m$ and $h$ to be invertible in $\varepsilon$ and $U$ in order to derive at our point-identification result.\footnote{This is the main difference to other attempts in the literature like kasy2014instrumental, who sought to identify the nonseparable triangular model whilst allowing for a possibly infinite-dimensional unobservable error term of the first stage. We do need to make a dimensionality restriction for our result to hold.}

assumption[Dimensions] The supports of $Y, \varepsilon$, $X$, $U$, and $Z$ satisfy $\text{dim}(\mathcal{Y}) = \text{dim}(\mathcal{E}) = d$, $\text{dim}(\mathcal{X}) = \text{dim}(\mathcal{U}) = k$, and $\mathcal{Z}\subseteq\mathbb{R}^m$ for finite integers $d,k,m$.

We allow for the whole or part of the vector $X\in\mathbb{R}^k$ to be endogenous. In the case where only a part of $X$ is endogenous, it is customary to write the second stage of (ref) as $Y=m(X,W,\varepsilon)$, where $W$ are the exogenous covariates, and condition all results on $W$. In this case the dimension of $U$ has to be reduced to match the dimension of the endogenous variables. For the sake of conciseness we consider all elements of $X$ endogenous and suppress $W$ throughout.

The idea to allow for multivariate first- and second stages is to find a sufficiently general functional requirement for $m$ and $h$, which seminal results like imbens2009identification, torgovitsky2015identification, and d2015identification require to be strictly increasing and continuous. The issue is that there is no complete order in higher dimensions, so that those classical identification results based on matzkin2003nonparametric are not applicable in this setting. We therefore argue that the appropriate generalization of a strictly increasing and continuous function in these settings is a measure preserving isomorphism.

definitionA map $T:\mathcal{E}\to\mathcal{Y}$ transporting a probability measure $P_\varepsilon$ onto another probability measure $P_Y$ is measure preserving if it is measurable\footnote{Measurability of $T$ means that $\mathscr{B}_{\mathbb{R}^d}=T^{-1}\mathscr{B}_{\mathbb{R}^d}$. $\mathcal{E}$ denotes the support of $\varepsilon$.} and \begin{equation} P_Y(E)=P_\varepsilon(T^{-1}(E)) \end{equation} for every set $E$ in the Borel $\sigma$-algebra $\mathscr{B}_{\mathbb{R}^d}$ corresponding to $Y$.\footnote{$T^{-1}(E)$ denotes the set of points $e\in \mathcal{E}$ such that $T(e)\in E$.} If $T$ is invertible and its inverse is also measure preserving, it is called a measure-preserving isomorphism.

A way to check whether a transformation is measure preserving is by checking this property for all half-open rectangles of the form $(a,b]$, $a,b\in\mathbb{R}^d$.\footnote{Half-open rectangles in $\mathbb{R}^d$ are the $d$-fold cartesian product of half-open intervals, i.e. $(a,b]\coloneqq \bigtimes_{i=1}^d (a_i,b_i]$, $a_i,b_i\in\mathbb{R}$ and $a = (a_1,\ldots,a_d)'$, $b=(b_1,\ldots,b_d)'$.}

proposition$T:\mathcal{E}\to\mathcal{Y}$ with $y=T(\varepsilon)$ transporting $P_\varepsilon$ onto $P_Y$ is measure preserving if and only if \begin{equation} P_Y((a,b]) = P_\varepsilon(T^{-1}((a,b])). \end{equation}
proofThis immediately follows from Theorem A.8 in einsiedler2013ergodic and the fact that all rectangles of the form $(a,b]$ form a semi-ring in $\mathscr{B}_Y$ and $\mathscr{B}_X$.

The set of all measure preserving isomorphisms between two distribution functions is large. In particular, notice that a strictly increasing and continuous $T$ in the univariate case must be measure preserving since

equation[equation omitted — 182 chars of source]

where the third equality follows from the fact that $T$ is continuous and strictly increasing. Strictly increasing and continuous functions are special measure preserving maps because they map every interval of the form $(-\infty,y_1]$ to an interval of the same form $T^{-1}((-\infty,y_1])\equiv(-\infty,\varepsilon_1]$ such that both intervals have the same probability (Figure (ref)). This requirement makes strictly increasing and continuous functions unique in the class of measure preserving isomorphisms\footnote{A strictly increasing and continuous function is invertible.} between two fixed distributions $F_X$ and $F_Y$, a property which has first been exploited in matzkin2003nonparametric. On the other hand, measure preservation only requires that an interval $(-\infty,y_1]$ gets mapped to some combination of intervals, a much weaker restriction (Figure (ref)).\footnote{Completely formally, the image of a measure preserving set need not even be an interval, but could be more general, like a Cantor set for example.}

figure[figure omitted — 2,279 chars of source]
figure[figure omitted — 2,885 chars of source]

To make these concepts more intuitive consider $\mathcal{E}$ as well as $\mathcal{Y}$ to be a continuum of individuals, respectively. The probability measure $P_\varepsilon$ then gives the “size” of each (Borel-) subset $E\subset\mathcal{E}$ of individuals, and analogous for $P_Y$. The classical idea for identification using monotonicity is that---in one dimension---requiring the production function $f:\mathcal{E}\to\mathcal{Y}$ to be strictly increasing and continuous means that $f$ preserves the ordering of individuals when mapping from $\varepsilon$ to $Y$. This immediately implies that the map $f$ preserves measure, as a group of people $E\subset\mathcal{E}$ with size $P_\varepsilon(E)$ gets mapped to a group of people $f(E)\subset\mathcal{Y}$ of size $P_Y(f(E))=P_\varepsilon(E)$, simply by the fact that the order of individuals needs to be preserved, which follows from (ref). Since a strictly increasing and continuous $f$ is invertible, this is our characterization of measure preservation. As will become clear, the main requirement for general identification is measure preservation; if $m$ is to be point-identified, we in addition need to require $m$ to be unique. Uniqueness comes from functional form restrictions like (multivariate generalizations of) monotonicity, but monotonicity is simply a sufficient condition for measure preservation and uniqueness, which in turn are a sufficient condition for point-identification of $m$.

On the outset it may seem like a tautology that we require uniqueness of the production function in order to obtain point-identification. Note, however, that there is a distinction between uniqueness and statistical identification, and that the latter does not follow from the former in general. A production function which is theoretically unique need not be identifiable; in fact our identification result below provides the machinery to go from uniqueness to identification in the setting of nonseparable triangular models. In particular, we briefly show below that a linear first-stage relationship of the form $X = \beta Z+U$ with $\mathcal{U}=\mathbb{R}^k$ cannot be identified by our methods for $\beta\neq 0$, which is perfectly analogous to a comment in torgovitsky2015identification.

In this respect note that monotonicity and continuity are also structural assumptions on $m$ which ensure its uniqueness and measure-preservation. It is in this sense that our framework generalizes the framework in matzkin2003nonparametric as we allow for more general nonparametric function classes than just (univariate) strict monotonicity and continuity. So what we require is simply some nonparametric structural assumption which guarantees that the function $m$ is unique under this. There are plenty of these structural assumptions. In fact, one general class of functions comes from the theory of optimal transport and can be used for the identification of Hedonic models with multivariate heterogeneity and endogenous characteristics, our main application. We call those production functions determinable.

definitionA measurable production function $m:\mathcal{E}\to \mathcal{Y}_x$ is determinable if the class of functional form assumptions on $m$ intersected with the class of measure preserving isomorphisms between $P_\varepsilon$ and $P_{Y|X=x}$ contains a unique element for all $x\in\mathcal{X}$.

Requiring $m(x,\varepsilon)$ to be strictly continuous and increasing in $\varepsilon$ is a functional form restriction which makes it determinable, if we in addition rule out the existence of strictly increasing and continuous transformations $g\circ m$ of $m$, because for given $F_\varepsilon$ $m$ and $g\circ m$ are observationally equivalent matzkin2003nonparametric. Let us give some other examples of determinable measure-preserving isomorphisms. The following example is in the univariate case.

example[A matching example on the real line with concave cost functions] mccann1999exact introduces a matching model on the real line $\mathbb{R}$ between a distribution $P_\varepsilon$ of suppliers (e.g. coal mines) of some product and a distribution $P_Y$ of demanders (e.g. factories) of this product. The problem is this model consists of minimizing the transport cost between coal mines and factories. The cost of transportation is modeled as $c(y-e)$, $y\in\mathcal{Y}$, $e\in\mathcal{E}$ for strictly concave $c$. McCann argues that a concave cost function of the transport distance provides a reasonable model for applications in which shipping occurs along a single route, because the resulting shipping routes display economies of scale. He goes on to show that solving this optimal transport problem under the specified concave cost function has a unique measure preserving isomorphism as a solution $Y = m(\varepsilon)$, which is not monotone. Assuming that $m$ is the solution of the above optimal matching for concave costs is a functional form restriction which makes $m$ determinable.

Since $m$ is not strictly increasing, it cannot be identified by current approaches in the literature. As another example, note that multivariate monotonicity also leads to determinability in higher dimensions.

example[Multivariate monotonicity] For absolutely continuous probability measures $P_\varepsilon$ and $P_{Y|X=x}$ with finite second order moments there exists a unique measure preserving isomorphism $m(x,\cdot)$ transporting $P_\varepsilon$ onto $P_{Y|X=x}$ for all $x$, which takes the form of a gradient of a convex function by a famous result in brenier1991polar, i.e. $m(x,\cdot)\coloneqq \nabla\varphi_x(\cdot)$ for convex $\varphi_x$. mccann1995existence generalized Brenier's theorem to show that this result holds even if the probability measures do not have finite second order moments. We have included this result in the appendix (Theorem (ref)) for the sake of completeness. Therefore, assuming that $m(x,\cdot)$ is the gradient of a convex function for all $x\in\mathcal{X}$ is a functional form restriction which makes $m$ determinable.

Gradients of convex functions are the most natural generalization of increasing and continuous functions. In particular, if $T=\nabla\varphi$ for some convex $\varphi$, then it is monotone in the following sense villani2003topics: \[\langle T(x)-T(z),x-z\rangle\geq 0,\] where $\langle\cdot,\cdot\rangle$ denotes the inner product on $\mathbb{R}^d$. Here it is easy to see that if $x>z$ in the partial ordering induced by the positive cone on $\mathbb{R}^d$, then this definition implies that $T(x)\geq T(z)$ or that $T(x)$ and $T(z)$ are not comparable. This monotonicity property has been exploited recently by carlier2016vector who use it to generalize the concept of univariate quantiles.

In general, functional form restrictions will most likely come from solving functional equations in economic theory, like a minimization problem for demand functions, for example. If these functional equations admit a unique invertible solution, then they induce a functional form restriction which make their solution determinable.

example[Demand function] Following matzkin2007heterogeneous, we denote by $V(y,x,\varepsilon)$ the indirect utility function of a consumer over bundles of goods $y$, where $x$ are observable and $\varepsilon$ are unobservable characteristics. Then a demand function can be obtained by \[d(p,I,x,\varepsilon)\coloneqq \operatorname*{\arg\!\min}_{y}\{V(y,x,\varepsilon) : p\cdot y\leq I\},\] for $p$ the price vector and $I$ the initial endowment. The standard assumption is then that $d$ is the unique solution and invertible in $\varepsilon$ (e.g. berry2014identification berry2014identification, p. 1757), which makes $d$ determinable.

Note that in order to identify a determinable $m$ in practice, one needs to make a normalization assumption, usually on the unobservable distribution $F_\varepsilon$, to guarantee that there is only a unique set of $(m,F_\varepsilon)$ which can generate the distribution $F_{Y|X=x}$ for all $x\in\mathcal{X}$. This should be intuitively clear as one in principle needs to identify two things, the unobservable $F_\varepsilon$ as well as the corresponding production function $m$.

Main assumptions for the theoretical main result

We can now lay out the main assumptions for the theoretical identification result. A convenient property of our approach is that one can use the assumptions from torgovitsky2015identification for the first stage $X=h(Z,U)$, i.e. requiring that $h$ can be written element-wise as $h(Z,U) = [h_1(Z,U_1),\ldots,h_k(Z,U_k)]'$ and require the univariate functions $h_i$ to be strictly increasing and continuous in each univariate $U_i$, which makes our approach a direct generalization. This, however, requires the strong assumption of a compact and rectangular support for $\mathcal{X}|Z=z$ for all $z\in\mathcal{Z}$ (torgovitsky2015supplement torgovitsky2015supplement and d2015identification d2015identification).

We propose a general multivariate approach which allows for weaker assumptions on the support of $F_{X|Z=z}$ and $h$, but requires an additional normalization assumption on the first stage, which fixes the distribution of $U$. We assume that the distribution of $U$ is known.\footnote{An alternative approach would be the general control variable approach, which we introduce in gunsilius2017identification.} With this assumption we can make nonparametric functional form assumptions on $h$, like requiring it to be the gradient of a convex function itself.\footnote{It will be important to work with the gradients of convex functions for our main identification result. In particular, the whole identification result rests on an apparently new insight into properties of Brenier's theorem about gradients of convex functions mapping between two multivariate probability distributions, which we prove in Lemma (ref) in the appendix.}

assumptionThe production function $h(z,u)$ in first stage of model (ref) takes the form of the gradient of a convex function between $U$ and $X$ for all $z\in\mathcal{Z}$. Moreover, $U$ has an absolutely continuous distribution $F_U$.

The following assumption is the normalization assumption on $U$.

assumptionThere is a known $\bar{z}\in\mathcal{Z}$ such that $h(\bar{z},u)=u$ for all $u\in\mathcal{U}$.

Assumption (ref) is a standard assumption made in the literature and was first proposed in matzkin2003nonparametric. It requires $h$ to be the identity mapping for a known $\bar{z}$ between $X$ and $U$ and hence fixes the distribution of $U$. This assumption has also been used in other areas, most notably the measurement error literature, where hu2008instrumental require a known functional which fixes the unobservables in a certain rotation. Assumption (ref) allows us to relax the strong assumption of a rectangular and compact support on the observable $F_{X|Z=z}$:

assumptionThe distributions $F_{X|Z=z_i}$ are absolutely continuous with convex support $\mathcal{X}_{z_i}\subseteq\mathbb{R}^k$ for all $z_i\in\mathcal{Z}$ and are strictly increasing.

We allow for $Z$ to be a discrete and even binary instrument.\footnote{We have included the statement and the proof of the theorem for continuous $Z$ in the appendix, because it is not relevant for the identification of multi-market Hedonic models.} In the following we focus on the binary case, i.e. $\mathcal{Z}=z,z'$, because in the proof of the main result we assume $h$ to be the identity for one $z$ and the gradient of a convex function for the other $z'$. Also, for the identification of Hedonic models, two markets are sufficient by definition, and adding more markets would not change the result in any way.

assumption$Z$ is a valid instrument for $X$, i.e. (i) it generates exogenous variation in $X$ such that $F_{X|Z=z}(x)\neq F_{X|Z=z'}(x)$ for at least one $x\in\mathcal{X}$ and (ii) is independent of $\varepsilon$ and $U$, denoted by $(\varepsilon,U)\protect\mathpalette{\protect\independenT}{\perp} Z$.

Note that we do not require the dimensionality of $Z$ to be at least the dimensionality of $X$. In fact, and this is analogous to the result in torgovitsky2015supplement, $Z$ can be lower- and even one-dimensional for multivariate $X$ and the approach still works. This is important for the application to Hedonic models, because the idea is to consider each market to be the realization of a one-dimensional instrument. Intuitively, we can allow for lower-dimensional $Z$ because we make use of the nonlinear and nonparametric structure, i.e. taking into account the information of all higher order and not just first-order moments; this will become apparent momentarily when we introduce Assumption (ref).

The ultimate goal for identification is to point-identify the (multivariate) second stage production function $m$.

assumption$m(x,\varepsilon)$ is a determinable measure preserving isomorphism between $Y$ and $\varepsilon$ and is continuous in $X$ for all $x$. Moreover, $P_\varepsilon$ is known.\footnote{In the univariate case following matzkin2003nonparametric, $P_\varepsilon$ is usually normalized to be the uniform distribution. Since we work with general and multivariate distributions, we only assume it is known. Alternatively, instead of assuming $P_\varepsilon$ is known, one could assume that $m(\bar{x},e)=e$ at some known $\bar{x}\in\mathcal{X}$ for all $e\in\mathcal{E}$, i.e. assuming that it is the identity map for $\bar{x}$, just as for the first stage. Then $P_\varepsilon$ is also fixed, because by measure preservation of $m$ we have $P_{Y|X=\bar{x}}(E) = P_{\varepsilon}(m^{-1}(\bar{x},E)) = P_\varepsilon(E)$ for every Borel set $E$ in $\mathscr{B}_{\mathbb{R}^d}$.}

It turns out, however, that assuming continuity of $m(x,\varepsilon)$ in $x$ is a rather strong assumption. In particular, we will not be able to guarantee this in our efforts to identify the multi-market Hedonic model in the next section. Therefore, we need to make a weaker assumption. This requires the definition of convergence in measure, see for instance bogachev2007measure1.

definitionSuppose we are given a measure space $(\mathcal{X},\mathscr{A}_X)$ with a probability measure $P$ and a sequence of $P$-measurable functions $\{f_n\}_{n\in\mathbb{N}}$. Then the sequence $\{f_n\}_{n\in\mathbb{N}}$ is said to converge in measure to a $P$-measurable function $f$ if for every $c>0$ one has \[\lim_{n\to\infty} P(x:|f_n(x)-f(x)|\geq c)=0.\]

Based on this, we only require $m(x,\varepsilon)$ to be continuous in $x$ in the following weaker sense.

definitionIf for every sequence $\{x_n\}_{n\in\mathbb{N}}\in\mathcal{X}$ which converges to some $x\in\mathcal{X}$ the corresponding sequence $m(x_n,\varepsilon)$ converges in measure to $m(x,\varepsilon)$, we say that $m$ is continuous in measure.

The weaker assumption we need to require for $m$ is hence that it is continuous in measure based on Definition (ref).

assumptions{6}{'} $m(x,\varepsilon)$ is a determinable measure preserving isomorphism between $Y$ and $\varepsilon$ and is continuous in measure in $x$. Moreover, $P_\varepsilon$ is known.

We also have to make a large support assumption on $\varepsilon$. This is the same assumption that torgovitsky2015identification had to impose.

assumptionThe support $\mathcal{E}$ is convex and independent of $X=x$ and $Z=z$, i.e. $\mathcal{E}_{x,z}$ coincides with $\mathcal{E}$ for all $(x,z)\in\mathcal{X}\times\mathcal{Z}$.

Now, in order to identify the model with discrete or binary instruments, the main assumption needs to be made on the distribution functions $F_{X|Z=z_i}$, $i=1,\ldots,k_z$. It is the analogous assumption to the ones made in torgovitsky2015identification, who requires all univariate distribution functions to intersect in at least one point. In the following we consider binary $Z$ with realizations $z$ and $z'$; furthermore, for each pair $z,z'\in\mathcal{Z}$ we denote as $\mathcal{I}(z,z')\subseteq \mathcal{X}_z\cup\mathcal{X}_{z'}$ the set where $F_{X|Z=z}$ and $F_{X|Z=z'}$ intersect. That is, \[\mathcal{I}(z,z')\coloneqq\{x\in\mathcal{X}_z\cup\mathcal{X}_{z'}: 0<F_{X|Z=z}(x)=F_{X|Z=z'}(x)<1\}.\] Moreover, for every point $x_0\in\mathcal{X}_z=\mathcal{X}_{z'}$, we let $I_z(x_0)$ denote the isoquant or level set of $F_{X|Z=z}$ at $x_0$, which is the set of all $x\in\mathcal{X}_z$ which have the same probability as $x_0$ under $F_{X|Z=z}$, formally: \[I_z(x_0)\coloneqq\{x\in\mathcal{X}_z:F_{X|Z=z}(x)=F_{X|Z=z}(x_0)\}.\] Since we assumed that $F_{X|Z=z}$ is absolutely continuous with convex support and strictly increasing in a multivariate sense, we can regard the distribution functions $F_{X|Z=z}$ and $F_{X|Z=z'}$ as utility functions, in which case the isoquants $I_z(\cdot)$ and $I_{z'}(\cdot)$ can be interpreted as the indifference curves (or in higher dimensions: indifference manifolds) of $F_{X|Z=z}$ and $F_{X|Z=z'}$, respectively. This analogy is the key in proving the result as it allows us to work with the indifference curves instead of the distribution functions.

Now for stating the main assumption which gives us identification, we need to introduce the concept of transversal intersection of manifolds. The following definition is adapted from milnor1997topology.

definitionTwo submanifolds $N$ and $N'$ of an ambient manifold $M$ intersect transversally if for each $x\in N\cap N'$ their tangent spaces at $x$, denoted by $T_x N$ and $T_x N'$, together generate the tangent space $T_x M$ in the sense that $T_x N+T_xN'=T_xM$.

Transversal intersection of indifference curves of different utility functions is a standard assumption made in economic theory mas1989theory and is very weak since it is a generic property in the sense that basically all indifference curves between different preferences intersect transversally by a result from Ren\'e Thom (see ekeland2004identification ekeland2004identification for a discussion of generic properties). We are now in the position to state the main assumption on the distributions $F_{X|Z=z}$.

assumptionLet $Z$ be binary with $\mathcal{Z}=\{z,z'\}$. Then the following properties of $F_{X|Z=z}$ and $F_{X|Z=z'}$ hold: \begin{enumerate} • $\mathcal{X}_{z}=\mathcal{X}_{z'}$. • The epigraphs\footnote{The epigraph of a real valued function $f:X\to\mathbb{R}$ for the level $\alpha\in\mathbb{R}$ is defined by $\text{epi}(f;\alpha)\equiv\{(x,\alpha)\in\mathcal{X}\times\mathbb{R}:\alpha\geq f(x)\}$, see aliprantis2006infinite.} of $F_{X|Z=z}$ and $F_{X|Z=z'}$, denoted by $\text{epi}(F_{X|Z=z};\alpha)$ and $\text{epi}(F_{X|Z=z'};\alpha)$ for $\alpha\in[0,1]$, are convex sets. The set of all points where the isoquants meet, $\mathcal{I}(z,z')$, consists of at least one connected manifold $\mathcal{M}(z,z')$. At all points $x_0\in\mathcal{I}(z,z')$ the isoquants either intersect transversally or coincide in a neighborhood $\mathcal{N}(x_0)$ around $x_0$. • The manifold $\mathcal{M}(z,z')$ is such that (i) for each $x\in\mathcal{X}_{z}=\mathcal{X}_{z'}$ there is an $m\in\mathcal{M}(z,z')$ with $0<F_{X|Z=z}(m)=F_{X|Z=z'}(m)<1$ which either dominates or is dominated by $x$. (ii) All points $x\in\mathcal{X}_{z}$ with $F_{X|Z=z}(x)=0$ lie on one side of the manifold and all points where $F_{X|Z=z}(x)=1$ lie on the other, and analogously for $F_{X|Z=z'}$; “lying on one side of the manifold” means that there are no two points $x_1,x_2\in\mathcal{X}_{z}$ with $F_{X|Z=z}(x_1)=F_{X|Z=z}(x_2)=0$ (respectively: $F_{X|Z=z'}(x_1)=F_{X|Z=z'}(x_2)=1$) such that the line $(1-t)x_1+tx_2$ for $t\in[0,1)$ intersects $\mathcal{M}(z,z')$. \end{enumerate}

Part 1 of Assumption (ref) is restrictive as it requires that the supports of all conditional distribution functions coincide. This assumption is slightly stronger than the assumption in the univariate case of torgovitsky2015identification as the distribution functions there only need to intersect but the supports need not coincide. On the other hand, we allow for the supports to be convex and unbounded, a much weaker assumption. Parts 2 and 3 of Assumption (ref) appear to be high level, but are actually rather natural, weak, and easy to check in practice.

To see that Assumption (ref) is a reasonable assumption to make in practice, consider Figure (ref) as an example, where we display the intersection of a bivariate $t$ distribution $F_{X|Z=z}$ with density function

align*[align* omitted — 409 chars of source]

and a bivariate Normal distribution $F_{X|Z=z'}$ with density function \[f_{X|Z=z} = (2\pi)^{-1}|\Sigma'|^{-\tfrac{1}{2}}\exp(-\tfrac{1}{2}(x-\mu)^{T}(\Sigma')^{-1}(x-\mu))\quadfor\thickspace\medspace \mu=(0,0)'\thickspace\medspaceand\thickspace\medspace \Sigma'=

pmatrix[pmatrix omitted — 26 chars of source]

.\]

figure[figure omitted — 181 chars of source]

In this case, the set where all isoquants intersect, $\mathcal{I}(z,z')$, consist of two separate manifolds. The one on the bottom does not satisfy part 4 of Assumption (ref), but the one on top, $\mathcal{M}(z,z')$, does. This is all we need since we only need one manifold to satisfy Assumption (ref). To check part 3 of Assumption (ref) we need to look at the contour plot. Figure (ref) shows that the respective isoquants either intersect transversally (in the interior of the graph) or converge towards one another (at the boundaries) so that they coincide there. Assumption (ref) is hence satisfied which would guarantee point-identification of $m$ in a nonseparable triangular model where $F_{X|Z=z}$ and $F_{X|Z=z'}$ are those two distributions.

We want to mention in this respect that even though the form of $I(z,z')$ depends on $\Sigma$ and $\Sigma'$, identification holds for all combinations of $\Sigma$ and $\Sigma'$ and degrees of Freedom $v$ simply by the fact that both $t$- and Normal distribution have infinite support. That is, in all cases there is a manifold $\mathcal{M}(z,z')\subseteq I(z,z')$ with the required properties simply because the CDFs need to intersect at some point in the infinite support. Note that we have chosen an example where the variance of $F_{X|Z=z}$ does not exist since $v=2$. Our approach still works in this case since the gradient of convex functions mapping one distribution to the other exists (see Theorem (ref) in the appendix). The same holds if we choose two Normal distributions. This makes us confident that Assumption (ref) is satisfied in many important practical applications, not just in two but also higher dimensions. This is why our result holds more generally than the results of torgovitsky2015identification and d2015identification, as the authors there need to assume a compact support in their multivariate settings. That said, there are certainly cases which do not satisfy Assumption (ref). In particular, we cannot allow for the fact that one distribution first-order stochastically dominates the other distribution in a multivariate sense, but this requirement is the same as in torgovitsky2015identification or d2015identification, only in the multivariate case.

figure[figure omitted — 169 chars of source]

In the univariate case, i.e. when $X,U\in\mathbb{R}$, the manifold $\mathcal{M}(z,z')$ reduces to a point in the special case where $\mathcal{X}_z$ and $\mathcal{X}_{z'}$ coincide. In this case we can simply use Torgovitsky's assumption on the first stage which makes our approach a direct generalization of torgovitsky2015identification.

To get a better idea of when Assumption (ref) holds in general, consider Figure (ref) and suppose that the sheet of paper represents $\mathbb{R}^2$ with the standard partial order $x\geq y$ if and only if $x_1\geq y_1$ and $x_2\geq y_2$. Depicted there are three different scenarios for the supports $\mathcal{X}_z$ and $\mathcal{X}_{z'}$ as well as the manifold $\mathcal{M}(z,z')$ in these supports. These pictures are schematic versions of Figure (ref) in that they only depict the supports $\mathcal{X}_z$ and $\mathcal{X}_{z'}$ and the respective manifolds $\mathcal{M}(z,z')$, but not the contour plots. Consider the picture on the left first. Any one of those three depicted manifolds in this example satisfies Assumption (ref) as all three are of dimension 1 and connected; moreover, for each point $x\in\mathcal{X}_z=\mathcal{X}_{z'}$, there is a point in every manifold which either dominates or is dominated by $x$.\footnote{Recall that a point $x$ dominates $x'$, i.e. $x>x'$, if it lies to the “north-east” of $x'$.} Also all points in the support with $F_{X|Z=z}(x)=0$ lie on the same side of the manifolds, and analogously for $F_{X|Z=z'}$. The example in the center violates Assumption (ref) since $x_1$ with $F_{X|Z=z}(x_1)=0$ and $x_2$ with $F_{X|Z=z}(x_2)=0$ lie on opposite sides of $\mathcal{M}(z,z')$.\footnote{Note that $F_{X|Z=z}(x_1)=0=F_{X|Z=z'}(x_2)$, because the rectangles $(-\infty,x_1]$ and $(-\infty,x_2]$ do not intersect the support, so that $F_{X|Z=z}(x_1) = P_{X|Z=z}((-\infty,x_1])=0=P_{X|Z=z}((-\infty,x_2])=F_{X|Z=z}(x_2)$.} Finally, the example on the right violates Assumption (ref) even though there are two connected manifolds which together are such that each point $x$ either dominates or is dominated by some $m$ in one of the manifolds. The problem here is that there is not one manifold alone for which this holds. In fact, there is no point in the top left manifold which either dominates or is dominated by $x_4$, i.e. lies to the north-east or south-west; similarly, there is no point on the bottom right manifold which either dominates or is dominated by $x_3$. In addition, both manifolds violate the requirement that all points with $F_{X|Z=z}(x)=0$ and $F_{X|Z=z'}(x)=0$ lie on one side.

figure[figure omitted — 1,857 chars of source]

Assumption (ref) is also weaker than the currently existing assumptions in another respect. In particular, we only require one manifold $\mathcal{M}(z,z')$ for binary $Z$. In contrast, torgovitsky2015supplement requires that for every endogenous variable $X_i$, $F_{X_i|Z=z}$ and $F_{X_i|Z=z'}$ have to intersect in at least one point, and this for every $i=1,\ldots,d_x$. So in a $k$-variate case, this would actually require $k$-linear curves, each orthogonal to one of the $k$ dimensions, instead of one general curve from Assumption (ref).

All Assumptions are rather weak and can even be checked by estimating the respective distribution functions $\hat{F}_{X|Z=z}$ and $\hat{F}_{X|Z=z'}$ and examining wether their intersection satisfy Assumption (ref). Note, however, that, analogous to torgovitsky2015identification, this support assumption excludes linear relationships like $X=\beta Z+U$ with $\mathcal{U} = \mathbb{R}^k$ for $\mathcal{Z}=\{0,1\}$ with bounded supports, because in those relationships the two conditional distribution functions would simply be a shift of each other and would not intersect for $\beta\neq 0$.

Still, by our above reasoning and the fact that the assumption is satisfied by Normal distributions, we are convinced that Assumption (ref) holds in many practical settings.

The theoretical main result and intuition of the proof

Under the above mentioned assumptions, we can now state the main theoretical result of this article.

theoremLet Assumptions (ref) -- (ref) hold and let $Z$ be discrete with at least two points in its support. Then $m(x,\varepsilon)$ is identified for almost every $x\in\mathcal{X}$ and all $\varepsilon\in\mathcal{E}$ in model (ref). The identified set \[\mathbb{I}\coloneqq \{m\in\mathcal{H}(Y_{xz},\varepsilon_{xz}):(m^{-1}(X,Y),U)\protect\mathpalette{\protect\independenT}{\perp} Z\}\] hence contains an $X$-almost everywhere unique element. Here, $\mathcal{H}(Y_{xz},\varepsilon_{xz})$ denotes the set of all measure preserving isomorphisms between $P_{Y|X,Z}$ and $P_{\varepsilon|X,Z}$ satisfying Assumption (ref). If Assumption (ref) holds in place of (ref), the analogous result holds, but $m$ is then only identified for almost every $\varepsilon\in\mathcal{E}_{xz}$ instead of all $\varepsilon$.

The importance of Assumption (ref) is that it basically does not weaken the result (identification for almost every $\varepsilon$ compared to identification for every $\varepsilon$) compared to Assumption (ref), while being a rather substantial weakening of Assumption (ref). In particular, we can guarantee continuity in probability of $m$ but not full continuity in the next section for identification of Hedonic models.

We have relegated the proof of this result to the appendix, but let us give an outline of the idea. Intuitively, the problem of identification results from the fact that $m$ is the map between $P_\varepsilon$ and $P_{Y|X}$ for exogenous $X$. If $X$ were actually exogenous, we would not need a first stage relationship, because in this case the observable distribution $F_{Y|X}$ is exactly the distribution corresponding to $m$ and we could simply use the observable distribution and a normalization of $F_\varepsilon$ to identify $m$. This is the underlying idea for identification of single market Hedonic models with exogenous characteristics.

Since $X$ is endogenous, however, the observable distribution $F_{Y|X}$ is not the right distribution for identifying $m$. We therefore need an instrument $Z$ which has a nonzero influence on $X$ and is independent of $\varepsilon$. Then for a binary (or discrete) $Z$ with values $z$ and $z'$ and the first stage relationships $X=h(z,U)$ and $X=h(z',U)$, we can use $U$ with known distribution $F_U$ as a control variable in the sense of imbens2009identification, only in a multivariate setting and not requiring $U$ to have a uniform distribution. In fact, note that by the assumption that $h$ is the gradient of a convex function mapping $U$ to $X$ for $z'$ and the identity map between $X$ and $U$ for $z$ and the fact that both $P_{X|Z}$ and $P_U$ are absolutely continuous, $h$ establishes a bijective relation between $X$ and $U$ for $z$ as well as $z'$. That is, for each $u\in\mathcal{U}$ there are two $x,x'\in\mathcal{X}$, possibly coinciding, corresponding to it: $u=h^{-1}(x,z)$ and $u=h^{-1}(x',z')$. If we fix $Z=z$, then the relation is bijective. In the other direction, for every $x\in\mathcal{X}$ there are two $u,u'\in\mathcal{U}$: $x=h(z,u)$ and $x=h(z',u')$. The second crucial ingredient is the measure preservation of $h$. In fact, for every Borel set $E_u\in\mathscr{B}_{\mathbb{R}^k}$ we have $P_U(E_u) = P_{X|Z=z}(h^{-1}(E_u,z))$ and similarly for $z'$, so that the distributions do not change if we condition on $U$ or $X$.

Therefore, for fixed $Z$, conditioning on $U$ is the same as conditioning on $X$. Now the crucial assumption guaranteeing that $Z$ is independent of $\varepsilon$ and $U$ allows us to use the fact that for each $u$ there are two $x,x'$, depending on which realization of $Z$ we use for the map $h$. Then the idea---for all $x\in\mathcal{X}$---is to construct a sequence from $x$ to $u$ via $h^{-1}(\cdot,z)$, and then change to $x'$ via $x'=h(z',u)$, then change to $u'$, and so forth. The key here is that for this sequence starting with any $x$ the distributions $F_{Y|X}$ and $F_\varepsilon$ of the second stage do not change because of the independence of $Z$ and $\varepsilon$. So this sequence induces an exogenous change in $X$ by changing $Z$ which does not affect the distribution of $F_\varepsilon$. By assumption (ref), this sequence must converge and cannot go on forever, because at some point it must be that $x=x'$. This holds for every starting point $x$, so that we can in principle identify $m$ by exogenously varying $Z$.

In practice, we do not observe $F_{Y|X}$ for exogenous $X$, even using the instrument $Z$. Theorem (ref) hence only shows that $m$ is identifiable, and gives us the identification set, but not a constructive way to obtain $m$. This reasoning is perfectly analogous to the result in torgovitsky2015identification and also d2015identification. Note again that the dimension of $Z$ can be smaller and even one-dimensional for this, as long as $Z$ is a valid instrument for each variable in the vector $X$. This works since we require nonlinearities of $m$ by way of Assumption (ref) and hence implicitly take into account the information of all higher order moments as mentioned.

To be slightly more formal: the basic idea is to prove uniqueness of $m$. So it is natural to assume that there are $m$ and $m^*$ as well as corresponding $\varepsilon$ and $\varepsilon^*$ satisfying the assumptions and $Y=m(X,\varepsilon)$ as well as $Y=m^*(X,\varepsilon^*)$. Identifiability can then be proved if $m=m^*$. This is done by showing that the isomorphism \[q(x,z,\cdot)=q(x,\cdot)=m^{-1}(x,m^*(x,\cdot))\] is actually the identity, i.e. that $q(x,e)=e$ for all $x\in\mathcal{X}_z\cup\mathcal{X}_{z'}$ and $e\in\mathcal{E}$, which would imply $m=m^*$ for every $x$. To show that $q(x,e)$ is the identity with respect to $\varepsilon$, it is actually sufficient to show that it is not a function of $x$ by the fact that $F_\varepsilon$ is known and $m$ is determinable. In fact, as $m(x,e)$ is determinable between $F_\varepsilon$ and $F_{Y|X=x}$ for each $x\in\mathcal{X}$, there can be no other $m(x,\cdot)$ of this functional form by definition for each $x$. Therefore, if $q(x,e)$ is only a function of $e$, say $f(e)$, this means that the functional form of $m^{-1}\circ m^*$ does not change with $x$ so that both $m$ and $m^*$ have the same functional form. But since $m$ is determinable, it must be that $m=m^*$. This is the same reasoning as in torgovitsky2015identification, only put in a more general framework. In order to achieve this, Assumption (ref) is crucial, as it guarantees that $P_{\varepsilon|X,Z}=P_{\varepsilon|X}$ and analogously for $\varepsilon^*$.

figure[figure omitted — 662 chars of source]

Therefore, varying $Z$ does not affect the distributions of the unobservables, but does affect $X$, which means that one can get at the exogenous effect of $X$ on $Y$. This is the same reasoning as in the case where $Z$ is absolutely continuous. In the binary case one can use a general sequencing argument as mentioned above. Let us be more specific about this sequencing argument now.

Recall that in the univariate case torgovitsky2015identification uses the measure preserving isomorphism $(x,z)\mapsto (F_{X|Z=z}(x),z)$ to condition $\varepsilon$ on $F_{X|Z}$ instead of $X,Z$ and then applies the monotone rearrangement $T(x)=F^{-1}_{X|Z=z'}(F_{X|Z=z}(x))$ as a map between $F_{X|Z=z}$ and $F_{X|Z=z'}$; this ensures that for every point $x_0\in\mathcal{X}_z\cup\mathcal{X}_{z'}$ $F_{\varepsilon|X=x_0,Z=z'}=F_{\varepsilon|X=Tx_0,Z=z}$ and analogously for $\varepsilon^*$, so that $q$ is the same for all iterations $T^n$. He then shows that in one dimension this iteration converges to a fixed point and can hence show that $q$ is constant for all starting points $x_0$, which by the assumed normalization implies that $m=m^*$.

Now, there are mainly two reasons for why this simple reasoning does not work in a higher dimensional setting. First, the map $(x,z)\mapsto (F_{X|Z=z}(x),z)$ is only invertible in the one-dimensional case, so this simple argument does not work. Our solution for this is Assumption (ref). With this we can write $U=h^{-1}(X,Z)$ since both $P_{X|Z}$ and $P_U$ are absolutely continuous, so that $h$ is invertible, which gives

equation[equation omitted — 131 chars of source]

The second equality in (ref) follows from Assumption (ref). The first equality follows from the following reasoning: the map $\phi: (X,Z) \mapsto (h^{-1}(X,Z),Z)$ is a measure-preserving isomorphism since $z\mapsto z$ is a measure preserving isomorphism and $x\mapsto h^{-1}(x,z)$ is a measure preserving isomorphism for all $z$, so that for every rectangle $E_x\times E_z\equiv(-\infty,x]\times (-\infty,z]\in\mathscr{B}_{\mathbb{R}^{k+m}}$ \[P_{X,Z}(E_x\times E_z) = P_{U,Z}(\phi^{-1}(E_x\times E_z))\equiv P_{U,Z}(h^{-1}(E_x,z)\times E_z).\] Analogously, the map $(\varepsilon,x,z)\mapsto (\varepsilon,h^{-1}(x,z),z)$ is a measure preserving isomorphism for the same reasoning so that for every rectangle $E_\varepsilon\times E_x\times E_z\equiv (-\infty,\varepsilon]\times (-\infty,x]\times (-\infty,z] \in\mathscr{B}_{\mathbb{R}^{d+k+m}}$ \[P_{\varepsilon,X,Z}(E_\varepsilon\times E_x\times E_z) = P_{\varepsilon,U,Z}(\phi^{-1}(E_\varepsilon\times E_x\times E_z))=P_{\varepsilon,U,Z}(E_\varepsilon\times h^{-1}(E_x,z)\times E_z).\] Thus \[P_{\varepsilon|X,Z}(E_\varepsilon)=\frac{P_{\varepsilon,X,Z}(E_\varepsilon\times E_x\times E_z)}{P_{X,Z}(E_x\times E_z)} =\frac{P_{\varepsilon,U,Z}(E_\varepsilon\times h^{-1}(E_x,z)\times E_z)}{P_{U,Z}(h^{-1}(E_x,z)\times E_z)}=P_{\varepsilon|U,Z}(E_\varepsilon).\] The last thing to notice is that conditioning on measure zero events does not cause issues, because $(X,Z) \mapsto (h^{-1}(X,Z),Z)$ is measurable with measurable inverse by definition of a measure preserving isomorphism, so that their $\sigma$-algebras coincide, i.e. $\sigma(U,Z) = \sigma(X,Z)$. We give another formal proof of this fact in Lemma (ref) in the appendix, using disintegrations.\footnote{Note that this conditioning is different from the approach in kasy2014instrumental. Kasy used the mapping $\psi:(X,Z)\mapsto (X,h^{-1}(X,U))$, where he defined the inverse of $h$ is with respect to $X$, which is not invertible since in his case it is a map from $\mathbb{R}^2$ to $\mathbb{R}\times \mathbb{U}$, where $\mathbb{U}$ is the (in Kasy's case possiby infinite dimensional) metric space containing $U$. Therefore, the respective $\sigma$-algebras $\sigma(X,Z)$ and $\sigma(X,h^{-1}(X,U))$ need not coincide. In our case, however, we use the measure preserving isomorphism $\phi(X,Z)= (h^{-1}(X,Z),Z)$, which is measurable with measurable inverse, so that $\sigma(X,Z)$ and $\sigma(h^{-1}(X,Z),Z)$ coincide.}

Second, we need to use a general sequencing argument which is more intricate in higher dimensions. By Assumption (ref) the distribution of $U$ is fixed to be $F_U=F_{X|Z=z}$ for one $z\in\mathcal{Z}$ and we require the map $h^{-1}(x,z)$ to be the identity and the other map $h^{-1}(x,z')$ to be the gradient of a convex function transporting $F_U$ onto $F_{X|Z=z'}$. They key step here then is a new result for the dynamics of measure preserving isomorphisms which take the form of the gradient of a convex function, which we prove in the appendix. Intuitively, we use Assumption (ref) and the strict monotonicity of the $F_{X|Z}$ to show that the gradient of a convex function $T$ mapping $F_{X|Z=z}$ onto $F_{X|Z=z'}$ never crosses $\mathcal{M}(z,z')$ in the sense that for each $x\in\mathcal{X}_z\cup\mathcal{X}_{z'}$ the curve $(1-t)x+tTx$, $t\in(0,1)$ never intersects $\mathcal{M}(z,z')$. With this we can show that $\mathcal{M}(z,z')$ is a fixed set of the iteration $T^nx_0$ for the map $T$ between $F_{X|Z=z}$ and $F_{X|Z=z'}$. The key here is Assumption (ref) which requires that the manifold $\mathcal{M}(z,z')$ lies in the supports $\mathcal{X}_z=\mathcal{X}_{z'}$ in such a way that an iteration of this map converges to $\mathcal{M}(z,z')$ for every point. The proof of Theorem (ref) contains the details.

Note in this respect that if we were to make a different functional form assumption on $h$ we would have to work out the dynamics of a different measure preserving isomorphism which in turn would lead to a different Assumption (ref). Gradients of convex functions are very general and well-behaved as transport maps, however, and it is not likely that one will find a measure preserving map with better properties. Even more importantly, the assumption that $h$ is the gradient of a convex function is the most natural generalization of a strictly increasing and continuous $h$ to the multivariate setting. Lastly, Theorem (ref) in the appendix shows identification in the case where $Z$ is absolutely continuous under weaker assumptions on the supports $\mathcal{X}_z$, but requiring $m$ to be a measure preserving $C^{1}$-diffeomorphism instead of simply being a measure preserving isomorphism, which is much stronger. Its statement is a straightforward generalization of the result in torgovitsky2015supplement.

To conclude this section, we also want to stress again that Theorem (ref) and its absolutely continuous counterpart from the appendix are not constructive identification results in the sense that they do not provide us with the function $m$. They just provide the identified set $\mathbb{I}$ which we prove to contain a single element $m$, just as the univariate result torgovitsky2015identification and d2015identification. There are ways to estimate the function $m$ semi-parametrically like komunjer2010semi or torgovitsky2016minimum, but a fully nonparametric approach is still lacking. This is especially important to keep in mind in the following section where we show identification of the Hedonic model in multiple markets and identification of the BLP-model without index restrictions.

Applications: BLP-, and Hedonic models

In order to showcase the applicability of Theorem (ref) we apply it in two different settings. First to the BLP berry1995automobile model, where we complement the point-identification result of berry2014identification. Second to Hedonic models with multivariate heterogeneity and endogenous characteristics, providing the first identification result in this setting and answering an open question posed in chernozhukov2014single in the process. Let us start with the former.

The BLP model

The BLP model berry1995automobile was introduced for identifying and estimating utility- and cost functions of participants in demand and supply systems of differentiated product markets when only aggregate market share data are available to the researcher. The only article providing results on identification of the BLP model to date is the seminal berry2014identification. In this article the authors need to make a somewhat artificial index restriction, because their identification result relies on univariate identification results from the literature, in particular the identification result in chernozhukov2005iv. Using Theorem (ref) we can generalize their result directly to prove nonparametric identification without the need for the index restriction, complementing their result. Let us start with the demand side.

\paragraph{Demand side} The demand side in this model is obtained by aggregating a continuum of individual discrete choice models in the following way, where we adapt the notation from berry2014identification. Each consumer $i$ in market $t$ chooses a good $j$ from a market $\mathcal{J}_t\coloneqq \{0,1,\ldots,J_t\}$, which consists of a continuum of consumers with total measure $M_t$. A market is formally defined by $(\mathcal{J}_t,\chi_t)$ with $\chi_t\coloneqq (x_t,p_t,\xi_t)$. Here, $x_t=(x_{1t},\ldots,x_{J_tt})$ is a $K\times J_t$ matrix containing the observed and exogenous characteristics of the products in the market. $\xi_t\coloneqq (\xi_{1t},\ldots,\xi_{J_tt})$ contains all of the unobservable characteristics at the product or market level and $p_t\coloneqq(p_{1t},\ldots,p_{J_tt})$ contains observable endogenous characteristics, i.e. those characteristics which are correlated with $\xi_t$ like the price.

Consumer preferences in the BLP model are determined by indirect utilities in the sense that consumer $i$ in market $t$ has conditional indirect utilities $v_{i0t},\ldots,v_{iJ_tt}$. Following berry2014identification, we normalize the outside option $v_{i0t}$ to be zero, i.e. $v_{i0t}=0$ for all $i$ and $t$ and assume that the utilities are independent and identically distributed across consumers and markets with joint distribution function $F_{v}(v_{i1t},\ldots,v_{iJ_tt}|\chi_t).$ Then the standard assumption is that $\operatorname*{\arg\!\max}_{j\in\mathcal{J}}v_{ijt}$ is unique with probability 1, which leads to the following definition of the market shares $s_{jt}$ for each product $j$ in market $t$:

equation[equation omitted — 163 chars of source]

under the normalization $s_{0t} = 1-\sum_{k=1}^J s_{kt}$. Now here is where berry2014identification are forced to introduce the index restriction assumption, because it enables them to write the demand function element-wise for every $j$. In particular, they define a univariate index $\delta_{jt} = \delta_j(x_{jt},\xi_{jt})$ for each product $j$, where $\delta_j$ is a function which is strictly increasing and continuous in the unobservable $\xi_{jt}$ for $x_{jt}$.\footnote{In the main text, they even assume that $\delta_j$ is linear, a much stronger assumption, but they relax this assumption to allow for strictly increasing and continuous $\delta_j$ in the appendix, and this is the result we focus on.} The idea then is to write the demand function $\sigma_j(\chi_t)$ for each $j$ only in terms of $x_{jt}$, $\xi_{jt}$, and $p_t$, element-wise for every $j$, and then identify $\xi_{jt}$ by inverting $\sigma_j$ as well as the index $\delta_j$ to get \[\xi_{jt} = \delta_j^{-1}\left(\sigma_j^{-1}(s_t,p_t),x_{jt}\right),\] requiring both $\sigma_j$ and $\delta_j$ to be strictly increasing and continuous functions in $\xi_{jt}$.

It is exactly here where we can apply the framework from section (ref). In fact, the functional form assumption of strict monotonicity and continuity on the index $\delta_j$ and the demand function $\sigma_j$ is the standard assumption from matzkin2003nonparametric, and simply serves as a tool in order to work with a univariate unobservable for every distribution. Using Theorem (ref), we are able to prove identification of the model without being forced to make any index restrictions, therefore complementing the result in berry2014identification. We do so as follows.

Firstly, we allow for a general demand function $\sigma$ solving the demand problem (ref) over all products $j = 1,\ldots, J$ in the market simultaneously and for all individuals $i$ in the continuum $M_t$, so that $s_{t}=\sigma(\chi_t).$ This is the main difference to the index restriction: berry2014identification allow for multiple products like we do, but they only do so element-wise, i.e. they treat every good separately. We on the other hand can allow for general interactions between the products. The outside option is still fixed as above. Analogous to berry2014identification, we assume that this maximization problem has a unique solution. Note that this is the uniqueness assumption we need to make $\sigma$ determinable. In addition we have to assume that $\sigma$ is invertible between $\xi_t$ and $s_t$, a condition which might be hard to satisfy in practice, but has been the standard assumption in this literature (see matzkin2007heterogeneous for further discussion of this point). berry2014identification require monotonicity and continuity in every element, which is a sufficient condition for invertibility of $\sigma$ and might be even harder to satisfy in practice.

Measure preservation is a natural assumption for BLP models. Recall that $s_t$ is the vector of market shares of every product, which possesses a certain (conditional) probability distribution $P_{s_t|x_t,p_t}$. This distribution is just the distribution of choices of individuals $i$ in the market $t$, so that one can view each point in the support of $s_t$ conditional on $x_t$ and $p_t$, denoted by $\mathcal{S}_{x_t,p_t}$, as the purchase plan of an individual $i$, determining the probability with which this individual is to buy which product $j$ in the market. Then this individual $i$ needs to have a certain evaluation of the products $j=1,\ldots,J$, which is unobservable to the econometrician, i.e. a distribution over the unobservables $x_{jt}$ analogous to the purchase plan of the individual; this can be thought of as giving for each product $j$ a probability of how “important” the respective unobservable $x_{jt}$ of the product is for the individual's choice.

Since all individuals lie on a continuum, it makes more sense to talk about sets of individuals instead of unique individuals. Therefore, every (Borel-) set $E\in \mathcal{S}_{x_t,p_t}$ of individuals with purchase plans $P_{s_t|x_t,p_t}(E)$ must have a corresponding set in $P_{\xi_t}$ which is of the same size, because all evaluations and purchase plans are based on the same set of individuals. But this is exactly the definition of a measure preserving demand function $\sigma$, i.e. we require

equation[equation omitted — 165 chars of source]

Now, again, as in the example of Hedonic models, the need for Theorem (ref) arises from the fact that the $p_t$ are endogenous. To the best of our knowledge, this provides the first instance where nonseparable triangular models can be used for identification of the BLP model. Those results were not possible previously, because they required that the second stage in those models be univariate for point-identification, as argued in berry2014identification. Providing complete identification of the BLP model therefore provides another instance proving how important a multivariate generalization of these identification results really is. In order to apply Theorem (ref), we need to model the first stage relationship

equation[equation omitted — 55 chars of source]

where $z_t$ is a set of instruments which are excluded from the demand model, $u_t$ is a vector of unobservables with non distribution and of the same dimension as $p_t$ and $h$ is a measure preserving isomorphism, which we assume to be the gradient of a convex function transporting the distribution of $u_t$ onto the distribution of $p_t$ for all $t$.

In our setting it is very natural to let $p_{jt}$ be multivariate for every $j$, hence letting $p_t$ be a matrix. All we have to do to make this work is to vectorize the matrix $p_t$ by stacking each column onto one another, i.e. identifying the matrix space $\mathbb{R}^{J_t\times K}$ with $\mathbb{R}^{J_t\cdot K}$, where $K$ is the number of columns for every $p_{jt}$. Note again, that we can allow for instruments $z_t$ which are discrete and lower dimensional than the endogenous variables $p_t$, allowing for binary policy changes. All the instruments need to satisfy is $z_t\perp (u_t,\xi_t)$ and that they have an influence on each element of $p_t$. Let us now state the identification result for the demand side. The observables of the market are $(M_t,x_t,p_t,s_t,z_t)$. The demand side is modeled through (ref) and (ref). Then the following holds.

proposition[General identification of the demand side in the BLP model] In the case where the instruments $z_t$ are absolutely continuous, let the regularity assumptions hold as stated in Theorem (ref) in the appendix. In the case where the instruments $z_t$ are discrete, let Assumptions (ref) -- (ref) hold for $p_t$, $z_t$, $u_t$, and $\xi_t$. Then the model (ref) and (ref) is identified in the sense that the identified set \[\mathbb{I}\coloneqq \{\sigma\in\mathcal{H}(S_{x_t,p_t},\xi_t):(\sigma^{-1}(x_t,p_t,s_t),u_t)\perp z_t\}\] contains an almost everywhere unique element $\sigma$. $\mathcal{H}(S_{x_t,p_t},\xi_t)$ is the set of all isomorphisms between $\xi_t$ and $s_t$ for exogenous $x_t$ and $p_t$.

The proof of this proposition follows immediately from Theorem (ref) or Theorem (ref). Proposition (ref) therefore provides nonparametric identification for the demand side of the BLP model in the most general case, only requiring $\sigma$ to be a measure preserving isomorphism. Note that Assumption (ref) requires a normalization of the demand function. This can be done by assuming a multivariate uniform distribution for $\xi_t$, in the sense that all $\xi_{jt}$ are uniformly distributed, which is the analogue to the normalization in berry2014identification who assume a univariate uniform distribution for every $\xi_{jt}$. In our case, one is actually free to model the dependency structure between the $\xi_{jt}$, however, i.e. one is not required to assume that they are all independently distributed. As for $\sigma$ being a measure preserving isomorphism, this is satisfied as soon as the utility maximization problem has a unique and invertible solution.

The nice thing about Proposition (ref) is that the assumptions of a unique and measure preserving demand function $\sigma$ are natural and can be implied by the set-up of the model. Also notice how our approach allows for multivariate $p_t$ even from the set-up. The last interesting and also important thing to recall is that we can allow for instruments to be of lower dimension than the endogenous variables $p_t$. This is especially important in practice. In fact, it might often be the case that there is a dichotomous shock introduced into the model, possibly through a policy change, which can serve as an instrument. If this policy change is truly independent and is such that it influences all $p_t$, then it alone can serve as a single instrument to identify the whole demand side of the model, under the restriction that Assumption (ref) on the supports of $P_{p_t|z_t=z}$ and $P_{p_t|z_t=z'}$ is satisfied; but this assumption can be checked in higher dimensions and simply be eyeballed in the case where $p_t$ is two-dimensional, an important special case.

\paragraph{Supply side} Having identified the demand side, one can model the supply side in basically two ways. First, one can simply assume that one knows the oligopolistic structure of the supply side in which case one can immediately deduce the vector of marginal costs $mc_t\coloneqq (mc_{1t},\ldots,mc_{J_tt})$ by

equation[equation omitted — 67 chars of source]

since all quantities on the right hand side are observed ($s_t$, $M_t$, $p_t$) or identified ($\sigma$). Note that (ref) is more general than the function proposed in berry2014identification, which is, again, only defined element-wise, i.e. one $\psi_j$ for every product $j$, analogous to their element-wise definition of the demand function $\sigma_j$. In the case for known $\psi$, there is nothing to do from an econometric perspective, as one simply assumes away the problem of identifying the respective monopoly structure, i.e. the function $\psi$. This can be warranted in some cases, where one has additional knowledge on the oligopoly structure. Based on this, one can identify the cost functions

equation[equation omitted — 61 chars of source]

with some instrument (i.e. supply shifter) and Theorem (ref) or Theorem (ref). Here, $w_t\coloneqq(w_{1t},\ldots,w_{J_tt})$ are observable and exogenous cost shifters, and $\omega_t\coloneqq(\omega_{1t},\ldots,\omega_{J_tt})$ are unobservable cost shifters. Let $\mathcal{J}_j$ denote the set of products produced by the firm producing product $j$. Let $q_{jt}=M_ts_{jt}$ be the quantity produced of good $j$ in equilibrium and let $Q_{jt}$ be the vector of quantities of all goods $k\in\mathcal{J}_j$. This setting is completely analogous to the demand side if we replace $w_t\equiv x_t$, $\omega_t\equiv\xi_t$, and $Q_t=p_t$.

The important and more realistic way to model the supply side, however, is to allow for an unknown function $\psi$. Note that in this case, there are several approaches towards identification. In principle, there are four things to identify in the model: $mc_t$, $\psi$, $c$, and the unobservable shocks $\omega_t$. The approach in berry2014identification is to combine (ref) and (ref) into one equation

equation[equation omitted — 140 chars of source]

eliminating $mc_t$ in the process. This approach enables them to identify the unobservable shocks $\omega_t$ as well as the function $\pi$ which incorporates both $c$ and $\psi$. Their approach consists of assuming that both $c$ and $\psi$ can be written element-wise, with each $c_j$ being linear in $w_{jt}$ and $\omega_{jt}$. We can do the same but much more generally with our approach, simply requiring $\psi$ and $c$ to be measure preserving isomorphisms. Then $\pi^{-1}$ must be a measure preserving isomorphism, too. Now realize that (ref) is perfectly analogous to (ref). Therefore, we can apply Theorem (ref) again in order to establish identification. We only need some instrument $z_t$ for the supply side. Then a first stage for the endogenous $p_t$ is

equation[equation omitted — 57 chars of source]

where, again, $u_t$ is unobservable and of the same dimension as $p_t$ and $h$ is the gradient of a convex function, exactly as in (ref). Then we can state the following identification result for the supply side, which we only state in terms of Theorem (ref)---the statement for Theorem (ref) is of course analogous.

propositionIn the case where the instruments $z_t$ are discrete, let Assumptions (ref) -- (ref) hold for $p_t$, $z_t$, $u_t$, and $\omega_t$. Then the model consisting of (ref) and (ref) is identified in the sense that the identified set \[\mathbb{I}\coloneqq \{\pi\in\mathcal{H}(S_{w_t,M_t,p_t},\omega_t):(\pi^{-1}(w_t,M_t,p_t,s_t),u_t)\perp z_t\}\] contains an almost everywhere unique element $\pi$. $\mathcal{H}(S_{w_t,M_t,p_t},\omega_t)$ is the set of all isomorphisms between $\omega_t$ and $s_t$ for exogenous $w_t$ and $p_t$ and fixed $M_t$.

The proof again follows immediately from Theorem (ref). Again, one has to make a normalization assumption, but assuming that $mc_t$ is multivariate uniform is not a strong restriction.

Now, in many cases one is actually interested in identifying both functions $c$ and $\psi$ separately, because $\psi$ gives information about the oligopoly structure of the market. This can again be done in our setting if one has proper instruments for both equations, which in general is not a strong requirement, because one can use standard exogenous shifters as argued in berry2014identification. The idea is to use our identification approach first on (ref) and then on (ref) once we have identified $mc_t$. Let us assume that there are appropriate sets of shifters (i.e. instruments) $z_t^1$ and $z_t^2$ for (ref) and (ref), respectively. Note that we do admit the likely case $z^1_t=z^2_t$. We also need two first stage equations:

equation[equation omitted — 63 chars of source]

and

equation[equation omitted — 63 chars of source]

for unobservables $u_t^1$ and $u_t^2$ with the same dimension as $p_t$ and $Q_t$, respectively, and $h_1$ and $h_2$ are gradients of convex functions. Since $Q_t$ is a matrix, we rely on the simple trick of writing it as a vector, stacking the columns upon one another, as mentioned. The main identification result for the supply side is then as follows. Again, we only state the result for Theorem (ref), but the result for Theorem (ref) is perfectly analogous.

proposition[General identification of the supply side of the BLP model] In the case where the instruments $z^1_t$ and $z^2_t$ are discrete, let the first stages (ref) and (ref) satisfy Assumptions (ref) -- (ref), let $c$ and $\psi$ satisfy Assumption (ref), and let Assumptions (ref) as well as (ref) -- (ref) hold for $p_t$, $Q_t$, $z^1_t$, $z^2_t$, $u^1_t$, $u_t^2$, $mc_t$, and $\omega_t$. Then the two models consisting of (ref) and (ref) as well as (ref) and (ref) are identified in the sense that the identified sets \[\mathbb{I}_1\coloneqq \{\psi^{-1}\in\mathcal{H}(S_{w_t,M_t,p_t},mc_t):(\psi(\sigma,M_t,p_t,s_t),u_t^1)\perp z^1_t\}\] and \[\mathbb{I}_2\coloneqq \{c\in\mathcal{H}(S_{w_t,M_t,Q_t},\omega_t):(c^{-1}(w_t,M_t,Q_t,s_t),u_t^2)\perp z^2_t\}\] contain almost everywhere unique elements $\psi$ and $c$, respectively. As before, $\mathcal{H}(S_{w_t,M_t,p_t},mc_t)$ and $\mathcal{H}(S_{w_t,M_t,Q_t},\omega_t)$ are the respective sets of isomorphisms.
proofFirst, one needs to identify $mc_t$ and $\psi$ in (ref) and (ref). This follows immediately from Theorem (ref) and the fact that $\sigma$ has already been established to be identified on the demand side by Proposition (ref). Then once $mc_t$ is identified, one can turn to the identification of $c$ and $\omega_t$ in (ref) and (ref), which also follows immediately from Theorem (ref) and the fact that $mc_t$ has already been identified.

Propositions (ref) and (ref) together establish general identification of the BLP model, without the need to make index restriction assumptions and hence allowing for the most general set-up. The main assumptions, in addition to regularity assumptions like continuity and invertibility, guaranteeing this result are measure preservation and uniqueness of the demand function $\sigma$ as well as the cost functions $\psi$ and $c$. Both are very natural and fundamental, being a direct result of the model set-up. As a result of the endogeneity of quantity and price, we rewrote both sides of the model as a nonseparable triangular system and applied our two results from the previous section to guarantee point identification of the respective functions. There are certainly other ways for identification of this model than using Theorem (ref) or Theorem (ref), but both are very general and it is rather unlikely that even more general theorems will hold for this setting. Note that this application was just an outline that a different identification result to berry2014identification holds, where one does not have to make index restriction assumptions or assume that the demand function is element-wise strictly increasing and continuous, but can allow for more general results. The discussion in this section was very high level. A next step from here is to find appropriate low level assumptions on the demand function (other than monotonicity) which imples that it is measure-preserving and invertible. Those should come from economic theory. Let us now turn to the main application of our main result where we prove nonparametric identification of Hedonic models with multivariate unobservable heterogeneity and endogenous characteristics.

Hedonic models with endogenous characteristics

Theorem (ref) enables us to prove nonparametric identification of general Hedonic models with multivariate unobservable heterogeneity and endogenous observable characteristics in a multi-market setting, the main result of this whole section.

Seminal identification results in ekeland2004identification and heckman2010nonparametric have focused on single-market Hedonic models with univariate unobservable heterogeneity and stronger functional form assumptions on the utility function. The recent working paper chernozhukov2014single extended these results to single-market Hedonic models with multivariate unobservable heterogeneity using optimal transport theory. We complement these results in this subsection by for the first time allowing for endogenous observable characteristics in a multi-market setting under multivariate unobservables.

We adapt the notation in chernozhukov2014single slightly to make it compatible with the notation from the previous section. We consider an environment where consumers and producers trade a good or contract which is fully characterized by its type or quality $y\in\mathcal{Y}\subseteq\mathbb{R}^d$. Its price, $p(y)$ is determined endogenously in equilibrium. Producers $\operatorname{\widetilde{W}}$ and consumers $\operatorname{\widetilde{X}}$ are characterized by their respective types, $\operatorname{\widetilde{w}}\in\widetilde{\mathcal{W}}\subseteq\mathbb{R}^{m+d}$ and $\operatorname{\widetilde{x}}\in\widetilde{\mathcal{X}}\subseteq\mathbb{R}^{k+d}$. They are price takers and maximize quasi-linear utility $U(\operatorname{\widetilde{x}},y)-p(y)$ and profit $p(y)-C(\operatorname{\widetilde{w}},y)$, respectively, where $U$ is upper semicontinuous and bounded and $C$ is lower semicontinuous and bounded. Both are normalized to zero in case of nonparticipation. We use the following equilibrium concept.

assumption[Equilibrium concept] The pair $(\gamma,p)$, for $\gamma$ a probability measure on $\widetilde{\mathcal{X}}\times\mathcal{Y}\times\widetilde{\mathcal{W}}$ and $p$ a function on $\mathcal{Y}$, is a hedonic equilibrium in the sense that $\gamma$ has marginals $P_{\operatorname{\widetilde{X}}}$ and $P_{\operatorname{\widetilde{W}}}$ and for $\gamma$-almost all $(\operatorname{\widetilde{x}},y,\operatorname{\widetilde{w}})$ \begin{align*} U(\operatorname{\widetilde{x}},y)-p(y) = \max_{y'\in\mathcal{Y}}(U(\operatorname{\widetilde{x}},y')-p(y)),\\ p(y)-C(\operatorname{\widetilde{w}},y) = \max_{y'\in\mathcal{Y}}(p(y)-C(\operatorname{\widetilde{w}},y')). \end{align*} In addition, observed qualities $y=y(\operatorname{\widetilde{x}},\operatorname{\widetilde{w}})$, which maximize the joint surplus $U(\operatorname{\widetilde{x}},y)-C(\operatorname{\widetilde{w}},y)$ for all $(\operatorname{\widetilde{x}},\operatorname{\widetilde{w}})\in\widetilde{\mathcal{X}}\times\widetilde{\mathcal{W}}$, lie in the interior of $\mathcal{Y}$.

For absolutely continuous distributions $P_{\operatorname{\widetilde{X}}}$ and $P_{\operatorname{\widetilde{W}}}$ ekeland2010existence and chiappori2010hedonic show that a pure equilibrium exists and is unique under the twist condition from optimal transport villani2008optimal which requires that the gradients $\nabla_{\operatorname{\widetilde{x}}}U(\operatorname{\widetilde{x}},y)\coloneqq\frac{\partial}{\partial \operatorname{\widetilde{x}}} U(\operatorname{\widetilde{x}},y)$ and $\nabla_{\operatorname{\widetilde{w}}} C(\operatorname{\widetilde{w}},y)\coloneqq\frac{\partial}{\partial \operatorname{\widetilde{w}}} C(\operatorname{\widetilde{w}},y)$ exist and are injective as functions of quality $y$. To obtain an estimable model, unobservable heterogeneity is introduced in the following way.\footnote{We only focus on the consumer's side throughout, because identification of the supplier's side is analogous.}

assumption[Unobservable heterogeneity] The observable type $\operatorname{\widetilde{x}}$ consists of an observable portion $x\in\mathbb{R}^k$ as well as a random unobservable portion $\varepsilon\in\mathbb{R}^d$, i.e. $\operatorname{\widetilde{x}}=(x,\varepsilon)$.\footnote{Recall that $\varepsilon$ is assumed to be of the same dimension as $Y$.} Furthermore, utility $U(\operatorname{\widetilde{x}},y)$ can be decomposed as $U(\operatorname{\widetilde{x}},y) = \bar{U}(x,y) + \xi(x,\varepsilon,y)$.

In this setting, the object of interest for identification is the deterministic component of the utility function, $\bar{U}(x,y)$. Denote $V(x,y) = p(y)-\bar{U}(x,y).$ Here is where chernozhukov2014single need to make the rather strong assumption that the observable characteristics $X$ are exogenous, i.e. $X\protect\mathpalette{\protect\independenT}{\perp}\varepsilon$. The idea for identification then proceeds roughly as follows. The Monge-Kantorovich problem leads to natural functional form restrictions in this setting as the consumer's optimization program is to choose $y$ such that

equation[equation omitted — 408 chars of source]

Identification of $\bar{U}(x,y)$ requires identification of $V(x,y)$; to identify the latter one can use two approaches. One way is to show that the pair $(V(x,y),V^\gamma(x,y))$ uniquely solves the dual problem of the Monge-Kantorovich problem under the cost function $\xi(x,\varepsilon,y)$ in the general case under some regularity assumptions. The second route to identify $V(x,y)$ is via the optimal planner's problem, which takes the form of the Monge-Kantorovich problem with general cost function $\xi(x,\varepsilon,y)$. In the special case $\xi(x,\varepsilon,y)=y'\varepsilon$ considered in ekeland2004identification for example, the cost function for the Monge-Kantorovich problem is the squared Euclidean norm, in which case the inverse demand function $m^{-1}(x,y)$ to take the form of the Brenier map as the determinable measure preserving map between $P_{Y|X}$ and $\varepsilon$ under regularity assumptions. To root this set-up in our notation from above, note that $m^{-1}(x,y)$ is the inverse of the measure preserving map $m:\mathcal{E}\to\mathcal{Y}_x$ from the setting in section (ref).

Allowing for general $\xi(x,\varepsilon,y)$ is a generalization of the concept of convex duality, leading to (ref), see villani2003topics. (ref) possesses a dual problem of the form

equation[equation omitted — 125 chars of source]

analogous to convex duality. For standard convex duality, if $V$ is convex and proper, it holds that $(V^*)^*=V$ rockafellar1997convex. Therefore, if the solution $(V^\gamma)^\gamma$ to the dual problem (ref) coincides with $V$, we say that $V$ is $\gamma$-convex. This idea is also used in chernozhukov2014single in the following assumption, which we also require.

assumption$V(x,y)$ is $\gamma$-convex.

Under some further regularity assumptions chernozhukov2014single then use the fact that the Monge-Kantorovich problem has a unique measure preserving solution, so that $m^{-1}(x,y)$ is determinable in the sense of Definition (ref). Note that it is the assumption of exogenous $X$ which enables them to identify $m^{-1}(x,y)$, because they use the functional form restrictions imposed by the Monge-Kantorovich on $m^{-1}(x,y)$ to identify it, generalizing the identification result from matzkin2003nonparametric in the process. They cannot allow for endogenous $X$, however, and explicitly leave open the problem of identifying it. This is where we can apply Theorem (ref).

For the endogenous case instruments in the form of exogenous shifters are required. One very convenient setting exists if the researcher has access to data in several, i.e. at least two, disjoint markets $z$ and $z'$. This is where Theorem (ref) shines, because one can consider the two markets as realizations of a binary instrument $Z$ under the assumption that $\varepsilon$ does not change between markets, i.e. $Z \protect\mathpalette{\protect\independenT}{\perp} \varepsilon$. Then if the distribution of $X$ changes between the two markets, $Z$ is an instrument for $X$, as needed for Theorem (ref). Formally, we have to include a first stage $X=h(Z,\eta)$, where $h$ satisfies Assumptions (ref) and (ref) for every market $z_i$ and the unobservable $\eta$ is of the same dimension as $X$.

assumption[Multiple markets as a discrete instrument] The distribution of the unobservable $\varepsilon$ does not change between different markets $z$, $z'$, but the distribution of $X$ does.

As before, we assume that $F_\eta$ is known and $h$ is the gradient of a convex function, as done in the proof of Theorem (ref). Now in order to apply Theorem (ref), we need to not only guarantee that $m^{-1}(x,y)$ is the determinable measure preserving map between $P_{Y|X=x,Z=z_i}$ and $P_{\varepsilon|X=x,Z=z_i}$, but also that it is continuous in all $x$. It turns out that under reasonably weak regularity assumptions we cannot guarantee this strong form of continuity; we can guarantee continuity in measure, however, so that we have to use Assumption (ref).

assumption[Regularity assumptions] The following hold: \begin{itemize} • $U(\tilde{x},y)$ and $C(\tilde{y},y)$ are differentiable with respect to $x$ and $w$, respectively, for all $\tilde{x}$ and $\operatorname{\widetilde{w}}$. • $\xi(x,\varepsilon,y)$ is continuously differentiable with respect to $y$ for all $x$, $\varepsilon$, and $y$. • $\det\left(\nabla_{y\varepsilon}^2\xi(x,\varepsilon,y)\right)\neq0$ for all $x,\varepsilon,y$, where $\nabla_{y\varepsilon}^2\equiv (\partial^2\xi / \partial\varepsilon_i\partial y_j)_{ij}$ denotes the Hessian. • Twist condition: For all $x$ and $y$, the gradient $\nabla_y\xi(x,\varepsilon,y)$ of $\xi(x,\varepsilon,y)$ in $y$ is injective as a function of $\varepsilon$. • $h$ is the gradient of a convex function transporting $\eta$ onto $X$. • $P_{\varepsilon}$ is known. • Assumptions (ref) -- (ref) hold for this model. \end{itemize}

Parts (i) -- (iv) of Assumption (ref) guarantee existence and uniqueness of an optimal transport map between $P_{Y|X=x,Z=z_i}$ and $P_{\varepsilon|X=x,Z=z_i}$ for all $x$ and $z_i$ (see e.g. Chapter 10 in villani2008optimal villani2008optimal) and are the same as in chernozhukov2014single. Parts (v) -- (vii) are the assumptions we need to require for Theorem (ref). Overall, our regularity assumptions are only slightly more demanding than the ones of chernozhukov2014single in the exogenous case, but we can allow for endogenous characteristics. Of course, we need to require that $F_{X|Z=z_i}$ satisfy Assumption (ref), which requires that their supports are convex, do not change between markets, and admit an appropriate manifold $\mathcal{M}(z_i,z_j)$. This leads to the following result.

theorem[Identification of utility functions] In the Hedonic model defined in Assumptions (ref) and (ref) assume $\varepsilon\not\protect\mathpalette{\protect\independenT}{\perp} X$. Let the researcher have access to data in at least two disjoint markets $z$ and $z'$ and assume a first stage of the form $X=h(Z,\eta)$ for $\eta,X\in\mathbb{R}^k$. Furthermore, let Assumptions (ref), (ref), and (ref) hold. Then the utility function $\bar{U}(x,y)$ is identified for almost every $(x,y)\in\mathcal{X}\times\mathcal{Y}$ up to an additive constant.

The idea for the proof of Theorem (ref) is to apply Theorem (ref). As stated, the first four parts of Assumption (ref) guarantee existence and uniqueness of the optimal transport map $m^{-1}(x,y)$ as proved in chernozhukov2014single, which is the determinable measure preserving map between $P_{Y|X=x,Z={z_i}}$ and $P_{\varepsilon|X=x,Z=z_i}$ required by Theorem (ref). The additional requirement that this measure preserving map be continuous in measure for every $x$ is guaranteed by the existence of a continuous disintegration and a slight generalization of stability results for the Monge-Kantorovich problem villani2008optimal. The other regularity assumptions then guarantee that $m^{-1}(x,y)$ can be identified for almost every $y$ and $x$ through Theorem (ref). Assumption (ref) is needed to guarantee differentiability of $V(x,y)$, which in turn leads to the identification of $\bar{U}(x,y)$ for almost every $x$ and $y$ through a first order condition since $m^{-1}(x,y)$ is identified for almost every $x$ and $y$. The detailed proof is in the appendix.

Theorem (ref) is the first result in the literature to use data from multiple markets to identify Hedonic models with multivariate unobservable and endogenous characteristics; it therefore answers the open question stated in chernozhukov2014single who ask under which conditions one can identify Hedonic models with multivariate unobservables that are not independent of the observables.

Conclusion

In this article we have proposed a framework for nonparametric point-identification of nonseparable triangular models with a multivariate first- and second stage. The main result is a direct generalization of the seminal results from torgovitsky2015identification and d2015identification for point-identification of nonseparable triangular models with discrete instruments. In particular, we derive primitive conditions on the data generating process under which a nonseparable triangular model with a multivariate first and second stage is identified.

This result is widely applicable. In fact, we use it to derive assumptions under which both the supply and the demand side of the BLP model are nonparametrically identified, even under general heterogeneity. Previously, one had to uphold index restrictions berry2014identification. As a second and main application, we prove the first nonparametric identification result for Hedonic models with endogenous characteristics and multivariate heterogeneity, treating different markets as realizations of a discrete instruments. In particular, this answers an open question in chernozhukov2014single showing when Hedonic models with general heterogeneity are identified under endogeneity. Other possible applications not addressed in this article are to competing risk models lee2013nonparametric or generalized random coefficient models lewbel2017unobserved.

We were able to obtain the theoretical result, because we use and derive some apparently new results in the theory of optimal transport, in particular, we derive some apparently new results about properties of transport maps which take the form of gradients of convex functions, the arguably most natural generalization of a strictly increasing and continuous function to the multivariate setting. In particular, we prove a new result on their dynamics between two absolutely continuous measures whose supports coincide, and defining a criterion for the existence of a fixed set of measure zero. This result builds the heart of our third main theoretical result, but is also of interest in itself as it for instance can be applied to provide a new explanation for how equilibria are formed in General Equilibrium theory.

Finally, our main theoretical identification result is non-constructive as it is a direct generalization of the seminal result of torgovitsky2015identification. There are ways to use this identification result for semi-parametric identification, but a fully nonparametric approach is still lacking. The next important step will hence be to find slightly stronger assumptions than the current ones which would enable nonparametric estimation and inference in these models. Moreover, since the identification result rely upon optimal transport theory, it is imperative to derive further statistical properties of these maps, in particular their large sample distributions.