Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
65,482 characters · 18 sections · 121 citation commands
Identification of hedonic equilibrium and nonseparable simultaneous equations
\address{MIT, NYU and Sciences-Po, Penn State, University of Alberta}
\thispagestyle{empty}
{ Keywords: Hedonic equilibrium, multivariate quantile identification, multidimensional unobserved heterogeneity, cyclical monotonicity, optimal transport.
JEL subject classification: C14, C39, C61, C78 }
Hedonic models were initially introduced by Court:39 to price highly differentiated goods in terms of their attributes. The vast subsequent literature on hedonic regressions of prices on attributes aimed at measuring the marginal willingness of consumers to pay for the attributes of the good they acquired, or the marginal willingness of workers to accept compensation for the attributes of their occupations. When unobservable taste for attributes drives the consumers' choices, however, a simple regression of price on attributes cannot inform us on the willingness to pay for different quality levels from the ones characterizing the good actually acquired. Nor can they inform us on the willingness to pay for characteristics of the good they would acquire under counterfactual market conditions, with different endowments, preferences and technology.
The willingness to pay for counterfactual transactions, together with structural parameters of preferences and technology, can be recovered with a general equilibrium theory of hedonic models, dating back to Tinbergen:1956 and Rosen:1974 (see Heckman:2019 for an account of their respective contributions). The common underlying framework, which we also adopt here, is that of a perfectly competitive market with heterogeneous buyers and sellers and traded product quality bundles and prices that arise endogenously in equilibrium\footnote{When preferences are quasi-linear in price and under mild semicontinuity assumptions, Ekeland:2010 and CMN:2010 show that equilibria exist, in the form of a joint distribution of product and consumer types (who consumes what), a joint distribution of product and producer types (who produces what) and a price schedule such that markets clear for each traded product.}. Rosen:1974 proposes a two-step procedure to estimate general hedonic models and thereby analyze general equilibrium effects of changes in buyer-seller compositions, preferences and technology on qualities traded at equilibrium and their price (see Heckman:99). The first step is a regression of prices on attributes, and the second is a simultaneous equations estimation of the demand and supply system with marginal prices estimated in the first step as endogenous variables.
Brown:83 and BR:1982 point out that changes in consumers' unobserved taste for attributes would lead them to source goods from different suppliers, so that exclusion restrictions from the supply side cannot be justified\footnote{See also Epple:1987, Bartik:1987 and KL:1988.}. The literature on recovering marginal willingness to pay for counterfactual transactions has since followed three strategies: relying on multiple markets across space or time (see Brown:83, BR:1982, KL:1988, TW:2001 and many references in KST:2013); relying on specific functional forms for utility (BB:2005, BT:2011 BT:2011,BT:2019 and references therein); or assume consumers care about a single dimension of good heterogeneity, via a quality index (EHN:2004 EHN:2002,EHN:2002a,EHN:2004, HMN:2010 HMN:2003,HMN:2010, EPS:2010 and EQS:2020). Each of these strategies has drawbacks. Multimarket strategies rely on the assumption of no leakage between markets, or consumption substitution across time and space. They also rely on the assumption that preferences and the distribution of preference types are stable across markets, so that variation comes from the supply side and preferences are identified, but not technology (or vice versa with symmetric assumptions, see EHN:2004). Identification strategies based on specific parameterizations of preferences and technology cannot distinguish features of the specification that are crucial to identification and features that are convenient approximations. Identification proofs must also be repeated for each new parameterization, the suitability of which depends on the application (see the discussion in Yinger:2014). Finally, it is important in many applications to account for heterogeneity in consumers' (or workers') relative valuations of different attributes, and hence move beyond the case of a scalar index of attributes.
This paper proposes an identification strategy based on a single market, where agents have heterogeneous relative valuations for different attributes, without relying on a specific parametric specification of preferences and technology. A leading case in the class of specifications we entertain is $U(x,\varepsilon,z)=\bar U(x,z)+z'\varepsilon,$ where $U$ is the valuation of the bundle of attributes $z$ as a function of the vectors of observable and unobservable consumer characteristics $x$ and $\varepsilon$ respectively. We show that for each choice of distribution of types $\varepsilon$, the function $\bar U$ is recovered nonparametrically (and the same result holds for the supply side). Our contribution is a direct generalization of the main identification strategy in HMN:2010. In the latter, under a single crossing condition\footnote{Also known as Spence-Mirlees or supermodularity condition.} on the utility function, the first order condition of the consumer problem yields an increasing demand function, i.e., quality demanded by the consumer as an increasing function of her unobserved type, interpreted as unobserved taste for quality. Assortative matching guarantees uniqueness of demand, as the unique increasing function that maps the distribution of unobserved taste for quality, which is specified a priori, and the distribution of qualities, which is observed. Hence demand is identified as a quantile function, as in Matzkin:2003. Identification, therefore, is driven by a shape restriction on the utility function.
The main achievement of this paper is to show that a suitable multivariate extension (called {\em twist condition}) of the single crossing shape restriction delivers the same identification result in hedonic equilibrium with multiple good quality dimensions. Heuristically, the proof mirrors HMN:2010 in that is first involves showing identification of inverse demand, which then allows the identification of marginal utility from the first order condition of the consumer's program. The identification of inverse demand, i.e., a single valued mapping from a vector of good qualities to a vector of unobserved consumer type, involves the twist shape restriction and cyclical monotonicity of the hedonic equilibrium solution. The recovery of marginal utility from the first order condition of the consumer's program is involved, because differentiability of the hedonic equilibrium price function is not guaranteed. Known conditions for differentiability of transport potentials in general optimal transport problems, due to MTW:2005 and which would yield differentiability of the hedonic equilibrium price in our context, are very strong and rule out many simple forms of the matching surplus (see Chapter 12 of Villani:2009). We are able to by-pass the MTW:2005 conditions using the special structure of the hedonic equilibrium, and to show approximate differentiability (Definition 10.2 page 218 of Villani:2009) of the price function, for which we need absolute continuity of the distribution of good qualities traded at equilibrium. To that end, we provide a set of mild conditions on the primitives under which the endogenous distribution of qualities traded at equilibrium is absolutely continuous. The proof of absolute continuity of the distribution of qualities traded at equilibrium is based on an argument from FJ:2008, also applied in KP:2014\footnote{FJ:2008 and KP:2014 focus on quadratic distance cost, i.e., $\zeta(x,\varepsilon,z)=-d^2(z,\varepsilon)$ in the notation of Assumption (ref), and work in more exotic geometric spaces.}.
An important special case of our main identification theorem is the case, where the consumer's utility depends on consumer unobserved heterogeneity $\varepsilon$ only through the index $z'\varepsilon$, where $z$ is the vector of good qualities. This case has the appealing interpretation that each dimension of unobservable taste is associated with a good quality dimension and the appealing feature that marginal utility is characterized as the solution of a convex program. However, choosing the dimension of unobserved consumer heterogeneity to be equal to the dimension of the vector of good qualities is a somewhat arbitrary modelling choice, and we provide an extension of our main identification theorem to cases where the dimension of unobserved consumer heterogeneity is lower, including a model with scalar unobserved heterogeneity. We derive a local identification result under mild conditions, but for global identification, we need a shape restriction on the endogenous price function, for which we know of no sufficient conditions on primitives. Another restrictive aspect of our main result is the necessary normalization of the distribution of unobserved heterogeneity, when identifying primitives from a single market. We provide some relaxation of this constraint when data from multiple markets is available, but the results are still fragmentary.
The analysis of identification of inverse demand in hedonic equilibrium reveals that inverse demand satisfies a multivariate notion of monotonicity, whose definition depends on the form of the utility function that is maximized. In the univariate case, this notion reduces to monotonicity of inverse demand. In the case, where the consumer's utility depends on consumer unobserved heterogeneity $\varepsilon$ only through the index $z'\varepsilon$, where $z$ is the vector of good qualities, inverse demand is the gradient of a convex function. We show that this notion of multivariate monotonicity is a suitable shape restriction to identify nonseparable simultaneous equations models, generalizing the quantile identification method of Matzkin:2003 and complementing results in Matzkin:2015, where monotonicity is imposed equation by equation\footnote{Not all results in Matzkin:2003 and HMN:2010 require normalization of the distribution of unobserved heterogeneity, as we do here.}.
On the identification of multi-attribute hedonic equilibrium models, EHN:2004 require marginal utility (resp. marginal product) to be additively separable in unobserved consumer (resp. producer) characteristic. HMN:2010 show that demand is nonparametrically identified under a single crossing condition and that various additional shape restrictions allow identification of preferences without additive separability (see also HMN:2003 HMN:2003,HMN:2005). EHN:2004 emphasize the one-dimensional case but argue that their results can be extended to multivariate attributes under a separability assumption in the utility (without restricting the distribution of unobserved heterogeneity). Our paper directly follows HMN:2010 and generalizes the insight therein to allow for heterogeneity in response to different dimensions of amenities or qualitites. However, the analyses in HMN:2010 and ours are non nested. HMN:2010 considers several specifications with Barten scales that are outside the scope of our generalization. Nesheim:2013 (developed independently and concurrently) is the most closely related paper and complements our work. He achieves identification under an additive separability restriction, but without restricting the distribution of unobserved heterogeneity. He also imposes conditions from MTW:2005 to obtain differentiability of the price. CMN:2010 derive a matching formulation of hedonic models and thereby highlight the close relation between empirical strategies in matching markets and in hedonic markets. CMP:2016 have a section on identification (posterior to this paper), where they use a similar strategy. Two main differences arise from the difference between matching and hedonic models. On the one hand, CMP:2016 need an extra step to account for the fact that in their matching model the price function is not observed. On the other hand, they do not have to worry about regularity of endogenous objects such as the price and distribution of goods in the hedonic equilibrium setting. GS:2012 extend the work of CS:2006 and identify preferences in marriage markets, where agents match on discrete characteristics, as the unique solution of an optimal transport problem, but unlike the present paper, they are restricted to the case with a discrete quality space. The strategy is extended to the set-valued case by CGS:2013, who use subdifferential calculus to identify dynamic discrete choice problems. DGH:2014 use network flow techniques to identify discrete hedonic models.
On the identification of nonlinear simultaneous equations models, Matzkin:2015 uses equation by equation monotonicity in the one dimensional unobservables and exclusion restrictions. BH:2018 consider transformation models, as in Matzkin:2008 (see also Matzkin:2013). These strategies do not require normalization of the distribution of unobserved heterogeneity. SSS:2018 also (independently and in a very different context) use cyclical monotonicity for identification in panel discrete choice models. EGH:2012 and GH:2012 propose a notion of multivariate quantile based on Brenier's Theorem. CCG:2014 (coetaneous with the present paper) propose a conditional version of the optimal transport quantiles of EGH:2012 and GH:2012 and apply it to quantile regression, whereas CGHH:2017 (also coetaneous with this paper) apply optimal transport quantiles to the definition of statistical depth, ranks and signs. However, these papers do not consider identification. The present work is, to the best of our knowledge, the first to apply the notion of multivariate quantiles based on optimal transport results to the identification of simultaneous equations, thus providing a multivariate extension of Matzkin:2003's quantile identification idea.
The remainder of the paper is organized as follows. Section (ref) sets the hedonic equilibrium framework out. Section (ref) gives a brief account of the methodology and main results on nonparametric identification of preferences in single attribute hedonic models, mostly drawn from EHN:2004 and HMN:2010. Section (ref) shows how these results and the shape restrictions that drive them can be extended to the case of multiple attribute hedonic equilibrium markets. Section (ref) derives multivariate shape restrictions to identify nonseparable simultaneous equations models. The last section concludes. Proofs of the main results are relegated to the appendix, as are necessary background definitions and results on optimal transport theory and a list of our notational conventions.
We consider a competitive environment, where consumers and producers trade a good or contract, fully characterized by its type or quality $z$. The set of feasible qualities $Z\subseteq \mathbb{R}^{d_z}$ is assumed compact and given a priori, but the distribution of the qualities actually traded arises endogenously in the hedonic market equilibrium, as does their price schedule $p(z)$. Producers are characterized by their type $\tilde y\in \tilde Y\subseteq \mathbb{R}^{d_{\tilde y}}$ and consumers by their type $\tilde x\in \tilde X\subseteq \mathbb{R}^{d_{\tilde x}}$. Type distributions $ P_{\tilde x}$ on $\tilde X$ and $P_{\tilde y}$ on $\tilde Y$ are given exogenously, so that entry and exit are not modelled. Consumers and producers are price takers and maximize quasi-linear utility $U(\tilde x,z)-p(z)$ and profit $p(z)-C(\tilde y,z)$ respectively. Utility $U(\tilde x,z)$ (respectively cost $C(\tilde y,z)$) is upper (respectively lower) semicontinuous and bounded. In addition, the set of qualities $Z(\tilde x,\tilde y)$ that maximize the joint surplus $U(\tilde x,z)-C(\tilde y,z)$ for each pair of types $(\tilde x,\tilde y)$ is assumed to have a measurable selection. Then, Ekeland:2010 and CMN:2010 show that an equilibrium exists in this market, in the form of a price function $p$ on $Z$, a joint distribution $P_{\tilde xz}$ on $\tilde X\times Z$ and $P_{\tilde yz}$ on $ \tilde Y\times Z$ such that their marginal on $Z$ coincide, so that market clears for each traded quality $z\in Z$. Uniqueness is not guaranteed, in particular prices are not uniquely defined for non traded qualities in equilibrium. Purity is not guaranteed either: an equilibrium specifies a conditional distribution $P_{z|\tilde x}$ (respectively $P_{z|\tilde y}$) of qualities consumed by type $\tilde x$ consumers (respectively produced by type $\tilde y$ producers). The quality traded by a given producer-consumer pair $(\tilde x,\tilde y)$ is not uniquely determined at equilibrium without additional assumptions.
Ekeland:2010 and CMN:2010 further show that a pure equilibrium exists and is unique, under the following additional assumptions. First, type distributions $P_{\tilde x}$ and $P_{\tilde y}$ are absolutely continuous. Second, gradients of utility and cost, $\nabla_{\tilde x}U(\tilde x,z)$ and $\nabla_{\tilde y}C(\tilde y,z)$ exist and are injective as functions of quality $z$. The latter condition, also known as the {\em Twist Condition} in the optimal transport literature, ensures that all consumers of a given type $\tilde x$ (respectively all producers of a given type $ \tilde y$) consume (respectively produce) the same quality $z$ at equilibrium.
The identification problem consists in the recovery of structural features of preferences and technology from observation of traded qualities and their prices in a single market. The solution concept we impose in our identification analysis is the following feature of hedonic equilibrium, i.e., maximization of surplus generated by a trade.
Given observability of prices and the fact that producer type $\tilde y$ (respectively consumer type $ \tilde x$) does not enter into the utility function $U(\tilde x,z)$ (respectively cost function $C(\tilde y,z)$) directly, we may consider the consumer and producer problems separately and symmetrically (see EHN:2002a). We focus on the consumer problem and on identification of utility function $U(\tilde x,z)$. Under assumptions ensuring purity and uniqueness of equilibrium, the model predicts a deterministic choice of quality $z$ for a given consumer type $ \tilde x$. We do not impose such assumptions, but we need to account for heterogeneity in consumption patterns even in case of unique and pure equilibrium. Hence, we assume, as is customary, that consumer types $\tilde x$ are only partially observable to the analyst. We write $\tilde x=(x,\varepsilon)$, where $x\in X\subseteq\mathbb{R} ^{d_x}$ is the observable part of the type vector, and $\varepsilon\in \mathbb{R}^{d_\varepsilon}$ is the unobservable part. We shall make a separability assumption that will allow us to specify constraints on the interaction between consumer unobservable type $\varepsilon$ and good quality $z$ in order to identify interactions between observable type $x$ and good quality $z$.
The main primitive object of interest is the deterministic component of utility $\bar U(x,z)$. For convenience, we shall use the transformation $V(x,z):=p(z)-\bar U(x,z)$. The latter will be called the consumer's {\em potential}, in line with the optimal transport terminology. Since the price is assumed to be identified, identification of $V$ is equivalent to identification of $\bar U$. To achieve identification, we require fixing a choice of function $\zeta$: a leading example, discussed in Section (ref), is the case $\zeta(x,\varepsilon,z)=z'\varepsilon$. Identification also requires fixing the conditional distribution $P_{\varepsilon\vert x}$ of unobserved heterogeneity. This corresponds to the normalization of the distribution of scalar unobservable utility in existing quantile identification strategies. Discrete choice models also typically rely on a fixed distribution for unobservable heterogeneity (generally extreme valued). The requirements to fix both $\zeta$ and $P_{\varepsilon\vert x}$ will be relaxed to some extent in Section (ref) which entertains the possibility of further identification power using information from multiple markets.
In this section, we recall and reformulate results of HMN:2010 on identification of single attribute hedonic models. Suppose, for the purpose of this section, that $ d_z=d_\varepsilon=1$, so that unobserved heterogeneity is scalar, as is the quality dimension. Suppose also that $\zeta$ is twice continuously differentiable in $z$ and $\varepsilon$.
Suppose further (for ease of exposition) that $V$ is twice continuously differentiable in $z$. The main identifying assumption is a shape restriction on utility called single crossing, Spence-Mirlees or supermodularity, depending on the context.
The first order condition of the consumer problem yields
which, under Assumption (ref), implicitly defines an inverse demand function $z\mapsto\varepsilon(x,z)$, which specifies which unobserved type consumes quality $z$. Combining the second order condition $ \zeta_{zz}(x,\varepsilon,z)<V_{zz}(x,z)$ and further differentiation of ((ref)), i.e., $\zeta_{zz}(x,\varepsilon,z)+\zeta_{\varepsilon z}(x,\varepsilon,z)\varepsilon_z(x,z)=V_{zz}(x,z)$, yields
Hence the inverse demand is increasing and is therefore identified as the unique increasing function that maps the distribution $P_{z|x}$ to the distribution $P_{\varepsilon\vert x}$, namely the quantile transform. Denoting $F$ the cumulative distribution function corresponding to the distribution $P$, we therefore have identification of inverse demand according to the strategy put forward in Matzkin:2003 as:
The single crossing condition of Assumption (ref) on the consumer surplus function $\zeta(x,\varepsilon,z)$ yields positive assortative matching, as in the Becker:1973 classical model. Consumers with higher taste for quality $\varepsilon$ will choose higher qualities in equilibrium and positive assortative matching drives identification of demand for quality. The important feature of Assumption (ref) is injectivity of $\zeta_z(x,\varepsilon,z)$ relative to $\varepsilon$ and a similar argument would have carried through under $\zeta_{z\varepsilon}(x,\varepsilon,z)<0 $, yielding negative assortative matching instead.
Once inverse demand is identified, the consumer potential $V(x,z)$, hence the utility function $\bar U(x,z)$, can be recovered up to a constant by integration of the first order condition ((ref)):
We summarize the previous discussion in the following identification statement, originally due to HMN:2010.
Unlike the demand function, which is identified without knowledge of the surplus function $\zeta$, as long as the latter satisfies single crossing (Assumption (ref)), identification of the preference function $\bar U(x,z)$ does require a priori knowledge of the function $\zeta$.
This section develops the multivariate analogue of identification results in Section (ref). The strategy follows the same lines. First, a shape restriction on the utility function, analogue to Assumption (ref), and a consequence of maximization behavior, called {\em cyclical monotonicity}, will identify the inverse demand: we will show that a single type $\varepsilon(x,z)$ chooses good quality $z$ at equilibrium. Second, formally, the utility is then recovered from the first order condition of the consumer's program. The latter step, however, involves significant difficulties, due to the possible lack of differentiability of the endogenous price function.
In the one dimensional case, identification of inverse demand was shown under the single crossing Assumption (ref). We noted that the sign of the single crossing condition was not important for the identification result. Instead, what is crucial is the following, weaker, condition, which is commonly known as the {\em Twist Condition} in the optimal transport literature. The crucial condition, maintained throughout, is Assumption (ref). As shown in CMN:2010, it holds when for each distinct pair of consumer types $(\varepsilon_1,\varepsilon_2)$, the function $z\mapsto\zeta(x,\varepsilon_1,z)-\zeta(x,\varepsilon_2,z)$ has no critical point. It is satisfied for example in the case $\zeta(x,\varepsilon,z)=z'\varepsilon$, in the case $\zeta(x,\varepsilon,z)=F(x,z'\varepsilon),$ when $z'\varepsilon>0,$ and $F$ is increasing and convex in its second argument, or in the case $\zeta(x,\varepsilon,z)=\sum_{k=1}^{d}F_k(x,\varepsilon_k,z_k),$ where, for each $k$ and $x$, $F_k$ is supermodular in $(\varepsilon_k,z_k)$. All these examples are discussed in Sections (ref) and (ref) below.
From GN:65, it is sufficient that $D^2_{z\varepsilon}\zeta(x,\varepsilon,z)$ be positive definite everywhere for Assumption (ref) to be satisfied. Alternative sets of sufficient conditions are given in Theorem 2 of Mas-Colell:79\footnote{Relatedly, BGH:2013 propose injectivity results under a gross substitutes condition.}. Assumption (ref), unlike single crossing, is well defined in the multivariate case, and we shall show, using recent developments in optimal transport theory, that it continues to deliver the desired identification in the multivariate case.
An important implication of Assumption (ref) is that traded quality $z$ maximizes the joint surplus\footnote{This observation is the basis for the characterization of hedonic models as transferable utility matching models in CMN:2010.} $U(\tilde x,z)-C(\tilde y,z)$. Let, therefore, $S(x,\varepsilon,\tilde y):=\sup_{z\in Z}[U((x,\varepsilon),z)-C(\tilde y,z)]$ be the surplus of a consumer-producer pair $((x,\varepsilon),\tilde y)$ at equilibrium. Suppose consumer $(x,\varepsilon_0)$ (resp. $(x,\varepsilon_1)$) is paired at equilibrium with producer $\tilde y_0$ (resp. $\tilde y_1$) to exchange good quality $z$. Then, as shown in the proof of Lemma (ref) in the Appendix, their total surplus is at least as large in the current consumer-producer matching as it would be, were they to switch partners. This is a property of the optimal allocation called {\em cyclical monotonicity}. The total surplus cannot be improved by a cycle of reallocations of consumers and producers. Applied here to cycles of length two, cyclical monotonicity yields:
where the first inequality holds because of cyclical monotonicity and the second inequality holds by definition of the surplus function $S$ as a supremum. Hence, equality holds throughout in the previous display. Therefore, choice $z$ maximizes both $z\mapsto U((x, \varepsilon_0), z) -C(\tilde y_0,z)$ and $z\mapsto U((x, \varepsilon_1), z) -C(\tilde y_0, z)$. It follows that $\nabla_z\zeta(x,\varepsilon_1,z)=\nabla_z\zeta(x,\varepsilon_0,z)$. The Twist Assumption (ref) then yields equality of $\varepsilon_0$ and $\varepsilon_1$ and the following lemma (proved formally in the Appendix).
We see from Lemma (ref) that identification of inverse demand holds under conditions that are analogous to the scalar case, where the Twist condition replaces Spence-Mirlees as a shape restriction. We also see that the identification proof also relies on a notion of monotonicity. We will push this analogy further in the Section (ref) and show that the inverse demand function $z\mapsto\varepsilon(x,z)$ itself satisfies a generalized form of monotonicity, we call $\zeta$-monotonicity, since its definition involves the function $\zeta$.
Heuristically, once identification of inverse demand is established, and $\varepsilon(x,z)$ is uniquely defined in equilibrium, the first order condition of the consumer's problem $\sup_{z\in Z}\{\zeta(x,\varepsilon,z)-V(x,z)\}$ delivers identification of marginal utility $\nabla_z\bar U(x,z)$. However, using the first order condition presupposes smoothness of the potential $V$, hence of the endogenous price function $z\mapsto p(z)$. Conditions for differentiability of the potential $V$ in optimal transportation problems are given by MTW:2005. They are applied to identification in hedonic models in Nesheim:2013. However, the conditions in MTW:2005 are not transparent and they are known to be very strong, excluding simple forms of the surplus, such as $S(\tilde x,\tilde y):=\vert \tilde x-\tilde y\vert^p$, for all $p\ne2$ for instance. The remainder of this section, therefore, will be devoted to proving a weaker form of differentiability of the price function, which can be used to identify marginal utility from the first order condition of the consumer's problem.
The consumer's problem yields the expression for indirect utility:
Equation ((ref)) defines a generalized notion of convex conjugation, in which the consumer's indirect utility is the conjugate of $V(x,z)=p(z)-\bar U(x,z)$. This notion of conjugation can be inverted, similarly to convex conjugation, into:
where $V^{\zeta\zeta}$ is called the {\em double conjugate} of $V$. In the special case $\zeta(x,\varepsilon,z)=z'\varepsilon$, $\zeta$-conjugate simplifies to the ordinary convex conjugate of convex analysis and convexity of a lower semi-continuous function is equivalent to equality with its double conjugate. Hence, by extension, equality with its double $\zeta$-conjugate defines a generalized notion of convexity (see Definition 2.3.3 page 86 of Villani:2003).
We establish $\zeta$-convexity of the potential as a step towards a notion of differentiability, in analogy with convex functions, which are locally Lipschitz, hence almost surely differentiable, by Rademacher's Theorem (see for instance Villani:2009, Theorem 10.8(ii)). It also delivers a notion of $\zeta$-monotonicity for inverse demand, discussed in Section (ref).
This result only provides information for pairs $(x,z)$, where type $x$ consumers choose good quality $z$ in equilibrium. In order to obtain a global smoothness result on the potential $V$, we need conditions under which the endogenous distribution of good qualities traded in equilibrium is absolutely continuous with respect to Lebesgue measure on $\mathbb R^{d_z}$. They include absolute continuity of the distribution of unobserved heterogeneity, additional smoothness conditions on preferences and technology and the Twist Assumption (ref)(B), which requires the dimension of unobserved heterogeneity to be the same as the dimension of the good quality space, i.e., $d_\varepsilon=d_z$ (this will be relaxed in Section (ref)).
We then obtain our principal intermediate lemma, which is of independent interest for the theory of hedonic equilibrium. The proof can be found in the technical appendix CGHP:2020app.
Lemma 4 in the technical appendix CGHP:2020app shows everywhere differentiability of the double conjugate potential $V^{\zeta\zeta}$. This doesn't imply differentiability of $V$ everywhere, since $V$ is $\zeta$-convex (i.e., $V=V^{\zeta\zeta}$) only $P_{z\vert x}$ almost everywhere. However, combined with Lemma (ref), this yields approximate differentiability of $z\mapsto V(x,z)$ as defined in Definition 6 of the technical appendix CGHP:2020app. Using uniqueness of the approximate gradient of an approximately differentiable function then yields identification of marginal utility $\nabla_z\bar U(x,z)$ from the first order condition
where $\nabla_{ap,(z)}$ denotes the approximate gradient (with respect to $z$) of Definition 6 of the technical appendix CGHP:2020app, and where the inverse demand function $\varepsilon(x,z)$ is uniquely determined, by Lemma (ref). This yields our main identification theorem.
Theorem (ref)(1) provides identification of marginal utility without any restriction on the distributions of observable characteristics of producers and consumers. The latter may include discrete characteristics. Regularity conditions in Assumption (ref) are satisfied in the cases $\zeta(x,\varepsilon,z)=\varepsilon'z$ and $\zeta(x,\varepsilon,z)=\exp(\varepsilon'z)$ as we discuss in Sections (ref) and (ref) respectively. They preclude bunching of consumers at equilibrium, as shown in Lemma (ref). The result is driven by the shape restriction (ref) and the strong normalization assumption on the distribution of unobserved heterogeneity. This assumption is inevitable in a single market identification based on a generalized quantile identification strategy. Section (ref) discusses (partial) identification without knowledge a priori of the distribution of unobserved heterogeneity. Theorem (ref)(2) provides a framework for estimation and inference on the identified marginal utility based on new developments in computational optimal transport (see for instance the survey in PC:2019).
A leading special case of the identification result in Theorem (ref) is the choice $\zeta(x,\varepsilon,z)=z^{\prime}\varepsilon$, where marginal utility is linear in unobservable taste. A natural interpretation of this specification is that each quality dimension $z_j$ is associated with a specific unobserved taste intensity $\varepsilon_j$ for that particular quality dimension. Assumptions (ref) and (ref)(2,3) are automatically satisfied when $\zeta(x,\varepsilon,z)=\varepsilon'z$. In addition, $\zeta$-convexity reduces to traditional convexity, so that we have the following corollary of Theorem (ref):
A significant computational advantage of Corollary (ref)(2) over the general case is that the potential solves a convex program. The first order condition ((ref)) also simplifies to $\varepsilon(x,z)=\nabla_{ap,z}\;V(x,z)$, where $\nabla_{ap,z}$ denotes the approximate gradient with respect to $z$ (Definition 6 of the technical appendix CGHP:2020app). Hence, inverse demand in this case is the approximate gradient of a $P_{z\vert x}$ almost surely convex function. This can be interpreted as a multivariate version of monotonicity of inverse demand and assortative matching, since in the univariate case, we recover monotonicity of inverse demand as in Section (ref).
In the special case of Section (ref), the curvature of marginal willingness to pay for qualities varies only with observable characteristics, but is independent of unobservable type $\varepsilon$. Two ways Specification $\zeta(z,\varepsilon,z)=z'\varepsilon$ can be generalized to allow the curvature of marginal utility to vary with unobserved type are the following.
Our main identification result, Theorem (ref), is obtained under conditions that force the dimension of unobserved heterogeneity to be the same as the dimension $d_z$ of the good quality space. The conditions that impose $d_\varepsilon=d_z$ are Assumption (ref) which requires the distribution $P_{\varepsilon\vert x}$ to be absolutely continuous with respect to Lebesgue measure on $\mathbb R^{d_z}$, and Assumption (ref)(B), which requires injectivity of $z\mapsto\nabla_\varepsilon\zeta(x,\varepsilon,z)$. In the special case $\zeta(x,\varepsilon,z)=\varepsilon'z$, the interpretation of each dimension of $\varepsilon$ as a quality dimension specific taste is appealing. However, the choice of dimension of unobserved heterogeneity remains an arbirary modelling choice.
In this section, we relax these assumptions and analyze identification with unobserved heterogeneity of lower dimension, including $d_\varepsilon=0$ and $d_\varepsilon=1$. First recall that inverse demand is identified in Lemma (ref) under Assumptions (ref), (ref), (ref)(A) and (ref), which only require $d_\varepsilon\leq d_z$. We also know from Lemma (ref) that the potential $z\mapsto V(x,z)$ is $P_{z\vert x}$ almost everywhere $\zeta$-convex. It is therefore $\zeta$-convex on any open subset of the support of $P_{z\vert x}$, which we show implies local identification. To obtain a global identification result, we need this $\zeta$-convexity everywhere.
Unfortunately, this global constraint implies a constraint on the endogenous price function, for which we do not have sufficient conditions in the general case. Under Assumption (ref), we show differentiability of the potential function $z\mapsto V(x,z)$, hence identification of marginal utility.
Local identification therefore holds under very weak assumptions, as seen in Theorem (ref)(1a). However, in cases with lower dimensional unobserved heterogeneity, consumer choices may be concentrated on a lower dimensional manifold, so that there are no open subsets in the support of $P_{z\vert x}$. In such cases, Theorem (ref)(1b) tells us that we can only identify marginal willingness to pay along the support of good attributes actually traded at equilibrium. To illustrate the idea, suppose consumers are acquiring housing. The latter is differentiated along two dimensions, size and air quality, say. Suppose consumers with identical observable characteristics x are heterogeneous along a single scalar dimension of unobserved heterogeneity $\varepsilon$. We would then expect equilibrium housing choices to be concentrated on a curve in the (size $\times$ air quality) space, implicitly defining a scalar quality index that is monotonic in unobserved type $\varepsilon$. Our result says that we can identify counterfactual marginal willingness to pay for size and air quality along that curve of observed equilibrium choices only. To obtain global identification with lower dimensional unobserved heterogeneity, we need a global $\zeta$-convexity assumption on the potential (which implies a shape restriction on the endogenous price function).
We investigate special cases:
All identification results so far, require fixing the distribution of unobserved consumer heterogeneity a priori. In this section, we derive identifying information from multiple markets, and the possibility of jointly (partially) identifying the utility function $\bar U(x,z)$ and the distribution $P_{\varepsilon\vert x}$. Suppose $m_1$ and $m_2$ index two separate markets, in the sense that producers, consumers or goods cannot move between markets. Markets differ in the distributions of producer and consumer characteristics $(P_x^{m_1},P_{\tilde y}^{m_1})$ and $(P_x^{m_2},P_{\tilde y}^{m_2})$. Suppose, however, that the distribution of unobserved tastes $P_{\varepsilon\vert x}$ and the utility function $U(x,\varepsilon,z)=\bar U(x,z)+z'\varepsilon$ is identical in both markets. Both markets are at equilibrium. The equilibrium price schedule in market $m$ is $p^{m}(z)$. The equilibrium distribution of traded qualities in market $m$ is $P^m_{z\vert x}$.
Under the assumptions of Corollary (ref), in each market, we recover a nonparametrically identified utility function $\bar U^{m}(x,z;P_{\varepsilon\vert x})$, where the dependence in the unknown distribution of tastes $P_{\varepsilon\vert x}$ is emphasized. For each fixed $P_{\varepsilon\vert x}$, Corollary (ref) tells us that $\nabla_z\bar U^{m}(x,z;P_{\varepsilon\vert x})$ is uniquely determined. In each market, the first order condition of the consumer's problem is $\varepsilon(x,z;m)=\nabla_{ap,z}p^m(z)-\nabla_z\bar U(x,z;P_{\varepsilon\vert x})$. Differencing across markets therefore yields:
which is an identifying equation for $P_{\varepsilon\vert x}$. The right-hand side of ((ref)) is identified, since the price functions are observed. Moreover, for each market $m$, the inverse demand $\varepsilon(x,z;m)$ uniquely determines $P_{\varepsilon\vert x}$, since it pushes the identified $P^m_{z\vert x}$ forward to $P_{\varepsilon\vert x}$. Point identification would require conditions under which the difference $\varepsilon(z,x;m_1)-\varepsilon(x,z;m_2)$ uniquely determines $P_{\varepsilon\vert x}$, which is beyond the scope of the present work.
The analysis of hedonic equilibrium models in Section (ref) motivates a new approach to the identification of nonseparable simultaneous equations models of the type $H(x,z)=\varepsilon$, where $x\in\mathbb R^{d_x}$ is an observed vector of covariates, $z\in\mathbb R^{d_z}$ is the vector of dependent variables, $P_{z\vert x}$ is identified from the data, $H$ is an unknown function and $\varepsilon\in\mathbb R^{d_\varepsilon}$ is a vector of unobservable shocks with distribution $P_{\varepsilon\vert x}$. In the case $d_\varepsilon=d_z=1$, $H$ is identified by Matzkin:2003 subject to the normalization of $P_{\varepsilon\vert x}$ and monotonicity of $z\mapsto H(x,z)$ for all $x$. This section develops a class of shape restrictions that allows identification of $H$ in the multivariate case $1\leq d_\varepsilon\leq d_z$.
As in the scalar case, we fix the conditional distribution $P_{\varepsilon|x}$ of errors a priori. This is justified by the fact that for any vector of dependent variables $Z\sim P_{z\vert x}$ and any pair of absolutely continuous error distributions $(P_{\varepsilon\vert x},\tilde P_{\tilde\varepsilon\vert x})$, there is an invertible mapping $T$ such that $H(x,Z)\sim P_{\varepsilon\vert x}$ and $\tilde H(x,Z):=H(x,T(Z))\sim \tilde P_{\tilde\varepsilon\vert x}$ (by McCann:95), so that $(H,P_{\varepsilon\vert x})$ and $(\tilde H,\tilde P_{\tilde\varepsilon\vert x})$ are observationally equivalent.
For identification, we rely on a shape restriction that emulates monotonicity in $z$ of $H(x,z)$ in the scalar case. This generalized monotonicity notion is inherited from utility maximizing choices of good quality $z$ by consumers with characteristics $(x,\varepsilon)$. As such, it is indexed by the utility function.
Two special cases help clarify the concept of $\zeta$-monotonicity:
The class of $\zeta$-monotone functions has structural underpinnings as demand functions resulting from the maximization of a utility function over good qualities $z\in\mathbb R^{d_z}$. Suppose a consumer with characteristics $(x,\varepsilon)$ chooses $z$ based on the maximization of $\bar U(x,z)+\zeta(x,\varepsilon,z)-p(z)$. Suppose $\zeta$ satisfies Assumption (ref) and $V(x,z):=p(z)-\bar U(x,z)$ is $\zeta$-convex. Then, the demand function $H$ that satisfies $\nabla V(x,z)=\nabla_z\zeta(x,H(x,z),z)$, $P_{z\vert x}$ a.s., is $\zeta$-monotonic by Theorem 10.28(b) page 243 of Villani:2009. Theorem (ref) below then shows identification of demand when utility is of the form $\bar U(x,z)+\zeta(x,\varepsilon,x)$ and $\zeta$ is fixed.
Theorem (ref) is a relatively straightforward application of classical results in optimal transport theory, in particular Theorem 10.28 page 243 of Villani:2009. Brenier's polar factorization theorem, in Brenier:91, was, to the best of our knowledge, first used to define multivariate quantile functions by EGH:2012 and GH:2012 with decision theoretic applications. CCG:2014 and CGHH:2017 (both coetaneous with the present paper) apply McCann:95 to multivariate quantile regression and multivariate depth, quantiles, ranks and signs respectively. This section relies on an extension of these optimal transport results to more general transport costs and interprets it as an identification result, thus extending scalar quantile identification strategies.
If we revisit the two special cases of $\zeta$-monotonicity above, we obtain the classical quantile identification of Matzkin:2003, and a result on the identification of nonseparable simultaneous equations systems within the class of gradients of convex functions.
Although the previous result is presented as a corollary of Theorem (ref), it holds under weaker conditions than would be implied by Theorem (ref) in case $\zeta(x,\varepsilon,z)=z'\varepsilon$ and is a direct application of the {\em Main Theorem} in McCann:95. The only constraint is the absolute continuity of the distribution of $\varepsilon$, so that the outcome vector $z$ and the covariate vector $x$ are unrestricted.
Beyond the special cases of Corollary (ref), we revisit the examples of Section (ref). First consider the case of exponential transform $\zeta(x,\varepsilon,z):=\exp(z'\varepsilon)$. Given marginal utility $\bar U(x,z)+\exp(z'\varepsilon)$ and price $p(z)$, Theorem (ref) tells us that the solution $\varepsilon=H(x,z)$ to the consumer's problem $\nabla p(z)-\nabla_z\bar U(x,z)=\exp(z'H(x,z))H(x,z)$ is unique. In the case consumers are maximizing utility of the form $\bar U(x,z)+\sum_{k=1}^dF_k(x,z_k\varepsilon_k)$ , the corresponding system of differential equations with a unique solution is $p_{z_k}(z)-\bar U_{z_k}(x,z)=f_k(x,z_kH_k(x,z))H_k(x,z)$, each $k$, where $p_{z_k}$ and $\bar U_{z_k}$ are the partial derivatives with respect to the $k$-th variable, $f_k$ is the derivative of $F_k$ with respect to the second argument, and $H_k$ is the $k$-th component of $H$.
This paper proposed a set of conditions under which utilities and costs in a hedonic equilibrium model are identified from the observation of a single market outcome. The proof strategy extends EHN:2004 and HMN:2010 (hereafter EHMN) to the case of goods characterized by more than one attribute. The proposed shape restriction on the utility function, called twist condition, extends the single crossing condition in EHMN. The proof of identification mirrors that of (one of the strategies in) EHMN. First, inverse demand is identified from the twist condition and cyclical monotonicity (a feature of equilibrium). Then the first order condition of the consumer's problem allows the recovery of the utility function, once a suitable form of weak differentiability of the endogenous price function is ensured. The identification proof highlights another parallel with EHMN, which is (generalized) monotonicity of inverse demand. In the scalar case, this generalized monotonicity reduces to monotonicity, whereas in the special case, where utility takes the form $U((x,\varepsilon),z)=\bar U(x,z)+z'\varepsilon$, inverse demand is the gradient of a convex function. We then show that this generalized form of monotonicity is a suitable shape restriction to identify nonseparable simultaneous equations models with a strategy that extends the quantile identification of Matzkin:2003. Most of our results involve fixing the distribution of unobserved consumer heterogeneity a priori, as in the original quantile identification method. Although we provide some discussion of the case, where data from multiple distinct markets can provide additional identifying equations to (partially) identifying $(\bar U(x,z),P_{\varepsilon\vert x})$ jointly, more research is needed to develop point identification conditions for the latter.