EconBase
← Back to paper

Multivariate ordered discrete response models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

120,318 characters · 16 sections · 57 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Multivariate ordered discrete response models with two layers of dependence

\spacing{1}

abstractWe develop a class of multivariate ordered discrete response models featuring general rectangular structures, which allow for functionally interdependent thresholds across dimensions, extending beyond traditional (lattice) models that assume threshold independence. The new models incorporate two layers of dependence: one arising from the interdependence of decision rules (capturing broad bracketing behaviors) and another from the correlation of latent utilities conditional on observables. We provide microfoundations, explore semiparametric and parametric specifications, and establish identification conditions under {logical} consistency in decision-making. An empirical application to health insurance markets demonstrates the advantages of this new framework, showing how it disentangles moral hazard (captured via threshold dependence) from adverse selection (isolated in unobservable correlations), offering insights into behavioral responses obscured by lattice models. \begin{description} • Ordered response, multiple dimensions, broad bracketing, narrow bracketing, coherency, insurance, moral hazard, adverse selection • C14, C31, C35, D9 \end{description}

\spacing{1.4}

Introduction

This paper examines ordered discrete response models in which individuals make simultaneous decisions across multiple categorical dimensions, each with a meaningful ordering. A natural way to extend univariate ordered response models to these multivariate contexts -- and the one adopted by most empirical work -- is to define latent utilities and threshold decision rules for each dimension. However, such extensions, while intuitive, often oversimplify the decision-making process by assuming complete functional independence of threshold decision structures across dimensions, as illustrated in the left panel in Figure (ref).\footnote{In an auction context, this assumption is akin to suggesting that a firm bidding in two simultaneous auctions for complementary objects would employ functionally independent equilibrium strategies in each, a notion contradicted by auction theory literature. See discussion in gentrykomarovasch2023 and gentrykomarovasch2019monotone for more detail.} From a behavioral economics perspective, these models align with agents exhibiting narrow bracketing, prioritizing simpler choice rules over joint utility optimization. The structure of these models reveals that threshold intersections across dimensions create a lattice in multidimensional space, once again illustrated in the left panel in Figure (ref), prompting us to term them lattice models.\footnote{This is the term we use ourselves for this model, it is not a commonly accepted terminology}

figure[figure omitted — 1,233 chars of source]

We introduce and explore a comprehensive class of models for selecting ordered categories across multiple dimensions. These models retain familiar features of ordered response models, such as (a) reliance on latent utilities for each dimension, as in lattice models, and (b) threshold-based decision rules. However, the models innovate by allowing thresholds across dimensions to be functionally interdependent. This flexibility enables our models to capture more complex and nuanced economic behavior compared to lattice models.

Drawing on behavioral economics, our framework accommodates broad bracketing as a general case, while encompassing narrow bracketing as a special case, since lattice models are nested within our broader class. Like lattice models, we focus on a single economic agent making simultaneous decisions across multiple dimensions. In other words, different dimensions will not represent different economic agents interacting strategically, This does not mean that our setting is completely irrelevant to to game-theoretical contexts. For instance, drawing on the auction analogy from above, our decision structure determined ex-ante can represent a single bidder’s equilibrium strategies across multiple auctions.

The right panel in Figure (ref) illustrates the models we propose. We refer to them as models with general rectangular structures. Sometimes to distinguish them from lattice models and to signify the fact that intersections of thresholds across dimensions no longer form a lattice, we may refer to them as non-lattice models.\footnote{When using the term non-lattice, one has to keep in mind that lattice models are a special case of such models.}

Models with general rectangular structures pose both theoretical and practical challenges. On the theoretical side, understanding economic behavior of agents making decisions and estimating latent utility parameters or their joint dependence requires disentangling two distinct elements: the functional dependence of decision thresholds and the interdependence of latent utilities across dimensions, conditional on observables. On the practical side, one must impose restrictions on the thresholds to ensure the internal coherence of the decision structure (a concept we formalize later) and estimate a greater number of parameters from the data.

While these challenges are nontrivial, they come with significant rewards, as models with general rectangular structures allow us to more accurately uncover the true decision structure and identify the underlying economic primitives.

Specifically, general rectangular structures come with two key advantages over existing multivariate ordered choice models. First, non-lattice models permit richer forms of interaction across dimensions by allowing two distinct layers of complementarity or substitutability. For instance, the threshold decision rule might display substitutability, while the dependence structure of unobservables in the latent utilities might reflect complementarities. Furthermore, within the threshold decision structure itself, patterns of complementarity or substitutability can vary across different regions of the latent utility space. Our application to health insurance exploits this feature of two distinct layers of complementarity/substitutability to disentangle moral hazard from advantageous/adverse selection.

Second, lattice models force the sign of any partial effect on the conditional probability of exceeding a given level in one dimension to be constant across the domain (for which we will provide a concrete example in the text). By contrast, models with general rectangular structures allow these partial effects to change sign, capturing more nuanced and flexible behavioral patterns. Lattice models also restrict the conditional probability of exceeding a given level in one dimension (given all covariates across processes) from depending on covariates that do not belong to its own latent process. Models with general rectangular structures relax this restriction, allowing indirect effects from other covariates. For instance, a price subsidy meant to encourage health insurance enrollment can indirectly influence the probability of exceeding a given level in the pension plan dimension.

After reviewing the related work and situating our contribution within the broader literature, we begin our analysis in Section (ref). There, we formally define models with general rectangular structures (which we sometimes refer to as non-lattice models), lattice models\footnote{For clarity, we emphasize that our non-lattice models encompass lattice models as a special case, even though the terminology might suggest otherwise.}, and develop our concept of coherency. In Section (ref) we provide two microfoundations for general rectangular structures. The first includes explicit synergies/crowding out effects in the joint utility specification. The second microfounds general rectangular models as the outcome from discretizing a continuous joint utility maximization problem.

Section (ref) develops a semiparametric specification of multivariate ordered response models with general rectangular structures and examines their properties in detail. In particular, Section (ref) highlights how these models capture richer economic behavior than lattice models, focusing on the more flexible patterns of partial effects discussed above. Section (ref) then turns to identification. We proceed step by step, starting from the parameters associated with exclusive covariates and ending with the identification of thresholds, which is more involved. The analysis assumes independence between unobservables and covariates, and that each process includes at least one exclusive covariate with a meaningful effect. Section (ref) discusses possible estimation approaches for semiparametric models with general rectangular structures. The main method builds on coppejans2007 subject to coherency constraints expressed as equalities involving thresholds, though we do not provide formal inference results.

Section (ref) focuses on the parametric case, where the joint distribution of unobservables is assumed to follow a multivariate normal distribution. We illustrate identification under much less stringent assumptions than in the semiparamatric case through a numerical exercise using a bivariate modelm and discuss estimation via maximum likelihood, subject to the same coherency constraints.

Section (ref) presents Monte Carlo simulations assessing the performance of the proposed non-lattice probit estimator under normal errors, compared with the standard bivariate ordered probit estimators which effectively estimates a mis-specified lattice model.

Section (ref) contains applications. The first explores the relationship between cryptocurrency familiarity and optimism. It shows how the non-lattice approach helps uncover differences in how individuals form and express opinions about bitcoin’s value obscured by the lattice model. The second application concerns insurance markets, where we show how the non-lattice model can disentangle moral hazard from selection (adverse or advantageous) by allowing functionally dependent thresholds that capture moral hazard in a coherent and data-driven way leaving selection to be fully captured by correlation of unobservables.

Section (ref) concludes and the online supplement contains more details on coherency and proofs of all formal results.

Our contributions and literature review

Our paper contributes to the literature on the economic foundations of ordered choice models by extending the analysis from univariate to multivariate decision problems. We study an agent making several ordered choices whose decisions interact through functionally dependent fixed thresholds across dimensions. Most existing work focuses on univariate models. For example, cunha2007 develop a “generalized ordered choice” model with thresholds that depend on observables and unobservables, showing how this framework captures a wide range of economic settings, including dynamic ones such as schooling decisions. Earlier contributions include cameron1998, heckman1999, carneiro2003, and lewbel2003, who study ordered models with random or sequentially determined thresholds.

In contrast, our model allows for complex interactions across multiple ordered dimensions while maintaining fixed thresholds. Here, thresholds depend on the realizations of other endogenous variables rather than regressors or unobservables, requiring a joint model of all endogenous processes. This extension offers a more flexible structure on thresholds than the one implied by fixed thresholds and univariate stochastic thresholds determined by regressors and errors.

From a more foundational point of view, two main approaches have emerged in the literature on univariate threshold-based ordered response models. One treats thresholds as reduced-form tools that aid estimation but have limited behavioral interpretation (e.g., greene2010 and boes2006). For example, greene2010 notes that thresholds may capture psychological attitudes with “bunched” cut points suggesting strong preferences, and dispersed ones reflecting indifference. In economic contexts like schooling or job satisfaction, thresholds are often viewed as cost-benefit barriers, though this link is typically conceptual rather than derived from optimization. Anderson1984 extends this idea with “stereotype” ordered regressions, where thresholds relate to category proportions rather than absolute utility levels.\footnote{In the Anderson1984 model, ordinal categories are not tied to a single latent variable with fixed thresholds.}

A second strand grounds thresholds in explicit optimization problems, offering clearer microfoundations. For instance, bhat1998 derive thresholds from a range-based utility model, while ApesteguiaBallester2023 propose type-ordered random utility models, where ordered choices arise from heterogeneous preference types without restrictive distributional assumptions. Structural approaches such as cunha2007 also belong to this line of work, though none extend to multivariate settings.

Our paper takes a first step toward microfoundations for multivariate ordered response models with general rectangular structures. Section (ref) presents two approaches. The first, illustrated with a bivariate example, interprets higher discrete responses as bundles of lower outcomes that may be complements or substitutes, depending on latent utilities and thresholds. The second characterizes marginal utilities along each dimension and shows that a rectangular structure naturally arises as the optimal discrete response satisfying discrete analogues of first-order conditions. Together, these provide a foundation for understanding multivariate ordered decision-making.

Another related line of research concerns choice bracketing also referred to as sequential vs.\ simultaneous choice simonson1992, narrow vs. broad decision frames kahneman1993, local vs. overall value functions heyman1996, and isolated vs. distributed choice herrnstein1991. This literature is largely theoretical and experimental tversky1981,read1999,thaler1999,rabin2009,lian2020,camara2021,zhang2021, with only a few descriptive or structural empirical studies \citep*{camerer1997,thakral2021}.\footnote{tversky1981 provides a classic example of narrow bracketing in experimental settings.} To date, econometric work has not explicitly modeled narrow versus broad bracketing behavior. Our framework of general rectangular structures offers a natural way to do so with lattice models corresponding to narrow bracketing and the broader rectangular structure capturing \textit{broad bracketing} decisions. This setup also enables formal testing for broad bracketing by examining whether thresholds in the latent space conform to a lattice structure.

A tangentially related literature is the discrete choice framework with strategic interactions, where outcomes for one player depend on the actions of others tamer2003,berry2007,ciliberto2009,honor2010,chesher2017,chesher2020,aradillas2022. In these models, each agent represents a distinct dimension, and best responses can lead to incoherent or incomplete outcomes. In contrast, our paper focuses on a single economic agent making decisions along multiple dimensions. For such an agent, the decision problem is internally consistent by construction, and therefore models with general rectangular structures are coherent. As we detail in the following section, by coherency we mean logical consistency in decision-making that ensures that rectangular regions representing different discrete responses do not overlap and together cover the entire latent space $\mathbb{R}^D$.\footnote{For more on coherency, see heckman1978 and tamer2003. tamer2003 distinguish between incoherency and incompleteness in games with strategic interactions, a distinction followed by later studies. In our setting, we use the term “coherency” to refer more generally to overall logical consistency.}

Model with a general rectangular structure

We formally define general rectangular structures and lattice multivariate ordered response models for an agent making decisions across $D \geq 2$ dimensions. These models map a $D$-variate latent continuous metric $(Y^{*c_1}, \ldots, Y^{*c_D})$ to a discrete metric $(Y^{c_1}, \ldots, Y^{c_D})$, with ordered responses in dimension $d$ denoted as $y^{(d)}_j$, $j=1, \ldots, M_d$, and satisfying $y^{(d)}_1 < \cdots < y^{(d)}_{M_d}$.

definition[General rectangular structure model] A model has a general rectangular structure (or sometimes we refer to it as a non-lattice model) if \[ (Y^{c_1}, \ldots, Y^{c_D}) = (y^{(1)}_{j_1}, \ldots, y^{(D)}_{j_D}) \quad \iff \quad (Y^{*c_1}, \ldots, Y^{*c_D}) \in R_{j_1, \ldots, j_D}, \text{ where} \] \begin{equation} R_{j_1, \ldots, j_D} = \bigtimes_{d=1}^D \left({\alpha}^{(d)}_{j_1, \ldots, j_{d-1}, {\color{red} j_{d} \: - \: 1}, j_{d+1}, \ldots, j_D}, {\alpha}^{(d)}_{j_1, \ldots, j_{d-1}, {\color{red} j_d}, j_{d+1}, \ldots, j_D}\right], \end{equation} with thresholds ${\alpha}^{(d)}_{j_1, \ldots, j_{d-1}, {\color{red} j_d}, j_{d+1}, \ldots, j_D}$ increasing in $j_d$ for given other indices and normalized at the boundary as \[ \alpha^{(d)}_{j_1, \ldots, j_d, \ldots, j_D} = +\infty \text{ for } j_d = M_d, \quad \alpha^{(d)}_{j_1, \ldots, j_d, \ldots, j_D} = -\infty \text{ for } j_d = 0. \]

Threshold intersections in Definition (ref) do not necessarily form a lattice in $\mathbb{R}^D$, reflecting functionally interdependent decision rules, akin to broad bracketing in behavioral economics.

definition[Lattice model] A lattice model is a special case of a general rectangular structure (non-lattice) model in which each threshold $\alpha^{(d)}_{j_1, \ldots, j_d, \ldots, j_D} $ depends only on the index in its own dimension: \[ (Y^{c_1}, \ldots, Y^{c_D}) = (y^{(1)}_{j_1}, \ldots, y^{(D)}_{j_D}) \quad \iff \quad Y^{*c_d} \in \left( \alpha^{(d)}_{j_d-1}, \alpha^{(d)}_{j_d} \right] \quad \forall d, \text{ with } \] \[ \alpha^{(d)}_{j_d} = +\infty \text{ for } j_d = M_d, \quad \alpha^{(d)}_{j_d} = -\infty \text{ for } j_d = 0. \]

Here, thresholds are functionally independent across dimensions, forming a lattice in $\mathbb{R}^D$ in their intersections. Lattice models correspond to a decision maker who narrowly brackets, since decisions can now be seen as made dimension-by-dimension, as opposed to jointly. Lattice models will misspecify a decision maker who broadly brackets.

Thus, the distinction between broad and narrow bracketing is fully captured by functional interdependence or independence of decision rules, determined by the thresholds. Both lattice and non-lattice models permit correlated decisions through latent processes, but the latter distinguish correlation in unobservables from interdependent decision rules.

\paragraph*{Coherency}

The flexibility of general rectangular structure models is achieved by allowing thresholds for each dimension to depend on the full vector of response indices. But this comes at a cost, as Definition (ref) does not guarantee that the division of latent space into regions $R_{j_1, \ldots, j_D}$ is to be exhaustive or mutually exclusive. As a result, the condition that each latent profile maps to exactly one observed response (which is to us is associated with logical consistency in decision-making) and to which we refer to as coherency is not satisfied by Definition (ref) construction. Since the model aims to describe the behavior of a logically consistent decision-maker, it should always satisfy coherency, especially when we take it to the data. In general rectangular structure models ensuring this requires explicit constraints on the thresholds. Lattice models, by contrast, are coherent by construction.

We now examine coherency in a bivariate general rectangular structure ordered response model, providing a formal condition on thresholds for the model to be coherent. The characterization of coherency for $D>2$ is more involved and is given in the online supplement.

prop[Coherency for $D=2$] Consider a bivariate general rectangular structure ordered response model, defined by a set of thresholds $\left\{\alpha^{(1)}_{j_1, j_2}, \alpha^{(2)}_{j_1, j_2}\right\}_{j_1=1, j_2=1}^{M_1-1, M_2-1}$. Given thresholds normalizations at the boundary, the model is coherent -- i.e., the latent space is partitioned into mutually exclusive and exhaustive rectangular regions $R_{j_1, j_2} = (\alpha^{(1)}_{j_1-1, j_2}, \alpha^{(1)}_{j_1, j_2}] \times (\alpha^{(2)}_{j_1, j_2-1}, \alpha^{(2)}_{j_1, j_2}]$ each corresponding to a unique observed outcome -- if and only if, for all $(j_1, j_2)$, \begin{equation} \left( \alpha^{(1)}_{j_1+1, j_2} - \alpha^{(1)}_{j_1, j_2} \right) \cdot \left( \alpha^{(2)}_{j_1, j_2+1} - \alpha^{(2)}_{j_1, j_2} \right) = 0. \end{equation}
figure[figure omitted — 1,107 chars of source]

In other words, for the model to be coherent, the thresholds must satisfy a local condition for each $2 \times 2$ block of adjacent cells. Specifically, within each block, at least one of the dimensions must have constant thresholds across that block. This requirement prevents ambiguity in decision-making by ensuring that when an agent faces a choice within a $2 \times 2$ block, they make their decision sequentially: first along one dimension, where the thresholds remain fixed, and then along the other dimension. An illustration of a local problem and the coherency requirement is given in Figure (ref). Thus, in every local decision problem one dimension is leading and the leading dimension may be different across different parts of the domain (e.g., when one considers health insurance levels vs retirement contribution level, it may very well be the case that for lower levels the insurance decision is the leading one whereas for higher levels of both the leading decision is the retirement contributions as at those levels long-run financial planning may be of more relevance).

Microfoundations

There are several ways to approach a general rectangular structure model from a microeconomic foundations perspective. We propose two such approaches.

The first approach directly models complementarities and substitutabilities in the joint utility across different pairings of options. We illustrate how it can be done in a simple bivariate model with two discrete options (1 and 2) in each dimension. The left panel in Figure (ref) shows substitutability in the decision structure reflected in the larger threshold in the second dimension when $Y^{c_1}=1$. It shows that choosing a higher level in dimension 1 makes it harder to choose a higher level in the other, often due to resource constraints. The right panel in Figure (ref) shows complementarity as choosing a higher level in dimension 1 facilitates a higher level in the other dimension.

figure[figure omitted — 1,864 chars of source]

Consider the following utilities across four pairs of discrete choices: for constant $v<0$,

align[align omitted — 169 chars of source]

where $D=\mathbf{1}(Y^{*c_1} > 0 )\mathbf{1}(Y^{*c_2} > v)$. Then the $\text{argmax}_{j_1,j_2} U(j_1,j_2)$ is (i) (1,1) when $Y^{*c_1}, Y^{*c_2}\leq 0$; (ii) $(1,2)$ when $Y^{*c_1} \leq 0, Y^{*c_2} > 0$; (iii) (2,1) when $Y^{*c_1} > 0, Y^{*c_2} \leq v$; and (iv) (2,2) when $Y^{*c_1} > 0$, $Y^{*c_2} > v$. The lower threshold $v<0$ facilitates choosing $ Y^{c_2} = 2 $ when $ Y^{c_1} = 2 $, suggesting complementarity (as in the right panel in Figure (ref)). The utility boost $-v > 0 $ in $ U(2,2) $ when $ D = 1 $ acts like a synergy term, incentivizing (2,2) over (2,1) or (1,2) when propensities are sufficient. This reflects scenarios where choosing one high option reduces the marginal cost of the other (e.g., economies of scale, subsidies).

To obtain substitutes in the decision structure (as in the left panel in Figure (ref)), consider $w>0$ and replace $U(2,2)$ in ((ref)) with the following definition: $$U(2,2) =D(Y^{*c_1} + Y^{*c_2})+ (1-D)(Y^{*c_1} + Y^{*c_2}-w),$$ with $D$ defined in the same way as before. The higher threshold for $ Y^{*c_2}$ when $ Y^{c_1} = 2 $ indicates that choosing $ Y^{c_1} = 2 $ raises the level needed for $ Y^{c_2} = 2 $, reflecting substitutability. This captures resource competition (e.g., budget, time) where pursuing one high choice increases the cost of the other. The penalty $-w < 0$ when $ D = 0 $ reinforces the trade-off. An analogous construct could be employed for any number of ordered choices in each dimension.

The second approach to providing microeconomic foundations for general rectangular structure models is to view discrete options as the result of discretizing an underlying continuous space, whether due to survey design, categorical reasoning, or similar factors. If one had a smooth function $U(y^{(1)},y^{(2)})$ of continuous responses $(y^{(1)},y^{(2)})$ then the global maximum $(\overline{y}^{(1)},\overline{y}^{(2)})$ would have necessarily satisfied $\frac{\partial U(\overline{y}^{(1)},\overline{y}^{(2)})}{\partial y^{(1)}}=0, $ $\frac{\partial U(\overline{y}^{(1)},\overline{y}^{(2)})}{\partial y^{(2)}}=0$. With the discrete grid $(y^{(1)}_{j_1},y^{(2)}_{j_2})$ to find a maximum on the grid we have to ensure the function value does not increase when moving to any neighboring grid points in each coordinate: that is, $(y^{(1)}_{j_1},y^{(2)}_{j_2})$ is the maximizer of $U$ on the grid only if

align[align omitted — 353 chars of source]

A general rectangular structure model is obtained when for any $j_1>1$ and any $j_2>1$,

align*[align* omitted — 226 chars of source]

These specifications of the utility differences (analogues of marginal utilities) ensure that for any realization $(Y^{*c_1},Y^{*c_2})$ of latent processes, there is only one pair $(j_1,j_2)$ that satisfies ((ref))-((ref)). Namely, this is the pair $(j_1,j_2)$ such that $(Y^{*c_1},Y^{*c_2}) \in R_{j_!,j_2}$.

Semiparametric specification, partial effects, identification, estimation

To take our model to the data, we adopt the standard approach in the discrete response literature, specifying each $d$-th continuous latent process as a linear index:

equation[equation omitted — 101 chars of source]

where $x_d$ is a row vector of observable covariates, $\beta_d$ is a column vector of unknown parameters, and $\varepsilon_d$ is an unobservable error term. This structure interprets $x_d \beta_d$ as the systematic component of an agent's latent propensity to choose an ordered category in dimension $d$, with $\varepsilon_d$ capturing random shocks to the latent utility. The errors $\varepsilon_1, \ldots, \varepsilon_D$ may be dependent, allowing latent processes $Y^{*c_d}$ to be correlated conditional on covariates. This linear index, standard in discrete choice models, supports estimation of threshold-based interdependence in general rectangular structure models while maintaining parsimony. While more general functions of $x_d$ could be used without compromising identifiability, the linear form offers practical simplicity with minimal loss of flexibility.

Given the complexity of the model, particularly with regard to the two-layer dependence structure, we should be prepared for fairly stringent requirements on the data to ensure identification of the following objects of interest : $\beta_d$, $d=1,...,D$, the joint c.d.f of $\boldsymbol{\varepsilon}=(\varepsilon_1, \ldots, \varepsilon_D)'$, and the thresholds. We start by imposing Assumption (ref).

assumption$\boldsymbol{\varepsilon}=(\varepsilon_1, \ldots, \varepsilon_D)'$ is independent of $x=(x_1,\ldots, x_D)$ and has a convex support in $\mathbf{R}^D$.
notationFor any $d=1, \ldots, D$, let $\kappa_d $ denote either $\leq$ or $>$ sign. Denote $$F_{\kappa_1, \kappa_2, \ldots, \kappa_D}(t_1,...,t_D)=P(\cap_{d=1}^D(\varepsilon_d \, \kappa_d \, t_d )) $$ for any $\kappa_d \in \{\leq, >\}$, $d=1,..,D$. Functions $F_{\leq, ..., \leq}$ and $F_{>, ..., >}$ are the joint c.d.f. and the joint survival function of $\boldsymbol{\varepsilon}$, respectively, and for simplicity we will interchangeably denote them as $F$ and $\overline{F}$, respectively. All the other cases correspond to hybrid forms of c.d.f.s and survival functions. Similar notation apply to subsets of dimensions. Also denote $F_{d,\kappa_d}(t_d) = P(\varepsilon_d \, \kappa_d \, t_d)$ for $\kappa_d \in \{\leq, >\}$. Thus, $F_{d,\leq}$ is the marginal c.d.f. and $F_{d,>}$ is the marginal survival function of $\varepsilon_d$.

In our related paper, km2025lattice, we outline the increasing degree of restrictions required to ensure identification in semiparametric lattice models. There, we discuss why identification of parameters and thresholds can rely on a weaker version of Assumption (ref) -- namely, the independence of $\varepsilon_d$ from $x_d$ for each $d=1, \ldots, D$. However, as we argue there, the identification of the joint distribution of $\boldsymbol{\varepsilon}$ does rely on Assumption (ref) even in lattice models. Thus, when one of the objectives is to identify the joint distribution of errors (motivation for this is discussed later) then our Assumption (ref) is no more restrictive than what is required in lattice models.

Advantages of non-lattice models over lattice models

Next, we describe how general rectangular structure models offer significant advantages over lattice models in capturing complex interdependencies between discrete choices, Even in the simplest possible case of a bivariate model with 2 discrete choices in each dimension we can talk about at least four different empirical structures. These arise from the interactions between complementarity or substitutability in decision thresholds (the first layer) and complementarity or substitutability in unobservables (the second layer), the latter captured through positive or negative dependence that reflects whether shocks reinforce or offset one another.

In many applications, substitutability/complementarity relationships at each layer may be unknown a priori and need to be identified from the available data.

\paragraph*{Cross-partial effects: Partial effects in one dimension can depend on covariates exclusive to other processes} In another aspect, under Assumption (ref), models with general rectangular structures allow the likelihood $P(Y^{c_d} \geq y^{(d)}_j | x)$ that an individual selects at least a certain level of commitment in one decision area, given all relevant personal and contextual factors, depend on covariates exclusive to other processes (e.g., $x_h, h \neq d$), enabling analysis of partial effects $\frac{\partial P(Y^{c_d} \geq y^{(d)}_j | x)}{\partial x_{h,m}}$. This effect is indirect -- covariates exclusive to processes in other dimensions affect the probabilities of decision in a given dimension indirectly through their influence on latent utilities in other dimensions. In lattice models under Assumption (ref) such partial effects are absent (that is, zero).

Consider, for example, high school students choosing academic effort ($Y^{c_1}$: low = 1, medium = 2, high = 3) obtained from hours of study and extracurricular activity participation ($Y^{c_2}$: low = 1, moderate = 2, high = 3) e.g. expressed through involvement in sports or clubs. These choices are interdependent at the layer of the decision structure with the direction of that interdependence unknown a priori. Indeed, high academic effort may limit time for extracurriculars, but at the same time extensive extracurricular involvement may encourage academic effort for college applications. Covariates exclusive to $Y^{*c_1}$ could be parental education and access to tutors, covariates exclusive to $Y^{*c_2}$ could be school resources and peer involvement. Covariates shared by both $Y^{*c_1}$ and $Y^{*c_2}$ could be socioeconomic status and school quality. A non-lattice model would e.g. allow school resources (exclusive to $Y^{*c_2}$) affect the probability of high academic effort. It may show for example that better school resources (such as sports facilities) increase the probability of high academic effort by motivating students to balance high academic effort for college admissions. This would be informative for resource allocation policies. A lattice model would miss that effect.

To illustrate this theoretically, take the model in Figure (ref) and note that

multline*[multline* omitted — 314 chars of source]

With the lattice structure ($\alpha^{(1)}_{1,1}= \alpha^{(1)}_{1,2}$ ), only the middle term in this representation is left. Therefore, with the lattice stricture, $\frac{\partial P(Y^{c_1} \leq 1 | x)}{\partial x_{2,m}}=0$, where $x_{2,m}$ is a covariate exclusive to the process $Y^{*c_2}$. Thus, cross-partial effects are not possible in the lattice model.

With the non-lattice structure, {$$\frac{\partial P(Y^{c_1} \leq 1 | x)}{\partial x_{2,m}}=-\beta_{2,m} \frac{\partial F\left( \alpha^{(1)}_{1,1} -x_1\beta_1, \alpha^{(2)} -x_2\beta_2\right)}{\partial e_2}+\beta_{2,m} \frac{\partial F\left( \alpha^{(1)}_{1,2} -x_1\beta_1, \alpha^{(2)} -x_2\beta_2\right)}{\partial e_2}$$}is not necessarily 0 because of $\alpha^{(1)}_{1,1}\neq \alpha^{(1)}_{1,2}$. In contrast, under Assumption (ref), lattice models restrict $P(Y^{c_d} \geq y^{(d)}_j | x) = P(Y^{c_d} \geq y^{(d)}_j | x_d)$, ignoring cross-process effects.

\paragraph*{Partial and cross-partial effects can vary in sign across the domain.}

figure[figure omitted — 1,562 chars of source]

Let us first illustrate this property for cross-partial effects. Using the model in Figure (ref) and the formula for $\frac{\partial P(Y^{c_1} \leq 1 | x)}{\partial x_{2,m}}$ above, we can see that this partial effect is guaranteed to be positive with a positive probability (and non-negative a.e.) iff $\beta_{2,m} ( \alpha^{(1)}_{1,2} - \alpha^{(1)}_{1,1})>0$. In a completely analogous way, we can consider $\frac{\partial P(Y^{c_1} \leq 2 | x)}{\partial x_{2,m}}$ and obtain this partial effect is guaranteed to be positive with a positive probability (and non-negative a.e.) iff $\beta_{2,m} ( \alpha^{(1)}_{2,2} - \alpha^{(1)}_{2,1})>0$. Since $\alpha^{(1)}_{2,2} - \alpha^{(1)}_{2,1}<0$ and $\alpha^{(1)}_{1,2} - \alpha^{(1)}_{1,1}>0$, partial effects $\frac{\partial P(Y^{c_1} \leq 1 | x)}{\partial x_{2,m}}$ and $\frac{\partial P(Y^{c_1} \leq 2 | x)}{\partial x_{2,m}}$ will have different signs. In contrast, in lattice models the sign of such cross-partial effect would remain consistent across the whole domain,

Let us now look at the partial effects with respect to own covariates. If $x_{d.m}$ is exclusive to $Y^{*c_d}$, then the own partial effect $\frac{\partial P(Y^{c_d} \leq y^{(d)}_j | x)}{\partial x_{d,m}}$ will have the same sign in a non-lattice model for any $y^{(d)}_j$, $j=1,\ldots, M_d$. The situation is different if $x_{d,m}$ is shared with another latent process. Using our model in Figure (ref), suppose $x_{1,m}=x_{2,m_2}$ and obtain that in this case, {

multline*[multline* omitted — 562 chars of source]

} As an example, consider a special case of independent $\varepsilon_1$ and $\varepsilon_2$. Then {

multline*[multline* omitted — 360 chars of source]

} If $\beta_{2,m_2}(\alpha^{(1)}_{j,2}-\alpha^{(1)}_{j,1})$ and $\beta_{1,m}$ have the same sign, then we have the difference of either two positive or two negative terms. Each of them can potentially dominate the other depending on the values of indices (and, hence, $x_1$ and $x_2$), thresholds, as well as $\beta_{1,m}$ and $\beta_{2,m_2}$.

The feature of own partial effects with respect to a shared covariate or cross-partial facts not having the same sign across the domain complicates the identification analysis.

Identification in the semiparametric model

Next, we establish identification in semiparametric models with non-lattice structures. The knowledge of index parameters and both layers of dependence -- through threshold parameters and the joint c.d.f. -- is central to counterfactual analysis and policy design involving joint outcomes such as household decisions on, say, healthcare and education investments. The double-layer dependence structure will e.g. determine whether bundled interventions reinforce or crowd out each other.

The approach used in lattice models in km2025lattice for the identification of index parameters and threshold differences relies on the ability to isolate different dimensions and consider one dimension at a time. This is not going to work here as we cannot isolate different dimensions. E.g., from Figure (ref) one can see that $P\left(Y^{(1)}\leq y_j^{(1)}|x\right)$ for $j=1,2$, cannot be expressed just in terms of the index $x_1\beta_1$ and the marginal c.d.f. $F_{1,\leq}$. General rectangular structure cases therefore require a different approach to identification. Intuitively, the identification of parameters $\beta_d$ and the threshold structure in these models should be more demanding on the data compared to lattice models, especially given an unknown dependence structure of unobservables. This is indeed the case until we get to the stage of identifying the joint c.d.f. of unobservables. At that stage, as follows from Theorem (ref) here and Theorem 4 in km2025lattice both lattice and non-lattice models become similarly demanding on the data, which is an interesting result.

After introducing a helpful definition and notations, we proceed to prove identification in several steps as follows:

itemize• step: identification of parameters corresponding to exclusive covariates in each process (Theorem (ref)). • step: identification of parameters corresponding to non-exclusive covariates (Theorem (ref)). • step: identification of the joint c.d.f. (Theorem (ref)). • step: identification of the thresholds (Theorem (ref)).
definitionCovariate $x_{d,i}$ is exclusive to process $d$ if $ x_{d,i} \, | x_{-d}$ has a non-degenerate distribution almost everywhere for $x_{-d} \equiv (x_1, \ldots, x_{d-1}, x_{d+1}, \ldots, x_D)$.
notationFor each $d=1, \ldots, D$, let $x_{d, 1:L_d}$ denote the subvector of $x_d$ that consists of all the covariates in $x_d$ that are exclusive to the process $Y^{*c_d}$.

Intuitively, an exclusive covariate is one that contains information unique to the $d$-th process $Y^{*c_d}$ and cannot be perfectly predicted from the covariates associated with other processes. It is without a loss of generality that these exclusive covariates are arranged to be the first few covariates within $x_d$.

notationDue to the presence of shared covariates among processes, the vector $x=(x_1,...,x_D)$ may contain several identical variables. Its effective dimension will count each of these shared covariates only once. Let us denote the effective dimension of $x$ as $K$. In other words, $K$ is the number of exclusive covariates across all processes plus the number of covariates shared by at least two processes (and counted only once in this effective dimension).

Theorem (ref) gives sufficient conditions for the identification of $\beta_{d; 1;L_d}$, $d=1, \ldots, D$, which are the parameters corresponding to the exclusive covariates in each process.

theoremConsider a $D$-variate ordered discrete response model with the index structure ((ref)). Suppose Assumption (ref) holds for each $d=1, \ldots, D$, and the model has a coherent general rectangular structure. Suppose that the following conditions are satisfied: \begin{enumerate} • $L_d \geq 1$ for each $d=1, \ldots, D$. • The coefficient $\beta_{d,1}$ corresponding to $x_{d,1}$ in $x_{d} \beta_{d}$ is 1, $d=1, \ldots, D$. • For each $d=1, \ldots, D$, there exists $j_d=1,\ldots, M_d-1$, such that $P(S_d(j_d))>0$, $$\text{ where } \quad S_d(j_d)= \left\{x: 0<P(Y^{c_d} \leq y^{(d)}_{j_d}|x) < P(Y^{c_d} \leq y^{(d)}_{j_d+1}|x) \right\}.$$ In addition, $S_d(j_d)$ contains a Cartesian product $(\underline{x}_{d,1}, \overline{x}_{d,1}) \times {S}_{d,-1}(j_d)$, where $\overline{x}_{d,1}>\underline{x}_{d,1}$ and ${S}_{d,-1}(j_d) \subset \mathbf{R}^{K-1}$, such that $P((\underline{x}_{d,1}, \overline{x}_{d,1}) \times {S}_{d,-1}(j_d))>0$ (the order of covariates in this Cartesian product coincides with the order of covariates in $x$) and ${S}_{d,-1}(j_d)$ in not contained in any proper linear subspace of $ \mathbf{R}^{K-1}$. \end{enumerate} Then parameters $\beta_{d, 1:L_d}$, $d=1, \ldots, D$, corresponding to the exclusive covariates in each process are identified.

Condition (a) states that each process has at least one exclusive covariate, and condition (b) effectively states that the first (and potentially only) exclusive covariate in process $Y^{*c_d}$ has a non-zero coefficient; it further normalizes it to 1 (alternatively, could be normalized to $-1$ if the impact is negative). Normalization restrictions like these are standard in semiparametric models where parameter vectors generally can only be identified up to scale. These normalizations can be different across $d$ (some normalzied to to 1, some to $-1$). Condition (c) is a version of the rank condition and, intuitively, requires that for $d=1, \ldots, D$, there is some some continuous variation in at least one exclusive covariate in $x_d$, conditional on other covariates, at least in that part of domain that gives non-trivial (and, thus, informative) probabilities of choice with respect to dimension $d$. One of the requirements is that $Y^{c_d}$ takes at least two different values with positive probabilities. In condition (c) for simplicity we took them to be two different consecutive values $y^{(d)}_{j_d}$ and $y^{(d)}_{j_d+1}$ (more generally, they don't need to be consecutive).

Our next result is on the identification of those parameters components that correspond to shared regressors. It is given in Theorem (ref) and relies on strengthening conditions on exclusive covariates to have a large enough support. The result leverages the multidimensional nature of the problem and the ability to consider probabilities $P(\cap_{d=1}^D(Y^{c_d} \, \kappa_d \, y^{(d)}_{j_d})|x )$, where $\kappa_d \in \{\leq,>\}$. In bivariate models the regions inside these probabilities are easy to visualize in the latent space as constructed using rectangles starting from one “corner” of partitioning structure. Note that large (or large enough) support assumptions are common in the semiparametric literature and, in particular, in semiparametric univariate ordered response models (see e.g. manski1985,manski1988,horowitzbook, lewbel2000, lewbel2003).

theoremSuppose all the conditions of Theorem (ref) hold. Also assume that: \begin{itemize} • there is a collection of indices $(j_1,...,j_d)$ such that the intersection $S=\cap_{d=1}^D S_d(j_d)$ has positive probability measure and full affine dimension $K$; • for each $d$ either $\underline{x}_{d,1}$ is small enough to guarantee that $\alpha^{(d)}_{j_1,...,j_D}-\underline{x}_{d,1} - x_{d,-1} \beta_{d,-1}$ is at the upper support point of the $\varepsilon_d$ distribution, or $\overline{x}_{d,1}$ is large enough to guarantee that $\alpha^{(d)}_{j_1,...,j_D}-\overline{x}_{d,1}-x_{d,-1} \beta_{d,-1}$ is at the lower support point of the $\varepsilon_d$ distribution for $x_{d,-1} \in {S}_{d,-1}(j_d)$. \end{itemize} Then $\beta_d$, $d=1, \ldots, D$, are identified.

Condition (b) Theorem (ref) can be reformulated in terms of observed choice probabilities (and, thus, verified in practice) where one would need to check that some of them can attain one of its natural bounds (either 0 or 1) through the variation in an exclusive covariates with other covariates taking values in some subset of a positive measure. Note that Theorem (ref) does not rely on the result of Theorem (ref) as its proof does not use the fact that all parameters for exclusive covariates have been identified and establishes the identification of the whole vector $\beta_d$, including $\beta_{d,1:L_1}$ independently of Theorem (ref). It is instructive though to have Theorem (ref) as a separate result to emphasize that the identification of exclusive covariates' parameters requires weaker conditions.

Our next result is on the identification of the joint distribution of $\boldsymbol{\varepsilon}$. It is enough to identify one function $F_{\kappa_1 \ldots \kappa_D}$ to fully characterize this distribution. We can identify the distribution from one of the “corners” in our partitioning that gives us enough variation in probabilities.

theoremSuppose all the conditions of Theorem (ref) hold for a collection of indices $(j_1,...,j_D)$ such that $j_d=1$ or $j_d+1=M_d$ for each $d$. Additionally, in condition (b) of Theorem (ref) suppose that for each $d=1,...,D$ both of the following conditions hold: $\underline{x}_{d,1}$ is small enough to guarantee that $\alpha^{(d)}_{j_1,...,j_D}-\underline{x}_{d,1} - x_{d,-1} \beta_{d,-1}$ is at the upper support point of the $\varepsilon_d$ distribution, and $\overline{x}_{d,1}$ is large enough to guarantee that $\alpha^{(d)}_{j_1,...,j_D}-\overline{x}_{d,1}-x_{d,-1} \beta_{d,-1}$ is at the lower support point of the $\varepsilon_d$ distribution for $x_{d,-1} \in {S}_{d,-1}(j_d)$ (this is in contrast with either/or required in Theorem (ref)). Then, under the normalization $F_{d,\leq}(e_{0d})=c_{0d}$ for each marginal c.d.f. $F_{d,\leq}$ for some known $e_{0d}$ and $c_{0d} \in (0,1)$, $d=1, \ldots, D$, the distribution of $\varepsilon$ is identified.

Conditions on covariates in Theorem (ref) first identify marginal distributions up to a shift and then, coupled with the normalization restrictions, fully identify them. In addition, threshold parameters of the “corner” of the partitioning structure in the latent space implicitly specified in the formulation of Theorem (ref) (through the values of indices $j_d$, $d=1,\ldots,D$,) are identified. Then the observed probabilities of that “corner” region together with the knowledge of thresholds defining it identify the joint distribution of $\boldsymbol{\varepsilon}$.

Note that conditions in Theorems (ref)-(ref) are increasingly more restrictive. This is not surprising as, first, when thinking of Theorem (ref) vs Theorem (ref) conditions, it is intuitive that the identification of the parameters corresponding to shared covariates is harder due to multiple effects occurring together when this covariate varies. When comparing Theorem (ref) with Theorem (ref) conditions, we see that in the former conditions are more restrictive in requiring that there is sufficient variation in covariates at the boundary of the partitioning of the latent space (namely, in one of the “corners”).

Our final result is on the identification of threshold parameters. This result allows us to find out whether decision-making is consistent with broad bracketing or narrow bracketing. Identification comes from variation in covariates and consideration of probabilities of various rectangular regions, which can be expressed in terms of $F_{\kappa_1,\ldots,\kappa_D}$. Theorem (ref) gives a formal identification result for the thresholds. It strengthens previous conditions by essentially requiring that for any rectangle $R_{j_1,...,j_D}$ there is a positive mass of $x$ that delivers a strictly positive choice probability for this rectangle (in contrast, Theorem (ref) only required that to apply in one of the “corner” regions).

commentsuch as \begin{small} \begin{eqnarray*} &P&\left((Y^{c_1}, \ldots, Y^{c_D}) = \left(y^{(1)}_{j_1}, \ldots, y^{(D)}_{j_D} \right) \left. \right| x\right) = P\left((Y^{*c_1}, \ldots, Y^{*c_D}) \in R_{j_1, \ldots, j_D} | x \right) \\ &=& \sum_{\substack{ m_1 \in \{0,1\} \dots \\ m_D \in \{0,1\}}} (-1)^{m_1+\ldots + m_D} P\left( \bigcap_{{d}=1}^D \left(\varepsilon_d < {\alpha}^{(d)}_{j_1-m_1, \ldots, j_{d-1}-m_{d-1}, j_d-m_d, j_{d+1}-m_{d+1}, \ldots, j_D-m_D} -x_d \beta_d \right) \right) \end{eqnarray*} \end{small}
theoremSuppose all the conditions of Theorem (ref) hold for any collection of indices $(j_1,\ldots,j_D)$ with $j_d\in \{1,M_d-1\}$, $d=1,\ldots, D$. Then all the thresholds ${\alpha}^{(d)}_{q_1, \ldots, q_{d-1}, q_d, q_{d+1}, \ldots, q_D}$ are identified.
figure[figure omitted — 3,315 chars of source]

We prove the identification of thresholds in Theorem (ref) sequentially in a manner somewhat consistent with solving a puzzle and it is best illustrated in the bivariate case in Figure (ref). In Stage 1, thresholds are identified that define “corner” regions $R_{q_1, \ldots, q_D}$ with $q_d \in \{1,M_{d-1}\}$. The result of this step is given in Panel A. In Stage 2, border regions are considered and for any $d=1,.., D$, the thresholds $a^{(d)_{q_1,...,q_{d-1}, q_d,q_{d+1}, ...q_D}}$ are identified when $q_{h}$, $h \neq d$, remains fixed at its value 1 or $M_h-1$ whereas $q_{d}$ varies from 2 to $M_h-2$. In the bivariate case the thresholds identified after Stage 2 are in Panel B in Figure (ref) as dotted lines (dotted because their length is not known). Stage 3 continues to consider border regions for each $d=1,.., D$ and identifies thresholds $a^{(d)}_{q_1,...,q_{d-1}, q_d,q_{d+1}, ...q_D}$ for when $q_d$ remains fixed at its value 1 or $M_d-1$ whereas $q_{h}$, $h \neq 2,$ vary from 2 to $M_h-2$. In the bivariate case the thresholds identified after Stage 3 are in Panel C in Figure (ref) (the thresholds from Stage 2 are now in solid lines at their lengths are known). Stage 4 identifies all the thresholds “in the middle” proceeding sequentially from the rectangular regions close to the border further into the depth of partitioning. Each stage builds on the results of the previous stage. Importantly, the sequential nature of the threshold identification process ensures that at each phase there are at most $D$ unknown thresholds (out of the overall $2D$ thresholds forming a rectangle of interest) that need to be identified. Identification of yet unknown thresholds is obtained through a variation in $D$ indices $x_d\beta_d$, $d=1,\ldots,D$, and the knowledge of $F_{\kappa_1,...,\kappa_D}$ for a suitable $\kappa_1,...,\kappa_D$ (Theorem (ref) implies the knowledge of $F_{\kappa_1,...,\kappa_D}$ for any $\kappa_1,...,\kappa_D$). Which $F_{\kappa_1,...,\kappa_D}$ is suitable for the identification task of a particular rectangle depends on which thresholds forming this rectangle are already known and which still need to be identified (up to $D$ of these).

Note that the only assumption we make about the support of $\boldsymbol{\varepsilon}$ is that it is convex. This support can be bounded, partially bounded (that is, bounded in some directions but not others), or extend over the entire $\mathbb{R}^D$. The geometry of this support is closely related to what we require from the support of the exclusive covariates. For example, if the support of $\boldsymbol{\varepsilon}$ is bounded, then it is sufficient for the exclusive covariates to vary only within a finite range as well. Regardless of its specific form, our identification argument remains flexible as it can rely on areas near the finite boundary of the support, or on directions in which the support is unbounded, as long as those directions form a sufficiently large cone.

commentSuppose we have already identified $\alpha^{(1)}_{q_1-1.q_2}$ and $\alpha^{(2)}_{q_1.q_2-1}$ from earlier phases. Then we only need to identify $\alpha^{(1)}_{q_1.q_2}$, $\alpha^{(2)}_{q_1.q_2}$ and the direction in which to proceed os determined by $\kappa_1$ and $\kappa_2$ being both $\leq$. Fix indices $x_1\beta_1$ and $x_2\beta_2$ and denote $z_0\equiv (\alpha^{(1)}_{q_1-1.q_2}-x_1\beta_1, \alpha^{(2)}_{q_1.q_2-1} - x_2\beta_2)'$. If we suppose that there are two pairs $(\alpha^{(1)}_{q_1.q_2}, \alpha^{(2)}_{q_1.q_2})$ and $(\widetilde{\alpha}^{(1)}_{q_1.q_2}, \widetilde{\alpha^{(2)}_{q_1.q_2}}($ of unknown thresholds consistent with observables. then we can deal with two different pairs $(\Delta_1, \Delta_2)$ and $(\widetilde{\Delta}_1, \widetilde{\Delta}_2)$ with $\Delta_1\equiv \alpha^{(1)}_{q_1.q_2}-\alpha^{(1)}_{q_1-1.q_2}$, $\Delta_2\equiv \alpha^{(2)}_{q_1.q_2}-\alpha^{(2)}_{q_1.q_2-1}$ and $\Delta_1\equiv \widetilde{\alpha}^{(1)}_{q_1.q_2}-\alpha^{(1)}_{q_1-1.q_2}$, $\Delta_2\equiv \widetilde{\alpha}^{(2)}_{q_1.q_2}-\alpha^{(2)}_{q_1.q_2-1}$ such that in Panel A in Figure (ref) the probability masses of the dark and light rectangles are exactly the same for any corner point $z_0$ in the two-dimensional space. The idea is then to possibly vary $z_0$ enough to show that this is not possible. An alternative consideration would be to proceed in different direction as the
commentAt each step in the second stage we proceed sequentially in a way to only having to identify two unknown threshold parameters at a time (rather than three thresholds). We then proceed in an analogous sequential way to cases of rectangular regions “in the middle” having, once again, only two unknown thresholds at a time (rather than three of four). With two known and two unknown thresholds defining a rectangular region $R_{j_1, j_2}$, the idea behind identification is to hypothetically suppose that there are two pairs of unknown thresholds consistent with observables. We can then think about from which “corner” (corresponds to four combinations of extreme value indices) we want to start the sequential identification process. This would imply that in one of the four situations depicted in Figure (ref) the probability masses of the red and green rectangles are exactly the same for any corner point $z_0$ in the two-dimensional space. The identification comes from the fact that no matter what the support of unobservable $(\varepsilon_1, \varepsilon_2)$ is, at least one of these four depicted situation will be inconsistent with observable probabilities of choice. Thus, there is always a “corner” from which we can start to establish the identification of thresholds. We fill in the details of this identification idea in Appendix (ref). \begin{figure}[!t] \caption{Illustration of the identification idea in the bivariate case with two know thresholds.} \begin{tikzpicture}[scale=0.7] \fill [opacity = 0.8, pattern=north west lines, pattern color=orange] (2,3) rectangle (0,0); \draw [very thick] (0,0) -- (0,3); \draw [very thick] (2,0) -- (2,3); \draw [very thick] (0,3) -- (2,3); \draw [very thick] (0,0) -- (2,0); \node[right] at (2,3) {$z_0+({\Delta}_1,\Delta_2)$}; \node at (2,3) [circle,fill,inner sep=2.5pt]; \fill [opacity = 0.8, pattern=crosshatch, pattern color=black] (5,1) rectangle (0,0); \draw [very thick] (0,0) -- (0,1); \draw [very thick] (5,0) -- (5,1); \draw [very thick] (0,1) -- (5,1); \draw [very thick] (0,0) -- (5,0); \node[above right] at (5,1) {$z_0+(\widetilde{\Delta}_1,\widetilde{\Delta}_1)$}; \node at (5,1) [circle,fill,inner sep=2.5pt]; \node at (0,0) [circle,fill,inner sep=2.5pt]; \node[below, left] at (0,0) { $z_0$}; \node [below=0.5cm, align=flush center,text width=2cm] { Panel A }; \end{tikzpicture} \hskip 0.4in \begin{tikzpicture}[scale=0.7] \fill [opacity = 0.8, pattern=north west lines, pattern color=orange] (3,3) rectangle (5,0); \draw [very thick] (3,3) -- (5,3); \draw [very thick] (5,0) -- (5,3); \draw [very thick] (3,0) -- (3,3); \draw [very thick] (3,0) -- (5,0); \fill [opacity = 0.8, pattern=crosshatch, pattern color=black] (0,1) rectangle (5,0); \draw [very thick] (5,0) -- (5,1); \draw [very thick] (5,0) -- (0,0); \draw [very thick] (5,1) -- (0,1); \draw [very thick] (0,1) -- (0,0); \node at (5,0) [circle,fill,inner sep=2.5pt]; \node[below, right] at (5,0) { $z_0$}; \node [below=0.5cm, align=flush center,text width=2cm] { Panel B }; \end{tikzpicture} \begin{tikzpicture}[scale=0.7] \fill [opacity = 0.8, pattern=north west lines, pattern color=orange] (2,3) rectangle (0,0); \draw [very thick] (0,0) -- (0,3); \draw [very thick] (2,0) -- (2,3); \draw [very thick] (0,3) -- (2,3); \draw [very thick] (0,0) -- (2,0); \fill [opacity = 0.8, pattern=crosshatch, pattern color=black] (5,2) rectangle (0,3); \draw [very thick] (0,2) -- (0,3); \draw [very thick] (5,2) -- (5,3); \draw [very thick] (0,2) -- (5,2); \draw [very thick] (0,3) -- (5,3); \node at (0,3) [circle,fill,inner sep=2.5pt]; \node[above, left] at (0,3) { $z_0$}; \node [below=0.5cm, align=flush center,text width=2cm] { Panel C }; \end{tikzpicture} \hskip 1.4in \begin{tikzpicture}[scale=0.7] \fill [opacity = 0.8, pattern=north west lines, pattern color=orange] (5,3) rectangle (3,0); \draw [very thick] (5,3) -- (5,0); \draw [very thick] (5,0) -- (3,0); \draw [very thick] (3,0) -- (3,3); \draw [very thick] (3,3) -- (5,3); \fill [opacity = 0.8, pattern=crosshatch, pattern color=black] (5,3) rectangle (0,2); \draw [very thick] (5,3) -- (5,2); \draw [very thick] (5,2) -- (0,2); \draw [very thick] (0,2) -- (0,3); \draw [very thick] (0,3) -- (5,3); \node at (5,3) [circle,fill,inner sep=2.5pt]; \node[above, right] at (5,3) { $z_0$}; \node [below=0.5cm, align=flush center,text width=2cm] { Panel D }; \end{tikzpicture} \caption*{Notes: in each picture, the red and green rectangles have the same probability mass when conjecturing that two different sets of thresholds can generate observables and when proceeding from one of the four “corners”.} \end{figure}

Estimation in a semiparametric model

The majority of existing estimation approaches for univariate semiparametric ordered response models (either under full stochastic independence of the unobservable from covariates or under a slightly more general formulation with a multiplicative scedastic function like in chenkhan2003) do not extend to multivariate models with general rectangular structures.\footnote{Most of these methods, however, can be extended to multivariate lattice models, as discussed in km2025lattice.} For example, in the two-stage approach of klein2002, which first estimates the index parameter using kernel density estimates of the conditional choice probabilities and then identifies the threshold parameters through shift restrictions, adapting the shift restrictions to non-lattice settings proves challenging. The same limitation applies to liu2019. Similarly, the approaches proposed in lewbel2000, lewbel2003 do not generalize to non-lattice frameworks. Moreover, the strategy of chenkhan2003 cannot be directly applied to models with general rectangular structures to estimate all finite-dimensional parameters of interest. In such settings, equality of joint probabilities $ P(Y^{(1)}=y^{(1)}_{j_1}, Y^{(2)}=y^{(2)}_{j_2} \ | \ x_1,x_2) = P(Y^{(1)} = y^{(1)}_{1}, Y^{(2)}=y^{(2)}_{1} \ | \ \tilde{x}_1, \tilde{x}_2)$ only implies that $\tilde{x}_d \beta_d = x_d \beta_d$ if and only if $x_{-d} =\tilde{x}_{-d}$, for $d=1,2$. Hence, their method may only be suitable for estimating parameters associated with exclusive covariates.

If we were interested in the estimation of parameters on exclusive covariates, we could proceed in many ways. We could combine pairwise differences \citep*{honore2005} with maximum rank correlation (MRC) estimation \citep*{han87} or any other method suitable for single-index models to estimate $\beta_{d,1:L_d}$ for exclusive regressors. Pairwise differencing would restrict attention to comparisons where non-exclusive covariates and all other dimensions are similar, while MRC would exploit the stochastic dominance which is behind the result in Theorem (ref).

We find that the only existing approach that can be extended to non-lattice model and which permits the estimation of all the index parameters as well as all the thresholds and the unknown joint distribution of unobservables is coppejans2007, which, analogously to us, relies on the assumption of independence of the unobservable from covariates as well as enough variation in covariates. Focusing on the bivariate case for clarity, let us describe how we can extend coppejans2007 to our setting.

Consider a random sample $\left\{(y^{(1)(i)}, y^{(2)(i)}, x_1^{(i)}, x_2^{(i)}) \right\}_{i=1}^N$. First, create an estimate of the probability of the bivariate latent process falling into the rectangle $R_{j_1,j_2}$:

align*[align* omitted — 216 chars of source]

coppejans2007 deals with the univariate $\widehat{F}$ and models it using a quadratic B-spline whose coefficients are estimated jointly with the index and threshold parameters. In the multivariate case, we can model $\widehat{F}$ using tensor-product B-splines and estimate their coefficients jointly with thresholds and index parameters. The estimation proceeds by \[ \max_{\theta, \widehat{F}} \mathcal{L}(\theta) = \frac{1}{N} \sum_{i=1}^{N} \sum_{j_{1}=1}^{M_{1}} \sum_{j_{2}=1}^{M_{2}} 1\left[(y^{(1)(i)}, y^{(2)(i)}) = (y_{j_{1}}^{(1)}, y_{j_{2}}^{(2)})\right] \log(\widehat{\ell}^{(i)}_{j_{1},j_{2}}), \] where $\theta$ combines all the index and threshold parameters. In the bivariate setting, the basis functions in tensor-product representation for $\widehat{F}$ consist of $S_1 \cdot S_2$ products $\mathcal{R}_{1;s_1,S_1}(e_1;q_1) \cdot \mathcal{R}_{2;s_2,S_2}(e_2;q_2)$, $s_1 = 1,\dots,S_1$, $s_2 = 1,\dots,S_2$, of univariate B-splines evaluated at specific $(e_1,e_2)$. Here $q_d$ is the degree in dimension $d=1,2$. There is a system of knots in each dimension which is not explicitly incorporated in our notation. The full tensor-product B-spline for $\widehat{F}(e_1,e_2)$ is a linear combination of these basis functions: \[ \sum_{s_1=1}^{S_1} \sum_{s_2=1}^{S_2} h_{s_1 s_2} \mathcal{R}_{1;s_1,S_1}(e_1;q_1) \mathcal{R}_{2;s_2,S_2}(e_2;q_2), \] with coefficients $\{h_{s_1 s_2}\}$ constrained to ensure valid c.d.f. properties. Specifically, (a) monotonicity in each dimension is enforced by

align*[align* omitted — 192 chars of source]

(b) c.d.f. bounds are maintained by $0 \leq h_{s_1 s_2} \leq 1 \quad \text{for all } s_1, s_2$.\footnote{For more details on shape constraints in tensor-product B-splines, see bk2022.} Linear equality constraints on some $h_{s_1 s_2}$ can also impose normalization restrictions on marginal distributions of unobservables. Coherency requires additional constraints on thresholds. In the bivariate case, these can be imposed via a penalty term added to the objective: e.g. in the form of

equation[equation omitted — 201 chars of source]

for a large $\lambda_N > 0$.

The distribution theory in coppejans2007 generalize as well with some obvious modifications required to make it applicable in multivariate non-lattice setting: regularity conditions and conditions on the growth of the B-spline based would need to be adjusted.

Parametric specification

The semiparametric framework offers a solid approach for achieving identification results despite model complexity, while also providing an estimation method that generalizes the approach of coppejans2007. However, in practice, researchers may opt for a parametric family for the joint distribution of unobserved $\boldsymbol{\varepsilon}$ due to computational convenience. This choice eliminates the need for nonparametric estimation of the joint distribution and typically reduces the data requirements for identifying all unknown parameters. These parametric families must, however, remain flexible to accommodate correlations among unobservables in the latent processes. Following established traditions in statistics and econometrics, natural choices for the joint distribution of unobservables are the multivariate normal distribution and a multivariate extension of a logistic distribution.

Even though parametric versions of the model will be identified under the conditions outlined in Section (ref), intuitively, much weaker conditions ensuring sufficient variation (potentially discrete and/or exclusive covariates) should suffice for identification in the parametric case. A useful analogy is the comparison between a univariate single-index model with an unknown monotone link function and the probit model. The probit model requires only a finite number of points satisfying a rank condition, whereas the single-index model demands richer variation, often guaranteed by the presence of a continuous covariate, which is not required in the probit model.

Adopting parametric assumptions within our general rectangular structure setting creates the following theoretical challenge for identification. While they may exist in theory, deriving clear-cut weaker conditions that do not involve exclusive or continuous covariates and are sufficient to identify all model parameters is complex. This difficulty mirrors the lack of straightforward identification conditions in multinomial probit or logit models that allow unknown correlations among different latent processes. We illustrate parametric identification without exclusive covariates through a numerical identification exercise.

Consider the case \(D=2\). Let \(\boldsymbol{\varepsilon}=(\varepsilon_1,\varepsilon_2)'\) be jointly normal with mean \((0,0)'\), unit variances, and correlation \(\rho\).\footnote{As usual, we normalize means and variances because shifts and scale changes yield observationally equivalent parameter vectors.} Denote its bivariate normal cumulative distribution function by \(\Phi_2(\cdot,\cdot;\rho)\). Focusing on one of the “corner” regions (for concreteness, the south-west quadrant), identification of the $k_1+k_2+3$ parameters $\beta_1, \beta_2, \rho,\alpha^{(1)}_{11},\alpha^{(2)}_{11}$ can be studied using \(T\geq k_1+k_2+3\) distinct covariate values \(x_i=(x_{i1},x_{i2})'\) by writing a system of \(T\) equations in the \(k_1+k_2+3\) unknowns: for \(i=1,\dots,T\)

equation[equation omitted — 210 chars of source]

where the left-hand side \(p_{\mathrm{obs}}(x_i)\) is observable. Larger \(T\) or the presence of covariates exclusive to one of the processes can aid identification, but identification can proceed even when all covariates are shared, provided the \(x_i\) vary sufficiently.

To illustrate this, consider the case of each latent processes having a single covariate which is shared between the two processes. Then $\theta=(\beta_1,\beta_2,\alpha^{(1)}_{11},\alpha^{(2)}_{11},\rho)$, has five unknowns. We perform a numerical identifiability exercise around the true \(\theta_0=(1.2,\,-0.8,\;0.5,\,-0.2,\;0.4)\). Define the objective function for a candidate parameter \(\theta\) as \[ Q_T(\theta) = \sum_{i=1}^T \left( p_{\mathrm{obs}}(x_i) - \Phi_2\big(\alpha^{(1)}_{11} - x_{i1}'\beta_1,\; \alpha^{(2)}_{11} - x_{i2}'\beta_2,\; \rho\big) \right)^2. \] We compute \(p_{\mathrm{obs}}(x_i)\) from the known \(\theta_0\). The experiment is conducted for two designs: (i) \(T=5\) design points drawn from the interval \([-2,2]\); and (ii) \(T=10\) points obtained by adding five more draws from the same interval.

To explore the local geometry of \(Q_T\) around \(\theta_0\), we reparametrize \(\rho\) as \(z=\operatorname{atanh}(\rho)\) (so \(z\in\mathbb{R}\) and \(\rho=\tanh (z)\in[0,1]\)), and then consider hyperspheres in this transformed parameter space. For a given radius \(r\) we sample 5000 directions on the sphere of Euclidean radius \(r\) about \(\theta_0\); for each sampled point \(\theta\) we evaluate \(Q_T(\theta)\) and record the minimum value found on that sphere. Two direction-sampling schemes are applied: fixed-direction sampling in which we draw random directions once and scale them to each radius \(r\), and random-sphere sampling in which we draw new random directions independently for each radius \(r\).

Figure (ref) shows the resulting plots (log scale) of the minimum objective \(Q_T\) versus radius \(r\) for both sampling schemes (left: fixed-direction; right: random-sphere). These plots display how well the model discriminates the true parameter vector from alternatives at varying distances, and the extent to which this discrimination improves with the number of design points \(T\) (shown through the fact that $Q_{T}(\theta)$ is increasing in $T$ for all radii $r$). The pronounced jitter visible in the random-sphere plot reflects sampling variability across radii.

\paragraph*{Estimation} To outline an estimation approach in the case of parametric assumptions on the distribution of unobservables, we continue with $D=2$ and $\boldsymbol{\varepsilon}=(\varepsilon_1,\varepsilon_2)'$ being jointly normal with zero mean, unit variances, and correlation \(\rho\). Given a random sample $\left\{(y^{(1)(i)}, y^{(2)(i)},x_1^{(i)},x_2^{(i)}) \right\}_{i=1}^N$ and collecting $\beta_1,\beta_2,\rho$ and all the thresholds in $\alpha$ in one parameter vector $\theta$, we construct the log-likelihood function

eqnarray*[eqnarray* omitted — 274 chars of source]

$$\text{ with } \quad \ell^{(i)}_{j_{1},j_{2}} = \sum_{t_1=0}^{1} \sum_{t_2=0}^{1} (-1)^{t_1 + t_2} \Phi_{2}\left( \alpha_{j_1 - t_1, j_2}^{(1)} - x_1^{(i)}\beta_1, \alpha_{j_1, j_2 - t_2}^{(2)} - x_2^{(i)}\beta_2; \rho \right).$$ Analogously to semiparametric model case, this log-likelihood needs to be maximized subject to a linear inequality constraints that describe ordering of thresholds in each dimensions, normalization constraints on the the thresholds, and non-linear equality constraints $q(\alpha)=0$ that collect coherency constraints ((ref)) across all the local models.

figure[figure omitted — 544 chars of source]

The constrained maximum likelihood estimator (MLE) $\hat{\theta}$ solves the optimization problem $\max_{\theta} \mathcal{L}(\theta)$ subject to the described constraints on thresholds. The coherency constraints are differentiable and so under the typical MLE regularity conditions (e.g. newey1994), we have $\sqrt{N}(\hat{\theta}-\theta_{0}) \overset{d}{\longrightarrow} \mathcal{N}(0,V)$, where $V = B J B'$, with $J = \mathbb{E}\left[\frac{\partial \log(\ell^{(i)}(\theta_0))}{\partial \theta}\frac{\partial \log(\ell^{(i)}(\theta_0))}{\partial \theta'}\right]$, $B=J^{-1} - J^{-1}Q'(QJ^{-1}Q')^{-1}QJ^{-1}$, and $Q = \dfrac{\partial q(\theta_{0})}{\partial \theta'}$. The natural plug-in sample-analogue estimator of $V$ provides a consistent estimator for the variance-covariance matrix.

Monte Carlo experiments

We now examine Monte Carlo simulations for the parametric case with normal errors, as outlined in Section (ref). We compare the constrained maximum likelihood estimator described in Section (ref) for models with general rectangular structures (we will refer to it from now as non-lattice probit) to the standard bivariate ordered probit estimator (that is, the estimator of a lattice model under normal errors). The baseline model is $$Y^{*c_1} = x\beta_{1} + w_{1}\gamma_{1} + \varepsilon_{1}, \qquad Y^{*c_2} = x\beta_{2} + w_{2}\gamma_{2} + \varepsilon_{2},$$ where unobservables are independent of regressors and jointly normal with zero means and unit variances. This setup distinguishes exclusive and non-exclusive covariates. We explore scenarios with no exclusive covariates ($\gamma_1=\gamma_2=0$) and with an exclusive covariate in one latent process. The key findings are: (i) non-lattice model parameters may be estimated well without exclusive covariates, and (ii) using lattice instead of non-lattice models can yield inconsistent estimators for all parameters by ignoring broad bracketing in the decision process. Our simulations and applications show cases in which an expected positive correlation between unobservables is estimated to be significantly negative under a lattice model. The inconsistency in index parameters depends on how well the lattice model approximates the true non-lattice model.

Each simulation design uses 250 independent random samples of size $N=5,000$, with the penalty term in ((ref)) set to $N$. Alternative $N$ and $\lambda_N$ values yield similar results.

figure[figure omitted — 1,667 chars of source]

Design 1: 2$\times$2 structure, no excluded regressors

We investigate parameter estimation in non-lattice probit models without exclusive covariates by setting $\gamma_1=\gamma_2=0$, thereby removing $w_1$ and $w_2$. We set $\beta_1=1$, $\beta_2=0.5$, $\rho=0.33$, and use a $2\times2$ non-lattice structure with thresholds $\alpha^{(2)}_{11}=\alpha^{(2)}_{21}=1$, $\alpha_{11}^{(1)}=-2$, and $\alpha_{12}^{(1)}=1.5$ (see Figure (ref)). We consider three distributions for the common regressor $x$: (Design 1A) uniform $[-5,5]$, (Design 1B) discrete on 10 points $\{\pm 5, \pm 3.5, \pm 2.5, \pm 1.5, 0, 0.5\}$ with equal probabilities, and (Design 1C) discrete on five points $\{\pm 5, \pm 2.5, 0\}$ with equal probabilities.

Table (ref) reports mean and standard deviation of parameter estimates for Design 1A. The non-lattice method estimates all parameters with minimal bias. The bivariate lattice ordered probit method estimates $\beta_1$ and the first-dimension threshold well, but performs poorly for $\beta_2$, $\rho$, and the second-dimension thresholds. Estimates of $\rho$ are not very precise due to the absence of excluded regressors.

For Designs 1B and 1C, we estimate using only the non-lattice probit method. Results in Table (ref) show that all parameters are estimated well, even with the minimal discrete variation in $x$ in Design 1C. As expected, as a result of the limited variation in the common regressor, Design 1C exhibits higher standard deviations.

table[table omitted — 1,548 chars of source]
table[table omitted — 1,335 chars of source]

Design 2: 4$\times$3 with one excluded covariate

In the second simulation design, we extend the number of discrete values $M_{d}$ in both dimensions. The discrete dependent variable $Y^{c_1}$ can take four values and $Y^{c_2}$ can take three values. This generates a 4$\times$3 non-lattice structure, illustrated in Figure (ref). The common covariate $x$ follows a uniform $[-3,3]$ distribution, though continuity here is not necessary. The covariate $w_{1}$ is a discrete random variable taking values -2.5, -1.5, -0.5 and 0.5 with equal probability. We set $\gamma_2=0$ thus effectively removing $w_{2}$ in the second equation. The parameter values are $\beta_{1} = 1.5, \gamma_{1} = -4, \beta_{2} = 3$ and $\rho = 0.5$.

Table (ref) lists the across-simulation means and standard deviations of the index parameters and the correlation coefficient. Table (ref) in the online supplement provides the values of the thresholds, together with their estimated means and standard deviations. The non-lattice bivariate ordered probit method estimates all the parameters with almost no bias. On the contrary, the lattice bivariate ordered probit method estimates the parameters with a relatively large bias. The mean squared errors in the non-lattice method are far lower than those in the lattice method for all of the parameters. Assuming a lattice structure makes estimating the correlation parameter $\rho$ decidedly difficult, with the method failing to estimate the correct sign for $\rho$, let alone an approximately close value.

figure[figure omitted — 3,272 chars of source]

Cryptocurrency Familiarity and Optimism

In the first of two applications, we use data from the Survey of Consumer Payment Choice (SCPC) foster2021 to study opinions on future movements in cryptocurrency prices.\footnote{See also benetton2022 and kahn2016.} Conducted annually by the Federal Reserve Banks of Atlanta, Boston, and San Francisco, the SCPC tracks U.S. consumers’ payment method adoption, recently noting a shift toward online and mobile payments due to the COVID-19 pandemic. Our sample includes 4,600 individuals (2015–2020), with data on demographics (e.g., income, age, gender, education), payment method use (e.g., credit cards, cryptocurrencies, mobile platforms like Google Pay), and perceptions of safety, convenience, and cost. Additional data cover fraud exposure, FICO score ranges, and household financial roles. See foster2021 for details.

We examine whether opinions on bitcoin’s future value are interdependent with cryptocurrency familiarity. $Y^{c_1}$ is an ordered variable for familiarity with bitcoin (-1: not familiar, 0: slightly familiar, 1: somewhat familiar, 2: moderately/extremely familiar).\footnote{Moderate and extremely familiar are combined due to few respondents reporting extreme familiarity.} For $Y^{c_2}$, we use an ordered variable for bitcoin’s expected value in one year (-1: decrease, 0: no change, 1: increase). We use a bivariate normal distribution specification for the vector of unobservables with zero mean and unit variances and an unknown correlation $\rho$.

table[table omitted — 802 chars of source]

Applications

Figure (ref) shows the estimated threshold structure for a non-lattice model. It reveals varied thresholds. For example, individuals with low familiarity (blue dots) expect no value change at a lower threshold in the opinion dimension (i.e., lower $Y^{*Opinion}$ values yield a "no change" opinion compared to other familiarity groups). Conversely, those with high familiarity (red crosshatch) have closely spaced thresholds, defining a narrow "no change" region. For this group, most $Y^{*Opinion}$ values reflect strong opinions, either decreasing or increasing. These findings describe decision-making structures, not probabilistic outcomes.

Table (ref) provides estimates of $\beta$ and $\rho$. The correlation $\rho$ ranges from 0.03 (lattice model) to 0.84 (non-lattice model), with coefficients differing by over 20% in magnitude. Notably, the coefficient on male changes sign -- it is negative (and statistically significant at the 5% level) in the lattice model and positive (not statistically significant at the 5% level) in the non-lattice model. The lattice model thus suggests males are more pessimistic about bitcoin’s value as the effect as $P(Y^{Optimisim} \geq j|x_{-male}, x_{male}=1)- P(Y^{Optimisim} \geq j|x_{-male}, x_{male}=0)$ is negative for any level $j$ and any $x_{-male}$. As discussed earlier, a non-lattice model allows for changes in the signs of partial effects across the domain, therefore, the positive coefficient for males obtained there does not directly imply that this model suggests males are more optimistic for any level $j$ and any $x_{-male}$. Additional post-estimation analysis we have conducted does confirm, however, that given the estimated thresholds the non-lattice model gives $P(Y^{Optimisim} \geq j|x_{-male}, x_{male}=1)- P(Y^{Optimisim} \geq j|x_{-male}, x_{male}=0)$ as positive for any level $j$ and any $x_{-male}$ in the data, thus, giving a stable sign of this partial effect across the domain.\footnote{A potentially different estimated threshold structure could have resulted in the switching of signs for this partial effect.}.

small\begin{table}[tbp] \caption{Estimation coefficients: bitcoin familiarity and optimism} \begin{threeparttable} \begin{tabular}{lccccccccccccc} \toprule Variable && O-Probit && O-Probit && Non-lattice && Lattice \\ \hline Familiarity with Bitcoin\\ \hline \textsc{Low Income} && -0.16 (0.06) && && -0.11 (0.05) && -0.16 (0.06) \\ \textsc{Age} && -0.02 (0.00) && && -0.02 (0.00) && -0.01 (0.00) \\ \textsc{Male} && 0.55 (0.06) && && 0.42 (0.06) && 0.55 (0.06) \\ \textsc{Low Education} && -0.54 (0.09) && && -0.40 (0.09) && -0.54 (0.09) \\ \\ \textit{Bitcoin “optimism”}\\ \hline \textsc{Low Income} && && 0.07 (0.06) && -0.00 (0.06) && 0.07 (0.06) \\ \textsc{Age} && && -0.01 (0.00) && -0.01 (0.00) && -0.01 (0.00) \\ \textsc{Male} && && -0.13 (0.05) && 0.02 (0.06) && -0.13 (0.05) \\ \textsc{Low Education} && && 0.13 (0.07) && -0.00 (0.08) && 0.13 (0.07) \\ \hline $\rho$ && NA && NA && 0.84 (0.23) && 0.03 (0.03) \\ N && 1818 && 1818 && 1818 && 1818 \\ \bottomrule \end{tabular} \begin{tablenotes}[flushleft] • Notes: Columns labeled “O-probit” provide estimates from univariate ordered probit models. The “Non-lattice” column provides estimates from using non-lattice bivariate ordered probit model. The “Lattice” column assumes a lattice structure, but estimates the two equations jointly. \end{tablenotes} \end{threeparttable} \end{table}

Another aspect illustrated by the thresholds in Figure (ref) regarding the decision-making process is that individual decisions can be modeled sequentially using a binary decision tree, where each node represents a decision based on a single latent process. This structure can be termed a hierarchical non-lattice model, a specific subset of non-lattice models.\footnote{Any model with a general rectangular structure can be represented by a decision tree, but typically, each node may involve multiple or all latent processes.} Hierarchical non-lattice models maintain coherence, as each node in the decision tree further refines the partitioning of the latent space. Figure (ref) in the online supplement depicts the binary decision tree that outlines the estimated hierarchical decision-making process in this cryptocurrency application.

{\tiny

figure[figure omitted — 3,423 chars of source]

}

Identifying moral hazard and adverse selection in insurance markets

In the empirical analysis of asymmetric information in insurance markets, a highly influential framework was introduced by chiappori2000testing, which proposed testing for adverse selection by estimating a bivariate probit model linking insurance coverage decisions and ex post risk realizations and taking a positive correlation between the latent errors of the insurance and risk equations as evidence of asymmetric information. Subsequent research has emphasized the limitations of this framework, pointing out its inability to disentangle (adverse) selection from moral hazard because both forces contribute to the estimated positive correlation, but with different policy implications. As emphasized in cohen2010, “the disentanglement of adverse selection and moral hazard is probably the most significant and difficult challenge that empirical work on adverse selection in insurance markets faces.” Recent work has attempted to move beyond these limitations. On the theoretical side, a large literature has developed richer models of contract choice and risk response (e.g. EinavFinkelsteinLevin2010, hendren2013ECTA). Empirically, researchers have leveraged quasi-experimental variation or structural models to separately identify selection and moral hazard (e.g. Handel2013AER, Einav_etal2013AERselectiononmoralhazard. HackmannKolstadKowalski2015, among many other). However, these approaches often require highly specific data environments or strong structural assumptions.

Our general rectangular structure framework provides an alternative empirical strategy. First, it allows utilization thresholds to vary with insurance status, thereby directly incorporating individuals' behavioral (“moral hazard”) response to coverage into the econometric model. At the same time, it permits insurance-choice thresholds to depend on anticipated behavioral responses to insurance, accommodating what Einav_etal2013AERselectiononmoralhazard and others term “selection on moral hazard.” Of course, coherency still needs to be satisfied.

We apply this general rectangular structure to U.S. health insurance markets using the Medical Expenditure Panel Survey (MEPS), which offers nationally representative data on insurance coverage and healthcare utilization. Our sample includes approximately 60,000 individuals from 2005 to 2010, prior to the Affordable Care Act. Following the standard framework, we specify latent insurance and utilization equations as

equation[equation omitted — 185 chars of source]

Common covariates include demographics (logged income and its square, dividend payments, family size, logged hourly wage, age, education, gender, marital status, race, region, and year) and pre-existing conditions (diabetes, asthma, high blood pressure, high cholesterol, angina, heart attack, stroke, emphysema, arthritis). An excluded covariate for $x_{\text{ins}}$ is partner’s job-provided coverage.

The insurance outcome $Y_{\text{ins}} = 1$ if the individual holds private health insurance in January and is 0 otherwise. Utilization $Y_{\text{use}}$ is measured categorically (0 for no charges, 1 for below-median charges, 2 for above-median charges)

Moral hazard is isolated in the general rectangular structure by allowing utilization thresholds to depend on coverage. E.g., $Y_{\text{ins}} = 1$ lowers these thresholds if moral hazard is present. Insurance-dimension thresholds may potentially differ across utilization responses, provided coherency is maintained. One might argue, however, that due to the natural sequencing, where insurance coverage is selected first and utilization occurs second, the thresholds in the insurance dimension are invariant to utilization status. This restriction can be directly incorporated into the identification and estimation processes, automatically ensuring coherency while simplifying both identification and estimation. In this case, moral hazard is fully captured by differences in utilization thresholds across coverage levels. Once this behavioral effect is accounted for, the remaining correlation between $\varepsilon_{\text{ins}}$ and $\varepsilon_{\text{use}}$ can be interpreted as evidence of adverse or advantageous selection.

If, in contrast, the thresholds in the insurance-coverage dimension are permitted to vary with utilization status (in addition to utilization thresholds varying with insurance coverage), this accommodates the aforementioned “selection on moral hazard” where individuals may select coverage partly based on their anticipated behavioral (“moral hazard”) response to insurance. Then, the correlation between $\varepsilon_{\text{ins}}$ and $\varepsilon_{\text{use}}$ represents residual adverse selection, while moral hazard continues to be captured by lower utilization thresholds when $Y_{\text{ins}} = 1$.

We estimate the model with a general rectangular structure subject only to coherency constraints, thus potentially allowing for “selection on moral hazard” and letting the model and the data reveal to us in particular if that phenomenon exists or whether the choices can be considered to be sequential and consistent with the widely perceived natural timing of things.

figure[figure omitted — 2,164 chars of source]

Estimation results are presented in Figure (ref). First, the estimated general rectangular structure model does reveal moral hazard through coverage-dependent thresholds. Utilization thresholds shift downward when $Y_{\text{ins}} = 1$: low-to-medium usage shifts from $-0.02$ (0.04) uninsured to -0.45 (0.07) insured, and medium-to-high from $0.68$ (0.04) to 0.46 (0.08). Second, the results reveal no “selection on moral hazard” as the insurance coverage thresholds are estimated as invariant to utilization levels. The adverse selection is captured by the correlation coefficient $\rho$ which drops from 0.21 (0.01) in the lattice model to 0.04 (0.02) in the non-lattice model, suggesting minimal adverse selection after accounting for moral hazard. Our findings align with a growing body of evidence from structural and experimental studies (e.g. Einav_etal2013AERselectiononmoralhazard, Handel2013AER), which generally find limited adverse selection once moral hazard is accounted for. Note that additionally the estimated structure allows for a cross-partial (indirect) effect of partner’s job-provided coverage on the degree of utilization, even though this variable does not enter directly the latent process in the utilization dimension.

Thus, our general rectangular structure approach disentangles moral hazard from adverse selection while avoiding strong assumptions, using only observational data and flexible thresholds. It outperforms traditional lattice models and, because it can be applied with standard survey or administrative data, provides a powerful way to revisit much of the empirical literature built on lattice-based or ordered probit models.

Conclusion

This paper proposes a framework with general rectangular structures that extends the reach of traditional ordered response models, thereby providing new insights into economic decision-making processes. The framework captures interactions across dimensions in both decision rules and latent factors, with traditional lattice models as a special case. By formalizing coherency, deriving utility-based microfoundations, and proving identification, our framework enables realistic modeling of complex multidimensional choices. The approach reveals sign-changing partial effects, indirect covariate influences, and varying complementarity or substitutability, all of which are absent in lattice models. In doing so, it expands the econometric toolkit for studying multidimensional decisions and improves interpretability. As our empirical examples demonstrate, existing important economic contexts such as selection markets can be revisited (and new environments can be traversed) with the models we have introduced.

\setcounter{page}{1} \setcounter{section}{0}

center[center omitted — 53 chars of source]