Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
81,487 characters · 13 sections · 65 citation commands
Learning Preferences from Conjoint Data: A Structural Deep Learning Approach
\thispagestyle{empty} \setcounter{page}{1}
Conjoint experiments hold great promise for helping researchers understand how individuals trade off different dimensions of preference. By presenting respondents with multidimensional profiles and asking them to choose, they directly elicit the multi-attribute tradeoffs inherent in real-world decisions. This insight was recognized early: greenhalgh1981 used conjoint analysis to study negotiator preferences over contract terms, and shamir1995 explicitly framed conjoint designs as a way to measure value tradeoffs and interactions in mass opinion. Since hainmueller2014causal introduced conjoint experiments to political science, applications have proliferated across a number of areas from immigration policy hainmueller2015hidden, bansak2016europeans to candidate evaluation saha2022welfare, democratic accountability graham2020democracy, and tax policy preferences ballard2017structure.
The ability of conjoints to elicit multi-attribute tradeoffs connects tightly to theoretical quantities at the core of our understanding of political choice---quantities that arise from specifying an underlying utility function that governs choice behavior. These quantities, such as the marginal rate of substitution across preference dimensions, willingness to pay for a policy attribute, or the fraction of a population that prefers one alternative over another, are central to theories of voting, representation, and political economy, and conjoint experiments are, in principle, ideally suited to estimate them. But because estimating these quantities requires imposing parametric structural assumptions on preferences, and because the recent turn toward design-based causal inference in political methodology has been skeptical of such structural assumptions, the development of conjoint analysis in political science focused primarily on nonparametric causal estimands. The average marginal component effect (AMCE) introduced by hainmueller2014causal and subsequent extensions using BART robinson2024detect, causal forests abramson2022, Bayesian mixtures goplerud2025, and design-based hypothesis tests ham2024 have dominated applied work in the field. This has meant that the underlying ability of conjoint designs to speak to structural quantities of interest is underdeveloped and under-explored.
Skepticism of structural assumptions is well-motivated---structural models can be wrong and the consequences of misspecification can be severe. In this paper, we address this concern by introducing a method that combines the conjoint design with deep neural networks (DNNs) farrell2021deep, farrell2025deep and double/debiased machine learning (DML) chernozhukov2018double. We embed a DNN within a random utility logit model: the network takes each respondent's observed characteristics as input and outputs that respondent's complete preference vector, the full set of marginal utilities over all attribute levels in the conjoint design. This preference vector enters a structural logit choice model, so the predicted choice probability is determined by the randomly assigned profile contrast and the respondent's estimated preferences jointly. Because the logit structure is preserved, all structural quantities remain well-defined and computable from the estimated preference vectors. The DNN replaces the distributional assumptions of standard mixed logit with a fully flexible mapping from respondent characteristics to preferences, addressing the concern that any particular parametric specification may not capture the data generating process. For inference on population-average preference parameters, we use cross-fitting and the influence-function correction of DML, which provides valid confidence intervals (CIs) despite the flexibility of the first-stage estimator. We think of this approach as lean structure---the upside of a structural model without its usual parametric downside.
The payoff is substantial, and it speaks directly to an active debate about what conjoint experiments can and cannot deliver. bansak2023amce provide a systematic treatment of what conjoint experiments can---and cannot---aggregate, showing that different quantities of interest (the AMCE, the effect of an attribute on the probability of winning, the fraction of voters who prefer a given attribute level) carry different substantive meanings and that the choice of quantity should be driven by the underlying question; a central implication of their analysis is that many theoretically natural quantities---in particular, the fraction of voters who prefer an attribute level---are essentially infeasible to recover using standard reduced-form methods, because those methods identify only marginal population averages and therefore cannot point-identify individual-level distributional features of preferences without additional structure. Our paper develops the structural machinery to recover individual-level preferences directly, and therefore to deliver precisely those quantities of interest that bansak2023amce and others have identified as substantively desirable but out of reach under existing reduced-form approaches. The power of our approach comes from combining two sources of leverage: the flexible structure of a deep neural network, which lets each respondent's preference vector depend on covariates in an essentially unrestricted way, and the identification power of randomization, since the conjoint design randomly assigns profile contrasts to each respondent, pinning down the structural preferences without the strong functional-form assumptions that traditional structural models typically invoke.
We apply our method to three prominent conjoint studies, showing how the structural approach reveals deep preference heterogeneity. In the saha2022welfare candidate-choice experiment, we recover the full distribution of preferences across respondents and find that the vast majority prefer female candidates---near-consensus in one direction, masked by a near-zero average effect. Candidate empathy shows the opposite pattern: its AMCE is near zero, but this masks a sharp partisan cleavage---Democrats reward empathy while Republicans penalize it---so the positive and negative individual-level preferences cancel almost exactly in the average. In the graham2020democracy democracy conjoint, the structural model reveals that virtually all voters oppose undemocratic behavior, but many weight party and policy more heavily---the disagreement is about intensity, not direction. In the ballard2017structure tax-plan experiment, we recover an individual-level preferred rate schedule for each respondent and find that nearly all have progressive revealed preferences---a near-universal feature across every partisan, income, and ideological subgroup---while Democrats and Republicans differ primarily in which brackets drive their choices. The same application provides an individual-level external validation: the DNN-recovered progressivity slope correlates well with each respondent's self-reported ideal tax rates---a measure collected independently of the conjoint task and never seen by the model during training. Our results show that by focusing on nonparametric causal estimands, prior conjoint studies have been unable to uncover this rich preference heterogeneity.
In the marketing literature, it has long been known that conjoint studies can identify the structural parameters of an underlying utility model. The conjoint measurement foundations laid by luce1964simultaneous and green1971conjoint were linked to random utility models by mcfadden1974conditional, and the connection between conjoint experiments and structural discrete choice has been a staple of marketing research green1978conjoint, green1990conjoint, louviere2000stated, train2009discrete.\footnote{The introduction of conjoint analysis to political science by hainmueller2014causal does discuss the random utility foundation, citing luce1964simultaneous, green1971conjoint, and mcfadden1974conditional; see bansak2021conjoint for a comprehensive handbook treatment. Our contribution is to develop this structural approach and the quantities it enables.} Random utility models have also been used extensively in political science---notably the probabilistic spatial-utility frameworks of poole1985spatial and palfrey1987relationship, the heterogeneous voter-choice models of rivers1988heterogeneity and alvarez1998multiparty, and the Bayesian ideal-point models of martin2002dynamic and clinton2004statistical---and the formal literature on spatial voting and preference aggregation enelow1984spatial, hinich1997analytical, austensmith1999positive provides the theoretical foundations for the structural quantities we recover. But these traditions have not been linked to conjoint experiments. Our paper makes this connection, and by using deep neural networks to flexibly estimate preference heterogeneity, addresses the main worry that political scientists have about parametric identification: that the specified model may not be capturing the underlying data generating process.
This section develops the structural model that underlies our approach. We begin with the conjoint experimental setup and the random utility model that connects profile attributes to respondent choices. We then define the structural quantities of interest that the model enables and show that these quantities are identified by the conjoint randomization design.
Consider a conjoint experiment in which each of $M$ respondents evaluates $T$ choice tasks.\footnote{We use constant-$T$ notation because the number of tasks per respondent is fixed by design in most conjoint studies; all formulas extend immediately to varying $T_i$.} In each task $t$, respondent $i$ views two profiles (alternatives) and selects one. Each profile $j \in \{1, 2\}$ is characterized by a vector of randomly assigned attribute levels, encoded as dummy variables relative to a reference category: $\mathbf{X}_{ijt} \in \mathbb{R}^p$, where $p$ is the total number of non-reference attribute levels across all attributes. To fix ideas, consider the candidate-choice experiment of saha2022welfare, where respondents evaluate hypothetical candidates for public office described by five attributes---policy agenda, talent, number of children, gender, and prior candidacy---yielding $p = 13$ dummy-coded attribute levels.
We define the profile-pair difference
which captures the contrast between the two profiles on all attribute dimensions simultaneously. Let $Y_{it} \in \{0,1\}$ denote whether respondent $i$ chose profile 1 in task $t$. The total number of choice observations is $N = MT$.
Respondent characteristics are collected in a vector $\mathbf{Z}_i \in \mathbb{R}^{p_Z}$, which may include demographics (age, gender, education, income), attitudinal measures (political ideology, party identification), and contextual variables (region, urban/rural residence). For example, in the candidate-choice experiment, $\mathbf{Z}_i$ includes the respondent's party identification, ideology, gender, age, education, income, race, and region---15 variables in total. Crucially, $\mathbf{Z}_i$ is constant across all tasks for respondent $i$.
We ground our framework in the canonical random utility model of mcfadden1974conditional. Respondent $i$ derives utility from profile $j$ in task $t$ according to
where $\boldsymbol{\beta}(\mathbf{Z}_i) \in \mathbb{R}^p$ is a vector of marginal utilities that varies as an unknown function of respondent characteristics, and $\varepsilon_{ijt}$ are independent and identically distributed Type I Extreme Value (Gumbel) taste shocks.
Each coefficient $\beta_k(\mathbf{Z}_i)$ has a direct interpretation: it is the marginal utility that respondent $i$ assigns to attribute level $k$ relative to the reference level of the corresponding attribute. For instance, in the candidate-choice experiment, $\beta_{\text{Female}}(\mathbf{Z}_i)$ measures how much respondent $i$ values a female candidate relative to a male candidate---and this valuation may differ across respondents with different party identification, education, or ideology. The key modeling choice is that $\boldsymbol{\beta}(\cdot)$ is treated as a fully nonparametric function: we impose no distributional assumptions on how preferences relate to respondent characteristics. This is what distinguishes our approach from standard mixed logit, which assumes $\boldsymbol{\beta}_i \sim N(\boldsymbol{\mu}, \boldsymbol{\Sigma})$ or restricts the dependence on covariates to be linear.
The respondent chooses profile 1 if $U_{i1t} > U_{i2t}$. Since the difference of two independent and identically distributed Gumbel random variables follows a logistic distribution, the forced-choice probability is:
where $G(v) = (1 + e^{-v})^{-1}$ is the logistic cumulative distribution function. The profile-pair differencing (ref) eliminates any alternative-specific constant, so we need only estimate the single parameter function $\boldsymbol{\beta}(\cdot)$.
The key assumption of this model is an additive utility that is linear in attribute levels, which rules out attribute interactions unless explicitly included. The logistic link from the Gumbel error distribution is a secondary modeling choice---common alternatives such as probit are nearly indistinguishable in practice. While $\boldsymbol{\beta}(\mathbf{Z})$ is left fully nonparametric, the additive utility structure is what distinguishes the approach from purely design-based methods and is the source of its additional identifying power. We discuss its limitations in Section (ref), where we also discuss some possible extensions to address these limitations.
The structural model defined by equations (ref) and (ref) enables a set of quantities that are inaccessible to reduced-form analysis. We focus here on the quantities that we estimate in our applications below; all require the individual-level preference vector $\boldsymbol{\beta}(\mathbf{Z}_i)$ and cannot be recovered from non-parametric causal estimands or from heterogeneous treatment effect estimators that operate one attribute at a time.
The preference vectors also support a number of additional structural quantities that may be of interest depending on the application, including: consumer surplus via the closed-form welfare measure of mcfadden1981econometric, indifference sets that map individual-specific substitution patterns geometrically, preference clustering via $k$-means or other algorithms applied to $\hat{\boldsymbol{\beta}}(\mathbf{Z}_i) \in \mathbb{R}^p$, preference inequality measures such as $\mathrm{Var}(\hat\beta_k(\mathbf{Z}_i))$ or Gini coefficients, decisiveness measures $|(\mathbf{X}_A - \mathbf{X}_B)^\top \hat{\boldsymbol{\beta}}(\mathbf{Z}_i)|$ capturing the strength of preference for a given profile comparison, and the majority preference function---for any profile pair $(A, B)$, the fraction of respondents whose deterministic utility index $(\mathbf{X}_A - \mathbf{X}_B)^\top \hat{\boldsymbol{\beta}}(\mathbf{Z}_i) > 0$.\footnote{Unlike logit choice probabilities, the majority preference function abstracts from idiosyncratic taste shocks and directly addresses the concern raised by abramson2022 that the AMCE can indicate the opposite of the true majority preference. See also our discussion of this criticism below.} We refer the reader to train2009discrete, louviere2000stated, and green1990conjoint for detailed treatments of these quantities.
\paragraph{Relations to the AMCE.} As noted above, the structural AME recovers the AMCE under correct specification, so the structural model complements rather than replaces what bansak2023amce show is a valuable estimand that maps directly to aggregate vote shares. The key limitation of the AMCE is that it estimates the effect of one attribute marginalized over all others and therefore cannot recover the joint preference vector for a given respondent. This means it cannot deliver marginal rates of substitution, counterfactual choice probabilities for complete profiles, compensating differentials, attribute importance decompositions, or preference polarization---each of which is a functional of $\boldsymbol{\beta}(\mathbf{Z}_i)$. Conditional AMCEs interact treatment indicators with covariates or estimate the AMCE within subgroups hainmueller2014causal, leeper2020measuring, but they still average within subgroups rather than recovering individual-level preferences. Relatedly, delacuesta2022 note that under preference heterogeneity the AMCE depends on the profile-randomization distribution and is not a structural parameter of the utility model, and abramson2022 show that the AMCE can indicate the opposite of the true majority preference---a quantity the structural approach point-identifies for each respondent. Section (ref) illustrates these distinctions in practice.
All of the quantities of interest defined above---average preference parameters, average marginal effects, individual preference vectors, preference polarization, attribute importance, marginal rates of substitution, compensating differentials, and counterfactual choice probabilities---are derived from the individual-level preference function $\boldsymbol{\beta}(\mathbf{Z})$. It is therefore sufficient to show that $\boldsymbol{\beta}(\mathbf{Z})$ itself is identified, since every quantity of interest listed above is a known functional of $\boldsymbol{\beta}(\mathbf{Z})$ and inherits identification from it. We now show that $\boldsymbol{\beta}(\mathbf{Z})$ is identified whenever (i) the profile contrast $\boldsymbol{\Delta}\mathbf{X}_{it}$ is randomly assigned independently of respondent characteristics $\mathbf{Z}_i$, and (ii) the random utility model (ref)--(ref) is correctly specified. Both conditions are built into the conjoint design.
By construction, $\boldsymbol{\Delta}\mathbf{X}_{it}$ is randomly assigned and therefore independent of $\mathbf{Z}_i$: \[ \boldsymbol{\Delta}\mathbf{X}_{it} \perp \!\!\! \perp \mathbf{Z}_i. \] This exogeneity condition, guaranteed by the experimental design, means that the conditional choice probability (ref) identifies the preference function $\boldsymbol{\beta}(\mathbf{Z})$ without concerns about omitted variable bias or endogenous sorting that plague observational discrete choice.
To see what this buys us, consider a simplified version of the candidate-choice experiment with two binary attributes (candidate gender and party) and a binary respondent characteristic (college degree). Randomization means that within each education group, differences in choice probabilities across profile contrasts identify the preference coefficients directly: a comparison of female versus male candidates (holding party fixed) isolates $\beta_{\text{Female}}(z)$, and a comparison of Democrat versus Republican candidates (holding gender fixed) isolates $\beta_{\text{Party}}(z)$. Once these two coordinates are identified, the structural logit model implies the probability of choosing a female Democrat over a male Republican as $G(\beta_{\text{Female}}(z) + \beta_{\text{Party}}(z))$---the logit model links the isolated effects into a joint preference vector. Randomization identifies each attribute's marginal utility, and the structural model combines them.
We emphasize that identification here has two components: experimental randomization nonparametrically identifies the conditional choice probability $\mathrm{Pr}(Y_{it}=1 \mid \boldsymbol{\Delta}\mathbf{X}_{it}, \mathbf{Z}_i)$, eliminating the endogeneity concerns that dominate observational discrete choice, and the logit functional form inverts these probabilities to recover the preference function $\boldsymbol{\beta}(\mathbf{Z})$. The logit link is therefore the remaining structural assumption.
Two features, however, limit its bite. The DNN is fully flexible in how it maps respondent characteristics to preferences, so the parametric restriction applies only to how preferences map to choice probabilities, not to the heterogeneity itself. In practice, common link functions (logit, probit) are nearly indistinguishable in the interior of the probability space, and the DNN can partially compensate for link misspecification by rescaling the preference vector; the binding structural assumption is additive utility, not the choice of error distribution.
We parameterize $\boldsymbol{\beta}(\mathbf{Z})$ as a deep neural network (DNN) with a structural output layer, following farrell2025deep. The estimation procedure has three stages: (i) train the DNN to learn the preference function, (ii) cross-fit to avoid overfitting bias, and (iii) apply the DML influence-function correction for valid inference.
The architecture consists of a feature network that maps respondent characteristics to preference parameters, and a model layer that embeds these parameters in the structural logit model. The feature network takes $\mathbf{Z} \in \mathbb{R}^{p_Z}$ as input and applies $L$ hidden layers with ReLU activations: \[ \mathbf{h}_\ell = \text{ReLU}(\mathbf{W}_\ell \, \mathbf{h}_{\ell-1} + \mathbf{b}_\ell), \quad \ell = 1, \ldots, L, \] with $\mathbf{h}_0 = \mathbf{Z}$, where $\text{ReLU}(x) = \max(0, x)$ is a standard nonlinear activation function applied elementwise, and $\mathbf{W}_\ell$, $\mathbf{b}_\ell$ are weight matrices and bias vectors learned during training. The final hidden layer maps to the $p$-dimensional preference vector: \[ \boldsymbol{\beta}(\mathbf{Z}) = \mathbf{W}_{L+1} \, \mathbf{h}_L + \mathbf{b}_{L+1} \;\in\; \mathbb{R}^p, \] with no activation function on the output, since preference parameters are unrestricted in sign and magnitude. The model layer then computes the logit index $\boldsymbol{\Delta}\mathbf{X}^\top \boldsymbol{\beta}(\mathbf{Z})$, which enters the logistic choice probability. This is the key architectural choice: the model layer enforces the linear-in-parameters utility structure of the logit model, so the DNN learns the preference function $\boldsymbol{\beta}(\mathbf{Z})$ rather than a free-form mapping from $(\boldsymbol{\Delta}\mathbf{X}, \mathbf{Z})$ to choice probabilities. The network is trained by minimizing the binary cross-entropy loss: \[ \mathcal{L} = -\frac{1}{N} \sum_{i=1}^M \sum_{t=1}^{T} \left[ Y_{it} \log G\!\left(\boldsymbol{\Delta}\mathbf{X}_{it}^\top \boldsymbol{\beta}(\mathbf{Z}_i)\right) + (1 - Y_{it}) \log\!\left(1 - G\!\left(\boldsymbol{\Delta}\mathbf{X}_{it}^\top \boldsymbol{\beta}(\mathbf{Z}_i)\right)\right) \right], \] which is exactly the negative log-likelihood of the heterogeneous logit model---the same objective function a researcher would maximize when estimating a standard logit, except that the coefficients now depend on respondent characteristics through the network. This structural loss ensures that the network learns parameters with the economic interpretation of marginal utilities, not merely prediction-optimal features. The estimation pipeline can be summarized schematically as: \[ \mathbf{Z} \;\xrightarrow[\text{(nonparametric)}]{\;\text{hidden layers}\;}\; \hat{\boldsymbol{\beta}}(\mathbf{Z}) \;\xrightarrow[\text{(parametric)}]{\;\text{model layer}\;}\; \boldsymbol{\Delta}\mathbf{X}^\top \hat{\boldsymbol{\beta}}(\mathbf{Z}) \;\xrightarrow[\text{(link function)}]{\;G(\cdot)\;}\; \hat{Y}. \] Naively plugging in $\hat{\boldsymbol{\beta}}(\mathbf{Z}_i)$ to estimate $\theta_k = \mathbb{E}[\beta_k(\mathbf{Z})]$ yields biased estimates due to overfitting. We address this via $K$-fold cross-fitting chernozhukov2018double: each of $M$ respondents is assigned to one of $K$ folds (all tasks from the same respondent go to the same fold), the DNN is trained on $K-1$ folds, and $\hat{\boldsymbol{\beta}}(\mathbf{Z}_i)$ is computed for held-out respondents. After cycling through all folds, every observation has a predicted preference vector estimated without using that observation's data.
Even with cross-fitting, a first-order bias remains from the estimation error of $\hat{\boldsymbol{\beta}}$. The DML framework corrects this via an influence function. For attribute $k$, respondent $i$, and task $t$, define:
where $\hat{G}_{it} = G(\boldsymbol{\Delta}\mathbf{X}_{it}^\top \hat{\boldsymbol{\beta}}(\mathbf{Z}_i))$ and $[\cdot]_k$ denotes the $k$-th element. The first term is the plug-in estimate; the second is the debiasing correction, which projects the logit residual $(Y_{it} - \hat{G}_{it})$ through the treatment $\boldsymbol{\Delta}\mathbf{X}_{it}$ and the inverse local information matrix: \[ \Lambda_{jk}(\mathbf{Z}) = \mathbb{E}\!\left[G'\!\left(\boldsymbol{\Delta}\mathbf{X}^\top \boldsymbol{\beta}(\mathbf{Z})\right) \cdot \Delta X_j \cdot \Delta X_k \;\middle|\; \mathbf{Z}\right], \] where $G'(v) = G(v)(1 - G(v))$ is the logistic density. The matrix $\boldsymbol{\Lambda}(\mathbf{Z})$ measures how informative the typical profile contrast is about the preference parameters for a respondent of type $\mathbf{Z}$: it is large when profile contrasts vary widely and when choice probabilities are near one-half (where the logit model is most informative about the underlying preferences). Intuitively, if the DNN underpredicts the choice probability for certain profile contrasts, the correction adjusts the preference estimate upward for the corresponding attributes. The debiased estimator of the average preference parameter $\theta_k = \mathbb{E}[\beta_k(\mathbf{Z})]$ is then:
The influence function $\psi_{ikt}(\theta, \boldsymbol{\beta})$ satisfies the Neyman orthogonality condition: \[ \frac{\partial}{\partial r} \mathbb{E}\!\left[\psi_{ikt}(\theta_k^0, \boldsymbol{\beta}_0 + r\,\mathbf{h})\right]\bigg|_{r=0} = 0 \quad \text{for all } \mathbf{h}, \] where $\mathbf{h}$ is an arbitrary perturbation direction. This means that first-order errors in $\hat{\boldsymbol{\beta}}$ do not propagate into $\hat{\theta}_k$, permitting $\sqrt{N}$-consistent inference despite the slower convergence rate of the nonparametric first stage.
The variance is estimated with clustering at the respondent level: \[ \widehat{\mathrm{Var}}(\hat{\theta}_k) = \frac{M}{M-1} \cdot \frac{1}{N^2} \sum_{m=1}^{M} \left(\sum_{t: i(t)=m} \psi_{mkt}\right)^2. \] Clustering at the respondent level is essential: a single respondent contributes multiple observations (one per task), and these observations are correlated through the shared preference vector $\boldsymbol{\beta}(\mathbf{Z}_i)$; ignoring this correlation would understate the standard errors (SEs), as we confirm empirically in each application below. Under regularity conditions given in farrell2025deep and chernozhukov2018double---which require the DNN to converge to the true $\boldsymbol{\beta}(\mathbf{Z})$ at a rate faster than $N^{-1/4}$---the DML estimator is $\sqrt{N}$-consistent and asymptotically normal: \[ \frac{\hat{\theta}_k - \theta_k}{\sqrt{\widehat{\mathrm{Var}}(\hat{\theta}_k)}} \;\xrightarrow{\;d\;}\; N(0, 1). \] As a basic consistency check, Supplementary Materials (ref) shows that the DNN's average $\hat{\theta}_k$ reproduces the coefficient of a pooled homogeneous logit on the same data almost exactly ($r = 0.998$ across 28 attribute levels), and that its within-subgroup conditional means reproduce subgroup-specific logit coefficients with comparable accuracy ($r = 0.955$ across 15 country subgroups $\times$ 28 attributes), so any finding a reduced-form AMCE analysis would deliver is recoverable as a special case of the structural estimator. We further validate the finite-sample performance of the estimator---including coverage, individual-level preference recovery, and comparison to alternative methods---via Monte Carlo simulations in Supplementary Materials (ref). Table (ref) summarizes the methodological landscape.
We illustrate the structural DNN estimator with three published conjoint experiments that highlight distinct strengths of the approach. The first application recovers the distribution of preferences across respondents---including preference polarization and compensating differentials---in a candidate-choice experiment. The second reframes a prominent finding about democratic accountability by recovering the direction--intensity decomposition and the marginal rate of substitution between democracy and partisanship. The third demonstrates willingness to pay as a structural quantity, enabled by a conjoint design with a monetary attribute.
saha2022welfare study how voters evaluate candidates for public office using a conjoint experiment fielded to 1,191 respondents. Each respondent evaluated three pairs of hypothetical candidates described by five attributes: policy agenda (3 levels), talent (7 levels), number of children (4 levels), gender (2 levels), and prior candidacy (2 levels), yielding $p = 13$ dummy-coded attribute levels. We specify the structural DNN with architecture $\mathbf{Z} \to 32 \to 32 \to 16 \to \boldsymbol{\beta} \in \mathbb{R}^{13}$, using $K = 10$ cross-fitting folds and 2,000 training epochs. The covariate vector $\mathbf{Z}_i$ includes 15 respondent characteristics (party identification, ideology, gender, age, education, income, race, and region).
Panel A of Figure (ref) reports the DML-corrected average preference estimates $\hat{\theta}_k$ with 95% confidence intervals on the logit scale. The dominant attributes are policy agenda: both Moderate Changes ($\hat{\theta} = 0.81$, $p < 0.001$) and Complete Overhaul ($\hat{\theta} = 0.81$, $p < 0.001$) are strongly preferred over the reference category (Very Few Changes). Among talent attributes, only Hard-Working is statistically significant ($\hat{\theta} = 0.45$, $p < 0.001$). The gender coefficient---Male relative to Female---is negative but not significant ($\hat{\theta} = -0.11$, 95% CI $[-0.24, 0.01]$), consistent with the original AMCE finding. Clustered standard errors exceed unclustered standard errors by a factor of 1.78, confirming the importance of respondent-level clustering.
Panel B converts these estimates to the probability scale---average marginal effects---and overlays the AMCE from a standard linear probability model. The two sets of estimates are nearly indistinguishable, confirming that the structural model's population average $\hat{\theta}_k = n^{-1}\sum_i \hat{\beta}_k(\mathbf{Z}_i)$ reproduces the standard reduced-form AMCE as a special case. This equivalence is a feature of the method: any finding that a conventional AMCE analysis would deliver remains fully recoverable, while the structural model additionally provides the individual-level preference vectors $\hat{\beta}_k(\mathbf{Z}_i)$ that underlie these averages.
Panel C reveals what lies behind the averages. Because the structural model recovers $\hat{\beta}_k(\mathbf{Z}_i)$ for each respondent, we can compute the fraction with $\hat{\beta}_k(\mathbf{Z}_i) > 0$ (favoring the attribute level) versus $< 0$ (opposing it). Three patterns are striking.
First, despite the near-zero average effect for gender, 83% of respondents have $\hat{\beta}_{\text{Male}} < 0$, meaning they prefer female candidates all else equal. This is not consensus indifference---it is near-consensus in one direction.
Second, Empathetic has a near-zero average preference parameter ($\hat{\theta} = -0.02$), but the structural model reveals an almost perfectly polarized electorate: 49% favor empathetic candidates while 51% oppose them. This polarization is sharply partisan: Democratic respondents have a mean $\hat{\beta}_{\text{Empathetic}}$ of $+0.42$, while Republicans average $-0.28$. The near-zero average is not indifference---it is a cancellation of opposing preferences across party lines.
Third, some attributes show genuine consensus. Hard-Working is unanimously favored (100% positive), and both Moderate Changes and Complete Overhaul are unanimously favored. For these consensus attributes, the average preference parameter and the polarization measure tell the same story; the structural model's additional power becomes evident precisely when preferences are heterogeneous.
Figure (ref) shows the full density of $\hat{\beta}_k(\mathbf{Z}_i)$ across respondents for each attribute level, ordered by variance. Complete Overhaul is the most heterogeneous: while nearly all respondents favor it, the intensity ranges from near-zero to over $1.5$ in logit units. The Empathetic distribution straddles zero, while Hard-Working is tightly concentrated in positive territory. These densities illustrate the direction-versus-intensity distinction: attributes can agree in direction yet differ sharply in intensity (Hard-Working), or can split in both dimensions simultaneously (Empathetic).
Figure (ref) drills further into the two most striking attributes by cutting the individual-level preference parameters along two respondent dimensions simultaneously---gender and party. The top row shows the preference for a female over a male candidate. The positive skew is nearly universal across all six subgroups, consistent with the 83% figure in the right panel of Figure (ref), but the gender gap is visible throughout: female respondents reward female candidates roughly twice as heavily as male respondents within each party ($+0.19$ vs.\ $+0.07$ among Democrats, $+0.18$ vs.\ $+0.08$ among Independents, $+0.12$ vs.\ $+0.03$ among Republicans). Party differences are small by comparison. The bottom row shows the preference for empathetic candidates, and the picture is qualitatively different. Here the party cleavage is the dominant axis: Democrats and Independents are centered in positive territory while Republicans are centered in negative territory, with female Republicans essentially indifferent ($-0.01$) and male Republicans actively penalizing empathy ($-0.08$). The fraction of respondents with $\hat{\beta}_{\text{Empathetic}} > 0$ falls from 56% among Democrats to 38% among Republicans.
The sign and spread of $\hat{\beta}_k(\mathbf{Z}_i)$ tell us which way voters lean on each attribute, but not how much each attribute actually drives their vote. Two attributes can generate equally large coefficients yet contribute very differently to the decision once we account for how far the attribute levels are spread apart in the design. We quantify this with an individual-level importance share, defined for respondent $i$ and attribute group $g$ as $\text{Imp}_{i,g} = \sum_{k \in g} \hat{\beta}_{i,k}^2 \cdot \text{Var}(X_k) \,/\, \sum_{g'} \sum_{k \in g'} \hat{\beta}_{i,k}^2 \cdot \text{Var}(X_k)$, the share of the variance in the respondent's utility that each attribute accounts for. Aggregated at the respondent level, these shares map directly onto the question a campaign strategist cares about: given the range of candidate profiles a voter might actually see, which attributes are moving their decision? Figure (ref) shows the distribution of these shares across respondents. Policy agenda dominates decisively (mean 66%), consuming two-thirds of the decision variance for the median voter---substantially more than its $\hat{\theta}$ alone would suggest, because the agenda attribute has three highly differentiated levels whose effects compound. Talent and children each average around 15%, and gender and prior candidacy contribute less than 3% on average. The spread of the agenda share is wide: for some respondents it accounts for over 80% of utility variance, while for others it drops below 40% as talent or children attributes carry more weight. The substantive point is that Saha and Weeks' near-zero average for gender is not just a canceled preference---it is an attribute that the electorate, with only a handful of exceptions, treats as a minor tiebreaker behind policy and competence. The structural model lets us see that the 83% female-preference finding of Figure (ref) and the low importance share in Figure (ref) are both true simultaneously: voters do prefer female candidates, but for most of them this preference is worth only a small fraction of the weight they place on the candidate's policy agenda.
\FloatBarrier
graham2020democracy study how American voters trade off democratic principles against policy and partisan considerations. Their conjoint experiment presents 1,605 respondents with pairs of hypothetical state legislative candidates described by eight attribute groups: party (co-partisan vs.\ not), two policy positions (economic and social), six good-governance practices, seven undemocratic actions, two valence violations, sex, race, and profession---yielding $p = 30$ attribute levels. Each respondent evaluated approximately 13 choice tasks (20,657 total). We specify the DNN with architecture $\mathbf{Z} \to 64 \to 64 \to 32 \to \boldsymbol{\beta} \in \mathbb{R}^{30}$, $K = 10$ folds, and 2,000 epochs. The covariate vector $\mathbf{Z}_i$ includes 12 variables: ideology (7-point), party ID (7-point), Trump approval, age, education, household income, authoritarianism, political knowledge, gender, and race indicators.
Co-partisanship has the largest positive effect ($\hat{\theta} = 0.66$, $p < 0.001$). All seven undemocratic actions carry significant negative coefficients, ranging from $-0.46$ (executive orders) to $-0.80$ (prosecuting journalists). DML-clustered standard errors exceed unclustered standard errors by a factor of 3.72, reflecting the 13 tasks per respondent.
The central claim of graham2020democracy's paper is that only 11.7% of Americans prioritize democratic principles in their vote choice. Our structural model reframes this finding. Figure (ref) shows that for every undemocratic action, fewer than 5% of respondents in this sample have $\hat{\beta}_k(\mathbf{Z}_i) > 0$. For four of seven actions---ignoring courts, gerrymandering (10 seats), prosecuting journalists, and closing polling stations---the fraction is exactly 0%. There is near-universal rejection of undemocratic behavior in this sample.
The apparent contradiction dissolves once we distinguish between direction and magnitude. graham2020democracy's 11.7% counts respondents who chose the more democratic candidate regardless of all other attributes---a measure that conflates the strength of democratic preferences with the strength of competing considerations. Our structural model shows that, in this sample, virtually all voters dislike undemocratic behavior (direction), but many voters weight party and policy considerations more heavily (magnitude). The disagreement across the electorate is not about whether democracy matters---it is about how much.
Figure (ref) displays the full density of $\hat{\beta}_k(\mathbf{Z}_i)$ for each attribute level. The undemocratic actions are entirely concentrated in negative territory, but the spread varies considerably, indicating that while virtually everyone opposes these actions, the intensity of opposition varies substantially.
We define democracy sensitivity for respondent $i$ as the mean absolute value of their undemocratic-action coefficients, $\text{DemSens}_i = (1/7)\sum_{k \in \mathcal{U}} |\hat{\beta}_k(\mathbf{Z}_i)|$. Figure (ref) plots average democracy sensitivity across the 7-point ideology scale, revealing a striking U-shape. Liberals are most sensitive (0.78), sensitivity drops to its minimum for moderates (0.53), and rises again for extreme conservatives (0.60). This contradicts a simple partisan narrative: the “democracy deficit” is concentrated in the disengaged middle, not among strong partisans of either stripe.
The marginal rate of substitution between undemocratic actions and co-partisanship quantifies how much partisan benefit a voter must sacrifice to avoid a democratic violation. Table (ref) reports the MRS for three undemocratic actions. On average, voters would sacrifice 93% of the partisan benefit to avoid journalist prosecution, 76% to avoid ignoring courts, and 66% to avoid a 10-seat gerrymander. These ratios vary sharply by ideology: liberals sacrifice more than one full unit of partisan benefit for journalist prosecution ($\text{MRS} = -1.05$), while conservatives sacrifice only 70%.
The MRS measures how much partisan benefit a voter would have to sacrifice on the margin to avoid a democratic violation. A complementary and more discrete question---closer in spirit to the headline defection rates in graham2020democracy's original analysis---asks whether a specific compensator is sufficient to induce a voter to accept a given undemocratic action. For each voter $i$ and undemocratic action $j$, we compute the compensating-differentials indicator $\mathbf{1}\{\hat{\beta}_{u_j}(\mathbf{Z}_i) + \text{benefit}_k(\mathbf{Z}_i) \geq 0\}$ under three alternative compensators: (i) co-partisanship (benefit $= +\hat{\beta}_{\text{party}}(\mathbf{Z}_i)$), (ii) a full-range policy swing on both economic and social policy ($3(|\hat{\beta}_{p_1}(\mathbf{Z}_i)| + |\hat{\beta}_{p_2}(\mathbf{Z}_i)|)$), and (iii) the voter's single favorite good-governance feature ($\max_k \hat{\beta}_{g_k}(\mathbf{Z}_i)$). The quantity of interest is the fraction of respondents for whom compensation holds, reported in Figure (ref) by ideology tercile.
Three features stand out. First, the “None” column confirms near-universal resistance to democratic erosion in this sample: across all 7 actions and all ideology groups, fewer than 5% of respondents have a positive coefficient on any undemocratic action. Second, co-partisanship alone is a remarkably powerful compensator: for 5 of 7 actions, a majority of respondents in every ideology tercile would accept the violation in exchange for voting for their own party. Third, and most strikingly, the co-partisanship column reveals a sharp ideological asymmetry on the hardest cases. For prosecuting journalists, only 40% of liberals would accept the violation in exchange for a co-partisan, compared to 88% of conservatives---a 48 percentage-point gap. For closing polling stations the gap is 22 points (71% vs.\ 93%); for ignoring courts it is 26 points (67% vs.\ 93%). These three actions---the ones most closely tied to free elections, a free press, and judicial independence---are also where liberals are most willing to cross party lines to defend democracy. The full-range policy-swing column shows that a sufficiently large policy concession can compensate essentially everyone for essentially any violation, which re-emphasizes why policy alignment dominates democratic principles in observational voting patterns: the policy stakes are simply very large.
\FloatBarrier
ballard2017structure study American mass preferences over federal income-tax policy, asking 2{,}000 US adults to choose between pairs of hypothetical tax plans in which each plan specifies the marginal rate on six income brackets together with a revenue indicator describing how much revenue the plan raises relative to current policy. Their central reduced-form finding is that Americans have generally progressive preferences---opposing higher rates on the poor and supporting higher rates on the rich---but that support for a plan responds much more elastically to rates on the poor than to rates on the rich. The analysis sample contains 2,000 US respondents, each evaluating 8 paired comparisons (16,000 tasks). The seven continuous attributes produce $p = 7$ treatment variables: marginal rates on six income brackets (entering in percentage points) and a five-level revenue indicator rescaled to $\{-2,-1,0,1,2\}$. Every attribute is continuous, so the DNN recovers an individual-level slope---utility per percentage point of rate---for every bracket and every respondent, producing a full preferred tax schedule $\hat{\boldsymbol{\beta}}(\mathbf{Z}_i) \in \mathbb{R}^7$ per person rather than a handful of average level effects. We specify the DNN with architecture $\mathbf{Z} \to 32 \to 32 \to 16 \to \boldsymbol{\beta} \in \mathbb{R}^7$, $K = 10$ folds, and 2,000 epochs. The covariate vector $\mathbf{Z}_i$ includes 12 respondent characteristics: age, gender, party ID, education, race, income, inequality aversion, and beliefs about the tax-economy relationship. Because each respondent provides 8 tasks, within-respondent correlation is substantial: the DML/independent and identically distributed SE ratio is 3.00---the highest of any application we run---so naive independent and identically distributed inference would overstate precision by a factor of three.
Figure (ref) reports the DML-corrected average effects. A 1 percentage-point increase in the bottom-bracket rate lowers log-odds of plan support by 0.043; the same increase on the top bracket raises log-odds by only 0.011. This fourfold absolute asymmetry is the quantitative form of the central ballard2017structure finding: plan support is highly elastic with respect to rates on low incomes but nearly inelastic with respect to rates on high incomes. The \$85--175k bracket is indistinguishable from zero on average---the “dead zone”---and revenue has a small positive but non-significant average effect.
Because each respondent has a full 7-element vector $\hat{\boldsymbol{\beta}}(\mathbf{Z}_i)$, we can compute an individual-level progressivity slope---the within-respondent regression of bracket-specific marginal utilities on the log of the bracket midpoint: \[ s_i \;=\; \frac{\sum_{k=1}^{6}(\log m_k - \overline{\log m}) \cdot \hat{\beta}_{i,k}}{\sum_{k=1}^{6}(\log m_k - \overline{\log m})^2}. \] A positive $s_i$ means respondent $i$ prefers higher rates on higher incomes; zero means a flat tax; negative means regressive. Remarkably, in this sample 99.9% of respondents have a positive slope, and $\hat{\beta}_{i,\text{top}} > \hat{\beta}_{i,\text{bottom}}$ for 99.1%. The progressive direction is essentially universal here: in every partisan, income, ideological, and inequality-aversion subgroup of this sample, fewer than 0.4% of respondents have regressive revealed preferences. Democrats' mean slope is $+0.0131$, Republicans' is $+0.0100$---a real but modest 31% gap---and 99.7% of Republican respondents still reveal progressive preferences. This is the sharpest possible individual-level form of the ballard2017structure finding that progressive preferences cut across partisan lines.
Figure (ref) exposes the within-party heterogeneity that the structural model recovers. The upper panel overlays each party's median schedule with its inter-quartile ($25$--$75$) and $10$--$90$ percentile bands. Three facts are immediately visible: the party-level medians are essentially indistinguishable at the bottom brackets but fan out at the top, with Democrats above Republicans on the top two brackets and below on the bottom-middle brackets; the within-party $25$--$75$ bands overlap substantially across parties at every bracket, so within-party heterogeneity dwarfs the between-party gap; and the $10$--$90$ band is wide and straddles zero in every middle bracket, meaning each party contains both bracket-level “hawks” and “doves.” The lower panel shows a random sample of 200 individual respondent schedules per party as semi-transparent lines, with the party mean overlaid. Almost every individual line in this sample slopes upward, so the progressive direction is near-universal here, but the levels vary enormously---some Republicans want a schedule steeper than the Democratic median, and some Democrats want schedules nearly as flat as the Republican median. The between-party difference is a shift in central tendency, not a separation of populations.
The structural model also enables a variance decomposition by subgroup that is invisible to reduced-form subgroup AMCEs. Figure (ref) reports the share of plan-choice variance attributable to each attribute, by party. Democrats allocate 34.6% of variance to the top bracket; Republicans only 12.5%---a ratio of nearly three-to-one. Republicans, in turn, weight the \$10--35k and \$35--85k brackets substantially more than Democrats do ($26.3\%$ and $16.3\%$ vs.\ $14.4\%$ and $5.2\%$). Democrats' plan-choice variance is driven 35% by what happens to the rich; Republicans' variance is driven 43% by what happens to working-class and middle-class incomes. This is a substantive reframing of the standard partisan narrative: Republican opposition to progressive taxation is primarily opposition to taxing the working class, not primarily protection of the rich.
Alongside the conjoint, ballard2017structure collected each respondent's self-reported ideal marginal tax rate for each of the six brackets. These ideal rates were elicited outside the conjoint task and are not used by the DNN at any stage of training, so they provide an independent benchmark against which the structural model's recovered individual preferences can be validated. No reduced-form AMCE estimator can be validated at the individual level, because it does not produce an individual-level preference parameter. Using the 401 respondents who answered both the bottom- and top-bracket ideal questions, the correlation between the revealed progressivity slope $s_i$ and its self-reported analog is $r = 0.37$, and the correlation between the revealed and self-reported top-minus-bottom rate gap is $r = 0.43$. Using all $n = 2{,}000$ respondents, the DNN's individual-level top-bracket coefficient correlates $r = 0.37$ with the self-reported ideal top rate. These correlations survive disaggregation by party ($r \approx 0.26$--$0.30$ within each partisan subgroup), confirming that the DNN recovers the correct individual ordering of tax progressivity within, not just across, parties. A correlation of $r = 0.43$ between a DNN-recovered preference parameter and an independent self-report is substantively strong---comparable to typical cross-measure correlations in political psychology and well above what one would expect from noise.
The individual-level preference vector $\hat{\boldsymbol{\beta}}(\mathbf{Z}_i)$ also enables evaluation of any hypothetical tax schedule. We compare four stylized plans: a revenue-neutral flat 15% tax, a steeply progressive plan with rates rising $0/5/15/25/35/45$ and “much more” revenue, a regressive plan with the reverse rate pattern and “much less” revenue, and a revenue-neutral status-quo analog with rates $5/15/25/25/35/35$. The progressive plan yields the only positive mean systematic utility in the set. Head-to-head probabilistic comparisons are striking: the progressive plan beats the flat plan for $78.7\%$ of respondents overall and for $73.1\%$ of Republicans, while it beats the status-quo plan for roughly $70\%$ of respondents. Americans prefer the status quo to a flat or regressive plan, but they prefer an explicitly more progressive alternative to the status quo even more strongly. The ballard2017structure conclusion that Americans' preferences lie “close to current policy” is consistent with the status-quo plan ranking second-best, but it understates how much the public would prefer a more aggressively progressive schedule if offered one. A purely partisan model of tax preferences---“Republicans want flat, Democrats want progressive”---is not consistent with the data: nearly three-quarters of Republicans prefer a steeply progressive schedule with more revenue to a flat, revenue-neutral tax.
Our method rests on several assumptions, which we now discuss alongside a set of possible extensions that addresses these limitations.
The first consideration is what the method estimates when preference heterogeneity is not fully captured by observed respondent characteristics. If the true process is $\beta_i = \boldsymbol{\beta}(\mathbf{Z}_i) + \boldsymbol{\eta}_i$ with a respondent-specific latent component, the method targets $\mathbb{E}[\beta_i \mid \mathbf{Z}_i]$---the conditional mean preference given observables. This is a well-defined estimand, not a source of bias: omitting covariates from $\mathbf{Z}_i$ does not bias the average preference parameters $\theta_k = \mathbb{E}[\beta_k(\mathbf{Z}_i)]$, and the individual-level $\hat{\boldsymbol{\beta}}(\mathbf{Z}_i)$ consistently estimates the conditional expectation $\mathbb{E}[\boldsymbol{\beta}_i \mid \mathbf{Z}_i]$, which is the best prediction of a respondent's preferences given what is observed about them. What changes is the interpretation: $\hat{\boldsymbol{\beta}}(\mathbf{Z}_i)$ recovers the average preference of respondents who share $i$'s observable profile, rather than $i$'s exact personal preference. Richer covariates narrow this gap, which is why we recommend collecting a rich set of $\mathbf{Z}_i$ variables.
The practical consequence is that nonlinear functionals of the conditional mean can differ from the conditional mean of those functionals: the distribution of $\hat{\boldsymbol{\beta}}(\mathbf{Z}_i)$ across respondents understates the true dispersion of $\beta_i$ (because it misses the within-$\mathbf{Z}$ component $\boldsymbol{\eta}_i$), and the MRS computed from conditional means need not equal the conditional mean of the MRS, since $\mathbb{E}[\beta_1 \mid \mathbf{Z}] / \mathbb{E}[\beta_2 \mid \mathbf{Z}] \neq \mathbb{E}[\beta_1 / \beta_2 \mid \mathbf{Z}]$ by Jensen's inequality. These smoothed quantities remain valid and interpretable as properties of $\mathbb{E}[\boldsymbol{\beta} \mid \mathbf{Z}]$, but researchers should bear in mind that they provide a lower bound on the true degree of preference heterogeneity. A natural extension is a hybrid model that combines the DNN mean function with latent random coefficients $\boldsymbol{\eta}_i \sim N(\mathbf{0}, \boldsymbol{\Sigma}(\mathbf{Z}_i))$, using within-respondent dependence across repeated tasks to identify the latent covariance structure.
A second key assumption is an additive random utility model with a linear-in-parameters utility index. If respondents care about attribute interactions not included in $\mathbf{X}_{ijt}$, use noncompensatory decision rules, or evaluate profiles relative to task-specific reference points, the model is misspecified. Interactions can be addressed by including interaction terms in $\mathbf{X}_{ijt}$, but noncompensatory rules are fundamentally inconsistent with the linear utility framework and would require a different modeling strategy. Relatedly, the Gumbel distribution of taste shocks fixes the error scale, so preferences are identified only up to this normalization: if the true utility includes respondent-specific scale $\sigma_i$, the data identify $\boldsymbol{\beta}_i / \sigma_i$ rather than $\boldsymbol{\beta}_i$, and what appears as taste heterogeneity may partly reflect scale heterogeneity. Modeling scale heterogeneity jointly with preference heterogeneity is a natural and straightforward extension of the current framework.
As the utility index is linear in $\mathbf{X}$, it rules out complementarities between attributes unless interaction terms are explicitly included. Manual inclusion of pairwise interactions is straightforward but quickly becomes high-dimensional ($p(p+1)/2$ terms for $p$ attributes). A more systematic extension is a low-rank interaction layer in the DNN architecture, in which the network outputs both the main-effect vector $\boldsymbol{\beta}(\mathbf{Z}) \in \mathbb{R}^p$ and a low-rank factor $\mathbf{V}(\mathbf{Z}) \in \mathbb{R}^{p \times r}$ with $r \ll p$, so that the utility index becomes $\boldsymbol{\Delta}\mathbf{X}^\top \boldsymbol{\beta}(\mathbf{Z}) + \|\mathbf{V}(\mathbf{Z})^\top \boldsymbol{\Delta}\mathbf{X}\|^2$. The rank constraint regularizes the interaction space while preserving structural interpretability; $r$ can be chosen by cross-validation. More broadly, the structural AME provides a built-in diagnostic for these functional-form assumptions: as discussed in Section (ref), a discrepancy between the structural AME and the nonparametric AMCE would signal misspecification of either the logit link or the additive utility structure.
A third set of caveats concerns inference. Formal DML guarantees apply to the average parameters $\theta_k$, but many of the structural quantities reported in the applications---including the MRS, WTP, and the AME---are smooth functionals of these parameters, for which valid inference extends naturally via the delta method applied to the DML influence functions, or equivalently, via Neyman orthogonal moment conditions for the functionals themselves. What remains genuinely open is inference on distributional quantities that depend on the full shape of $\boldsymbol{\beta}(\mathbf{Z})$ rather than its mean---polarization fractions, importance-share distributions, and individual-level preference rankings, for instance. Our simulations in Supplementary Materials (ref) show that these quantities are recovered with high accuracy when the sample is large enough, and Supplementary Materials (ref) provides design guidance on the sample sizes required. Developing formal inferential tools for these distributional quantities is an important direction for future work.
This paper develops a structural estimator for conjoint experiments that combines a random utility logit model with a deep neural network mean function, yielding an individual-level preference vector $\hat{\boldsymbol{\beta}}(\mathbf{Z}_i)$ for every respondent together with valid double/debiased machine learning inference on population averages. The approach nests the reduced-form AMCE as a special case---our average $\hat{\theta}_k$ reproduces a pooled homogeneous-logit coefficient almost exactly both overall and within subgroups (Supplementary Materials (ref))---so nothing a standard AMCE analysis would deliver is lost. What is gained is access to the full distribution of preferences and to structural quantities---polarization, marginal rates of substitution, willingness to pay, compensating differentials, counterfactual welfare---that an average effect cannot express.
The three applications show what this buys. In the candidate-choice experiment, a near-zero average preference for gender coexists with the vast majority of respondents preferring female candidates, and a near-zero average preference for empathetic candidates dissolves into a sharp partisan split; the structural model recovers the direction--intensity decomposition that are obscured by reduced-form averages. In the democracy-versus-partisanship tradeoff, virtually every respondent has a negative individual-level coefficient on every undemocratic action, resolving the apparent paradox of the paper's headline finding: opposition to democratic erosion is near-universal in direction, but many voters weight partisan and policy considerations more heavily in magnitude, and liberals are substantially more willing than conservatives to cross party lines on the hardest cases. In the tax-plan experiment, we recover a full preferred tax schedule for each respondent and find that nearly all respondents have progressive revealed preferences, including nearly all Republicans; the recovered progressivity slope correlates meaningfully with independently elicited self-reported ideal rates, providing external validation at the individual level.
Structural heterogeneity is a complement to reduced-form AMCE analysis rather than a replacement. When the research question is about the average effect of an attribute, the AMCE is appropriate and remains the natural quantity to report. But when the research question is about polarization, cross-attribute tradeoffs, individual welfare, or counterfactual choice---questions about whether preferences are shared, about how much one attribute is worth relative to another, or about what a population would choose under a policy that was never part of the experimental design---an individual-level preference vector is required, and the structural DNN approach developed here makes this practical at the sample sizes that are typical of applied conjoint research.
For researchers planning to use this approach, our simulation results (Supplementary Materials (ref) and (ref)) suggest several practical guidelines. First, collect a rich set of respondent covariates $\mathbf{Z}_i$. The DNN learns heterogeneity entirely through observed characteristics; richer covariates mean finer-grained preference recovery. Demographics, ideology, party identification, attitudes, and behavioral measures all contribute. Second, ensure a sufficient total sample size. Population-level quantities (average preferences, polarization, importance shares) are reliably recovered at $NT \approx 3{,}000$--$5{,}000$ total observations, while accurate individual-level preference vectors require $NT \approx 8{,}000$--$10{,}000$. The total number of observations $NT$ is the key design parameter: our factorial simulations show that neither the number of respondents $N$ nor the number of tasks $T$ alone is a sufficient statistic, though $N$ contributes more than $T$ (53% vs.\ 30% of variance in recovery accuracy). Third, include enough tasks per respondent. The marginal gain from additional tasks is largest at the bottom: going from one to two tasks per respondent yields the single largest improvement in individual-level recovery. We recommend at least three to five tasks per respondent, consistent with standard conjoint practice. Fourth, the method scales straightforwardly with the number of attributes $p$: all three applications in this paper have different attribute counts ($p = 7$, 13, and 30), and the estimator performs well across this range without manual tuning beyond the network architecture.
The method is implemented in the R package, sconjoint, which provides functions for estimation, inference, and visualization.\footnote{A user tutorial is available at \url{https://yiqingxu.org/packages/sconjoint/}. The package was developed using \href{https://statsclaw.ai/}{StasClaw}, an AI-collaborative workflow described in qinxu2026statsclaw.}