Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
175,424 characters · 75 sections · 0 citation commands
Sensitivity to Subjective Expected Utility Maximization: A Methodological Study, with an Illustrative Application to LLM Decision-Making
{ \hypersetup{linkcolor=} \setcounter{tocdepth}{3} }
\addcontentsline{toc}{section}{Motivation}
\addcontentsline{toc}{subsection}{1.1 Two approaches to assessing a decision maker}
There are two broad approaches to assessing a decision maker from a body of choices. The first is external, or label-based: do the agent's choices agree with a labeled set of decisions held to be correct --- a key that records, for each decision problem, which alternative an informed judge would select; would such a judge have forwarded the same claim, taken the same gamble, made the same call? The criterion is agreement with a curated answer key, whoever supplies it. The second is internal: are the choices coherent with one another in the way a stated standard of consistency requires --- responsive to the probabilities, the stakes, and the tradeoffs in the way the standard prescribes? We assess this second kind of conformity from the agent's choices as observed, without attributing to the agent any particular internal computation or deliberative procedure: “internal” qualifies the standard --- coherence among the agent's own choices --- not the agent's cognition. The standard constrains behavior, and the question is whether the behavior conforms to it.
The two assessments answer to different things and can come apart. The consistency standard constrains how an agent's choices hang together, not which alternative an external key marks as correct: an agent whose beliefs and values differ from those the key encodes can be fully consistent yet disagree with it, and an agent can match the key on a given decision while its choices violate the standard elsewhere. The internal-consistency assessment is the one this paper develops, and it is especially relevant where a correct-choice key is unavailable or costly to construct --- in assessing a human analyst, an institution, or, increasingly, a large language model deployed to make or recommend choices.
Both approaches have a long precedent in the study of decision makers. One strand of internal-consistency critique assesses choices against expected utility theory as a standard for decision making under risk: the Allais paradox (Allais 1953) exhibits a pattern of preferences standardly read as violating the independence axiom (roughly, that a component common to two options should not affect the preference between them),\footnote{The diagnosis is analysis-relative rather than forced. Levi (1986) accounts for the same pattern within a generalization of subjective expected utility theory that admits indeterminacy in both probabilities and utilities: relaxing the assumption that any two alternatives are comparable in terms of preference accommodates the Allais choices with no violation of independence. Unlike the descriptive theories built on the pattern, Levi's theory is not reducible to a binary preference relation; we return to it in §\hyperref[sec-seu-standard]{1.4}.} a tension Kahneman and Tversky took up in developing prospect theory (Kahneman and Tversky 1979; Tversky and Kahneman 1992). These demonstrations, and the large-scale replication that subsequently confirmed their patterns (Ruggeri et al. 2020), were built on pairwise choices. Such choices reveal a pattern inconsistent with subjective expected utility maximization, but they do not by themselves fix which assumption of the standard is responsible: the same Allais pattern reads as a violation of the independence axiom or, on Levi's analysis, as a relaxation of the ordering assumption with no independence violation (see the footnote above and §\hyperref[sec-seu-standard]{1.4}). The instrument developed here, by contrast, places no restriction on the number of alternatives in a decision problem, which opens the possibility of probing how sensitivity itself varies with menu size (an extension we return to in §\hyperref[sec-disc-limitations]{8.5}). A second strand --- the heuristics-and-biases program --- assesses judgments about uncertainty against the probability axioms, construed as a coherence standard on degrees of belief in the Dutch-book tradition (Ramsey 1926; Finetti 1937), on which a set of degrees of belief is faulted if it would sanction a combination of bets guaranteeing a sure loss: the conjunction fallacy of the “Linda problem” (Tversky and Kahneman 1983) exhibits credences that violate that calculus, evidence used to motivate the representativeness heuristic. In each case the decision maker is faulted not against a labeled key but against a coherence requirement that its own choices or judgments fail to meet. The contrasting tradition scores choices against an externally curated set of good options: the choice architecture of Thaler and Sunstein (2008), for instance, assesses diners against a menu whose “healthy” items have been labeled by dieticians. The instrument developed here belongs to the internal-consistency tradition --- it measures conformity to an internal standard --- with subjective expected utility maximization, which couples a coherence requirement on degrees of belief with one on preferences, in the role the axioms of rational choice play in those precedents.
\addcontentsline{toc}{subsection}{1.2 Challenges for label-based assessment of decision quality}
The label-based approach scores choices against a key, drawn either from an adjudicated set of correct choices or from the realized outcomes of the decisions. Two difficulties limit what such a key can tell us about the quality of the decisions: a key of correct choices is often impractical to obtain, and a key of outcomes measures the wrong thing.
The first concerns obtaining the labels at all. A key of correct choices records, decision by decision, which alternative is the right one to select, and assembling one is often impractical. The insurance-claims study we report in §\hyperref[sec-application]{7} is a case in point: there is no off-the-shelf key that says which choice was correct, and constructing one would mean asking subject-matter experts what they would choose across the decision problems --- an exercise that is costly and, because expert judgment varies, sensitive to which experts are consulted. We could not produce such a key credibly under realistic experimental conditions.
The second difficulty is more basic, and remains even where a key is at hand --- in particular where it records realized outcomes rather than adjudicated choices. A good outcome is not the same as a good decision: the outcome of a choice typically depends on how the relevant uncertainty happens to resolve, which lies outside the decision maker's control. A score computed purely from outcomes also ignores the way the context in which a decision is made informs the beliefs and desires of the decision maker. Treating outcomes as a direct measure of decision quality is defensible only under strong assumptions --- for instance, that the decision maker is properly incentivized and has enough information to learn the “objective” probabilities of the relevant states. Those assumptions are reasonable in some controlled settings, but we would not want to impose them in a study like the insurance experiment of §\hyperref[sec-application]{7}, where neither condition can be taken for granted. A decision that is well structured given the agent's information may still be scored poorly by an outcome key that was generated under different information, and vice versa.
\addcontentsline{toc}{subsection}{1.3 The need for consistency-based evaluation}
What remains available when neither kind of key can be had is the structure of the choices themselves. If we are willing to state a standard of good decision making explicitly, we can ask whether an agent's choices are responsive to probabilities, utilities, and tradeoffs in the way that standard prescribes --- without a labeled set of correct choices, and without reference to which outcome happened to be realized. Evaluation of this internal-consistency kind replaces a claim about outcomes with a claim about conformity to a stated standard. It does not require any labeled set of decisions; it requires that the standard be named explicitly and --- as the rest of the paper makes concrete --- that a statistical model operationalizing conformity to it be specified, since any graded measure of conformity is read off a fitted model and is therefore conditional on that model's assumptions, prior, and design (§\hyperref[sec-disc-alpha]{8.4}). In return it yields an evaluation that is well defined even where no key of correct choices and no outcome record exist, and that does not depend on which outcome happened to be realized.
\addcontentsline{toc}{subsection}{1.4 SEU as the stated standard}
We take subjective expected utility (SEU) maximization as the reference standard. An SEU-committed agent holds a subjective probability over the states of the world and a utility over consequences, and prefers an alternative whose expected utility is maximal (Savage 1954; Anscombe and Aumann 1963; Neumann and Morgenstern 1947). The standard does not require a unique maximizer: when several alternatives attain the maximal expected utility they are all admissible, and picking among them lies beyond what the standard prescribes (Ullmann-Margalit and Morgenbesser 1977). We do not argue that SEU is the uniquely correct theory of rational choice, and nothing below depends on that stronger claim. The argument is conditional: if SEU is the reference standard, then a graded, statistically rigorous measure of conformity to it is methodologically useful. SEU is the natural first case for such an instrument: it is the most widely shared normative benchmark in decision theory, so a conformity measure built on it reaches the largest audience. The identifiability questions that occupy §§\hyperref[sec-m0-identifiability]{3} and \hyperref[sec-m1-identifiability]{5} are then taken up for that instrument in turn.
Two qualifications bound the sense in which other standards could be substituted for SEU. First, several theories often mentioned alongside SEU are not normative standards of decision quality at all but descriptive accounts of how people actually choose; cumulative prospect theory (Kahneman and Tversky 1979; Tversky and Kahneman 1992) is the clearest example. Substituting such a theory is therefore more than a change of functional form: the quantity being measured changes from conformity to a norm to fit to a behavioral model, and the interpretation of the instrument changes with it. Put plainly: a descriptive substitute would still return a number, but that number would measure fit to observed behavior, not conformity to a standard of good decision making. Maxmin expected utility (Gilboa and Schmeidler 1989) is more plausibly read as a normative alternative, and it supplies an admissible value function for the softmax construction; but because that value function is a minimum over a set of priors, it is not smooth, and substituting it into the gradient-based (Hamiltonian Monte Carlo) implementation of §§\hyperref[sec-m0-implementation]{4} and \hyperref[sec-m1-implementation]{6} is not a drop-in replacement --- the kinks of the min would require either a smoothing that introduces a precision parameter whose finite-sample estimate would be entangled with \(\alpha\) (the two are separately defined, but a smoothed-min sharpness and a choice sensitivity press on overlapping features of the data), or a non-gradient sampler. Smooth ambiguity models (Klibanoff, Marinacci, and Mukerji 2005) avoid the non-smoothness by construction and would substitute more readily. Cumulative prospect theory, by contrast, could not be substituted at all without abandoning the normative reading. Second, the template applies only to standards that rank alternatives by a single real-valued evaluation. Standards that decline to reduce choice to such an evaluation --- for instance the decision theory of Levi (1980, 1986), in which admissibility does not reduce to maximizing a preference ordering: two admissible options may be incomparable rather than equally preferred --- fall outside the framework entirely, not merely outside its current scope. Parallel instruments for the real-valued non-SEU standards are possible by the same template and are out of scope here.
\addcontentsline{toc}{subsection}{1.5 Why a graded measure rather than a binary verdict}
Real decision makers, human or machine, only ever approximately satisfy any rationality axiom (Simon 1955; Lieder and Griffiths 2020). A binary verdict --- “rational” or “not” --- discards exactly the information that an evaluation should preserve, namely how closely and how reliably an agent's choices track the standard. We therefore measure conformity on a continuum. The formal vehicle, developed in §\hyperref[sec-abstract-model]{2}, is a softmax (Luce--McFadden) choice model (Luce 1959; McFadden 1974; Train 2009) carrying a single scalar sensitivity parameter \(\alpha \geq 0\) applied to SEU-valued alternatives. As \(\alpha\) ranges from \(0\) to \(\infty\), the implied behavior runs from uniform-random choice (insensitive to expected utility) to deterministic SEU maximization (perfectly sensitive to it). The empirically interesting agents live strictly between these poles, and \(\alpha\) is the coordinate that locates them there. We call \(\alpha\) the agent's SEU sensitivity: the disposition of an agent committed to SEU maximization to act in accordance with that commitment, held separate from both the content of the commitment (the agent's beliefs and utilities) and the \emph{outcomes} it happens to realize.
\addcontentsline{toc}{subsection}{1.6 Scope}
This paper is methodological. Its focus is a calibrated measurement instrument and the identifiability analysis that licenses it; the decision-theoretic and philosophical material frames that analysis. §§\hyperref[sec-abstract-model]{2}--\hyperref[sec-m1-implementation]{6} specify, motivate, and validate the instrument: the abstract choice model and its three characterizing properties (§\hyperref[sec-abstract-model]{2}), the identifiability of \(\alpha\) together with the finding that the belief and utility parameters \((\beta, \delta)\) are only weakly informed in the finite-sample Bayesian workflow for an uncertain-choice-only model (§\hyperref[sec-m0-identifiability]{3}), a Stan implementation validated by prior predictive checks, parameter recovery, and simulation-based calibration (SBC) (§\hyperref[sec-m0-implementation]{4}), and the same arc for an extended model that adds risky choices (§§\hyperref[sec-m1-identifiability]{5}--\hyperref[sec-m1-implementation]{6}). §\hyperref[sec-application]{7} then runs the full workflow end-to-end on real LLM choice data: a \(2\times2\) illustrative application crossing GPT-4o and Claude 3.5 Sonnet with insurance-claims triage and Ellsberg-style urn gambles (Ellsberg 1961), using sampling temperature as the experimental lever. The application is included to show that the methodology does real evaluative work on real choice data (Binz and Schulz 2023; Hagendorff, Fabi, and Kosinski 2023); it is not a substantive contribution to LLM behavioral science, and §\hyperref[sec-app-limits]{7.6} draws that line explicitly.
Several things are deliberately out of scope, and are flagged once at their natural locations rather than revisited: an information-optimal lottery design for sharpening the utility parameter (future work, §\hyperref[sec-disc-limitations]{8.5}); the hierarchical extension h_m01 and the multi-LLM alignment study it was built for (a planned companion paper, §\hyperref[sec-app-followup]{7.7}); generalized-sensitivity models with block-specific or context-specific \(\alpha\); a test for ambiguity aversion in the Ellsberg setting (the SEU-plus-softmax model is, by assumption, of the wrong form for that, §\hyperref[sec-app-limits]{7.6}); and any absolute-rationality ranking between LLMs (only within-design comparative claims about \(\alpha\) are supported).
\addcontentsline{toc}{subsection}{1.7 What is new}
Multinomial logit is McFadden's (McFadden 1974); Stan, SBC, and the prior-predictive/recovery/SBC workflow have been standard practice for a decade (Carpenter et al. 2017; Talts et al. 2018; Gelman et al. 2020). The contribution here is not any one of these tools. It is, first, a conceptual contribution --- item (a) below: a principled explication of what it means for a decision maker to be sensitive to, or aligned with, the requirements of SEU maximization across a range of decision contexts --- and, resting on that explication, a four-part methodological package, items (b)--(e), that makes the explication measurable and validates the measurement.
(a) A principled explication of SEU sensitivity. Within the softmax/Luce--McFadden specification, three characterizing properties of the choice rule (§\hyperref[sec-three-properties]{2.3}, proved in Appendix A) license reading its single scalar parameter \(\alpha\) as the degree to which an agent's choices track SEU maximization, with \(\alpha = 0\) at indifference to expected utility and \(\alpha \to \infty\) at deterministic maximization. This gives a graded, well-defined meaning to “alignment with expected-utility reasoning” across the family of decision problems the estimable model accommodates (the experiments expressible in the data block of the Stan models, e.g., m_01), and separates that disposition from the content of the agent's commitments (its beliefs and utilities) and from the outcomes it realizes (§\hyperref[sec-conceptual-payoff]{2.5}). The generalization across many agents and contexts via a hierarchical model is left to a companion paper (§\hyperref[sec-app-followup]{7.7}). In plain terms: \(\alpha\) assigns a precise number to how closely an agent's choices follow expected-utility reasoning, holding apart what the agent believes, what it wants, and how things happen to turn out.
(b) A clean identifiability decomposition. We separate the identifiability of \(\alpha\) from the expected-utility vector \(\eta\) --- stated as a proposition under an explicit genericity condition (§\hyperref[sec-alpha-from-eta]{3.3}, Appendix B.1) --- from the question of whether the data also pin down the ingredients of \(\eta\), namely beliefs (\(\beta\)) and utilities (\(\delta\)). At realistic sample sizes they do not: uncertain choices leave \((\beta, \delta)\) only weakly informed, as the Bayesian workflow shows directly --- wide, correlated \((\beta, \delta)\) posteriors against tight \(\alpha\) intervals (§\hyperref[sec-bd-weak]{3.4}, §\hyperref[sec-m0-recovery]{4.3}). The two are one structural picture (§\hyperref[sec-m0-link]{3.5}): given the \(\bm{\upsilon}\)-endpoint convention that fixes the utility scale, the within-menu choice-probability contrasts inform \(\alpha\) through the products \(\alpha(\eta_r - \eta_s)\), while beliefs and utilities enter only through \(\eta\) and trade off against each other in a way that uncertain-choice data barely resolve. Because the data inform \(\alpha\) only through those products, what “recovered” means is always relative to the utility spreads a particular design and prior induce --- a design-conditionality the application makes quantitative (§\hyperref[sec-app-validation]{7.4}). In plain terms: under a given design and prior, the alignment number \(\alpha\) can be recovered from choice data, but the agent's underlying beliefs and utilities often cannot be told apart from one another by that same data.
(c) A worked demonstration that in-principle identifiability does not predict finite-sample estimability. The m_0/m_1 contrast makes the gap concrete in two distinct ways. For the utility parameter \(\delta\), adding the risky block supplies a \(\beta\)-free route that secures its identifiability in principle (§\hyperref[sec-delta-id]{5.5}) yet buys almost no precision at the design sample size --- a matched-count credible-interval-width reduction under 1% (§\hyperref[sec-m1-recovery]{6.4}). For the sensitivity parameter \(\alpha\), a matched recovery study shows that finite-sample precision, in this design, is governed by data quantity rather than the type of choice that an in-principle argument might privilege. \emph{In plain terms:} a quantity that is recoverable in theory may still not be recoverable from a realistic amount of data.
(d) The marginal-SBC demarcation. Marginal rank uniformity is necessary but not sufficient for joint posterior calibration. Separately, the correlated, barely-contracted joint \((\beta, \delta)\) posterior --- which can itself be perfectly calibrated --- reflects weak joint informativeness, and marginal SBC does not reveal it: per-parameter ranks pass for both models even where beliefs and utilities are jointly only weakly resolved (§§\hyperref[sec-m0-sbc]{4.4}, \hyperref[sec-m1-sbc]{6.5}, \hyperref[sec-discussion]{8.3}). In plain terms: a standard calibration check can look healthy one parameter at a time while still concealing that two parameters are jointly unresolved.
(e) An illustrative full-pipeline application of the instrument to LLM choice data, in which the \(\alpha\) inference and the posterior-predictive summaries come in clean --- the only flagged diagnostics being confined to the weakly-informed nuisance parameters (§\hyperref[sec-app-validation]{7.4}) --- and the framework reports posterior support for a structured comparative effect in two of four LLM\(\times\)task cells while withholding it in the other two, which it reads as inconclusive on their own diagnostics --- with the design's resolution quantified explicitly for one of them (§\hyperref[sec-application]{7}). In plain terms: run end-to-end on real language-model choices, the instrument's sensitivity estimates behave as designed and report an effect in two of four settings, with the non-detections reflecting limited resolution rather than established absence.
\addcontentsline{toc}{subsection}{1.8 Roadmap}
The paper alternates between abstract/identifiability sections and their computational realizations, then applies the result. §\hyperref[sec-abstract-model]{2} fixes the abstract model. §\hyperref[sec-m0-identifiability]{3} establishes the identifiability picture for the uncertain-choice-only model m_0, and §\hyperref[sec-m0-implementation]{4} realizes and validates it in Stan. §\hyperref[sec-m1-identifiability]{5} extends the model with risky choices and revisits identifiability, and §\hyperref[sec-m1-implementation]{6} realizes and validates the extension --- inheriting both the design and the identifiability questions from the earlier pair. §\hyperref[sec-application]{7} runs the validated workflow on real LLM data, and §\hyperref[sec-discussion-top]{8} draws the methodological lessons together. The formal core --- the proofs that close §§\hyperref[sec-m0-identifiability]{3} and \hyperref[sec-m1-identifiability]{5} --- is in Appendix B; the body states results and points to it. Concretely, the dependency structure is
\[ \underbrace{\S2}_{\text{model}} \;\longrightarrow\; \underbrace{\S3 \to \S4}_{\texttt{m\_0}} \;\longrightarrow\; \underbrace{\S5 \to \S6}_{\texttt{m\_1}} \;\longrightarrow\; \underbrace{\S7}_{\text{application}} \;\longrightarrow\; \underbrace{\S8}_{\text{discussion}}, \]
with §\hyperref[sec-m0-identifiability]{3} feeding both §\hyperref[sec-m0-implementation]{4} and §\hyperref[sec-m1-identifiability]{5}, the two implementation sections (§§\hyperref[sec-m0-implementation]{4}, \hyperref[sec-m1-implementation]{6}) converging on the application, and Appendix B supplying the proofs that close the two identifiability sections.
How to read this paper. Readers approaching from statistics or the Bayesian workflow may treat §§\hyperref[sec-abstract-model]{2}--\hyperref[sec-m0-identifiability]{3} and \hyperref[sec-m1-identifiability]{5} as the modeling setup and concentrate on the validation and application (§§\hyperref[sec-m0-implementation]{4}, \hyperref[sec-m1-implementation]{6}, \hyperref[sec-application]{7}); readers approaching from decision theory or formal epistemology will find the conceptual claims in §§\hyperref[sec-motivation]{1}--\hyperref[sec-abstract-model]{2} and \hyperref[sec-discussion-top]{8} and can take the identifiability propositions (§§\hyperref[sec-m0-identifiability]{3}, \hyperref[sec-m1-identifiability]{5}) on their statements, with proofs in Appendix B. Terms of art from each field are glossed at first use for the others.
\addcontentsline{toc}{section}{The Abstract Model: Softmax Choice and Sensitivity}
This section specifies the choice model abstractly, before any commitment to where values come from. The payoff of the abstraction is that the three properties that license reading \(\alpha\) as a sensitivity parameter (§\hyperref[sec-three-properties]{2.3}) hold for any value function; the subjective expected utility (SEU) specialization of §\hyperref[sec-seu-specialization]{2.4} then inherits them as corollaries.
\addcontentsline{toc}{subsection}{2.1 Alternatives, values, and the sensitivity parameter}
Let \(\mathcal{R} = \{1, \dots, R\}\) be a finite set of distinct alternatives. A value function \(V : \mathcal{R} \to \mathbb{R}\) assigns a real number \(V(r)\) to each alternative. A decision problem \(m\) presents a subset of alternatives, encoded by availability indicators \(I_{m,r} \in \{0,1\}\). A sensitivity parameter \(\alpha \geq 0\) governs how sharply choices track value differences.
\addcontentsline{toc}{subsection}{2.2 The softmax choice rule}
The probability of selecting alternative \(r\) from the available set in problem \(m\) is the softmax (Luce--McFadden--Boltzmann) rule
The functional form has independent precedents: Luce's (1959) ratio-scale choice axiom, McFadden's (1974) random-utility derivation of multinomial logit, and the Boltzmann distribution of statistical mechanics (where \(\alpha\) is inverse temperature). These three routes reach the same functional form --- a choice likelihood exponential in value differences --- from independent starting points, which is part of what recommends it here. We use \(\alpha\) rather than a temperature \(T = 1/\alpha\) because higher \(\alpha\) corresponds to higher sensitivity, the more intuitive direction for our purposes. (This model temperature \(T = 1/\alpha\) is distinct from the LLM sampling temperature that appears as an experimental covariate in §\hyperref[sec-application]{7}; the application asks, in part, how the estimated \(\alpha\) responds to that sampling temperature.)
Why softmax rather than probit or another stochastic-choice wrapping? Three features recommend it. Luce's (1959) independence-of-irrelevant-alternatives axiom yields a ratio form \(P(r \mid \mathcal{A}) = w(r)/\sum_{j \in \mathcal{A}} w(j)\) over positive weights; specifying an exponential link \(w(r) = \exp(\alpha V(r))\) from the real-valued value scale to those weights --- equivalently, log-odds linearity in value differences --- selects the softmax (Equation (ref)). Given that link, softmax embeds the optimization and uniform-choice limits as the two endpoints of a single scalar \(\alpha\) (Properties 2 and 3 of §\hyperref[sec-three-properties]{2.3}, proved as Theorems A.2 and A.3 in Appendix A), and its log-likelihood is concave in the index \(\alpha V\) --- a standard multinomial-logit property (McFadden 1974; Train 2009) --- so that, with the value scale \(V\) held fixed, the \(\alpha\)-likelihood is unimodal and free of spurious local optima. (The concavity is in the index \(\alpha V\), i.e. the surface in \(\alpha\) with \(V\) fixed; it is not joint concavity in \(\alpha\) together with the value parameters, because \(\alpha V\) is bilinear in the two. Nor does concavity guarantee a finite maximizer: if the observed choices happen to select a value-maximal alternative in every problem --- the analogue of complete separation in logistic regression --- the likelihood increases in \(\alpha\) without bound, and it is the prior on \(\alpha\) that restores a proper posterior. The Bayesian workflow of §\hyperref[sec-m0-implementation]{4} never relies on a maximum-likelihood point estimate.) These are two distinct virtues, and we keep them apart: the log-odds linearity and single-scalar limit structure above make \(\alpha\) interpretable, while the concavity is what makes it well-behaved to estimate. Probit, by contrast, replaces the Luce ratio with Gaussian latent errors: a probit scale parameter carries the same two limits (uniform choice as it grows, deterministic optimization as it vanishes), but probit offers no closed-form IIA/log-odds structure and no comparably tractable single-scalar link, so it is the softmax's closed-form log-odds linearity that singles it out here (Train 2009).
Related uses of a softmax precision parameter. The scalar \(\alpha\) has close relatives across several literatures. In experimental game theory it is the precision parameter \(\lambda\) of quantal response equilibrium (McKelvey and Palfrey 1995), which plays for strategic choice the role \(\alpha\) plays here for single-agent choice. In the econometrics of risky choice, Hey and Orme (1994) estimate expected-utility and generalized-EU models wrapped in stochastic choice, initiating a literature on how the noise specification shapes inference about the deterministic core. In machine learning, the same exponential-in-value rule appears as Boltzmann rationality, the standard observation model in inverse reinforcement learning (Ziebart et al. 2008). Two cautions from the econometric branch carry over. Wilcox (2011) argues that a logit-style precision is not comparable across contexts unless the utility scale is normalized per context (“contextual utility”), and Apesteguia and Ballester (2018) show that fixed-precision random-utility wrappings can order risk attitudes non-monotonically. We take both seriously rather than claiming immunity: the \(\bm{\upsilon}\)-endpoint convention of §\hyperref[sec-notation]{2.6} fixes the utility scale (best consequence \(= 1\), worst \(= 0\)) uniformly across problems. This is a global normalization over a fixed consequence space, not the per-context renormalization Wilcox's contextual utility prescribes; the two coincide here only because every problem in a given design draws its consequences from the same \(K\)-element space, so the within-context utility range is the same fixed range in every problem. Within such fixed-consequence-space designs the convention supplies the scale comparability the critique asks for; across designs with different consequence spaces it does not, which is one reason the paper licenses only within-design comparative readings of \(\alpha\) (§\hyperref[sec-app-limits]{7.6}, §\hyperref[sec-disc-alpha]{8.4}), not cross-instrument cardinal ones.
\addcontentsline{toc}{subsection}{2.3 Three characterizing properties of softmax}
The following hold for any value function \(V\), within each decision problem \(m\) and relative to its available set \(\mathcal{A}_m = \{r : I_{m,r} = 1\}\). Write \(V^\ast_m = \max_{r \in \mathcal{A}_m} V(r)\) for the maximal available value, \(\mathcal{R}^\ast_m = \{r \in \mathcal{A}_m : V(r) = V^\ast_m\}\) for the value-maximizing set, and \(\mathcal{R}^-_m = \mathcal{A}_m \setminus \mathcal{R}^\ast_m\). Proofs are in Appendix A.
Together the three properties give exactly the conceptual content needed to read \(\alpha\) as “sensitivity to value maximization,” with no reference yet to where values come from: \(\alpha\) interpolates monotonically (Property 1) between indifference to value (Property 3) and exclusive pursuit of the maximum (Property 2).
\addcontentsline{toc}{subsection}{2.4 SEU specialization}
Specialize the value function to a subjective expected utility. Fix \(K\) consequences with an (ordered) utility vector \(\bm{\upsilon} \in \mathbb{R}^K\). Each alternative \(r\) carries a subjective probability distribution \(\bm{\psi}_r \in \Delta^{K-1}\) over consequences, and its value is its expected utility
Because \(\eta_r\) is just a particular value function, Properties 1--3 apply verbatim, and \(\alpha\) now measures sensitivity to SEU maximization specifically: the disposition to choose the expected-utility-maximizing alternative.
\addcontentsline{toc}{subsection}{2.5 The conceptual payoff}
Sensitivity to SEU maximization is the disposition of an SEU-committed agent to act in accordance with its commitments. The contrast between an agent's commitment to a standard and its performance relative to that standard is Levi's (Levi 1980, chap. 1); we return to it, and to the sense in which \(\alpha\) measures a tendency to perform, in §\hyperref[sec-disc-meaning]{8.1}. The disposition is distinct from (a) the content of those commitments --- the agent's beliefs \(\bm{\psi}\) and utilities \(\bm{\upsilon}\) --- and (b) the outcomes the agent happens to realize. The SEU construction thus decomposes choice behavior conceptually into beliefs, utilities, and sensitivity. The three pieces are conceptually separable in this way, but they are not equally separable \emph{empirically}: as §\hyperref[sec-bd-weak]{3.4} and §\hyperref[sec-delta-id]{5.5} show, \(\alpha\) separates cleanly from the rest, whereas \(\beta\) and \(\delta\) are only weakly informed by uncertain choices and are recovered jointly rather than each in isolation.
\addcontentsline{toc}{subsection}{2.6 Notation, the utility-scale convention, and the \(\beta\) gauge}
For the parameterized models of §\hyperref[sec-m0-identifiability]{3} and §\hyperref[sec-m1-identifiability]{5} we generate subjective probabilities from \(D\)-dimensional feature vectors \(w_r\) via a linear-softmax belief map \(\bm{\psi}_r = \mathrm{softmax}(\beta\, w_r)\) with \(\beta \in \mathbb{R}^{K \times D}\), in the multinomial-logit tradition (McFadden 1974; Train 2009). The utility vector is built from increments \(\delta \in \Delta^{K-2}\) via cumulative sums, with endpoints fixed by convention. (The \(K-1\) increments \(\delta_j = \upsilon_{j+1} - \upsilon_j\) are nonnegative and sum to \(\upsilon_K - \upsilon_1 = 1\), so \(\delta\) lies in the \((K-2)\)-simplex \(\Delta^{K-2}\), matching the Stan declaration simplex{[}K\ -\ 1{]}\ delta.)
Two facts about this parameterization recur throughout and are stated once here, then referenced rather than redefined. The first fixes the zero and unit of the utility scale, which is what makes \(\alpha\) a well-posed target; the second records an indeterminacy of the belief map.
The consolidated glossary is collected once here (Table (ref)) and referenced, not redefined, elsewhere.
\addcontentsline{toc}{subsection}{2.7 The identifiability question}
The choice-probability function is determined by \((\alpha, \bm{\psi}, \bm{\upsilon})\). Whether the parameterizations of \(\bm{\psi}\) and \(\bm{\upsilon}\) --- namely \((\beta, \delta)\) --- are recoverable from observed choices is the question §\hyperref[sec-m0-identifiability]{3} and §\hyperref[sec-m1-identifiability]{5} take up.
\addcontentsline{toc}{section}{Choice Under Uncertainty Alone: Identifiability of \(\alpha\), and the Weakly-Informed \((\beta,\delta)\)}
\addcontentsline{toc}{subsection}{3.1 The parameterization}
Recall from §\hyperref[sec-notation]{2.6} the uncertain-choice (“m_0”) model. Subjective probabilities are generated from \(D\)-dimensional features \(w_r\) via \(\bm{\psi}_r = \mathrm{softmax}(\beta\, w_r)\) with \(\beta \in \mathbb{R}^{K\times D}\). Utilities are constructed from increments \(\delta\) by cumulative summation, yielding \(0 = \upsilon_1 \le \upsilon_2 \le \cdots \le \upsilon_K = 1\). There are \(K-1\) nonnegative increments, summing to one, so \(\delta\) ranges over the \((K-2)\)-dimensional simplex \(\Delta^{K-2}\) --- the sum-to-one constraint removes one degree of freedom --- and Stan declares it simplex{[}K\ -\ 1{]}\ delta, counting the \(K-1\) entries. Expected utilities are \(\eta_r = \bm{\psi}_r^\top \bm{\upsilon}\), and choices follow Equation (ref) with \(V(r) = \eta_r\). The \(\beta\) gauge (§\hyperref[sec-notation]{2.6}) is the only \(\beta\) indeterminacy we use.
\addcontentsline{toc}{subsection}{3.2 Why this parameterization}
Linear-softmax for \(\bm{\psi}\) aligns with the multinomial-logit / discrete-choice tradition (McFadden 1974; Train 2009); ordered utilities with endpoints fixed by convention remove the affine scale-and-shift indeterminacy of vNM utility (Theorem A.4). The two conventions together leave \(\alpha\) as a scalar with the limit interpretation of §\hyperref[sec-three-properties]{2.3}.
\addcontentsline{toc}{subsection}{3.3 Identifiability of \(\alpha\) from \(\eta\)}
Sketch (full proof in Appendix B.1). On any menu with non-constant \(\eta\) the log-odds between two alternatives with \(\eta_r \neq \eta_s\) is \(\log[P(r)/P(s)] = \alpha(\eta_r - \eta_s)\), strictly monotone in \(\alpha\); the softmax choice-probability map is therefore injective in \(\alpha\) on such a menu.
The structure of the claim matters: \(\alpha\) is identified from \(\eta\), not from \((\beta, \delta)\) directly. Whether \((\beta, \delta)\) themselves are recoverable from uncertain choices is a separate question --- the subject of §\hyperref[sec-bd-weak]{3.4}.
\addcontentsline{toc}{subsection}{3.4 Why \((\beta,\delta)\) are weakly informed by uncertain choices}
Proposition 3.1 identifies \(\alpha\) from \(\eta\). The remaining question is whether the data also pin down the two ingredients that make up \(\eta\) --- beliefs (\(\beta\)) and utilities (\(\delta\)). They do not, at realistic sample sizes, and the reason is structural. Expected utility composes the two multiplicatively, \[ \eta_r \;=\; \mathrm{softmax}(\beta\, w_r)^\top \bm{\upsilon}(\delta), \] the product of a \(\beta\)-driven simplex and a \(\delta\)-driven utility vector. Uncertain choices see \((\beta, \delta)\) only through this scalar \(\eta_r\), so a family of \((\beta, \delta)\) pairs that trades a change in beliefs against a compensating change in utilities leaves the implied expected utilities --- and hence the choice probabilities --- almost unchanged. The compensation is approximate rather than exact: the data can pin down \(\eta\) (and through it \(\alpha\)) sharply while still failing to separate the two factors whose product forms it. The \(\bm{\upsilon}\)-endpoint convention (§\hyperref[sec-notation]{2.6}) fixes the zero and unit of the utility scale, but it does not touch this multiplicative \((\beta,\delta)\) coupling, which is not an indeterminacy that any convention removes; the coupling is what leaves \(\beta\) and \(\delta\) only weakly informed by uncertain-choice data.
This shows up directly in the Bayesian workflow rather than as a separate formal claim. Parameter recovery in m_0 (§\hyperref[sec-m0-recovery]{4.3}) returns wide marginal credible intervals for both \(\beta\) and \(\delta\) --- in contrast to the tight, well-calibrated \(\alpha\) intervals --- together with correlated \(\beta\)--\(\delta\) estimation errors: concretely, across recovery replicates the posterior-mean estimation errors of a representative weight entry \(\beta_{1,1}\) and a representative increment \(\delta_1\) are correlated (Pearson; §\hyperref[sec-m0-recovery]{4.3}). This is an across-replicate error correlation, not a within-posterior one; because raw \(\beta\) entries carry the row-shift gauge we read this representative-component diagnostic as an illustrative signature of the coupling rather than a gauge-invariant statistic, and we do not attach significance to its sign, which is not stable across simulation settings --- it shifts, for instance, as the true \(\alpha\) used to generate the recovery data is varied. The recovered posteriors concentrate not on \((\beta,\delta)\) individually but on a compensating trade-off between them. That correlated, barely-contracted joint posterior is the practical signature of the multiplicative coupling, and it is what we mean throughout by calling \((\beta,\delta)\) weakly informed by uncertain choices.
We state this as an empirical feature of the posterior, not as a theorem: no strict non-identifiability is claimed, and no gauge group on \((\beta,\delta)\) beyond the \(\beta\) row-shift of §\hyperref[sec-notation]{2.6} is asserted. The point is that uncertain-choice data alone supply little information for separating beliefs from utilities --- which is exactly what motivates the extended model of §\hyperref[sec-m1-identifiability]{5}.
\addcontentsline{toc}{subsection}{3.5 The two facts are one structural picture}
\(\alpha\) is identified from \(\eta\) (Proposition 3.1), whereas the two ingredients of \(\eta\) --- beliefs and utilities --- are only weakly informed by uncertain choices (§\hyperref[sec-bd-weak]{3.4}). These are two sides of one fact: the map \((\beta,\delta) \to \eta\) collapses beliefs and utilities multiplicatively into a single scalar per alternative, and choices act on those scalars through \(\alpha\). Given the design and prior, the data therefore speak clearly about \(\alpha\) --- the within-menu \(\eta\)-contrasts identify it up to the scale fixed by the \(\bm{\upsilon}\)-endpoint convention --- while saying little about how a given \(\eta\) splits into its \((\beta,\delta)\) parts. Strictly, the observed log-odds identify the products \(\alpha(\eta_r - \eta_s)\); it is the endpoint convention that fixes the utility scale, together with the prior and design, that turns this into a sharp statement about \(\alpha\) itself --- as the posterior-contraction, parameter-recovery, and SBC diagnostics of §\hyperref[sec-m0-implementation]{4} confirm. The dependence on the design is not a formality: how sharply the data speak about \(\alpha\) is governed by the within-menu \(\eta\)-gaps the design induces, and the application reports these gaps, alongside the resulting \(\alpha\)-posterior contraction, as explicit per-design diagnostics (§\hyperref[sec-app-validation]{7.4}).
Why \(\alpha\) recovers cleanly anyway. Empirically, the two effects do not interfere: in parameter recovery (§\hyperref[sec-m0-recovery]{4.3}) \(\alpha\) is recovered with low bias and well-calibrated intervals regardless of the wide, correlated spread in the \((\beta,\delta)\) posterior. Proposition 3.1 does not by itself guarantee this: it identifies \(\alpha\) given \(\eta\), and since \((\beta,\delta)\) determine the \(\eta\)-spreads, a shift in that spread can in principle trade off against \(\alpha\) inside the products \(\alpha(\eta_r - \eta_s)\). What the recovery study shows is that under this design and prior the trade-off is not exercised: the indeterminacy in how \(\eta\) splits into beliefs and utilities leaves the \(\alpha\) estimate essentially untouched. We report this as an observed, design- and prior-dependent feature of the recovered posterior rather than derive it from a separability argument.
This is a structural feature of decisions under uncertainty: beliefs and utilities enter only through expected utility, and uncertain-choice data alone supply little information for separating them.
\addcontentsline{toc}{subsection}{3.6 From the weakly-informed \((\beta,\delta)\) to the extended model}
That uncertain choices barely inform \((\beta,\delta)\) motivates the extended model of §\hyperref[sec-m1-identifiability]{5}. Before getting there, §\hyperref[sec-m0-implementation]{4} shows that the practical consequences --- wide credible intervals on \((\beta,\delta)\), tight intervals on \(\alpha\) --- are directly visible in a Stan implementation, and that marginal simulation-based calibration nonetheless passes for all three parameters (the marginal-SBC demarcation).
\addcontentsline{toc}{section}{A Basic Computational Implementation: Model m_0 in Stan}
This section turns the abstract m_0 model into an estimable Bayesian model and shows that the identifiability picture of §\hyperref[sec-m0-identifiability]{3} is directly visible in a Stan implementation: \(\alpha\) recovers well, \((\beta, \delta)\) recover poorly, and marginal simulation-based calibration (SBC) nonetheless passes for all three --- the marginal-SBC demarcation.
\addcontentsline{toc}{subsection}{4.1 Model specification}
Data. \(M\) decision problems; \(R\) distinct alternatives with \(D\)-dimensional feature vectors \(w_r\); availability indicators \(I_{m,r} \in \{0,1\}\); observed choices \(y_m\). A problem \(m\) presents \(N_m = \sum_r I_{m,r}\) alternatives.
Parameters and priors. \[ \alpha \sim \mathrm{Lognormal}(0, 1), \qquad \beta_{k,d} \sim \mathcal{N}(0,1), \qquad \delta \sim \mathrm{Dirichlet}(1, \dots, 1). \] The \(\mathrm{Lognormal}(0,1)\) prior on \(\alpha\) (median \(1\), roughly 90% of mass in \((0.19, 5.18)\)) is a substantive choice: it spans the full near-uniform-to-near-deterministic sensitivity range while concentrating mass near \(\alpha = 1\). The prior shape co-determines which \(\alpha\) values are well-measured in finite samples; we defend it via prior-predictive coverage in §\hyperref[sec-m0-prior]{4.2}.
Transformed parameters. \(\bm{\psi}_r = \mathrm{softmax}(\beta\, w_r)\); the ordered utilities \(\upsilon_k = \sum_{j<k}\delta_j\) via a cumulative sum (equivalently upsilon\ =\ cumulative_sum(append_row(0,\ delta))), so \(\upsilon_1 = 0\) and \(\upsilon_K = 1\); and \(\eta_r = \bm{\psi}_r^\top \bm{\upsilon}\).
Likelihood. \(y_m \sim \mathrm{Categorical}\big(\mathrm{softmax}(\alpha\, \eta_{[m]})\big)\), where \(\eta_{[m]}\) collects the expected utilities of the alternatives available in problem \(m\).
\addcontentsline{toc}{subsection}{4.2 Prior predictive analysis}
The goal is to confirm that the priors permit a sensible range of choice behaviors before any data are seen. The summary statistic we monitor is the per-dataset SEU-maximizer rate: the fraction of choices that select the expected-utility-maximal available alternative under the simulated parameters. Drawing \((\alpha, \beta, \delta)\) from the prior, simulating data via m_0_sim.stan, and tabulating this rate shows its prior-predictive distribution covering the full range from near-uniform choice (rate \(\approx M^{-1}\sum_m 1/N_m\), the chance baseline --- equal to \(1/\bar N\) only when menu sizes are equal) to near-deterministic SEU maximization (rate \(\approx 1\)). The \(\mathrm{Lognormal}(0,1)\) prior on \(\alpha\) places substantial mass at both low (near-random) and high (near-deterministic) sensitivities, so neither extreme of the sensitivity range is ruled out a priori.
\addcontentsline{toc}{subsection}{4.3 Parameter recovery}
Paradigm. Draw parameters from the prior via m_0_sim.stan, simulate data, fit m_0.stan, and compare posterior summaries to the known true values. We evaluate bias (\(\approx 0\) ideal), credible-interval coverage (\(\approx\) nominal 90%), and interval width / RMSE as precision measures.
\(\alpha\) recovery. \(\alpha\) recovers with low aggregate bias, well-calibrated 90% intervals, and useful precision, as Figure (ref) shows: a true-versus-estimated scatter with 90% credible intervals, and per-replicate intervals whose empirical coverage matches the nominal 90%. Two qualifications keep this from overstatement. First, the low bias is an aggregate (across-replicate) summary; the scatter shows the shrinkage a Bayesian point estimate under a proper prior should show --- posterior means pulled toward the prior center, visibly so for the largest true \(\alpha\), so the signed error is conditionally negative in the upper range and positive at the bottom even though it is near zero on average. Second, precision is not uniform: intervals widen materially with \(\alpha\), so “useful precision” describes the bulk of the prior range, not its upper tail. This is the computational counterpart of Proposition 3.1 --- which identifies \(\alpha\) given \(\eta\) --- made empirical: under this design and prior the data supply enough information to pin \(\alpha\) over the range the prior concentrates on. The clean recovery is thus a design- and prior-dependent fact, not a corollary of the proposition alone.
\(\beta\) and \(\delta\) recovery. By contrast \((\beta, \delta)\) show detectably wider intervals and intervals that narrow more slowly with sample size, together with correlated posterior-mean estimation errors --- computed across recovery iterations between a representative weight entry \(\beta_{1,1}\) and a representative increment \(\delta_1\). We report this last item as an illustrative, non-gauge-invariant signature rather than a signed quantity: its sign is not robust (it depends on the true \(\alpha\) used to simulate the recovery data, since \(\beta_{1,1}\) carries the row-shift gauge of §\hyperref[sec-notation]{2.6}), and it is the presence of across-replicate error coupling, not its direction, that matters here. This is the computational manifestation of the weakly-informed \((\beta,\delta)\) of §\hyperref[sec-bd-weak]{3.4}: the multiplicative \((\beta,\delta)\) coupling realized as a correlated joint posterior and wide marginal intervals. Crucially, the poor \((\beta, \delta)\) recovery does not contaminate \(\alpha\) --- exactly as the empirical separation noted in §\hyperref[sec-m0-link]{3.5}.
\addcontentsline{toc}{subsection}{4.4 Simulation-based calibration}
Simulation-based calibration (SBC) checks whether the model-and-sampler together return posteriors that are calibrated on average: if a true parameter is drawn from the prior, data are simulated from it, and the model is refit, then a well-calibrated 90% credible interval should contain that truth 90% of the time. The rank test below checks this across all quantiles at once.
Method. SBC draws a parameter from the prior, simulates data, fits the model, and computes the rank of the true value within the posterior draws; under correct calibration these ranks are uniform (Talts et al. 2018) --- a truth drawn from the prior is equally likely to fall at any rank within a correctly calibrated posterior, so systematic departures from uniformity flag a mismatch among model, data, and sampler. We diagnose uniformity with rank histograms and the ECDF-difference plot with a simultaneous confidence band (Modrák et al. 2025), at \(N_{\mathrm{sbc}} = 999\), thinning \(4\), single chain, with the \(L = N_{\mathrm{sbc}}\) convention (Appendix D.3).
Result. At \(N_{\mathrm{sbc}} = 999\) the marginal rank distributions are consistent with uniformity for \(\alpha\), \(\beta\), and \(\delta\). This run does not independently reproduce the weakly-informed \((\beta,\delta)\) finding of §§\hyperref[sec-m0-identifiability]{3}--\hyperref[sec-m0-recovery]{4.3}, because that weakness is one of joint informativeness and geometry --- a correlated, barely-contracted \((\beta,\delta)\) posterior that can itself be perfectly calibrated --- to which marginal SBC is blind.
\addcontentsline{toc}{subsection}{4.5 Section summary}
m_0 succeeds at recovering \(\alpha\) --- the parameter of primary interest --- but the data leave \((\beta,\delta)\) only weakly informed, leaving the expected-utility content underdetermined at realistic \(n\). This motivates the extended model of §\hyperref[sec-m1-identifiability]{5}, whose implementation we take up in §\hyperref[sec-m1-implementation]{6}.
\addcontentsline{toc}{section}{The Extended Abstract Model with Risky Choices}
\addcontentsline{toc}{subsection}{5.1 Enriching the choice domain --- a two-step justification}
The extended model (“m_1”) augments the uncertain-choice setup with risky alternatives whose probabilities are objectively given. Step 1 alone does not identify \(\delta\); the two steps below together do.
History (brief). Knight (1921) distinguished risk from uncertainty; von Neumann and Morgenstern (1947) axiomatized expected utility under risk; Savage (1954) gave SEU under uncertainty. Anscombe and Aumann (1963) combined the two by considering acts mapping states to lotteries, which lets one identify utilities from the risky margin and beliefs from the state margin. Our m_1 uses a related empirical strategy: the connection is conceptual, not a direct application of the Anscombe--Aumann representation theorem. The important part is Steps 1+2, not the representation theorem.
\addcontentsline{toc}{subsection}{5.2 The extended model}
Augment the \(M\) uncertain problems with \(N\) risky problems built from \(S\) distinct lotteries \(\pi_s\) (objective simplexes over the \(K\) consequences). The model shares the same \(\alpha\) and the same \(\delta \to \bm{\upsilon}\) map across both blocks; risky expected utilities \(\eta^{(r)}_s = \pi_s^\top \bm{\upsilon}\) depend only on \(\delta\), not on \(\beta\). (In the Stan code the lottery simplices enter the data block as x and are compacted per problem into x_risky; we use \(\pi\) in the body.)
\addcontentsline{toc}{subsection}{5.3 Preservation of the three sensitivity properties}
The softmax properties of §\hyperref[sec-three-properties]{2.3} are properties of the choice rule given any value function, so they apply unchanged to both uncertain and risky sub-problems. Therefore the interpretation of \(\alpha\) as sensitivity-to-SEU-maximization is preserved in the extended model.
\addcontentsline{toc}{subsection}{5.4 Identifiability of \(\alpha\) in the extended model}
As in §\hyperref[sec-alpha-from-eta]{3.3}, this is identification of \(\alpha\) given \(\eta\). Jointly --- with \(\delta\) (and, for uncertain alternatives, \(\beta\)) unknown --- non-constant values pin down only the products \(\alpha(\eta_r - \eta_s)\); it is the additional structure of m_1, namely the lottery-diversity condition of §\hyperref[sec-delta-id]{5.5} that recovers the utility scale from the risky block, that turns the product into a statement about \(\alpha\) itself. In m_0 the same role is played by the endpoint convention together with the prior and design (§\hyperref[sec-m0-link]{3.5}).
\addcontentsline{toc}{subsection}{5.5 Identifiability of \(\delta\) in the extended model}
Intuition (full proof in Appendix B.3). A risky menu's softmax depends on the linear functional \(\pi_s^\top \bm{\upsilon}\). Pinning down \(K\) utilities requires \(K-1\) independent lottery directions (one degree of freedom is fixed by the endpoint convention); with enough diverse lotteries those directions are present. Formally, when the lottery differences span the simplex tangent space, the map \(\bm{\upsilon} \mapsto ((\pi_s - \pi_1)^\top \bm{\upsilon})_s\) has rank \(K-1\), so the risky log-odds recover \(\bm{\upsilon}\) --- and jointly \(\alpha\) --- once the endpoint convention fixes the scale, hence \(\delta\). \(S \geq K\) is only the necessary cardinality condition; the spanning condition itself concerns the realized lotteries, so we verify it directly for the \(K = 3\), \(S = 15\) lottery set of Appendix D.0: the \(14 \times 3\) lottery-difference matrix (differences of the other 14 lotteries against the first) has rank \(2 = K - 1\), with nonzero singular values \(1.08\) and \(0.67\) (spikes/report_design_diagnostics_spike.py) --- the condition holds for the implemented design as a checked fact, not an inference from the count. The proof is a linear-algebra inversion.
\addcontentsline{toc}{subsection}{5.6 What the extended model does not deliver: \(\beta\)}
The pointwise observation that a single scalar \(\eta_r\) does not determine the \((K-1)\)-dimensional simplex \(\bm{\psi}_r\) when \(K \geq 3\) does not settle the question for \(\beta\), because the belief vectors are linked across alternatives by the shared map \(\bm{\psi}_r = \mathrm{softmax}(\beta w_r)\); what matters is the rank of the composite map, and that is design-specific.
Some intuition for the rank condition. The map \(\beta \mapsto (\eta\text{-contrasts})\) sends the \((K-1)D\) free coordinates of the gauge-fixed weight matrix to the list of expected-utility contrasts the design produces. Its Jacobian is the matrix of first partial derivatives of those output contrasts with respect to the input coordinates --- a local linear picture of how the observable contrasts respond to a small change in \(\beta\). We call it the contrast-Jacobian to emphasize that it is assembled from the \(\eta\)-contrasts (which are gauge-invariant), not from the raw expected utilities \(\eta_r\) (which are defined only up to the per-menu additive constant of §\hyperref[sec-notation]{2.6}). By the inverse function theorem, where this Jacobian has full column rank \((K-1)D\) the map is locally invertible: two nearby \(\beta\) that differ modulo the gauge cannot produce identical contrasts, so \(\beta\) is locally identified. Where the rank is deficient, some direction in \(\beta\)-space changes no contrast --- and hence no choice probability --- at all, so \(\beta\) can move undetected and is not identified. Because the rank can in principle vary with \(\beta\), we do not read it off at a single point but evaluate it numerically at a spread of \(\beta\) drawn from the prior.
The two design regimes in this paper land on opposite sides of the condition (spikes/report_design_diagnostics_spike.py). At the foundational design (\(K = 3\), \(D = 5\), \(R = 15\)) the gauge-fixed \(\beta\) has \((K-1)D = 10\) degrees of freedom against up to \(R - 1 = 14\) contrasts, and the computed contrast-Jacobian rank is exactly \(10\) at every one of 20 prior draws checked --- so \(\beta\) is locally identified modulo the gauge there in principle, which sharpens rather than undercuts the §\hyperref[sec-m0-recovery]{4.3} finding: the weak \(\beta\) recovery at that design is a finite-sample information problem, not a structural non-identification. At the application designs (\(D = 32\), \(R = 30\); §\hyperref[sec-app-design]{7.2}) the inequality reverses decisively --- \((K-1)D = 64\) (\(K = 3\)) or \(96\) (\(K = 4\)) degrees of freedom against a computed contrast rank of \(29 = R - 1\) --- so there \(\beta\) is genuinely not identified, and no amount of per-design data volume can change that. We accordingly reserve “irreducible nuisance” for designs where the rank condition fails, as it does in the application; at designs where it holds, \(\beta\) is better described as identified-but-weakly-informed. Either way, the contribution of the risky block is to sharpen \(\bm{\upsilon}\), not to recover \(\beta\) (Appendix B.4).
\addcontentsline{toc}{subsection}{5.7 The conceptual upshot --- and its finite-sample caveat}
Enriching the choice domain with risky alternatives yields, in principle, identification of \(\alpha\) (from either block) and of the utilities \(\delta\) (from the \(\beta\)-free risky block, Step 1, plus lottery diversity, Step 2). It does not deliver the beliefs: \(\beta\) enters uncertain choices only through the per-menu \(\eta\)-contrasts, so whether it is identified at all is a design-rank question (Proposition 5.3) --- the condition fails outright at the application designs, and even at the foundational design, where it holds, the finite-sample information is meager (§\hyperref[sec-m0-recovery]{4.3}). And even the utility gain is an in-principle one --- it implies estimability only in the limit. §\hyperref[sec-m1-implementation]{6} asks whether the gain is realized at realistic finite sample sizes, and finds that for \(\delta\) it largely is not --- the cleanest illustration in the paper of the gap between identifiability and precise estimability.
\addcontentsline{toc}{section}{A Basic Computational Implementation of the Extended Model: m_1 in Stan}
\addcontentsline{toc}{subsection}{6.1 Model specification}
m_1 keeps the m_0 parameters and priors --- \(\alpha \sim \mathrm{Lognormal}(0,1)\), \(\beta_{k,d} \sim \mathcal{N}(0,1)\), \(\delta \sim \mathrm{Dirichlet}(1,\dots,1)\) --- and adds \(N\) risky problems built from \(S\) distinct objective lotteries \(\pi_s \in \Delta^{K-1}\) (Stan data: x). The log-likelihood is the sum of the m_0 uncertain-choice term and a risky-choice term that uses \(\eta^{(r)}_s = \pi_s^\top \bm{\upsilon}\) directly --- in Stan, eta_risky{[}i{]}\ =\ dot_product(x_risky{[}i{]},\ upsilon). The key structural contrast is that risky expected utilities depend only on \(\bm{\upsilon}\) (hence \(\delta\)), not on \(\beta\) (§\hyperref[sec-m1-twostep]{5.1}, Step 1).
\addcontentsline{toc}{subsection}{6.2 Study design and the matched comparison}
Design. We use the existing risky-choice configuration (\(K = 3\), \(S = 15\) lotteries; full constants in Appendix D.0), not a \(\delta\)-optimal design. The matched design fixes one study design across conditions and slices the same simulated choices four ways, so that precision differences across conditions reflect the estimand and the choice type rather than random variation between separate simulations:
The central test is B vs C: same total choice count, same true parameters per iteration, with only the model and the type of choice differing. The informative control is A vs B: doubling the uncertain block alone, holding the model fixed. Condition D doubles both blocks and completes the factorial slicing; it plays no role in the two named contrasts, so §\hyperref[sec-m1-recovery]{6.4} reports no separate D comparison.
What each reported magnitude is. Each of the \(n = 100\) iterations runs the recovery loop of §\hyperref[sec-m0-recovery]{4.3} on the matched design: draw one set of true parameters from the prior, simulate the four sliced choice sets (conditions A--D) under those parameters, fit the corresponding model in each condition, and record a per-iteration recovery metric --- a posterior-mean RMSE against the known truth, or a posterior credible-interval width --- for each condition. Because every condition within an iteration shares the same true parameters, the conditions are paired: a contrast such as B vs C or A vs B is formed within each iteration and then summarized across iterations by its paired-iteration median. The interval we report is a bootstrap 90% CI over the \(n = 100\) iterations (10,000 resamples): we resample iterations with replacement, recompute the paired median on each resample, and take its central 90% range, so the CI measures how stable the aggregate contrast is across simulated datasets.
Sign convention. Each §6.4 magnitude is a signed percentage change in a precision measure (a posterior-mean RMSE or a credible-interval width), signed so that a positive value denotes an improvement --- a reduction in RMSE or in CI width --- and a negative value denotes degraded precision. Its bootstrap 90% CI is expressed in these same signed-percentage units --- it is the interval for that one signed quantity, not a separate scale --- so a CI lying wholly above zero indicates a reliable improvement, whereas a CI straddling zero leaves the direction undetermined.
What we do not claim. We do not engineer a design that makes the \(\delta\) CI “materially” smaller than m_0's; at \(n = 100\) the matched-count \(\delta\) CI-width gain is \(\approx 0.8\%\) and \(\delta\) RMSE is statistically unchanged. A \(\delta\)-information-optimal lottery design --- pitting, e.g., the certain intermediate consequence against a 50/50 mix of the extremes to triangulate each \(\delta_k\) --- is named as future work (§\hyperref[sec-disc-limitations]{8.5}).
\addcontentsline{toc}{subsection}{6.3 Prior predictive analysis}
On the combined (uncertain + risky) prior predictive, the SEU-maximizer rate again covers the full sensitivity range. Because m_1 assumes a shared \(\alpha\) across blocks (§\hyperref[sec-m1-properties]{5.3}), one might want to confirm that risky and uncertain choices are mutually compatible with a single \(\alpha\); but that is a claim about the posterior, not the prior. A genuine check would require a posterior-predictive block comparison, which we flag as the appropriate diagnostic for the single-\(\alpha\) caveat (§\hyperref[sec-m1-alpha-recovery]{6.4.1}) rather than claim to have performed here.
\addcontentsline{toc}{subsection}{6.4 Parameter recovery --- what the matched comparison actually shows}
\addcontentsline{toc}{subsubsection}{6.4.1 \(\alpha\): no per-choice advantage from the risky block}
At matched total choice count (B vs C), swapping in the \(\beta\)-free risky block yields no \(\alpha\)-RMSE improvement: \(-8.1\%\), bootstrap 90% CI \([-22.0, +6.8]\) --- the point estimate is slightly worse (it favors m_0) and the interval straddles zero. The CI-width measure agrees (paired median \(-1.4\%\), improved in only 44% of iterations; Wilcoxon signed-rank \(p \approx 0.39\)). Here and below the Wilcoxon signed-rank test is a paired, distribution-free test of whether the per-iteration paired differences between conditions are centered at zero; we report it alongside the bootstrap CI because it makes no normality assumption about those differences and a small \(p\) would indicate a systematic shift in one direction across iterations.
The informative control is A\(\to\)B: doubling the uncertain block alone cuts \(\alpha\) RMSE by 27.1% (90% CI \([+15.0, +38.5]\); CI-width paired median \(+14.0\%\), improved in 78% of iterations). Read together, the two contrasts show that, in this matched design, \(\alpha\) precision at realistic \(n\) tracks data quantity rather than the type of choice: the \(\beta\)-free risky block confers no detected per-choice advantage --- the B\(\to\)C interval straddles zero, so a modest risky-block advantage remains compatible with the data --- whereas more uncertain choices sharpen \(\alpha\) directly. This is a statement about the realized, unoptimized design, not a general claim that choice type cannot matter. We draw out the methodological reading in §\hyperref[sec-m1-lesson]{6.4.4}.
\addcontentsline{toc}{subsubsection}{6.4.2 \(\delta\): the payoff is essentially nil}
At matched choice count (B vs C), m_1 narrows the \(\delta\) CI width by only \(\approx 0.8\%\) (paired-iteration median; bootstrap 90% CI \([0.6, 1.2]\); narrower in 72% of iterations; Wilcoxon signed-rank \(p \approx 1.2\times10^{-6}\)) and leaves \(\delta\) RMSE statistically unchanged (\(0.3\%\), 90% CI \([-4.0, +4.5]\)). The CI-width effect is real but practically negligible. The reading: \(\delta\) is identifiable in m_1 (§\hyperref[sec-delta-id]{5.5}, hence estimable in the limit) but not precisely estimable at realistic sample sizes in this design.
\addcontentsline{toc}{subsubsection}{6.4.3 \(\beta\)}
\(\beta\) recovery in m_1 is improved relative to m_0 but remains subject to the additive row-shift gauge (§\hyperref[sec-notation]{2.6}; Appendix B.4); summaries target contrasts and induced \(\bm{\psi}\), not absolute row levels.
\addcontentsline{toc}{subsubsection}{6.4.4 The methodological lesson --- two phenomena}
The matched comparison yields two separable substantive lessons, which we keep apart.
(i) For \(\delta\) --- the canonical lesson, sharpened. The principle is that identifiability (hence estimability in the limit) does not entail precise estimability at realistic sample sizes. \(\delta\) is identifiable in m_1 (§\hyperref[sec-delta-id]{5.5}), but the information about \(\delta\) contributed by 25 risky choices at moderate \(\alpha\) is small: the matched B\(\to\)C comparison shrinks the \(\delta\) CI width by \(\approx 0.8\%\) and moves \(\delta\) RMSE not at all. With unboundedly many risky choices \(\delta\) would be recovered exactly; with the design's 25 it is not. In-principle identification buys essentially nothing at the design sample size --- the identifiability-versus-estimability gap in its starkest form.
(ii) For \(\alpha\) --- finite-\(n\) precision tracks quantity, not type. One might expect the m_1 \(\alpha\) gain to be a per-choice advantage of the \(\beta\)-free risky block; at \(n = 100\) there is no such advantage (B\(\to\)C null). What sharpens \(\alpha\) is simply more data of the same kind (A\(\to\)B: a \(27.1\%\) RMSE reduction). The correct lesson is not “a special structure sharpens an already-identified parameter”; it is that in-principle identifiability tells you nothing about where finite-\(n\) precision comes from --- here, \emph{in this design}, it comes from quantity, not type (the matched contrast detects no risky-block advantage but does not rule out a smaller one). Phenomena (i) and (ii) are distinct.
Figure (ref) summarizes the three matched conditions visually.
\addcontentsline{toc}{subsection}{6.5 SBC for m_1}
Rank histograms and ECDF-difference plots for \(\alpha\), \(\beta\), \(\delta\) are consistent with uniformity. We do not claim an m_0\(\to\)m_1 \(\delta\)-calibration improvement (e.g., a flatter histogram): marginal SBC is uniform for \(\delta\) in both models, because the relevant \((\beta,\delta)\) weakness is in the joint posterior (the marginal-SBC demarcation again, §\hyperref[sec-m0-sbc]{4.4}). Figure (ref) makes this concrete for \(\delta\): the empirical-CDF diagnostic keeps both models inside the same simultaneous band. The appropriate diagnostic is a joint or projection-based rank statistic (§\hyperref[sec-discussion]{8.3}). The usual single-chain SBC caveats apply: thinning, single-chain Monte-Carlo error, and sample-size driven rank-resolution all bound what the histograms can detect.
\FloatBarrier
\addcontentsline{toc}{section}{Illustrative Application: SEU Sensitivity in LLM Decisions}
\addcontentsline{toc}{subsection}{7.1 Why include an application here}
The methodology of §§\hyperref[sec-abstract-model]{2}--\hyperref[sec-m1-implementation]{6} is only useful insofar as it does real evaluative work on real choice data. This section runs the full workflow end-to-end on a \(2\times2\) factorial design --- two LLMs crossed with two task families --- to demonstrate three things: (i) the workflow scales from the 25--100 simulated choices per condition of the §§\hyperref[sec-m0-implementation]{4}/\hyperref[sec-m1-implementation]{6} studies to \(N \approx 300\) real-LLM choices per condition without redesign; (ii) prior recalibration is a routine, principled per-study step rather than a redesign of the model; and (iii) the framework registers a structured comparative effect when the posterior supports one and withholds support when it does not --- where “withholds support” means inconclusive at the achievable resolution, not an established zero (§\hyperref[sec-app-claude-insurance]{7.5.2}).
\addcontentsline{toc}{subsection}{7.2 The \(2\times2\) design}
LLMs. GPT-4o (OpenAI) and Claude 3.5 Sonnet (Anthropic).
Task families. (i) Insurance claims triage over a 30-claim pool, with \(K = 3\) consequences --- the three investigator-agreement outcomes (neither, one, or both investigators agree the claim warrants investigation); and (ii) Ellsberg-style urn gambles over a 30-gamble pool, with \(K = 4\) consequences (\$0, \$1, \$2, \$3 payouts), organized into three ambiguity tiers (Tier 1 unambiguous, Tier 2 moderately ambiguous, Tier 3 high-ambiguity in the spirit of Ellsberg (1961)). In both tasks the alternatives presented in a problem are the pool items themselves (claims or gambles). Each task has a single fixed design, built once and shared across all temperature conditions; within that design the menu size of each problem was drawn at random from \(N_m \in \{2, 3, 4\}\).
Per cell. Five sampling-temperature conditions; \(M \approx 100\) base problems \(\times\) 3 position-counterbalanced presentations \(= 300\) choices per condition (constants in Appendix D.5). Temperature grids differ by provider: GPT-4o \(\{0.0, 0.3, 0.7, 1.0, 1.5\}\); Claude \(\{0.0, 0.2, 0.5, 0.8, 1.0\}\) (Anthropic-API range constraint). The unequal grids complicate direct slope-magnitude comparison across LLMs, though not the qualitative reading; we flag this again at §\hyperref[sec-app-results]{7.5.2a}.
Feature pipeline. Both task families use the same two-stage construction. Insurance: a per-claim LLM assessment, then choice over those assessments; assessment text embedded via text-embedding-3-small (embedding turns each assessment's text into a numeric vector), then reduced by pooled PCA (a linear projection onto the leading directions of variation) to \(D = 32\) features --- a compromise between explained variance and keeping the fit tractable. Ellsberg: the identical pipeline applied to gambles --- a per-gamble LLM assessment of each urn (a short free-text analysis of the payoff probabilities and ambiguity), then choice over those assessments, embedded and pooled-PCA-projected to \(D = 32\). The objective urn composition is not entered into the model directly: in both tasks the feature vector \(w_r\) is an embedding of an LLM-generated assessment, and the choice is made over those assessments. Both applications therefore measure the \emph{whole assessment-and-choice pipeline}, not the choice stage in isolation; the two cells differ in the decision domain (claims vs. gambles) and in \(K\) (3 vs. 4), not in which stages the temperature lever touches (we return to this at §\hyperref[sec-app-limits]{7.6.5}). Position counterbalancing addresses a position-bias problem identified in preliminary elicitation runs; unparseable responses are recorded as missing rather than coerced to a default, and missing rates are negligible across all 20 conditions.
One dimensional fact about this pipeline deserves explicit acknowledgment: with \(D = 32\) features and only \(R = 30\) distinct items, every realized feature matrix has full row rank (rank 30, verified for all 20 conditions; spikes/report_design_diagnostics_spike.py). The linear map \(\beta\) can therefore assign an arbitrary pattern of belief weights across the 30 items --- the feature construction imposes no cross-item restriction of its own, and the discipline on \(\beta\) comes from its prior rather than from feature-space parsimony. None of the comparisons below rest on \(\beta\) being identified (§\hyperref[sec-bd-weak]{3.4}); we state the rank fact so the \(D = 32\) choice is not misread as a binding structural constraint.
\addcontentsline{toc}{subsection}{7.3 The model fit (m_01 / m_02)}
The fitted model is structurally identical to m_0 of §\hyperref[sec-m0-implementation]{4} but with a \(\mathrm{Lognormal}(\mu, \sigma)\) prior on \(\alpha\) calibrated to each application's design and consequence space: m_01 for the insurance task (\(K = 3\)) and m_02 for the Ellsberg task (\(K = 4\)). The two programs differ in the consequence count \(K\) --- and hence in the dimensions of \(\upsilon\), \(\delta\), and the per-alternative belief objects \(\psi\) --- as well as in the \(\alpha\) prior; the likelihood form is identical.
Prior calibration anchor. Why does the \(K = 4\) Ellsberg task need a different \(\alpha\) prior than the \(K = 3\) insurance task? Not because random choice differs --- the menu size is the same --- but because with four consequences a larger \(\alpha\) is empirically required to reach the same prior-implied rate of good choice; the calibration below determines it directly rather than from a closed-form softmax argument. The calibration target is the prior predictive SEU-maximizer selection rate --- the fraction of prior-simulated agents that pick the SEU-maximal available alternative, obtained by drawing choice problems from the design and parameters from the prior and tabulating the outcome. At the application's \(R\) and \(K\) the foundational \(\mathrm{Lognormal}(0,1)\) prior of §\hyperref[sec-m0-spec]{4.1} places most of its mass on near-random sensitivities, so the prior predictive concentrates well below the SEU-max rates we take to be plausible for this application --- a range we posit rather than derive, and whose influence on the findings is examined in §\hyperref[sec-app-prior-sensitivity]{7.6.6}. (Numerical overflow in \(\exp(\alpha\eta)\) is handled by the softmax implementation and is not the reason for recalibration --- indeed the calibrated priors place more mass at large \(\alpha\).) The \(76\)--\(78\%\) target places the prior mode in the broad interior between chance and determinism: informative enough to move mass off the near-random end of the range that dominates under \(\mathrm{Lognormal}(0,1)\), yet loose enough --- a wide 90% prior interval on \(\alpha\) --- to let the likelihood move the posterior. A grid search over twelve Lognormal hyperparameter pairs scans \((\mu, \sigma)\) to hit this interior target, selecting \(\mathrm{Lognormal}(3.0, 0.75)\) for the insurance task (\(K = 3\)): median \(\alpha \approx 20\), 90% interval \(\approx [5.5, 67]\), implied SEU-max rate \(\approx 78\%\). For the Ellsberg task (\(K = 4\)) the prior is recalibrated by the same procedure, selecting \(\mathrm{Lognormal}(3.5, 0.75)\) (median \(\alpha \approx 33\), 90% interval \(\approx [10, 124]\), implied SEU-max rate \(\approx 76\%\)). (The 90% intervals quoted here are the empirical 5%--95% quantiles of the 200 \(\alpha\) draws used in the grid search, which is why they differ slightly from the analytic lognormal quantiles.) The recalibration is needed not because the random-choice baseline changes --- that baseline is set by the menu size, not the consequence count \(K\), and the menu size \(N_m \in \{2, 3, 4\}\) is identical across the two tasks, so a uniform-random chooser selects the SEU-maximizer with the same \(1/N_m\) probability either way --- but because the larger consequence space (\(K = 4\)) changes the distribution of expected-utility contrasts across alternatives, so the same grid search selects a higher \(\alpha\) to reach the same prior-implied SEU-max rate --- a direction we read off the calibration empirically rather than from a closed-form softmax argument.
\addcontentsline{toc}{subsection}{7.4 Validation at the application's scale}
The workflow of §§\hyperref[sec-m0-prior]{4.2}--\hyperref[sec-m0-sbc]{4.4} is re-run at the application's scale (\(M = 300\), \(K = 3\) or \(4\), \(D = 32\), \(R = 30\), 20 recovery iterations per cell).
Parameter recovery for \(\alpha\). Anchored on relative metrics, because true \(\alpha\) lies on a wide multiplicative scale under the calibrated prior: relative bias within \(\pm 10\%\) of the mean true value, relative RMSE well below 25%, and 90% CI coverage at nominal. With 20 iterations per cell these recovery checks are reassurance against gross miscalibration rather than precise coverage estimates --- the binomial uncertainty on a coverage proportion estimated from 20 draws is wide (a standard error of roughly 7 percentage points at nominal 90%). As throughout, \((\beta, \delta)\) recovery remains weak --- the §\hyperref[sec-m0-identifiability]{3} picture is unchanged --- but the cross-condition comparison is a claim about \(\alpha\), and \(\alpha\) is fit for purpose.
SBC for \(\alpha\). SBC (simulation-based calibration, defined in §\hyperref[sec-m0-sbc]{4.4}) validates the (prior, likelihood, sampler) triple and is conditionally independent of which agent generated the held-out empirical data.
MCMC diagnostics. The split-\(\hat{R}\) statistic compares between- and within-chain variance and should sit at or below \(1.01\) for a well-mixed chain; the effective sample size (ESS) reports how many independent draws the correlated chain is worth. Across all 20 factorial fits, \(\alpha\) --- the parameter the §\hyperref[sec-app-results]{7.5} temperature analysis rests on --- is never \(\hat{R}\)-flagged, and ESS is satisfactory in every fit, so the per-condition \(\alpha\) posteriors are a sound basis for the cross-condition comparison. Divergent transitions are negligible (\(\leq 0.15\%\) per fit, \(\leq 6/4000\); 34 total), though spread across 12/20 fits in both tasks and both providers rather than confined to the highest GPT-4o temperatures. \(\hat{R} > 1.01\) appears in 4/20 fits as marginal exceedances on a minority of high-dimensional components, confined to the weakly-informed nuisance parameters --- \(\beta\), \(\delta\), and the global length-\(K\) utility vector \(\upsilon\) (whose scale is fixed by the endpoint convention but whose interior levels are only weakly informed, tracking \(\delta\)) --- together with the per-trial latents \(\eta\) and \(\psi\) in the harder \(K = 4\) GPT-4o \(\times\) Ellsberg fits (plus one Claude \(\times\) insurance fit, \(\eta\) only). This corroborates the §\hyperref[sec-bd-weak]{3.4}/§\hyperref[sec-m1-recovery]{6.4} weakly-informed \((\beta,\delta)\) finding --- the poorly-mixing parameters are exactly the ones recovery flags as weakly informed --- and is not a defect in \(\alpha\) inference. As a direct check, all 20 conditions were refit under stricter sampler settings (\(\mathrm{adapt\_delta} = 0.99\), maximum tree depth \(12\), warmup \(2000\); seed 42): the divergences vanish entirely (0 across all 20 fits, with no tree-depth saturations), the per-condition \(\alpha\) posterior medians move by at most \(0.06\) committed-posterior standard deviations, and the four §\hyperref[sec-app-results]{7.5} temperature slopes are unchanged to within Monte-Carlo noise (e.g., GPT-4o \(\times\) insurance \(-24.6 \to -24.8\), \(P(\beta_{\text{temp}} < 0)\) \(0.99 \to 0.99\); Claude \(\times\) Ellsberg \(-15.0 \to -14.7\), \(0.77 \to 0.77\)) --- confirming the divergences were benign step-size artifacts rather than symptoms of a biased posterior (Appendix \hyperref[sec-d2]{D.2}).
Posterior predictive checks. Three complementary summaries (log-likelihood, modal-choice frequency, mean predicted probability of the chosen alternative) yield posterior predictive \(p\)-values in \([0.32, 0.66]\) across all 20 fits (60/60 in band, mean \(\approx 0.46\)) --- no evidence of systematic misfit. PPC adequacy is a necessary (not sufficient) condition for the per-condition \(\alpha\) posteriors to be a credible basis for the cross-condition comparison; absent it, the comparison should be set aside.
\addcontentsline{toc}{subsection}{7.5 Results --- the \(2\times2\) cross-LLM \(\times\) cross-task picture}
\addcontentsline{toc}{subsubsection}{7.5.1 GPT-4o \(\times\) insurance}
Posterior medians of \(\alpha\) are highest at \(T = 0.0\) and lowest at \(T = 1.5\), intermediate temperatures monotone-ish in between. The global slope \(\Delta\alpha/\Delta T\) has median \(\approx -24.6\), 90% CI \(\approx [-52.4, -6.7]\), \(P(\text{slope} < 0) \approx 0.99\). The probability of strict monotonic decrease across all five levels is only \(\approx 0.12\) (driven by overlap between \(T = 0.3\) and \(T = 0.7\)); collapsing those two raises it to \(\approx 0.38\). The headline directional claim is well-supported; fine-grained adjacent-step orderings are not.
\addcontentsline{toc}{subsubsection}{7.5.2 Claude \(\times\) insurance}
The framework does not detect a monotonic \(\alpha\)--temperature pattern (note the careful phrasing: does not detect, not establishes the absence of). Posterior medians are \(\approx \{71, 53, 73, 71, 55\}\) at \(\{0.0, 0.2, 0.5, 0.8, 1.0\}\). Global slope: median \(\approx -2.9\), 90% CI \(\approx [-42.9, 30.8]\), \(P(\text{slope} < 0) \approx 0.56\); \(P(\text{strict monotone decrease}) < 0.01\). The oscillation in the medians is consistent with posterior noise around a roughly flat function, not a substantive non-monotonic response.
\addcontentsline{toc}{subsubsection}{7.5.2a Cross-LLM comparison}
A formal cross-study comparison gives \(P(\text{GPT-4o slope} < \text{Claude slope}) = 0.817\) on the full grids; restricting GPT-4o to Claude's \(T \leq 1.0\) ceiling (re-summarizing with \(T = 1.5\) dropped) gives \(0.816\) --- the directional LLM contrast is robust to the unequal-grid confound (GPT-4o spans \(T \in [0.0, 1.5]\); Claude spans \([0.0, 1.0]\)). We report the Claude-grid-restricted number alongside the full-grid one and make the qualitative pattern (monotone-decline vs. no-monotone) the headline, not the numeric inequality. The between-LLM probabilities (\(\sim 0.82\)) are lower than GPT-4o's own within-cell \(P(\text{slope} < 0) > 0.98\) because they fold in the independent uncertainty of both cells; they answer a different question from either within-cell probability (Appendix D.6(4)). “Temperature” is not the same instrument across providers (§\hyperref[sec-app-limits]{7.6.2}(a)).
\addcontentsline{toc}{subsubsection}{7.5.3 GPT-4o \(\times\) Ellsberg}
The temperature--\(\alpha\) decline is reproduced on the Ellsberg task. \(\alpha\) medians \(\{110.4, 106.9, 99.5, 84.0, 52.2\}\) at \(T = \{0.0, 0.3, 0.7, 1.0, 1.5\}\) (90% CIs \([74.4, 167.2]\) / \([72.6, 163.8]\) / \([65.4, 154.6]\) / \([57.1, 126.1]\) / \([35.5, 80.3]\)); global slope \(\Delta\alpha/\Delta T\) median \(-38.4\), 90% CI \([-72.1, -10.0]\), \(P(\text{slope} < 0) \, 0.984\); \(P(\text{strict monotone} \downarrow) \, 0.090\) (a noisy decline, not step-wise). In behavioral terms, the fitted decline means the model selects the expected-utility-maximizing gamble progressively less often as temperature rises, drifting toward the menu-uniform choice of the §\hyperref[sec-three-properties]{2.3} low-\(\alpha\) limit without reaching it.
\addcontentsline{toc}{subsubsection}{7.5.4 Claude \(\times\) Ellsberg}
No monotonic \(\alpha\)--temperature pattern. \(\alpha\) medians \(\{85.4, 56.1, 82.3, 53.0, 66.4\}\) at \(T = \{0.0, 0.2, 0.5, 0.8, 1.0\}\) (90% CIs \([58.9, 127.2]\) / \([38.2, 84.1]\) / \([52.3, 135.3]\) / \([36.0, 80.5]\) / \([45.3, 99.8]\)) --- the medians oscillate rather than decline; global slope \(\Delta\alpha/\Delta T\) median \(-15.0\), 90% CI \([-52.2, 19.6]\), \(P(\text{slope} < 0) \, 0.766\); \(P(\text{strict monotone} \downarrow) \, 0.0085\). As with Claude \(\times\) insurance, this is does not detect, not establishes absence. We do not transfer the insurance minimum-detectable-effect number here --- the Ellsberg cell differs in \(K\), in the \(\alpha\) prior, and in its wider per-condition posteriors --- so we read this cell as an inconclusive null on its own diagnostics (oscillating, non-monotone medians and a slope CI spanning zero), the narrower claim the resolution discipline of §\hyperref[sec-app-claude-insurance]{7.5.2} licenses.
\addcontentsline{toc}{subsubsection}{7.5.5 The \(2\times2\) reading}
In this \(2\times2\), the descriptive pattern lines up with the LLM axis rather than the task axis: GPT-4o exhibits the temperature--\(\alpha\) relationship across both task families, Claude exhibits it in neither. But two of the four cells are inconclusive nulls rather than measured zeros, so the pattern is suggestive of an LLM-level difference, not an established LLM effect. With only two task families and two LLMs this is a pattern statement, not a factorial generalization (§\hyperref[sec-app-limits]{7.6.2}(d)). Figure (ref) presents it as a \(2\times2\) forest plot of the per-cell global-slope posteriors.
\addcontentsline{toc}{subsection}{7.6 What the application demonstrates --- and what it does not}
\addcontentsline{toc}{subsubsection}{7.6.1 What it demonstrates}
\addcontentsline{toc}{subsubsection}{7.6.2 What it does not demonstrate}
(a) That sampling temperature is a portable behavioral instrument. The Claude cells do not reproduce the GPT-4o pattern under the same task and choice model; but the minimum-detectable-effect analysis of §\hyperref[sec-app-claude-insurance]{7.5.2} shows that even a GPT-4o-sized slope would sit below the resolution floor of the Claude \(\times\) insurance cell --- the one cell for which an MDE was computed; the Claude \(\times\) Ellsberg cell is read as an inconclusive null on its own diagnostics --- so the non-reproduction is inconclusive: it does not establish that temperature behaves differently across providers, only that any Claude-side effect, if present, is below what these data can resolve. (b) That GPT-4o is “more rational” than Claude in any context-free sense. The supported reading of \(\alpha\) here is a within-design comparative one, not an absolute-rationality ranking. (c) Anything about ambiguity aversion in the Ellsberg study. The SEU+softmax model is, by assumption, of the wrong form to distinguish ambiguity-driven choice from EU-driven choice --- but not because ambiguity aversion must show up as a lower \(\alpha\). An ambiguity-averse agent can be perfectly accommodated by the model with a high \(\alpha\) and pessimistically distorted beliefs: the belief parameters \((\beta, \psi)\) are free to place extra weight on bad outcomes of ambiguous gambles, in which case ambiguity aversion is absorbed into the \emph{belief} side of the model rather than the sensitivity side. A departure that the belief side cannot mimic is instead absorbed into a lower \(\alpha\) (more randomness). Either way, genuine ambiguity attitude, distorted beliefs, and plain noise are not separable within this model class. Testing for ambiguity aversion requires a model with explicit ambiguity-attitude parameters --- for instance maxmin expected utility (Gilboa and Schmeidler 1989), which adds a parameter governing how heavily the agent weights the worst-case prior over an ambiguous event, or its \(\alpha\)-MEU generalization (Ghirardato, Maccheroni, and Marinacci 2004) (whose mixing weight, conventionally also written \(\alpha\), is unrelated to the sensitivity parameter \(\alpha\) of this paper); and pooling across the three ambiguity tiers as the application does would yield a mixture rather than identify a mechanism. Tier-stratified \(\alpha\) is named as future work, not a paper claim. \textbf{(d) That LLM identity is the only relevant axis in general.} With only two task families and two LLMs, the \(2\times2\) supports a pattern statement, not a factorial generalization.
\addcontentsline{toc}{subsubsection}{7.6.3 Construct-validity layering (reaffirmed)}
The three-layer reading guide of §\hyperref[sec-app-why]{7.1} --- (1) model adequacy, (2) comparative claims under shared design, (3) absolute claims about EU rationality, with only (1) and (2) supported --- is the single most important reading instruction for this section. The §\hyperref[sec-app-results]{7.5} cross-condition contrast is a layer-(2) claim; it is not a layer-(3) certification of either LLM's context-free rationality.
\addcontentsline{toc}{subsubsection}{7.6.4 Ellsberg: do not adjudicate the normative question}
The Ellsberg stimuli carry historical weight that does not reduce to the behavioral “ambiguity aversion” reading: alternative readings include Ellsberg's own normative defense (Ellsberg 1961, 2001) and Levi's reading as a violation of the completeness axiom rather than the sure-thing principle (Levi 1980, 1986). A study that takes SEU as its measurement device cannot use the resulting estimates to vindicate or refute SEU as a normative standard. (See Camerer and Weber (1992) and Trautmann and Kuilen (2015) for the broader behavioral-ambiguity literature.) The paper's contribution lies elsewhere.
\addcontentsline{toc}{subsubsection}{7.6.5 Both applications measure the whole pipeline}
Neither application isolates a single cognitive stage, and it is worth being precise about what \(\alpha\) actually measures. Each trial --- insurance or Ellsberg --- is produced by a composed system: a first LLM call assesses the stimulus (a free-text claim, or an urn gamble) in natural language, an embedding model maps the resulting text to a vector, a PCA step reduces it to the feature space the choice model reads, and a second LLM call then chooses among the assessed alternatives (§\hyperref[sec-app-design]{7.2}). The fitted sensitivity \(\alpha\) is a property of that whole assembly --- the assessment and choice calls together with the deterministic feature pipeline --- and not of the choice stage taken on its own. The estimand is therefore the SEU sensitivity of the pipeline as configured, and the temperature lever acts on every LLM-mediated component of it at once. This is the appropriate object of study for the methodological claim --- the workflow recovers a well-calibrated \(\alpha\) for a realistic end-to-end system --- but it means that a within-cell change in \(\alpha\) cannot, by itself, be attributed to any one stage. The GPT-4o monotonicity is thus consistent with temperature affecting (i) the assessment stage, (ii) the choice stage, or (iii) both. A further caveat concerns the feature geometry itself: the belief-feature vectors \(w_r\) are produced by text-embedding-3-small for both providers, so Claude's choices are read through an external embedding space that need not preserve the distinctions driving its selections. A within-cell change in \(\alpha\) could therefore reflect the assessment stage, the choice stage, or upstream embedding/PCA mismatch. Fixed-feature and alternative-embedding robustness checks --- holding features constant across temperature, or re-embedding with a provider-matched model --- are the natural way to separate these, and we leave them to future work.
The Ellsberg cell does not remove the assessment layer. It applies the same assess--embed--choose construction to a different decision domain (\(K = 4\) monetary gambles in place of \(K = 3\) claim-triage outcomes); the objective urn composition enters only through the LLM's free-text assessment, never as a structural input to the model. It is therefore a cross-domain replication rather than a stage-isolation design. What it establishes is correspondingly more modest, but still useful: that the GPT-4o temperature pattern is not an artifact of the particular insurance content or its assessment prompts, since a comparable effect reappears when the same pipeline is pointed at a structurally different task. It does not decompose the effect by stage --- because the domain and the assessment prompt change together with everything else, the Ellsberg cell cannot say whether the insurance effect originates in the assessment stage, the choice stage, or both. The two cells together show that the effect travels across domains; isolating its locus within the pipeline would require dedicated stage-ablation designs (assessment-only versus choice-only manipulations) that lie beyond this paper's scope.
\addcontentsline{toc}{subsubsection}{7.6.6 The cross-condition findings are robust to the \(\alpha\) prior}
Because each cell's headline object is a temperature-slope of a parameter carrying an informative, prior-predictive-calibrated prior (§\hyperref[sec-app-model]{7.3}), it is fair to ask whether the slope findings are an artifact of that prior. They are not. We re-estimated every condition of all four cells under three alternative \(\alpha\) priors --- shifting the lognormal location by \(\pm 0.5\) (\(\mathrm{Lognormal}(\mu_{\text{base}} \mp 0.5,\, 0.75)\)) and widening the scale to \(\mathrm{Lognormal}(\mu_{\text{base}},\, 1.25)\) --- holding the choice data and the feature matrices fixed (the committed per-condition Stan data), and recomputing each cell's slope and \(P(\text{slope} < 0)\) with the same draw-wise population-OLS functional. Across all four priors the qualitative reading of every cell is unchanged (Figure (ref)): both GPT-4o cells retain a clearly negative slope --- insurance median in \([-30.7, -22.6]\) with \(P(\text{slope}<0) \in [0.987, 0.991]\), Ellsberg median in \([-44.8, -36.8]\) with \(P(\text{slope}<0) \in [0.984, 0.994]\) --- while both Claude cells remain inconclusive nulls, insurance median in \([-4.2, -2.7]\) with \(P(\text{slope}<0) \in [0.55, 0.58]\) and Ellsberg median in \([-17.0, -14.0]\) with \(P(\text{slope}<0) \in [0.75, 0.77]\). The wider-prior variant moves the slope magnitudes the most, as expected, but never flips a sign or crosses a resolution threshold. The prior calibrates the scale on which \(\alpha\) is read; it does not manufacture the cross-condition contrasts. (Sweep code: spikes/report_prior_sensitivity_spike.py; claims-ledger row C17.)
\addcontentsline{toc}{subsection}{7.7 What the application motivates for follow-up work}
Design-induced cross-condition correlation. The five temperature conditions in each cell draw from a single fixed set of \(R = 30\) alternatives --- the same claims (or gambles) at every temperature, though their LLM assessments and embeddings are regenerated independently at each temperature (§\hyperref[sec-app-design]{7.2}). Whatever is idiosyncratic about that particular set --- its embedding geometry, EU spread, or typical best-vs-second-best gap --- is a property of this pool rather than of the population of pools the design might have drawn, so the cell's contrasts speak to the realized pool and need not transport to another. Two concerns follow that independent per-condition m_01 fits do not address: the cross-condition contrasts may not generalize beyond the realized pool, and the pool may interact with temperature --- a pool-by-temperature effect that a single-pool design cannot separate from the main effect. The principled fix is the hierarchical extension h_m01: \(\log \alpha_c = \gamma_0 + \gamma_1 T_c + \varepsilon_c\) with a small cell-level random effect \(\varepsilon_c\), which estimates the temperature--sensitivity relationship as a single regression slope with one calibrated uncertainty statement and partially pools the five conditions; its multi-pool generalization (the companion paper) additionally lets the item pool be modeled as a source of variation rather than treated as a fixed idiosyncrasy.
Companion paper. A planned alignment-study companion paper (multi-LLM \(\times\) multi-prompt factorial) will introduce h_m01 as the analysis vehicle and apply it to the alignment-study data. That paper is out of scope here; the present paper references it once, as the natural next step for which h_m01 was developed.
Tier-stratified \(\alpha\) for the Ellsberg studies, and a non-SEU comparator (e.g., a small \(\alpha\)-MEU instrument) that could be paired with the SEU instrument to triangulate ambiguity-driven vs. EU-driven choice, are flagged as future work in the application reports and remain out of scope here.
\addcontentsline{toc}{section}{Discussion}
\addcontentsline{toc}{subsection}{8.1 What “sensitivity to SEU maximization” means, restated}
Sensitivity to SEU maximization is the disposition of an agent committed to SEU maximization to act in accordance with that commitment. The model decomposes choice behavior into three conceptually distinct objects: (i) the agent's beliefs (\(\beta\), hence the subjective probabilities \(\psi\)); (ii) the agent's utilities (\(\delta\), hence the ordered utilities \(\bm{\upsilon}\)); and (iii) the agent's sensitivity \(\alpha\) to the expected-utility ranking those beliefs and utilities induce. What the agent is committed to is the SEU standard itself, and that standard places requirements on how the agent's beliefs and values --- represented by a subjective probability and a utility --- bear on choice; beliefs and utilities are those inputs, while \(\alpha\) is how reliably the agent's choices track the expected-utility ranking the standard derives from them. The three softmax-choice properties of §\hyperref[sec-three-properties]{2.3} --- monotonicity in \(\alpha\), the optimization limit, and the uniform-choice limit --- are exactly what license reading \(\alpha\) as a graded measure that runs from near-uniform choice (\(\alpha \to 0\)) to deterministic SEU maximization (\(\alpha \to \infty\)).
This reading returns to Isaac Levi's distinction between an agent's commitment to a standard and the agent's performance relative to it (Levi 1980, chap. 1), introduced in §\hyperref[sec-conceptual-payoff]{2.5}. One can be committed to a standard one fails to live up to: most of us are committed to the laws of arithmetic yet occasionally make calculation errors, and those errors are lapses of performance rather than rejections of the standard. Taking SEU theory to specify the agent's normative commitment, \(\alpha\) is exactly the agent's tendency to perform in accordance with that commitment.
Two points are worth keeping in mind regarding our borrowing of Levi's distinction. First, Levi was no defender of SEU as the standard of rational choice; on the contrary, he developed an influential generalization of it that admits indeterminate (imprecise) probabilities and utilities (Levi 1986), and his substantive decision theory therefore falls outside the real-valued template of this paper (§\hyperref[sec-seu-standard]{1.4}). What we borrow is his meta-level distinction between commitment and performance, not his account of the standard itself. Second, that distinction is of interest even once SEU is granted as the relevant standard. Consider the SEU violations documented by Kahneman and Tversky (Kahneman and Tversky 1979; Tversky and Kahneman 1992): do they reveal failures of performance by agents committed to SEU, or an absence of commitment to SEU as a standard? Our framework does not settle that question, but it does make it tractable: under the assumption that the agent is committed to SEU, \(\alpha\) estimates how reliably the agent's choices conform to that commitment.
\addcontentsline{toc}{subsection}{8.2 The methodological role of identifiability analysis}
Identifiability is a property of the likelihood: it tells us which parameters could in principle be recovered from choice data. It does not tell us how much data, or how good a design, is needed to recover them precisely. The governing principle is:
The m_0 / m_1 contrast illustrates both sides of this distinction, and the paper is careful to keep the two sides apart (§\hyperref[sec-m1-lesson]{6.4.4}):
\addcontentsline{toc}{subsection}{8.3 The methodological role of computational validation}
Within the Bayesian workflow this paper adopts, three checks are standard practice, and the paper applies all three. Prior predictive checks (§\hyperref[sec-m0-prior]{4.2}, §\hyperref[sec-m1-prior]{6.3}) expose what a prior implies before any data are seen; in this setting the SEU-maximizer-rate statistic makes the \(\alpha\) prior's behavioral content legible and turns per-study prior recalibration into a principled, repeatable step (§\hyperref[sec-app-model]{7.3}). Parameter recovery and simulation-based calibration (SBC) together test whether the posterior recovers known truths --- recovery for bias, precision, and interval coverage; SBC for calibration of the whole (prior, likelihood, sampler) triple.
A specific caution emerges from our results, which we name so it is citable rather than discursive:
The positive contribution of SBC here --- sampler/implementation validation, per-parameter marginal calibration, and the demarcation itself --- is stated in full in §\hyperref[sec-m0-sbc]{4.4}. Two follow-ups should be kept distinct. A joint or projection-based rank statistic would extend SBC's reach from marginal to joint calibration --- a concrete next step for the SBC literature. But no rank statistic, marginal or joint, measures informativeness: whether a posterior has contracted enough to be useful is diagnosed by recovery, contraction, and interval-width summaries (§\hyperref[sec-m0-recovery]{4.3}, §\hyperref[sec-app-validation]{7.4}), not by calibration checks. The paper's \((\beta,\delta)\) finding is of the second kind, which is why no SBC variant --- however joint --- would have surfaced it on its own.
\addcontentsline{toc}{subsection}{8.4 Why \(\alpha\) is the primary quantity of interest}
A central methodological claim of the paper is that \(\alpha\) can be measured precisely even when \((\beta, \delta)\) cannot be --- a separation the recovery study and application deliver empirically, under the paper's priors and designs, rather than one guaranteed by the conditional identification proposition alone. The mechanism is that beliefs and utilities enter the choice probabilities only through the expected-utility vector \(\eta\), and the log-odds scale with differences in \(\eta\) at rate \(\alpha\) (§\hyperref[sec-m0-link]{3.5}). Proposition 3.1 identifies \(\alpha\) only conditional on \(\eta\); in the uncertain-choice model \(\eta\) is itself generated by the unknown \((\beta, \delta)\), and near the linear region of the belief softmax a change in the spread of \(\eta\) can trade off with \(\alpha\) in the observed log-odds. What makes \(\alpha\) sharply estimable in practice is therefore an empirical fact about the calibrated prior and the realized designs --- the posterior concentrates the \(\alpha\)-bearing log-odds scale while leaving the \((\beta, \delta)\) trade-off broad (§\hyperref[sec-bd-weak]{3.4}) --- not a structural guarantee read off the conditional proposition. The §\hyperref[sec-application]{7} application instantiates the claim: the cross-condition comparison is a claim about \(\alpha\), \(\alpha\) is sharply and calibratedly recovered in m_0 and the calibrated-prior variants m_01/m_02, and the weak \emph{informativeness} of \((\beta, \delta)\) does not undermine that comparison. This comparative discipline is also the paper's answer to the cross-context comparability critiques of softmax-precision parameters (Wilcox 2011; Apesteguia and Ballester 2018): the \(\bm{\upsilon}\)-endpoint convention (§\hyperref[sec-notation]{2.6}) fixes one global utility scale shared by every problem in a design --- which supplies the within-context normalization the contextual-utility critique asks for only because all problems draw on the same consequence space (§\hyperref[sec-softmax-rule]{2.2}) --- and the supported claims compare \(\alpha\) only within a fixed design.
\addcontentsline{toc}{subsection}{8.5 Limitations and extensions}
Precise \(\delta\) estimation. The modest finite-sample \(\delta\) gain (§\hyperref[sec-m1-delta-recovery]{6.4.2}) is a limitation, not a barrier in principle. Three approaches could improve it --- a \(\delta\)-information-optimal lottery design (e.g., triangulating contrasts that pit the certain intermediate consequence against a \(50/50\) mix of the extremes), substantially larger samples, and prior regularization on \(\delta\) --- but they are not independent: a \(\delta\)-optimal design helps less when \(\alpha\) is small (a shallow softmax at moderate \(\alpha\) dampens the same contrasts), so precise \(\delta\) recovery likely requires improvements on several of these fronts at once rather than any one of them alone. A \(\delta\)-optimal lottery-design study is named here as future work rather than attempted in this paper.
The single-\(\alpha\) assumption. A single \(\alpha\) governing both the uncertain and risky blocks (§\hyperref[sec-m1-properties]{5.3}) is substantive and testable. It is essential to the §\hyperref[sec-m1-recovery]{6.4} quantity-versus-type reading, which is why that reading is stated as explicitly conditional on it. The natural relaxation is a block-specific model fitting separate \(\alpha_{\mathrm{unc}}, \alpha_{\mathrm{risky}}\) (our internal m_2); it is named here, not pursued.
Menu size and sensitivity. The model places no restriction on the number of alternatives in a decision problem: the available set \(\mathcal{A}_m\) may be of any size, the softmax of §\hyperref[sec-softmax-rule]{2.2} is taken over it, and the Stan implementation already accepts a per-problem alternative count (Appendix C). The classical prospect-theory paradigm and its large-scale replication (Kahneman and Tversky 1979; Tversky and Kahneman 1992; Ruggeri et al. 2020) restricted attention to pairwise choice, appropriately for their purpose; the present formulation does not, which makes the relationship between sensitivity and menu size an estimable object rather than a fixed feature of the design. Treating menu size \(N_m\) as a covariate on \(\log \alpha\) --- e.g. \(\log \alpha = \gamma_0 + \gamma_1 N_m\) in the hierarchical vehicle above --- would let one ask whether sensitivity rises or falls as alternatives are added; one illustrative, model-agnostic hypothesis is that a larger menu raises the cognitive load on the decision maker and thereby lowers \(\alpha\), but the sign is left open. This is named, not pursued: the §\hyperref[sec-application]{7} designs randomize and position-counterbalance \(N_m\) and fit a single \(\alpha\) pooled across menu sizes, so estimating a menu-size effect cleanly would require a purpose-built design that varies menu size as a factor rather than a nuisance.
Functional form. Linear-softmax belief formation (\(\psi_r = \mathrm{softmax}(\beta w_r)\)) is a strong functional form; extensions to nonlinear belief formation and richer feature maps are possible.
Hierarchical extensions. The §\hyperref[sec-application]{7} application motivates the hierarchical extension h_m01 to address the design-induced cross-condition structure of §\hyperref[sec-app-followup]{7.7}: because the five conditions reuse a single fixed pool of alternatives (with the LLM assessments and embeddings regenerated independently at each temperature), fitting m_01 independently per condition leaves two things on the table --- partial pooling of information across the five temperature conditions, and any treatment of the item pool as a modeled source of variation. The two are distinct and should not be conflated: partial pooling stabilizes the per-condition \(\alpha\) estimates and yields one calibrated slope statement, but pooling across temperature conditions that all reuse the same pool cannot, by itself, generalize the findings beyond that realized pool. Modelling \(\log \alpha_c = \gamma_0 + \gamma_1 T_c + \varepsilon_c\) with condition-level effects \(\varepsilon_c\) partially pools the five temperature conditions and turns between-condition contrasts into estimated regression effects on \(\log \alpha\). Because a single pool is reused across all conditions, however, h_m01 still cannot separate a pool \emph{main} effect (confounded with the intercept \(\gamma_0\)) or a pool-by-temperature interaction (absorbed into the slope \(\gamma_1\)); doing so --- and with it any generalization over item pools --- requires several distinct pools. That multi-pool generalization is the analysis vehicle of a planned alignment-study companion paper and is out of scope here.
Further application studies. A tier-stratified Ellsberg analysis, a non-SEU comparator instrument (e.g., a small \(\alpha\)-MEU model paired with the SEU instrument to triangulate ambiguity-driven versus EU-driven choice), and the alignment study itself are pursued in companion work and are out of scope here.
\addcontentsline{toc}{subsection}{8.6 Closing}
A precise definition of sensitivity, a formal identifiability analysis, a computational validation pipeline, and an end-to-end illustrative application together provide a transparent, reusable framework for procedurally evaluating decision makers --- human or machine --- against a stated rationality standard expressible as expected-value maximization, with SEU as the canonical case. The framework reports where a parameter is identifiable but not precisely estimable, where in-principle arguments fail to predict finite-sample behavior, where marginal calibration is blind to a jointly weakly-informed posterior, and where an instrument declines to support an effect its own diagnostics cannot resolve. Reporting these limits is part of what makes the framework usable as a measurement instrument rather than a procedure that returns an effect by construction.