Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
90,818 characters · 17 sections · 45 citation commands
The Innate Economic Preferences of Language Models
JEL Classification: C25, C45, D81, G11.
Language models (i.e. LLMs) are increasingly delegated authority over decisions that allocate real resources. They are used to negotiate on the behalf of a principal araujo2026how, to synthesize the information that determines the outcomes of lending eisfeldt2024ai, and to advise households on saving and investing during the life cycle desilva2026ai. Initially, language models gained attention for drafting text and meeting technical benchmarks. Now, especially with the introduction of AI agents, language models are resolving economic tradeoffs without direct input from a human principal. The economic question is no longer whether they are intelligent but whose preferences they impose. Importantly, a principal delegating an allocation problem to a language model cannot ex ante specify her economic preferences for every unforeseen decision it will encounter. imas2025agentic show experimentally that this gap is consequential. When many such delegated choices clear the same market, the risk attitude a model brings to an underspecified instruction stops being a private feature and shapes how resources are allocated, a transition shahidi2025coasean argues is imminent as agents transact on behalf of consumers. That attitude is the object we measure. We study not a model reasoning toward a specified mandate but the default preference it reveals when the instruction leaves the choice over economic tradeoffs open, and we ask whether that default is stable enough to be treated as a preference at all.
We argue that a language model capable of reasoning need not be a coherent chooser. Typically, a model is judged as capable if it produces the known correct answer to a factual question. Choosing among competing economic ends is different. When those ends conflict, no choice is correct in the abstract because the tradeoff among them is settled by preference rather than by analysis, and what a principal delegates is precisely the authority to choose. A model can therefore reason well and still choose in ways that reveal no stable preference at all. This lack of stability becomes an explicit economic liability when choices violate, for example, the axiom of transitivity. If a language model's choices cycle according to a standard money pump, a counterparty can use it to systematically extract wealth from the principal. This is the distinction that matters for delegation. Therefore, our central question follows. What preferences govern a language model's choices over economic tradeoffs, and do those choices satisfy the revealed-preference restrictions required for a stable utility interpretation? Our answer, for the risk tradeoffs we study, is that they do, within limits we make precise. When the instruction leaves the tradeoff open, every model brings a systematic and measurable default risk attitude to the choice. However, we find the models' preferences are not invariant to how the options are presented to them.
We reach this answer with a single measurement program built on restricting how a language model chooses. Specifically, we require the model to express its choice by emitting a single unit of output, called a token, as its revealed preference. The model selects that token through a fixed decoding rule that maps a vector of internal scores, its logits, into a probability over the available tokens. Our first contribution is to establish that this decoding rule admits an exact random utility representation, one in which the logits are the systematic utility index and an internal parameter called temperature is the scale of an additive extreme-value shock. The representation is precisely the conditional logit of mcfadden1972conditional, but its standing here is different. In the classical model, the systematic index is latent and must be inferred from choices under an assumed distribution for the unobserved shock. In a language model, the index is the logit vector itself, recorded directly from a single forward pass of the model, so the same quantity the discrete-choice econometrician estimates is here observed. The stochastic part of the representation imposed by temperature is only the decoding noise. Because the systematic index is observed, the identification of preferences is not based on the shock distribution that the classical estimator leans on.
Our experimental program that generated these insights has two steps. We first test whether a stable preference exists or whether the model's choices obey the revealed-preference restrictions required for a utility interpretation. Six diagnostics operationalize a set of standard revealed-preference axioms: completeness, reflexivity, monotonicity, transitivity, continuity, and independence of irrelevant alternatives. In the second step, we estimate preferences structurally in a controlled portfolio choice environment. In both steps, the model faces menus of assets, each described by an expected return and standard deviation, and selects one in a single forced choice, with the attributes varied exogenously across menus so that the return-risk tradeoff is entirely under the researcher's control. A maintained quadratic utility over the two attributes reduces the model's preference to a single risk-aversion parameter, which the estimator recovers. For open-weight models, for which internal computation can be readily inspected, we read the utility index directly, so we observe the utility-relevant signal itself rather than reconstructing it from choices as in a human experiment ludwig2025large. For proprietary frontier models such as those from OpenAI, Anthropic and Google, whose internal computation is hidden, we recover the same quantities from repeated sampled choices using standard maximum likelihood, exactly as one would with human subjects. The replication code we provide is a protocol we recommend for testing agents along these dimensions of rationality.
Our findings across models demonstrate their choices respect the economic content of a menu but not its irrelevant features. Monotonicity, continuity, and transitivity hold at or near perfect rationality, so the models honor mean-risk dominance, respond smoothly to gradual changes in return and risk, and rank assets consistently. However, they fail to respond in a perfectly rational manner to the presentation of irrelevant alternatives, which is a form of changing the context of the decision. Moving an option's position in a menu can change the value a model attaches to it, and adding a strictly dominated third asset can shift the strength of preference between two unchanged options, thus reflexivity and independence of irrelevant alternatives are the weakest diagnostics throughout. The economic consequence of this error is that a model's preference ordering remains relatively stable while its intensity of preference does not, and only near indifference does this instability become observable. Furthermore, our structural estimations reveal that every model we evaluate is risk averse, and risk aversion varies across the major AI labs (OpenAI, Anthropic, Alibaba, etc.). Thus, the same menu implies materially different chosen portfolios depending on which lab's model is asked.
Although these findings broadly support treating language models as agents with stable and measurable preferences, they may not be optimal for a principal's delegated task. This raises the question of whether a desired economic preference can be installed in a model rather than merely measured in it. We show that it can. By fine-tuning, the standard procedure for adjusting a model's weights toward a specified objective,\footnote{This fine-tuning exercise differs from preference-based alignment pipelines that learn from human comparisons or optimize policies against learned preference models christiano2017deep,ouyang2022training,rafailov2023direct. Related economic work fine-tunes language model agents toward explicit rational and moral preference structures in economic games and moral dilemmas lu2025aligning.} we write the maintained utility into a specific open-sourced language model at preset risk-preference targets, and the model's induced choices match risk tolerance targets with high precision. Fine-tuning is therefore a direct lever for inducing a specified economic preference, expressed through the utility function, into the agent that will act on it.
Our contribution sits among several literatures. The first studies language models as economic agents and behavioral subjects. A growing body treats models as simulated decision-makers that asks how closely their behavior follows that of people horton2023large,mei2024turing,park2023generative, akata2025playing, deploys them as synthetic respondents in applied research brand2023using, and measures their psychological traits directly serapio2025psychometric. Part of this work documents that model choices respond to surface features such as the order in which options are listed pezeshkpour2024large, a sensitivity that our reflexivity and invariance diagnostics formalize and quantify. bini2026behavioral study language model responses in preference-based tasks. Most of this literature describes behavior, whereas we estimate and test a preference. A second strand is the structural discrete choice and revealed-preference tradition on which our estimator rests afriat1967construction,mcfadden1972conditional,manski1977structure,varian1992microeconomic,berry1994estimating,train2009discrete,charness2013experimental,matvejka2015rational,demuynck2023computing, which we apply to the directly observed internals of a generative model rather than to choices alone. A third is the experimental estimation of risk preferences from choices over alternatives with known payoff distributions holt2002risk,choi2007consistency,cohen2007estimating,harrison2008risk,andersen2008eliciting,bruhin2010risk,apesteguia2018monotone, of which kim2024learning, liu2025evaluating and ouyang2024ethical are the nearest applications to language models. ellis2026state approach the same alignment problem from the principal's side. Our framing also connects to work on the delegation of decisions to algorithms and on when people accept or resist machine judgment dietvorst2015algorithm,logg2019algorithm,chugunova2022we, to which we add a structural account of the preferences an algorithm brings once a decision is delegated to it.
The remainder of the paper proceeds as follows. Section (ref) develops the random utility representation and the revealed-preference diagnostics. Section (ref) specializes the framework to the portfolio choice environment and derives the two estimators. Section (ref) describes the experimental design and implementation. Section (ref) reports the diagnostics, the structural estimates, and the fine-tuning validation. Section (ref) concludes.
Suppose a researcher gave the following decision problem to a human subject.
From the perspective of classical choice theory, the exact mechanism driving the choice made by the subject is an unobserved latent process. For instance, it can be difficult to measure the strength of preference one has over the available options. Consequently, whatever the subject's underlying utility function is (if any), all aspects of preference must be identified through the observed decisions from the experiment. In contrast, because language models possess an explicit computational architecture that maps an option menu input to a choice, their decision-making apparatus is directly observable. In this sense, the preferences of a language model are easier to inspect than those of a human. This paper leverages this architectural transparency to identify not only the economic preferences revealed by a language model, but also specific components of the structural mechanism mapping information to choice. We begin by describing the relevant components of this architecture and how they relate to economic models of utility.
Modern transformer-based language models generate text autoregressively as a sequence of discrete tokens drawn from a finite vocabulary set $\mathcal{V}$ vaswani2017attention,brown2020language.\footnote{Tokens are the fundamental units of text parsed by the model, encompassing words, subwords, and punctuation marks. The vocabulary set $\mathcal{V}$ is fixed by the model's architecture.} At each step in the sequence, the model evaluates the accumulated context of preceding tokens through a highly parameterized non-linear function comprising attention mechanisms and feed-forward networks. This evaluation yields a vector of latent unnormalized scores over the entire vocabulary. In the machine learning literature, these raw output values are referred to as “logits.” The realization of the subsequent token is then drawn stochastically from a categorical probability distribution over $\mathcal{V}$. The model derives this distribution by applying a softmax transformation to the vector of logits, mapping the unbounded latent scores into proper probabilities that sum to one.
In the discrete choice experiments that follow, we impose a one-step forced-choice problem in which a single output token from the model constitutes a revealed preference. This constrained decision problem invites comparison with the System 1 mode of cognition described by stanovich2000individual and kahneman2003maps. The resulting observed choices represent fast and associative heuristics rather than deliberative and sequential reasoning. Under this restricted choice architecture, we demonstrate that the model's preference rankings are completely determined by the logit differentials over the permissible action space.
For single-token revealed-preference problems, logit gaps provide the primitive object for identification. These gaps determine the relative odds of the menu labels under the decoding rule introduced below. With an affine utility specification, the same gaps become scaled utility differences. This maps the machine-learning object of a logit gap into the econometric object of a utility gap, placing our identification strategy in the structural choice tradition thurstone1927law,marschak1960binary,manski1975maximum. We formalize the mapping by showing that the induced choice rule admits an exact random utility representation nested within the multinomial logit framework of mcfadden1972conditional. We then define six revealed-preference indices that measure conformance to utility axioms and odds-invariance restrictions on the model's forced-choice outputs. The portfolio-choice problem in Section (ref) derives the two estimators used for our structural recovery exercise.
For a given discrete choice problem, let $x$ denote the prompt context and let $C(x) \subseteq \mathcal{V}$ denote the set of label tokens (e.g. Option “A” vs “B”) comprising the feasible options in the choice menu. Let $u_i(x)$ denote the raw logit on token $i \in \mathcal{V}$. Under standard softmax decoding bridle1989training, the language model induces a next-token probability distribution over the full vocabulary,
where the parameter $\tau > 0$ is temperature, which scales the dispersion of the induced choice distribution. As $\tau \downarrow 0$ the distribution concentrates all mass on the single token with the highest logit across $\mathcal{V}$. As $\tau \rightarrow \infty$ the distribution approaches the uniform over $\mathcal{V}$, spreading mass evenly across all vocabulary tokens and effectively destroying the model's inherent coherence.
Since the shocks $\epsilon_i$ are i.i.d.\ Type I extreme value across all $i \in \mathcal{V}$, the softmax rule in Equation (ref) is the full-vocabulary multinomial logit representation of the random utility model from mcfadden1972conditional. Under this interpretation, a model's raw logit values serve as a systematic utility index, and the temperature parameter $\tau$ is the scale of the stochastic component. Consequently, as $\tau \downarrow 0$, choice converges to the label with the highest logit, and the distribution recovers only the ordinal ranking of the systematic utilities (i.e. the noise vanishes and the cardinal magnitudes of utility differences become irrelevant).
To connect logits to structural preferences, we specify the systematic utility index as a positive affine transformation of a cardinal utility index $V_i$ over the full vocabulary,\footnote{Under expected utility, any two cardinal representations of the same preferences differ by a positive affine transformation von1947theory. Equation (ref) therefore preserves the underlying vNM preference ordering and the associated risk attitudes.}
The scale parameter $\kappa > 0$ maps utility differences into logit units and is assumed constant across prompt contexts, so that scaled utility gaps remain comparable across the menus used in our experiments. The term $\zeta(x)$ is a prompt-specific common shifter that is the same for all tokens within a context and therefore cancels from every within-menu comparison. We leave the functional form of $V_i$ is unrestricted here and explore a quadratic form in Section (ref). Nonetheless, its dependence on preference parameters already bears on identification, since apesteguia2018monotone show that an additive random utility model of this form can fail to identify risk preferences when $V_i$ is a constant-risk-aversion expected utility whose between-prospect difference vanishes as risk aversion grows. We therefore require $V_i$ to be affine in the risk-preference parameter, which keeps that difference monotone in it.
Having embedded the language model's token probabilities in a random utility framework, we now formulate the restrictions under which the token-level utility index can be interpreted as stable revealed preferences, and propose concomitant measures of restriction satisfaction. The transformer architecture supplies the logits over the vocabulary, and the random utility representation maps those logits into systematic utilities and induced choice probabilities. The axioms below are therefore not mechanical consequences of the architecture, but restrictions we impose and evaluate to determine whether the induced choices admit a stable revealed-preference interpretation. Throughout this subsection, for each token $k \in \mathcal{V}$, let \[ U_k(x,\tau) = \kappa V_k(x) + \zeta(x) + \tau \epsilon_k \] denote its realized random utility.
Following this utility specification, we define rational behavior according to the classic axioms of completeness, reflexivity, monotonicity, transitivty, and continuity presented in varian1992microeconomic. Additionally, we define an axiom for the independence of irrelevant alternatives that is implied by the random utility model hausman1984specification. Each definition is subsequently paired with a measure of adherence to the axiom that we employ in the experimental protocol that follows.
This means that any token chosen with positive probability is in the choice menu. For example, if the menu choices are A and B ($C(x)=\{A,B\}$) then perfect completeness would imply that the model puts no weight on any other tokens except for those representing “A" and “B". Traditionally, the completeness axiom dictates that the agent can provide a ranking of feasible options, allowing for indifference between bundles but not for answers that would fail to provide a preference ranking. Our condition operationalizes this requirement as menu adherence, in that the model must confine its response to the feasible options rather than produce a token that conveys no ranking of them.
The index $\mathcal{C}(x,\tau)$ lies in $[0,1]$ and equals one if and only if menu completeness holds on $C(x)$. Its dependence on temperature is only through the effective scale $\kappa/\tau$. Lower values of $\tau$ concentrate probability on tokens with maximal utility, while higher values of $\tau$ move the distribution toward the vocabulary share of $C(x)$. Because the full-vocabulary softmax places positive probability on every token whenever logits are finite, exact menu completeness is unattainable at any positive temperature. The axiom therefore describes an ideal, and the index measures proximity to that ideal. Exact completeness obtains only when the logits outside the menu are masked, or in the zero-temperature limit provided the token with the maximal logit lies in the intended menu. Given feasibility of the intended menu, the next restriction concerns equality. Reflexivity requires labels representing the same alternative to receive the same utility and hence the same choice probability.
This condition operationalizes the reflexivity axiom as label invariance, in that the utility assigned to an alternative may not depend on the token label under which it is presented. If a language model is indifferent between option A and option B, and if the model is not biased to choose the first option, then each label should receive one half of the choice probability, which we can observe directly in the logits.
The coordinate space in Figure (ref) represents a choice space of portfolios that increase in risk from left to right and have higher expected return moving up the horizontal axis. Thus, relative to the reference portfolio (gold star), points directly above it should be preferable to points below. The color gradient shows the progression of logit differences. Dark blue represents an area where the portfolios are intensely preferred to the reference star because those points offer both a much higher return and lower risk. The solid black line maps an indifference curve, or all the points that are valued equally as the reference point. A risk averse decision maker would exhibit an upward sloping indifference curve in this choice space, as shown in the picture. Panel (a) of Figure (ref) shows that if A and B label the same portfolio, then they should receive the same utility regardless of which label is assigned.
The index $\mathcal{R}(i,j,x,\tau)$ lies in $[0,1]$ and equals one if and only if $V_i(x)=V_j(x)$, equivalently $P(i \mid x,\tau)=P(j \mid x,\tau)$. Temperature enters only through the scaled utility difference. If $V_i(x)=V_j(x)$, exact reflexivity is invariant to temperature. If $V_i(x)\neq V_j(x)$, lowering $\tau$ increases the scaled gap and drives $\mathcal{R}(i,j,x,\tau)$ toward zero, while raising $\tau$ compresses the gap and moves $\mathcal{R}(i,j,x,\tau)$ toward one.
The next restriction concerns strict improvement. Given an exogenous mean-risk dominance order, monotonicity requires the model's induced order to respect every dominance comparison. Panel (b) of Figure (ref) shows that portfolio A has both a higher expected return, $\mu$, and less uncertainty, relative to porfolio B. Therefore, any agent whose preferences are increasing in expected return and decreasing in risk should prefer A to B. (Rationality alone does not deliver this ranking, since a risk-seeking expected utility maximizer may prefer the riskier portfolio.) We use $\succ_D$ to denote this mean-risk dominance relation. The benchmark utility index maintained in Section (ref) satisfies this monotonicity requirement for every positive value of its risk preference parameter on the mean-risk dominance menus used in our design, where the dominating portfolio has higher expected return and lower raw second moment $R$.
Our measure requires that the probability of choosing a preferable portfolio ($i$) be higher than the probability of choosing a less desirable portfolio ($j$). A language model can place positive probability on the mean-risk dominated option without violating the monotonicity axiom because the axiom restricts the ordering of systematic utilities and choice probabilities, not the probability level assigned to each option. For example, considering panel (b) of Figure (ref), the probability assigned to portfolio B can be positive as long as the probability assigned to A is higher. Monotonicity therefore is a pairwise restriction relative to the exogenous order $\succ_D$. It requires the model's induced order to agree with $\succ_D$ on each dominance pair, but it imposes no condition on the joint consistency of distinct pairwise comparisons. The next axiom imposes that consistency condition by requiring the strict relation induced by the model to be closed under composition (represented in panel (c) of Figure (ref)).
Within a single prompt context the index $V_i(x)$ assigns a real number to every token, so the ordering induced inside one menu is transitive by construction. The empirical content of the axiom therefore concerns whether the rankings elicited from separate binary menus can be rationalized by a single stable ordering over the underlying alternatives.
Panel c of Figure (ref)) shows that if A lies above the indifference curve for bundle B (the gold star) and C lies below that indifference curve, then it follows that this decision maker would prefer A to C, if they have rational preferences. Therefore, transitivity is a finite consistency requirement on a closed triad. If the first two comparisons induce a strict chain, then the third comparison must agree with the order obtained by composing the first two.
The next axiom is continuity, which rules out discontinuous jumps in preference. Standard preference continuity implies that, between alternatives ranked on opposite sides of a reference alternative, a continuous path must cross the reference alternative's indifference contour.
Continuity is the corresponding path-existence restriction. It requires an indifference crossing of the reference alternative along the interpolation path, equivalently a zero crossing of the log probability odds. The grid index records whether the probed comparisons bracket such a crossing, without imposing monotonic movement or uniqueness.
The final restriction concerns the stability of pairwise odds when alternatives outside the pair are varied. Within a fixed prompt context, odds invariance is a direct consequence of the model's equivalence to the standard random utility model, where the ratio of choice probabilities for any two options depends solely on their systematic utility difference luce1959individual. Across the binary and expanded prompts used below, the restriction becomes an empirical test of whether that systematic utility difference remains stable when the menu changes.
Independence of irrelevant alternatives is an odds-invariance restriction for an irrelevant alternative introduced by construction. It requires the pairwise log odds of $a$ against $b$ to be unchanged when the menu is expanded by an alternative $c$ that is strictly mean-risk dominated by both $a$ and $b$.
Together, the six indices state the restrictions under which the random utility representation and the affine logit specification admit a revealed-preference interpretation in one-step token choice. We next apply this framework to the portfolio choice problem and derive the two structural estimators.
The structural framework in Section (ref) connects the model's logit vector to a cardinal utility index $V_i$ through the scale parameter $\kappa$, but leaves the functional form of $V_i$ unrestricted. We now specialize that index to a portfolio choice environment in which each alternative is a one-period asset with observed expected return $\mu_i$ and standard deviation $\sigma_i$, so that the relevant tradeoff between return and risk is transparent, fully observed by the researcher, and directly presented to the model. The researcher then recovers variance and the raw second moment mechanically from the reported moments. This environment gives the structural parameters an economically interpretable domain while keeping the objects of choice under complete experimental control.
In this environment, each prompt context $x_j$ presents a binary feasible portfolio set $C(x_j) = \{A,B\}$, where the two label tokens encode portfolios with observed central moments $(\mu_{Aj}, \sigma_{Aj})$ and $(\mu_{Bj}, \sigma_{Bj})$. Holding the rest of the prompt template fixed while varying these moments across menus gives a direct implementation of the measurement objects from Section (ref) in a setting with economically interpretable alternatives. Consequently, the associated axioms reveal whether language models exhibit preferences consistent with a single utility function. Relatedly, we note the risk preferences recovered in this setting reflect a preference index over a well-defined class of risky allocations, but one that need not summarize behavior outside it. It is well-known that risk preferences are not globally consistent across domains among human subjects einav2012general, harrison2007naturally, schildberg2018risk, and we explore the extent to which this is true for language models as well.
Our maintained benchmark is quadratic expected utility, which serves as a useful approximation to the expected utilities of a diverse set of utility functions markowitz2014mean. To express this compactly, let $R_i = E[x_i^2]$ denote the raw second moment of portfolio returns. The moments shown in the prompt then imply $R_i = \sigma_i^2 + \mu_i^2$, since $\sigma_i^2 = E[x_i^2] - E[x_i]^2$. Recovering structural risk preference parameters from discrete choices over financial alternatives with known payoff distributions has clear precedent in the empirical literature mester1996study, choi2007consistency, cohen2007estimating, list2011ceos.\footnote{Richer alternatives such as prospect theory, probability weighting, and disappointment aversion have been studied extensively kahneman1979prospect, gul1991theory, prelec1998probability, bruhin2010risk.} For the label tokens that encode portfolios, we model the corresponding expected utility index as
where $\beta > 0$ indicates risk aversion, $\beta < 0$ indicates risk seeking, and $\beta = 0$ collapses to expected-return maximization. Under Equation (ref), whether portfolio $i$ is preferred to portfolio $i'$ when $\mu_i > \mu_{i'}$ and $R_i > R_{i'}$ depends entirely on whether $\beta \leq (\mu_i - \mu_{i'})/(R_i - R_{i'})$, so differences in $\beta$ translate directly into different portfolio rankings and, once the model selects among feasible allocations, different economic outcomes.
Substituting Equation (ref) into Equation (ref) for the portfolio labels maps the observed portfolio moments into the pre-temperature logit index. Specifically, for alternative $i$ in menu $j$,
where the additive term $\zeta_j$ is common within menu $j$ and drops from within-menu comparisons. By imposing strategic variation in $(\mu_{ij}, R_{ij})$ across alternatives and menus, we can identify the structural parameters of interest. We therefore consider two observation regimes using this framework. When raw pre-softmax logits on the menu labels are observed, within-menu logit gaps can be used to identify the structural parameters directly. When only sampled choices are observed, we employ traditional random utility model estimation methods via maximum likelihood. We explore each of these in turn below. Appendix (ref) summarizes the identification conditions and design implications under both observation regimes.
When logits are observed, the data for menu $j$ are the two label logits together with the corresponding portfolio moments. Because the additive menu-specific level in Equation (ref) is common to both options, estimation proceeds on the within-menu logit gap and does not require menu fixed effects. An equivalent stacked representation is \[ u_{ij} = \theta_{0j} + \alpha \mathbf{1}\{i=A\} + \theta_1 \mu_{ij} + \theta_2 R_{ij} + \varepsilon_{ij} \] where $u_{ij}$ is the raw logit for alternative $i$ in menu $j$, $\theta_{0j}$ absorbs the menu-specific level including the common term $\zeta_j$ from Equation (ref), and $\alpha$ captures any systematic additive position bias of the first option labeled A. Position bias is a well-documented phenomenon of current language models, which we discuss in Section (ref) in further detail. With two observations per menu, this stacked equation is algebraically equivalent to the first-difference regression used for estimation.
For a binary menu $j$ with options A and B, within-menu differencing yields
where $\Delta u_j = u_{Aj} - u_{Bj}$, $\Delta \mu_j = \mu_{Aj} - \mu_{Bj}$, $\Delta R_j = R_{Aj} - R_{Bj}$, and $\eta_j = \varepsilon_{Aj} - \varepsilon_{Bj}$. Estimation therefore runs on the logit gap with a pooled constant and the within-menu differences in moments. Because the dependent variable is the raw pre-temperature logit gap, decoding temperature does not affect the identifying equation or the precision of this estimator. In this sense, the estimator keeps the component of the logit gap that is systematically explained by portfolio attributes and treats the remaining gap as either positional bias, absorbed by $\alpha$ or removed by mirroring, or specification error.
Identification requires that the slope mapping from utility units to logit units be common across menus, that any positional effect enter as a common additive constant, and that the design matrix with rows $(\Delta \mu_j,\Delta R_j)$ have full rank. The reduced-form slopes imply \[ \hat{\kappa} = \hat{\theta}_1 \] and \[ \hat{\beta} = -\frac{\hat{\theta}_2}{\hat{\theta}_1}. \] Inference can use heteroskedasticity-robust standard errors for $(\hat{\theta}_1,\hat{\theta}_2)$. Since $\hat{\beta}$ is a ratio estimator, inference for $\hat{\beta}$ follows from the Delta method applied to the estimated covariance matrix.
When only sampled choices are observed, we restrict attention to valid A/B responses and let $y_{jr}=1$ if option A is chosen on valid draw $r$ from menu $j$ and $y_{jr}=0$ if option B is chosen. With $\Delta \mu_j = \mu_{Aj} - \mu_{Bj}$ and $\Delta R_j = R_{Aj} - R_{Bj}$, the binary choice probability is
where $\Lambda(t) = \exp(t) / (1 + \exp(t))$, $\gamma_1 = \kappa / \tau$, and $\gamma_2 = -\kappa\beta / \tau$. Note that varying $\tau$ across menus does not add identifying content because it only rescales the common latent index. Therefore, in the benchmark design we hold $\tau = 1$ fixed across estimation menus, which preserves the native logit scale.
Full-rank variation in $(\Delta \mu_j,\Delta R_j)$ identifies the reduced-form slopes. Choice data therefore identify $\beta = -\gamma_2/\gamma_1$ whenever $\gamma_1 \neq 0$. In the frontier regime (i.e. ChatGPT from OpenAI) of our experiments, let $n_j$ denote the number of valid A/B responses from menu $j$. Let $Y_j=\sum_{r=1}^{n_j} y_{jr}$ denote the number of those responses choosing option A. Estimation maximizes the grouped binomial log-likelihood
The omitted binomial coefficient is constant in $(\gamma_1,\gamma_2)$ and therefore does not affect the maximizer. Under these conditions, together with the standard identification and regularity conditions stated above, maximum likelihood yields consistent estimates $\hat{\gamma}_1$ and $\hat{\gamma}_2$ newey1994large. Because $\tau$ is exogenously imposed, the structural parameters can be recovered as \[ \hat{\kappa} = \tau \hat{\gamma}_1 \] and \[ \hat{\beta} = -\frac{\hat{\gamma}_2}{\hat{\gamma}_1}. \] Inference for $\hat{\kappa}$ follows from the inverse Hessian, while for $\hat{\beta}$ the Delta method once again is convenient away from weak-ratio cases.
The experiment is built to address two measurement problems. The first is to determine whether a model's single-token choices satisfy the revealed-preference restrictions required for the utility representation defined by our measures of completeness, reflexivity, monotonicity, transitivity, continuity, and independence-of-irrelevant-alternatives formalized in Section (ref). The second is to estimate the shape of preferences under the quadratic mean-risk benchmark of Section (ref), recovering the scale parameter $\kappa$ and the risk preference parameter $\beta$.
The unit of observation throughout is a model-menu prompt, a single decision context $x$ that presents a labeled portfolio menu and to which the model responds by placing probability on one label token. What the researcher observes from that prompt differs across the two regimes studied in the paper. For open-weight models we record the raw pre-softmax logits on the label tokens from a single forward pass.\footnote{All open-weight models are accessed through their publicly released weights. The full raw logit vector is recorded at the label token positions before any softmax decoding or sampling rule is applied. No provider-side post-processing intervenes between the final layer output and the recorded logit vector, so the observed-logit regime is exact.} For private frontier models we typically do not observe logits and instead query the model repeatedly under softmax decoding, recording the sampled choices from which empirical choice probabilities are formed.
Our subject pool spans representative samples of non-thinking models from the open- and closed-source regimes. Within each regime we include models that were publicly available during the window of our study in June 2026, exposed stable version identifiers, and could be queried under a one-token forced-choice protocol with fixed decoding controls. Table (ref) lists the resulting sample. Panel A covers open-weight models, and Panel B covers private frontier models. Full details on the experimental code, hardware, vendor metadata, prompt and design hashes, and run identifiers are collected in the replication materials in Appendix (ref).
The main experiment administers a common menu set to every model. The decision prompt follows a fixed three-part template as presented in Figure (ref) below. The first part frames the decision as an annual financial portfolio choice and instructs the model to reply with its preferred label from the listed options. The second part presents the options by specifying the expected return and standard deviation of each. The third part is the cue suffix “I choose Option”. For private frontier models this suffix is appended to the end of the user prompt. For open-weight models it is prefilled at the start of the assistant response block. In both cases the model is induced to emit a single token indicating its choice. In all prompts the labels are the capital letters A and B (and C for the trinary independence menu), which are single tokens in the vocabulary of every model in the subject pool.\footnote{Some tokenizers include leading-space variants of the A and B tokens. When more than one token realization maps to a label, we aggregate at the label level by summing the softmax probabilities of the associated token realizations. The equivalent grouped logit at temperature $\tau$ is $\tau\log\sum_r\exp(u_r/\tau)$ over realizations $r$ of the label. We verify this property before running the experiment and report the token IDs in Appendix (ref).}\footnote{We explored the potential effects of using these specific option labels by randomly selecting tokens as option labels. We found no meaningful difference in these experimental results.}
The two observation regimes differ not only in what is recorded but in how choice probabilities are formed. In the open-weight regime, one forward pass per menu yields the exact label logits, and no replication is required. In the frontier regime, choice probabilities are formed either from repeated sampled choices or, where the provider exposes them, from first-token logprobs. Across both regimes we fix the decoding controls at temperature $\tau = 1$ and top-$p = 1$.\footnote{Top-$p$ sampling (also known as nucleus sampling) restricts the candidate token set at each step of decoding to the smallest subset of tokens whose cumulative probability exceeds the threshold $p$ holtzman2019curious. Setting top-$p=1$ keeps the entire vocabulary and avoids truncation, which would break the full-vocabulary softmax that underlies the random utility representation in Proposition (ref) and the indices that rest on it mcfadden1972conditional, train2009discrete.}
All menus draw choice options from a discrete grid of portfolio bundles with expected returns $\mu \in \{2, 3, \ldots, 18\}$ percent and standard deviations $\sigma \in \{10, 15, 20, \ldots, 55\}$ percent, yielding $17 \times 10 = 170$ distinct bundles, each with raw second moment $R = \sigma^2 + \mu^2$. This range spans low-return, low-risk profiles through high-return, high-risk profiles, and the spacing is fine enough to generate meaningful variation in the risk-return tradeoff. The $170$ bundles admit $\binom{170}{2} = 14{,}365$ distinct binary pairings, of which $13{,}596$ form the admissible working pool after removing pairs that are degenerate for the relevant first differences. From this pool we strategically construct the diagnostic and estimation menus described below, holding the menu sets constant across all models.
To control for a well known position bias exhibited by many language models zheng2024large,pezeshkpour2024large,guan2025order, and for the broader possibility that choices depend on the frame in which a fixed alternative is presented salant2008f, every canonical menu is also administered in a mirrored presentation that swaps which option carries each label. The reflexivity placebos described below show that a label's position can carry a systematic logit premium orthogonal to the systematic utility index in Equation (ref). For each menu we therefore form the position-corrected log-odds gap
where $\Delta u$ and $\Delta u^{\text{mirror}}$ denote the A-minus-B log-odds gap under the canonical and mirrored presentations. Because the positional premium enters both presentations with the same sign while the utility difference reverses, the half difference removes the positional component and retains the systematic component. Unless stated otherwise, the monotonicity, transitivity, continuity, and independence diagnostics are evaluated on these position-corrected gaps. The reflexivity index is computed within a single presentation by construction, since the placebo's two options are identical and the asymmetry is itself the positional signal, and the trinary independence menus use the analogous correction described below.
Table (ref) summarizes the diagnostic battery by utility axiom, reporting the menu construction used to evaluate each measure and the corresponding open-weight prompt-level observations. Completeness differs from the other diagnostics because it is evaluated on the union of administered binary prompts rather than on a separate menu class. For each open-weight prompt we extract the full logit vector from a single forward pass and compute the completeness index $\mathcal{C}(x,\tau)$ from Measure (ref) as the softmax mass on the valid label tokens $\{A,B\}$. For frontier models, completeness is the fraction of sampled responses that parse as a valid label.
The diagnostic menus are designed to isolate distinct failures of revealed preference (see Figure (ref) for reference). Placebo menus hold the portfolio fixed and therefore attribute any A-versus-B asymmetry to label position. Dominance menus test whether the model respects unambiguous mean-risk improvements. Triad menus test whether pairwise rankings close consistently across three bundles, while interpolation menus ask whether rankings vary continuously along a path between endpoint bundles. Independence menus add a dominated third option to test whether the relative odds between the two focal portfolios are stable to irrelevant menu expansion.
Our quadratic utility benchmark identifies preferences from variation in the within-menu differences $\Delta\mu$ and $\Delta R$. Consequently, the estimation menus are selected to maximize the statistical efficiency of these two regressors. After enumerating the admissible binary pairs from the $170$-bundle grid, we exclude identical bundles and pairs in which one option strictly mean-risk dominates the other, and from the remainder we select menus using a greedy forward D-optimal algorithm that sequentially maximizes the determinant of the information matrix over $(\Delta\mu, \Delta R)$ kiefer1959optimum,mitchell2000algorithm,rose2009constructing. \footnote{The D-optimal criterion is invoked here, for the parametric estimation, rather than for the axiom diagnostics, because its purpose is to break the collinearity between the return and risk differences and sharpen identification of the structural slopes.} The open-weight estimation sample comprises $680$ D-optimal canonical menus together with their $680$ mirrors. The frontier estimation sample uses $400$ canonical estimation menus, administered with mirrors for $800$ total presentations.
For open-weight models, the data for each menu are the two label logits, and estimation proceeds on the position-corrected within-menu logit gap. We fit the first-difference regression in Equation (ref) with paired corrected differencing, which removes the additive positional advantage $\alpha$ before recovering $\hat{\kappa} = \hat{\theta}_1$ and $\hat{\beta} = -\hat{\theta}_2/\hat{\theta}_1$. Inference for the ratio follows from the Delta method greene2003econometric. For frontier models, we observe sampled choices rather than logits and estimate the binary logit index in Equation (ref) by maximizing the Bernoulli log-likelihood in Equation (ref), recovering $\hat{\kappa}$ and $\hat{\beta}$ from the reduced-form slopes as in Section (ref). For frontier models, choice probabilities are estimated from repeated queries, with replication counts varying by provider and menu class.\footnote{Structural estimation menus and sampling-based diagnostics receive the highest replication, while Google diagnostics that expose first-token logprobs require only a single query. Some large diagnostic pools are capped for cost control; Appendix (ref) reports the full schedule.}
Design construction and query scheduling use fixed seeds (design and trial seeds set to $42$), and the menu sets are held constant across all subjects. Open-weight logits are extracted deterministically in a single forward pass per menu, so decoding temperature does not enter the open-weight identifying equation. Each query is issued in an isolated session to prevent cross-prompt contamination, and we assess sensitivity to prompt wording through the prompt-variant robustness checks reported in Appendix (ref).
For the open-weight observed-logit estimator we report heteroskedasticity-robust standard errors and treat the residual structure as a specification diagnostic, since the logits are deterministic given the prompt and any dispersion reflects departures from the maintained affine specification rather than sampling noise. For the frontier sampled-choice estimator, by contrast, the repeated draws carry a genuine sampling interpretation, and inference follows from the likelihood in the usual way.
We organize the results in two parts. We first evaluate each model against the six axiom diagnostics defined in Section (ref) (completeness, reflexivity, monotonicity, transitivity, continuity, and independence of irrelevant alternatives) using the corresponding indices $\mathcal{C}$, $\mathcal{R}$, $\mathcal{M}$, $\mathcal{T}$, $\mathcal{K}_G$, and $\mathcal{I}$ computed from the observed logit vectors, and report analogous sampled-choice diagnostics for the frontier runs where available. These diagnostics establish whether each model's single-token choices satisfy the conditions for a stable utility-theoretic interpretation before any structural parameters are estimated. We then turn to the structural estimation of the risk preference parameter $\beta$ and the scale parameter $\kappa$ under the maintained benchmark utility $V_i = \mu_i - \beta R_i$.
Table (ref) reports the model-level averages of each diagnostic index across the open-weight and frontier diagnostic runs. Four patterns stand out. First, completeness is near-perfect for most models. Qwen 3, Gemma 4, and Ministral 3 concentrate effectively all softmax mass on the intended label tokens, with $\mathcal{C} > 0.999$. The Llama 3.1 family is the exception, with the 70B variant allocating roughly one quarter of its probability mass to tokens outside the menu.\footnote{This result is robust to variations in the prompting of the model. For example, some models are trained to produce output in markdown format, which can lead to leading characters for emphasis (e.g. “*” or other symbols in front of the labels).} Second, reflexivity reveals substantial heterogeneity. Models that achieve high completeness do not necessarily treat identical portfolios symmetrically. Qwen 3 (14B) and Gemma 4 exhibit near-zero reflexivity. Llama 3.1 (8B Instruct) shows the strongest adherence to the reflexivity axiom among the base models, achieving $\mathcal{R} = 0.836$. All but one of the open weight models we tested (Gemma 4) expressed varying degrees of preference for the first option in the menu, which had to be controlled for in the structural estimation procedure that follows. It is only when the portfolios are exactly the same that the models indulge this position bias, thus there is no economic consequence for this violation of the reflexivity axiom because, regardless of the label that is chosen, the final allocation is the same. See Table (ref) in the appendix for more details on this result. Third, monotonicity is satisfied at ceiling across the entire subject pool, and continuity is satisfied at ceiling wherever the premise-verified continuity index is available.
These results indicate that the models uniformly respect stochastic dominance and exhibit smooth, monotone movement of the log-odds along interpolation paths. Transitivity is high throughout, ranging from $0.957$ for Qwen 3 (8B) to $1.000$ for Llama 3.1 (70B) and the OpenAI frontier models, implying that preference cycles are rare but not entirely absent among directional triads.
Fourth, independence of irrelevant alternatives is the weakest axiom across the entire subject pool. No model approaches the ceiling scores observed for monotonicity or continuity. The index $\mathcal{I}$ ranges from $0.122$ for Gemma 4 to $0.920$ for Llama 3.1 (8B Instruct), indicating that adding a strictly dominated third option systematically shifts the log-odds between the original pair. The IIA scores covary strongly with reflexivity. Models exhibiting severe positional label bias, such as Gemma 4 ($\mathcal{R} = 0.000$, $\mathcal{I} = 0.122$) and Qwen 3 14B ($\mathcal{R} = 0.001$, $\mathcal{I} = 0.322$), show the largest violations. This suggests that the same positional sensitivity that prevents symmetric treatment of identical options also makes the pairwise odds fragile to menu expansion.
\noindentFinding 1: The open-weight and frontier models in our sample largely adhere to our measures on the axioms of utility theory. However, we find that a position bias emerges among nearly identical options which leads to violations in reflexivity. Models also do not adhere to the independence of irrelevant alternatives assumption.
Having established which axiom conditions each model satisfies, we now estimate the structural parameters of the maintained quadratic expected utility $V_i = \mu_i - \beta R_i$. Table (ref) reports estimates of the scale parameter $\hat{\kappa}$ and the risk aversion parameter $\hat{\beta}$. The open-weight rows use the observed-logit estimator in Equation (ref), with paired corrected differencing that removes the additive positional bias $\alpha$ before fitting. The frontier rows use the sampled-choice maximum likelihood estimator. Heteroskedasticity-robust standard errors appear in parentheses, and inference for $\hat{\beta}$ follows from the Delta method applied to the ratio $-\hat{\theta}_2 / \hat{\theta}_1$.
Every model in the subject pool yields a positive and statistically significant estimate of $\beta$, indicating that all models exhibit risk aversion. The point estimates vary by a factor of nearly four, ranging from $\hat{\beta} = 0.0031$ for Qwen 3 (8B) to $\hat{\beta} = 0.0114$ for Ministral 3 (8B Instruct). This range implies economically meaningful variation in risk attitudes across model families. Models at the upper end penalize risk more aggressively and will forgo higher expected returns in favor of lower dispersion at comparatively modest risk differentials.
Every model yields a strictly positive estimate for the scale parameter $\hat{\kappa}$, confirming that they naturally interpret the value ordering of the portfolios in the experiment by preferring higher expected returns. If, in contrast, $\kappa$ estimates had been negative then models would have preferred portfolios with lower expected returns, ceteris paribus. The magnitude of $\hat{\kappa}$ exhibits wide cross-model variation. Gemma 4 (31B) produces the largest estimate at $\hat{\kappa} = 1.482$, implying that a one-unit increase in the cardinal utility index moves the pre-softmax logit by nearly 1.5 units. By contrast, Llama 3.1 (8B) yields $\hat{\kappa} = 0.039$, roughly forty times smaller.
To provide a richer visualization of these preferences beyond the point estimates in Table (ref), we plot each model's structural $\hat{\beta}$ estimate (except Ministral) in the $\mu$--$\sigma$ plane. Starting from a base bundle with expected return $\mu = 10$ and risk $\sigma = 30$, we sweep over the dense grid of alternative portfolios from our experimental protocol and record the position-corrected logit gap between each grid portfolio and the base. The yellow dashed line overlays the indifference curve implied by the mean-variance structural estimates, and the heatmap colors encode the sign and magnitude of the logit difference, with blue regions preferred to the base and red regions dispreferred.
Several patterns emerge from these plots. All six models produce upward-sloping indifference curves, confirming that higher risk must be compensated by higher expected return. The curves are roughly convex in the $\mu$--$\sigma$ plane, consistent with the quadratic utility specification maintained in the structural estimation. However, the steepness of the curves varies considerably across models, reflecting the heterogeneity in $\hat{\beta}$ documented above. Models with larger risk aversion estimates, such as Gemma 2 (9B), trace out steeper curves that demand more return compensation per unit of additional risk. Models with smaller $\hat{\beta}$, such as Qwen 3 (14B), produce flatter curves that tolerate substantially more risk for modest return increments.
\noindentFinding 2: All open-weight and frontier models in our sample exhibit varying degrees of risk aversion that can be approximated by a quadratic utility function.
Can specific economic preferences be induced into the language model? This question matters because a principal who delegates an allocation decision to a language model delegates authority over the tradeoff that the instruction leaves unresolved. If the model's default risk attitude differs from the principal's, then the delegated choice implements the model's objective rather than the principal's. The preceding results show that these default preferences can be measured. We now ask whether the preference-misalignment component of this agency problem can also be addressed directly by inducing the model to reveal a target economic preference in subsequent choices.
We study this question by fine-tuning Llama 3.1 (8B Instruct) toward two principal-specified risk preferences. The principal's desired ordering over portfolios is represented by the same maintained utility index, \[ V_i^* = \mu_i - \beta^* R_i, \] where $\beta^*$ is chosen by the principal. We consider two target specifications. The first is approximately risk neutral, with $\beta^* = 0.0001$. The second is more risk averse, with $\beta^* = 0.0045$. These targets are not prompts asking the model to describe a preference. They are the preferences the delegated agent is trained to implement in its own choice probabilities.\footnote{This differs from prompt engineering. Prompt engineering conditions the model on instructions supplied in the input context and leaves the model weights unchanged, so any induced behavior persists only while that context is present. Fine-tuning updates the weights that generate the choice rule itself, allowing the principal-specified preference to be embedded in the model's revealed choice probabilities.}
For each training menu, we compute the utility difference implied by the principal's target preference and translate it into the log odds that the delegated model should assign to the two options.
For the same menu, let $\Delta z_m(\theta)$ denote the model's actual logit gap between the two answer labels, where $\theta$ collects the model weights. The preference component of fine-tuning chooses weights to close the distance between the actual logit gap and the target logit gap, \[ \mathcal{L}(\theta) = \frac{1}{M}\sum_{m=1}^{M} \left(\Delta z_m(\theta) - \Delta z_m^*\right)^2 . \] Minimizing $\mathcal{L}(\theta)$ updates the model toward the principal's choice rule. If the model places too much relative logit weight on option $A$, the residual is positive and the objective rewards changes that reduce the $A$ versus $B$ gap. If the model places too little relative logit weight on option $A$, the objective rewards changes that increase that gap. Repeating this comparison across the training menus moves the model's logit surface toward the logit surface implied by the principal's utility. The scale parameter $\kappa$ governs the strength of the probabilistic response but does not determine the recovered risk preference, since $\hat{\beta}$ is identified by the ratio of the risk and return slopes. Appendix (ref) gives the formal multi-term objective, the training corpus, the logit-scale choice, and the implementation details.
After fine-tuning, we evaluate the delegated models exactly as before, on held-out menus not used in training. The test is whether the model now reveals the principal's target preference rather than the base model's default preference, and whether that induced preference passes the same diagnostic battery. Table (ref) reports the rationality diagnostics. Relative to the baseline Llama 3.1 (8B Instruct), both induced-preference models substantially improve completeness and reflexivity, with completeness above $0.99$ and reflexivity above $0.98$. Monotonicity remains perfect in both cases. The near risk-neutral target achieves perfect transitivity and raises IIA to $0.9484$, while the more risk-averse target maintains high diagnostic performance, with IIA equal to $0.8000$ and transitivity equal to $0.9583$.
Table (ref) reports the structural estimates and Figure (ref) plots the implied indifference curves over the portfolio space. The induced models reveal the principal-specified preferences with high precision. The near risk-neutral specification targets $\beta^* = 0.0001$, and the fine-tuned model yields $\hat{\beta} = 0.000087$ with a standard error of $0.000003$. The more risk-averse specification targets $\beta^* = 0.0045$, and the fine-tuned model yields $\hat{\beta} = 0.004560$ with a standard error of $0.000017$. In both cases, the target value lies within the 95% confidence interval. The scale parameter estimates are similarly precise. The low-$\beta$ model yields $\hat{\kappa} = 1.001301$ with a standard error of $0.000720$, and the high-$\beta$ model yields $\hat{\kappa} = 0.718940$ with a standard error of $0.002206$. The restricted quadratic surface provides a tight fit to the indifference data, producing $R^2$ values of $0.9992$ and $0.9935$.
The economic interpretation is that the principal need not accept the model's default risk attitude as an unavoidable agency cost within this portfolio environment. A principal can specify a target utility function, fine-tune the model toward the choice odds implied by that utility, and then verify out of sample that the delegated model reveals the intended preference. The exercise therefore addresses the preference-misalignment part of the principal-agent problem studied here. The agent's objective can be measured, deliberately shifted, and audited with the same structural tools.
\noindentFinding 3: A principal can install its preferred risk attitude in a language model through fine-tuning, embedding the preference in the model's weights so that it governs how the model resolves tradeoffs.
The starting point of this paper is an equivalence. When a language model is required to state its choice with a single token, the rule by which it selects that token is exactly the conditional logit random utility model of mcfadden1972conditional. Using this equivalence, we develop a structural methodology for recovering the economic preferences of these models directly from their generative outputs. The raw pre-softmax logits assigned to menu labels can therefore be interpreted as a systematic utility index and place the choices of a language model directly within the scope of standard structural discrete choice methods. It yields one estimator when the model's internal scores are observable, as with open-weight models, and the standard maximum likelihood estimator when only realized choices are observed, as with proprietary frontier models (i.e. ChatGPT). The same framework also motivates six revealed-preference diagnostics for completeness, reflexivity, monotonicity, transitivity, continuity, and independence of irrelevant alternatives. These diagnostics allow the stability of the utility interpretation to be evaluated before any structural parameter is estimated.
Applying this method to analyze the choices of several language models, our results reveal a clear distinction between coherence about economic content and invariance to economically irrelevant features. Across the models we study, monotonicity holds for every model, continuity holds wherever the premise-verified index is available, and transitivity is high. The models therefore respect mean-risk dominance, move smoothly along the interpolation paths in the experiment, and rarely generate cycles across separately elicited menus. In this sense, the models behave like textbook rational agents. The main failures arise from features of the elicitation environment that should be economically irrelevant. Reflexivity frequently fails because models assign different weights to identical portfolios upon label or position permutation. Independence of irrelevant alternatives proves to be the weakest diagnostic, as introducing a dominated option routinely alters the pairwise odds between the original alternatives. These failures covary, which suggests that positional sensitivity and menu sensitivity are closely related in single-token model choice.
Conditional on these limits, the structural estimates recover economically interpretable risk preferences. Every open-weight and frontier model in our sample is risk averse, with a positive and statistically significant risk-aversion parameter under the maintained quadratic specification. The estimates vary by nearly a factor of four across model families. This variation is large enough to imply different portfolio rankings and therefore different allocations when a model is asked to choose among risky alternatives. The estimated indifference curves are upward sloping and approximately convex in the mean-variance plane, so the familiar quadratic apparatus provides a useful first-order description of the preference surface these models reveal.
The fine-tuning exercise provides a structural validation of this interpretation. We install a target utility function in an open-source model at two known risk-aversion targets, and our estimator recovers both targets with high precision on menus the model never saw in training. The exercise also shows that the main invariance failures are not fixed properties of the base model. In the controlled environment studied here, the relevant preference object can therefore be measured and deliberately shifted, and the same intervention can improve the invariance properties required for its interpretation.
Two boundaries delimit what we claim. Each choice is a single token produced in one forward pass rather than the outcome of extended deliberation, and the domain is a controlled portfolio environment rather than the full space of economic decisions. These boundaries are what make the measurement exact, since within them the mapping from logits to utility follows from the softmax rule rather than from an auxiliary assumption. They also locate the object precisely. What we recover is the default preference a model brings before it reasons, which is the relevant object when a principal delegates under an incomplete instruction and the model answers directly. Whether that default persists when the model deliberates across many tokens, when the prompt is enriched toward a fuller mandate, or when the domain moves beyond portfolio choice is the natural next question, as is whether the same preference governs the longer chains of reasoning through which these models increasingly act. While we briefly entertain these extensions in our robustness checks, we treat the single-token measurement as our primary interest and the foundation those extensions build on. Whether structurally measured preferences persist under deliberation and across economic environments is the natural next question for this line of work.
These findings matter for the use of language models as economic agents. Delegating allocation decisions to a model also delegates authority to the preferences that govern its choices. Those preferences need not be observed by the principal, and the diagnostics in this paper show that they need not be invariant to irrelevant features of the menu. A coherent deployment strategy therefore requires more than evidence that a model can correctly rank dominated alternatives. It requires precise measurement of the preference governing choice. It requires diagnostics to verify the stability of that preference. Finally, it requires tools to shift the preference when a different objective is intended. The evidence presented here supplies a controlled demonstration of exactly that program.