EconBase
← Back to paper

Inflation Attitudes of Large Language Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

81,708 characters · 21 sections · 45 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inflation Attitudes of Large Language Models

abstractThis paper investigates the ability of Large Language Models (LLMs), specifically GPT-3.5-turbo (GPT), to form inflation perceptions and expectations based on macroeconomic price signals. We compare the LLM's output to household survey data and official statistics, mimicking the information set and demographic characteristics of the Bank of England's Inflation Attitudes Survey (IAS). Our quasi-experimental design exploits the timing of GPT's training cut-off in September 2021 which means it has no knowledge of the subsequent UK inflation surge. We find that GPT tracks aggregate survey projections and official statistics at short horizons. At a disaggregated level, GPT replicates key empirical regularities of households' inflation perceptions, particularly for income, housing tenure, and social class. A novel Shapley value decomposition of LLM outputs suited for the synthetic survey setting provides well-defined insights into the drivers of model outputs linked to prompt content. We find that GPT demonstrates a heightened sensitivity to food inflation information similar to that of human respondents. However, we also find that it lacks a consistent model of consumer price inflation. More generally, our approach could be used to evaluate the behaviour of LLMs for use in the social sciences, to compare different models, or to assist in survey design.

Introduction

Agent expectations have been a crucial component of many approaches in economic analysis, at least since Lucas1975. Economic agents who are forward looking, and the systems comprising them, behave substantially differently from those which merely react to current or past observations. However, the details of real-world expectations formation processes, like those of households and firms, are not yet well understood. They likely exhibit considerable heterogeneity DAcuntoWeber2024, and common modelling approaches like rational or adaptive expectations each face challenges.\\ The release of ChatGPT in the end of 2022 led to a surge in the interest and the development of large language models (LLMs) and their applications. LLMs have shown impressive results across a variety of tasks, often reaching or exceeding human performance levels. Examples are the abilities to converse and reason, recall knowledge, answer logical questions, or write computer code christie2024judgement,galatzer2024benchmark,luo2024neuro,licorish2025comparing. In short, LLMs can behave like, and master tasks comparable to, humans. Economics is no exception to this development. Early work uses LLMs as simulated economic agents who are given endowments, information and preferences so that their behaviour can be studied in various scenarios via simulations Horton2023.\\ We contribute to this nascent literature on using LLMs for economic analysis by addressing the problem of understanding agents' economic perception and expectation formation, and by providing general approaches for analysing LLM outputs in the treatment setting. In particular, we investigate the ability of a version of OpenAI's GPT model\footnote{Slightly confusingly, the series of GPT models brown2020language created by OpenAI, and used to power ChatGPT, are themselves examples of a generative pre-trained transformer (GPT), the generic term for the core technology underlying modern LLMs. We will use GPT or LLM interchangeably as representing this larger class of models. Explicitly, we are not specifically discussing the use, advantages or disadvantages of OpenAI's models versus other flagship large language models, or endorsing, or not, their use. Rather the model in this work is used as a representative model with which to test our methods and provide a subject for analysis.} to assess consumer price inflation in the present and future when provided with different price signals. Our contributions can be summarised in four distinct points.\\ One, we use a quasi-experimental design around a real-word scenario probing out-of-sample and out-of-distribution model behaviour. Following the synthetic survey setting Argyle2023simulate,arora2023probing, we replicate two samples of the Bank of England's Inflation Attitudes Survey (IAS, a quarterly survey which tracks inflation perceptions and expectations of UK households) around the peak of consumer price inflation in late 2022. While relatively brief, this shock to consumer prices was unseen alongside several dimensions. We investigate how GPT reacts to such an extreme event testing the limits of the model, but also how information leakage from this event may have affected subsequently released models.\\ Two, we provide novel machine learning interpretability tools for LLMs, which allow us to measure the effects of multiple information treatments and account for their interactions. We frame our work in the experimental treatment context. An LLM prompt consists of a synthetic persona based on the demographic characteristics of real-world survey respondents, economic conditioning information (treatment or a scenario), the IAS survey question, and instructions.\\ The discrete nature of the information treatment bits allows us to formulate a Shapley value decomposition StrumbeljKononenko2010 of the LLM's survey responses, so leveraging one of the most widely used and accepted machine learning explainability tools. The synthetic survey setting with information treatments is thus particularly well suited for this approach.\\ Three, we provide a multi-way evaluation of synthetic survey responses on the micro-, demographic meso-, and aggregate macro-level between simulated responses, actual responses, and official statistics. The IAS will serve as our “human benchmark” for model tuning and testing. The results are encouraging for the use of LLMs in our context. Aggregate response distributions can be matched quite well with the help of temperature tuning. On the meso-level GPT's outputs are often aligned with survey results along demographic groups, and are seen to be closer to official statistics than human responses. Interestingly, we find that GPT exhibits human-like biases including an oversensitivity to salient inflation components.\\ However, we also find reasons for caution when using LLMs. The micro-level correspondence between LLM and human responses is rather weak and partially unstable. This could be caused by the rudimentary economic conditioning environment and leaves plenty of scope for future research. Furthermore, the LLM we use exhibits inconsistencies which point to a lack of internal logic or the missing of a consistent world model of the concepts studied in this paper.\\ Four, we provide a discussion around the use of LLMs to inform decisions. Researchers and decision makers are ultimately interested in the behaviour of actual humans and so results from artificial intelligence (AI) experiments always require empirical validation. We provide approaches to this, but also highlight ethical considerations. LLMs can be biased, for instance with respect to demographic characteristics which we are investigating here bai2025bias. In our setting, the situation can arise where LLM outputs match either the human benchmark or official statistics, but not both. When using LLM outputs in downstream tasks the original purpose of the analysis should be considered, and a strategy to handle the inevitable trade-offs devised.\\ Overall, we believe that our experimental design and approach consisting of validation, sensitivity analysis, and model explainability, all in the context of a three-way comparison between LLM, human, and official data, can be easily transferred to both other applications and different models, and as such contributes to an accepted framework for using LLMs in social science contexts. For instance, having studied a particular model and being aware of its strengths and weaknesses then opens the door for its use within the limits drawn out by the analysis, or may raise red flags for when a model cannot be used for a specific task. \\ Hurdles for and concerns about the use of LLM for research and decision making will likely need addressing for each situation separately, because, by their nature, the precise behaviour of a LLM cannot be known a priori in a given situation. Once overcome, one may (cautiously) use a model for the simulation of human subjects in social science, economics, or financial settings. Possible applications include the generation and cross-check of official statistics, including “counterfactual statistics”, and the more efficient and effective design of surveys.\\ The remainder of this paper is structured as follows: Section\ (ref) summarises the literature on household inflation expectations, and the nascent use of LLMs in economics and surveys. Section\ (ref) describes the experimental setting and introduces the methodology. Section\ (ref) presents the main results. We conclude with the discussion in Section\ (ref).

Literature

Household Inflation Expectations and their determinants

Inflation expectations play a special role in many areas in economics. For instance, in designing monetary policy because they determine households' savings and consumption decisions via the consumption Euler equation. Inflation expectations also drive agents' wage bargaining, durable investment including housing and mortgage choices, and portfolio choices Bernanke2007. Because inflation expectations affect the decisions and actions of many economic actors, they also affect aggregate economic outcomes. Central Banks around the world actively try to manage inflation expectations, and understanding their key determinants and how they transmit into economic decisions is a topic of high policy relevance.\\ The key building block, across both theoretical and empirical strands, is the role of information frictions faced by households, with particular focus on financial literacy levels de2011measuring, cognitive abilities d2019cognitive, levels of attention sims2010rational,cavallo2017inflation, sources of information lamla2015information, subjective model of the economy macaulay2022heterogeneous, transmission of policy communication coibion2022monetary,COIBION2020103297,d2020effective,ehrmann2022central,mcmahon2023getting and personal inflation experiences.\\ Within that, most of the literature so far points to the significant role of everyday price signals observed by individuals MankiwReis2002, MackowiakWiederholt2009, CoibionGorodnichenko2015a. Households also focus on the price changes of goods they purchase frequently, such as grocery items, rather than the price changes of a representative consumption bundle when forming their inflation expectations VanderKlaauw2012, deBruin2011, DAcunto2021. Moreover, households tend to put a higher weight on positive than negative price changes when forming inflation expectations, which helps explain the persistent upward bias that has been documented in the literature Mankiw2003.\\ Over the years, a growing literature has additionally focused on documenting and explaining empirical regularities of households' inflation expectations. The most salient feature is the substantial cross-sectional dispersion, which is systematically correlated with a set of demographic characteristics arioli2017eu,del2008s,jonung1981perceived. This finding provides support to the argument that highlights the importance of individual-level drivers of inflation expectations, and suggests that traditional models of beliefs formation, which target the mean, median, or otherwise representative household expectations, fail to account for the most notable empirical regularities of the inflation expectations. For example, hobijn2009household and kaplan2017inflation study the variation in personal inflation rates experienced between households, with the latter documenting higher inflation rates among lower-income families. \\

LLMs in economic analysis

Like the personal computer, AI is special in the sense that it affects the structure of the economy but also provides a tool for analysis. The former is an active field of research (e.g. Acemoglu2022labour,hui2024labour,chen2024displacement). For us, AI, in the form of LLMs, is a tool which can be used for economic research korinek2023ai,korinek2025ai,charness2025next.\\ LLMs cannot only be used as tools to perform tasks as summarising literature, writing code, or helping with ideation, but also to simulate economic subjects themselves Argyle2023simulate,arora2023probing,manning2024automatedsocialsciencelanguage. A range of work supports the idea that on individual, self-contained questions, value judgements and actions taken by LLMs align with human behaviours across psychological, philosophical, economic and political tests Aher2022, Brookins2023, FariaeCastro2023. Horton2023 is an example of early work simulating economic agents, “homo silicus”. Human behaviour often deviates from rational choice theory as investigated by behavioural economics. In light of this, it is of interest whether LLMs and the agents they represent behave rationally or rather “human-like” bounded-rationally. Early evidence suggests that it may depend on the context and the information provided ross2024llmeconomicusmappingbehavioral,henning2025llmagentsreplicatehuman. At the same time, several studies show that LLMs know more than their human-readable outputs implies orgad2025llmsknowshowintrinsic,buckmann2025improving,buckmann2025, suggesting that researchers may not yet have learned to fully utilise LLMs in different contexts. It has also been shown that LLMs memorise information selectively and partially from their training data lopezlira2025memorizationproblemtrustllms,crane2025, making a clean out-of-sample evaluation essential for valid downstream inference ludwig2025llm.\\[.2cm] We will be using LLMs in an experimental setting simulating human survey subjects using conditioning information in a macroeconomic context Bybee2023,Hansen2025Simulating,Zarifhonarvar2025survey. Our findings relate to the strand of the literature that examines the importance of personal inflation experiences via observed price signals, meaning that we will focus on the items like food and energy that households are exposed to more frequently, and will try to establish whether there is any association with simulated survey inflation expectations. The evidence so far has suggested that grocery prices bias inflation expectations systematically to the upside DAcunto2021. The rationale being that grocery prices are more volatile than other prices and price changes typically revert more quickly. These features are one reason why many central banks focus on measures of core inflation, which exclude the price changes of groceries (and gas and energy), when examining inflationary pressures in the economy. However, by doing so, central banks risk overlooking inflation expectations picking up due to households' frequent exposure to higher than usual grocery and energy prices. The techniques presented below offer a laboratory to assess this for a given scenario.

Methodology

Prompting Strategy

We use OpenAI's GPT to answer the Bank's Inflation Attitudes Survey (IAS) in the context of the 2022 inflation surge, which peaked in October 2022. Specifically, we use OpenAI's chat completions application programming interface (API) to query its GPT models (specifically gpt-3.5-turbo-0613) programmatically creating 'synthetic IAS samples'. For price measures, we focus on consumer price index inflation including owner occupiers' housing costs (CPIH) and its subcomponents in the UK.\footnote{This is a more comprehensive measure than consumer price index inflation excluding owner occupiers housing costs (CPI). Both measures peaked at the same time in 2022.}\\ As has been proposed in previous work Jiang2022, LLMs can be conditioned on representing a particular political or demographic group. We will include gender, age, income, housing tenure, social class, UK region in our analysis which are collected on an individual bases as part of the IAS (see Appendix (ref)). Additionally, we introduce economic conditions in the form of inflation of sub-components of the CPIH, in particular food, restaurants & cafes, energy, and everything else (other).\footnote{The first three correspond to the components 01, 11.1.1, 04.5 of the CPIH, respectively. Together these have about 21% of index weight and the remainder is covered by the other component. In much of our analysis, we will treat food & restaurants jointly.} More precisely, we provide the three-month average of year-on-year (yoy) inflation of each component preceding the survey month to approximate the information set actual survey respondents have with respect to consumer price inflation. The first three components are often more volatile and seen as more salient by households, who overweight them when forming inflation expectations repec:boe:boeewp:1125. We will investigate whether GPT shows biases to any of the components we use for economic conditioning. Iterating through the survey sample at a given point in time will return UK representative synthetic sets of survey responses on inflation perceptions and expectations and allow us to investigate their drivers.\\ Perceptions here relate to the current rate of price inflation and expectations to expected year-on-year price changes in one, two, and five years.\footnote{Inflation perceptions can be seen as short-term expectations given the lag of six to seven weeks until price data are available after the end of the reference month.} We instruct the API through both the system and user prompts. The system prompt remains the same in all cases. An example prompt for inflation perceptions is the following:\\[.4cm] System: You are pretending to be the person described given your best guess as to their personal, social and economic situation.\\[.4cm] We then alter the \texttt{user} prompt as\\[.4cm] \texttt{You are male, aged 16-24, live in the North of England or Northern Ireland, are upper-middle class and are not working with an income of >£45000. You got your A-levels but not a degree and live in a house you rent.}\\[.2cm] \texttt{In the last few months, food inflation has been 17% (9.8% in restaurants and cafes), energy price inflation was about 88%. On average the rate of inflation on other goods was about %}\footnote{Numbers bigger than ten in absolute terms are rounded to the next integer. One digit is given otherwise.}\\[.2cm] \texttt{You are going to be asked questions about your perception of current and future inflation. Which of these options best describes how prices have changed over the last 12 months?}\footnote{Answer options are mapped to the midpoint of each interval or 0.5 percentage points beyond the reference value. For example, “risen by more than 15% ” is taken to be 15.5%.}

enumerate• gone down by less than 1% • gone down by 1-2% • ... [The other options in the IAS] • risen by 13-14% • risen by 14-15% • risen by more than 15%

Please choose one option, no explanation.\\[.2cm] LLMs have been shown to be sensitive to the order in which response choices are presented pezeshkpour2024large. To address potential bias from a particular choice presentation, we scramble the response options for each respondent in each sample with its own random seed.

Experimental Setting

figure[figure omitted — 355 chars of source]

In terms of survey timing, and the corresponding samples to concentrate on, we consider two samples at the peak of consumer price inflation in the end of 2022 and early 2023, namely 2022Q4 (November 2022) and 2023Q1 (February 2023). There are three main reasons for this. One, this brings us outside GPT's training period which ends in September 2021. As such it does not know about the following inflation surge or particular drivers contributing to it, like the Russian invasion of Ukraine and the subsequent spike in the energy prices among others\footnote{The Appendix contains a set of validation questions used to verify GPT's knowledge cut-off, and test whether there has been leakage into the model throughout the analysis. The answers to these questions have been stable over time with the latest test performed on 1.\ December\ 2025.}. This means that an information treatment (conditioning of GPT) related to the subsequent inflation surge can be interpreted as quasi-experimental. Two, this is around the time that aggregate inflation measures peaked but there still was uncertainty about their actual paths and how temporary this shock might have been. This means we are able to gauge the maximum impact of the inflation shock. Third, the size of the shock, as measured by economic conditioning information, was well outside the data ranges GPT would have observed in the past. This allows us to probe the model in a real, yet extreme, situation, and to evaluate its limits in the current context.\\ The experimental setting is depicted in Figure (ref), which shows year-on-year inflation of the overall index, and the subcomponents we use in our analysis. The end of GPT's training period is given by the dashed vertical line where all inflation measures have still been well within historical ranges. The two solid vertical lines correspond to the two survey samples we use. These coincide with the peaks of the different series. The details of the economic conditions used in either sample are given in Table (ref). These are broadly similar to each other, so we can feel confident carrying over insights gained in one sample to the other.\\ In line with common practice in statistical learning, we use the first sample (2022Q4) for cross-validation (CV) and the second sample (2023Q1) as our main test sample. In particular, we will calibrate GPT's temperature parameters in the range $T\in\{0,0.25,0.5,0.75,1,1.25,1.5\}$.\footnote{Theoretically OpenAI's API allows to increase the temperature up to $T=2$. However, for $T>1.5$ we observe that GPT often is either not able to follow the instructions returning random strings unrelated to the prompt, or throws an exception after a considerable delay. We consider this a model breakdown making it basically unusable.} This affects how deterministic (smaller $T$) or random (higher $T$) GPT's answers are by affecting the width of its softmax output probability distribution. This gives us some control over the moments of GPT's response distribution.

table[table omitted — 872 chars of source]

LLM treatment effects

One of the biggest advantages of the use of LLMs in the current context is that the analysis of treatment effects does not need to rely on the potential outcomes framework rubin2005outcomes, which stipulates the impossibility of observing the treated and untreated at the same time. By virtue of simulating subjects, survey respondents in our case, we are always able to observe the same subject under any treatment state.\\ Let $t_i\in\{0,1\}$ refer to a subject $i$ receiving a treatment $t$ or not. Here, $t=0$ may be a reference treatment, like a placebo. In our case, $t=1$ corresponds to including the economic conditions for the two survey samples listed in Table (ref) in the user prompt. No treatment, $t=0$, is either the omission of economic conditions in the prompt, or the inclusion of some reference values, e.g.\ the pre-training cut-off historic averages (see Table (ref)). With $x_i$ being the vector of demographic characteristics of subject $i$, their {\it individual treatment effect} simulated by GPT can be written as

equation[equation omitted — 76 chars of source]

where $g(\cdot)$ is the GPT output with the assumption that we can perform a meaningful difference operation. This will be trivial in our case as we map all survey responses to numbers. With a sample size $N$, the {\it average treatment effect} simulated by GPT is

equation[equation omitted — 116 chars of source]

There are two major concerns regarding the validity of (ref). First, the sample over which it has been calculated. Second, potential bias coming from the use of LLMs instead of human subjects. The first is common to the treatment and survey literature and is addressed by generating a nationally representative sample stemming from the underlying IAS. Addressing the second is one of the contributions of the paper, where we will perform a three way comparison between GPT, the IAS, and official out-turns.

LLM explainability

Machine learning models, including LLMs, are often subject to the black box critique: there are no clear input-output relations which can be used to explain a model's predictions based on its inputs. Such relations are simple to obtain in a linear regression model where a variable's coefficient is the measure of the input-output relationship. However, since machine learning models do not specify an explicit functional form, there is no corresponding concept of a coefficient, making model explanation, interpretation and investigation challenging. Additionally, the black box critique is particularly severe for LLMs because of the high-dimensional and unstructured nature of their inputs and outputs (such as text) and the fact that a user of a commercial LLM will not have direct access to a fitted model's (very many) parameters.\\

Shapley values are a well-established tool for explaining machine learning model predictions StrumbeljKononenko2010. Shapley values are a concept borrowed from game theory, where they describe the contributions of players to a cooperative game's group payoff. In the modelling setting, they can be used to decompose model predictions based on the contributions from each input variable. This information can then be used to identify model drivers and potentially complex non-linear relationships learned by a model. The Shapley value for a feature $k$ and observation or subject $i$ for a model $g(\cdot)$ can be written as

equation[equation omitted — 185 chars of source]

where the variable set $x'$ runs over all sets $\mathcal{C}(x)\setminus k$, which is the set of all possible variable combinations of $K-1$ variables when excluding $k$. The combinatorial weighting factor $|x'|!(K-|x'|-1)!/K!$ sums to one over $\mathcal{C}(x)\setminus k$. Eq.\ (ref) can be interpreted as the marginal contribution of variable $k$ to all possible coalitions excluding it taking all possibilities of complementing or substituting any other individual or group of variables into account.\\

We propose a general framework to address the black box critique based on explainable machine learning approaches BuckmanJoseph2022. In particular, we will adapt the Shapley value framework to the survey and treatment setting Joseph2019. The application of Shapley values to LLMs in the general case is difficult, because of the difficulty parsing inputs into discrete variables. While general text inputs are encoded into lists of discrete variables using byte-pair encoding (BPE, brown2020language) and we could perhaps extract active tokens from that through considering semantics or syntax, any approach will be complex in itself and quickly run into the curse of dimensionality given the computational complexity of $K!$ in (ref).

However, this situation is considerably simplified in the survey and treatment setting. The conditioning information (inputs) can be readily separated, ex ante, into discrete parts, e.g.\ demographic categories, and the response is a single number (inflation perceptions or expectations). This means that LLM predictions can be decomposed similarly to the conventional case of supervised learning with a single target to model.

An interesting question is what the relation between the Shapley value (ref) of a treatment $t$ and its treatment effect is. Based on (ref), these are the same for a single treatment. However, we have multiple treatments $t=(t_1,\dots,t_d)$ in our case corresponding to information on the several CPIH price components we include in the LLM's prompt, i.e.\ $t=(t_f,t_r,t_e,t_o)$ for the food, restaurants & cafes, energy, and other components, respectively. A na\"ive way for evaluating a single treatment, say, the effect of high energy prices on inflation perceptions would be to set the remaining information treatment values to some neutral values, like long-run averages or null, or excluding it altogether, and then subtract that model prediction from the full treatment case.\\ The choice of the untreated or control reference depends on the question being answered. For example, if one wants to know how high energy prices affect inflation perception {\it all other things being normal}, one can take long-run averages $\bar{t}=(\bar{t}_f,\bar{t}_r,\cdot,\bar{t}_o)$ for the other treatment values and write down the {\it na\"ive treatment effect}

equation[equation omitted — 175 chars of source]

where the only difference between the two terms on the right-hand side is in the value of the energy information treatment. The above expression is na\"ive in the sense that for a particular multi-treatment scenario, we expect the joint set of inputs to matter, i.e.\ there are potentially important interactions between the individual treatment components. Exactly this situation is taken into account in the calculation of Shapley values in in (ref) by the consideration of all possible subsets of variables not including the variable of interest, $t_e$ in the current example. Following (ref), the Shapley value for energy is calculated as

align[align omitted — 558 chars of source]

where we treated the food and restaurant & cafes components as a single variable which is either active (scenario value) or passive (average value). This can be done because of the linearity of (ref) allowing us to considerably reduce the computational burden by bunching variables.\footnote{The details regarding a low-dimensional subset of variables is often of interest in these high-dimensional settings, where one can either create lower-dimensional `factors' leveraging domain knowledge or an algorithmic approach, or take a small number of variables of interest and treat all others jointly as `other' as we do. Both approaches allow us to considerably reduce the computational complexity of (ref) while still being exact. Further approximations can be made by sampling coalitions from $\mathcal{C}(x)$.} We will follow the bunching approach treating “food” and “restaurant & cafes” as a single “food & restaurants” variable. This also is an example of how Shapley values can be used to consistently represent complex quantities by the grouping variables. \\ We will compare results from the na\"ive and Shapley treatment evaluations and see that they can differ indicating important interactions between treatment subcomponents.

Results

Temperature calibration

We investigate how the temperature parameter ($T$) affects GPT's response distribution using values in the range $T\in\{0,0.25,0.5,0.75,1,1.25,1.5\}$ for inflation perceptions in the IAS cross-validation sample (2022Q4). We track the mean and standard deviation of the resulting GPT response distributions and compare them to those of human responses. We summarise this comparison in the equally weighted loss function

equation[equation omitted — 167 chars of source]

where $MN(\cdot,w)$ and $SD(\cdot,w)$ are weighted mean and standard deviation of the input vector with survey weights $w$, and $l\in\{1,2\}$ corresponding to a linear or a quadratic loss. This allows us to (try to) match aggregate survey responses for validation. We also consider the relation between GPT and IAS on the individual or micro-level by tracking the Pearson correlation coefficient between the two for different temperature values.\\ The results for this exercise are summarised in Table (ref), with temperature values increasing from the top to the bottom. We make several observations. First, both loss measures decrease monotonically with increasing temperature, meaning that a higher temperatures leads to a better match of GPT to human responses. Second, GPT tends to predict higher inflation values than humans (the difference between means is always positive), while the width of GPT's response distribution surpasses that of human response for $T=1.25$ and above. This can be seen in Figure (ref) which shows histograms for GPT and IAS inflation perceptions for $T=0$ (upper part) and $T=1.5$ (lower part) for the cross-validation sample. Visually, GPT responses match IAS responses well for $T=1.5$. Additionally, mean GPT responses are also close to the true value of aggregate consumer price inflation in November\ 2022.

table[table omitted — 1,196 chars of source]

A third observation is that GPT-IAS micro-level correlations are stable but arguably quite low across the temperature range. This means that despite good matches on the aggregate level, GPT is not necessarily a good model for individual subject responses.\footnote{The response order randomisation on the individual level will be an additional reason for this, which we will not further investigate here.} This does, however, not mean that a lack of micro-level agreement means that higher-level results will not be accurate or useful. Collective or aggregate decision making has been shown to potentially be more accurate compared to individual estimates Galton1907VP,Krause2011SI. We will, however, see that the results for the micro-level comparison may not be robust for high temperatures, and therefore continue our analysis by considering the cases of $T=1.5$ and $T=1$ side by side.

figure[figure omitted — 433 chars of source]

The corresponding test histograms for the main sample (2023Q1) are given in Figure (ref). We see again that the $T=1.5$ GPT response distribution overlaps well with IAS responses. Additionally, the mean value is again close to the actual realised official value (solid line). We will investigate these results in more detail in the next section alongside the response profile across time horizons.

figure[figure omitted — 442 chars of source]

Time profile of inflation expectations

For both economic theory and policy making, inflation expectations are paramount. A crucial question is whether these are `anchored' at about the central bank's inflation target on longer horizons. We investigate GPT inflation expectations time profile for the IAS 2023Q1 sample with the economic conditioning given in Table (ref).

figure[figure omitted — 482 chars of source]

The time profile of aggregate GPT expectations for horizons of up to five years is shown in Figure (ref) for $T=1.5$.\footnote{The profile for $T=0$ is qualitatively very similar and is given in the Appendix.} The close match of inflation perceptions (horizon zero) between GPT, IAS, and ONS out-turn matches the lower part of Figure (ref). Both GPT (blue) and IAS (orange) expectations decrease with the horizon. However, there are a major discrepancies between future horizons from two to five years out. Though somewhat elevated, IAS mean expectations are in line with the historical distribution (IAS swath) and at the upper bounds of historically observed year-on-year inflation (CPIH swath).\footnote{Historical back data always end in Sep-21 coinciding with GPT's knowledge cut-off if not stated otherwise.} In contrast, GPT expectations stay roughly constant and elevated beyond the one year horizon. This suggests caution when using GPT infer inflation expectations.

sidewaystable\begin{tabular}{c|rrcrrccrr|ccrc} \toprule horizon & n$_{miss}$ & $MN$ & diff$_{MN}$ & SD & diff$_{SD}$ & L1-loss & L2-loss & pcc & pval & $MN_{uc}$ & effect$_{MN}$ & SD$_{uc}$ & effect$_{SD}$ \\ \midrule \multicolumn{14}{c}{GPT ($T=1.5$)}\\ \midrule 0 & 13 & 9.39 & 0.82 & 5.04 & 0.10 & 0.46 & 0.34 & -0.01 & 0.49 & 2.68 & 6.71 & 3.23 & 1.80 \\ 1 & 44 & 6.14 & 1.07 & 3.03 & -1.81 & 1.44 & 2.21 & 0.04 & 0.01 & 3.52 & 2.62 & 1.67 & 1.36 \\ 2 & 84 & 5.58 & 1.79 & 2.94 & -1.62 & 1.71 & 2.92 & 0.02 & 0.18 & 3.47 & 2.10 & 1.69 & 1.25 \\ 5 & 102 & 5.42 & 1.50 & 2.52 & -2.24 & 1.87 & 3.62 & 0.04 & 0.01 & 3.32 & 2.11 & 1.51 & 1.01 \\ \midrule \multicolumn{14}{c}{GPT ($T=0$)}\\ \midrule 0 & 0 & 10.76 & 2.19 & 3.79 & -1.15 & 1.67 & 3.07 & 0.03 & 0.06 & 3.18 & 7.58 & 1.75 & 2.04 \\ 1 & 0 & 6.26 & 1.19 & 2.09 & -2.75 & 1.97 & 4.48 & 0.05 & 0.00 & 3.23 & 3.03 & 1.13 & 0.96 \\ 2 & 0 & 5.65 & 1.85 & 1.81 & -2.76 & 2.31 & 5.52 & 0.07 & 0.00 & 3.17 & 2.48 & 1.05 & 0.77 \\ 5 & 0 & 5.36 & 1.42 & 1.76 & -3.01 & 2.21 & 5.53 & 0.07 & 0.00 & 3.10 & 2.26 & 0.97 & 0.78 \\ \bottomrule \end{tabular} \caption{Test statistics for GPT responses for different expectation horizons at $T=1.5$ (upper panel) and $T=0$ (lower panel): horizon in years (zero is current inflation perceptions), GPT missing values (NA responses), survey weighted mean ($MN$), difference to IAS mean, weighted sample standard deviation (SD), difference to IAS SD, $L1$-loss of GPT difference to IAS mean and SD (equally weighted), same $L2$-loss, Pearson correlation coefficient between GPT and IAS responses, corresponding $p$-value, weighted mean of unconditioned GPT responses, difference in mean GPT responses, weighted SD of unconditioned GPT responses, and difference in SD of GPT responses. Sources: IAS, authors' calculations.}

Comprehensive summary statistics comparing GPT and IAS at different horizons and for a high ($T=1.5$) and low ($T=0$) temperature are given in Table (ref). The right part of the table investigates the effect of including economic conditioning information in the prompt compared to only including demographics. The effect columns show the average treatment (ref) for the mean and standard deviation of the GPT response distributions. The inclusion of economic effects has shifted the mean considerably upwards. This was expected given that all treatment components have been considerably above their historic averages, such that this is a sense check.\\ Looking at the last column of Table (ref), we see that the width of the GPT response distribution has also increased as a consequence of the information treatment. This suggests that there are interactions between the demographics of the different personas given in the prompts and the economic conditioning. We will investigate this in the next section. However, we will first analyse the effect of temperature on GPT's responses in more detail.

High versus low temperature

We see in Table (ref) that there are considerable differences between the results for high and low GPT temperatures. In line with the cross-validation results, the losses (ref) are lower for $T=1.5$. in almost all case, the higher temperature setting better matches both the mean and standard deviation of the IAS response distribution.\\ However, setting a high temperature also has considerable drawbacks. First, there can be a considerable number of invalid responses, in our case particularly for longer horizons. It is not clear what is driving this result. A possible explanation may be that the questions for longer horizons are logically more challenging as they refer to year-on-year changes after a certain time has passed. This may complicate following the instructions.\\ Second, the micro-level relation between GPT and IAS responses as measured by Pearson correlations is weak, volatile across horizons, and even partly breaks down. In contrast, the GPT-IAS relations are considerably stronger and actually increasing with the horizon for the zero temperature setting.\\ Lastly, the cross-horizon correlations are mostly weak for the high temperature case with patterns very different to human responses. The corresponding cross-correlations are listed in Table (ref). GPT at zero temperature shows patterns much more similar to those of human responses, especially between expectations at longer horizons.\\ Because of these observations, we will focus on the $T=0$ case below. This has the additional advantage that the interpretation of results does not carry uncertainty from the temperature setting which still can be investigated separately.

table[table omitted — 872 chars of source]

Model time trends

Consumer price inflation and the “cost of living crisis” have been persistent topics of debate in the UK since the 2022 inflation spike\footnote{\url{https://commonslibrary.parliament.uk/research-briefings/cbp-9428/}}. That is, inflation has become a more salient topic in the public discourse. At the same time, LLM fine-tuning using reinforcement learning based on human user preferences has become common ouyang2022training. However, such fine tuning can lead to an unpredictable level of information leakage from beyond the stated training cut-off. To test whether there may have been such leakages or salience of high inflation in ChatGPT's model family, we test unconditioned inflation perceptions for the main sample for different models, where we remove the economic conditioning information from the prompt only leaving demographics and the survey instructions. \\ The results of this exercise are summarised in Figure (ref), which plots average inflation perceptions against the release dates of the different models. We see that there is a clear time trend of an increase of about 3 p.p. unconditional inflation perceptions per year potentially creating large biases in later models as inflation numbers quickly fell after the spike.\\ Models with the same stated knowledge cut-off in September\ 2021 but later release dates seem to be affected by this trend suggesting the use of reinforcement learning and subsequent information leakage. Models released in early 2023 seem not to be affected by this, which highlights the value of the experimental setting for model validation and testing presented here.\\ This also highlights challenges for the use of closed-weight proprietary models like those from the ChatGPT family. The subsequent analyses will provide ways to either account for such “offsets” or to analyse the salience of different pieces of information to a model.

figure[figure omitted — 444 chars of source]

Demographic drivers

When driven by food and energy price growth, high inflation may be more concerning for lower income households, as they spend a larger proportion of their consumption basket on necessities like food and energy. By similar arguments, the inflation rate experienced by different demographic groups may vary.\\ Now, for a category $D_c$ we hypothesise the relative ordering of experienced inflation values $\mathcal{O}_c=\{\pi_{c_1}>\dots>\pi_{c_j}>\dots>\pi_{c_K}\}$ for $j\in\{1,\dots,K\}$ different classes, for example based on assumptions about their respective consumption baskets in a given economic environment. Taking $j=1$ as the reference class for each $D_c$, we can formulate the following hypothesis tests for jointly assessing $\mathcal{O}_c$ for categories $c\in\{1,\dots,C\}$ based on the response model

equation[equation omitted — 182 chars of source]

where $s$ denotes the source of the $y_i$ (IAS or GPT in our case), $b$ is a constant, and $d_{ic_j}$ is a vector of dummies encoding the demographic profile of subject $i$. The coefficient vector $\beta=(\beta_{1_2},\dots,\beta_{C_K})'$ captures the joint relation between demographics and survey responses given a scenario. The interpretation of the elements of $\beta$ is intuitive, namely by how many percentage points (p.p.) inflation perceptions or expectations are on average higher or lower (depending on the sign of an element) if a subject belongs to a certain demographic group, i.e.\ $d_{ic_j}=1$.\\ An appealing property of (ref) is that it adjusts for a different location and scale of the response distribution $y_s$ via $b$ and the magnitude the $\beta$s, respectively. This means that, despite potentially poorer aggregate fits of a response distribution to a benchmark, we still can make inference about its drivers. This will be the case for the GPT ($T=0$) cases, which showed poorer aggregate matches to the IAS distribution but stronger correlations on the micro level (see Table (ref)).\\ In the context of the main sample scenario in Table (ref), we hypothesise that the following reference groups within our demographic categories have experienced the highest levels of inflation: income: less than £9999 (lowest income), housing: council house, age: 16-24 (youngest), social class: working class, education: GCSEs but not A-levels (lowest formal education), region: Scotland. Consequently, when fitting (ref) we expect all components of $\beta$ to be {\it negative}.\\

Inflation perceptions

Table (ref) summarises the results for the IAS and again for GPT for $T=1.5$ and $T=0$ for inflation perceptions. We also include estimates of the actually experienced inflation by the different demographics categories at the time of the survey based on official statistics (ONS reference). Focusing on the $T=0$ case, we see that the majority of the coefficients are indeed negative and statistically highly significant indicating that GPT's responses are in line with the economic intuition guiding the choice of the reference classes, i.e.\ our hypotheses about which demographic groups may be more affected by the inflation spike around the time of the survey.\\ We can further validate this intuition and GPT's responses by comparing the $\beta$ estimates with the ONS reference. We indeed see that GPT responses are very much aligned with actual realisations for most categories when comparing directions: most entries are negative.\footnote{A notable exception is the coefficient of the ONS reference point for the lowest income group which is most likely is due to changes in the definition of lowest income group within the IAS survey itself over the years.} The exception to this are regional differences, where GPT thought they would be large, while there where almost none, perhaps because energy prices are regulated and determined mostly on the national and not regional level. \\ Comparing GPT responses to the IAS, we see that most coefficients are negative again. However, there is a major discrepancy between the two for age. While GPT thinks that there is a mostly negative and marginally increasing effect with age, IAS respondents believe that the effect is clearly positive and strongly increasing. A possible effect is the observed cohort effect of inflation perceptions MalmendierNagel2013: older cohorts may have experienced more “scarring” by periods of high inflation and, as a consequence, have on average more pessimistic views on inflation in an heightened inflation environment.\\

table[table omitted — 3,592 chars of source]
figure[figure omitted — 797 chars of source]

The demographic regression results are graphically summarised for the first three categories of Table (ref) in Figure (ref). The left-hand side panels compare IAS and GPT coefficients against ONS out-turns as measured by the difference against the reference class in p.p.. The right-hand side panels compare GPT and IAS coefficients against each other on a scatter plot: the more dots align along the diagonal as summarised by the fit line, the more similar human and GPT responses are alongside this demographics dimension.\\ We see that there is reasonable agreement between GPT, IAS, and ONS out-turns for income and especially for housing. GPT's response patterns closely match that of the outcome. This is interesting and encouraging as there are various channels how different housing tenures can affect realised inflation values suggesting that GPT can be useful in modelling complex and potentially unknown relationships.\\ However, our analysis also suggests caution when looking at the results for income in the bottom part of Figure (ref). Here the three-way comparison is far less aligned. As pointed out before, GPT and IAS response patterns are mostly opposed to each other with some similarity in trends with increasing age. The alignment with the out-turn is weak at best.\footnote{The corresponding figure is given in the Appendix for $T=1.5$, where see that micro-level relations generally weaken with higher temperature.}\\ This example raises the important question how to interpret and use GPT as a modelling tool: should we prefer GPT to be more closely aligned with the human sample we emulate or with the survey quantities we elicit? We will address such concerns in the discussion at the end of the paper.

table[table omitted — 3,601 chars of source]
figure[figure omitted — 819 chars of source]

Inflation expectations

Perceptions are easier to compare with realised values as there is less noise affecting the analysis between eliciting survey responses and measuring the reference quantities. However, many economic concepts relate to expectations at longer horizons, often at one quarter or one year out. We therefore repeat the exercise from the previous subsection for inflation expectations at the one-year horizon. The results are summarised in Table (ref), where we will focus on the $T=0$ case. \\ Looking at explanatory power ($R^2$), it is interesting to observe that GPT now shows more alignment with demographics on average compared to the IAS. The comparison with realised values (last column of Table (ref)) is less interesting now because aggregate inflation decreased markedly from 9.2% (February 2023) to 3.8% (February 2024), and the variation within demographic groups is much smaller compared to when inflation peaked.\\ The patterns of signs persist between inflation perceptions and one-year expectations: most signs are negative, again with exception for age for the IAS. Looking at individual demographics in Figure (ref), we see that there is good alignment between GPT and the IAS for income, housing, and social class. This congruence is particularly strong for the income distribution, a main quantity of interest when investigating heterogeneity in economic analyses.

Economic drivers

We investigate how the details of the economic conditioning affect GPT's ($T=0$) predictions for inflation perceptions in the main sample.

Marginal sensitivity

Aggregate inflation scales approximately linearly with inflation of the subcomponents of the indeces.\footnote{Approximately because the basket-weighted sum of component price growth does not need to be the same as overall index growth but will be a good approximation in most cases.} As such, it is an interesting question whether GPT's inflation perceptions scale linearly with the different components of the economic conditioning information treatment. To do so, we track the na\"ive treatment effect (ref) for food & restaurants, energy, and the main other component with the offset being the respective aggregate prediction value at the zero input for that component.\footnote{For example, we subtract $MN(g(t=(0,0,\bar{t}_e,\bar{t}_o);x_i);w)$ for the food & restaurants component.} In each case, we vary the respective component covering the historically observed values including the inflation spike in 2023Q1.\footnote{For food & restaurants, we varied food price inflation keeping the ratio between this and restaurants & cafes inflation constant at the value of the main scenario.}\\

figure[figure omitted — 739 chars of source]

The results of this exercise are shown in Figure (ref). We again make several observations. First, GPT's aggregate perceptions indeed scale linearly with subcomponent inflation over large parts of the input space. However, this scaling does not consistently cover the historically observed input ranges (shaded areas) for some components (energy and other), but may extend far beyond them (food and other). GPT's behaviour is particularly puzzling for the main other component. It is mostly insensitive to this component between zero and two percent inflation but then scales linearly far beyond historically observed values.\\ Second, the slopes within the linear response ranges of GPT for the different components can be directly compared to the corresponding basket weights. This comparison is shown in the upper part of Table (ref), where the ratio between the slope and the weight measures whether GPT over- or under-reacts to a component corresponding to a value bigger or smaller than one, respectively. GPT strongly over-reacts to information from the food & restaurant and the energy components. This behaviour may be said to be similar to that of humans, since it is known that they over-rely on salient components such as groceries to make judgments about inflation DAcunto2021. Interestingly, GPT's responsiveness to the main other component is very much proportional to its share in the basket beyond the 2% point in the input range, going far beyond the historically observed range. This suggest that GPT can partially be used to make realistic inference about inflation once corrections for the observed bias at the origin are included.\\ Lastly, the actual and real scenario from the main sample is rather extreme. The vertical purple lines which mark the economic conditioning are far beyond the historically observed data ranges and also outside GPT's linear scaling range for food & restaurants and energy. It is somewhat interesting that GPT fails to extrapolate monotonically with respect to food & restaurants inflation. As such, our analysis sketches out the boundaries within which GPT's responses may be judged consistent but also highlights that GPT does not seem to have a complete internal model of consumer price inflation.

table[table omitted — 875 chars of source]

Shapley decomposition

In each panel of Figure (ref), the value on the vertical axis where the 2023Q1 conditioning line intersects with the GPT response curve measures the na\"ive treatment effect (ref) for that component. As discussed, this decomposition does not take into account possible interactions between the treatment components. These are captured by the Shapley components (ref). A comparison between the two indicates how important such interactions are. Both are shown in the lower part of Table (ref). The Shapley decomposition measures all components against the zero baseline, i.e.\ where non-active component values are set to zero.\\ The ranking of components in the main scenario is the same. However, we do see sizeable changes for the importance of the energy treatment, which increases by more than 50% in absolute terms for the Shapley case. This suggests that GPT's inflation perceptions are particularly elevated due to jointly high values of food and energy price inflation. This is confirmed by considering a scenario with only these two components at non-zero values as in the main scenario. The effect for this scenario is 11.09% of aggregate GPT inflation perceptions. The remainder compared to the full main scenario is accounted for by the other component and its interactions with the other two components.\footnote{The fact that neither the na\"ive nor the Shapley decomposition with zero values for inactive components perfectly sums up to the aggregate total of inflation perceptions in the main scenario suggests the existence of additional, though minor, inconsistencies in GPT's inflation attitudes.}\\

figure[figure omitted — 376 chars of source]

The relevance of a component for aggregate inflation is the product of its magnitude and weight. For instance, energy inflation was very ran very high in the main scenario while its weight is moderate at about 4% in the consumption basket. By comparing GPT aggregate components with realised contributions we can assess whether it over- or under-weights individual components for a given scenario independent of its response function. This comparison is shown for the main scenario in Figure (ref). In line with previous results, GPT heavily over-weights food & restaurant inflation and under-weights the main other component. Its perception of the relevance of energy price inflation may be said to be roughly in line with its actual importance.\\ Finally, an advantage of the GPT decompositions approaches presented here is that we can make such inference in the first place. To get to something comparable for the actual survey, we need to build a model of how human respondents react to economic conditions. A possible approach to do so is to regress aggregate survey responses on the economic conditions at the time of the survey across a certain time period. We can then measure the importance of each component for a given scenario for this model by multiplying the coefficient values with the respective scenario input values, which corresponds to the Shapley value of a variable in a linear model.\\ Modelling results for different fit ranges are shown in Table (ref). We see that such relations can be rather unstable over time. Only the relevance of the food & restaurant component is stable over time.\footnote{All coefficients turn insignificant for a short modelling period of only five years.} The Shapley decomposition corresponding to the first row of Table (ref) is shown in orange in Figure (ref). A possibly robust statement we can make from this comparison is that humans and GPT overreact to food & restaurant inflation in line with previous findings providing some trust in the use of GPT in this context.

table[table omitted — 779 chars of source]

Discussion

We investigated an LLM's ability to form perceptions of current and expectations of future consumer price inflation based on a set of economic conditioning information in a representative survey setting. Our analysis provided a set of tools to assess model sensitivity and explainability with respect to the conditioning information. We found that the used GPT model can reproduce key characteristics of human responses and official statistics. This is an impressive feat, as GPT models have not been explicitly designed to perform the tasks investigated here and no specific fine tuning was applied to the model: we achieved partially satisfactory performance out-of-the box for a relatively complex problem.\\ However, our analysis also revealed several shortcomings of the used model in the given context. Similarly to human responses, GPT overweights the importance of salient price components, like those for food. This could be an advantage depending on the application. However, it also partially shows no sensitivity to the main other component. Its sensitivity exhibits strong non-linearities even within historically observed ranges of a kink-type nature for no apparent reason. As such, GPT does not appear to have a fully consistent world model for consumer price inflation. Additionally, GPT responses' correlation with those of actual survey respondents is relatively weak. This latter observation may largely be caused by the very limited conditioning information we provided to the model, which will certainly be insufficient to capture the idiosyncratic state of human respondents well. Achieving better alignment on the micro level is an interesting research problem in itself leaving plenty of scope for future work.\\ We conclude our analysis with a discussion of ethical consideration when using LLMs to simulate or investigate human behaviour. One concern may be the simulation of individual persons, here survey subjects, without their explicit consent. This aspect needs to be addressed by data governance and is independent of the model type being used. The same concern applies to a linear regression model as much as it does to a more complex neural network, such as an LLM.\\ However, the more complex nature of LLMs leads us to a second concern which is more difficult to address. Namely, the black box critique of machine learning models, which is even more aggravated in the LLM case compared to supervised learning approaches. Model complexity (like the number of parameters in the neural network) is orders of magnitude larger compared to those used in supervised problems. While the training data (often vast corpora of unstructured data) are mostly unknown and inaccessible to the user, as are the details of the training algorithms used to build the model.\\ This second point puts more burden on model validation and testing. Much of this paper was dedicated to just that. However, as we have seen the comparison with human benchmark data (actual survey responses) and official statistics can still lead to difficult questions, because these three may not be aligned. LLM outputs may match human benchmark data, but fail to reproduce other statistics, or vice versa. There are many reasons for why this can happen, and it may not be possible to address these in the analysis leaving us with the difficult decision of whether and how to use LLM results, for instance to inform decisions.\\ Considering the goal of the analysis can serve as guidance here. If the aim is to replicate or analyse official statistics, such as macroeconomic aggregates, mismatch on more disaggregate levels can be tolerated if validation has been sufficiently satisfactory by some criteria. As pointed out, mismatch on the micro level does not preclude match on the aggregate level, and it may actually be a source of better aggregate performance.\\ However, if the goal of the analysis is to infer characteristics of humans or subgroups of the population, like demographic groups, potentially informing decisions affecting them, mismatch may not be tolerable if it cannot be accounted for in some way, such as via bias correction when using LLM outputs in a downstream model ludwig2025llm. A decision on the usefulness of LLM outputs will most likely face trade-offs even after validation and possible corrections, necessitating a robust governance structure for the use of AI in decision making.\\ Finally, most of our results where obtained for a single model without a guarantee that these generalise to other models or settings. Given the complexity and diversity of LLMs, it is therefore essential to validate a model whenever either the model or setting changes as long as we do not have a general understanding of LLM behaviour in a given context. However, if validation has been satisfactory, the use of LLMs (or a particular LLM) opens new doors for research given their ability to simulate complex entities, like survey subjects, in-situ, cheaply and flexibly in a way which would be impossible in real-world settings.