Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
78,239 characters · 11 sections · 134 citation commands
Estimation and inference in models with multiple behavioural equilibria
\thispagestyle{empty}
The rational expectations hypothesis, proposed in the seminal work by Muth1961 and Lucas1972, is a benchmark for the study of macroeconomic models, as it allows for a dynamic system in which agents' expectations are incorporated in a manner internally consistent with the model itself. Starting with marsar:89 and Sargent93,Sargent99, however, the orthodox view of agents forming full-information rational expectations has been increasingly questioned; see, e.g., evans:01, MankiwReis2002, and Bernanke07. With a particular focus on inflation expectations, a large body of recent empirical studies that exploit survey data corroborates the finding that agents deviate from the benchmark of full rationality; see, e.g., CoibionGorodnichenko2012, BachmannBergSims15, mn:2016, CoibionGorodnichenkoRopele20, coibion:20, and BordaloGennaioliEtAl2020.
Several deviations from the pure rational expectations paradigm have since been proposed. Notable examples include models with frictions, such as sticky information (e.g., MankiwReis2002) and rational inattention (e.g., Sims2003). Another popular route, drawing on bounded rationality and popularized by, among others, bray1982learning, bray1986rational, marsar:89, and evans:01, treats agents as statisticians who form expectations using simple, possibly misspecified, econometric models.
This is also the approach taken in this paper, which closely follows the specifications studied in HommesSorger1998, Lansing2009, hz14, and HommesMavromatisOzdenZhu2023. In particular, we consider a New Keynesian Phillips curve (NKPC) in which agents update their beliefs by recursively estimating the parameters of a first-order autoregression, AR(1), that they (incorrectly) perceive to be the correct model of the inflation process. As hz14 show theoretically, this framework provides a parsimonious account of several features of observed inflation dynamics.
Importantly, because agents remain even in the long run boundedly rational without correctly perceiving all stochastic aspects of the model, multiple so-called {\it behavioural equilibria} can coexist: self-confirming inflation paths that are fixed points of the perceived law of motion, rather than rational-expectations solutions of the full model. Recently, HommesMavromatisOzdenZhu2023 and davide:25 take variations of this model to the data and, using Bayesian techniques, show its usefulness both in-sample and out-of-sample.
As we discuss in detail below, however, because of the model’s nonlinear features and the presence of multiple behavioural equilibria, a priori little is known about the properties of (frequentist) estimators for the structural parameters in these models. A growing literature on the econometrics of models with adaptive learning shows that the statistical analysis is indeed a challenging task; see, e.g., chev:10, adam:16, chev:17, chrismass:18,chrismass:19, mayer:22,mayer:23, christiano:24, and mm:25.
To close this gap, we {\it first} provide a thorough statistical analysis of a plain-vanilla NKPC model with constant-gain learning. Specifically, we $(i)$ establish geometric ergodicity of the dynamic system; $(ii)$ prove strong consistency and asymptotic normality of the nonlinear least-squares (NLS) estimator of the structural parameters and the associated behavioural equilibria; and $(iii)$ provide detailed guidance on conducting inference. Substantively, this delivers an estimable NKPC with learning and a characterization of the location of its behavioural equilibria. Methodologically, we develop a general strategy for establishing ergodicity and NLS inference in nonlinear learning models with potentially multiple equilibria, and highlight when and why standard limit theory breaks down. In doing so, we draw on recent results for nonlinear time series models to derive the model’s probabilistic properties (chot:19) and to justify point estimation (francq:24). Moreover, several less-standard results emerge that connect to the broader econometric literature. For example, akin to hansen:1996, we show how to conduct inference when certain parameters (viz., the learning gain) are not {\it jointly} identified. In this case, following saikkonen:95 and seo:11, we can still verify $\sqrt n$-consistency of the remaining structural parameters. Our analysis of the (multiple) equilibria is also conceptually related to kasy:15, who develops a nonparametric procedure for inference on the number of roots of an unknown function. In contrast, we work in a parametric setup where the equilibria are the roots of a polynomial function that is known up to a finite dimensional parameter. This allows us to study the asymptotic distribution of equilibrium locations, its number, and uniform inference on the complete function. When inference concerns the (multiple) equilibria, convergence can be slow, akin to the second-order identification mechanism in dovonon:18. {\it Secondly}, we verify the theoretical findings in finite samples, using both simulated and real-world data. The empirical application to US data illustrates how our methods can be used to characterize uncertainty about equilibrium locations and learning dynamics in practice.
The remainder of this paper is organized as follows: Section (ref) introduces the model. The time series properties of the resulting data generating process (dgp) are established in Section (ref). Sections (ref) and (ref) discuss, respectively, consistency and asymptotic normality of the NLS estimator. The asymptotic distribution of the resulting estimator of the behavioural equilibria is derived in Section (ref), while Section (ref) is devoted to various inference procedures. Finite sample evidence in the form of a Monte Carlo experiment and an estimation exercise using US data are presented in Sections (ref) and (ref), respectively, before Section (ref) concludes. All proofs are delegated to the appendix.
Consider a standard version of a forward-looking NKPC:
where $t=1,2,\dots$ indexes time, $\pi_t$ denotes inflation, $\pi_{t+1}^e$ is the agents', not necessarily rational, expectation of future inflation, the driving variable $y_t$ is a proxy for real marginal costs (e.g. an output gap measure or real unit labour costs), and $u_t$ is an error term (see, e.g., Woodford2003 for a theoretical discussion or MavroeidisPlagborgMollerStock2014 for a more empirical treatment). Following, among others, Lansing2009 and hz14, the driving variable is assumed to follow an AR(1) process,
with intercept $a$, persistence parameter $\rho$, and innovation $\varepsilon_t$. The parameters $\rho$ and $\delta$ are in the interval $(-1,1)$, whereas the variances of the shocks $\varepsilon_t$ and $u_t$ are given by $0<\sigma^2_\varepsilon < \infty$ and $0<\sigma^2_u < \infty$, respectively. The relevant economic parameter is $\psi \in {\mathbb R}$, usually called the slope, that measures how inflation responds to changes in economic `{slack}'.
Under the full information rational expectations (RE) hypothesis, agents have complete knowledge of the model structure (e.g. they know the functional form of Eq. (ref)) and its history and make optimal use of it by setting $\pi_{t+1}^e = \textnormal{\textsf{E}}[\pi_{t+1} \mid \mathcal{F}_t]$, where $\mathcal{F}_t \coloneqq \sigma(\{(u_i,\varepsilon_i): i \leq t\})$. It can easily be shown (see e.g. hz14) that the RE equilibrium is given by a linear projection of $\pi_t$ on $(1,y_t)^{\textnormal{\textsf{T}}}$:
The RE paradigm, although providing an important benchmark, has been challenged over the years. One popular alternative, proposed by the macroeconomic learning literature, is to assume that agents themselves act like statisticians when forming their inflation expectations. In this literature, agents usually know the functional form of the RE equilibrium (ref) but not its parameters, which they estimate; hence the term bounded rationality (see, e.g., evans:01 and evan:20 for overviews).
Here, we deviate from complete rationality a bit further: Following hz14, we assume that agents even lack knowledge of the functional form. Instead, expectations of tomorrow's inflation are formed according to a perceived law of motion (PLM) that incorrectly stipulates AR(1) dynamics for $\pi_t$:
with $e_t$ denoting an error, while $(\alpha, \beta)^{\textnormal{\textsf{T}}}$ are parameters unknown by the agent, who has to estimate them. The PLM in Eq. (ref) is thus an auxiliary model employed by the agent to update her beliefs about future inflation. Note that, similar to the discussion in evans:01, the PLM is misspecified in the sense that it does not nest the RE equilibrium Eq. (ref).
Two questions arise: (1) How do agents forecast? and (2) What do agents actually forecast? To address the first question, assume that expectations are formed at the end of period $t\!-\!1$ using a mean-squared-error-optimal two-period-ahead rule based only on information available at $t\!-\!1$\footnote{This mechanism--where forecasts are based on lagged rather than current information--is standard in the adaptive-learning literature (e.g., DeGrauweMarkiewicz2013 or LansingMa2017) and is consistent with the timeline of events in elias2016; see also adam2003learning for a theoretical justification.}. This implies $\pi_{t+1}^e = \alpha + \beta^2 (\pi_{t-1}-\alpha)$ and, therefore, the implied actual law of motion (ALM) for inflation is
which would obtain {\it if} agents knew the parameters $(\alpha,\beta)^{\textnormal{\textsf{T}}}$ governing their forecasting rule in Eq. (ref).
To answer the second question, hz14 introduce the notion of {\it behavioural learning equilibria}. With agents misspecifying the functional form (note that they disregard the regressor $y_t$ entirely), hz14 discipline behaviour from becoming overly irrational by imposing that ($a$) the unconditional mean and ($b$) the first-order autocorrelation of $\pi_t$ are correctly perceived. These two restrictions pin down the equilibrium values of $(\alpha,\beta)^{\textnormal{\textsf{T}}}$. In particular, hz14 equate the PLM and Eq. (ref) on their unconditional moments, which yields that expected inflation under the ALM satisfies
whereas the first-order autocorrelation $\beta$ solves the following quartic equation
where $\lambda \coloneqq (\delta,\psi,\rho,\sigma_u^2,\sigma_\varepsilon^2)^{\textnormal{\textsf{T}}}$. Put differently, the unconditional mean and autocorrelation agents believe in must equal the unconditional mean and autocorrelation actually generated by the NKPC under those beliefs. It is worth noting that, in this case, expected inflation corresponds to the rational expectation case. Furthermore, Eq. (ref) might deliver multiple solutions, representing different expectational equilibria, depending on the value assigned to the structural NKPC parameters. For instance, the solution is unique when $\psi>0$ is large enough and $\sigma^2_u$ is small. As discussed by hz14, multiple equilibria correspond to different long-run inflation persistence regimes, and policy or shocks can move the economy between them. Specifically, drawing on the findings presented in blanchard:16 and JorgensenLansing2025, an equilibrium characterized by a low $\beta$ parameter implies strongly anchored expectations, a scenario that aligns with the original specification of the Phillips curve, namely, one defined in terms of inflation levels and the output gap. Conversely, a high $\beta$ parameter would be more consistent with an accelerationist formulation of the curve.
Agents typically do not know the parameters $(\alpha,\beta)^{\textnormal{\textsf{T}}}$ of their forecasting model Eq. (ref). Instead, and in line with the adaptive learning literature (evans:01), we assume agents employ a recursive updating scheme to estimate them; formally, a stochastic approximation algorithm (ben:90). In particular, agents are presumed to use a constant gain version of the sample autocorrelation learning (SAC) algorithm proposed by hz14 (see also HommesMavromatisOzdenZhu2023) to estimate the subjective parameters $(\alpha,\beta)^{\textnormal{\textsf{T}}}$:
for some so-called `gain' or `learning' parameter $\gamma > 0$ and initial value $\pi_0=\alpha_0$. Note that $\gamma > 0$ implies that more weight is attached to recent data and that the recursive estimators do not converge in probability but, as discussed below, converge in distribution to non-degenerate weak limits. If, in contrast, $\gamma \rightarrow 0$, then the adaptive learning estimates can converge under mild conditions to one of the possible behavioural equilibria computed from Eq. (ref), and in this respect the equilibrium is called locally stable. Constant-gain learning with $\gamma>0$ implies, in turn, persistent deviations from rationality and is often employed in empirical applications where it is the preferred choice due to its tracking ability in nonstationary environments (see, e.g. milani:07 or chev:10).\footnote{markiewicz2014adaptive find that constant-gain least squares typically yields lower mean squared forecast errors than decreasing-gain learning, particularly for macroeconomic series such as inflation, GDP growth, and unemployment.} Importantly, in their empirical estimation of an extended version of the NKPC discussed here and in hz14, HommesMavromatisOzdenZhu2023 employ the constant gain specification of Eq. (ref).
To sum up, under learning, the resulting data generating process (dgp) is thus given by Eqs. (ref) (for $y_t$) and (ref) (for $\alpha_t$, $\beta_t$) in conjunction with the true ALM:
The unknown parameters in this model are $\theta \coloneqq (\gamma,\delta,\psi)^{\textnormal{\textsf{T}}}$ together with $(a,\rho,\sigma_u,\sigma_\varepsilon)$, that determine the potentially multiple equilibria $\beta$, all of which we aim to estimate. This task is nontrivial because the dgp defined by (ref), (ref), and (ref) is highly nonlinear. \footnote{Note that, although the dgp can be seen as an `{\it observation driven}' model in the sense of cox:81 (see also blasques:24), commonly used results in that literature to establish the time series properties do not apply in a straightforward way. For example, as discussed in more detail below, results based on stochastic recurrence equations like strau:06 require a Lipschitz condition that fails here.} Nevertheless, as discussed in the next section, it can be cast as a non-linear state-space model, in the sense of meyn:12, and shown to possess sufficient regularity needed for our later analysis.
Before turning to estimation and inference for the structural parameters, we first establish the time series properties needed to ensure that this exercise is meaningful. In particular, we establish that the model, summarized by the ALM in Eq. (ref), viewed as a nonlinear Markov chain, is geometrically ergodic. This matters because, irrespective of initialization, this guarantees that the chain converges to a unique stationary distribution and insures the {\it invertibility} of the time series model (see, e.g. strau:06 or bla:18). Moreover, this notion of ergodicity justifies the use of law of large numbers (LLN) and central limit theorem (CLT) results for sample averages and smooth functionals which are key ingredients for our later analysis of estimation and inference (for discussions of these concepts in the context of non-linear time series models see, e.g. carrasco2002mixing or kristensen2005geometric).
To set the stage, note that we can represent $\beta_t$ also recursively i.e. \[ \beta_t = \beta_{t-1}+ \frac{\gamma}{r_t} ((\pi_t-\alpha_{t-1})(\pi_{t-1}-\alpha_{t-1})-\beta_{t-1}(\pi_t-\alpha_{t-1})^2), \] and \[ r_t = r_{t-1}+\gamma ((\pi_t-\alpha_{t-1})^2-r_{t-1}). \] We follow a convention from hz14, and assume that the agents' initial value equals the observed inflation; i.e. $\alpha_0 \stackrel{!}{=} \pi_0 \eqqcolon \operatorname*{\mathfrak{a}} \in \mathbb{R}$. Thus setting $r_1 = \gamma(\pi_1-\alpha_0)^2$, ensures $\beta_1 = 0$ for any $r_0, \beta_0 \in \mathbb{R}$. Now, using these two recursions, we define to that end the $5\times 1$ state vector $s_t \coloneqq (\pi_t,\alpha_t,y_t,\beta_t,r_t)^{\textnormal{\textsf{T}}}$. In Appendix (ref) we show that the dynamic system $s_t$ can be viewed as a nonlinear Markov recursion $s_t = \mathcal{G}(s_{t-1},v_t),$ for a continuous mapping $\mathcal{G}(s,v)$ and innovation vector $v_t \coloneqq (u_t,\varepsilon_t)^{\textnormal{\textsf{T}}}$. It is worth noting that $s_t$ belongs to the class of nonlinear state space models, as defined in meyn:12 and that, by Assumption (ref) below, $s_t$ is weak Feller (meyn:12).
We cannot, however, directly rely on off-the-shelf results typically used in the nonlinear time series literature to verify geometric ergodicity of $s_t$. For example, the otherwise very general results in tong1990non do not apply because, among other reasons, the underlying skeleton difference equation fails to be Lipschitz, owing to the highly nonlinear dependence induced by $\beta_t$. For a similar reason, we cannot invoke strau:06 or the ergodicity result from ds:93 used, for instance, by adam:16; see Remark (ref) for details. Instead, we derive the stochastic properties of the model from first principles, following the exposition in chot:19.
In doing so, we impose the following assumption:
While some mild deviations from the {\sf IID} assumption might in principle be possible, existence of lower semicontinous error-densities with full support is essential for both our ergodicity proof as well as parameter identification. Under Assumption (ref), we can show that $s_t$ explores all meaningful regions of the state space ($\varphi$-irreducibility) and avoids periodic behaviour (aperiodicity). To obtain ergodicity we also need to preclude divergence to infinity. This can be ensured by a so-called {\it drift condition}, for which we require the following parameter constraints.
The bounds $|\delta| <1$ and $|\rho|<1$ guarantee stability and ergodicity of the dynamic system, while $\psi$ can be left unrestricted for many of our core results. The gain is by construction positive, while $\gamma < 1$ is not restrictive as it covers the range of estimates of $\gamma$ typically found in empirical work (see, e.g. berardi2017empirical). We occasionally impose additional parameter restrictions to obtain further insights: for instance, $\rho > 0$ ensures existence of at least one equilibrium, $\delta \neq 0$ guarantees that $\gamma$ is {\it jointly} identified with $(\delta,\psi)$, while $\psi > 0$ is imposed to obtain valid uniform confidence bands. In line with economic theory, hz14 assume $\rho \in [0,1)$, $\psi>0$, and $\delta \in [0,1)$, where $\rho > 0$ is imposed to ensure existence of at least one equilibrium.
Intuitively, geometric ergodicity ensures exponentially fast forgetting of initial conditions, a crucial requirement for valid estimation and inference covered in the following section.
Observing a sample $\{\pi_t,y_t\}_{t=1}^n$ of length $n$, we estimate the $3 \times 1$ parameter vector $\theta \coloneqq (\gamma,\delta,\psi)^{\textnormal{\textsf{T}}}$, using the nonlinear least squares (NLS) estimator $\theta_n \coloneqq (\theta_{\gamma,n},\theta_{\delta,n},\theta_{\psi,n})^{\textnormal{\textsf{T}}}$, which minimizes, over a suitable parameter space $\Theta \subset {\mathbb R}^3$, the sample objective
where, for a given initial value $\operatorname*{\mathfrak{a}}$, $f_t(\cdot,\operatorname*{\mathfrak{a}})$ is the nonlinear regression function
Importantly, by Proposition (ref) the choice of initial condition is asymptotically irrelevant: in terms of bougerol:93, strau:06, or blasques:22, the {\it filter} forgets $\operatorname*{\mathfrak{a}}$ at a geometric rate and we suppress in the following the dependence on $\operatorname*{\mathfrak{a}}$.
Given $\theta_n$, we can estimate $\sigma_u^2$ using $\hat\sigma_u^2 \coloneqq \frac1{n}Q_n(\theta_n)$. Moreover, let $\rho_n$ and $\hat\sigma_e^2$ denote the OLS estimators of the AR(1) process in Eq. (ref) and its innovation variance $\sigma_e^2$, respectively. We can then define for $$\lambda_n \coloneqq (\theta_{\delta,n},\theta_{\psi,n},\rho_n,\hat\sigma_u^2,\hat\sigma_\varepsilon^2)^{\textnormal{\textsf{T}}},$$ the estimator(s) of the (potentially multiple) equilibria
implicitly given as the solution of Eq. (ref).
Consistency of the NLS estimator requires identification of the true parameter vector $\theta_0$. Typically (e.g.\ jennrich:1969) one shows that the population objective \(Q(\theta)\coloneqq\textnormal{\textsf{E}}[Q_n(\theta)/n]\) is uniquely minimized at \(\theta_0\). In our setting this route is impractical: the nonlinearity induced by \(\beta_t\)—itself a ratio of discounted partial sums that depend on \(\theta\)—renders \(Q(\theta)\) analytically intractable. Instead, we adapt the strong-consistency argument in francq:24, which extends francq:04 and does not require an explicit evaluation of the population objective.
To illustrate, note that $$\pi_t = a\psi+\delta h_{t-1}(\gamma)+\rho\psi y_{t-1}+w_t, \quad h_t(\gamma) \coloneqq \alpha_t(\gamma)+\beta_t(\gamma)^2(\pi_t-\alpha_t(\gamma)),$$ where \(w_t \coloneqq \psi\varepsilon_t+u_t\) drives the conditional law of \(s_t\mid\mathcal F_{t-1}\). A key ingredient for identification is to show that $h_t(\gamma)$ is well separated from $h_t(\gamma_0)$ whenever $\gamma \neq \gamma_0$. While the nonlinearity in $\beta_t(\gamma)$ complicates this task, on tail events \(\{|w_t|\ge c\}\) (which have strictly positive probability by the full-support part of Assumption (ref)), the easier-to-handle $\alpha_t(\gamma)$ dominates. This yields the following separation result:
This result, combined with compactness of $\Theta$, yields strong consistency via the arguments of francq:24. We therefore impose the additional restriction of the parameter space.
We are now ready to state the first core result:
It is immediate that $\delta = 0$ implies that $\gamma$ is no longer {\it jointly} identified with $(\delta,\psi)$. However, we can still show that the remaining components of the NLS estimator remain (weakly) $\sqrt n$-consistent. In particular, following saikkonen:95 and seo:11, the idea is to show that under $\delta_0=0$
uniformly outside any shrinking neighbourhood of $\eta_0$ whose radius goes to zero slower than $1/\sqrt{n}$. This, in turn, forces the NLS minimizer to lie within an $1/\sqrt{n}$ neighbourhood of $\eta_0$. Our proof builds, amongst others, on a uniform LLN and CLT for $Z_n(\gamma) \coloneqq \frac1{n}\sum_{t=1}^n z_t(\gamma)z_t(\gamma)^{\textnormal{\textsf{T}}}$ and $S_n(\gamma) \coloneqq \frac1{\sqrt n}\sum_{t=1}^n z_t(\gamma) u_t$, respectively, which--following andrews:92 and hansen:1996b--can be shown by establishing stochastic Lipschitz bounds. Technically, these Lipschitz conditions are a result of the following uniform moment bounds that will be used throughout:
Intuitively, bounding moments of the derivatives of $r_t$ and $\beta_t$ requires more stringent moment conditions on the errors $v_t = (u_t,\varepsilon_t)^{\textnormal{\textsf{T}}}$ because $r_t$ and $\beta_t$ are functions of sample second moments and thus already involve squared data.
This leads us to impose the following restriction on the errors:
As will be shown in the following section, if $\delta_0 \neq 0$, then the joint NLS estimator (weakly) $\sqrt n$-consistent for the full parameter vector $(\gamma,\delta,\psi)^{\textnormal{\textsf{T}}}$ under a slightly strengthened set of assumptions.
Having established the consistency of $\theta_n$, we turn to the question of its limiting distribution. To aid the discussion, define the gradient vector and the Hessian matrix of the nonlinear regression function $f_t(\theta) = \delta h_{t-1}(\gamma)+\psi y_{t}$ with respect to $\theta = (\delta,\gamma,\psi)^{\textnormal{\textsf{T}}}$ by \[ \dot f_t(\theta) = (h_{t-1}(\gamma),\delta_0 \dot{h}_{t-1}(\gamma),y_{t})^{T} \quad and \quad H_{t}(\theta) \coloneqq
, \] where $ \dot{h}_t(\gamma) \coloneqq (1-\beta_t(\gamma)^2)\dot{\alpha}_t(\gamma)+2\beta_t(\gamma)\dot{\beta}_t(\gamma)(\pi_t-\alpha_t(\gamma)) $ and
The notation $\dot \alpha$ and $\ddot \alpha$ ($\dot \beta$ and $\ddot \beta$) is used to indicate the first and second derivative wrt $\gamma$ of $\alpha$ ($\beta$). Using the definition of $\theta_n$ in conjunction with the mean-value theorem, ensures that there exists a $\bar\theta_n$ on the line segmenting connecting $\theta_n$ and $\theta_0$ such that \[ \sqrt{n}(\theta_n-\theta_0) = - M_n(\bar\theta_n)^{-1}\frac1{\sqrt n}\sum_{t=1}^n \dot f_tu_t, \quad \dot f_t \coloneqq \dot f_t(\theta_0), \] where \( M_n(\theta) \coloneqq -\frac1{n}\sum_{t=1}^n \dot f_t(\theta)\dot f_t(\theta)^{\textnormal{\textsf{T}}} + \frac1{n}\sum_{t=1}^n(\pi_t-f_t(\theta))H_t(\theta). \) Given the consistency of $\theta_n$, standard arguments (see, e.g. amemiya1985advanced) yield limiting normality if the following holds true:
Often, ($a$) can be ensured by direct evaluation of the expectation $A$. Here, due to the highly nonlinear nature of the $\dot f_t(\theta)$, deriving $A$ in closed form is not feasible. Still, we are able to verify ($a$) by a delicate tail argument, which, as an immediate consequence of Lemma (ref) and the CLT for homoskedastic martingale difference sequences, yields part ($c$). Moreover, as argued in andrews:92, part ($b$) follows from a stochastic Lipschitz bound. These results are summarized by the following lemma:
In view of Lemma (ref), it is interesting to note that, although part ($b$) involves bounds on expected moments of up to the third derivatives of the agents' recursive estimates, we do not require more than the eight error-moments already imposed by Assumption (ref).
We are now ready to derive the limiting distribution of the estimator:
As is common for extremum estimators (see, e.g. newmc:94), we restrict the true value $\theta_0$ to fall within the interior of $\Theta$, excluding boundary cases such like $\delta = 0$, which--as already discussed--require additional attention.
Recall from Eq. (ref) that an equilibrium is defined via $\vartheta_n = F(\vartheta_n;\lambda_n)$ for $\lambda_n = (\theta_{\delta,n},\theta_{\psi,n},\rho_n,\hat\sigma_u ^2,\hat\sigma_\varepsilon^2)^{\textnormal{\textsf{T}}}$, where $\theta_n =(\theta_{\gamma,n},\theta_{\delta,n}, \theta_{\psi,n})^{\textnormal{\textsf{T}}}$, while $\rho_n$ and $\hat\sigma_\varepsilon^2$ denote the OLS estimator of $\rho$ and the associated error variance $\sigma_\varepsilon^2$, respectively. In order to study the properties of our estimator of the equilibria, we thus need the following intermediate result:
If $u_t$ is symmetric, $\Omega$ has a simple block-diagonal structure: $$ \Omega =
, $$ where $\Sigma_{23}$ is the $2 \times 2$ $(2,3)$-block of $\Sigma$ and $D \coloneqq diag(1-\rho^2,E[u^4]-\sigma_u^4,E[\varepsilon^4]-\sigma_\varepsilon^4)$.
Let $r(\lambda)$ denote the number of roots of $G(\beta;\lambda)=\beta - F(\beta;\lambda)$ for a given $\lambda$, and define $$r_0 \coloneqq r(\lambda_0),\quad r_n \coloneqq r(\lambda_n),\quad G(\beta) \coloneqq G(\beta;\lambda_0),\quad G_n(\beta) \coloneqq G(\beta;\lambda_n).$$ By Assumption (ref), $G(\beta)$ is continuous on $[0,1]$, and, if in addition $\rho>0$ and $\psi \neq 0$, satisfies $G(0)<0<G(1)$. This implies $r_0 \in \{1,2,3\}$ and that the number of simple roots is odd. In particular, if $r_0=2$, then there is one simple root and one repeated root, at which $G(\cdot)$ does not change sign; this case is illustrated by the solid black line in Figure (ref). Moreover, when $r_0=2$, the event $\{r_n=r_0\}$ has probability zero; in terms of Figure (ref), one observes with equal probability either the dashed curve ($r_n=1$) or the dotted curve ($r_n=3$) for the realised $G_n(\beta)$. Because the derivative of $G(\beta)$ at the repeated root is zero, the equilibria are not identified at first order but only at second order, echoing the discussion in dovonon:18. This leads to a slower convergence rate and is summarised by the following corollary.
To set the stage, recall that $\sqrt{n}(\lambda_n-\lambda_0) \rightarrow_d \mathbb{Z}\sim {\mathcal N}(0,\Omega)$, let $\vartheta_{i,n}$, $1 \leq i \leq 3$, denote the potential roots of $\beta \mapsto G_n(\beta)$ and impose the following:
Assumption (ref) introduces the additional restriction $\psi \neq 0$. Since $\|F_\lambda(\beta;\lambda)\| = \beta^2$ if the true $\psi_0 = 0$, this is important to ensure that $F_\lambda(\beta;\lambda)$ is uniformly bounded away from zero at $\lambda = \lambda_0$; a crucial condition for the later analysis. The sign restriction $\rho > 0$ together with $\psi \neq 0$ ensures $G(0) < 0 < G(1)$ (e.g. existence of at least one root).
Before proceeding, let us briefly comment on the case $r_0 = 2$ where $G(\beta)$ merely touches the horizontal axis. Using an argument similar to dovonon:18, a second-order Taylor expansion yields, for $\beta$ around the double root $\beta_0^\dagger$, \[ \frac1{a}G(\beta;\lambda_n) = (\beta - \beta_0^\dagger)^2+{\mathbb Z}_n+ o_p(n^{-1/2}), \quad {\mathbb Z}_n \coloneqq \frac1{a}B^{\textnormal{\textsf{T}}}(\lambda_n-\lambda_0), \] with $\sqrt{n}\mathbb{Z}_n \rightarrow_d \frac1{a}B^{\textnormal{\textsf{T}}}{\mathbb Z}$. That is, we have a parabola $\beta \mapsto (\beta-\beta_0^\dagger)^2$ shifted by a random intercept $\mathbb{Z}_n = O_p(n^{-1/2})$. As shown in Figure (ref), a random downward shift (${\mathbb Z}_n < 0$) yields, in addition to the simple root on the left, the dotted line ($r_n = 3$) with two more crossings at $\vartheta_n^\pm$, while (${\mathbb Z}_n > 0$) shifts the parabola above the zero line so that the dashed line with no additional crossing emerges ($r_n = 1$). Because ${\mathbb Z}$ is symmetric, both these events occur asymptotically with equal probability of $\frac1{2}.$ Importantly, although the roots are consistent the number of roots is not. Intuitively, consistency of the repeated roots $\vartheta_n^\pm$ depends on the {\it magnitude} of $\sqrt{|\mathbb{Z}_n|}$, which is $O_p(n^{-1/4})$. On the other hand, the root count $r_n$ is determined by the {\it sign} of $\mathbb{Z}_n$, which, by limiting Gaussianity of $\sqrt{n}\mathbb{Z}_n$, remains a fair coin asymptotically. Put differently, $\lambda \mapsto r_n(\lambda)$ is discontinuous at $\lambda_0$ under multiple roots, so $r(\lambda_n)$ need not to converge to $r_0 = 2$ even though $\vartheta_n^\pm$ are consistent on the $\{r_n = 3\}$ event. An interesting by-product is that the double-roots $\vartheta_n^\pm$ have, with opposite sign, the same leading term of order $n^{-1/4}$. For large $n$, they therefore split symmetrically around the repeated root $\beta_0^\dagger$. Akin to a two-sample Jackknife, averaging thus eliminates this leading {\it bias} term.
The variance-covariance matrix $\Sigma = \sigma_u^2A^{-1}$, $A = \textnormal{\textsf{E}}[\dot f_t \dot f_t^{\textnormal{\textsf{T}}}]$, of the limiting distribution of $\theta_n$ is not known in closed-form. We therefore resort to numerical derivatives to estimate $A$. For a step size $\ell_n$, the estimator $A_n \coloneqq A_n(\theta_n)$ with $i$,$j$-th element ($1\leq i,j \leq 3$) is given by
for \(\ell_n \searrow 0\) and $e_i$, $i \in \{1,2,3\}$, denoting some sequence of step-sizes and the $3 \times 1$ unit vector, respectively. Recalling $\hat\sigma_u^2 = \frac1{n} Q_n(\theta_n)$, the following result is obtained:
The rate imposed on the step-size is common in the literature (see, e.g. newmc:94, oh2013simulated or mm:25).
We can use Corollary (ref) directly to test hypotheses about $\theta_0 \in \Theta$ using Wald-type tests. An important exception is the boundary case $\delta = 0$, where $\gamma$ is a non-identified nuisance parameter; this is similar to andpol:1994 and hansen:1996. As a solution, we propose the following procedure similar to hansen:17 (see also mm:25): Under the null $H_0$: $\delta = 0$, the ALM reduces to $\pi_t = \psi y_t+u_t$, where $\psi$ can be estimated by the OLS estimator $\tilde\psi_n$, say. Let $\tilde\sigma_n^2 = \frac1{n}\sum_t^n(\pi_t-\tilde\psi_n y_t)^2$ and define $\hat\sigma_n^2 \coloneqq \hat\sigma_n(\gamma_n)$, $\hat\sigma_n(\gamma) \coloneqq \frac1{n}Q_n^\star(\gamma)$, where $Q_n^\star(\cdot)$ is the profiled objective (see Remark (ref) above). The proposed test statistic is the following {\it sup}{\sf F} statistic:
Note that, by definition of the profiled $\gamma_n$, the statistic takes it maximum at ${\it sup}{\sf F} = n(\tilde\sigma_n-\hat\sigma_n)/\hat\sigma_n$, while, by Frisch-Waugh-Lovell, $F_n(\gamma)$ equals the squared $t$-statistic for the exclusion of $h_{t-1}(\gamma)$ in a regression of $\pi_t$ on $(h_{t-1}(\gamma),y_t)$.
Since the limiting process depends on unknown quantities, we follow hansen:17 and propose the use of a Gaussian multiplier bootstrap to obtain critical values: For $b \in \{1,\dots,B\}$, let $\{\pi_{1,b},\dots,\pi_{n,b}\}$, be a draw of standard normal variates. Use $\pi_{t,b}$ in place of $\pi_t$ to compute the bootstrap equivalent $sup{\sf F}^b$, say, of Eq. (ref). For $B$ large enough, we can then approximate the bootstrap $p$-value using $p_{n,B} \coloneqq \frac1{B}\sum_{b=1}^B1\{sup{\sf F} \leq sup{\sf F}^b\}$, and reject at a significance level $\alpha \in (0,1)$ if $p_{n,B} \leq \alpha$. As shown in the appendix $p_{n,B}$ is consistent for the true $p$-value as $n$ and $B$ diverge.
Another interesting question is how to conduct inference with respect to the equilibria $\beta$. Here, we propose uniform confidence bands for the unknown curve $\beta \mapsto G(\beta)$. To fix ideas, note that, by the functional Delta method, we have in $\ell^\infty([0,1])$, $\sqrt{n}(G(\beta,\lambda_n) - G(\beta,\lambda_0)) \rightsquigarrow {\mathbb G}(\beta),$ where ${\mathbb G}(\beta)$ is a Gaussian process with kernel $\textnormal{\textsf{cov}}[{\mathbb G}(\beta_1),{\mathbb G}(\beta_2)] = F_\lambda(\beta_1)^{\textnormal{\textsf{T}}}\Omega F_\lambda(\beta_2)$. We can therefore compute uniform $1-\alpha$ confidence intervals as summarized by the following Corollary:
If preferred, a non-studentized, {\it percentile} version of the confidence bands could be obtained by letting $[G_n(\beta) \pm \tilde c_\alpha n^{-1/2}]$, with $\tilde c_{\alpha}$ such that $\mathbb{P}\{\operatorname*{\textnormal{\textsf{sup}}}_{\beta} |{\mathbb G}(\beta)| \leq \tilde c_\alpha\} = 1-\alpha$. Note that, contrary to pointwise confidence intervals at selected roots $\beta$, our confidence bands are constructed for the function $\beta \mapsto G_n(\beta;\lambda_0)$ itself, exploiting smoothness in $\lambda$, and are thus unaffected by potentially repeated roots where $G_\beta$ vanishes.
Since the limiting distribution is unknown and non-pivotal, to obtain $c_\alpha$, we suggest to use a multiplier bootstrap based on the first order approximation: \[ \sqrt{n}(G_n(\beta)-G(\beta)) = \frac1{\sqrt n}\sum_{t=1}^n m_t(\beta;\kappa_0)+o_p(1), \] where \[ m_t(\beta;\kappa) \coloneqq -F_\lambda(\beta;\lambda)^{\textnormal{\textsf{T}}}\phi_t(\kappa), \quad \kappa \coloneqq (\lambda^{\textnormal{\textsf{T}}},\gamma)^{\textnormal{\textsf{T}}}, \] and $\frac1{\sqrt n}\sum_t \phi_t(\kappa_0)$ is the influence function of $\sqrt{n}(\lambda_n-\lambda_0)$. Thus, for a draw $b \in \{1,\dots,B\}$ of standard Gaussian multipliers $\{\xi_{1,b},\dots,\xi_{n,b}\}$, and, conditionally on the data, we have \[ \frac1{\sqrt n}\sum_{t=1}^n \xi_{t,b} m_t(\beta;\kappa_0) \rightsquigarrow_{\mathbb P} {\mathbb G}(\beta) \quad \textup{in} \quad \ell^\infty([0,1]), \] where $ \rightsquigarrow_{\mathbb P}$ means that weak convergence holds in probability with respect to the data-generating measure. In practice, neither $\kappa_0$ nor $\phi_t(\kappa_0)$ is observed. As shown in Appendix (ref), we can however construct $\phi_{t,n}(\kappa_n)$ such that $\frac{1}{n}\sum_{t=1}^n\|\phi_{t,n}(\kappa_n)-\phi_t(\kappa_0)\|^2=o_p(1)$, and define the feasible linearization $m_{t,n}(\beta)\coloneqq -F_\lambda(\beta;\lambda_n)^\top\phi_{t,n}(\kappa_n)$. We then simulate the critical value from the studentized bootstrap statistic \[ T_{n,b} \coloneqq \operatorname*{\textnormal{\textsf{sup}}}_{\beta\in[0,1]}\frac{\big|\frac{1}{\sqrt n}\sum_{t=1}^n \xi_{t,b} m_{t,n}(\beta)\big|}{s_n(\beta)}, \qquad s_n(\beta)^2\coloneqq \frac{1}{n}\sum_{t=1}^n m_{t,n}(\beta)^2, \] and take $c_{\alpha,n,B}$ as the empirical $1-\alpha$ quantile of $\{T_{n,b}\}_{b=1}^B$. As verified in the appendix, if $B \rightarrow \infty$ and $n \rightarrow \infty$, $c_{\alpha,n,B}\rightarrow_p c_\alpha$, yielding the uniform bands in Corollary (ref).
In this section, we present finite-sample evidence based on simulated data. We consider two scenarios: Scenario (A) with three behavioural equilibria ($r_0 = 3$) and Scenario (B) with two equilibria ($r_0 = 2$). The Monte Carlo dgp is defined by Eqs. (ref), (ref), and (ref), with {\sf IID} Gaussian innovations $v_t = (u_t,\varepsilon_t)^{\textnormal{\textsf{T}}} \sim \mathcal N(0,\textup{\sf diag}(\sigma_u^2,\sigma_\varepsilon^2))$. For Scenario (A), the parameter values are inspired by our empirical analysis. In particular, the parameter vector estimated by NLS is set to $\theta = (0.076,0.998,0.090)^{\textnormal{\textsf{T}}}$, the AR(1) dynamics (ref) are parametrised via $(a,\rho) = (-0.02,0.93)$, and, for the innovations, we set $(\sigma_u,\sigma_\varepsilon) = (0.44,0.76)$. This configuration gives rise to three behavioural equilibria $\beta = (0.174,0.856,0.999)^{\textnormal{\textsf{T}}}$. The implied values of $\lambda = (\delta,\psi,\rho,\sigma_u^2,\sigma_\varepsilon^2)^{\textnormal{\textsf{T}}}$ are broadly in line with those in hz14. For Scenario (B), we keep the value of the gain and the AR(1) dynamics, and choose $\lambda = (0.95,0.2734,0.9,1,1)^{\textnormal{\textsf{T}}}$, which gives rise to a simple root at $\beta = 0.9766$ and a repeated root at $\beta^\dagger = 0.5551$.
We report, for the NLS estimator $\theta_n$ of $\theta = (\gamma,\delta,\psi)^{{\textnormal{\textsf{T}}}}$, the mean, bias, standard deviation, and the empirical size of two-sided $t$-tests (using, in accordance with Corollary (ref), numerical standard errors) at a nominal 5% level. We also present the empirical coverage of the percentile ({\it cover}) and studentized ({\it cover-s}) confidence bands (see Corollary (ref)) for the curves $\beta \mapsto G(\beta)$ plotted in Figure (ref), based on $B = 4{,}999$ bootstrap replications at a 95% coverage level. Finally, we report the mean and standard deviation of the estimated equilibria, together with the observed frequency at which the true number of roots is recovered, i.e.\ $\mathbb{P}\{r = r_0\}$.
The findings are summarized in Table (ref). In both scenarios, the NLS estimator appears consistent and the normal approximation for the $t$-statistics is fairly accurate, while the empirical coverage of the studentized confidence bands is slightly below the nominal 95% level for smaller $n$. Moreover, in line with Corollary (ref), we observe that the estimated number of roots $r_n$ either converges to the true value $r_0$ or, in Scenario (B), oscillates between $r = 1$ and $r = 3$ with approximately equal probability. In the latter case, the convergence of the estimates associated with the repeated root appears to be an order of magnitude slower than the $\sqrt{n}$-rate observed for simple roots, while averaging the two nearby estimates corresponding to the repeated root improves performance.
To illustrate the empirical relevance of our theoretical results, we estimate the NKPC given by Eqs. (ref), (ref), and (ref) using quarterly US\ data from 1960:Q1 to 2019:Q4 ($n=240$). Inflation $\pi_t$ is measured by CPI inflation. Since marginal costs are proportional to the output gap or the real labour share (see, e.g., Woodford2003 or MavroeidisPlagborgMollerStock2014), we consider two standard slack measures: ({\sf A}) the output gap measure from the Congressional Budget Office and ({\sf B}) the percentage change in real unit labour costs (see also gali1999inflation or sbordone2002prices). The three series are plotted in Figure (ref).
We employ a preliminary step to obtain starting values: In particular, we use an initial grid search over $\gamma \in (0,0.1)$ in order to obtain starting values, treating the initial value $\operatorname*{\mathfrak{a}}$ of the learning recursion as a parameter to be estimated. NLS estimation is carried out using the {\sf R} {\sf optim}-routine with default settings, restricting the parameter space for $\delta$ to $(0,1)$. The upper panel of Table (ref) reports the NLS estimation results, while the lower panel summarizes OLS estimation of the AR(1) models for the two different proxies of the driving variable; 95% confidence intervals based on numerical derivatives are given beneath. The NLS estimates of specification {\sf A} (output gap) and {\sf B} (unit labour costs) are in line with the empirical macro literature (see, e.g., milani:07, chev:10, Lansing2009, or HommesMavromatisOzdenZhu2023). The estimated learning gains imply that agents place most weight on roughly the last 3 to 5 years of data for specifications {\sf A} and {\sf B}, respectively. The estimate of the discount factor on expected inflation (viz. $\delta$) is large. This confirms that forward-looking expectations play a quantitatively important role in the Phillips curve and that inflation dynamics are highly persistent. The slope parameter is positive and statistically significant, but small in magnitude, consistent with the `flattened' NKPC found in much of the empirical literature.
In specification {\sf A}, the estimated parameters imply three behavioural equilibria $\beta_1 = 0.1914$, $\beta_2 = 0.8316$, $\beta_3= 0.9995$; in specification {\sf B} we find a unique equilibrium $\beta = 0.01341$. The corresponding $\beta \mapsto G_n(\beta)$ functions are shown in Figure (ref). The key difference is the persistence of the driving variable: $\rho$ is close to one in {\sf A} but relatively small in {\sf B}. This pattern is consistent with the theoretical properties of the map $\beta \mapsto G(\beta;\lambda)$, which shows that, holding other parameters fixed, $\rho \rightarrow 0$ leads to a single equilibrium with $\beta \rightarrow 0$. The low persistence of unit labour costs is sufficient to support only a single, less persistent equilibrium, and the feedback between expectations and realized inflation is too weak to generate a high-persistence belief regime. Finally, in case of specification {\sf A}, $\beta_2$ is under the criterion of hz14 not stable; i.e. an agent that uses asymptotically all data ($\gamma \rightarrow 0$) is not able to learn $\beta_2$ in the limit. The empirical model thus points to two stable belief regimes for inflation persistence: a low and a high persistence equilibrium.
We have developed estimation and inference methods for a New Keynesian Phillips curve with constant-gain learning and potentially multiple behavioural equilibria. Under mild conditions the resulting model is geometrically ergodic, and a nonlinear least squares estimator of the structural parameters is strongly consistent and asymptotically normal. We explain how to conduct inference with respect to structural parameters, including the equilibria, and apply these techniques to U.S.\ inflation data, revealing multiple belief regimes in the data. This paper is only a starting point and several extensions are left for future work: A natural next step is to develop tools for multivariate versions of the model, along the lines of milani:07 or HommesMavromatisOzdenZhu2023, in order to allow for richer joint dynamics of macroeconomic variables (see also the discussion in MavroeidisPlagborgMollerStock2014). It would also be useful to relax some of our regularity conditions, to exploit panel data to increase effective sample size, and to adapt the analysis to decreasing-gain learning rules using results from earlier work on the econometrics of models with adaptive learning like chev:10, chev:17, chrismass:18, mayer:22, or mm:25.
\addcontentsline{toc}{section}{References}