EconBase
← Back to paper

Least squares estimation in nonstationary nonlinear cohort panels with learning from experience

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

90,125 characters · 11 sections · 125 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

-2cmLeast squares estimation in nonstationary nonlinear cohort panels with learning from experience

\thispagestyle{empty}

abstractWe discuss techniques of estimation and inference for nonstationary nonlinear cohort panels with learning from experience, showing, inter alia, the consistency and asymptotic normality of the nonlinear least squares estimator used in empirical practice. Potential pitfalls for hypothesis testing are identified and solutions proposed. Monte Carlo simulations verify the properties of the estimator and corresponding test statistics in finite samples, while an application to a panel of survey expectations demonstrates the usefulness of the theory developed. {\bf Keywords:} adaptive learning, inflation expectations, nonlinear least squares with nonsmooth objective function, cohort panel data, asymptotic theory, nuisance parameters

Introduction

Following Sargent93,Sargent99, the literature has seen a renewed interest in how economic agents form expectations. In particular, researchers and policy makers alike increasingly question the orthodox framework of viewing agents as forming full-information rational expectations, see e.g.\ evans:01, MankiwReis02, and Bernanke07. Especially expectations about future inflation are relevant for understanding economic outcomes. Recent empirical work on the formation process of inflation expectations includes BachmannBergSims15 and CoibionGorodnichenkoRopele20 who emphasise the importance of expected inflation for consumption and investment decisions, respectively, and coibion:20 who analyse how inflation expectations can be used as a policy tool by monetary authorities.

A concomitant development is the increasing recognition that representative agent theory, the predominant approach to modelling in economics, may be insufficient for explaining economic fluctuations and that the heterogeneity between agents needs to be taken account of, see e.g.\ HeathcoteStoreslettenViolante09 and KaplanViolante18 for surveys and Yellen16 for the view of a policy maker. One of the driving forces behind this development has been the growing availability and analysis of surveys of both households' and firms' beliefs, see for instance WeberD’AcuntoGorodnichenkoCoibion22 and DAcuntoMalmendierWeber23 for overviews, and link:23 for a recent investigation into the heterogeneity of housholds' and firms' expectations. The Michigan Survey of Consumers (MSC) is one of the longest-running surveys that contains information on agents' inflation expectations. Early work on the MSC concentrated on analysing aggregates of the data, see e.g.\ MankiwReisWolfers04, Branch04, and CoibionGorodnichenko12. Recently, however, the focus has shifted to taking full advantage of the entire panel of survey respondents, as do, for instance, BachmannBergSims15, mn:2016, and meekmonti:23.

At the confluence of these two strands of the literature stand CoibionGorodnichenkoKamdar18 who forcefully argue for “a careful (re-)consideration of the expectations formation process and a more systematic inclusion of real-time expectations through survey data” (p.\ 1447). One recent line of such research is on so-called `experience effects', stipulating that exposure to personal or public economic or political outcomes tends to shape agents' behaviour, see Malmendier21 for a current survey. In particular, in their seminal paper on how inflation expectations are determined by individual experiences, mn:2016 (MN, henceforth) depart from the rational expectations paradigm by making use of an adaptive learning framework in which agents entertain their own --potentially mis-specified-- model of how inflation is determined and estimate it recursively to form their expectations. Similarly, in their empirical analysis, MN make full use of the MSC, in both the cross-sectional and the time dimensions. The two main findings of MN are ($i$) substantial heterogeneity between individuals of different age and ($ii$) what MN call recency bias, a concept related to the availability heuristic by TverskyKahneman74. The heterogeneity is manifest in the weight that individuals give to new data as they update their inflation forecasts and that depends on their age. The recency bias is captured by the magnitude of the so-called `gain parameter' $\gamma$ in the estimated updating equation and indicates that individuals' recent experiences have a stronger impact on their expectations than distant ones.

These results have spawned a string of papers on `learning-from-experience' that either re-use MN's parameter estimates in similar models for different empirical applications or that extend MN's specification for describing the MSC data. Examples of the former category are naknun:15 on stock prices and dividends, and Acedanski17 on the wealth distribution; both papers calibrate their models with MN's estimated gain parameter of $\hat \gamma = 3.044$. In the latter category fall MadeiraZafar15 who extend MN's analysis of the MSC dataset by allowing for the heterogeneous use of private information, and Gwak22 who builds a model similar to MN's yet includes in the specification a Markov-switching component to distinguish between learning-from-experience in high and low volatile inflation regimes. Recently, MalmendierNagelYan21 analyse the voting behaviour of the Fed's FOMC members using individuals' learnt-from-experience inflation expectations as given input variable, and nagel:24 explores the implications on real interest rates of learning from experience.

What all these papers have in common is that the models they consider are highly complex and that neither their microfoundation nor their econometrics is yet fully understood. Indeed, from an economic theory point of view, DuffyShin23 take a step back and use the concept of learning-from-experience in a microfounded demography-based model to rationalise constant gain learning. In the present paper, we follow their example, go back to square one, and derive the econometric theory of a learning-from-experience model. In fact, the full complexity of the empirical models estimated by MadeiraZafar15, mn:2016, MalmendierNagelYan21, and Gwak22 is beyond the scope of the present paper. Instead, we consider a special case of theirs that is analytically tractable. It nevertheless allows us to estimate a plausible learning-from-experience model empirically and engage in statistically well-founded inference. Doing so, we make progress on two fronts: Empirically, we shed new light on the question of heterogeneous inflation expectations and recency bias, obtaining conclusions that are in line with the aforementioned empirical papers on the MSC. Theoretically, we establish novel econometric results for the analysis of learning-from-experience models that set the scene for future work analysing the econometrics of even more complex models.

The econometric specification we adopt in the present paper is a special case of the model used in MN and can be viewed as nonstationary nonlinear cohort panel data model with time fixed effects. The nonlinearity in the regression function stems from the recursively generated expectations, while the nonstationarity arises due to a stochastic evaporating trend component, reminiscent of the linear regressions with deterministic evaporating trends studied by phillips:07. Related research demonstrates that the statistical analysis of estimators in macroeconomic models with similar adaptive learning schemes provides a challenging task, see e.g. chev:10, chev:17, chrismass:18,chrismass:19, mayer:22,mayer:23, or christiano:24. This literature shows that estimation of and inference in models with adaptive learning is far from standard and often marred by weak-identification, asymptotic collinearity, or non-standard convergence rates. As will be discussed below, one of MN's contributions is to sidestep these issues to some extent by exploiting the cross-sectional variation across individuals. However, as mentioned earlier, the theoretical properties of the nonlinear econometric methods employed in the empirical learning-from-experience papers have not been examined to date, and implicit or explicit claims that the model parameters are identified or that certain statistics have some given asymptotic distribution call for verification.

The aim of the present paper is thus to bridge the existing gap between econometric theory and empirical practice. Our contributions are twofold: {\it First}, we derive new asymptotic results for point estimation and inference in a nonlinear cohort panel data model with learning from experience, thereby extending the established econometric results in the literature in terms of ($i$) a panel dimension, ($ii$) a heterogeneous gain sequence, and ($iii$) multivariate estimation by nonlinear least squares (NLS, henceforth). As a {\it second} contribution, we apply our results to an empirical model of the MSC dataset. Our model is akin to the baseline specification employed by MadeiraZafar15, mn:2016, MalmendierNagelYan21, Gwak22 and nagel:24. The present paper is therefore in the tradition of milani:07, chev:10, adam:16, and hommes:23 who derive rigorous econometric results for the modelling of substantive empirical problems in the economics of adaptive learning. Yet while all three papers use constant gain learning specifications, estimated by Bayesian methods in milani:07 and by continuously updated GMM in chev:10 and by the method of simulated moments in adam:16, our paper is the first to provide well-founded econometric insights for a decreasing gain model specification.

In the theory part of the paper, we consider different asymptotic regimes that depend on whether the number of cohorts is fixed or not. One important conclusion for point estimation is that, albeit consistent in all scenarios, the NLS estimator might not be asymptotically normal if the number of cohorts is fixed due to an objective function that is not differentiable everywhere. As argued below, this problem can be overcome if the number of cohorts diverges. However, in this case, we are confronted with the additional challenge of asymptotic collinear regressors. This technical hurdle notwithstanding, asymptotic normality, albeit at a nonstandard convergence rate, is established by combining results from analytical number theory with seminal results for extremum estimators with nonsmooth objective function, see e.g.\ newmc:94.

A further focus is placed on hypothesis testing. The inference conducted in the aforementioned empirical learning-from-experience papers is, by and large, classical. Yet our asymptotic analysis identifies potential pitfalls due to slow convergence and parameter non-identification. To address the latter issue, we propose a solution that builds on the results of hansen:1996 regarding hypotheses that involve non-identified nuisance parameters. We investigate the properties of the NLS estimator and the test statistics in finite samples by use of Monte Carlo simulations. In the empirical part of the paper we revisit the MSC dataset. Using the estimation and inference procedures developed in the theory part of our paper we confirm, on the whole, the findings of previous empirical papers on the MSC. Yet we note conflicting evidence on the weight given by agents to private experiences when they form expectations about future inflation.

The remainder of the paper is organised as follows. Sections (ref) and (ref) introduce the model and the NLS estimator, respectively. We lay out assumptions, discuss consistency and asymptotic normality of the estimator in Section (ref). Standard errors and inferential methods are addressed in Section (ref). Section (ref) contains a Monte Carlo study, while the empirical application is presented in Section (ref). Additional results are relegated to the Supplementary Material.

Model

Consider the following model that relates observed survey expectations, denoted by $z_{t,s}$, to the learnt expectation about future macro-level inflation $y_{t+1}$, denoted by $a_{t,s}$, via

equation[equation omitted — 91 chars of source]

with $t$ and $s$ indexing the time and birth period, respectively. Here, \(\alpha_t\) represents a time fixed effect capturing a component in the expectations formation process that is common to all cohorts, while $\beta$ is the weight given to private, i.e.\ cohort-specific, experiences. The term \(\varepsilon_{t,s}\) is some error term further specified below. The model in Eq.\ (ref) is supplemented by an equation specifying the way in which different cohorts update their beliefs over time. In particular, it will be assumed that, at the end of period $t$, individuals born in period \(s\) form expectations about future aggregate inflation based on the available inflation history according to an adaptive learning rule

equation[equation omitted — 123 chars of source]

where \(\mathfrak{a}_s \in \mathbb R\) is some initial value. The updating scheme in Eq.\ (ref) is a stochastic approximation algorithm that can be viewed as a generalisation of recursive least squares (see, e.g., ben:90), where the so-called `gain sequence' \(\gamma_{t,s}\) measures the responsiveness to previous prediction mistakes.

Implicit in Eq.\ (ref) is the assumption that individuals use a constant level model as their so-called perceived law of motion (PLM) for prediction. Within the macroeconomic learning literature, the PLM is the model individuals use to forecast $y_{t+1}$ and which in general does not coincide with the true data generating process of $y_t$. Note that the model in Eqs.\ ((ref)) and ((ref)) is not of the self-referential type found in the classical adaptive learning literature surveyed by evans:01 since the dependent variable $z_{t,s}$ is different from the covariate $y_t$ upon which the learning recursion is based.

Following the empirical learning-from-experience literature, we assume that the weight individuals attach to previous observations depends on their age \(t-s\) according to \(\gamma_{t,s} \coloneqq \gamma_{t,s}(\gamma)\) such that

equation[equation omitted — 168 chars of source]

where \(\gamma>0\) is some unknown gain parameter. This specification creates heterogeneity in the way different cohorts form their expectations. The economic interpretation of \(\gamma\) is that of a `forgetting factor', where \(\gamma > 1\) (\(\gamma < 1\)) means that agents attach more (less) weight to recent revisions of the data while \(\gamma=1\) results in ordinary least squares learning with equally weighted observations. The choice of gain sequence $\gamma_{t,s}$ in Eq.\ ((ref)) extends the classical least-squares recursion of, e.g., marsar:89 by making the updating weight a function of age $t-s$ rather than merely time $t$.

It follows from Eqs.\ (ref) and (ref) that the forecast of an individual of age $t-s$ is simply a weighted average of past and present information. That is, $a_{t,s} \coloneqq a_{t,s}(\gamma)$ where, for any $s>\gamma > 0$,

align[align omitted — 223 chars of source]

see Lemma (ref) of the appendix for details. The `floor' function is defined as \(\floor{x} = \{m \in \mathbb{Z}: m \leq x\}\), and we use the conventions $\prod_{i = s+1}^{s}(1-\gamma/i) \coloneqq 1$, and $\kappa_{\floor{\gamma},s}(\gamma) \coloneqq \prod_{i = \floor{\gamma}+1}^{s}(1-\gamma/i).$ Importantly, all information before birth plus \(\gamma\) is discarded, thereby ensuring an economically reasonable sequence of non-negative weights $\kappa_{j,s}(\gamma)$ .

The complete model is thus given by Eqs.\ (ref), (ref), and (ref). It is a nonlinear cohort panel data model with time fixed effects. The age effect is captured by the nonlinear age-dependent belief updating mechanism $a_{t,s}(\gamma)$, while the time effects enter the model linearly. The parameters to be estimated are thus $\theta \coloneqq (\beta,\gamma)^{\textnormal{\textsf{T}}}$. This model is akin to MN's baseline specification; see also Eq.\ (6) in MN. Note that Eq.\ (ref) is a special case of the updating scheme that MN equip individuals with, in that they consider an AR(1) as PLM. Agents in our setup are therefore assumed to be less sophisticated when compared to their counterparts in MN, yet given the prominence of `simple' forecasting rule for inflation in publications by the Federal Reserve Banks\footnote{See AtkesonOhania01, PasaogullariMeyer2010, and BauerMcCarthy15 for examples.} a more restricted perception of how inflation is generated can arguably be seen as more realistic for boundedly rational agents. A comparison of both specifications in terms of a Monte Carlo study and an extended empirical application is included in the Supplementary Material; a theoretical treatment is, however, beyond the scope of the paper.

Estimation

The parameters in the model given by Eqs.\ (ref), (ref), and (ref) will be estimated by nonlinear least squares (NLS). Before proceeding with a discussion of the estimator, some additional notation is needed. In particular, we denote by $l$ and $u$ the first and the last age group to be considered, and by \(n\) the last time period. Consequently, the time index $t$ and the birth year $s$ take on values in

equation*[equation* omitted — 100 chars of source]

respectively. Defining \[ m \coloneqq u-l+1, \quad 1 \leq l < u < n, \] to be the number of cohorts, the pooled data set is seen to consist of a total of \[N \coloneqq (n-u)m\] observations. The structure of the dataset, illustrated in Table (ref), is similar to age-period-cohort panels covered elsewhere in the literature (e.g. HarnauNielsen18 or FannonNielsen19). Our specification is, however, fundamentally different from the aforementioned literature due to the nature of the nonlinearly and recursively generated cohort effects (see also the discussion in maletal:21).

\colorlet{myred}{white}

table[table omitted — 1,604 chars of source]

We are now ready to introduce the NLS estimator of the true parameter vector \(\theta_0 \coloneqq (\beta_0,\gamma_0)^{\textnormal{\textsf{T}}}\). In a first step, the time fixed effects are eliminated by subtracting from Eq.\ (ref) cohort-means, yielding

equation[equation omitted — 93 chars of source]

where a tilde indicates deviations from cohort-means:

equation[equation omitted — 194 chars of source]

For economy of notation, the dependence of \(\tilde e_{t,s}\) on the number of cohorts \(m\) is implicitly understood. In terms of the data structure illustrated in Table (ref), the cohort-means are given by row-wise averages. The NLS estimator \(\theta_n \coloneqq (\theta_{\beta,n},\theta_{\gamma,n})^{\textnormal{\textsf{T}}}\), say, then minimises the objective

equation[equation omitted — 234 chars of source]

over \(\theta = (\beta, \gamma)^{\textnormal{\textsf{T}}} \in {\mathbb R} \times \Gamma \), $\Gamma \coloneqq [\ubar\gamma,\bar\gamma]$ for $0<\ubar\gamma<\bar\gamma<\infty$ specified below. Note that since \(\tilde{a}_{t,s}(\gamma) \coloneqq a_{t,s}(\gamma)-\bar a_t(\gamma)\), \(Q_n(\theta)\) depends on \(\gamma\) through both \(a_{t,s}(\gamma)\) and its cohort-mean. Clearly, the objective is highly nonlinear in $\gamma$, necessitating the use of numerical routines for estimation. However, we note that the computational burden of numerical optimisation can be reduced significantly by profiling $Q_n (\theta)$ in Eq.\ (ref) further w.r.t.\ $\beta$. Put differently, upon exploiting that the model is linear in \(\beta\) we get $(\beta_n,\gamma_n)^{\textnormal{\textsf{T}}}$ where $\beta_n \coloneqq \beta_n(\gamma_n)$, and $\gamma_n$ minimises

align*[align* omitted — 338 chars of source]

The same approach is commonly taken in threshold models (see, e.g. hansen:17).

Assumptions, consistency, and asymptotic normality

The statistical analysis of the NLS estimator requires some care. Upon inspecting Eq.\ (ref), it is evident that the sample objective function $Q_n(\theta)$ is continuous in $\theta$ but has “kinks”, i.e.\ lacks differentiability, if $\gamma \in \mathbb{Z}$. Importantly, even its suitably standardised population counterpart might not be differentiable everywhere, a crucial prerequisite for the asymptotic normality of $\theta_n$ as $n \rightarrow \infty$, see e.g. newmc:94. To ensure differentiability --at least in the limit-- we must allow for the number of cohorts $m$ to diverge. However, the resulting setting poses two further challenges: First, contraction (or invertibility) arguments, that turn out to be very convenient when establishing the required uniform results in similar settings like GARCH models, score driven models, or, more generally, nonlinear time series models (e.g.\ jensen:04, strau:06, blasques:18, or potscher:13), fall short in our setting. In view of Eq.\ (ref), this follows readily by recognising that the mapping $f_{t,s}(\cdot)$ defined via $a_{t,s} = f_{t,s}(a_{t-1,s},y_t;\gamma)$, $$f_{t,s}(a,y;\gamma) = (1-\gamma_{t,s}(\gamma))a+\gamma_{t,s}(\gamma)y $$ is not contracting in its first argument because $$\operatorname*{\textnormal{\textsf{sup}}}\limits_{l \leq t-s \leq u, \gamma \in \Gamma}\operatorname*{\textnormal{\textsf{ln}}}\left|\frac{\partial f_{t,s}(a,y;\gamma)}{\partial a}\right| = 0$$ if $m$ (and thus $u$) diverges. Second, the nonlinear regression function degenerates if the number of cohorts diverges. To see this, it is helpful to recognise that under certain regularity conditions laid down below

align[align omitted — 123 chars of source]

as age grows large (i.e \(t-s \rightarrow \infty\)), with

align[align omitted — 195 chars of source]

Put differently, our setting can be described as a nonlinear panel specification with stochastic, potentially evaporating trend component. The analysis of linear models with deterministic (e.g. phillips:07) and stochastic (e.g., chrismass:19 or mayer:22) evaporating trends leads us to expect nonstandard limiting behaviour also in the present case. More specifically, the evaporating component of the recursively generated expectations causes time-varying moments (e.g. $\textnormal{\textsf{var}}[a_{t,s}] = O(t-s)$ cf. Eq.\ (ref)) and thus nonstationarity.

To summarize, we cannot derive the properties of $\theta_n$ from standard asymptotic theory. Instead, we must derive them from first principles. In particular, the recursive solution in Eq.\ (ref) enables us, by leveraging insights from analytical number theory, to directly apply seminal results from $M$-estimation. This approach allows us to establish, among other things, the consistency, convergence rates, limiting distribution, and inferential methods for $\theta_n$. In doing so, we impose the following assumptions.

Assumptions

First, we distinguish between two asymptotic regimes:

assumption\(u\) and \(l\) are fixed constants independent of \(n\).
assumption\(u \coloneqq u(n) \rightarrow \infty\) as \(n \rightarrow \infty\) such that \(m = u-l+1\rightarrow \infty\), where \(l\) is either fixed or \(l \coloneqq l(n) \rightarrow \infty\), with \[ (\operatorname*{\textnormal{\textsf{ln}}}(u)-\operatorname*{\textnormal{\textsf{ln}}}(l))/\operatorname*{\textnormal{\textsf{ln}}}(n) \rightarrow \lambda_{1} \in (0,\infty), \quad u/n \rightarrow \lambda_2 \in [0,1), \]

Assumption (ref) refers to a fixed cohort length (\(m\)), while Assumption (ref) allows \(m = m(n)\) to diverge pathwise as a function of the sample size \(n\). Note, that the lower bound \(l\) can be fixed, while \(u = u(n) \rightarrow \infty\) ensures that \(m\) diverges with $n$.\footnote{The assumption allows the (finite) limit of the ratio $m/n$ to be either strictly positive or zero. For example, to see that the latter case is covered, suppose that $l$ is fixed and $u \sim n^{\kappa}$, $\kappa \in (0,1)$, so that $\operatorname*{\textnormal{\textsf{ln}}}(u)/\operatorname*{\textnormal{\textsf{ln}}}(n) \sim \kappa > 0$ but $m/n \rightarrow 0$.} Intuitively, under Assumption (ref), we are able to consistently estimate the fixed effects so that (asymptotically) the additional estimation error due to cohort-demeaning --present under Assumption (ref)-- vanishes. More importantly, however, $m = m(n) \rightarrow \infty$ ensures a sufficiently smooth population objective such that we can hope to derive the limiting distribution of $\theta_n$.

The next assumption specifies the distributional characteristics of $y_t$ and the error term:

assumption\textcolor[rgb]{1,1,1}{.} \begin{enumerate}[label=B.\arabic*,ref=B.\arabic*] • \(\{y_t\}_t\) is fourth-order stationary with continuous spectral density bounded away from zero, autocovariance function $c(\cdot)$ such that $\operatorname*{\textnormal{\textsf{sup}}}_{\tau \geq 1}(1+|\tau|)^2|c(\tau)| \leq \infty$, and absolutely summable cumulants up to order four. • For each \(s\leq t\), \(\{\varepsilon_{t,s}\}_{t}\) form martingale difference sequences with respect to \(\mathcal{F}_{t} \coloneqq \sigma(\{y_{i+1},\varepsilon_{i,j}: i \leq t, j\leq i\})\) such that \(\textnormal{\textsf{E}}[\varepsilon_{t,s}^r \mid \mathcal{F}_{t-1}]\), \(r \in \{2,3,4\}\) are finite constants a.s. so that \(\textnormal{\textsf{E}}[\varepsilon_{t,s}\varepsilon_{t,k} \mid \mathcal{F}_{t-1}] = \sigma^21\{k=s\}\) a.s.. \end{enumerate}

Assumption (ref) places some structure on the dependence of the process \(\{y_t\}_t\) using a fourth-order cumulant condition. Any stationary Gaussian process with $\operatorname*{\textnormal{\textsf{sup}}}_{\tau \geq 1}(1+|\tau|)^2|c(\tau)| \leq \infty$ satisfies Assumption (ref) because higher order cumulants are zero in this case. More generally, stationary processes under (strong) mixing conditions (see, e.g., doukhan:89) as well as linear processes with absolutely summable Wold coefficients and IID innovations that have finite fourth moments (see, e.g., hannan:70) can be shown to have absolutely summable fourth cumulants, i.e. $\sum_{i,j,k=-\infty}^\infty |c(i,j,k)| < \infty,$ $c(i,j,k) \coloneqq \textsf{cum}[y_t,y_{t+i},y_{t+j},y_{t+k}].$ These summability conditions restrict the memory of \(\{y_t\}_t\) to be short and allow us to evaluate higher order moments of the recursion in Eq.\ (ref) based on arguments borrowed from dem:2008. Assumption (ref) assumes that the error term is a homoskedastic martingale difference sequence with finite homokurtosis; importantly, the assumption rules out serial correlation among time (\(t\)) and birth period (\(s\)). We leave any weakening of Assumption (ref) for future research, but return to this issue briefly as part of a Monte Carlo study in the Supplementary Material.

Finally, Assumption (ref) restricts the parameter space:

assumption\(\theta_0 \in \textnormal{\sf int}(\Theta)\), where \(\Theta \coloneqq \Xi \times \Gamma\), $\Xi \coloneqq [\ubar\beta,\bar\beta]$, $\Gamma \coloneqq [\ubar\gamma,\bar\gamma]$ for $-\infty<\ubar\beta\leq \bar\beta < \infty$ and \(\frac2{3} < \ubar\gamma < \bar\gamma <\infty\).

Assuming a compact and convex parameter space is a standard assumption for nonlinear regression (e.g. jennrich:1969 or chan:15). In particular, we impose compactness on the parameter space of $\beta$. We follow hansen:17, who argues that, in principle, the restriction could be relaxed such that $\Xi = \mathbb{R}$ at the expense of more technical detail, as, for instance, discussed in newmc:94. Whenever interest lies in identifying jointly the parameter vector $\theta$, we impose the additional identification restriction \(\beta_0 \neq 0\). Fortunately, as shown in Section (ref), it is still possible to draw statistical inferences involving the hypothesis \(\beta = 0\), provided a suitable test statistic is used. Although it seems possible to relax the constraint \(\ubar\gamma > 2/3\) and allow for \(\gamma \in (1/2,2/3]\), this comes with a substantial increase in additional technicalities and is thus left for future research. The boundary point $\gamma = 1/2$, in particular, presents several difficulties as already discussed in chrismass:18 for a linear regression model. Importantly, Assumption (ref) allows for “recency bias” ($\gamma > 1$) as well as updating schemes where distant data points are weighted more heavily than recent ones ($\gamma < 1$). As we will see in Section (ref), this allows us to empirically test the hypothesis of “recency bias” put forward by MN.

Consistency

Inspired by the analysis of the NLS estimator with trending data by park:01, we make use of the following seminal result of jennrich:1969: If the `identification criterion' $ D_n(\theta) \coloneqq Q_n(\theta)-Q_n(\theta_0), $ scaled suitably by some sequence \(\nu_n \rightarrow \infty\), converges uniformly in probability to a continuous (deterministic) function that is uniquely minimised at \(\theta = \theta_0\), then, \(\theta_n \rightarrow_p \theta_0\). This allows us to establish the consistency of the NLS estimator as summarized below:

proposition{\textcolor{white}{.}} \begin{enumerate} • If Assumptions (ref), (ref), and (ref) are satisfied, then $ \operatorname*{\textnormal{\textsf{sup}}}_{\theta \in \Theta}|\frac1{n}D_n(\theta) - {\sf D}_{m}(\theta)| \rightarrow_p 0,$ where ${\sf D}_{m}(\theta) \coloneqq \sum_{s,k=l}^u\left[1\{s=k\}-\frac{1}{m}\right]{\sf D}_{s,k,m}(\theta)$, with \begin{align*} \normalfont D_{s,k,m}(\theta) \coloneqq &\ \beta^2\sum_{i=\floor{\gamma}}^s\sum_{j=\floor{\gamma}}^k\kappa_{i,s}(\gamma)\kappa_{j,k}(\gamma)c(k-s+i-j)\\ &+\beta_0^2\sum_{i=\floor{\gamma_0}}^s\sum_{j=\floor{\gamma_0}}^k\kappa_{i,s}(\gamma_0)\kappa_{j,k}(\gamma_0)c(k-s+i-j)\\ &-2\beta\beta_0\sum_{i=\floor{\gamma_0}}^s\sum_{j=\floor{\gamma}}^k\kappa_{i,s}(\gamma_0)\kappa_{j,k}(\gamma)c(k-s+i-j). \end{align*} • If Assumptions (ref), (ref), and (ref) are satisfied, then $ \operatorname*{\textnormal{\textsf{sup}}}_{\theta \in \Theta}|\frac1{n\operatorname*{\textnormal{\textsf{ln}}}(n)}D_n(\theta) - {\sf D}(\theta)| \rightarrow_p 0, $ where \({\sf D}(\theta) \coloneqq (\lambda\omega)^2(\beta^2\varphi(\gamma,\gamma)+\beta_0^2\varphi(\gamma_0,\gamma_0)-2\beta\beta_0\varphi(\gamma,\gamma_0))\), with $\omega^2$ and \(\varphi(\cdot,\cdot)\) defined in Eq.\ (ref) while \( \lambda^2 \coloneqq \lambda_1(1-\lambda_2). \) \end{enumerate} If, in addition, \(\beta_0 \neq 0\), then \(\normalfont\theta \mapsto \textsf{D}_{m}(\theta)\) and \(\normalfont\theta \mapsto \textsf{D}(\theta)\) are uniquely minimised at \(\theta = \theta_0\). Consequently, \(\theta_n \rightarrow_p \theta_0\).

Under Assumption (ref), the population criterion function \(\theta \mapsto \textsf{D}_m(\theta)\) can be viewed as a quadratic form of the (Toeplitz) covariance matrix \(\{c(|j-i|)\}_{0\leq i \leq j \leq u}\), that is positive definite under Assumption (ref). Note how the population objective function for fixed $m$ is smooth in $\beta$ but viewed as a function of the gain parameter $\gamma \mapsto {\sf D}_m(\beta,\gamma)$ lacks differentiability. Also under the asymptotic regime of Assumption (ref), the map \(\theta \mapsto \textsf{D}(\theta)\) is non-negative as inspection of the function \(\varphi(\gamma_1,\gamma_2)\) reveals. Moreover, \(\textsf{D}(\cdot)\) can be viewed as the smooth limit of the rescaled \(\textsf{D}_m(\cdot)\), i.e. \({\operatorname*{\textnormal{\textsf{ln}}}}^{-1}(n)\textsf{D}_m(\theta) \rightarrow \textsf{D}(\theta)/(1-\lambda_2)\) for \(m \rightarrow \infty\) as \(n \rightarrow \infty\). Finally, the factor $\lambda^2 = \operatorname*{\textnormal{\textsf{lim}}}\limits_{n \rightarrow \infty} \operatorname*{\textnormal{\textsf{ln}}}(u/l)/\operatorname*{\textnormal{\textsf{ln}}}(n)(1-u/n)$ in Proposition (ref), Part 2, can be interpreted as the asymptotic relative proportion of cohorts to time periods.

An intriguing aspect of Proposition (ref) is the different scaling of the objective function. Specifically, the scaling depends on whether $m$ is fixed or $m$ diverges. In the former case, the scaling is $\nu_n = n$, while in the latter, it is $\nu_n = n\operatorname*{\textnormal{\textsf{ln}}}(n)$. The following example is intended to provide further intuition on this point:

exampleConsider the case of the scalar (OLS) estimator $\theta_n$ of $\theta = \beta$ in case $\gamma = 1$ is known. Here, it is known that the convergence rate of the estimator is determined by the scaling $\nu_n$, say, needed to stabilize the regresser second sample-moment, which, when scaled by $\nu_n$, is given by $H_n \coloneqq \frac1{\nu_n}\sum_{t=u+1}^n\sum_{s=l}^u \tilde a_{t,t-s}^2$. Assuming $\textnormal{\textsf{var}}[\varepsilon] = 1$ and $c(\tau) = 1\{\tau = 0\}$, then we get under Assumption (ref) with $\nu_n = n\operatorname*{\textnormal{\textsf{ln}}}(n)$ \begin{align} E[H_n] = \,& \left(1-\frac{u}{n}\right)\frac1{\operatorname*{ln}(n)} \bigg[\left(1-\frac1{m}\right)(\psi(u+1)-\psi(l)) - 2\left(1-\frac1{m}\right)\\ \,& \qquad \qquad \quad + \frac{2l}{m}(\psi(u+1)-\psi(l+1))\bigg] = \left(1-\frac{u}{n}\right)\frac{\operatorname*{ln}(u/l)}{\operatorname*{\textnormal{\textsf{ln}}}(n)}+O\left(\frac1{\operatorname*{\textnormal{\textsf{ln}}}(n)}\right) \nonumber, \end{align} for the digamma function $\psi(\cdot)$ (see the appendix for details). That is, the scaling by $\nu_n = n\operatorname*{\textnormal{\textsf{ln}}}(n)$ ensures that the expected regressor second-moment stabilizes and converges to the asymptotic relative sample-size $\lambda^2 = \operatorname*{\textnormal{\textsf{lim}}}\limits_{n \rightarrow \infty} \operatorname*{\textnormal{\textsf{ln}}}(u/l)/\operatorname*{\textnormal{\textsf{ln}}}(n)(1-u/n)$ as $m=m(n) \rightarrow \infty$ with $n \rightarrow \infty$. If, on the other hand, under Assumption (ref) only $n$ diverges and $m$ is fixed, then inspection of Eq. (ref) reveals that scaling by $\nu_n = n$ suffices to obtain a nondegenerate limit.

Beyond the special case treated in the preceding example, one might want to make a more general statement about the convergence rate of the $\theta_n$. While the convergence rate of the estimator can be directly deduced from the limiting distribution derived in the next section under the asymptotic regime (ref), the same approach cannot be taken if $m$ is fixed. The reason is that under asymptotic regime (ref), the {\it population} objective is not differentiable, which, however, is an indispensable requirement to derive the limiting distribution. Instead, to derive the rate of convergence under Assumption (ref), we make use of van:1996, exploiting a stochastic Lipschitz bound on $\gamma \mapsto a_{t,t-s}(\gamma)$. Due to the non-differentiability of $\gamma \mapsto {\sf D}_m(\beta,\gamma)$ the remaining difficulty lies in verifying $-{\sf D}_m(\theta) \leq - C\Vert \theta-\theta_0 \Vert^2$ for all $\theta$ in a neighbourhood of $\theta_0$ and some finite $C>0$. Because a standard (Taylor) expansion approach fails, our argument instead rests on deriving the subgradient. In doing so, as we saw in Proposition (ref) already, we have to exclude the case $\beta_0 = 0$ to {\it jointly} identify $\beta$ and $\gamma$. However\footnote{We are grateful to a reviewer for pointing this out.}, {\it individually}, the first element $\theta_{\beta,n}$ of $\theta_n \coloneqq (\theta_{\beta,n},\theta_{\gamma,n})^{\textnormal{\textsf{T}}}$, still estimates $\beta_0 = 0$ consistently. In showing this, we adapt the discussion in saikkonen:95 and seo:11 that both build on wu:81. The preceding discussion can be summarized as follows:

corollaryUnder the conditions of Proposition (ref), $\sqrt{\nu_n}\Vert \theta_n-\theta_0\Vert = O_p(1)$, while, if $\beta_0 = 0$, $\sqrt{\nu_n}|\theta_{n,\beta}| = O_p(1)$, where $\nu_n = n$ or $\nu_n = n\operatorname*{\textnormal{\textsf{ln}}}(n)$ depending on whether Assumption (ref) or (ref) holds, respectively.

When comparing the convergence rates under both asymptotic regimes, we note that the rather slow additional \(\operatorname*{\textnormal{\textsf{ln}}}(n)\)-factor of the convergence rate \(\nu_n = n\operatorname*{\textnormal{\textsf{ln}}}(n)\) under Assumption (ref) arises due to the trending behaviour of the data mentioned earlier and captures the variation as cohorts grow older (\(t-s\)); this is similar in nature to the convergence rates featuring in earlier related work in models with macroeconomic time series (see, e.g., chrismass:18 or mayer:22). It is thus only the variation across cohorts (\(s\)) that leads to the convergence factor \(n\), thereby making the NLS estimator practically appealing. Or, in the words of mn:2016 the “cross-sectional heterogeneity\ ...\ provides a new source of identification”. Our results therefore provide a rigorous justification for this assertion.

Asymptotic normality

As discussed before, the (centred) objective $D_n(\cdot)$ is not differentiable on the set of “kink points” where the gain is integer-valued. This complicates the proof of asymptotic normality. However, as discussed in newmc:94, a less restrictive notion of smoothness called stochastic differentiability can bypass the common requirement that the sample objective is differentiable twice, provided the population objective is sufficiently smooth (see e.g. srisuma2013supplement, oh2013simulated, or mayerwied23 for similar arguments). As Proposition (ref) reveals, this requires that $m \rightarrow \infty$, because even the population objective ${\sf D}_m(\cdot)$, that obtains for $m$ fixed, is not differentiable when viewed as a function of $\gamma.$

To that end, we approximate in a first step $D_n(\cdot)$ with a smooth counterpart $D^\dagger_n(\cdot)$, say. More specifically, under the asymptotic regime of Assumption (ref), results from analytical number theory can be used to obtain a smooth approximation: \[ \operatorname*{\textnormal{\textsf{sup}}}\limits_{\theta\in \Theta}\nu_n^{-1}|D_n(\theta)-D_n^\dagger(\theta)| = o_p(1), \quad \nu_n = n \operatorname*{\textnormal{\textsf{ln}}}(n), \] where $D^\dagger_n(\cdot)$, satisfying the stochasticly differentiability mentioned above and whose probability limit coincides with the smooth function $\theta \mapsto \textsf{D}(\theta)$, is given by

align[align omitted — 249 chars of source]

with $r_{t,t-s}(\gamma) \coloneqq \sum_{i=1}^s h_{i,s}(\gamma)y_{t-s+i}$, $h_{i,s}(\gamma) \coloneqq \frac{\gamma}{s^\gamma}i^{\gamma-1}$. Note that $D_n(\beta,1) = D_n^\dagger(\beta,1)$; see Lemma (ref) of the appendix for details. Based on seminal results for extremum estimators with non-smooth objective function collected in newmc:94, we can then derive the following proposition.

propositionIf Assumptions (ref), (ref) and (ref) are satisfied and $\beta_0 \neq 0$, then \[\normalfont \sqrt{\nu_n}(\theta_n-\theta_0) = -\left(\frac1{2}\frac{\partial^2}{\partial\theta\partial\theta^{\textnormal{\textsf{T}}}}\textsf{D}(\theta)\bigg\rvert_{ \theta=\theta_{\scalebox{.45}{0}}}\right)^{-1} \frac1{\sqrt{4\nu_n}}\frac{\partial}{\partial\theta} D^\dagger_n(\theta)\bigg\rvert_{ \theta=\theta_{\scalebox{.45}{0}}} + o_p(1), \] where \[\normalfont \frac1{\sqrt{4\nu_n}}\frac{\partial}{\partial\theta} D^\dagger_n(\theta)\bigg\rvert_{ \theta=\theta_{\scalebox{.45}{0}}} \rightarrow_d \mathcal{N}_2(0,(\sigma \omega\lambda)^2{\sf H}), \quad \normalfont\textsf{H} \coloneqq \textsf{H}(\theta_0), \] with $\normalfont\frac1{2}\frac{\partial^2}{\partial\theta\partial\theta^{\textnormal{\textsf{T}}}}\textsf{D}(\theta) = (\omega\lambda)^2{\sf H}(\theta)$, \[\normalfont \normalfont\sqrt{\nu_n}(\theta_n-\theta_0) \rightarrow_d \mathcal{N}_2\left(0,\left(\frac{\sigma}{\omega\lambda}\right)^2\textsf{H}^{-1}\right), \quad \textsf{H}(\theta)\coloneqq \varphi(\gamma,\gamma) \begin{bmatrix} 1 & \frac{\beta}{\gamma}\frac{\gamma-1}{(2\gamma-1)}\\ \frac{\beta}{\gamma}\frac{\gamma-1}{(2\gamma-1)} & \frac{\beta^2}{\gamma}\frac{1/\gamma+2(\gamma-1)}{(2\gamma-1)^2}\end{bmatrix}. \]

In general, the variance-covariance matrix depends on the $2 \times 1$ parameter vector of interest $\theta$ through ${\sf H}(\theta)$, which might lead to non-similar inference (see, e.g. nankervis:85). Indeed, as discussed in the following section, the local power of a $t$-test for $H_0:$ $\gamma = \gamma_0$ can be arbitrarily close to the nominal significance level for values of $\gamma_0$ found in empirical studies. If, however, $\gamma_0 = 1$, then ${\sf H}(\theta)$ is a diagonal matrix\footnote{We are grateful to a reviewer for pointing this out.} so that the limiting marginal distributions of the elements of $\theta_n = (\theta_{n,\beta},\theta_{n,\gamma})^{\textnormal{\textsf{T}}}$ are independent of each other and free of the respective parameters itself, i.e. $\nu_n \textnormal{\textsf{var}}[\theta_{\beta,n}] \rightarrow (\sigma/\gamma_0)^2/(\omega\lambda)^{2}$ and $\nu_n\textnormal{\textsf{var}}[\theta_{\gamma,n}]\rightarrow (\sigma/\beta_0)^2/(\omega\lambda)^{2}$. Intuitively, $\gamma_0 = 1$ reduces the weighted least-squares recursion with data-dependent weights in Eq.\ (ref) to an on-line ordinary least-squares estimator free of the nuisance parameter $\gamma$. Finally, we note that the relative asymptotic sample size $\lambda^2$, featuring before in Proposition (ref) and Example (ref), enters inversely the limiting variance-covariance matrix. We could thus also directly scale the estimator with the relative sample size to get a limiting distribution free of $\lambda^2$, i.e. $\sqrt{(n-u)\operatorname*{\textnormal{\textsf{ln}}}(u/l)}(\theta_n-\theta_0) \rightarrow_d \mathcal{N}_2(0,(\sigma/\omega)^2{\sf H}^{-1}),$ see also the discussion of Example (ref).

Standard errors and inference

Standard errors require consistent estimators of the error variance and the Hessian. While the former is consistently estimated using $s_n^2 \coloneqq s_n^2(\theta_n)$, $s^2_n(\theta) \coloneqq \frac1{N}Q_n(\theta)$, estimators of the latter can alternatively be based on any of the following three expressions:

align[align omitted — 249 chars of source]

with, see Eq. (ref) above, $r^{(1)}_{t,t-s}(\gamma) \coloneqq \sum_{j=1}^sh^{(1)}_{j,s}(\gamma)$, $h^{(1)}_{j,s}(\gamma) \coloneqq h_{j,s}(\gamma)(\operatorname*{\textnormal{\textsf{ln}}}(j/s)+1/\gamma)$; or

align[align omitted — 313 chars of source]

with $e_{j,s}(\theta) \coloneqq (h_{j,s}(\gamma),\beta h_{j,s}^{(1)}(\gamma))^{\textnormal{\textsf{T}}}$; or

align[align omitted — 252 chars of source]

where ${\sf H}(\theta_0)$ is given in Proposition (ref). Akin to the discussion of the relative accuracy of observed and expected Fisher information (see e.g. efron:78 or lind:97), we may refer to (ref) and (ref) as `observed' Hessian and `expected' Hessian, respectively, while (ref) is referred to as the `asymptotic' Hessian. We find that, although computationally attractive, an estimator based on the asymptotic Hessian is in finite samples inferior. It is instructive to illustrate these quantities by returning to our discussion of the OLS estimator in Example (ref): Here, the observed Hessian is $H_n(\theta) = H_n = \frac1{\nu_n}\sum_{t=u+1}^n\sum_{s=l}^u \tilde a_{t,t-s}^2$. An analytical expression of the expected Hessian $\textnormal{\textsf{E}}[H_n]$ is given by Eq. (ref), which shows that the asymptotic Hessian --given here by ${\sf H}(\theta) = \lambda^2$-- differs from $\textnormal{\textsf{E}}[H_n]$ by an order of magnitude of ${\operatorname*{\textnormal{\textsf{ln}}}}^{-1}(n)$. Therefore, the use of the latter provides even in large samples only a poor approximation of the finite sample variance $\nu_n\textnormal{\textsf{var}}[\theta_n]$.

Turning back to the general case, we can readily construct estimators from the three different Hessians (ref), (ref), and (ref) using appropriate sample counterparts. We call theses estimators $H_{j,n}$, $j \in \{1,2,3\},$ respectively. Specifically, one obvious estimator is the observed Hessian in (ref) evaluated at $\theta_n$, i.e. $H_{1,n} = H_n(\theta_n)$. From the expected Hessian in (ref) an estimator $H_{2,n} \coloneqq H_{2,n}(\theta_n)$ obtains by replacing the unknown quantities $(\theta_0,\{c(\tau)\}_{\tau=0}^u)$ entering $\textnormal{\textsf{E}}[H_n(\theta_0)]$ with the sample counterparts $(\theta_n,\{c_n(\tau)\}_{\tau=0}^u)$, $c_n(\tau) \coloneqq \frac{1}{n}\sum_{t=1}^{n-\tau}(y_t-\bar y)(y_{t+\tau}-\bar y)$, i.e.

equation[equation omitted — 308 chars of source]

Similarly, we can make (ref) operational via $H_{3,n} \coloneqq H_{3,n}(\theta_n)$, $H_{3,n}(\theta) \coloneqq (\omega_n\lambda_n)^2 {\sf H}(\theta)$, where $\lambda_n^2 \coloneqq \operatorname*{\textnormal{\textsf{ln}}}(u/l)/\operatorname*{\textnormal{\textsf{ln}}}(n)(1-u/n)$ and $\omega_n^2$ is an estimator of the long-run variance $\omega^2$.

Finally, we propose an additional estimator $H_{4,n} \coloneqq H_{4,n}(\theta_n)$, say, that does not exploit the analytical expressions of the Hessian. In particular, because $Q_n(\cdot)$ is not differentiable, this estimator is simply based on a second-order numerical derivative of the objective function with $i$,$j$-th element ($1\leq i,j \leq 2$) given by

equation[equation omitted — 273 chars of source]

for \(\ell_n \searrow 0\) and $e_i$, $i \in \{1,2\}$, denoting some sequence of step-sizes and the $2 \times 1$ unit vector, respectively.

corollarySuppose the assumptions of Proposition (ref) hold, then ($i$) $H_{1,n} \rightarrow_p (\omega\lambda)^2{\sf H}$, ($ii$) $H_{2,n} \rightarrow_p (\omega\lambda)^2{\sf H}$ if $\operatorname*{\textnormal{\textsf{max}}}\limits_{1 \leq \tau \leq u} \sqrt{n}|c_n(\tau)-c(\tau)|= O_p(1)$ and $m = o({\operatorname*{\textnormal{\textsf{ln}}}}(n)\sqrt{n})$, ($iii$) $H_{3,n} \rightarrow_p (\omega\lambda)^2{\sf H}$ if $\omega_n^2 \rightarrow_p \omega^2$, and ($iv$) $H_{4,n} \rightarrow_p (\omega\lambda)^2{\sf H}$ if $\sqrt{\nu_n}\ell_n\rightarrow \infty$, where convergence in probability holds elementwise.

Two comments seem warranted: First, as discussed in xiao:14, the condition in ($ii$) holds under $u = O(n^{\iota})$, $\iota \in (0,1)$, for a wide range of stationary processes $\{y_t\}_t$; the condition in ($iii$) on the long-run variance estimator can be verified for various candidates of $\omega_n^2$ (see, e.g. andrews:91); the condition in ($iv$) on the step-size is common in the literature (see, e.g. newmc:94 or oh2013simulated). Second, we expect $H_{3,n}$, i.e. the estimator based on the asymptotic Hessian, to perform worst among the four estimators in finite, medium, and even large samples. As already suggested by Example (ref), the reason is that ${H}_{3,n}$ provides only a poor approximation of the finite sample Hessian due to a bias of order $O(\operatorname*{\textnormal{\textsf{ln}}}^{-1}(n))$.

As the following corollary reveals, hypotheses of the form \(H_0\): \(R\theta_0 = \rho_0\), for some \(q\times 2\), \(q \leq 2\), restriction matrix \(R\) and \(\rho_0 \in \mathbb{R}^q\), can be tested using the Wald statistic

align[align omitted — 212 chars of source]
corollaryUnder the assumptions of Corollary (ref), \(\normalfont\textsf{W}_{j,n} \rightarrow_d \chi^2(q)\), $j \in \{1,2,3,4\},$ given $H_0$ is true.

Under a sequence $\theta_{0,n} \coloneqq \theta_0 + \Delta/\sqrt{\nu_n}$, $\Delta \coloneqq (\Delta_\beta,\Delta_\gamma)^{\textnormal{\textsf{T}}} \in \mathbb{R}^2$, of so-called `Pitman drifts' the limiting distribution is $\chi^2(q)$ with non-centrality parameter $\kappa \coloneqq (\kappa_\beta,\kappa_\gamma)^{\textnormal{\textsf{T}}} = (\omega\lambda/\sigma)^{2}\Delta^{\textnormal{\textsf{T}}} R^{\textnormal{\textsf{T}}} R {\sf H} R^{\textnormal{\textsf{T}}} R\Delta$, i.e. the test has non-trivial local power in a $\nu_n^{-1/2}$-vicinity around the null. It is instructive to consider the special case of the (squared) $t$-statistics for the hypotheses $H_0$: $\beta = \beta_0$ and $H_0$: $\gamma = \gamma_0$. From the above we then get for the former $\kappa_\beta = (\omega\lambda\gamma\Delta_\beta/\sigma)^2/(2\gamma-1)$, which is independent of $\beta$, and, when viewed as function of $\gamma$, decreasing on $(1/2,1]$ but increasing on $[1,\infty)$. On the other hand, we obtain $\kappa_\gamma = (\omega\lambda\beta\Delta_\gamma/\sigma)^2(1+2\gamma(\gamma-1))/(2\gamma-1)^3$, which is an increasing function of $\beta$ but decreasing in $\gamma$ on the interval $(1/2,\infty)$. Moreover, for $\gamma = 1$, $\kappa_\beta = (\Delta_\beta/\Delta_\gamma)^2\kappa_\gamma/\beta^2 = (\omega\lambda\Delta_\beta/\sigma)^2$. This is illustrated in Figure (ref), depicting the power curves of the $t$-statistics as a function of $\gamma$.

figure[figure omitted — 1,332 chars of source]

Importantly, Corollary (ref) does not allow for hypotheses containing the restriction \(\beta = 0\). Since this testing problem involves a nuisance parameter (viz. $\gamma$) that is not identified under the null, see e.g. andpol:1994 or hansen:1996 and the references therein. More specifically, we follow hansen:1996, hansen:17 and consider the following `supF' statistic

align[align omitted — 268 chars of source]

where $\tilde\sigma_n^2 \coloneqq \frac{1}{N}\sum_{t=u+1}^n\sum_{s=t-u}^{t-l} \tilde z_{t,s}^2$ and $\sigma_n^2(\gamma) \coloneqq \frac{1}{N}Q_n^\star(\gamma)$, with $Q^\star_n(\cdot)$ being the profiled objective defined at the end of Section (ref). The following corollary summarises the limiting behaviour of ${\it sup}{\sf F}$ under the null.

corollarySuppose Assumptions (ref), (ref) are satisfied and \(H_0\): \(\beta = 0\). \begin{itemize} • If Assumption (ref), holds, then \(\normalfont {\it sup}\textsf{F} \rightarrow_d T \coloneqq \operatorname*{\textnormal{\textsf{sup}}}_{\gamma \, \in \, \Gamma}\mathbb{S}_m^2(\gamma)/\varphi_m(\gamma,\gamma), \) where $\mathbb{S}_m(\cdot)$ is a Gaussian process with covariance kernel \[ \varphi_m(\gamma_1,\gamma_2) \coloneq \sum_{s,k=l}^u\left(1\{s=k\}-\frac{1}{m}\right)\sum_{i=\floor{\gamma_1}}^s\sum_{j=\floor{\gamma_2}}^kc(k-s+i-j) \kappa_{i,s}(\gamma_1)\kappa_{j,k}(\gamma_2). \] • If Assumption \textnormal{(ref)}, holds, then \(\normalfont {\it sup}\textsf{F} \rightarrow_d T \coloneqq \operatorname*{\textnormal{\textsf{sup}}}_{\gamma \, \in \, \Gamma}\mathbb{S}^2(\gamma)/\varphi(\gamma,\gamma), \) where $\mathbb{S}(\gamma)$ is a Gaussian process with covariance kernel $\varphi(\cdot,\cdot)$ of Eq. (ref). \end{itemize}

Three aspects of Corollary (ref) are worth exploring. Firstly, as part ($a$) of Corollary (ref) reveals, for this testing problem to be operational it is not necessary to require $m$ to diverge with $n$. To see this, note that, under the null $\beta = 0$, \[ \frac{N(\tilde\sigma_n^2-\sigma^2_n(\gamma))}{\sigma^2_n(\gamma)} = \frac{\sigma^2}{\sigma_n^2(\gamma)}\frac{(\frac1{\sigma\sqrt{\nu_n}}\sum_{t=u+1}^nS_t(\gamma))^2}{\frac1{\nu_n}\sum_{t=u+1}A_t(\gamma,\gamma)}, \] where $S_t(\gamma) \coloneqq \sum_{s=l}^u \tilde a_{t,t-s}(\gamma)\varepsilon_{t,t-s}$ and $A_{t}(\gamma_1,\gamma_2)\coloneqq \sum_{s=l}^u \tilde a_{t,t-s}(\gamma_1) \tilde a_{t,t-s}(\gamma_2)$. Due to a stochastic Lipschitz condition verified in the appendix, we can deduce that (upon scaling by sample size and $\sigma$) $\sum_{t=u+1}^n S_t(\gamma)$ converges weakly to a Gaussian process with kernel coinciding with the limit of the process $\sum_{t=u+1}^n A_t(\gamma,\gamma)$, while, by the same arguments and the LLN, we get $ \sigma_n^2(\gamma) = \frac{1}{N}\sum_{t=u+1}^n\sum_{s=l}^u \varepsilon_{t,t-s}^2(1+O_p(N^{-1})) \rightarrow_p \sigma^2 $ uniformly in $\gamma$. This yields the claim without the need of differentiability with respect to $\gamma$.

Secondly, it follows readily that under a sequence of local alternatives $\beta_{0,n} = \Delta_\beta/\sqrt{\nu_n}$, $\Delta_\beta\in \mathbb R$, Corollary (ref) holds with ${\mathbb S}_m(\cdot)$ and ${\mathbb S}(\cdot)$ replaced by ${\mathbb S}^\star_m(\cdot) = {\mathbb S}_m(\cdot)+\varphi_m(\gamma_0,\cdot)\Delta_\beta/\sigma$ and ${\mathbb S}^\star(\cdot) = {\mathbb S}(\cdot)+\varphi(\gamma_0,\cdot)\Delta_\beta/\sigma$, respectively; thus implying that tests based on ${\sf F}_n$ have non-trivial power in a $\nu_n^{-1/2}$-neighbourhood of the null.

Thirdly, the process in ($b$) can be seen as the limit of that in ($a$) for $m = m(n) \rightarrow \infty$. However, akin to our discussion of standard errors, the convergence rate is slow, i.e. $|\operatorname*{\textnormal{\textsf{ln}}}^{-1}(n)\varphi_m(\gamma_1,\gamma_2)-\lambda_1\varphi(\gamma_1,\gamma_2)|=O(\operatorname*{\textnormal{\textsf{ln}}}^{-1}(n))$. This means that even in large samples it might not be a good idea to use critical values obtained by simulating the limiting process in ($b$). As an alternative, we could simulate the process in ($a$). However, its generation would depend on the autocovariances, implying that such a procedure becomes computationally highly expensive. Instead, to implement the test, we adopt a simple Gaussian multiplier bootstrap proposed by hansen:1996: For each $b \in \{1,\dots,B\}$, let ${\sf F}_{n,b}$ be the test statistic in (ref), where $\tilde z_{t,s}$ is replaced by $z_{t,s,b}$, with $\{z_{t,s,b}: u < t \leq n, t-u \leq s \leq t-l\}$ denoting a sample of {\sf IID} standard normal variates. Next, define the bootstrap $p$-value $p_n \coloneqq 1- {\sf G}_n({\sf F}_n)$, where ${\sf G}_n(t) \coloneqq {\sf P}({\sf F}_{n,b} \leq t \mid {\mathcal S}_n)$ denotes the cumulative distribution function (cdf) of ${\sf F}_{n,b}$, conditional on the data $\mathcal{S}_n \coloneqq \sigma(\{\{z_{t,s}\}_{s=t-u}^{t-l}\}_{t=u+1}^n, \{y_t\}_{t=1}^n)$. As revealed by the following corollary, this bootstrap replicates correctly the first-order asymptotic distribution of the test statistic.

corollaryUnder the conditions of Corollary $\textnormal{\ref{cor:3}}$, $p_n = 1-{\sf G}(T)+o_p(1)$, where $ 1-{\sf G}(T)\sim {\sf Unif}[0,1]$, ${\sf G}(\cdot)$ is the cdf of the random variable $T$ defined in Corollary (ref).

Because ${\sf G}_n(\cdot)$ is unobservable, we simulate it via $G_{n,B}(t) \coloneqq \frac1{B} \sum_{b=1}^B1\{{\sf F}_{n,b} \leq t\}$ and define the simulated $p$-value $p_{n,B} \coloneqq 1-{G}_{n,B}({\sf F}_n).$ The approximation can be made arbitrarily accurate by letting $B \rightarrow \infty$. Thus, for a large value of $B$ and some predefined significance level $\alpha$, we reject the null $\beta = 0$ if $p_{n,B} \leq \alpha$.

Monte Carlo simulation

In our simulation exercise we simulate from the model given by the three equations \ (ref), (ref), and (ref). In particular, the survey expectation at time $t$ formed by individuals born in period $s$, that is, the dependent variable in the nonlinear regression model in (ref), is simulated according to $z_{t,s} = \alpha_t + \beta a_{t,s}(\gamma) + \varepsilon_{t,s}$, with error term $\varepsilon_{t,s} \stackrel{\textsf{IID}}{\sim} \mathcal{N}(0,1/2)$ and time-specific effect $\alpha_t = \xi_t+y_t/2$, $\xi_t \stackrel{\textsf{IID}}{\sim} {\sf UNIF}[0,1].$ The recursion for the nonlinear regression function $a_{t,s}(\gamma)$ evolves according to Eq. (ref) using \(\{y_t\}\), which, in turn, is generated as an AR(1) process $y_t = \rho y_{t-1} + v_t,$ $v_t \stackrel{\textsf{IID}}{\sim} \mathcal{N}(0,1-\rho^2),$ for which we consider a mildly ($\rho=0.50$) and a highly ($\rho=0.99$) dependent scenario. We set the learning parameter in Eq. (ref) to $\gamma \in \{0.8,3.0\}$ and distinguish between the case of identification ($\beta = 0.6$) and non-identification ($\beta = 0$) of the joint parameter vector $\theta$. The case $(\beta,\gamma) = (0.6, 3.0)$ corresponds to our empirical findings. Simulations with more general specifications, including correlated $\varepsilon$ and additional predetermined regressors, do not yield substantial differences when appropriate standard errors are used and are therefore relegated to the Supplementary Material.

Inspired by the empirical application and the analysis in MN, we consider three different sample sizes indexed by $k \in \{2,3,4\}$: \[ n = k\times 150,\quad u = k\times 75, \quad l = 25. \] Numerical optimisation over $\theta$ is based on the optimize routine of the statistical software R (rcovre:21). More specifically, we obtain \(\gamma_n\), the minimizer of the profiled NLS objective discussed in Section (ref) on \(\Gamma = [2/3,10]\), which yields $\beta_n = \beta_n(\gamma_n)$.\footnote{The results numerically very close to jointly minimising $Q_n(\theta)$ based on the BFGS algorithm with $(\beta_n,\gamma_n)$ as starting values. Also, extending \(\Gamma = [2/3,10]\) to \(\Gamma = [1/10,10]\) did not change the results substantially.} Using 1,000 Monte Carlo repetitions, we report the mean and variance of the estimator as well as rejection frequencies of two-sided \(t\)-tests for $H_0:$ \(\gamma = \gamma_0\) and for $H_0:$ \(\beta = \beta_0\) based on asymptotic critical values derived from Corollary (ref). The $t$-statistics, labelled $t_j$, $j \in \{1,2,3,4\}$, are equipped with the four different standard errors discussed in in Section (ref), two of which require specification of additional nuisance parameters: First, in case of (ref), the estimator of the long-run variance $\omega^2$ is chosen to be the estimator in newwest:94 with Bartlet kernel and automated bandwidth selection. Second, the numerical derivative in (ref) is calculated using a tuning parameter \(\ell_n = \delta_{n}(\gamma+\delta_n)\), where $\delta_{n} = \nu_n^{-2/5}$ in accordance with the requirement $\sqrt{\nu_n}\ell_n \rightarrow \infty$ of Corollary (ref). Note that $\delta_n$ is larger by at least one order of magnitude than step-sizes for numerical derivatives typically encountered in statistical software. We also report the empirical rejection frequencies of the {\it sup}{\sf F} statistic using $p$-values obtained from the Gaussian multiplier bootstrap with $B = 99$ (see Corollary (ref)). All test decisions are executed at a nominal significance level of five per cent.

{4.66pt}

table[table omitted — 4,146 chars of source]

In line with Propositions (ref) and (ref), Table (ref) shows that estimation precision increases with sample size. For $\beta = 0.6$, the empirical size of all $t$-tests but $t_3$ becomes reasonably close to the nominal size of five per cent. As expected, $t_3$ performs very poorly even in large samples. The “supF” test, using the Gaussian multiplier bootstrap, appears to consistently reject the alternative $\beta = 0.6$ of the null $\beta = 0$. Next, turn to the scenario under $\beta = 0$, where identification of $\gamma$ breaks down. This is reflected by the poor performance of $\gamma_n$. However, the small sample evidence suggests that we can still consistently estimate $\beta = 0$, thereby corroborating the theoretical result from Corollary (ref). Moreover, we observe that the $t$-statistics for $\beta = 0$ are oversized because of the non-identified gain $\gamma$ under the null. This problem is solved by the use of ${\sf F}_n$.

Empirical application

Reassured by our Monte Carlo evidence in the previous section, we now turn to the empirical analysis of the MSC dataset. The model we consider explains surveyed inflation expectations by age-specific inflation forecast that are learnt from experience, and is given by Eqs.\ (ref), (ref) and (ref) in Section (ref), i.e.\

align[align omitted — 169 chars of source]

and

align[align omitted — 213 chars of source]

The two observed variables in this model are ($i$) the survey expectation of next period's inflation, as recorded in the MSC, and ($ii$) the U.S.\ consumer price index (CPI). Inflation expectations formed by cohort $s$ in time period $t$ are derived from the underlying raw data of the MSC. Numerical MSC micro data are available at a monthly frequency from 1978 onwards\footnote{See the MSC website at \url{https://data.sca.isr.umich.edu/}.}. Following MN, we aggregate these data to quarterly frequency for cohorts that are between 25 and 74 years of age, yielding the cohort-specific inflation expectations $z_{t,s}$ used in Eq.\ (ref) and displayed in Figure (ref).

figure[figure omitted — 2,257 chars of source]

The recursively generated quarterly age-specific inflation forecasts $a_{t,s}$ in Eq.\ (ref) are based, for a given value of $\gamma$, on quarterly U.S.\ CPI $y_t$ which, in turn, is derived from the monthly CPI series originally published\footnote{See Robert Shiller's website at \url{http://www.econ.yale.edu// shiller/data.htm}.} by shiller:00. The result is a sample of in total 8,800 pairs of quarterly observations $(z_{t,s},a_{t,s})$ between 1978Q1 and 2023Q3.

We also consider two variant datasets. One includes pre-1978 archive data of the MSC. These exist, however, only for some intermittent time periods and are not always in the form of quantitative inflation expectations, yet Curtin96 suggests a procedure for making them comparable to post-1978 data. We use the archive data as made available by MN\footnote{See the homepage of Stefan Nagel: \url{https://voices.uchicago.edu/stefannagel/files/2021/06/InflExpCode.zip}.} since they no longer seem to be available for download to the same extent at the MSC\footnote{Only 35 pre-1978 surveys are available for download, see the ICPSR website at \url{https://www.icpsr.umich.edu/web/ICPSR/series/54}.}. Grafting MN's pre-1978 data on the publicly available post-1978 data yields a sample that stretches back to 1953Q4 and comprises 10,615 observation pairs $(z_{t,s}, a_{t,s})$. The second variant dataset, considered for reasons of comparison, is the one used in the analysis of MN, containing 8,215 observations between 1953Q4 and 2009Q4.

The two unknown parameters in model (ref)--(ref), viz. $ \theta = (\beta,\gamma)^{\textnormal{\textsf{T}}} $ are estimated by NLS, as discussed in Section (ref). We use Ox version 9.3, see Doornik07, for the computations. The results for our prime sample from 1978 to 2023 are reported in the first two columns of Table\ (ref). The point estimate $\hat\gamma = 3.16$ of the gain parameter is of a similar order of magnitude as the values found by MadeiraZafar15, mn:2016, Gwak22 and nagel:24, and the 95% confidence interval for the true gain effectively covers the competing estimates despite the different model specification: $[2.74,3.57]$. However, the 95% confidence interval for $\beta$ is $[0.76,0.91]$, comprising values that are statistically different from the estimates reported in MN, Gwak22, and, in particular, MadeiraZafar15. This is evidence indicating that the r\^{o}le of personal experience in forecasting inflation may be higher than has been indicated by the literature so far.

table[table omitted — 2,977 chars of source]

Recall from the discussion in Section (ref) that caution needs to be exercised when testing the null hypothesis $H_0 \colon \beta = 0$. In particular, we argued that the nuisance parameter $\gamma$ is not identified under the null, implying that the size of the usual Wald test is not controlled, see $\textsf W_n$ in Eq.\ (ref) and the Monte Carlo evidence in Section (ref). Instead, we use the `supF'-test by hansen:1996 discussed in Corollary (ref) for testing $H_0$, yielding ${\it sup}{\sf F}n = 476.82$. The distribution of {\it sup}{\sf F} is again approximated by the Gaussian multiplier discussed before, based on 99 bootstrap replications, resulting in a $p$-value of 0.00. This corroborates, for our model, the statement by mn:2016 that “$\beta$ is significantly different from zero”.

Given the joint asymptotic normality of the NLS estimator that we establish in Proposition (ref), it is now also possible to put the hypothesis of `no recency bias' to a test. As explained in Section (ref), this hypothesis can be parameterised as

align*[align* omitted — 108 chars of source]

Computing the corresponding $t$-statistic yields a one-sided $p$-value of $0.00$. Thus, in our model, there is strong empirical evidence to reject the null in favour of MN's conjecture that economic agents weight more recent observations more heavily than distant ones when forecasting inflation.

Estimating the model using the extended dataset from 1953 to 2023 yields 95% confidence intervals for $\beta$ and $\gamma$ comprising parameter values that are considerably lower than for our prime sample, viz.\ $[0.63,0.77]$ and $[2.49, 3.31]$, respectively. Note that the interval for $\beta$ effectively does not overlap with that based on our prime dataset. The intervals based on the MN dataset are similar. It appears that these results are driven by the pre-1978 data, yet we leave it to future research to look in detail at issues such as structural change. Importantly, the tests of $H_0 \colon \beta=0$ and $H_0 \colon \gamma \leq 1$ continue to be soundly rejected when the two variant datasets are used.

Several lessons can thus be learnt from the empirical application: First, using our model to describe post-1978 MSC data, we obtain a confidence interval for the gain parameter $\gamma$, viz.\ $[2.74, 3.57]$, that comprises most parameter estimates found in the literature. This is re-assuring, given that MN's parameter estimate $\hat \gamma= 3.044$ has become something like a yardstick for calibrating `learning from experience' models. Secondly, the hypothesis that there is no recency bias is rejected in our model, which lends support to MN's theory that recent experiences weigh more heavily when agents forecast the future. Thirdly, in line with the thrust of the `learning from experience' literature, our model does not support the conjecture that private experiences do not matter in forecasting inflation. Our empirical evidence indicates that private experiences carry statistically significantly more weight than has been previously reported in the literature by MN, Gwak22, and, in particular, MadeiraZafar15. Finally, it appears as if the aforementioned conclusions are sensitive to the inclusion of pre-1978 MSC archive data. The question of whether this is data issue or a model issue is, however, left to future work.

Concluding remarks

This paper contributes to the burgeoning literature on analysing the heterogeneity in the expectations formation process. In particular, we establish the econometric theory for NLS estimation and inference in nonlinear panels with learning from experience. We show that the estimator is consistent and derive its rate of convergence. However, we find that asymptotic normality may not be obtained when the number of cohorts is small. If, on the other hand, the number of cohorts diverges, we prove that the NLS estimator is asymptotically normal, albeit at a nonstandard convergence rate, using seminal results on extremum estimation with nonsmooth objective functions. Building on this rigorous econometric foundation, we apply our findings to an empirical model of the Michigan Survey of Consumers (MSC) data and confirm conjectures made in the learning-from-experience literature on the gain parameter as well as on the contribution of private experiences.

Our analysis can be seen as a starting point for future extensions. One such research avenue would be to consider the econometric theory of more elaborate forms of belief updating as in the empirical applications of mn:2016, Acedanski17, Gwak22, and nagel:24. For example, as mentioned above, mn:2016 assume that agents use a more general PLM for forecasting inflation, in the sense that the information contained in an additional regressor $x_t$ is taken into account. Specifically, an agent born in period \(s\) estimates in each period \(t\) the parameter of a linear regression $\phi$, say, according to the general stochastic recursive algorithm:

align*[align* omitted — 249 chars of source]

with \(\gamma_{t,s} = \gamma_{t,s}(\gamma_0)\) as in Eq.\ (ref). Similar to ordinary least squares with stochastic regressors, this updating scheme is a Newton-type algorithm that utilizes information on second moments through $r_{t,s}$. Clearly, it is a multivariate generalisation of our setup in that our learning rule in Eq.\ (ref) obtains with \(x_t=1\). Given the recursive estimate of $\phi_{t-1,s}$, the learnt expectation is defined as $ a_{t,s} \coloneqq \phi_{t-1,s}^{\textnormal{\textsf{T}}} x_t$ so that the data generating process of the dependent variable obtains as $z_{t,s} = \alpha_t + \beta a_{t,s} + \varepsilon_{t,s}$. A technical treatment of the NLS estimator of $\beta$ and $\gamma$ would involve analysing the counterpart of the expressions in Eq.\ (ref), with the crucial difference that the weights $\kappa_{j,s}$ are now ($i$) stochastic and ($ii$) dependent on the recursion of $r_{t,s}$. The analytical examination of this generalised model is, however, non-trivial. Some preliminary results are contained in the Supplementary Material, where we investigate the small sample behaviour of the NLS estimator empirically and by simulation. Several of the results we established above seem to carry over.

\singlespacing \addcontentsline{toc}{section}{References}