Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Two-step estimation in linear regressions with adaptive learning
\newgeometry{margin=1in}
titlepage\thispagestyle{empty}
\begin{abstract}
Weak consistency and asymptotic normality of the ordinary least-squares estimator in a linear regression with adaptive learning is derived when the crucial, so-called, `gain' parameter is estimated in a first step by nonlinear least squares from an auxiliary model.
\end{abstract}
Keywords: Stochastic approximation, nonlinear least squares, asymptotic collinearity, generated regressor.
\restoregeometry
Introduction
Consider a linear regression with adaptive learning
equation[equation omitted — 94 chars of source]
where \(\varepsilon_t\) is an error term, \(\delta_0 \in \mathbb{R}\) and the parameter space of \(\beta_0\) is specified below. The variable \(a_{t} \coloneqq a_t(\theta_0)\) is updated according to a stochastic approximation algorithm
equation[equation omitted — 113 chars of source]
with \(a_0(\theta_0) = \mathfrak{a}_0\) for some initial value \(\mathfrak{a}_0\). The so-called `gain' sequence \(\{\theta_0/t\}_{t\, \geq \,1}\) of positive constants captures the responsiveness to previous prediction mistakes and depends on an unknown gain parameter \(\theta_0\). For \(\theta_0 = 1\), recursion (ref) reduces to recursive least-squares, which updates \(a_t\) by weighing all observations equally. If, in turn, \(\theta_0 > 1\) (\(\theta_0 < 1\)), then earlier observations receive less (more) weight. The dynamics of (ref) with this gain specification have been investigated by masa89 who interpret \(\theta_0\) as a `forgetting factor'; see also nanu15, malmnag16, or malmetal21 for recent economic applications. Typically, more general formulations of the type \(y_t = \beta_0 a_{t-1}x_t + \delta_0x_t + \varepsilon_t\) are considered within this literature, where \(x_t\) is some stochastic regressor and \(y_{t|t-1}^\textnormal{\textsf{e}} \coloneqq a_{t-1}x_t\) is viewed as the subjective expectation an economic agent forms about \(y_t\) at time \(t-1\) by estimating the so-called rational expectations equilibrium \(y_t = \alpha_0x_t + \varepsilon_t\), with \(\alpha_0 \coloneqq \delta_0/(1-\beta_0)\); see evho01. For brevity, we leave this extension with stochastic \(x_t\) to future research.
It is well known from the literature on stochastic approximation that the decreasing-gain specification ensures \(a_t \rightarrow \alpha_0\) in an appropriate probabilistic sense; see bmp90 or lai03 for an introduction to this literature. Although the resulting asymptotic collinearity makes estimation of the model parameters \(\lambda_0 \coloneqq (\beta_0,\delta_0)^\textsf{T}\) difficult, limiting normality and/or consistency of the ordinary least-squares estimator of \(\lambda_0\) can be established under certain regularity conditions; see CM18, CM19, and may21. The major limitation of these papers is, however, that \(\theta_0\) is treated as a known constant. In what follows, I will thus weaken this assumption and treat the case where the gain parameter \(\theta_0\) is first estimated by nonlinear techniques from a noisy sample of \(\{a_t\}_{t \geq 1}\), an exercise commonly encountered in the empirical macroeconomic learning literature where survey data of the subjective expectations \(y_{t|t-1}^\textnormal{\textsf{e}} = a_{t-1}x_t\) are used to infer the unknown gain parameter \(\theta_0\); see, e.g., braev6, mapi14, malmnag16, or bega17. It is, however, so far unclear which (if any) statistical properties any such estimator possesses.
The primary objective of this note is thus to derive weak consistency and asymptotic normality of the prominent nonlinear least-squares estimator of \(\theta_0\). This is achieved by combining classical results from jenn69 and lai94 with more recent arguments due to chawa15 and wang21. In addition, the following section also derives the limiting distribution of the feasible ordinary least-squares estimator of \(\lambda_0\) based on the estimated gain and highlights asymptotic equivalence with its unfeasible counterpart under certain parameter combinations, thereby complementing earlier results from CM18 and may21 (CMM, henceforth). All proofs are contained in the online appendix.
Asymptotic Analysis
Nonlinear Least-Squares Estimation of \(\theta_0\)
Suppose a sample \(\{z_1,\dots,z_n\}\) is observed, where
equation[equation omitted — 70 chars of source]
\(u_t\) is an error term and the notational convention \(a_t = a_t(\theta_0)\) is recalled. This setting corresponds to the empirical macroeconomic learning literature cited above, where \(\{z_1,\dots,z_n\}\) is interpreted as noisy survey data of \(y_{t|t-1}^\textnormal{\textsf{e}} = a_{t-1}x_t\) (with the simplifying assumption \(x_t = 1\)), from which the gain parameter is estimated using the least squares method
\[
\hat{\theta}_n \coloneqq \operatorname*{\textsf{arg\,min}}\limits_{\theta \, \in \, [
\underaccent{\bar}{\theta},\bar{\theta}]} Q_n(\theta),\;\;Q_n(\theta) \coloneqq \sum_{t\,=\,1}^n(z_{t}- a_{t-1}(\theta))^2, \quad 0 <
\underaccent{\bar}{\theta} \leq \bar{\theta} < \infty.
\]
Here \(a_{t}(\theta)\) denotes the counterpart of Eq. (ref) that obtains by recursion from Eq. (ref) upon replacing the true \((\theta_0,\mathfrak{a}_0)^\textsf{T}\) with a candidate value \((\theta,\mathfrak{a})^\textsf{T} \in [
\underaccent{\bar}{\theta},\bar{\theta}] \times \mathbb{R}\).\footnote{It is shown in the appendix that the choice of the initial value is asymptotically inconsequential for a suitably restricted parameter space. Thus, in what follows, the dependence of various quantities on the initial value is suppressed.}
Weak Consistency
The consistency result for \(\hat{\theta}_n\) uses a strategy proposed initially by jenn69. Define \(D_n(\theta,\theta_0) \coloneqq Q_n(\theta)-Q_n(\theta_0)\). If, for some sequence \(\nu_n \rightarrow \infty\), \(\nu_n^{-1}D_n(\theta,\theta_0) \stackrel{p}{\rightarrow} D(\theta,\theta_0)\) uniformly in \(\theta \in [
\underaccent{\bar}{\theta},\bar{\theta}]\), where \(\theta \mapsto D(\theta,\theta_0)\) is continuous and has unique minimum \(\theta_0\) $a.s.$, then \(\hat{\theta}_n \stackrel{p}{\rightarrow} \theta_0\). Uniform convergence, in turn, follows from pointwise convergence and a stochastic Lipschitz condition for \(\theta \mapsto Q_n(\theta)\); see, e.g., and92. In order to establish weak consistency of \(\hat{\theta}_n\) along these lines, the following assumptions are imposed.
assumption\(\theta_0 \in (
\underaccent{\bar}{\theta},\bar{\theta})\) so that ($a$) \( 1<
\underaccent{\bar}{\theta} < \bar{\theta} < \infty\) and ($b$) \(\theta_0(1-\beta_0) > 1/2\).
assumption\textcolor[rgb]{1,1,1}{.}
\begin{enumerate}
• \(\{\varepsilon_t,\mathcal{F}_t\}_{t \geq 1}\), \(\mathcal{F}_t \coloneqq \sigma(\{\varepsilon_j\}_{j \, \leq \, t})\), forms a stationary martingale difference sequence such that \(\textnormal{\textsf{E}}[\varepsilon_t^4 \mid \mathcal{F}_{t-1}] < \infty\) a.s. and set \(\sigma_\varepsilon^2 \coloneqq \textnormal{\textsf{E}}[\varepsilon_t^2 \mid \mathcal{F}_{t-1}]\).
• \(\{u_t,\mathcal{I}_t\}_{t \geq 1}\), \(\mathcal{I}_t \coloneqq \sigma(\{u_j, a_{j}\}_{j \, \leq \, t})\), forms a stationary martingale difference sequence such that \(\textnormal{\textsf{E}}[u_t^4 \mid \mathcal{I}_{t-1}] < \infty\) a.s. and set \(\sigma_u^2 \coloneqq \textnormal{\textsf{E}}[u_t^2 \mid \mathcal{I}_{t-1}]\).
• The initial value \(\mathfrak{a}_0\) is independently distributed of \(\{u_t\}_{t \geq 1}\) and \(\{\varepsilon_t\}_{t \geq 1}\) such that \(\textnormal{\textsf{E}}[\mathfrak{a}_0^4] < \infty.\)
\end{enumerate}
remark\normalfont
Part ($a$) of Assumption (ref) implies that recent observations are weighed more heavily, a scenario of practical importance as illustrated by, among others, nanu15, malmnag16, or malmetal21. Part ($b$) of Assumption \textnormal{(ref)} is not uncommon for the analysis of decreasing gain stochastic approximation; see, e.g., \textnormal{evho01} and \textnormal{bmp90}. Importantly, \textnormal{($b$)} is equivalent to \(\beta_0 < 1-1/(2\theta_0)\) thereby implying that \(\beta_0 < 1\). More specifically, Assumption \textnormal{(ref)} translates to a short-memory condition for \(y_t\); see \textnormal{chema17} and \textnormal{may21}. Finally, Assumption \textnormal{(ref)} imposes convenient regularity conditions.
As a first result, weak consistency is obtained under these assumptions:
propositionLet \(\nu_n = \textnormal{\textsf{log}}\,n\) and suppose Assumptions (ref) and (ref) are satisfied.
\begin{enumerate}
• For each \(\theta \in [
\underaccent{\bar}{\theta},\bar{\theta}]\), \(\nu_n^{-1}D_n(\theta,\theta_0) \stackrel{p}{\rightarrow} D(\theta,\theta_0) \coloneqq (\theta-\theta_0)^2 d(\theta,\theta_0)\), where
\[\normalfont
d(\theta,\theta_0)\coloneqq \left(\frac{\sigma_a}{\theta_0}\right)^2\frac{(2(1-\beta_0)\theta_0-1)\theta-((1-\beta_0)\theta_0-1)}{(2\theta-1)(\theta_0(1-\beta_0)+\theta-1)},
\]
with \(\normalfont\sigma_a^2 \coloneqq \sigma_\varepsilon^2\theta_0^2/(2(1-\beta_0)\theta_0-1)\); \(D(\theta,\theta_0)\) is continuous and uniquely minimized at \(\theta = \theta_0\).
• For each \(\theta_1,\theta_2 \in [
\underaccent{\bar}{\theta},\bar{\theta}]\), \(|D_n(\theta_1,\theta_0)-D_n(\theta_2,\theta_0)| \leq h(|\theta_1-\theta_2|)L_n\) for some nonrandom function \(h(x)\) such that \(h(x) \rightarrow 0\) as \(x \rightarrow 0\) and where \(\nu_n^{-1}L_n = O_p(1)\).
\end{enumerate}
Consequently, \(\hat{\theta}_n \stackrel{p}{\rightarrow} \theta_0\).
remark\normalfont
The identifying criterion \(D(\theta,\theta_0)\) is non-negative due to the restrictions imposed by Assumption (ref). The factor \(\sigma_a^2\) coincides with the limiting variance of \(n^{1/2}(a_n-\alpha_0)\), \(\alpha_0 \coloneqq \delta_0/(1-\beta_0)\); see, e.g., CMM. Using a recent result from wang21 in conjunction with arguments from lai94 and CMM, the stochastic Lipschitz condition is shown to be satisfied for
\[\normalfont
L_n = \operatorname*{\textnormal{\textsf{sup}}}\limits_{\theta \, \in \, [
\underaccent{\bar}{\theta},\bar{\theta}]}\operatorname*{\textnormal{\textsf{max}}}\left\{\left\vert\sum_{t\,=\,1}^n \dot{a}_{t-1}(\theta)u_t\right\vert,\,\sum_{t\,=\,1}^n \dot{a}_{t-1}^2(\theta)\right\},\;\; \dot{a}_t(\theta) \coloneqq \frac{\textsf{d}\,a_t(\theta)}{\textsf{d}\,\theta}.
\]
Limiting Normality
In order to derive the limiting distribution, the framework of chawa15 is adopted. They show that if
equation[equation omitted — 573 chars of source]
and
equation[equation omitted — 217 chars of source]
holds on \(\Theta_n \coloneqq \{\theta \in [
\underaccent{\bar}{\theta},\bar{\theta}]: \textnormal{\textsf{log}}^{1/2}n|\theta-\theta_0| \leq k_n\}\) for some sequence \(k_n \rightarrow \infty\) as \(n \rightarrow \infty\), then \(\textnormal{\textsf{log}}^{1/2}n(\hat{\theta}_n-\theta_0) = Z_n+o_p(1)\),
with
\[
Z_n \coloneqq \textnormal{\textsf{log}}^{1/2}n\left[\sum_{t\,=\,1}^n \dot{a}_t^2(\theta_0)\right]^{-1}\sum_{t\,=\,1}^n \dot{a}_t(\theta_0)u_t.
\]
The latter quantity, in turn, can be shown to converge weakly to a Gaussian random variable. Proposition (ref) summarizes this argument:
propositionIf the conditions of Proposition (ref) are satisfied, then \(\normalfont\textnormal{\textsf{log}}^{1/2}n(\hat{\theta}_n-\theta_0) = Z_n + o_p(1)\), with \( \normalfont Z_n \stackrel{d}{\rightarrow}\textnormal{\textsf{N}}(0,(\sigma_u/\sigma_{\dot{a}})^2)\), where \(\sigma_{\dot{a}}^2 \coloneqq d(\theta_0,\theta_0)\) and \(d(\theta,\theta_0)\) has been defined in Proposition (ref) so that
\[\normalfont
\sigma_{\dot{a}}^2 = d(\theta_0,\theta_0) = \left(\frac{\sigma_a}{\theta_0}\right)^2 \frac{\theta_0(2-\beta_0)-2\theta_0^2(1-\beta_0)-1}{(2\theta_0-1)(1-\theta_0(2-\beta_0))}.
\]
remark\normalfont The limit distribution follows from a CLT due to dav93 that allows for degenerate variances. Note that \(\sigma_{\dot{a}} > 0\) is ensured by Assumption (ref).
Two-Step Estimation of \(\lambda_0\)
Since (ref) is linear in \(\lambda_0\), it is reasonable to estimate \(\lambda_0\) based on the least-squares estimator \(\hat{\lambda}_n(\hat{\theta}_n)\), where \(\hat{\lambda}_n(\theta) \coloneqq \left[\sum_{t\,=\,1}^nw_t(\theta)w_t(\theta)^\textsf{T}\right]^{-1}\sum_{t\,=\,1}^n w_t(\theta)y_t\), with \(w_t(\theta) \coloneqq (1, a_{t-1}(\theta))^\textsf{T}\).
To gain some intuition for the following result, it is instructive
to note that a second-order Taylor series expansion yields
equation[equation omitted — 272 chars of source]
with \(Z_n\) from Proposition (ref), \(S_n(\theta) \coloneqq \textnormal{\textsf{log}}^{-1/2}n\sum_{t\,=\,1}^n (a_{t-1}(\theta)-\alpha_0)(\varepsilon_t+\beta_0(a_{t-1}(\theta)-a_{t-1}))\), and \(\dot{S}_n(\theta) = \textsf{d}S_n(\theta)/(\textsf{d}\theta)\). Since CMM have shown that \(S_n(\theta_0)\) is asymptotically normal, the following limiting results are mainly due to Proposition (ref) and the fact that the first-order term of the Taylor series expansion obeys
equation[equation omitted — 355 chars of source]
Proposition (ref) formalizes this argument:
propositionSuppose the conditions of Proposition (ref) are satisfied. If, in addition, \(\{u_t\}_{t \geq 1}\) and \(\{\varepsilon_t\}_{t \geq 1}\) are independent, then
\[\normalfont
\textnormal{\textsf{log}}^{1/2}n(\hat{\lambda}_n(\hat{\theta}_n)-\lambda_0) \stackrel{d}{\rightarrow} \textnormal{\textsf{N}}(0_2,(1+B(\theta_0,\beta_0))V_0), \;\; V_0 \coloneqq \frac{2(1-\beta_0)\theta_0-1}{\theta_0^2} \begin{bmatrix} \alpha_0^2 & -\alpha_0 \\ -\alpha_0 & 1 \end{bmatrix},
\]
where \(\normalfont0_2 \coloneqq (0,0)^\textsf{T}\) and
\[
B(\theta_0,\beta_0) \coloneqq \frac{\beta_0^2(2\theta_0-1)((1-\beta_0)\theta_0-1)^2}{((2-\beta_0)\theta_0-1)(1 + \theta_0((1-\beta_0)(2\theta_0-1)-1 ))}\left(\frac{\sigma_u}{\sigma_\varepsilon}\right)^2.
\]
remark\normalfont The same logarithmic convergence rate has been found by CMM for the infeasible ordinary least-squares estimator treating the gain as known. The factor \(B(\theta_0,\beta_0) \geq 0\) is due to the presence of the generated covariate. This generated regressor issue vanishes if either (a) \(\beta_0 = 0\) or (b) \((1-\beta_0)\theta_0 = 1 \Leftrightarrow \beta_0 = 1-\theta_0^{-1}\), because in both cases \(B(\theta_0,\beta_0) = 0\); i.e., the singular limiting distribution coincides with that of the infeasible estimator \(\hat{\lambda}_n(\theta_0)\); see CMM. The asymptotic equivalence is well known in linear regression if the coefficient on the generated covariate is zero; see, e.g., \textnormal{pag84}. Condition \textnormal{(\textit{b})} \((1-\beta_0)\theta_0 = 1\) ensures, in an appropriate probabilistic sense, orthogonality between \(a_t\) and \(\dot{a}_t(\theta_0)\). Thus, in view of Eqs. \textnormal{(ref)} and \textnormal{(ref)}, conditions \textnormal{(\textit{a})} and \textnormal{(\textit{b})} have the same effect on the limiting distribution in that the first-order term in Eq. \textnormal{(ref)} vanishes if they hold true. Under both conditions, the limiting distributions of \(\textnormal{\textsf{log}}^{1/2}n(\hat{\theta}_n-\theta_0)\), \(n^{1/2}(a_{n}-\alpha_0)\), and \(\textnormal{\textsf{log}}^{1/2}n(\hat{\lambda}_n(\theta_0)-\lambda_0)\) do not depend \textnormal{(\textit{directly})} on \(\beta_0\), the coefficient on \(a_{t-1}\) in model (ref). Finally, independence between \(\{u_t\}_{t \geq 1}\) and \(\{\varepsilon_t\}_{t \geq 1}\), although not essential for deriving Proposition \textnormal{(ref)}, greatly simplifies notation with little loss of insight.
Joint Estimation
Let me conclude by briefly commenting on two alternative estimation procedures that do not require an auxiliary model based on a noisy sample like Eq. (ref). More specifically, instead of estimating \(\lambda_0 = (\delta_0,\beta_0)^\textsf{T}\) and \(\theta_0\) sequentially, one could also try to estimate \(\gamma_0 \coloneqq (\delta_0,\beta_0,\theta_0)^\textsf{T}\) jointly based directly on Eq. (ref); see, e.g., mi07, chemama10, or bega17b for empirical applications of joint estimation procedures. In the current context, this would mean to use the following nonlinear least-squares estimator
\[
\hat{\gamma}_n \coloneqq \operatorname*{\textsf{arg\,min}}_{\gamma \, \in \, \Gamma} Q_{1,n}(\gamma),\;\; Q_{1,n}(\gamma) \coloneqq \sum_{t\,=\,1}^n(y_t-r_t(\gamma))^2,\;\; r_t(\gamma) \coloneqq \delta+\beta a_{t-1}(\gamma),
\]
with \(\Gamma \coloneqq [
\underaccent{\bar}{\delta},\bar{\delta}] \times [
\underaccent{\bar}{\beta},\bar{\beta}] \times [
\underaccent{\bar}{\theta},\bar{\theta}]\), where \(-\infty <
\underaccent{\bar}{\delta} < \bar{\delta} < \infty\) and \(-\infty <
\underaccent{\bar}{\beta} < \bar{\beta} < 1\). This joint estimator fails, however, important sufficient conditions for nonlinear regression. First, there does not exist a sequence \(\{\tau_n\}_{n \geq 1}\), with \(\tau_n \rightarrow \infty\), such that the sample second moment matrix \(M_n(\gamma) \coloneqq \sum_{t\,=\,1}^n\dot{r}_t(\gamma)\dot{r}_t(\gamma)^\textsf{T}\) of the \(3\times 1\) gradient vector \(\dot{r}_t(\gamma) = (1,a_{t-1}(\theta),\beta \dot{a}_{t-1}(\theta))^\textsf{T}\), evaluated at the true parameter values and scaled by \(\tau_n^{-1}\), converges in probability to an invertible matrix, violating, for example, Eq. (2.4) in lai94. Second, jenn69's jenn69 condition for weak consistency is violated, as there does not exist a sequence \(\{\nu_n\}_{n \geq 1}\), with \(\nu_n \rightarrow \infty\), such that \(D_{1,n}(\gamma,\gamma_0) \coloneqq Q_{1,n}(\gamma)-Q_{1,n}(\gamma_0)\) converges, when scaled by \(\nu_n^{-1}\), uniformly in probability to a deterministic function with unique minimum. More specifically, we have on \(\Lambda \coloneqq \{\gamma \in \Gamma: \delta = \alpha_0(1-\beta)\} \subset \Gamma\)
\[
\operatorname*{\textnormal{\textsf{sup}}}\limits_{\gamma \, \in \, \Lambda}|\textnormal{\textsf{log}}^{-1}n D_{1,n}(\gamma,\gamma_0)-D_1(\kappa,\kappa_0)| = o_p(1),\;\; \kappa \coloneqq (\beta,\theta)^\textsf{T},
\]
where
equation[equation omitted — 394 chars of source]
is a continuous non-negative function of \(\kappa\) that attains its minimum of zero if either (1) \(\kappa = \kappa_0\) or, irrespective of whether \(\theta \neq \theta_0\), (2) \(\beta = \beta_0 = 0\). The non-uniqueness due to case (2) reflects an identification failure resulting from the non-separability of \(\beta\) and \(\theta\) in Eq. (ref). On the other hand,
\(
\operatorname*{\textnormal{\textsf{sup}}}\limits_{\gamma \, \in \, \Gamma}n^{-1}|D_{1,n}(\gamma,\gamma_0)-\tilde{D}_1(\lambda,\lambda_0)|=o_p(1),
\)
where \(\tilde{D}_1(\lambda,\lambda_0) \coloneqq (\delta-\alpha_0(1-\beta))^2\), with \(\tilde{D}_1(\lambda,\lambda_0)= 0\) uniformly on \(\Lambda\).\footnote{Interestingly, the estimator of chemama10 is subject to a similar underidentification issue under deacreasing-gain learning; see may21.} Even though these conditions are not necessary, the question of whether weak consistency of \(\hat{\gamma}_n\) or even asymptotic normality can be shown in spite of their violation is beyond the scope of this note.
A little more can be said about the following estimator: Again, the starting point is Eq. (ref), which, after a few rearrangements, yields \(y_t^\star = \beta_0 a_{t-1}^\star + \varepsilon_t\), with \(y_t^\star \coloneqq y_t - \alpha_0\), \(a_t^\star(\theta) \coloneqq a_t(\theta) - \alpha_0\), and \(a_t^\star \coloneqq a_t^\star(\theta_0)\). Treating, for the sake of argument, \(\alpha_0\) as known, we could first estimate \(\kappa_0 = (\beta_0,\theta_0)^\textsf{T}\) using
equation[equation omitted — 268 chars of source]
where \(\Xi \coloneqq [
\underaccent{\bar}{\beta},\bar{\beta}]\times[
\underaccent{\bar}{\theta},\bar{\theta}]\). Contrary to the preceding discussion, the suitably scaled second moment matrix of the \(2 \times 1\) gradient vector \(\dot{\ell}_t(\kappa) = (a_{t-1}^\star(\theta),\beta\dot{a}_{t-1}(\theta))^\textsf{T}\) is positive definite with probability approaching one; i.e.
\[
log^{-1}n \sum_{t\,=\,1}^n \dot{\ell}_t(\kappa_0)\dot{\ell}_t(\kappa_0)^T \stackrel{p}{\rightarrow} K \coloneqq
bmatrix[bmatrix omitted — 110 chars of source]
,
\]
where \(K\) is positive definite as long as \(\beta_0 \neq 0\). Moreover, jenn69's jenn69 condition for weak consistency is satisfied if \(\beta_0 \neq 0\) as \(\textnormal{\textsf{log}}^{-1}n D_{2,n}(\kappa,\kappa_0)\) converges uniformly in probability to \(D_1(\kappa,\kappa_0)\), which, as mentioned previously, is uniquely minimized at \(\kappa_0\) as long as \(\beta_0 \neq 0\). Thus, by ruling out \(\beta_0 = 0\), we obtain from the above and the results contained in the appendix that
\(
\textnormal{\textsf{log}}^{1/2}n(\hat{\kappa}_n-\kappa_0) \stackrel{d}{\rightarrow} \mathcal{N}(0_2,V_2),\;\; V_2 \coloneqq \sigma_\varepsilon^2K^{-1}.
\)
Since, by may21, \(\hat{\alpha}_n = \alpha_0 + O_p(n^{-1/2})\), replacing \(\alpha_0\) with \(\hat{\alpha}_n\) in (ref) is asymptotically inconsequential, making \(\hat{\kappa}_n\) a fully feasible estimator. Moreover, in view of \(\alpha_0 = \delta_0/(1-\beta_0)\), an estimator of \(\delta_0\) can be given by \(\hat{\kappa}_{n,\delta} \coloneqq \hat{\alpha}_n(1-\hat{\kappa}_{n,\beta})\), where \(\hat{\kappa}_{n,\beta}\) is the first component of \(\hat{\kappa}_n\). Importantly, if \((1-\beta_0)\theta_0 = 1\), then \(M\) is diagonal and the limiting distributions of both the infeasible and the feasible least-squares estimators from the previous section coincide with that of \(\textnormal{\textsf{log}}^{1/2}n((\hat{\kappa}_{n,\delta},\hat{\kappa}_{n,\beta})^\textsf{T}-\lambda_0)\). Recall that \((1-\beta_0)\theta_0 = 1\) played a crucial role within the context of Proposition (ref). Moreover, the limiting variance of the second component of \(\hat{\kappa}_n\) associated with \(\theta_0\), \(\hat{\kappa}_{n,\theta}\), say, is given by \((2\theta_0-1)(\theta_0(2-\beta_0)-1)^2/(\beta_0\theta_0)^2\),
which, contrary to the limiting variance of \(\hat{\theta}_n\) from Proposition (ref), does not depend on innovation variances.
Conclusion
This note derives the singular limiting distribution of the two-step least-squares estimator in a linear regression with estimated gain. Although the limiting distribution is contaminated by the presence of a generated covariate, this generated-regressor problem disappears for certain parameter combinations.
\nolinenumbers
\addcontentsline{toc}{section}{References}
thebibliography\bibitem[\citeauthoryear{Andrews}{Andrews}{1992}]{and92}
Andrews (1992).
\newblock Generic uniform convergence.
\newblock {\em Econometric Theory\/} {\em 8}, 241--257. \doi{10.1017/S0266466600012780}
\bibitem[\citeauthoryear{Benveniste et al.}{Benveniste et al.}{1990}]{bmp90}
Benveniste, A., M. M\'{e}tivier, and P. Priouret (1990).
\newblock {\em Adaptive Algorithms and Stochastic Approximation\/}.
\newblock Applications of Mathematics, Volume 22. Berlin: Springer.
\bibitem[\citeauthoryear{Berardi and Galimberti}{Berardi and Galimberti}{2017a}]{bega17}
Berardi, M., and J. K. Galimbeti (2017a).
\newblock Empirical calibration of adaptive learning.
\newblock {\em Journal of Economic Behavior & Organization\/} {\em 144}, 219--237. \doi{10.1016/j.jebo.2017.10.004}
\bibitem[\citeauthoryear{Berardi and Galimberti}{Berardi and Galimberti}{2017b}]{bega17b}
Berardi, M., and J. K. Galimbeti (2017b).
\newblock On the initialization of adaptive learning in macroeconomic models.
\newblock {\em Journal of Economic Dynamics & Control.\/} {\em 78}, 26--53. \doi{10.1016/j.jedc.2017.03.002}
\bibitem[\citeauthoryear{Branch and Evans}{Branch and Evans}{2006}]{braev6}
Branch, W. A., and G. W. Evans (2006).
\newblock A simple recursive forecasting model
\newblock {\em Economics Letters\/} {\em 91}, 158--166. \doi{10.1016/j.econlet.2005.09.005}
\bibitem[\citeauthoryear{Chan and Wang}{Chan and Wang}{2015}]{chawa15}
Chan, N., and Q. Wang (2015).
\newblock Nonlinear regressions with nonstationary time series.
\newblock {\em Journal of Econometrics\/} {\em 185}, 182--195. \doi{10.1016/j.jeconom.2014.04.025}
\bibitem[\citeauthoryear{Chevillon et al.}{Chevillon et al.}{2010}]{chemama10}
Chevillon, G., M. Massmann, and S. Mavroeidis (2010).
\newblock Inference in models with adaptive learning.
\newblock {\em Journal of Monetary Economy\/} {\em 57}, 341--351. \doi{10.1016/j.jmoneco.2010.02.003}
\bibitem[\citeauthoryear{Chevillon and Mavroeidis}{Chevillon and Mavroeidis}{2017}]{chema17}
Chevillon, G., and S. Mavroeidis (2017).
\newblock Learning can generate long memory.
\newblock {\em Journal of Econometrics\/} {\em 198}, 1--7. \doi{10.1016/j.jeconom.2017.01.001}
\bibitem[\citeauthoryear{Christopeit and Massmann}{Christopeit and Massmann}{2018}]{CM18}
Christopeit, N., and M. Massmann (2018).
\newblock Estimating structural parameters in regression models with adaptive learning.
\newblock {\em Econometric Theory\/} {\em 34}, 68--111. \doi{10.1017/S0266466616000529}
\bibitem[\citeauthoryear{Christopeit and Massmann}{Christopeit and Massmann}{2019}]{CM19}
Christopeit, N., and M. Massmann (2019).
\newblock Strong consistency of the least squares estimator in regression models with adaptive learning.
\newblock {\em Electronic Journal of Statistics\/} {\em 13}, 1646--1693. \doi{10.1214/19-EJS1558}
\bibitem[\citeauthoryear{Davidson}{Davidson}{1993}]{dav93}
Davidson, J. (1993).
\newblock The central limit theorem for globally nonstationary near-epoch dependent functions of
mixing processes: the asymptotically degenerate case.
\newblock {\em Econometric Theory\/} {\em 9}, 402--412. \doi{10.1017/S0266466600007738}
\bibitem[\citeauthoryear{Evans and Honkapohja}{Evans and Honkapohja}{2001}]{evho01}
Evans, G. W., and S. Honkapohja (2001).
\newblock {\em Learning and Expectations in Macroeconomics\/}.
\newblock Princeton: Princeton University Press.
\bibitem[\citeauthoryear{Jennrich}{Jennrich}{1969}]{jenn69}
Jennrich, R. I. (1969).
\newblock Asymptotic properties of non-linear least squares estimators.
\newblock {\em The Annals of Mathematical Statistics\/} {\em 40}, 633--643. \doi{aoms/1177697731}
\bibitem[\citeauthoryear{Lai}{Lai}{1994}]{lai94}
Lai, T. L. (1994).
\newblock Asymptotic properties of nonlinear least squares estimates in stochastic regression models.
\newblock {\em The Annals of Statistics\/} {\em 22}, 1917--1930. \doi{10.1214/aos/1176325764}
\bibitem[\citeauthoryear{Lai}{Lai}{2003}]{lai03}
Lai, T. L. (2003).
\newblock Stochastic approximation.
\newblock {\em The Annals of Statistics\/} {\em 31}, 391--406. \doi{10.1214/aos/1051027873}
\bibitem[\citeauthoryear{Malmendier and Nagel}{Malmendier and Nagel}{2016}]{malmnag16}
Malmendier, U., and S. Nagel (2016).
\newblock Learning from inflation experiences.
\newblock {\em The Quarterly Journal of Economics\/} {\em 131}, 53--87. \doi{110.1093/qje/qjv037}
\bibitem[\citeauthoryear{Malmendier et al.}{Malmendier et al.}{2021}]{malmetal21}
Malmendier, U., S. Nagel, and Z. Yan (2021).
\newblock The making of hawks and doves.
\newblock {\em Journal of Monetary Economics\/} {\em 117}, 19--42. \doi{10.1016/j.jmoneco.2020.04.002}
\bibitem[\citeauthoryear{Marcet and Sargent}{Marcet and Sargent}{1989}]{masa89}
Marcet, A., and T. Sargent (1989).
\newblock Convergence of least squares learning mechanisms in self-referential linear stochastic models.
\newblock {\em Journal of Economic Theory\/} {\em 48}, 337--368.
\bibitem[\citeauthoryear{Markiewicz and Pick}{Markiewicz and Pick}{2014}]{mapi14}
Markiewicz, A., and A. Pick (2014).
\newblock Adaptive learning and survey data.
\newblock {\em Journal of Economic Behavior & Organization\/} {\em 107}, 685--707. \doi{10.1016/j.jebo.2014.04.005}
\bibitem[\citeauthoryear{Mayer}{Mayer}{2022}]{may21}
Mayer, A. (2022).
\newblock Estimation and inference in adaptive learning models with slowly decreasing gains.
\newblock {\em Journal of Time Series Analysis\/} {\em 107}, 720--749. \doi{10.1111/jtsa.12636}
\bibitem[\citeauthoryear{Milani}{Milani}{2007}]{mi07}
Milani, F. (2007).
\newblock Expectations, learning and macroeconomic persistence.
\newblock {\em Journal of Monetary Economics\/} {\em 54}, 2065--2082. \doi{10.1016/j.jmoneco.2006.11.007}
\bibitem[\citeauthoryear{Nakov and Nu\ {n}o}{Nakov and Nu\ {n}o}{2015}]{nanu15}
Nakov, A. and G. Nu\ {n}o (2015).
\newblock Learning from experience in the stock market.
\newblock {\em Journal of Economic Dynamics and Control\/}. {\em 52}, 224--239. \doi{10.1016/j.jedc.2014.11.017}
\bibitem[\citeauthoryear{Pagan}{Pagan}{1984}]{pag84}
Pagan, A. (1984).
\newblock Econometric issues in the analysis of regressions with generated regressors.
\newblock {\em International Economic Review\/} {\em 25}, 221--247.
\bibitem[\citeauthoryear{Wang}{Wang}{2021}]{wang21}
Wang, Q. (2021).
\newblock Least squares estimation for nonlinear regression models with heteroskedasticity.
\newblock {\em Econometric Theory\/} {\em 37}, 1267--1289. \doi{10.1017/S0266466620000493}
\newgeometry{margin=1.3in}