EconBase
← Back to paper

The Spectral Approach to Linear Rational Expectations Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

119,972 characters · 15 sections · 150 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

The Spectral Approach to Linear Rational Expectations Models

\newtheorem{lem}{Lemma} \newtheorem{thm}{Theorem} \newtheorem{cor}{Corollary} \newtheorem{prop}{Proposition} \theoremstyle{definition} \newtheorem{exmp}{Example} \newtheorem{defn}{Definition} \newtheorem{alg}{Algorithm}

\abstract{ This paper considers linear rational expectations models in the frequency domain. The paper characterizes existence and uniqueness of solutions to particular as well as generic systems. The set of all solutions to a given system is shown to be a finite dimensional affine space in the frequency domain. It is demonstrated that solutions can be discontinuous with respect to the parameters of the models in the context of non-uniqueness, invalidating mainstream frequentist and Bayesian methods. The ill-posedness of the problem motivates regularized solutions with theoretically guaranteed uniqueness, continuity, and even differentiability properties.

JEL Classification: C10, C32, C62, E32.

Keywords: Linear rational expectations models, frequency domain, spectral representation, Wiener-Hopf factorization, regularization, Gaussian likelihood function.}

Introduction

Spectral (or frequency domain) analysis of covariance stationary processes is a cornerstone of time series analysis. Since its beginnings in the late 1930s, it has benefited from being at the intersection of a number of fundamental mathematical subjects including probability, functional analysis, and complex analysis rozanov,nikolski,bingham2,bingham1. Almost concurrently, economic theory began to focus on expectations of future earnings, prices, interest rates etc.\ as determinants of present economic activity knight,keynes,cagan. This led to the pioneering work of muth who proposed that expectations, rather than being arbitrary inputs into models or arbitrarily determined within models, could be made both endogenous and model-consistent, hence “rational” (see pesaran1987 for further historical context). Today, rational expectations models are the mainstay of business cycle research canova,dd,hs. This paper attempts to bridge the gap between the two strands of literature.

Classical linear systems, such as vector autoregressive moving average (VARMA) models, are linear transformations from an input process to an output process, with present values of the output depending linearly on present and past values of the input as well as past values of the output in every time period. The linear rational expectations model (LREM) class extends classical linear systems by allowing linear dependence on expectations of future values of the output as well. Such models arise naturally from the inter-temporal optimization problems of households and firms in economic modeling. Spectral analysis has focused almost entirely on classical linear systems bd,pourahmadi,lp. Although a number of works have considered LREMs in the frequency domain whiteman,onatski,tanwalker,tan,meyer, none have attempted a general account that parallels the aforementioned textbook treatments. Thus, the first aim of this paper is to situate LREMs in the frequency domain literature at a high level of generality that extends the aforementioned textbook treatments.

Using the spectral representation of covariance stationary processes due to kolmogorov1,kolmogorov2,kolmogorov3 and cramer1,cramer2, this paper recasts the LREM problem in the classical Hilbert space of the frequency domain literature. This is demonstrated concretely on simple scalar models before generalizing to multivariate models. Whereas classical spectral analysis focused entirely on the backward shift operator, the new spectral analysis of LREMs requires the introduction of a new operator associated with expectations. The LREM problem then reduces to a linear system in Hilbert space. However, unlike the special case of VARMA models, LREMs cannot be solved by simply inverting a polynomial matrix due to the presence of the new expectation operators. Solving the LREM problem requires using the method of Wiener-Hopf factorization wh,gf,cg. This paper characterizes existence and uniqueness of solutions to particular as well as generic LREMs, generalizing results by onatski. The set of all solutions to a given LREM is shown to be a finite dimensional affine space in the frequency domain. The dimension of this space is expressed much more simply than in funovits,funovits2020. It is important to note that the underlying assumptions in this paper are weaker than in the previous literature whiteman,onatski,tanwalker,tan,meyer, which requires the exogenous process to have a purely non deterministic wold representation, an assumption that is demonstrably unnecessary. The weaker assumptions of this paper also permit a clear answer as to why unit roots must be excluded, an aspect of the theory absent from the previous literature.

The main results of the paper concern the ill-posedness of the LREM problem in macroeconometrics. hadamard defines a problem to be well-posed if its solutions satisfy the conditions of existence, uniqueness, and continuous dependence on its parameters. The LREM problem is ill-posed because it violates not just the second condition but also the third. Indeed, it has long been accepted that non-uniqueness is a feature of the LREM problem and many techniques have been developed to select a solution for any given LREM exhibiting non-uniqueness (e.g.\ taylor, msv, sunspots, fanelli, farmeretal, bianchinicolo). This paper highlights the fact that selections that have been proposed in the literature are not guaranteed to be continuous with respect to parameters of the LREM in the context of non-uniqueness. This problem seems to be either not well understood or not fully appreciated so far; to the author's knowledge, generic is the only acknowledgement of this problem.

The problem of discontinuity is quite serious because it invalidates mainstream econometric methodology as reviewed, for example, in canova, dd, or hs. This methodology takes a black-box approach to estimation and inference, whereby data is used to construct certain objective functions (likelihood functions, GMM criterion functions, or posterior distributions), which are then either numerically optimized by the Newton-Raphson algorithm or explored by the random walk Metropolis-Hastings algorithm (or the many variations of such algorithms) to produce estimates of parameters as well as confidence or credible intervals. When the discontinuous selected solutions are fed into this methodology, the aforementioned objective functions become discontinuous, invalidating assumptions required for the aforementioned methodologies to work. This paper illustrates concretely how things can go wrong with a simple example in Section (ref).

Fortunately, the literature on ill-posed problems offers an immediate solution to the problem: regularization. The idea here is that when theory is insufficient to pin down a unique solution, other information can be brought to bear. The method can be interpreted in at least two ways: (i) penalizing economically unreasonable solutions or shrinking towards economically reasonable ones or (ii) imposing prior beliefs on the frequencies of fluctuations that solutions ought to exhibit. For example, we may like to avoid solutions where certain variables vary too wildly or impose the prior that solutions to a business cycle model should exhibit fluctuations of period between 4 and 32 quarters in quarterly data. The paper provides conditions for existence and uniqueness of regularized solutions and proves that they are continuously (even differentiably) dependent on their parameters. Thus, under non-uniqueness, regularization selects solutions that can be used in any mainstream econometric method, frequentist or Bayesian.

This work is related to several recent strands in the literature. kn, tq17, kk18, ident, and kk23 study the identification of LREMs based on the spectral density of observables. cv, tq12qe, sala15 utilize spectral domain methods for estimating LREMs using ideas that go back to hs80. recoverability study the problem of subordination (what they call “recoverability”) in the context of macroeconometric models. linsys utilizes a generalization of Wiener-Hopf factorization in order to study unstable and non-stationary solutions ot LREMs. ephermidze provide recent results on the continuity of spectral factorization. Finally, regular provides numerical algorithms for computing regularized solutions.

This paper is organized as follows. Section (ref) sets up the notation and reviews the fundamental concepts of spectral analysis of time series. Section (ref) considers the solution of simple LREMs with elementary frequency domain methods. Section (ref) introduces Wiener-Hopf factorization. Section (ref) sets up the LREM problem, its existence and uniqueness properties, and establishes its ill-posedness. Section (ref) introduces regularized solutions. Section (ref) provides an application of the theory. Section (ref) concludes. Sections (ref)-(ref) comprise the Appendix.

Notation and Review

We will denote by $\mathbb{Z}$, $\mathbb{R}$, and $\mathbb{C}$ the sets of integers, real numbers, and complex numbers respectively. We will denote by $\mathbb{T}=\{z\in\mathbb{C}:|z|=1\}$ the unit circle in $\mathbb{C}$. The letters $i$, $j$, $n$, and $m$ will always stand for natural numbers. By $\mathbb{C}^{m\times n}$ we will denote the set of $m\times n$ matrices of complex numbers. We will use $I_n$ to denote the identity $n\times n$ matrix. For $M\in\mathbb{C}^{m\times m}$, $\mathrm{tr}(M)=\sum_{i=1}^m M_{ii}$. For $M\in\mathbb{C}^{m\times n}$, $M^\ast\in\mathbb{C}^{n\times m}$ is the conjugate transpose of $M$, and $\|M\|_{\mathbb{C}^{m\times n}}=\left(\mathrm{tr}(MM^\ast)\right)^{1/2}=\left(\sum_{i=1}^m\sum_{j=1}^n|M_{ij}|^2\right)^{1/2}$.

All random variables in this paper are defined over a single probability space $(\Omega,\mathscr{F},P)$. For a complex random variable $x$, the expectation is denoted by $\boldsymbol{E}x=\int_\Omega x(\omega)dP(\omega)$. The space $\mathscr{L}_2$ is defined as the Hilbert space of complex valued random variables $x$, modulo $P$-almost sure equality, such that $\boldsymbol{E}|x|^2<\infty$ with inner product and norm

align*[align* omitted — 139 chars of source]

Similar to other Hilbert spaces we will consider in this paper, $\mathscr{L}_2$ is a set of equivalence classes of functions and not a set of functions. However, mathematical convention “relegate[s] this distinction to the status of a tacit understanding” xrudin. Thus, we will write “$x=y$” instead of “$x(\omega)=y(\omega)$ for $P$-almost all $\omega\in\Omega$” and similarly for elements of all of the other Hilbert spaces we consider in this paper. The Hilbert space $\mathscr{L}_2^n$ is defined as the $n$-fold Cartesian product, $\mathscr{L}_2\times\cdots\times \mathscr{L}_2$ consisting of column vectors $x=\left[

smallmatrixx_1\\ \vdots\\ x_n

\right]$, $x_j\in \mathscr{L}_2$, $j=1,\ldots, n$, with the inner product and norm

align*[align* omitted — 177 chars of source]

If $\mathscr{S}\subset\mathscr{L}_2$ is a closed subspace and $x\in\mathscr{L}_2^n$, then the minimum of

align*[align* omitted — 84 chars of source]

with respect to $y\in\mathscr{S}^n=\mathscr{S}\times\cdots\times\mathscr{S}$ is attained by minimizing each term on the right hand side independently with respect to $y_j\in\mathscr{S}$, $j=1,\ldots, n$. It follows that the orthogonal projection of $x\in\mathscr{L}_2^n$ onto $\mathscr{S}^n$ in $\mathscr{L}_2^n$, denoted by $\boldsymbol{P}(x|\mathscr{S}^n)$, is given by $\left[

smallmatrix\boldsymbol{P}(x_1|\mathscr{S})\\ \vdots\\ \boldsymbol{P}(x_n|\mathscr{S})

\right]$, where $\boldsymbol{P}(x_j|\mathscr{S})$ is the orthogonal projection of $x_j$ onto $\mathscr{S}$ in $\mathscr{L}_2$, $j=1,\ldots, n$.

Let $\ldots,\xi_{-1},\xi_0,\xi_1,\ldots$ be a doubly infinite sequence in $\mathscr{L}_2^n$. We will refer to this stochastic process simply as $\xi$. The $j$-th element of $\xi_t$ will be denoted by $\xi_{jt}$. If for all $t,s\in\mathbb{Z}$, $\boldsymbol{E}\xi_t=\boldsymbol{E}\xi_s$ and $\boldsymbol{E}\xi_t\xi_s^\ast$ depends on $t$ and $s$ only through $t-s$, we say that $\xi$ is a covariance stationary process. Given $\xi$, we may define a number of useful objects.

Let $\mathscr{H}$ be the closure in $\mathscr{L}_2$ of the set of all finite complex linear combinations of the set $\{\xi_{jt}: j=1,\ldots,n, t\in\mathbb{Z}\}$. We define $\mathscr{H}^n\subset \mathscr{L}_2^n$ to be the $n$-fold Cartesian product of $\mathscr{H}$.

Traditionally, spectral analysis utilizes the forward (rather than backward) shift operator,

align*[align* omitted — 98 chars of source]

This operator extends uniquely to a unitary operator on $\mathscr{H}$ rozanov.

We may next define the unique spectral measure $F$ on the unit circle, $\mathbb{T}$, which satisfies

align*[align* omitted — 83 chars of source]

rozanov. Many textbooks express the integral above as $\int_{-\pi}^\pi e^{\mathrm{i}(t-s)\lambda}dF(\lambda)$; the notation here is substantially more compact. When each element of $F$ is absolutely continuous with respect to normalized Lebesgue measure on $\mathbb{T}$, defined by

align*[align* omitted — 121 chars of source]

then the Radon-Nikod\'{y}m derivative $dF/d\mu$ is the spectral density matrix of $\xi$. Note that the spectral density of $n$-dimensional standardized white noise is $I_n$. Another important example of a spectral measure, is the Dirac measure at $w\in\mathbb{T}$ defined as

align*[align* omitted — 101 chars of source]

for Borel subsets $\Lambda\subset\mathbb{T}$. Note that $F=\delta_w I_n$ is the spectral measure of the purely deterministic process $\xi_t=\xi_0 w^t$ for $t\in\mathbb{Z}$, with $\boldsymbol{E}\xi_0\xi_0^\ast=I_n$.

We may also define the Hilbert space $H$ of all Borel-measurable mappings $\phi:\mathbb{T}\rightarrow\mathbb{C}^{1\times n}$ (i.e.\ $\phi(z)$ is a row vector for $z\in\mathbb{T}$) such that $\int \phi dF\phi^\ast<\infty$ with inner product and norm

align*[align* omitted — 110 chars of source]

where we identify $\phi=\varphi$ if $\|\phi-\varphi\|_H^2=\int(\phi-\varphi)dF(\phi-\varphi)^\ast=0$ rozanov. We define the Hilbert space $H^m=H\times\cdots\times H$ to be the $m$-fold Cartesian product of $H$ consisting of the $m\times n$ matrices $\phi=\left[

smallmatrix\phi_1\\ \vdots\\ \phi_m

\right]$, $\phi_j\in H$, $j=1,\ldots, m$ endowed with the inner product and norm

align*[align* omitted — 126 chars of source]

By the same argument as used earlier, if $S\subset H$ is a closed subspace, the orthogonal projection of $\phi\in H^m$ onto $S^m$ in $H^m$, denoted by $\boldsymbol{P}(\phi|S^m)$, is given by $\left[

smallmatrix\boldsymbol{P}(\phi_1|S)\\ \vdots\\ \boldsymbol{P}(\phi_m|S)

\right]$, where $\boldsymbol{P}(\phi_j|S)$ is the orthogonal projection of $\phi_j$ onto $S$ in $H$, $j=1,\ldots, m$.

The spectral representation theorem states that

align*[align* omitted — 56 chars of source]

for a column vector of random measures $\Phi$ with $\Phi(\Lambda)\in \mathscr{H}^n$ and $\boldsymbol{E}\Phi(\Lambda)\Phi^\ast(\Lambda)=F(\Lambda)$ for Borel subsets $\Lambda\subset\mathbb{T}$ rozanov. Many textbooks express the integral above equivalently as $\int_{-\pi}^\pi e^{\mathrm{i}t\lambda}dZ(\lambda)$, where $Z(\lambda)=\Phi(\{e^{\mathrm{i}\tau}:\tau\in(-\pi,\lambda]\})$ for $\lambda\in(-\pi,\pi]$; again, the notation here is substantially more compact. The spectral representation theorem establishes a unitary mapping $H^m\rightarrow \mathscr{H}^m$ defined by $\phi\mapsto h=\int \phi d\Phi$ rozanov. We call $\phi$ the spectral characteristic of $h$. Thus,

align*[align* omitted — 219 chars of source]

Denote by $e_j\in\mathbb{C}^{1\times n}$ the row vector whose elements are all equal to zero except for the $j$-th which is equal to one. The spectral representation theorem implies that $\mathscr{H}_t$, the closure in $\mathscr{L}_2$ of the set of all finite complex linear combinations of $\{\xi_{js}: j=1,\ldots,n, s\leq t\}$, is in correspondence with the closure of the set of all finite complex linear combinations of $\{z^se_j: j=1,\ldots,n, s\leq t\}$ in $H$, denoted $H_t$. Let $\mathscr{H}^m_t$ and $H_t^m$ be $m$-fold Cartesian products of $\mathscr{H}_t$ and $H_t$ respectively. This implies that

align*[align* omitted — 137 chars of source]

Thus, the best linear prediction of $\xi_{t+s}$ in terms of $\xi_t, \xi_{t-1},\ldots$ is given by

align*[align* omitted — 123 chars of source]

It is easily established that $\mathscr{H}_t=\boldsymbol{U}^t\mathscr{H}_0$ and $H_t=z^tH_0$ for all $t\in\mathbb{Z}$ so that $\boldsymbol{P}(\boldsymbol{U}^t\xi|\mathscr{H}_t)=\boldsymbol{U}^t\boldsymbol{P}(h|\mathscr{H}_0)$ for all $t\in\mathbb{Z}$ and $h\in \mathscr{H}$ and likewise $\boldsymbol{P}(z^t\phi|H_t)=z^t\boldsymbol{P}(\phi|H_0)$ for all $t\in\mathbb{Z}$ and $\phi\in H$ rozanov. Thus

align*[align* omitted — 134 chars of source]

If $\zeta$ is an additional covariance stationary processes such that for every $t,s\in\mathbb{Z}$, $\boldsymbol{E}\zeta_s\xi_t^\ast$ depends on $t$ and $s$ only through $t-s$, we will say that $\zeta$ is causal in $\xi$ if $\zeta_0\in \mathscr{H}_0^m$. Finally, if $\nu$ is causal in $\xi$ and satisfies $\boldsymbol{P}(\nu_{t+1}|\mathscr{H}_t)=0$ for all $t\in\mathbb{Z}$, we call $\nu$ an innovation process.

Examples

Armed with the basic machinery above, we now make a first attempt at solving LREMs in the frequency domain. We will see that solving simple univariate LREMs involves only elementary spectral domain techniques as discussed in textbook treatments of spectral analysis such as bd or pourahmadi. The methods also provide strong hints to the general approach to solving LREMs. In this section, $\xi$ is a scalar covariance stationary process with $\boldsymbol{U}$, $\Phi$, $F$, $\mathscr{H}$, and $H$ defined as in the previous section.

The Autoregressive Model

We begin on familiar territory with the stationary autoregression

align[align omitted — 74 chars of source]

where $|\alpha|<1$. The frequency-domain analysis of this model is available in many textbooks. For completeness, we provide a treatment here that is geared towards understanding the more general cases to come.

We require a covariance stationary solution causal in $\xi$ (further motivation of this restriction can be found in Section (ref)). Thus, we require

align*[align* omitted — 59 chars of source]

for some spectral characteristic $\phi\in H_0$.

Notice that we may restrict attention to the equation

align*[align* omitted — 54 chars of source]

If a solution $X_0$ exists, then the rest of the process can be generated as

align*[align* omitted — 60 chars of source]

and clearly satisfies (ref).

Thus, we must solve for $\phi$ in

align*[align* omitted — 58 chars of source]

Since integration with respect to $\Phi$ is a unitary mapping $H\rightarrow\mathscr{H}$, it has an inverse and the equation above is equivalent to

align[align omitted — 54 chars of source]

in the frequency domain.

The linear mapping $\phi\mapsto \alpha z^{-1} \phi$ on $H_0$ is bounded in norm by $|\alpha|<1$, thus the mapping $\phi\mapsto (1-\alpha z^{-1})\phi$ is invertible and

align*[align* omitted — 52 chars of source]

is the unique solution to (ref), where the summation is understood to converge with respect to $H$ norm basic. Thus, we arrive at the unique solution to $X_0$,

align*[align* omitted — 144 chars of source]

The interchange of the summation and the stochastic integral is admissible because the stochastic integral is a bounded linear operator mapping from $H$ to $\mathscr{H}$ and the inner summation converges in $H$. It follows that the unique covariance stationary solution is

align*[align* omitted — 150 chars of source]

The operator $\boldsymbol{U}^t$ is interchangeable with the summation because it is a bounded linear operator on $\mathscr{H}$ and the summation converges in $\mathscr{L}_2$.

The Cagan Model

The Cagan model is given as,

align[align omitted — 105 chars of source]

with $|\beta|<1$. Again, we look for a covariance stationary solution causal in $\xi$ and we restrict attention to the equation

align*[align* omitted — 65 chars of source]

and, following a similar argument to that used above, we arrive at the underlying frequency domain problem,

align[align omitted — 68 chars of source]

Since the linear mapping $\phi\mapsto \beta \boldsymbol{P}(z\phi|H_0)$ on $H_0$ is bounded in norm by $|\beta|<1$, we see that the mapping $\phi\mapsto \phi-\beta \boldsymbol{P}(z\phi|H_0)$ is invertible so that

align*[align* omitted — 68 chars of source]

is the unique solution to (ref), where the summation is understood to converge with respect to $H$ norm basic. Thus, we arrive at the unique solution to $X_0$,

align*[align* omitted — 201 chars of source]

The interchange of the summation and the stochastic integral is admissible because the stochastic integral is a bounded linear operator mapping $H$ to $\mathscr{H}$ and the inner summation converges in $H$. It follows that the unique covariance stationary solution is

align*[align* omitted — 205 chars of source]

The operator $\boldsymbol{U}^t$ is interchangeable with the summation because it is a bounded linear operator on $\mathscr{H}$ and the summation converges in $\mathscr{L}_2$.

The Mixed Model

Now suppose we have the more general model

align[align omitted — 117 chars of source]

where $a,b,c\in\mathbb{C}$ and $ac\neq0$. This leads to the frequency-domain equation

align[align omitted — 79 chars of source]

As noted by sargent79, the solution to this system depends on the factorization of

align*[align* omitted — 61 chars of source]

We assume, without loss of generality, that $|\gamma|\leq|\delta|$. There are four cases to consider.

Suppose $|\gamma|<1<|\delta|$. We can then write $M(z)=a(z-\delta)(1-\gamma z^{-1})$ and express the system as

align*[align* omitted — 105 chars of source]

The first equation can be solved as in the Cagan model (since $|\delta^{-1}|<1$), while the second can be solved as in the autoregressive model (since $|\gamma|<1$). This procedure leads us to the unique solution

align*[align* omitted — 228 chars of source]

In the time domain, we obtain the following solution

align*[align* omitted — 135 chars of source]

which then gives us the general solution,

align*[align* omitted — 160 chars of source]

Thus, when $|\gamma|<1<|\delta|$, there exists a unique solution.

Next, suppose that $|\delta|<1$. Then we may write $M(z)=az(1-\delta z^{-1})(1-\gamma z^{-1})$ and express our system as

align*[align* omitted — 100 chars of source]

The first equation does not have a unique solution in general. For example, when $F = \mu$, then $\varphi=z^{-1}+\psi$ solves the equation for any $\psi\in\mathbb{C}$. More generally, every solution is of the form

align*[align* omitted — 34 chars of source]

where $\psi\in H_0$ with $z\psi$ orthogonal to $H_0$. We can then solve for $\phi$ as

align*[align* omitted — 102 chars of source]

In the time domain, this leads to the general solution

align*[align* omitted — 130 chars of source]

where $\nu_t=\int z^t\psi d\Phi\in \mathscr{H}_t$ for $t\in\mathbb{Z}$ is an arbitrary innovation process. Indeed, $\nu_t \in \mathscr{H}_t$ for all $t \in \mathbb{Z}$ because $\psi \in H_0$ and $\nu_{t+1}$ is orthogonal to $\mathscr{H}_t$ for all $t \in \mathbb{Z}$ because $\boldsymbol{P} (\nu_{t+1}|\mathscr{H}_t) = \int\boldsymbol{P} (z^{t+1}\psi|H_t)d\Phi = \int z^t\boldsymbol{P} (z\psi|H_0)d\Phi = 0$ as $z\psi$ is orthogonal to $H_0$. Thus, when $|\delta|<1$, there exist potentially infinitely many solutions.

Now suppose $|\gamma|>1$. Then we may write $M(z)=a\delta\gamma z^{-1}(1-\delta^{-1}z)(1-\gamma^{-1}z)$ and

align*[align* omitted — 115 chars of source]

Clearly $\varphi$ can be solved as in the Cagan model to produce

align*[align* omitted — 130 chars of source]

This implies that

align*[align* omitted — 133 chars of source]

However, this equation cannot hold generally. To see this, let $\xi$ be a standard white noise so that $F=\mu$. Then $\{z^s: s \in \mathbb{Z}\}$ is an orthonormal set and the equation above reduces to

align*[align* omitted — 49 chars of source]

Since $\phi$ is in the closure of the linear span of $\{z^s:s\leq0\}$, the zero-th Fourier coefficient of the left hand side is equal to zero, while that of the right hand side is equal to $\frac{1}{a\delta\gamma}\neq0$. So there is no solution in general when $|\gamma|>1$.

Finally, suppose either $|\gamma|=1$ or $|\delta|=1$ so that $M(w)=0$ for some $w\in\mathbb{T}$. Then there is no solution in general, in the sense that there exist processes $\xi$ for which no covariance stationary solution $X$ can be found. To see this, let $F=\delta_w$, the Dirac measure at $w$. If a solution $\phi\in H_0$ to (ref) exists, then it must satisfy

align*[align* omitted — 66 chars of source]

This then implies that $M\phi=M(w)\phi(w)=0$ in $H$, which implies that $\boldsymbol{P}(M\phi|H_0)=0$, contradicting the fact that $\boldsymbol{P}(M\phi|H_0)=1$. Thus, when $|\gamma|=1$ or $|\delta|=1$ there exists no solution in general.

We will see that the analysis of the simple models above contains most of the elements necessary for the general multivariate case with more than one lead and/or lag and arbitrary covariance stationary $\xi$. The reader interested in a more complete elementary analysis of the simple models above (e.g.\ considering the case $|\beta|>1$ of the Cagan model) is directed to previous versions of this paper available on arXiv.

Wiener-Hopf Factorization

The approach we have taken in the last section is well understood in the theory of convolution equations gf. The requisite factorization of $M(z)$ into two parts, a part to solve like the Cagan model and a part to solve like the autoregressive model, is known as a Wiener-Hopf factorization wh. In this section, we state the basic concepts and properties of Wiener-Hopf factorization that we will need.

defnLet $\mathcal{W}$ be the class of functions $M:\mathbb{T}\rightarrow\mathbb{C}$ defined by \begin{align} M(z)=\sum_{s=-\infty}^\infty M_sz^s, \end{align} where $M_s\in\mathbb{C}$ for $s\in\mathbb{Z}$ and $\sum_{s=-\infty}^\infty |M_s|<\infty$. Define $\mathcal{W}_\pm\subset\mathcal{W}$ to be the class of functions (ref) with $M_s=0$ for $s\lessgtr0$. The sets $\mathcal{W}^{m\times n}$, $\mathcal{W}^{m\times n}_\pm$ are defined as the sets of matrices of size $m\times n$ populated by elements of $\mathcal{W}$ and $\mathcal{W}_\pm$ respectively.

The class of functions $\mathcal{W}$ is known as the Wiener algebra in the functional analysis literature classes2. Note that every function analytic in a neighborhood of $\mathbb{T}$ defines an element of $\mathcal{W}$ but the opposite inclusion does not hold (e.g.\ $\sum_{s=0}^\infty s^{-2}z^s\in\mathcal{W}_+$ diverges outside $\mathbb{T}$). For $M\in\mathcal{W}^{m\times n}$, we also have that

align*[align* omitted — 94 chars of source]

where $\|M\|_\infty$ is the $\mu$--essential supremum of $\|M(z)\|_{\mathbb{C}^{m\times n}}$. We will also need the operator

align*[align* omitted — 85 chars of source]

defined over $\mathcal{W}^{m\times n}$ and, more generally, over the class of square summable series satisfying $\sum_{s=-\infty}^\infty \|M_s\|_{\mathbb{C}^{n\times m}}^2<\infty$ classes2.

defnA Wiener-Hopf factorization of $M\in\mathcal{W}^{m\times m}$ is a factorization \begin{align} M=M_+M_0M_-, \end{align} where $M_+\in\mathcal{W}^{m\times m}_+$, $M_+^{-1}\in\mathcal{W}^{m\times m}_+$, $M_-\in\mathcal{W}^{m\times m}_-$, $M_-^{-1}\in\mathcal{W}^{m\times m}_-$, $M_0$ is a diagonal matrix with diagonal elements $z^{\kappa_1},\ldots,z^{\kappa_m}$, and $\kappa_1\geq\cdots\geq\kappa_m$ are integers, called partial indices.

We remark that the factorization given in Definition (ref) is termed a left factorization in the Wiener-Hopf factorization literature gf. It differs from the right factorization utilized in onatski and linsys, where the roles of $M_\pm(z)$ are reversed. The difference is due to the fact that the present analysis works with the forward shift operator, whereas onatski and linsys work with the backward shift operator. A left factorization of $M$ is obtained from a right factorization of $M^\ast$.

thmLet $M\in\mathcal{W}^{m\times m}$ and suppose $\det(M(z))\neq0$ for all $z\in\mathbb{T}$. Then $M$ has a Wiener-Hopf factorization and its partial indices are unique.
proofSee Theorems VIII.1.1 and VIII.2.2 of gf.

The condition in Theorem (ref) is minimal for existence of a Wiener-Hopf factorization and uniqueness of the $M_0$ part. The $M_\pm$ parts are not unique but their general form is well understood gf. The non-uniqueness of $M_\pm$ has no bearing on any of our results. Wiener-Hopf factorizations can be computed in a variety of ways (see constructive for a recent survey).

The partial indices allow us to identify an important subset of $\mathcal{W}^{m\times m}$.

defnLet $\mathcal{W}^{m\times m}_\circ$ be the subset of $M\in\mathcal{W}^{m\times m}$ such that $\det(M(z))\neq0$ for all $z\in\mathbb{T}$ and $\kappa_1-\kappa_m\leq1$.

When $m=1$, $0=\kappa_1-\kappa_m\leq1$ so that $\mathcal{W}^{1\times 1}_\circ$ is the set of $M\in\mathcal{W}$ such that $M(z)\neq0$ for all $z\in\mathbb{T}$. Notice that for every element of $\mathcal{W}^{m\times m}_\circ$, the partial indices are either all non-negative, all non-positive, or all zero. This implies that for every element of $\mathcal{W}^{m\times m}_\circ$,

align*[align* omitted — 105 chars of source]

where $\mathrm{sign}$ is equal to 1, $-1$, or 0 according to whether the argument is positive, negative, or zero respectively. Since $\sum_{i=1}^m\kappa_i$ is the winding number of $\det(M(z))$ as $z$ traverses the unit circle counter-clockwise gf, we can easily determine the sign of the partial indices of elements of $\mathcal{W}^{m\times m}_\circ$ by looking at the winding number. The importance of this fact will become clear in the next section when we combine it with the following fact.

thmIf $\mathcal{W}^{m\times m}$ is endowed with the $\mu$--essential supremum norm then $\mathcal{W}^{m\times m}_\circ$ is open and dense in $\{M\in\mathcal{W}^{m\times m}: \det(M(z))\neq0\text{ for all }z\in\mathbb{T}\}$.
proofSee Theorems 1.20 and 1.21 of gks.

Theorem (ref) implies that the generic or typical element of $\mathcal{W}^{m\times m}$ that admits a Wiener-Hopf factorization is an element of $\mathcal{W}_\circ^{m\times m}$. Said differently, the elements of $\mathcal{W}^{m\times m}\backslash\mathcal{W}_\circ^{m\times m}$ are non-generic or exceptional in the space of Wiener-Hopf factorizable elements of $\mathcal{W}^{m\times m}$.

To see the role played by Wiener-Hopf factorization, consider system (ref) again. In the first case, $M$ factorized as

align*[align* omitted — 339 chars of source]

In the fourth case, $M(w)=0$ for some $w\in\mathbb{T}$ and there does not exist a solution in general. Notice that the case of existence and uniqueness is associated with a partial index of zero, the case of existence and non-uniqueness is associated with a positive partial index, and the cases of non-existence in general are associated with a negative partial index and/or a zero of $M(z)$. These associations are not accidental as we will see in the next section.

Existence and Uniqueness

Given the Wiener-Hopf factorization techniques, we can now address the general LREM problem in the frequency domain. Our first task is to define existence and uniqueness of solutions to the LREM problem in the time domain.

defnAn LREM is a pair $(M,N)\in\mathcal{W}^{m\times m}\times\mathcal{W}^{m\times n}$, expressed formally as \begin{align} \sum_{s=-\infty}^\infty M_sE_tX_{t+s}=\sum_{s=-\infty}^\infty N_sE_t\xi_{t+s},\qquad t\in\mathbb{Z} \end{align} where $\xi$ is exogenous and $X$ is endogenous.

Expression (ref) is understood as a formal set of structural equations whose mathematical meaning we will postpone to Definition (ref) below. The structural equations relate current, expected, and lagged values of the endogenous process $X$ to current, expected, and lagged values of an exogenous process $\xi$. In particular, $E_tX_{t+s}$ is understood to mean the expectation of $X_{t+s}$ at time $t$ if $s>0$ and $X_{t-|s|}$ if $s\leq0$; $E_t\xi_{t+s}$ is understood similarly. The notion of expectation pertinent to us will be made clear in Definition (ref). A textbook example is the New Keynesian LREM,

align[align omitted — 206 chars of source]

where $X_1$, $X_2$, and $X_3$ consists of inflation, output, and the interest rate respectively; $\xi$ is interpreted as a policy shock; and $\theta=(\theta_1,\ldots,\theta_6)$ is a parameter. The first equation of (ref) is the New Keynesian Phillips curve, relating current inflation to one-period-ahead expected inflation and current output, and results from the inter-temporal optimization of firms; the second equation relates current output to one-period-ahead expected output and the interest rate and results from the inter-temporal optimization of households; finally, the third equation is the policy rule by which the monetary authority sets the interest rate (see e.g. galibook). LREM (ref) is represented by the pair,

align*[align* omitted — 241 chars of source]

Occasionally, researchers also consider solutions driven by sunspots, economic forces that do not appear explicitly in the system (their associated columns of $N_s$ are equal to zero) but are nevertheless considered to influence the behavior of the solution; we discuss these solutions further in Section (ref).

The class of models considered in Definition (ref) includes the class of structural equation models ($M$ and $N$ are constant), the class of VARMA models ($M$ and $N$ are matrix polynomials in $z^{-1}$), and all of the LREMs considered in canova, dd, and hs. For example, the model of sw, studied in Section 6.2 of hs, consists of $m=14$ endogenous variables and $n=7$ exogenous shocks.

Now in order to endow (ref) with mathematical meaning, we must first note that LREMs take as inputs not only realizations of exogenous inputs $\xi$ but also a spectral measure $F$.

center[center omitted — 365 chars of source]

That is because the output $X$ of an LREM solves a system of equations in past, present, as well as expected values of $X$. In the time domain perspective on LREMs, the role of $F$ is played by a filtration with respect to which conditional expectations can be computed linsys. Here, conditional expectations are substituted by linear projections, the more convenient analogue to expectations in the frequency domain and, in order to compute projections, $F$ needs to be specified as well. We will see that the transfer function of solutions requires the triple $(M,N,F)$, while the realizations of the output require, additionally, the realizations of the inputs. In following this system-theoretic approach to LREMs, therefore, we will often use phrases such as “for every spectral measure $F$” or “for every covariance stationary process $\xi$”.

defnLet $\xi$ be a zero-mean, $n$-dimensional, covariance stationary process with spectral measure $F$ and let $(M,N)$ be an LREM. A solution to $(M,N)$ (or equation (ref)) is an $m$-dimensional covariance stationary process $X$, causal in $\xi$, and satisfying \begin{align} \sum_{s=-\infty}^\infty M_s\boldsymbol{P}(X_{t+s}|\mathscr{H}_t^m)=\sum_{s=-\infty}^\infty N_s\boldsymbol{P}(\xi_{t+s}|\mathscr{H}_t^n),\qquad t\in\mathbb{Z}, \end{align} where the series converge in $\mathscr{H}^m$. We say that $(M,N)$ has no solution in general if it is possible to find a $\xi$ such that no solution to $(M,N)$ exists. A solution $X$ is unique if whenever $Y$ is also a solution, then $X_t=Y_t$ for all $t\in\mathbb{Z}$.

By Definition (ref), if a unique solution (or at least a unique selection from among all possible solutions) exists for every covariance stationary $\xi$, then the LREM is a mathematical system that transforms arbitrary covariance stationary inputs into covariance stationary outputs kalmanetal,kailath,hd,sontag,caines. If no solution exists for a given $\xi$, then the LREM is said to have no solution in general (it will often be convenient to consider purely deterministic $\xi$ to show that no solution exists in general). The restriction to covariance stationarity involves no loss of generality as $X$ and $\xi$ describe deviations away from steady state in modern LREMs (see e.g. galibook; non-stationary LREMs are studied in linsys).

Definition (ref) also restricts attention to causal solutions. This is because the purpose of an LREM is to explain the behavior of economic variables in terms of past and present economic shocks and to obtain impulse responses. It makes little sense to consider models where current inflation is determined by a future shock to technology, for example. Note, however, that we do not impose invertibility of $\xi$ in terms of $X$ as it is not necessary for our purposes, although it does become necessary for estimation purposes (see ident).

As in the previous section, it suffices to solve the $t=0$ equation of (ref),

align*[align* omitted — 142 chars of source]

For if $X$ is a covariance stationary process causal in $\xi$ and satisfies this system, applying the forward shift operator $t$ times to each equation we obtain (ref). Of course, the forward shift operator commutes with the summation because the sum converges in $\mathscr{H}^m$. Let $X_0$ have the spectral characteristics $\phi\in H_0^m$. Then the frequency domain equivalent is

align[align omitted — 145 chars of source]

Classical frequency domain theory is built on the backwards shift operator $\phi\mapsto z^{-1}\phi$. In order to solve for $\phi$ in (ref), we will need an additional, closely related, operator.

defnDefine $\boldsymbol{V},\boldsymbol{V}^{(-1)}:H_0\rightarrow H_0$ to be the operators \begin{align*} \boldsymbol{V}:\phi\mapsto\boldsymbol{P}(z\phi|H_0),& & \boldsymbol{V}^{(-1)}:\phi\mapsto z^{-1}\phi. \end{align*} If $\kappa\geq1$ (resp.\ $\kappa\leq-1$), we denote by $\boldsymbol{V}^\kappa$ the operator $\boldsymbol{V}$ (resp.\ $\boldsymbol{V}^{(-1)}$) composed with itself $|\kappa|$ times and we define $\boldsymbol{V}^0=\boldsymbol{I}$, the identity mapping.

The following lemma (proof omitted) lists the most important properties of $\boldsymbol{V}$ and $\boldsymbol{V}^{(-1)}$.

lemThe operators $\boldsymbol{V}$ and $\boldsymbol{V}^{(-1)}$ have the following properties: \begin{enumerate} • $\boldsymbol{V}^\ast=\boldsymbol{V}^{(-1)}$. • $(\boldsymbol{V}\phi,\boldsymbol{V}\phi)\leq(\phi,\phi)$ for all $\phi\in H_0$. • $\left(\boldsymbol{V}^{(-1)}\phi,\boldsymbol{V}^{(-1)}\varphi\right)=(\phi,\varphi)$ for all $\phi,\varphi\in H_0$. • $\boldsymbol{V}^{(-1)}$ is a right inverse of $\boldsymbol{V}$. • $\ker(\boldsymbol{V})=H_0\ominus z^{-1}H_0$. • $\dim\ker(\boldsymbol{V})\leq n$. • For all $\kappa\in\mathbb{Z}$ and $\phi\in H_0$, $\boldsymbol{V}^\kappa\phi=\boldsymbol{P}(z^\kappa\phi|H_0)$. • For $\kappa>1$, $\ker(\boldsymbol{V}^\kappa)=\ker(\boldsymbol{V})\oplus \boldsymbol{V}^{(-1)}\ker(\boldsymbol{V})\oplus\cdots\oplus \boldsymbol{V}^{1-\kappa}\ker(\boldsymbol{V})$. \end{enumerate}

Lemma (ref) (i) states that the backwards shift operator $\boldsymbol{V}^{(-1)}$ is the adjoint of $\boldsymbol{V}$. Lemma (ref) (ii) implies that the operator norm of $\boldsymbol{V}$ is bounded above by 1. Lemma (ref) (iii) implies that $\boldsymbol{V}^{(-1)}$ is an isometry basic. This implies that the operator norm of $\boldsymbol{V}^{(-1)}$ is equal to 1. Lemma (ref) (iv) establishes that $\boldsymbol{V}$ is right-invertible. Lemma (ref) (v) clarifies the obstruction to left-invertibility of $\boldsymbol{V}$ as $\ker(\boldsymbol{V})$ may be non-trivial. For example, when $F=\mu$ and $\phi\in\mathbb{C}$, then $\boldsymbol{V}^{(-1)}\boldsymbol{V}(\phi)=\boldsymbol{P}(\,\phi\,|z^{-1},z^{-2},\ldots\,)=0$. Since $H_0\ominus z^{-1}H_0$ is the set of spectral characteristics associated with innovations to $\xi$, $\ker(\boldsymbol{V})=\{0\}$ if and only if $\xi$ is purely deterministic. Lemma (ref) (vi) expresses the intuitive fact that the dimension of the innovation space of a covariance stationary process is bounded above by the dimension of the process. Lemma (ref) (vii) is a convenient expression for iterates of the $\boldsymbol{V}$ operator. Finally, Lemma (ref) (viii) decomposes the kernel of $\boldsymbol{V}^\kappa$ into a direct sum generated by the kernel of $\boldsymbol{V}$. It follows, since $\boldsymbol{V}^{(-1)}$ is an isometry, that $\dim\ker(\boldsymbol{V}^\kappa)=\kappa\dim\ker(\boldsymbol{V})$. The time-domain analogue of the decomposition in Lemma (ref) (viii) is the familiar one from brozeetal85,brozeetal95, where a process $\nu$ causal in $\xi$ satisfies

align*[align* omitted — 85 chars of source]

if and only if

align*[align* omitted — 173 chars of source]

That is, if and only if $\nu_{t+\kappa}$ is representable as the sum of the prediction revisions between $t+1$ and $t+\kappa$ for all $t\in\mathbb{Z}$.

The fact that the operators $\boldsymbol{V}^s$ are uniformly bounded in the operator norm by 1 (Lemma (ref) (i) and (ii)) ensures that $\sum_{s=-\infty}^\infty M_s\boldsymbol{V}^s$ is a bounded linear operator on $H_0$ whenever $\sum_{s=-\infty}^\infty M_sz^s\in\mathcal{W}$ classes1. More generally, we have the following definition, adopted from gf.

defnFor $M\in\mathcal{W}^{m\times n}$ with $ij$-th element $M_{ij}$, define $\boldsymbol{M}_{ij}:H_0\rightarrow H_0$ as \begin{align*} \boldsymbol{M}_{ij}&=\sum_{s=-\infty}^\infty M_{sij}\boldsymbol{V}^s,\qquad i=1,\ldots, m,\qquad j=1,\ldots, n, \end{align*} where the series converges in the operator norm, and $\boldsymbol{M}:H_0^n\rightarrow H_0^m$ as \begin{align*} \boldsymbol{M}\phi&=\left[\begin{array}{ccc} \boldsymbol{M}_{11}& \cdots & \boldsymbol{M}_{1n}\\ \vdots & & \vdots\\ \boldsymbol{M}_{m1}& \cdots & \boldsymbol{M}_{mn} \end{array}\right]\left[\begin{array}{c} \phi_1\\ \vdots\\ \phi_n \end{array}\right]=\left[\begin{array}{c} \sum_{j=1}^n \boldsymbol{P}(M_{1j}\phi_j|H_0)\\ \vdots\\ \sum_{j=1}^n \boldsymbol{P}(M_{mj}\phi_j|H_0) \end{array}\right]=\boldsymbol{P}(M\phi|H_0^m). \end{align*}

For $M\in\mathcal{W}^{m\times n}$, $\|\boldsymbol{M}\phi\|_{H^m}=\|\boldsymbol{P}(M\phi|H_0^m)\|_{H^m}\leq\|M\phi\|_{H^m}\leq\|M\|_\infty\|\phi\|_{H^n}$, thus the operator norm of $\boldsymbol{M}$ is bounded by $\|M\|_\infty$. Note that by Lemma (ref) (i), the adjoint of $\boldsymbol{M}$ is

align*[align* omitted — 84 chars of source]

where $M^\ast\in\mathcal{W}^{n\times m}$ is given by $(M(z))^\ast$ for $z\in\mathbb{T}$.

Definition (ref) allows us to express (ref) more compactly as,

align[align omitted — 76 chars of source]

Recall that the spectral characteristic of $\xi$ is $I_n$. Thus, we have arrived at a linear equation in the Hilbert space $H_0^m$. Equations (ref), (ref), and (ref) are clearly special cases of (ref). It is also instructive to consider the special case where $(M,N)$ is a VARMA model, so $\boldsymbol{M}\phi=\boldsymbol{P}(M\phi|H_0^m)=M\phi$ and $\boldsymbol{N}I_n=\boldsymbol{P}(N|H_0^m)=N$, and the system reduces to $M\phi=N$, a problem that is very well understood hd.

A more delicate analysis is required for the general case. Luckily Section (ref) hints towards a solution. First, we obtain a Wiener-Hopf factorization,

align*[align* omitted — 26 chars of source]

By Theorem (ref), this factorization exists if $\det(M(z))\neq0$ for all $z\in\mathbb{T}$. Then $\boldsymbol{M}_+$, $\boldsymbol{M}_0$, and $\boldsymbol{M}_-$ can be defined as in Definition (ref) and it is easily checked that

align*[align* omitted — 78 chars of source]

a fact that at first seems trivial until one recalls that $\boldsymbol{V}^i\boldsymbol{V}^j$ is not generally equal to $\boldsymbol{V}^{i+j}$ when $i<0<j$ (Lemma (ref) (iv) and (v)). Then (ref) can be broken up into the system of equations,

align[align omitted — 141 chars of source]

Solving each of these systems is completely understood in the Wiener-Hopf factorization literature gf,cg. Indeed, the first system of equations involves only the $\boldsymbol{V}$ operator and is solved as in the Cagan model, while the third system of equations involves only the $\boldsymbol{V}^{(-1)}$ operator and is solved as in the autoregressive model. To see this, note that since $M_+^{-1},M_-^{-1}\in\mathcal{W}^{m\times m}$, the operators

align*[align* omitted — 151 chars of source]

are bounded and linear on $H_0^m$. By Lemma (ref) (vii), since $M_+, M_+^{-1}\in\mathcal{W}_+^{m\times m}$, we have that for all $\phi\in H_0^m$,

align*[align* omitted — 362 chars of source]

On the other hand, since $M_-,M_-^{-1}\in\mathcal{W}_-^{m\times m}$, $\boldsymbol{M}_-:\phi\mapsto M_-\phi$ and $\boldsymbol{M}_-^{-1}:\phi\mapsto M_-^{-1}\phi$ for every $\phi\in H_0^m$ (i.e.\ they are multiplication operators). It follows that

align*[align* omitted — 144 chars of source]

We have established that

align*[align* omitted — 182 chars of source]

where $\boldsymbol{I}$ is the identity mapping on $H_0^m$. The first and third equations of (ref) are therefore uniquely solvable as

align*[align* omitted — 95 chars of source]

It remains to solve for $\varphi$ in (ref). Since $\boldsymbol{M}_0$ is a diagonal operator matrix with $i$-th diagonal entry equal to $\boldsymbol{V}^{\kappa_i}$, existence and uniqueness of solutions to (ref) depend on the partial indices of $M$. Notice that if all of the partial indices of $M$ are equal to zero, then $\boldsymbol{M}_0=\boldsymbol{I}$, and (ref) has the unique solution $\phi=\boldsymbol{M}_-^{-1}\boldsymbol{M}_+^{-1}\boldsymbol{N}I_n$. If all the partial indices of $M$ are non-negative, the operators $\boldsymbol{V}^{\kappa_i}$ above are right invertible, which implies that $\boldsymbol{M}_0$ is right invertible; in that case, a right inverse of $\boldsymbol{M}_0$ is given as

align[align omitted — 278 chars of source]

If $\kappa_1>0$ and $\kappa_m\geq0$, there are generally infinitely many other right inverses of $\boldsymbol{M}_0$ and therefore infinitely many right inverses of $\boldsymbol{M}$; indeed, every right inverse of $\boldsymbol{M}$ is of the form $\boldsymbol{M}_-^{-1}\boldsymbol{M}_0^{(-1)}\boldsymbol{M}_+^{-1}$ for some right inverse $\boldsymbol{M}_0^{(-1)}$ of $\boldsymbol{M}_0$.

The theoretical foundations for existence and uniqueness are now complete and all that remains is to apply well-known cookie-cutter results from functional analysis along with the simple techniques we employed in Section (ref).

lemIf $(M,N)$ is an LREM, $\det(M(z))\neq0$ for all $z\in\mathbb{T}$, and $M$ has a Wiener-Hopf factorization $M_+M_0M_-$, then for every spectral measure $F$, $\boldsymbol{M}$ is one-to-one if the partial indices of $M$ are non-positive and $\boldsymbol{M}$ is onto if the partial indices of $M$ are non-negative. In the latter case, the general form of solutions to (ref) is \begin{align} \phi=\boldsymbol{M}^{(-1)}\boldsymbol{N}I_n+\boldsymbol{M}_-^{-1}\psi, \end{align} where \begin{align} \boldsymbol{M}^{(-1)}=\boldsymbol{M}_-^{-1}\boldsymbol{M}_0^{(-1)}\boldsymbol{M}_+^{-1}, \end{align} $\boldsymbol{M}_0^{(-1)}$ is a right inverse of $\boldsymbol{M}_0$, $\psi\in \ker(\boldsymbol{M}_0)$, and $\dim(\ker(\boldsymbol{M}))=\dim(\ker(\boldsymbol{V}))\sum_{i=1}^m\kappa_i$.
proofProposition $2^\circ$ of Section VIII.4 of gf applied to $M^\ast$ together with Lemma (ref) imply that \begin{align*} \dim(\mathrm{coker}(\boldsymbol{M}))=-\dim(\ker(\boldsymbol{V}))\sum_{\kappa_i<0}\kappa_i,& &\dim(\ker(\boldsymbol{M}))=\dim(\ker(\boldsymbol{V}))\sum_{\kappa_i>0}\kappa_i. \end{align*} Thus, $\boldsymbol{M}$ is one-to-one if $\kappa_1\leq 0$ and onto if $\kappa_m\geq0$. In the latter case, (ref) is clearly a solution to (ref). On the other hand, if $\phi$ is a solution to (ref), define \begin{align*} \psi=\boldsymbol{M}_-\phi-\boldsymbol{M}_0^{(-1)}\boldsymbol{M}_+^{-1}\boldsymbol{N}I_n. \end{align*} Clearly $\psi\in H_0^m$. Finally, \begin{equation*} \boldsymbol{M}_0\psi=\boldsymbol{M}_0\boldsymbol{M}_-\phi-\boldsymbol{M}_+^{-1}\boldsymbol{N}I_n=\boldsymbol{M}_+^{-1}(\boldsymbol{M}\phi-\boldsymbol{N}I_n)=0.\qedhere \end{equation*}

Basic linear algebra implies that there exists a solution to (ref) if and only if $\boldsymbol{N}I_n$ is in the image of $\boldsymbol{M}$ and a solution is unique if and only if $\ker(\boldsymbol{M})=\{0\}$. Thus, Lemma (ref) provides a sufficient condition for existence and necessary and sufficient conditions for uniqueness of solutions to (ref) irrespective of the exogenous inputs. Recall that we insist on conditions invariant with respect to $F$ in keeping with our system-theoretic approach to LREMs. Of course, restricting attention to a particular $F$, we can say slightly more. For example, if the partial indices are non-negative and $F$ has the property that $\ker(\boldsymbol{V})=\{0\}$ then by Lemma (ref) (viii), $\ker(\boldsymbol{M}_0)=\{0\}$ and there is a unique solution to (ref). That is to say, there can be no multiplicity of solutions for a perfectly predictable $\xi$. This point is made in a different context in linsys, p.\ 641.

In the special case of a VARMA model, Lemma (ref) reduces to the classical result that if $\det(M(z))\neq0$ for all $|z|\geq1$, there exists a unique causal covariance stationary solution for every spectral measure $F$ hd. This is due to the fact that in that case $M_+=M_0=I_m$ and $M_-=M$ is a Wiener-Hopf factorization of $M$. Note that the problem of uniqueness does not arise in the case of VARMA.

Consider the case of existence and non-uniqueness ($\kappa_1>0$ and $\kappa_m\geq0$) more closely. Lemma (ref) (v) and (viii) imply that $\ker(\boldsymbol{M}_0)$ is made up of spectral characteristics corresponding to arbitrary innovation processes (see Appendix (ref) for a detailed analysis). Thus, Lemma (ref) states that the set of all solutions to (ref) is generated by spectral characteristics corresponding to arbitrary innovation processes. The dimension of $\ker(\boldsymbol{M})$ is the dimension of the innovation space of $\xi$ multiplied by the winding number of $\det(M)$. This number is obtained in funovits,funovits2020 based on the sims framework (see also sorge).

Lemma (ref) begs the question of what happens if zeros are present on the unit circle. This is addressed in the next result.

lemIf $(M,N)$ is an LREM, $\det(M(w))=0$ for some $w\in\mathbb{T}$, and $\mathrm{rank}[\;M(w)\quad N(w)\;]=m$, there exists a spectral measure $F$ such that (ref) has no solution in $H_0^m$.
proofLet $0\neq x\in\mathbb{C}^{1\times m}$ satisfy $x M(w)=0$ and choose $F=\delta_w I_n$. If a solution $\phi\in H_0^m$ to (ref) exists, it must satisfy $\|\phi\|^2_{H^m}=\sum_{j=1}^m\int \phi_jdF\phi_j^\ast=\|\phi(w)\|^2_{\mathbb{C}^{m\times n}}<\infty$. Since $x M(z)\phi(z)=0$ for $F$--a.e.\ $z\in\mathbb{T}$, it follows that $x \boldsymbol{M}\phi=\boldsymbol{P}(x M\phi|H_0)=0$. This implies that $0=x\boldsymbol{N}I_n=\boldsymbol{P}(x N|H_0)$. Since $x N(z)=x N(w)$ for $F$--a.e.\ $z\in\mathbb{T}$ and $x N(w)\in H_0$, $x N(w)=\boldsymbol{P}(x N|H_0)=0$. This implies that $x [\;M(w)\quad N(w)\;]=0$, a contradiction.

The basic idea behind Lemma (ref) is that when $\det(M(w))=0$ for some $w\in\mathbb{T}$, the system has a form of instability akin to that of a resonance frequency in a mechanical system. When such a system is subjected to an input oscillating at frequency $\arg(w)$, its output cannot be covariance stationary. In particular, a mechanical system will oscillate with increasing amplitude until failure arnold. In the parlance of system theory, the system is said to have an unstable mode sontag.

The rank condition on $[\;M(w)\quad N(w)\;]$ in Lemma (ref) permits inputs at frequency $\arg(w)$ to excite the instability in the system. In the VARMA literature it is typically assumed that $\mathrm{rank}[\;M(z)\quad N(z)\;]=m$ for all $z\in\mathbb{C}$ hd. In the systems theory literature, similar conditions characterize controllability of the output in terms of the input kailath. Without a condition of this sort, there may be no input that can excite the system's instability. For example, the LREM $(M,N)=(1-z^{-1},1-z^{-1})$ has a solution $\phi=1$ for any $F$. Note that this condition permits oscillatory inputs to excite instability but not necessarily white noise inputs. For example, the LREM $(M,N)=(-z+3-2z^{-1},3-2z^{-1})$ has $M(1)=0$ and $N(1)=1$ so that the conditions of Lemma (ref) are satisfied but the instability of the system cannot be excited by a white noise input because, as is easily checked, $\phi=1$ is a solution to (ref) when $F=\mu$. It is possible to formulate a different condition on $N$ that will permit white noise inputs to excite instability in the system and, indeed, much more can be said about stability in LREMs. However, a general analysis of stability of LREMs is outside the scope of this paper and is left for future research.

We are now ready to state the following result.

thm[Onatski's First Theorem] If $(M,N)$ is an LREM, $\det(M(z))\neq0$ for all $z\in\mathbb{T}$, and $M$ has a Wiener-Hopf factorization $M_+M_0M_-$, then \begin{enumerate} • If the partial indices of $M$ are non-negative, then for every covariance stationary process $\xi$ there exists a solution $X$ to $(M,N)$. The general form of solutions to $(M,N)$ is \begin{align*} X_t=\int z^t \boldsymbol{M}^{(-1)}\boldsymbol{N}I_nd\Phi+\int z^t \boldsymbol{M}_-^{-1}\psi d\Phi,\quad t\in\mathbb{Z}, \end{align*} where $\boldsymbol{M}^{(-1)}$ is a right inverse of $\boldsymbol{M}$, $\psi\in \ker(\boldsymbol{M}_0)$, and the dimension of the solution space is the dimension of the innovation space of $\xi$ times the winding number of $\det(M)$. • If the partial indices of $M$ are all equal to zero, there is a unique solution for every covariance stationary process $\xi$ given by \begin{align*} X_t=\int z^t \boldsymbol{M}_-^{-1}\boldsymbol{M}_+^{-1}\boldsymbol{N}I_nd\Phi,\quad t\in\mathbb{Z}, \end{align*} • If $M$ has a negative partial index and the zero-th Fourier coefficient of the last row of $[M_+^{-1}N]_-$ is non-zero, then there exists no solution to $(M,N)$ in general. \end{enumerate}
proof(i) and (ii) follow immediately from Lemma (ref) and the spectral representation theorem. (iii) We claim that for $F=\mu I_n$, there exists no solution to $(M,N)$. For any solution must satisfy $\boldsymbol{M}_0\boldsymbol{M}_-\phi=\boldsymbol{M}_+^{-1}\boldsymbol{N}I_n=[M_+N]_-$ as $\{z^se_j: j=1,\ldots, n,s\in\mathbb{Z}\}$ is orthonormal for our choice of $F$. But the last row of this equation is $z^{\kappa_m}e_mM_-\phi=e_m[M_+^{-1}N]_-$, which evidently cannot hold if $\kappa_m<0$, as the zero-th Fourier coefficient of the left hand side is zero, while that of the right hand side is not by assumption.

The relationship between partial indices and existence and uniqueness of solutions to LREMs was first discovered by onatski. Theorem (ref) generalizes onatski by allowing for arbitrary covariance stationary input and non-constant $N$. The general expression of solutions in (i) is similar to the one obtained in Theorem 4.1 (ii) of linsys; sunspots, farmeretal, bianchinicolo, and regular provide algorithms for computing these solutions.

Note that the analogue of (iii) in onatski is incorrect: if $N$ ($\Gamma$ in Onatski's notation) is constant and non-zero, and $M$ ($A$ in Onatski's notation) has a negative (positive in Onatski's framework) partial index, it does not follow that a solution fails to exist. A counterexample is $(M,N)=\left(\left[

smallmatrix1 & 0 \\ 0 & z^{-1}

\right],\left[

smallmatrix1 \\ 0

\right]\right)$, which has the unique solution $\left[

smallmatrix1 \\ 0

\right]$ for any choice of $F$. It is easy to check, however, that Onatski's statement is correct in the scalar case $m=1$. Our additional condition on $[M_+^{-1}N]_-$, like the coprimeness condition in Lemma (ref), permits us to use white noise as an input to the system for which there cannot be an admissible output (different conditions can also be given). We will see shortly that this condition is generic.

The unit root case can be stated as

thmIf $(M,N)$ is an LREM, $\det(M(w))=0$ for some $w\in\mathbb{T}$, and $\mathrm{rank}[\;M(w)\quad N(w)\;]=m$, then there exists no solution to $(M,N)$ in general.
proofFollows from Lemma (ref) and the spectral representation theorem.

We can also state the following result for generic systems.

thm[Onatski's Second Theorem] For a generic LREM $(M,N)$ with $\det(M(z))\neq0$ for all $z\in\mathbb{T}$, there exists possibly infinitely many solutions, a unique solution, or no solution in general according to whether $\det(M(z))$ winds around the origin a positive, zero, or negative number of times as $z$ traverses $\mathbb{T}$ counter-clockwise.
proofBy Theorem (ref), $\mathcal{W}_\circ^{m\times m}$ is generic in the space of Wiener-Hopf factorizable elements of $\mathcal{W}^{m\times m}$. For fixed $M\in\mathcal{W}_\circ^{m\times m}$ with Wiener-Hopf factorization $M=M_+M_0M_-$, the generic element of $\mathcal{W}^{m\times n}$ has a zero-th Fourier coefficient of $e_m[M_+^{-1}N]_-$ that is non-zero (i.e.\ $\int e_m[M_+^{-1}N]_-d\mu\neq0$). Indeed, when $\mathcal{W}^{m\times n}$ is endowed with the $\mu$--essential supremum norm, $N\mapsto\int e_m[M_+^{-1}N]_-d\mu$ is a continuous mapping from $\mathcal{W}^{m\times n}$ to $\mathbb{C}^{1\times n}$. The result then follows from Theorem (ref) (iii).

In closing, it is important to note that the assumptions underlying existence and uniqueness in this paper are weaker than in all of the previous literature on frequency domain solutions of LREMs, namely the work of whiteman, onatski, tanwalker, tan, and meyer. These works require the exogenous process to have a purely non deterministic wold representation, an assumption that is demonstrated to be unnecessary. The weaker assumptions of this paper have also facilitated discussion of zeros of $M$, an aspect of the theory absent from the previous literature. The advantage of these stronger assumptions, however, is that they do permit closed form expressions of solutions. See Appendix (ref) for a more detailed discussion.

Ill-Posedness and Regularization

We have seen in the previous section that LREMs may have infinitely many solutions. In order to estimate and conduct inference, various selection mechanisms have been proposed in the macroeconometrics literature to associate a unique solution to any given set of parameters. We will review these proposals and show that, unfortunately, continuity of the selected solutions is not guaranteed. This leads to the development of a new regularized solution with guaranteed regularity properties.

Non-Uniqueness

The first approach to non-uniqueness, proposed by taylor, selects the solution that minimizes the variance of the price variable in the model if one exists. Taylor motivates this solution by arguing that collective rationality of economic agents will naturally lead them to coordinate their activities to achieve this equilibrium. Unfortunately, this provides no guidance for models in which indeterminacy afflicts non-price variables.

The second approach to non-uniqueness commits to a particular algorithm for obtaining a solution and ignores all other solutions. That is, it commits to a particular choice of right inverse to $\boldsymbol{M}$, say $\boldsymbol{M}^{(-1)}$, and selects the solution $\boldsymbol{M}^{(-1)}\boldsymbol{N}I_n$, ignoring all other solutions $\boldsymbol{M}^{(-1)}\boldsymbol{N}I_n+\ker(\boldsymbol{M})$. This is the “minimum state variable” approach of msv. Although this may appear to solve the problem of selecting a unique solution, there are generally infinitely many algorithms for obtaining solutions as there are infinitely many right inverses of $\boldsymbol{M}$. Thus, the concept of minimum state variable solution is not well-defined without specifying the particular algorithm or right inverse to be used for the solution. Geometrically, there is a minimum state variable solution at every single point of $\boldsymbol{M}^{(-1)}\boldsymbol{N}I_n+\ker(\boldsymbol{M})$ (see Figure (ref)). Without sound economic reasoning for the choice of algorithm or right inverse, it cannot be argued that one minimum state variable solution obtained by one algorithm is a more appropriate choice than another, obtained by a different algorithm.

figure[figure omitted — 786 chars of source]

Finally, the modern approach to non-uniqueness, as exemplified by sunspots, farmeretal, and bianchinicolo parametrizes the full set of solutions. Like the minimum state variable approach, it commits to a particular algorithm or right inverse of $\boldsymbol{M}$, say $\boldsymbol{M}^{(-1)}$, and represents each solution as a sum $\boldsymbol{M}^{(-1)}\boldsymbol{N}I_n+\chi$. The first part, $\boldsymbol{M}^{(-1)}\boldsymbol{N}I_n$, is interpreted as generated by fundamental economic forces (the ones that appear explicitly in the model (ref)), while the second part, $\chi\in\ker(\boldsymbol{M})$, generated as we have seen by arbitrary innovation processes, is interpreted as non-fundamental sunspots farmer. The coordinates on $\ker(\boldsymbol{M})$ are appended to the parameters of $(M,N)$ and both estimated by either frequentist or Bayesian methods. The geometry of the spectral approach allows us to see very clearly a serious conceptual problem with this approach. Various papers in the literature claim to present evidence of the importance of sunspots as drivers of macroeconomic activity by showing that the contribution of sunspots, as measured by the size of the estimated $\chi$, is not insignificant empirically. However, it is clear from Figure (ref) that the size of $\chi$ very much depends on the choice of algorithm or right inverse of $\boldsymbol{M}$: by one representation, $\boldsymbol{M}^{(-1)}\boldsymbol{N}I_n+\chi_1$, sunspots play a large role, by another representation, $\boldsymbol{M}^{[-1]}\boldsymbol{N}I_n+\chi_2$, sunspots play a small role. Without a sound economic reason for the choice of algorithm or right inverse of $\boldsymbol{M}$, it is unclear that the contribution of sunspots is being measured correctly. Thus, the modern approach suffers from a similar difficulty as the minimum state variable approach.

figure[figure omitted — 737 chars of source]

Discontinuity

We have established in Lemma (ref) that the set of all solutions to (ref) is of the form $\boldsymbol{M}^{(-1)}\boldsymbol{N}I_n+\ker(\boldsymbol{M})$ for some right inverse $\boldsymbol{M}^{(-1)}$ of $\boldsymbol{M}$. In the course of proving Theorem (ref) below we will see that small variations of $M$ in the $\mu$--essential supremum norm lead to small variations in the orthogonal projection onto $\ker(\boldsymbol{M})$ in the operator norm. Unfortunately, small changes in $M$ in the $\mu$--essential supremum norm are not guaranteed to lead to small changes in $\boldsymbol{M}^{(-1)}\boldsymbol{N}I_n$ in the $H^m$ norm. Discontinuity can occur when $M$ falls inside the non-generic set $\mathcal{W}^{m\times m}\backslash\mathcal{W}^{m\times m}_\circ$. We will illustrate this discontinuity with the simplest possible example. It is, of course, possible to illustrate discontinuity using a more realistic example similar to (ref) but doing so comes at the cost of analytic tractability, as solving even the simplest multivariate LREMs can be quite cumbersome. It is important to keep in mind that LREMs are typically parametrized purely on theoretical considerations, without any regard to statistical properties, so that elements of $\mathcal{W}^{m\times m}\backslash\mathcal{W}^{m\times m}_\circ$ are not expressly excluded from the parametrization onatski,generic.

Consider the following example,

align[align omitted — 200 chars of source]

with $\theta\in\mathbb{R}$. Thus, $\xi$ is a standardized white noise process.

In order to compute solutions we will need to obtain a Wiener-Hopf factorization,

align*[align* omitted — 474 chars of source]

Thus, there exists a (non-unique) solution for every $\theta\in\mathbb{R}$. Define $\boldsymbol{M}(\theta)$ as in Definition (ref), $\phi\mapsto\boldsymbol{P}(M(\,\cdot\,,\theta)\phi|H_0^2)$. Define $\boldsymbol{M}(\theta)^{(-1)}$ as in (ref) with $\boldsymbol{M}_0(\theta)^{(-1)}$ defined as in (ref). We will restrict attention to minimum state variable solutions (i.e.\ $\psi=0$ in (ref)),

align*[align* omitted — 296 chars of source]

This solution is clearly discontinuous at $\theta=0$. In particular,

align*[align* omitted — 141 chars of source]

tends to infinity as $\theta\rightarrow0$. Thus, discontinuities can arise in the process of solving an LREM.

While this discontinuity may be quite jarring to readers acquainted with the LREM literature, from the Wiener-Hopf factorization literature point of view there is nothing surprising about this discontinuity. Indeed, (ref) is a minor modification of the example given in Section 1.5 of gks. The fact that only multivariate systems with non-unique solutions can exhibit this discontinuity may explain why discontinuity has not received sufficient attention in the LREM literature, as solving such models is difficult to do analytically. The heart of the problem is that, when $m>1$, there are points in $\mathcal{W}^{m\times m}$ at which partial indices are discontinuous, these are precisely the non-generic set $\mathcal{W}^{m\times m}\backslash\mathcal{W}^{m\times m}_\circ$ gks. That is, if $M\in\mathcal{W}^{m\times m}$ has partial indices satisfying $\kappa_1-\kappa_m>1$, then there are small (in the $\mu$--essential supremum norm) changes in $M$ that lead to a jump in $M_0$ and since $M$ has undergone only a small change, $M_\pm$ must also jump. If we then choose a fixed right inverse $\boldsymbol{M}_0^{(-1)}$ in the solution (ref) (e.g.\ choosing $\boldsymbol{M}_0^{(-1)}$ as in (ref)), then the discontinuity can affect the solution. This is, unfortunately, what is done in all existing solution algorithms linsys,regular. In our example, the partial indices of $M(z,\theta)$ are $(2,0)$ for $\theta=0$ and $(1,1)$ for $\theta\neq0$; $M(z,0)\in \mathcal{W}^{2\times 2}\backslash\mathcal{W}^{2\times 2}_\circ$ and $M(z,\theta)\in \mathcal{W}^{2\times 2}_\circ$ for $\theta\neq0$. The discontinuity in the partial indices is what generates the discontinuity of the Wiener-Hopf factors, which in turn generates a discontinuity in the solution.

It bears emphasizing that this discontinuity is a feature of the mathematical problem, it is not a feature of the particular algorithm used to solve the problem; regular shows how it arises in the sims framework, linsys shows how it arises in a linear systems framework, and Appendix (ref) shows how it arises when solving the system by hand. It is also important to note that the discontinuity of Wiener-Hopf factorization implies that there can exist no numerically stable way to compute it generally (see Example C.2 of the online supplement to linsys for an illustration). While the elements of $\mathcal{W}^{m\times m}_\circ$ can be factorized using finite precision arithmetic, the elements of $\mathcal{W}^{m\times m}\backslash\mathcal{W}^{m\times m}_\circ$ cannot be factorized without infinite precision. Thus, it is a non-starter to verify numerically whether a given system is generic or not in the process of estimation and inference. Note that an analogous problem arises in the computation of the Jordan canonical form, which is discontinuous at a non-generic set of matrices in $\mathbb{C}^{m\times m}$ when $m>1$ hj1; there, the recommendation is to compute the Schur canonical form instead; and in the next section we will propose, similarly, to compute different solutions to LREMs than the ones proposed in the literature.

Regularization

Having established that the LREM problem in macroeconometrics is ill-posed, we now consider how to obtain economically meaningful solutions amenable to mainstream econometric techniques.

Perhaps the most natural solution to the ill-posedness problem is to avoid it altogether and restrict attention to systems with unique solutions (i.e.\ LREMs $(M,N)$, where the partial indices of $M$ are equal to zero). ident have shown that unique solutions are not only continuous in the parameters of an LREM but also analytic (see the proof of Theorem 6.2). A less restrictive solution is to allow for non-uniqueness but restrict attention to generic systems (i.e.\ $\mathcal{W}^{m\times m}_\circ$). However, genericity cannot be taken for granted as both onatski and generic have warned and some models may be parametrized to always fall inside $\mathcal{W}^{m\times m}\backslash\mathcal{W}^{m\times m}_\circ$ (interestingly, generic is widely but erroneously considered to be a critique of onatski, see blog). Moreover, we need differentiability, not just continuity, in order to ensure asymptotic normality of extremum estimators pp,hd as it facilitates the construction of confidence intervals as well as hypothesis testing nm. Therefore, we opt for a more straightforward solution, regularization.

With the geometry of the spectral approach in view, one is led inexorably to consider the Tykhonov-regularized solution to (ref), which minimizes the total variance, $\|X_0\|^2_{\mathscr{H}^m}=\|\phi\|^2_{H^m}$ among all solutions to (ref),

align*[align* omitted — 114 chars of source]

It can be shown that $\phi_{\boldsymbol{I}}=\boldsymbol{M}^\dag\boldsymbol{N}I_n$, where $\boldsymbol{M}^\dag$ is the Moore-Penrose inverse of $\boldsymbol{M}$ groetsch. This takes a particularly simple form in our context.

lemIf $M\in\mathcal{W}^{m\times m}$, $\det(M(z))\neq0$ for all $z\in\mathbb{T}$, and the partial indices of $M$ are non-negative then, \begin{align*} \boldsymbol{M}^\dag=\boldsymbol{M}^\ast(\boldsymbol{M}\boldsymbol{M}^\ast)^{-1}. \end{align*}
proofBy Lemma (ref), $\boldsymbol{M}$ is onto. It follows that $\boldsymbol{M}^\dag=\boldsymbol{M}^\ast(\boldsymbol{M}\boldsymbol{M}^\ast)^\dag$ groetsch. Since $\boldsymbol{M}$ is onto, $\boldsymbol{M}^\ast$ is one-to-one and so $(\boldsymbol{M}\boldsymbol{M}^\ast)^\dag=(\boldsymbol{M}\boldsymbol{M}^\ast)^{-1}$.

Lemma (ref) implies a very simple technique for computing the Tykhonov-regularized solution. Whenever a solution to (ref) exists, the Tykhonov-regularized solution is obtained from the unique solution to the auxiliary block triangular frequency domain system,

align*[align* omitted — 274 chars of source]

This implies that whenever a solution exists, the regularized solution is obtained uniquely by solving an auxiliary LREM. For example, in the mixed model from Section (ref), the Tykhonov-regularized solution is the solution $X$ to the auxiliary LREM,

gather*[gather* omitted — 243 chars of source]

This regularized solution exists and is unique whenever the original system satisfies the conditions for existence.

Note that when the partial indices of $M$ are all equal to zero, $\boldsymbol{M}^\dag=\boldsymbol{M}^\ast(\boldsymbol{M}\boldsymbol{M}^\ast)^{-1}=\boldsymbol{M}^\ast(\boldsymbol{M}^\ast)^{-1}(\boldsymbol{M})^{-1}=(\boldsymbol{M})^{-1}$. In other words, Tykhonov-regularization has no effect if the solution is unique.

figure[figure omitted — 952 chars of source]

More generally, we may consider

align*[align* omitted — 130 chars of source]

where $\boldsymbol{L}:H^m\rightarrow H^l$ is a bounded linear operator chosen by the researcher. Next, we consider economic motivation for a number of choices of $\boldsymbol{L}$.

If the researcher wishes to shrink the variance of the $i$-th component of $X$, they can set

align*[align* omitted — 48 chars of source]

Note that when the $i$-th component is the price variable, we obtain the taylor solution.

One can also shrink expected values of $X$. For example, if certain solutions obtained by the methods reviewed above yield expectations of output that are too variable relative to what one expects empirically, then one can impose this prior by using

align*[align* omitted — 62 chars of source]

where $j$ is the coordinate corresponding to output in $X$.

Linear combinations of lagged, current, and expected values of coordinates of $X$ can also be shrunk similarly. The operator

align*[align* omitted — 220 chars of source]

shrinks the second difference of all variables, imposing smoothness on solutions, similar to the idea of the hp filter.

Note that the operators above belong to the class of operators defined in Definition (ref). Thus, we can more generally use any $L\in\mathcal{W}^{l\times m}$ to construct the weight

align*[align* omitted — 172 chars of source]

More importantly, regularization can allow the researcher to shrink not just across time but across frequencies. For example, the researcher may wish to shrink the spectrum of the solution towards frequencies of between $2\pi/32$ and $2\pi/4$ corresponding to business cycle fluctuations of period 4-32 quarters in quarterly data, i.e.\ using

align*[align* omitted — 141 chars of source]

Finally, we can consider solutions that minimize a finite weighted sum of individual criteria as reviewed above,

align*[align* omitted — 101 chars of source]

where $\boldsymbol{L}_i:H^m\rightarrow H^{l_i}$ and $a_i>0$ for $i=1,\ldots,d$. This would allow the researcher to impose $d$ different criteria according to the weights $a_1,\ldots,a_d$. This can be achieved by using

align*[align* omitted — 152 chars of source]

The argument for regularization is the same as employed throughout the inverse and ill-posed problems literatures: if theory is insufficient to pin down a unique continuous solution, other information can be employed. In our case, regularization allows economically meaningful shrinkage criteria to select the best among all possible solutions or, what amounts to the same, it allows the researcher's priors about economic behavior to select the most appropriate solution. Regularization resolves the problem of selecting an economically grounded solution and, as we will soon see, also ameliorates the discontinuity problem.

Existence of regularized solutions is guaranteed whenever the given LREM satisfies the already minimal conditions for existence of solutions. That is because, according to Lemma (ref), the set of all solutions is finite dimensional and so the problem of minimizing $\|\boldsymbol{L}\phi\|^2_{H^l}$ subject to $\boldsymbol{M}\phi=\boldsymbol{N}I_n$ reduces to a problem of minimizing a non-negative quadratic form (see Figure (ref)). Uniqueness of regularized solutions, on the other hand, is the subject of the next result.

lemIf $(M,N)$ is an LREM, $\det(M(z))\neq0$ for all $z\in\mathbb{T}$, the partial indices of $M$ are non-negative, and $\boldsymbol{L}:H^m\rightarrow H^l$ is a bounded linear operator, then \begin{align*} \phi_{\boldsymbol{L}}=(\boldsymbol{I}-(\boldsymbol{L}|_{\ker(\boldsymbol{M})})^\dag \boldsymbol{L})\boldsymbol{M}^\dag \boldsymbol{N}I_n, \end{align*} where $\boldsymbol{L}|_{\ker(\boldsymbol{M})}$ is the restriction of $\boldsymbol{L}$ to $\ker(\boldsymbol{M})$, is a regularized solution to (ref). $\phi_{\boldsymbol{L}}$ is the only element of $\arg\min\left\{\|\boldsymbol{L}\phi\|^2_{H^l}: \boldsymbol{M}\phi=\boldsymbol{N}I_n\right\}$ in $H_0^m$ if and only if $\ker(\boldsymbol{L})\cap\ker(\boldsymbol{M})=\{0\}$.
proofSee Appendix (ref).

In the course of proving Lemma (ref) we find that the set of regularized solutions is $\phi_{\boldsymbol{L}}+\ker(\boldsymbol{M})\cap\ker(\boldsymbol{L})$. Thus, regularization produces a unique solution if and only if $\ker(\boldsymbol{L})\cap\ker(\boldsymbol{M})=\{0\}$. Geometrically, this condition requires the operator $\boldsymbol{L}$ to put weight on all directions of indeterminacy of the LREM; if $\ker(\boldsymbol{L})\cap\ker(\boldsymbol{M})\neq\{0\}$, it will be possible to perturb $\phi_{\boldsymbol{L}}$ in any direction in $\ker(\boldsymbol{L})\cap\ker(\boldsymbol{M})$ to arrive at another regularized solution and uniqueness will fail. A host of other representations of $\phi_{\boldsymbol{L}}$ are obtained in Appendix (ref).

Some special cases of Lemma (ref) are particularly instructive. When $\ker(\boldsymbol{L})=\{0\}$, the regularized solution is unique regardless of $\boldsymbol{M}$. That is, when $\boldsymbol{L}$ puts weight on all directions in the solution space, the regularized solution is unique. An example of this is $\boldsymbol{L}=\boldsymbol{I}$, which produces the Tykhonov-regularized solution $\phi_{\boldsymbol{I}}=\boldsymbol{M}^\dag\boldsymbol{N}I_n$. When $\ker(\boldsymbol{M})=\{0\}$ (i.e.\ the solution to the LREM is unique), the regularized solution is the unique solution regardless of $\boldsymbol{L}$. This is due to the fact that $(\boldsymbol{L}|_{\ker(\boldsymbol{M})})^\dag$ is the zero operator (because $\boldsymbol{L}|_{\ker(\boldsymbol{M})}$ is the zero operator) and $\boldsymbol{M}^\dag=(\boldsymbol{M})^{-1}$ so that $\phi_{\boldsymbol{L}}=(\boldsymbol{M})^{-1}\boldsymbol{N}I_n$, the unique solution.

Recalling that $\phi_{\boldsymbol{L}}$ is a solution to the LREM, note how the mapping from parameters to the set of all solutions, $(M,N)\mapsto\phi_{\boldsymbol{L}}+\ker(\boldsymbol{M})$, compares to the mapping from parameters to the set of regularized solutions $(M,N)\mapsto\phi_{\boldsymbol{L}}+\ker(\boldsymbol{M})\cap\ker(\boldsymbol{L})$. The difference is simply a matter of dimension reduction. It is important to note that neither mapping is generally one-to-one and solution sets associated with different parameters may intersect.

thmIf $(M,N)$ is an LREM, $\det(M(z))\neq0$ for all $z\in\mathbb{T}$, the partial indices of $M$ are non-negative, and $\boldsymbol{L}:H^m\rightarrow H^l$ is a bounded linear operator, then for every covariance stationary process $\xi$ there exists a solution $X$ minimizing $\|\boldsymbol{L}\phi\|_{H^l}$ given by \begin{align*} X_t=\int z^t (\boldsymbol{I}-(\boldsymbol{L}|_{\ker(\boldsymbol{M})})^\dag \boldsymbol{L})\boldsymbol{M}^\dag\boldsymbol{N}I_n d\Phi,\quad t\in\mathbb{Z}. \end{align*} The solution is unique if and only if $\ker(\boldsymbol{M})\cap\ker(\boldsymbol{L})=\{0\}$.
proofFollows from Lemma (ref) and the spectral representation theorem.

The expression for regularized solutions in Theorem (ref) is primarily of theoretical interest. It will allow us to study continuity and differentiability with respect to underlying parameters. For estimation and inference, on the other hand, regular provides a numerical algorithm for computing regularized solutions in the sims framework, which leads to equivalent algebraic rather than geometric criteria for existence and uniqueness.

We have shown that, with an appropriate choice of $\boldsymbol{L}$, the regularized solution can overcome the non-uniqueness problem. Our next result shows that regularized solutions also overcome the discontinuity problem. The continuity guarantee that we require for mainstream econometric methodology is not with respect to the $H^m$ norm but with respect to the $\mu$--essential supremum norm; see e.g.\ the continuity results for bounded spectral densities and the Gaussian likelihood functions in hannan,dp,anderson. If we parametrize $M$ and $N$ as $M(z,\theta)$ and $N(z,\theta)$, then it is clear that we need $M(z,\theta)$ and $N(z,\theta)$ to be jointly continuous in $z$ and $\theta$. However, ga note that this is not sufficient to ensure continuity of the Wiener-Hopf factors in the $\mu$--essential supremum norm. Thus, we use Green and Anderson's idea of imposing control over $\frac{d}{dz}M(z,\theta)$ and $\frac{d}{dz}N(z,\theta)$.

thmLet $M:\mathbb{T}\times\Theta\rightarrow\mathbb{C}^{m\times m}$ and $N:\mathbb{T}\times\Theta\rightarrow\mathbb{C}^{m\times n}$. Under the conditions \begin{enumerate} • $F=\mu I_n$. • $\Theta\subset\mathbb{R}^d$ is an open set. • $M(\,\cdot\,,\theta)$ and $N(\,\cdot\,,\theta)$ are analytic in a neighborhood of $\mathbb{T}$ for every $\theta\in\Theta$. • $M(z,\theta)$, $N(z,\theta)$, $\frac{d}{dz}M(z,\theta)$, and $\frac{d}{dz}N(z,\theta)$ are jointly continuous at every $(z,\theta)\in\mathbb{T}\times\Theta$. • $\ker(\boldsymbol{L})\cap\ker(\boldsymbol{M}(\theta))=\{0\}$ for all $\theta\in\Theta$. • $\det(M(z,\theta))\neq0$ for all $z\in\mathbb{T}$, and the partial indices of $M(z,\theta)$ are all non-negative for all $\theta\in\Theta$. \end{enumerate} Then $\phi_{\boldsymbol{L}}(\theta)=(\boldsymbol{I}-(\boldsymbol{L}|_{\ker(\boldsymbol{M}(\theta))})^\dag \boldsymbol{L})\boldsymbol{M}(\theta)^\dag\boldsymbol{N}(\theta)I_n $, is continuous in the $\mu$--essential supremum norm.
proofSee Appendix (ref).

The assumptions of Theorem (ref) are quite strong relative to the discussion so far. However, the relevant case for most macroeconometric applications is the case where $F=\mu I_n$, while $M(z,\theta)$ and $N(z,\theta)$ are Laurent matrix polynomials of uniformly bounded degree. In this case, the continuity of the coefficients of $M(z,\theta)$ and $N(z,\theta)$ is sufficient to ensure conditions (iii) and (iv) of Theorem (ref). The New-Keynesian model (ref) satisfies these conditions, as do all of the LREMs in canova, dd, or hs.

For the purpose of establishing asymptotic normality, we typically need not just continuity in the essential supremum norm but also differentiability in the essential supremum norm (i.e. the finite differential in $\theta$ converges to the infinitesimal differential in the essential supremum norm over $z\in\mathbb{T}$). This stronger form of differentiability allows us to differentiate under integrals, which appear in the asymptotics of maximum likelihood and generalized method of moments estimators. The following result provides exactly what we need.

thmLet $p$ be a positive integer. Under assumptions (i) -- (vi) of Theorem (ref) and, \begin{enumerate} • $M(z,\theta)$, $N(z,\theta)$, $\frac{d}{dz}M(z,\theta)$, and $\frac{d}{dz}N(z,\theta)$ are jointly continuously differentiable of all orders up to $p$ for all $(z,\theta)\in\mathbb{T}\times\Theta$. \end{enumerate} Then $\phi_{\boldsymbol{L}}(\theta)=(\boldsymbol{I}-(\boldsymbol{L}|_{\ker(\boldsymbol{M}(\theta))})^\dag \boldsymbol{L})\boldsymbol{M}(\theta)^\dag\boldsymbol{N}(\theta)I_n$, is continuously differentiable of order $p$ with respect to $\theta$ in the $\mu$--essential supremum norm.
proofSee Appendix (ref).

Condition (vii) of Theorem (ref) is a direct strengthening of condition (iv) of Theorem (ref). When $M(z,\theta)$ and $N(z,\theta)$ are Laurent matrix polynomials of uniformly bounded degree, as is usually the case in macroeconomic models, the $p$-th order continuous differentiability of the coefficients of $M(z,\theta)$ and $N(z,\theta)$ is sufficient to ensure this condition is satisfied.

To see Theorem (ref) in action, consider the regularized solution to the example from the previous subsection. This can be obtained analytically in just a handful of steps.

align*[align* omitted — 931 chars of source]

To invert the operator in the expression above, we note that it is of the form defined in Definition (ref), with underlying $\mathcal{W}^{2\times 2}$ element, $\left[

smallmatrix1 & \theta z \\ \theta z^{-1} & (\theta^2+1)

\right]$. This matrix has the Wiener-Hopf factorization $\left[

smallmatrix1 & \frac{\theta}{\theta^2+1}z \\ 0 & 1

\right]\left[

smallmatrix1 & 0 \\ 0 & 1

\right]\left[

smallmatrix\frac{1}{\theta^2+1} & 0 \\ \theta z^{-1} & \theta^2+1

\right]$. Thus,

align*[align* omitted — 1,535 chars of source]

where the last equality follows from the fact that $\ker(\boldsymbol{V})=\mathbb{C}^{1\times 2}$ when $F=\mu I_2$. As guaranteed by Theorems (ref) and (ref), this solution is not just continuous as a function of $\theta$ but also smooth. Note that the implicit right inverse of $\boldsymbol{M}_0$ in $\boldsymbol{M}^\dag$ is

align*[align* omitted — 428 chars of source]

Thus, it is only by using a special right inverse of $\boldsymbol{M}_0(\theta)$ that we can absorb the effect of discontinuity in the partial indices at $\theta=0$ on the solution.

To summarize, the LREM problem in macroeconometrics is ill-posed. However, regularization produces solutions that are unique, continuous, and even smooth under very general regularity conditions.

Application

As an application of the spectral approach to LREMs, we apply mainstream methodology to estimate and draw inference on the non-generic model of Section (ref). We will see that the Gaussian likelihood function displays very irregular behavior, invalidating underlying assumptions of mainstream frequentist and Bayesian analysis. In turn, regularized solutions can avoid some of these anomalies.

Consider the non-generic system of Section (ref) in the time domain,

align*[align* omitted — 173 chars of source]

with $\xi$ an i.i.d.\ Gaussian sequence of mean zero and covariance matrix $I_2$. The objective is to estimate and draw inference on $\theta_0\in\mathbb{R}$ from an observed dataset $x_T=(X_1^\ast,\ldots,X_T^\ast)^\ast$. We will follow mainstream methodology in disregarding uniqueness, identifiability, and invertibility in parametrizing the model (see ident for a careful parametrization that addresses all of these considerations). The solution we computed in the previous section is

align*[align* omitted — 271 chars of source]

Let $\Sigma_T(\theta)$ be the covariance of $x_T$ according to the model parametrized by $\theta$. Then the likelihood function is

align*[align* omitted — 140 chars of source]

It is more convenient to work with

align*[align* omitted — 109 chars of source]

We will refer to $\ell_T(\theta)$ as the log-likelihood function. Observe that

align*[align* omitted — 387 chars of source]

where

align*[align* omitted — 166 chars of source]

Then $\Sigma_T(0)^{-1}=I_{2T}$ and it is easily checked that for $\theta\neq0$,

align*[align* omitted — 433 chars of source]

where

align*[align* omitted — 371 chars of source]

It is also easily checked that $\det(\Sigma_T(\theta))=1+\theta^2$ for $\theta\neq0$ for all $T\geq1$. Substituting into the log-likelihood function we find that

align*[align* omitted — 59 chars of source]

and, for $\theta\neq0$,

align*[align* omitted — 306 chars of source]

It is clear that

align*[align* omitted — 126 chars of source]

almost surely because $x_T$ has a continuous distribution regardless of the value of $\theta_0$. Since $\ell_T(\theta)$ is continuous on $\mathbb{R}\backslash\{0\}$ and bounded from below, $\ell_T(\theta)$ must attain local minima somewhere to the left and to the right of $\theta=0$ with probability 1. Thus, almost surely, $\ell_T(\theta)$ is $W$ shaped, diverging as $0\neq|\theta|\rightarrow0$ and as $|\theta|\rightarrow\infty$, although it has a finite value at $\theta=0$. See Figure (ref) for an illustration with simulated data. It follows that the likelihood function $p(x_T|\theta)$ almost surely has at least three local maxima, one of which is $\theta=0$. Moreover, almost surely $\lim_{0\neq|\theta|\rightarrow0}p(x_T|\theta)=0$ and $p(x_T|\theta)|_{\theta=0}>0$ so the likelihood function almost surely has a simple discontinuity at $\theta=0$.

figure[figure omitted — 226 chars of source]

Taking the limit $T\rightarrow\infty$ is straightforward as most terms vanish. By the strong law of large numbers, almost surely,

align*[align* omitted — 210 chars of source]

The moments above can be read from $G_0$ and $G_1$ evaluated at $\theta=\theta_0$ so the right hand side is

align*[align* omitted — 296 chars of source]

See Figure (ref) for plots of $\ell(\theta,0)$ and $\ell(\theta,1)$. It is easy to check that $\ell(\theta,0)$ has local minima at $0$, $+1$, and $-1$ and if $\theta_0\neq0$, $\ell(\theta,\theta_0)$ has local minima at $0$, $\theta_0$, and a third point of the opposite sign to $\theta_0$ (the expression is complicated so we omit it). Note that for $\theta\in\mathbb{R}\backslash\{0\}$, $\ell_T(\theta)-\ell(\theta,\theta_0)$ is expressible as a sum of eight terms, each of which is a product of a continuous function in $\theta$ and a quantity that converges almost surely to zero. Thus, $\lim_{T\rightarrow\infty}\sup_{\theta\in\Theta}|\ell_T(\theta)-\ell(\theta,\theta_0)|=0$ for any compact subset $\Theta\subset\mathbb{R}\backslash\{0\}$.

Consider the case $\theta_0=0$. It is clear from the discussion above that the maximum likelihood estimator is consistent, although it is not asymptotically normal, and any prior that has an atom at the point $\theta=0$ will yield a posterior distribution that concentrates at the true value. However, this analytical approach cannot be carried over to a general methodology because analytical descriptions of solutions are generally infeasible for all but the simplest LREMs; the mapping from $(M,N)$ to any solution is highly non-linear. The practitioner who seeks to estimate this model on a computer may suspect a discontinuity at $\theta=0$ because they will be able to see the solution behaving erratically near that point but being short of an analytical description of the solution: (i) they will not have any theoretical guarantees as to what happens when $\theta\rightarrow0$ because none exist and (ii) they will not be able to compute the solution at $\theta=0$ reliably due to the numerical instability of computing Wiener-Hopf factorizations of non-generic elements of $\mathcal{W}^{m\times m}$.

Now consider the likely outcome of applying mainstream econometric methodology to the data when $\theta_0=0$.

Let us begin with the most basic of empirical analyses: plotting the likelihood function. Each evaluation of the likelihood function is obtained by solving the model numerically then using the Kalman filter canova,dd,hs. Without knowing a priori the importance of the likelihood function at $\theta=0$, the researcher will not know to include that point specifically in the list plot of the likelihood function. Even if they did include it, the plot is not guaranteed to reveal the value of the likelihood function at this point because of the numerical instability of the computation at this point. Thus, the plot of the likelihood function will appear to have modes near $\pm1$, while the population parameter, $\theta_0=0$, may (and most likely will) appear to be the least likely point of the parameter space because $\lim_{0\neq|\theta|\rightarrow0}p(x_T|\theta)=0$.

Mainstream frequentist analysis utilizes variations of the Newton-Raphson algorithm to minimize $\ell_T(\theta)$, then obtains confidence intervals from derivatives of $\ell_T(\theta)$ canova,dd. This exercise is almost certain to fail in our setting because, again, without knowing a priori that the point $\theta=0$ is crucial to the analysis, the algorithm will be initialized in $\mathbb{R}\backslash\{0\}$ and it will converge almost surely to (approximately) either $+1$ or $-1$ depending on where it is initialized. The standard error of the estimate, computed using the second derivative of $\ell_T(\theta)$ at the estimated $\theta$, will be of order $O(T^{-1/2})$ almost surely so as the sample size gets larger it will appear to the researcher that the estimate is precise when it is in fact completely wrong.

Mainstream Bayesian analysis utilizes variations on the random walk Metropolis-Hastings algorithm to sample from the posterior distribution function,

align*[align* omitted — 58 chars of source]

where $p(\theta)$ is a prior distribution function, before reporting estimates of the mean or median of the posterior distribution as well as credibility regions canova,dd,hs. Here, again, estimation is certain to fail because, being unaware of the significance of the point $\theta=0$, the researcher is likely to use a continuous prior that places probability zero at the set $\{0\}$. The result is that the posterior distribution will concentrate around $\pm1$. The only way to allow the prior to be informative about $\theta$ is to allow for an atom in the prior at $\theta=0$ but none of the aforementioned textbooks present algorithms that can accommodate posterior distributions with atoms.

In contrast to the above, the regularized solution,

align*[align* omitted — 149 chars of source]

displays no such anomalies. If $\tilde{x}_T=(\tilde{X}_1^\ast,\ldots,\tilde{X}_T^\ast)^\ast$ and $\tilde{\Sigma}_T(\theta)$ is the covariance of $\tilde x_T$ according to the model parametrized by $\theta$. Then the $ij$-th block of $\tilde{\Sigma}_T(\theta)$ is of the form $\int z^{i-j}\phi_{\boldsymbol{I}}(\theta)\phi_{\boldsymbol{I}}(\theta)^\ast d\mu$. Since $\phi_{\boldsymbol{I}}(\theta)$ is continuously differentiable of any order with respect to $\theta$ in the $\mu$--essential supremum norm, we can exchange the integration with differentiation with respect to $\theta$ so $\tilde{\Sigma}_T(\theta)$ is continuously differentiable of any order with respect to $\theta$. This implies that the same is true of the likelihood function at every point $\theta\in\mathbb{R}$ where $\tilde{\Sigma}_T(\theta)^{-1}$ exists. It is easy to check that $\tilde{\Sigma}_T(\theta)$ and $\tilde{\Sigma}_T(\theta)^{-1}$ have the same block structure as $\Sigma_T(\theta)$ and $\Sigma_T(\theta)^{-1}$ respectively. The only difference is that now

align*[align* omitted — 984 chars of source]

Figure (ref) provides a plot of the log-likelihood function of the regularized solution with simulated data along with the almost sure limit in $T$. It is clear now that frequentist and Bayesian analyses can be carried out using the regularized solution. Note that the first, second, and third derivatives of the limiting log-likelihood function vanish at $\theta=0$ when $\theta_0=0$. Thus a slightly more delicate frequentist analysis is necessary (see e.g.\ rotnitzky). This is, again, made possible by the existence of derivatives of all orders, thanks to regularization.

figure[figure omitted — 222 chars of source]

Conclusion

This paper has extended the LREM literature in the direction of spectral analysis. It has done so by relaxing common assumptions and developing a new regularized solution. The spectral approach has allowed us to study examples of limiting Gaussian likelihood functions of simple LREMs, which demonstrate the advantages of the new regularized solution as well as highlighting weaknesses in mainstream methodology. For the remainder, we consider some implications for future work.

The regularized solution proposed in this paper is the natural one to consider for the frequency domain. However, its motivation has been entirely econometric in nature and this begs the question of whether it can be derived from decision-theoretic foundations as proposed by taylor. Regularization has already made inroads into decision theory gabaix. This line of inquiry may yield other forms of regularization which may have more interesting dynamic or statistical properties.

The parameter space in LREMs can be disconnected and it can matter a great deal where one initializes their optimization routine to find the maximum likelihood estimator or their exploration routine for sampling from the posterior distribution. Therefore, it would be useful to develop simple preliminary estimators of LREMs analogous to the results for VARMA (e.g.\ Sections 8.4 and 11.5 of bd) that can provide good initial conditions for frequentist and Bayesian algorithms.

Wiener-Hopf factorization theory has been demonstrated here and in previous work (onatski, linsys, ident) to be the appropriate mathematical framework for analysing LREMs. This begs the question of what is the appropriate framework for non-linear rational expectations models. The hope is that the mathematical insights from the theory of linear models will allow for important advances in non-linear modeling and inference.

Finally, researchers often rely on high-level assumptions as tentative placeholders when a result seems plausible but a proof from first principles is not apparent. Continuity of solutions to LREMs with respect to parameters has for a long time been one such high-level assumption in the LREM literature. The fact that it is generally false, should give us pause to reflect on the prevalence of this technique. At the same time, the author hopes to have conveyed a sense of optimism that theoretical progress from first principles is possible.

Appendix