EconBase
← Back to paper

Tilting Approximate Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

92,344 characters · 28 sections · 61 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Tilting Approximate Models

\email{[email removed]}

\setcounter{page}{1} \pagenumbering{arabic} \pagestyle{plain} \thispagestyle{empty}

\onehalfspacing

abstractModel approximations are common practice when estimating structural or quasi-structural models. The paper proposes utilizing projections to re-impose information about the exact model in the form of conditional moments. The resulting estimator efficiently combines the information provided by the approximate model and the moment conditions. The paper develops the corresponding asymptotic theory and provides simulation evidence that tilting substantially reduces the mean squared error for parameter estimates. It applies the methodology to pricing long-run risks in aggregate consumption in the US, whereas the model is solved using the doi:10.1111/j.1540-6261.1988.tb04598.x approximation. Tilting improves empirical fit and results suggest that approximation error is a source of upward bias in estimates of risk aversion and downward bias in the elasticity of intertemporal substitution. \newline \newline Keywords: Information Projections, Approximation Bias, Non-linearity, Asset Pricing.\newline JEL Classification: C10, C51, E44

\onehalfspacing

Introduction

Model approximations are quite common in structural estimation as they are regularly used to obtain the equilibrium law of motion, or more generally, the reduced form of a model. It is well documented in the literature that alternative methods deliver approximation errors that may not be uniform across the state space, while reducing the error can be quite costly in terms of running time (see e.g. FERNANDEZVILLAVERDE2016527 for structural macroeconomic models). The quality of approximation employed for estimation can induce a form of misspecification that can be consequential for judging the importance of different mechanisms and structures. For example, JOFI:JOFI12615 show that for asset pricing models with long run risks, the error in moments when using a log-linear approximation can exceed $70\%$, while the equity risk premium can be overestimated by 100 basis points. doi:10.3982/ECTA12791 find that errors from the first and second order perturbation solutions of the New Keynesian model can exceed $100\%$.\footnote{Approximations also matter for parameter identification, as they can effectively limit the model's empirical content. For example, Canova2009431 shows that in log-linearized DSGE models the mapping between the structural parameters and the coefficients of the solution (which map to first and second moments in the data) is ill-behaved.}

The trade-off between tractability and misspecification becomes very relevant to parameter estimation, as the solution must be repeatedly evaluated at alternative points in the parameter space. This paper considers the use of conditional density projections (see Komunjer_Ragusa and Czizar for the unconditional case) and shows that when applied to structural models, they effectively combine the information produced by the approximate model with the one contained in the exact equilibrium conditions. The approximate model is represented by a conditional probability measure with density $f(X|Z,\varphi)$, indexed by parameter vector $\varphi$ and conditioned on $Z$. The projection tilts the approximate model by obtaining a probability distribution that is as close as possible to the base measure and satisfies the conditional moments of the exact model, that is $\mathbb{E}(m(X,\vartheta )|Z)=0$, where $\vartheta$ is the structural parameter of interest. For example, the approximate density can be the one generated by a complete log-linearized dynamic equilibrium model for $X$ given $Z$ (which includes lagged values of $X$), while the exact moment could be the non-linear Euler equation for consumption.\footnote{Potential applications extend beyond macroeconomics. Most structural microeconomic models that are based on optimizing agents have to satisfy analogous conditions. Other examples include models of network formation, where the conditional density of an exponential random graph model is usually approximated, and the pseudo-likelihood may not correspond to a complete set of sufficient statistics (see e.g. paula_2017,10.1214/13-AOS1155 and references therein). The information projection can potentially re-impose this information on the pseudo-likelihood.)}

More intuition can be built by previewing one of the results of this paper. Under a sequence of approximate models local to the data generating process, maximizing the likelihood of the tilted model is equivalent to exploiting the following moment for estimation, which automatically combines the efficient GMM first order condition with the (orthogonalized) score of the approximate model:\footnote{$(M_{z},V_{m,z})$ signify the Jacobian and Variance of the conditional moments.}

eqnarray[eqnarray omitted — 322 chars of source]

By construction, the implied estimate for $\vartheta$ is going to be more efficient than the one that only exploits $\mathbb{E}(m(X,\vartheta )|Z)=0$ for estimation. Symmetrically, the bias arising from employing the approximate model will be much lower once it is combined with the exact moment condition. The result is a substantial reduction in the mean squared error for estimating $\vartheta$.\footnote{In a very different setting, DITRAGLIA2016187 proposes using invalid instruments on purpose in an IV framework. In this paper the moment conditions that are based on "invalid instrument" are the scores of the approximate model. The difference is that the additional moment conditions are not arbitrary, but pinned down by the approximate model, which achieves the best possible efficiency in the correct specification limit. }

Information projections have been considered in different contexts, such as forecasting (see e.g. Giacomini2014145, 10.2307/3839160) and parameter estimation using conditional or unconditional moment restrictions (see e.g. the use of exponential tilting in SusanneSchennah,KitamuraStutzer, 10.2307/2998561, and Generalized Empirical Likelihood (GEL) i.e. ECTA:ECTA482). Within the latter literature, a non-parametric model of the (joint) distribution of the data is forced to satisfy a set of (conditional) moment restrictions in the first step, giving rise to a non-parametric likelihood function which can be used to estimate and do inference on the parameters indexing the moment condition. This paper departs from this strand of the literature, and in particular SusanneSchennah, by considering a generalized version of exponential tilting in the first step, where the form of $f(X|Z,\varphi)$ is parametrically specified. This expands the range of probability measures that can be considered, as it allows us to investigate the performance of estimators based on tilting structural or quasi-structural models. Within the macroeconomic literature, information projections have been considered by 10.3982/QE737 as a way to improve the numerical accuracy of discretization of Markov processes such as income, which are then used to price disaster risk. This paper considers the econometric properties of employing projections to restore information lost by approximating the law of motion in the generic case.\footnote{Similar in spirit, KRISTENSEN2017189 consider Newton-Raphson adjustments of approximate estimators to remove the first-order bias that arises from stochastic and non-stochastic approximations. In a broadly related paper, https://doi.org/10.3982/QE1413 propose composite estimators to deal with misspecification in macroeconomic models. }

The paper provides consistency and asymptotic distribution results for a generic $(\vartheta,\varphi)$, where $\varphi$ is finite dimensional. The main results concern a density which is induced by a structural approximation, such as in models with optimizing agents, and hence there is a mapping from $\varphi$ to $\vartheta$. When misspecification vanishes at a fast enough rate, $\hat{\vartheta}$ attains the parametric lower bound (Cramer-Rao bound). Analogous rate conditions have been considered by Geweke_Hahn_Ackerberg when estimating approximated equilibrium models using likelihood methods. The paper also derives the asymptotic distribution along drifting parameter sequences that are local to the true model, where misspecification is in the form of improper finite dimensional restrictions to the true density and the error vanishes at a root-n rate. The pseudo maximum likelihood estimator (PMLE) of the tilted approximate model can dominate in terms of asymptotic risk the MLE of the exact model. This dominance is the result of both higher efficiency and low bias of the PMLE based on the tilted approximate model. Simulation evidence shows that tilting significantly lowers the MSE as compared to the "naive" MLE based on the approximate model, which has a higher bias. For completeness and ease of exposition, the paper first provides results when the base density is a non-structural model, and hence $\varphi$ and $\vartheta$ are independent. When misspecification vanishes at a fast enough rate, $\hat{\vartheta}$ attains the semi-parametric lower bound (SLB) (see CHAMBERLAIN1987305).\footnote{This is the same bound achieved by other semi-parametric estimators such as GMM and GEL and its variants. In simulation evidence based on the Mean Squared Error (MSE), the estimator is compared to the Continuously Updated GMM estimator (CU-GMM) and the Empirical Likelihood (EL) estimator in the conditional moment case. Additional evidence is provided regarding CU-GMM and the Exponentially Tilted EL (ETEL) in the unconditional moment case as well. In all cases, the estimator behaves favorably in finite samples. }

The estimation approach proposed in the paper can be also casted as an adversarial estimator. Recent work has highlighted the connections between adversarial estimation and GEL estimation (see e.g. https://doi.org/10.48550/arxiv.2007.06169,https://doi.org/10.48550/arxiv.2204.10495,Gigliutti) while in the machine learning literature this problem has been casted as a zero sum game between the adversary (generator) and a discriminator. In Appendix \hyperref[AppA]{A} , I discuss how the estimator can be formulated as the outcome of a sequential, non-zero sum game.

The empirical application contributes to the literature that is related to non-linearities in macro-finance models with long run risk, and the misspecification that arises from linearization techniques, as the latter ignore the level effects of consumption growth on risk premia (see i.e. JOFI:JOFI12615 and references there in.). Information projections are applied to pricing long run risks in aggregate consumption doi:10.1111/j.1540-6261.2004.00670.x, where the base density is generated using the doi:10.1111/j.1540-6261.1988.tb04598.x approximation. Using US data, the paper investigates the empirical importance of non-linearity by re-imposing the non-linear equilibrium condition. The tilted model is strongly preferred by the data. Parameter estimates suggest that the quality of the approximation is yet another reason for the downward bias to estimates of the intertemporal elasticity of substitution and the upward bias in risk aversion, which have been puzzling in the literature, as they imply that asset prices are increasing in uncertainty.

The rest of the paper is organized as follows. Section 2 introduces information projections and provides an asset pricing example. Section 3 presents the econometric properties, computational aspects and supportive simulation evidence. Section 4 applies the methodology to the long run risk model and provides additional simulation evidence. Section 5 concludes. Appendix \hyperref[AppA]{A} contains the main proofs and derivations. Appendix \hyperref[AppB]{B} (intended for online publication) contains supplementary results.

Finally, a word on notation. The set of parameters $\psi$ is decomposed in $\vartheta\in\Theta$, the set of structural (economic) parameters, and $\varphi$ the parameters indexing the density $f(X|Z,\varphi)$. Let $N$ denote the length of the data and $N_{s}$ the length of simulated series. $X$ is an $n_{x}\times1$ vector of the variables of interest while $Z$ is an $n_{z}\times1$ vector of conditioning variables. Both $X$ and $Z$ induce a probability space $(\Omega,\mbox{\ensuremath{\mathcal{F}}},\mathbb{P})$. Three different probability measures are used, the true measure $\mathbb{P}$ (with $\mathbb{P}_{N}$ the corresponding empirical measure), the base measure $F_{\varphi}$ which is indexed by parameters $\varphi$ and the ${H}_{(\varphi,\vartheta)}$ measure which is obtained after the information projection. These measures are absolutely continuous with respect to a dominating measure $v$, where $v$ is the Lebesgue measure, and possess the corresponding density functions $p,f$ and $h$. Conditional measures and densities are denoted by an additional index that specifies the conditioning variable, i.e. $F_{\varphi,z}$. $\mathbb{E}_{P} $ is the mathematical expectations operator with respect to measure $P$; $\mathbb{E}_{P_{z}}$ is shorthand for conditional expectations respectively. $q^{l}(X,Z,\psi)$ is a general $X\otimes Z$ measurable function and ${q}(X,Z,\psi)$ is an $n_{q}\times1$ vector containing these functions. Moreover, $q_{\psi}$ abbreviates the Jacobian matrix of $q$ and $q_{\psi\psi'}$ the Hessian with respect to $\psi$. Unless otherwise stated, $||.||$ signifies the Euclidean norm. For any (matrix) function the subscript $i$ denotes the evaluation at datum $(x_{i},z_{i})$. Subscript $j$ is for simulated data using the base density. The operator $\to_{p}$ signifies convergence in probability and $\to_{d}$ convergence in distribution; $\mathcal{N}(.,.)$ is the Normal distribution. For any $q$, we may abbreviate its mathematical expectation given measure $P$ by $q_{P}$, that is, $q_{P}\equiv \mathbb{E}_{P}q$ while $V_{P,q}$ is the corresponding variance. Conditional variance is denoted by $V_{P_{z},q}$. If $P\equiv\mathbb{P}$ then $V_{\mathbb{P}}\equiv\mathbb{V}$. $V_{ll'}$ signifies the $(l,l')$ component of a matrix $V$.

Tilting the Approximate Model

The main idea can be illustrated as follows. Economic theory implies a set of restrictions:\[\int m(X,Z,\vartheta)d\mathbb{P}(X|Z,\vartheta)=0 \] where $m(X,Z,\theta)$ is a set of conditional moment functions characterized by $\vartheta$, which includes all the economically relevant parameters. The researcher is able to obtain an approximation to $\mathbb{P}(X|Z,\vartheta)$ e.g. the density implied by the solution to the log-linearized model, $f(X|Z,\varphi)$, which by construction does not satisfy the original moment condition: $\int m(X,Z,\vartheta)f(X|Z,\varphi)dX\neq 0$. This implies that matching the approximated model to the data violates the main economic implications of the original specification, and this violation is state (time) and parameter dependent.

This paper proposes to use a conditional information projection, that is, to obtain a density that is as close as possible to the approximate density $f(X|Z,\varphi)$ but satisfies the original moment conditions. This is equivalent to solving the following infinite dimensional optimization problem:\footnote{One could also consider optimizing with respect to $h(X|Z)g(Z)$ as $g(Z)$ is also unknown, yet $h^{\star}$ is the same as the moment restrictions hold for any $Z$, and therefore any $g(Z)$.}

eqnarray[eqnarray omitted — 113 chars of source]
eqnarray*[eqnarray* omitted — 79 chars of source]

where $g(Z)$ is the marginal distribution of $Z$. In the information projections literature minimization (ref) is called exponential tilting it minimizes the Kullback-Leibler ($\mathbf{KL}$) distance, whose convex conjugate has an exponential form. The solution to the above problem, if it exists\footnote{ See Komunjer_Ragusa and section 3 where I relate my assumptions to theirs.}, is given by

equation[equation omitted — 124 chars of source]

where $\psi=(\vartheta',\varphi')'$, $\mu$ is the vector of the Lagrange multiplier functions enforcing the conditional moment conditions on $f(X|Z,\varphi)$ and $\lambda$ is a scaling function. Both are pinned down by:

eqnarray[eqnarray omitted — 257 chars of source]

Given (ref), $\psi$ is estimated using the (limited information) log likelihood function:

eqnarray[eqnarray omitted — 130 chars of source]

Had we used an alternative objective function to (ref), e.g. another particular case from the general family of divergences in cressie_read, this would result to a different form for $h^{\star}(X|Z,\psi)$.\footnote{For brevity, in the rest of the paper I drop the $\star$ notation on $h$.} Under correct specification for $f(X|Z,\varphi)$, this choice does not matter asymptotically, while it matters in finite samples. Exponential tilting nevertheless ensures a positive density function $h$.

Relation to other methods

The projection in the first step (i.e. imposing the conditional moment conditions) is done using a different divergence measure than the estimation objective. A consequence is that the first order conditions of the estimator are mathematically different than a constrained estimator such as the one employed by 1989.\footnote{In other words, the primal problem of maximum tilted likelihood is not equivalent to the primal problem of constrained maximum likelihood estimation. In terms of implementation, tilting replaces the nonlinear constraint on the coefficients of the approximating density with a set of convex optimizations.} Furthermore, the projection results in exponential tilting of the base density, which preserves its properties as a density.\footnote{This is in contrast to using other divergences which might result in the concentrated objective not having a density interpretation such as empirical likelihood (EL), i.e. the EL weights can be negative, do not satisfy the restrictions for arbitrary parameters and do not provide an estimated density 2007ch.} The closest estimation method is that of SusanneSchennah. In the case of unconditional moment restrictions and when the base measure is entirely uninformative, that is $F(X;\varphi)$ is the empirical cdf, the proposed estimator collapses to the ETEL estimator as follows:

eqnarray[eqnarray omitted — 77 chars of source]

with $\hat{w}_{i}=\underset{w_{i}}{\arg\min} \sum_{i=1..N} w_{i}log\left(Nw_{i}\right)$, s.t. $\sum_{i=1..N}w_{i}m(X_{i};\vartheta)=0,\sum_{i=1..N}w_{i}=1$.

SusanneSchennah argues that ETEL behaves favorably as it combines the lower higher order bias of EL and the robustness of ET to moment misspecification, as implied probabilities tend to distribute weights more evenly. In this paper, the tilting in the first step makes sure that $h$ has a density interpretation (similar to EL weights). Moreover, the moment conditions are not satisfied with respect to the approximate model $(f)$. Compared to other divergences, the KL distance $\int h ln(\frac{h}{f})dv$ is expected to perform better under this type of misspecification (similar to the robustness of ET).\footnote{In particular, there is no need to have extreme values of $h$ on sample points at which the moment function $(m)$ is close to zero, because $h$ can be close to zero at points that have large values of $m$, even if $f>0$. This would not be possible with the reverse KL distance, which would require $f$ to be zero when $h=0$, making computation cumbersome.}

Moreover, the parametric nature of the base density employed in this paper allows us to consider a wider class of probability measures that may be derived from structural models, expanding therefore the applicability of information projections to this class of models. As already mentioned, the base density can be generated by an explicit approximate solution to a set of optimality and equilibrium conditions, which could arise from both macroeconomic and microeconomic models.\footnote{When the approximate density is non-structural, then its choice can be informed by utilizing possible knowledge of the reduced form of the structural model and directly using such a form in constructing the base density without explicitly obtaining an approximate solution. In the (log) linearized DSGE case, for example, this corresponds to using a $VAR(p)$ or $VARMA(p,q)$ where $(p,q)$ can increase with the sample size, or a state space model in general. } It is also worth commenting on why tilting using a conditional density (and a conditional moment restriction) might be preferable. In some potential applications, a time varying conditional density is needed to accommodate possible exogenous structural shifts in the economy. Correspondingly, the conditional moment itself can be time varying e.g. when agents display a time preference shift, which also implies endogenous time variation in the underlying density. In order to estimate the parameters governing the structural change jointly with the rest of the parameters, the tilted likelihood function has to be constructed sequentially.

Finally, in the case in which $f(X|Z,\varphi)$ belongs to the exponential family and the moment conditions are linear, exponential tilting is a convenient choice as theoretical conditions imposes additional structure to the moments of the base density. I present below an illustrative example of projecting on densities that satisfy moment conditions that arise from economic theory. Due to linearity, the resulting distribution after the change of measure is conjugate to the base measure.

An Example from Asset Pricing

The consumption - savings decision of the representative household implies an Euler equation restriction on the joint stochastic process of consumption, $C_{t}$, and gross interest rate, $R_{t}$, where $\mathcal{F}_{t}$ is the information set of the agent at time $t$ and $\mathbb{E}_{\mathbb{P}}$ signifies rational expectations :

equation*[equation* omitted — 100 chars of source]

Suppose that the base model is a bivariate VAR for consumption and the interest rate which, for analytical tractability, are not correlated. Their joint density conditional on $\mathcal{F}_{t}$ is therefore: \[ \left(

array[array omitted — 33 chars of source]

\mid\mathcal{F}_{t}\right)\sim N\left(\left(

array[array omitted — 33 chars of source]

\right),\left(

array[array omitted — 30 chars of source]

\right)\right) \] where I have set $\rho_{R}=0$ for analytical convenience.\footnote{Allowing for a non-zero mean in the interest rate is entirely possible but unnecessary. Importantly, the Euler equation has no information on $\mathbb{E}_{t}R_{t+1}$. This is an example of an otherwise testable restriction on $\varphi$ that would not be undone by the projection.} For a quadratic utility function, that is $U(C_{t})=C_{t}-\frac{\gamma}{2} C_{t}^{2}$, the Euler equation implies that $Cov(R_{t+1},C_{t+1}|\mathcal{F}_{t})=\frac{1}{\beta\gamma}(\gamma C_{t}-1)$. The set of structural parameters is $\vartheta:=(\beta,\gamma)$. Since the two variables are now correlated, the distorted density $h(C_{t+1},R_{t+1}|\mathcal{F}_{t})$ will feature a new covariance structure: \[\left(

array[array omitted — 96 chars of source]

\right)\] where $Var_{t}(C_{t+1})=Var_{t}(R_{t+1})=\frac{1}{1-\mu_{t}^2}$, and ${\mu_{t}}$ has the interpretation of the coefficient of projecting consumption on the interest rate, $Var^{-1}_{t}(R_{t+1})Cov(R_{t+1},C_{t+1}|\mathcal{F}_{t})$. In Appendix \hyperref[AppA]{A} I illustrate how the same expression can be obtained formally using a conditional density projection\footnote{More precisely, what is obtained is the density conditional on $Z=z$.}, that is, solving (ref). Notice that the projection is not simply a restriction on the reduced form parameters as it adds information. This is evident from a change in the conditional covariance matrix of $(C_{t+1},R_{t+1})$ from zero to a function of $(\mathcal{F}_{t},\vartheta,\varphi)$. The only feature of the base model that changes after the projection is the one that does not agree with the moment conditions e.g. the covariance matrix. Other features, such as the conditional means, do not change after the projection as they are consistent with the moment condition. Moreover, the fact that in this example the Euler equation is a direct transformation of the parameters of the base density is an artifact of the form of the utility function assumed, and is therefore a special case. In general, an analytical solution cannot be easily obtained and we therefore resort to solving a simulated version of (ref).

Computational Details

$ $ With conditional moment restrictions, the projection involves computing Lagrange multipliers which are functions defined on $\Psi\times Z$. Thereby, the projection has to be implemented at each point $z_{i}$ when the tilted-likelihood function is constructed.

The general algorithm for the inner loop is therefore as follows:\footnote{In estimation, the tilted likelihood function is constructed at every proposal for the vector $\psi$. I experimented with pre-estimating the unknown functions $\mu(Z,\psi)$ and $\lambda(Z,\psi)$ by simulating at different points of the support of $Z$ and $\psi$ and using function approximation methods i.e. splines to compute multipliers at the sample points $z_{i}$ and proposals for parameter vector $\psi$ but I found this to be an inefficient way of implementing the algorithm; it is far better to compute the projection "online".}

enumerate• Given proposal for $(\varphi,\vartheta)$ and datum $z_{i}$, simulate $N_{s}$ observations from $F(x;z_{i},\varphi)$ • Compute: \begin{itemize} • $\hat{\mu}(x;z_{i},\vartheta)=\arg\min\frac{1}{N_{s}}\sum_{j=1:N_s}\exp(\mu(z_{i},\psi)'m(x_{j};z_{i},\vartheta))$ and • $\hat{\lambda}(x;z_{i},\vartheta)=-\log(\frac{1}{N_{s}}\sum_{j=1:N_{s}}\exp(\mu(z_{i},\psi)'m(x_{j};z_{i},\vartheta)))$ \end{itemize} • Evaluate $log\left(h(x,z_{i};\psi)\right)$.

The pseudo log-likelihood function can be constructed using $\sum_{i=1..N}log\left(h(x_{i},z_{i};\psi)\right)$. To give a sense of how much additional time it takes to compute a tilted likelihood function, for every likelihood point, with a single moment condition, computation time increases by an average of $1\%$ of a second in the simulation experiments I will refer to later on in the paper.\footnote{All computing times reported in the paper correspond to an Intel i5-8365U, $1.6$ Ghz processor.}

In the next section I analyze the frequentist properties of using the tilted density to estimate $\psi\equiv(\vartheta,\varphi)$. The main challenge is the fact that we project on a possibly misspecified density. Explicitly acknowledging for estimating the parameters of the density yields some useful insight to the behavior of the estimator. Moreover, although tilting delivers a density, as we will see the information contained in the resulting estimating equations for the structural parameter of interest is of semi-parametric nature and does not always lend itself to proper Bayesian analysis. The tilted density is thus best viewed as a statistical criterion function with an analog interpretation. When there is an explicit mapping between the structural parameters and the parameters of the base density, then the scores of the latter do provide additional information. In this sense the statistical criterion function nests the conventional parametric likelihood function as a special case. Monte Carlo simulation techniques are nevertheless applicable in all cases, as in Chernozhukov2003293 and doi:10.3982/ECTA14525. The frequentist results presented below determine the limiting behavior of this criterion function and are typically assumed to hold for the monte carlo simulation techniques, which can involve prior parameter distributions as well.

Setup and Assumptions

Recall that we maximize the empirical analogue to (ref), which, abstracting from simulation error that comes from computing $(\mu,\lambda)$, is equivalent to choosing $\hat{\psi}$ to solve the following program:

eqnarray*[eqnarray* omitted — 222 chars of source]
eqnarray*[eqnarray* omitted — 240 chars of source]

where for notational brevity I substituted $Z=z_{i}$ for $z_{i}$ and set $\mu_{i}=\mu(z_{i}),\lambda_{i}=\lambda(z_{i})$. Denoting the Jacobian of the moment conditions by ${M}$, the first order conditions are the following, where $\mu_{\psi}(z_{i}),\lambda_{\psi}(z_{i})$ denote derivatives with respect to the corresponding parameter vector:

eqnarray[eqnarray omitted — 346 chars of source]

where $\mathfrak{s}(.)\equiv\frac{\partial}{\partial \varphi}log( f(X|z_{i},\varphi))$ is the score function of the base density and

eqnarray[eqnarray omitted — 237 chars of source]

ASSUMPTIONS I

Assume a stationary ergodic sequence of vectors of random variables $\{x_{i}',z_{i}',\}_{i=1}^{N}$ with summable autocovariances together in an appropriate filtration $\mathcal{F}_{i}$ so that $(x_{i}',z_{i}')$ are measurable with respect to $(\mathcal{F}_{i},\mathcal{F}_{i-1})$ respectively. Furthermore, define innovations $\{\sigma^{l}_{i-1}\epsilon^{l}_{i}\}_{l=1..n_{m}}$ where $\epsilon^{l}_{i}$ a martingale difference sequence, $\sigma^{l}_{i-1}$ is $\mathcal{F}_{i-1}$-measurable and $\mathbb{E}_{P_{z}}\epsilon^{{l}^{2}}_{i}=1$. Moreover:

enumerate[label=\roman*.] • (COMP) $\Theta\subset\mathbb{R}^{n_{\vartheta}},\Phi\subset\mathbb{R}^{n_{\varphi}}$ are compact. • (ID)$\exists ! \psi^{\star}\in int(\Psi):\psi^{\star}=\arg\underset{\psi\in\Psi}{\max}\mathbb{E}\log h(X|Z,\psi)$ • (BD-1a)$\forall l \in {1..n_{m}}\text{ and for \ensuremath{d\geq4}}, P\in\left\{F_{\varphi},\mathbb{P}\right\}: \\ \mathbb{E}_{P_{z}}\sup_{\psi}\|m^{l}(x,z,\vartheta)\|^{d}, \mathbb{E}_{P_{z}}\sup_{\psi}\|m^{l}_{\vartheta}(x,z,\vartheta)\|^{d} \mbox{\text{}} \text{and } \mathbb{E}_{P_{z}}\sup_{\psi}\|m^{l}_{\vartheta\vartheta}(x,z,\vartheta)\|^{d} $ are finite, $\mathbb{P}({z})-a.s$\footnote{Here, the subtlety is that it has to hold for the base measure and the true measure. Given absolute continuity of $d\mathbb{P}(X|Z)$ with respect to $dF(X|Z)$, the existence of moments under $\mathbb{P}(X|Z)$ is sufficient for the existence of moments under $F(X|Z)$.}. • (BD-1b)$\sup_{\psi}\mathbb{E}_{\mathbb{P}_{z}}\|e^{\mu(z)'|m(x,z,\vartheta)|}\|^{2+\delta}<\infty$ for $\delta>0$, $\forall \mu(z)>0, \mathbb{P}({z})-a.s. $ \footnote{ Note that BD-1a and BD-1b imply that $\sup_{\psi}\mathbb{E}_{\mathbb{P}_{z}}\|e^{\mu(z)'m(x,z,\vartheta)+\lambda(z,\vartheta)}m(x,z,\vartheta_{0})\|^{1+\delta}<\infty$ for $\delta>0$ and $\forall z$.} • (\textbf{BD-2})$\mathbb{E}_{\mathbb{P}}sup_{\psi}\mid\log f(x|z,\psi)\mid^{2+\tilde{\delta}}<\infty$ where $\tilde{\delta}>0$. • \textbf{(PD-1) }For any non zero vector $\xi$ and closed $\mathcal{B}_{\delta}(\psi^{\star})$ , $\delta>0$ and for all $P_{z}$ such that $\mathbf{KL}(P_{z},F_{\varphi,{z}})\leq \mathbb{E}_{P_{z}}e^{\mu(z)'m(x,z,\vartheta)}$, the following holds, $\mathbb{P}({z})-a.s$: \\ $0<\inf_{\xi\times\mathcal{B}_{\delta}(\psi^{\star})}\xi'\mathbb{E}_{P_{z}}{m}(x,\vartheta){m}(x,z,\vartheta)'\xi<\sup_{\xi\times\mathcal{B}_{\delta}(\psi^{\star})}\xi'\mathbb{E}_{P_{z}}{m}(x,z,\vartheta){m}(x,\vartheta)'\xi<\infty$

Assumptions (i)-(ii) correspond to typical compactness and identification assumptions found in Newey19942111 where (ii) is assumed to also hold for the objective function in the case in which $h(X|Z)$ does not collapse to $f(X|Z)$ (thus $\psi^{\star}$ is the pseudo-true value\footnote{It is a theoretical possibility that misspecification can push pseudo-true values to the boundary but catering for this in the econometric analysis goes beyond the scope of this paper.}) while (iii) assumes uniform boundedness of conditional moments and their first and second derivatives, up to a set of measure zero. Assumption (iv) assumes existence of exponential absolute $2+\delta$ moments and (v) assumes the existence of $2+\delta$ absolute moments of the log-likelihood based on the base density.\footnote{Since $log(h(x|z,\psi))=\mu(z,\psi)'m(x,z,\vartheta)+\lambda(z,\psi)+\log f(x|z,\psi)$, existence of $\mu(z,\psi)$ and $\lambda(z,\psi)$, together with BD-1a,b and BD-2 imply that $\mathbb{E}_{\mathbb{P}}\sup_{\psi}\mid\log h(x|z,\psi)\mid^{2+\tilde{\delta}}<\infty$.} Finally, (vi) assumes away pathological cases of perfect correlation between moment conditions under any measure $P_{z}$ that is contiguous and within a certain distance to $F_{\varphi,{z}}$.

Existence and feasibility of the projection at $z_{i}$

Komunjer_Ragusa provide primitive conditions for the case of projecting using a divergence that belongs to the $\phi-$ divergence class and moment restrictions that have unbounded moment functions. Assumptions BD-1a and BD-1b are sufficient for their primitive conditions (Theorem 3). More particularly, the feasibility assumption of Komunjer_Ragusa i.e. there exists at least one measure $H\in \mathcal{H}$ such that $\int log\left(\frac{dH}{dF}\right)dH<\infty$ is trivially satisfied by $H=\mathbb{P}$, the true probability measure, as long as the Kullback Leibler distance of the approximate measure $F$ to the true measure $\mathbb{P}$ is well defined. In this case, the Kullback - Leibler distance minimized at each projection step becomes as follows: \[\int \log\left(\frac{dH(.;z_{i},\vartheta)}{dF(.;z_{i},\varphi)}\right)dH(.;z_{i},\vartheta) = \int \log\left(\frac{d\mathbb{P}(.;z_{i},\vartheta)}{dF(.;z_{i},\varphi)}\right)d\mathbb{P}_{z_{i}}\] This distance will be finite as long as the base density has a finite KL distance with respect to $\mathbb{P}(.)$ for some $\varphi\in \Phi$. If $F(.;z_{i},\varphi)$ is induced by a structural approximation, the distance will be finite if the approximation error is uniformly controlled over $z$. To see why, consider the density $f(x_{t};z_{t},\hat{\gamma}(z;\vartheta),\vartheta)$ induced by a set of approximations $\hat{\mathbf{\gamma}}(z;\varphi)$ to functions $\gamma(z)$, such as policy functions or value functions. The latter induces the true density, ${p}(x_{t};z_{t},{\gamma}(z),\vartheta)$. Then, using a functional Taylor expansion

eqnarray*[eqnarray* omitted — 281 chars of source]

Hence, if the (log) density $f$ is $K$ times pathwise differentiable and the approximation method controls the error uniformly in $z_{i}$ i.e. $\sup_{z}||\hat{\gamma}-\gamma||<\infty$, then all of the terms on the RHS are bounded. Existing work has established uniform upper bounds on errors in value function iteration 10.2307/2998564 and policy function iteration doi:10.1137/S0363012902399824, as well as bounds on invariant distributions of simulations of approximated models RePEc:ecm:emetrp:v:73:y:2005:i:6:p:1939-1976.

In practice, at each parameter proposal by the estimation algorithm it is easy to check whether the projection exists, as the Lagrange multiplier vector $\mu$ takes implausibly high absolute values, and hence the parameter draw is rejected. To see this, notice that in (ref) if $m(X,Z,\vartheta)$ is uniformly positive or negative across $X$ for any $Z$, which implies that the moment condition cannot be satisfied, then optimal $\mu$ is either positive or negative infinity.

Asymptotic Properties

This section contains the consistency and asymptotic distribution results for $\hat{\psi}$. Appendix \hyperref[AppA]{A} ((ref)) provides expressions for the first and second order derivatives of $(\mu(z_i),\lambda(z_i))$ which determine the behavior of $\hat\psi$ in the neighborhood of $\psi^{\star}_0$. These expressions will be useful for the characterization of the properties of the estimator. I first outline a result which is useful in understanding the properties of the estimator.

LemmaFor a base conditional density $F_{z}$ with distance to $\mathbb{P}_{z_{i}}$ equal to $\chi^{2}(F_{z_{i}},\mathbb{P}_{z_{i}})$: \\ (a) $\mu_{i}=O_{p_{z}}(\chi^{2}(F_{z_{i}},\mathbb{P}_{z_{i}})^{\frac{1}{2}})$\\ (b) $ \max_{i}\sup_{\vartheta}|\mu_{i}'m(\vartheta,x_{i})|=O_{p_{z}}(\max_{i}\chi^{2}(F_{z_{i}},\mathbb{P}_{z_{i}})^{\frac{1}{2}}N^{\frac{1}{d}}) \label{mu} $
proofSee Appendix \hyperref[Proofs]{A}.

When $\sup_{i}\chi^{2}(F_{z_{i}},\mathbb{P}_{z_{i}}) \sim N^{-\xi}$, $\mu_{i}=o_{p_{z}}(1)$. If $\frac{2}{d}<\xi$, $\max_{i}\sup_{\vartheta}|\mu_{i}'m(\vartheta,x_{i})|=o_{p}(1)$. The Lagrange multipliers from imposing the equilibrium conditions will converge to zero under vanishing approximation error.

Consistency

The uniform consistency of the estimator is shown by first proving pointwise consistency and then stochastic equicontinuity of the objective function. Details of the proof are in Appendix \hyperref[AppA]{A}. Under fixed misspecification, the estimator is consistent for $\psi_{0}^{\star}$, defined below.

DefinitionThe true value is denoted by $\psi_{0}$. The pseudo-true value $\psi_{0}^{\star}$ is the $\psi\in\Psi$ that minimizes the Kullback-Leibler $(\mathbf{KL})$ distance between $H(X,Z,\psi)$ and $\mathbb{P}(X,Z)$, decomposed as follows: \begin{eqnarray} 0&\leq& \mathbb{E}_{\mathbb{P}}\log\left(\frac{d\mathbb{P}_{z}}{dF_{(\varphi,z)}}\right) - \mathbb{E}_{\mathbb{P}}\left[\mu'({\psi,z})\mathbb{E}_{\mathbb{P}_{z}}m(X,\vartheta)\right]+\mathbb{E}_{\mathbb{P}}\left[\log\left(\mathbb{E}_{F_{(\varphi,z)}}\exp(\mu'({\psi,z})m(X,\vartheta))\right)\right] \end{eqnarray} Correspondingly, since $\mu({\psi,z}):=\underset{\mu}{\arg\min}\left(\mathbb{E}_{F_{(\varphi,z)}}\exp(\mu'm(X,\vartheta))\right)$, $\varphi^{\star}_{0}$ is the value of $\varphi\in\Phi$ such that $F(X|Z,\varphi)$ is as close as possible (in $\mathbf{KL}$ units) to $\mathbb{P}(X|Z,\varphi)$ {and} satisfies \[\mathbb{E}_{F_{(\varphi,z)}}\exp(\mu'({\psi,z})m(X,\vartheta^{\star}_{0}))m(X,\vartheta^{\star}_{0})=0\] where $\vartheta^{\star}_{0}$ is the value of $\vartheta\in\Theta$ such that $\mathbb{E}_{F_{(\varphi^{\star},z)}}\exp(\mu'({\psi,z})m(X,\vartheta))m(X,\vartheta)=0 \label{def_psistar}$.

At the pseudo-true value, $H_{(\psi_{0}^{\star},z)}$ is the closest parametric distribution to $\mathbb{P}_{z}$, while both distributions satisfy a common moment restriction, $\mathbb{E}_{H_{(\psi^{\star},z)}}m(X,\vartheta^{\star})=\mathbb{E}_{\mathbb{P}}m(X,\vartheta_{0})=0$. The tilted distribution will satisfy the exact moment conditions, and will be -by construction- closer to the distribution implied by the economic model. The smaller $\mathbf{KL}(F,\mathbb{P})$ is, the closer to zero are the last two terms in (ref).\footnote{Note that there is a similarity between our definition of $\psi_{0}^{\star}$ to the definition of 10.2307/2646780, but in our case the moment restriction is satisfied by the sampled population, asymptotically.} The interpretation of $\psi_{0}^{\star}$ can be qualified further by examining the profiled $\mathbf{KL}$ distance.

LemmaThe profiled (over $\mu_{z}$) distance $\mathbf{KL}(H_{\psi},\mathbb{P})$ in (ref) is approximately equivalent to \begin{eqnarray*} \mathbb{E}_{\mathbb{P}}\log\left(\frac{d\mathbb{P}_{z}}{dF_{(\varphi,z)}}\right)+ \mathbb{E}_{\mathbb{P}}\left[\mathbb{E}_{F_{(\varphi,z)}}m(X,\vartheta)V^{-1}_{F_{(\varphi,z)}}\left(\mathbb{E}_{\mathbb{P}_{z}}m(X,\vartheta)-\mathbb{E}_{F_{(\varphi,z)}}m(X,\vartheta) \right)\right] \end{eqnarray*}
proofFollows from (ref), the implicit map of $\mu$ in the proof of (ref) and a first order approximation of the third term in (ref).

If $(\varphi,\vartheta)$ are independent, then minimizing this criterion is equivalent to finding the value of $\varphi$ that makes the approximate model as close as possible to the true data generating process for any level of $\vartheta$, which simultaneously minimizes both the first and second term, while for a fixed value of $\varphi$, minimizing over $\vartheta$ seeks to bring $\mathbb{E}_{F_{(\varphi,z)}}m(X,\vartheta)$ as close as possible to zero, that is, the approximate model satisfying the moment condition. A similar interpretation holds when there is a mapping from the reduced form to the structural parameters, which arises when we consider approximate equilibrium models. As mentioned in the introduction (see (ref)) and explicitly shown later in the paper, the first order condition will be equivalent to setting to zero the orthogonalized combinations of the scores of the approximate model and the efficient GMM first order condition.

Having defined the pseudotrue values, I present below the asymptotic results.

theoremConsistency for $\psi_{0}^{\star}$\\ Under Assumptions I : $ (\hat{\vartheta},\hat{\varphi})\underset{p}{\rightarrow}(\vartheta_{0}^{\star},\varphi_{0}^{\star}) \label{consstar} $
proofSee Appendix \hyperref[Proofs]{A}.
Definition{Vanishing Approximation Error.}$ $\\ Assuming that densities are twice differentiable, define the metric \[\Delta(F,P)=\sup_{k\in\{0,1,2\}}\mathbb{E}_{{P}}\left( D^{k}\left(\frac{f}{p}-1\right)\right)^{2}\] which measures how well density $f$ approximates $p$ up to the second derivative.\footnote{Define $D^{1}:=\sup_{l}\mid\frac{\partial (.)}{\partial x_{l}}\mid$ and $D^{2}:=\sup_{ll'}\mid\frac{\partial^{2} (.)}{\partial x_{l}\partial x_{l'}}\mid$. Uniform convergence of $\Delta$ to zero implies that the approximation error vanishes up to the second derivative, which ensures uniform convergence of the scores and the Hessian (elementwise). Note that for $k=0$, $\mathbb{E}_{{P}}\left( D^{0}\left(\frac{f}{p}-1\right)\right)^{2}$ is the $\chi^{2}$ distance between $f$ and $p$. } An asymptotically correct density is defined as the density $f(X|Z,\varphi)$ that converges uniformly to the true conditional density, such that $\sup_{\varphi}\sup_{i}\Delta(F_{z_{i}},\mathbb{P}_{z_{i}})\to 0$.

Therefore, under vanishing approximation error, consistency is for $\vartheta_0$:

CorollaryIf $\sup_{\phi}\sup_{i}\Delta(F_{z_{i}},\mathbb{P}_{z_{i}})\to 0$, then $\vartheta_{0}^{\star}=\vartheta_0$.
proofSee Appendix \hyperref[Proofs]{A}.

Asymptotic Distribution

The limiting distribution of the estimator is derived using a first order approximation around the pseudotrue parameter $\psi_{0}^{\star}$. The results that follow hold for a generic base density $f(X|Z,\varphi)$, where I distinguish between non-structural approximations where $\varphi$ and $\vartheta$ are not functionally related, and structural approximations, where the two are related. Despite that the main focus in the case of a structural approximation, in each set of results it is instructive first to deal with the case of a non-structural approximation.

Non-Structural Approximations

The estimator for $\vartheta$ that results from a non-structural approximation of the density converges to the same distribution as estimators employed in the semi-parametric literature, where the density is treated as a nuisance parameter, estimated using a (semi)non-parametric approximation. The next set of results demonstrate asymptotic Normality while semi-parametric efficiency for $\vartheta$ is retained if $N^{-\frac{1}{2}}\Delta\to 0$.

theoremAsymptotic Distribution when $\Psi:=(\Phi,\Theta)\in \mathbb{R}^{dim \varphi}\times \mathbb{R}^{dim \vartheta}$\\ Under Assumption I and $N_{s}$,$N\rightarrow\infty$ such that $\frac{N}{N_{s}}{\to} 0$ \begin{eqnarray*} N^{\frac{1}{2}}(\psi-\psi_{0}^\star)\underset{d}{\rightarrow} \mathcal{N}\left(0,\bar{V}(\psi_0^\star)\right) \end{eqnarray*}
proofSee proof of Theorem (ref) in Appendix \hyperref[Proofs]{A} .

The assumption $\frac{N}{N_{s}}{\to} 0$ obviates the need to account for simulation error in the variance covariance matrix. This can be relaxed i.e. it can converge to a positive number but it adds no new insights. In Appendix \hyperref[AppA]{A} I derive the exact form of the covariance matrix of the estimator. Under (asymptotic) correct specification $\sup_{\varphi}\sup_{i}\Delta(F_{z_{i}},\mathbb{P}_{z_{i}})\equiv \kappa^{-1}_{n}\sim N^{-\xi}$, $\mathfrak{s}_{i}\to s_{i}$ where $s_{i}$ is the score of the true density. For $\xi>1$, $ G(\psi)\equiv -\mathbb{V}_{g}(\psi)$, $\bar{V}(\psi_0^\star)=\bar{V}(\psi_0)$. In particular, the variance is block diagonal (I drop dependence on $(\vartheta,\varphi)$):

eqnarray*[eqnarray* omitted — 302 chars of source]

where $\mathfrak{B}_{i}'=\mathbb{V}^{-1}_{m,i}\mathbb{E}(m{s}'|z_{i})$ is the coefficient of projecting the scores on the moment conditions, where I abbreviate $V_{\mathbb{P}_{z},m}\equiv\mathbb{V}_{m,i}$. The upper left component of $\bar{V}(\psi_0)$ is the same as the (inverse) information matrix corresponding to $\vartheta$ when the conventional optimally weighted GMM criterion is employed. What this implies is that lack of knowledge of $\varphi$ does not affect the efficiency of estimating the structural parameters $\vartheta$, at least asymptotically. Also, note that $\hat\vartheta$ attains the SLB even if $\xi>\frac{1}{2}$, as the off-diagonal block of the Jacobian (cross-derivatives) converges to zero as long as $\xi>0$.

The expressions above have an intuitive interpretation. If the moment conditions used span the same space spanned by the scores of the density, then $\bar{V}$ trivially attains the Cramer - Rao bound in the upper left component as $m\equiv \mathfrak{s}$ and the covariance matrix becomes singular as both $m$ and $\mathfrak{s}$ give the same information. Conversely, the less predictable is the score from the additional moment conditions used (that is, $\|\mathfrak{B}_{i}\|$ is close to zero), the higher the efficiency attained for estimating $\varphi$, where $\left(\mathbb{E}{s}_{i}(\varphi){s}_{i}(\varphi)'\right)^{-1}$, is the lowest variance possible among regular estimators. This also corroborates the claim that the estimator is not equivalent to a constrained MLE as in the latter case, constraints typically increase efficiency for estimating the reduced form. In this case, by using the moment conditions which involve unknown parameters ($\vartheta$) can only increase the variance for estimating $\varphi$ as the former are not essential to the identification to the latter when $(\varphi,\vartheta)$ are independent.\footnote{ The best analogy can be made to including irrelevant regressors in multivariate regression: as long as the additional regressors are not essential to the identification of the causal effect, the result is an increase in the variance of its estimate.}

Structural Approximations

If the base density is obtained by solving a structural model, there is an explicit mapping between $\varphi$ and $\vartheta$, and the first order conditions are a linear combination of the conditions analyzed above as $ \frac{dQ_{N}}{d\vartheta} = \frac{\partial Q_{N}}{\partial \vartheta} + \frac{\partial Q_{N}}{\partial \varphi}\frac{d \varphi}{d \vartheta} =0 $. When the solution is exact, $f(X|Z)$ is pinned down by a unique parametric sub-model that automatically satisfies the moment conditions and $(\mu_{i},\lambda_{i})$ are zero for all $i$, which implies that both $\frac{\partial Q_{N}}{\partial \vartheta}(\hat{\vartheta}) $ and $\frac{\partial Q_{N}}{\partial \vartheta}(\hat{\vartheta}) $ will tend to zero as $\mathbb{E}M_{i}'V^{-1}_{{i},m}m_{i}=0$ and $\mathbb{E}{s}_{i}=0$ at $\vartheta_{0}$, while the second term in the variance of $\varphi$ vanishes as well. Moreover, $\mathbb{E}{s}_{i}(\varphi){s}_{i}(\varphi)'=\mathbb{E}\frac{d\vartheta}{d\varphi}'\mathbb{E}{s}_{i}(\vartheta){s}_{i}(\vartheta)'\frac{d\vartheta}{d\varphi}>\mathbb{E}\frac{d\vartheta}{d\varphi}'\mathbb{E}({M}|z_{i})'V^{-1}_{m,i}\mathbb{E}({M}|z_{i})\frac{d\vartheta}{d\varphi}$. The moment restrictions are trivial and add nothing to the information embedded in the scores:

theoremAsymptotic Distribution when $\Upsilon \in \mathbb{C}^{2} : \Theta \to \Phi$. \\ Under Assumption I and $N_{s}$,$N\rightarrow\infty$ such that $\frac{N}{N_{s}}{\to} 0$ : \begin{eqnarray*} N^{\frac{1}{2}}(\vartheta-\vartheta_{0}^\star)\underset{d}{\rightarrow} \mathcal{N}\left(0,\bar{V}(\vartheta_0^\star)\right) \end{eqnarray*} where for $\sup_{\varphi}\sup_{i}\Delta(F_{z_{i}},\mathbb{P}_{z_{i}})\equiv \kappa^{-1}_{N}\sim N^{-\xi}$ with $\xi>1$, $\bar{V}(\vartheta_0)^{-1} \equiv \left(\frac{d\varphi}{d\vartheta}\right)'\mathbb{E}{s}_{i}{s}_{i}'\frac{d\varphi}{d\vartheta}$.
proofSee proof of Theorem (ref) in Appendix \hyperref[Proofs]{A} .

Rate(s) of Convergence

With a non-structural base density, the discrepancy with the true model does not have first order effects for $\hat{\vartheta}$ if $\xi>\frac{1}{2}$, and for $\hat{\varphi}$ if $\xi>1$.\footnote{ Geweke_Hahn_Ackerberg obtains a similar type of result for approximated likelihoods (corresponding to $\xi>1$). Note that $\xi=1$ corresponds to the case when the approximate density is defined by parameters that are within a root-n neighborhood to the true density, and hence $\Delta(F_{z_{i}},\mathbb{P}_{z_{i}})=O_{p_{z}}(N^{-1})$.} As shown in the proof of Theorem (ref),

eqnarray*[eqnarray* omitted — 277 chars of source]

When $\xi<\frac{1}{2}$, the first term in $g_{1,N}$ converges to a normal random vector while the last term diverges. Hence, the whole term diverges. Similarly, when $\xi<1$, the whole term $g_{2,N}$ diverges. To have a meaningful limit in this case, we have to normalize with $N^{\xi}$ and $N^{\frac{\xi}{2}}$ respectively:

eqnarray*[eqnarray* omitted — 275 chars of source]

and therefore the rate of convergence for parameter estimates is slower than root-n. It is important to note that $\xi>\frac{1}{2}$ allows for a wide range of cases for local-misspecification, with approximate densities converging at a much slower rate than the standard parametric rate $(\xi=1)$. This is more interesting in the case of structural approximation: The information added by tilting the model to satisfy the moment condition will indeed be driven by the efficient GMM first order condition and will not depend to first order on the degree of local misspecification, even if $\xi\leq 1$: \[N^{\frac{\xi}{2}}\left(g_{1,N}+\frac{d\varphi}{d\vartheta}'g_{2,N}\right)=N^{\frac{\xi}{2}-1}\sum_{i} M_{f_{i}}'V^{-1}_{f_{i},m}m_{i} + O_{p}(N^{-\frac{\xi}{2}})+ \frac{d\varphi}{d\vartheta}'N^{\frac{\xi}{2}-1}\underset{i}{\sum}\left({s}_{i} -{s}_{f_{i}}+\mu_{i,\varphi}'m_{i}\right))+ O_{p}(1)\]

Further Properties under Local Misspecification

The previous analysis considered the asymptotic distribution under a fixed data generating process. I next analyze the asymptotic distribution under drifting parameter sequences, which is more useful to understand the behavior of the estimator under standard local misspecification $(\xi=1)$. In particular, I consider approximation errors that can be represented by non-linear restrictions imposed by the approximate model on the true data generating process as follows: $r(\varphi)=0$, where $r$ is some non-linear function of the reduced form parameters.

One example of restrictions, which is also the one I will consider in the application, is ignoring non-linear effects in dynamic forward looking models. For illustration, consider two popular methods of approximating the solutions to dynamic equilibrium models, perturbation and projection. Suppose that $dim(X)=1$. In the case of perturbation around $\bar{X}$, the Taylor approximation of a policy function results in $C(X,J;\vartheta) = C(\bar{X};\vartheta) + C_{1}(\bar{X};\vartheta)(X-\bar{X}) +...+ O(\|X-\bar{X}\|^{J+1})$ while in the projection case, the approximation is of the form $C(X,J,\alpha,\beta;\vartheta)=\sum^{J}_{j=1..J}\alpha_{j}b_{j}(X,\beta)$ where $b_{j}(X,b)$ is a basis function and $\alpha$ are the projection coefficients. Assuming there exist finite $({J}^{\star},b^{\star},\alpha^{\star})$ such that approximation errors are negligible, restricting the number of approximating terms or the functional form of the basis $b_{j}(X,b)$ restricts the moments of $C$ (and hence $\varphi$) as $\mathbb{E}C^{l}(X,J,\beta;\vartheta)$ is a function of $(J,\beta)$.

Adopting the local asymptotic experiment approach, see for example BHansen_2016, I investigate convergence in distribution along sequences $\psi_N$ where $\psi_N= \psi^{\star}_0 + c N^{-\frac{1}{2}}$ for $\psi_N$ the true value, $\psi^{\star}_0\in\Psi$ the centering value and $c$ the localizing parameter. The true parameter is therefore ”close” to the restricted parameter space up to arbitrary $c$, and hence misspecification is arbitrary.

theoremAsymptotic Distribution when $\Psi:=(\Phi,\Theta)\in \mathbb{R}^{dim \varphi}\times \mathbb{R}^{dim \vartheta}$\\ For $R(\varphi)\equiv \frac{\partial}{\partial \varphi}r(\varphi)$, $\bar{G}^{-1}\equiv\left(\begin{array}{cc} G^{11} & G^{12}\\ G^{21} & G^{22} \end{array}\right)$, $S_1\equiv[I_{n_\vartheta},0_{n_\vartheta\times n_\varphi}]$, $S_2\equiv[ 0_{n_\varphi\times n_\vartheta}, I_{n_\varphi}]$ Under assumptions I such that $ N^{\frac{1}{2}}\bar{G}(\tilde\psi)^{-1}g(\psi_N)\underset{d}{\to} \mathcal{Z}\sim \mathcal{N}(0,\Omega)$: \begin{enumerate} • $N^{\frac{1}{2}}(\hat\vartheta-\vartheta_N)\underset{d}{\to} S_1\mathcal{Z}$$N^{\frac{1}{2}}(\hat\varphi-\varphi_N)\underset{d}{\to} \mathcal{Z}_r \equiv S_2\mathcal{Z}-G^{22}(\psi^{\star}_0)R(\varphi^{\star}_0)(R(\varphi^{\star}_0)'G^{22}(\psi^{\star}_0)R(\varphi^{\star}_0))^{-1}R(\varphi^{\star}_0)'(S_2(\mathcal{Z}+c))$ • For any non zero vector $\xi$, $\xi'(\mathbb{V}(S_2\mathcal{Z})-\mathbb{V}(\mathcal{Z}_r))\xi \geq 0$ \end{enumerate}
proofSee proof of Theorem (ref) in Appendix \hyperref[Proofs]{A} .

There are two main implications of (ref) for $\hat\vartheta$. First, for $c\neq0$, the asymptotic distribution of $\hat{\varphi}$ is non regular i.e. the distribution depends on $c$ (see p. 115 in VDV). Second, the variance of $\hat\varphi$ is lower than the Cramer-Rao bound for regular estimators. For $\varphi_N$ arbitrarily close to the restricted subspace of $\varphi_0$, efficiency increases. The distribution of $\hat\vartheta$ is regular and its variance reaches the SLB. With a structural approximation, $\hat\vartheta$ inherits the bias-variance trade-off due to the restrictions imposed by the approximation. The next result clarifies how this trade-off arises.

LemmaEstimating equation when $\Upsilon \in \mathbb{C}^{2} : \Theta \to \Phi$. \\ Since $\xi=1$, the estimating equation is equal to \begin{eqnarray}\frac{1}{N}\sum^{N}_{i=1}\left({M}_{f_{i}}'V^{-1}_{f_{i},m}m_{i}\right)+ \frac{d\varphi}{d\vartheta}'\frac{1}{N}\sum^{N}_{i=1}\left(\mathfrak{s}_{i}-\mathbb{E}_{h}(\mathfrak{s}|z_{i})- \mathbb{E}_{h}(\mathfrak{s}m'|z_{i})V^{-1}_{h_{i},m}m_{i}\right)=O_{p}(N^{-\frac{1}{2}}) \end{eqnarray} which in the limit is equivalent to \begin{eqnarray}\mathbb{E}_{\mathbb{P}}\left({M}_{i}'V^{-1}_{m_{i}}\mathbb{E}_{\mathbb{P}}(m|z_{i})+ \frac{d\varphi}{d\vartheta}'\left({s}_{i}-\mathbb{E}_{\mathbb{P}}({s}|z_{i})- \mathbb{E}_{\mathbb{P}}({s}m'|z_{i})V^{-1}_{m_{i}}\mathbb{E}_{\mathbb{P}}(m|z_{i})\right)\right)=0 \end{eqnarray}
proofThe result is a direct application of Lemmas (ref) in Appendix \hyperref[AppA]{A} and (ref) in Appendix \hyperref[AppB]{B} to the first order conditions in the proof of Theorem (ref).

The condition in Lemma (ref) can be read in two ways: The first is that in the absence of the projection, the corresponding estimating equation would simply be $\frac{d\varphi}{d\vartheta}'\mathbb{E}_{\mathbb{P}}\mathbb{E}_{\mathbb{P}}({s}|z_{i})=0$. The projection adds the moment conditions that are key to restoring the information lost by the approximation that is relevant for identification. Conversely, adding the score based on the approximate model as an additional moment condition can improve the efficiency of estimating $\vartheta$ using only $\mathbb{E}_{\mathbb{P}}(m|z_{i})=0$. Nevertheless, in the limit, it is the variance of the approximate score that determines efficiency.

theoremAsymptotic Distribution when $\Upsilon \in \mathbb{C}^{2} : \Theta \to \Phi$. \\ For $R(\varphi)\equiv \frac{\partial}{\partial \varphi}r(\varphi)$, $\tilde{R}\equiv \left(\frac{d\vartheta}{d\varphi}\right)'R$. Under assumptions I such that $ N^{\frac{1}{2}}\check{\bar{G}}(\tilde\psi)^{-1}\check{g}(\psi_N)\underset{d}{\to} \mathcal{Z}\sim \mathcal{N}(0,\Omega)$: \begin{enumerate} • $N^{\frac{1}{2}}(\hat\vartheta-\vartheta_N)\underset{d}{\to} \mathcal{Z}_r\equiv \mathcal{Z}-\check{G}(\psi^{\star}_0)\tilde{R}(\varphi^{\star}_0)(\tilde{R}(\varphi^{\star}_0)'\check{G}(\psi^{\star}_0)\tilde{R}(\varphi^{\star}_0))^{-1}\tilde{R}(\varphi^{\star}_0)'(\mathcal{Z}+\check{c})$ • For any non zero vector $\xi$, $\xi'(\mathbb{V}(\mathcal{Z})-\mathbb{V}(\mathcal{Z}_r))\xi \geq 0$ \end{enumerate}
proofSee proof of Theorem (ref) in Appendix \hyperref[Proofs]{A} .

The main insight from Theorem (ref) is that under local misspecification with $\xi=1$, the efficiency gain arising from employing the tilted approximate model is the same as the efficiency gain obtained by employing the approximate model. Nevertheless, note that the localizing vector $\check{c}$ is not necessarily the same as ${c}$ in Theorem (ref), as the bias is different. This distinction will be important for the asymptotic risk of the estimator based on the tilted model and will be explored further in Theorem (ref).

The next set of results characterize the estimator performance in terms of Mean Squared Error and asymptotic risk. I first introduce a loss function and the corresponding definition of risk.

Asymptotic Risk

DefinitionFor any estimator $\hat{\vartheta}_{N}$ of a parameter $\vartheta_{N}$, let $l(\vartheta_{N},\hat{\vartheta}_{N})$ be the loss function that characterizes the accuracy of the estimator. The risk of the estimator is the expected loss $\mathbb{E}_{\vartheta_{N}}l(\vartheta_{N},\hat{\vartheta}_{N})$.

ASSUMPTIONS II

enumerate• Locally Quadratic Loss: $l(\vartheta,\hat{\vartheta})\geq 0$, $l(\vartheta,\vartheta)= 0$ and $W(\vartheta):=\frac{1}{2}l_{\vartheta,\vartheta'}(\vartheta,\hat{\vartheta})\mid_{\hat{\vartheta} ={\vartheta} }$ continuous in a neighborhood of $\vartheta_{0}$
DefinitionDefine $\rho(c,\hat{\vartheta})$ as the asymptotic risk of the estimator sequence $\left\{\hat{\vartheta}_{N}\right\}_{N=1,2..\infty}$ for parameter sequence $\left\{\vartheta_{N}\right\}_{N=1,2..\infty}$:\[\rho(c,\hat{\vartheta}):=\lim_{\zeta\to\infty}\liminf_{N\to\infty}\mathbb{E}_{\vartheta_{N}}\min[N l(\vartheta_{N},\hat{\vartheta}_{N}),\zeta]\] where $\zeta$ is an arbitrary trimming level that is asymptotically negligible.

The following lemma enables us to compute the asymptotic risk given a smooth loss function.

Lemma(BHansen_2016, Lemma 1) For an estimator $\hat{\vartheta}_{N}$ such that $N^{\frac{1}{2}}\hat{\vartheta}_{N}\underset{\vartheta_{N}}{\to}\mathbb{Z}$, if $\vartheta_{N}\underset{N\to\infty}{\to}\vartheta_{0}$, then for any loss function that satisfies Assumption II, $\rho(c,\hat{\vartheta})=\mathbb{E}(\mathbb{Z}'W \mathbb{Z})$.

I next characterize the Mean Squared Error of the proposed estimator in a local asymptotic experiment under a stronger set of conditions, while armed with (ref), we can deduce the corresponding result for asymptotic risk.

Structural Approximation

theoremMean Squared Error and Asymptotic Risk when $\Upsilon \in \mathbb{C}^{2} : \Theta \to \Phi$\\ For $D:=\tilde{R}'\check{G}\tilde{R}$, \begin{enumerate} • $ \mathbf{MSE}(\hat{\vartheta}) = \check{G}\tilde{R}'D^{-1}\tilde{R}\check{c}\check{c}'\tilde{R}'D^{-1}\tilde{R}'\check{G}'+(I-\check{G}\tilde{R}D^{-1}\tilde{R}')\Omega(I-\check{G}\tilde{R}D^{-1}\tilde{R}')' $ • If in addition to (1), $\xi'(\tilde{R}D^{-1}\tilde{R}'\check{c}\check{c}'- I)\xi \leq 0$ then \begin{itemize} • For any non zero vector $\xi$ and $\mathbf{MSE}(\vartheta^{MLE})=\left(\left(\frac{d\varphi}{d\vartheta}\right)'\mathbb{E}{s}_{i}{s}_{i}'\left(\frac{d\varphi}{d\vartheta}\right)\right)^{-1}$, \[\xi'\left(\mathbf{MSE}(\vartheta^{MLE})-\mathbf{MSE}(\hat{\vartheta})\right)\xi\geq0\quad \textit{and }\rho(\check{c},\hat{\vartheta})<\rho({\vartheta}^{MLE})\] \end{itemize} \end{enumerate}
proofSee proof of Theorem (ref) in Appendix \hyperref[Proofs]{A} .

The efficiency gains in Theorem (ref) can potentially translate to lower MSE if condition (2) in Theorem (ref) is met. The reason why this condition is much more likely to be met by the tilted model as compared to the approximate model is the improved approximation delivered by tilting. I next illustrate that within the class of approximated equilibrium models, the information projection alleviates, at least partially, the misspecification caused by the approximation evaluated at $\vartheta_{0}$.

Improved Approximation

Proposition\[\mathbb{E}_{\mathbb{P}_{z}}\log\left(\frac{d\mathbb{P}_{z}}{dH_{z}(\vartheta_{0})}\right)<\mathbb{E}_{\mathbb{P}_{z}}\log\left(\frac{d\mathbb{P}_{z}}{dF_{z}(\vartheta_{0})}\right)\]
proofSee Appendix \hyperref[AppA]{A}

(ref) implies that if one obtains an approximate solution, tilting the density to satisfy the non-linear conditions results in a more accurate approximation (in KL units). This is yet another important property of this method, since if one cannot obtain an accurate approximation due to computational and time limits, tilting will increase accuracy. A caveat of the this result is that it holds at the true parameter value. Nevertheless, more can be said about the use of information projections at any $\vartheta$. The next result explicitly shows that adding information by tilting the model reduces the distance to the true value, for any degree of local missspecification in the non-tilted model.

theoremConsider a sequence of dgp's $\vartheta_N= \vartheta_0 + c_{0}N^{-\frac{1}{2}}$ with $\hat{\vartheta}_0\overset{p}{\to}\vartheta_0$ the restricted MLE estimator and $c_{0}$ the localizing vector. Similarly, define $\vartheta_{1}$ as the probability limit of the tilted PMLE estimator. Then, there exists a vector $c_{1}:=N^{\frac{1}{2}}(\vartheta_N-\vartheta_1)$ such that for some $\eta>0$, \begin{eqnarray}\mid\mid\vartheta_{1}-\vartheta_{N}\mid\mid^{2}=\mid\mid\vartheta_{0}-\vartheta_{N}\mid\mid^{2}-\frac{\eta}{N}+O\left(\frac{1}{N^{\frac{3}{2}}}\right) \end{eqnarray}
proofSee proof of Theorem (ref) in Appendix \hyperref[Proofs]{A} .

In Theorem (ref), the true parameter sequence is defined to be within a $\check{c}N^{-\frac{1}{2}}$ of the restricted space, which is the manifold consistent with the restrictions implied by the approximate model and the moment restrictions. In Theorem (ref), the true parameter sequence is defined to be within a $c_{0}N^{-\frac{1}{2}}$ of the restricted space consistent with the approximate model without further information. Tilting the model transforms the restricted parameter space such that it is closer to the true parameter sequence. The next example explicitly shows how the MSE is reduced due to lower bias.

Analytical Example using the Stochastic Growth model

Assuming a logarithmic utility function for the representative agent and a Cobb-Douglas production function, the stochastic growth model is characterized by the following conditions:

eqnarray*[eqnarray* omitted — 295 chars of source]

where $(\alpha,\delta)$ are the capital share of output and depreciation rate, while $(\rho_{z},\sigma_{\epsilon})$ are the parameters of the exogenous productivity shock. In order to obtain analytical results, let us assume that $\delta=1-\kappa\frac{Y_{t+1}}{K_{t+1}}$.\footnote{This assumption can be microfounded by assuming a depreciation rate that is a function of capital utilization, i.e. $\delta(u_{t})=1-\frac{\delta_{0}}{u_{t}}$. At the optimum choice of $u_{t}$, $\delta(u_{t}^{\star})=1-\alpha\frac{Y_{t}}{K_{t}}$, and hence $\kappa=\alpha$. } For a log-Normal productivity shock, it leads to the following DGP: \[ \left(

array[array omitted — 81 chars of source]

\right)\sim N\left(\left(

array[array omitted — 103 chars of source]

\right)+\underset{{\left(

array[array omitted — 68 chars of source]

\right)}}\left(

array[array omitted — 72 chars of source]

\right),\left(

array[array omitted — 57 chars of source]

\right)\right) \] where I have defined $\tilde{\alpha}:=\alpha+\kappa$. What we are after here is to investigate how incorrectly imposing that $\kappa=0$ but tilting the model to satisfy the (unconditional) Euler equation can reduce the bias. The first order conditions for $\beta$ are as follows:

eqnarray*[eqnarray* omitted — 389 chars of source]

where the first term is the Euler equation and the second term is the score function, with weights $(w_{1},w_{2})$ respectively which depend on structural parameters.\footnote{Specifically, $w_{1}$ corrresponds to $\mathbb{E}\left((M-s'm)V^{-1}\right)$ in the first order condition in (ref) and $w_{2}=\frac{1}{\beta \sigma^{2}_{z}}$.} Taking expectations and setting $\alpha\approxeq0$ to simplify derivations yields that to first order, \[\left(\frac{\hat{\beta}^{tilt}}{\beta}-1\right)^{2}=\left(\frac{w_{2}\kappa}{w_{1}+{w_{2}}{}}\right)^{2}<\kappa^{2}=\left(\frac{\hat{\beta}^{PMLE}}{\beta}-1\right)^{2}\] Hence, if $\kappa^{2}\sim \frac{\eta}{{N}}$, then \[\left(\frac{\hat{\beta}^{tilt}}{\beta}-1\right)^{2}=\left(\frac{\hat{\beta}^{PMLE}}{\beta}-1\right)^{2}-\frac{\eta}{{N}}\]as in Theorem (ref).

Monte Carlo Experiments

In this section, I provide simulation evidence for the finite sample performance of this method in terms of $MSE(\hat{\vartheta})$, where I consider sample sizes relevant for macroeconomic data at quarterly frequency. I perform several experiments using (a) the Consumption Euler equation that prices a risk free bond and (b) the Euler equation in a production economy. In the first model, I compare the estimator's performance to other limited information methods using a correctly specified and misspecified reduced form density (non-structural approximation case), using both the conditional and the unconditional moments. In the second model, I investigate the estimator's performance when the base model is a log-linear structural approximation to the stochastic growth model. Note that for these experiments, misspecification is kept fixed. For brevity, the Monte Carlo experiment with the model used in the empirical application is presented directly in that section.

Estimating the Consumption Euler equation

For this model, I make assumptions such as the solution and the resulting DGP are analytically tractable. Assuming a logarithmic utility function for the representative agent, the consumption Euler equation is as follows:

eqnarray*[eqnarray* omitted — 84 chars of source]

The resulting DGP is a Bivariate log-Normal VAR for consumption and the interest rate: \[ \left(

array[array omitted — 70 chars of source]

\right)\sim N\left(\left(

array[array omitted — 57 chars of source]

\right)+{\left(

array[array omitted — 52 chars of source]

\right)}\left(

array[array omitted — 66 chars of source]

\right),\left(

array[array omitted — 58 chars of source]

\right)\right) \] I use the following parameterization: $\rho_{CR}=1,\rho_{R} = 0.95,\beta=0.85$, $\sigma^{2}_{R} = 0.5$ and I calibrate the measurement error variance to $\sigma^{2}_{C,me} = 0.05$. Below, I plot the MSE comparisons for estimating the discount factor $\beta$ with a varying sample size, $N=\left\{10..200\right\}$. Smaller samples are relevant for typical sub-sample analysis done with macroeconomic data.

In Figure (ref), I compare the performance of the CU-GMM and EL (Empirical Likelihood) estimators (with data generated with no measurement error) to this paper's estimator for $\hat\beta$, both in the case of estimating $(\rho_{CR},\rho_{R},\sigma^{2}_{R})$ in the correctly specified case and when - incorrectly- setting $\rho_{CR}=\rho_{R}$. To avoid using more information than necessary, I do not make use of the knowledge that the mean of the density depends on $\beta$.\footnote{ For the CU-GMM estimator, to construct the unconditional moments without losing any information I used optimal instruments estimated using kernel weights with bandwidth equal to $N^{-0.2}$. For the empirical likelihood estimator, I follow 2004 that use a localized version of EL instead of optimal instruments, again with bandwidth equal to $N^{-0.2}$.}

figure[figure omitted — 217 chars of source]

The performance of CU-GMM and EL is worse in all cases. Interpreting CU-GMM and EL as plug-in estimators using the conditional empirical CDF, where the latter is the most basic infinite dimensional model for the true conditional CDF, it is not surprising that estimating a few parameters has better MSE performance in small samples, even in the case of misspecification.\footnote{It can be shown that $\mathbb{E}_{t}M - \mathbb{E}_{f,t}M=e^{-(2-\rho_{R})log(\beta)-(1-\rho_{R})log \tilde{R}_{t}+\frac{1}{2}(\sigma^2_{R}+\sigma^2_{C})}(1-e^{\rho_{CR}\log \widetilde{\Delta C}_{t}})$. The restriction that $\rho_{CR}=\rho_{R}$ implies that $c_{1}$ is not zero.}We observe similar performance in the unconditional case (please see Appendix \hyperref[AppB]{B} for these additional Monte Carlo simulations).

In the next experiment, I extend the analysis of the estimator's performance when the base model is a log-linear approximation to the equilibrium law of motion. In section (ref), it has already been shown that tilting reduces bias. The next numerical exercise extends the analysis of the stochastic growth model to the more general case that cannot be analytically solved. This can shed light on whether tilting actually improves on the performance of the PMLE in terms of MSE.

Estimating the Stochastic Growth Model

Assuming a Constant Relative Risk Aversion (CRRA) utility function for the representative agent and a Cobb-Douglas production function, the stochastic growth model is characterized by the following conditions:

eqnarray*[eqnarray* omitted — 305 chars of source]

where $\gamma$ is the risk aversion coefficient and $(\alpha,\delta)$ are the capital share of output and depreciation rate, while $(\rho_{z},\sigma_{\epsilon})$ are the parameters of the exogenous productivity shock. The resulting DGP is a non-linear process for productivity, consumption, the capital stock, output and the real interest rate. I employ collocation methods to obtain an accurate solution using the following parameterization: $(\alpha,\beta,\delta,\gamma,\rho_{z},\sigma_{\epsilon})=(\frac{1}{3},0.98,0.05,2,0.95,0.1)$.\footnote{I employ 100 grid points for the capital stock and 30 grid points for the productivity shock and piecewise 3rd order splines as basis functions. Collocation coefficients are computed using Newton's method.} Correspondingly, the log-linearized model is solved using QZ decomposition and the resulting density is a Gaussian state space model, whose likelihood can be readily constructed using the Kalman filter.

Figure (ref) plots the MSE comparisons for estimating the risk aversion coefficient $\gamma$.\footnote{I focus on the risk aversion coefficient as it is the only structural parameter that cannot be estimated from steady state values given enough observables.} Again the CU-GMM estimator is estimated with feasible optimal instruments, while the base density is constructed using observations on consumption growth and the capital stock. Due to misspecification, the log-linearized model delivers a higher MSE than CU-GMM. However, tilting the likelihood to satisfy the Euler equation significantly improves performance. Given our theoretical results, the simulation evidence corroborates the claim that tilting alleviates the bias induced by the approximation to the non-linear law of motion.

Computation Times

Computing the log-linear law of motion and tilting it to satisfy the non-linear moment condition takes a total of $0,017697$ seconds. The non-linear solution takes $27$ seconds.

figure[figure omitted — 209 chars of source]

Application to Pricing Macroeconomic Risk

Low frequency fluctuations in aggregate consumption have been shown to be important in explaining several asset pricing facts. The long run risk model of doi:10.1111/j.1540-6261.2004.00670.x and its subsequent variations impose cross equation restrictions that link asset prices to consumption growth, where recursive preferences RePEc:eee:moneco:v:24:y:1989:i:3:p:401-421,10.2307/2937817, 10.2307/1913778 differentiate between risk aversion and the intertemporal elasticity of substitution (IES). These restrictions can be summarized by the Euler equation that involves the unobserved aggregate consumption dividend $R_{\alpha,t+1}$, consumption growth $G_{c,t+1}$ and stochastic variation in the discount factor, $G_{l,t+1}$, as in doi:10.1111/jofi.12437:

eqnarray[eqnarray omitted — 143 chars of source]

where $\theta = \frac{1-\gamma}{1-\frac{1}{\psi}}$, $\gamma$ is the relative risk aversion coefficient, $\psi$ the IES and $\beta$ the discount factor.

Relatively recent attempts to estimate this model using standard non-durable consumption data have stressed several issues that need to be taken into account. As argued by BANSAL201652, time aggregation is an important source of bias when low frequency data is used. In fact, BANSAL201652 find empirical support for a monthly decision interval. Quarterly or yearly data are therefore likely to be responsible for the downward bias to estimates of the IES and upward bias in risk aversion, which have been puzzling in the literature, as they imply that asset prices are increasing in uncertainty.

A possible resolution of this issue is to use monthly data, which are nevertheless contaminated by measurement error. A recent paper doi:10.3982/ECTA14308, SSY hereafter, provides a mixed frequency approach to make optimal use of a long span of consumption data while keeping measurement error under control. In this application I investigate the empirical implications of an equally important aspect of empirical macro-finance, which is the quality of the underlying approximation to the equilibrium value of the unobserved $R_{\alpha,t+1}$. What has been standard up to now was to use the doi:10.1111/j.1540-6261.1988.tb04598.x log linear approximation, which has been recently criticized by JOFI:JOFI12615 as being too crude when the underlying dynamics are persistent. I take the linear approximation as given, and impose (ref) using the methodology in this paper.

To isolate the informational content of imposing the Euler equation, I employ similar data and specification to doi:10.3982/ECTA14308 using only monthly data from Feb 1959 - Dec 2015 on non durable per capita consumption growth (Quantity Index for Personal Consumption Expenditures) and the risk free rate.\footnote{Please see appendix A5 of doi:10.3982/ECTA14308 on how the ex-ante real risk free rate is computed.} The underlying approximating model for consumption growth ($\Delta c_{t+1}=\log(G_{c,t+1})$) and the risk free rate ($r_{f,t}$) is summarized as follows:

eqnarray[eqnarray omitted — 445 chars of source]

where $(B_{0},B_{1},B_{2,x},B_{2,c})$ are functions of the deep parameters $(\gamma,\psi,\beta)$, the risk aversion, the elasticity of intertemporal substitution and discount factor respectively.\footnote{For details on the underlying solution mapping I encourage the reader to consult doi:10.3982/ECTA14308.} Moreover, $x_{t}$ is the persistent component of consumption growth and the time preference shock $x_{l,t+1}:=\log(G_{l,t+1})$ is also possibly persistent. Regarding the observation equation, I calibrate the measurement error to the values estimated by SSY\footnote{For monthly consumption growth, I set $\sigma^2_{me}=2(\sigma^2_{\epsilon}+\sigma^2_{q})$ where $\sigma^2_{\epsilon}$ and $\sigma^2_{q}$ are the variances of the measurement error in monthly and quarterly consumption respectively, as estimated by SSY. This specification is approximately equal to SSY's specification of the measurement error in consumption growth from the third month to the first month of the next quarter, so $\sigma^2_{me}$ is actually an upper bound to the measurement errors of the rest of the months.}.

I perform estimation in two steps. A subset of the reduced form dynamics, that is, equations 17,19,21-23, are identified without using the long run risk model (and hence (ref) is irrelevant to their identification). The cash flow parameters $\phi\equiv (\rho,\chi_{x},\sigma,\rho_{v,x},\sigma_{v,c})$ are therefore estimated by using Markov Chain Monte Carlo and the particle filter.\footnote{As in {doi:10.3982/ECTA14525}, I rely on quantiles of the posterior draws to construct confidence sets.} I then estimate the deep parameters $(\psi,\gamma)$ conditional on the posterior mode, $\phi^{\star}_{post}$.\footnote{Note that pre-estimating $\phi$ is consistent with the theory developed earlier as the moment condition (ref) does not provide identifying information for $(\rho,\chi_{x},\sigma,\rho_{v,x},\sigma_{v,c})$. Moreover, since the moment condition (ref) does not provide a measurement density for $x_{l,t}$, I would have to rely on the approximate model to identify the process parameters. I thus calibrate the time preference risk parameters $\rho_{l}$ and $\sigma_{l}$ to the posterior median estimates of SSY. A Bayesian approach has been recently proposed to deal with the lack of measurement density by GALLANT2017198. } This implies that $\varphi\equiv (B_{0},B_{1},B_{2,x},B_{2,c})$, where $(B_{0},B_{1},B_{2,x},B_{2,c})$ are functions of $(\psi,\gamma)$. The base density used is the predictive density of the non-Gaussian state space model (Gaussian conditional on the identified volatility states) for $(\Delta c_{t+1},r_{f,t+1})$ \footnote{The are obviously alternative more efficient algorithms to deal with stochastic volatility i.e. Metropolis within Gibbs algorithm.}. Correspondingly, the conditionally Gaussian model is a bivariate Normal distribution for $(\Delta c_{t+1},r_{f,t+1})$ with conditional means $\mu_{c} + x_{t}$ and $B_{0} + B_{1}x_{t+1} + B_{2,x}\sigma^{2}_{x,t+1} + B_{2,c}\sigma^{2}_{c,t+1}$ respectively. Conditional linearity is achieved by using the doi:10.1111/j.1540-6261.1988.tb04598.x approximation to asset returns $r_{a,t+1}$ and solving for the price consumption ratio.

I next investigate the usefulness of tilting the approximating density to satisfy the non linear condition (ref), both in terms of parameter identification and model prediction. I find that this is both empirically relevant, and economically significant.\footnote{This also corroborates the numerical results of JOFI:JOFI12615, who have shown that this approximation can be too crude when consumption growth is persistent.}

figure[figure omitted — 150 chars of source]

Figure (ref) presents the benchmark posterior distributions together with the prior distributions, both in the log-linear approximate model (AM) and the tilted model (TM). As evident, correcting for the underlying non-linearity leads to improved identification as the posterior mode and MLE estimates are closer to what are considered more plausible values for risk aversion and the elasticity of intertemporal substitution. A direct implication is that measurement error is not the only source of upward bias in estimates for the former and downward bias in estimates of the latter. Approximation errors is clearly another one. Moreover, posteriors are narrower. I also report the Maximum Likelihood estimates for the tilted model, to give a sense of how much the prior information matters for both exercises.

In terms of relative fit, the tilted model is strictly preferred by the data. At their respective modes ($\psi^{\star}$), the log-posterior and log-likelihood of the tilted model are $7194.5$ and $7206.3$ respectively, while for the non-tilted approximate model they equal $7188.7$ and $7196.7$. The tilted model dominates by 5.8 log-posterior and 9.6 log-likelihood units respectively. The corresponding Bayes factor when computing the marginal likelihood using the harmonic mean ranges from 5.62 to 5.66.\footnote{Please see Geweke99 and a5de6c49 for details on this method.}

Monte Carlo Simulation

The Monte Carlo evidence presented in Section 6 was based on more stylized models than the one employed in the application. Nevertheless, the same basic observations follow through when simulating a model similar to the one employed in this section. I next present MSE results from estimating the model, abstracting away from preference shocks while volatility is treated as observed.

The linear and non-linear DGP's are simulated using the code provided by JOFI:JOFI12615. The DGP is based in setting $(\psi=1.5, \gamma= 6, \rho=0.999,\chi_{x}=3,\sigma=0.2, \sigma_{v,c} =0.004, \sigma_{v,x}=0.01, \rho_{v,x}=\rho_{v,c}=0.99)$. As evident from Figure (ref), while increasing the sample size does eliminate a significant amount of the MSE in estimating both parameters, and in particular risk aversion, tilting yields significant reductions in the MSE at all sample sizes.

Computation Times

Finally, from the practitioner's point of view, solving the model non-linearly takes $75$ minutes, while the log-linear solution takes $0.30$ seconds. Tilting the log-linear model to satisfy the non-linear condition takes $0.008$ seconds. For estimation purposes at least, tilting is a computationally attractive method.

figure[figure omitted — 205 chars of source]

Conclusion

Model approximations are often necessary in applied settings, but they can sometimes eliminate features of the exact model that are consequential for inference, and naturally involve a bias-variance tradeoff. The paper has considered the econometric properties of tilting the approximate model to re-impose features such as conditional moment restrictions, and finds that both in theory, simulation and empirical results that tilting substantially improves econometric performance. The paper's application and additional monte carlo evidence also contributes to the empirical macro-finance literature by demonstrating that employing the widely used linear approximations to returns can affect parameter estimates, and hence it complements existing computational evidence by JOFI:JOFI12615. The quality of the approximation is yet another reason for the downward bias to estimates of the intertemporal elasticity of substitution and the upward bias in risk aversion.