EconBase
← Back to paper

The Empirical Saddlepoint Estimator

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

38,230 characters · 12 sections · 52 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

The Empirical Saddlepoint Estimator

center[center omitted — 51 chars of source]
abstractWe define a moment-based estimator that maximizes the empirical saddlepoint (ESP) approximation of the distribution of solutions to empirical moment conditions. We call it the ESP estimator. We prove its existence, consistency and asymptotic normality, and we propose novel test statistics. We also show that the ESP estimator corresponds to the MM (method of moments) estimator shrunk toward parameter values with lower estimated variance, so it reduces the documented instability of existing moment-based estimators. In the case of just-identified moment conditions, which is the case we focus on, the ESP estimator is different from the MM estimator, unlike the recently proposed alternatives, such as the empirical-likelihood-type estimators. { Keywords: Empirical Saddlepoint Approximation; Method of Moments; Kullback-Leibler Divergence Criterion; Maximum-probability Estimator; Variance Penalization. }

\setcounter{tocdepth}{3}

Introduction

The saddlepoint (SP) approximation has been developed to approximate distributions. Because of its accuracy it is regularly used in several fields, such as numerical analysis (e.g., 2000Loader's algorithm to approximate binomial distributions, and which is notably used in the statistical software R) and actuarial sciences (e.g., 1932Ess's approximation for distributions tails). In statistics, the SP approximation and its empirical version ---the empirical saddlepoint (ESP) approximation--- have been used to approximate finite-sample distributions 1954Dan,1988DavisonHinkley.\footnote{Standard monographs and introductions about the ESP and the SP approximation for statistics include 1990FieRon, 1994Kolassa, 1995Jen, 1999GouCas and is 2007Butler.\\ $^*$University of Luxembourg, 6 rue Coudenhove-Kalergi, L-1359 Luxembourg\\ ${ }^{\#}$Carnegie Mellon University, 5000 Forbes Ave, Pittsburgh, PA 15213, USA. }

In the present paper, we propose to use the ESP approximation to define a point estimator $\hat{\theta}_T$. We call it the ESP estimator. It maximizes the 1994RonWel's ESP approximation, i.e.,

eqnarray[eqnarray omitted — 105 chars of source]

where $ \hat{f}_{\theta^{*}_T}(.)$ is the ESP approximation of the distribution of solutions to empirical moment conditions

eqnarray[eqnarray omitted — 100 chars of source]

and where $\psi(.,.)$ denotes the moment function s.t. $\mathbb{E} [ \psi(X_1, \theta_0) ]=0_{m \times 1}$ an $m$-dimensional vector of zeros, $(X_t)_{t=1}^T$ i.i.d. data, $\theta_{0}\negmedspace\in\negmedspace \mathbf{\Theta}\negthickspace\subset \negthickspace\mathbf{R}^m$ the unknown parameter of interest, and $T$ the sample size. The exact formula for $\hat{f}_{\theta^{*}_T}(.)$ is reminded below in equation (ref) on p. (ref).

The ESP estimator is a moment-based estimator. Since 1894Pea,1902Pearson's method of moment (MM), moment-based estimators have been found useful in a variety of applications (e.g., covariance structure analysis in psychology, and asset pricing in economics). Their two main advantages are (i) they do not require a parametric family of probability distributions for the data so they are less prone to model misspecification, and (ii) they allow complex models for which the likelihood function is intractable.

Nevertheless, the increase use of the MM and its extensions has revealed that they can be unstable and perform poorly in finite samples (e.g., July 1996 special issue of JBES). The idea of the ESP estimator to improve on the MM estimator is the following. By definition, the MM estimate $\theta^*_T(\omega)$ solves a realization of the empirical moment conditions (ref), but it typically does not solve the empirical moment condition for another realization $\omega$ of the data. Thus, we might want an estimate that does only take into account the realized empirical moment conditions, but also their other potential realizations. More precisely, we want an estimate that accounts for all the potential realizations of the empirical moment conditions according to their probability weight of occurrence. This leads to the ESP estimate, which is a maximum-probability estimate. The ESP estimate maximizes the estimated probability weights of solving the empirical moment conditions.\footnote{This is in contrast to the traditional motivation for ML estimators, which maximize the probability weights of obtaining a sample equal to the observed sample. In other words, the support of the ESP distribution is the parameter space, while the support of the distribution associated with a likelihood is the data space. Thus, if we are looking for relevant parameter values instead of data values ---as it is typically the case---, a maximum-probability motivation appears more appealing than the traditional ML motivation. } If the empirical moment conditions (ref) have a unique solution with a continuous distribution, the ESP estimator maximizes the ESP approximation of a probability density function of the solution $\theta^*_T$.\footnote{Another motivation for maximum probability estimators is decision theoretic. Maximum probability estimators follows from the minimization of the expectation of a loss “function” that equals zero when $\theta$ solves the empirical moment conditions and one otherwise by normalization. This motivation is similar to the decision-theoretic justification for the Bayesian maximum a posteriori estimator 1994Robert. As in Bayesian analysis, the choice of other loss functions is possible. It is left for future research.} We rely on the ESP approximation because simulation and theoretical evidence shows the ESP approximation can be very accurate in small sample 1988DavisonHinkley,1994RonWel.

Besides the maximum-probability motivation, we show that the ESP estimator corresponds to an MM estimator shrunk toward parameter values with lower estimated variance. More precisely, we decompose the logarithm of the ESP approximation as the sum of a term, which is maximized at the MM estimator, and a variance penalty, which discounts parameter values with high estimated variance. Under assumptions adapted from the entropy literature, we establish the ESP estimator has the same good asymptotic properties as the MM estimator, so the variance penalization is a finite-sample correction. We also derive the ESP counterparts of the Wald, Lagrange multiplier (LM), analogue likelihood-ratio (ALR) test statistics, as well as another test statistic. Then, we investigate the ESP estimator through Monte-Carlo simulations. We compare its performance with the exponential tilting (ET) estimator, which is equal to the MM estimator in the just-identified case (i.e., when the number of parameters is the same as the number of moment conditions). Results show that the variance penalization of the ESP estimator reduces the finite-sample instability of the ET estimator (or equivalently, of the MM estimator). An empirical application illustrates the gain from this greater stability in terms of inference.

The ESP estimator is not the first proposal to improve on the MM and its extensions. Alternative moment-based approaches have been proposed such as the empirical likelihood approach of Owen 1994QinLawless, the continuously updating approach 1996HanHeaYar, the already-mentioned exponential tilting (ET) approach 1997KitStu,1998ImbSpaJoh, and combinations of the aforementioned approaches 2007Schennach. All these approaches yield an estimator closely related to the empirical likelihood estimator, so we call them empirical-likelihood-type estimators. In the just-identified case, when well-defined, all of these empirical-likelihood-type estimators are numerically equal to the original Pearson's MM estimator $\theta^*_T$. Because we focus on the just-identified case, it thus is sufficient for us to compare the ESP estimator with the MM estimator, or with one of any of these more recent estimators.

In addition to the already cited papers, the present paper, which supersedes the unpublished manuscript 2009Sow, is related to many other ones. We clarify these relations in Section (ref) (p. (ref)). To the best of our knowledge, none of the prior papers use the SP or the ESP to propose a novel moment-based point estimator. Overall, the present paper brings together the literature on the saddlepoint approximation and the literature on moment-based estimation.

Finite-sample analysis

In the present section, we remind the formula for the ESP approximation, and analyze its finite-sample structure. Then, we decompose the log-ESP into two terms and show that the ESP estimator is a MM estimator shrunk toward parameter values with lower estimated variance.

The ESP approximation

Formalizing and generalizing prior works 1988DavisonHinkley,1989Feu,1990Wang,1990YoungDaniels, 1994RonWel propose the following ESP approximation to estimate the distribution of a solution to the empirical moment conditions (ref)

eqnarray[eqnarray omitted — 279 chars of source]

where $|.|_{\det}$ denotes the determinant function, $\theta^*_T$ a solution to (ref), $\psi_t(.):=\psi (X_t,.)$, and

eqnarray[eqnarray omitted — 761 chars of source]

The ESP approximation (ref) is the empirical counterpart of the SP approximation of 1982Fie. From a computational point of view, the ESP approximation (ref) is not complicated.\footnote{We do not claim that the ESP estimator is as easy to compute as the MM estimator, but that its additional complexity is similar to the recently proposed empirical-likelihood-type estimators (e.g., ET estimator), and that it is worthwhile in several applications (e.g., Section (ref)). Moreover, it seems to make sense to develop novel estimation methods that take advantage of the increasingly available computational power.} The only implicit quantity is $\tau_T(\theta)$, which solves the tilting equation (ref), which, in turn, is just the FOC (first-order condition) of the unconstrained convex problem $\min_{\tau\in \mathbf{R}^m}\sum_{t=1}^T\mathrm{e}^{\tau ' \psi_t(\theta)} $. A full understanding of the ESP approximation (ref) arguably requires to work through higher-order asymptotic expansions along the lines of 1982Fie. However, direct inspection of the ESP approximation (ref) also provides insight for how it incorporates information from the data through two channels.

The first channel is the ET (exponential tilting) term $\exp\left\{T\ln\left[ \frac{1}{T}\sum_{t=1}^T\mathrm{e}^{\tau_T(\theta) ' \psi_t(\theta)}\right]\right\} $. In equation (ref), for any $\theta\in \mathbf{\Theta} $, the terms $\frac{\exp\left[\tau_T(\theta) ' \psi_t(\theta)\right]}{\sum_{i=1}^T\exp\left[\tau_T(\theta) ' \psi_i(\theta)\right]} $ tilt (i.e., reweight) the empirical weights $ 1/T$, so the finite-sample moment conditions (ref) holds. This tilting determines, through equation (ref), the multinomial distribution $(w_{t,\theta})_{t=1}^T$ that is the closest to the empirical distribution ---in the sense of the Kullback-Leibler divergence criterion--- s.t. the finite-sample moment conditions (ref) holds: The tilting equation (ref) is the FOC w.r.t. (with respect to) $\tau$ of the Lagrangian dual problem of the minimization problem

eqnarray[eqnarray omitted — 287 chars of source]

where $\sum_{t=1}^T w_{t,\theta}\log[w_{t,\theta}/(1/T)]$ is the Kullback-Leibler divergence criterion between the empirical distribution and the multinomial distribution $(w_{t,\theta})_{t=1}^T$ with the same support 1981Efron,1997KitStu. Then, for the given $\theta\in \mathbf{\Theta} $, in the ESP approximation (ref), the ET term $\exp\left\{T\ln\left[ \frac{1}{T}\sum_{t=1}^T\mathrm{e}^{\tau_T(\theta) ' \psi_t(\theta)}\right]\right\} $ indicates the extent of the tilting needed to set the finite-sample moment conditions (ref) (or equivalently, equation (ref)) to zero. The bigger is the tilting of the empirical distribution, the less compatible are the data with $\theta$ solving the empirical moment conditions, and the smaller should be the ET term $\exp\left\{T\ln\left[ \frac{1}{T}\sum_{t=1}^T\mathrm{e}^{\tau_T(\theta) ' \psi_t(\theta)}\right]\right\} $. It can be easily seen that $\frac{1}{T}\sum_{t=1}^T\mathrm{e}^{\tau_T(\theta) ' \psi_t(\theta)}$ reaches its maximum when $\theta$ is a solution $\theta^*_T$ of the empirical moment conditions (ref), i.e., when $\tau_T(\theta^*_T)=0_{m \times 1} $ and no tilting is needed.\footnote{For a complete proof, one can follow the same reasoning as in the proof of Lemma (ref) in 2019HolcblatSowell with the empirical distribution in lieu of $\mathbb{P}$.}

In the ESP approximation on equation (ref), the second term $\left(\frac{T}{2\pi}\right)^{m/2} $ comes from the multivariate Gaussian distribution that is the leading term of the Edgeworth's asymptotic expansions underlying ESP approximations. However, because it is constant w.r.t. $\theta$, it does not affect the maximization of the ESP approximation, so it is not an information channel for the ESP estimator. The remaining term $\left|\Sigma_T(\theta) \right|_{\det}^{-\frac{1}{2}}$ , which we call the variance term, is the second channel through which the ESP approximation incorporates information from data. The variance term discounts the ET term according to the tilted estimated variance of the solution to the finite-sample moment conditions. Under standard assumptions, a consistent estimator of the asymptotic variance of $\sqrt{T}(\theta^*_T-\theta_0)$ is $\left[ \frac{1}{T}\sum_{t=1}^T \frac{\partial \psi_{t} (\theta^*_T)}{\partial \theta}\right]^{-1}\left[\frac{1}{T}\sum_{t=1}^T \psi_{t}(\theta^*_T)\psi_{t}(\theta^*_T)'\right]\left[ \frac{1}{T}\sum_{t=1}^T \frac{\partial \psi_{t} (\theta^*_T)'}{\partial \theta}\right]^{-1} $. The bigger the variance term is, the less plausible a solution takes exactly this value, and the smaller is $\left|\Sigma_T(\theta) \right|_{\det}^{-\frac{1}{2}}$ ---note the negative power. Therefore, overall, for a given $\theta\in \mathbf{\Theta} $, the bigger the tilting or the estimated variance, the smaller the ESP approximation, i.e., the estimated probability weight that $\theta$ solves the empirical moment conditions (ref).

The ESP estimator as a shrinkage estimator

As explained in the introduction, the recently proposed moment-based estimators are numerically equal to the Pearson's MM estimator in the just-identified case. Thus, it is sufficient to compare the ESP estimator with one of them in order to understand the difference between the former and the other proposed moment-based estimators. The ET estimator of 1997KitStu and 1998ImbSpaJoh is particularly convenient for this purpose. Taking the logarithm of the ESP approximation (ref), and removing the terms constant w.r.t. $\theta$, it can be seen that, $\mathbb{P}$-a.s. for $T$ big enough, the ESP estimator $\hat{\theta}_T$ maximizes the objective function

eqnarray[eqnarray omitted — 172 chars of source]

where $\ln \left[\frac{1}{T}\sum_{t=1}^T \mathrm{e}^{\tau_T(\theta)'\psi_t(\theta)}\right]$ is an increasing transformation of the objective function of the ET estimator. Thus, the difference between the ESP estimator and the ET estimators comes only from the log-variance term $-\frac{1}{2T} \ln \vert \Sigma_T(\theta)\vert_{\det}$. The latter does not only incorporates additional information from data as explained in Section (ref), but it also penalizes parameter values with higher estimated variance. Thus, the ESP estimator is an ET estimator ---or equivalently, a MM estimator--- shrunk toward parameter values with lower estimated variance. Now, as the factor $\frac{1}{2T} $ suggests and the proofs of Section (ref) show, the log-variance term vanishes asymptotically, so the shrinkage is a finite-sample correction.

Asymptotic properties

In the present section, we investigate the asymptotic properties of the ESP estimator. Good asymptotic properties can be regarded as a minimal requirement for the ESP estimator, which is based on a small-sample asymptotic approximation. All the proofs and assumptions are in the online Appendix 2019HolcblatSowell.

Existence, consistency and asymptotic normality

Under assumptions adapted from the entropy literature, the following theorem establishes the existence, the strong consistency, and the asymptotic normality of the ESP estimator $\hat{\theta}_T$.

theorem[Existence, consistency and asymptotic normality] Under Assumption (ref), $\mathbb{P}$-a.s. for $T$ big enough, there exists $\hat{\theta}_T$ s.t. \begin{enumerate} • $\mathbb{P}$-a.s. as $T \rightarrow \infty$, $\hat{\theta}_T \rightarrow \theta_0$ ; and • under the additional Assumption (ref), as $T \rightarrow \infty$, $\sqrt{T} ( \hat{\theta}_T - \theta_0 ) \underset{}{\stackrel{D}{\longrightarrow}} \mathcal{N}\left(0, \Sigma(\theta_0)\right)$. \end{enumerate} where $\Sigma(\theta_0):=\negthickspace \left[\mathbb{E} \frac{\partial \psi(X_{1},\theta_{0})}{\partial \theta' }\right]^{-1}\negthickspace\negthickspace \mathbb{E}\left[\psi(X_{1},\theta_{0}) \psi(X_{1},\theta_{0})' \right] \left[ \mathbb{E} \frac{\partial \psi(X_{1},\theta_{0})'}{\partial \theta }\right]^{-1} $, $\stackrel{D}{\rightarrow} $ denotes the convergence in distribution.

Theorem (ref) shows that the ESP estimator has the same first-order asymptotic properties as the MM and hence the recently proposed moment-based estimators. Although the asymptotic properties of the ESP estimator are standard, the proof of Theorem (ref) is quite involved. The crux of the proof is to show that the variance penalization $-\frac{1}{2T} \ln \vert \Sigma_T(\theta)\vert_{\det}$ vanishes sufficiently quickly asymptotically, so it does not distort the first-order asymptotic.

More on inference\,: The trinity$+1$

The ESP estimator provides different ways to test parameter restrictions

eqnarray[eqnarray omitted — 93 chars of source]

where $r: \mathbf{\Theta} \rightarrow \mathbf{R}^q$ with $q \in [\![ 1, \infty[\![$. More precisely, within the ESP framework, there exist the usual trinity of Wald, LM and ALR tests statistics, plus another test statistic, which we call the exponential tilting (ET) test statistic. Our ET test has a structure similar to a test for over-identifyied moment conditions in 1998ImbSpaJoh.

Under a mild standard additional assumption, the following theorem shows that the Wald, LM ALR, and ET statistics asymptotically follow a chi-squared distribution with $q$ degrees of freedom.

theorem[The trinity$+1$: Wald, LM, ALR and ET tests] Define $R(\theta):=\frac{\partial r(\theta)}{\partial \theta'}$, and the following Wald, LM, ALR and ET test statistics \begin{eqnarray*} \mathrm{Wald}_T& :=& Tr(\hat{\theta}_T)'[R(\hat{\theta}_T)\widehat{\Sigma(\theta_0)}_TR(\hat{\theta}_T)']^{-1}r(\hat{\theta}_T)\\ \mathrm{LM}_T &:=& T \check{\gamma}_T'[R(\check{\theta}_T) \widehat{\Sigma(\theta_0)}_TR(\check{\theta}_T)']\check{\gamma}_T=\frac{\partial\ln[\hat{f}_{\theta^{*}_T}(\check{\theta}_T)]}{\partial \theta' }\widehat{\Sigma(\theta_0)}_T^{-1}\frac{\partial\ln[\hat{f}_{\theta^{*}_T}(\check{\theta}_T)]}{\partial \theta }\\ \mathrm{ALR}_T & := & 2\{\ln[\hat{f}_{\theta^{*}_T}(\hat{\theta}_T)]-\ln[ \hat{f}_{\theta^{*}_T}(\check{\theta}_T)]\}\\ \mathrm{ET}_T&:=& T \tau_T(\check{\theta}_T)'\widehat{V}_T\tau_T(\check{\theta}_T) \end{eqnarray*} where $\widehat{\Sigma(\theta_0)}_T$ and $\widehat{V}_T$ are symmetric matrices that converge in probability to $\Sigma(\theta_0)$ and $ \mathbb{E}[ \psi(X_1, \theta_0)\psi(X_1, \theta_0)'] $, respectively; and where $\check{\gamma}_T$ and $\check{\theta}_T$ respectively denote the Lagrange multiplier and a solution to the maximization of $\hat{f}_{\theta^{*}_T}(\theta)$ w.r.t. $\theta\in \mathbf{\Theta}$ under the constraint that $r(\theta)=0_{q \times 1}$.\footnote{In mathematical terms, $\check{\theta}_T\in \arg \max_{\theta \in \check{\Theta}} \hat{f}_{\theta^{*}_T}(\theta)$ where $\check{\Theta}:=\{\theta \in \Theta: r(\theta)=0_{q \times 1}\}$ and $\check{\gamma}_T$ is the Lagrangian multiplier s.t. $\frac{1}{T} \frac{\partial \ln[\hat{f}_{\theta^{*}_T}(\check{\theta}_T)]}{\partial \theta} + \frac{\partial r(\check{\theta}_T)'}{\partial \theta}\check{\gamma}_T=0_{m \times 1}$.} Under Assumptions (ref), (ref) and (ref), if the test hypothesis (ref) holds, as $T \rightarrow \infty$, \begin{eqnarray*} \mathrm{Wald}_T, \mathrm{LM}_T, \mathrm{ALR}_T, \mathrm{ET}_T \stackrel{D}{\rightarrow } \chi^2_{q}. \end{eqnarray*}

Theorem (ref) can also be used to obtain valid confidence regions by the inversion of the test statistics with $\check{\theta}_T=\theta_0$. Our Wald, LM and ALR test statistics share some similarity with the test statistics proposed in 1997KitStu, 1998ImbSpaJoh and 2003RobRonYou. The main difference is that the latter are built around (possibly constrained) maximizers of the ET term, while our tests statistics are based on the (possibly constrained) ESP estimator, which maximizes the whole ESP approximation including the variance term.

Examples

In the present section, we further investigate and illustrate the finite-sample properties of the ESP estimator.\footnote{In addition to our finite-sample analysis of the ESP objective function (Section (ref)), our derivation of the first-order asymptotic properties (Section (ref)), our Monte-Carlo simulations and empirical application (present section), another way to shed light on the finite-sample properties of the ESP estimator would be to derive its higher-order asymptotic properties such as its second-order bias 1996RilstoneSrivastavaUllah. In the present paper, we do not follow this way because it would add several dozens of pages of proofs without much insight: Our preliminary derivations yield a long and complicated structure for the second-order bias, from which we struggle to gain insight. The length and the complexity of the second-order bias mainly comes from (i) the derivatives of the variance $\left|\Sigma_T(\theta) \right|_{\det}^{-1/2}$; and (ii) the reliance on the exact FOCs instead of approximate FOCs. A mild preview of this complexity can be seen in 2019HolcblatSowell.} We focus on the comparison with the ET estimator, as previously noted, (i) in the just-identified case, which is the case addressed in the present paper, the MM estimator and the recently proposed moment-based estimators are equal to the ET estimator so there is no loss of generality in terms of point estimation, and (ii) the ESP objective function nests the ET objective function, so that the source of the difference between the two is easily understood ---it necessarily comes from the variance term (see Section (ref)). For brevity, we present the main results for a numerical and an empirical example that are known to be challenging for moment-based estimation.

Numerical example\,: Monte-Carlo simulations

We simulate the just-identified version of the 1996HalHor model, which has become a standard benchmark to compare the performance of moment-based estimators in statistics 2007Schennach,2012LoRonchetti and econometrics 1998ImbSpaJoh,2001Kitamura. This model can be interpreted as a simplified consumption-based asset pricing model where $\beta $ is the relative risk aversion (RRA) parameter 2002GreLamSmi. In the simulations, we estimate the two parameters $(\mu, \beta)$ with the moment function

eqnarray*[eqnarray* omitted — 315 chars of source]

where $\mu_0 = -.72 $, $ \beta_0 = 3$, and $X_{t}$ and $Y_{t}$ are jointly i.i.d. random variables with distribution $\mathcal{N}(0, .16)$.

table[table omitted — 1,632 chars of source]

Table (ref) reports the mean-squarred error (MSE), bias and variance of the ESP and ET estimators for different sample sizes. The MSE, the variance and the bias of the ESP estimator are always smaller than for the ET estimator, and the differences are notable, especially for small sample sizes. In fact, Table (ref) understates the improvement delivered by the variance penalization of the ESP objective function. We help the ET estimator (or equivalently, the MM estimator),\footnote{We numerically check that they deliver the same estimates even for the small sample sizes $T=25$ and $50$. } by restricting its parameter space to $\beta< 15$. Without this parameter restriction, the behaviour of the ET estimator is very unstable. An analysis of the typical shape of the objective functions for small sample size explains this phenomenon. The typical ET objective function has a ridge that follows from around the population parameter values ($\beta_{0}=3$, $\mu_{0}=-.72$) towards $(1000, -600)$. The ridgeline is not totally flat, and it often has a gentle downward slope as we move away from the area near the population parameter values. However, regularly, for some simulated samples, the very top of the ridge is extremely far from the population parameter values, so that ET estimates are very far from the population parameter values. This does not happen for the ESP estimator. The variance term of the ESP objective function ensures that the ridge drops sufficiently as we move away from the maximum that is near the population parameter value. Thus, in line with our finite-sample analysis of the ESP objective (Section (ref)), the ESP estimator is much more stable.

Empirical example

In this section, we present an empirical example from asset pricing. Since 1982HanSin, moment-based estimation is standard in consumption-based asset pricing. For brevity, we focus on the key features of the example. See 2019HolcblatSowell for additional information and comparisons.

table[table omitted — 1,613 chars of source]

We estimate the relative risk aversion (RRA) $\theta$ of a representative agent of the US economy. Previous studies have shown that existing moment-based estimation approaches often produce unstable RRA parameter estimates. We rely on the following moment condition

eqnarray[eqnarray omitted — 127 chars of source]

where $\frac{C_{t}}{C_{t-1}}$ is the growth consumption and $(R_{m,t}-R_{f,t})$ the market return in excess of the risk-free rate. The moment condition, which is common to many consumption-based asset pricing models, and the data are similar to 2012JulGho corresponding to standard US data at yearly frequency from Shiller's website spanning from 1890 to 2009. We report ET and ESP estimates as well as confidence regions based on the inversion of the ALR test statistics of Theorem (ref) (p. (ref)) with $\check{\theta}_T=\theta_0$. The latter have the advantage to take into account the whole shape of the objective function unlike $t$-statistics-based confidence regions, which only account for the shape of the objective function in a neighborhood of the estimate through its standard errors.

In Table (ref), Figures (A) and (B) respectively display the ET term and the ESP approximation. For ease of comparison, the scale is the same, and we normalize both of them so they integrate to one. The normalized ET term is much flatter around its maximum than the normalized ESP approximation. Flatness of the objective function around the estimate has been documented for other existing moment-based estimators, and it has often been regarded as one of the main sources of the instability of the RRA estimates 2000StoWri,2001NeeRoyWhi. Figure (B) shows that the normalized ESP is sharp around the ESP estimator. The relative sharpness of the ESP yields sharper confidence regions\,: The ESP confidence region is less than half its ET counterpart. In light of the variance penalization term in the ESP objective function (Section (ref) on p. (ref)) and the shrinkage-like behavior of the ESP estimator in the Monte-Carlo simulations (Section (ref)), the relative sharpness of the ESP inference is not surprising. In 2019HolcblatSowell, additional empirical evidences corroborate the increased stability and precision of the ESP estimator w.r.t. the ET estimator (or equivalently, MM estimator).

Connection to the literature and further research directions

The present paper demonstrates a previously unknown connection between the SP approximation and moment-based estimation, and hence it is related to many papers in addition to the ones already cited. Following 1954Dan, the literature in statistics 1986EastonRonchetti,1991Spady,1992Jen,2012LaVecchiaRonchettiTrojani, 2015BrodaKan,2018FasioloEtAl and econometrics 1978Phi,1979HolPhi,1982Phi,1994Lieberman,2006AitSahYu has used the SP (saddlepoint) and ESP approximations to obtain accurate approximations of distributions, especially in the tails. The strand of the SP literature that is closest to our paper derives SP approximations to the distribution of statistics that correspond to solutions of nonlinear estimating equations. The latter strand of literature started with 1982Fie and continued with 1990Sko,1993MontiRonchetti,1997Imb,1998JenWoo,2000AlmFieRob,2003RobRonYou, and 2003RonTro, among others. More recently, 2010CzeRon, 2011MaRon, and 2012LoRonchetti,2013KunRil,2015KundhiRilstone propose more accurate tests for indirect inference, functional measurement error models, moment condition models, nonlinear estimators and GEL (generalized empirical likelihood) estimators, respectively. To the best of our knowledge, unlike the present paper, none of the prior papers use the SP or the ESP to develop an estimation method that yields a novel moment-based estimator. In ongoing work, we generalize the ESP approximation to the over-identified case, and establish further good mathematical properties.

Notes and acknowledgements

Parts of the present paper have previously circulated under the title “The Empirical Saddlepoint Likelihood Estimator Applied to Two-Step GMM" 2009Sow. Some proofs of the present paper also borrow technical results from 2012Hol. Helpful comments were provided by Philipp Ketz (discussant), Eric Renault, Aman Ullah and seminar/conference participants at Carnegie Mellon University, CFE-CMStatistics 2017, Swiss Finance Institute (EPFL and the University of Lausanne), 10th French Econometrics Conference (Paris School of Economics), at the Econometric Society European Winter Meeting 2018 (University of Naples Federico II), and at the University of Luxembourg.