EconBase
← Back to paper

The Empirical Saddlepoint Estimator

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

38,230 characters

The Empirical Saddlepoint Estimator



\begin{center} Version with online appendix included.\end{center}

\medskip

\begin{abstract}
We define a moment-based estimator that maximizes the empirical saddlepoint (ESP) approximation of the distribution of solutions to empirical moment conditions. We call it the ESP estimator. We prove its existence, consistency and asymptotic normality, and we propose novel test statistics. We also show that the ESP estimator corresponds to the MM (method of moments) estimator shrunk toward parameter values with lower estimated variance, so it reduces the documented instability of existing moment-based estimators. In the case of just-identified moment conditions, which is the case we focus on, the ESP estimator is different from the MM estimator, unlike the recently proposed alternatives, such as the empirical-likelihood-type estimators.


\bigskip



{\noindent \textit{Keywords}: Empirical Saddlepoint Approximation; Method of Moments;  Kullback-Leibler Divergence Criterion; Maximum-probability Estimator; Variance Penalization.
}

\medskip


\end{abstract}

\date{\today}

\maketitle
\setcounter{tocdepth}{3}



\section{Introduction }\label{sec1}
The saddlepoint (SP)  approximation has been developed to  approximate distributions. Because of its accuracy it is regularly used in several fields, such as  numerical analysis  (e.g., \cite{2000Loader}'s algorithm to approximate  binomial distributions, and which is notably used in the statistical software R) and actuarial sciences (e.g., \cite{1932Ess}'s approximation for distributions tails). In statistics, the SP approximation and its empirical version ---the empirical saddlepoint (ESP) approximation---  have been used to approximate  finite-sample distributions  \citep[e.g.,][]{1954Dan,1988DavisonHinkley}.\footnote{Standard monographs and introductions about the ESP and the SP approximation for statistics include \cite{1990FieRon}, \cite{1994Kolassa},  \cite{1995Jen},   \cite{1999GouCas}  and is \cite{2007Butler}.\\
$^*$University of Luxembourg, 6 rue Coudenhove-Kalergi, L-1359 Luxembourg\\
${ }^{\#}$Carnegie Mellon University, 5000 Forbes Ave, Pittsburgh, PA 15213, USA.
}

In the present paper, we propose to use the ESP  approximation to define a \textit{point} estimator $\hat{\theta}_T$.
We call it the ESP estimator. It  maximizes the   \cite{1994RonWel}'s ESP approximation, i.e.,
\begin{eqnarray}
\hat{\theta}_T \in \arg \max_{\theta \in \mathbf{\Theta}}  \hat{f}_{\theta^{*}_T}(\theta)
\end{eqnarray}
where  $ \hat{f}_{\theta^{*}_T}(.)$  is the ESP approximation of  the distribution of  solutions to  empirical  moment conditions  \begin{eqnarray}
\frac{1}{T}\sum_{t=1}^T \psi\left(X_t,\theta\right) =0
\label{Eq:EstimatingEquation}
\end{eqnarray}
and where $\psi(.,.)$ denotes the moment function s.t. $\mathbb{E} [ \psi(X_1, \theta_0) ]=0_{m \times 1}$  an $m$-dimensional vector of zeros,   $(X_t)_{t=1}^T$  i.i.d. data, $\theta_{0}\negmedspace\in\negmedspace \mathbf{\Theta}\negthickspace\subset \negthickspace\mathbf{R}^m$ the  unknown parameter of interest, and $T$ the sample size. The exact formula for  $\hat{f}_{\theta^{*}_T}(.)$ is reminded below in equation \eqref{Eq:ESPApproximationDefnMText} on p. \pageref{Eq:ESPApproximationDefnMText}.

The ESP estimator  is a moment-based estimator.  Since \cite{1894Pea,1902Pearson}'s  method of moment (MM),   moment-based estimators   have been found useful in a variety of applications (e.g., covariance structure analysis in psychology, and   asset pricing in economics). Their two main advantages are  (i) they do not require a parametric family of probability distributions for the data so they are less prone to model misspecification, and (ii) they allow complex models for which the likelihood function is  intractable.


Nevertheless, the  increase use of the MM and its extensions has     revealed that they  can be unstable   and  perform poorly    in  finite samples (e.g., July 1996 special issue of JBES). The idea of the ESP estimator to improve on the MM estimator is the following. By definition, the MM estimate $\theta^*_T(\omega)$  solves  a realization of the empirical moment conditions \eqref{Eq:EstimatingEquation}, but it typically does not solve  the empirical moment condition for another realization $\omega$ of the data. Thus, we might want an estimate that does only take into account the realized empirical moment conditions, but also their  other potential realizations. More precisely, we want an estimate that accounts for all the potential realizations of the empirical moment conditions according to  their probability weight of occurrence.   This leads to the ESP estimate, which is  a \textit{maximum-probability} estimate. The ESP estimate maximizes the estimated probability weights of solving the  empirical moment conditions.\footnote{This is in contrast to the traditional motivation for ML estimators, which  maximize the  probability weights  of obtaining a sample equal to the observed sample. In other words, the support of the ESP distribution is the parameter space, while the support of the distribution associated with a likelihood is the data space. Thus, if we are looking for relevant  parameter values instead of data values ---as it is typically the case---, a maximum-probability motivation appears more appealing than the traditional ML motivation. }   If the empirical  moment conditions \eqref{Eq:EstimatingEquation}  have a unique solution with a continuous distribution,  the ESP estimator maximizes the ESP approximation of a probability density function of the solution $\theta^*_T$.\footnote{Another motivation for maximum probability estimators is decision theoretic. Maximum probability estimators follows from the minimization of the expectation of   a loss ``function'' that equals zero  when $\theta$ solves the empirical moment conditions and one otherwise by normalization. This motivation is similar to the  decision-theoretic justification for the  Bayesian maximum a posteriori estimator \citep[e.g.,][sec. 4.1.2]{1994Robert}.  As in Bayesian analysis, the choice of other loss functions is possible. It is left for future research.} We rely on the ESP approximation  because  simulation and theoretical evidence shows the  ESP approximation can be  very accurate in small sample \citep[e.g.,][]{1988DavisonHinkley,1994RonWel}.




Besides the maximum-probability motivation, we show that the ESP estimator corresponds to an MM estimator shrunk toward parameter values with lower estimated variance. More precisely, we decompose the logarithm of the ESP approximation as the sum of a term, which is maximized at the  MM  estimator, and a variance penalty, which discounts  parameter values with high estimated variance.   Under   assumptions adapted from the entropy literature, we establish  the ESP estimator has the same good asymptotic properties as the MM estimator, so the variance penalization is a  finite-sample correction. We also    derive the ESP counterparts of the Wald, Lagrange multiplier (LM),  analogue likelihood-ratio (ALR) test statistics, as well as another test statistic. Then, we investigate  the ESP estimator through Monte-Carlo simulations.
We compare its performance with the exponential tilting (ET) estimator, which is equal to the MM estimator  in the just-identified case (i.e., when the number of parameters is the same as the number of moment conditions).   Results show that the variance penalization of the ESP estimator reduces the finite-sample instability of the ET estimator (or equivalently, of the MM estimator). An empirical application  illustrates the gain from this greater stability in terms of inference.

The ESP estimator is not the first proposal   to improve on the MM and its extensions. Alternative   moment-based approaches have been proposed    such as the empirical likelihood approach of Owen  \citep[][]{1994QinLawless}, the continuously updating approach \citep{1996HanHeaYar}, the already-mentioned exponential tilting (ET) approach \citep{1997KitStu,1998ImbSpaJoh}, and combinations of the aforementioned approaches \cite[e.g.,][]{2007Schennach}. All these approaches yield an estimator  closely related to the empirical likelihood estimator, so we call them empirical-likelihood-type estimators.  In  the just-identified case, when well-defined, all of these empirical-likelihood-type estimators are numerically equal to the original Pearson's  MM estimator  $\theta^*_T$.
  Because we focus on  the  just-identified case,  it thus is sufficient for us to compare the ESP estimator with the MM estimator, or with one of any   of these more recent estimators.








In addition to the already cited papers, the present paper, which supersedes the unpublished manuscript \cite{2009Sow}, is related to many other ones. We clarify these relations in Section \ref{Sec:Literature} (p. \pageref{Sec:Literature}). To the best of our knowledge, none of the prior papers use the SP or the ESP to propose a novel moment-based point estimator. Overall,  the present paper brings together the literature on the saddlepoint approximation and the literature on moment-based estimation.


\section{Finite-sample analysis}
In the present section, we remind the formula for the ESP approximation, and   analyze its finite-sample structure. Then, we decompose the log-ESP into two terms and show that  the ESP estimator is a MM estimator shrunk toward parameter values with lower estimated variance.
\subsection{The ESP approximation}\label{Sec:InsightsAboutESP}




  Formalizing and generalizing  prior works \citep{1988DavisonHinkley,1989Feu,1990Wang,1990YoungDaniels}, \cite{1994RonWel} propose the following ESP approximation to estimate the distribution of a solution to  the empirical moment conditions \eqref{Eq:EstimatingEquation}
\begin{eqnarray}
\hat{f}_{\theta^{*}_T}(\theta) :=
\exp\left\{T\ln\left[ \frac{1}{T}\sum_{t=1}^T\mathrm{e}^{\tau_T(\theta) ' \psi_t(\theta)}\right]\right\}\left(\frac{T}{2\pi}\right)^{m/2}\left|\Sigma_T(\theta) \right|_{\det}^{-\frac{1}{2}}
   \label{Eq:ESPApproximationDefnMText}
\end{eqnarray}
where $|.|_{\det}$ denotes the determinant function, $\theta^*_T$ a solution to  \eqref{Eq:EstimatingEquation},  $\psi_t(.):=\psi (X_t,.)$, and
\begin{eqnarray}
\Sigma_T(\theta) & := & \left[ \sum_{t=1}^T w_{t,\theta}\frac{\partial \psi_{t} (\theta)}{\partial \theta'}\right]^{-1}\left[\sum_{t=1}^T w_{t,\theta}\psi_{t}(\theta)\psi_{t}(\theta)'\right]\left[ \sum_{t=1}^T w_{t,\theta}\frac{\partial \psi_{t} (\theta)'}{\partial \theta}\right]^{-1}, \label{Eq:TiltedSigma_TMText}\\
w_{t,\theta} & := & \frac{\exp\left[\tau_T(\theta) ' \psi_t(\theta)\right]}{\sum_{i=1}^T\exp\left[\tau_T(\theta) ' \psi_i(\theta)\right]}  \text{ , }\label{Eq:TiltedWeightMText} \\
\tau_T(\theta) &\text{such that }&\sum_{t=1}^T \psi_{t}(\theta)\frac{\exp\left[\tau_T(\theta) ' \psi_t(\theta)\right]}{\sum_{i=1}^T\exp\left[\tau_T(\theta) ' \psi_i(\theta)\right]}\times \frac{1}{T}=0 \text{.  }\label{Eq:ESPTiltingEquationMText}
\end{eqnarray}

The ESP approximation \eqref{Eq:ESPApproximationDefnMText} is the empirical counterpart of the SP approximation  of \cite{1982Fie}.  From  a computational point of view, the ESP approximation \eqref{Eq:ESPApproximationDefnMText} is not  complicated.\footnote{We do not claim that the ESP estimator is as easy to compute as the MM estimator, but that its additional complexity is similar to the recently proposed empirical-likelihood-type estimators (e.g., ET estimator), and that it is worthwhile in several applications (e.g., Section \ref{Sec:Examples}). Moreover, it seems to make sense to develop novel estimation methods that take advantage of  the increasingly available computational power.} The only implicit quantity is $\tau_T(\theta)$, which solves the tilting equation \eqref{Eq:ESPTiltingEquationMText}, which, in turn,  is just the FOC (first-order condition) of the unconstrained convex problem  $\min_{\tau\in \mathbf{R}^m}\sum_{t=1}^T\mathrm{e}^{\tau ' \psi_t(\theta)}  $. A full understanding of the ESP approximation \eqref{Eq:ESPApproximationDefnMText} arguably requires to work  through higher-order asymptotic expansions along the lines of \cite{1982Fie}. However, direct inspection of the ESP approximation \eqref{Eq:ESPApproximationDefnMText}  also provides  insight for how it incorporates  information from the data through   two channels.

The first channel is the   \textit{ET (exponential tilting) term} $\exp\left\{T\ln\left[ \frac{1}{T}\sum_{t=1}^T\mathrm{e}^{\tau_T(\theta) ' \psi_t(\theta)}\right]\right\}  $.  In equation  \eqref{Eq:ESPTiltingEquationMText}, for  any $\theta\in \mathbf{\Theta} $, the terms $\frac{\exp\left[\tau_T(\theta) ' \psi_t(\theta)\right]}{\sum_{i=1}^T\exp\left[\tau_T(\theta) ' \psi_i(\theta)\right]} $ tilt (i.e., reweight) the empirical weights $ 1/T$, so the finite-sample moment conditions \eqref{Eq:ESPTiltingEquationMText} holds.  This tilting determines, through equation \eqref{Eq:TiltedWeightMText}, the multinomial distribution $(w_{t,\theta})_{t=1}^T$ that is the closest to the empirical distribution ---in the sense of the Kullback-Leibler divergence criterion--- s.t. the finite-sample moment conditions \eqref{Eq:ESPTiltingEquationMText} holds: The tilting equation \eqref{Eq:ESPTiltingEquationMText} is the FOC w.r.t. (with respect to) $\tau$ of the Lagrangian dual problem of  the minimization problem
\begin{eqnarray}
\min_{(w_{1,\theta},w_{2,\theta} ,\cdots,w_{T,\theta}) \in ]0,1]^T} \sum_{t=1}^T w_{t,\theta}\log\left(\frac{w_{t,\theta}}{1/T}\right) \nonumber
 \\\text{ s.t.} \sum_{t=1}^Tw_{t,\theta} \psi_t(\theta)=0 \text{ and } \sum_{t=1}^T w_{t,\theta}=1, \label{Eq:KLEmpMomentCond}
\end{eqnarray}
where $\sum_{t=1}^T w_{t,\theta}\log[w_{t,\theta}/(1/T)]$ is the Kullback-Leibler divergence criterion
between
the empirical distribution and the multinomial distribution $(w_{t,\theta})_{t=1}^T$ with the same support \cite[e.g.,][]{1981Efron,1997KitStu}. Then,
    for  the given $\theta\in \mathbf{\Theta} $, in the ESP approximation \eqref{Eq:ESPApproximationDefnMText}, the   ET  term $\exp\left\{T\ln\left[ \frac{1}{T}\sum_{t=1}^T\mathrm{e}^{\tau_T(\theta) ' \psi_t(\theta)}\right]\right\}  $ indicates the extent  of the tilting   needed to set the  finite-sample moment conditions \eqref{Eq:KLEmpMomentCond} (or equivalently,  equation \eqref{Eq:ESPTiltingEquationMText}) to zero.  The bigger  is the tilting of the empirical distribution, the less compatible are the data   with $\theta$ solving  the empirical moment conditions,  and the smaller should be the ET term $\exp\left\{T\ln\left[ \frac{1}{T}\sum_{t=1}^T\mathrm{e}^{\tau_T(\theta) ' \psi_t(\theta)}\right]\right\}  $. It can be easily seen that $\frac{1}{T}\sum_{t=1}^T\mathrm{e}^{\tau_T(\theta) ' \psi_t(\theta)}$ reaches its maximum when $\theta$ is a solution $\theta^*_T$ of  the empirical moment conditions \eqref{Eq:EstimatingEquation}, i.e., when      $\tau_T(\theta^*_T)=0_{m \times 1} $ and no tilting is needed.\footnote{For a complete proof, one can follow the same reasoning as in the proof of Lemma \ref{Lem:AsTiltingFct} in \citet[p. \pageref{Lem:AsTiltingFct}]{2019HolcblatSowell}  with the empirical distribution in lieu of $\mathbb{P}$.}

In the ESP approximation on equation \eqref{Eq:ESPApproximationDefnMText}, the second term $\left(\frac{T}{2\pi}\right)^{m/2} $ comes from the multivariate Gaussian distribution that  is the leading term of the Edgeworth's asymptotic expansions underlying ESP approximations. However,  because it is constant w.r.t. $\theta$, it does not affect the maximization of the ESP approximation, so it  is \textit{not} an information channel for the ESP estimator. The  remaining term $\left|\Sigma_T(\theta) \right|_{\det}^{-\frac{1}{2}}$   , which we call the \textit{variance term}, is the second channel through which the ESP approximation incorporates information from data. The variance term discounts the ET term according to the tilted estimated variance of the solution to the finite-sample moment conditions. Under standard assumptions, a consistent estimator of the asymptotic variance of   $\sqrt{T}(\theta^*_T-\theta_0)$  is $\left[ \frac{1}{T}\sum_{t=1}^T \frac{\partial \psi_{t} (\theta^*_T)}{\partial \theta}\right]^{-1}\left[\frac{1}{T}\sum_{t=1}^T \psi_{t}(\theta^*_T)\psi_{t}(\theta^*_T)'\right]\left[ \frac{1}{T}\sum_{t=1}^T \frac{\partial \psi_{t} (\theta^*_T)'}{\partial \theta}\right]^{-1} $.  The bigger the variance term is, the less plausible a solution takes exactly this value, and the smaller is $\left|\Sigma_T(\theta) \right|_{\det}^{-\frac{1}{2}}$  ---note the negative power. Therefore, overall, for a given  $\theta\in \mathbf{\Theta} $, the bigger the tilting  or  the estimated variance,  the smaller the ESP approximation, i.e., the estimated probability weight that $\theta$ solves the empirical moment conditions \eqref{Eq:EstimatingEquation}.

\subsection{The ESP estimator as a shrinkage estimator}\label{Sec:ESPvsMMandOthers} As explained in the introduction, the recently proposed moment-based estimators  are numerically equal to the Pearson's MM estimator in the just-identified case. Thus, it is sufficient to compare the ESP estimator with  one of them in order to understand the difference between the former and the other proposed moment-based estimators. The ET estimator of \cite{1997KitStu} and \cite{1998ImbSpaJoh} is particularly convenient for this purpose.
Taking the logarithm of the ESP approximation \eqref{Eq:ESPApproximationDefnMText}, and removing the terms constant w.r.t. $\theta$, it can be seen that,   $\mathbb{P}$-a.s. for $T$ big enough, the ESP estimator $\hat{\theta}_T$ maximizes the  objective function
\begin{eqnarray} \ln  \left[\frac{1}{T}\sum_{t=1}^T \mathrm{e}^{\tau_T(\theta)'\psi_t(\theta)}\right]- \frac{1}{2T} \ln \vert \Sigma_T(\theta)\vert_{\det}, \label{Eq:LogESP}
\end{eqnarray}
 where $\ln  \left[\frac{1}{T}\sum_{t=1}^T \mathrm{e}^{\tau_T(\theta)'\psi_t(\theta)}\right]$ is an increasing transformation of the objective function of the ET estimator. Thus, the difference between the ESP estimator and  the ET estimators  comes only from the log-variance term $-\frac{1}{2T} \ln \vert \Sigma_T(\theta)\vert_{\det}$. The latter does not only incorporates additional information from data as explained in Section \ref{Sec:InsightsAboutESP}, but  it also  penalizes parameter values with higher estimated variance. Thus, the ESP estimator is an ET estimator ---or equivalently, a MM estimator--- shrunk toward parameter values with lower estimated variance. Now, as the factor $\frac{1}{2T} $ suggests and the proofs of Section \ref{Sec:Asymptotic} show, the log-variance term vanishes asymptotically, so the shrinkage is a finite-sample correction.
\section{Asymptotic properties}\label{Sec:Asymptotic}

In the present section, we investigate the asymptotic properties of the ESP estimator.  Good asymptotic properties can be regarded as a
minimal requirement for the
ESP estimator, which is based on a small-sample asymptotic approximation.
 All the proofs and assumptions are in the online Appendix \cite{2019HolcblatSowell}.
\subsection{Existence, consistency and asymptotic normality }
  Under  assumptions adapted from the entropy literature,  the following theorem establishes the existence,  the strong consistency, and the asymptotic normality of the ESP estimator $\hat{\theta}_T$.

\begin{theorem}[Existence, consistency and asymptotic normality]\label{theorem:ConsistencyAsymptoticNormality} Under Assumption \ref{Assp:ExistenceConsistency}, $\mathbb{P}$-a.s. for $T$ big enough, there exists $\hat{\theta}_T$ s.t.
\begin{enumerate}
\item[(i)] $\mathbb{P}$-a.s. as $T \rightarrow \infty$, $\hat{\theta}_T \rightarrow \theta_0$ ; and
\item[(ii)] under the additional Assumption \ref{Assp:AsymptoticNormality}, as $T \rightarrow \infty$,  $\sqrt{T} ( \hat{\theta}_T  - \theta_0 )
 \underset{}{\stackrel{D}{\longrightarrow}}   \mathcal{N}\left(0, \Sigma(\theta_0)\right)$.

\end{enumerate}
where $\Sigma(\theta_0):=\negthickspace  \left[\mathbb{E}  \frac{\partial \psi(X_{1},\theta_{0})}{\partial \theta' }\right]^{-1}\negthickspace\negthickspace \mathbb{E}\left[\psi(X_{1},\theta_{0}) \psi(X_{1},\theta_{0})' \right] \left[ \mathbb{E}  \frac{\partial \psi(X_{1},\theta_{0})'}{\partial \theta }\right]^{-1} $, $\stackrel{D}{\rightarrow} $ denotes the convergence in distribution.
\end{theorem}

Theorem \ref{theorem:ConsistencyAsymptoticNormality} shows that the ESP estimator has the same first-order asymptotic properties  as the MM and hence the recently proposed moment-based estimators. Although the asymptotic properties of the ESP estimator are standard, the  proof of Theorem \ref{theorem:ConsistencyAsymptoticNormality} is quite involved. The crux of the proof is to show that the variance penalization  $-\frac{1}{2T} \ln \vert \Sigma_T(\theta)\vert_{\det}$  vanishes sufficiently quickly asymptotically, so it does not distort the first-order asymptotic.









\subsection{More on inference\,: The trinity$+1$}


The ESP estimator provides different ways to test  parameter restrictions
\begin{eqnarray}
\mathrm{H}_0: r(\theta_{0})=0_{q \times 1} \label{Eq:HypParameterRestriction}
\end{eqnarray}
where $r: \mathbf{\Theta} \rightarrow \mathbf{R}^q$ with $q \in [\![ 1, \infty[\![$.
More precisely, within the ESP framework,  there exist the usual trinity of  Wald, LM and ALR tests statistics, plus  another test statistic, which we call the exponential tilting (ET)  test statistic.  Our ET test has a structure similar to a test for over-identifyied moment conditions in \cite{1998ImbSpaJoh}.



Under a mild  standard additional assumption, the following theorem shows that the Wald, LM  ALR, and ET statistics asymptotically follow a chi-squared distribution with $q$ degrees of freedom.


\begin{theorem}[The trinity$+1$: Wald, LM, ALR  and ET tests] \label{theorem:TrinityPlus1} Define $R(\theta):=\frac{\partial r(\theta)}{\partial \theta'}$, and the following Wald, LM,  ALR and ET test statistics
\begin{eqnarray*}
\mathrm{Wald}_T& :=& Tr(\hat{\theta}_T)'[R(\hat{\theta}_T)\widehat{\Sigma(\theta_0)}_TR(\hat{\theta}_T)']^{-1}r(\hat{\theta}_T)\\
\mathrm{LM}_T &:=& T \check{\gamma}_T'[R(\check{\theta}_T)  \widehat{\Sigma(\theta_0)}_TR(\check{\theta}_T)']\check{\gamma}_T=\frac{\partial\ln[\hat{f}_{\theta^{*}_T}(\check{\theta}_T)]}{\partial \theta' }\widehat{\Sigma(\theta_0)}_T^{-1}\frac{\partial\ln[\hat{f}_{\theta^{*}_T}(\check{\theta}_T)]}{\partial \theta }\\
\mathrm{ALR}_T & := & 2\{\ln[\hat{f}_{\theta^{*}_T}(\hat{\theta}_T)]-\ln[ \hat{f}_{\theta^{*}_T}(\check{\theta}_T)]\}\\
\mathrm{ET}_T&:=& T \tau_T(\check{\theta}_T)'\widehat{V}_T\tau_T(\check{\theta}_T)
\end{eqnarray*}
where $\widehat{\Sigma(\theta_0)}_T$ and $\widehat{V}_T$ are symmetric matrices that converge in probability to  $\Sigma(\theta_0)$ and  $ \mathbb{E}[ \psi(X_1, \theta_0)\psi(X_1, \theta_0)'] $, respectively; and where $\check{\gamma}_T$ and $\check{\theta}_T$ respectively denote the Lagrange multiplier and a solution to the maximization of  $\hat{f}_{\theta^{*}_T}(\theta)$ w.r.t. $\theta\in \mathbf{\Theta}$ under the constraint that  $r(\theta)=0_{q \times 1}$.\footnote{In mathematical terms, $\check{\theta}_T\in \arg \max_{\theta \in \check{\Theta}} \hat{f}_{\theta^{*}_T}(\theta)$ where $\check{\Theta}:=\{\theta \in \Theta: r(\theta)=0_{q \times 1}\}$ and $\check{\gamma}_T$ is the Lagrangian multiplier s.t.  $\frac{1}{T} \frac{\partial \ln[\hat{f}_{\theta^{*}_T}(\check{\theta}_T)]}{\partial \theta} + \frac{\partial r(\check{\theta}_T)'}{\partial \theta}\check{\gamma}_T=0_{m \times 1}$.}
Under Assumptions \ref{Assp:ExistenceConsistency}, \ref{Assp:AsymptoticNormality} and \ref{Assp:Trinity}, if the test hypothesis \eqref{Eq:HypParameterRestriction} holds,  as $T \rightarrow \infty$,
\begin{eqnarray*}
\mathrm{Wald}_T, \mathrm{LM}_T, \mathrm{ALR}_T, \mathrm{ET}_T \stackrel{D}{\rightarrow } \chi^2_{q}.
\end{eqnarray*}

\end{theorem}



  Theorem \ref{theorem:TrinityPlus1} can also be used to obtain valid confidence regions by  the inversion of the  test statistics with  $\check{\theta}_T=\theta_0$.   Our  Wald, LM  and ALR test statistics share some similarity with the test statistics proposed in \cite{1997KitStu}, \cite{1998ImbSpaJoh} and \cite{2003RobRonYou}. The main difference is that the latter are built around (possibly constrained) maximizers of the ET term, while our tests statistics are based on the (possibly constrained) ESP estimator, which maximizes the whole ESP approximation including the variance term.






\section{Examples}\label{Sec:Examples}

In the present section, we  further investigate and illustrate the finite-sample properties of the ESP estimator.\footnote{In addition to our finite-sample analysis of the ESP objective function (Section \ref{sec1}), our derivation of the first-order asymptotic properties (Section \ref{Sec:Asymptotic}),    our Monte-Carlo simulations and  empirical application (present section),  another way  to shed light on the finite-sample properties of the ESP estimator would be to derive its higher-order asymptotic properties such as its second-order bias  \citep[e.g.,][]{1996RilstoneSrivastavaUllah}.
In the present paper, we do not follow this way because it would add several dozens of pages of proofs without much insight: Our preliminary derivations yield a long and complicated structure for the second-order bias, from which we struggle to gain insight. The length and the complexity of the second-order bias mainly comes from (i) the derivatives of the variance $\left|\Sigma_T(\theta) \right|_{\det}^{-1/2}$; and (ii) the reliance on the exact FOCs instead of approximate FOCs.   A mild preview of this complexity can be seen in \citet[Appendix \ref{Sec:LTAndDerivatives}]{2019HolcblatSowell}.}  We focus on the comparison  with the   ET estimator, as previously noted, (i) in the just-identified case, which is the case addressed in the present paper, the MM estimator and the recently proposed moment-based estimators are  equal to the ET estimator so  there is no loss of generality in terms of point estimation, and (ii) the ESP objective function nests the ET objective function, so that the source of the difference between the two is easily understood ---it necessarily comes from the variance term (see Section \ref{Sec:ESPvsMMandOthers}). For brevity,  we  present the main results for a numerical and an empirical example that are known to be challenging for moment-based estimation.


\subsection{Numerical example\,: Monte-Carlo simulations}\label{Sec:NumericalExp}

We simulate the just-identified version of the \cite{1996HalHor}  model, which has become  a standard benchmark to compare  the performance of moment-based estimators in statistics \citep[e.g.,][]{2007Schennach,2012LoRonchetti}   and econometrics \citep[e.g.,][]{1998ImbSpaJoh,2001Kitamura}. This model can be interpreted as
a simplified consumption-based asset pricing model where $\beta $ is the relative risk aversion (RRA) parameter \citep{2002GreLamSmi}.
  In the simulations, we estimate the two parameters $(\mu, \beta)$ with the  moment function
\begin{eqnarray*}
\psi_t(\beta , \mu ) = \left[ \begin{array}{c}
                        \exp\left\{\mu- \beta \left( X_{t} + Y_{t} \right) + 3 Y_{t}  \right\} - 1 \\
    Y_t  \left( \frac{}{} \exp\left\{\mu  - \beta \left( X_t + Y_t \right) + 3 Y_t  \right\} - 1  \right)
                      \end{array}
  \right]
\end{eqnarray*}
where  $\mu_0 = -.72 $, $ \beta_0 = 3$,
and
$X_{t}$ and $Y_{t}$ are jointly i.i.d. random variables with distribution $\mathcal{N}(0, .16)$.

\begin{table}[htp]\caption{\textbf{ESP vs. ET  estimator  for the just-identified Hall and Horowitz model.}     }\label{table:HH}
\begin{center}

\begin{tabular}{c r r r r r}
\hline
\hline
$T$ & & \multicolumn{2}{c}{$\beta$} & \multicolumn{2}{c}{$\mu$} \\
\hline
& & \multicolumn{1}{c}{ET} & \multicolumn{1}{c}{ESP}  & \multicolumn{1}{c}{ET} & \multicolumn{1}{c}{ESP}  \\
\cline{2-6}
 & MSE   & 3.6228  &  0.7065  &  1.5391  &  0.2319    \\
25 & Bias   & 0.4782  &  -0.0048  &  -0.1855  &   0.1089      \\
 & Var.   &  3.3941  &  0.7065  &  1.5047  &  0.2200      \\
\cline{2-6}
 & MSE    &   1.7024  &  0.3344  &  0.9959  &  0.1292     \\
50 & Bias   & 0.2670  &  -0.0160  &  -0.1330  &   0.0619       \\
 & Var.   &  1.6311  &  0.3342  &  0.9782  &  0.1254     \\
\cline{2-6}
 & MSE    & 0.6812  &  0.1742  &  0.4780  &  0.0645    \\
100 & Bias   &   0.1429  &  -0.0119  &  -0.0735  &   0.0388      \\
 & Var.   &  0.6608  &  0.1741  &  0.4726  &  0.0630      \\
\cline{2-6}
 & MSE    &  0.2162  &  0.0830  &  0.1457  &  0.0324    \\
200 & Bias   &  0.0684  &  -0.0113  &  -0.0340  &   0.0223     \\
 & Var.   & 0.2115  &  0.0829  &  0.1445  &  0.0319    \\ \hline
\end{tabular}

\end{center}
\raggedright \begin{flushleft}\begin{tiny} Note: The reported statistics are based on 10,000 simulated samples of sample size equal to the indicated $T$. For ET, the parameter space is restricted to $\beta <15$ in order to limit  the erratic behaviour of the estimator  at sample sizes $T=25$ and $50$. No  such parameter restriction is imposed for ESP.      \end{tiny}\end{flushleft}\begin{center}\begin{tiny}     \end{tiny} \end{center}
\end{table}


Table \ref{table:HH} reports the mean-squarred error (MSE), bias and variance of the ESP and ET estimators for different sample sizes.  The   MSE, the variance and the bias of the ESP estimator are always smaller than for the ET   estimator, and the differences are notable, especially for small sample sizes. In fact,
Table \ref{table:HH} understates the improvement delivered by the variance penalization of the ESP objective function. We help the ET estimator (or equivalently,  the MM estimator),\footnote{We numerically check that they deliver the same estimates even for the small sample sizes $T=25$ and $50$.  } by restricting its parameter space to $\beta< 15$. Without this parameter restriction, the behaviour of the ET estimator is  very unstable.
  An analysis of the typical shape of the objective functions for small sample size explains this phenomenon.    The typical ET objective function has a ridge that follows from around the population parameter values ($\beta_{0}=3$, $\mu_{0}=-.72$) towards $(1000, -600)$.  The  ridgeline is not totally flat, and it often   has a gentle downward slope as we move away from the area near the population parameter values.  However, regularly, for some simulated samples, the very top  of the ridge   is  extremely far from the population parameter values, so that  ET estimates  are very far from the population parameter values. This does not happen for the ESP estimator. The variance term of the ESP objective function  ensures that the ridge drops sufficiently as we move away from the maximum that is near the population parameter value. Thus, in line with our finite-sample analysis of the ESP objective (Section \ref{Sec:ESPvsMMandOthers}), the ESP estimator is much more stable.






\subsection{Empirical example} \label{Sec:EmpExample}



In this section, we present an empirical example from   asset pricing. Since \cite{1982HanSin}, moment-based estimation is standard in consumption-based asset pricing.  For brevity, we focus on the key features of the example.  See \citet[Appendix \ref{Ap:EmpiricalExample} on p. \pageref{Ap:EmpiricalExample}]{2019HolcblatSowell}  for  additional information and comparisons.
\begin{table}[ht!] \caption{\textbf{ET  vs. ESP inference (1890--2009)}} \label{Tab:ETvsESP1890Short}
 \medskip
 \begin{tabular}{ll}
\hline\multicolumn{2}{l}{Empirical moment condition: $\frac{1}{2009-1889}\sum_{t=1890}^{2009}\left[  \left(\frac{C_{t}}{C_{t-1}}\right)^{-\theta}(R_{m,t}-R_{f,t})\right]=0 $, where} \\
\multicolumn{2}{l}{$R_{m,t}:=$ gross market  return,\, $R_{f,t}:=$risk-free asset gross return,\, $C_t:=$ consumption, }\\
\multicolumn{2}{l}{ and $\theta:=$relative risk aversion;}\\
\multicolumn{2}{l}{\text{Normalized ET:=}$\exp\negthickspace\left\{T\ln\left[ \frac{1}{T}\sum_{t=1}^T\mathrm{e}^{\tau_T(.) ' \psi_t(.)}\right]\right\}\negthickspace/\negthickspace\int_{\Theta} \exp\negthickspace\left\{T\ln\left[ \frac{1}{T}\sum_{t=1}^T\mathrm{e}^{\tau_T(\theta) ' \psi_t(\theta)}\right]\right\}\mathrm{d} \theta$;    }\\
\multicolumn{2}{l}{\text{Normalized ESP:=}$ \hat{f}_{\theta^*_T}(.)/\negthickspace\int_{\Theta} \hat{f}_{\theta^*_T}(\theta)\mathrm{d} \theta$;    }\\
\multicolumn{2}{l}{$\hat{\theta}_{\mathrm{ET},T}=\hat{\theta}_{\mathrm{MM},T}=50.3$ (bullet) and $\hat{\theta}_{\mathrm{ESP},T}=32.21$ (bullet); }\\
\multicolumn{2}{l}{  ET and ESP  support $=[-218.2,289.0]$; 95\% ET ALR conf. region=$[18.3, 289.0] $ (stripe);  }\\
\multicolumn{2}{l}{   95\% ESP ALR conf. region=$[15.0,112.7]$ (stripe). }
  \\
\hline
\includegraphics[scale=.351]{1890_2009_ET_Norm_ALR.png} & \includegraphics[scale=.351]{1890_2009_ESP_Norm_ALR.png} \\
  \begin{small}(A) ET est. and ALR conf. region.\end{small} &  \begin{small}(B) ESP est. and ALR conf. region.\end{small} \\
\hline
\hline
\end{tabular}
\end{table}

We  estimate the relative risk aversion  (RRA) $\theta$ of a representative agent of the US economy.   Previous studies have shown that existing moment-based estimation approaches often produce   unstable  RRA  parameter estimates. We rely on the following   moment condition
\begin{eqnarray}
\mathbb{E}\left[\left(\frac{C_{t}}{C_{t-1}}\right)^{-\theta}(R_{m,t}-R_{f,t})\right]=0,\label{Eq:KeyEmpMomCond}
\end{eqnarray}
where $\frac{C_{t}}{C_{t-1}}$ is the growth consumption and $(R_{m,t}-R_{f,t})$ the market return in excess of the risk-free rate.
 The moment condition, which is common to many consumption-based asset pricing models, and the data are similar to \cite{2012JulGho} corresponding to standard US data at yearly frequency from Shiller's website spanning from 1890 to 2009. We report ET and ESP estimates as well as confidence regions based on the inversion of the ALR test statistics of Theorem \ref{theorem:TrinityPlus1} (p. \pageref{theorem:TrinityPlus1}) with $\check{\theta}_T=\theta_0$.  The latter have the advantage to take into account the whole shape of the objective function unlike $t$-statistics-based confidence regions, which only account
for the shape of the objective function in a neighborhood of the estimate  through its standard errors.





In Table \ref{Tab:ETvsESP1890Short}, Figures (A) and (B)  respectively display the ET term and the ESP approximation. For ease of comparison, the scale is the same, and we normalize both of them so they integrate to one. The normalized ET term is much   flatter around its maximum than the normalized ESP approximation. Flatness of the objective function  around the estimate has  been documented for other existing moment-based estimators, and it has often been regarded as one of the main sources of the instability of the RRA estimates  \citep[e.g.,][]{2000StoWri,2001NeeRoyWhi}. Figure (B)  shows that the normalized ESP is  sharp around the ESP estimator. The relative sharpness of the ESP yields  sharper confidence regions\,: The ESP confidence region is less than half its ET counterpart.  In light of the variance penalization term in the ESP objective function (Section \ref{Sec:ESPvsMMandOthers} on p. \pageref{Sec:ESPvsMMandOthers}) and the shrinkage-like behavior of the ESP estimator in  the Monte-Carlo simulations (Section \ref{Sec:NumericalExp}), the relative sharpness of the ESP inference is not surprising. In \cite{2019HolcblatSowell}, additional empirical evidences  corroborate  the increased stability and precision of the ESP estimator w.r.t. the ET estimator (or equivalently, MM estimator).







\section{Connection to the literature and further research directions}\label{Sec:Literature}


 The present paper demonstrates a previously unknown connection between  the SP approximation and   moment-based estimation, and hence it is related to many papers in addition to the ones already cited.
   Following  \cite{1954Dan},  the literature in
statistics
\citep[e.g.,][]{1986EastonRonchetti,1991Spady,1992Jen,2012LaVecchiaRonchettiTrojani, 2015BrodaKan,2018FasioloEtAl}   and econometrics \citep[e.g.,][]{1978Phi,1979HolPhi,1982Phi,1994Lieberman,2006AitSahYu}   has
used  the SP (saddlepoint) and ESP approximations  to obtain  accurate  approximations  of  distributions, especially in the tails.  The strand of the SP literature that is closest to our paper derives SP approximations to   the distribution of statistics that correspond to solutions of nonlinear estimating equations. The latter strand of literature started with \cite{1982Fie} and continued with \cite{1990Sko,1993MontiRonchetti,1997Imb,1998JenWoo,2000AlmFieRob,2003RobRonYou},  and
\cite{2003RonTro}, among others.
More recently, \cite{2010CzeRon}, \cite{2011MaRon}, and  \cite{2012LoRonchetti,2013KunRil,2015KundhiRilstone}   propose more accurate  tests for indirect inference, functional measurement error models, moment condition models, nonlinear estimators and GEL (generalized empirical likelihood) estimators, respectively.
To the best of our knowledge, unlike the present paper, none of the prior papers use the SP or the ESP to develop an estimation method that yields a novel moment-based estimator.
In ongoing work, we  generalize the ESP approximation  to the over-identified case, and  establish further good mathematical properties.


\bibliographystyle{kluwer}
\bibliography{general}



\section*{Notes and acknowledgements}
Parts of the present paper have previously circulated under the title ``The Empirical Saddlepoint Likelihood Estimator Applied to Two-Step GMM" \citep{2009Sow}. Some proofs of the present paper also borrow technical results from \cite{2012Hol}. Helpful comments were provided by Philipp Ketz (discussant), Eric Renault, Aman Ullah and seminar/conference participants at Carnegie Mellon University, CFE-CMStatistics 2017,  Swiss Finance Institute (EPFL and the University of Lausanne), 10th French Econometrics Conference (Paris School of Economics),  at the Econometric Society European Winter Meeting 2018 (University of Naples Federico II), and at the University of Luxembourg.
\newpage