EconBase
← Back to paper

An MCMC Approach to Classical Estimation

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

95,058 characters · 0 sections · 140 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

\lhead \rhead{\thepage} \cfoot \cfoot{\fancyplain} \thispagestyle{empty}

center[center omitted — 376 chars of source]
center[center omitted — 13 chars of source]
center[center omitted — 337 chars of source]

\hrule

center[center omitted — 31 chars of source]

\baselineskip=.88\baselineskip { This paper studies computationally and theoretically attractive estimators called the Laplace type estimators (LTE), which include means and quantiles of Quasi-posterior distributions defined as transformations of general (non-likelihood-based) statistical criterion functions, such as those in GMM, nonlinear IV, empirical likelihood, and minimum distance methods. The approach generates an alternative to classical extremum estimation and also falls outside the parametric Bayesian approach. For example, it offers a new attractive estimation method for such important semi-parametric problems as censored and instrumental quantile, nonlinear GMM and value-at-risk models. The LTE's are computed using Markov Chain Monte Carlo methods, which help circumvent the computational curse of dimensionality. A large sample theory is obtained for regular cases.}

{\it JEL Classification:} C10, C11, C13, C15

{\it Keywords:} Laplace, Bayes, Markov Chain Monte Carlo, Sandwich Formula, GMM, instrumental regression, censored quantile regression, instrumental quantile regression, empirical likelihood, value-at-risk

\baselineskip=1.1\baselineskip

\hrule

\@startsection{section}{1} {\z@}{-2.5ex plus -0.5ex minus -0.1ex}{0.5ex plus 0.1ex}{Introduction}

A variety of important econometric problems pose not only a theoretical but a serious computational challenge, cf. andrews:comput. A small (and by no means exhaustive) set of such examples include (1) Powell's censored median regression for linear and nonlinear problems, (2) nonlinear IV estimation, e.g in the blp1 model, (3) the instrumental quantile regression, (4) the continuous-updating GMM estimator of cue, and related empirical likelihood problems. These problems represent a formidable practical challenge as the extremum estimators are known to be difficult to compute due to highly nonconvex criterion functions with many local optima (but well pronounced global optimum). Despite extensive efforts, see notably Andrews (1997), the problem of extremum computation remains a formidable impediment in these applications.

This paper develops a class of estimators referred to as the Laplace type estimators (LTE) or Quasi-Bayesian estimators (QBE),\footnote{ A preferred terminology is taken to be the `Laplace Type Estimators', since the term `Quasi-Bayesian Estimators' is already used to name Bayesian procedures that use either 'vague' or 'data-dependent' prior or multiple priors, cf. berger:review.} which are defined similarly to Bayesian estimators but use general statistical criterion functions in place of the parametric likelihood function. This formulation circumvents the curse of dimensionality inherent in the computation of the classical extremum estimators by instead focusing on LTE which are functions of integral transformations of the criterion functions and can be computed using Markov Chain Monte Carlo methods (MCMC), a class of simulation techniques from Bayesian statistics. This formulation will be shown to yield both computable and theoretically attractive new estimators to such important problems as (1)-(4) listed above. Although the aforementioned applications are mostly microeconometric, the obtained results extend to many other models, including GMM and quasi-likelihoods in the nonlinear dynamic framework of gallant:white.

The class of LTE's or QBE's aim to explore the use of the Laplace approximation (developed by Laplace to study large sample approximations of Bayesian estimators and for use in other non-statistical problems) outside of the canonical Bayesian framework -- that is, outside of parametric likelihood settings when the likelihood function is not known. Instead, the approach relies upon other statistical criterion functions of interest in place of the likelihood, transforms them into proper distributions -- Quasi-posteriors -- over a parameter of interest, and defines various moments and quantiles of that distribution as the point estimates and confidence intervals, respectively. It is important to emphasize that the underlying criterion functions are mainly motivated by the analogy principle in place of the likelihood principle, are not the likelihoods (densities) of the data, and are most often semi-parametric.\footnote{In this paper, the term `semi-parametric' refers to the cases where the parameters of interest are finite-dimensional but there are nonparametric nuisance parameters such as unspecified distributions.}

The resulting estimators and inference procedures possess a number of good theoretical and computational properties and yield new, alternative approaches for the important problems mentioned earlier. The estimates are as efficient as the extremum estimates; and, in many cases, the inference procedures based on the quantiles of the Quasi-posterior distribution yield asymptotically valid confidence intervals, which also perform notably well in finite samples. For example, in the quantile regression setting, those intervals provide valid large sample and excellent small sample inference without requiring nonparametric estimation of the conditional density function (needed in the standard approach). The obtained results are general and useful -- they cover the examples listed above under general, non-likelihood based conditions that allow discontinuous, non-smooth semi-parametric criterion functions, and data generating processes that range from iid settings to the nonlinear dynamic framework of gallant:white. The results thus extend the theoretical work on large sample theory of Bayesian procedures in econometrics and statistics, e.g. bickel2, ibragimov, andrews:odds, kim:ecca.

The LTE's are computed using MCMC, which simulates a series of parameter draws such that the marginal distribution of the series is (approximately) the Quasi-posterior distribution of the parameters. The estimator is therefore a function of this series, and may be given explicitly as the mean or a quantile of the series, or implicitly as the minimizer of a smooth globally convex function.

As stated above, the LTE approach is motivated by the estimation and inference efficiency as well as computational attractiveness. Indeed, the LTE approach is as efficient as the extremum approach, but generally may not suffer from the computational curse of dimensionality (through the use of MCMC). The reason is that the computation of LTE's is itself statistically motivated. LTE's are typically means or quantiles of a quasi-posterior distribution, hence can be estimated (computed) at the parametric rate $1/\sqrt{B}$, where $B$ is the number of draws from that distribution (functional evaluations). In contrast, the mode (extremum estimator) is estimated (computed) by the MCMC and similar grid-based algorithms at the nonparametric rate $(1/B)^{\frac{p}{d+2p}}$, where $d$ is the parameter dimension and $p$ is the smoothness order of the objective function.

Another useful feature of LT estimation is that, by using information about the overall shape of the objective function, point estimates and confidence intervals may be calculated simultaneously. It also allows incorporation of prior information, and allows for a simple imposition of constraints in the estimation procedure.

The remainder of the paper proceeds as follows. Section (ref) formally defines and further motivates the Laplace type estimators with several examples, reviews the literature, and explains other connections. The motivating examples, which are all semi-parametric and involve no parametric likelihoods, will justify the pursuit of a more general theory than is currently available. Section (ref) develops the large sample theory, and Sections 3 and 4 further explore it within the context of the econometric examples mentioned earlier. Section 4 briefly reviews important computational aspects and illustrates the use of the estimator through simulation examples. Section 5 contains a brief empirical example, and Section 6 concludes.

Notation. Standard notation is used throughout. Given probability measure $P$, ${\rightarrow}_p \ $ denotes the convergence in (outer) probability with respect to the outer probability $P^*$; ${\rightarrow}_d \ $ denotes the convergence in distribution under $P^*$, etc. See e.g. vaart for definitions. $|x|$ denotes the Euclidean norm $\sqrt{x'x}$; $B_{\delta}(x)$ denotes the ball of radius $\delta$ centered at $x$. A notation table is given in the appendix.

\@startsection{section}{1} {\z@}{-2.5ex plus -0.5ex minus -0.1ex}{0.5ex plus 0.1ex}{Laplacian or Quasi-Bayesian Estimation: Definition and Motivation}

\@startsection{subsection}{2} {\z@}{-2.25ex plus -0.3ex minus -0.2ex}{0.05ex plus 0.05ex}{Motivation} Extremum estimators are usually motivated by the analogy principle and defined as maximizers of random average-like criterion functions $L_n\left(\theta\right)$, where $n$ denotes the sample size. $n ^{-1} L_n\left(\theta\right)$ are typically viewed as transformations of sample averages that converge to criterion functions $M\left(\theta\right)$ that are maximized uniquely at some $\theta_0$. Extremum estimators are usually consistent and asymptotically normal, cf. amen, gallant:white, newey_mcfadden, potscher. However, in many important cases, actually computing the extremum estimates remains a large problem, as discussed by andrews:comput.

Example 1: Censored and Nonlinear Quantile Regression. A prominent model in econometrics is the censored median regression model of powell_lad. Powell's censored quantile regression estimator is defined to maximize the following nonlinear objective function

align[align omitted — 187 chars of source]

where $\rho_{\tau}(u) = \left( \tau - 1(u< 0)\right) u$ is the check function of koenker:1978, $\omega_i$ is a weight, and $Y_i$ is either positive or zero. Its conditional quantile $q(X_i, \theta)$ is specified as $\max(0, g(X_i,\theta))$. The censored quantile regression model was first formulated by powell_lad as a way to provide valid inference in Tobin-Amemiya models without distributional assumptions and with heteroscedasticity of unknown form. The extremum estimator based on the Powell's criterion function, while theoretically elegant, has a well-known computational difficulty. The objective function is similar to that plotted in Figure 1 -- it is nonsmooth and highly nonconvex, with numerous local optima, posing a formidable obstacle to the practical use of this extremum estimator; see buchinsky, buchinsky:phd, fitzenberger1, and khan:1998 for related discussions. In this paper, we shall explore the use of LT estimators based on Powell's criterion function and show that this alternative is attractive both theoretically and computationally.

Example 2: Nonlinear IV and GMM. amemiya:IV, hansen:gmm, cue introduced nonlinear IV and GMM estimators that maximize

align[align omitted — 203 chars of source]

where $m_i\left(\theta\right)$ is a moment function defined such that the economic parameter of interest solves $$ E m_i(\theta_0) =0. $$ The weighting matrix may be given by $ W_n(\theta)=\left[\frac{1}{n}\sum_{i=1}^nm_i\left(\theta\)m_i\left(\theta\right)'\right]^{-1} + o_p(1)$ or other sensible choices. Note that the term “$o_p(1)$" in $L_n$ implicitly incorporates generalized empirical likelihood estimators, which will be discussed in section 4. Up to the first order, objective functions of empirical likelihood estimators for $\theta$ (with the Lagrange multiplier concentrated out) locally coincide with $L_n$. Applications of these estimators are numerous and important (e.g. blp1, cue, imbens:el), but while global maxima are typically well-defined, it is also typical to see many local optima in applications. This leads to serious difficulties with applying the extremum approach in applications where the parameter dimension is high. As in the previous example, LTE's provide a computable and theoretically attractive alternative to extremum estimators. Furthermore, Quasi-posterior quantiles provide a valid and effective way to construct confidence intervals and explore the shape of the objective function.

Example 3: Instrumental and Robust Quantile Regression. Instrumental quantile regression may be defined by maximizing a standard nonlinear IV or GMM objective function\footnote{Early variants based on the Wald instruments go back to mood and hogg, cf. koenker:1998. }

align[align omitted — 197 chars of source]

where

align[align omitted — 129 chars of source]

$Y_i$ is the dependent variable, $D_i$ is a vector of possibly endogeneous variables, $X_i$ is a vector of regressors, $Z_i$ is a vector of instruments, and $W_n\left(\theta\right)$ is a positive definite weighting matrix, e.g. $$ W_n(\theta)=\left[\frac{1}{n}\sum_{i=1}^nm_i\left(\theta\)m_i\left(\theta\right)'\right]^{-1} + o_p(1) \ \text{ or } \ W_n(\theta) = \frac{1}{\tau (1-\tau)} \left[\frac{1}{n}\sum_{i=1}^n Z_iZ_i'\right]^{-1}, $$ or other sensible versions. Motivations for estimating equations of this sort arise from traditional separable simultaneous equations, cf. amen, and also more general nonseparable simultaneous equation models and heterogeneous treatment effect models.\footnote{See iqr for the development of this direction. }

Clearly, a variety of huber2 type robust estimators can be defined in this way. For example, suppose in the absence of endogeneity $$ q(X,\theta) = X'\beta(\tau), $$ then $Z= f(X)$ can be constructed to preclude the influence of outliers in $X$ on the inference. For example, choosing $Z_{ij} =1 ( X_{ij} < x_j ), \ j=1,...,\dim(X)$, where $x_j$ denotes the median of $\{X_{ij}, i \leq n \}$ produces an approach that is similar in spirit to the maximal regression depth estimator of regdepth, whose computational difficulty is well known, as discussed in depreg:comput. The resulting objective function $L_n(\beta)$ is highly robust to both outliers in $X_{ij}$ and $Y_i$. In fact, it appears that the breakdown properties of this objective function are similar to those of the objective function of regdepth.

Despite a clear appeal, the computational problem is daunting. The function $L_n$ is highly non-convex, almost everywhere flat, and has numerous discontinuities and local optima.\footnote{macurdy_timmins propose to smooth out the edges using kernels, however this does not eliminate non-convexities and local optima; see also abadie:1995.} (Note that the global optimum is well pronounced.) Figure 1 illustrates the situation. Again, in this case the LTE approach will yield a computable and theoretically attractive alternative to the extremum-based estimation and inference.\footnote{Another computationally attractive approach, based on an extension of koenker:1978 quantile regression estimator to instrumental problems like these, is given in Chernozhukov and Hansen(2001). } Furthermore, we will show that the Quasi-posterior confidence intervals provide a valid and effective way to construct confidence intervals for parameters and their smooth functions without non-parametric estimation of the conditional density function evaluated at quantiles (needed in standard approach). $\square$\\

figure[figure omitted — 947 chars of source]

The LTE's studied in this paper can be easily computed through Markov Chain Monte Carlo and other posterior simulation methods. To describe these estimators, note that although the objective function $L_n\left(\theta\right)$ is generally not a log-likelihood function, the transformation

align[align omitted — 212 chars of source]

is a proper distribution density over the parameter of interest, called here the Quasi-posterior. Here $\pi\left(\theta\right)$ is a weight or prior probability density that is strictly positive and continuous over $\Theta$, for example, it can be constant over the parameter space. Note that $p_n$ is generally not a true posterior in the Bayesian sense, since it may not involve the conditional data density or likelihood, and is thus generally created through non-Bayesian statistical learning.

The Quasi-posterior mean is then defined as

align[align omitted — 288 chars of source]

where $\Theta$ is the parameter space. Other quantities such as medians and quantiles will also be considered. A formal definition of LTE's is given in Definition (ref).

In order to compute these estimators, using Markov Chain Monte Carlo methods, we can draw a Markov chain (see Figure 1), $$S = \left(\theta^{(1)}, \theta^{(2)}, ..., \theta^{({\scriptscriptstyle B})} \right),$$ whose marginal density is approximately given by $p_n(\theta)$, the Quasi-posterior distribution. Then the estimate $\widehat \theta$, e.g. the Quasi-posterior mean, is computed as

align[align omitted — 103 chars of source]

Analogously, for a given continuously differentiable function $g: \Theta \to \Bbb{R}$, the $90\%$-confidence intervals are constructed simply by taking the $.05$-th and $.95$-th quantiles of the sequence $$ g(S) = \(g(\theta^{(1)}), ..., g(\theta^{({\scriptscriptstyle B})}) \right), $$ see Figure (ref). Under the information equality restrictions discussed later, such confidence regions are asymptotically valid. Under other conditions, it is possible to use other Quasi-posterior quantities such as the variance-covariance matrix of the series $S$ to define asymptotically valid confidence regions, see Section 3. It shall be emphasized repeatedly that the validity of this approach does not depend on the likelihood formulation.

\@startsection{subsection}{2} {\z@}{-2.25ex plus -0.3ex minus -0.2ex}{0.05ex plus 0.05ex}{Formal Definitions}

Let $\rho_n\(u\right)$ be a penalty or loss function associated with making an incorrect decision. Examples of $\rho_n\(u\right)$ include

itemize$\rho_n\(u\right) = \vert \sqrt{n} u\vert^2$, \ \ the squared loss function, • $\rho_n\(u\right) = \sqrt{n}\sum_{j=1}^d \vert u_j\vert$, \ \ the absolute deviation loss function, • $\rho_n\(u\right) = \sqrt{n} \sum_{j=1}^d \left(\tau_j-1\(u_j\leq 0\right)\)u_j$, for $\tau_j \in (0,1)$ for each $j$, the check loss function of koenker:1978.

The parameter is assumed to belong to the subset $\Theta$ of Euclidean space. Using the Quasi-posterior $p_n$ density in ((ref)), define the Quasi-posterior risk function as:

align[align omitted — 372 chars of source]
definitionThe class of LTE minimize the function $Q_n(\zeta)$ in ((ref)) for various choices of $\rho_n$: \begin{align}\begin{split} \hat\theta &= \arg\inf_{\zeta\in\Theta} \Big [ Q_n(\zeta) \Big ]. \end{split}\end{align}

The estimator $\widehat \theta $ is a decision rule that is least unfavorable given the statistical (non-likelihood) information provided by the probability measure $p_n$, using the loss function $\rho_n$. In particular, the loss function $\rho_n$ may asymmetrically penalize deviations from the truth, and $\pi$ may give differential weights to different values of $\theta$. The solutions to the problem ((ref)) for loss functions i-iii include the Quasi-posterior means, medians, and marginal $\tau_j$-th quantiles, respectively.\footnote{ This formulation implies that conditional on the data, the decision $\widehat \theta$ satisfies Savage's axioms of choice under uncertainty with subjective probabilities given by $p_n$ (these include the usual asymmetry and negative transitivity of strict preference relationship, independence, and some other standard axioms).}

\@startsection{subsection}{2} {\z@}{-2.25ex plus -0.3ex minus -0.2ex}{0.05ex plus 0.05ex}{Related Literature} Our analysis will rely heavily on the previous work on Bayesian estimators in the likelihood setting. The initial large sample work on Bayesian estimators was done by Laplace (see stigler:laplace for a detailed review). Further early work of bernstein1 and vonmises1 has been considerably extended in both econometric and statistical research, cf. ibragimov, bickel2, andrews:odds, phillips:ploberger, and kim:ecca, among others.

In general, Bayesian asymptotics require very delicate control of the tail of the posterior distribution and were developed in useful generality much later than the asymptotics of extremum estimators. The treatments of bickel2 and ibragimov are most useful for the present setting, but are inevitably tied down to the likelihood setting. For example, the latter treatment relies heavily on Hellinger bounds that are firmly rooted in the objective function being a likelihood of iid data. However, the general flavor of the approach is suited for the present purposes. The treatment of bickel2 can be easily extended to smooth, possibly incorrect iid likelihoods,\footnote{See pseudoBayes for an extension to the more than three times differentiable smooth misspecified iid likelihood case. The conditions do not apply to GMM or even Example 1. } but does not apply to censored median regression or any of the GMM type settings. andrews:odds and phillips:ploberger study the large sample approximation of posteriors and posterior odds ratio tests in relation to the classical Wald tests in the context of smooth, correctly specified likelihoods. kim:ecca derives the limit behavior of posteriors in likelihood models over shrinking neighborhood systems. Kim's approach and related approaches have been important in describing the essence of posterior behavior, but the limit behavior of point estimates like ours does not follow from it.\footnote{E.g. the approach does not characterize the behavior of the posterior mean: $\int_{-\infty}^{\infty} \theta p_n(\theta) d \theta$, which also requires the study of the complete $L_n(\theta)$ and the analysis of convergence of the posterior in the total variation of moments norm.}

Formally and substantively, none of the above treatments apply to our motivating examples and the estimators given in Definition 1. These examples do not involve likelihoods, deal mostly with GMM type objective functions, and often involve discontinuous and non-smooth criterion functions to which the above mentioned results do not apply. In order to develop the theory of LTE's for such examples, we extend the previous arguments. The results obtained here enable the use of Bayesian tools outside of the Bayesian framework -- covering models with non-likelihood-based criterion functions, such as examples listed earlier and other semi-parametric objective functions that may, for example, depend on preliminary estimates of infinite-dimensional nuisance parameters. Moreover, our results apply to general forms of data generating processes -- from the cross-sectional framework of amen to the nonlinear dynamic framework of gallant:white and potscher.

Our motivating problems are all semi-parametric, and there are several pure Bayesian approaches to such problems, see notably doksum, freedman, hahn:bb, imbens:chamberlain, kottas. Semi-parametric models have some parametric and nonparametric components, e.g. the unspecified nonparametric distribution of data in Examples 1-3. The mentioned papers proceed with the pure Bayesian approach to such problems, which involves Bayesian learning about these two components via a two-step process. In the first step, Bayesian non-parametric learning with Dirichlet priors is used to form beliefs about the joint non-parametric density of data, and then draws of the non-parametric density (“Bayesian bootstrap") are made repeatedly to compute the extremum parameter of interest. This approach is purely Bayesian, as it fully conforms to the Bayes learning model. It is clear that this approach is generally quite different from LTE's or QBE's studied in this paper, and in applications, it still requires numerous re-computations of the extremum estimates in order to construct the posterior distribution over the parameter of interest. In sharp contrast, the LT estimation takes a “shortcut" by essentially using the common criterion functions as posteriors, and thus entirely avoids both the estimation of the nonparametric distribution of the data and the repeated computation of extremum estimates.

Finally note that the LTE approach has a limited-information or semi-parametric nature in the sense that we do not know or are not willing to specify the complete data density. The limited-information principle is powerfully elaborated in the recent work of zellner:maxent, who starts with a set of moment conditions, calculates the maximum entropy densities consistent with the moment equations, and uses those as formal (misspecified) likelihoods. While in the present framework, calculation of the maximum entropy densities is not needed, the large sample theory obtained here does cover Zellner's (zellner:maxent) estimators as one fundamental case. Related work by kim2 derives a limited information likelihood interpretation for certain smooth GMM settings.\footnote{kim2 also provided some useful asymptotic results for $\exp(L_n(\theta) )$ using the shrinking neighborhood approach. However, Kim's (kim2) approach does not cover the estimators and procedures considered here, see previous footnote. } In addition, the LTE's based on the empirical likelihood are introduced in Section 4 and motivated there as respecting the limited information principle.

\@startsection{section}{1} {\z@}{-2.5ex plus -0.5ex minus -0.1ex}{0.5ex plus 0.1ex}{Large Sample Properties}

This section shows that under general regularity conditions the Quasi-posterior distribution concentrates at the speed $1/\sqrt{n}$ around the true parameter $\theta_0$ as measured by the “total variation of moments" norm (and total variation norm as a special case), that the LT estimators are consistent and asymptotically normal, and that Quasi-posterior quantiles and other relevant quantities provide asymptotically valid confidence intervals.

\@startsection{subsection}{2} {\z@}{-2.25ex plus -0.3ex minus -0.2ex}{0.05ex plus 0.05ex}{Assumptions} We begin by stating the main assumptions. In addition, it is assumed without further notice that the criterion functions $L_n(\theta)$ and other primitive objects have the standard measurability properties. For example, given the underlying probability space $(\Omega, \mathcal{F}, P)$, for any $\omega \in \Omega$, $L_n(\theta)$ is a measurable function of $\theta$, and for any $\theta \in \Theta$, $L_n(\theta)$ is a random variable, that is a measurable function of $\omega$.

assumption[Parameter] The true parameter $\theta_0$ belongs to the interior of a compact convex subset $\Theta$ of Euclidean space $\Bbb{R}^d$.
assumption[Penalty Function] The loss function $\rho_n: \Bbb{R}^d \rightarrow \Bbb{R}_+$ satisfies: \begin{itemize} • $\rho_n(u)=\rho( \sqrt{n} u )$, where $\rho(u) \geq 0$ and $\rho(u) = 0$ iff $u=0$, • $\rho$ is convex and $\rho(h) \leq 1 + | h |^p$ for some $p \geq 1$, • $\varphi(\xi) = \int_{\Bbb{R}^d} \rho(u-\xi) e^{\scriptscriptstyle -u'a u} du$ is minimized uniquely at some $\xi^* \in \Bbb{R}^d$ for any finite $a >0$, • the weighting function $\pi: \Theta \rightarrow \Bbb{R}_+$ is a continuous, uniformly positive density function. \end{itemize}
assumption[Identifiability] For any $\delta > 0$, there exists $\epsilon > 0$, such that \begin{align}\begin{split}\nonumber \liminf_{n \rightarrow \infty }P_*\left \{\sup_{\vert \theta - \theta_0\vert \geq \delta} \frac1n \Big (L_n\left(\theta\right) - L_n\left(\theta_0\right)\Big ) \leq -\epsilon \right \} = 1. \end{split}\end{align}
assumption[Expansion] For $\theta$ in an open neighborhood of $\theta_0$, \begin{itemize} • $L_n\left(\theta\right) - L_n\left(\theta_0\right) = \left(\theta-\theta_0\right)'\Delta_n\left(\theta_0\right) -\frac12 \left(\theta - \theta_0\right)' \[n J_n\left(\theta_0\right)\right]\left(\theta - \theta_0\right) + R_n\left(\theta\right)$, • $\Omega_n^{\scriptscriptstyle -1/2}(\theta_0) \Delta_n\left(\theta_0\right)/\sqrt{n} \ {\rightarrow}_d \ \mathcal{N}\left(0, I\right)$, • $J_n\left(\theta_0\right)=O(1)$ and $\Omega_n(\theta_0)=O(1)$ are uniformly in $n$ positive-definite constant matrices, • for each $\epsilon > 0$ there is a sufficiently small $\delta >0$ and large $M>0$ such that \begin{align}\begin{split}\nonumber \textrm{\textbf{(a)}} \ \ \limsup_{n \to \infty} \ &P^* \left \{ \sup_{ M/\sqrt{n}\leq |\theta - \theta_0 | \leq \delta } \frac{ | R_n\left(\theta\right) |} { n | \theta - \theta_0 |^2}> \epsilon \right \} < \epsilon, \\ \textrm{\textbf{(b)}} \ \ \limsup_{n \to \infty} \ &P^* \left \{ \sup_{ |\theta - \theta_0 | \leq M/\sqrt{n}}\ \ \ \ | R_n\left(\theta\right) | \ > \epsilon \ \right \} = 0. \end{split}\end{align} \end{itemize}

\@startsection{subsection}{2} {\z@}{-2.25ex plus -0.3ex minus -0.2ex}{0.05ex plus 0.05ex}{Discussion of Assumptions} In the following we discuss the stated assumptions under which Theorem 1-4 stated below will be true. We argue that these assumptions are simple but encompass a wide variety of econometric models -- from cross-sectional models to nonlinear dynamic models. This means that Theorems 1-4 are of wide interest and applicability.

In general, Assumptions 1-4 are related to but different from those in bickel2 and ibragimov. The most substantial differences appear in Assumption 4, and are due to the general non-likelihood setting. Also in Assumption 4 we introduce Huber type conditions to handle the tail behavior of discontinuous and non-smooth criterion functions. In general, the early approaches are inevitably tied to the iid likelihood formulation, which is not suited for the present purposes.

The compactness Assumption (ref) is conventional. It is shown in the proof of Theorem 1 that it is not difficult to drop compactness. For example, in the case of Quasi-posterior quantiles in Theorem 3, it is only required that $\pi$ is a proper density; in the case of Quasi-posterior variances in Theorem 4, it is only required that $\int_{\Theta}|\theta|^2 \pi(\theta) d\theta < \infty$; and for the general loss functions considered in Theorem 2 it is required that $\int_{\Theta}|\theta|^p \pi(\theta) d\theta < \infty$. Of course, compactness guarantees all of the above given Assumption 2. Also, that the parameter is on the interior of the parameter space rules out some non-regular cases; see for example andrews_est.

Assumption (ref) imposes convexity on the penalty function. We do not consider non-convex penalty functions for pragmatic reasons. One of the main motivations of this paper is the generic computability of the estimates, given that they solve well-defined convex optimization problems. The domination condition, $\rho(h) \leq 1 + | h |^p$ for some $1 \leq p < \infty$, is conventional and is satisfied in all examples of $\rho$ we gave.

The assumption that $\varphi(\xi)= \int \rho(u-\xi) e^{\scriptscriptstyle -u'a u} du \propto E \rho( \mathcal{N}(0, a^{-1}) -\xi) $ attains a unique minimum at some finite $\xi^*$ for any positive definite $a$ is required, and it clearly holds for all of examples of $\rho$ we mentioned. In fact, when $\rho$ is symmetric, $\xi^*=0$ by anderson_lemma's lemma.

Assumption (ref) is implied by the usual uniform convergence and unique identification conditions as in amen. The proof of Lemma (ref) can be found in amen and white:QML.

lemmaGiven Assumption 1, Assumption (ref) holds if there is a function $M_n(\theta)$ that \begin{itemize} • is nonstochastic, continuous on $\Theta$, for any $\delta > 0$, $\limsup_n ( \sup_{|\theta - \theta_0|>\delta} M_n(\theta) - M_n(\theta_0) ) < 0$, • $ L_n\left(\theta\right)/n -M_n(\theta) $ converges to zero in (outer) probability uniformly over $\Theta$. \end{itemize}

Assumption 4 is satisfied under the conditions of Lemma 2, which are known to be mild in nonlinear models. Assumption (ref).ii requires asymptotic normality to hold, and is generally a weak assumption for cross-sectional and many time-series applications. Assumption (ref).iii rules out the cases of mixed asymptotic normality for some non-stationary time series models (which can be incorporated at a notational cost with different scaling rates).

Assumption (ref).iv easily holds when there is enough smoothness.

lemmaGiven Assumptions 1 and 3, Assumption (ref) holds with $$ \Delta_n(\theta_0) = \nabla_{\theta} L_n(\theta_0) \text{ and } J_n(\theta_0) = - \nabla_{\theta \theta'} M_n\left(\theta_0\right) =O(1), \text{ if } $$ \begin{itemize} • for some $\delta > 0$, $L_n\left(\theta\right)$ and $M_n\left(\theta\right)$ are twice continuously differentiable in $\theta$ when $\vert \theta - \theta_0\vert < \delta$, • there is $\Omega_n(\theta_0)$ such that $\Omega_n^{\scriptscriptstyle -1/2}(\theta_0)\nabla_{\theta} L_n(\theta_0)/\sqrt{n} \ {\rightarrow}_d \ \mathcal{N}(0, I)$, \ \ $J_n(\theta_0)=O(1)$ and $\Omega_n(\theta_0) =O(1)$ are uniformly positive definite, and • for some $\delta>0$ and each $\epsilon >0$ $$\limsup_{n \rightarrow \infty} P^*\left \{ \sup_{| \theta -\theta_0| < \delta} | \nabla_{\theta \theta'} L_n\left(\theta\right)/n - \nabla_{\theta \theta'} M_n\left(\theta\right) | > \epsilon\right \} =0.$$ \end{itemize}

Lemma 2 is immediate, hence its proof is omitted. Both Lemmas 1 and 2 are simple but useful conditions that can be easily verified using standard uniform laws of large numbers and central limit theorems. In particular, they have been proven to hold for criterion functions corresponding to

itemize• Most smooth cross-sectional models described in amen; • The smooth nonlinear stationary and dynamic GMM and Quasi-likelihood models of hansen:gmm, gallant:white and potscher, covering Gordin(mixingale type) conditions and near-epoch dependent processes such as ARMA, GARCH, ARCH, and other models alike; • General empirical likelihood models for smooth moment equation models studied by imbens:el, kitamura_stutzer, neweysmith, Owen (owen:1989,owen:1990,owen:1991, owen_book), qin_lawless, and the recent extensions to the conditional moment equations.

Hence the main results of this paper, Theorems 1-4, apply to these fundamental econometric and statistical models. Moreover, Assumption (ref) does not require differentiability of the criterion function and thus holds even more generally. Assumption (ref).iv is a Huber-like stochastic equicontinuity condition, which requires that the remainder term of the expansion can be controlled in a particular way over a neighborhood of $\theta_0$. In addition to Lemma 2, many sufficient conditions for Assumption 4 are given in empirical process literature, e.g. amen, andrews1, newey:eq, pakes_pollard, and vaart. Section 4 verifies Assumption 4 for the leading models with nonsmooth criterion functions, including the examples discussed in the previous section.

\@startsection{subsection}{2} {\z@}{-2.25ex plus -0.3ex minus -0.2ex}{0.05ex plus 0.05ex}{Convergence in the Total Variation of Moments Norm} Under Assumptions 1-4, we show that the Quasi-posterior density concentrates around $\theta_0$ at the speed $1/\sqrt{n}$ as measured by the total variation of moments norm, and then use this preliminary result to prove all other main results.

Define the local parameter $h$ as a normalized deviation from $\theta_0$ and centered at the normalized random “score function":

align[align omitted — 160 chars of source]

Define by the Jacobi rule the localized Quasi-posterior density for $h$ as

align[align omitted — 177 chars of source]

Define the total variation of moments norm for a real-valued measurable function $f$ on $S$ as $$ \| f \|_{\scriptscriptstyle TVM(\alpha)} \equiv \int_{S} ( 1 + |h|^{\alpha}) |f(h)| dh. $$

theorem[Convergence in Total Variation of Moments Norm] Under Assumptions (ref) - (ref), for any $0 \leq \alpha < \infty$, \begin{align}\begin{split}\nonumber & \| p_n^*(h) - p^*_{\infty}(h) \|_{\scriptscriptstyle TVM(\alpha)} \equiv \int_{H_n} (1 + | h |^\alpha) | p_n^*(h) - p^*_{\infty}(h) | dh {\rightarrow}_p \ 0, \end{split}\end{align} where $H_n = \{ \sqrt{n}(\theta - \theta_0)- J_n\left(\theta_0\right)^{-1} \Delta_n\left(\theta_0\right)/\sqrt{n}: \theta \in \Theta \}$ and $$ p^*_{\scriptscriptstyle \infty}(h) = {\sqrt{\scriptstyle \frac{\det J_n(\theta_0)}{(2 \pi)^d}}} \cdot \exp \left( - \frac{1}{2} h'J_n(\theta_0)h\right). $$

Theorem (ref) shows that $p_n\left(\theta\right)$ is concentrated at a $1/\sqrt{n}$ neighborhood of $\theta_0$ as measured by the total variation of moments norm. For large $n$, $p_n\left(\theta\right)$ is approximately a random normal density with the random mean parameter $\theta_0+ J_n\left(\theta_0\right)^{-1} \Delta_n\left(\theta_0\right)/n,$ and constant variance parameter $J_n(\theta_0)^{-1}/n.$

Theorem (ref) applies to general statistical criterion functions $L_n(\theta)$, hence it covers the parametric likelihood setting as a fundamental case, in particular implying the Bernstein-Von Mises theorems, which state the convergence of the likelihood posterior to the limit random density in the total variation norm. Note also that the total variation norm results from setting $\alpha =0$ in the total variation of moments norm. The use of the latter is needed to deduce the convergence of LTE's such as the posterior means or variances in Theorems 2-4.

\@startsection{subsection}{2} {\z@}{-2.25ex plus -0.3ex minus -0.2ex}{0.05ex plus 0.05ex}{Limit Results for Point Estimates and Confidence Intervals} As a consequence of Theorem (ref), Theorem 2 establishes $\sqrt{n}$- consistency and asymptotic normality of LTE's. When the loss function $\rho\left(\cdot\right)$ is symmetric, LTE's are asymptotically equivalent to the extremum estimators.

Recall that the extremum estimator $\sqrt{n} ( \widehat \theta_{ex} - \theta_0)$, where $\widehat \theta_{ex} = \arg \sup_{\theta \in \Theta} L_n(\theta)$, is first order-equivalent to $$ U_n \equiv \frac{1}{\sqrt{n}} J_n(\theta_0)^{-1}\Delta_n(\theta_0). $$ Given that the $p_n^*$ approaches $p^*_{\scriptscriptstyle \infty}$, it may be expected that the LTE $\sqrt{n}(\widehat \theta - \theta_0)$ is asymptotically equivalent to $$ Z_n = \arg\inf_{z \in \Bbb{R}^d} \left \{ \ \ \int_{\Bbb{R}^d} \rho\(z-u\right) p^*_{\scriptscriptstyle \infty} ( u - U_n ) \ du \ \right \}. $$ To see a relationship between $Z_n$ and $U_n$, define $$ \xi_{\scriptscriptstyle J_n(\theta_0)} \equiv \arg\inf_{z \in \Bbb{R}^d} \left \{ \ \ \int_{\Bbb{R}^d} \rho\(z-u\right) p^*_{\scriptscriptstyle \infty} (u) \ du \ \right \}, $$ which exists by Assumption 2.\footnote{For example, in the scalar parameter case, if $\rho(h) = (\alpha - 1(h <0))h$, the constant $\xi_{\scriptscriptstyle J(\theta_0)} = q_{\alpha} J_n(\theta_0)^{-1/2}$, where $q_{\alpha}$ is the $\alpha$-quantile of $\mathcal{N}(0,1)$.} If $\rho$ is symmetric, i.e. $\rho(h) = \rho(-h)$, then by Anderson's lemma $ \xi_{\scriptscriptstyle J_n(\theta_0)} = 0. $ Hence $$ Z_n = \xi_{\scriptscriptstyle J_n(\theta_0)} + U_n, $$ and we are prepared to state the result.

theorem[LTE in Large Samples] Under Assumptions (ref)-(ref), \begin{align}\begin{split}\nonumber & \sqrt{n}(\hat\theta - \theta_0) = \xi_{\scriptscriptstyle J_n(\theta_0)} + U_n +o_p(1), \ \ \Omega_n^{\scriptscriptstyle -1/2}(\theta_0) J_n(\theta_0) U_n {\rightarrow}_d \ \mathcal{N}(0 \ , I ). \end{split}\end{align} Hence $$\Omega_n^{\scriptscriptstyle -1/2}(\theta_0) J_n(\theta_0) \left( \sqrt{n}(\hat\theta - \theta_0) - \xi_{\scriptscriptstyle J_n(\theta_0)}\right) {\rightarrow}_d \ \mathcal{N}( 0, I). $$ If loss function $\rho_n$ is symmetric, i.e. $\rho_n\left( h\right) = \rho_n(-h)$ for all $h$, $\xi_{\scriptscriptstyle J_n(\theta_0)}=0$ for each $n$.

In order for the Quasi-posterior distribution to provide valid large sample confidence intervals, the density of $W_n=\mathcal{N}\left(0, J_n(\theta_0)^{-1} \Omega_n(\theta_0) J_n(\theta_0)^{-1}\right)$ should coincide with that of $p^*_{\scriptscriptstyle \infty}\(h\right)$. This requires

align[align omitted — 223 chars of source]

or equivalently $$ \Omega_n(\theta_0) \sim J_n(\theta_0), $$ which is a generalized information equality. The information equality is known to hold for regular, correctly specified likelihoods. It is also known to hold for appropriately constructed criterion functions of generalized method of moments, minimum distance estimators, generalized empirical likelihood estimators, and properly weighted extremum estimators; see Section 4.

Consider construction of the confidence intervals for the quantity $ g(\theta_0)$, and suppose $g$ is continuously differentiable. Define

align[align omitted — 204 chars of source]

Then a LT confidence interval is given by $\left[ c_{g,n}(\alpha/2) , c_{g,n}(1-\alpha/2) \right].$ As previously mentioned, these confidence intervals can be constructed by using the $\alpha/2$ and $1-\alpha/2$ quantiles of the MCMC sequence $$ \left( g(\theta^{(1)}), ..., g(\theta^{\scriptscriptstyle (B)})\right) $$ and thus are quite simple in practice. In order for the intervals to be valid in large samples, one needs to ensure the generalized information equality, which can be done easily through the use of optimal weighting in GMM and minimum-distance criterion functions or the use of generalized empirical likelihood functions; see Section 4.

Consider now the usual asymptotic intervals based on the $\Delta$- method and any estimator with the property

align[align omitted — 138 chars of source]

Such intervals are usually given by

align[align omitted — 423 chars of source]

where $q_{\alpha}$ is the $\alpha$-quantile of the standard normal distribution. The following theorem establishes the large sample correspondence of the Quasi-posterior confidence intervals to the above intervals.

theorem[Large Sample Inference I] Suppose Assumptions 1-4 hold. In addition suppose that the generalized information equality holds: \begin{align}\begin{split}\nonumber \lim_{n \to \infty} J_n( \theta_0) \Omega_n(\theta_0)^{-1} = I. \end{split}\end{align} Then for any $\alpha \in (0,1)$ $$c_{g, n}(\alpha) - g(\widehat \theta) - q_{\alpha}{\displaystyle \frac{\sqrt{ \nabla_{\theta} g( \theta_0)^{'} J_n(\theta_0)^{-1}\nabla_{\theta}g(\theta_0) }}{\sqrt{n}}}= o_p\left(\frac{1}{\sqrt{n}}\right),$$ and \begin{align}\begin{split}\nonumber \lim_{n \rightarrow \infty} P^* \Big \{ c_{g, n}(\alpha/2) \leq g(\theta_0) \leq c_{g, n}(1-\alpha/2) \Big \} = 1-\alpha. \\ \end{split}\end{align}

One practical limitation of this result arises in the case of regression criterion functions (M-estimators), where achieving the information equality may require nonparametric estimation of appropriate weights, e.g. as in censored quantile regression discussed in Section 4. This may entirely be avoided by using a different method for construction of confidence intervals. Instead of the Quasi-posterior quantiles, we can use the Quasi-posterior variance as an estimate of the inverse of the population Hessian matrix $J^{-1}_n(\theta_0)$, and combine it with any available estimate of $\Omega_n(\theta_0)$ (which typically is easier to obtain) in order to obtain the $\Delta$-method style intervals. The usefulness of this methods is particularly evident in the censored quantile regression, where direct estimation of $J_n(\theta_0)$ requires use of nonparametric methods.

theorem[Large Sample Inference II] Suppose Assumptions 1-4 hold. Define for $\widehat \theta = \int_{\Theta} \theta p_n(\theta) d \theta$, $$ \widehat J_n^{-1}( \theta_0) \equiv \int_{\Theta} n (\theta- \widehat \theta )(\theta - \widehat \theta)' p_n(\theta) d \theta, $$ and $$c_{g,n}(\alpha)\equiv g(\widehat \theta) + q_{\alpha}\cdot {\displaystyle \frac{\sqrt{ \nabla_{\theta} g(\widehat \theta)^{'}\widehat J_n(\theta_0)^{-1} \widehat \Omega_n (\theta_0) \widehat J_n(\theta_0)^{-1} \nabla_{\theta}g(\widehat \theta) }}{\sqrt n}},$$ where $\widehat \Omega_n(\theta_0) \Omega_n^{-1}(\theta_0) {\rightarrow}_p \ I$. Then $\widehat J_n(\theta_0)J_n(\theta_0)^{-1} {\rightarrow}_p \ I$, and \begin{align}\begin{split}\nonumber \lim_{n \rightarrow \infty} P_* \Big \{ c_{g, n}(\alpha/2) \leq g(\theta_0) \leq c_{g, n}(1-\alpha/2) \Big \} = 1-\alpha. \end{split}\end{align}

In practice $\widehat J_n( \theta_0)^{-1}$ is computed by multiplying by $n$ the variance-covariance matrix of the MCMC sequence $ S = \left(\theta^{(1)}, \theta^{(2)}, ..., \theta^{({\scriptscriptstyle B})} \right) $.

Remark (Sandwich; added in 2022): Note that the above result establishes that should use the Huber sandwich $$ \widehat J_n(\theta_0)^{-1} \widehat \Omega_n (\theta_0) \widehat J_n(\theta_0)^{-1}$$ to perform frequentist inference using Quasi-Bayesian methods (and Bayesian methods) when the generalized information equality does not hold. Several more recent papers arrive at a similar point. $\square$

\@startsection{section}{1} {\z@}{-2.5ex plus -0.5ex minus -0.1ex}{0.5ex plus 0.1ex}{Applications to Selected Problems}

This section further elaborates the approach through several examples. Assumptions 1-4 cover a wide variety of smooth econometric models (by virtue of Lemma 1 and Lemma 2). Thus, what follows next is mainly motivated by models with non-smooth moment equations, such as those occurring in Examples 1 - 3. Verification of the key Assumption 4 is not immediate in these examples, and Propositions 1-3 and the forthcoming examples show how to do this in a class of models that are of prime interest to us.

\@startsection{subsection}{2} {\z@}{-2.25ex plus -0.3ex minus -0.2ex}{0.05ex plus 0.05ex}{Generalized Method of Moments and Nonlinear Instrumental Variables} Going back to Example 2, recall that a typical model that underlies the applications of GMM is a set of population moment equations:

align[align omitted — 131 chars of source]

Method of moment estimators involve maximizing an objective function of the form

eqnarray[eqnarray omitted — 554 chars of source]

The choice ((ref)) of the weighting matrix implies the generalized information equality under standard regularity conditions.

Generally, by a Central Limit Theorem $\sqrt{n} g_n(\theta_0) {\rightarrow}_d \ \mathcal{N}(0, W^{-1}(\theta_0))$, so that the objective function can be interpreted as the approximate log-likelihood for the sample moments of the data $g_n(\theta)$. Thus we can think of GMM as an approach that specifies an approximate likelihood for selected moments of the data without specifying the likelihood of the entire data.\footnote{This does not help much in terms of providing formal asymptotic results for the GMM model.}

We may impose Assumptions 1-4 directly on the GMM objective function. However, to highlight the plausibility and elaborate on some examples that satisfy Assumption 4 consider the following proposition.

proposition[Method-of-Moments and Nonlinear IV] Suppose that Assumptions 1-2 hold, and that for all $\theta$ in $\Theta$, $m_i(\theta)$ is stationary and ergodic, and \begin{itemize} • conditions ((ref))-((ref)) hold, • $J(\theta) \equiv G(\theta)' W(\theta) G(\theta) >0 $ and is continuous, $G(\theta) = \nabla_{\theta} E m_i(\theta)$ is continuous, • $\Delta_n(\theta_0)/\sqrt{n} = -\sqrt{n} g_n(\theta_0)' W(\theta_0) G(\theta_0) {\rightarrow}_d \ \mathcal{N}(0, \Omega(\theta_0))$, $\Omega(\theta_0) \equiv G(\theta_0)' W(\theta_0) G(\theta_0)$, • for any $\epsilon > 0$, there is $\delta>0$ such that \begin{align}\begin{split} \limsup_{n \rightarrow \infty} P^*\left \{ \sup_{|\theta - \theta'| \leq \delta} \frac{\sqrt{n} | \(g_n(\theta) - g_n(\theta')\right) - \(E g_n(\theta) - Eg_n(\theta')\right) | }{1 + \sqrt{n} | \theta - \theta' |} > \epsilon \right\} < \epsilon. \end{split}\end{align} \end{itemize} Then Assumption (ref) holds. In addition the information equality holds by construction. Therefore the conclusions of Theorems (ref)-(ref) hold with $\Delta_n(\theta_0)$, $\Omega_n(\theta_0)\equiv \Omega\left(\theta_0\right)$ and $J_n(\theta_0)\equiv J\left(\theta_0\right)$ defined above, where the condition ((ref)) is only needed for the conclusions of Theorem (ref) to hold.

Therefore, for symmetric loss functions $\rho_n$, the LTE is asymptotically equivalent to the GMM extremum estimator. Furthermore, the generalized information equality holds by construction, hence Quasi-posterior quantiles provide a computationally attractive method of “inverting" the objective function for the confidence intervals.

For twice continuously differentiable smooth moment conditions, the smoothness conditions on $\nabla_{\theta} L_n(\theta) $ and $\nabla_{\theta \theta'} L_n (\theta)$ stated in Lemma (ref) trivially imply condition iv in Proposition (ref). More generally, andrews1, pakes_pollard and vaart provide many methods to verify that condition in a wide variety of method-of-moments models.

\paragraph{\bf Example 3 Continued.} Instrumental median regression falls outside of both the classical Bayesian approach and the classical smooth nonlinear IV approach of amemiya:IV. Yet the conditions of Proposition (ref) are satisfied under mild conditions:

itemize$(Y_i, D_i, X_i, Z_i)$ is an iid data sequence, $E [m_i (\theta_0)Z_i] = 0$, and $\theta_0$ is identifiable, • $\{ m_i(\theta) = \left( \tau - 1( Y_i \leq q(D_i, X_i, \theta))\right) Z_i, \theta \in \Theta \}$ is a Donsker class,\footnote{This is a very weak restriction on the function class, and is known to hold for all practically relevant functional forms, see vandervaart. } $E \sup_{\theta} | m_i(\theta) |^2 < \infty$, • $ G(\theta) = \nabla_{\theta} E m_i(\theta)= - E f_{\scriptscriptstyle Y|D,X,Z}(q(D,X, \theta)) Z \nabla_{\theta} q(D,X, \theta)'$ is continuous, • $ J(\theta) = G(\theta)'W(\theta)G(\theta)$ $>0$ and is continuous in an open ball at $\theta_0$.

In this case the weighting matrix can be taken as

align[align omitted — 138 chars of source]

so that the information equality holds. Indeed, in this case $$ \Omega(\theta_0) = G(\theta_0)'W(\theta_0) G(\theta_0) = J(\theta_0), $$ where $$ W(\theta_0) = \text{plim} \ W_n(\theta_0)= \left[\text{Var} \ m_i(\theta_0) \right]^{-1}, \ \ \text{Var} \ m_i(\theta_0) = \tau (1- \tau) E Z_i Z_i'.$$

When the model $q$ is linear and the dimension of $D$ is small, there are computable and practical estimators in the literature.\footnote{These include e.g. the “inverse" quantile regression approach in iqr, which is an extension of koenker:1978's quantile regression to endogenous settings.} In more general models, the extremum estimates are quite difficult to compute, and the inference faces the well-known difficulty of estimating sparsity parameters.

On the other hand, the Quasi-posterior median and quantiles are easy to compute and provide asymptotically valid confidence intervals. Note that the inference does not require the estimation of the density function. The simulation example given in Section 5 strongly supports this alternative approach.

Another important example which poses computational challenge is the estimation problem of blp1. This example is similar in nature to the instrumental quantile regression, and the application of the LT methods may be fruitful there.

\@startsection{subsection}{2} {\z@}{-2.25ex plus -0.3ex minus -0.2ex}{0.05ex plus 0.05ex}{Generalized Empirical Likelihood} A class of objective functions that are first-order equivalent to optimally weighted GMM (after recentering) can be formulated using the generalized empirical likelihood framework.

A class of generalized empirical likelihood functions(GEL) are studied in imbens_spady_johnson, kitamura_stutzer, and neweysmith. For a set of moment equations {$ E m_i(\theta_0) = 0$} that satisfy the conditions of section (ref), define

align[align omitted — 150 chars of source]

Then set

align[align omitted — 132 chars of source]

where $ \widehat \gamma (\theta)$ solves

align[align omitted — 153 chars of source]

and $p = \dim( m_i)$.

The scalar function $s(\cdot)$ is a strictly convex, finite, and three times differentiable function on an open interval of $\Bbb{R}$ containing $0$, denoted $\mathcal V$, and is equal to $+ \infty$ outside such an interval. $s\left(\cdot\right)$ is normalized so that both $\nabla s\left(0\right)=1$ and $\nabla^2 s\left(0\right) = 1$. The choices of the function $s(v) = - \ln (1-v)$, $\exp(v)$, and $(1+v)^2/2$ lead to the well-known empirical likelihood, exponential tilting, and continuous-updating GMM criterion functions.

Simple and practical sufficient conditions for Lemma (ref) are given in qin_lawless, imbens_spady_johnson, kitamura:annals, kitamura_stutzer, including stationary weakly dependent data, neweysmith, and hahnexp. Thus, the application of LTE's to these problems is immediate.

To illustrate a further use of LTE's we state a set of simple conditions geared towards non-smooth microeconometric applications such as the instrumental quantile regression problem. These regularity conditions imply the first order equivalence of the GEL and GMM objective functions. The Donskerness condition below is a weak assumption that is known to hold for all reasonable linear and nonlinear functional forms encountered in practice, as discussed in vandervaart.

proposition[Empirical Likelihood Problems] Suppose that Assumptions 1- 2 hold, and that the following conditions are satisfied: for some $\delta>0$ and all $\theta \in \Theta$ \begin{itemize} • condition ((ref)) holds and that $m_i(\theta)$ is iid, • $E m_i(\theta)$ is continuously differentiable in $\theta$ with derivative $G(\theta)$, and $V(\theta)=E m_i(\theta) m_i(\theta)'$ is continuous in $\theta$; with $G(\theta_0)$ having full rank; • $\sup_{|\theta - \theta_0|<\delta} |m_i (\theta)| < K$ a.s., for some constant $K$$\{ m_i (\theta), \theta \in \Theta\}$ is Donsker class, where $$\sqrt{n} g_n(\theta_0) = \frac{1}{\sqrt{n}}\sum_{i=1}^n m_i\left(\theta_0\right) {\rightarrow}_d \ \mathcal{N}\left(0, V\left(\theta_0\right)\right), \ \ V\left(\theta_0\right) = E \{ m_i(\theta_0) m_i(\theta_0)'\}>0,$$ \end{itemize} then Assumptions (ref) and (ref) hold, and thus the conclusions of Theorems (ref) - (ref) are true with \begin{align}\begin{split}\nonumber &\Delta_n\left(\theta_0\right) / \sqrt{n} = \sqrt{n} g_n(\theta_0) V(\theta_0)^{-1} G(\theta_0) {\rightarrow}_d \ \mathcal{N}( 0, \Omega(\theta_0)), \\ &\Omega(\theta_0) = G(\theta_0)' V(\theta_0)^{-1} G(\theta_0), \\ &J(\theta_0) = G(\theta_0)' V(\theta_0)^{-1} G(\theta_0), \ \ G(\theta_0) = \nabla_{\theta} E m_i(\theta_0). \end{split}\end{align} The information equality holds in this case.

Remark. We have corrected condition ii. relative to the published version. In particular, the full rank condition $G(\theta_0)$ was implicitly used in the proof but missing in the statement of assumptions. $\square$

Another (equivalent) way to proceed is through the dual formulation. Consider the following criterion function

align[align omitted — 202 chars of source]

where $h$ is the Cressie-Reid divergence criterion, cf. neweysmith $$ h(\pi) = \frac{1}{\gamma(\gamma +1)} \frac{(\frac{\pi}{1/n})^{\gamma+1} -1}{n}. $$ The function $L_n(\theta)$ in ((ref)) is the generalized empirical likelihood function for $\theta$ with the concentrated out probabilities. In fact, ((ref)) corresponds to ((ref)) by the argument given in qin_lawless p.303-304, so that Proposition 2 covers ((ref)) as a special case up to renormalization.

Empirical probabilities $\widehat \pi_i(\theta)$'s are obtained in ((ref)) using the extremum method. The case $\gamma = -1$ yields the empirical likelihood case, where $\widehat \pi_i(\theta)$'s are obtained through the maximum likelihood method. Taking $\gamma=0$ yields the exponential tilting case, where $\widehat \pi_i(\theta)$'s are obtained through minimization of the Kullback-Leibler distance from the empirical distribution. Taking $\gamma=1$ yields the continuous-updating case, where $\widehat \pi_i(\theta)$'s are obtained through the minimization of the Euclidean distance from the empirical distribution. Each approach generates the implied probabilities $\widehat \pi_i(\theta)$ given $\theta$. qin_lawless and neweysmith provide the formulas: $$ \widehat \pi_i(\theta) = \nabla s(\widehat \gamma(\theta)' m_i (\theta))/\sum_{i=1}^n \nabla s(\widehat \gamma(\theta)' m_i (\theta)).$$

The Quasi-posterior for $\theta$ and $\widehat \pi_i(\theta)$ can be used for predictive inference. Suppose $m_i(\theta) = m(X_i, \theta)$ for some random vector $X_i$. Then the Quasi-posterior predictive probability is given by

align[align omitted — 236 chars of source]

which can be computed by averaging over the MCMC sequence evaluated at $h_n$, $ \(h_n(\theta^{(1)}),..., h_n(\theta^{(B)})\right).$ It follows similarly to the proof of Theorem 1 in qin_lawless that

align[align omitted — 185 chars of source]

where $\Pi_{\scriptscriptstyle A} \equiv P\{X_i \in A\} (1 - P\{X_i \in A\}) - E m_i(\theta_0)' 1\{X_i \in A\} \cdot U \cdot E m_i(\theta_0) 1\{X_i \in A\} $, $U = V(\theta_0)^{-1} \{ I - G(\theta_0) J(\theta_0)^{-1} G(\theta_0)V(\theta_0)^{-1} \}$.

\@startsection{subsection}{2} {\z@}{-2.25ex plus -0.3ex minus -0.2ex}{0.05ex plus 0.05ex}{M-estimation} M-estimators, which include many linear and nonlinear regressions as special cases, typically maximize objective functions of the form $$L_n\left(\theta\right) = \sum_{i=1}^n m_i\left(\theta\right).$$ $m_i\left(\theta\right)$ need not be the log likelihood function of observation $i$, and may depend on preliminary non-parametric estimation. Assumptions (ref)--(ref) usually are satisfied by uniform laws of large numbers and by unique identification of the parameter; see for example amen and newey_mcfadden. The next proposition gives a simple set of sufficient conditions for Assumption (ref).

proposition[M-problems] Suppose Assumptions (ref)-(ref) hold for the criterion function specified above with the following additional conditions: Uniformly in $\theta$ in an open neighborhood of $\theta_0$, $m_i(\theta)$ is stationary and ergodic, and for $\overline{m}_n(\theta)= \sum_{i=1}^n m_i(\theta)/n$, \begin{itemize} • there exists $\dot m_i(\theta_0)$ such that $E\dot m_i(\theta_0) =0$ for each $i$ and, for some $\delta >0$, \begin{align}\begin{split}\nonumber &\bigg \{ \frac{m_i(\theta)- m_i(\theta_0) - \dot m_i(\theta_0)' (\theta - \theta_0)}{| \theta - \theta_0 |}, \ \ \theta: | \theta - \theta_0 | < \delta \bigg \} is a Donsker class, \\ & E [ \overline{m}_n(\theta) -\overline{m}_n (\theta_0) - \overline{\dot m}_n(\theta_0)'(\theta - \theta_0)]^2 = o( | \theta- \theta_0|^2), \\ & \Delta_n\left(\theta_0\right)/\sqrt{n} = \sum_{i=1}^n \dot m_i(\theta_0)/\sqrt{n} {\rightarrow}_d \ \mathcal{N}(0, \Omega(\theta_0)). \end{split}\end{align} • $ J(\theta) = - \nabla_{\theta \theta'} E[m_i(\theta)] $ is continuous and nonsingular in a ball at $\theta_0$. \end{itemize} Then Assumption 4 holds. Therefore, the conclusions of Theorems (ref), (ref), and 4 hold. If in addition $J(\theta_0) = \Omega(\theta_0)$, then the conclusions of Theorem (ref) also hold.

The above conditions apply to many well known examples such as LAD, see for example vaart. Therefore, for many nonlinear regressions, Quasi-posterior means, modes, and medians are asymptotically equivalent, and Quasi-posterior quantiles provide asymptotically valid confidence statements if the generalized information equality holds. When the information equality fails to hold, the method of Theorem 4 provides valid confidence intervals.

\paragraph{\bf Example 1 Continued.} Under the conditions given in powell_lad or newey:powell for the censored quantile regression, the assumptions of Proposition (ref) are satisfied. Furthermore, it is not difficult to show that when the weights $\omega^*_i$ are nonparametrically estimated, the conditions of newey:powell imply Assumption (ref). Under iid sampling, the use of efficient weighting

align[align omitted — 95 chars of source]

where $f_i = f_{\scriptscriptstyle Y_i|X_i}(q_i), \ q_i = q(X_i;\theta_0),$ validates the generalized information equality, and the Quasi-posterior quantiles form asymptotically valid confidence intervals. Indeed, since

align[align omitted — 175 chars of source]

and

align[align omitted — 209 chars of source]

with

align[align omitted — 110 chars of source]

we have

align[align omitted — 88 chars of source]

For this class of problems, the Quasi-posterior means and medians are asymptotically equivalent to the extremum estimators. The Quasi-posterior quantiles provide asymptotically valid confidence intervals when the efficient weights are used. However, estimation of efficient weights requires preliminary estimation of parameter $\theta_0$. When other weights are used, the method of Theorem 4 provides valid confidence intervals.

\@startsection{section}{1} {\z@}{-2.5ex plus -0.5ex minus -0.1ex}{0.5ex plus 0.1ex}{Computation and Simulation Examples} In this section we briefly discuss the MCMC method and present simulation examples.

\@startsection{subsection}{2} {\z@}{-2.25ex plus -0.3ex minus -0.2ex}{0.05ex plus 0.05ex}{Markov Chain Monte Carlo }

The Quasi-posterior density is proportional to {

align[align omitted — 130 chars of source]

} In most cases we can easily compute $e^{L_{\scriptscriptstyle n}(\theta)} \pi(\theta)$. However, computation of the point estimates and confidence intervals typically requires evaluation of integrals like

align[align omitted — 220 chars of source]

for various functions $g$. For problems for which no analytic solution exists for ((ref)), especially in high dimensions, MCMC methods provide powerful tools for evaluating integrals like the one above. See for example chib_handbook, gewekehandbook, and montecarlo for excellent treatments.

MCMC is a collection of methods that produce an ergodic Markov chain with the stationary distribution $p_n$. Given a starting value $\theta^{(0)}$, a chain $(\theta^{(t)}, 1 \leq t \leq B)$ is generated using a transition kernel with stationary distribution $p_n$, which ensures the convergence of the marginal distribution of $\theta^{(B)}$ to $p_n$. For sufficiently large $B$, the MCMC methods produce a dependent sample $(\theta^{(0)}, \theta^{(1)},..., \theta^{(B)})$ whose empirical distribution approaches $p_n$. The ergodicity and construction of the chains usually imply that as $B\rightarrow \infty$,

align[align omitted — 155 chars of source]

We stress that this technique does not rely on the likelihood principle and can be fruitfully used for computation of LTE's. (Appendix B provides the formal details.)

One of the most important MCMC methods is the Metropolis-Hastings algorithm.

Metropolis-Hastings (MH) algorithm with Quasi-Posteriors. Given the Quasi-posterior density $p_n(\theta) \propto e^{L_{\scriptscriptstyle n}(\theta)} \pi(\theta)$, known up to a constant, and a prespecified conditional density $q(\theta'|\theta)$, generate $\left(\theta^{(0)},...,\theta^{\scriptscriptstyle (B)}\right)$ in the following way,

itemize• Choose a starting value $\theta^{(0)}$. • Generate $\xi$ from $q(\xi|\theta^{(j)})$. • Update $\theta^{(j+1)}$ from $\theta^{(j)}$ for $j=1,2,...$, using {\begin{align}\begin{split}\nonumber \theta^{(j+1)} = \left \{ \begin{array}{ccc} \xi & with probability &\rho(\theta^{(j)}, \xi) \\ \theta^{(j)} & with probability &1- \rho(\theta^{(j)}, \xi) \end{array} \right ., \end{split}\end{align}} where {\begin{align}\begin{split}\nonumber \rho(x,y) = \min \left( \frac{ e^{L_{\scriptscriptstyle n}(y)} \pi(y) q(x| y)} { e^{L_{\scriptscriptstyle n}(x)} \pi(x) q(y|x)} , 1 \right ). \end{split}\end{align}}

Note that the most important quantity in the algorithm is the probability $\rho(x,y)$ of the move from an “old" point $x$ to the “new" point $y$, which depends on how much of an improvement in $e^{L_{\scriptscriptstyle n}(y)} \pi(y)$ a possible “new" value of $y$ yields relative to $e^{L_{\scriptscriptstyle n}(x)} \pi(x)$ at the “old" value $x$. Thus, the generated chain of draws spends a relatively high proportion of time in the higher density regions and a lower proportion in the lower density regions. Because such proportions of times are balanced in the right way, the generated sequence of parameter draws has the requisite marginal distribution, which we then use for computation of means, medians, and quantiles. (How closely the sequence travels near the mode is not relevant.)

Another important choice is the transition kernel $q$, also called the instrumental density. It turns out that a wide variety of kernels yield Markov chains that converge to the distribution of interest. One canonical implementation of the MH algorithm is to take $$ q (x|y) = f ( | x - y|), $$ where $f$ is a density symmetric around $0$, such as the Gaussian or the Cauchy density. This implies that the proposals $ \left(\theta^{(j)}\right)$ follow a random walk. This is the implementation we used in this paper. chib_handbook, gewekehandbook and montecarlo can be consulted for important details concerning the implementation and convergence monitoring of the algorithm.

It is now worth repeating that the main motivation behind the LTE approach is based on its efficiency properties (stated in sections 3 and 4) as well as computational attractiveness. Indeed, the LTE approach is as efficient as the extremum approach, but may avoid the computational curse of dimensionality through the use of MCMC. LTE's are typically means or quantiles of a Quasi-posterior distribution, hence can be computed (estimated) at the parametric rate $1/\sqrt{B}$,\footnote{Note that the rates are used for the informal motivation. We fix $d$ in the discussion, but the rate may typically increase linearly or polynomially in $d$ if $d$ is allowed to grow. } where $B$ is the number of MCMC draws (functional evaluations). Indeed, under canonical implementations, the MCMC chains are geometrically mixing, so the rates of convergence are the same as under independent sampling. In contrast, the extremum estimator (mode) is computed (estimated) by the MCMC and similar grid-based algorithms at the nonparametric rate $(1/B)^{\frac{p}{d+2p}}$, where $d$ is the parameter dimension and $p$ is the smoothness order of the objective function.

We used an optimistic tone regarding the performance of MCMC. Indeed, in the problems we study, the objective functions have numerous local optima, but all exhibit a well pronounced global optimum. These problems are important, and therefore the good performance of MCMC and the derived estimators are encouraging. However, various pathological cases can be constructed, see montecarlo. Functions may have multiple separated global modes (or approximate modes), in which case MCMC may require extended time for convergence. Another potential problem is that the initial draw $\theta^{(0)}$ may be very far in the tails of the posterior $p_n(\theta)$. In this case, MCMC may also take extended time to converge to the stationary distribution. In the problems we looked at, this may be avoided by choosing a starting value based on economic considerations or other simple considerations. For example, in the censored median regression example, we may use the starting values based on an initial Tobit regression. In the instrumental median regression, we may use the two stage least squares estimates as the starting values.

\@startsection{subsection}{2} {\z@}{-2.25ex plus -0.3ex minus -0.2ex}{0.05ex plus 0.05ex}{Monte Carlo Example 1: Censored Median Regression}

As discussed in Section 2, a large literature has been devoted to the computation of Powell's censored median regression estimator. In the simulation example reported below, we find that both in small and large samples with high degree censoring, the LT estimation may be a useful alternative to the popular iterated linear programming algorithm of buchinsky:phd. The model we consider is

align[align omitted — 198 chars of source]

The true parameter $(\beta_0, \beta_1, \beta_2, \beta_3 )$ is $\left(-6, 3, 3, 3\right)$, which produces about $40\%$ censoring.

The LTE is based on the Powell's objective function $L_n\left(\theta\right) = -\sum_{i=1}^n \vert Y_i - \max\left(0, \beta_0 + X_i'\beta\right) \vert.$ The initial draw of the MCMC series is taken to be the ordinary least squares estimate, and other details are summarized in Appendix B.

Table (ref) reports the results. The number in parentheses in the iterated linear programming (ILP) results indicates the number of times that this algorithm converges to a local minimum of $0$. The first row for the ILP reports the performance of the algorithm among the subset of simulations for which the algorithm does not converge to the local minimum at 0. The second row reports the results for all simulation runs, including those for which the ILP algorithm does not move away from the local minimum. The LTE's (Quasi-posterior mean and median) never converge to the local minimum of $0$, and they compare favorably to the ILP even when the local minima are excluded from the ILP results, as can be seen from Table (ref). When the local minima are included in the ILP results, LTE's do markedly better.

center[center omitted — 41 chars of source]

\@startsection{subsection}{2} {\z@}{-2.25ex plus -0.3ex minus -0.2ex}{0.05ex plus 0.05ex}{Monte Carlo Example 2: Instrumental Quantile Regression}

We consider a simulation example similar to that in quantint. The model is

align[align omitted — 222 chars of source]

The true parameter $\left(\alpha_0,\beta_0\right)$ equals $0$, and we consider the instrumental moment conditions

align[align omitted — 270 chars of source]

In simulations, the initial draw of the MCMC series is taken to be the ordinary least squares estimate, and other details are summarized in Appendix B.

While instrumental median regression is designed specifically for endogenous or nonlinear models, we use a classical exogenous example in order to provide a contrast with a clear undisputed benchmark -- the standard linear quantile regression. The benchmark provides a reliable and high-quality estimation method for the exogenous model. In this regard, the performance of the LT estimation and inference, reported in Table (ref) and Table (ref), is encouraging.

Table (ref) summarizes the performance of LTE's and the standard quantile regression estimator. Table (ref) compares the performance of the LT confidence intervals to the standard inference method for quantile regression implemented in S-plus 4.0. The reported results are averaged across the slope parameters. The root mean square errors of the LTE's are no larger than those of quantile regression. Other criteria demonstrate similar performance of two methods, as predicted by the asymptotic theory. The coverage of Quasi-posterior quantile confidence intervals is also close to the nominal level of $90 \%$ in both small and large samples. It is also noteworthy that the intervals do not require nonparametric density estimation, as the standard method requires.

center[center omitted — 46 chars of source]

\@startsection{section}{1} {\z@}{-2.5ex plus -0.5ex minus -0.1ex}{0.5ex plus 0.1ex}{An Illustrative Empirical Application}

The following illustrates the use of LT estimation in practice. We consider the problem of forecasting the conditional quantiles or value-at-risk (VaR) of the Occidental Petroleum (NYSE:OXY) security returns. The problem of forecasting quantiles of return distributions is not only important for economic analysis, but is fundamental to the real-life activities of financial firms. We offer an econometric analysis of a dynamic conditional quantile forecasting model, and show that the LTE approach provides a simple and effective method of estimating such models (despite the difficulties inherent in the estimation).

The dataset consists of 2527 daily observations (September, 1986 -- November, 1998) on

itemize$Y_t$, the one-day returns of the Occidental Petroleum (NYSE:OXY) security, • $X_t$, a vector of returns and prices of other securities that affect the distribution of $Y_t$: a constant, lagged one-day return of Dow Jones Industrials (DJI), the lagged return on the spot price of oil (NCL, front-month contract on crude oil on NYMEX), and the lagged return $Y_{t-1}$.

The choice of variables follows a general principle in which the relevant conditioning information for estimating value-at-risk of a stock return, $X_t$, may contain such variables as a market index of corresponding capitalization and type (for instance, the $S\&P 500$ returns for a large-cap value stock), the industry index, a price of a commodity or some other traded risk that the firm is exposed to, and lagged values of its stock price.

Two functional forms of predictive $\tau$-th quantile regressions were estimated:

align[align omitted — 431 chars of source]

where $Q_{\scriptscriptstyle Y_{t+1}}(\tau|I_t, \theta(\tau))$ denotes the $\tau$-th conditional quantile of $Y_{t+1}$ conditional on the information $I_t$ available at time $t$. In other words, $Q_{\scriptscriptstyle Y_{t+1}}(\tau|I_t, \theta(\tau))$ is the value-at-risk at the probability level $\tau$. The idea behind the dynamic models is to better incorporate the entire past information and better predict risk clustering, as introduced by engle:caviar. The nonlinear dynamic models described by engle:caviar are appealing, but appear to be difficult to estimate using conventional extremum methods, see engle:caviar for discussion. An extended empirical analysis of the linear model is given in vchern2.

The LT estimation and inference strategy is based on the koenker:1978 criterion function,

align[align omitted — 180 chars of source]

where $\rho_{\tau}(u) = (\tau -1(u<0))u$. This criterion function is similar to that described in Example 1, with the exception that there is no censoring. The starting value $s =100$ initializes the recursive specification so that the imputed initial conditions (taken to be the marginal quantiles) have a numerically negligible effect.

In the first step, we constructed the LT estimates using the flat weights $$ w_t(\tau) =1/\tau(1- \tau) \text{ for each } \ t=s,...,T. $$ The results of the first step are not presented here, but they are very similar to those reported below. Because the weights are not optimal, the information equality does not hold, hence Quasi-posterior quantiles are not valid for confidence intervals. However, the confidence intervals suggested in Theorem 4 lead to asymptotically valid inference. Under the assumption of correct dynamic specification, stationary sampling, and the conditions specified in Proposition 3, the LTE's are consistent and asymptotically normal

align[align omitted — 207 chars of source]

where for $\nabla q_t(\tau) = \partial Q_{\scriptscriptstyle Y_{t}}(\tau|I_{t-1}, \varrho(\tau), \theta(\tau))/ \partial ( \varrho , \theta')'$ and $q_t(\tau) = Q_{\scriptscriptstyle Y_{t}}(\tau|I_{t-1}, \varrho(\tau), \theta(\tau))$, $$ J(\theta_0)= E f_{\scriptscriptstyle Y_{t}|I_{t-1}}(q_t(\tau) ) \nabla q_t(\tau) \nabla q_t(\tau)', $$ and for $\frac{\Delta_n(\theta_0)}{\sqrt{T-s}} = \frac{1}{\sqrt{T-s}} \sum_{t=s}^T \left[\tau - 1(Y_t < q_t(\tau))\right] \nabla q_t(\tau)$, $$ \Omega(\theta_0) = \lim_{\scriptscriptstyle T \to \infty} \frac{1}{T-s} E \Delta_n(\theta_0) \Delta_n(\theta_0)' = \tau(1-\tau) E \nabla q_t(\tau) \nabla q_t(\tau)'. $$ If the model is not correctly specified, then, for example, the newey_west estimator provides a consistent and robust procedure for estimation of the limit variance $\Omega(\theta_0)$.

The estimation of the matrix $J(\theta_0)^{-1}$ can be done through the use of nonparametric methods as in powell_lad. Alternatively, as suggested in Theorem 4, we can use the variance-covariance matrix of the MCMC sequence of parameter draws multiplied by $n=(T-s)$ as a consistent estimate of $J(\theta_0)^{-1}$. Plugging the estimates into the variance expression ((ref)), we obtain the standard errors and confidence intervals that are qualitatively similar to those reported in Figures 4-7.

In order to illustrate the use of Quasi-posterior quantiles (Theorem 3) and improve estimation efficiency, we also carried out the second step estimation using the Koenker-Bassett criterion function ((ref)) with the weights $$ \widehat w_t(\tau) = \frac{h}{[ Q_{\scriptscriptstyle Y_{t}}(\tau+h/2|I_{t-1}, \widehat \varrho(\tau), \widehat \theta(\tau)) - Q_{\scriptscriptstyle Y_{t}}(\tau- h/2 |I_{t-1}, \widehat \varrho(\tau), \widehat \theta(\tau)) ]} \cdot \frac{1}{\tau(1-\tau)},$$ where $h \propto C n^{-1/3}$ and $C>0$ is chosen using the rule given in quantint. Under the assumption of correct dynamic specification, these weights imply the generalized information equality, which validates Quasi-posterior quantiles for inference purposes, as in ((ref))-((ref)). The following analysis is based on the second step estimates. The $.05$-th, $.5$-th, and $.95$-th Quasi-posterior quantiles are computed for each coefficient $\theta_j(\tau)$ $(j=1,...,4)$ and $\varrho(\tau)$, and then used to form the point estimates and the $90\%$-confidence intervals, which are reported Figures 4-7 for $\tau = .2, .4,...,.8$.

Figures 2 and 3 present the estimated surfaces of the conditional VaR functions of the dynamic model and linear models, respectively, plotted in the time-probability level coordinates, $(t, p),$ ($ p=\tau$ is the quantile index.) We report VaR for many values of $\tau$. The conventional VaR reporting typically involves the probability levels at a given $\tau$. Clearly, the whole VaR surface formed by varying $\tau$ represents a more complete depiction of conditional risk.

The dynamics depicted in Figures 2 and 3 unambiguously indicate certain dates on which market risk tends to be much higher than its usual level. The difference between the linear and the recursive model is also striking. The risk surface generated by the recursive model is much smoother and is much more persistent. Furthermore, this difference is statistically significant, as Figure 7 shows.

Focusing on the recursive model, let us examine the economic and statistical interpretation of the slope coefficients $\widehat \theta_2(\cdot)$, $\widehat \theta_3(\cdot)$, $\widehat \theta_4(\cdot)$, $\widehat \varrho\left(\cdot\right)$, plotted in Figures 4-7.

The coefficient on the lagged oil price return, $\widehat \theta_2(\cdot)$, is insignificantly positive in the left and right tails of the conditional return distribution. It is insignificantly negative in the middle part. The coefficient on the lagged DJI return, $\widehat \theta_3(\cdot)$, in contrast, is significantly positive for all values of $\tau$. We also notice a sharp increase in the middle range. Thus, in addition to the strong positive relation between the individual stock return and the market return (DJI) (dictated by the fact that $\theta_2(\cdot) > 0$ on $(0.2,0.8)$) there is also additional sensitivity of the median of the security return to the market movements.

The coefficient on the own lagged return, $\widehat \theta_4(\cdot)$, on the other hand, is significantly negative, except for values of $\tau$ close to $0$. This may be interpreted as a reversion effect in the central part of the distribution. However, the lagged return does not appear to significantly shift the quantile function in the tails. Thus, the lagged return is more important for the determination of intermediate risks.

Most importantly, the dynamic coefficient $\widehat \varrho\left(\cdot\right)$ on the lagged VaR is significantly negative in the low quantiles and in the high quantiles, but is insignificant in the middle range. The significance of $\widehat \varrho\left(\cdot\right)$ is a strong evidence in favor of the recursive specification. The magnitude and sign of $\widehat \varrho\left(\cdot\right)$ indicates both the reversion and significant risk clustering effects in the tails of the distribution (see Figure 7). As expected, there is zero effect over the middle range, which is consistent with the random walk properties of the stock price. Thus, the dynamic effect of lagged VaR is much more important for the tails of the quantile function, that is for risk management purposes.

\@startsection{section}{1} {\z@}{-2.5ex plus -0.5ex minus -0.1ex}{0.5ex plus 0.1ex}{Conclusion}

In this paper, we study the Laplace-type Estimators or Quasi-Bayesian Estimators that we define using common statistical, non-likelihood based criterion functions. Under mild regularity conditions these estimators are $\sqrt{n}$-consistent and asymptotically normal, and Quasi-posterior quantiles provide asymptotically valid confidence intervals. A simulation study and an empirical example illustrate the properties of the proposed estimation and inference methods. These results show that in many important cases the Quasi-Bayesian estimators provide useful alternatives to the usual extremum estimators. In ongoing work, we are extending the results to models in which $\sqrt{n}$-convergence rate and asymptotic normality do not hold, including the maximum score problem.

\@startsection{section}{1} {\z@}{-2.5ex plus -0.5ex minus -0.1ex}{0.5ex plus 0.1ex}*{ Acknowledgments} We thank the editor for the invitation of this paper to Journal of Econometrics and an anonymous referee for prompt and highest quality feedback. We thank Xiahong Chen, Shawn Cole, Gary Chamberlain, Ivan Fernandez, Ronald Gallant, Jinyong Hahn, Bruce Hansen, Chris Hansen, Jerry Hausman, James Heckman, Bo Honore, Guido Imbens, Roger Koenker, Shakeeb Khan, Sergei Morozov, Whitney Newey, Ziad Nejmeldeen, Stavros Panageas, Chris Sims, George Tauchen, and seminar participants at Brown University, Duke-UNC Triangle Seminar, MIT, MIT-Harvard, University of Chicago, Princeton University, University of Wisconsin at Madison, University of Michigan, Michigan State University, Texas-AM University, the Winter meeting of the Econometric Society, the 2002 European Econometric Society Meeting in Venice for insightful comments. We gratefully acknowledge the financial support provided by the U.S. National Science Foundation grants SES-0214047 and SES-0079495.