EconBase
← Back to paper

Smoothed GMM for quantile models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

77,915 characters · 19 sections · 115 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Smoothed GMM for quantile models

abstractThis paper develops theory for feasible estimators of finite-dimensional parameters identified by general conditional quantile restrictions, under much weaker assumptions than previously seen in the literature. This includes instrumental variables nonlinear quantile regression as a special case. More specifically, we consider a set of unconditional moments implied by the conditional quantile restrictions, providing conditions for local identification. Since estimators based on the sample moments are generally impossible to compute numerically in practice, we study feasible estimators based on smoothed sample moments. We propose a method of moments estimator for exactly identified models, as well as a generalized method of moments estimator for over-identified models. We establish consistency and asymptotic normality of both estimators under general conditions that allow for weakly dependent data and nonlinear structural models. Simulations illustrate the finite-sample properties of the methods. Our in-depth empirical application concerns the consumption Euler equation derived from quantile utility maximization. Advantages of the quantile Euler equation include robustness to fat tails, decoupling of risk attitude from the elasticity of intertemporal substitution, and log-linearization without any approximation error. For the four countries we examine, the quantile estimates of discount factor and elasticity of intertemporal substitution are economically reasonable for a range of quantiles above the median, even when two-stage least squares estimates are not reasonable. {\em Keywords}: instrumental variables, nonlinear quantile regression, quantile utility maximization. {\em JEL classification codes:} C31, C32, C36

\thispagestyle{empty}

\baselineskip19.5pt \pagenumbering{arabic}

Introduction

Since the seminal work of KoenkerBassett78, quantile regression (QR) has attracted considerable interest in statistics and econometrics. QR estimates conditional quantile functions that provide insight into heterogeneous effects of policy variables. This is especially valuable for program evaluation studies, where these methods help analyze how treatments or social programs affect the outcome's distribution. Nevertheless, endogeneity has been a pervasive concern in economics due to simultaneous causality, omitted variables, measurement error, self-selection, and estimation of equilibrium conditions, among other causes. Extending the standard QR, ChernozhukovHansen05, ChernozhukovHansen06, ChernozhukovHansen08 present results on identification, estimation, and inference for an instrumental variables QR (IVQR) model that allows for endogenous regressors.\footnote{We refer to ChernozhukovHansenWuthrich17 for an overview of IVQR. They discuss alternative, complementary QR models with endogeneity in Section 1.2.5, specifically the triangular system and local quantile treatment effect (LQTE) model. Even in the LQTE model, the IVQR estimator has a meaningful interpretation; see Wuthrich16b.} However, computational difficulties have limited practical estimators to linear models with iid data (discussed below).

Under weaker conditions than prior IVQR papers, we develop theory around feasible smoothed estimators. \footnote{The methods developed in this paper are also related to those for semiparametric and nonparametric models. Identification, estimation, and inference of general (non-smooth) conditional moment restriction models have received much attention in the econometrics literature, as in NeweyMcFadden94, ChenLintonvanKeilegom03, ChenPouzo09,ChenPouzo12, and ChenLiao15, for example. However, theoretical results are only for unsmoothed estimators that are often not computationally feasible in practice.} We consider a set of unconditional moments implied by a general parametric conditional quantile restriction and study exactly identified and over-identified models. Under misspecification of the conditional model, our results still hold for the pseudo-true parameter solving the unconditional moments, complementing the results in AngristChernozhukovFernandez-Val06 for QR. For identification, we provide sufficient conditions for local identification based on these moments. For estimation, since using unsmoothed sample moments is generally intractable, we study smoothed estimators that compute quickly and may have improved precision KaplanSun17. Specifically, we develop smoothed method of moments (MM) and smoothed generalized method of moments (GMM) quantile estimators for exactly identified and over-identified models, respectively.\footnote{QR and IVQR have been discussed as GMM by Buchinsky98 and ChernozhukovHansenWuthrich17, among others.} Unlike prior IVQR estimation papers, we allow for weakly dependent data and nonlinear structural models when establishing the large sample properties of the estimators, namely, consistency and asymptotic normality.

Our in-depth empirical study estimates a quantile Euler equation using aggregate time series data. This equation is derived from a quantile utility maximization model. This model is an interesting alternative to the standard expected utility model because it is robust to fat tails, allows heterogeneity through the quantiles, decouples the elasticity of intertemporal substitution from risk attitude, and results in an Euler equation that does not suffer from any approximation error when log-linearized.\footnote{Heavy tails in consumption data have been documented recently by TodaWalsh15,TodaWalsh17.} Quantile preferences were first studied by Manski88 and were axiomatized by Chambers09 and Rostek10. \Citet{deCastroGalvao17} use quantile preferences in a dynamic economic setting and provide a comprehensive analysis of a dynamic rational quantile model. They derive the policy function (Euler equation) as a nonlinear conditional quantile restriction. Consequently, we may use smoothed GMM to estimate the structural parameters, including the elasticity of intertemporal substitution (EIS). Numerous papers have estimated the EIS, e.g., HansenSingleton83, Hall88, CampbellMankiw89, OgakiReinhart98, and Yogo04. For the four countries we study, the smoothed quantile estimates of the discount factor and EIS are economically reasonable for a range of quantiles above the median, including cases where the 2SLS estimates are not reasonable.

For IVQR estimation of the model in ChernozhukovHansen05, the literature lacks results for feasible estimators allowing nonlinear structural models and dependent data.\footnote{For nonlinear QR (no IV), see Powell94, OberhoferHaupt15, and references therein.} The following are iid sampling assumptions: Condition (i) on p.\ 310 in ChernozhukovHong03, Assumption 2.R1 in ChernozhukovHansen06, and Assumption 1 in KaplanSun17. A nonlinear structural model is allowed by the computationally demanding\footnote{\Citet{ChernozhukovHansenWuthrich17} comment, “This approach bypasses the need to optimize a non-convex and non-smooth criterion at the cost of needing to design a sampler that adequately explores the quasi-posterior in a reasonable amount of computation time.”} Markov Chain Monte Carlo estimator in ChernozhukovHong03, but linear-in-parameters models are required in (3.4) in ChernozhukovHansen06 and Assumption 1 in KaplanSun17. Additionally, ChernozhukovHansen06 note, “The computational advantages of our estimator rapidly diminish as the number of endogenous variables increases” (p.\ 501). Even if only one observed variable is endogenous, this restriction limits the use of interactions and transformations (like polynomial terms). The results from ChernozhukovHansen06 have been extended in unpublished work by SuYang11 to non-iid data for use with a correctly specified linear spatial autoregressive model, treating regressors and instruments as nonstochastic. \Citet{ChenLee17} propose an estimator for linear-in-parameters IVQR models using mixed integer quadratic programming, but computation is very slow: with only four parameters and $n=100$ observations, their Table 1 shows average computation times for IV median regression exceeding five minutes. \Citet{Wuthrich17a} proposes an estimator without assuming linearity but only for a binary treatment (and iid data). From a Bayesian perspective, LancasterJun10 allow nonlinear models but only iid data (and require computation over a grid of coefficient values or else by Markov Chain Monte Carlo). We relax both iid sampling and linearity in our formal results, while maintaining the computational simplicity, speed, and scalability of the method in KaplanSun17, and adding the efficiency of two-step GMM.

Historically, smoothing IVQR moment conditions was proposed first in unpublished notes by MaCurdyHong99, mentioned later in (also unpublished) MaCurdyTimmins01 and the handbook chapter by MaCurdy07. Whang06 and Otsu08 use moment smoothing for empirical likelihood QR. The related idea of smoothing non-differentiable objective functions goes back to Amemiya82, if not earlier, and has been used for QR by Horowitz98, GalvaoKato16, and FernandesGuerreHorta17, among others.

(ref) presents the model and discusses identification. (ref) develops the smoothed MM and GMM estimators, whose asymptotic properties are provided in (ref). (ref) contains simulation results. In (ref) we illustrate the new approach empirically. (ref) suggests directions for future research. The appendix collects all proofs.

We conclude this introduction with some remarks about the notation. Random variables and vectors are uppercase ($Y$, $X$, etc.), while non-random values are lowercase ($y$, $x$); for vector/matrix multiplication, all vectors are treated as column vectors. Also, $\operatorname{\mathds{1}}\mathopen{}\mathclose\bgroup\originalleft\{\cdot\aftergroup\egroup\originalright\}$ is the indicator function, $\operatorname{E}(\cdot)$ expectation, $\operatorname{Q}_\tau(\cdot)$ the $\tau$-quantile, $\operatorname{P}(\cdot)$ probability, and $\mathrm{N}\mathopen{}\mathclose\bgroup\originalleft(\mu,\sigma^2\aftergroup\egroup\originalright)$ the normal distribution. For vectors, $\mathopen{}\mathclose\bgroup\originalleft\lVert\cdot\aftergroup\egroup\originalright\rVert$ is the Euclidean norm. Acronyms used include those for central limit theorem (CLT), continuous mapping theorem (CMT), elasticity of intertemporal substitution (EIS), generalized method of moments (GMM), mean value theorem (MVT), probability density function (PDF), uniform law of large numbers (ULLN), and weak law of large numbers (WLLN).

Model and identification

We consider the following nonlinear conditional quantile model

equation[equation omitted — 115 chars of source]

where $Y_i\in\mathcal{Y}\subseteq{\mathbb R}^{d_Y}$ is the endogenous variable vector, $Z_i\in\mathcal{Z}\subseteq{\mathbb R}^{d_Z}$ is the full instrument vector that contains $X_i\in\mathcal{X}\subseteq{\mathbb R}^{d_X}$ as a subset, $\Lambda(\cdot)$ is the “residual function” that is known up to the finite-dimensional parameter of interest $\beta_{0\tau} \in \mathcal{B} \subseteq {\mathbb R}^{d_\beta}$, and $\tau \in (0,1)$ is the quantile index. The model in (ref) can be represented by conditional moment restrictions as

equation[equation omitted — 300 chars of source]

where $\operatorname{\mathds{1}}\mathopen{}\mathclose\bgroup\originalleft\{\cdot\aftergroup\egroup\originalright\}$ is the indicator function.

To estimate $\beta_{0\tau}$, we use unconditional moments implied by (ref):

equation[equation omitted — 372 chars of source]

\Citet{KaplanSun17} consider a special case of (ref) with $\Lambda\bigl(Y,X,\beta\bigr)=Y_1-Y_{-1}^{\top} \beta_1-X^{\top} \beta_2$, where $Y_1$ is the outcome and $Y_{-1}=(Y_2,\ldots,Y_{d_Y})$ are endogenous regressors. We take $Z_i$ as given; see KaplanSun17 for discussion of optimal instruments. Our asymptotic results assume only (ref), so they are robust to misspecification of the structural model in (ref), treating $\beta_{0\tau}$ as the pseudo-true parameter satisfying (ref).

Given (ref), $\beta_{0\tau}$ is “locally identified” if there exists a neighborhood of $\beta_{0\tau}$ within which only $\beta_{0\tau}$ satisfies (ref). This holds if the partial derivative matrix of the right-hand side of (ref) with respect to the $\beta$ argument is full rank; see, e.g., ChenChernozhukovLeeNewey14. This full rank condition is formally stated below in (ref)(ii). The following proposition states the local identification result.

propositionGiven (ref) and (the full rank) (ref)(ii), $\beta_{0\tau}$ is locally identified.

Global identification is notoriously more difficult to establish, although ChernozhukovHansen05 provide some results for IVQR.

To fix ideas, we discuss two examples of structural models in the form of (ref). The first example is a random coefficient model as in ChernozhukovHansen05,ChernozhukovHansen06. Let $D$ be an endogenous “treatment” (like education), $U\sim\textrm{Unif}(0,1)$ an unobserved variable (like ability), $X$ exogenous regressors, and $\tilde{Y}_d = q\bigl( d, x, \beta_0(U) \bigr)$ potential outcomes (like wage), where $q(\cdot)$ is known and $\beta_0(U)$ is a random coefficient vector depending on $U$. Let $\beta_{0\tau} = \beta_0(\tau)$, $Y = (\tilde{Y}_D, D)$, and $\Lambda(Y,X,\beta) = \tilde{Y}_D - q(D,X,\beta)$. Under their Assumptions A1--A5, Theorem 1 in ChernozhukovHansen05 yields a conditional quantile restriction like in (ref):

equation*[equation* omitted — 168 chars of source]

Another example of a structural model applies to our empirical application in (ref). Under certain assumptions, if individuals maximize the $\tau$-quantile of utility instead of expected utility, then the resulting consumption Euler equation can be written in the form

equation*[equation* omitted — 107 chars of source]

where $\beta$ is the discount factor, $r_t$ is real interest rate, $C_t$ is consumption, $U(\cdot)$ is the utility function, $\Omega_{t}$ is the information set, and $\operatorname{Q}_{\tau}[ W_t \mid \Omega_{t}]$ denotes the conditional $\tau$-quantile of $W_t$ given $\Omega_{t}$. With isoelastic utility and instruments $Z_t$ chosen from $\Omega_t$, we obtain (ref):

equation*[equation* omitted — 118 chars of source]

The smoothed {MM} and {GMM} estimators

This section presents smoothed estimators based on the moment conditions in (ref). The smoothed MM and smoothed GMM estimators are designed for exactly identified and over-identified models, respectively. We now introduce notation, followed by the estimators.

Let the population map $M \colon \mathcal{B} \times \mathcal{T} \mapsto {\mathbb R}^{d_Z}$ be

align[align omitted — 540 chars of source]

where superscript “$u$” denotes “unsmoothed.” The population moment condition (ref) is

equation[equation omitted — 64 chars of source]

Without smoothing, the corresponding sample moments simply replace population expectation $\operatorname{E}(\cdot)$ with sample expectation $\hat{\operatorname{E}}(\cdot)$, i.e., the sample average. Analogous to the population map $M(\cdot)$ in (ref), the unsmoothed sample map is

equation[equation omitted — 315 chars of source]

The well-known computational difficulty ChernozhukovHong03 of minimizing a GMM criterion based on $\hat{M}_n^u( \beta, \tau )$ comes from the discontinuous indicator function $\operatorname{\mathds{1}}\mathopen{}\mathclose\bgroup\originalleft\{\cdot\aftergroup\egroup\originalright\}$ inside $g^u_i(\beta,\tau)$. To address this difficulty, we smooth the indicator function.

With smoothing (no “$u$” superscript), the sample analogs of (ref) are

align[align omitted — 355 chars of source]

where $h_n$ is a bandwidth (sequence) and $\tilde{I}(\cdot)$ is a smoothed version of the indicator function $\operatorname{\mathds{1}}\mathopen{}\mathclose\bgroup\originalleft\{\cdot \ge 0\aftergroup\egroup\originalright\}$. The $\tilde{I}(\cdot)$ in (ref) has been used by Horowitz98, Whang06, and KaplanSun17, who use the fact that its derivative is a fourth-order kernel to establish higher-order improvements in the linear iid setting. The double subscript on $g_{ni}$ is a reminder that we have a triangular array setup because $g_{ni}$ depends on the bandwidth sequence $h_n$ in addition to $(Y_i,X_i,Z_i)$.

figure[figure omitted — 619 chars of source]

Method of moments (exact identification)

With exact identification ($d_Z=d_\beta$), our estimator solves the smoothed sample moment conditions\footnote{The right-hand side of (ref) can be relaxed to $o_p(n^{-1/2})$.}

equation[equation omitted — 85 chars of source]

Numerically, as long as the bandwidth is not too near zero and $\Lambda(\cdot)$ is differentiable in $\beta$, (ref) is easy to solve since the Jacobian exists. Further, it is easy to check whether the estimate indeed satisfies (ref). In contrast, with over-identification, it is impossible to know if the numerical solution is the global (not just local) minimum of the GMM criterion function.\footnote{See footnote 5 in ChernozhukovHong03.} Consequently, combining moments and using (ref) provides a reliable initial value for the GMM minimization.

“One-step” {GMM} (over-identification)

With over-identification ($d_Z>d_\beta$), (ref) has no solution. Thus, a natural GMM estimator is the “one-step” estimator proposed in NeweyMcFadden94 that takes one Newton--Raphson-type step from an initial consistent (but not efficient) estimator. \Citet[Thm.\ 3.5, p.\ 2151]{NeweyMcFadden94} show that this is sufficient for asymptotic efficiency, although they assume $g(\cdot)$ is smooth and fixed. We use the one-step estimator only for intermediate computation, focusing on the more common two-step estimator (in (ref)) for asymptotic theory.

Let

equation[equation omitted — 118 chars of source]

where $\nabla_{\beta^{\top}}$ denotes the partial derivative with respect to $\beta^{\top}$, and $\bar{\beta}$ is an initial estimator consistent for $\beta_{0\tau}$. With iid data, let

equation[equation omitted — 133 chars of source]

be an estimator of $\Omega = \operatorname{E}[ g_{ni}( \beta_{0\tau}, \tau ) g_{ni}( \beta_{0\tau}, \tau )^{\top} ]$. As in (3.11) of NeweyMcFadden94, the one-step estimator is

equation[equation omitted — 209 chars of source]

For $\bar{\beta}$, we use (ref). For $\bar{ \Omega}$ with dependent data, (ref) is replaced by a long-run variance estimator as in NeweyWest87 and Andrews91, which we use in our code. However, this lacks formal justification; see (ref).

Two-step {GMM} estimator (over-identification)

We also consider the two-step GMM estimator to achieve asymptotic efficiency in over-identified models ($d_Z>d_\beta$).\footnote{This is not always true with time series under fixed-smoothing asymptotics HwangSun15.} Let $\hat{W}$ be a symmetric, positive definite weighting matrix. The smoothed GMM estimator minimizes a weighted quadratic norm of the smoothed sample moment vector:

align[align omitted — 471 chars of source]

The usual optimal weighting matrix is an estimator of the inverse long-run variance of the sample moments:\footnote{There are other approaches to achieve efficiency without explicitly estimating the long-run variance, like the (Bayesian) exponentially tilted empirical likelihood of Schennach07 and Schennach05, on which LancasterJun10 is based.} $\hat{W}^* = \bar{\Omega}^{-1} \xrightarrow{p} \Omega^{-1}$, where $\bar{\Omega}$ depends on an initial estimate $\bar{\beta}$ as in (ref). The resulting efficient two-step GMM estimator is

equation[equation omitted — 195 chars of source]

Computing (ref) is difficult because the function may be non-convex. To find the global minimum, we use the simulated annealing algorithm from the GenSA package in R R.GenSA, which is suited to such problems. Despite its strengths, simulated annealing cannot reliably solve (ref) without an initial value reasonably close to the solution. Thankfully, such an initial value is provided by (ref) or (ref).

After computing (ref), one could run simulated annealing again with the unsmoothed objective function. For linear iid IVQR, KaplanSun17 suggest that smoothing improves (pointwise in $\tau$) mean squared error but may reduce estimated heterogeneity (across $\tau$), so the benefit of such a final step is ambiguous.

Large sample properties

We now establish consistency and asymptotic normality of both the smoothed MM and GMM estimators.

Assumptions

Different subsets of the following assumptions are used for different results.

assumptionFor each observation $i$ among $n$ in the sample, endogenous vector $Y_i\in\mathcal{Y}\subseteq{\mathbb R}^{d_Y}$ and instrument vector $Z_i\in\mathcal{Z}\subseteq{\mathbb R}^{d_Z}$; a subset of $Z_i$ is $X_i\in\mathcal{X}\subseteq{\mathbb R}^{d_X}$, with $d_X\le d_Z$. The sequence $\{Y_i,Z_i\}$ is strictly stationary and weakly dependent.
assumptionThe function $\Lambda \colon \mathcal{Y}\times\mathcal{X}\times\mathcal{B} \mapsto {\mathbb R}$ is known and has (at least) one continuous derivative in its $\mathcal{B}$ argument for all $y\in\mathcal{Y}$ and $x\in\mathcal{X}$.
assumptionThe parameter space $\mathcal{B}\in{\mathbb R}^{d_\beta}$ is compact; $d_\beta \le d_Z$. Given $\tau\in(0,1)$, the population parameter $\beta_{0\tau}$ is in the interior of $\mathcal{B}$ and uniquely satisfies the moment condition \begin{equation} 0 = \operatorname{E}\bigl[ Z_i \bigl( \operatorname{\mathds{1}}\mathopen\mathclose\bgroup\originalleft\{\Lambda\bigl(Y_i,X_i,\beta_{0\tau}\bigr) \le 0\aftergroup\egroup\originalright\} - \tau \bigr) \bigr] . \end{equation}
assumptionThe matrix $\operatorname{E}\mathopen{}\mathclose\bgroup\originalleft(Z_iZ_i^{\top} \aftergroup\egroup\originalright)$ is positive definite (and finite).
assumptionThe function $\tilde{I}(\cdot)$ satisfies $\tilde{I}(u)=0$ for $u\le-1$, $\tilde{I}(u)=1$ for $u\ge1$, and $-1\le\tilde{I}(u)\le2$ for $-1<u<1$. The derivative $\tilde{I}'(\cdot)$ is a symmetric, bounded kernel function of order $r\ge2$, so $\int_{-1}^{1}\tilde{I}'(u)\,du=1$, $\int_{-1}^{1}u^k\tilde{I}'(u)\,du=0$ for $k=1,\ldots,r-1$, and $\int_{-1}^{1}\lvert u^r\tilde{I}'(u) \rvert \,du<\infty$ but $\int_{-1}^{1}u^r\tilde{I}'(u)\,du\ne0$, and $\int_{-1}^{1} \lvert u^{r+1} \tilde{I}'(u) \rvert \,du<\infty$.
assumptionThe bandwidth sequence $h_n$ satisfies $h_n=o(n^{-1/(2r)})$.
assumptionGiven any $\beta\in\mathcal{B}$ and almost all $Z_i=z$ (i.e.,\ up to a set of zero probability), the conditional distribution of $\Lambda(Y_i,X_i,\beta)$ given $Z_i=z$ is continuous in a neighborhood of zero.
assumptionFor a fixed $\tau\in(0,1)$, using the definition in (ref), \begin{equation} \sup_{\beta\in\mathcal{B}} \lVert \hat{M}_n(\beta,\tau) - \operatorname{E}[ \hat{M}_n(\beta,\tau) ] \rVert = o_p(1) . \end{equation}
assumptionLet $\Lambda_i\equiv \Lambda\bigl(Y_i,X_i,\beta_{0\tau}\bigr)$ and $D_i\equiv \nabla_{\beta}\Lambda\bigl(Y_i,X_i,\beta_{0\tau}\bigr)$, using the notation \begin{equation} \nabla_{\beta} \Lambda\mathopen\mathclose\bgroup\originalleft(y,x,\beta_0\aftergroup\egroup\originalright) \equiv \frac{\partial }{\partial \beta} \Lambda\mathopen\mathclose\bgroup\originalleft(y,x,\beta\aftergroup\egroup\originalright) \Bigr\rvert_{\beta=\beta_0} , \end{equation} for the $d_\beta\times1$ partial derivative vector. Let $f_{\Lambda|Z}(\cdot\midz)$ denote the conditional PDF of $\Lambda_i$ given $Z_i=z$, and let $f_{\Lambda|Z,D}(\cdot\midz,d)$ denote the conditional PDF of $\Lambda_i$ given $Z_i=z$ and $D_i=d$. (i) For almost all $z$ and $d$, $f_{\Lambda|Z}(\cdot\midz)$ and $f_{\Lambda|Z,D}(\cdot\midz,d)$ are at least $r$ times continuously differentiable in a neighborhood of zero, where the value of $r$ is from (ref). For almost all $z\in\mathcal{Z}$ and $u$ in a neighborhood of zero, there exists a dominating function $C(\cdot)$ such that $\lvertf_{\Lambda|Z}^{(r)}(u\midz)\rvert\le C(z)$ and $\operatorname{E}\bigl[C(Z)\lvertZ\rvert\bigr] < \infty$. (ii) The matrix \begin{equation} G \equiv \mathopen\mathclose\bgroup\originalleft. \frac{\partial }{\partial \beta^{\top} } \operatorname{E}\bigl[ Z_i \operatorname{\mathds{1}}\mathopen\mathclose\bgroup\originalleft\{ \Lambda(Y_i,X_i,\beta) \le 0 \aftergroup\egroup\originalright\} \bigr] \aftergroup\egroup\originalright\rvert_{\beta=\beta_{0\tau}} = -\operatorname{E}\mathopen\mathclose\bgroup\originalleft\{ Z_i D_i^{\top} f_{\Lambda|Z,D}(0\midZ_i,D_i) \aftergroup\egroup\originalright\} \end{equation} has rank $d_\beta$.
assumptionA pointwise CLT applies: \begin{equation} \sqrt{n} \bigl\{ \hat{M}_n(\beta_{0\tau},\tau) - \operatorname{E}[\hat{M}_n(\beta_{0\tau},\tau)] \bigr\} \xrightarrow{d} \mathrm{N}\mathopen\mathclose\bgroup\originalleft(0,\Sigma_{\tau}\aftergroup\egroup\originalright) . \end{equation}
assumptionLet $Z_i^{(k)}$ denote the $k$th element of $Z_i$, and similarly $\beta^{(k)}$. Let $G_{kj}$ denote the row $k$, column $j$ element of $G$ (from (ref)). Assume \begin{equation} -\frac{1}{nh_n} \sum_{i=1}^{n} \tilde{I}'\bigl(-\Lambda(Y_i,X_i,\tilde{\beta}_{k})/h_n \bigr) Z_i^{(k)} \frac{\partial }{\partial \beta^{(j)}}\Lambda(Y_i,X_i,\beta) \Bigr\rvert_{\beta=\tilde{\beta}_{\tau,k}} \xrightarrow{p} G_{kj} . \end{equation} for each $k=1,\ldots,d_\beta$ and $j=1,\ldots,d_\beta$, where each $\tilde{\beta}_{k}$ lies between $\beta_{0\tau}$ and $\hat{\beta}_{\mathrm{MM}}$ (defined in (ref) and (ref), respectively).
assumptionFor the weighting matrix, $\hat{W} \xrightarrow{p} W$, and both are symmetric, positive definite matrices.

For transparency, (ref) includes sampling assumptions that help establish the high-level assumptions (ref), which may require additional restrictions on dependence (mixing conditions); see (ref). (ref) is stronger than a nonparametric model but more general than a linear-in-parameters model. (ref) assumes global identification, following the GMM tradition going back to Hansen82 due to “the difficulty of specifying primitive identification conditions for GMM” NeweyMcFadden94, although ChernozhukovHansen05 have some such results for IVQR. (ref) matches Assumption 2(ii) in KaplanSun17; it is relatively weak, imposing only a finite second moment on $Z_i$ (and, unlike 2SLS, no moment restrictions on $Y_i$. (ref) is essentially Assumption 4(i,ii) of KaplanSun17.

(ref) ensures that the asymptotic effect of smoothing is negligible; it is relatively weak given $r=4$ (as in (ref)), and given the optimal $n^{-1/(2r-1)}$ rate for $h_n$ from KaplanSun17 for linear IVQR with iid data. With weakly dependent data, other bandwidth restrictions are needed to establish (ref); see (ref) for details. (ref) contains suggestions for practical bandwidth selection.

(ref) can be checked easily in most cases. For example, if $\Lambda(Y_i,X_i,\beta)=\tilde{Y}_i-(D_i,X_i^{\top} )\beta$ and $Z^{\top} =(X^{\top} ,\tilde{Z})$, then (ref) is satisfied if the outcome $Y$ has a continuous distribution conditional on almost all $(X^{\top} ,\tilde{Z})=(x^{\top} ,\tilde{z})$. (ref) generally requires some restriction on dependence and moments, but it is much weaker than the iid sampling assumption of ChernozhukovHansen06 or KaplanSun17. (ref) is used for the asymptotic normality result; it generalizes parts of Assumptions 3 and 7 in KaplanSun17 to our nonlinear model. The full rank of $G$ is also sufficient for local (but not global) identification ChenChernozhukovLeeNewey14. The CLT in (ref) is a high-level assumption, similar to condition (iv) in Theorem 7.2 of NeweyMcFadden94, for example; like (ref), it requires some restriction on dependence and moments but does not require iid sampling. Examples of more primitive sufficient conditions for (ref) and (ref) are given later in (ref). (ref) is actually a generalization of the consistency of Powell's estimator for the asymptotic covariance matrix of the usual quantile regression estimator, as detailed in (ref). (ref) embodies the stochastic equicontinuity that is often separately assumed, as in Theorem 7.2(v) in NeweyMcFadden94; it also involves interrelated restrictions of dependence, moments, and the bandwidth rate, as described in (ref).

(ref) is standard for GMM. For two-step GMM, it is satisfied given a consistent estimator of the (inverse) asymptotic covariance matrix; see discussion in (ref).

Finally, note the lack of a conditional quantile restriction. Only the unconditional moments in (ref) are assumed to be satisfied by the (pseudo) true parameter. Thus, all our results hold even under misspecification (of a conditional model).

Consistency

To establish consistency, we use Theorem 5.9 in vanderVaart98, showing the two required conditions are satisfied here. One condition is an identification condition. The other requires uniform (in $\beta\in\mathcal{B}$) convergence in probability of the sample maps $\hat{M}_n(\beta,\tau)$ to the population map $M(\beta,\tau)$. No iid sampling assumption is required; the second assumption may be established under weak dependence.

A detailed example of primitive conditions for the high-level uniform weak law of large numbers assumed in (ref) is given in (ref).

In addition to (ref), we must show that the sequence of (non-random) maps $\operatorname{E}[ \hat{M}_n(\beta,\tau) ]$ converges to the desired population map $M(\beta,\tau)$, as in (ref).

lemmaUnder (ref), for a fixed $\tau$, using definitions in (ref), \begin{equation} \sup_{\beta\in\mathcal{B}} \mathopen\mathclose\bgroup\originalleft\lvert \operatorname{E}[ \hat{M}_n(\beta,\tau) ] - M(\beta,\tau) \aftergroup\egroup\originalright\rvert = o(1) . \end{equation}

(ref) is intuitive. Without smoothing, $M(\cdot)=\operatorname{E}[\hat{M}^u_n(\cdot)]$ for all $n$. With smoothing, we should expect this to hold asymptotically if the smoothing is asymptotically negligible. The next result establishes consistency.

theoremUnder (ref) for smoothed MM, and additionally (ref) for smoothed GMM, the estimators from (ref) are consistent: \begin{equation} \hat{\beta}_{\mathrm{MM}} - \beta_{0\tau} = o_p(1) , \quad \hat{\beta}_{\mathrm{GMM}} - \beta_{0\tau} = o_p(1) . \end{equation}

Asymptotic normality

To establish asymptotic normality, smoothing facilitates the usual approach of expanding the sample moments around $\beta_{0\tau}$ because the smoothed sample moments are differentiable. That is, we may take a mean value expansion of the first-order condition, rearrange, and take limits.

The following lemma aids the proof of (ref). It relies on (ref) and a proof that the asymptotic “bias” is negligible, i.e., $\sqrt{n}\operatorname{E}[\hat{M}_n(\beta_{0\tau},\tau)] \to 0$.

lemmaUnder (ref), \begin{equation} \sqrt{n}\hat{M}_n(\beta_{0\tau},\tau) \xrightarrow{d} \mathrm{N}\mathopen\mathclose\bgroup\originalleft(0,\Sigma_{\tau}\aftergroup\egroup\originalright) , \quad \Sigma_{\tau} = \lim_{n\to\infty} \operatorname{Var}\mathopen\mathclose\bgroup\originalleft( n^{-1/2}\sum_{i=1}^{n} g_{ni}(\beta_{0\tau},\tau) \aftergroup\egroup\originalright) . \end{equation} With iid data and the conditional quantile restriction $\operatorname{P}\bigl( \Lambda(Y_i,X_i,\beta_{0\tau})\le0 \mid Z_i\bigr)=\tau$, then $\Sigma_{\tau} = \tau(1-\tau) \operatorname{E}(Z_iZ_i^{\top})$.

The asymptotic normality of our estimators can now be stated. We also show their asymptotically linear (influence function) representations.

theoremUnder (ref) for smoothed MM, and additionally (ref) for smoothed GMM, for the estimators from (ref), \begin{equation*} \begin{split} \sqrt{n} ( \hat{\beta}_{\mathrm{MM}} - \beta_{0\tau} ) &= -G^{-1} \frac{1}{\sqrt{n}} \sum_{i=1}^{n} g_{ni}(\beta_{0\tau},\tau) +o_p(1) , \\ \sqrt{n} ( \hat{\beta}_{\mathrm{MM}} - \beta_{0\tau} ) &\xrightarrow{d} \mathrm{N}\mathopen\mathclose\bgroup\originalleft(0,(G^{^{\top}} \Sigma_{\tau}^{-1} G)^{-1}\aftergroup\egroup\originalright) , \\ \sqrt{n} ( \hat{\beta}_{\mathrm{GMM}} - \beta_{0\tau} ) &= -\{ G^{\top} W G \}^{-1} G^{\top} W \frac{1}{\sqrt{n}} \sum_{i=1}^{n} g_{ni}(\beta_{0\tau},\tau) +o_p(1) , \\ \sqrt{n} ( \hat{\beta}_{\mathrm{GMM}} - \beta_{0\tau} ) &\xrightarrow{d} \mathrm{N}\bigl( 0 , (G^{\top} W G )^{-1} G^{\top} W \Sigma_{\tau} W G (G^{\top} W G )^{-1} \bigr) , \end{split} \end{equation*} where $G$ is from (ref), $W$ is from (ref), and $\Sigma_{\tau}$ is from (ref).

As usual, choosing a weighting matrix such that $\hat{W} \xrightarrow{p} W = \Sigma_{\tau}^{-1}$ is asymptotically efficient in the sense that the resulting asymptotic covariance matrix $(G^{^{\top}} \Sigma_{\tau}^{-1} G)^{-1}$ minus the above GMM covariance matrix is negative semidefinite. This is the sense in which the two-step estimator in (ref) is efficient.

(ref) can also be used to construct Wald tests in the usual way. The result and its proof are also helpful for constructing “distance metric” hypothesis tests and over-identification tests. We detail these in a separate paper on inference (in progress).

A consistent long-run variance estimator for quantile models is currently lacking in the literature, as lamented in other quantile papers. For example, Kato12 notes that results in Andrews91 do not apply because they assume smoothness; specifically, Assumptions B(iii) and C(ii) are violated for unsmoothed quantile models. Similarly, Assumptions 4 and 5 in NeweyWest87 are violated, as is Assumption 4 in deJongDavidson00. Our smoothing yields differentiability, but also a triangular array, violating (2.2) in Andrews91, for example. However, deJongDavidson00 allow triangular arrays and generally have very weak conditions. In future work, we hope to verify their Assumption 4 for our smoothed quantile GMM setting.

Monte {Carlo} simulations

This section reports Monte Carlo simulation results to illustrate the finite-sample performance of the proposed methods. Replication code is available on the third author's website.\footnote{It is written in R R.core and uses packages from R.pracma and R.GenSA.}

The following DGPs/models are used; details are in (ref) and the code. DGP 1 has a binary treatment, binary IV, and iid sampling, as with a randomized treatment offer but self-selection into treatment, and the treatment effect increases with the quantile index $\tau$. DGPs 2 and 3 are time series regressions with measurement error and either Gaussian (DGP 2) or Cauchy (DGP 3) errors in the outcome equation, but no slope heterogeneity. The first three models are exactly identified, so the smoothed MM and GMM estimators are identical; we compare these with the usual QR and IV estimators. DGP 4 is for estimating a log-linearized quantile Euler equation using time series data, as in our empirical application; the model is over-identified, so we can compare MM with different GMM estimators.

To quantify precision, instead of root mean squared error (RMSE), we report “robust RMSE.” This replaces bias with median bias and replaces standard deviation with interquartile range (divided by $1.349$); it equals RMSE if the sampling distribution is normal. The primary reason to use the “robust” version is that sometimes the usual IV estimator does not even possess a first moment in finite samples, let alone finite variance Kinal80. \footnote{If one really cares about finite-sample RMSE per se, then OLS should be preferred to IV in the cases where the IV RMSE is infinite but the OLS RMSE is finite.}

table[table omitted — 4,566 chars of source]

(ref) shows the precision of our smoothed estimator from (ref) (“S(G)MM”) with a very small bandwidth $h=0.0001$ (for simplicity), as well as the usual quantile regression (“QR”) estimator (ignoring endogeneity) and the usual (mean) IV estimator.

(ref) shows that for all DGPs, the smoothed estimator's robust RMSE declines toward zero as $n$ increases. In contrast, the QR estimator's robust RMSE never goes to zero due to endogeneity, and the IV estimator's robust RMSE fails to go to zero in DGP 1 where there is heterogeneity across quantiles. This reflects the theoretical result that only our estimator is consistent for $\gamma_\tau$ when there is endogeneity and heterogeneity.

(ref) also shows important finite-sample differences not captured by first-order asymptotics. With $n=20$, for DGP 2 or 3, the lowest robust RMSE is actually that of QR: despite its (median) bias being the largest due to ignoring the endogeneity, its dispersion is so much smaller than the other estimators' dispersions that its overall robust RMSE is the smallest. This advantage persists to $n=50$, but eventually $n$ is large enough for the (median) bias to dominate. In DGPs 2 and 3 that lack slope heterogeneity, the IV estimator is the most efficient when errors are Gaussian (and $n$ is large enough), but not with Cauchy errors, reflecting the greater efficiency of the median (over the mean) when errors are heavy-tailed.

(ref) compares simulated robust RMSE of three smoothed estimators (all with $h=0.1$) of log-linearized quantile Euler equations, specifically the EIS parameter. Compared to the GMM estimator with identity weighting matrix, the two-step GMM estimator is always more efficient. The biggest such advantage is with the smallest sample size, $n=50$; this seems surprising since the two-step GMM's estimated weighting matrix has the largest variance in that case. Two-step GMM is not always better (or always worse) than the MM estimator that takes the linear projection of the (lone) endogenous regressor onto the vector of (five) instruments to be the second excluded instrument (in addition to the constant). Two-step GMM has a smaller robust RMSE in some cases, even half that of MM with $\tau=0.25$ and $n=500$, but in other cases MM has smaller robust RMSE, especially when $n=50$. Perhaps the additional variance of the two-step estimator due to its use of a long-run variance estimator (for the weighting matrix) makes it less efficient in these cases, a phenomenon explored in non-quantile GMM by HwangSun15. Alternatively, perhaps future work can improve the long-run variance estimator's precision, in turn improving the two-step estimator's precision.

table[table omitted — 1,939 chars of source]

Application: quantile {Euler} equation

This section illustrates the usefulness of the proposed methods through an empirical example: the estimation of a quantile Euler equation. We apply the proposed methodology to an economic model of intertemporal allocation of consumption and estimate the elasticity of intertemporal substitution (EIS). The EIS is a parameter of central importance in macro\-economics and finance. We refer to Campbell03, Cochrane05, and LjungqvistSargent12, and the references therein, for a comprehensive overview.

There is a large empirical literature that attempts to estimate the EIS; among others, HansenSingleton83, Hall88, CampbellMankiw89, CampbellViceira99, Campbell03, and Yogo04. The majority of the literature relies on the traditional expected utility framework. The purpose of this application is to estimate and make inference on the EIS for selected developed countries in Campbell03's (Campbell03) data set using the quantile utility maximization model. The quantile model has useful advantages, such as robustness, ability to capture heterogeneity, and separation the notion of risk attitude from the intertemporal substitution.

(ref) describes in detail the model that leads to the quantile Euler equation, establishing parallels with the standard expected utility Euler equation. (ref) describes the estimation procedure, and (ref) discusses log-linearization for the quantile model. (ref) discusses an interpretation of the parameters in question. In (ref) we review the data, and finally (ref) presents the empirical results.

Description of the economic model

\Citet{deCastroGalvao17} employ a variation of the standard economy model of Lucas78. The economic agents decide on the intertemporal consumption and savings (assets to hold) over an infinity horizon economy, subject to a linear budget constraint. The decision generates an intertemporal policy function, which is used to estimate the parameters of interest for a given utility function. Their work is related to that of Giovannetti13, who works with a similar model but restricts the analysis to two periods, whereas deCastroGalvao17 consider an infinite horizon.

The specific model is as follows. Let $C_{t}$ denote the amount of consumption good that the individual consumes in period $t$. At the beginning of period $t$, the consumer has $x_{t}$ units of the risky asset, which pays dividend $d_{t}$. The price of the consumption good is normalized to one, while the price of the risky asset in period $t$ is $p(d_{t})$. Then, the consumer decides its consumption $C_{t}$ and how many units of the risky asset $x_{t+1}$ to save for the next period, subject to the budget constraint

align[align omitted — 168 chars of source]

and positivity restriction

align[align omitted — 59 chars of source]

In equilibrium, we have that $x_{t}^{\ast}=1, \forall t,k$.

So far, the model is exactly the same as the standard Lucas' model, but the objective function will differ. In the standard model, the consumer maximizes

equation[equation omitted — 195 chars of source]

subject to (ref), where $\beta \in (0,1)$ is the discount factor, $U \colon {\mathbb R}_{+} \mapsto {\mathbb R}$ is the utility function, and $\Omega_0$ is the information set at time $t=0$. For the expected utility choice problem, dynamic consistency and the principle of optimality for (ref) imply that at time $s \ge 1$, the consumer chooses $\{C_{t},x_{t}\}_{t \ge s}$ to maximize

equation[equation omitted — 201 chars of source]

subject to (ref). The connection between problems (ref) is made explicit by the linearity of the expectation operator and the law of iterated expectations:

align[align omitted — 499 chars of source]

\Citet{deCastroGalvao17} replace the operator $\operatorname{E}$ in (ref) with $\operatorname{Q}_{\tau}$, such that the consumer maximizes the following quantile objective function

align[align omitted — 862 chars of source]

again subject to (ref), where $\beta_{\tau}\in(0,1)$ is the discount factor for the quantile $\tau$.

Unfortunately, linearity and the law of iterated expectations do not hold for the $\tau$-quantile operator, $\operatorname{Q}_{\tau}$. Thus, in order to preserve dynamic consistency and the principle of optimality, we need to maintain the structure developed in (ref). \Citet{deCastroGalvao17} show that the limit above exists and is well defined. Moreover, they show that the quantile preferences are dynamically consistent, the principle of optimality holds, and the corresponding dynamic problem yields a value function, via a fixed-point argument. They further provide conditions so that the value function is differentiable and concave.

When comparing the expected and quantile utility models, we note that the structure in the right-hand side of (ref) reflects the following associated value function,

equation[equation omitted — 278 chars of source]

The value function for the quantile problem is the same as (ref) but with $\operatorname{Q}_\tau$ replacing $\operatorname{E}$:

equation[equation omitted — 287 chars of source]

In addition, deCastroGalvao17 derive the corresponding Euler equation, using the fact that in equilibrium, the holdings are $x_{t}=1$ for all $t$:

equation[equation omitted — 168 chars of source]

Defining the asset's return by

equation*[equation* omitted — 96 chars of source]

the Euler equation in (ref) simplifies to

equation[equation omitted — 221 chars of source]

After parameterizing the utility function, (ref) is a conditional quantile restriction in the form of (ref), as in our econometric model.

The quantile Euler equation in (ref) looks similar to the standard Euler equation from expected utility maximization,

equation[equation omitted — 205 chars of source]

The expressions inside the conditional quantile and conditional expectation are identical.

For obtaining the mentioned results, deCastroGalvao17 assume the following.

assumption\begin{enumerate} • The dividends assume values in $\mathcal{Z} \subseteq {\mathbb R}$, which is a bounded interval, and $\mathcal{X} =[0,\bar{x}]$ for some $\bar{x}>1$; • $\{d_{t}\}$ is a Markov process with PDF $f \colon \mathcal{Z} \times \mathcal{Z} \mapsto {\mathbb R}_{+}$, which is continuous, symmetric ($f(a,b)=f(b,a)$), $f(d_{t},d_{t+1})>0$ for all $(d_{t},d_{t+1}) \in \mathcal{Z} \times \mathcal{Z}$, and satisfies the property that if $h \colon \mathcal{Z} \mapsto {\mathbb R}$ is weakly increasing and $z \le z'$, then \begin{equation} \int_{\mathcal{Z}} h(\alpha) f(\alpha \mid z) \, d \alpha \le \int_{\mathcal{Z}} h(\alpha) f(\alpha \mid z') \, d \alpha ; \end{equation} • $U \colon {\mathbb R}_{+} \mapsto {\mathbb R}$ is given by $U(c)= \frac{1}{1-\gamma}c^{1-\gamma}$, for $\gamma>0$; • $z \mapsto z+p(z)$ is $C^{1}$ and non-decreasing, with $\frac{d }{d z} z\mathopen{}\mathclose\bgroup\originalleft[\ln (z+p(z)) \aftergroup\egroup\originalright] \ge \gamma$. \end{enumerate}

Assumptions (ref)(ref) are standard in economic applications. In (ref)(ref), it is natural to expect that the price $p(z)$ is non-decreasing with the dividend $z$, and $z+p(z)$ being non-decreasing is an even weaker requirement.

Estimation

We follow a large body of the literature Campbell03 and use isoelastic utility,

equation[equation omitted — 99 chars of source]

The ratio of marginal utilities is

equation[equation omitted — 174 chars of source]

From (ref), the Euler equation is thus

equation[equation omitted — 210 chars of source]

The quantile Euler equation in (ref) is a conditional quantile restriction with finite-dimensional parameter vector $(\beta_\tau,\gamma_\tau)$, as in (ref), so we may use smoothed GMM estimation.

Log-linearization

One benefit of the quantile Euler equation is that it may be log-linearized with no approximation error, unlike the standard Euler equation. One may rewrite (ref) as

equation[equation omitted — 166 chars of source]

For general random variable $W$, $\operatorname{Q}_\tau[\ln(W)]=\ln(\operatorname{Q}_\tau[W])$ (“equivariance”) since $\ln(\cdot)$ is strictly increasing and continuous. In contrast, $\operatorname{E}[\ln(W)] \le \ln(\operatorname{E}[W])$ by Jensen's inequality. Continuing from (ref),

align[align omitted — 236 chars of source]

If $\gamma>0$, then since $\operatorname{Q}_\tau(W)=-\operatorname{Q}_{1-\tau}(-W)$ (and $0=-0$),

align*[align* omitted — 300 chars of source]

Thus, $\ln(\beta)/\gamma$ and $1/\gamma$ are the intercept and slope (respectively) of the $1-\tau$ IV quantile regression of $\ln(C_{t+1}/C_t)$ on a constant and $\ln(1+r_{t+1})$, with instruments from $\Omega_t$.

Similarly, in the sample, the $g_{ni}$ should be equivalent for nonlinear and log-linear estimation since $\hat{\epsilon}_{t+1} \le 1 \iff \ln(\hat{\epsilon}_{t+1}) \le 0$. Thus, the corresponding estimators should be identical. In our application, this is generally true (matching 2+ significant figures), although sometimes there are differences due to the numerical methods, especially simulated annealing, since the log-linear minimization is done in a transformed parameter space. Additionally, the nonlinear and log-linear estimators do not match when the latter is negative (implying misspecification) since the above arguments assume $\gamma>0$.

Interpretation

The parameters of interest in (ref) are $\beta_{\tau}$ and $\gamma_{\tau}$. The former is the usual discount factor. The parameter $1/\gamma_{\tau}$ is the standard measure of EIS implicit in the CRRA utility function in (ref). The EIS is a measure of responsiveness of the consumption growth rate to the real interest rate. As in Hall88, in a model with uncertainty, the interpretation is similar, and a high value of EIS means that when the real interest rate is expected to be high, the consumer will actively defer consumption to the later period.

The interpretation of $1/\gamma_{\tau}$ as the EIS remains valid for the quantile maximization model.\footnote{Hall88's (Hall88) argument that $\gamma$ fundamentally represents the EIS rather than risk aversion applies here, too.} Most directly, this can be seen in equation (ref), where $1/\gamma$ is the derivative of $\ln(C_{t+1}/C_{t})$ with respect to $\ln(1+r_{t+1})$, holding $\epsilon_{t+1}$ constant.

Data

We use data originally from Campbell03 and provided by Yogo04.\footnote{\url{https://sites.google.com/site/motohiroyogo/research/EIS_Data.zip}} It consists of aggregate level quarterly data for the United States (US), United Kingdom (UK), Australia (AUS), and Sweden (SWE). The sample period for the US is 1947Q3--1998Q4, UK is 1970Q3--1999Q1, Australia is 1970Q3--1998Q4, and Sweden is 1970Q3--1999Q2. Consumption is measured at the beginning of the period, consisting of nondurables plus services for the US and total consumption for the other countries, in real, per capita terms. The real interest rate deflates a proxy for the nominal short-term rate by the consumer price index. Instruments are lags of log real consumption growth, nominal interest rate, inflation, and a log dividend-price ratio for equities. For a complete description of the data, see Campbell03.

Results

Other than using quantiles, our estimation follows Table 2 of Yogo04. \Citet{Yogo04} uses 2SLS to estimate the (structural) log-linearized model $\ln(C_{t+1}/C_t) = \delta_0 + \delta_1 \ln(1+r_{t+1}) + u_{t+1}$, where $r_{t+1}$ is the real interest rate, instrumenting for $\ln(1+r_{t+1})$ with twice lagged measures of nominal interest rate, inflation, consumption growth, and log dividend-price ratio, where $\delta_1=1/\gamma$ is the EIS and $\delta_0=\ln(\beta)/\gamma$. \Citet{Yogo04} emphasizes that these are strong instruments that predict the real interest rate well, although formally characterizing “strong” for IVQR remains an open question. \footnote{In contrast, when trying to estimate the EIS by 2SLS of $\ln(1+r_{t+1})$ on $\ln(C_{t+1}/C_t)$, or replacing the real interest rate with a real stock index return, the instruments are weak because it is difficult to predict consumption growth or stock returns.}

(ref) show the quantile Euler equation estimates for $\beta_\tau$ and $\gamma_\tau$ (respectively), using the results from (ref), with the smoothed MM estimator in (ref) and the smoothed two-step GMM estimator in (ref), for the deciles $\tau=0.1,\ldots,0.9$. For two-step GMM, the long-run variance estimator follows Andrews91 with a quadratic spectral kernel. For both estimators, the plug-in bandwidth from KaplanSun17 was used. For comparison, 2SLS estimates are in each table's bottom row.

\sisetup{round-precision=2,round-mode=places,table-format=1.2}

table[table omitted — 2,020 chars of source]

\sisetup{round-precision=1,round-mode=places,table-format=-2.1}

table[table omitted — 2,030 chars of source]

(ref) generally show economically unrealistic estimates at lower $\tau$ but plausible estimates at larger $\tau$. For $\tau \le 0.4$, some of the $\beta_\tau$ estimates are unrealistically far from one, and many of the $\gamma_\tau$ estimates are negative. For $\tau \ge 0.5$, in contrast, most of the $\beta_\tau$ estimates are close to one, and most $\gamma_\tau$ estimates seem plausible. For $\tau \in \{0.7,0.8\}$ in particular, looking across all four countries and both MM and GMM estimates, the estimates are all contained within the ranges $\hat\beta_\tau \in [0.97, 1.02]$ and $\hat\gamma_\tau \in [3.2, 13.9]$.

The differences between MM and two-step GMM estimates are often relatively small, especially when the estimates are reasonable. However, the table shows some economically significant differences, such as for the US (even with $\tau \ge 0.6$).

The differences between the quantile and 2SLS estimates can be economically significant. This includes the case of Sweden, where the 2SLS estimates are entirely unrealistic: $\hat\beta_\mathrm{2SLS}=0.27$ and $\hat\gamma_\mathrm{2SLS}=-544.4$. Although smaller $\tau$ produce unrealistic estimates, the Sweden quantile estimates for $\tau=0.7$ and $\tau=0.8$ have $\hat\beta_\tau=0.97$ and $\hat\gamma_\tau$ in the range $[5.3, 13.9]$, all perfectly reasonable. \footnote{Further, among the other seven countries whose data Yogo04 examined (Netherlands, Canada, France, Germany, Italy, Japan, Switzerland), all seven had negative 2SLS estimates $\hat\gamma_\mathrm{2SLS}<0$, but five of the seven had positive $\hat\gamma_\tau>0$ with $\tau=0.9$ (and plausible $\hat\beta_\tau$).} For the other countries, the $\tau=0.5$ estimates are most similar to 2SLS, but $\tau \ge 0.7$ leads to more realistic $\hat{\beta}_\tau$ and smaller $\hat{\gamma}_\tau$.

In all, this empirical application illustrates that the quantile utility maximization model and new smoothed estimators serve as important tools to study economic behavior.

Conclusion

For finite-dimensional parameters defined by general quantile-type restrictions, we have developed smoothed MM and GMM estimation and asymptotic theory, for exactly and over-identified models, respectively, allowing for weakly dependent data and nonlinear models. This includes nonlinear IV quantile regression and quantile Euler equations as special cases, and our theory is robust to misspecification of the structural models.

The empirical results suggest that quantile utility maximization combined with our smoothed estimation can provide a useful, economically meaningful alternative to estimation based on expected utility. A bonus feature is the ability to log-linearize the quantile Euler equation without any approximation error, unlike the standard Euler equation. Future work may apply our methods to household panel data or carefully consider how to determine $\tau$.

There is more to explore econometrically, too: quantile GMM inference (in progress), IVQR averaging estimators (in progress), optimal bandwidth choice, non/semiparametric models, fixed-smoothing asymptotic approximations, higher-order bootstrap refinements, formally establishing (ref) when $D_i$ depends on $\beta_{0\tau}$, and results uniform in $\tau$, among other topics.