EconBase
← Back to paper

Posterior Inference in Curved Exponential Families under Increasing Dimensions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

36,204 characters · 9 sections · 39 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Posterior Inference in Curved Exponential Families under Increasing Dimensions

abstractThis work studies the large sample properties of the posterior-based inference in the curved exponential family under increasing dimension. The curved structure arises from the imposition of various restrictions on the model, such as moment restrictions, and plays a fundamental role in econometrics and others branches of data analysis. We establish conditions under which the posterior distribution is approximately normal, which in turn implies various good properties of estimation and inference procedures based on the posterior. In the process we also revisit and improve upon previous results for the exponential family under increasing dimension by making use of concentration of measure. We also discuss a variety of applications to high-dimensional versions of the classical econometric models including the multinomial model with moment restrictions, seemingly unrelated regression equations, and single structural equation models. In our analysis, both the parameter dimension and the number of moments are increasing with the sample size.

Introduction

The main motivation for this paper is to obtain large sample results for posterior inference in the curved exponential family under increasing dimension. In the exponential family, the log of a density is linear in the parameters $\theta \in \Theta$; in the curved exponential family, the parameters $\theta$ are restricted to lie on a curve $\eta \mapsto \theta(\eta)$ parameterized by a lower dimensional parameter $\eta \in \Psi$. There are many classical examples of densities that fall in the curved exponential family; see for example Ef78, lehmann, and bandorff. Curved exponential densities have also been extensively used in applications Ef78,Heckman1974,hh05,hunter2007. An example of the condition that creates a curved structure in an exponential family is a moment restriction of the type: $$ \int m(x, \nu) f(x,\theta) d x =0, $$ that restricts $\theta$ to lie on a curve that can be parameterized as $\{\theta(\eta), \eta \in \Psi\}$, where component $\eta=(\nu, \beta)$ contains $\nu$ and other parameters $\beta$ that are sufficient to parameterize all parameters $\theta \in \Theta$ that solve the above equation for some $\nu$. In econometric applications, often moment restrictions represent Euler equations that result from the data being an outcome of an optimization by rational decision-makers; see e.g. hansen:singleton, chamberlain, imbens, CH, and Donald-Imbens-Newey. In the last section of the paper we discuss in more details other econometric models that fit this framework, such as multivariate linear models, seemingly unrelated regressions, single equation structural models, as in Zellner and zellner:text. We also discuss multinomial model with moment restrictions. Thus, the curved exponential framework is a fundamental complement to the exponential framework.

Under high-dimensionality, despite of its applicability, theoretical properties of the curved exponential family are not as well understood as the corresponding properties of the exponential family. We contribute to the theoretical analysis of the posterior inference in curved exponential families under high dimensionality. We provide sufficient conditions under which consistency and asymptotic normality of the posterior is achieved when both the dimension of the parameter space and the sample size are large. Our framework only requires weak conditions on the prior distribution, which allows for improper priors. In particular, the uninformative prior always satisfies our assumptions. We also study the convergence of moments and the rates with which we can estimate them. We then apply these results to a variety of models where both the parameter dimension and the number of moments are increasing with the sample size.

The present analysis of the posterior inference in the curved exponential family builds upon the work of G2000 who studied posterior inference in the exponential family under increasing dimension. Under sufficient growth restrictions on the dimension of the model, it was shown that the posterior distributions concentrate in neighborhoods of the true parameter and can be approximated by an appropriate normal distribution. Such analysis extended in a fundamental way the classical results of Portnoy for maximum likelihood methods for the exponential family with increasing dimensions.

In addition to a detailed treatment of the curved exponential family, we also revisit the exponential family setting under increasing dimension. We present several new results that complement the results in G2000. First, we amend the conditions on priors to allow for a larger set of priors, for example, improper priors; second, we use concentration inequalities for logconcave densities to sharpen the conditions under which the normal approximations apply; and third, we show that the approximation of $\alpha$-th order moments of the posterior by the corresponding moments of the normal density becomes exponentially difficult in the moment order $\alpha$.

We also note that by establishing the asymptotic normality of the posterior distribution we can invoke results in BelloniChernozhukov2009 that guarantees good computational properties for MCMC methods. Moreover, new results on sampling from manifolds (see Diaconis2012) permits the implementation of different random walk schemes that are useful for implementing inference in curved exponential families.

This work allows for increasing dimension, so it can be thought as a sieve technique. However, this paper does not formally account for the approximation errors resulting from using approximate functional forms as opposed to exact functional forms. Approximation errors can be introduced into the model and our results can also be shown to hold under more stringent conditions (approximations errors need to vanish at rates faster than the sampling errors), a sharp analysis of the impact of the approximation error can be delicate and is outside of the scope of the present paper. An example where approximation errors are controlled is the work Bontemps2011 where a (non-parametric) Gaussian regression framework with an increasing number of regressors is studied. We view the extension of the current (non-Gaussian) setting to non-parametric cases under sharp conditions as an important direction for future work.

The rest of the paper is organized as follows. In Section (ref) we formally define the framework, assumptions, and develop results for the exponential family. In Section (ref), the main section, we develop the results for the curved exponential family. In Section (ref) we apply our results on a variety of applications. Appendices collect proofs of the main results and technical lemmas.

{\bf Notation.} For $a, b \in {\rm I\kern-0.18em R}^d$, their (Euclidean) inner product is denoted by $\sp{a}{b}$, and $\|a\| = \sqrt{\sp{a}{a}}$. The unit sphere in ${\rm I\kern-0.18em R}^d$ is denoted by $S^{d-1}= \{ v \in {\rm I\kern-0.18em R}^d : \|v \| = 1 \}$ and the $\ell_2$-ball centered at $\bar \theta$ with radius $\varepsilon>0$ is denoted by $B_{d}(\bar\theta,\varepsilon) = \{ \theta \in {\rm I\kern-0.18em R}^{d} : \|\theta-\bar\theta \| \leq \varepsilon \}$. For a linear operator $A$, the operator norm is denoted by $\|A\|_{op} = \sup \{ \|Aa\| : \|a\| = 1\}$. Let $\phi_d(\cdot; \mu; V)$ denote the $d$-dimensional Gaussian density function with mean $\mu$ and covariance matrix $V$. Throughout the paper we have a triangular array of random samples $\{ X_1^{(n)} \ \ X_2^{(n)} \cdots \ X_n^{(n)}, n\geq 1\}$. For notational convenience we will suppress the superscript $^{(n)}$ but it is understood that ($\theta^{(n)}, \psi^{(n)}, d^{(n)}, \Theta^{(n)})$ are changing with $n$.

Exponential Family Revisited

We assume that the data $\{X_i$, $i=1,\ldots,n$\} are independent ${d}$-dimensional vectors each drawn from a ${d}$-dimensional exponential family whose density is defined by

equation[equation omitted — 122 chars of source]

where $\theta \in \Theta$ an open convex set of ${\rm I\kern-0.18em R}^{{d}}$, $\psi$ is the normalizing convex function, and $h$ depends only $x$. Let $\theta_0 \in \Theta$ denote the (sequence of) true parameter and let $\mu = \psi'(\theta_0)$ and $F = \psi''(\theta_0)$ be the mean and covariance matrix of $X_i$ (with $J=F^{1/2}$ denoting its square root, i.e., $JJ = F$). We further assume that $\theta_0$ is suitable away from the boundary of $\Theta$, namely $B_{d}(\theta_0,\sqrt{{d}\|F^{-1}\|_{op}/n}\log n) \subset \Theta$. Throughout we assume $d\to \infty$ as $n\to\infty$.

Under this framework, the posterior density of $\theta$ given the observed data $\left\{X_i\right\}_{i=1}^n$ is defined as

equation[equation omitted — 344 chars of source]

where $\pi(\cdot)$ denotes a prior distribution on $\Theta$.

Our results are stated in terms of a re-centered Gaussian distribution in the local parameter space $\mathcal{U} = \sqrt{n}J( \Theta - \theta_0)$. The re-centering is $\Delta_n:= \sqrt{n}J^{-1}\left( \frac{1}{n}\sum_{i=1}^nX_i - \mu \right)$; it follows that $E[\Delta_n] = 0$, and $E[ \Delta_n\Delta_n' ]=I_d$ where $I_d$ denotes the ${d}$-dimensional identity matrix. Moreover, the posterior in the local parameter space is defined for $u \in \mathcal{U}$ as

equation[equation omitted — 238 chars of source]

In the same lines of Portnoy and G2000, conditions on the growth rates of the third and fourth moments are imposed. The following quantities play an important role in the analysis:

eqnarray[eqnarray omitted — 477 chars of source]

where $V$ is a random variable distributed as $J^{-1}(U-E_\theta[U])$ and $U$ has density $f(\cdot;\theta)$ as defined in ((ref)).

remarkAlthough we focus on the i.i.d. framework the analysis can be directly extended to the case that observations are independent but not necessarily identically distributed provided the joint likelihood can still be written in the exponential family form ((ref)). In this case the normalizing function satisfies $\psi = \frac{1}{n}\sum_{i=1}^n\psi_i$ where the function $\psi_i$ is induced by the $i$ observation. Similarly, the quantities $B_{1n}$ and $B_{2n}$ are defined as the average of across $i$ of $B_{1ni}(c)$ and $B_{2ni}(c)$ which are defined as in ((ref)) and ((ref)) for the $i$th observation.
remarkWe note that $\lambda_n(c)$ is related but different from the quantity with the same notation defined in G2000. We provide a technical discussion about the differences in the Appendix. For now we note that ((ref)) is always smaller than its counterpart in G2000 which leads to weaker requirements. In the specific applications of interest we develop bounds on ((ref)). In the case that the density ((ref)) is logconcave in the data, we provide generic bounds for ((ref)) in Appendix D.

Next we impose some regularity conditions on the prior $\pi$. \\

{\bf Assumption P($c_n$).} For the specified positive sequence $c_n$, the prior density function $\pi$ satisfies: $$ {\rm } \sup_{\theta\in \Theta} \ln [\pi(\theta)/\pi(\theta_0)] \leq O({d}) \ \mbox{and} \ \ | \ln \pi(\theta) - \ln \pi(\theta_0) | \leq K_n(c_n) \| \theta - \theta_0 \| $$for any $\theta$ s.t. $\|\theta - \theta_0\| \leq \sqrt{c_n\|F^{-1}\|_{op}{d}/n}$, with $ K_n(c_n) \sqrt{c_n\|F^{-1}\|_{op}{d}/n} =o(1)$.

\\ In what follows the sequence $c_n$ will typically remain uniformly bounded in $n$. These conditions differ from the ones imposed in G2000. Although the same Lipschitz condition is assumed, we require only a relative lower bound on the value of the prior on the true parameter instead of an absolute bound. Thus this condition requires that the true parameter does not have an exponentially small prior value relative to other parameter values. We note that such conditions allow for improper priors which were not allowed in G2000. Importantly, the uninformative prior trivially satisfies Assumption P.

Next we state the main results of this section.

theoremFor any fixed value $c>0$, suppose that $(i)$ $B_{1n}(c) \sqrt{{d}/n} = o(1)$, $(ii)$ $\lambda_n(c) {d} = o(1)$, $(iii)$ $\|F^{-1}\|_{op}{d}/n = o(1)$, and $(iv)$ Assumption P($c$) hold. Then we have asymptotic normality of the posterior density function $$ \int_{\mathcal{U}} | \pi_n^*(u) - \phi_{d}(u;\Delta_n,I_{d})| du = o_p(1). $$

Theorem (ref) establishes the asymptotic normality of the posterior density function. It has different assumptions on the prior relative to Theorem 3 of G2000. However, Theorem (ref) does not require additional technical assumptions used in G2000, as discussed in Appendix A, and the growth condition of ${d}$ with relative to the sample size $n$ is improved by $\ln {d}$ factors.

In some applications stronger convergence properties for the posterior distribution can be required. The following theorem provides sufficient conditions for the $\alpha$-moment convergence. In what follows, for sequences of $\alpha$ and $d$, let $M_{{d},\alpha} := (d+\alpha) \left( 1 + \frac{\alpha \ln(d+\alpha)}{d+\alpha} \right).$

theoremIn addition to the conditions (i) and (iii) of Theorem (ref), suppose that the following hold for any fixed $\bar{c}$: $(ii')$ $\lambda_n\left(\bar{c} M_{{d},\alpha}/{d}\right) [\bar{c}M_{{d},\alpha}]^{1+\alpha/2} =o(1)$; $(iv')$ Assumption P($\bar{c}M_{{d},\alpha}/{d}$) and { $K_n\left(\bar{c}M_{{d},\alpha}/{d}\right)\sqrt{\|F^{-1}\|_{op}\big[\bar{c}M_{{d},\alpha}\big]^{1+\alpha}/n} =o(1)$}. Then we have \begin{equation} \int_{\mathcal{U}} \|u\|^\alpha | \pi_n^*(u) - \phi_{d}(u;\Delta_n,I_{d})| du = o_p(1). \end{equation}

Conditions $(ii')$ and $(iv')$ are strengthening of conditions $(ii)$ and $(iv)$ of Theorem (ref) respectively. We emphasize that Theorem (ref) allows for $\alpha$ and ${d}$ to grow as the sample size increases. Our conditions highlight the polynomial trade off between $n$ and ${d}$ which contrasts with an exponential trade off between $n$ and $\alpha$. This suggests that the estimation of higher moments in increasing dimensions applications could be very delicate. Conditions $(ii')$ and $(iv')$ simplify significantly if $\alpha\ln d = o(d)$, in which case $M_{{d},\alpha} = d(1+o(1))$.

remarkSuppose that we are interested in allowing $\alpha$ to grow with the sample size as well. If ${d}$ is growing in a polynomial rate with respect to $n$, our results do not allow for $\alpha = O(\ln n)$. Some limitation along these lines should be expected since there is an exponential trade off between $\alpha$ and $n$. However, it is possible to have $\alpha = O(\sqrt{\ln n})$. Such slow growth conditions illustrate the potential limitations for the practical estimation of higher order moments.

Curved Exponential Family

Next we consider the curved exponential family. Let $X_1, X_2,\ldots,X_n$ be i.i.d. observations from a $d$-dimensional curved exponential family with density given by $$ f(x;\theta) = h(x)\exp\left( \sp{x}{\theta(\eta)} - \psi(\theta(\eta))\right), $$ where $\eta \in \Psi \subset {\rm I\kern-0.18em R}^{{d}_1}$, $\theta:\Psi\to \Theta$, an open subset of ${\rm I\kern-0.18em R}^d$, and $d \to \infty$ as $n \to \infty$ as before.

The parameter of interest is $\eta$, whose true value $\eta_0$ is suitably bounded away from the boundary of $\Psi \subset {\rm I\kern-0.18em R}^{{d}_1}$ (see Assumption A). The true value of $\theta$ induced by $\eta_0$ is given by $\theta_0 = \theta(\eta_0)$. The mapping $\eta \mapsto \theta(\eta)$ takes values from ${\rm I\kern-0.18em R}^{{d}_1}$ to ${\rm I\kern-0.18em R}^{d}$ where $ {d}_1 \leq {d}$. Moreover, assume that $\eta_0$ is the unique solution to the system $\theta(\eta) = \theta_0$. Thus, the parameter $\theta$ corresponds to a high-dimensional parametrization, and $\eta$ describes the lower-dimensional parametrization of the density.

We require the following regularity conditions on the mapping $\theta(\cdot)$ and on the prior.

itemize• Assumption A. For every fixed $\kappa$, there exists a linear operator $G:{\rm I\kern-0.18em R}^{d_{1}}\to {\rm I\kern-0.18em R}^d$ such that uniformly in $\gamma \in B_{{d}_1}(0,\kappa N_n) \subset \sqrt{n}(\Psi-\eta_0)$, where $N_n = \sqrt{{d}}+\{\sqrt{d_{prior}} + \sqrt{{d}_1\|(G'FG)^{-1}\|_{op}} +\sqrt{{d}_1\log ({d}_1+\|J^{-1}\|_{op}/\varepsilon_0)}\}\max\{1, \|J^{-1}\|_{op}/\varepsilon_0\}$, where $\varepsilon_0$ is defined in Assumption B, we have \begin{equation} \sqrt{n}\left( \theta(\eta_0 + \gamma/\sqrt{n})-\theta(\eta_0) \right) = R_{1n} + (I+R_{2n})G\gamma, \end{equation} where \begin{equation} \begin{array}{l} \|R_{1n}\| \|J\|_{op}\{\sqrt{d}+ \|JG\|_{op} N_n\} =o(1) \ \ and\\ \{\|R_{2n}\|_{op}\sqrt{d} + \|JR_{2n}G\|_{op}N_n\}\|JG\|_{op}N_n =o(1). \end{array}\end{equation} • Assumption B. Uniformly in $n$, there exist a positive constant $\varepsilon_0$ bounded away from zero, such that for every $\eta \in \Psi$ we have \begin{equation} \| \theta(\eta) - \theta(\eta_0) \| \geq \varepsilon_0 \| \eta - \eta_0\|.\end{equation} • {\bf Assumption P'.} The prior density function $\pi$ satisfies: $$\sup_{\eta \in \Psi} \ln [\pi(\theta(\eta))/\pi(\theta(\eta_0))] \leq O({d}_{prior}).$$

Thus the mapping $\eta \mapsto \theta(\eta)$ is allowed to be nonlinear and discontinuous. For example, the additional condition of $R_{1n}=0$ implies the continuity of the mapping in a neighborhood of $\eta_0$. More generally, condition ((ref)) impose that the map admits an (uniform) approximate linearization in the neighborhood of $\eta_0$.

A prior $\pi$ on $\Theta$ induces a prior over $\Psi$ as $\pi(\eta) = \pi(\theta(\eta))/\int_\Psi \pi(\theta(\tilde \eta))d\tilde \eta$. Alternatively the prior can be placed directly over $\Psi$. Assumption P' also bounds the maximum log-likelihood given by the prior to any $\eta$ different than $\eta_0$ to be of the order ${d}_{prior}$ which can grow with $n$. If Assumption P holds we have ${d}_{prior} \leq {d}$. However, if the prior is placed directly on $\Psi$ we typically have ${d}_{prior} = {d}_1$. Finally, if the prior is uninformative we trivially have ${d}_{prior} = 1$. The posterior of $\eta$ given the data is denoted by $$ \pi_n(\eta) \propto \pi(\theta(\eta)) \cdot \prod_{i=1}^n f(x_i;\theta(\eta)) = \pi(\theta(\eta)) \cdot \exp\left( n \sp{\bar X}{\theta(\eta)} - n\psi(\theta(\eta))\right) $$ where $\bar X = (1/n)\sum_{i=1}^n X_i$.

Under this framework, we also define the local parameter space to describe contiguous deviations from the true parameter as $$ \gamma = \sqrt{n}(\eta - \eta_0), \ \ \mbox{and let} \ \ s = (G'FG)^{-1}G'\sqrt{n}(\bar X - \mu) $$ be a first order approximation to the normalized maximum liklelihood/ex\-tre\-mum estimate. Under this setting, the following relations hold for $s$: $$ E[s]= 0, \ \ E [s s'] = (G'FG)^{-1}, \ \ \mbox{and} \ \ \|s\| = O_p(\sqrt{{d}_1\|(G'FG)^{-1}\|_{op}}).$$ The posterior density evaluated at $\gamma \in \Gamma := \sqrt{n}(\Psi-\eta_0)$ is given by $$ \pi^*_n(\gamma) = \ell(\gamma)/\int_{\Gamma} \ell(\gamma) d \gamma, $$ where $${

array[array omitted — 288 chars of source]

}$$

In order to formally state our results we use the following additional definition $$ a_n = \sup \{ c : \lambda_n(c) \leq 1/16 \}. $$ The sequence $a_n$ characterizes a neighborhood of size $\sqrt{a_n{d}}$ around the true local parameter for which the posterior $\ell(\cdot)$ is bounded above by a proper Gaussian density. This is useful since this neighborhood grows and controlling Gaussian tails is typically easier. In turn, this allows for weaker conditions in the next results.

Next we address the consistency question for the maximum likelihood estimator associated with the curved exponential family.

theoremSuppose that Assumptions A, B and P' hold. Then, provided the condition $\|(G'FG)^{-1}\|_{op}\{{d}_1+{d}_{prior}\}/d = o(a_n)$ holds, we have that the maximum likelihood estimator $\widehat \eta$ satisfies $$\| \widehat \eta - \eta_0 \| = O_p\left(\sqrt{\frac{{d}_1+{d}_{prior}}{n}\|(G'FG)^{-1}\|_{op}}\right). $$

Two remarks regarding Theorem (ref) are worth mentioning. First, we note that the condition $\lambda_n(c)=o(1)$ implies $a_n\to \infty$. However, $\lambda_n(c)=o(1)$ is stronger than the condition $\sqrt{d/n}B_{1n}(c)=o(1)$ used for consistency obtained in G2000 for the exponential family case. Second, the consistency result in Theorem (ref) can be substantially impacted by the choice of prior. If the prior used is defined over the full space, it might place an exponentially small (in the dimension $d$) weight in $\eta_0$ relative to other points. In the case ${d}_1\sim {d}$ this is not problematic but it could impact the rates if ${d}_1 = o({d})$. In cases a prior can be placed directly over $\Gamma$ so that $d_{prior} = O({d}_1)$, we obtain the standard rate of convergence of $\sqrt{{d}_1/n}$.

Finally, we state the asymptotic normality result for the curved exponential family. In what follows, let $M_n:=\{1+2\|JG\|_{op}\}^2\|(G'FG)^{-1}\|_{op} \{{d}_1+d_{prior}\}$.

theoremSuppose that Assumptions A, B, and P' hold. Further suppose conditions (i)-(iii) of Theorem (ref) hold, and for each fixed $c>0$, P($cM_n/d$), $d_1\lambda_n(cM_n/d)=o(1)$ and $\{1+2\|JG\|_{op}\}^2N_n^2/d = o(a_n)$. Then, asymptotic normality for the posterior density associated with the curved exponential family holds, $$ \int | \pi_n^*(\gamma) - \phi_{{d}_1}(\gamma; s,(G'FG)^{-1})| d\gamma=o_p(1).$$

Applications to Selected Econometric Models

In this section we verify the conditions that lead to asymptotic normality in a variety of econometric problems covering both exponential and curved exponential families under increasing dimension. Most examples are motivated by the classical work of zellner:text on Bayesian econometrics.

Multivariate Linear Model

In this section we consider a multivariate linear model. The response variable $y_i$ is a $d_y$-dimensional vector, the disturbances $u_i$ are normally distributed with mean zero and covariance matrix $\Sigma_0$. The covariates $z_i$ are $d_z$-dimensional and the parameter matrix of interest $\Pi_0$ is $d_z\times d_y$,

equation[equation omitted — 69 chars of source]

For notational convenience, let $Y$ and $Z$ denote the matrices whose rows are given by $y_i$ and $z_i$ respectively. Note that the dimension of the model is $d = d_y^2 + d_zd_y$.

Conditioning on the covariates $Z$, this model can be cast as an exponential family model by the following parametrization

equation[equation omitted — 356 chars of source]

and using the (trace) inner product $\sp{\theta}{X} = {\rm trace}(X_1'\theta_1) + {\rm trace}(X_2'\theta_2)$, see for instance Garderen. This parametrization leads to the normalizing function

equation[equation omitted — 149 chars of source]

We make the following assumptions on the design. Uniformly in $n$, the covariates satisfy $\max_{i\leq n} \|z_i\| = O(d_z^{1/2})$, the matrices $Z'Z/n$ and $\Sigma_0$ have eigenvalues bounded away from zero and from above, and the matrix $\Pi_0$ has full rank with singular values also bounded away from zero and from above.

Under these assumptions, by Lemma (ref) in the Appendix it follows that $\|F^{-1}\|_{op} = O(1)$. Also, Lemma (ref) in the Appendix bounds the quantities $B_{1n}(c) = O(d_z)$ and $B_{2n}(c) = O(d_z^2)$. Therefore we have asymptotic normality by Theorem (ref) provided that the condition $d(d_z\sqrt{d/n}+d_z^2d/n) = o(1)$ holds.

Seemingly Unrelated Regression Equations

The seemingly unrelated regression model (Zellner) considers a collection of models

equation[equation omitted — 100 chars of source]

where the dimension of $\beta_k$ is $d_k$. Let $d_z$ denote the total number of distinct covariates. The $d_y$-dimensional vector of disturbances $u$ has zero mean and covariance $\Sigma_0$ where $c<\lambda_{min}(\Sigma_0)\leq \lambda_{max}(\Sigma_0)\leq C$ for some fixed constants $c>0$ and $C<\infty$ independent of $n$. This model can be written in the form of ((ref)) by setting $\Pi = [ \pi_1(\beta_1) ; \pi_2(\beta_2) ; \cdots ; \pi_{d_y}(\beta_{d_y}) ]$. Note that the vector $ \pi_k(\beta_k)$ has zeros for regressors that do not appear in the $k$th model. Garderen shows that this model is a curved exponential model provided that the matrix $\Pi$ has some zero restrictions.

Consider the assumptions of Section (ref). In this case we have that

equation[equation omitted — 346 chars of source]

We restrict the space of $\Sigma$ to consider $\lambda_{min}(\Sigma) > \lambda_{min}$ a fixed constant (note that this induces $\lambda_{max}(\Sigma^{-1}) < 1/ \lambda_{min}$ which leads to a convex region in the parameter space), and that operator norm of $\Pi$ is bounded by a constant, $\|\Pi\|_{op}\leq M$ a fixed constant.

The mapping $\theta(\cdot)$ is twice differentiable and Lemma (ref) establishes that condition ((ref)) holds with $R_{2n} = 0$ and $\|R_{1n}\| \leq O(N_n^2/\sqrt{n})$. Provided $\|G\|_{op}\leq C$, this implies that the requirement of $d^3\log^3 d = O(\{d_y^6 + d_z^3d_y^3\}\log^3(d_y+d_z))=o(n)$ suffices for Assumption A to hold.

Next we verify Assumption $B$ for $\varepsilon_0 = \min\{\frac{1}{4},\lambda_{min}(\Sigma_0^{-1}) / [8(1+M)]\}$. Direct calculations provide$$

array[array omitted — 151 chars of source]

$$We can assume that $\|\eta_1 - \eta_{01}\| < 2\varepsilon_0 \|\eta - \eta_0\|$ otherwise Assumption B holds. In turn, this implies that $\|\eta_2 - \eta_{02}\| \geq \frac{1}{2}\|\eta - \eta_0\|$. In this case, since the operator norm of $\eta_2$ satisfies $\|\eta_2\|_{op}\leq M$, we have $$

array[array omitted — 498 chars of source]

$$ which verifies Assumption B.

Single Structural Equation Model

Next we consider the single structural equation,

equation[equation omitted — 103 chars of source]

for which the associated reduced form system, given by the multivariate linear model in ((ref)), can be partitioned as

equation[equation omitted — 221 chars of source]

We assume full column rank of $Z$ and ${\rm rank}(\pi_{21} \ \ \pi_{22}) = {\rm rank}(\pi_{22}) = d_y-1$ where $d_y$ is the dimension of $(y^{(1)}_i \ \ {y^{(2)}_i}')$. The compatibility between the models ((ref)) and ((ref)) requires that $$ \pi_{11} = \pi_{12}\beta + \gamma, \ \ \pi_{21} = \pi_{22} \beta, \ \ \mbox{and} \ \ u^{(1)}_i = {u^{(2)}_i}'\beta + v_i. $$

The model can also be embedded in ((ref)) as follows

equation[equation omitted — 504 chars of source]

Similar arguments to those used in Section (ref) show that Assumptions A and B hold under the standard strong instrument asymptotics.

Multinomial Model

This example of multinomial model was also analyzed in G2000. Our goal is to weaken some of the conditions required previously using the techniques proposed here.

Let $\mathcal{X} = \{ x^0, x^1,\ldots, x^{d}\}$ be the known finite support of a multinomial random variable $X$ where $d$ is allowed to grow with sample size $n$. For each $j$ denote by $p_j$ the probability of the event $\{ X = x^j \}$ which is assumed to satisfy $\max_{0\leq j\leq {d}} 1/p_j = O({d})$. The parameter space is given by $\theta = (\theta_1,\ldots,\theta_{d})$ where $\theta_j = \log(p_j/(1-\sum_{k=1}^{{d}}p_k))$. It follows that under the assumption on the $p_j$'s the true value of $\theta_j$'s are bounded. The Fisher information matrix is given by $F = P - pp'$ where $P={\rm diag}(p)$. In this case we have $B_{1n}(c) = O({d}^{3/2})$ and $B_{2n}(c) = O({d}^2)$. We refer to G2000 for detailed calculations.

The growth condition ${d}^6(\log {d})/n \to 0$ was imposed in G2000 to obtain the asymptotic normality results (the case of $\alpha=0$). We weaken this growth requirement by combining the derivation in G2000 with the analysis in Section (ref) with an uninformative (improper) prior. In this case we have $K_n(c) = 0$ and our definition of $\lambda_n(c)$ remove the logarithmic factors. As a result, Theorem (ref) leads to a weaker growth condition ${d}^4/n\to 0$. For $\alpha$-moment estimation, the conditions of Theorem (ref) are satisfied with the condition that ${d}^{4+\alpha+\delta}/n \to 0$ for any strictly positive value of $\delta$. Recently another approach based on Le Cam's proof that is specific to discrete probability distributions allows for further improvements, see BoucheronGassiat2009.

Multinomial Model with Moment Restrictions

In this subsection we provide a high-level discussion of the multinomial model with moment restrictions. Let $\mathcal{X} = \{ x^0, x^1, x^2, \ldots, x^d\}$ be the known finite support of a multinomial random variable $X$ which was described in Section (ref). Conditions $(i)-(iv)$ are verified as in Section (ref).

As discussed in the introduction, it is of interest to incorporate moment restrictions into this model, see chamberlain and imbens for discussions. This will lead to a curved exponential model as studied in Section (ref).

The parameter of interest is $\eta \in \Psi \subset {\rm I\kern-0.18em R}^{{d}_1}$ a compact set. Consider a (twice continuously differentiable) vector-valued moment function $m:\mathcal{X} \times \Psi \to {\rm I\kern-0.18em R}^M$ such that $$E[m(X,\eta)] = 0\ \ \mbox{for a unique} \ \eta_0 \in \Psi.$$ The log-likelihood function associated with this model

equation[equation omitted — 287 chars of source]

and $l(q,\eta)=-\infty$ if the probability distribution $q$ violates any of the moments conditions. The log-likelihood function ((ref)) induces the mapping $q:\Psi \to \Delta^{d-1}$ formally defined as

equation[equation omitted — 206 chars of source]

In this case, the function $\theta_j(\eta) = \log( q_j(\eta) / q_0(\eta))$ (for $j=1,\ldots,d$) is the mapping from $\Psi \to \Theta$ discussed in Section (ref). Assuming that the matrix $E\left[ m(X,\eta)m(X,\eta)'\right]$ is uniformly positive definite over $\eta$, Qin-Lawless use the inverse function theorem to show that $\theta(\cdot)$ is a twice continuous differentiable mapping of $\eta$ in a neighborhood of $\eta_0$. In particular this implies that we can take $R_{2n}=0$ and $\|R_{1n}\| = O( {d} {d}_1^2 (N_n^2/n) )$. Thus, provided $G'FG$ has eigenvalues bounded away from zero and from above, $d_{prior}\leq d$, and because $\|F^{-1}\|_{op}=O({d})$, Assumption A holds provided that $\|R_{1n}\| N_n = O({d}^3{d}_1^3\log {d} /n) = o(1)$.

Assumption B is satisfied if $\eta$ belongs in a compact set $\Psi$ and that the mapping $\theta(\cdot)$ is injective (over a set that contains $\Psi$ in its interior). We refer to Newey-McFadden for a discussion of primitive assumptions for identification with moment restrictions.