The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
36,204 characters
Posterior Inference in Curved Exponential Families under Increasing Dimensions
\begin{abstract}
This work studies the large sample properties of the
posterior-based inference in the curved exponential family under
increasing dimension. The curved structure arises from the
imposition of various restrictions on the model, such as moment restrictions, and plays a fundamental role in econometrics and others branches of
data analysis. We establish conditions under which the posterior
distribution is approximately normal, which in turn implies
various good properties of estimation and inference procedures
based on the posterior. In the process we also revisit and improve upon previous results for
the exponential family under increasing dimension by making use of concentration of measure. We also discuss a variety of applications to high-dimensional versions of the classical econometric models including the multinomial model
with moment restrictions, seemingly unrelated regression equations, and single structural equation models. In our analysis, both the parameter dimension and
the number of moments are increasing with the sample size.
\keywords{curved exponential family, Bernstein-Von Mises theorems, increasing dimension,
single-equation structural equations, seemingly unrelated regression, multivariate linear models,
multinomial model with moment restrictions.}
\end{abstract}
\section{Introduction}\label{Sec:Intro}
The main motivation for this paper is to obtain large sample
results for posterior inference in the curved exponential family
under increasing dimension. In the exponential family,
the log of a density is linear in the parameters $\theta \in \Theta$;
in the curved exponential family, the parameters $\theta$ are
restricted to lie on a curve $\eta \mapsto \theta(\eta)$
parameterized by a lower dimensional parameter $\eta \in \Psi$.
There are many classical examples of densities that fall in the
curved exponential family; see for example \cite{Ef78},
\cite{lehmann}, and \cite{bandorff}. Curved exponential densities have also been
extensively used in applications \cite{Ef78,Heckman1974,hh05,hunter2007}. An example of the condition that creates a curved structure in an exponential
family is a moment restriction of the type:
$$
\int m(x, \nu) f(x,\theta) d x =0,
$$
that restricts $\theta$ to lie on a curve that can be
parameterized as $\{\theta(\eta), \eta \in \Psi\}$, where
component $\eta=(\nu, \beta)$ contains $\nu$ and other
parameters $\beta$ that are sufficient to parameterize all
parameters $\theta \in \Theta$ that solve the above equation for
some $\nu$. In econometric applications, often moment
restrictions represent Euler equations that result from the data
being an outcome of an optimization by rational
decision-makers; see e.g.
\cite{hansen:singleton}, \cite{chamberlain}, \cite{imbens}, \cite{CH}, and \cite{Donald-Imbens-Newey}. In the last section of the paper we discuss in more details other econometric models that fit this framework, such as multivariate linear models, seemingly unrelated regressions,
single equation structural models, as in \cite{Zellner} and \cite{zellner:text}. We also discuss
multinomial model with moment restrictions. Thus, the curved exponential framework
is a fundamental complement to the exponential framework.
Under high-dimensionality, despite of its applicability,
theoretical properties of the curved exponential family are not as
well understood as the corresponding properties of the exponential
family. We contribute to the theoretical analysis
of the posterior inference in curved exponential families under high
dimensionality. We provide sufficient conditions under which
consistency and asymptotic normality of the posterior is achieved
when both the dimension of the parameter space and the sample size
are large. Our framework only requires weak conditions on the
prior distribution, which allows for improper priors. In particular, the uninformative prior always satisfies our assumptions. We also study the convergence of moments and the rates with
which we can estimate them. We then apply these results to a variety of models where both the
parameter dimension and the number of moments are increasing with
the sample size.
The present analysis of the posterior inference in the curved
exponential family builds upon the work of
\cite{G2000} who studied posterior inference in the exponential
family under increasing dimension. Under sufficient growth
restrictions on the dimension of the model, it was shown that the
posterior distributions concentrate in neighborhoods of the true
parameter and can be approximated by an appropriate normal
distribution. Such analysis extended in a fundamental way the
classical results of \cite{Portnoy} for maximum likelihood
methods for the exponential family with increasing dimensions.
In addition to a detailed treatment of the curved exponential
family, we also revisit the exponential family setting under increasing dimension. We
present several new results that complement the results in \cite{G2000}. First, we amend the conditions on priors to allow
for a larger set of priors, for example, improper priors; second,
we use concentration inequalities for logconcave densities to
sharpen the conditions under which the normal approximations
apply; and third, we show that the approximation of $\alpha$-th
order moments of the posterior by the corresponding moments of the
normal density becomes exponentially difficult in the moment order
$\alpha$.
We also note that by establishing the asymptotic normality of the posterior distribution we can invoke results in \cite{BelloniChernozhukov2009} that guarantees good computational properties for MCMC methods. Moreover, new results on sampling from manifolds (see \cite{Diaconis2012}) permits the implementation of different random walk schemes that are useful for implementing inference in curved exponential families.
This work allows for increasing dimension, so it can be thought as a sieve technique. However, this paper does not formally account for the approximation errors resulting from using approximate functional forms as opposed to exact functional forms. Approximation errors can be introduced into the model and our results can also be shown to hold under more stringent conditions (approximations errors need to vanish at rates faster than the sampling errors), a sharp analysis of the impact of the approximation error can be delicate and is outside of the scope of the present paper. An example where approximation errors are controlled is the work \cite{Bontemps2011} where a (non-parametric) Gaussian regression framework with an increasing number of regressors is studied. We view the extension of the current (non-Gaussian) setting to non-parametric cases under sharp conditions as an important direction for future work.
The rest of the paper is organized as follows. In Section
\ref{Sec:Exp} we formally define the framework, assumptions, and
develop results for the exponential family. In Section
\ref{Sec:Curved}, the main section, we develop the results for the
curved exponential family. In Section \ref{Sec:Applications} we apply our results on
a variety of applications. Appendices
collect proofs of the main results and technical lemmas.
{\bf Notation.} For $a, b \in {\rm I\kern-0.18em R}^d$, their (Euclidean) inner product is denoted
by $\sp{a}{b}$, and $\|a\| = \sqrt{\sp{a}{a}}$. The unit sphere in
${\rm I\kern-0.18em R}^d$ is denoted by $S^{d-1}= \{ v \in {\rm I\kern-0.18em R}^d : \|v \| = 1 \}$ and the $\ell_2$-ball centered at $\bar \theta$ with radius $\varepsilon>0$ is denoted by $B_{d}(\bar\theta,\varepsilon) = \{ \theta \in {\rm I\kern-0.18em R}^{d} : \|\theta-\bar\theta \| \leq \varepsilon \}$.
For a linear operator $A$, the operator norm is denoted by $\|A\|_{op}
= \sup \{ \|Aa\| : \|a\| = 1\}$. Let $\phi_d(\cdot; \mu; V)$
denote the $d$-dimensional Gaussian density function with mean
$\mu$ and covariance matrix $V$. Throughout the paper we have a triangular array of random samples $\{ X_1^{(n)} \ \ X_2^{(n)} \cdots \ X_n^{(n)}, n\geq 1\}$. For
notational convenience we will suppress the superscript $^{(n)}$ but
it is understood that ($\theta^{(n)}, \psi^{(n)}, d^{(n)}, \Theta^{(n)})$ are changing with $n$.
\section{Exponential Family Revisited}\label{Sec:Exp}
We assume that the data $\{X_i$, $i=1,\ldots,n$\} are independent ${d}$-dimensional
vectors each drawn from a ${d}$-dimensional exponential
family whose density is defined by
\begin{equation}\label{Def:Expens} f\left(x;\theta\right) = h(x){\rm exp}\Big(
\sp{x}{\theta} - \psi\big(\theta\big)\Big),
\end{equation}
where $\theta \in \Theta$ an open convex set of
${\rm I\kern-0.18em R}^{{d}}$, $\psi$ is the normalizing convex
function, and $h$ depends only $x$. Let $\theta_0 \in \Theta$ denote the
(sequence of) true parameter and let $\mu =
\psi'(\theta_0)$ and $F = \psi''(\theta_0)$ be the mean and
covariance matrix of $X_i$ (with $J=F^{1/2}$ denoting its square root, i.e., $JJ = F$). We further assume that $\theta_0$ is suitable away
from the boundary of $\Theta$, namely $B_{d}(\theta_0,\sqrt{{d}\|F^{-1}\|_{op}/n}\log n) \subset \Theta$. Throughout we assume $d\to \infty$ as $n\to\infty$.
Under this framework, the posterior density of $\theta$ given the
observed data $\left\{X_i\right\}_{i=1}^n$ is defined as
\begin{equation}\label{Def:Post}
\pi_n(\theta) = \frac{\pi(\theta) \prod_{i=1}^n f(X_i;\theta)}{\int_{\Theta} \pi(\xi) \prod_{i=1}^n f(X_i;\xi) d\xi}
= \frac{\pi(\theta) {\rm exp}\left( \sp{\sum_{i=1}^{n} X_i}{\theta} - n \psi(\theta)
\right)}{\int_{\Theta} \pi(\xi) {\rm exp}\left( \sp{\sum_{i=1}^{n} X_i}{\xi} - n \psi(\xi)
\right)d\xi} ,
\end{equation} where $\pi(\cdot)$ denotes a prior
distribution on $\Theta$.
Our results are stated in terms of a re-centered Gaussian
distribution in the local parameter space $\mathcal{U} =
\sqrt{n}J( \Theta - \theta_0)$. The
re-centering is $\Delta_n:= \sqrt{n}J^{-1}\left(
\frac{1}{n}\sum_{i=1}^nX_i - \mu \right)$; it follows that $E[\Delta_n]
= 0$, and $E[ \Delta_n\Delta_n' ]=I_d$ where $I_d$ denotes the ${d}$-dimensional identity matrix. Moreover, the posterior
in the local parameter space is defined for $u \in \mathcal{U}$ as
\begin{equation}\label{Def:PostLPS} \pi^*(u) = \frac{\pi(\theta_0
+ n^{-1/2} J^{-1}u) \prod_{i=1}^n f(X_i;\theta_0 + n^{-1/2}
J^{-1}u)}{\int_\mathcal{U} \pi(\theta_0 + n^{-1/2} J^{-1}u)
\prod_{i=1}^n f(X_i;\theta_0 + n^{-1/2} J^{-1}u) du}.
\end{equation}
In the same lines of \cite{Portnoy} and
\cite{G2000}, conditions on the growth rates of the third and
fourth moments are imposed. The following
quantities play an important role in the analysis:
\begin{eqnarray}
\label{Def:B1n} B_{1n}(c) &=& \sup_{\theta, a} \left\{ |E_{\theta}\left[ \sp{a}{V}^3 \right] | : a
\in S^{{d}-1}, \|J(\theta-\theta_0)\|^2 \leq \frac{c{d}}{n}
\right\}, \\
\label{Def:B2n}
B_{2n}(c) &=& \sup_{\theta, a} \left\{ E_{\theta}\left[ \sp{a}{V}^4 \right] : a
\in S^{{d}-1}, \|J(\theta-\theta_0)\|^2 \leq \frac{c{d}}{n}
\right\},\\
\label{Def:lambda}
\lambda_n(c) &=& \frac{1}{6} \left( \sqrt{\frac{c{d}}{n}}B_{1n}(0)
+ \frac{c{d}}{n} B_{2n}(c) \right),
\end{eqnarray}
where $V$ is a random variable distributed
as $J^{-1}(U-E_\theta[U])$ and $U$ has density $f(\cdot;\theta)$
as defined in (\ref{Def:Expens}).
\begin{remark}
Although we focus on the i.i.d. framework the analysis can be directly extended to the case that observations are independent but not necessarily identically distributed provided the joint likelihood can still be written in the exponential family form (\ref{Def:Expens}). In this case the normalizing function satisfies $\psi = \frac{1}{n}\sum_{i=1}^n\psi_i$ where the function $\psi_i$ is induced by the $i$ observation. Similarly, the quantities $B_{1n}$ and $B_{2n}$ are defined as the average of across $i$ of $B_{1ni}(c)$ and $B_{2ni}(c)$ which are defined as in (\ref{Def:B1n}) and (\ref{Def:B2n}) for the $i$th observation.
\end{remark}
\begin{remark}
We note that $\lambda_n(c)$ is related but different from
the quantity with the same notation defined in \cite{G2000}. We provide a technical discussion about the differences in the Appendix. For now we note that (\ref{Def:lambda}) is always smaller than its counterpart in \cite{G2000} which leads to weaker requirements. In the specific applications of interest we develop bounds on (\ref{Def:lambda}). In the case that the density (\ref{Def:Expens}) is logconcave in the data, we provide generic bounds for (\ref{Def:lambda}) in Appendix D.
\end{remark}
Next we impose some regularity conditions on the prior $\pi$.
~\\
{\bf Assumption P($c_n$).} For the specified positive sequence $c_n$, the prior density function $\pi$ satisfies:
$$ {\rm } \sup_{\theta\in \Theta} \ln
[\pi(\theta)/\pi(\theta_0)] \leq O({d}) \ \mbox{and} \ \ | \ln \pi(\theta) - \ln \pi(\theta_0) | \leq K_n(c_n) \| \theta - \theta_0 \|
$$for any $\theta$ s.t. $\|\theta - \theta_0\| \leq
\sqrt{c_n\|F^{-1}\|_{op}{d}/n}$, with $ K_n(c_n)
\sqrt{c_n\|F^{-1}\|_{op}{d}/n} =o(1)$.
~\\
\indent In what follows the sequence $c_n$ will typically remain uniformly bounded in $n$.
These conditions differ from the ones imposed in \cite{G2000}. Although the same Lipschitz
condition is assumed, we require only a relative lower bound on
the value of the prior on the true parameter instead of an
absolute bound. Thus this condition requires that the true parameter does not have an exponentially small prior value relative to other parameter values. We note that such
conditions allow for improper priors which were not allowed in
\cite{G2000}. Importantly, the uninformative prior trivially satisfies
Assumption P.
Next we state the main results of this section.
\begin{theorem}\label{Thm:Main:improv}
For any fixed value $c>0$, suppose that
$(i)$ $B_{1n}(c) \sqrt{{d}/n} = o(1)$,
$(ii)$ $\lambda_n(c) {d} = o(1)$,
$(iii)$ $\|F^{-1}\|_{op}{d}/n = o(1)$, and
$(iv)$ Assumption P($c$) hold. Then we have asymptotic normality of the posterior density
function
$$ \int_{\mathcal{U}} | \pi_n^*(u) - \phi_{d}(u;\Delta_n,I_{d})| du = o_p(1). $$
\end{theorem}
Theorem \ref{Thm:Main:improv} establishes the asymptotic normality of the posterior density
function. It has different
assumptions on the prior relative to Theorem 3 of \cite{G2000}. However, Theorem \ref{Thm:Main:improv} does not require additional technical
assumptions used in \cite{G2000}, as discussed in Appendix
A, and the growth condition of ${d}$ with relative
to the sample size $n$ is improved by $\ln {d}$ factors.
In some applications stronger
convergence properties for the posterior distribution can be required. The
following theorem provides sufficient conditions for the
$\alpha$-moment convergence. In what follows, for sequences of $\alpha$ and $d$, let
$M_{{d},\alpha} := (d+\alpha) \left( 1 + \frac{\alpha
\ln(d+\alpha)}{d+\alpha} \right).$
\begin{theorem}\label{Thm:Main:Alpha}
In addition to the conditions (i) and (iii) of Theorem \ref{Thm:Main:improv}, suppose that the following hold for any fixed $\bar{c}$: $(ii')$ $\lambda_n\left(\bar{c} M_{{d},\alpha}/{d}\right)
[\bar{c}M_{{d},\alpha}]^{1+\alpha/2}
=o(1)$; $(iv')$ Assumption P($\bar{c}M_{{d},\alpha}/{d}$) and {\small $K_n\left(\bar{c}M_{{d},\alpha}/{d}\right)\sqrt{\|F^{-1}\|_{op}\big[\bar{c}M_{{d},\alpha}\big]^{1+\alpha}/n} =o(1)$}. Then we have
\begin{equation}\label{Result:Alpha}
\int_{\mathcal{U}} \|u\|^\alpha | \pi_n^*(u) - \phi_{d}(u;\Delta_n,I_{d})| du = o_p(1).
\end{equation}
\end{theorem}
Conditions $(ii')$ and $(iv')$ are strengthening of conditions $(ii)$ and $(iv)$ of Theorem \ref{Thm:Main:improv} respectively.
We emphasize that Theorem \ref{Thm:Main:Alpha} allows for $\alpha$
and ${d}$ to grow as the sample size increases. Our conditions
highlight the polynomial trade off between $n$ and ${d}$ which contrasts with an
exponential trade off between $n$ and $\alpha$. This suggests that
the estimation of higher moments in increasing dimensions
applications could be very delicate. Conditions $(ii')$ and
$(iv')$ simplify significantly if $\alpha\ln d = o(d)$, in which case $M_{{d},\alpha} = d(1+o(1))$.
\begin{remark}
Suppose that we are interested in allowing $\alpha$ to grow with
the sample size as well. If ${d}$ is growing in a polynomial rate
with respect to $n$, our results do not allow for $\alpha = O(\ln
n)$. Some limitation along these lines should be expected since
there is an exponential trade off between $\alpha$ and $n$.
However, it is possible to have $\alpha = O(\sqrt{\ln n})$.
Such slow growth conditions illustrate the potential limitations
for the practical estimation of higher order moments.
\end{remark}
\section{Curved Exponential Family}\label{Sec:Curved}
Next we consider the curved exponential family. Let $X_1, X_2,\ldots,X_n$ be i.i.d. observations from a $d$-dimensional curved exponential family with density given by
$$
f(x;\theta) = h(x)\exp\left( \sp{x}{\theta(\eta)} -
\psi(\theta(\eta))\right),
$$
where $\eta \in \Psi \subset {\rm I\kern-0.18em R}^{{d}_1}$, $\theta:\Psi\to \Theta$,
an open subset of ${\rm I\kern-0.18em R}^d$, and $d \to \infty$ as $n \to \infty$ as
before.
The parameter of interest is $\eta$, whose true value $\eta_0$
is suitably bounded away from the boundary of $\Psi \subset {\rm I\kern-0.18em R}^{{d}_1}$ (see Assumption A). The true value of $\theta$ induced by $\eta_0$ is given by $\theta_0 =
\theta(\eta_0)$. The mapping $\eta \mapsto \theta(\eta)$ takes
values from ${\rm I\kern-0.18em R}^{{d}_1}$ to ${\rm I\kern-0.18em R}^{d}$ where $ {d}_1 \leq
{d}$. Moreover, assume that $\eta_0$ is the unique
solution to the system $\theta(\eta) = \theta_0$.
Thus, the parameter $\theta$ corresponds to a high-dimensional
parametrization, and $\eta$ describes
the lower-dimensional parametrization of the density.
We require the following regularity conditions on the mapping
$\theta(\cdot)$ and on the prior.
\begin{itemize}
\item[] \textbf{Assumption A}. For every fixed $\kappa$, there exists a linear
operator $G:{\rm I\kern-0.18em R}^{d_{1}}\to {\rm I\kern-0.18em R}^d$ such that uniformly
in $\gamma \in B_{{d}_1}(0,\kappa N_n) \subset \sqrt{n}(\Psi-\eta_0)$, where $N_n = \sqrt{{d}}+\{\sqrt{d_{prior}} + \sqrt{{d}_1\|(G'FG)^{-1}\|_{op}} +\sqrt{{d}_1\log ({d}_1+\|J^{-1}\|_{op}/\varepsilon_0)}\}\max\{1, \|J^{-1}\|_{op}/\varepsilon_0\}$,
where $\varepsilon_0$ is defined in Assumption B, we have
\begin{equation}\label{Curved:Rep}
\sqrt{n}\left( \theta(\eta_0 + \gamma/\sqrt{n})-\theta(\eta_0)
\right) = R_{1n} + (I+R_{2n})G\gamma, \end{equation} where
\begin{equation}\label{Curved:Cond1}
\begin{array}{l}
\|R_{1n}\| \|J\|_{op}\{\sqrt{d}+ \|JG\|_{op} N_n\} =o(1) \ \ \mbox{and}\\
\{\|R_{2n}\|_{op}\sqrt{d} + \|JR_{2n}G\|_{op}N_n\}\|JG\|_{op}N_n =o(1).
\end{array}\end{equation}
\item[] \textbf{Assumption B}. Uniformly in $n$, there exist a
positive constant $\varepsilon_0$ bounded away from zero, such that for every $\eta \in \Psi$
we have
\begin{equation}\label{Curved:Cond2} \| \theta(\eta) -
\theta(\eta_0) \| \geq \varepsilon_0 \| \eta -
\eta_0\|.\end{equation}
\item[] {\bf Assumption P'.} The prior density function $\pi$ satisfies:
$$\sup_{\eta \in \Psi} \ln
[\pi(\theta(\eta))/\pi(\theta(\eta_0))] \leq O({d}_{prior}).$$
\end{itemize}
Thus the mapping $\eta \mapsto \theta(\eta)$ is allowed to be
nonlinear and discontinuous. For example, the additional condition
of $R_{1n}=0$ implies the continuity of the mapping in a
neighborhood of $\eta_0$. More generally, condition
(\ref{Curved:Cond1}) impose that the map admits an (uniform)
approximate linearization in the neighborhood of $\eta_0$.
A prior $\pi$ on $\Theta$ induces a prior over $\Psi$ as $\pi(\eta) = \pi(\theta(\eta))/\int_\Psi \pi(\theta(\tilde \eta))d\tilde \eta$. Alternatively the prior can be placed directly over $\Psi$. Assumption P' also bounds the maximum log-likelihood given by the prior to any $\eta$ different than $\eta_0$ to be of the order ${d}_{prior}$ which can grow with $n$. If Assumption P holds we have ${d}_{prior} \leq {d}$. However, if the prior is placed directly on $\Psi$ we typically have ${d}_{prior} = {d}_1$. Finally, if the prior is uninformative we trivially have ${d}_{prior} = 1$. The posterior of $\eta$
given the data is denoted by
$$
\pi_n(\eta) \propto \pi(\theta(\eta)) \cdot \prod_{i=1}^n
f(x_i;\theta(\eta)) = \pi(\theta(\eta)) \cdot \exp\left( n \sp{\bar
X}{\theta(\eta)} - n\psi(\theta(\eta))\right)
$$ where $\bar X = (1/n)\sum_{i=1}^n X_i$.
Under this framework, we also define the local parameter space to
describe contiguous deviations from the true parameter as
$$
\gamma = \sqrt{n}(\eta - \eta_0), \ \ \mbox{and let} \ \ s =
(G'FG)^{-1}G'\sqrt{n}(\bar X - \mu)
$$
be a first order approximation to the normalized
maximum liklelihood/ex\-tre\-mum estimate. Under this setting, the following relations hold
for $s$: $$ E[s]= 0, \ \ E [s s'] = (G'FG)^{-1}, \ \ \mbox{and} \ \ \|s\| =
O_p(\sqrt{{d}_1\|(G'FG)^{-1}\|_{op}}).$$ The posterior density evaluated at $\gamma \in \Gamma := \sqrt{n}(\Psi-\eta_0)$ is given by $$ \pi^*_n(\gamma) =
\ell(\gamma)/\int_{\Gamma} \ell(\gamma) d \gamma, $$ where
$${\small \begin{array}{rl} \ell(\gamma) & = \exp\left( n\sp{\bar
X}{\theta(\eta_0 + n^{-1/2}\gamma)-\theta(\eta_0)} -n\left[
\psi(\theta
(\eta_0 + n^{-1/2}\gamma)) - \psi(\theta (\eta_0))\right] \right) \\
&\ \ \ \ \ \ \ \ \ \ \ \ \times \pi\left(\theta\left(\eta_0 + n^{-1/2}\gamma\right)\right).
\end{array}}$$
In order to formally state our results we use the following additional definition
$$ a_n = \sup \{ c : \lambda_n(c) \leq 1/16 \}. $$
The sequence $a_n$ characterizes a neighborhood of size
$\sqrt{a_n{d}}$ around the true local parameter for which the posterior $\ell(\cdot)$ is bounded above by a proper Gaussian density. This is useful since this neighborhood grows and controlling Gaussian tails is typically easier. In turn, this allows for weaker conditions in the next results.
Next we address the consistency question for the maximum
likelihood estimator associated with the curved exponential
family.
\begin{theorem}\label{Thm:CurvedConsistency}
Suppose that Assumptions A, B and P' hold. Then, provided the condition $\|(G'FG)^{-1}\|_{op}\{{d}_1+{d}_{prior}\}/d = o(a_n)$ holds, we have that the maximum likelihood estimator $\widehat \eta$ satisfies
$$\| \widehat \eta - \eta_0 \| = O_p\left(\sqrt{\frac{{d}_1+{d}_{prior}}{n}\|(G'FG)^{-1}\|_{op}}\right). $$
\end{theorem}
Two remarks regarding Theorem \ref{Thm:CurvedConsistency} are
worth mentioning. First, we note that the condition
$\lambda_n(c)=o(1)$ implies $a_n\to \infty$. However, $\lambda_n(c)=o(1)$ is stronger than the condition
$\sqrt{d/n}B_{1n}(c)=o(1)$ used for consistency obtained in \cite{G2000} for the exponential family case. Second, the consistency
result in Theorem \ref{Thm:CurvedConsistency} can be substantially impacted by the choice of prior. If the prior used is defined over the full space, it might place an exponentially small (in the dimension $d$) weight in $\eta_0$ relative to other points. In the case ${d}_1\sim {d}$ this is not problematic but it could impact the rates if ${d}_1 = o({d})$. In cases a prior can be placed directly over $\Gamma$ so that $d_{prior} = O({d}_1)$, we obtain the standard rate of convergence of $\sqrt{{d}_1/n}$.
Finally, we state the asymptotic normality result for the
curved exponential family. In what follows, let $M_n:=\{1+2\|JG\|_{op}\}^2\|(G'FG)^{-1}\|_{op} \{{d}_1+d_{prior}\}$.
\begin{theorem}\label{Thm:Cexpo}
Suppose that Assumptions A, B, and P' hold. Further suppose conditions (i)-(iii) of Theorem \ref{Thm:Main:improv} hold, and for each fixed $c>0$, P($cM_n/d$), $d_1\lambda_n(cM_n/d)=o(1)$ and $\{1+2\|JG\|_{op}\}^2N_n^2/d = o(a_n)$. Then, asymptotic
normality for the posterior density associated with the curved
exponential family holds,
$$ \int | \pi_n^*(\gamma) - \phi_{{d}_1}(\gamma; s,(G'FG)^{-1})| d\gamma=o_p(1).$$
\end{theorem}
\section{Applications to Selected Econometric Models}\label{Sec:Applications}
In this section we verify the conditions that lead to asymptotic normality in a variety of econometric problems covering both exponential and curved exponential families under increasing dimension. Most examples
are motivated by the classical work of \cite{zellner:text} on Bayesian econometrics.
\subsection{Multivariate Linear Model}\label{Sec:MLM}
In this section we consider a multivariate linear model. The response variable $y_i$ is a $d_y$-dimensional vector, the disturbances $u_i$ are normally distributed with mean zero and covariance matrix $\Sigma_0$. The covariates $z_i$ are $d_z$-dimensional and the parameter matrix of interest $\Pi_0$ is $d_z\times d_y$,
\begin{equation}\label{MLM}
y_i = z_i\Pi_0 + u_i \ \ \ i = 1,\ldots,n.
\end{equation} For notational convenience, let $Y$ and $Z$ denote the matrices whose rows are given by $y_i$ and $z_i$ respectively. Note that the dimension of the model is $d = d_y^2 + d_zd_y$.
Conditioning on the covariates $Z$, this model can be cast as an exponential family model by the following parametrization
\begin{equation}\label{Par:MLM}
\theta = \left( \begin{array}{c} \theta_1 \\ \theta_2 \end{array}\right) = \left( \begin{array}{c} -\frac{1}{2} \Sigma^{-1} \\ \Pi \Sigma^{-1} \end{array}\right), \ \ \bar{X} = \frac{1}{n}\left( \begin{array}{c} \bar{X}_{1} \\ \bar{X}_{2} \end{array}\right) = \frac{1}{n}\left( \begin{array}{c} Y'Y \\ Z'Y \end{array}\right)
\end{equation} and using the (trace) inner product $\sp{\theta}{X} = {\rm trace}(X_1'\theta_1) + {\rm trace}(X_2'\theta_2)$, see for instance \cite{Garderen}. This parametrization leads to the normalizing function
\begin{equation}\label{Par:MLM2}
\psi(\theta) = -\frac{1}{4n} {\rm trace}(Z\theta_2\theta_1^{-1}\theta_2'Z') - \frac{1}{2} \log \det ( -2\theta_1 ).
\end{equation}
We make the following assumptions on the design. Uniformly in $n$, the covariates satisfy $\max_{i\leq n} \|z_i\| = O(d_z^{1/2})$, the matrices $Z'Z/n$ and $\Sigma_0$ have eigenvalues bounded away from zero and from above, and the matrix $\Pi_0$ has full rank with singular values also bounded away from zero and from above.
Under these assumptions, by Lemma \ref{lemma:MultivariateLinearModelF} in the Appendix it follows that $\|F^{-1}\|_{op} = O(1)$.
Also, Lemma \ref{MultivariateLinearModelB1nB2n} in the Appendix bounds the quantities $B_{1n}(c) = O(d_z)$ and $B_{2n}(c) = O(d_z^2)$. Therefore we have asymptotic normality by Theorem \ref{Thm:Main:improv} provided that the condition
$d(d_z\sqrt{d/n}+d_z^2d/n) = o(1)$ holds.
\subsection{Seemingly Unrelated Regression Equations}\label{Sec:SUR}
The seemingly unrelated regression model (\cite{Zellner}) considers a collection of models
\begin{equation}\label{SUR}
y_{ik} = x_{ik}'\beta_k + u_{ik}, \ \ \ k=1,\ldots,d_y, \ \ i=1,\ldots,n,
\end{equation} where the dimension of $\beta_k$ is $d_k$. Let $d_z$ denote the total number of distinct covariates. The $d_y$-dimensional vector of disturbances $u$ has zero mean and covariance $\Sigma_0$ where $c<\lambda_{min}(\Sigma_0)\leq \lambda_{max}(\Sigma_0)\leq C$ for some fixed constants $c>0$ and $C<\infty$ independent of $n$. This model can be written in the form of (\ref{MLM}) by setting $\Pi = [ \pi_1(\beta_1) ; \pi_2(\beta_2) ; \cdots ; \pi_{d_y}(\beta_{d_y}) ]$. Note that the vector $ \pi_k(\beta_k)$ has zeros for regressors that do not appear in the $k$th model.
\cite{Garderen} shows that this model is a curved exponential model provided that the matrix $\Pi$ has some zero restrictions.
Consider the assumptions of Section \ref{MLM}. In this case we have that
\begin{equation}\label{SUR:c}
\eta = \left( \begin{array}{c} \eta_1 \\ \eta_2 \end{array}\right)= \left( \begin{array}{c} \Sigma^{-1} \\ \Pi \end{array}\right), \ \ \theta(\eta) = \left( \begin{array}{c} -\frac{1}{2}\eta_1 \\ \eta_2\eta_1 \end{array}\right)= \left( \begin{array}{c} -\frac{1}{2} \Sigma^{-1} \\ \Pi \Sigma^{-1} \end{array}\right).
\end{equation}
We restrict the space of $\Sigma$ to consider $\lambda_{min}(\Sigma) > \lambda_{min}$ a fixed constant (note that this induces $\lambda_{max}(\Sigma^{-1}) < 1/ \lambda_{min}$ which leads to a convex region in the parameter space), and that operator norm of $\Pi$ is bounded by a constant, $\|\Pi\|_{op}\leq M$ a fixed constant.
The mapping $\theta(\cdot)$ is twice differentiable and Lemma \ref{Lemma:SURtheta} establishes that condition (\ref{Curved:Rep}) holds with $R_{2n} = 0$ and $\|R_{1n}\| \leq O(N_n^2/\sqrt{n})$. Provided $\|G\|_{op}\leq C$, this implies that the requirement of $d^3\log^3 d = O(\{d_y^6 + d_z^3d_y^3\}\log^3(d_y+d_z))=o(n)$ suffices for Assumption A to hold.
Next we verify Assumption $B$ for $\varepsilon_0 = \min\{\frac{1}{4},\lambda_{min}(\Sigma_0^{-1}) / [8(1+M)]\}$. Direct calculations provide$$
\begin{array}{rcl}
\|\theta(\eta) - \theta(\eta_0)\| & \geq & \max \{ \frac{1}{2}\| \eta_1 - \eta_{01} \|, \|\eta_2\eta_1 - \eta_{02}\eta_{01}\| \}. \\
\end{array}
$$We can assume that $\|\eta_1 - \eta_{01}\| < 2\varepsilon_0 \|\eta - \eta_0\|$ otherwise Assumption B holds. In turn, this implies that $\|\eta_2 - \eta_{02}\| \geq \frac{1}{2}\|\eta - \eta_0\|$. In this case, since the operator norm of $\eta_2$ satisfies $\|\eta_2\|_{op}\leq M$, we have
$$
\begin{array}{rcl}
\|\theta(\eta) - \theta(\eta_0)\| & \geq & \|\eta_2\eta_1 - \eta_{02}\eta_{01}\| \\
& = & \| \eta_2 ( \eta_1 - \eta_{01}) + ( \eta_{2} - \eta_{02})\eta_{01}\| \\
& \geq & \| \eta_{2} - \eta_{02} \| \lambda_{min}(\Sigma_0^{-1}) - \|\eta_2\|_{op}\| \eta_1 - \eta_{01} \| \\
& \geq & \| \eta - \eta_{0} \| \lambda_{min}(\Sigma_0^{-1})/2 - M2\varepsilon_0\|\eta - \eta_{0} \|\\
& \geq & \| \eta - \eta_{0} \| \lambda_{min}(\Sigma_0^{-1})/4 \geq \varepsilon_0 \| \eta - \eta_{0} \| \\
\end{array}
$$ which verifies Assumption B.
\subsection{Single Structural Equation Model}\label{Sec:SSEM}
Next we consider the single structural equation,
\begin{equation}\label{SSEM}
y^{(1)}_i = {y^{(2)}_i}'\beta + {z^{(1)}_i}'\gamma + v_i, \ \ i=1,\ldots,n,
\end{equation}
for which the associated reduced form system, given by the multivariate linear model in (\ref{MLM}), can be partitioned as
\begin{equation}\label{SSEM-red}
(y^{(1)}_i \ \ {y^{(2)}_i}' ) = ({z^{(1)}_i}' \ \ {z^{(2)}_i}') \left( \begin{array}{cc} \pi_{11} & \pi_{12} \\ \pi_{21} & \pi_{22}\\ \end{array}\right) + ( u^{(1)}_i \ \ {u^{(2)}_i}' ).
\end{equation}
We assume full column rank of $Z$ and ${\rm rank}(\pi_{21} \ \ \pi_{22}) = {\rm rank}(\pi_{22}) = d_y-1$ where $d_y$ is the dimension of $(y^{(1)}_i \ \ {y^{(2)}_i}')$. The compatibility between the models (\ref{SSEM}) and (\ref{SSEM-red}) requires that
$$ \pi_{11} = \pi_{12}\beta + \gamma, \ \ \pi_{21} = \pi_{22} \beta, \ \ \mbox{and} \ \ u^{(1)}_i = {u^{(2)}_i}'\beta + v_i. $$
The model can also be embedded in (\ref{MLM}) as follows
\begin{equation}\label{SUR:c}
\eta = \left( \begin{array}{c} \eta_1 \\ \eta_2 \\ \eta_3 \end{array}\right)= \left( \begin{array}{c} \Sigma^{-1} \\ \left( \begin{array}{c} \pi_{12} \\ \pi_{22} \end{array}\right) \\ \left( \begin{array}{c} \gamma\\ \beta \end{array}\right) \\ \end{array}\right), \ \ \theta(\eta) = \left( \begin{array}{c} -\frac{1}{2} \Sigma^{-1} \\ \left( \begin{array}{cc} \gamma + \pi_{12}\beta & \pi_{12} \\ \pi_{22}\beta & \pi_{22}\\ \end{array}\right) \Sigma^{-1} \end{array}\right).
\end{equation}
Similar arguments to those used in Section \ref{Sec:SUR} show that Assumptions A and B hold
under the standard strong instrument asymptotics.
\subsection{Multinomial Model}\label{Sec:MM}
This example of multinomial model was also analyzed in \cite{G2000}. Our goal is to weaken some of the conditions required previously using the techniques proposed here.
Let $\mathcal{X} =
\{ x^0, x^1,\ldots, x^{d}\}$ be the known finite support of a
multinomial random variable $X$ where $d$ is allowed to grow with
sample size $n$. For each $j$ denote by $p_j$ the probability of
the event $\{ X = x^j \}$ which is assumed to satisfy $\max_{0\leq j\leq {d}}
1/p_j = O({d})$. The parameter space is given by $\theta =
(\theta_1,\ldots,\theta_{d})$ where $\theta_j =
\log(p_j/(1-\sum_{k=1}^{{d}}p_k))$. It follows that under the assumption on the
$p_j$'s the true value of $\theta_j$'s are bounded. The Fisher
information matrix is given by $F = P - pp'$ where $P={\rm
diag}(p)$. In this case we have $B_{1n}(c)
= O({d}^{3/2})$ and $B_{2n}(c) = O({d}^2)$. We refer to \cite{G2000} for detailed calculations.
The growth condition ${d}^6(\log {d})/n \to 0$ was imposed
in \cite{G2000} to obtain the asymptotic normality results (the case of
$\alpha=0$). We weaken this growth requirement by combining
the derivation in \cite{G2000} with the analysis in Section \ref{Sec:Exp} with an uninformative
(improper) prior. In this case we have $K_n(c) = 0$ and our
definition of $\lambda_n(c)$ remove the logarithmic factors.
As a result, Theorem \ref{Thm:Main:improv} leads to a weaker growth
condition ${d}^4/n\to 0$. For
$\alpha$-moment estimation, the conditions of Theorem \ref{Thm:Main:Alpha} are
satisfied with the condition that ${d}^{4+\alpha+\delta}/n \to 0$
for any strictly positive value of $\delta$. Recently another approach based on Le Cam's proof that is specific to discrete probability distributions allows for further improvements, see \cite{BoucheronGassiat2009}.
\subsection{Multinomial Model with Moment Restrictions}\label{Sec:MMMR}
In this subsection we provide a high-level discussion of the
multinomial model with moment restrictions. Let $\mathcal{X} = \{ x^0,
x^1, x^2, \ldots, x^d\}$ be the known finite support of a
multinomial random variable $X$ which was described in Section
\ref{Sec:MM}. Conditions $(i)-(iv)$ are verified as in Section \ref{Sec:MM}.
As discussed in the introduction, it is of interest to incorporate
moment restrictions into this model, see \cite{chamberlain} and \cite{imbens} for discussions. This will lead to a curved exponential model as
studied in Section \ref{Sec:Curved}.
The parameter of interest is $\eta \in \Psi \subset {\rm I\kern-0.18em R}^{{d}_1}$ a
compact set. Consider a (twice continuously differentiable)
vector-valued moment function $m:\mathcal{X} \times \Psi \to {\rm I\kern-0.18em R}^M$ such
that $$E[m(X,\eta)] = 0\ \ \mbox{for a unique} \ \eta_0 \in
\Psi.$$ The log-likelihood function
associated with this model
\begin{equation}\label{Def:MMlog}
\begin{array}{l}
\displaystyle l(q,\eta) = \sum_{i=1}^n \sum_{j =0}^d I\{X_i = x_j\} \ln
q_j \\
\displaystyle \mbox{for some} \ q \ \mbox{and}\ \eta \ \mbox{such that} \
\sum_{j=0}^d q_j m(x_j,\eta) = 0, \ \sum_{j=0}^d q_j = 1, \ q\geq 0,
\end{array}
\end{equation} and $l(q,\eta)=-\infty$ if the probability distribution $q$ violates any of the
moments conditions. The log-likelihood function (\ref{Def:MMlog}) induces the mapping $q:\Psi \to \Delta^{d-1}$ formally defined as
\begin{equation}\label{Def:ThetaMM}
\begin{array}{rl}
\displaystyle q(\eta) = \arg\max_{q}& l(q,\eta) \\
& \displaystyle \sum_{j=0}^d q_j m(x_j,\eta) = 0, \ \sum_{j=0}^d
q_j = 1, \ q \geq 0.
\end{array}
\end{equation}
In this case, the function $\theta_j(\eta)
= \log( q_j(\eta) / q_0(\eta))$ (for $j=1,\ldots,d$) is the
mapping from $\Psi \to \Theta$ discussed in Section \ref{Sec:Curved}. Assuming that
the matrix $E\left[ m(X,\eta)m(X,\eta)'\right]$ is uniformly positive
definite over $\eta$, \cite{Qin-Lawless} use the
inverse function theorem to show that $\theta(\cdot)$ is a twice
continuous differentiable mapping of $\eta$ in a neighborhood of
$\eta_0$. In particular this implies that we can take $R_{2n}=0$ and $\|R_{1n}\| = O( {d} {d}_1^2 (N_n^2/n) )$.
Thus, provided $G'FG$ has eigenvalues bounded away from zero and from above, $d_{prior}\leq d$, and because $\|F^{-1}\|_{op}=O({d})$, Assumption A holds provided that $\|R_{1n}\| N_n = O({d}^3{d}_1^3\log {d} /n) = o(1)$.
Assumption B is satisfied if $\eta$
belongs in a compact set $\Psi$ and that the mapping $\theta(\cdot)$ is
injective (over a set that contains $\Psi$ in its interior). We
refer to \cite{Newey-McFadden} for a discussion
of primitive assumptions for identification with moment
restrictions.