Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
90,614 characters · 18 sections · 38 citation commands
Optimal Estimation under a Semiparametric Density Ratio Model
{\bf \em Keywords}: Biased sampling; Empirical likelihood; Exponential family; Exponential tilting; Statistical efficiency.
In numerous applications, scientists may independently collect a single sample from each of the connected and similar populations. The dataset hence contains independent multiple samples. For example, to gauge the economic state of a country, econometricians may study the evolution of income distribution over time roine2015long. They may hence collect cross-sectional or panel survey data from year to year wooldridge2010econometric, and the underlying populations for these multiple samples are naturally connected. There are various approaches to data analysis of this nature. One may postulate a parametric model such as normal or gamma on these populations davison2003statistical. However, some statistical inference procedures such as quantile estimation may have poor performance when the model assumptions are merely mildly violated. For instance, statistical analyses of a dataset under two similar models, log-normal and gamma models, can yield markedly different results wiens1999log. To mitigate the risks of model misspecification, one may take a nonparametric approach wasserman2006all by not imposing any parametric assumption. However, the nonparametric approaches ignore the latent structures shared by the multiple populations from which the samples are collected, and hence fail to utilize the potential gain in statistical efficiency. As a balanced trade-off between the model misspecification risks and the statistical efficiency, we advocate semiparametric approaches hardle2004nonparametric, tsiatis2006semiparametric.
Recently, there have been many research studies on the semiparametric density ratio model (DRM) anderson1979multivariate. Let $ g_0 (x), \ldots, g_m (x) $ be the density functions for the multiple populations. The DRM postulates that
where $ \mathbf{q} (x) $ is a prespecified vector-valued function and $ \bm{\theta}_k $ is an unknown vector-valued parameter. The DRM effectively utilizes the latent structures shared by the multiple populations. Not surprisingly, methods developed under the DRM are often statistically more efficient than the nonparametric methods that ignore the latent structures and use the samples separately for individual populations. In particular, the DRM-based empirical likelihood (EL) approaches owen2001empirical are found to have nice asymptotic properties and superior performances qin1993empirical, chen2013quantile, cai2017hypothesis. Under the framework of the DRM and EL, the multiple samples are used together to draw inferences on each population.
These results lead to an interesting question: what is the limit of the efficiency gain through the DRM-based approach? A gold-standard target to consider is the parametric efficiency. As a semiparametric model, a DRM naturally contains many parametric models, with each forming a so-called parametric submodel. For every such parametric submodel, there is a classical Cram\'{e}r--Rao lower bound. As a result, the asymptotic variance for any consistent and asymptotically normal semiparametric estimator is no smaller than the supremum of the Cram\'{e}r--Rao lower bounds for all parametric submodels. Such a supremum is famously known as the semiparametric efficiency bound newey1990semiparametric. When a parametric submodel contains the true distributions that generate the data and is regular, the parametric maximum likelihood estimator (MLE) has the smallest asymptotic variance that attains the Cram\'{e}r--Rao lower bound casella2002statistical. Therefore, no methods can achieve higher efficiency under the DRM than it can under its parametric submodel that the true distributions reside. This gives an efficiency upper bound for any semiparametric approach including the DRM-based estimators. Motivated by this observation, our research question becomes: is it likely that the DRM-based approaches achieve parametric efficiency under a parametric submodel? If yes, when would that happen?
In this article, we focus on the scenario where the sample size from one population is significantly larger than the sample sizes from the other populations. We are interested in the efficiency of the DRM-based estimators for the populations with smaller samples. In other words, we want to quantify how much help the large sample population can offer to improve the efficiency of the estimators for the small sample populations. The motivation for us to consider this scenario is as follows. Suppose the sample size from one population is infinite. The corresponding population distribution can therefore be regarded as known. Consequently, treating the infinite-sample population density as $ g_0 (x) $, the semiparametric DRM in (ref) would reduce to a parametric model for the other populations. We therefore expect that the DRM-based estimators for the small sample populations will achieve parametric efficiency under the corresponding parametric models. This article establishes and proves this claim for the model parameters, distribution functions, and quantiles. Although we prove these results mostly in a two-sample scenario for clarity, our discoveries are general for the multi-sample scenario. More specifically, when one population has an extra large sample, the DRM-based estimators for the other populations are as efficient as the best possible under the true parametric submodel. For convenience, we refer to this scenario as the two-sample scenario throughout the article.
In the existing literature, there have been some discussions on the asymptotic efficiency of the DRM-based approaches, such as zhang2000quantile and chen2013quantile. They focus on situation where the sample sizes all go to infinity at the same rate and find that the DRM-based quantile estimators are asymptotically more efficient than the nonparametric empirical quantiles. However, none of these studies systematically investigate the boundary of the efficiency gain under the DRM. This article tackles this important research problem by examining the aforementioned two-sample scenario. Furthermore, their theoretical results rely on the assumption that the ratio of any two sample sizes remains fixed, positive, and finite as the sample sizes aproach infinity. In contrast, we study the asymptotic efficiency of the DRM-based approaches in a different and more general context that allows the sample sizes to grow at different rates and allows the ratios of sample sizes to evolve.
The rest of the article is organized as follows. In Section (ref), we elaborate more on the DRM and state the research problem. For comparison, in Section (ref) we present the gold-standard parametric efficiency under a specific parametric submodel of the DRM. In Section (ref), we give an overview of the EL-based inferences under the DRM, and study the asymptotic efficiency of the EL-DRM estimators of the model parameters, distribution functions, and quantiles. For illustration, we show in Section (ref) how the asymptotic variance of the DRM-based quantile estimator evolves as a function of the ratio of the sample sizes when the true population distributions $ G_k $ are all identical but not assumed to be identical in the estimation. We conduct some simulation studies in Section (ref) by generating data from normal and exponential distributions, whose results support our theoretical findings. In Section (ref), we report results from a real-data analysis on the revenues of U.S. collegiate sports. Finally, Section (ref) concludes our contributions in the article with a discussion. Proofs of the theoretical results are provided in \hyperref[sec:appendix]{Appendix}.
Suppose we have $ m+1 $ independent sets of independent and identically distributed (i.i.d.) samples respectively from population distributions $ G_0, \ldots, G_m $:
Let $ g_0 (\cdot), \ldots, g_m (\cdot) $ be the corresponding density functions with respect to some common $ \sigma $-finite measure. They satisfy the DRM if
with $ \mathbf{q} (x) $ a prespecified vector-valued function of dimension $ d $ and $ \bm{\theta}^{\top} \coloneqq (\bm{\theta}_1^{\top}, \ldots. \bm{\theta}_m^{\top}) $ an unknown vector-valued parameter. The first element of $ \mathbf{q} (x) $ is set to be 1 so that the corresponding coefficient in $ \bm{\theta}_k $ is a normalization constant. Therefore, we may write $ \mathbf{q} (x) $ as
for some vector-valued function $ \mathbf{q}_{-}^{\top} (x) $ with dimension $ d-1 $. Correspondingly, we may also decompose the parameter $ \bm{\theta}_k $ into $ (\alpha_k, \bm{\beta}_k) $. We also require the elements of $ \mathbf{q} (x) $ to be linearly independent; otherwise, some elements of $ \mathbf{q} (x) $ are redundant. By convention, we call $ G_0 $ the base distribution and $ \mathbf{q} (x) $ the basis function. The DRM and its inference, which we introduce later, are invariant to the choice of the base distribution. In this article, we select the population distribution from which we have a much larger sample as the base distribution.
The DRM covers many well-known parametric distribution families, with different choices of the base distribution $ G_0 $ and the basis function $ \mathbf{q} (x) $. For example, the Poisson distribution family satisfies the DRM with $ G_0 $ being Poisson and $ \mathbf{q} (x) = (1, x)^{\top} $; the normal distribution family satisfies the DRM with $ G_0 $ being normal and $ \mathbf{q} (x) = (1, x, x^{2})^{\top} $; and the gamma distribution family also satisfies the DRM with $ G_0 $ being gamma and $ \mathbf{q} (x) = (1, x, \log x)^{\top} $. In fact, any exponential family model satisfies the DRM. On the other hand, when using DRM, we do not assume that each population distribution belongs to any parametric distribution family. Instead of directly modelling the distribution of each population, the DRM takes a semiparametric approach to modelling the relationship between the multiple populations. Therefore, we may see that the DRM is a very flexible model and has a low risk of model misspecification. A more appealing feature of the DRM is that, with an appropriately chosen basis function $ \mathbf{q} (x) $, the DRM allows us to utilize the combined data to estimate each distribution $ G_k $. This would lead to efficiency gain compared with estimating $ G_k $ individually using the sample from itself. The DRM has been widely applied in many fields; see the books sugiyama2012density and qin2017biased. It is also closely connected with other semiparametric models such as models for biased sampling vardi1982nonparametric, vardi1985empirical, the exponential tilting model rathouz2009generalized, and the proportional likelihood ratio model luo2012proportional. Many studies on a variety of inference problems in the literature have confirmed this efficiency gain through either theoretical or empirical analyses, including zhang2000quantile, chen2013quantile, cai2017hypothesis.
The research problem we consider in this article is: when the sample from $ G_0 $ has a much larger size than the samples from the other $ G_k $, would the DRM-based estimators of some functionals of $ G_k $ achieve parametric efficiency under a parametric submodel? Without loss of generality and for clarity, hereafter we present our results mostly in the two-sample scenario, namely $ m = 1 $. Our results are general for the multiple samples as long as we have a much larger sample from one population than the others. We denote by $ n_k $ the size of the sample from $ G_k $, and let $ n = \sum_{k=0}^{m} n_k $ be the total sample size. Formally, when $ G_0, G_1 $ satisfy the DRM, we study the asymptotic efficiency of the estimators of some functionals of $ G_1 $ when
The functionals of $ G_1 $ we consider include the model parameters $ \bm{\theta} $, the cumulative distribution function $ G_1 (x) $ for every $ x $, and the quantiles of $ G_1 $.
The research problem has many applications in statistics and econometrics. For example, an important aspect of econometrics is to analyze cross-sectional and panel data with some structures ng2006testing, chen2012semiparametric, lin2012estimation, boneva2015semiparametric, su2016identifying, gao2020heterogeneous, su2023identifying, which is useful for areas of economic research such as understanding the economic state of a population at specific time points and evaluating the effects of a policy across different groups, regions, or time periods. The DRM has been found successful for the analysis of such data in the long-term monitoring of lumber strength properties zidek2018statistical, chen2022permutation. This is because the DRM-based methods are able to borrow strength across the multiple cross-sectional samples to exploit the shared latent structures between different populations over years and regions. Therefore, it is of scientific significance to quantify the resulting efficiency gain when analyzing such type of data. Our contributions in this article achieve this aim by focusing on the situation where one wishes to make inferences on a population with a small data set, aided by a large data set from another connected population. In the real-data analysis section of this article, we study the efficiency of the DRM-based quantile estimators under such scenario based on a revenue data from U.S. collegiate sports.
To study the efficiency of the DRM-based inferences, we first reveal the efficiency limit under its parametric submodels. A parametric submodel is a parametric family of distributions we select for the data that satisfy the semiparametric model (in our case, the DRM) assumptions. In this article, the parametric submodel we consider for $ G_0, G_1 $ is an exponential family model:
where $ \bm{\eta}_0, \bm{\eta}_1 $ are the unknown natural parameters to be estimated, $ \mathbf{q}_{-} (x) $ is the sufficient statistic for a single observation $ x $ that is the same as the one in (ref) in the DRM, $ B (\cdot) $ is a known nonnegative function, and $ A (\bm{\eta}_k) $ is uniquely determined by $ \bm{\eta}_k $: \[ A (\bm{\eta}_k) = -\log \int B (x) \exp \{ \bm{\eta}_k^{\top} \mathbf{q}_{-} (x) \} \mathrm{d} x, \] such that $ g_k $ is indeed a density function for $ k = 0, 1 $.
{\bf Remark}: the results presented in the article are still valid if we consider another set of parametric submodel: $ g_0 (x) $ being completely known and $ g_1 (x) $ from an exponential family model. This submodel consists of joint distributions of two independent set of i.i.d. samples in the form of \( \prod_{i=1}^{n_0} g_0 (x_{0i}) \prod_{j=1}^{n_1} g_0 (x_{1j}) \exp \{ \alpha + \bm{\beta}^{\top} \mathbf{q}_{-} (x_{1j} ) \}. \)
With the decomposition of $ \mathbf{q} (x) $ and $ \bm{\theta} $ in (ref), we get an equivalent form of the DRM in (ref):
where $ \alpha $ is a normalization constant determined by \[ \alpha = -\log \int g_0(x) \exp \{ \bm{\beta}^\top \mathbf{q}_{-} (x) \} \mathrm{d} x. \] Therefore, the exponential family model in (ref) satisfies the DRM in (ref) with \[ \bm{\theta} =
=
. \] In other words, the exponential family model in (ref) is indeed a parametric submodel of the DRM. In this section, we obtain the asymptotic results for $ G_1 $ under this parametric submodel.
We first study the efficiency of the parametric MLE of $ \bm{\theta} $ under the parametric submodel (ref). This will be our gold-standard target with which we will compare the efficiency of the DRM estimator of $ \bm{\theta} $ to be presented later. We obtain a parametric estimator of $ \bm{\theta} $ using the maximum likelihood method. The parametric log-likelihood function given the two independent i.i.d. samples is
From this, we get the score function: for $ k = 0, 1 $, \[ \frac{\partial \ell_{\mathrm{para}} (\bm{\eta}_0, \bm{\eta}_1)}{\partial \bm{\eta}_k} = n_k \frac{\partial A (\bm{\eta}_k)} {\partial \bm{\eta}_k} + \sum_{j=1}^{n_k} \mathbf{q}_{-} (x_{k j}). \] We let $ \tilde \bm{\eta}_k $ denote the parametric MLE of $ \bm{\eta}_k $, and they satisfy
We remark that for convenience, with a generic function $ f (\boldsymbol{y}) $, we have used the notation \[ \frac{\partial f (\boldsymbol{y}^{*}) }{\partial \boldsymbol{y}} = \left. \frac{\partial f (\boldsymbol{y}) }{\partial \boldsymbol{y}} \right \rvert_{\boldsymbol{y} = \boldsymbol{y}^{*}} \] in the preceding expression, and we retain the use of such notation hereafter.
By the invariance property of maximum likelihood estimation (see Theorem (ref) below), which states that any function of MLE is still an MLE of the function of the parameter, we must have that the MLE of $ \bm{\theta} $, denoted as $ \tilde \bm{\theta} $, is given by
Before we formally state the invariance property of the MLEs, we first introduce the so-called induced likelihood casella2002statistical, which is also known as the profile likelihood in the literature barndorff1988parametric. Suppose a distribution family is indexed by a parameter $ \theta $, and we want to find the MLE of some function of $ \theta $, denoted by $ \tau (\theta) $. When $ \tau (\cdot) $ is one-to-one, then given the MLE of $ \theta $ being $ \hat \theta $, it can be seen that the MLE of $ \tau (\theta) $ is $ \tau (\hat \theta) $. When $ \tau (\cdot) $ is not one-to-one, there are some technical issues under the current definition of likelihood. To overcome these problems, we need a more general definition of the likelihood function of $ \tau (\theta) $, called the induced likelihood function, and correspondingly a more general definition of the MLE. They are defined as follows.
With this extended definition of the MLE, we are now ready to formally state the invariance property of MLEs in the following theorem.
After obtaining the parametric MLE $ \tilde \bm{\theta} $, we now study its asymptotic properties. By the standard results on maximum likelihood estimation theory, under the exponential family model in (ref) and suppose some regularity conditions are satisfied, the MLE $ \tilde \bm{\eta}_k $ is asymptotically normal:
as $ n_k \to \infty $, where $ \bm{\eta}_k^{*} $ denotes the true value of $ \bm{\eta}_k $. Applying the delta method casella2002statistical, we have that $ (A (\tilde \bm{\eta}_k), \tilde \bm{\eta}_k) $ is also asymptotically normal: \[ \sqrt {n_k}
\overset {d} \to N \left ( \bm {0},
{\mathrm {Var}}_k^{-1} [\mathbf{q}_{-} (X)]
\right ), \] as $ n_k \to \infty $. Note that here we have used a standard result on the exponential family model:
Taking advantage of the nice mathematical properties of exponential families, we forego other otherwise necessary verifications.
Finally, because the two samples are independent of each other, we have
where $ n_1/n_0 \to 0 $, as $ n_0, n_1 \to \infty $. We will compare the asymptotic variance in (ref) of the parametric MLE $ \tilde \bm{\theta} $ with the asymptotic variance of the DRM-based estimator of $ \bm{\theta} $ later.
We next study the inference on the distribution function $ G_1 (x) $ under the two-sample exponential family model in (ref). By Theorem (ref), the MLE of $ G_1 (x) = G_1 (x; \bm{\eta}_1^{*}) $ is given by \[ \tilde G_1 (x) = G_1 (x; \tilde \bm{\eta}_1) = \int_{-\infty}^{x} B (t) \exp \{ \tilde \bm{\eta}_1^{\top} \mathbf{q}_{-} (t) + A (\tilde \bm{\eta}_1) \} \mathrm {d} t, \] where we recall that $ \tilde \bm{\eta}_1 $ is the MLE of $ \bm{\eta}_1 $.
The limiting distribution of such a parametric distribution estimator $ \tilde G_1 (x) $ can be straightforwardly obtained, as follows. Recall that the MLE $ \tilde \bm{\eta}_1 $ is asymptotically normal: \[ \sqrt {n_1} (\tilde \bm{\eta}_1 - \bm{\eta}_1^{*}) \overset {d} \to N (\bm {0}, {\mathrm {Var}}_1^{-1} [\mathbf{q}_{-} (X)]), \] as $ n_1 \to \infty $. Applying the delta method, we have that for every fixed $ x $, the MLE $ \tilde G_1 (x) $ is also asymptotically normal:
Under regularity conditions, we can further simplify the variance matrix in (ref) as follows. First, we note that
by the definition of $ \mathbf{Q} (x) $ given in Theorem (ref). Further, recall from (ref) that $ \partial A (\bm{\eta}_1^{*}) / \partial \bm{\eta}_1 = - {\mathbbm E}_1 [\mathbf{q}_{-} (X)]. $ Therefore, the preparation results in (ref) and (ref) together lead to a simpler expression for the variance matrix in (ref):
We will examine later that the DRM-based estimator of $ G_1 (x) $ attains the same asymptotic efficiency as above.
Quantiles are very important population parameters in many applications. In this section, we study the asymptotic efficiency of the parametric quantile estimator under the exponential family model in (ref). We first define, for any quantile level $ p \in (0, 1) $, the $ p $th quantile of $ G_1 $ as \[ \xi_p = G_1^{-1} (p) \coloneqq \inf \{ t: G_1 (t) \geq p \}, \] where the inverse function $ G_1^{-1} (\cdot) $ is also known as the quantile function. We assume that the distribution $ G_1 (x) $ has a density function $ g_1 (x) $ that is positive and continuous at $ x = \xi_p $. Then, by Theorem (ref), the MLE of $ \xi_p = G_1^{-1} (p; \bm{\eta}_1^{*}) $ is given by \[ \tilde \xi_p = G_1^{-1} (p; \tilde \bm{\eta}_1), \] where $ \tilde \bm{\eta}_1 $ is the MLE of $ \bm{\eta}_1 $ under the parametric model in (ref).
Because the MLE $ \tilde \bm{\eta}_1 $ is asymptotically normal (see (ref)), the quantile MLE $ \tilde \xi_p $ is also asymptotically normal by the delta method:
In addition, the variance matrix in (ref) has a more specific expression, as follows. Firstly,
Then, from previous derivations given in (ref) and (ref), we have
Hence, the variance matrix in (ref) can be written as
Our objective, to be realized later, is to investigate whether the DRM-based estimator of $ \xi_p $ has the asymptotic variance as low as the variance in (ref) and equivalently in (ref).
Under the DRM, the base distribution $ G_0 $ is left unspecified. If we impose a parametric form on $ G_0 $, the DRM would reduce to a parametric model. In this case, we gain model simplicity but bear a higher risk of model misspecification. We thus leave $ G_0 $ unspecified and use the nonparametric EL of owen1988empirical as a platform for statistical inference under the DRM. There has been a rich literature on the coupling of the EL and the DRM; see, for example, qin1993empirical, qin1997goodness, qin1998inferences, fokianos2001semiparametric.
We first review the EL method under the DRM. For convenience, let $ p_{kj} = \mathrm{d} G_0 (x_{k j}) = P (X = x_{k j}; G_0) $, the probability of observing $ x_{k j} $ under $ G_0 $ for all applicable $ k, j $. Applying the likelihood principle, we obtain the EL based on the multiple sample under the DRM:
with $ \bm{\theta}_0 \coloneqq \bm{0} $ by convention. We observe that $L_{n} (\cdot) = 0$ if $ G_k $ are continuous distribution functions. This seemingly devastating property does little harm to the usefulness of the EL. As we will see, in the EL we search for the distribution estimator within the space of discrete distributions that assign positive probability mass to the observed data. This does not eliminate much generality because every distribution can be precisely approximated by such discrete distributions when the sample size grows. Since $ L_{n} (\cdot) $ in (ref) is also a function of the parameters $ \bm{\theta} $ and base distribution $ G_0 $, we may also write its logarithm as
The EL-based inferences on the population parameters are usually carried out through a profile likelihood function. We first observe that the DRM assumption in (ref) implies \[ 1 = \int \mathrm{d} G_0 (x), \hspace{5mm} 1 = \int \mathrm{d} G_r = \int \exp \{ \bm{\theta}_r^{\top} \mathbf{q} (x) \} \mathrm{d} G_0 (x), \hspace{3mm} r = 1, \ldots, m. \] Confining the common support of $ G_0, \ldots, G_m $ to the observed data $ \{ x_{k j} \}_{k, j} $ yields the EL-version constraints \[ \sum_{k, j} p_{k j} = 1, \hspace{5mm} \sum_{k, j} p_{k j} \exp \{ \bm{\theta}_r^{\top} \mathbf{q} (x_{k j}) \} = 1, \hspace{3mm} r = 1, \ldots, m. \] With these preparations and following the foundational work by qin1994empirical, we define the profile log-EL function of $ \bm{\theta} $ as the supremum of the log-EL $ \ell_{n} (\bm{\theta}, G_{0}) $ in (ref) over $ G_0 $ subject to the above constraints:
The above optimization problem has a simple solution by the Lagrange multiplier method. For the best relevance, we present the results when $ m = 1 $. We have
for some $ \hat \lambda_1 $ satisfying $ \sum_{k, j} 1/[ n + \hat \lambda_1 ( \exp \{ \bm{\theta}^{\top} \mathbf{q} (x_{k j}) \} - 1)] = 1 $. One may estimate $ \bm{\theta} $ by the maximum empirical likelihood estimator (MELE):
At $ \bm{\theta} = \hat \bm{\theta} $, some algebra gives $ \hat \lambda_1 = n_1 $, and we naturally get another function by replacing $ \hat \lambda_1 $ with $ n_1 $ in the profile log-EL $ \tilde \ell_{n} (\bm{\theta}) $:
The profile log-EL $ \tilde \ell_{n} (\bm{\theta}) $ and $ \ell_{n} (\bm{\theta}) $ have the same maximum value and maximizer. Because of this, we study the asymptotic properties of the MELE $ \hat \bm{\theta} $ through the analytically simpler $ \ell_{n} (\bm{\theta}) $. By convention, we call this function a {\em dual function} of the profile log-EL. Following chen2013quantile, we regard this dual function $ \ell_{n} (\bm{\theta}) $ as if it is the profile log-EL.
Our ultimate goal is to study the asymptotic efficiency of the DRM-based inferences of $ G_1 $ that has a smaller sample under the two-sample scenario when $ n_0/n_1 \to \infty $. To achieve this goal, we first show that the MELE $ \hat \bm{\theta} $ under the DRM is asymptotic normal with the same low variance as the MLE $ \tilde \bm{\theta} $ in (ref) derived under the parametric submodel (ref). Therefore, when the sample sizes $ n_0/n_1 \to \infty $ and the DRM holds, the inference on $ G_1 $ {\color{blue} \bf is asymptotically as efficient} as it attains under a corresponding correct parametric model for $ G_1 $. We further prove that the DRM-based estimator of the cumulative distribution function $ G_1 (x) $ and the quantiles of $ G_1 $ both achieve parametric efficiency under the two-sample scenario. Proofs of the main results in this sectino are provided in \hyperref[sec:appendix]{Appendix}.
In this section, we show that the MELE $ \hat \bm{\theta} $ in (ref) under the DRM is asymptotic normal with the same low asymptotic variance as the parametric MLE $ \tilde \bm{\theta} $ in (ref) under the corresponding parametric submodel in (ref). Therefore, when the sample sizes $ n_0/n_1 \to \infty $ and the semiparametric DRM holds, the inference on $G_1$ is as efficient as we postulate a correct parametric model for $G_1$.
Our theories are established under the mild conditions as follow. We use $ {\mathbbm E}_0 [X] $ and $ {\mathbbm E}_1 [X] $ for expectations calculated when $X$ has distributions $G_0$ and $G_1$ respectively. We denote the true parameter value by $\bm{\theta}^*$.
Condition (ref) implies that the matrix $ {\mathbbm E}_{1} [ \mathbf{q} (X) \mathbf{q}^{\top} (X) ] $ is also positive definite. With respect to both $G_0$ and $G_1$, Condition (ref) states that the moment generating function of $ \mathbf{q} (X) $ exists in a neighbourhood of $ \bm {0} $. Hence, all finite-order moments of $ \| \mathbf{q} (X) \| $ are finite. Similarly, $ \| \mathbf{q} (X) \|^{L} \exp \{ \bm{\theta}^{* \top} \mathbf{q} (X) \} $ has finite expectation for all positive $L$. Furthermore, when $ n $ is large enough and $ \bm{\theta} $ is in a small neighbourhood of the truth $ \bm{\theta}^{*} $, the derivatives of the log-EL $ \ell_{n} (\bm{\theta}) $ are all bounded by some polynomials of $ \| \mathbf{q} (X) \| $, and therefore they are all integrable.
The main goal of this section is the efficiency of MELE $ \hat \bm{\theta} $. The following intermediate results are helpful to comprehend the main result.
Since the profile log-EL $ \ell_n (\bm{\theta}) $ has items related to observations from both $G_0$ and $G_1$, the first expectation ${\mathbbm E}$ in the above lemma is with respect to these distributions item by item. The main result is as follows.
We can now answer the question: is the asymptotic variance as low as it can be under the parametric submodel in (ref)? Recall that $\mathbf{q}^{\top} (x) = (1, \mathbf{q}_{-}^{\top} (x))$, and let $ \mathbf{I}_d $ be the identity matrix of dimension $ d $. After some matrix algebra harville1997matrix, we have
With the help from the identity in (ref), we can see that the asymptotic variance in Theorem (ref) is exactly equal to the asymptotic variance of the parametric MLE of $ \bm{\theta} $ in (ref). In other words, when $ n_0/n_1 \to \infty $, the DRM-based estimator of $ \bm{\theta} $ achieves the same asymptotic efficiency as the parametric estimator under the submodel (ref), namely the highest possible efficiency. Estimating the parameter $ \bm{\theta} $ and hence also the density ratio is an interesting problem by itself. We, however, focus more on the efficiency of estimating the cumulative distribution function and quantiles of $ G_1 $. The following sections are devoted to these tasks. As a note for notation, hereafter we may drop the dummy variable inside the expectation and variance operators for a better presentation: \[ {\mathbbm E}_1 [\mathbf{q}_{-}] = {\mathbbm E}_1 [\mathbf{q}_{-} (X)], \hspace {5mm} {\mathrm {Var}}_1 [\mathbf{q}_{-}] = {\mathrm {Var}}_1 [\mathbf{q}_{-} (X)]. \]
In this section, we investigate the asymptotic efficiency of the DRM-based estimator of $ G_1 (x) $ under the two-sample scenario. We first define this estimator. With the DRM-based MELE $ \hat \bm{\theta} $ defined in (ref), we have the fitted values of $ p_{k j} = P (X = x_{k j}; G_0) $ that characterize $ G_0 $: \[ \hat p_{k j} = [ n_0 + n_1 \exp \{ \hat \bm{\theta}^{\top} \mathbf{q} (x_{k j}) \} ]^{-1}. \] We then naturally obtain an estimator of $ G_0 (x) $ as $ \hat G_0 (x) = \sum_{k, j} \hat p_{k j} \mbox{$\mathbbm{1}$} (x_{k j} \leq x) $, and an estimator of $ G_1 (x) $ under the DRM:
where $ \mbox{$\mathbbm{1}$} (\cdot) $ is the indicator function.
The following theorem states that $ \hat G_1 (x) $ is also asymptotically normal as $ n_0/n_1 \to \infty $.
We observe that the asymptotic variance in the above theorem is the same as that in (ref), which further equals the asymptotic variance of the parametric MLE $ \tilde G_1 (x) $ in (ref). Therefore, same as the inference on $ \bm{\theta} $, the DRM-based estimator of the distribution function $ G_1 (x) $ is also as efficient as the parametric estimator under the submodel model in (ref) when $ n_0/n_1 \to \infty $.
In this section, we derive the asymptotic variance of DRM-based quantile estimator for $ G_1 $. The $ p $th quantile of $ G_1 $ is denoted by $ \xi_p $. Estimator of $ \xi_p $ under the DRM is constructed based on the distribution estimator $ \hat G_1 (x) $ defined in (ref):
Because of the relationship between the quantile and distribution function, we often study the asymptotic properties of quantile estimator through distribution estimator. With the help from Theorem (ref) for the asymptotic normality of the DRM-based distribution estimator, we are able to derive the limiting distribution of the DRM-based quantile estimator $ \hat \xi_p $, as shown in the following theorem.
Evidently, the asymptotic variance in Theorem (ref) equals the variance in (ref), and equivalently is as low as the asymptotic variance in (ref) of the parametric MLE $ \tilde \xi_p $. Hence, under the two-sample scenario where $ n_1/n_0 \to 0 $, we have proved that the DRM-based quantile estimator for $ G_1 $ attains parametric efficiency as if the parametric model in (ref) is assumed.
Beyond our core contribution, we also present a useful result that reveals a linear relationship between quantile estimator and distribution estimator. This type of result is famously known as the Bahadur representation.
The conclusion in Theorem (ref) is stronger than that in Theorem (ref).
For illustration, in this section we demonstrate how the asymptotic variance of the DRM-based quantile estimator $ \hat \xi_p $ evolves as a function of the sample size ratio $ k = n_0/n_1 $ when two population distributions are actually identical. That is, we study the efficiency when the true model parameter $ \bm{\theta}^* = \mathbf {0} $ in the DRM (ref). However, we do not assume the knowledge of $ G_0 = G_1 $ when fitting the DRM to the data. Although our main focus is on the situation when $ k \to \infty $, we can learn a lot from the case when $ k $ is finite but large. Applying the results from zhang2000quantile and chen2013quantile, the following corollary explicitly quantifies the efficiency gap between the DRM-based quantile estimator $ \hat \xi_p $ and the parametric MLE $ \tilde \xi_p $ of the quantile.
The asymptotic variance in (ref) depends on $ k = n_0/n_1 $, providing an insight on how the efficiency of the DRM quantile estimator evolves as $ n_0/n_1 \to \infty $. We now take a closer look at the first part in the right hand side of (ref). Let $ F_{n, 1} (x) = n_1^{-1} \sum_{j=1}^{n_1} \mbox{$\mathbbm{1}$} (X_{1 j} \leq x) $ denote the empirical distribution of the sample from $ G_1 $. The empirical quantile with level $ p $ for $ G_1 $ is $ \bar \xi_{p} = \inf \{ x: F_{n, 1} (x) \geq p \}. $ It is well known that if $ g_1 (\xi_p) > 0 $, $ \bar \xi_{p} $ is asymptotically normal serfling2000approximation:
Therefore, the asymptotic variance of the DRM-based quantile estimator given in (ref) is a weighted average of the asymptotic variance of the empirical quantile and the asymptotic variance of the MLE. Interestingly, as $ k = n_0/n_1 $ increases, the efficiency of the DRM-based quantile estimator $ \hat \xi_{p} $ approaches the efficiency of the MLE $ \tilde \xi_p $. Although this intriguing observation is developed under a special situation where $ G_0 = G_1 $ and $ k $ is assumed to be fixed and positive, it enlightens us on the efficiency of the DRM-based quantile estimator in addition to our theoretical result in Theorem (ref) that is established under general conditions and evolving $ k $.
In this section, we report some simulation results on the efficiency of the DRM-based estimators of quantiles. We are particularly interested in their efficiency in estimating lower or higher quantiles, compared to the empirical quantiles and the parametric MLEs. Because the probability density function often has small value at $\xi_p$ when $p$ is close to zero or one, the corresponding empirical quantile has very large variance according to (ref). In contrast, the parametric quantile estimators are largely free from this deficiency but suffer from serious bias when the model is misspecified. Based on our theoretical results, DRM-based quantile estimator achieves the efficiency of the parametric estimator with low model misspecification risk in the presence of an extra sample with a much larger size. In the following sections, we use 1000 repetitions to obtain the simulated biases and variances of the three quantile estimators: the DRM-based estimator, the parametric MLE, and the empirical quantile.
We first examine the performance of the DRM-based quantile estimator when data are from normal distributions. The family of normal distributions has many nice statistical properties. For example, the MLEs of the model parameters and quantiles under the normal family have closed forms. Such properties make it easier to compare the efficiency between the DRM-based estimator and the parametric MLE. Also, the normal distributions collectively satisfy the DRM with basis function $ \mathbf{q} (x) = (1, x, x^2)^{\top} $. We assume the knowledge of this basis function but not the parametric form in the DRM approach in simulations.
We generate the first sample from $ N (\mu_{0}, \sigma_{0}^{2}) $ and the second sample from $ N (\mu_{1}, \sigma_{1}^{2}) $. We observe the performance of the DRM-based quantile estimators for various choices of the means $ \mu_0, \mu_1 $, standard deviations $ \sigma_0, \sigma_1 $, and sample sizes $ n_0, n_1 $. Under normal model, the parametric MLE quantile estimator of $\xi_p$ is given by
where $\Phi^{-1} (p)$ is the $p$th quantile of the standard normal distribution, and $\tilde \mu_1$ and $\tilde \sigma_1$ are the sample mean and standard deviation based on $\{x_{1 j}: j=1, \ldots, n_1\}$. It can be shown that the asymptotic variance of the MLE $ \tilde \xi_{p} $ is given by
We first generate both samples from standard normal distribution, namely, $ \mu_0 = \mu_1 = 0 $ and $ \sigma_0 = \sigma_1 = 1 $. Table (ref) contains the simulated biases and variances of the three quantile estimators after being properly scaled based on sample size. We consider 4 quantile levels $ p \in \{ 0.01, 0.05, 0.10, 0.50 \} $ and 4 sample size combinations $ n_1 \in \{ 100, 1000 \} $ and $ n_0 \in \{ 10 n_1, 100 n_1 \} $. Due to symmetry of the normal distribution, we do not include the higher level quantiles. The biases and variances in the table are inflated by a factor $ \sqrt {n_1} $ and $ n_1 $ respectively. We observe that when $ n_0/n_1 = 100 $, the asymptotic variances of the DRM-based quantile estimators approximately equal the weighted averages of the variances of the parametric MLEs (with weight $ k/(k+1) $) and empirical quantiles (with weight $ 1/(k+1) $). This is consistent with our theoretical finding in Corollary (ref). Further, when $n_0/n_1$ increases from $10$ to $100$, the variances of the DRM-based quantile estimators $ \hat \xi_{p} $ rapidly approaches the variances of the parametric MLEs $ \tilde \xi_{p} $. The improvement in efficiency due to increased $n_0/n_1$ is especially significant for quantiles at levels $p$ close to zero.
We next experiment with data from normal distributions with $ \mu_0 = 1, \mu_1 = 2 $ and $ \sigma_0^2 = 1.5, \sigma_1^2 = 2 $. The performance of the three quantile estimators is summarized in Table (ref). We note that our previous comments on the efficiency of the DRM-based estimators when $ n_0/n_1 $ increases are also applicable here, and therefore we do not offer more interpretations. We anticipate other combinations of normal distribution, if not far different, will still lead to similar results.
We then examine the performance of the three quantile estimators based on data from exponential distributions. The exponential distributions collectively fit into the DRM with the basis function $ \mathbf{q} (x) = (1, x)^{\top} $. We assume the knowledge of this $ \mathbf{q} (x) $ but not the parametric form when fitting the DRM to the data in simulations.
The quantile of the exponential distribution with mean $ \mu $ has a simple form: $\xi_{p} = -\mu \log(1-p)$. In the two-sample setting of this article, the parametric MLE of the second population mean is the corresponding sample mean: $\tilde \mu= n_1^{-1} \sum_{j = 1}^{n_1} x_{1 j}$. This leads to the parametric MLE quantile estimator:
with variance
We first generate both samples from exponential distributions with $ \mu= 1 $. We include the three quantile estimators at levels $p \in \{ 0.5, 0.9, 0.95, 0.99 \}$. The exponential distribution has low density values at higher level quantiles. Therefore, the efficiency gain at higher level quantiles are more meaningful for our investigation. We use the same sample size combinations as in the previous section, and the results are given in Table (ref). We again observe the phenomenon illustrated in Corollary (ref): the asymptotic variances of the DRM-based quantile estimators are close to the weighted averages of the variances of the parametric MLEs (with weight $ k/(k+1) $) and empirical quantiles (with weight $ 1/(k+1) $). Same as in the normal data example, when $n_0/n_1$ increases from $10$ to $100$, the efficiency of the DRM-based quantile estimator quickly approaches the efficiency of the parametric MLE. The improvement is particularly obvious for quantiles at levels $p$ close to one.
We next simulate data from exponential distributions with $ \mu_0 = 1/0.3$ and $\mu_1 = 2 $. The performance of the quantile estimators is summarized in Table (ref). The same conclusions regarding the DRM efficiency with growing $ n_0/n_1 $ can be drawn in this situation. We expect similar findings for other combinations of exponential distributions.
In this section, we study the efficiency of the DRM-based quantile estimator and its parametric and nonparametric competitors with a real-world data. We use the collegiate sports budgets dataset from the TidyTuesday data project tidytuesday, which is accessible from the Github repository (\url {https://github.com/rfordatascience/tidytuesday/tree/master/data/2022/2022-03-29}). The dataset contains yearly samples concerning some demographics of collegiate sports in the U.S. from 2015 to 2019. The variable we consider is the total revenue (in USD) per sports team for both men and women, for which we have approximately 17,000 observations each year, and we log-transform the values to make the scale more suitable for numerical computation. As an exploratory data analysis, we plot in Figure (ref) the histograms of the log-transformed revenue data for the years 2015--2019. The population distributions of the revenues in these years look similar. This is also reflected in their kernel density estimators silverman1986density as depicted in the solid curves in Figure (ref). On one hand, the density estimates are close to bell-shaped, suggesting quantile estimation based on normal model is bearable. On the other hand, the estimated densities sufficiently deviate normal density as they have two or more modes, which can be better illustrated by comparing the estimated densities with the fitted normal densities depicted in the dashed curves. However, the estimated densities appear to have some common structures. Therefore, a DRM-based approach is more convincing and may work well in this situation.
We conduct a real data-based simulation as follows. We regard the yearly samples from 2015--2019 as five populations, and sample with replacement from these populations to form multiple samples repeatedly. As remarked previously, our two-sample results are applicable to the multi-sample situation when $n_0/n_j \to \infty$ for all $j \neq 0$. We hence make the sample size from 2015 substantially larger to mimic the situation of $n_0/n_j$ being very large. This is meaningful as we can see how a large historical dataset could help predict the future with small datasets under the DRM. Specifically, we sample from 2016--2019 with equal sizes $ n \in \{ 200, 500 \} $, and sample from 2015 with size $ n_0 \in \{ 200, 1000, 5000 \} $. To apply the DRM-based approach, the user needs to specify a basis function $ \mathbf{q} (x) $. In this simulation, we use i) data-adaptive basis function $ \mathbf{q} (x) $ zhang2022density learned using the full data with $d=2$; ii) $ \mathbf{q} (x) = (1, x)^{\top} $ whose DRM contains the normal model with equal variances; iii) $ \mathbf{q} (x) = (1, x, x^2)^{\top} $ whose DRM contains the normal model without equal variance assumption. The simulation also correspondingly includes iv) the parametric MLE of quantile derived under the normal model with common variance assumption; v) the parametric MLE under the normal model with no assumption on equal variances; and vi) the empirical quantiles. By regarding the empirical quantiles based on the full data (treated as the populations) as the truth, we compute the simulated absolute biases, variances, and mean squared errors (MSEs) of these quantile estimators at some selected levels $ p $ for each population from 2016 to 2019. To save space, we report only the average values of the aforementioned performance measures across 2016--2019 in Tables (ref)--(ref).
We first observe that when $ n $ is fixed at 200 and 500 while $ n_0 $ increases from 200 to 5000, the variances of the DRM-based quantile estimators with various $ \mathbf{q} (x) $ decrease in nearly all the cases. This observation supports our theoretical results. Second, the DRM-based quantile estimators are more efficient than the nonparametric empirical quantiles, which suggests that data pooling via the DRM works well for this data. Third, when $ n_0 = 5000 $, the quantile estimators derived under the DRM with $ \mathbf{q} (x) = (1, x)^{\top} $ overall have comparable MSEs with the parametric estimators under the normal model with common variance assumption, except for the $90\%$ quantile level. The same also applies to the comparison between the estimators under DRM with $ \mathbf{q} (x) = (1, x, x^2)^{\top} $ and the parametric estimators under the normal model with no equal variance assumption. This phenomenon can partly be explained by seeing in Figure (ref) that the true distributions deviate noticeably from the fitted normal distributions at the the right tails, which suggests normal model is unsatisfactory there. Finally, the DRM quantile estimators using the data-adaptive basis function $ \mathbf{q} (x) $ generally beat the DRM estimators using the two prespecified $ \mathbf{q} (x) $ in all three performance measures, except for very few cases where the results are still comparable. This is expected because when the basis function $ \mathbf{q} (x) $ is adaptively learned from the data, the resulting DRM should fit the data better than a predetermined DRM. In fact, when $ n_0 = 5000 $, the adaptive DRM produces overall the most accurate quantile estimators, indicating that an appropriate DRM is suitable to the data.
\endinput
The DRM for multi-sample data is generating growing research interest and has wide applications in statistics and econometrics. It has proven particularly useful in situations where the populations may possess shared underlying structures. The DRM provides a good trade-off between low model misspecification risk and satisfactory statistical efficiency. Its effectiveness primarily stems from its ability to enable users to draw inferences on each population using pooled data, which leads to efficiency gain compared to using individual samples that overlooks the shared latent structures. The literature has engaged in discussions regarding the efficiency of the DRM approach in comparison to the nonparametric approach. However, none of these discussions have systematically explored the limit of the efficiency gain through the DRM, which is an important research problem. This article addresses this problem by considering a scenario where one of the samples significantly outweighs the others in size. We establish through theoretical analysis that within this context, the DRM-based estimators of model parameters, distribution functions, and quantiles for the smaller sample populations attain the efficiency as if a parametric model is assumed. In essence, for these estimands and in this scenario, we identify their highest achievable efficiency under a specific parametric model, investigate their asymptotic efficiency under the DRM, and demonstrate the equivalence between the two aforementioned efficiencies. Our simulation experiments and analyses of real-world data, with a particular focus on quantile estimation, support our theoretical discoveries. The significance of this article's contribution extends to practical scenarios where researchers aim to make inferences about a population with limited data, but can rely on a substantial dataset from a related population for support.
While the scenario examined in this article addresses many real-world applications, there exist situations where this sampling scheme is not applicable. In such cases, we may encounter a growing number of samples, each of relatively similar sizes. For example, when economists investigate the evolution of income distribution over time, they may collect income samples year after year, and these yearly samples often have comparable sizes. Consequently, it becomes intriguing to investigate the asymptotic efficiency of the DRM approach when the sample sizes $ n_i/n_j \to 1 $ and the number of populations $ m \to \infty $ as $ n_i \to \infty $. Additionally, it is worthwhile to study the asymptotic efficiency of other population parameters than the ones we investigated. We leave these interesting problems for future work, anticipating that the technical methods and theoretical results in this article may be valuable for such inquiries.
This research was partially supported by the Natural Sciences and Engineering Research Council of Canada (Grants RGPIN-2018-06484, RGPIN-2019-04204, and RGPIN-2020-05897), the Canadian Statistical Sciences Institute (Grant 592307), and the Department of Statistical Sciences in the University of Toronto. We also thank the Digital Research Alliance of Canada for computing support. This work was partially completed when Archer Gong Zhang was a PhD student at the University of British Columbia and a Postdoctoral Fellow at the University of Toronto.