EconBase
← Back to paper

Optimal Estimation under a Semiparametric Density Ratio Model

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

90,614 characters · 18 sections · 38 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Optimal Estimation under a Semiparametric Density Ratio Model

abstractIn many statistical and econometric applications, we gather individual samples from various interconnected populations that undeniably exhibit common latent structures. Utilizing a model that incorporates these latent structures for such data enhances the efficiency of inferences. Recently, many researchers have been adopting the semiparametric density ratio model (DRM) to address the presence of latent structures. The DRM enables estimation of each population distribution using pooled data, resulting in statistically more efficient estimations in contrast to nonparametric methods that analyze each sample in isolation. In this article, we investigate the limit of the efficiency improvement attainable through the DRM. We focus on situations where one population's sample size significantly exceeds those of the other populations. In such scenarios, we demonstrate that the DRM-based inferences for populations with smaller sample sizes achieve the highest attainable asymptotic efficiency as if a parametric model is assumed. The estimands we consider include the model parameters, distribution functions, and quantiles. We use simulation experiments to support the theoretical findings with a specific focus on quantile estimation. Additionally, we provide an analysis of real revenue data from U.S. collegiate sports to illustrate the efficacy of our contribution.

{\bf \em Keywords}: Biased sampling; Empirical likelihood; Exponential family; Exponential tilting; Statistical efficiency.

Introduction

In numerous applications, scientists may independently collect a single sample from each of the connected and similar populations. The dataset hence contains independent multiple samples. For example, to gauge the economic state of a country, econometricians may study the evolution of income distribution over time roine2015long. They may hence collect cross-sectional or panel survey data from year to year wooldridge2010econometric, and the underlying populations for these multiple samples are naturally connected. There are various approaches to data analysis of this nature. One may postulate a parametric model such as normal or gamma on these populations davison2003statistical. However, some statistical inference procedures such as quantile estimation may have poor performance when the model assumptions are merely mildly violated. For instance, statistical analyses of a dataset under two similar models, log-normal and gamma models, can yield markedly different results wiens1999log. To mitigate the risks of model misspecification, one may take a nonparametric approach wasserman2006all by not imposing any parametric assumption. However, the nonparametric approaches ignore the latent structures shared by the multiple populations from which the samples are collected, and hence fail to utilize the potential gain in statistical efficiency. As a balanced trade-off between the model misspecification risks and the statistical efficiency, we advocate semiparametric approaches hardle2004nonparametric, tsiatis2006semiparametric.

Recently, there have been many research studies on the semiparametric density ratio model (DRM) anderson1979multivariate. Let $ g_0 (x), \ldots, g_m (x) $ be the density functions for the multiple populations. The DRM postulates that

align[align omitted — 122 chars of source]

where $ \mathbf{q} (x) $ is a prespecified vector-valued function and $ \bm{\theta}_k $ is an unknown vector-valued parameter. The DRM effectively utilizes the latent structures shared by the multiple populations. Not surprisingly, methods developed under the DRM are often statistically more efficient than the nonparametric methods that ignore the latent structures and use the samples separately for individual populations. In particular, the DRM-based empirical likelihood (EL) approaches owen2001empirical are found to have nice asymptotic properties and superior performances qin1993empirical, chen2013quantile, cai2017hypothesis. Under the framework of the DRM and EL, the multiple samples are used together to draw inferences on each population.

These results lead to an interesting question: what is the limit of the efficiency gain through the DRM-based approach? A gold-standard target to consider is the parametric efficiency. As a semiparametric model, a DRM naturally contains many parametric models, with each forming a so-called parametric submodel. For every such parametric submodel, there is a classical Cram\'{e}r--Rao lower bound. As a result, the asymptotic variance for any consistent and asymptotically normal semiparametric estimator is no smaller than the supremum of the Cram\'{e}r--Rao lower bounds for all parametric submodels. Such a supremum is famously known as the semiparametric efficiency bound newey1990semiparametric. When a parametric submodel contains the true distributions that generate the data and is regular, the parametric maximum likelihood estimator (MLE) has the smallest asymptotic variance that attains the Cram\'{e}r--Rao lower bound casella2002statistical. Therefore, no methods can achieve higher efficiency under the DRM than it can under its parametric submodel that the true distributions reside. This gives an efficiency upper bound for any semiparametric approach including the DRM-based estimators. Motivated by this observation, our research question becomes: is it likely that the DRM-based approaches achieve parametric efficiency under a parametric submodel? If yes, when would that happen?

In this article, we focus on the scenario where the sample size from one population is significantly larger than the sample sizes from the other populations. We are interested in the efficiency of the DRM-based estimators for the populations with smaller samples. In other words, we want to quantify how much help the large sample population can offer to improve the efficiency of the estimators for the small sample populations. The motivation for us to consider this scenario is as follows. Suppose the sample size from one population is infinite. The corresponding population distribution can therefore be regarded as known. Consequently, treating the infinite-sample population density as $ g_0 (x) $, the semiparametric DRM in (ref) would reduce to a parametric model for the other populations. We therefore expect that the DRM-based estimators for the small sample populations will achieve parametric efficiency under the corresponding parametric models. This article establishes and proves this claim for the model parameters, distribution functions, and quantiles. Although we prove these results mostly in a two-sample scenario for clarity, our discoveries are general for the multi-sample scenario. More specifically, when one population has an extra large sample, the DRM-based estimators for the other populations are as efficient as the best possible under the true parametric submodel. For convenience, we refer to this scenario as the two-sample scenario throughout the article.

In the existing literature, there have been some discussions on the asymptotic efficiency of the DRM-based approaches, such as zhang2000quantile and chen2013quantile. They focus on situation where the sample sizes all go to infinity at the same rate and find that the DRM-based quantile estimators are asymptotically more efficient than the nonparametric empirical quantiles. However, none of these studies systematically investigate the boundary of the efficiency gain under the DRM. This article tackles this important research problem by examining the aforementioned two-sample scenario. Furthermore, their theoretical results rely on the assumption that the ratio of any two sample sizes remains fixed, positive, and finite as the sample sizes aproach infinity. In contrast, we study the asymptotic efficiency of the DRM-based approaches in a different and more general context that allows the sample sizes to grow at different rates and allows the ratios of sample sizes to evolve.

The rest of the article is organized as follows. In Section (ref), we elaborate more on the DRM and state the research problem. For comparison, in Section (ref) we present the gold-standard parametric efficiency under a specific parametric submodel of the DRM. In Section (ref), we give an overview of the EL-based inferences under the DRM, and study the asymptotic efficiency of the EL-DRM estimators of the model parameters, distribution functions, and quantiles. For illustration, we show in Section (ref) how the asymptotic variance of the DRM-based quantile estimator evolves as a function of the ratio of the sample sizes when the true population distributions $ G_k $ are all identical but not assumed to be identical in the estimation. We conduct some simulation studies in Section (ref) by generating data from normal and exponential distributions, whose results support our theoretical findings. In Section (ref), we report results from a real-data analysis on the revenues of U.S. collegiate sports. Finally, Section (ref) concludes our contributions in the article with a discussion. Proofs of the theoretical results are provided in \hyperref[sec:appendix]{Appendix}.

The research problem

Suppose we have $ m+1 $ independent sets of independent and identically distributed (i.i.d.) samples respectively from population distributions $ G_0, \ldots, G_m $:

align[align omitted — 129 chars of source]

Let $ g_0 (\cdot), \ldots, g_m (\cdot) $ be the corresponding density functions with respect to some common $ \sigma $-finite measure. They satisfy the DRM if

align*[align* omitted — 110 chars of source]

with $ \mathbf{q} (x) $ a prespecified vector-valued function of dimension $ d $ and $ \bm{\theta}^{\top} \coloneqq (\bm{\theta}_1^{\top}, \ldots. \bm{\theta}_m^{\top}) $ an unknown vector-valued parameter. The first element of $ \mathbf{q} (x) $ is set to be 1 so that the corresponding coefficient in $ \bm{\theta}_k $ is a normalization constant. Therefore, we may write $ \mathbf{q} (x) $ as

align[align omitted — 79 chars of source]

for some vector-valued function $ \mathbf{q}_{-}^{\top} (x) $ with dimension $ d-1 $. Correspondingly, we may also decompose the parameter $ \bm{\theta}_k $ into $ (\alpha_k, \bm{\beta}_k) $. We also require the elements of $ \mathbf{q} (x) $ to be linearly independent; otherwise, some elements of $ \mathbf{q} (x) $ are redundant. By convention, we call $ G_0 $ the base distribution and $ \mathbf{q} (x) $ the basis function. The DRM and its inference, which we introduce later, are invariant to the choice of the base distribution. In this article, we select the population distribution from which we have a much larger sample as the base distribution.

The DRM covers many well-known parametric distribution families, with different choices of the base distribution $ G_0 $ and the basis function $ \mathbf{q} (x) $. For example, the Poisson distribution family satisfies the DRM with $ G_0 $ being Poisson and $ \mathbf{q} (x) = (1, x)^{\top} $; the normal distribution family satisfies the DRM with $ G_0 $ being normal and $ \mathbf{q} (x) = (1, x, x^{2})^{\top} $; and the gamma distribution family also satisfies the DRM with $ G_0 $ being gamma and $ \mathbf{q} (x) = (1, x, \log x)^{\top} $. In fact, any exponential family model satisfies the DRM. On the other hand, when using DRM, we do not assume that each population distribution belongs to any parametric distribution family. Instead of directly modelling the distribution of each population, the DRM takes a semiparametric approach to modelling the relationship between the multiple populations. Therefore, we may see that the DRM is a very flexible model and has a low risk of model misspecification. A more appealing feature of the DRM is that, with an appropriately chosen basis function $ \mathbf{q} (x) $, the DRM allows us to utilize the combined data to estimate each distribution $ G_k $. This would lead to efficiency gain compared with estimating $ G_k $ individually using the sample from itself. The DRM has been widely applied in many fields; see the books sugiyama2012density and qin2017biased. It is also closely connected with other semiparametric models such as models for biased sampling vardi1982nonparametric, vardi1985empirical, the exponential tilting model rathouz2009generalized, and the proportional likelihood ratio model luo2012proportional. Many studies on a variety of inference problems in the literature have confirmed this efficiency gain through either theoretical or empirical analyses, including zhang2000quantile, chen2013quantile, cai2017hypothesis.

The research problem we consider in this article is: when the sample from $ G_0 $ has a much larger size than the samples from the other $ G_k $, would the DRM-based estimators of some functionals of $ G_k $ achieve parametric efficiency under a parametric submodel? Without loss of generality and for clarity, hereafter we present our results mostly in the two-sample scenario, namely $ m = 1 $. Our results are general for the multiple samples as long as we have a much larger sample from one population than the others. We denote by $ n_k $ the size of the sample from $ G_k $, and let $ n = \sum_{k=0}^{m} n_k $ be the total sample size. Formally, when $ G_0, G_1 $ satisfy the DRM, we study the asymptotic efficiency of the estimators of some functionals of $ G_1 $ when

align[align omitted — 110 chars of source]

The functionals of $ G_1 $ we consider include the model parameters $ \bm{\theta} $, the cumulative distribution function $ G_1 (x) $ for every $ x $, and the quantiles of $ G_1 $.

The research problem has many applications in statistics and econometrics. For example, an important aspect of econometrics is to analyze cross-sectional and panel data with some structures ng2006testing, chen2012semiparametric, lin2012estimation, boneva2015semiparametric, su2016identifying, gao2020heterogeneous, su2023identifying, which is useful for areas of economic research such as understanding the economic state of a population at specific time points and evaluating the effects of a policy across different groups, regions, or time periods. The DRM has been found successful for the analysis of such data in the long-term monitoring of lumber strength properties zidek2018statistical, chen2022permutation. This is because the DRM-based methods are able to borrow strength across the multiple cross-sectional samples to exploit the shared latent structures between different populations over years and regions. Therefore, it is of scientific significance to quantify the resulting efficiency gain when analyzing such type of data. Our contributions in this article achieve this aim by focusing on the situation where one wishes to make inferences on a population with a small data set, aided by a large data set from another connected population. In the real-data analysis section of this article, we study the efficiency of the DRM-based quantile estimators under such scenario based on a revenue data from U.S. collegiate sports.

Estimation efficiency under a parametric submodel

To study the efficiency of the DRM-based inferences, we first reveal the efficiency limit under its parametric submodels. A parametric submodel is a parametric family of distributions we select for the data that satisfy the semiparametric model (in our case, the DRM) assumptions. In this article, the parametric submodel we consider for $ G_0, G_1 $ is an exponential family model:

align[align omitted — 214 chars of source]

where $ \bm{\eta}_0, \bm{\eta}_1 $ are the unknown natural parameters to be estimated, $ \mathbf{q}_{-} (x) $ is the sufficient statistic for a single observation $ x $ that is the same as the one in (ref) in the DRM, $ B (\cdot) $ is a known nonnegative function, and $ A (\bm{\eta}_k) $ is uniquely determined by $ \bm{\eta}_k $: \[ A (\bm{\eta}_k) = -\log \int B (x) \exp \{ \bm{\eta}_k^{\top} \mathbf{q}_{-} (x) \} \mathrm{d} x, \] such that $ g_k $ is indeed a density function for $ k = 0, 1 $.

{\bf Remark}: the results presented in the article are still valid if we consider another set of parametric submodel: $ g_0 (x) $ being completely known and $ g_1 (x) $ from an exponential family model. This submodel consists of joint distributions of two independent set of i.i.d. samples in the form of \( \prod_{i=1}^{n_0} g_0 (x_{0i}) \prod_{j=1}^{n_1} g_0 (x_{1j}) \exp \{ \alpha + \bm{\beta}^{\top} \mathbf{q}_{-} (x_{1j} ) \}. \)

With the decomposition of $ \mathbf{q} (x) $ and $ \bm{\theta} $ in (ref), we get an equivalent form of the DRM in (ref):

align[align omitted — 167 chars of source]

where $ \alpha $ is a normalization constant determined by \[ \alpha = -\log \int g_0(x) \exp \{ \bm{\beta}^\top \mathbf{q}_{-} (x) \} \mathrm{d} x. \] Therefore, the exponential family model in (ref) satisfies the DRM in (ref) with \[ \bm{\theta} =

pmatrix[pmatrix omitted — 35 chars of source]

=

pmatrix[pmatrix omitted — 77 chars of source]

. \] In other words, the exponential family model in (ref) is indeed a parametric submodel of the DRM. In this section, we obtain the asymptotic results for $ G_1 $ under this parametric submodel.

Estimation of \texorpdfstring{$\boldsymbol{\theta}$} under the parametric submodel

We first study the efficiency of the parametric MLE of $ \bm{\theta} $ under the parametric submodel (ref). This will be our gold-standard target with which we will compare the efficiency of the DRM estimator of $ \bm{\theta} $ to be presented later. We obtain a parametric estimator of $ \bm{\theta} $ using the maximum likelihood method. The parametric log-likelihood function given the two independent i.i.d. samples is

align*[align* omitted — 345 chars of source]

From this, we get the score function: for $ k = 0, 1 $, \[ \frac{\partial \ell_{\mathrm{para}} (\bm{\eta}_0, \bm{\eta}_1)}{\partial \bm{\eta}_k} = n_k \frac{\partial A (\bm{\eta}_k)} {\partial \bm{\eta}_k} + \sum_{j=1}^{n_k} \mathbf{q}_{-} (x_{k j}). \] We let $ \tilde \bm{\eta}_k $ denote the parametric MLE of $ \bm{\eta}_k $, and they satisfy

align*[align* omitted — 130 chars of source]

We remark that for convenience, with a generic function $ f (\boldsymbol{y}) $, we have used the notation \[ \frac{\partial f (\boldsymbol{y}^{*}) }{\partial \boldsymbol{y}} = \left. \frac{\partial f (\boldsymbol{y}) }{\partial \boldsymbol{y}} \right \rvert_{\boldsymbol{y} = \boldsymbol{y}^{*}} \] in the preceding expression, and we retain the use of such notation hereafter.

By the invariance property of maximum likelihood estimation (see Theorem (ref) below), which states that any function of MLE is still an MLE of the function of the parameter, we must have that the MLE of $ \bm{\theta} $, denoted as $ \tilde \bm{\theta} $, is given by

align[align omitted — 174 chars of source]

Before we formally state the invariance property of the MLEs, we first introduce the so-called induced likelihood casella2002statistical, which is also known as the profile likelihood in the literature barndorff1988parametric. Suppose a distribution family is indexed by a parameter $ \theta $, and we want to find the MLE of some function of $ \theta $, denoted by $ \tau (\theta) $. When $ \tau (\cdot) $ is one-to-one, then given the MLE of $ \theta $ being $ \hat \theta $, it can be seen that the MLE of $ \tau (\theta) $ is $ \tau (\hat \theta) $. When $ \tau (\cdot) $ is not one-to-one, there are some technical issues under the current definition of likelihood. To overcome these problems, we need a more general definition of the likelihood function of $ \tau (\theta) $, called the induced likelihood function, and correspondingly a more general definition of the MLE. They are defined as follows.

definition[Extended definitions of likelihood and MLE casella2002statistical] Denote by $ L (\theta) $ the likelihood function of $ \theta $. The induced likelihood function for $ \eta = \tau (\theta) $ is defined as \begin{align*} L^{*} (\eta) = \sup_{\{ \theta: \tau (\theta) = \eta \}} L (\theta). \end{align*} Further, $ \hat \eta $ that maximizes the induced likelihood $ L^{*} (\eta) $ is defined as the MLE of $ \eta = \tau (\theta) $.

With this extended definition of the MLE, we are now ready to formally state the invariance property of MLEs in the following theorem.

theorem[Invariance property of MLE casella2002statistical] If $ \tilde \theta $ is the MLE of $ \theta $, then for any function $ \tau (\theta) $, $ \tau (\tilde \theta) $ is an MLE of $ \tau (\theta) $.

After obtaining the parametric MLE $ \tilde \bm{\theta} $, we now study its asymptotic properties. By the standard results on maximum likelihood estimation theory, under the exponential family model in (ref) and suppose some regularity conditions are satisfied, the MLE $ \tilde \bm{\eta}_k $ is asymptotically normal:

align[align omitted — 161 chars of source]

as $ n_k \to \infty $, where $ \bm{\eta}_k^{*} $ denotes the true value of $ \bm{\eta}_k $. Applying the delta method casella2002statistical, we have that $ (A (\tilde \bm{\eta}_k), \tilde \bm{\eta}_k) $ is also asymptotically normal: \[ \sqrt {n_k}

pmatrix[pmatrix omitted — 99 chars of source]

\overset {d} \to N \left ( \bm {0},

pmatrix[pmatrix omitted — 75 chars of source]

{\mathrm {Var}}_k^{-1} [\mathbf{q}_{-} (X)]

pmatrix[pmatrix omitted — 69 chars of source]

\right ), \] as $ n_k \to \infty $. Note that here we have used a standard result on the exponential family model:

align[align omitted — 135 chars of source]

Taking advantage of the nice mathematical properties of exponential families, we forego other otherwise necessary verifications.

Finally, because the two samples are independent of each other, we have

align[align omitted — 643 chars of source]

where $ n_1/n_0 \to 0 $, as $ n_0, n_1 \to \infty $. We will compare the asymptotic variance in (ref) of the parametric MLE $ \tilde \bm{\theta} $ with the asymptotic variance of the DRM-based estimator of $ \bm{\theta} $ later.

Distribution estimation under the parametric submodel

We next study the inference on the distribution function $ G_1 (x) $ under the two-sample exponential family model in (ref). By Theorem (ref), the MLE of $ G_1 (x) = G_1 (x; \bm{\eta}_1^{*}) $ is given by \[ \tilde G_1 (x) = G_1 (x; \tilde \bm{\eta}_1) = \int_{-\infty}^{x} B (t) \exp \{ \tilde \bm{\eta}_1^{\top} \mathbf{q}_{-} (t) + A (\tilde \bm{\eta}_1) \} \mathrm {d} t, \] where we recall that $ \tilde \bm{\eta}_1 $ is the MLE of $ \bm{\eta}_1 $.

The limiting distribution of such a parametric distribution estimator $ \tilde G_1 (x) $ can be straightforwardly obtained, as follows. Recall that the MLE $ \tilde \bm{\eta}_1 $ is asymptotically normal: \[ \sqrt {n_1} (\tilde \bm{\eta}_1 - \bm{\eta}_1^{*}) \overset {d} \to N (\bm {0}, {\mathrm {Var}}_1^{-1} [\mathbf{q}_{-} (X)]), \] as $ n_1 \to \infty $. Applying the delta method, we have that for every fixed $ x $, the MLE $ \tilde G_1 (x) $ is also asymptotically normal:

align[align omitted — 445 chars of source]

Under regularity conditions, we can further simplify the variance matrix in (ref) as follows. First, we note that

align[align omitted — 802 chars of source]

by the definition of $ \mathbf{Q} (x) $ given in Theorem (ref). Further, recall from (ref) that $ \partial A (\bm{\eta}_1^{*}) / \partial \bm{\eta}_1 = - {\mathbbm E}_1 [\mathbf{q}_{-} (X)]. $ Therefore, the preparation results in (ref) and (ref) together lead to a simpler expression for the variance matrix in (ref):

align[align omitted — 440 chars of source]

We will examine later that the DRM-based estimator of $ G_1 (x) $ attains the same asymptotic efficiency as above.

Quantile estimation under the parametric submodel

Quantiles are very important population parameters in many applications. In this section, we study the asymptotic efficiency of the parametric quantile estimator under the exponential family model in (ref). We first define, for any quantile level $ p \in (0, 1) $, the $ p $th quantile of $ G_1 $ as \[ \xi_p = G_1^{-1} (p) \coloneqq \inf \{ t: G_1 (t) \geq p \}, \] where the inverse function $ G_1^{-1} (\cdot) $ is also known as the quantile function. We assume that the distribution $ G_1 (x) $ has a density function $ g_1 (x) $ that is positive and continuous at $ x = \xi_p $. Then, by Theorem (ref), the MLE of $ \xi_p = G_1^{-1} (p; \bm{\eta}_1^{*}) $ is given by \[ \tilde \xi_p = G_1^{-1} (p; \tilde \bm{\eta}_1), \] where $ \tilde \bm{\eta}_1 $ is the MLE of $ \bm{\eta}_1 $ under the parametric model in (ref).

Because the MLE $ \tilde \bm{\eta}_1 $ is asymptotically normal (see (ref)), the quantile MLE $ \tilde \xi_p $ is also asymptotically normal by the delta method:

align[align omitted — 461 chars of source]

In addition, the variance matrix in (ref) has a more specific expression, as follows. Firstly,

align*[align* omitted — 181 chars of source]

Then, from previous derivations given in (ref) and (ref), we have

align*[align* omitted — 215 chars of source]

Hence, the variance matrix in (ref) can be written as

align[align omitted — 496 chars of source]

Our objective, to be realized later, is to investigate whether the DRM-based estimator of $ \xi_p $ has the asymptotic variance as low as the variance in (ref) and equivalently in (ref).

Estimation efficiency under the DRM

Under the DRM, the base distribution $ G_0 $ is left unspecified. If we impose a parametric form on $ G_0 $, the DRM would reduce to a parametric model. In this case, we gain model simplicity but bear a higher risk of model misspecification. We thus leave $ G_0 $ unspecified and use the nonparametric EL of owen1988empirical as a platform for statistical inference under the DRM. There has been a rich literature on the coupling of the EL and the DRM; see, for example, qin1993empirical, qin1997goodness, qin1998inferences, fokianos2001semiparametric.

EL-based inference under the DRM

We first review the EL method under the DRM. For convenience, let $ p_{kj} = \mathrm{d} G_0 (x_{k j}) = P (X = x_{k j}; G_0) $, the probability of observing $ x_{k j} $ under $ G_0 $ for all applicable $ k, j $. Applying the likelihood principle, we obtain the EL based on the multiple sample under the DRM:

align[align omitted — 190 chars of source]

with $ \bm{\theta}_0 \coloneqq \bm{0} $ by convention. We observe that $L_{n} (\cdot) = 0$ if $ G_k $ are continuous distribution functions. This seemingly devastating property does little harm to the usefulness of the EL. As we will see, in the EL we search for the distribution estimator within the space of discrete distributions that assign positive probability mass to the observed data. This does not eliminate much generality because every distribution can be precisely approximated by such discrete distributions when the sample size grows. Since $ L_{n} (\cdot) $ in (ref) is also a function of the parameters $ \bm{\theta} $ and base distribution $ G_0 $, we may also write its logarithm as

align[align omitted — 162 chars of source]

The EL-based inferences on the population parameters are usually carried out through a profile likelihood function. We first observe that the DRM assumption in (ref) implies \[ 1 = \int \mathrm{d} G_0 (x), \hspace{5mm} 1 = \int \mathrm{d} G_r = \int \exp \{ \bm{\theta}_r^{\top} \mathbf{q} (x) \} \mathrm{d} G_0 (x), \hspace{3mm} r = 1, \ldots, m. \] Confining the common support of $ G_0, \ldots, G_m $ to the observed data $ \{ x_{k j} \}_{k, j} $ yields the EL-version constraints \[ \sum_{k, j} p_{k j} = 1, \hspace{5mm} \sum_{k, j} p_{k j} \exp \{ \bm{\theta}_r^{\top} \mathbf{q} (x_{k j}) \} = 1, \hspace{3mm} r = 1, \ldots, m. \] With these preparations and following the foundational work by qin1994empirical, we define the profile log-EL function of $ \bm{\theta} $ as the supremum of the log-EL $ \ell_{n} (\bm{\theta}, G_{0}) $ in (ref) over $ G_0 $ subject to the above constraints:

align*[align* omitted — 254 chars of source]

The above optimization problem has a simple solution by the Lagrange multiplier method. For the best relevance, we present the results when $ m = 1 $. We have

align*[align* omitted — 228 chars of source]

for some $ \hat \lambda_1 $ satisfying $ \sum_{k, j} 1/[ n + \hat \lambda_1 ( \exp \{ \bm{\theta}^{\top} \mathbf{q} (x_{k j}) \} - 1)] = 1 $. One may estimate $ \bm{\theta} $ by the maximum empirical likelihood estimator (MELE):

align[align omitted — 100 chars of source]

At $ \bm{\theta} = \hat \bm{\theta} $, some algebra gives $ \hat \lambda_1 = n_1 $, and we naturally get another function by replacing $ \hat \lambda_1 $ with $ n_1 $ in the profile log-EL $ \tilde \ell_{n} (\bm{\theta}) $:

align[align omitted — 217 chars of source]

The profile log-EL $ \tilde \ell_{n} (\bm{\theta}) $ and $ \ell_{n} (\bm{\theta}) $ have the same maximum value and maximizer. Because of this, we study the asymptotic properties of the MELE $ \hat \bm{\theta} $ through the analytically simpler $ \ell_{n} (\bm{\theta}) $. By convention, we call this function a {\em dual function} of the profile log-EL. Following chen2013quantile, we regard this dual function $ \ell_{n} (\bm{\theta}) $ as if it is the profile log-EL.

Our ultimate goal is to study the asymptotic efficiency of the DRM-based inferences of $ G_1 $ that has a smaller sample under the two-sample scenario when $ n_0/n_1 \to \infty $. To achieve this goal, we first show that the MELE $ \hat \bm{\theta} $ under the DRM is asymptotic normal with the same low variance as the MLE $ \tilde \bm{\theta} $ in (ref) derived under the parametric submodel (ref). Therefore, when the sample sizes $ n_0/n_1 \to \infty $ and the DRM holds, the inference on $ G_1 $ {\color{blue} \bf is asymptotically as efficient} as it attains under a corresponding correct parametric model for $ G_1 $. We further prove that the DRM-based estimator of the cumulative distribution function $ G_1 (x) $ and the quantiles of $ G_1 $ both achieve parametric efficiency under the two-sample scenario. Proofs of the main results in this sectino are provided in \hyperref[sec:appendix]{Appendix}.

Estimation of \texorpdfstring{$\boldsymbol{\theta}$} under the DRM

In this section, we show that the MELE $ \hat \bm{\theta} $ in (ref) under the DRM is asymptotic normal with the same low asymptotic variance as the parametric MLE $ \tilde \bm{\theta} $ in (ref) under the corresponding parametric submodel in (ref). Therefore, when the sample sizes $ n_0/n_1 \to \infty $ and the semiparametric DRM holds, the inference on $G_1$ is as efficient as we postulate a correct parametric model for $G_1$.

Our theories are established under the mild conditions as follow. We use $ {\mathbbm E}_0 [X] $ and $ {\mathbbm E}_1 [X] $ for expectations calculated when $X$ has distributions $G_0$ and $G_1$ respectively. We denote the true parameter value by $\bm{\theta}^*$.

enumerate[(i)] • As $n_0, n_1 \to \infty$, $n_0/n_1 \to \infty$. • The matrix $ {\mathbbm E}_{0} [ \mathbf{q} (X) \mathbf{q}^{\top} (X) ] $ is positive definite. • For $ \bm{\theta}$ in a neighbourhood of the true parameter value $ \bm{\theta}^{*} $ or of $ \bm {0} $, we have \[ {\mathbbm E}_{0} \left [ \exp (\bm{\theta}^{\top} \mathbf{q} (X)) \right ] < \infty, \hspace {5mm} {\mathbbm E}_{1} \left [ \exp (\bm{\theta}^{\top} \mathbf{q} (X)) \right ] < \infty. \]

Condition (ref) implies that the matrix $ {\mathbbm E}_{1} [ \mathbf{q} (X) \mathbf{q}^{\top} (X) ] $ is also positive definite. With respect to both $G_0$ and $G_1$, Condition (ref) states that the moment generating function of $ \mathbf{q} (X) $ exists in a neighbourhood of $ \bm {0} $. Hence, all finite-order moments of $ \| \mathbf{q} (X) \| $ are finite. Similarly, $ \| \mathbf{q} (X) \|^{L} \exp \{ \bm{\theta}^{* \top} \mathbf{q} (X) \} $ has finite expectation for all positive $L$. Furthermore, when $ n $ is large enough and $ \bm{\theta} $ is in a small neighbourhood of the truth $ \bm{\theta}^{*} $, the derivatives of the log-EL $ \ell_{n} (\bm{\theta}) $ are all bounded by some polynomials of $ \| \mathbf{q} (X) \| $, and therefore they are all integrable.

The main goal of this section is the efficiency of MELE $ \hat \bm{\theta} $. The following intermediate results are helpful to comprehend the main result.

lemmaUnder Conditions (ref) to (ref), we have \begin{enumerate} • ${\mathbbm E} \big \{ \partial \ell_{n} (\bm{\theta}^*)/{\partial \bm{\theta}} \big \} = \bm {0}; $$ n_1^{-1/2} \{ \partial \ell_{n} (\bm{\theta}^*)/ {\partial \bm{\theta}} \} \overset {d} {\to} N (\bm {0}, {\mathrm {Var}}_1 [ \mathbf{q} (X) ] ) $ as $ n_0, n_1 \to \infty $; • $ - n_1^{-1} \{ \partial^{2} \ell_{n} (\bm{\theta}^*) /\partial \bm{\theta} \partial \bm{\theta}^{\top} \} \overset {p} {\to} {\mathbbm E}_{1} [ \mathbf{q} (X) \mathbf{q}^{\top} (X) ] $ as $ n_0, n_1 \to \infty $. \end{enumerate}

Since the profile log-EL $ \ell_n (\bm{\theta}) $ has items related to observations from both $G_0$ and $G_1$, the first expectation ${\mathbbm E}$ in the above lemma is with respect to these distributions item by item. The main result is as follows.

theoremUnder Conditions (ref) to (ref), as $ n_0, n_1 \to \infty $, the MELE $ \hat \bm{\theta} $ defined in (ref) is asymptotically multivariate normal: \[ \sqrt {n_1} (\hat \bm{\theta} - \bm{\theta}^{*}) \overset {d} {\to} N \left (\bm {0}, \left \{ \mathbbm {E}_{1} \left [ \mathbf{q} (X) \mathbf{q}^{\top} (X) \right ] \right \}^{-1} - \begin{pmatrix} 1 & \bm {0} \\ \bm {0} & \bm {0} \end{pmatrix} \right ). \]

We can now answer the question: is the asymptotic variance as low as it can be under the parametric submodel in (ref)? Recall that $\mathbf{q}^{\top} (x) = (1, \mathbf{q}_{-}^{\top} (x))$, and let $ \mathbf{I}_d $ be the identity matrix of dimension $ d $. After some matrix algebra harville1997matrix, we have

align[align omitted — 817 chars of source]

With the help from the identity in (ref), we can see that the asymptotic variance in Theorem (ref) is exactly equal to the asymptotic variance of the parametric MLE of $ \bm{\theta} $ in (ref). In other words, when $ n_0/n_1 \to \infty $, the DRM-based estimator of $ \bm{\theta} $ achieves the same asymptotic efficiency as the parametric estimator under the submodel (ref), namely the highest possible efficiency. Estimating the parameter $ \bm{\theta} $ and hence also the density ratio is an interesting problem by itself. We, however, focus more on the efficiency of estimating the cumulative distribution function and quantiles of $ G_1 $. The following sections are devoted to these tasks. As a note for notation, hereafter we may drop the dummy variable inside the expectation and variance operators for a better presentation: \[ {\mathbbm E}_1 [\mathbf{q}_{-}] = {\mathbbm E}_1 [\mathbf{q}_{-} (X)], \hspace {5mm} {\mathrm {Var}}_1 [\mathbf{q}_{-}] = {\mathrm {Var}}_1 [\mathbf{q}_{-} (X)]. \]

Distribution estimation under the DRM

In this section, we investigate the asymptotic efficiency of the DRM-based estimator of $ G_1 (x) $ under the two-sample scenario. We first define this estimator. With the DRM-based MELE $ \hat \bm{\theta} $ defined in (ref), we have the fitted values of $ p_{k j} = P (X = x_{k j}; G_0) $ that characterize $ G_0 $: \[ \hat p_{k j} = [ n_0 + n_1 \exp \{ \hat \bm{\theta}^{\top} \mathbf{q} (x_{k j}) \} ]^{-1}. \] We then naturally obtain an estimator of $ G_0 (x) $ as $ \hat G_0 (x) = \sum_{k, j} \hat p_{k j} \mbox{$\mathbbm{1}$} (x_{k j} \leq x) $, and an estimator of $ G_1 (x) $ under the DRM:

align[align omitted — 162 chars of source]

where $ \mbox{$\mathbbm{1}$} (\cdot) $ is the indicator function.

The following theorem states that $ \hat G_1 (x) $ is also asymptotically normal as $ n_0/n_1 \to \infty $.

theoremUnder Conditions (ref) to (ref), for every $ x $ in the support of $ G_1 $, we have that as $ n_0, n_1 \to \infty $, \begin{align*} \sqrt {n_1} & \{ \hat G_1 (x) - G_1 (x) \} \overset {d} {\to} \\ & N \left (0, \{ \mathbf{Q} (x) - {\mathbbm E}_1 [\mathbf{q}_{-}] G_1 (x) \}^{\top} {\mathrm {Var}}_1^{-1} [\mathbf{q}_{-}] \{ \mathbf{Q} (x) - {\mathbbm E}_1 [\mathbf{q}_{-}] G_1 (x) \} \right ), \end{align*} where we define \[ \mathbf{Q} (x) = \int_{-\infty}^{x} \mathbf{q}_{-} (t) \mathrm {d} G_1 (t). \] The variance matrix in the limiting distribution can also be written as \[ {\mathrm {Cov}}_1 [\mathbf{Q}^{\top} (X), \bm{\mathrm{1}} (X \leq x)] {\mathrm {Var}}_1^{-1} [\mathbf{q}_{-} (X)] {\mathrm {Cov}}_1 [\mathbf{Q} (X), \bm{\mathrm{1}} (X \leq x)], \] where $ {\mathrm {Cov}}_1 $ denotes the covariance calculated under distribution $ G_1 $.

We observe that the asymptotic variance in the above theorem is the same as that in (ref), which further equals the asymptotic variance of the parametric MLE $ \tilde G_1 (x) $ in (ref). Therefore, same as the inference on $ \bm{\theta} $, the DRM-based estimator of the distribution function $ G_1 (x) $ is also as efficient as the parametric estimator under the submodel model in (ref) when $ n_0/n_1 \to \infty $.

Quantile estimation under the DRM

In this section, we derive the asymptotic variance of DRM-based quantile estimator for $ G_1 $. The $ p $th quantile of $ G_1 $ is denoted by $ \xi_p $. Estimator of $ \xi_p $ under the DRM is constructed based on the distribution estimator $ \hat G_1 (x) $ defined in (ref):

align[align omitted — 77 chars of source]

Because of the relationship between the quantile and distribution function, we often study the asymptotic properties of quantile estimator through distribution estimator. With the help from Theorem (ref) for the asymptotic normality of the DRM-based distribution estimator, we are able to derive the limiting distribution of the DRM-based quantile estimator $ \hat \xi_p $, as shown in the following theorem.

theoremUnder Conditions (ref) to (ref), and suppose that the density function $ g_1 (\cdot) $ is continuous and positive at $ \xi_p $. The DRM-based quantile estimator is asymptotically normal: \begin{align*} \sqrt {n_1} (\hat \xi_p - \xi_p) \overset {d} \to N \left (0, \{ \mathbf{Q} (\xi_p) - p {\mathbbm E}_1 [\mathbf{q}_{-}] \}^{\top} \frac {{\mathrm {Var}}_1^{-1} [\mathbf{q}_{-}]} {g_1^2 (\xi_p)} \{ \mathbf{Q} (\xi_p) - p {\mathbbm E}_1 [\mathbf{q}_{-}] \} \right ), \end{align*} as $ n_0, n_1 \to \infty $.

Evidently, the asymptotic variance in Theorem (ref) equals the variance in (ref), and equivalently is as low as the asymptotic variance in (ref) of the parametric MLE $ \tilde \xi_p $. Hence, under the two-sample scenario where $ n_1/n_0 \to 0 $, we have proved that the DRM-based quantile estimator for $ G_1 $ attains parametric efficiency as if the parametric model in (ref) is assumed.

Beyond our core contribution, we also present a useful result that reveals a linear relationship between quantile estimator and distribution estimator. This type of result is famously known as the Bahadur representation.

theoremUnder Conditions (ref) to (ref), and suppose that the density function $ g_1 (\cdot) $ is continuous and positive at $ \xi_p $. The DRM-based quantile estimator has Bahadur representation: \begin{align*} \hat \xi_p = \xi_p + \frac {G_1 (\xi_p) - \hat G_1 (\xi_p)} {g_1 (\xi_p)} + O_p (n_1^{-3/4} \log^{1/2} n_1). \end{align*}

The conclusion in Theorem (ref) is stronger than that in Theorem (ref).

Efficiency of DRM-based quantile estimation when $ G_0 = G_1 $

For illustration, in this section we demonstrate how the asymptotic variance of the DRM-based quantile estimator $ \hat \xi_p $ evolves as a function of the sample size ratio $ k = n_0/n_1 $ when two population distributions are actually identical. That is, we study the efficiency when the true model parameter $ \bm{\theta}^* = \mathbf {0} $ in the DRM (ref). However, we do not assume the knowledge of $ G_0 = G_1 $ when fitting the DRM to the data. Although our main focus is on the situation when $ k \to \infty $, we can learn a lot from the case when $ k $ is finite but large. Applying the results from zhang2000quantile and chen2013quantile, the following corollary explicitly quantifies the efficiency gap between the DRM-based quantile estimator $ \hat \xi_p $ and the parametric MLE $ \tilde \xi_p $ of the quantile.

corAssume that the DRM holds with the true $ \bm{\theta}^* = \mathbf {0} $, and $ {\mathbbm E}_{0} \left [ \exp (\bm{\theta}^{\top} \mathbf{q} (X)) \right ] < \infty $ for $ \bm{\theta} $ in a neighbourhood of $ \bm{0} $. Further, assume that $ k = n_0/n_1 $ is fixed, finite, and positive as $ n_0, n_1 \to \infty $, and that the density function $ g_1 (x) $ is continuous and positive at $ x = \xi_p $. Then, the centralized DRM-based quantile estimator $ \sqrt{n_1} (\hat \xi_p - \xi_p) $ has a limiting normal distribution with variance \begin{align} \sigma_{\hat \xi_{p}}^2 = \frac {1} {k + 1} \frac {p (1 - p)} {g_{1}^{2} (\xi_{p})} + \frac {k} {k + 1} \sigma_{\tilde \xi_{p}}^2, \end{align} where $ \sigma_{\tilde \xi_{p}}^2 $ is the variance of the limiting normal distribution of the MLE $ \tilde \xi_p $ that has a matrix expression given in (ref).

The asymptotic variance in (ref) depends on $ k = n_0/n_1 $, providing an insight on how the efficiency of the DRM quantile estimator evolves as $ n_0/n_1 \to \infty $. We now take a closer look at the first part in the right hand side of (ref). Let $ F_{n, 1} (x) = n_1^{-1} \sum_{j=1}^{n_1} \mbox{$\mathbbm{1}$} (X_{1 j} \leq x) $ denote the empirical distribution of the sample from $ G_1 $. The empirical quantile with level $ p $ for $ G_1 $ is $ \bar \xi_{p} = \inf \{ x: F_{n, 1} (x) \geq p \}. $ It is well known that if $ g_1 (\xi_p) > 0 $, $ \bar \xi_{p} $ is asymptotically normal serfling2000approximation:

align[align omitted — 200 chars of source]

Therefore, the asymptotic variance of the DRM-based quantile estimator given in (ref) is a weighted average of the asymptotic variance of the empirical quantile and the asymptotic variance of the MLE. Interestingly, as $ k = n_0/n_1 $ increases, the efficiency of the DRM-based quantile estimator $ \hat \xi_{p} $ approaches the efficiency of the MLE $ \tilde \xi_p $. Although this intriguing observation is developed under a special situation where $ G_0 = G_1 $ and $ k $ is assumed to be fixed and positive, it enlightens us on the efficiency of the DRM-based quantile estimator in addition to our theoretical result in Theorem (ref) that is established under general conditions and evolving $ k $.

Simulation studies

In this section, we report some simulation results on the efficiency of the DRM-based estimators of quantiles. We are particularly interested in their efficiency in estimating lower or higher quantiles, compared to the empirical quantiles and the parametric MLEs. Because the probability density function often has small value at $\xi_p$ when $p$ is close to zero or one, the corresponding empirical quantile has very large variance according to (ref). In contrast, the parametric quantile estimators are largely free from this deficiency but suffer from serious bias when the model is misspecified. Based on our theoretical results, DRM-based quantile estimator achieves the efficiency of the parametric estimator with low model misspecification risk in the presence of an extra sample with a much larger size. In the following sections, we use 1000 repetitions to obtain the simulated biases and variances of the three quantile estimators: the DRM-based estimator, the parametric MLE, and the empirical quantile.

Data generated from normal distributions

We first examine the performance of the DRM-based quantile estimator when data are from normal distributions. The family of normal distributions has many nice statistical properties. For example, the MLEs of the model parameters and quantiles under the normal family have closed forms. Such properties make it easier to compare the efficiency between the DRM-based estimator and the parametric MLE. Also, the normal distributions collectively satisfy the DRM with basis function $ \mathbf{q} (x) = (1, x, x^2)^{\top} $. We assume the knowledge of this basis function but not the parametric form in the DRM approach in simulations.

We generate the first sample from $ N (\mu_{0}, \sigma_{0}^{2}) $ and the second sample from $ N (\mu_{1}, \sigma_{1}^{2}) $. We observe the performance of the DRM-based quantile estimators for various choices of the means $ \mu_0, \mu_1 $, standard deviations $ \sigma_0, \sigma_1 $, and sample sizes $ n_0, n_1 $. Under normal model, the parametric MLE quantile estimator of $\xi_p$ is given by

align*[align* omitted — 78 chars of source]

where $\Phi^{-1} (p)$ is the $p$th quantile of the standard normal distribution, and $\tilde \mu_1$ and $\tilde \sigma_1$ are the sample mean and standard deviation based on $\{x_{1 j}: j=1, \ldots, n_1\}$. It can be shown that the asymptotic variance of the MLE $ \tilde \xi_{p} $ is given by

align*[align* omitted — 121 chars of source]

We first generate both samples from standard normal distribution, namely, $ \mu_0 = \mu_1 = 0 $ and $ \sigma_0 = \sigma_1 = 1 $. Table (ref) contains the simulated biases and variances of the three quantile estimators after being properly scaled based on sample size. We consider 4 quantile levels $ p \in \{ 0.01, 0.05, 0.10, 0.50 \} $ and 4 sample size combinations $ n_1 \in \{ 100, 1000 \} $ and $ n_0 \in \{ 10 n_1, 100 n_1 \} $. Due to symmetry of the normal distribution, we do not include the higher level quantiles. The biases and variances in the table are inflated by a factor $ \sqrt {n_1} $ and $ n_1 $ respectively. We observe that when $ n_0/n_1 = 100 $, the asymptotic variances of the DRM-based quantile estimators approximately equal the weighted averages of the variances of the parametric MLEs (with weight $ k/(k+1) $) and empirical quantiles (with weight $ 1/(k+1) $). This is consistent with our theoretical finding in Corollary (ref). Further, when $n_0/n_1$ increases from $10$ to $100$, the variances of the DRM-based quantile estimators $ \hat \xi_{p} $ rapidly approaches the variances of the parametric MLEs $ \tilde \xi_{p} $. The improvement in efficiency due to increased $n_0/n_1$ is especially significant for quantiles at levels $p$ close to zero.

We next experiment with data from normal distributions with $ \mu_0 = 1, \mu_1 = 2 $ and $ \sigma_0^2 = 1.5, \sigma_1^2 = 2 $. The performance of the three quantile estimators is summarized in Table (ref). We note that our previous comments on the efficiency of the DRM-based estimators when $ n_0/n_1 $ increases are also applicable here, and therefore we do not offer more interpretations. We anticipate other combinations of normal distribution, if not far different, will still lead to similar results.

table[table omitted — 2,200 chars of source]
table[table omitted — 2,190 chars of source]

Data generated from exponential distributions

We then examine the performance of the three quantile estimators based on data from exponential distributions. The exponential distributions collectively fit into the DRM with the basis function $ \mathbf{q} (x) = (1, x)^{\top} $. We assume the knowledge of this $ \mathbf{q} (x) $ but not the parametric form when fitting the DRM to the data in simulations.

The quantile of the exponential distribution with mean $ \mu $ has a simple form: $\xi_{p} = -\mu \log(1-p)$. In the two-sample setting of this article, the parametric MLE of the second population mean is the corresponding sample mean: $\tilde \mu= n_1^{-1} \sum_{j = 1}^{n_1} x_{1 j}$. This leads to the parametric MLE quantile estimator:

align*[align* omitted — 53 chars of source]

with variance

align*[align* omitted — 71 chars of source]

We first generate both samples from exponential distributions with $ \mu= 1 $. We include the three quantile estimators at levels $p \in \{ 0.5, 0.9, 0.95, 0.99 \}$. The exponential distribution has low density values at higher level quantiles. Therefore, the efficiency gain at higher level quantiles are more meaningful for our investigation. We use the same sample size combinations as in the previous section, and the results are given in Table (ref). We again observe the phenomenon illustrated in Corollary (ref): the asymptotic variances of the DRM-based quantile estimators are close to the weighted averages of the variances of the parametric MLEs (with weight $ k/(k+1) $) and empirical quantiles (with weight $ 1/(k+1) $). Same as in the normal data example, when $n_0/n_1$ increases from $10$ to $100$, the efficiency of the DRM-based quantile estimator quickly approaches the efficiency of the parametric MLE. The improvement is particularly obvious for quantiles at levels $p$ close to one.

We next simulate data from exponential distributions with $ \mu_0 = 1/0.3$ and $\mu_1 = 2 $. The performance of the quantile estimators is summarized in Table (ref). The same conclusions regarding the DRM efficiency with growing $ n_0/n_1 $ can be drawn in this situation. We expect similar findings for other combinations of exponential distributions.

table[table omitted — 2,236 chars of source]
table[table omitted — 2,251 chars of source]

Real-data analysis

In this section, we study the efficiency of the DRM-based quantile estimator and its parametric and nonparametric competitors with a real-world data. We use the collegiate sports budgets dataset from the TidyTuesday data project tidytuesday, which is accessible from the Github repository (\url {https://github.com/rfordatascience/tidytuesday/tree/master/data/2022/2022-03-29}). The dataset contains yearly samples concerning some demographics of collegiate sports in the U.S. from 2015 to 2019. The variable we consider is the total revenue (in USD) per sports team for both men and women, for which we have approximately 17,000 observations each year, and we log-transform the values to make the scale more suitable for numerical computation. As an exploratory data analysis, we plot in Figure (ref) the histograms of the log-transformed revenue data for the years 2015--2019. The population distributions of the revenues in these years look similar. This is also reflected in their kernel density estimators silverman1986density as depicted in the solid curves in Figure (ref). On one hand, the density estimates are close to bell-shaped, suggesting quantile estimation based on normal model is bearable. On the other hand, the estimated densities sufficiently deviate normal density as they have two or more modes, which can be better illustrated by comparing the estimated densities with the fitted normal densities depicted in the dashed curves. However, the estimated densities appear to have some common structures. Therefore, a DRM-based approach is more convincing and may work well in this situation.

figure[figure omitted — 210 chars of source]
figure[figure omitted — 468 chars of source]

We conduct a real data-based simulation as follows. We regard the yearly samples from 2015--2019 as five populations, and sample with replacement from these populations to form multiple samples repeatedly. As remarked previously, our two-sample results are applicable to the multi-sample situation when $n_0/n_j \to \infty$ for all $j \neq 0$. We hence make the sample size from 2015 substantially larger to mimic the situation of $n_0/n_j$ being very large. This is meaningful as we can see how a large historical dataset could help predict the future with small datasets under the DRM. Specifically, we sample from 2016--2019 with equal sizes $ n \in \{ 200, 500 \} $, and sample from 2015 with size $ n_0 \in \{ 200, 1000, 5000 \} $. To apply the DRM-based approach, the user needs to specify a basis function $ \mathbf{q} (x) $. In this simulation, we use i) data-adaptive basis function $ \mathbf{q} (x) $ zhang2022density learned using the full data with $d=2$; ii) $ \mathbf{q} (x) = (1, x)^{\top} $ whose DRM contains the normal model with equal variances; iii) $ \mathbf{q} (x) = (1, x, x^2)^{\top} $ whose DRM contains the normal model without equal variance assumption. The simulation also correspondingly includes iv) the parametric MLE of quantile derived under the normal model with common variance assumption; v) the parametric MLE under the normal model with no assumption on equal variances; and vi) the empirical quantiles. By regarding the empirical quantiles based on the full data (treated as the populations) as the truth, we compute the simulated absolute biases, variances, and mean squared errors (MSEs) of these quantile estimators at some selected levels $ p $ for each population from 2016 to 2019. To save space, we report only the average values of the aforementioned performance measures across 2016--2019 in Tables (ref)--(ref).

We first observe that when $ n $ is fixed at 200 and 500 while $ n_0 $ increases from 200 to 5000, the variances of the DRM-based quantile estimators with various $ \mathbf{q} (x) $ decrease in nearly all the cases. This observation supports our theoretical results. Second, the DRM-based quantile estimators are more efficient than the nonparametric empirical quantiles, which suggests that data pooling via the DRM works well for this data. Third, when $ n_0 = 5000 $, the quantile estimators derived under the DRM with $ \mathbf{q} (x) = (1, x)^{\top} $ overall have comparable MSEs with the parametric estimators under the normal model with common variance assumption, except for the $90\%$ quantile level. The same also applies to the comparison between the estimators under DRM with $ \mathbf{q} (x) = (1, x, x^2)^{\top} $ and the parametric estimators under the normal model with no equal variance assumption. This phenomenon can partly be explained by seeing in Figure (ref) that the true distributions deviate noticeably from the fitted normal distributions at the the right tails, which suggests normal model is unsatisfactory there. Finally, the DRM quantile estimators using the data-adaptive basis function $ \mathbf{q} (x) $ generally beat the DRM estimators using the two prespecified $ \mathbf{q} (x) $ in all three performance measures, except for very few cases where the results are still comparable. This is expected because when the basis function $ \mathbf{q} (x) $ is adaptively learned from the data, the resulting DRM should fit the data better than a predetermined DRM. In fact, when $ n_0 = 5000 $, the adaptive DRM produces overall the most accurate quantile estimators, indicating that an appropriate DRM is suitable to the data.

table[table omitted — 6,442 chars of source]
table[table omitted — 6,450 chars of source]

\endinput

Conclusions and discussion

The DRM for multi-sample data is generating growing research interest and has wide applications in statistics and econometrics. It has proven particularly useful in situations where the populations may possess shared underlying structures. The DRM provides a good trade-off between low model misspecification risk and satisfactory statistical efficiency. Its effectiveness primarily stems from its ability to enable users to draw inferences on each population using pooled data, which leads to efficiency gain compared to using individual samples that overlooks the shared latent structures. The literature has engaged in discussions regarding the efficiency of the DRM approach in comparison to the nonparametric approach. However, none of these discussions have systematically explored the limit of the efficiency gain through the DRM, which is an important research problem. This article addresses this problem by considering a scenario where one of the samples significantly outweighs the others in size. We establish through theoretical analysis that within this context, the DRM-based estimators of model parameters, distribution functions, and quantiles for the smaller sample populations attain the efficiency as if a parametric model is assumed. In essence, for these estimands and in this scenario, we identify their highest achievable efficiency under a specific parametric model, investigate their asymptotic efficiency under the DRM, and demonstrate the equivalence between the two aforementioned efficiencies. Our simulation experiments and analyses of real-world data, with a particular focus on quantile estimation, support our theoretical discoveries. The significance of this article's contribution extends to practical scenarios where researchers aim to make inferences about a population with limited data, but can rely on a substantial dataset from a related population for support.

While the scenario examined in this article addresses many real-world applications, there exist situations where this sampling scheme is not applicable. In such cases, we may encounter a growing number of samples, each of relatively similar sizes. For example, when economists investigate the evolution of income distribution over time, they may collect income samples year after year, and these yearly samples often have comparable sizes. Consequently, it becomes intriguing to investigate the asymptotic efficiency of the DRM approach when the sample sizes $ n_i/n_j \to 1 $ and the number of populations $ m \to \infty $ as $ n_i \to \infty $. Additionally, it is worthwhile to study the asymptotic efficiency of other population parameters than the ones we investigated. We leave these interesting problems for future work, anticipating that the technical methods and theoretical results in this article may be valuable for such inquiries.

Acknowledgement

This research was partially supported by the Natural Sciences and Engineering Research Council of Canada (Grants RGPIN-2018-06484, RGPIN-2019-04204, and RGPIN-2020-05897), the Canadian Statistical Sciences Institute (Grant 592307), and the Department of Statistical Sciences in the University of Toronto. We also thank the Digital Research Alliance of Canada for computing support. This work was partially completed when Archer Gong Zhang was a PhD student at the University of British Columbia and a Postdoctoral Fellow at the University of Toronto.