Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
65,556 characters · 20 sections · 35 citation commands
Semiparametric inference for inequality measures under nonignorable nonresponse using callback data
In economics and official statistics, inequality measures play a central role in summarizing the distributions of economic variables such as income, wealth, and expenditure Atkinson1970. Let $F(y)$ denote the cumulative distribution function (CDF) of a continuous random variable $Y$ supported on $(0,\infty)$, which may represent, for example, household income. In this paper, we focus on commonly used inequality measures expressed as $\theta(F)$, where $\theta(\cdot)$ is a statistical functional of $F(y)$. These measures include the Gini index, quantiles, the Theil index, the coefficient of variation, and members of the generalized entropy and Atkinson classes (see Section (ref) for their formal definitions). For a comprehensive review of inequality measures and their theoretical properties, see Cowell2011.
There is a vast literature on statistical inference for inequality measures; see Davidson2007,Giorgi2017,Dufour2019,Yuan2023,Zalghout2025 and the references therein. Much of this literature, however, implicitly assumes that the observed data are representative of the target population. This assumption is often violated in practice, as survey samples used to measure inequality frequently suffer from nonresponse or missing data; see Groves1998.
According to Rubin1976, valid inference for $\theta(F)$ must account for the underlying missing-data mechanism. If the response probability depends on the unobserved value of $Y$, the mechanism is said to be nonignorable; otherwise, it is ignorable. Nonignorable nonresponse mechanisms are common in surveys on income and other sensitive information Korinek2006,Korinek2007,Bollinger2013,Bollinger2019, as high-income households are often less likely to respond. Under such mechanisms, the observed data constitute a biased sample of the target population, because the income distribution of respondents may differ systematically from the true distribution $F(y)$ Qin2017. Moreover, $F(y)$ is not nonparametrically identifiable based solely on the biased sample Wang2014,Miao2016,Morikawa2021. Consequently, many standard inference procedures for $\theta(F)$ are either inapplicable or yield inefficient and potentially misleading results Paraje2010,Alvarez2021. In general, without additional assumptions or auxiliary information, valid statistical inference for $\theta(F)$ is not possible when only a sample subject to nonignorable nonresponse is available.
In survey practice, additional contact attempts are commonly made when the initial attempt fails. Such callback strategies are widely used to increase response rates and to monitor survey quality DrewFuller1980,Couper1998,Bates2000,Durrant2012,Kreuter2013,Olson2013. The resulting callback data, which record the number of contact attempts for initially nonresponding units, are routinely collected and provide valuable auxiliary information for identifying $F(y)$ and for modeling and adjusting for nonignorable nonresponse Alho1990,Qin2014,Miao2024.
Let $m \ge 2$ denote the maximum number of contact attempts, and let $D$ be the callback variable, where $D=j$ if a household responds at the $j$th call, $j=1,\ldots,m$, and $D=m+1$ if it never responds. Thus, $Y$ is observed when $D \le m$ and missing otherwise. Define $\pi_j(y)=\Pr(D=j \mid Y=y, D \ge j)$ as the response probability at the $j$th call given nonresponse to earlier calls. The functions $\pi_j(y)$ characterize the nonignorable missing-data mechanism. Because income distributions are typically skewed and may exhibit complex shapes, we adopt the semiparametric model of Alho1990, which leaves $F(y)$ unspecified while modeling each $\pi_j(y)$ parametrically via a logistic regression,
Here, $\boldsymbol q(\cdot)$ is a prespecified $d$-dimensional function, with parameters $\boldsymbol\alpha=(\alpha_1,\ldots,\alpha_m)^\top$ and $\boldsymbol\beta$. When $\boldsymbol q(y)=y$, this reduces to Alho's original model, while allowing a flexible choice of $\boldsymbol q(y)$ improves adaptability. In this formulation, $\alpha_j$ captures call-specific response heterogeneity, and $\boldsymbol\beta$ is common across calls under the continuum of resistance assumption, which ensures identifiability of the semiparametric model Guan2018,Miao2024.
Under model (ref), Alho1990 proposed a conditional likelihood approach to estimate $\boldsymbol\alpha$ and $\boldsymbol\beta$, leading to an inverse probability weighting (IPW) estimator of the population mean. Building on the same model, Qin2014 developed a semiparametric full-likelihood method that estimates $\boldsymbol\alpha$, $\boldsymbol\beta$, and the population mean. Their approach yields efficient inference that is robust to assumptions on $F(y)$ and avoids the instability of IPW estimators Li2023, particularly when the estimated response probabilities are small. Recent extensions using callback data include generalized method of moments estimation of $\boldsymbol\alpha$, $\boldsymbol\beta$, and the finite population mean Kim2014, sensitivity analysis Daniels2015, and the incorporation of covariate information Guan2018,Chen2018.
To the best of our knowledge, no existing methods use callback data to develop valid inference procedures for inequality measures $\theta(F)$ under nonignorable nonresponse. Existing results based on callback data focus primarily on inference for the mean of $Y$ and do not extend to more general measures $\theta(F)$, which may involve complex linear, nonlinear, or nonsmooth functionals of $F$, such as quantiles or the Gini index. Unlike inference for the mean, inference for these measures is substantially more challenging, because establishing the theoretical properties of the associated estimators requires delicate asymptotic arguments. Moreover, existing methods rely on complex numerical optimization, and no efficient algorithm with theoretical guarantees is currently available.
In this paper, we develop efficient and reliable inference procedures for general inequality measures $\theta(F)$ under nonignorable nonresponse by incorporating callback data within the semiparametric model of Alho1990. Our contributions are summarized as follows.
The paper is organized as follows. Section (ref) proposes full-likelihood estimators for the model parameters and $\theta(F)$, establishes their asymptotic properties, and constructs valid confidence intervals for various inequality measures. Section (ref) presents an EM algorithm for practical implementation. Section (ref) reports the results of simulation studies, and Section (ref) analyzes a real income data set from the Consumer Expenditure Survey. Section (ref) concludes with some remarks. All proofs are provided in the Supplementary Material.
Suppose we observe a random sample $\{(Y_i, D_i): i=1,\ldots,N\}$ generated from Alho's semiparametric model (ref), where $F(y)$ is left unspecified. Let $n = \sum_{i=1}^N I(D_i \le m)$ denote the random number of respondents. Without loss of generality, we index the first $n$ observations as respondents and the remaining $N-n$ as nonrespondents, for whom the $Y_i$'s are unobserved. Hereafter, we refer to the first $n$ observations as the complete-case (CC) data.
Let
Then \[ \eta = \Pr(D \le m) = \int_0^\infty \rho(y;\boldsymbol\alpha,\boldsymbol\beta)\, dF(y) \] denotes the marginal probability that a generic household responds within $m$ callbacks. Denote the full parameter vector as $\boldsymbol{\varphi} = (\boldsymbol\alpha,\boldsymbol\beta,\eta,F)$.
We now derive the full likelihood for $\boldsymbol{\varphi}$. It factors into three components: (i) the likelihood of $D_1,\ldots,D_n$ conditional on $Y_1,\ldots,Y_n$ and on the event that the $n$ households respond within $m$ callbacks; (ii) the likelihood of $Y_1,\ldots,Y_n$ given that these $n$ households respond within $m$ callbacks; and (iii) the likelihood of the total number of respondents $n$.
First, given that the $i$th household responded within $m$ callbacks and provided the observed response $Y_i$, the conditional probability of responding on the $D_i$th callback is
Hence, the likelihood contribution from $D_1,\ldots,D_n$ conditional on $Y_1,\ldots,Y_n$ and on the event that the $n$ households responded within $m$ callbacks, also known as the conditional likelihood Alho1990, is
Second, given that the $i$th household responded within $m$ callbacks, the conditional probability of observing $Y_i$ is
Therefore, the likelihood contribution from $Y_1,\ldots,Y_n$ given that the $n$ households responded within $m$ callbacks is
Third, the number of respondents $n$ follows a $\mathrm{Binomial}(N,\eta)$ distribution, with likelihood
Combining (ref) and (ref)--(ref) yields the full likelihood of $\boldsymbol{\varphi}$ as
Following the empirical likelihood principle Owen2001, we model $F$ as \[ F(y) = \sum_{i=1}^{n} p_i I(Y_i \le y), \] where $p_i = dF(Y_i)$ for $i=1,\ldots,n$. Up to an additive constant that does not depend on $\boldsymbol{\varphi}$, the semiparametric full log-likelihood function is
where $\boldsymbol{\varphi}$ is subject to the constraints
The first two constraints ensure that $F$ is a valid CDF, while the last follows from the definition of $\eta$ and accounts for sample selection bias induced by nonignorable nonresponse.
The semiparametric maximum full-likelihood estimator of $\boldsymbol{\varphi}$ is defined as
Given $\hat{\boldsymbol{\varphi}}$, the corresponding semiparametric maximum full-likelihood estimator of $F(y)$ is
Since a closed-form expression for $\hat{\boldsymbol{\varphi}}$ is generally unavailable, we rely on numerical methods. In Section (ref), we develop a stable and computationally convenient EM algorithm to compute $\hat{\boldsymbol{\varphi}}$ and $\hat{F}(y)$. The resulting estimator $\hat{F}(y)$ provides the foundation for correcting sample selection bias and for estimating the inequality measures $\theta(F)$ in the presence of nonignorable nonresponse.
In this section, we propose a semiparametric approach for estimating inequality measures. Recall that the measures of interest, $\theta(F)$, can be expressed as statistical functionals of the income distribution $F$. Given the estimator $\hat{F}(y)$, a natural strategy is to adopt the plug-in principle by replacing $F(y)$ with $\hat{F}(y)$, which yields $\theta(\hat{F})$ as the semiparametric estimator of $\theta(F)$. This differs fundamentally from the IPW-type approach. For simplicity, we denote the resulting semiparametric full-likelihood estimator by $\hat{\theta} = \theta(\hat{F})$. We focus on three widely used classes of inequality measures, each corresponding to a specific functional form of $\theta(\cdot)$.
In applications, population quantiles and their functions are widely recognized as important measures of inequality Prendergast2017,Zalghout2025. Common examples include the median, quartiles, quintiles, percentiles, and the interquartile range of income. Moreover, following the implementation of new policies, trends in quantiles, as well as their differences or ratios, provide valuable information for monitoring changes in the income distribution and assessing policy impacts. In this section, we focus on the estimation of quantiles and their functions, which are nonlinear and nonsmooth functionals of $F$.
For $\tau \in (0,1)$, the $100\tau$th quantile of $F(y)$ is defined as
Replacing $F(y)$ with $\hat{F}(y)$ in (ref), the proposed semiparametric estimator of $\theta_{\mathrm{quan}}^{\tau}(F)$ is given by
If interest lies in functions of quantiles, such as the quantile difference $\theta_{\mathrm{quan}}^{\tau_1} - \theta_{\mathrm{quan}}^{\tau_2}$ or the quantile ratio $\theta_{\mathrm{quan}}^{\tau_1}/\theta_{\mathrm{quan}}^{\tau_2}$ for $\tau_1,\tau_2 \in (0,1)$, the corresponding semiparametric estimators are $\hat{\theta}_{\mathrm{quan}}^{\tau_1} - \hat{\theta}_{\mathrm{quan}}^{\tau_2}$ and $\hat{\theta}_{\mathrm{quan}}^{\tau_1}/\hat{\theta}_{\mathrm{quan}}^{\tau_2}$, respectively.
We next consider an important class of inequality measures that can be expressed as linear functionals of $F$ and their transformations. A $p$-dimensional linear functional of $F$ is defined as
for suitable choices of $\boldsymbol{u}(y) = \big(u_1(y),\ldots,u_p(y)\big)^\top$. Many commonly used inequality measures can be written in the unified form
where $h(\cdot): \mathbb{R}^p \to \mathbb{R}$ is a smooth function of $\boldsymbol\gamma_{\boldsymbol{u}}$. Examples in this class include population moments (centered or uncentered), the Theil index, the coefficient of variation, the mean logarithmic deviation, and members of the generalized entropy and Atkinson classes. Table (ref) summarizes these measures along with the corresponding specifications of $\boldsymbol{u}(\cdot)$ and $h(\cdot)$.
In our numerical studies, we focus on an important and widely used inequality measure, the Theil index Theil1967, defined as
It is immediate that $\theta_{\mathrm{Theil}}(F)$ is a special case of the class $\theta_{\boldsymbol{u},h}(F)$ with $\boldsymbol{u}(y) = \big(y, y\log(y)\big)^\top$ and $h(x_1,x_2)=x_2/x_1-\log(x_1)$. Moreover, $\theta_{\mathrm{Theil}}(F)$ corresponds to the limiting case of the generalized entropy class as the governing parameter $\zeta \to 1$. A useful property of the Theil index, and more generally of functionals of the form (ref), is their additive decomposability by subgroup Cowell2011, which facilitates the analysis of within-group and between-group inequality.
By definition (ref), $\theta_{\boldsymbol{u},h}(F)$ can be estimated by replacing $F(y)$ with $\hat{F}(y)$ in (ref). The proposed semiparametric estimator for this class of inequality measures is \[ \hat{\theta}_{\boldsymbol{u},h} = \theta_{\boldsymbol{u},h}(\hat{F}) = h(\hat{\boldsymbol\gamma}_{\boldsymbol{u}}), \quad \text{with} \quad \hat{\boldsymbol\gamma}_{\boldsymbol{u}} = \sum_{i=1}^n \hat{p}_i \boldsymbol{u}(Y_i). \] In particular, the proposed estimator of the Theil index is
Although $\theta_{\mathrm{quan}}^{\tau}(F)$ and $\theta_{\boldsymbol{u},h}(F)$ cover a wide range of inequality measures, the Gini index Gini1912 is a notable exception that cannot be expressed within these classes. In this section, we therefore consider the Gini index through a specific functional form of $\theta(\cdot)$. The Gini index has several equivalent representations; here, we adopt the following formulation Yuan2023:
where $\psi(F) = \int_0^\infty 2y F(y)\, dF(y)$ and $\mu(F) = \int_0^\infty y\, dF(y)$. Unlike the measures considered previously, $\theta_{\mathrm{Gini}}(F)$ is distinct in that $\psi(F)$ involves a U-statistic-type functional of $F$. The Gini index is among the most widely used measures of inequality and is routinely reported by major official statistical agencies worldwide. Its widespread use is also attributable to its close relationship with the Lorenz curve Lorenz1905, which provides an intuitive graphical interpretation.
Using $\hat{F}(y)$ in (ref), the proposed semiparametric estimator of $\theta_{\mathrm{Gini}}(F)$ is $\hat{\theta}_{\mathrm{Gini}} = \theta_{\mathrm{Gini}}(\hat{F})$, which can be written explicitly as
where $\hat{\mu} = \mu(\hat{F}) = \sum_{i=1}^{n} \hat{p}_i Y_i$ and $\hat{\psi} = \psi(\hat{F}) = 2\sum_{i=1}^{n} \hat{p}_i Y_i \hat{F}(Y_i)$. These quantities are introduced for notational convenience in subsequent theoretical developments.
In this section, we investigate the asymptotic behavior of the estimators of inequality measures proposed in Section (ref). Although the plug-in principle provides a natural approach for estimating inequality measures $\theta(F)$, establishing the asymptotic properties of the resulting estimators for statistical inference is nontrivial and requires a careful combination of appropriate asymptotic tools.
To facilitate the presentation, we introduce additional notation. Let $(\boldsymbol\alpha^*, \boldsymbol\beta^*, \eta^*)$ denote the true value of the model parameters $(\boldsymbol\alpha, \boldsymbol\beta, \eta)$. Define $\rho^*(y) = \rho(y; \boldsymbol\alpha^*, \boldsymbol\beta^*)$ and $\boldsymbol{\pi}^*(y) = \big(\pi(y;\alpha_1^*, \boldsymbol\beta^*), \ldots, \pi(y;\alpha_m^*, \boldsymbol\beta^*)\big)^\top$. Let $\mbox{\bf 1}_k$ and $\mbox{\bf 0}_k$ denote $k \times 1$ vectors of ones and zeros, respectively, and let $\mbox{\bf I}_k$ denote the $k \times k$ identity matrix. We further define \[ \boldsymbol{A}(y) = \big(\mbox{\bf I}_m,\ \mbox{\bf 1}_m \boldsymbol q(y)^\top\big)^\top \quad \text{and} \quad \boldsymbol{v}(y) = \rho^*(y)^{-1} \big(\{\rho^*(y) - 1\}\boldsymbol{\pi}^*(y)^\top \boldsymbol{A}(y)^\top,\ 1,\ (\eta^*)^2\big)^\top. \]
We begin by studying the asymptotic behavior of the proposed semiparametric estimator $\hat{F}(y)$ of $F(y)$.
The result of Theorem (ref) provides the foundation for deriving the asymptotic properties of the proposed estimators of inequality measures $\hat{\theta}_{\mathrm{quan}}^{\tau}$, $\hat{\theta}_{\boldsymbol{u},h}$, and $\hat{\theta}_{\mathrm{Gini}}$ in the subsequent theorems.
Building on Theorem (ref), we first apply the von Mises calculus and the functional delta method to derive a Bahadur representation for $\hat{\theta}_{\mathrm{quan}}^{\tau}$ Bahadur1966. This representation provides a convenient and tractable route for establishing the asymptotic properties of $\hat{\theta}_{\mathrm{quan}}^{\tau}$. Let $\theta_{\mathrm{quan}}^{\tau *}$ denote the true value of $\theta_{\mathrm{quan}}^{\tau}(F)$.
Combining the Bahadur representation of $\hat{\theta}_{\mathrm{quan}}^{\tau}$ in Lemma (ref) with the result of Theorem (ref), we establish the joint asymptotic normality of $\hat{\theta}_{\mathrm{quan}}^{\tau_1}$ and $\hat{\theta}_{\mathrm{quan}}^{\tau_2}$ for any $\tau_1, \tau_2 \in (0,1)$ in the following theorem. Throughout, we use “$\overset{d}{\to}$” to denote convergence in distribution.
Theorem (ref) also implies the marginal asymptotic normality of $\hat{\theta}_{\mathrm{quan}}^{\tau}$. In studies of income inequality, researchers are often interested in smooth functions of quantiles at different levels, such as $\hat{\theta}_{\mathrm{quan}}^{\tau_1} - \hat{\theta}_{\mathrm{quan}}^{\tau_2}$ and $\hat{\theta}_{\mathrm{quan}}^{\tau_1} / \hat{\theta}_{\mathrm{quan}}^{\tau_2}$. Using the result of Theorem (ref), the asymptotic distributions of such smooth functions can be conveniently derived via the delta method, as stated in the following corollary.
We next study the asymptotic properties of the proposed estimators of inequality measures in the class $\hat{\theta}_{\boldsymbol{u},h} = h(\hat{\boldsymbol\gamma}_{\boldsymbol{u}})$ for smooth functions $h$, which includes the Theil index as a special case. Let $\boldsymbol\gamma_{\boldsymbol{u}}^*$ and $\theta_{\boldsymbol{u},h}^*$ denote the true values of $\boldsymbol\gamma_{\boldsymbol{u}}$ and $\theta_{\boldsymbol{u},h}(F)$ defined in (ref) and (ref), respectively.
Theorem (ref) immediately implies the asymptotic normality of the proposed estimator $\hat{\theta}_{\mathrm{Theil}}$. Let $\theta_{\mathrm{Theil}}^*$ denote the true value of the Theil index $\theta_{\mathrm{Theil}}(F)$.
We now investigate the asymptotic properties of the proposed estimator of the Gini index, $\hat{\theta}_{\mathrm{Gini}}$, which requires additional asymptotic tools beyond those used in the previous section. Recall the definition of $\hat{\theta}_{\mathrm{Gini}}$ in (ref), where the numerator can be written as \[ \hat{\psi} = 2 \sum_{i=1}^{n} \hat{p}_i Y_i \hat{F}(Y_i) = \sum_{i=1}^{n} \sum_{j=1}^{n} 2 \hat{p}_i \hat{p}_j Y_i I(Y_j \le Y_i). \] Unlike $\hat{\boldsymbol\gamma}_{\boldsymbol{u}}$, the statistic $\hat{\psi}$ does not admit a linear representation, as it involves pairs of observations. As a result, Theorem (ref) is not directly applicable for deriving the asymptotic distribution of $\hat{\theta}_{\mathrm{Gini}}$. Instead, noting that $\hat{\psi}$ has the structure of a second-order V-statistic, we invoke asymptotic results from the theory of U- and V-statistics to establish the asymptotic normality of $\hat{\theta}_{\mathrm{Gini}}$. For notational convenience, let $\psi^*$, $\mu^*$, and $\theta_{\mathrm{Gini}}^*$ denote the true values of $\psi(F)$, $\mu(F)$, and $\theta_{\mathrm{Gini}}(F)$ defined in (ref), respectively.
In this section, we focus on constructing confidence intervals for the inequality measures under consideration. Based on the asymptotic normality results established in Section (ref), it is natural to construct Wald-type confidence intervals. To estimate the asymptotic variances of the proposed estimators, we replace the unknown parameters with their consistent estimators. Specifically, we estimate $(\boldsymbol\alpha^*, \boldsymbol\beta^*, \eta^*, F)$ by $(\hat{\boldsymbol\alpha}, \hat{\boldsymbol\beta}, \hat{\eta}, \hat{F})$, and estimate the density function $f(y)$ of $Y$ by the procedure as described below.
Since $Y$ is supported on $(0,\infty)$, we first estimate the density of the transformed variable $T=\log(Y)$, denoted by $k(t)$. Given $\hat{F}(y)$ in (ref), $k(t)$ is estimated by \[ \hat{k}(t) = \int K_b\big(t - \log(y)\big)\, d\hat{F}(y), \] where $K_b(x) = K(x/b)/b$ and $K(\cdot)$ is the standard normal kernel. The bandwidth $b$ is selected according to \[ b = 1.06\, n^{-1/5} \min\!\left(\widehat{\mathrm{IQR}}/1.34,\; \hat{\sigma}\right), \] where $\widehat{\mathrm{IQR}}$ and $\hat{\sigma}^2$ denote the interquartile range and variance estimators of $\log(Y)$, respectively, computed using $\hat{F}(y)$ Silverman1986. These quantities can be obtained using the methods described in Sections (ref) and (ref). Finally, the density of $Y$ is estimated by \[ \hat{f}(y) = \hat{k}\big(\log(y)\big)/y. \]
Let $\hat{\boldsymbol{\Omega}}(\tau_1,\tau_2)$ denote the estimator of $\boldsymbol{\Omega}(\tau_1,\tau_2)$ obtained by replacing $(\boldsymbol\alpha^*, \boldsymbol\beta^*, \eta^*, F, f)$ with $(\hat{\boldsymbol\alpha}, \hat{\boldsymbol\beta}, \hat{\eta}, \hat{F}, \hat{f})$. Similarly, let $(\hat{\sigma}_{\boldsymbol{u},h}^2, \hat{\sigma}_{\mathrm{Theil}}^2, \hat{\sigma}_{\mathrm{Gini}}^2)$ denote the estimators of $(\sigma_{\boldsymbol{u},h}^2, \sigma_{\mathrm{Theil}}^2, \sigma_{\mathrm{Gini}}^2)$ obtained by replacing $(\boldsymbol\alpha^*, \boldsymbol\beta^*, \eta^*, F)$ with $(\hat{\boldsymbol\alpha}, \hat{\boldsymbol\beta}, \hat{\eta}, \hat{F})$. These variance estimators can be shown to be consistent. By Slutsky's theorem, Wald-type confidence intervals can therefore be constructed for $\theta_{\mathrm{quan}}^{\tau}(F)$, $g(\theta_{\mathrm{quan}}^{\tau_1}, \theta_{\mathrm{quan}}^{\tau_2})$, $\theta_{\boldsymbol{u},h}(F)$, $\theta_{\mathrm{Theil}}(F)$, and $\theta_{\mathrm{Gini}}(F)$.
For example, a $100(1-\alpha)\%$ Wald-type confidence interval for $\theta_{\mathrm{quan}}^{\tau}(F)$ is given by \[ \left[ \hat{\theta}_{\mathrm{quan}}^{\tau} - \frac{z_{1-\alpha/2}\, \hat{\sigma}_{11}(\tau)}{\sqrt{N}}, \; \hat{\theta}_{\mathrm{quan}}^{\tau} + \frac{z_{1-\alpha/2}\, \hat{\sigma}_{11}(\tau)}{\sqrt{N}} \right], \] where $\hat{\sigma}_{11}(\tau)$ is the $(1,1)$ element of $\hat{\boldsymbol{\Omega}}(\tau,\tau)$, representing the estimated asymptotic standard deviation of $\hat{\theta}_{\mathrm{quan}}^{\tau}$, and $z_{1-\alpha/2}$ denotes the $100(1-\alpha/2)$th percentile of the standard normal distribution. Confidence intervals for $g(\theta_{\mathrm{quan}}^{\tau_1}, \theta_{\mathrm{quan}}^{\tau_2})$, $\theta_{\boldsymbol{u},h}(F)$, $\theta_{\mathrm{Theil}}(F)$, and $\theta_{\mathrm{Gini}}(F)$ can be constructed analogously. A simulation study evaluating the finite-sample performance of the proposed Wald-type confidence intervals for various inequality measures is presented in Section (ref).
Since the $Y_i$'s are unobserved for nonrespondents, the problem can be viewed as a special case of missing data. The EM algorithm thus provides a natural framework for numerically computing $\hat{\boldsymbol{\varphi}}$ defined in (ref).
Recall that the observed data consist of $\mathcal{O} = \{(Y_i, D_i): i=1,\ldots,n\} \cup \{D_k = m+1: k=n+1,\ldots,N\}$. Let $\mathcal{O}^* = \{(Z_k, D_k = m+1): k=n+1,\ldots,N\}$ denote the unobserved responses for the $N-n$ nonrespondents. If both $\mathcal{O}$ and $\mathcal{O}^*$ were available, the complete-data likelihood would take the form
Recall that $p_i = dF(Y_i)$ for $i=1,\ldots,n$, and that $F$ is modeled as $F(y) = \sum_{i=1}^{n} p_i I(Y_i \le y)$. Under this specification, each latent variable $Z_k$ can take values only in $\{Y_1,\ldots,Y_n\}$. The corresponding complete-data log-likelihood is therefore
subject to the constraint set $\mathcal{C}$ in (ref).
As in Dempster1977, each iteration of the EM algorithm consists of two steps, the E-step and the M-step. Suppose that $r$ iterations have been completed and the current parameter value is $\boldsymbol{\varphi}^{(r)}$. The $(r+1)$th iteration then proceeds as follows.
In the E-step, given the current parameter value $\boldsymbol{\varphi}^{(r)}$ and the observed data $\mathcal{O}$, we compute the conditional expectation of the complete-data log-likelihood
which is maximized in the subsequent M-step. To evaluate $Q(\boldsymbol{\varphi} \mid \boldsymbol{\varphi}^{(r)})$, it suffices to compute
where the final equality follows by an argument analogous to that leading to (ref). Combining (ref)--(ref), we can write \[ Q(\boldsymbol{\varphi} \mid \boldsymbol{\varphi}^{(r)}) = \ell_1^{(r)}(\boldsymbol\alpha,\boldsymbol\beta) + \ell_2^{(r)}(p_1,\ldots,p_n), \] where
and \[ w_i^{(r)} = (N-n)\, \frac{p_i^{(r)}\{1-\rho(Y_i;\boldsymbol\alpha^{(r)},\boldsymbol\beta^{(r)})\}} {1-\eta^{(r)}}. \]
In the M-step, we update the parameter vector from $\boldsymbol{\varphi}^{(r)}$ to $\boldsymbol{\varphi}^{(r+1)}$ by solving \[ \boldsymbol{\varphi}^{(r+1)} = \arg\max_{\boldsymbol{\varphi}} Q(\boldsymbol{\varphi} \mid \boldsymbol{\varphi}^{(r)}) \quad \text{subject to the constraints in \eqref{const}}. \] Owing to the additive structure of $Q(\boldsymbol{\varphi} \mid \boldsymbol{\varphi}^{(r)})$, the update $\boldsymbol{\varphi}^{(r+1)}$ can be obtained explicitly as
It is worth noting that the objective function $\ell_1^{(r)}(\boldsymbol\alpha,\boldsymbol\beta)$ is proportional to the weighted log-likelihood of a logistic regression model. Consequently, $(\boldsymbol\alpha^{(r+1)}, \boldsymbol\beta^{(r+1)})$ can be readily obtained by fitting a weighted logistic regression, which is supported by most standard statistical software. Additional implementation details are provided in Section 4 of the Supplementary Material.
We iterate the E-step and M-step until the increase in the log-likelihood $\ell(\boldsymbol{\varphi})$ in (ref) between successive iterations falls below a prespecified tolerance level, for example, $10^{-5}$. The following proposition establishes the monotonicity property of the proposed EM algorithm.
Based on extensive numerical experimentation, we find that the EM algorithm described above is not sensitive to the choice of initial values for $\boldsymbol{\varphi}$. Accordingly, in our numerical implementation, we initialize $\boldsymbol\alpha^{(0)} = \mbox{\bf 0}$, $\boldsymbol\beta^{(0)} = \mbox{\bf 0}$, and $p_i^{(0)} = 1/n$ for $i = 1,\ldots,n$, and set \[ \eta^{(0)} = \sum_{i=1}^n p_i^{(0)} \rho\!\left(Y_i; \boldsymbol\alpha^{(0)}, \boldsymbol\beta^{(0)}\right) = 1 - 2^{-m}. \]
In this section, we report simulation results assessing the finite-sample performance of the proposed semiparametric estimators and their associated confidence intervals for $\theta_{\text{quan}}^{\tau}(F)$ with $\tau = 0.25, 0.50,$ and $0.75$, $\theta_{\text{Theil}}(F)$ (as a representative member of the $\theta_{\mathbf{u},h}(F)$ class), and $\theta_{\text{Gini}}(F)$ under nonignorable nonresponse.
Three distributions are considered for $F(y)$: $\mathrm{Exp}(1)$, $\chi^2(1.5)$, and $\mathrm{Gam}(0.8,0.25)$, where $\mathrm{Exp}(a_1)$ denotes an exponential distribution with rate parameter $a_1$, $\chi^2(a_2)$ denotes a chi-squared distribution with $a_2$ degrees of freedom, and $\mathrm{Gam}(a_3,a_4)$ denotes a gamma distribution with shape parameter $a_3$ and rate parameter $a_4$. The callback indicator $D$ is generated according to Alho's callback model (ref) with $\boldsymbol q(y)=\log(y)$ and $m=2$. The parameter vector $(\boldsymbol\alpha^\top,\beta)$ is set to $(-1.5,\,0.5,\,-0.5)$, yielding moderate response rates across contacts. The negative value of $\beta$ implies that the response probability decreases as $Y$ increases.
The resulting true response probabilities for each contact attempt and the true values of the inequality measures are presented in Table (ref). Simulations are conducted for sample sizes $N = 1000$ and $2000$, with $M = 5000$ Monte Carlo replications for each design.
We compare the finite-sample performance of the following three estimators for the inequality measures:
The CC estimator discards all units with nonresponse or missing values and is therefore generally inconsistent under nonignorable nonresponse. The Ideal estimator is included solely as a benchmark and is not available in practice.
Finite-sample performance of a point estimator is evaluated using the relative bias (RB) and the root mean squared error (RMSE), defined as
where $\hat{x}^{(i)}$ denotes the estimate of a target parameter with true value $x$ in the $i$th replication $(i=1,\ldots,M)$. Simulation results are reported in Table (ref).
We summarize the main findings from Table (ref). First, as anticipated, the CC estimator performs poorly under nonignorable nonresponse, exhibiting substantial RB and RMSE across all simulation settings. Moreover, the CC estimator does not display consistency, as neither RB nor RMSE decreases when the sample size increases from $N=1000$ to $N=2000$. The CC estimator systematically underestimates the quantiles (negative RB) and overestimates the Theil and Gini indices (positive RB). Second, both the proposed semiparametric estimator and the Ideal estimator exhibit negligible RB across all settings. The RMSEs of the proposed estimator are comparable to those of the benchmark Ideal estimator and decrease as the sample size increases. Overall, the results indicate that the proposed semiparametric approach effectively corrects for bias induced by nonignorable nonresponse and achieves near-benchmark efficiency in finite samples.
The finite-sample performance of confidence intervals is evaluated using the coverage probability (CP), reported as a percentage, and the average length (AL), defined as
where $\mathcal{I}^{(i)}$ denotes the confidence interval for a target parameter $x$ in the $i$th replication, and $|\mathcal{I}^{(i)}|$ denotes its length. Simulation results for the nominal 95% confidence intervals of the quantiles, the Theil index, and the Gini index described in Section (ref), for sample sizes $N=1000$ and $2000$, are reported in Table (ref).
We summarize the main findings from Table (ref). Overall, the confidence intervals constructed using the proposed semiparametric method achieve coverage probabilities close to the nominal 95% level across most settings, with moderate average lengths. Some undercoverage is observed for certain parameters, such as $\theta_{\text{Theil}}(F)$, when the sample size is $N=1000$. As the sample size increases to $N=2000$, coverage probabilities approach the nominal level and interval lengths become shorter. These results indicate that the proposed semiparametric inference procedure provides reliable finite-sample inference under nonignorable nonresponse with callback data.
In this section, we apply the proposed estimation and inference procedure to data from the Consumer Expenditure Survey (CES), which is conducted by the United States Bureau of Labor Statistics to collect detailed information on household expenditures and related characteristics. A comprehensive description of the survey design is available at \url{https://www.census.gov/programs-surveys/ce.html}.
Our analysis uses the public-use adult data and associated paradata from the first quarter of the 2024 CES. For illustration, we consider the total amount of family resources after taxes over the previous 12 months (measured in \$10,000) as the study variable $Y$. The sample consists of 10,273 households with positive recorded values, with an overall nonresponse rate of 54.54%. The objective is to conduct efficient and reliable statistical inference for distributional and inequality measures, including quantiles, the Theil index, and the Gini index, using the CES data.
The maximum number of contact attempts recorded in the CES paradata is 18. Figure (ref) displays the response rates by number of contact attempts. The callback strategy is most effective during the initial contacts, with response rates declining as the number of calls increases. Specifically, 6.75% and 11.10% of households responded on the first and second calls, respectively, while an additional 27.61% responded in subsequent calls, yielding an overall response rate of 45.46%. Motivated by this pattern, we define four callback categories: $D=1$ and $D=2$ for households responding on the first and second calls, respectively; $D=3$ for those responding at later calls; and $D=4$ for households that did not respond.
To implement the proposed method, we use the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC) to select the functional form of $\boldsymbol q(y)$ in model (ref) from the candidate set $\{y, y^2, \log(y), \log^2(y)\}$ and their combinations. Based on the results in Table (ref), $\boldsymbol q(y)=\log(y)$ is selected, as it yields the smallest AIC and BIC values. The semiparametric full-likelihood estimates of the model parameters and their corresponding 95% confidence intervals are reported in Table (ref). In particular, the estimate of $\beta$ is $-0.191$, with a 95% confidence interval of $[-0.293,-0.090]$, indicating statistically significant evidence of nonignorable nonresponse in the CES data, with response propensity decreasing as the value of $Y$ increases.
The proposed semiparametric full-likelihood estimates of the quantiles, the Theil index, and the Gini index, together with their 95% confidence intervals, are reported in Table (ref). For comparison, the table also includes the corresponding estimates based on the 4{,}670 complete cases. The CC estimates differ markedly from the proposed semiparametric estimates, and all CC estimates fall outside the 95% confidence intervals constructed using the proposed method. This finding is consistent with the simulation results and indicates that the observed complete cases may not represent a random sample of the target distribution.
Figure (ref) compares the estimated distribution functions obtained from the proposed method and the CC method. A pronounced discrepancy between the two estimates is evident. Overall, these results suggest that ignoring nonignorable nonresponse can lead to substantial bias and unreliable inference for inequality measures, whereas the proposed semiparametric approach, which effectively incorporates callback information, yields more reliable estimates and valid inference for a range of commonly used inequality measures.
Nonignorable nonresponse has long been recognized as a major challenge in survey data analysis, yet its impact on inference for inequality measures remains relatively underexplored. In this paper, we develop a unified semiparametric framework for estimation and inference of inequality measures in the presence of nonignorable nonresponse. The proposed approach leverages callback data within a semiparametric full-likelihood formulation, leaving the underlying distribution $F(y)$ unspecified. We establish the asymptotic properties of the proposed estimators for a broad class of inequality measures, enabling valid inference under nonignorable nonresponse. From a practical perspective, we introduce an efficient and numerically stable EM algorithm with theoretical guarantees, facilitating straightforward implementation. Simulation studies and an empirical application illustrate the favorable finite-sample performance of the proposed estimation and inference procedures for commonly used inequality measures.
Several directions merit further investigation. First, in addition to nonignorable nonresponse, survey data often exhibit other forms of incompleteness, such as grouped reporting Gastwirth1972 and top-coding Feng2006. Extending the proposed framework to accommodate these features would be of practical interest. Second, it would be worthwhile to study inference for more general parameters defined through moment restrictions Qin1994 under nonignorable nonresponse. Finally, alternative sources of auxiliary information, such as regional response rates Korinek2006,Korinek2007, have been proposed as substitutes for callback data. Developing semiparametric inference theory under such settings represents another promising direction for future research.
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
C. Wang's work was supported by the Humanities and Social Sciences Foundation of the Ministry of Education of China (24YJA910005) and the National Natural Science Foundation of China (71988101); T. Yu's work is supported in part by the Singapore Ministry of Education Academic Research Tier 1 Fund (A-8000413-00-00); P. Li's work was supported by the Natural Sciences and Engineering Research Council of Canada (RGPIN-2020-0496).