EconBase
← Back to paper

Semiparametric inference for inequality measures under nonignorable nonresponse using callback data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

65,556 characters · 20 sections · 35 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Semiparametric inference for inequality measures under nonignorable nonresponse using callback data

center[center omitted — 477 chars of source]
abstractThis paper develops semiparametric methods for estimation and inference of widely used inequality measures when survey data are subject to nonignorable nonresponse, a challenging setting in which response probabilities depend on the unobserved outcomes. Such nonresponse mechanisms are common in household surveys and invalidate standard inference procedures due to selection bias and lack of population representativeness. We address this problem by exploiting callback data from repeated contact attempts and adopting a semiparametric model that leaves the outcome distribution unspecified. We construct semiparametric full-likelihood estimators for the underlying distribution and the associated inequality measures, and establish their large-sample properties for a broad class of functionals, including quantiles, the Theil index, and the Gini index. Explicit asymptotic variance expressions are derived, enabling valid Wald-type inference under nonignorable nonresponse. To facilitate implementation, we propose a stable and computationally convenient expectation-maximization algorithm, whose steps either admit closed-form expressions or reduce to fitting a standard logistic regression model. Simulation studies demonstrate that the proposed procedures effectively correct nonresponse bias and achieve near-benchmark efficiency. An application to Consumer Expenditure Survey data illustrates the practical gains from incorporating callback information when making inference on inequality measures.\\ {\bf Keywords:} Callback data; EM algorithm; Inequality measures; Missing not at random; Nonignorable nonresponse; Semiparametric full-likelihood

Introduction

In economics and official statistics, inequality measures play a central role in summarizing the distributions of economic variables such as income, wealth, and expenditure Atkinson1970. Let $F(y)$ denote the cumulative distribution function (CDF) of a continuous random variable $Y$ supported on $(0,\infty)$, which may represent, for example, household income. In this paper, we focus on commonly used inequality measures expressed as $\theta(F)$, where $\theta(\cdot)$ is a statistical functional of $F(y)$. These measures include the Gini index, quantiles, the Theil index, the coefficient of variation, and members of the generalized entropy and Atkinson classes (see Section (ref) for their formal definitions). For a comprehensive review of inequality measures and their theoretical properties, see Cowell2011.

There is a vast literature on statistical inference for inequality measures; see Davidson2007,Giorgi2017,Dufour2019,Yuan2023,Zalghout2025 and the references therein. Much of this literature, however, implicitly assumes that the observed data are representative of the target population. This assumption is often violated in practice, as survey samples used to measure inequality frequently suffer from nonresponse or missing data; see Groves1998.

According to Rubin1976, valid inference for $\theta(F)$ must account for the underlying missing-data mechanism. If the response probability depends on the unobserved value of $Y$, the mechanism is said to be nonignorable; otherwise, it is ignorable. Nonignorable nonresponse mechanisms are common in surveys on income and other sensitive information Korinek2006,Korinek2007,Bollinger2013,Bollinger2019, as high-income households are often less likely to respond. Under such mechanisms, the observed data constitute a biased sample of the target population, because the income distribution of respondents may differ systematically from the true distribution $F(y)$ Qin2017. Moreover, $F(y)$ is not nonparametrically identifiable based solely on the biased sample Wang2014,Miao2016,Morikawa2021. Consequently, many standard inference procedures for $\theta(F)$ are either inapplicable or yield inefficient and potentially misleading results Paraje2010,Alvarez2021. In general, without additional assumptions or auxiliary information, valid statistical inference for $\theta(F)$ is not possible when only a sample subject to nonignorable nonresponse is available.

In survey practice, additional contact attempts are commonly made when the initial attempt fails. Such callback strategies are widely used to increase response rates and to monitor survey quality DrewFuller1980,Couper1998,Bates2000,Durrant2012,Kreuter2013,Olson2013. The resulting callback data, which record the number of contact attempts for initially nonresponding units, are routinely collected and provide valuable auxiliary information for identifying $F(y)$ and for modeling and adjusting for nonignorable nonresponse Alho1990,Qin2014,Miao2024.

Let $m \ge 2$ denote the maximum number of contact attempts, and let $D$ be the callback variable, where $D=j$ if a household responds at the $j$th call, $j=1,\ldots,m$, and $D=m+1$ if it never responds. Thus, $Y$ is observed when $D \le m$ and missing otherwise. Define $\pi_j(y)=\Pr(D=j \mid Y=y, D \ge j)$ as the response probability at the $j$th call given nonresponse to earlier calls. The functions $\pi_j(y)$ characterize the nonignorable missing-data mechanism. Because income distributions are typically skewed and may exhibit complex shapes, we adopt the semiparametric model of Alho1990, which leaves $F(y)$ unspecified while modeling each $\pi_j(y)$ parametrically via a logistic regression,

equation[equation omitted — 183 chars of source]

Here, $\boldsymbol q(\cdot)$ is a prespecified $d$-dimensional function, with parameters $\boldsymbol\alpha=(\alpha_1,\ldots,\alpha_m)^\top$ and $\boldsymbol\beta$. When $\boldsymbol q(y)=y$, this reduces to Alho's original model, while allowing a flexible choice of $\boldsymbol q(y)$ improves adaptability. In this formulation, $\alpha_j$ captures call-specific response heterogeneity, and $\boldsymbol\beta$ is common across calls under the continuum of resistance assumption, which ensures identifiability of the semiparametric model Guan2018,Miao2024.

Under model (ref), Alho1990 proposed a conditional likelihood approach to estimate $\boldsymbol\alpha$ and $\boldsymbol\beta$, leading to an inverse probability weighting (IPW) estimator of the population mean. Building on the same model, Qin2014 developed a semiparametric full-likelihood method that estimates $\boldsymbol\alpha$, $\boldsymbol\beta$, and the population mean. Their approach yields efficient inference that is robust to assumptions on $F(y)$ and avoids the instability of IPW estimators Li2023, particularly when the estimated response probabilities are small. Recent extensions using callback data include generalized method of moments estimation of $\boldsymbol\alpha$, $\boldsymbol\beta$, and the finite population mean Kim2014, sensitivity analysis Daniels2015, and the incorporation of covariate information Guan2018,Chen2018.

To the best of our knowledge, no existing methods use callback data to develop valid inference procedures for inequality measures $\theta(F)$ under nonignorable nonresponse. Existing results based on callback data focus primarily on inference for the mean of $Y$ and do not extend to more general measures $\theta(F)$, which may involve complex linear, nonlinear, or nonsmooth functionals of $F$, such as quantiles or the Gini index. Unlike inference for the mean, inference for these measures is substantially more challenging, because establishing the theoretical properties of the associated estimators requires delicate asymptotic arguments. Moreover, existing methods rely on complex numerical optimization, and no efficient algorithm with theoretical guarantees is currently available.

In this paper, we develop efficient and reliable inference procedures for general inequality measures $\theta(F)$ under nonignorable nonresponse by incorporating callback data within the semiparametric model of Alho1990. Our contributions are summarized as follows.

itemize• We propose semiparametric maximum full-likelihood estimators for inequality measures $\theta(F)$ that adjust for sample selection bias arising from nonignorable nonresponse. • We establish the asymptotic properties of the proposed estimators and derive explicit analytical expressions for their asymptotic variances. In particular, we use the Bahadur representation to analyze estimators of quantiles and their functions, and apply the theory of U- and V-statistics to study the Gini index estimator. These results lead to valid Wald-type confidence intervals for a broad class of inequality measures. • We develop a novel and stable expectation-maximization (EM) algorithm for numerical implementation. The conditional expectations in the E-step admit closed-form expressions, and the constrained maximization in the M-step reduces to fitting a logistic regression model using standard software. We also establish the monotonicity property of the algorithm. • Through simulation studies and a real-data application, we demonstrate the favorable finite-sample performance of the proposed methods.

The paper is organized as follows. Section (ref) proposes full-likelihood estimators for the model parameters and $\theta(F)$, establishes their asymptotic properties, and constructs valid confidence intervals for various inequality measures. Section (ref) presents an EM algorithm for practical implementation. Section (ref) reports the results of simulation studies, and Section (ref) analyzes a real income data set from the Consumer Expenditure Survey. Section (ref) concludes with some remarks. All proofs are provided in the Supplementary Material.

Semiparametric estimation and inference of inequality measures

Semiparametric full-likelihood approach

Suppose we observe a random sample $\{(Y_i, D_i): i=1,\ldots,N\}$ generated from Alho's semiparametric model (ref), where $F(y)$ is left unspecified. Let $n = \sum_{i=1}^N I(D_i \le m)$ denote the random number of respondents. Without loss of generality, we index the first $n$ observations as respondents and the remaining $N-n$ as nonrespondents, for whom the $Y_i$'s are unobserved. Hereafter, we refer to the first $n$ observations as the complete-case (CC) data.

Let

equation*[equation* omitted — 460 chars of source]

Then \[ \eta = \Pr(D \le m) = \int_0^\infty \rho(y;\boldsymbol\alpha,\boldsymbol\beta)\, dF(y) \] denotes the marginal probability that a generic household responds within $m$ callbacks. Denote the full parameter vector as $\boldsymbol{\varphi} = (\boldsymbol\alpha,\boldsymbol\beta,\eta,F)$.

We now derive the full likelihood for $\boldsymbol{\varphi}$. It factors into three components: (i) the likelihood of $D_1,\ldots,D_n$ conditional on $Y_1,\ldots,Y_n$ and on the event that the $n$ households respond within $m$ callbacks; (ii) the likelihood of $Y_1,\ldots,Y_n$ given that these $n$ households respond within $m$ callbacks; and (iii) the likelihood of the total number of respondents $n$.

First, given that the $i$th household responded within $m$ callbacks and provided the observed response $Y_i$, the conditional probability of responding on the $D_i$th callback is

align*[align* omitted — 330 chars of source]

Hence, the likelihood contribution from $D_1,\ldots,D_n$ conditional on $Y_1,\ldots,Y_n$ and on the event that the $n$ households responded within $m$ callbacks, also known as the conditional likelihood Alho1990, is

equation[equation omitted — 185 chars of source]

Second, given that the $i$th household responded within $m$ callbacks, the conditional probability of observing $Y_i$ is

equation[equation omitted — 198 chars of source]

Therefore, the likelihood contribution from $Y_1,\ldots,Y_n$ given that the $n$ households responded within $m$ callbacks is

equation[equation omitted — 123 chars of source]

Third, the number of respondents $n$ follows a $\mathrm{Binomial}(N,\eta)$ distribution, with likelihood

equation[equation omitted — 75 chars of source]

Combining (ref) and (ref)--(ref) yields the full likelihood of $\boldsymbol{\varphi}$ as

equation*[equation* omitted — 188 chars of source]

Following the empirical likelihood principle Owen2001, we model $F$ as \[ F(y) = \sum_{i=1}^{n} p_i I(Y_i \le y), \] where $p_i = dF(Y_i)$ for $i=1,\ldots,n$. Up to an additive constant that does not depend on $\boldsymbol{\varphi}$, the semiparametric full log-likelihood function is

equation[equation omitted — 228 chars of source]

where $\boldsymbol{\varphi}$ is subject to the constraints

equation[equation omitted — 227 chars of source]

The first two constraints ensure that $F$ is a valid CDF, while the last follows from the definition of $\eta$ and accounts for sample selection bias induced by nonignorable nonresponse.

The semiparametric maximum full-likelihood estimator of $\boldsymbol{\varphi}$ is defined as

equation[equation omitted — 165 chars of source]

Given $\hat{\boldsymbol{\varphi}}$, the corresponding semiparametric maximum full-likelihood estimator of $F(y)$ is

equation[equation omitted — 110 chars of source]

Since a closed-form expression for $\hat{\boldsymbol{\varphi}}$ is generally unavailable, we rely on numerical methods. In Section (ref), we develop a stable and computationally convenient EM algorithm to compute $\hat{\boldsymbol{\varphi}}$ and $\hat{F}(y)$. The resulting estimator $\hat{F}(y)$ provides the foundation for correcting sample selection bias and for estimating the inequality measures $\theta(F)$ in the presence of nonignorable nonresponse.

Semiparametric estimation of inequality measures

In this section, we propose a semiparametric approach for estimating inequality measures. Recall that the measures of interest, $\theta(F)$, can be expressed as statistical functionals of the income distribution $F$. Given the estimator $\hat{F}(y)$, a natural strategy is to adopt the plug-in principle by replacing $F(y)$ with $\hat{F}(y)$, which yields $\theta(\hat{F})$ as the semiparametric estimator of $\theta(F)$. This differs fundamentally from the IPW-type approach. For simplicity, we denote the resulting semiparametric full-likelihood estimator by $\hat{\theta} = \theta(\hat{F})$. We focus on three widely used classes of inequality measures, each corresponding to a specific functional form of $\theta(\cdot)$.

Quantiles and their functions

In applications, population quantiles and their functions are widely recognized as important measures of inequality Prendergast2017,Zalghout2025. Common examples include the median, quartiles, quintiles, percentiles, and the interquartile range of income. Moreover, following the implementation of new policies, trends in quantiles, as well as their differences or ratios, provide valuable information for monitoring changes in the income distribution and assessing policy impacts. In this section, we focus on the estimation of quantiles and their functions, which are nonlinear and nonsmooth functionals of $F$.

For $\tau \in (0,1)$, the $100\tau$th quantile of $F(y)$ is defined as

equation*[equation* omitted — 124 chars of source]

Replacing $F(y)$ with $\hat{F}(y)$ in (ref), the proposed semiparametric estimator of $\theta_{\mathrm{quan}}^{\tau}(F)$ is given by

equation*[equation* omitted — 180 chars of source]

If interest lies in functions of quantiles, such as the quantile difference $\theta_{\mathrm{quan}}^{\tau_1} - \theta_{\mathrm{quan}}^{\tau_2}$ or the quantile ratio $\theta_{\mathrm{quan}}^{\tau_1}/\theta_{\mathrm{quan}}^{\tau_2}$ for $\tau_1,\tau_2 \in (0,1)$, the corresponding semiparametric estimators are $\hat{\theta}_{\mathrm{quan}}^{\tau_1} - \hat{\theta}_{\mathrm{quan}}^{\tau_2}$ and $\hat{\theta}_{\mathrm{quan}}^{\tau_1}/\hat{\theta}_{\mathrm{quan}}^{\tau_2}$, respectively.

A class of linear functionals and its functions

We next consider an important class of inequality measures that can be expressed as linear functionals of $F$ and their transformations. A $p$-dimensional linear functional of $F$ is defined as

equation[equation omitted — 137 chars of source]

for suitable choices of $\boldsymbol{u}(y) = \big(u_1(y),\ldots,u_p(y)\big)^\top$. Many commonly used inequality measures can be written in the unified form

equation[equation omitted — 130 chars of source]

where $h(\cdot): \mathbb{R}^p \to \mathbb{R}$ is a smooth function of $\boldsymbol\gamma_{\boldsymbol{u}}$. Examples in this class include population moments (centered or uncentered), the Theil index, the coefficient of variation, the mean logarithmic deviation, and members of the generalized entropy and Atkinson classes. Table (ref) summarizes these measures along with the corresponding specifications of $\boldsymbol{u}(\cdot)$ and $h(\cdot)$.

table[table omitted — 1,000 chars of source]

In our numerical studies, we focus on an important and widely used inequality measure, the Theil index Theil1967, defined as

equation*[equation* omitted — 186 chars of source]

It is immediate that $\theta_{\mathrm{Theil}}(F)$ is a special case of the class $\theta_{\boldsymbol{u},h}(F)$ with $\boldsymbol{u}(y) = \big(y, y\log(y)\big)^\top$ and $h(x_1,x_2)=x_2/x_1-\log(x_1)$. Moreover, $\theta_{\mathrm{Theil}}(F)$ corresponds to the limiting case of the generalized entropy class as the governing parameter $\zeta \to 1$. A useful property of the Theil index, and more generally of functionals of the form (ref), is their additive decomposability by subgroup Cowell2011, which facilitates the analysis of within-group and between-group inequality.

By definition (ref), $\theta_{\boldsymbol{u},h}(F)$ can be estimated by replacing $F(y)$ with $\hat{F}(y)$ in (ref). The proposed semiparametric estimator for this class of inequality measures is \[ \hat{\theta}_{\boldsymbol{u},h} = \theta_{\boldsymbol{u},h}(\hat{F}) = h(\hat{\boldsymbol\gamma}_{\boldsymbol{u}}), \quad \text{with} \quad \hat{\boldsymbol\gamma}_{\boldsymbol{u}} = \sum_{i=1}^n \hat{p}_i \boldsymbol{u}(Y_i). \] In particular, the proposed estimator of the Theil index is

equation*[equation* omitted — 199 chars of source]

Gini index

Although $\theta_{\mathrm{quan}}^{\tau}(F)$ and $\theta_{\boldsymbol{u},h}(F)$ cover a wide range of inequality measures, the Gini index Gini1912 is a notable exception that cannot be expressed within these classes. In this section, we therefore consider the Gini index through a specific functional form of $\theta(\cdot)$. The Gini index has several equivalent representations; here, we adopt the following formulation Yuan2023:

equation[equation omitted — 84 chars of source]

where $\psi(F) = \int_0^\infty 2y F(y)\, dF(y)$ and $\mu(F) = \int_0^\infty y\, dF(y)$. Unlike the measures considered previously, $\theta_{\mathrm{Gini}}(F)$ is distinct in that $\psi(F)$ involves a U-statistic-type functional of $F$. The Gini index is among the most widely used measures of inequality and is routinely reported by major official statistical agencies worldwide. Its widespread use is also attributable to its close relationship with the Lorenz curve Lorenz1905, which provides an intuitive graphical interpretation.

Using $\hat{F}(y)$ in (ref), the proposed semiparametric estimator of $\theta_{\mathrm{Gini}}(F)$ is $\hat{\theta}_{\mathrm{Gini}} = \theta_{\mathrm{Gini}}(\hat{F})$, which can be written explicitly as

equation[equation omitted — 213 chars of source]

where $\hat{\mu} = \mu(\hat{F}) = \sum_{i=1}^{n} \hat{p}_i Y_i$ and $\hat{\psi} = \psi(\hat{F}) = 2\sum_{i=1}^{n} \hat{p}_i Y_i \hat{F}(Y_i)$. These quantities are introduced for notational convenience in subsequent theoretical developments.

Asymptotic properties

In this section, we investigate the asymptotic behavior of the estimators of inequality measures proposed in Section (ref). Although the plug-in principle provides a natural approach for estimating inequality measures $\theta(F)$, establishing the asymptotic properties of the resulting estimators for statistical inference is nontrivial and requires a careful combination of appropriate asymptotic tools.

To facilitate the presentation, we introduce additional notation. Let $(\boldsymbol\alpha^*, \boldsymbol\beta^*, \eta^*)$ denote the true value of the model parameters $(\boldsymbol\alpha, \boldsymbol\beta, \eta)$. Define $\rho^*(y) = \rho(y; \boldsymbol\alpha^*, \boldsymbol\beta^*)$ and $\boldsymbol{\pi}^*(y) = \big(\pi(y;\alpha_1^*, \boldsymbol\beta^*), \ldots, \pi(y;\alpha_m^*, \boldsymbol\beta^*)\big)^\top$. Let $\mbox{\bf 1}_k$ and $\mbox{\bf 0}_k$ denote $k \times 1$ vectors of ones and zeros, respectively, and let $\mbox{\bf I}_k$ denote the $k \times k$ identity matrix. We further define \[ \boldsymbol{A}(y) = \big(\mbox{\bf I}_m,\ \mbox{\bf 1}_m \boldsymbol q(y)^\top\big)^\top \quad \text{and} \quad \boldsymbol{v}(y) = \rho^*(y)^{-1} \big(\{\rho^*(y) - 1\}\boldsymbol{\pi}^*(y)^\top \boldsymbol{A}(y)^\top,\ 1,\ (\eta^*)^2\big)^\top. \]

We begin by studying the asymptotic behavior of the proposed semiparametric estimator $\hat{F}(y)$ of $F(y)$.

theoremSuppose that Conditions C1--C5 in the Appendix are satisfied. Under model (ref), as $N \to \infty$, \( N^{1/2}\big\{\hat{F}(y) - F(y)\big\} \) converges weakly to a zero-mean tight Gaussian process with continuous sample paths and covariance function \begin{equation*} \begin{aligned} \Sigma(y_1, y_2) &= \mathbb{E}\!\left\{\frac{\xi_F(Y, y_1)\xi_F(Y, y_2)}{\rho^*(Y)}\right\} + \mathbb{E}\{\xi_F(Y, y_1)\boldsymbol{v}(Y)\}^{\top} \boldsymbol{\Gamma} \mathbb{E}\{\xi_F(Y, y_2)\boldsymbol{v}(Y)\}, \end{aligned} \end{equation*} for any $y_1, y_2 \in (0,\infty)$, where $\xi_F(Y,y) = I(Y \le y) - F(y)$ and $\boldsymbol{\Gamma}$ is defined in (ref) in the Appendix.

The result of Theorem (ref) provides the foundation for deriving the asymptotic properties of the proposed estimators of inequality measures $\hat{\theta}_{\mathrm{quan}}^{\tau}$, $\hat{\theta}_{\boldsymbol{u},h}$, and $\hat{\theta}_{\mathrm{Gini}}$ in the subsequent theorems.

Asymptotic properties of $\hat{\theta}_\text{quan}^{\tau}$

Building on Theorem (ref), we first apply the von Mises calculus and the functional delta method to derive a Bahadur representation for $\hat{\theta}_{\mathrm{quan}}^{\tau}$ Bahadur1966. This representation provides a convenient and tractable route for establishing the asymptotic properties of $\hat{\theta}_{\mathrm{quan}}^{\tau}$. Let $\theta_{\mathrm{quan}}^{\tau *}$ denote the true value of $\theta_{\mathrm{quan}}^{\tau}(F)$.

lemmaAssume that the conditions of Theorem (ref) are satisfied. For a given $\tau \in (0,1)$, the estimator $\hat{\theta}_{\mathrm{quan}}^{\tau}$ admits the following Bahadur representation: \begin{equation*} \begin{aligned} \hat{\theta}_{\mathrm{quan}}^{\tau} - \theta_{\mathrm{quan}}^{\tau *} = \frac{F(\theta_{\mathrm{quan}}^{\tau *}) - \hat{F}(\theta_{\mathrm{quan}}^{\tau *})} {f(\theta_{\mathrm{quan}}^{\tau *})} + o_p(N^{-1/2}), \end{aligned} \end{equation*} where $f(y)$ denotes the probability density function of $Y$, which is assumed to be positive and continuously differentiable in a neighborhood of $\theta_{\mathrm{quan}}^{\tau *}$.

Combining the Bahadur representation of $\hat{\theta}_{\mathrm{quan}}^{\tau}$ in Lemma (ref) with the result of Theorem (ref), we establish the joint asymptotic normality of $\hat{\theta}_{\mathrm{quan}}^{\tau_1}$ and $\hat{\theta}_{\mathrm{quan}}^{\tau_2}$ for any $\tau_1, \tau_2 \in (0,1)$ in the following theorem. Throughout, we use “$\overset{d}{\to}$” to denote convergence in distribution.

theoremAssume that the conditions of Lemma (ref) hold at levels $\tau_1, \tau_2 \in (0,1)$, with the corresponding true quantiles $\theta_{\mathrm{quan}}^{\tau_1 *}$ and $\theta_{\mathrm{quan}}^{\tau_2 *}$. Under model (ref), as $N \to \infty$, \[ N^{1/2} \big( \hat{\theta}_{\mathrm{quan}}^{\tau_1} - \theta_{\mathrm{quan}}^{\tau_1 *}, \hat{\theta}_{\mathrm{quan}}^{\tau_2} - \theta_{\mathrm{quan}}^{\tau_2 *} \big)^\top \overset{d}{\to} \mathcal{N}\!\left(\mbox{\bf 0},\; \boldsymbol{\Omega}(\tau_1,\tau_2)\right), \] where \[ \boldsymbol{\Omega}(\tau_1,\tau_2) = \boldsymbol{D}^{-1} \boldsymbol{\Sigma}\!\left(\theta_{\mathrm{quan}}^{\tau_1 *}, \theta_{\mathrm{quan}}^{\tau_2 *}\right) \boldsymbol{D}^{-1}, \] with \[ \boldsymbol{D} = \begin{pmatrix} f(\theta_{\mathrm{quan}}^{\tau_1 *}) & 0 \\ 0 & f(\theta_{\mathrm{quan}}^{\tau_2 *}) \end{pmatrix} \quad \text{and} \quad \boldsymbol{\Sigma}(y_1,y_2) = \begin{pmatrix} \Sigma(y_1,y_1) & \Sigma(y_1,y_2) \\ \Sigma(y_1,y_2) & \Sigma(y_2,y_2) \end{pmatrix}. \]

Theorem (ref) also implies the marginal asymptotic normality of $\hat{\theta}_{\mathrm{quan}}^{\tau}$. In studies of income inequality, researchers are often interested in smooth functions of quantiles at different levels, such as $\hat{\theta}_{\mathrm{quan}}^{\tau_1} - \hat{\theta}_{\mathrm{quan}}^{\tau_2}$ and $\hat{\theta}_{\mathrm{quan}}^{\tau_1} / \hat{\theta}_{\mathrm{quan}}^{\tau_2}$. Using the result of Theorem (ref), the asymptotic distributions of such smooth functions can be conveniently derived via the delta method, as stated in the following corollary.

corollaryAssume that the conditions of Theorem (ref) hold for $\theta_{\mathrm{quan}}^{\tau_1 *}$ and $\theta_{\mathrm{quan}}^{\tau_2 *}$. Let $g(\cdot,\cdot)$ be a twice continuously differentiable function. Then, as $N \to \infty$, \[ N^{1/2} \big\{ g(\hat{\theta}_{\mathrm{quan}}^{\tau_1}, \hat{\theta}_{\mathrm{quan}}^{\tau_2}) - g(\theta_{\mathrm{quan}}^{\tau_1 *}, \theta_{\mathrm{quan}}^{\tau_2 *}) \big\} \overset{d}{\to} \mathcal{N}\!\left(0,\; \sigma_{\mathrm{quan}}^2(\tau_1,\tau_2)\right), \] where \[ \sigma_{\mathrm{quan}}^2(\tau_1,\tau_2) = \boldsymbol{b}_g^\top \boldsymbol{\Omega}(\tau_1,\tau_2) \boldsymbol{b}_g, \quad \text{with} \quad \boldsymbol{b}_g = \left( \frac{\partial g (\theta_{\mathrm{quan}}^{\tau_1 *},\, \theta_{\mathrm{quan}}^{\tau_2 *})}{\partial \theta_{\mathrm{quan}}^{\tau_1}}, \frac{\partial g (\theta_{\mathrm{quan}}^{\tau_1 *},\, \theta_{\mathrm{quan}}^{\tau_2 *})}{\partial \theta_{\mathrm{quan}}^{\tau_2}} \right)^\top. \]

Asymptotic properties of $\hat\theta_{\boldsymbol{u},h}$

We next study the asymptotic properties of the proposed estimators of inequality measures in the class $\hat{\theta}_{\boldsymbol{u},h} = h(\hat{\boldsymbol\gamma}_{\boldsymbol{u}})$ for smooth functions $h$, which includes the Theil index as a special case. Let $\boldsymbol\gamma_{\boldsymbol{u}}^*$ and $\theta_{\boldsymbol{u},h}^*$ denote the true values of $\boldsymbol\gamma_{\boldsymbol{u}}$ and $\theta_{\boldsymbol{u},h}(F)$ defined in (ref) and (ref), respectively.

theoremSuppose that Conditions C1--C5 in the Appendix are satisfied and that $\int \rho^*(y)^{-1}\boldsymbol{u}(y)\boldsymbol{u}(y)^\top \, dF(y)$ is finite. Under model (ref), as $N \to \infty$, \begin{equation*} \begin{aligned} N^{1/2}\big(\hat{\theta}_{\boldsymbol{u},h} - \theta_{\boldsymbol{u},h}^*\big) \overset{d}{\to} \mathcal{N}\big(0, \sigma_{\boldsymbol{u},h}^2\big), \end{aligned} \end{equation*} where $\sigma_{\boldsymbol{u},h}^2 = \boldsymbol{b}_h^\top \boldsymbol{\Sigma}_{\boldsymbol{u}} \boldsymbol{b}_h$, with $\boldsymbol{b}_h = \partial h(\boldsymbol\gamma_{\boldsymbol{u}}^*) / \partial \boldsymbol\gamma_{\boldsymbol{u}}$, \[ \boldsymbol{\Sigma}_{\boldsymbol{u}} = \mathbb{E}\!\left\{\frac{\boldsymbol{\xi}_{\boldsymbol{u}}(Y)\boldsymbol{\xi}_{\boldsymbol{u}}^\top(Y)}{\rho^*(Y)}\right\} + \mathbb{E}\!\big\{\boldsymbol{\xi}_{\boldsymbol{u}}(Y)\boldsymbol{v}^\top(Y)\big\} \boldsymbol{\Gamma} \mathbb{E}\!\big\{\boldsymbol{v}(Y)\boldsymbol{\xi}_{\boldsymbol{u}}^\top(Y)\big\}, \] and $\boldsymbol{\xi}_{\boldsymbol{u}}(Y) = \boldsymbol{u}(Y) - \boldsymbol\gamma_{\boldsymbol{u}}^*$.

Theorem (ref) immediately implies the asymptotic normality of the proposed estimator $\hat{\theta}_{\mathrm{Theil}}$. Let $\theta_{\mathrm{Theil}}^*$ denote the true value of the Theil index $\theta_{\mathrm{Theil}}(F)$.

corollaryAssume that the conditions of Theorem (ref) hold for $\boldsymbol{u}(y) = \big(y, y\log(y)\big)^\top$. Then, as $N \to \infty$, \[ N^{1/2}\big(\hat{\theta}_{\mathrm{Theil}} - \theta_{\mathrm{Theil}}^*\big) \overset{d}{\to} \mathcal{N}\big(0, \sigma_{\mathrm{Theil}}^2\big), \] where $\sigma_{\mathrm{Theil}}^2 = \boldsymbol{b}_T^\top \boldsymbol{\Sigma}_{\boldsymbol{u}} \boldsymbol{b}_T$ with \[ \boldsymbol{b}_T = \frac{1}{\gamma_1^*} \big(-\gamma_2^*/\gamma_1^* - 1,\; 1\big)^\top, \quad \gamma_1^* = \int_0^\infty y\, dF(y), \quad \gamma_2^* = \int_0^\infty y\log(y)\, dF(y). \]

Asymptotic properties of $\hat{\theta}_\text{Gini}$

We now investigate the asymptotic properties of the proposed estimator of the Gini index, $\hat{\theta}_{\mathrm{Gini}}$, which requires additional asymptotic tools beyond those used in the previous section. Recall the definition of $\hat{\theta}_{\mathrm{Gini}}$ in (ref), where the numerator can be written as \[ \hat{\psi} = 2 \sum_{i=1}^{n} \hat{p}_i Y_i \hat{F}(Y_i) = \sum_{i=1}^{n} \sum_{j=1}^{n} 2 \hat{p}_i \hat{p}_j Y_i I(Y_j \le Y_i). \] Unlike $\hat{\boldsymbol\gamma}_{\boldsymbol{u}}$, the statistic $\hat{\psi}$ does not admit a linear representation, as it involves pairs of observations. As a result, Theorem (ref) is not directly applicable for deriving the asymptotic distribution of $\hat{\theta}_{\mathrm{Gini}}$. Instead, noting that $\hat{\psi}$ has the structure of a second-order V-statistic, we invoke asymptotic results from the theory of U- and V-statistics to establish the asymptotic normality of $\hat{\theta}_{\mathrm{Gini}}$. For notational convenience, let $\psi^*$, $\mu^*$, and $\theta_{\mathrm{Gini}}^*$ denote the true values of $\psi(F)$, $\mu(F)$, and $\theta_{\mathrm{Gini}}(F)$ defined in (ref), respectively.

theoremSuppose that Conditions C1--C5 in the Appendix are satisfied and that \( \int y^2 \rho^*(y)^{-1}\, dF(y) < \infty. \) Under model (ref), as $N \to \infty$, \[ N^{1/2}\big(\hat{\theta}_{\mathrm{Gini}} - \theta_{\mathrm{Gini}}^*\big) \overset{d}{\to} \mathcal{N}\big(0, \sigma_{\mathrm{Gini}}^2\big), \] where $\sigma_{\mathrm{Gini}}^2 = \boldsymbol{b}_G^\top \boldsymbol{\Sigma}_G \boldsymbol{b}_G$, with $\boldsymbol{b}_G = (\mu^*)^{-1}\big(-\psi^*(\mu^*)^{-1},\, 1\big)^\top$, \[ \boldsymbol{\Sigma}_G = \mathbb{E}\!\left\{\frac{\boldsymbol{\xi}_G(Y)\boldsymbol{\xi}_G(Y)^\top}{\rho^*(Y)}\right\} + \mathbb{E}\!\big\{\boldsymbol{\xi}_G(Y)\boldsymbol{v}^\top(Y)\big\} \boldsymbol{\Gamma} \mathbb{E}\!\big\{\boldsymbol{v}(Y)\boldsymbol{\xi}_G^\top(Y)\big\}, \] and \[ \boldsymbol{\xi}_G(Y) = \left( Y - \mu^*,\; 2\{YF(Y) + \int_Y^\infty t\, dF(t)\} - 2\psi^* \right)^\top. \]

Semiparametric inference of inequality measures

In this section, we focus on constructing confidence intervals for the inequality measures under consideration. Based on the asymptotic normality results established in Section (ref), it is natural to construct Wald-type confidence intervals. To estimate the asymptotic variances of the proposed estimators, we replace the unknown parameters with their consistent estimators. Specifically, we estimate $(\boldsymbol\alpha^*, \boldsymbol\beta^*, \eta^*, F)$ by $(\hat{\boldsymbol\alpha}, \hat{\boldsymbol\beta}, \hat{\eta}, \hat{F})$, and estimate the density function $f(y)$ of $Y$ by the procedure as described below.

Since $Y$ is supported on $(0,\infty)$, we first estimate the density of the transformed variable $T=\log(Y)$, denoted by $k(t)$. Given $\hat{F}(y)$ in (ref), $k(t)$ is estimated by \[ \hat{k}(t) = \int K_b\big(t - \log(y)\big)\, d\hat{F}(y), \] where $K_b(x) = K(x/b)/b$ and $K(\cdot)$ is the standard normal kernel. The bandwidth $b$ is selected according to \[ b = 1.06\, n^{-1/5} \min\!\left(\widehat{\mathrm{IQR}}/1.34,\; \hat{\sigma}\right), \] where $\widehat{\mathrm{IQR}}$ and $\hat{\sigma}^2$ denote the interquartile range and variance estimators of $\log(Y)$, respectively, computed using $\hat{F}(y)$ Silverman1986. These quantities can be obtained using the methods described in Sections (ref) and (ref). Finally, the density of $Y$ is estimated by \[ \hat{f}(y) = \hat{k}\big(\log(y)\big)/y. \]

Let $\hat{\boldsymbol{\Omega}}(\tau_1,\tau_2)$ denote the estimator of $\boldsymbol{\Omega}(\tau_1,\tau_2)$ obtained by replacing $(\boldsymbol\alpha^*, \boldsymbol\beta^*, \eta^*, F, f)$ with $(\hat{\boldsymbol\alpha}, \hat{\boldsymbol\beta}, \hat{\eta}, \hat{F}, \hat{f})$. Similarly, let $(\hat{\sigma}_{\boldsymbol{u},h}^2, \hat{\sigma}_{\mathrm{Theil}}^2, \hat{\sigma}_{\mathrm{Gini}}^2)$ denote the estimators of $(\sigma_{\boldsymbol{u},h}^2, \sigma_{\mathrm{Theil}}^2, \sigma_{\mathrm{Gini}}^2)$ obtained by replacing $(\boldsymbol\alpha^*, \boldsymbol\beta^*, \eta^*, F)$ with $(\hat{\boldsymbol\alpha}, \hat{\boldsymbol\beta}, \hat{\eta}, \hat{F})$. These variance estimators can be shown to be consistent. By Slutsky's theorem, Wald-type confidence intervals can therefore be constructed for $\theta_{\mathrm{quan}}^{\tau}(F)$, $g(\theta_{\mathrm{quan}}^{\tau_1}, \theta_{\mathrm{quan}}^{\tau_2})$, $\theta_{\boldsymbol{u},h}(F)$, $\theta_{\mathrm{Theil}}(F)$, and $\theta_{\mathrm{Gini}}(F)$.

For example, a $100(1-\alpha)\%$ Wald-type confidence interval for $\theta_{\mathrm{quan}}^{\tau}(F)$ is given by \[ \left[ \hat{\theta}_{\mathrm{quan}}^{\tau} - \frac{z_{1-\alpha/2}\, \hat{\sigma}_{11}(\tau)}{\sqrt{N}}, \; \hat{\theta}_{\mathrm{quan}}^{\tau} + \frac{z_{1-\alpha/2}\, \hat{\sigma}_{11}(\tau)}{\sqrt{N}} \right], \] where $\hat{\sigma}_{11}(\tau)$ is the $(1,1)$ element of $\hat{\boldsymbol{\Omega}}(\tau,\tau)$, representing the estimated asymptotic standard deviation of $\hat{\theta}_{\mathrm{quan}}^{\tau}$, and $z_{1-\alpha/2}$ denotes the $100(1-\alpha/2)$th percentile of the standard normal distribution. Confidence intervals for $g(\theta_{\mathrm{quan}}^{\tau_1}, \theta_{\mathrm{quan}}^{\tau_2})$, $\theta_{\boldsymbol{u},h}(F)$, $\theta_{\mathrm{Theil}}(F)$, and $\theta_{\mathrm{Gini}}(F)$ can be constructed analogously. A simulation study evaluating the finite-sample performance of the proposed Wald-type confidence intervals for various inequality measures is presented in Section (ref).

Expectation-Maximization algorithm

Since the $Y_i$'s are unobserved for nonrespondents, the problem can be viewed as a special case of missing data. The EM algorithm thus provides a natural framework for numerically computing $\hat{\boldsymbol{\varphi}}$ defined in (ref).

Recall that the observed data consist of $\mathcal{O} = \{(Y_i, D_i): i=1,\ldots,n\} \cup \{D_k = m+1: k=n+1,\ldots,N\}$. Let $\mathcal{O}^* = \{(Z_k, D_k = m+1): k=n+1,\ldots,N\}$ denote the unobserved responses for the $N-n$ nonrespondents. If both $\mathcal{O}$ and $\mathcal{O}^*$ were available, the complete-data likelihood would take the form

equation*[equation* omitted — 240 chars of source]

Recall that $p_i = dF(Y_i)$ for $i=1,\ldots,n$, and that $F$ is modeled as $F(y) = \sum_{i=1}^{n} p_i I(Y_i \le y)$. Under this specification, each latent variable $Z_k$ can take values only in $\{Y_1,\ldots,Y_n\}$. The corresponding complete-data log-likelihood is therefore

eqnarray[eqnarray omitted — 335 chars of source]

subject to the constraint set $\mathcal{C}$ in (ref).

As in Dempster1977, each iteration of the EM algorithm consists of two steps, the E-step and the M-step. Suppose that $r$ iterations have been completed and the current parameter value is $\boldsymbol{\varphi}^{(r)}$. The $(r+1)$th iteration then proceeds as follows.

In the E-step, given the current parameter value $\boldsymbol{\varphi}^{(r)}$ and the observed data $\mathcal{O}$, we compute the conditional expectation of the complete-data log-likelihood

equation[equation omitted — 232 chars of source]

which is maximized in the subsequent M-step. To evaluate $Q(\boldsymbol{\varphi} \mid \boldsymbol{\varphi}^{(r)})$, it suffices to compute

equation[equation omitted — 310 chars of source]

where the final equality follows by an argument analogous to that leading to (ref). Combining (ref)--(ref), we can write \[ Q(\boldsymbol{\varphi} \mid \boldsymbol{\varphi}^{(r)}) = \ell_1^{(r)}(\boldsymbol\alpha,\boldsymbol\beta) + \ell_2^{(r)}(p_1,\ldots,p_n), \] where

equation*[equation* omitted — 355 chars of source]

and \[ w_i^{(r)} = (N-n)\, \frac{p_i^{(r)}\{1-\rho(Y_i;\boldsymbol\alpha^{(r)},\boldsymbol\beta^{(r)})\}} {1-\eta^{(r)}}. \]

In the M-step, we update the parameter vector from $\boldsymbol{\varphi}^{(r)}$ to $\boldsymbol{\varphi}^{(r+1)}$ by solving \[ \boldsymbol{\varphi}^{(r+1)} = \arg\max_{\boldsymbol{\varphi}} Q(\boldsymbol{\varphi} \mid \boldsymbol{\varphi}^{(r)}) \quad \text{subject to the constraints in \eqref{const}}. \] Owing to the additive structure of $Q(\boldsymbol{\varphi} \mid \boldsymbol{\varphi}^{(r)})$, the update $\boldsymbol{\varphi}^{(r+1)}$ can be obtained explicitly as

equation*[equation* omitted — 404 chars of source]

It is worth noting that the objective function $\ell_1^{(r)}(\boldsymbol\alpha,\boldsymbol\beta)$ is proportional to the weighted log-likelihood of a logistic regression model. Consequently, $(\boldsymbol\alpha^{(r+1)}, \boldsymbol\beta^{(r+1)})$ can be readily obtained by fitting a weighted logistic regression, which is supported by most standard statistical software. Additional implementation details are provided in Section 4 of the Supplementary Material.

We iterate the E-step and M-step until the increase in the log-likelihood $\ell(\boldsymbol{\varphi})$ in (ref) between successive iterations falls below a prespecified tolerance level, for example, $10^{-5}$. The following proposition establishes the monotonicity property of the proposed EM algorithm.

propositionFor the EM algorithm described above, we have, for all $r \ge 0$, \[ \ell\!\left(\boldsymbol{\varphi}^{(r+1)}\right) \ge \ell\!\left(\boldsymbol{\varphi}^{(r)}\right), \] that is, the log-likelihood function $\ell(\boldsymbol{\varphi})$ defined in (ref) is nondecreasing across successive EM iterations.

Based on extensive numerical experimentation, we find that the EM algorithm described above is not sensitive to the choice of initial values for $\boldsymbol{\varphi}$. Accordingly, in our numerical implementation, we initialize $\boldsymbol\alpha^{(0)} = \mbox{\bf 0}$, $\boldsymbol\beta^{(0)} = \mbox{\bf 0}$, and $p_i^{(0)} = 1/n$ for $i = 1,\ldots,n$, and set \[ \eta^{(0)} = \sum_{i=1}^n p_i^{(0)} \rho\!\left(Y_i; \boldsymbol\alpha^{(0)}, \boldsymbol\beta^{(0)}\right) = 1 - 2^{-m}. \]

Simulation study

In this section, we report simulation results assessing the finite-sample performance of the proposed semiparametric estimators and their associated confidence intervals for $\theta_{\text{quan}}^{\tau}(F)$ with $\tau = 0.25, 0.50,$ and $0.75$, $\theta_{\text{Theil}}(F)$ (as a representative member of the $\theta_{\mathbf{u},h}(F)$ class), and $\theta_{\text{Gini}}(F)$ under nonignorable nonresponse.

Three distributions are considered for $F(y)$: $\mathrm{Exp}(1)$, $\chi^2(1.5)$, and $\mathrm{Gam}(0.8,0.25)$, where $\mathrm{Exp}(a_1)$ denotes an exponential distribution with rate parameter $a_1$, $\chi^2(a_2)$ denotes a chi-squared distribution with $a_2$ degrees of freedom, and $\mathrm{Gam}(a_3,a_4)$ denotes a gamma distribution with shape parameter $a_3$ and rate parameter $a_4$. The callback indicator $D$ is generated according to Alho's callback model (ref) with $\boldsymbol q(y)=\log(y)$ and $m=2$. The parameter vector $(\boldsymbol\alpha^\top,\beta)$ is set to $(-1.5,\,0.5,\,-0.5)$, yielding moderate response rates across contacts. The negative value of $\beta$ implies that the response probability decreases as $Y$ increases.

The resulting true response probabilities for each contact attempt and the true values of the inequality measures are presented in Table (ref). Simulations are conducted for sample sizes $N = 1000$ and $2000$, with $M = 5000$ Monte Carlo replications for each design.

table[table omitted — 748 chars of source]

Performance of point estimation

We compare the finite-sample performance of the following three estimators for the inequality measures:

itemize{0pt} {0pt} {0pt} • CC: a naive estimator based on complete cases only; • Proposed: the proposed semiparametric full-likelihood estimator; • Ideal: a benchmark estimator computed using the full data.

The CC estimator discards all units with nonresponse or missing values and is therefore generally inconsistent under nonignorable nonresponse. The Ideal estimator is included solely as a benchmark and is not available in practice.

Finite-sample performance of a point estimator is evaluated using the relative bias (RB) and the root mean squared error (RMSE), defined as

equation*[equation* omitted — 186 chars of source]

where $\hat{x}^{(i)}$ denotes the estimate of a target parameter with true value $x$ in the $i$th replication $(i=1,\ldots,M)$. Simulation results are reported in Table (ref).

table[table omitted — 2,841 chars of source]

We summarize the main findings from Table (ref). First, as anticipated, the CC estimator performs poorly under nonignorable nonresponse, exhibiting substantial RB and RMSE across all simulation settings. Moreover, the CC estimator does not display consistency, as neither RB nor RMSE decreases when the sample size increases from $N=1000$ to $N=2000$. The CC estimator systematically underestimates the quantiles (negative RB) and overestimates the Theil and Gini indices (positive RB). Second, both the proposed semiparametric estimator and the Ideal estimator exhibit negligible RB across all settings. The RMSEs of the proposed estimator are comparable to those of the benchmark Ideal estimator and decrease as the sample size increases. Overall, the results indicate that the proposed semiparametric approach effectively corrects for bias induced by nonignorable nonresponse and achieves near-benchmark efficiency in finite samples.

Performance of confidence intervals

The finite-sample performance of confidence intervals is evaluated using the coverage probability (CP), reported as a percentage, and the average length (AL), defined as

equation*[equation* omitted — 165 chars of source]

where $\mathcal{I}^{(i)}$ denotes the confidence interval for a target parameter $x$ in the $i$th replication, and $|\mathcal{I}^{(i)}|$ denotes its length. Simulation results for the nominal 95% confidence intervals of the quantiles, the Theil index, and the Gini index described in Section (ref), for sample sizes $N=1000$ and $2000$, are reported in Table (ref).

table[table omitted — 1,320 chars of source]

We summarize the main findings from Table (ref). Overall, the confidence intervals constructed using the proposed semiparametric method achieve coverage probabilities close to the nominal 95% level across most settings, with moderate average lengths. Some undercoverage is observed for certain parameters, such as $\theta_{\text{Theil}}(F)$, when the sample size is $N=1000$. As the sample size increases to $N=2000$, coverage probabilities approach the nominal level and interval lengths become shorter. These results indicate that the proposed semiparametric inference procedure provides reliable finite-sample inference under nonignorable nonresponse with callback data.

A real data example

In this section, we apply the proposed estimation and inference procedure to data from the Consumer Expenditure Survey (CES), which is conducted by the United States Bureau of Labor Statistics to collect detailed information on household expenditures and related characteristics. A comprehensive description of the survey design is available at \url{https://www.census.gov/programs-surveys/ce.html}.

Our analysis uses the public-use adult data and associated paradata from the first quarter of the 2024 CES. For illustration, we consider the total amount of family resources after taxes over the previous 12 months (measured in \$10,000) as the study variable $Y$. The sample consists of 10,273 households with positive recorded values, with an overall nonresponse rate of 54.54%. The objective is to conduct efficient and reliable statistical inference for distributional and inequality measures, including quantiles, the Theil index, and the Gini index, using the CES data.

The maximum number of contact attempts recorded in the CES paradata is 18. Figure (ref) displays the response rates by number of contact attempts. The callback strategy is most effective during the initial contacts, with response rates declining as the number of calls increases. Specifically, 6.75% and 11.10% of households responded on the first and second calls, respectively, while an additional 27.61% responded in subsequent calls, yielding an overall response rate of 45.46%. Motivated by this pattern, we define four callback categories: $D=1$ and $D=2$ for households responding on the first and second calls, respectively; $D=3$ for those responding at later calls; and $D=4$ for households that did not respond.

figure[figure omitted — 268 chars of source]

To implement the proposed method, we use the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC) to select the functional form of $\boldsymbol q(y)$ in model (ref) from the candidate set $\{y, y^2, \log(y), \log^2(y)\}$ and their combinations. Based on the results in Table (ref), $\boldsymbol q(y)=\log(y)$ is selected, as it yields the smallest AIC and BIC values. The semiparametric full-likelihood estimates of the model parameters and their corresponding 95% confidence intervals are reported in Table (ref). In particular, the estimate of $\beta$ is $-0.191$, with a 95% confidence interval of $[-0.293,-0.090]$, indicating statistically significant evidence of nonignorable nonresponse in the CES data, with response propensity decreasing as the value of $Y$ increases.

table[table omitted — 1,014 chars of source]
table[table omitted — 524 chars of source]

The proposed semiparametric full-likelihood estimates of the quantiles, the Theil index, and the Gini index, together with their 95% confidence intervals, are reported in Table (ref). For comparison, the table also includes the corresponding estimates based on the 4{,}670 complete cases. The CC estimates differ markedly from the proposed semiparametric estimates, and all CC estimates fall outside the 95% confidence intervals constructed using the proposed method. This finding is consistent with the simulation results and indicates that the observed complete cases may not represent a random sample of the target distribution.

table[table omitted — 667 chars of source]

Figure (ref) compares the estimated distribution functions obtained from the proposed method and the CC method. A pronounced discrepancy between the two estimates is evident. Overall, these results suggest that ignoring nonignorable nonresponse can lead to substantial bias and unreliable inference for inequality measures, whereas the proposed semiparametric approach, which effectively incorporates callback information, yields more reliable estimates and valid inference for a range of commonly used inequality measures.

figure[figure omitted — 353 chars of source]

Concluding remarks

Nonignorable nonresponse has long been recognized as a major challenge in survey data analysis, yet its impact on inference for inequality measures remains relatively underexplored. In this paper, we develop a unified semiparametric framework for estimation and inference of inequality measures in the presence of nonignorable nonresponse. The proposed approach leverages callback data within a semiparametric full-likelihood formulation, leaving the underlying distribution $F(y)$ unspecified. We establish the asymptotic properties of the proposed estimators for a broad class of inequality measures, enabling valid inference under nonignorable nonresponse. From a practical perspective, we introduce an efficient and numerically stable EM algorithm with theoretical guarantees, facilitating straightforward implementation. Simulation studies and an empirical application illustrate the favorable finite-sample performance of the proposed estimation and inference procedures for commonly used inequality measures.

Several directions merit further investigation. First, in addition to nonignorable nonresponse, survey data often exhibit other forms of incompleteness, such as grouped reporting Gastwirth1972 and top-coding Feng2006. Extending the proposed framework to accommodate these features would be of practical interest. Second, it would be worthwhile to study inference for more general parameters defined through moment restrictions Qin1994 under nonignorable nonresponse. Finally, alternative sources of auxiliary information, such as regional response rates Korinek2006,Korinek2007, have been proposed as substitutes for callback data. Developing semiparametric inference theory under such settings represents another promising direction for future research.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Acknowledgments

C. Wang's work was supported by the Humanities and Social Sciences Foundation of the Ministry of Education of China (24YJA910005) and the National Natural Science Foundation of China (71988101); T. Yu's work is supported in part by the Singapore Ministry of Education Academic Research Tier 1 Fund (A-8000413-00-00); P. Li's work was supported by the Natural Sciences and Engineering Research Council of Canada (RGPIN-2020-0496).