Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
40,837 characters · 10 sections · 64 citation commands
Empirical likelihood and uniform convergence rates for dyadic kernel density estimation
\address[H. D. Chiang]{ Department of Economics, University of Wisconsin-Madison\\ William H. Sewell Social Science Building, 1180 Observatory Drive, Madison, WI 53706, USA.} \email{[email removed]}
\address[B. Y. Tan]{ Global Asia Institute, National University of Singapore \\ Block S17, Level 3, \#03-01 10 Lower Kent Ridge Road \\ Singapore 11907} \email{[email removed]}
Consider a weighted undirected random graph consisting of vertices $i\in\{1,...,n\}$ and edges consisting of real-valued random variables $(X_{ij})_{i,j\in\{1,...n\}, i< j}$. Suppose $X_{ij}$ are identically distributed with density $f(\cdot)=f_{X_{12}}(\cdot)$. The dyadic kernel density estimator (KDE) of GrahamNiuPowell2019, evaluated at a design point $x\in \mathrm{supp}(X_{12})$, is defined as
Applications of dyadic data are numerous in different fields of the sciences and social sciences, such as bilateral trade, animal migrations, refugee diasporas, transportation networks between different cities, friendship networks between individuals, intermediate product sales between firms, research and development partnerships across different organizations, electricity grids across different states, and social networks between legislators, to list a few. See, e.g. Graham2019Handbook for a recent and comprehensive review. This paper investigates the asymptotic properties of and proposes new, improved inference methods for dyadic KDE with both complete and incomplete data.\\ The main contributions of this paper are two-fold. First, we establish uniform convergence rates for dyadic KDE under general conditions. These rates are new; no uniform rate is previously known for KDE under dyadic sampling in the literature. Second, we develop inference methods by adapting a modified version of the jackknife empirical likelihood (JEL) proposed by MatsushitaOtsu2020. The proposed modified JEL (mJEL) procedure is shown to be asymptotically valid regardless of the presence of dyadic clustering. In practice, dyadic clustered data are often incomplete or contain missing values. We extend our modified JEL procedure to cover the practically relevant case of incomplete data under the missing at random assumption. Extensive simulation studies show that this modified JEL inference procedure does not suffer from the unreliable finite sample performance of the analytic variance estimator, delivers precise coverage probabilities even under modest sample sizes, and is robust against both dyadic clustering and i.i.d. sampling for both complete and incomplete dyadic data. Finally, we illustrate the method by studying the distribution of congestion across different flight routes using data from the United States Department of Transportation (US DOT).\\ In an important recent work, GrahamNiuPowell2019 first propose and examine a nonparametric density estimator for dyadic random variables. Focusing on pointwise asymptotic behaviors, they demonstrate that asymptotic normality at a fixed design point can be established at a parametric rate with respect to number of vertices, an unique feature of dyadic KDE. This implies that its asymptotics are robust to a range of bandwidth rates. They also point out that degeneracy that leads to a non-Gaussian limit, such as the situation pointed out in Menzel2020, does not occur for dyadic KDE under some mild bandwidth conditions. For inference, they propose an analytic variance estimator that corresponds to the one proposed by FafchampsGubert2007 (henceforth FG) under parametric models and establish its asymptotic validity. { In their simulation, they discover that the FG variance estimator often leads to imprecise coverage with moderate sample sizes. Our proposed modified JEL procedure for dyadic KDE complements the findings of GrahamNiuPowell2019 by providing an alternative that has improved finite sample performance.\\
} In practice, the presence of missing edges is an important feature of network datasets, and a researcher working with network data often observes incomplete sets of edges. To our knowledge, none of the existing theoretical work in the literature that focuses on the asymptotics of dyadic KDE has tackled this issue. Under the missing at random assumption, we extend our theory for modified JEL to cover this practically relevant situation. Ever since the seminal papers of Owen1988,Owen1990, empirical likelihood has been extensively studied in the literature. For a textbook treatment of classical theory, see Owen200Book. Some recent developments under conventional asymptotics include HjortMcKeagueVan_Keilegom2009AoS and BravoEscancianoVan_Keilegom2020AoS, and many more. See chen2009review for a recent review. Jackknife empirical likelihood was first proposed in the seminal work of JingYuanZhang2009 for parametric $U$-statistics and has since been studied by GongPengQi2010JMA, PengOiVan_Keilegom2012, Peng2012CanJS, WangPengQi2013Sinica, ZhangZhao2013JMA, MatsushitaOtsu2018JER and ChenTabri2020, to list a few, for different applications. Most of the aforementioned works focus on actual $U$-statistics under i.i.d. sampling. MatsushitaOtsu2020 demonstrate that JEL can not only be modified for certain unconventional asymptotics of i.i.d. observations, but can also be applied to study edge probabilities for both sparse and dense dyadic networks studied in BickelChenLevina2011. This is achieved by incorporating the bias-correction idea of Hinkley1978, EfronStein1981. In a context where controlling bias and spikes - from the presence of the bandwidth in kernel estimation - presents extra challenges, our modified JEL for dyadic KDE provides a nonparametric counterpart to the network asymptotic results in MatsushitaOtsu2020. \\ { Theoretically speaking, the modified JEL is closely related to the growing literature that utilizes the idea of accounting for the contributions from both linear and quadratic terms in the Hoeffding-type decomposition of the statistics. This idea emerges in the literature that studies “small-bandwidth asymptotics," in the context of density-weighted average derivative by ,CattaneoCrumpJansson2014bootstrap,CattaneoCrumpJansson2014ET and CattaneoCrumpJansson2013JASA, for the sake of robustness over a wider range of bandwidths. This idea has since been explored in various other contexts such as an inference problem in partially linear models with many regressors in cattaneo2018alternative,cattaneo2018inference, as well as the various other examples explored in MatsushitaOtsu2020.}\\ Our results also complement the uniform convergence rates results for kernel type estimators under different dependence settings studied in EinmahlMason2000, EinmahlMason2005, DonyMason2008, Hansen2008ET, Kristensen2009ET and so forth. Theoretically, the $\sqrt{n}$-uniform rate of dyadic KDE share common underlying features with the kernel density estimator of EscancianoJacho-Chavez2012, as well as the residual distribution estimator proposed in AkritasVan_Keilegom2001, both under i.i.d. sampling.\\
{ In a seminal work, FafchampsGubert2007 proposed a dyadic robust estimator for regression models, which has since become the benchmark for inference under dyadic clustering. Related works include FrankSnijders1994 and SnijdersBorgatti1999, AronowSamiiAssenova2015 CameronMiller2014 in different contexts. Asymptotics for statistics of dyadic random arrays are studied by DDG2020 under a general empirical processes setting and by ChiangKatoSasaki2020 in high-dimensions. The former show the validity of a modified pigeonhole bootstrap (cf. Mccullagh2000 and Owen2007) while the latter utilize a multiplier bootstrap, both under a non-degeneracy condition. On the other hand, for two-way separately exchangeable arrays and linear test-statistics, Menzel2020 proposed a conservative bootstrap that is valid uniformly even under potentially non-Gaussian degeneracy scenarios. In this paper, we only need to focus on the non-degenerate and Gaussian degenerate cases since the non-Gaussian degeneracy scenario is not a concern for dyadic KDE; GrahamNiuPowell2019 pointed out that non-Gaussian degeneracy can be ruled out under mild bandwidth conditions.\\ More recently, cattaneo2022 study uniform convergence rate and uniform inference for dyadic KDE. Their uniform convergence rate results further improved upon the result by allowing uniformity over the whole support for compactly supported data and enabling boundary-adaptive estimation. They further derive the minimax rate of uniform convergence for density estimation with dyadic data and show that the dyadic KDE is, under appropriate conditions, minimax-optimal. For inference, they obtain uniform distribution theory and provide bias-corrected $t$-statistics-based uniform confidence bands. Their methods differ from the ones we propose, which are based on jackknife empirical likelihood.}
Let $|A|$ be the cardinality of a finite set $A$. Denote the uniform distribution on $[0,1]$ as $U[0,1]$. For $g:\mathcal{S}\to \mathbb{R}$ and $\mathcal{X}\subseteq \mathcal{S}$, denote $\|g\|_{\mathcal{X}}=\sup_{x\in\mathcal{X}}|g(x)|$. Write $a_n\lesssim b_n$ if there exists a constant $C>0$ independent of $n$ such that $a_n\le Cb_n$. Throughout the paper, the asymptotics should be understood as taking $n\to \infty$.\\ The rest of this paper is organized as follows. In Section (ref), we introduce our model and the dyadic KDE estimator. Our main theoretical results of uniform convergence rates and the validity of different JEL statistics for both complete and incomplete dyadic random arrays are then introduced in Section (ref). Section (ref) contains the simulation studies while the empirical application can be found in Section (ref). A practical guideline on bandwidth choice and proofs of all the theoretical results are contained in the Supplementary Appendix.
Consider a weighted undirected random graph. Let $I_{n} = \{ (i,j) : 1 \le i < j \le n\}$ and $I_{\infty} = \bigcup_{n=2}^{\infty} I_{n}$. The random graph consists of an array $( X_{ij})_{(i,j) \in I_{\infty}}$ of real-valued random variables that is generated by
for some Borel measurable map $\mathfrak{f}:[0,1]^3\to \mathbb{R}$ that is symmetric in the first two arguments. Note that $U_i$ and $U_{\{\iota,\jmath\}}$ are independent for all $i\in\mathbb{N}$, $(\iota,\jmath)\in I_{\infty}$. While ((ref)) may seem to be a specific structural assumption, it is implied by the following low-level condition of joint exchangeability via the Aldous-Hoover-Kallenberg representation, see e.g. Kallenberg2006.
Suppose that the researcher observes $(X_{ij})_{(i,j)\in I_{n}}$ for $n\ge 2$. The object of interest is the density function $f(\cdot)=f_{X_{12}}(\cdot)$ of $X_{12}$. We assume such a density function exists and is well-defined although some low-level sufficient conditions can be obtained. For a fixed design point $x\in\mathrm{supp}(X_{12})$ of interest, the dyadic KDE for $\theta=f(x)$ is defined as
where $h_n>0$ is a $n$-dependent bandwidth, $K_{ij,n}:=h_n^{-1}K((x-X_{ij})/h_n)$ and $K(\cdot)$ is a kernel function. We will be more specific about the requirements for $h_n$ and $K$ in the following Section.
{ In the rest of the paper, denote $f_{X_{12}|U_1}$ for the conditional density function of $X_{12}$ given $U_1$ and $f_{X_{12}|U_1,U_2}$ for the conditional density function of $X_{12}$ given $U_1,U_2$. In addition, define the shorthand notations $f'(\cdot)=\partial f(\cdot)/\partial x$, $f''(\cdot)=\partial^2 f(\cdot)/\partial x^2$, $f_{X_{12}|U_1}'(\cdot|u)=\partial f_{X_{12}|U_1}(\cdot|u)/\partial x$, $f_{X_{12}|U_1}''(\cdot|u)=\partial^2 f_{X_{12}|U_1}(\cdot|u)/\partial x^2$. Similarly, $f_{X_{12}|U_1,U_2}'(\cdot|u_1,u_2)=\partial f_{X_{12}|U_1,U_2}(\cdot|u_1,u_2)/\partial x$, and $f_{X_{12}|U_1,U_2}''(\cdot|u_1,u_2)=\partial^2 f_{X_{12}|U_1,U_2}(\cdot|u_1,u_2)/\partial x^2$.} We first study the uniform convergence rates of dyadic KDE under the following assumptions.
{ The support condition in Assumption (ref) restricts the scope to be on a compact interval $\mathcal X$ that is strictly contained in the support of $X_{12}$. Assumption (ref)(i) requires the observations to be generated following the dyadic structure discussed in Section (ref). Assumption (ref)(ii) assumes existence of the conditional densities conditional on vertex-specific latent shocks and imposes standard smoothness conditions on the unknown density and conditional densities. Assumption (ref)(iii) assumes a smooth second-order kernel. Finally, Assumption (ref)(iv) requires the sequence of bandwidths to be converging to zero in an adequate range. }
A proof can be found in Section D.1 of the Supplementary Appendix.
We now introduce JEL-based inference procedures for dyadic KDE. Denote the leave-one-out and leave-two-out index sets $I_n^{(i)}=\{(\iota,\jmath)\in I_n: \iota,\jmath\not\in \{i\}\}$, $I_n^{(i,j)}=\{(\iota,\jmath)\in I_n: \iota,\jmath\not\in\{ i,j\}\}$. { For a given $\theta$, define
and subsequently $S(\theta)=\widehat \theta-\theta$, $S^{(i)}(\theta)=\widehat \theta^{(i)}-\theta$, and $S^{(i,j)}(\theta)=\widehat \theta^{(i,j)}-\theta$. } Furthermore, define the pseudo true value for JEL as
For modified JEL, define
For each value of $\theta$, define the pseudo true value for modified JEL by
Now, define the modified JEL function for $\theta$ by
and $\ell(\theta)$ is defined analogously with $V_i^m$ replaced by $V_i$. The Lagrangian dual problems are
{ We say $f$ is non-degenerate at $x$ if $\mathrm{Var}\left(f_{X_{12}|U_1}(x|U_1)\right)\ge\underline L>0$, which means that the vertex-specific shock $U_1$ affects the conditional density $f_{X_{12}|U_1}$. Before stating the next result, which characterizes the asymptotics of JEL and modified JEL for inference, we make the following assumptions.}
{ Assumption (ref) requires the design point $x$ to be an interior point of the support of $X_{12}$. Assumption (ref) (ii) imposes constraints on the convergence rate of the sequence of bandwidths. Note that the lower bound $nh_n^2\to \infty $ is a sufficient condition for Lemma 2 in the appendix, which is subsequently used for bounding the linearization errors for the JEL functions $\ell$ and $\ell^m$. This is specific to empirical likelihood-based methods and is not required by procedures such as the Wald test with FG variance estimator proposed in GrahamNiuPowell2019. The restriction $nh_n^{5/2}\to 0$ here is used only in the degenerate case and can be relaxed to $nh_n^4\to 0$ in the non-degenerate case. Note that the bandwidth conditions accommodate the MSE optimal rate of $h_n=O(n^{-2/5})$ if the researcher is only concerned with the non-degenerate case.}
A proof can be found in Section D.2 of the Supplementary Appendix.
{
}
{
}
Let us now consider dyadic KDE with randomly missing incomplete data. Suppose the researcher observes
where the random variables $Z_{ij}\stackrel{d}{=}\text{Bernoulli}(p_n)$ that determine whether each edge is observed are i.i.d. and assumed to be independent from $(X_{ij})_{(i,j)\in {(i,j)\in I_{n}}}$ for some unknown probability of observation $p_n\in(0,1)$ { which can be but is not necessarily fixed in sample size.} Define $\widehat N=\widehat p_n {n\choose 2}$, $\widehat N_1=\widehat p_n {n-1\choose 2}$, and $\widehat N_2=\widehat p_n {n-2\choose 2}$, with $\widehat p_n={n\choose 2}^{-1}\sum_{{(i,j)\in I_{n}}} Z_{ij}$ being an estimate for $p_n$. Now let $\mathbb I_n=\{(i,j)\in I_{n}:Z_{ij}=1\}$, $\mathbb I_n^{(k)}=\{(i,j)\in I_{n}:Z_{ij}=1, i,j\ne k\}$ and $\mathbb I_n^{(k,\ell)}=\{(i,j)\in I_{n}:Z_{ij}=1, i,j\not\in\{ k,\ell\}\}$. { For a fixed $\theta=f(x)$, define the incomplete dyadic KDE estimator and its leave-out counterparts by
and the JEL pseudo true value for incomplete dyadic data by
where $\widehat S(\theta)=\widehat \theta_{\text{inc}}-\theta$, $\widehat S^{(i)}(\theta)=\widehat \theta_{\text{inc}}^{(i)}-\theta$.} In addition, for each $i< j$, define
For each $\theta$, define the modified JEL pseudo true value for incomplete dyadic data by
Now the modified JEL function for incomplete dyadic data for $\theta$ can be defined by
and $\widehat\ell(\theta)$ is defined analogously with $\widehat V_i^m$ replaced by $\widehat V_i$. The Lagrangian dual problems are now
Assumption (ref) imposes the random missing structure on the data and establishes a lower bound on the rate of $p_n$, the probability of observing an edge, to ensure the presence of enough data to establish asymptotic theory. Note that the this assumption accommodates sequences of $p_n$ that converge to zero slowly enough. { The condition $nh_n p_n\to \infty$ ensures that we have asymptotically increasing effective sample size. }
The following result is an incomplete data counterpart of Theorem (ref).
A proof can be found in Section D.3 of the Supplementary Appendix.
{ Theorem (ref) states the asymptotic distributions of the proposed JEL and modified JEL statistics for incomplete data. As in Theorem (ref), the modified JEL is asymptotically pivotal regardless of the asymptotic regime. However, with incomplete data, the asymptotic pivotality of the proposed JEL statistic happens not only under nondegeneracy but also when the proportion of observed data goes to zero slowly asymptotically. { In such a scenario, the term that consists of the randomness induced by the missing process (which is conditionally independent) dominates the original leading asymptotic term, which is potentially not pivotal. } Note that the construction of confidence intervals in Remark (ref) can be adapted with { $\ell^\bullet\in\{\widehat \ell^m,\widehat \ell\}$}. }
We consider four sets of data generating processes (DGP) in our simulations. The first two follow
with $U_{\{i,j\}}\stackrel{i.i.d.}{\sim}N(0,1)$ and $U_i=-1$ with probability $1/3$ and equals $1$ otherwise. We set $\beta\in\{0,1\}$, where $\beta=1$ is the sufficiently non-degenerate DGP considered in GrahamNiuPowell2019. This has the density:
where $\phi$ is the density of a standard normal random variable. Meanwhile, $\beta=0$ corresponds to the degenerate case where the true DGP is i.i.d. standard normal with density $f=\phi$. { The last two sets of DGPs are the incomplete data counterparts of the first two, with the same DGPs except that not all of the edges are observed following the set-up in Section (ref). For these cases, we set $p_n=0.5$.} We utilize the rule of thumb bandwidths $h_n^{S}$ and $h_n^{S,\text{inc}}$ proposed in (A.1) and (A.2) in Section A of the Supplementary Appendix, respectively. For complete dyadic data, we consider five alternative inference methods that are theoretically robust to dyadic clustering: (i) Wald statistic with FG variance estimator for dyadic KDE proposed in GrahamNiuPowell2019 (FG), (ii) Wald statistic with the leading term of FG (lFG), (iii) jackknife empirical likelihood (JEL), (iv) Wald statistic with modified jackknife variance estimator (mJK) from Remark (ref), and (v) modified jackknife empirical likelihood (mJEL). Note that (iii)-(v) are studied in this paper. { For incomplete dyadic data, we consider the incomplete counterparts of (iii) and (v) from Section (ref).} Throughout the simulation studies, we set the design point $x=1.675$, consistent with the simulation study in GrahamNiuPowell2019. Each simulation is iterated $5,000$ times.\\
Tables (ref), (ref), (ref), and (ref) show the simulation results under these different settings. In general, the simulation supports our theoretical findings. { For complete data, under both the dyadic and i.i.d. DGPs, both mJEL and mJK-based confidence intervals enjoy close to nominal coverage rates even with moderate sample sizes, while the FG estimator is oversized with small sample sizes and converges to the nominal size when $n$ is large. The mJK-based intervals are slightly more oversized than mJEL-based intervals for small sample sizes but they perform similarly when $n$ grows larger. The original JEL and lFG estimators are conservative but asymptotically valid under the dyadic DGP. Consistent with Remark 5, the JEL estimator is always severely conservative under the i.i.d. DGP, as is the lFG estimator. The behaviors of JEL and mJEL-based intervals under incomplete data are similar to those we observed with complete data; JEL is conservative but asymptotically converging to the correct size in the non-degenerate case while mJEL remains quite precise under both degenerate and non-degenerate DGPs.}
Flight delays caused by airport congestion are a familiar experience to the flying public. Aside from the frustration experienced by affected passengers, these delays are an important cause of fuel wastage and result in the release of criteria pollutants such as nitrogen oxides, carbon monoxide and sulfur oxides. As discussed in SchlenkerWalker2016, these releases harm the health of people living near airports. To mitigate these harms, it is important to understand the distribution of congestion across routes, or flights between pairs of airports.\\ In this empirical application, we apply the dyadic KDE and JEL and modified JEL procedure to understand the distribution of airport congestion, as measured by taxi time, across pairs of airports in the United States. The taxi time of a flight is defined as the sum of taxi-out time (the time between the plane leaving its parking position and taking off) and taxi-in time (the time between the plane landing and parking). In this context, the edges $X_{ij}$ and $X_{jk}$ represent flights between airport $i$ and airport $j$ and flights between airport $j$ and airport $k$ respectively; the fact that flights share a common airport $j$ induces dependence between $X_{ij}$ and $X_{jk}$. Therefore, this is a natural setting to apply dyadic KDE and modified JEL.\\ Data on flight taxi times are obtained from the US DOT Bureau of Transportation Statistics (BTS) Reporting Carrier On-Time Performance dataset, which compiles data on US domestic non-stop flights from all carriers which earn at least 0.5% of scheduled domestic passenger revenues. From the flight-level data for June 2020, we collapse taxi time into various summary statistics (mean, 95th percentile, maximum) for each origin-destination airport pair, with the summary statistics being taken over flights in both directions so that the resulting network is undirected. Noting that a missing edge means that there were no non-stop flights between the vertices the edge connects, we apply the KDE, JEL, and modified JEL to the resulting network. For simplicity, the sample is restricted to the 100 largest airports by number of departing flights.\\ The histograms, dyadic KDE estimates, pointwise JEL and pointwise modified JEL 95% confidence intervals for the mean, 95th percentile and maximum taxi times between airport pairs in June 2020 are shown in Figure 1A-1C. The intervals obtained by numerically inverting the test statistic for each design point (each whole minute in the range of each statistic). { Since both the JEL and modified JEL confidence intervals in Figure 1 are pointwise, a comparison between JEL and mJEL can only be made for each of the finitely many points at which the confidence intervals were calculated. Note that the design points are the same for the JEL and mJEL graphs, making a point-by-point comparison possible.}\\ For applied researchers, the results demonstrate how studying only the mean can be incomplete; unlike the mean, the 95th percentile and (particularly) maximum travel times exhibit positive skewness. Routes with taxi times lying in the positively skewed region cause the most problems, but this distinction is lost when only studying means. For all these statistics, the modified JEL procedure which we propose provides relatively precise estimates even under small sample sizes (100 vertices) and a large proportion of missing edges (72% missing).
This paper studies the asymptotic properties of dyadic kernel density estimation. We establish uniform convergence rates under general conditions. For inference, we propose a modified jackknife empirical likelihood procedure which is valid regardless of degeneracy of the underlying DGP. We further extend the results to cover incomplete or missing at random dyadic data. Simulation studies show robust finite sample performance of the modified JEL inference for dyadic KDE in different settings. We illustrate the method via an application to airport congestion.\\ Despite our focus on density estimation, the inference approach we take can be extended to cover the local regression as in graham2021minimax, the local polynomial density estimation of CattaneoMaJansson2020JASA or local polynomial distributional regression of CattaneoMaJansson2020wp under dyadic data. Another potential direction is to study incomplete data under unconfoundedness or other non-random missing models. { Finally, in light of the recent paper of cattaneo2022 which pioneers uniform inference for dyadic KDE using Wald statistics, it would be interesting to consider ways to adapt modified JEL for uniform inference as well. } Investigating the modified JEL under these assumptions provides interesting directions for future research.