EconBase
← Back to paper

Maximal Inequalities for Separately Exchangeable Empirical Processes

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

39,787 characters · 4 sections · 33 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Maximal Inequalities for Separately Exchangeable Empirical Processes

\address[Harold D. Chiang]{Department of Economics, University of Wisconsin-Madison, 1180 Observatory Drive Madison, WI 53706-1393, USA.} \email{[email removed]}

abstractThis paper derives new maximal inequalities for empirical processes associated with separately exchangeable random arrays. For fixed index dimension $K\ge 1$, we establish a global maximal inequality bounding the $q$-th moment ($q\in[1,\infty)$) of the supremum of these processes. We also obtain a refined local maximal inequality controlling the first absolute moment of the supremum. Both results are proved for a general pointwise measurable function class. Our approach uses a new technique partitioning the index set into transversal groups, decoupling dependencies and enabling more sophisticated higher moment bounds.

\allowdisplaybreaks

Introduction

This paper develops novel local and global maximal inequalities for empirical processes of separately exchangeable arrays, where the index dimension \(K\in\mathbb{N}\) is fixed but arbitrary and the empirical processes are defined on general classes of functions. Separately exchangeable (SE) arrays are widely utilised in modelling multiway-clustered random variables and/or \(K\)-partite networks in econometrics and statistics (see, for example, davezies2021empirical,mackinnon2021wild,menzel2021bootstrap,chiang2022multiway,graham2024sparse).

Maximal inequalities are powerful tools that are indispensable in the analysis of numerous econometric and statistical problems. They have proven crucial in areas such as semiparametric estimation and debiased machine learning (e.g. belloni2015uniform,chernozhukov2018double), quantile and instrumental variable quantile regression (see, for instance, kato2012asymptotics,chetverikov2016iv,galvao2020unbiased), conditional mode estimation (ota2019quantile), testing many moment inequalities chernozhukov2019inference, adversarial learning (kaji2023adversarial), targeted minimum loss-based estimation (van2017generally), and density estimation for dyadic data (cattaneo2024uniform), to name but a few.

Despite promising recent advances in the studies of SE arrays and its potential benefits for statistical inference and econometric applications, the theoretical literature on maximal inequalities for SE arrays remains rather scarce. In particular, no general global maximal inequality is available. While certain global maximal inequalities have been established for special cases—for instance, Theorem B.2 in chiang2023inference provides a bound for a setting with general \(K\) and any \(q\)-th moment (with \(q\in[1,\infty)\)) albeit only for a finite class of functions, and Lemma C.3 in liu2024estimation extends the result for a general class of functions but only for the first absolute moment (\(q=1\)) in two-way (\(K=2\)) settings —no general global maximal inequality applicable to arbitrary \(K,q\) and an infinite pointwise measurable class of functions has been developed for SE arrays.

The local maximal inequality is a crucial tool for obtaining sharp convergence rates. Apart from the classical i.i.d. case (\(K=1\)), no local maximal inequality for SE arrays exists in the literature. Establishing a local maximal inequality for SE arrays is especially challenging because their intricate multiway dependence structure induces complex interactions among observations. Unlike in the i.i.d. or \(U\)-statistics settings, the multidimensional dependencies inherent to SE arrays render classical techniques—such as symmetrisation and Hoeffding averaging—inapplicable. To address these challenges, we propose a novel proof strategy that carefully partitions the index set into transversal groups. This innovative construction effectively decouples the dependencies among observations within each group, thereby enabling the application of the Hoffmann–J\o rgensen inequality to derive sharp bounds for terms involving higher moments. Consequently, our work fills a critical gap by providing both global and local maximal inequalities for a potentially uncountable but pointwise measurable class of functions under SE sampling.

These methodological advances are underpinned by foundational results from empirical process theory. For textbook treatments, see e.g. vdVW1996,delaPenaGine1999,gine2016mathematical. In particular, the local maximal inequalities for i.i.d.\ random variables established in vanderVaartWellner2011 and CCK2014AoS, as well as maximal inequalities for \(U\)-statistics/processes presented in Chen2018,ChenKato2019b, provide the technical backbone for our arguments. Our results also built directly upon the symmetrisation and Hoeffding type decomposition for SE arrays developed in chiang2023inference.

We follow the fundamental notation for SE arrays as presented in davezies2021empirical, chiang2023inference. To fix ideas, let $K$ be a fixed positive integer and denote by \[ \bm{i} = (i_1, i_2, \dots, i_K) \in \mathbb{N}^K, \] a $K$-tuple index. Given a probability space $(S,\mathcal{S},P)$, suppose that \[\{ X_{\bm{i}} : \bm{i} \in \mathbb{N}^K \}\] is a collection of $\mathcal{S}$-valued random variables satisfying the separate exchangeability (SE) and dissociation (D) conditions defined below.

enumerate• For any $\pi =(\pi_1,\dots,\pi_K)$, a $K$-tuple of permutations of $\mathbb{N}$, $\{X_{\bm{i}}\}_{{\bm{i}}\in \mathbb{N}^K}$ and $\{X_{\pi(\bm{i})}\}_{{\bm{i}}\in \mathbb{N}^K}$ are identically distributed. • For any two set of indices $I,I'\subset \mathbb{N}^K$, $\{X_{\bm{i}}\}_{{\bm{i}}\in I}$ and $\{X_{\bm{i}}\}_{{\bm{i}}\in I'}$ are independent.

Under Conditions (SE) and (D), the Aldous--Hoover--Kallenberg (AHK) representation (see Corollary 7.35 in kallenberg2005probabilistic) guarantees the existence of the following representation:

align[align omitted — 149 chars of source]

where $\odot$ denotes the Hadamard (element-wise) product, the collection \[ \{ U_{\bm{i} \odot \bm{e}} : \bm{i} \in \mathbb{N}^K,\; \bm{e} \in \{0,1\}^K \setminus \{\bm{0}\} \} \] consists of mutually independent and identically distributed (i.i.d.) random variables, and $\tau$ is a Borel measurable map taking values in $\mathcal{S}$.

Let $\bm{N} = (N_1, N_2, \dots, N_K)$ and define \[ [\bm{N}] = \prod_{k=1}^{K} \{1,2,\dots,N_k\}. \] Also, denote \[ N = \prod_{k=1}^K N_k,\quad n = \min\{N_1, N_2, \dots, N_K\},\quad \text{and}\quad \overline{N} = \max\{N_1, N_2, \dots, N_K\}. \] We say a class of functions $\mathcal{F}:\mathcal{S}\to \mathbb{R}$ is pointwise measurable if there exists a countable subclass $\mathcal{F}'\subset \mathcal{F}$ such that for each $f\in \mathcal{F}$, there exists a sequence $(f_j)_j\subset \mathcal{F}'$ such that $f_j\to f$ pointwisely. Given the observed set of random variables \(\{X_{\bm{i}}:{\bm{i}} \in [\bm{N}]\}\) that satisfy Conditions (SE) and (D), and a pointwise measurable class of functions $\mathcal{F}$ with elements $f : \mathcal{S} \to \mathbb{R}$, define the sample mean process by \[ \mathbb{\mathbb{E}}_{N} f = \frac{1}{N} \sum_{\bm{i} \in [\bm{N}]} f(X_{\bm{i}}) \] and the empirical process by \[ \mathbb{G}_{n}(f) = \frac{\sqrt{n}}{N} \sum_{\bm{i} \in [\bm{N}]} \Bigl\{ f(X_{\bm{i}}) - \mathbb{E}\bigl[f(X_{\bm{1}})\bigr] \Bigr\}, \] where $\bm{1}=(1,...,1)$. Without loss of generality, assume that $\mathbb{E}[f(X_{\bm{1}})] = 0$ for all $f \in \mathcal{F}$. In this paper, we establish inequalities that control the $q$-th moment of the supremum of the empirical process, \( \mathbb{E}\bigl[\|\mathbb{G}_{n}\|_\mathcal{F}^q \bigr] , \) for some $q \in [1,\infty)$.

Notation

Let $\mathbb{N}$ denote the set of positive integers and $\mathbb{R}$ for the real line. For $a,b\in \mathbb{R}$, let $a\vee b=\max\{a,b\}$ and $a\wedge b = \min\{a,b\}$. Denote for $m\in \mathbb N$ that $ [m] = \{1,2,\ldots,n\}. $ For two real vectors $\bm{a}= (a_{1},\dots,a_{K})$ and $\bm{b} = (b_{1},\dots,b_{K})$, we denote $\bm{a} \le \bm{b}$ for $a_{j} \le b_{j}$ for all $1 \le j \le K$. Let $\mathrm{supp}(\bm{a}) = \{ j : a_j \ne 0\}$. We denote by $\odot$ the Hadamard product, i.e., for ${\bm{i}} = (i_1,\dots,i_K)$ and $\bm{j} = (j_1,\dots,j_K)$, ${\bm{i}} \odot \bm{j}= (i_1 j_1,\dots,i_K j_K)$. For each $k=1,2,...,K$, define $\mathcal{E}_k=\{{\bm{e}}\in \{0,1\}^K: {\bm{e}} \odot (1,...,1)=k\}$ and thus $\{0,1\}^K=\cup_{k=1}^K \mathcal{E}_k$. For $q\in [1,\infty]$, let $\|f\|_{Q,q}=(Q|f|^q)^{1/q}$. For a non-empty set $T$ and $f:T\to \mathbb{R}$, denote $\|f\|_T=\sup_{t\in T}|f(t)|$. For a pseudometric space $(T,d)$, let $N(T,d,\varepsilon)$ denote the $\varepsilon$-covering number for $(T,d)$. We say $F:\mathcal{S}\to \mathbb{R}_+$ is an envelope for a class of functions $\mathcal{F}\ni f:\mathcal{S}\to \mathbb{R}$ if $\sup_{f\in\mathcal{F}}|f(x)|\le F(x)$ for all $x\in\mathcal{S}$. For $0 < \beta < \infty$, let $\psi_{\beta}$ be the function on $[0,\infty)$ defined by $\psi_{\beta} (x) = e^{x^{\beta}}-1$. Let $\| \cdot \|_{\psi_{\beta}}$ denote the associated Orlicz norm, i.e., $\| \xi \|_{\psi_\beta}=\inf \{ C>0: \mathbb{E}[ \psi_{\beta}( | \xi | /C)] \leq 1\}$ for a real-valued random variable $\xi$.

Main Results

Before presenting the main results, let us first introduce the Hoeffding type decomposition from chiang2023inference. For any \(\bm{i}\in [\bm{N}]\), define

align*[align* omitted — 186 chars of source]

We then define recursively for \(k=1,2,\dots, K\) that

align*[align* omitted — 107 chars of source]

and for \(\bm{e}\in \bigcup_{k=2}^K\mathcal{E}_k\) set

align*[align* omitted — 330 chars of source]

Note that by the AHK representation (ref), for a fixed \(\bm{e}\) the distributions of \[ (P_{\bm{e}}f)\Bigl(\{U_{\bm{i}\odot \bm{e}'}\}_{\bm{e}'\le \bm{e}}\Bigr) \quad\text{and}\quad (\pi_{\bm{e}}f)\Bigl(\{U_{\bm{i}\odot \bm{e}'}\}_{\bm{e}'\le \bm{e}}\Bigr) \] do not depend on the index \(\bm{i}\). Hence, we shall write \(P_{\bm{e}}f\) and \(\pi_{\bm{e}}f\) for a generic \(\bm{i}\).

Now, fix any \(1\le k\le K\) and let \(\bm{e}\in \mathcal{E}_k\). Then, by Lemma 1 in chiang2023inference, for any \(\ell\in \mathrm{supp}(\bm{e})\) the random variable \( (\pi_{\bm{e}}f)\Bigl(\{U_{\bm{i}\odot \bm{e}'}\}_{\bm{e}'\le \bm{e}}\Bigr) \) has mean zero conditionally on \(\{U_{\bm{i}\odot \bm{e}'}\}_{\bm{e}'\le \bm{e}-\bm{e}_{\ell}}\). In addition, define \( I_{\bm{N},\bm{e}} = \{\bm{i}\odot \bm{e} : \bm{i}\in [\bm{N}]\}. \) Then, we have \[ \bigl|I_{\bm{N},\bm{e}}\bigr| = \prod_{k'\in\mathrm{supp}(\bm{e})} N_{k'}. \] Accordingly, define

align*[align* omitted — 195 chars of source]

we now obtain the Hoeffding-type decomposition

align*[align* omitted — 111 chars of source]

To bound \(\mathbb{E}\bigl[ \|\mathbb{G}_{n}(f)\|_{\mathcal{F}} \bigr] \), it thus suffices to control each individual term \(\mathbb{E}\bigl[\|H_{\bm{N}}^{\bm{e}}(f)\|_{\mathcal{F}}\bigr]\) separately.

Finally, fix any \(1\le k\le K\) and \(\bm{e}\in \mathcal{E}_k\). For a given class $\mathcal{F}$ with an envelope $F$, define for a $\delta>0 $ its uniform entropy integral by

align*[align* omitted — 218 chars of source]

where \[ P_{\bm{e}}\mathcal{F} := \{P_{\bm{e}}f : f\in \mathcal{F}\}, \] and the supremum is taken over all finite discrete distributions \(Q\).

The following result is a general global maximal inequality for SE empirical processes with an arbitrary index order $K$ and for a general order of moment $q\in[1,\infty)$. Its proof follows the arguments in the proof of Corollary B.1 in chiang2023inference with some modifications to account for a more general class of functions.

theorem[Global maximal inequality for SE processes] Suppose $\mathcal{F}:\mathcal{S}\to \mathbb{R}$ is a pointwise measurable class of functions with an envelope $F$. Let $(X_{\bm{i}})_{{\bm{i}}\in [\bm{N}]}$ be a sample from $S$-valued separately exchangeable random vectors $(X_{\bm{i}})_{{\bm{i}}\in \mathbb{N}^K}$. Pick any $1 \le k \le K$ and $\bm{e} \in \mathcal{E}_{k}$. Then, for any $q \in [1,\infty)$, we have \[ |I_{{\bm{N}},{\bm{e}}}|^{1/2}\left (\mathbb{E} \left [ \left \| H_{{\bm{N}}}^{\bm{e}}(f) \right \|_{\mathcal{F}}^{q} \right ] \right)^{1/q} \lesssim J_{\bm{e}}(1) \| F \|_{P,q\vee 2}. \]
proofBy symmetrisation inequality for SE processes (Lemma B.1 in chiang2023inference; note that it is dimension free), for independent Rademacher r.v.'s $(\varepsilon_{1,i_1})$,...,$(\varepsilon_{k,i_k})$ that are independent of $(X_{\bm{i}})_{{\bm{i}}\in\mathbb{N}^K}$, one has \begin{align*} |I_{{\bm{N}},{\bm{e}}}|^{1/2} \left(\mathbb{E}[\|H_{\bm{N}}^{\bm{e}}(f)\|_\mathcal{F}^q]\right)^{1/q} =&\left (\mathbb{E} \left [ \left \|\frac{1}{\sqrt{|I_{{\bm{N}},{\bm{e}}}|}} \sum_{{\bm{i}}\in I_{{\bm{N}},{\bm{e}}}} (\pi_{ {\bm{e}}}f )(\{U_{{\bm{i}} \odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}) \right \|_{\mathcal{F}}^{q} \right ] \right)^{1/q}\\ \lesssim& \left (\mathbb{E} \left [ \left \|\frac{1}{\sqrt{|I_{{\bm{N}},{\bm{e}}}|}} \sum_{{\bm{i}}\in I_{{\bm{N}},{\bm{e}}}}\varepsilon_{1,i_1}...\varepsilon_{k,i_k}\cdot(\pi_{ {\bm{e}}}f )(\{U_{{\bm{i}} \odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}) \right \|_{\mathcal{F}}^{q} \right ] \right)^{1/q}. \end{align*} By convexity of supremum and $\cdot \mapsto (\cdot)^q$, Jensen's inequality implies that the RHS above can be upperbounded up to a constant that depends only on $q$, $K$, and $k$ by \begin{align*} \left (\mathbb{E} \left [ \left \|\frac{1}{\sqrt{|I_{{\bm{N}},{\bm{e}}}|}} \sum_{{\bm{i}}\in I_{{\bm{N}},{\bm{e}}}}\varepsilon_{1,i_1}...\varepsilon_{k,i_k}\cdot(P_{\bm{e}} f )(\{U_{{\bm{i}} \odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}) \right \|_{\mathcal{F}}^{q} \right ] \right)^{1/q}. \end{align*} Denote $\mathbb P_{I_{{\bm{N}},{\bm{e}}}}=|I_{{\bm{N}},{\bm{e}}}|^{-1}\sum_{{\bm{i}} \in I_{{\bm{N}},{\bm{e}}}}\delta_{\{U_{{\bm{i}} \odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}}$, the empirical measure on the support of $\{U_{{\bm{i}} \odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}$. Observe that conditionally on $\{X_{\bm{i}}\}_{{\bm{i}}\in{\bm{N}}}$, the object \begin{align*} \frac{1}{\sqrt{|I_{{\bm{N}},{\bm{e}}}|}}\sum_{{\bm{i}}\in I_{{\bm{N}},{\bm{e}}}}\varepsilon_{1,i_1}...\varepsilon_{k,i_k}(P_{{\bm{e}}}f )(\{U_{{\bm{i}}\odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}) \end{align*} is a homogeneous Rademacher chaos processes of order $k$. By Lemma (ref), $L^q$ norm is bounded from above by $\psi_{2/k} $-norm up to a constant depends only on $(q,k)$, and thus by applying Corollary 5,1.8 in delaPenaGine1999, one has \begin{align*} &\left (\mathbb{E} \left [ \left \|\frac{1}{\sqrt{|I_{{\bm{N}},{\bm{e}}}|}} \sum_{{\bm{i}}\in I_{{\bm{N}},{\bm{e}}}}\varepsilon_{1,i_1}...\varepsilon_{k,i_k} \cdot(P_{ {\bm{e}}}f )(\{U_{{\bm{i}} \odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}) \right \|_{\mathcal{F}}^{q} \right ] \right)^{1/q}\\ \lesssim& \mathbb{E} \left [ \left\| \left \|\frac{1}{\sqrt{|I_{{\bm{N}},{\bm{e}}}|}} \sum_{{\bm{i}}\in I_{{\bm{N}},{\bm{e}}}}\varepsilon_{1,i_1}...\varepsilon_{k,i_k}\cdot(P_{ {\bm{e}}}f )(\{U_{{\bm{i}} \odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}) \right \|_{\mathcal{F}}\right\|_{\psi_{2/k}|(X_{\bm{i}})_{{\bm{i}}\in [{\bm{N}}]}}\right ] \\ \lesssim& \mathbb{E} \left [ \int_0^{\sigma_{I_{{\bm{N}},{\bm{e}}}}}\left[1+\log N\left(P_{ {\bm{e}}}\mathcal{F},\|\cdot\|_{\mathbb P_{I_{{\bm{N}},{\bm{e}}}},2},\tau\right)\right]^{k/2}d \tau \right ], \end{align*} where $\sigma_{I_{{\bm{N}},{\bm{e}}}}^2:=\sup_{f\in\mathcal{F}}\| P_{\bm{e}} f \|_{\mathbb P_{I_{{\bm{N}},{\bm{e}}}},2}^2 $. Using a change of variable and the definition of $J_{\bm{e}}$, the above bound becomes \begin{align*} &\mathbb{E} \left [ \int_0^{\sigma_{I_{{\bm{N}},{\bm{e}}}}}\left[1+\log N\left(P_{ {\bm{e}}} \mathcal{F},\|\cdot\|_{\mathbb P_{I_{{\bm{N}},{\bm{e}}}},2},\tau\right)\right]^{k/2}d \tau \right ] \\ =&\mathbb{E} \left [\|P_{ {\bm{e}}} F\|_{\mathbb P_{I_{{\bm{N}},{\bm{e}}}},2} \int_0^{\sigma_{I_{{\bm{N}},{\bm{e}}}}/\|P_{ {\bm{e}}} F\|_{\mathbb P_{I_{{\bm{N}},{\bm{e}}}},2}}\left[1+\log N\left(P_{ {\bm{e}}} \mathcal{F},\|\cdot\|_{\mathbb P_{I_{{\bm{N}},{\bm{e}}}},2},\tau\|P_{ {\bm{e}}} F\|_{\mathbb P_{I_{{\bm{N}},{\bm{e}}}},2}\right)\right]^{k/2}d \tau \right ]\\ \le& \mathbb{E} \left [\|P_{ {\bm{e}}} F\|_{\mathbb P_{I_{{\bm{N}},{\bm{e}}}},2} J_{\bm{e}}\left(\sigma_{I_{{\bm{N}},{\bm{e}}}}/\|P_{ {\bm{e}}} F\|_{\mathbb P_{I_{{\bm{N}},{\bm{e}}}},2}\right) \right ]\\ \le&J_{\bm{e}}\left(1\right) \|F\|_{P,q\vee 2} , \end{align*} where the last inequality follows from Jensen's inequality.

Although the global maximal inequality works for general $q$, in the case that the supremum of the first absolute moment is concerned, local maximal inequalities usually provides shaper bounds. The following is a novel local maximal inequality for SE empirical processes. Unlike the proof of the global maximal inequality, which is largely analogous to the corresponding results for $U$-processes that can be found in ChenKato2019b, its proof relies on a novel argument that utilises construction of a partition with the property that each block in the partition satisfies a transversality property. We present this construction in Lemma (ref) below.

theorem[Local maximal inequality for SE processes] Suppose $\mathcal{F}:\mathcal{S}\to \mathbb{R}$ is a pointwise measurable class of functions with an envelope $F$. Let $(X_{\bm{i}})_{{\bm{i}}\in [\bm{N}]}$ be a sample from $S$-valued separately exchangeable random vectors $(X_{\bm{i}})_{{\bm{i}}\in \mathbb{N}^K}$. Set ${\bm{e}}\in \{0,1\}^K$ and let $\sigma_{\bm{e}}$ be a constant such that $\sup_{f\in\mathcal{F}}\|P_{\bm{e}} f \|_{P,2}\le \sigma_{\bm{e}} \le \|P_{\bm{e}} F\|_{P,2}$, \[ \delta_{\bm{e}}=\sigma_{\bm{e}}/\|P_{\bm{e}} F\|_{P,2}\quad \text{ and } \quad M_{\bm{e}}=\max_{t\in [n]}(P_{\bm{e}} F)\left(\{U_{(t,...,t)\odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}\right),\] then \begin{align*} |I_{{\bm{N}},{\bm{e}}}|^{1/2}\mathbb{E} \left [ \left \| H_{{\bm{N}}}^{\bm{e}}(f) \right \|_{\mathcal{F}} \right ] \lesssim J_{\bm{e}}(\delta_{\bm{e}}) \|P_{\bm{e}} F\|_{P,2} +\frac{J_{\bm{e}}^2(\delta_{\bm{e}})\|M_{\bm{e}}\|_{P,2}}{\sqrt{n}\delta_{\bm{e}}^2}. \end{align*}
remarkAlthough our proof strategy broadly follows that of Theorem 5.1 in ChenKato2019b, a key divergence arises. In Theorem 5.1, the Hoffmann–J\o rgensen inequality (which requires independence) is applied via the classical \(U\)-statistic technique of Hoeffding averaging (see, for example, Section 5.1.6 in serfling1980approximation). However, the more intricate dependence structure inherent to separately exchangeable arrays renders Hoeffding averaging inapplicable in our context. To address this challenge, we introduce an alternative approach by establishing Lemma (ref), which partitions the index set \(I_{{\bm{N}},{\bm{e}}}\) into \(n\) transveral groups. Together with AHK representation (ref), this yields i.i.d. elements within each group, thereby facilitating the application of the Hoffmann–J\o rgensen inequality.
proofWe first state a crucial technical lemma which will be used in the following proof, a proof of this lemma is provided in the end of this section. \begin{lemma}[Partitioning into transversal groups] For any ${\bm{e}}\in \{0,1\}^K$, $I_{{\bm{N}},{\bm{e}}}$ can be partitioned into subsets $G$'s of size $n$ such that each $G$ is transversal, that is, any two distinct tuples $ (i_1,i_2,\dots,i_K),$ $(i_1',i_2',\dots,i_K')\in G $ satisfy \[ i_1\ne i_1',\quad i_2\ne i_2',\quad \dots,\quad i_K\ne i_K'. \] \end{lemma} We now present the proof of Theorem (ref). For an ${\bm{e}}\in \mathcal{E}_1$, the summands are i.i.d. and thus the desired result follows directly from Lemma (ref). Therefore, we assume $K\ge 2$ and ${\bm{e}}\in \mathcal{E}_k$ for a $k\in\{2,...,K\} $. Assume without loss of generality that ${\bm{e}}$ consists of $1$'s in its first $k$ elements and zero elsewhere. By applying the symmetrisation of Lemma B.1 in chiang2023inference, one has, for independent Rademacher r.v.'s $(\varepsilon_{1,i_1})$,...,$(\varepsilon_{k,i_k})$ that are independent of $(X_{\bm{i}})_{{\bm{i}}\in\mathbb{N}^K}$, that \begin{align*} |I_{{\bm{N}},{\bm{e}}}|^{1/2}\mathbb{E}[\|H_{\bm{N}}^{\bm{e}}(f)\|_\mathcal{F}] \lesssim& \mathbb{E}\left[\left|\frac{1}{\sqrt{|I_{{\bm{N}},{\bm{e}}}|}}\sum_{{\bm{i}}\in I_{{\bm{N}},{\bm{e}}}}\varepsilon_{1,i_1}...\varepsilon_{k,i_k}(\pi_{{\bm{e}}}f )(\{U_{{\bm{i}}\odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}})\right\|_\mathcal{F}\right]. \end{align*} Further, by convexity of supremum and Jensen's inequality, the RHS above can be upperbounded up to a constant that depends only on $K$ and $k$ by \begin{align*} \mathbb{E}\left[\left|\frac{1}{\sqrt{|I_{{\bm{N}},{\bm{e}}}|}}\sum_{{\bm{i}}\in I_{{\bm{N}},{\bm{e}}}}\varepsilon_{1,i_1}...\varepsilon_{k,i_k}(P_{{\bm{e}}}f )(\{U_{{\bm{i}}\odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}})\right\|_\mathcal{F}\right] \end{align*} Denote $\mathbb P_{I_{{\bm{N}},{\bm{e}}}}=|I_{{\bm{N}},{\bm{e}}}|^{-1}\sum_{{\bm{i}} \in I_{{\bm{N}},{\bm{e}}}}\delta_{\{U_{{\bm{i}} \odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}}$, the empirical measure on the support of $\{U_{{\bm{i}} \odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}$. Observe that conditionally on $\{X_{\bm{i}}\}_{{\bm{i}}\in{\bm{N}}}$, the object \begin{align*} R_{\bm{N}}^{\bm{e}}(f)=\frac{1}{\sqrt{|I_{{\bm{N}},{\bm{e}}}|}}\sum_{{\bm{i}}\in I_{{\bm{N}},{\bm{e}}}}\varepsilon_{1,i_1}...\varepsilon_{k,i_k}(P_{{\bm{e}}}f )(\{U_{{\bm{i}}\odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}) \end{align*} is a homogeneous Rademacher chaos processes of order $k$. Further, following Corollary 3.2.6 in delaPenaGine1999, for any $f,f'\in \mathcal{F}$ \begin{align*} \left\|R_{\bm{N}}^{\bm{e}}(f)-R_{\bm{N}}^{\bm{e}}(f')\right\|_{\psi_{2/k}|\{X_{\bm{i}}\}_{{\bm{i}}\in {\bm{N}}}}\lesssim \left\|R_{\bm{N}}^{\bm{e}}(f)-R_{\bm{N}}^{\bm{e}}(f')\right\|_{\mathbb P_{I_{{\bm{N}},{\bm{e}}}},2}. \end{align*} Hence the diameter of the function class $\mathcal{F}$ in $\|\cdot\|_{\psi_{2/k}|\{X_{\bm{i}}\}_{{\bm{i}}\in {\bm{N}}}}$-norm is upperbounded by $\sigma_{I_{{\bm{N}},{\bm{e}}}}^2$ up to a constant, where $\sigma_{I_{{\bm{N}},{\bm{e}}}}^2:=\sup_{f\in\mathcal{F}}\| P_{{\bm{e}}}f \|_{\mathbb P_{I_{{\bm{N}},{\bm{e}}}},2}^2 $. By applying Fubini's theorem, Corollary 5,1.8 in delaPenaGine1999, and a change of variables, we have \begin{align*} &\mathbb{E}\left[\left\|\frac{1}{\sqrt{|I_{{\bm{N}},{\bm{e}}}|}}\sum_{{\bm{i}}\in I_{{\bm{N}},{\bm{e}}}}\varepsilon_{1,i_1}...\varepsilon_{k,i_k}(P_{{\bm{e}}}f )(\{U_{{\bm{i}}\odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}})\right\|_\mathcal{F}\right]\\ \lesssim& \mathbb{E}\left[\left\|\left\|\frac{1}{\sqrt{|I_{{\bm{N}},{\bm{e}}}|}}\sum_{{\bm{i}}\in I_{{\bm{N}},{\bm{e}}}}\varepsilon_{1,i_1}...\varepsilon_{k,i_k}(P_{{\bm{e}}}f )(\{U_{{\bm{i}}\odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}})\right\|_\mathcal{F}\right\|_{\psi_{2/k}|\{X_{\bm{i}}\}_{{\bm{i}}\in{\bm{N}}}}\right]\\ \lesssim& \mathbb{E} \left [ \int_0^{\sigma_{I_{{\bm{N}},{\bm{e}}}}}\left[1+\log N\left(P_{ {\bm{e}}}\mathcal{F},\|\cdot\|_{\mathbb P_{I_{{\bm{N}},{\bm{e}}}},2},\tau\right)\right]^{k/2}d \tau \right ] \\ =&\mathbb{E} \left [ \|P_{\bm{e}} F\|_{\mathbb P_{I_{{\bm{N}},{\bm{e}}}},2} \int_0^{\sigma_{I_{{\bm{N}},{\bm{e}}}}/\|P_{\bm{e}} F\|_{\mathbb P_{I_{{\bm{N}},{\bm{e}}}},2}}\left[1+\log N\left(P_{ {\bm{e}}}\mathcal{F},\|\cdot\|_{\mathbb P_{I_{{\bm{N}},{\bm{e}}}},2},\tau\|P_{\bm{e}} F\|_{\mathbb P_{I_{{\bm{N}},{\bm{e}}}},2}\right)\right]^{k/2}d \tau \right ]\\ \le& \mathbb{E} \left [\|P_{\bm{e}} F\|_{\mathbb P_{I_{{\bm{N}},{\bm{e}}}},2} J_{\bm{e}}\left(\sigma_{I_{{\bm{N}},{\bm{e}}}}/\|P_{\bm{e}} F\|_{\mathbb P_{I_{{\bm{N}},{\bm{e}}}},2}\right) \right ]. \end{align*} By Lemma (ref), an application of Jensen's inequality yields \begin{align} |I_{{\bm{N}},{\bm{e}}}|^{1/2}\mathbb{E}[\|H_{{\bm{N}}}^{\bm{e}} (f)\|_\mathcal{F}]\lesssim& \|P_{\bm{e}} F\|_{P,2} J_{\bm{e}}\left(z\right), \end{align} where $z:=\sqrt{\mathbb{E}[\sigma_{I_{{\bm{N}},{\bm{e}}}}^2]/\|P_{\bm{e}} F\|_{P,2}^2}$. We now bound \begin{align*} \mathbb{E}[\sigma_{I_{{\bm{N}},{\bm{e}}}}^2]=&\mathbb{E}\left[\left\|\frac{1}{|I_{{\bm{N}},{\bm{e}}}|}\sum_{{\bm{i}} \in I_{{\bm{N}},{\bm{e}}}}(P_{\bm{e}} f )^2\left(\{U_{{\bm{i}} \odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}\right)\right\|_\mathcal{F}\right]. \end{align*} We aim to apply the Hoffmann–J\o rgensen inequality to handle the squared summands. However, because the summands are not independent, we invoke Lemma (ref). By applying this lemma, we obtain a partition \(\mathcal{G}\) of \(I_{{\bm{N}},{\bm{e}}}\) into \(|\mathcal{G}|=|I_{{\bm{N}},{\bm{e}}}|/n\) groups, each containing \(n\) i.i.d. observations. The i.i.d. property follows from the AHK representation (ref) and the fact that within each group, any two observations share no common indices \(i_1,\dots,i_K\). For each group \(G = \{{\bm{i}}_{1}(G), {\bm{i}}_{2}(G), \dots, {\bm{i}}_{n}(G)\} \in \mathcal{G}\), we define \[ D_{f,{\bm{e}}}(G) = \frac{1}{n}\sum_{t=1}^n \Bigl(P_{{\bm{e}}}f\Bigr)^2\!\left(\{U_{{\bm{i}}_t(G)\odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}\right), \] and let \[ D_{f,{\bm{e}}} = \frac{1}{n}\sum_{t=1}^n \Bigl(P_{{\bm{e}}}f\Bigr)^2\!\left(\{U_{(t,\dots,t)\odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}\right). \] Then we have \begin{align*} \frac{1}{|I_{{\bm{N}},{\bm{e}}}|}\sum_{{\bm{i}} \in I_{{\bm{N}},{\bm{e}}}}(P_{{\bm{e}}}f )^2\left(\{U_{{\bm{i}} \odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}\right)=\frac{1}{|I_{{\bm{N}},{\bm{e}}}|/n}\sum_{G\in \mathcal{G}} D_{f,{\bm{e}}}(G). \end{align*} Note that for each \(G\in\mathcal{G}\), the AHK representation in (ref) implies that \(D_{f,{\bm{e}}}\) and \(D_{f,{\bm{e}}}(G)\) are identically distributed. Consequently, by Jensen's inequality, we have \begin{align*} \mathbb{E}\Bigl[\sigma_{I_{{\bm{N}},{\bm{e}}}}^2\Bigr] &= \mathbb{E}\!\left[\left\|\frac{1}{|I_{{\bm{N}},{\bm{e}}}|/n}\sum_{G\in\mathcal{G}} D_{f,{\bm{e}}}(G)\right\|_\mathcal{F}\right] \le \mathbb{E}\!\left[\left\|D_{f,{\bm{e}}}\right\|_\mathcal{F}\right]. \end{align*} Let us denote this bound by \[ B_{n,{\bm{e}}} :=\mathbb{E}\!\left[\left\|D_{f,{\bm{e}}}\right\|_\mathcal{F}\right]= \mathbb{E}\!\left[\left\|\frac{1}{n}\sum_{t=1}^n \Bigl(P_{{\bm{e}}}f\Bigr)^2\!\Bigl(\{U_{(t,\dots,t)\odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}\Bigr)\right\|_\mathcal{F}\right]. \] Thus $z\le\widetilde z:=\sqrt{B_{n,{\bm{e}}}}/\|(\pi_{\bm{e}})F\|_{P,2}$. Note that by symmetrisation inequality for independent processes, the contration principle (Theorem 4.12. in ledoux1991probability), and the Cauchy-Schwarz inequality, one has \begin{align*} B_{n,{\bm{e}}}=&\mathbb{E}\left[\left\|\frac{1}{n}\sum_{t=1}^n (P_{{\bm{e}}}f )^2\left(\{U_{(t,...,t)\odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}\right)\right\|_{\mathcal{F}}\right]\\ \le& \sigma_{\bm{e}}^2 +\mathbb{E}\left[\left\|\frac{1}{n}\sum_{t=1}^n\left\{ (P_{{\bm{e}}}f )^2\left(\{U_{(t,...,t)\odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}\right)-\mathbb{E}\left[(P_{{\bm{e}}}f )^2\right]\right\}\right\|_{\mathcal{F}}\right]\\ \lesssim& \sigma_{\bm{e}}^2 +\mathbb{E}\left[\left\|\frac{1}{n}\sum_{t=1}^n\varepsilon_t\cdot (P_{{\bm{e}}}f )^2\left(\{U_{(t,...,t)\odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}\right)\right\|_{\mathcal{F}}\right]\\ \lesssim& \sigma_{\bm{e}}^2 +\mathbb{E}\left[M_{\bm{e}}\left\|\frac{1}{n}\sum_{t=1}^n\varepsilon_t\cdot (P_{{\bm{e}}}f )\left(\{U_{(t,...,t)\odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}\right)\right\|_{\mathcal{F}}\right]\\ \le& \sigma_{\bm{e}}^2 +\|M_{\bm{e}}\|_{P,2}\sqrt{\mathbb{E}\left[\left\|\frac{1}{n}\sum_{t=1}^n\varepsilon_t\cdot (P_{{\bm{e}}}f )\left(\{U_{(t,...,t)\odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}\right)\right\|_{\mathcal{F}}^2\right]}. \end{align*} An application of Hoffmann-J\o rgensen's inequality (Proposition A.1.6 in vdVW1996) gives \begin{align*} &\sqrt{\mathbb{E}\left[\left\|\frac{1}{n}\sum_{t=1}^n\varepsilon_t\cdot (P_{{\bm{e}}}f )\left(\{U_{(t,...,t)\odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}\right)\right\|_{\mathcal{F}}^2\right]}\\ \lesssim& \mathbb{E}\left[\left\|\frac{1}{n}\sum_{t=1}^n\varepsilon_t\cdot (P_{{\bm{e}}}f )\left(\{U_{(t,...,t)\odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}\right)\right\|_{\mathcal{F}}\right]+\frac{1}{n}\|M_{\bm{e}}\|_{P,2}. \end{align*} By employing analogous reasoning to that used in the initial part of the proof, we deduce that \begin{align*} &\mathbb{E}\left[\left\|\frac{1}{\sqrt{n}}\sum_{t=1}^n\varepsilon_t\cdot (P_{{\bm{e}}}f )\left(\{U_{(t,...,t)\odot {\bm{e}}'}\}_{{\bm{e}}'\le {\bm{e}}}\right)\right\|_{\mathcal{F}}\right]\\ \lesssim& \|P_{\bm{e}} F\|_{P,2}\int_{0}^{\widetilde z} \sup_{Q}\sqrt{1+\log N(P_{\bm{e}}\mathcal{F},\|\cdot\|_{Q,2},\epsilon \|P_{\bm{e}} F\|_{Q,2})}d\epsilon. \end{align*} Note that the integral on the RHS can be bounded by $J_{\bm{e}}(\widetilde z)$ and thus \begin{align*} B_{n,{\bm{e}}}\lesssim& \sigma_{\bm{e}}^2 + n^{-1}\|M_{\bm{e}}\|_{P,2}^2 + n^{-1/2} \|M_{\bm{e}}\|_{P,2} \|P_{\bm{e}} F\|_{P,2}J_{\bm{e}}(\widetilde z). \end{align*} Define $$\Delta=(\sigma_{\bm{e}}\vee n^{-1/2}\|M_{\bm{e}}\|_{P,2})/\|P_{\bm{e}} F\|_{P,2},$$ it then follows that \begin{align*} \widetilde z^2\lesssim \Delta^2 + \frac{\|M_{\bm{e}}\|_{P,2}}{\sqrt{n}\|P_{\bm{e}} F\|_{P,2}}J_{\bm{e}}(\widetilde z). \end{align*} By applying Lemma (ref) and Lemma 2.1 of vanderVaartWellner2011 with $J=J_{\bm{e}}$, $A=\Delta$, $B=\sqrt{\|M_{\bm{e}}\|_{P,2}/\sqrt{n}\|P_{\bm{e}} F\|_{P,2}}$ and $r=1$, it yields that \begin{align*} J_{\bm{e}}(z)\le J_{\bm{e}}(\widetilde z)\lesssim J_{\bm{e}}(\Delta)\left\{1+J_{\bm{e}}(\Delta)\frac{\|M_{\bm{e}}\|_{P,2}}{\sqrt{n}\|P_{\bm{e}} F\|_{P,2} \Delta^2}\right\}. \end{align*} Combining this with ((ref)), we obtain the bound \begin{align} |I_{{\bm{N}},{\bm{e}}}|^{1/2}\mathbb{E}[\|H_{{\bm{N}}}^{\bm{e}} (f)\|_\mathcal{F}]\lesssim& J_{\bm{e}}(\Delta)\|P_{{\bm{e}}} F\|_{P,2} +\frac{J_{\bm{e}}^2(\Delta)\|M_{\bm{e}}\|_{P,2}}{\sqrt{n}\Delta^2}. \end{align} Notice that $\delta_{\bm{e}}\le \Delta$ by their definitions. By Lemma (ref)(iii), one has \begin{align*} J_{\bm{e}}(\Delta)\le \Delta\frac{J_{\bm{e}}(\delta_{\bm{e}})}{\delta_{\bm{e}}}=\max\left\{J_{\bm{e}}(\delta_{\bm{e}}), \frac{\|M_{\bm{e}}\|_{P,2} J_{\bm{e}}(\delta_{\bm{e}})}{\sqrt{n}\|P_{\bm{e}} F\|_{P,2}\delta_{\bm{e}}}\right\}\le\max\left\{J_{\bm{e}}(\delta_{\bm{e}}), \frac{\|M_{\bm{e}}\|_{P,2} J_{\bm{e}}^2(\delta_{\bm{e}})}{\sqrt{n}\|P_{\bm{e}} F\|_{P,2}\delta_{\bm{e}}^2}\right\}, \end{align*} where the second inequality follows from the fact $J_{\bm{e}}(\delta_{\bm{e}})/\delta_{\bm{e}}\ge J_{\bm{e}}(1)\ge 1$. Finally, using Lemma (ref)(iii), \begin{align*} \frac{J_{\bm{e}}^2(\Delta)\|M_{\bm{e}}\|_{P,2}}{\sqrt{n}\Delta^2}\le \frac{J_{\bm{e}}^2(\delta_k)\|M_{\bm{e}}\|_{P,2}}{\sqrt{n}\delta_k^2} \end{align*} Combining the calculations with the bound in ((ref)), we have the desired inequality. \subsection*{Proof of Lemma (ref)} For $K=1$, the result is trivial. For $K\ge 2$, assume without loss of generality that $ N_1 \ge N_2 \ge \cdots \ge N_K. $ We prove for the case of ${\bm{e}}=(1,...,1)$ and $I_{{\bm{N}},{\bm{e}}}=[{\bm{N}}]$ since other cases follow exactly the same arguments. Our goal is to partition $[{\bm{N}}]$ into subsets (which we call groups) of size \(N_K\) that are transversal. For \(j=1,2,\dots,K-1\), define $\phi_j\colon [N_K]\times [N_j] \to [N_j]$ by \[ \phi_j(t,g)= ((t+g-2) \mod N_j) + 1. \] That is, for each \(t\in [N_K]\) and \(g\in [N_j]\) the value \(\phi_j(t,g)\) is computed by adding \(t\) and \(g-1\), reducing modulo \(N_j\) (so that the result lies in \(\{0,1,\dots,N_j-1\}\)), and then adding 1 to get an element of \([N_j]\). For the \(K\)-th coordinate we set $ \phi_K(t) = t \quad \text{for } t\in [N_K]. $ Index the groups by \[ (g_1,g_2,\dots,g_{K-1})\in [N_1]\times [N_2]\times \cdots \times [N_{K-1}]. \] Then, for each such \((g_1,\dots,g_{K-1})\), define \[ G_{(g_1,\dots,g_{K-1})} = \Bigl\{\, \Bigl( \phi_1(t,g_1),\, \phi_2(t,g_2),\, \dots,\, \phi_{K-1}(t,g_{K-1}),\, t \Bigr) : t\in [N_K] \,\Bigr\}. \] Thus, each group contains \(N_K\) elements. We now claim the transversality property in each group. For a fixed group \(G_{(g_1,\dots,g_{K-1})}\) and a fixed coordinate \(j\) (with \(1\le j\le K-1\)), the \(j\)th coordinate of an element is given by \[ \phi_j(t,g_j) = ((t+g_j-2) \mod N_j) + 1. \] Since the mapping \( t \mapsto ((t+g_j-2) \mod N_j) + 1 \) is injective (note that \(N_K\le N_j\) so that there is no collision in the range), it follows that the \(j\)-th coordinates of the elements of \(G_{(g_1,\dots,g_{K-1})}\) are all distinct. For the \(K\)-th coordinate, the identity mapping \(t \mapsto t\) is trivially injective. Next we show the covering of $[{\bm{N}}]$ and disjointness of the groups. Recall that the total number of groups is \( N_1\cdot N_2\cdots N_{K-1}. \) Each group has \(N_K\) elements; hence, the union of all groups has \( (N_1\cdot N_2\cdots N_{K-1})\cdot N_K = N \) elements. For surjectivity, let $x = (x_1, x_2, \ldots, x_K)$ be an arbitrary element of $[{\bm{N}}] = [N_1] \times [N_2] \times \cdots \times [N_K]$. We wish to show that there exist $(g_1, g_2, \ldots, g_{K-1}) \in [N_1] \times [N_2] \times \cdots \times [N_{K-1}]$ and $t \in [N_K]$ such that $$ x = \Bigl( \phi_1(t, g_1),\, \phi_2(t, g_2),\, \dots,\, \phi_{K-1}(t, g_{K-1}),\, t \Bigr). $$ Set $t = x_K$. Then for each $j = 1, 2, \ldots, K-1$, we must have $ \phi_j(x_K, g_j) = x_j, $ where by definition $\phi_j(x_K, g_j) = (((x_K + g_j - 2) \bmod N_j) + 1)$. Notice that for each fixed $x_K$, the mapping $$ g \mapsto ((x_K + g - 2) \bmod N_j) + 1 $$ is an affine function (with coefficient $1$) on the cyclic group $\mathbb{Z}/N_j$, and hence it is a bijection from $[N_j]$ onto $[N_j]$. Thus, for each $j$ there exists a unique $g_j \in [N_j]$ such that $\phi_j(x_K, g_j) = x_j$. Therefore, every $x \in S$ can be uniquely written in the form $$ \Bigl( \phi_1(x_K, g_1),\, \phi_2(x_K, g_2),\, \dots,\, \phi_{K-1}(x_K, g_{K-1}),\, x_K \Bigr), $$ which shows that the mapping $$ \tau: (g_1,\dots,g_{K-1},t) \mapsto \Bigl( \phi_1(t, g_1),\, \phi_2(t, g_2),\, \dots,\, \phi_{K-1}(t, g_{K-1}),\, t \Bigr) $$ is bijective. Thus, the collection \[ \mathcal{G} = \Bigl\{\, G_{(g_1,\dots,g_{K-1})} : (g_1,\dots,g_{K-1})\in [N_1]\times \cdots \times [N_{K-1}]\,\Bigr\} \] is a partition of $[{\bm{N}}]$ into groups of size \(N_K\), and in every group the entries in each coordinate are distinct.

Following Chapter 3.7 in gine2016mathematical, a function class $\mathcal{F}$ on $\mathcal{S}$ with envelope $F$ is called Vapnik–Chervonenkis-type (VC-type) with characteristics $(A,v)$ if

align*[align* omitted — 154 chars of source]

where the supremum is taken over all finite discrete distributions. By adapting the arguments used in the proofs of Corollaries 5.3 and 5.5 and Lemma 5.4 in ChenKato2019b, we derive the following local maximal inequality for VC-type function classes.

corollaryUnder the same setting as in Theorem (ref). In addition, suppose $\mathcal{F}$ is of VC-type with characteristics $A\ge (e^{2(K-1)}/16)\vee e$ and $v\ge 1$, then for each ${\bm{e}}\in \mathcal{E}_k$, one has \begin{align*} |I_{{\bm{N}},{\bm{e}}}|^{1/2}\mathbb{E} \left [ \left \| H_{{\bm{N}}}^{\bm{e}}(f) \right \|_{\mathcal{F}} \right ] \lesssim \sigma_{\bm{e}}\{v\log (A\vee \overline N)\}^{k/2} +\frac{\|M_{\bm{e}}\|_{P,2}}{\sqrt{n}}\{v\log (A\vee \overline N)\}^k. \end{align*}

Conclusion

In this paper, we have derived novel maximal inequalities for empirical processes associated with separately exchangeable (SE) arrays. Our contributions include a global maximal inequality that bounds the \(q\)-th moment of the supremum for any \(q\in[1,\infty)\), as well as a refined local maximal inequality controlling the first absolute moment, both established for a general pointwise measurable class of functions. These results extend the literature beyond the i.i.d. case and overcome the challenges posed by the intricate dependence structure of SE arrays.

A key innovation of our approach is the introduction of a new proof technique—partitioning the index set into transversal groups—which circumvents the limitations of classical tools such as Hoeffding averaging. This advancement not only fills an important gap in the theoretical framework for SE arrays, but also paves the way for more robust applications in econometric and statistical inference. Future research may build on these findings to further explore maximal inequalities under even broader conditions and to enhance their applicability in high-dimensional and machine learning contexts.