Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
243,429 characters · 58 sections · 18 citation commands
Supplementary Appendix to “Robust Inference for the Direct Average Treatment Effect with Treatment Assignment Interference”
For $n \in \mathbb{N}$, $[n] = \{1, \cdots, n\}$. For reals sequences $a_n = o(b_n)$ if $\limsup_{n\to\infty} \frac{|a_n|}{|b_n|} = 0$, $|a_n| \lesssim |b_n|$ if there exists some constant $C$ and $N > 0$ such that $n > N$ implies $|a_n| \leq C |b_n|$. For sequences of random variables $a_n = o_{\mathbb{P}}(b_n)$ if $\operatorname{plim}_{n \rightarrow \infty}\frac{|a_n|}{|b_n|} = 0$, $a_n = O_{\mathbb{P}}(b_n)$ if $\limsup_{M \rightarrow \infty} \limsup_{n \rightarrow \infty} \mathbb{P}[|\frac{a_n}{b_n}| \geq M] = 0$. For positive real sequences $a_n \ll b_n$ if $a_n = o(b_n)$. For a sequence of real-valued random variables $X_n$, we say $X_n = O_{\psi_p}(r_n)$ if there exists $N \in \mathbb{N}$ and $M > 0$ such that $\lVert X_n \rVert_{\psi_p} \leq M r_n$ for all $n \geq N$, where $\lVert \cdot \rVert_{\psi_p}$ is the Orlicz norm w.r.p $\psi_p(x) = \exp(x^p) - 1$. We say $X_n = O_{\psi_p, tc}(r_n)$, $tc$ stands for tail control, if there exists $N \in \mathbb{N}$ and $M > 0$ such that for all $n \geq N$ and $t > 0$, $\mathbb{P}(|X_n| \geq t) \leq 2 n\exp(-(t/(Mr_n))^{p}) + M n^{-1/2}$.
For a vector $\mathbf{v} \in \mathbb{R}^k$, the Euclidean norm is $\lVert \mathbf{v} \rVert = (\sum_{i = 1}^k \mathbf{v}_i^2)^{1/2}$,and the infinity norm is $\lVert \mathbf{v} \rVert_{\infty} = \max_{1 \leq i \leq k} |v_i|$. For a matrix $A = (a_{ij})_{i \in [m], j \in [n]} \in \mathbb{R}^{m \times n}$, the operator norm is $\lVert A \rVert = \lVert A \rVert_2 = \sup_{\lVert \mathbf{x} \rVert = 1} \lVert A\mathbf{x} \rVert$, the maximum absolute column sum norm is $\lVert A \rVert_1 = \sup_{1 \leq j \leq n} \sum_{i = 1}^m |a_{ij}|$, and the Frobenius norm is $\lVert A \rVert_F = \sqrt{\sum_{i = 1}^m \sum_{j = 1}^n a_{ij}^2}$. For sets $A$ and $B$, denote by $A \Delta B$ the set difference $(A \setminus B) \cup (B \setminus A)$.
$\operatorname{sgn}$ denotes the function such that $\operatorname{sgn}(x) = +$ if $x \geq 0$, and $\operatorname{sgn}(x) = -$ otherwise. $\Phi(x)$ denotes the standard Gaussian cumulative distribution function. For $\boldsymbol{\mu} \in \mathbb{R}^{k \times k}$ and $\boldsymbol{\Sigma} \in \mathbb{R}^{k \times k}$, $\mathsf{N}(\boldsymbol{\mu}, \boldsymbol{\Sigma})$ denotes the multivariate normal distribution with mean $\boldsymbol{\mu}$ and covariance matrix $\boldsymbol{\Sigma}$.
For notational simplicity, we consider $$\mathbf{W} = (W_i)_{1 \leq i \leq n}, \qquad W_i = 2 T_i - 1, 1 \leq i \leq n.$$ And we consider a more general setting compare to Assumption 3 in the main paper.
The Curie-Weiss model has a phase transition phenomena in different regimes. Let $\mca m = n^{-1}\sum_{i = 1}^n W_i$.
In the main paper, we focus on $(\beta, h)$ in $\mathscr{H}_1$. But for this section, we provide the results for all of $\mathscr{H}$, $\mathscr{C}$ and $\mathscr{L}$.
Suppose $\mathbf{X} = (X_1,\cdots,X_n)$ has i.i.d components such that $\mathbb{E} \left[|X_1|^3\right] < \infty$ independent to $\mathbf{W}$. The goal is to study the limiting distribution and the rate of convergence for
The magnetization $n^{-1}\sum_{i = 1}^n (W_i - \pi)$ has been studied using Stein's method eichelsbacher2010stein, chatterjee2010spin. Due to the multipliers, the Stein's method can not be directly applied for $\mca g_n$. We use a novel strategy based on the following de Finetti's lemma to show Berry Essseen results.
The de Finetti's theorem for exchangable sequences of random variable is a classical result diaconis1988recent, diaconis1980finetti, ellis1978statistics. For completeness, we include a short proof for the Curie-Weiss model in Section (ref).
Fix $\beta > 0$. We characterize the limiting distribution of $n^{-1}\sum_{i = 1}^n W_i X_i$ and the rate of convergence as $n \rightarrow \infty$ in the following lemma. In particular, we will see that the limiting distribution changes from a Gaussian distribution under high temperature, to a non-Gaussian distribution under critical temperature, to a Gaussian mixture under low temperature.
The magnetization $n^{-1}\sum_{i = 1}^n W_i$ has been studied using Stein's method eichelsbacher2010stein,chatterjee2010spin. Due to the multipliers, the Stein's method can not be directly applied to $n^{-1}\sum_{i = 1}^n X_i W_i$. We use a proof strategy based on the de Finetti's Lemma in Lemma (ref): There exists a latent variable $\mathsf{U}_n$ such that $W_1,\cdots,W_n$ are i.i.d condition on $\mathsf{U}_n$. Moreover, the density of $\mathsf{U}_n$ satisfies $f_{\mathsf{U}_n}(u) \propto \exp (- 1/2 u^2 + n \log \cosh( \sqrt{\beta/n} u)), u \in \mathbb{R}$.
We provide a proof sketch of Lemma (ref) (1) only. Throughout, take $c_{n, \beta} = \sqrt{n}(\beta - 1)$.
$W_i$'s are i.i.d condition on $\mathsf{U}_n$ with
Apply Berry-Esseen Theorem conditional on $\mathsf{U}_n$, and take $\mathsf{Z} \sim \mathsf{N}(0,1)$ independent to $\mathsf{U}_n$,
Lemma 2 in the supplementary material shows $\lVert \mathsf{U}_n \rVert_{\psi_2} \leq \mathtt{C}n^{1/4}$, hence by concentration arguments, $$\sup_{t \in \mathbb{R}}|\mathbb{P}(n^{-1}\sum_{i = 1}^n X_i W_i \leq t) - \mathbb{P}( \sqrt{v(\mathsf{U}_n)}\mathsf{Z} + \sqrt{n}e(\mathsf{U}_n)\leq t)|\leq K n^{-1/2}.$$
Consider $\mathsf{W}_n = n^{-1/4}\mathsf{U}_n$. By a change of variable from $\mathsf{U}_n$ and Taylor expand what is inside the exponent, we show $\mathsf{W}_n$ has density satisfying
where $g$ is a bounded smooth function. We show based on sub-Gaussianity of $\mathsf{W}_{n}$, with an upper bound of sub-Gaussian norm not depending on $\beta$, that the sixth order term is negligible and $$\sup_{t \in \mathbb{R}}|\mathbb{P}(\mathsf{W}_n \leq t) - \mathbb{P}(\mathsf{W} \leq t)| = O(\log^3 n n^{-1/2}),$$ where $\mathsf{W}$ has density proportional to $\exp (-c_{\beta,n}/2 w^2 - \beta_n^2 w^4 / 12)$.
Since $\mathsf{Z}$ is independent to $(\mathsf{U}_n, \mathsf{W}_n)$, we use data processing inequality and the previous two steps to show $n^{-1}\sum_{i = 1}^n X_i W_i$ is close to $n^{-1/4}v(n^{1/4}\mathsf{W}_{c_{\beta,n}})^{1/2}\mathsf{Z} + n^{1/4}e(n^{1/4}\mathsf{W}_{c_{\beta,n}}))$. Lemma 2 in the supplementary appendix imply $\lVert \mathsf{W}_{c_{\beta,n}} \rVert_{\psi_2} \leq \mathtt{K}$. By Taylor expanding $e(\cdot)$ and $v(\cdot)$ at $0$, we show $n^{1/4}e(\mathsf{U}_n)$ is close to $\mathbb{E}[X_i] \mathsf{W}_{c_{\beta,n}}$ and $n^{-1/4}\sqrt{v(\mathsf{U}_n)}\mathsf{Z}$ is close to $n^{-1/4}v(n^{1/4}\mathsf{W}_{c_{\beta,n}})^{1/2}\mathsf{Z}$.
The pseudo-likelihood estimator for Curie-Weiss regime with $h = 0$ is given by
Recall $\mathbf{W} = (W_1, \cdots, W_n)$ satisfies Assumption (ref). And for notational simplicity, let $g_i$ be the function such that
We denote $M_i = \sum_{j \neq i} E_{ij} W_i$, $N_i = \sum_{j \neq i} E_{ij}$. Then
Recall our definition of regimes: High temperature regime $\mathcal{A}_H = \{(\beta,h) \in [0,\infty) \times \mathbb{R}: h \neq 0 \text{ or } h = 0, \beta < 1\}$, critical temperature regime $\mathcal{A}_C = \{(1, 0)\}$, and low temperature regime $\mathcal{A}_L = \{(\beta,h) \in [0,\infty) \times \mathbb{R}: h = 0, \beta > 1\}$. Define the following rates that will be used in the convergence analysis:
and
Throughout Section (ref), we work with $(\beta, h) \in \mathcal{A}_H \cup \mathcal{A}_C$, and let $\pi$ be the unique solution to $x = \tanh(\beta x + h)$. Then friedli2017statistical implies $\mathbb{E}[W_i] = \pi + O(n^{-1})$. Let $\mca m = n^{-1} \sum_{i= 1}^n W_i$ and $\mca m_i = n^{-1} \sum_{j \neq i} W_j$.
Denote $p_i = \bb P_{\beta, h}(W_i = 1; \mathbf{W}_{-i}) = \left(\exp \left(-2\beta \mca m_i - 2 h\right) + 1\right)^{-1}$. We propose an unbiased estimator given by
We will show the followings have weak limits:
W.l.o.g, we analyse the error for treated data, the error for control data follows in the same way. First, decompose by
Now consider $\Delta_2$. Since $\frac{T_i}{p_i} = \frac{T_i - p_i}{p_i} + 1$, we have the decomposition,
where
where $\eta_i^{\ast}$ is some random quantity between $ \frac{M_i}{N_i}$ and $\pi$. Define $b_i = \sum_{j \neq i} \frac{E_{ij}}{N_j} Y_j^{\prime}\left(1, \pi \right)$. Then by reordering the terms,
For the term $\Delta_{2,3}$, we further decompose it into two parts:
where
Here we give a local-polynomial based learner $\wh f$ that satisfies requirements of Lemma (ref) (hence Theorem 4 in the main paper.)
This section presents the additional distributional results in the appendix. We continue to use the notations defined at the beginning of Section (ref).
Recall we consider a conditional estimand given by
where $\operatorname{sgn}(\mca m) = \operatorname{sgn}(2 n^{-1}\sum_{i = 1}^n T_i - 1)$. Let $\pi_*$ be the positive root of $x = \tanh(\beta x)$, and take $\pi_+ = 1/2 + \pi_*/2$, $\pi_- = 1/2 - \pi_*/2$.
Recall the following treatment assignment model from Section A.1: For $\beta \in [0,\infty)$ and $h \neq 0$, the treatment vector $\mathbf{T} = (T_1, \cdots, T_n)$ satisfies a distribution on $\{0,1\}^n$ such that
Let $\pi$ be the unique solution to $x = \tanh(\beta x + h)$.
Recall our notations: For block $k$ with $h_k \neq 0$ or $h_k = 0, 0 \leq \beta_k \leq 1$, $\pi_k$ denotes the unique solution to $x = \tanh(\beta_k x + h_k)$. For block $k$ with $h_k = 0, \beta_k > 1$, $\pi_{k,+}$ and $\pi_{k,-}$ denote the unique positive and negative solutions to $x = \tanh(\beta_k x + h_k)$, respectively.
Due to the potential existence of low temperature blocks, we use $\boldsymbol{sgn}$ to collect the average spins in all low temperature blocks, and fill in the positions for high and critical temperature blocks with zeros, that is,
And we use $\mathscr{S}$ to denote the collection of all possible configurations of $\boldsymbol{sgn}$, that is,
Also we denote the conditional fixed point based on $\boldsymbol{sgn} = \mathbf{s}$ by
We denote by $\mathscr{R}$ the collection of all hyperrectangles in $\mathbb{R}^K$.
The conclusion follows from the stochastic linearization result in Lemma (ref), and the Berry-Esseen result for Curie-Weiss magnetization with independent multipliers in Lemma (ref) (1) and (2).
The conclusion for Hajek estimator follows from the stochastic linearization result in Lemma (ref), and the (uniform in $\beta$) Berry-Esseen result for Curie-Weiss magnetization with independent multipliers in Lemma (ref) (1).
The conclusion for MPLE follows from Lemma (ref).
The conclusion follows from Lemma (ref) and Lemma (ref).
The uniform approximation for $\sqrt{n}(\wh \beta_n - 1)$ established in Lemma (ref) implies $$\inf_{\beta}\mathbb{P}_\beta(\beta \in \ca I(\alpha_1)) \geq \inf_{\beta}\mathbb{P}_\beta(\sqrt{n}(1 - \beta) \geq \mca q) \geq 1 - \alpha_1 + o_\mathbb{P}(1).$$ where $\mca q$ is the $\alpha_1$ quantile of $\min\{\max \{\mathsf{T}_{c_{\beta,n},n}^{-2} - \mathsf{T}_{c_{\beta,n},n}^2/(3n),0\},1\}$.
Then by a Bonferroni correction argument, the second step coverage can be lower bounded by
Observe that the event $\tau_n \in \widehat{\ca C}(\alpha_1,\alpha_2)$ conincides with the event $\wh \tau_n - \tau_n \in [\mathtt{L}, \mathtt{U}]$, where $\mathtt{U} = \sup_{\beta \in \ca I(\alpha_1)} H_n(1 - \frac{\alpha_2}{2};K_n, K_n,c_{\beta,n})$, $\mathtt{L} = \inf_{\beta \in \ca I(\alpha_1)} H_{n}(\frac{\alpha_2}{2};K_n, K_n, c_{\beta,n})$. Hence
Theorem 2 shows that the quantiles of the distributions of $\wh \tau_n - \tau_n$ can be uniformly approximated by quantiles from $H_n(\cdot;\kappa_1, \kappa_2,c_{\beta,n})$, if $\kappa_1$ and $\kappa_2$ are correctly specified, and the confidence interval is conservative, if we use upper bound $\mathtt{K}_n$ for $\kappa_1$ and $\kappa_2$. The conclusion then follows.
The conclusion follows from Theorem 4.1 and Lemma (ref).
The conclusion follows from Lemma (ref).
The conclusion follows from Lemma (ref).
The conclusion follows from Lemma (ref).
Using Gaussian integral identity $\exp(v^2/2) = \frac{1}{\sqrt{2 \pi}} \int_{-\infty}^{\infty} \exp \left( - u^2/2 + uv\right) du$,
Our proof is divided according to the different temperature regimes.
We introduce the handy notation given by $F(v):=-\frac{1}{2}v^2+\log\cosh(\sqrt{\beta}v+h)$. For the high temperature regime, we note that the term in the exponential can be expanded across its global minimum $v^*$ (which satisfies the first order stationary point condition given by $v^*=\sqrt{\beta}\tanh(\sqrt{\beta}v^*+h)$) by
Therefore, to obtain the limit of the expectation, we note that by the Laplace method given similar to the proof of Lemma (ref) and the definition of $\mathsf{V}_n:=n^{-1/2}\mathsf{U}_n$:
Then, we note that for $\ell\in\bb N$, when $h=0$ and $\beta<1$ we use the Laplace method again to obtain that for all $\ell\in\bb N$,
Then we can obtain that for all $t\in\bb R$, we have
which alternatively implies that
Then we study the critical temperature regime with $\beta=1$. Note that one has $\bb E[\mathsf{U}_n]=0$ and for all $\ell\in\bb N$ we have
Then we can obtain that $\ell\in\bb N$,
And we immediately obtain that
which finally leads to
We shall note that at the low temperature regime the function $F(v)$ has two symmetric global minima $v_1>0>v_2$, satisfying
Then we can check that by the Laplace method, for all $t>0$ (following the path given by the high temperature regime) we have
Then we similarly obtain that $ \bb E[\exp(t(\mathsf{V}_n-\bb E[\mathsf{V}_n|\mathsf{V}_n<0]))|\mathsf{V}_n<0]=\exp\left(\frac{(1+o(1))t^2}{2n(1-\sqrt{\beta}\operatorname{sech}^2(\sqrt{\beta}v_1))}\right)$. Hence we obtain that
Then we consider the drifting case.
First consider $\beta= 1-cn^{-\frac{1}{2}}$ with $c\in\bb R^+$ and $\beta \geq 0$. We will show that for any fixed $n$, $\lVert W_n \rVert_{\psi_2}$ is increasing in $\beta$ when $\beta \in [0,1]$. This will imply that in the drifting case, $\lVert W_n \rVert_{\psi_2}$ will be no larger than its value at the critical regime.
For a comparison argument, denote $F_{\beta}(v) = - \frac{1}{2}v^2 + \log \cosh(\sqrt{\beta} v)$. Let $0 < \beta_1 < \beta_2 \leq 1$. Then
where
Hence for any $n \in \mathbb{N}$ and $t > 0$,
increases as $\beta \in [0,1]$ increases. This shows that $\lVert W_n \rVert_{\psi_2}$ increases as $\beta \in [0,1]$ increases. Together with Equation (ref), we have under $\beta_n = 1 - \frac{c}{\sqrt{n}}$, $0 \leq c \leq \sqrt{n}$,
where $o(\cdot)$ is by an absolute constant.
Then we consider $\beta=1+cn^{-\frac{1}{2}}$. We shall note that under this situation it is not hard to check that
Then, under this case we have by Taylor expanding $F$ at $0$ and the fact that $\sup_{v \in \mathbb{R}}|F^{(5)}(v)| < \infty$,
Before we start to upper bound the moments, we first use the fact that $v_+=O(n^{-1/4})$ to obtain that
Then we obtain that
with $\mca C_3:=\frac{(3c)^{-1/2}}{3\int_{(-v_+,+\infty)}\exp\left(-cv^2-\frac{\sqrt{3c}}{3}v^3-\frac{1}{12}v^4\right)dv}$, $\mca C_4=\frac{1}{9\int_{(-v_+,+\infty)}\exp\left(-cv^2-\frac{\sqrt{3c}}{3}v^3-\frac{1}{12}v^4\right)dv}$,\\ and $\mca C_5=\frac{2^{-3/2}}{\int_{(-v_+,+\infty)}\exp\left(-cv^2-\frac{\sqrt{3c}}{3}v^3-\frac{1}{12}v^4\right)dv}$. Therefore, we can simply use the definition of the m.g.f. to obtain that
Then we use the fact that $\bb E[\mathsf{V}_n|\mathsf{V}_n>0]=v_+$ to obtain that (here we use proposition 2.5.2 in vershynin2018high)
Similarly one obtains that $ \bb E[\exp(t(\mathsf{V}_n-v_-))|\mathsf{V}_n<0]\leq\exp(18e^2n^{-1/2}\sigma^2t^2)$. And hence
We will leverage the representation of $\mathbf{W}$ as a mixture of independent Bernouli random variables after conditioning on some latent variable $\mathsf{U}_n$. We take $\mathsf{U}_n$ to be a random variable with density
Using Gaussian integral identity $\exp(v^2/2) = \frac{1}{\sqrt{2 \pi}} \int_{-\infty}^{\infty} \exp \left( - u^2/2 + uv\right) du$,
Hence condition on $\mathsf{U}_n$, $\mathbf{W}_i$ are i.i.d Bernouli with $\mathbb{P}(W_i = 1|\mathsf{U}_n) = \frac{1}{2} (\tanh (\sqrt{\frac{\beta}{n}} \mathsf{U}_n + h) + 1)$, and
Moreover,
\paragraph*{Step 1: Conditional Berry-Esseen} Apply Berry-Esseen Theorem conditional on $\mathsf{U}_n$,
Take $Z \sim N(0,1)$ independent to $\mathbf{W}$ and $X_i$'s. $\mathsf{U}_n$ is sub-Gaussian by Equation (ref), hence
\paragraph*{Step 2: Stabilization of Variance} By independence between $\mathsf{U}_n$ and $Z$, we have
where $v^{\ast}(\mathsf{U}_n)$ is some quantity between $\mathbb{E}[v(\mathsf{U}_n)]$ and $v(\mathsf{U}_n)$, and by Equation (ref), $v^{\ast}(\mathsf{U}_n) \geq C_2 \mathbb{V}[X_i]$. It follows from boundedness of $v(\mathsf{U}_n)$ and Lipshitzness of $\tanh$ in the expression of $v(\mathsf{U}_n)$ that
\paragraph*{Step 3: Reduction Through TV-distance Inequality}
where $b_n = \sqrt{n}v_0$. The first inequality is by relation between KS- and TV-distances. For the second inequality, denote $X = \sqrt{n}e(\mathsf{U}_n)$, $Y = \sqrt{n}e(b_n + \mathsf{U})$. Denote by $f_X, f_Y, f_Z$ the Lebesgue density of $X, Y, Z$ respectively. Then using $Z \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} X$ and $Z \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} Y$, by data processing inequality,
Above proves inequality (2). Inequality (3) is by scale-invariance of TV distance and data processing inequality.
\paragraph*{Step 4: Gaussian Approximation for $\mathsf{U}_n$} Consider $\mathsf{V}_n = n^{-1/2} \mathsf{U}_n$. Then
where $\phi(v) = - \frac{1}{2} v^2 + \log \cosh (\sqrt{\beta} v + h)$. $\phi$ is maximized at $v_0$ that solves
We will approximate the integral of $f_{\mathsf{V}_n}$ by Laplace method. We will introduce constants $c_0, c_1$ and $c_2$ that only depends on $\beta$ and $h$. By Equation (5.1.21) in bleistein1975asymptotic,
where the $O(n^{-1})$ term only depends on $n$ and $\phi$. It follows that
Then by a change of variable and the fact that $O(n^{-1})$ term does not depend on $v$,
Taylor expanding $\phi$ at $v_0 = n^{-1/2}u_0$ and using $\phi^{\prime}(v_0) = 0$, we get
where $v_{\ast}$ is some quantity between $v_0$ and $n^{-1/2}u$. Now take $b_n = u_0 = \sqrt{n}v_0$ and take $\mathsf{U} \sim N(0, (1 - \beta + v_0^2)^{-1})$, we have
where $v^{\ast}(u)$ is some random quantity between $v_0 = n^{-1/2}u_0$ and $n^{-1/2} u$. We will show that we can restrict the analysis to the region $[u_0 - c_0\sqrt{\log n}, u_0 + c_0 \sqrt{\log n}]$, which is where the bulk of mass lies. Since $\mathsf{U} \sim N(u_0, (1 - \beta + v_0^2)^{-1})$, for some constant $c$ only depending on $\beta$ and $h$, $\mathbb{P} \left(\left|b_n + \mathsf{U} - u_0 \right| \geq c \sqrt{\log n} \right) \leq n^{-1}$. Using a change of variable and concavity of $\phi$,
In the third line we used the fact that $\phi(v_0 + t) - \phi(v_0) = \int_{0}^t \phi^{\prime}(v_0 + s) d s$ and the first derivative is bounded by
where $w_0$ is the solution to $\tanh(\sqrt{\beta} v_0 + h) - \tanh(w_0) = (\sqrt{\beta}v_0 + h - w_0) \operatorname{sech}^2(w_0)$. It follows that $\phi(v) - \phi(v_0) \leq - \frac{1}{2} (1 - \operatorname{sech}(w_0)^2)(v - v_0)^2$. Using boundedness of $\tanh$ and $\operatorname{sech}$ and the Lipschitzness of $\exp$ when restricted to $[-1,1]$, we have
\paragraph*{Step 5: Gaussian Approximation for $\sqrt{n}e(b_n + \mathsf{U})$} In this step, we will show that $\sqrt{n}e(b_n + \mathsf{U})$ can be well-approximated by $\sqrt{\beta} \operatorname{sech}^2(\sqrt{\beta} v_0 + h)\mathsf{U}$ and hence $\frac{1}{\sqrt{n}} \sum_{i = 1}^n X_i(W_i - \pi)$ can be well-approximated by a Gaussian.
Since $d_{\operatorname{KS}}(\mathsf{U}_n, \mathsf{U}) = O(n^{-1/2})$ and $\pi = \mathbb{E} \left[\tanh \left( \sqrt{\frac{\beta}{n}}\mathsf{U}_n + h \right)\right]$, Taylor expanding $\tanh$ at $\sqrt{\beta}v_0 + h$,
It follows that $\mathbb{E} \left[\left|\sqrt{n}e(b_n + \mathsf{U}) - \sqrt{\beta} (1 - \frac{v_0^2}{\beta})\mathsf{U} \right|\right] = O(n^{-1/2})$ and hence
Recall $\mathsf{U} \sim N(0, (1 - \beta + v_0^2)^{-1})$, hence $\mathbb{E}[X_i] \sqrt{\beta}(1 - \frac{v_0^2}{\beta})\mathsf{U} \sim N(0, \mathbb{E}[X_i]^2\frac{(\beta - v_0)^2}{\beta (1 - \beta + v_0^2)})$. Moreover,
where the last line is because $\mathbb{E}[W_i|\mathsf{U}_n] = \tanh(\sqrt{\beta/n}\mathsf{U}_n)$ and $\mathsf{U}_n$ is sub-Gaussian. Since $Z \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} \mathsf{U}$,
Combining the previous five steps, we get
Throughout the proof, we denote by $\mathtt{C}$ an absolute constant, and $\mathtt{K}$ a constant that only depends on the distribution of $X_i$. The proofs for the critical temperature case will have a similar structure as the proof for the high temperature case, based the same $\mathsf{U}_n$ defined in Equation (ref).
Step 1: Conditional Berry-Esseen.
The same argument as in the high-temperature case gives
Step 2: Approximation for $\mathsf{U}_n$.
Take $\mathsf{W}$ to be a random variable with density function
independent to $\mathsf{Z}$. Take $\mathsf{W}_n = n^{-1/4}\mathsf{U}_n$ and $\mathsf{V}_n = n^{-1/2}\mathsf{U}_n$. Again $f_{\mathsf{V}_n}(v) \propto \exp(-n \phi(v))$, where $\phi(v) := - \frac{1}{2} v^2 + \log \cosh (v)$. In particular, $\phi^{(v)}(0) = 0$ for all $0 \leq v \leq 3$, and $\phi^{(4)}(0) = -2 < 0$, $\phi^{(5)}(0) = 0$, $\phi^{(6)}(0) = 16 > 0$. Example 5.2.1 in bleistein1975asymptotic leads to $$f_{\mathsf{V}_n}(v) = n^{\frac{1}{4}}\frac{\sqrt{2}}{3^{\frac{1}{4}}\Gamma(\frac{1}{4})}\exp(n \phi(v) - n \phi(0))(1 + o(1)),$$ which implies $f_{\mathsf{W}_n}(w) = f_{W}(w)(1 + o(1))$. Results in bleistein1975asymptotic do not give a rate, however. We will use a more cumbersome approach to obtain a slightly sub-optimal rate.
By a change of variable, $f_{\mathsf{W}_n}(w) = \frac{h_n(w)}{\int_{-\infty}^{\infty}h_n(u)du}$, where $h_n$ can be written as
The last equality follows from Taylor expanding the term in $\exp(\cdot)$ at $w = 0$, and $g$ is some bounded function.
Moreover, $\int_{[-10\sqrt{\log n}, 10\sqrt{\log n}]^c} h_n(w) d w = O(n^{-1/2}) = I_n [1 + O(n^{-\frac{1}{2}})]$. Hence for denominator, we have $\int_{-\infty}^{\infty} h_n(w)dw =I_n[1 + O((\log n)^3 n^{-\frac{1}{2}})]$. It follows that
Step 3: Data Processing Inequality.
We can use data processing inequality to get
\paragraph*{Step 4: Non-Gaussian Approximation for $n^{\frac{1}{4}}e(n^{\frac{1}{4}}\mathsf{W})$}
where we have use the fact that $\tanh^{(2)}(0) = 0$. Hence there exists $C >0$ such that for $n$ large enough, for any $t > 0$,
We have showed that there exists $c > 0$ such that
in which case $\mathsf{W}^2/\sqrt{n} \leq 1$ for large enough $n$. Hence for large enough $n$ if $t/\mathbb{E}[X_i] > c \sqrt{\log n} + 1$, then
If $0 < t/\mathbb{E}[X_i] < c\sqrt{\log n} + 1$, then
Now we study $g(x; \alpha) = (1 - \sqrt{1 - 4 x \alpha})/(2x), x >0$. Then $\sup_{\alpha \leq \frac{1}{4}} \sup_{0 \leq x \leq \frac{1}{2}}|\theta^{\prime}(x; \alpha)| \leq 2$ and $g(0;\alpha) = \alpha$. Since for large enough $n$, $0 < t/\mathbb{E}[X_i] < c\sqrt{\log n} + 1 \leq \frac{1}{4}$ and $0 \leq n^{-1/2} \leq \frac{1}{2}$, we have $\frac{1 - \sqrt{1 - 4 n^{-1/2} t/\mathbb{E}[X_i]}}{2 n^{-1/2}} \leq t/\mathbb{E}[X_i] + 2 n^{-1/2}$. Hence if $0 < t/\mathbb{E}[X_i] < c\sqrt{\log n} + 1$,
Combining (1), (2), (3),
By similar argument, we can show
Noticing that $W$ and $-W$ have the same distribution, the above two inequalities also hold for $t \leq 0$. Hence it follows from (0) that
\paragraph*{Step 5: Vanishing Variance Term.}
Denote by $f_{\mathsf{W} + n^{-1/4}\mathsf{Z}}$ the density of $\mathsf{W} + n^{-1/4}\mathsf{Z}$. Then
We will use Laplace method to show $f_{\mathsf{W} + n^{-1/4}\mathsf{\mathsf{Z}}}$ is close to $f_{\mathsf{W}}$. However, to get uniformity over $y$, we need to work harder than in the high temperature case. Define $\varphi(x) = x^2/2$ and $g_y(t) =\exp(-(t-y)^4/12)$. Consider
Following Section 5.1 in bleistein1975asymptotic, take $\tau > 0$ such that $\varphi(t) = \tau $, by a change of variable,
To get rate of convergence uniformly in $y$, we follow the proof of Watson's Lemma but consider only up to first order term. Taylor expanding $x \mapsto \exp(-x ^4)/12$ up to first order at $y$, we have
where $\tau^{\ast}$ is some quantity between $0$ and $\sqrt{2\tau}$ and
In particular, we have $\sup_{y \in \mathbb{R}} \sup_{u \in \mathbb{R}}|h_y(u)| < C$ for some absolute constant $C$. Then
Evaluating the first two terms, we get
Similarly, for $I_{y,-}$, change of variable by taking $\tau < 0$ such that $\varphi(t) = \tau$, we have
Combining the two parts, we get
Now take $\lambda = \sqrt{n}$ and multiply both sides by $\frac{n^{1/4}}{3^{1/4}\Gamma(\frac{1}{4})\sqrt{\pi}}$, we get
By a truncation argument, we have
Together with the fact that
we know
Putting together all previous steps, we have
Throughout the proof, we denote by $\mathtt{C}$ an absolute constant, and $\mathtt{K}$ a constant that only depends on the distribution of $X_i$. The proofs are based on essentially the same argument as in the high temperature case.
Instead of using sub-Gaussianity of $\mathsf{U}_n$, here we use $\mathsf{U}_n$ is sub-Gaussian condition on $\mathsf{U}_n \in \ca I_\ell$, $\ell \in\{-,+\}$. In particular, the previous step 2 by:
Step 2: Approximation for $\mathsf{U}_n$.
In case $\beta > 1$, $\phi(v) = \frac{1}{2}v^2 - \log(\cosh(\sqrt{\beta}v))$ has two global minimum $v_+$ and $v_-$, which are the two solutions of $v -\sqrt{\beta}\tanh(\sqrt{\beta}v) = 0$. We want to show $\phi^{(2)}(v_+) = \phi^{(2)}(v_-) = 1 - \beta + v_+^2 > 0$. It sufffices to show $v_+ > \sqrt{\beta - 1}$. Since $\phi^{\prime}(v) < 0$ for $v \in (0,v_+)$ and $\phi^{\prime}(v) > 0$ for $v \in (v_+,\infty)$, it suffices to show $\phi^{\prime}(\sqrt{\beta - 1}) < 0$. But
Hence $\phi^{(2)}(v_+) = \phi^{(2)}(v_-) > 0$. Observe that on $\ca I_- = (-\infty,0)$ and $\ca I_+ = (0,\infty)$ respectively, the absolute minimum of $\phi$ occurs at $v_-$ and $v_+$, and $\phi^{\prime}$ is non-zero on $\ca I_-$ and $\ca I_+$ except at $v_-$ and $v_+$. Hence we can apply Laplace method (Equation 5.1.21 in bleistein1975asymptotic) sperarately on $\ca I_-$ and $\ca I_+$ to get
It follows from the definition of $f_{\mathsf{V}_n}$ and a change of variable that the density of $\mathsf{U}_n = \sqrt{n} \mathsf{V}_n$ can be approximated by
where $u_l = \sqrt{n} v_l, l \in \{ +,- \}$. Since $\mathbb{P}(\mathsf{U}_n \in \ca I_+) = \mathbb{P}(\mathsf{U}_n \in \ca I_-) = \frac{1}{2}$, condition on $\mathsf{U}_n \in \ca I_+$,
It then follows from Equation (ref) that if we define $\mathsf{U}_+$ to be a random variable with density
then by Taylor expanding $\phi$ at $v_+ = n^{-1/2}u_+$ and a similar argument as in the proof for high temperature case,
The rest follows from the same argument as in the proof for high temperature case and is sub-Gaussianity of $\mathsf{U}_n$ condition on $\mathsf{U}_n \in \ca I_\ell$, $\ell \in\{-,+\}$.
Take $\mathsf{V}_n = n^{-1/2}\mathsf{U}_n$ and $\mathsf{Z}$ be a $\mathsf{N}(0,1)$ variable independent to $\mca m$ and $\mathsf{U}_n$. Based on the conditional mean and variance formulas in Equation (ref), using the conditional on $\mathsf{U}_n$ Berry-Esseen bound,
Using the fact that $v(\mathsf{U}_n)$ is bounded above and $\mathsf{Z}$ is Gaussian, and Taylor expanding $\tanh$, we get
The proof of Lemma (ref) (low temperature) shows that $ d_{\operatorname{TV}}(\mathsf{U}_n|\mathsf{U}_n \in \ca I_+, \mathsf{U}_+) = O(n^{-1/2})$ where $\mathsf{U}_+ \sim \mathsf{N}(\sqrt{n}\pi_+, (1 - \beta(1 - \pi_+^2))^{-1})$. Hence $\mathbb{P} (\mathsf{U}_n \leq \sqrt{\log n} | \mathsf{U}_n \geq 0) \lesssim \exp(-n)$. It follows that
By symmetry and the fact that $\mathbb{P}(\operatorname{sgn}(\mca m) = \ell) = \mathbb{P}(\operatorname{sgn}(\mathsf{U}_n) = \ell) = 1/2$ for $\ell = -, +$, we know
The conclusion then follows from Lemma (ref)(3).
Throughout the proof, we denote by $\mathtt{C}$ an absolute constant, and $\mathtt{K}$ a constant that only depends on the distribution of $X_i$.
Let $\mathsf{U}_n(c)$, $e(\mathsf{U}_n(c))$, $v(\mathsf{U}_n(c))$ be the latent variable, conditional mean, and conditional variance as previously defined when $\beta_n = 1 + c n^{-\frac{1}{2}}$, $c < 0$. For notational simplicity, we abbreviate the $c$, and call them $\mathsf{U}_n, e(\mathsf{U}_n), v(\mathsf{U}_n)$ respectively. By Lemma (ref), $\lVert \mathsf{U}_n \rVert_{\psi_2} \leq \mathtt{C} n^{1/4}$.
Step 1: Conditional Berry-Esseen.
Apply Berry-Esseen Theorem conditional on $\mathsf{U}_n$ in the same way as in the high temperature case, we get
Step 2: Non-Normal Approximation for $n^{-\frac{1}{4}}\mathsf{U}_n$.
Consider $\mathsf{W}_n = n^{-1/4}\mathsf{U}_n$. Then $f_{\mathsf{W}_n}(w) = I_n(c)^{-1} h_n(w)$, with $I_n(c) = \int_{-\infty}^{\infty}h_n(w)dw$, and
where by smoothness of $\log(\cosh(\cdot))$, $\lVert \theta \rVert_{\infty} \leq \mathtt{K}$. Then
Moreover, by a change of variable and the fact that $\beta_n \leq 1$,
Since $\lVert \mathsf{W}_n(c) \rVert_{\psi_2} \leq \mathtt{C}$, $I_n(c)^{-1}\int_{(-\mathtt{C}\sqrt{\log n},\mathtt{C} \sqrt{\log n})^c}h_n(w) dw \leq \mathtt{C} n^{-1/2}$. It follows that
Combining Equation (ref) and (ref), we have $I_n(c) =I(c)[1 + O(\mathtt{C}^6(\log n)^3 n^{-1/2})]$. It follows that
Step 3: A Reduction through TV-distance Inequality.
Since $\mathsf{Z} \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} (\mathsf{U}_n, \mathsf{W}_n)$, we can use data processing inequality to get
Step 4: Non-Gaussian Approximation for $n^{\frac{1}{4}}e(n^{\frac{1}{4}}\mathsf{W})$.
This is essentially the same as the proof for step 4 from the critical temperature case in Lemma (ref).
Step 5: Stabilization of Variance.
Using the same argument as Step 4 in the high temperature case for Lemma (ref), and $\lVert \mathsf{W} \rVert \leq \mathtt{K}$,
The conclusion then follows from putting together the previous five steps.
Consider the same $\mathsf{U}_n$ defined in Equation (ref). Recall $\phi(v) = \frac{v^2}{2} - \log \cosh(\sqrt{\beta_n}v)$, $\phi^{\prime}(v) = v - \sqrt{\beta_n} \tanh(\sqrt{\beta_n} v)$, $\phi^{(2)}(v) = 1 - \beta_n \operatorname{sech}^2(\sqrt{\beta_n}v)$. And we take $v_{n,+} > 0$, $v_{n,-} < 0$ to be the two solutions of $v - \sqrt{\beta_n}\tanh(\sqrt{\beta_n}v) = 0$.
Step 2': Non-Normal Approximation for $n^{-\frac{1}{4}}\mathsf{U}_n$.
Take $\mathsf{V}_n = n^{-1/2}\mathsf{U}_n$. Then $f_{\mathsf{V}_n}(v) \propto \exp(-n \phi(v))$. Taylor expanding $\phi^{\prime}$ at $0$, we know there exists some function $g$ that is uniformly bounded such that $\phi^{\prime}(v) = (1 - \beta_n) v + \frac{1}{3}\beta_n^2 v^3 + \beta_n^3 g(v) v^5$. Hence $$v_{n,+} = \sqrt{\frac{3(\beta_n - 1)}{\beta_n^2}} + O(\beta_n - 1) = \sqrt{3 c}n^{-1/4} + O(n^{-1/2}).$$ Taylor expand $\tanh$ and $\operatorname{sech}$ at $0$,
Take $\mathsf{W}_n = n^{1/4}\mathsf{V}_n = n^{-1/4}\mathsf{U}_n$, $\mca w_+ = n^{1/4}v_{n,+} = \sqrt{3 c} + O(n^{-1/4})$, and $\mca w_- = n^{1/4} v_{n,-}$. Define
By a change of variable and Taylor expansion, the density for $\mathsf{W}_n$ satisfies
By Lemma (ref), for $\ell \in \{-,+\}$, condition on $\mathsf{W}_n \in \ca I_{c,n,\ell}$, $\mathsf{W}_n - \mca w_\ell$ is sub-Gaussian with $\psi_2$-norm bounded by $\mathtt{C}$. Let $\mathsf{W}_{c,n}$ be a random variable with density at $w$ proportional to $\exp(h_{c,n}(w))$. By similar argument as Equations (ref) and (ref), $$d_{\operatorname{KS}}(\mathsf{W}_n| \mathsf{W}_n \in \ca I_{c,n,\ell}, \mathsf{W}_{c,n}| \mathsf{W}_{c,n} \in \ca I_{c,n,\ell}) \leq \mathtt{C} (\log n)^3 n^{-1/2}).$$
The other steps, conditional Berry-Esseen, reduction through TV-distance inequality, and non-Gaussian approximation for $n^{\frac{1}{4}}e(n^{\frac{1}{4}}\mathsf{W}_{c,n})$ can be proceeded in the same way as in the proof for Lemma (ref), with $\mathsf{W}_n - \mca w_\ell$ sub-Gaussian condition on $\mathsf{W}_n \in \ca I_{c,n,\ell}$ with $\psi_2$-norm bounded by $\mathtt{C}$, and respectively for $\mathsf{W}_{c,n}$.
Again we take $\mathsf{U}_n$ to be the latent variable from Lemma (ref), and $\mathsf{W}_n = n^{-1/4}\mathsf{U}_n$. From Step 2 in the proof of Lemma (ref), $f_{\mathsf{W}_n}(w) = I_n(c)^{-1} h_n(w)$, with $I_n(c) = \int_{-\infty}^{\infty}h_n(w)dw$, and
where by smoothness of $\log(\cosh(\cdot))$, $\lVert \theta \rVert_{\infty} \leq \mathtt{K}$.
Case 1: When $\sqrt{n}(\beta_n - 1) = o(1)$. We can apply Berry-Esseen conditional on $\mathsf{U}_n$ the same way as in the proof of Lemma (ref), and its Step 2 can also be applied here to show that if we take $\widetilde{\mathsf{W}}_c$ to be a random variable with density proportional to $\exp(-c_n^2/2 w^2 - \beta_n^2/12 w^4)$, then $d_{\operatorname{KS}}(\mathsf{W}_n,\widetilde{\mathsf{W}}_c) = O((\log n)^3 n^{-1/2})$. Moreover, $c_n = o(1)$ and $\beta_n = 1 - o(1)$. Hence $d_{\operatorname{KS}}(\mathsf{W}_n,\mathsf{W}_0) = o(1)$. The rest of the proof then follows from Step 3 to Step 5 in the proof for the critical regime case in Lemma (ref).
Case 2: When $\sqrt{n}(1 - \beta_n) \gg 1$. Again we still have $\lVert \mathsf{U}_n \rVert_{\psi_2} = O(n^{1/4})$. And we take $v_+ > 0$, $v_- < 0$ to be the two solutions of $v - \sqrt{\beta_n}\tanh(\sqrt{\beta_n}v) = 0$. Similarly as in the previous case, the first two steps in the proof of Lemma (ref) implies $d_{\operatorname{KS}}(\mathsf{W}_n,\widetilde{\mathsf{W}}_c) = o(1)$, where the density of $\mathsf{W}_c$ is proportional to $\exp(-c_n^2/2w^2 - \beta_n^2/12 w^4)$. Since $c_n \gg 1$, the first term in the exponent dominates, and we can show $d_{\operatorname{KS}}(\mathsf{W}_n, \mathsf{W}_c^{\dag}) = o(1)$, where $\mathsf{W}_c^\dag$ has density proportional to $\exp(-c_n^2/2 w^2)$. Again, we can Taylor expand to get $n^{1/4}e(n^{1/4}\mathsf{W})) =\mathbb{E}[X_i] n^{\frac{1}{4}}\tanh\left(n^{-\frac{1}{4}}\mathsf{W}\right) = \mathbb{E}[X_i] [\mathsf{W} - O (\frac{\mathsf{W}^2}{3\sqrt{n}})]$, and show $d_{\operatorname{KS}}(n^{1/4}e(n^{1/4}\mathsf{W}_c^\dag), \mathbb{E}[X_i]\mathsf{W}_c^\dag) = o(1)$. Combining with stablization of variance as in the proof of Lemma (ref) (high temperature case), we can show $$d_{\operatorname{KS}}(\mca g_n, n^{-1/4}\mathbb{E}[X_i^2]^{1/2}\mathsf{Z} + \mathbb{E}[X_i]\mathsf{W}_c^\dag) = o(1).$$ Since $\mathsf{Z}$ and $\mathsf{W}_c^\dag$ are independent Gaussian random variables, we also have $d_{\operatorname{KS}} (\mca g_n/\sqrt{\mathbb{V}[\mca g_n]}, \mathsf{Z}) = o(1)$.
Case 3: When $\sqrt{n}(\beta_n - 1) \gg 1$. By Lemma (ref) (2),
where $\mathsf{W}_{c,n}$ has density proportional to $\exp(h_{c,n}(w))$, with
and $\ca I_{c,n,-} = (-\infty,K_{c,n,-})$ and $\ca I_{c,n,+} = (K_{c,n,+},\infty)$ such that $\mathbb{E}[\mathsf{W}_{c,n}|\mathsf{W}_{c,n} \in \ca I_{c,n,\ell}] = w_{c,n,\ell}$ for $\ell \in \{-,+\}$. Now we calculate the order of the coefficients under $\sqrt{n}(\beta_n - 1) \gg 1$. First, suppose $\beta_n = 1 + c n^{\gamma}$ for some $\gamma \in (0,\infty)$ and $c$ not depending on $n$. Then $v_+ = \sqrt{\frac{3(\beta_n - 1)}{\beta_n^2}} + O(\beta_n - 1) = \sqrt{3 c}n^{-\gamma/2} + O(n^{-\gamma})$. Taylor expand $\tanh$ and $\operatorname{sech}$ at $0$,
We see when $\gamma = 1/2$, all of $\sqrt{n}\phi^{(2)}(v_+)$, $n^{1/4}\phi^{(3)}(v_+)$ and $\phi^{(4)}(v_+)$ are of order 1. And when $c_n = \sqrt{n}(\beta_n - 1) \gg 1$, we have $\sqrt{n}\phi^{(2)}(v_+) \gg n^{1/4}\phi^{(3)}(v_+) \gg \phi^{(4)}(v_+)$. Since $w_+ = n^{1/4} v_+ = \sqrt{3 c_n} \gg 1$, and similarly, $|w_-| \gg 1$, condition on $\mathsf{W}_{c,n} \in [n]$, $\mathsf{W}_{c,n} - \mathbb{E}[\mathsf{W}_{c,n}|\mathsf{W}_{c,n} \in [n]]$ is $\mathtt{C}$-sub-Gaussian, $\ell \in \{-,+\}$. By similar concentration arguments as in the proof for Step 2 in Lemma (ref) (1), we can show the second order term in $h_{c,n}$ dominates, and for $\ell \in \{-,+\}$,
The conclusion then follows from pluggin the (conditional) Gaussian approximation for $\mathsf{W}_{c_n,n}$ back into Equation (ref), and the fact that $\mathsf{Z}$ is independent to $\mathsf{W}_{c,n}$ and also Gaussian.
Throughout the proof, the Ising spins $\mathbf{W}=(W_i)_{i=1}^n$ are distributed according to Assumption (ref) with parameters $(\beta,h)$. For brevity, we write $\mathbb{P}$ in place of $\mathbb{P}_{\beta,h}$.
Let $\mathsf{U}_n$ be the latent random variable from Lemma (ref). Condition on $\mathsf{U}_n$, $\mathbf{X}_i W_i$'s are i.i.d random vectors. For $u \in \mathbb{R}$, define
and to save notations, we denote
Suppose $\mathsf{Z}_d \sim \mathsf{N}(\mathbf{0}, \mathbf{I}_{d \times d})$ independent to $\mathsf{U}_n$. By chernozhukov2017central
From the proofs of Lemma (ref), we know the term $t(\mathsf{U}_n)$ stabilizes, $$d_{\operatorname{KS}} \Big(t(\mathsf{U}_n), \sigma \mathsf{Z}\Big) = O(n^{-1/2}), \qquad \sigma = \Big(\frac{\beta(1 - \pi^2)^2}{1 - \beta(1 - \pi^2)}\Big)^{1/2}.$$ By Lemma (ref),
For each $\varepsilon$, define $A_{\varepsilon}$ to be the event $\{\lVert \Sigma(\mathsf{U}_n)^{1/2} - \Sigma^{1/2})\mathsf{Z}_d \rVert \leq \varepsilon\}$. Since $d$ is fixed, we can work with each dimension to get
where in the last line, we have chosen $\varepsilon = n^{-1/2}\sqrt{\log n}$ and used Nazarov's inequality (Lemma A.1 in chernozhukov2017central). Since $\mathsf{Z}_d$ and $\mathsf{U}_n$ are independent, we can show via data processing inequality that
Combining the previous results,
We still have conditional Berry-Esseen as in Equation (ref). The proof of Lemma (ref) implies $$d_{\operatorname{KS}}(n^{-1/4}t(\mathsf{U}_n),\mathsf{R}) = O(n^{ -1/2}).$$ Hence $\lVert \Sigma(\mathsf{U}_n) - \mathbb{E}[\Sigma(\mathsf{U}_n)] \rVert_{\operatorname{max}} = O_{\psi,2}(n^{-1/4})$. By concentration of $\mathsf{U}_n$, approximation of $n^{-1/4}t(\mathsf{U}_n)$ by $\mathsf{R}$, and anti-concentration of $\mathsf{R}$, we can use similar arguments as Equation (ref) to get
By independence between $\mathsf{Z}_d$ and $\mathsf{U}_n$, and approximation of $n^{-1/4}t(\mathsf{U}_n)$ by $\mathsf{R}$, we can use data processing inequality to get
It follows that
We still have conditional Berry-Esseen as in Equation (ref). From the proof of Lemma (ref) (3) and Remark (ref), for $\ell = -,+$, $$d_{\operatorname{KS}}(t(\mathsf{U}_n) - \sqrt{n} \pi_{\ell}|\operatorname{sgn}(\mca m) = \ell, \sigma \mathsf{Z}) = O(n^{-1/2}),$$ where $\sigma^2 = \frac{\beta(1 - \pi_+^2)^2}{1 - \beta(1 - \pi_+^2)}$. The rest of the proof follows from the arguments for I. High Temperature or Nonzero External Field, using conditional concentration of $\mathsf{U}_n$ given $\operatorname{sgn}(\mca m)$.
Our proof is constructive. We show that consistent estimate of $n\mathbb{V}[\widehat{\tau}_n]$ would imply that one can distinguish between two constructed hypotheses easily. Let $\mathcal{P}_n$ be the class of distributions of random vectors $(\mathbf{W} = (W_1, \cdots, W_n), \mathbf{Y} = (Y_1, \cdots, Y_n))$ taking values in $\mathbb{R}^{2n}$ that satisfies Assumptions 1,2,3. Consider the following two data generating processes:
where $0 < u < 1$, and in both cases $(\varepsilon_i: 1 \leq i \leq n)$ are i.i.d $\mathsf{N}(0,1)$ random variables, independent to $\mathbf{W}$. Denote by $\mathbb{P}_{0,n}$ and $\mathbb{P}_{1,n}$ the laws of $(\mathbf{W},\mathbf{Y})$ under $\text{DGP}_0$ and $\text{DGP}_1$. Then {
} the first line uses chain rule of $d_{\text{KL}}$, the second line uses $$d_{\text{KL}}(\mathbb{P}_{0,n}(\mathbf{Y}|\mathbf{W}),\mathbb{P}_{1,n}(\mathbf{Y}|\mathbf{W})) = d_{\text{KL}}(\mathbb{P}_{0,n}(\mathbf{Y}),\mathbb{P}_{1,n}(\mathbf{Y})) = 0.$$ From Theorem 2.3 (and its proof) in bhattacharya2018inference, $$M := \lim_{n \rightarrow \infty} d_{\text{KL}}(\mathbb{P}_{0,n}(\mathbf{W}),\mathbb{P}_{1,n}(\mathbf{W})) < \infty.$$ Hence for large enough $n$,
Le Cam's method (Section 15.2.1 in wainwright2019high) gives for large enough $n$,
in the last line we used Theorem 2 (1) to get $n \mathbb{V}_{\mathbb{P}_{n,0}}[\widehat \tau - \tau] - n \mathbb{V}_{\mathbb{P}_{n,1}}[\widehat \tau - \tau] = \varepsilon (1 + o(1))$.
The following discussions will be organized according to the three different cases: (1) When $\beta<1$. (2) When $\beta\geq 1$, $\mca m$ concentrates around $0$. (3) When $\beta\geq 1$ and $\mca m$ concentrates around two symmetric locations $w_{+}>0$ and $w_{-}<0$ with $|w_{+}|=|w_{-}|$.
We have required $\wh \beta \in [0,1]$. For analysis, consider an unrestricted pseud-likelihood estimator,
where $l(\beta;\mathbf{W})$ is the pseudo log-likelihood given by
We show that $l(\beta;\mathbf{W})$ is concave.
and
Hence $l(\cdot;\mathbf{W})$ is concave everywhere in $\mathbb{R}$. This shows $\wh \beta = \min\{\max\{\widehat{\beta}_\text{UR},0\},1\}$. Now we study limiting distribution of $\widehat{\beta}_\text{UR}$
To obtain a more precise distribution for $\widehat{\beta}_\text{UR}$, we use Fermat's condition to obtain that
here $O(\cdot)$'s are all up to an absolute constant. By Lemma (ref) with $X_i = 1$, we can show $\mathbb{E}[|(n \mca m)^{-1}|] \leq \mathtt{C} n^{-1/2}$. By Markov inequality, $(n \mca m)^{-1} = O_\mathbb{P}(n^{-1/2})$. Taylor expanding $\tanh$, we have
where in the above equation, both $O(\cdot)$ and $O_\mathbb{P}(\cdot)$ are up to absolute constants. The rest of the results are given according to the different temperature regimes.
(1) The High Temperature Regime. Using Lemma (ref) with $X_i = 1$, our result for the high temperature regime with $\beta<1$ implies that $n^{\frac{1}{2}}\mca m\overset{d}{\to}\mathsf{N}( 0,\frac{1}{1-\beta}) \Rightarrow (1-\beta)n\mca m^2\overset{d}{\to} \chi^2(1)$. Therefore we conclude that $\frac{1-\beta}{1-\widehat{\beta}_\text{UR}}\overset{d}{\to}\chi^2(1)$. The conclusion then follows from $\wh \beta = \min\{\max\{\widehat{\beta}_\text{UR},0\},1\}$.
(2) The Critical Temperature Regime. Using Lemma (ref) with $X_i = 1$, we have $d_{\operatorname{KS}}(n^{\frac{1}{4}}\mca m, \mathsf{W}_0) = o(1)$. This implies $n^{\frac{1}{2}}(\widehat{\beta}_\text{UR}-1) \overset{d}{\to} Law(\frac{\mathsf{W}_0^2}{3}-\frac{1}{\mathsf{W}_0^2}).$ Since $\mathsf{W}_0 = O_\mathbb{P}(1)$, $\mathbb{P}(\widehat{\beta}_\text{UR} < 0) = o(1)$. The conclusion then follows from $\wh \beta = \min\{\max\{\widehat{\beta}_\text{UR},0\},1\}$.
When $\mca m$ concentrates around $\pi_+$ and $\pi_-$ we have when $\mca m>0$, use the fact that $\pi_{\ell}=\tanh(\beta\pi_{\ell})$ for $\ell\in\{+,-\}$,
and the similar argument gives
The conclusion then Lemma (ref) (3) and the convergence of $\mca m$ to $\pi_+$ or $\pi_-$.
Again we consider the unrestricted PMLE given by
where $l(\beta;\mathbf{W})$ is the pseudo log-likelihood given by
For $\beta \in [0,1]$, that is $c_\beta = \sqrt{n}(\beta - 1) \leq 0$, Equation (ref) and the approximation of $\mca m$ by $n^{-1/2}\mathsf{Z} + n^{-1/4}\mathsf{W}_c$ from Lemma (ref) gives
The conclusion follows from the fact that $x \mapsto \max\{\min\{x,0\},1\}$ is $1$-Lipschitz.
Since we use the conditional probability $p_i$ in the inverse probability weight, we have
and the conclusion follows from $\mathbb{E}[T_i|\mathbf{T}_{-i},(f_i)_{i \in [n]}, \mathbf{E}] = p_i$.
First consider the treatment part.
For the second term, taylor expand $p_i^{-1}, p_i$ as follows:
where $\xi_i^{\ast}$ is some random quantity that lies between $4 \frac{\beta}{n} \sum_{j \neq i} W_j$ and $4 \frac{\beta}{n} \sum_{j \neq i} \pi$. Taking the parameters $c_i^+ = g_i \left(1, \pi \right) \left(1 + \exp(-2 \beta \pi - 2 h) \right)$, $d^+ = \beta(1 - \tanh(\beta \pi + h))\mathbb{E}[g_i(1,\pi)].$ Then
\paragraph*{Proof of (1):} By Lemma (ref), $\mca m - \pi = O_{\psi_{\beta, h}}(n^{-\mathtt{r}_{\beta, h}})$. The claim follows from Equation (ref) and a union bound argument. \paragraph*{Proof of (2):}
By Lemma (ref),
Taylor expand $\tanh(x)$ at $x = \beta \pi + h$, we have
hence
\paragraph*{Proof of (3):} The first line follows from a Taylor expansion of $p_i = (1 + \exp(2 \beta \mca m_i + 2 h))^{-1}$ at $\pi$, and $\mca m_i - \pi = O_{\psi_{\beta,h}}(n^{-\mathtt{r}_{\beta, h}})$, noticing that $c_i$, $\lVert \psi^{\prime \prime} \rVert_{\infty}$ are bounded. The second line follows by reordering the terms.
\paragraph*{Proof of (4):} By Lemma (ref), $\tanh(\beta \pi + h) = \pi + O(n^{-1})$. By boundedness and i.i.d of $g_i(1,\pi)$, $\frac{1}{n} \sum_{j \neq i} c_j = \overline{c} + O(n^{-1}) = \mathbb{E}[c_i] + O_{\mathbb{P}}(n^{-1/2}) + O(n^{-1})$. Similarly, for the control part, taking the parameters $c_i^- = g_i \left(-1, \pi \right) \left(1 + \exp(2 \beta \pi + 2 h) \right)$, $d^- = \beta(1 - \tanh(-\beta \pi - h))\mathbb{E}[g_i(-1,\pi)].$
Using Lemma (ref) again, we can show $(1 +\exp(- 2 \beta \pi - 2 h))/2 = 1/\pi + O(n^{-1})$ and $(1 + \exp(2 \beta \pi + 2 h))/2 = 1/(1 - \pi) + O(n^{-1})$, $\tanh(- \beta \pi - h) = - \pi + O(n^{-1})$. The result then follows from replacing these quantities in $c_i^{+}, c_i^{-}, d^+, d^-$ by corresponding ones using $\pi$.
We decompose by $\Delta_{2,2} = \Delta_{2,2,1} + \Delta_{2,2,2}$, where
Notice that the first term is a quadractic form. Define $\mathbf{H}$ such that $H_{ij} = \frac{g_i^{\prime}(1,\pi) E_{ij}}{2 \mathbb{E}[p_i] N_i}$. Then $\Delta_{2,2,1} = n^{-\mathtt{a}_{\beta, h}} (\mathbf{W} - \pi)^{\operatorname{T}} \mathbf{H} (\mathbf{W} - \pi)$. Take $\mathsf{U}_n$ to be the latent variable from Lemma (ref). Then we decompose
where
Since $\lVert \mathbf{H} \rVert_2 \leq \lVert \mathbf{H} \rVert_F \leq \frac{B}{2 \pi} \sqrt{n}(\min_i N_i)^{-1/2}$, we can apply Hanson-Wright inequality conditional on $\mathsf{U}_n, \mathbf{E}$,
Since $g_i^{\prime}(1,\pi)$'s are independent to $W_i$, by Lemma (ref), $$n^{-\mathtt{a}_{\beta, h}}\sum_{i = 1}^n (W_i - \pi) g_i^{\prime}(1,\pi) = O_{\psi_{\beta,h},tc}(1).$$ By Equation (ref), Lipschitzness of $\tanh$ and Lemma (ref), $\mathbb{E}[W_i|\mathsf{U}_n] - \pi = O_{\psi_{\beta,h}}(n^{-\mathtt{r}_{\beta, h}})$, hence
Then by concentration of $\frac{M_i}{N_i}$ from Lemma (ref), we have
The bound for $\Delta_{2,2,1,d}$ follows from the definition of $\mathbf{H}$ and $\mathsf{U}_n$,
For $\Delta_{2,2,2}$, a Taylor expansion of $p_i$ in terms of $\mca m_i$, and the concentration of $\frac{M_i}{N_i}$ in Lemma (ref) implies that
Throughout the proof, the Ising spins $\mathbf{W}=(W_i)_{i=1}^n$ are distributed according to Assumption (ref) with parameters $(\beta,h)$. For brevity, we write $\mathbb{P}$ in place of $\mathbb{P}_{\beta,h}$.
Take $\mathsf{U}_n$ to be the latent variable given in Lemma (ref). We further decompose by
where $\eta_i^{\ast}$ is some value between $\pi$ and $M_i/N_i$, and
Since $\mathbb{E}[W_i|\mathsf{U}_n,\mathbf{U}] = \tanh\left(\sqrt{\frac{\beta}{n}}\mathsf{U}_n + h \right)$, we have $\mathbb{E}[W_i|\mathsf{U}_n] - \pi = O_{\psi_{\beta,h}}(n^{-\mathtt{r}_{\beta, h}})$ and $(\mathbb{E}[W_i|\mathsf{U}_n] - \pi)^2 = O_{\psi_{\mathtt{p}_{\beta, h}/2}}(n^{-2\mathtt{r}_{\beta, h}})$. It then follows from boundness of $g_i^{(2)}(1, \eta_i^{\ast})$ that
Condition on $\mathsf{U}_n$, $W_i$'s are i.i.d. By Mc-Diarmid inequality conditional on $\mathsf{U}_n$ for each $\sum_{j \neq i}\frac{E_{ij}}{N_i}(W_j - \mathbb{E}[W_j|\mathsf{U}_n])$ and using a union bound over $i \in [n]$, for all $i \in [n]$, for all $t > 0$,
The tails for $n^{\mathtt{r}_{\beta, h}}(\mathbb{E}[W_j|\mathsf{U}_n] - \pi)$ are also controlled,
Integrate over the distribution of $\mathsf{U}_n$ and using a union bound, for large $n$, for all $t > 0$,
By Equation (ref), condition on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A}$, $\mathbb{P}(N_i \leq \mathbb{E}[N_i|\mathbf{U}]/3|\mathbf{U}) \leq n^{-100}$. Hence for such $\mathbf{U}$,
In other words, conditional on $\mathbf{U}$ s.t. $A(\mathbf{U}) \in \mathcal{A}$,
For notational simplicity, we will denote
and since we assume $g_i(\ell,\cdot)$ is $C^4$ for $\ell \in \{-1,1\}$, we know $\theta(\ell,\cdot)$ is $C^2$ for $\ell \in \{-1,1\}$. Then we can decompose $\Delta_{2,3,1,a} - \mathbb{E}[\Delta_{2,3,1,a}|\mathbf{E}]$ as
where $F$ is a function that possibly depends on $\beta(\mathbf{U})$ and $\mathbf{E}$.
First part of $\Delta_{2,3,1,a}$: The first two terms have a quadratic form in $W_j - \mathbb{E}[W_j|\mathsf{U}_n]$, except for the term $\theta(M_i/N_i)$. We will handle it via a generalized version of Hanson-Wright inequality. Fix $\mathsf{U}_n$ and $\mathbf{E}$, consider
Denoting by $D_k H$ the partial derivative of $H$ w.r.p to $W_k$ and $D_{k,l}$ the mixed partials, then
Since we have assumed $f$ is at least $4$-times continuously differentiable, we can apply standard concentration inequalities for $\sum_{j \neq i}\frac{E_{ij}}{N_i}(W_j - \mathbb{E}[W_j|\mathsf{U}_n])$ to get
Hence the gradient of $H$ is bounded by
Moreover, the mix partials are
Hence $\lVert D_{k,l}H(\mathbf{W}) \rVert_{\infty} \lesssim n^{-1/2}\sum_{i =1}^n \frac{E_{ik} E_{il}}{N_i^2}$. Hence
Moreover, since $H F$ is symmetric,
Hence by Theorem 3 from dagan2021learning, for all $t > 0$,
By Equation (ref) and a similar argument for upper bound, for each $i \in [n]$, conditional on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A}$, with probability at least $1 - n^{-100}$, $\mathbb{E}[N_i|\mathbf{U}]/2 \leq N_i \leq 2 \mathbb{E}[N_i|\mathbf{U}]$. Hence for each $t > 0$,
that is
Second part of $\Delta_{2,3,1,a}$: Next, we will show $n^{1 - \mathtt{a}_{\beta, h}} \left(\mathbb{E} \left[B_i \middle| \mathsf{U}_n, \mathbf{U}, \mathbf{E} \right] - \mathbb{E} \left[B_i | \mathbf{E} \right]\right)$, is small. There exists a function $F$ that possibly depends on $\beta$ and $\mathbf{E}$ such that
Define $p(u) = \mathbb{P}(W_j = 1|\mathsf{U}_n,\mathbf{U})$. Then
Using chain rule and product rule for derivatives,
where in the last line, we have used
and that fact that $\lVert p^{\prime} \rVert_{\infty} = O((2 \beta/n)^{0.5})$ and Hoeffiding's inequality for $\sum_{j \neq i}\frac{E_{ij}}{N_i}(W_j - \mathbb{E}[W_j|\mathsf{U}_n])$,
Since $\mathsf{U}_n = O_{\psi_{\beta, h}}(n^{\mathtt{a}_{\beta, h} - 1/2})$, we have
Combining Equations (ref) and (ref), conditional on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A}$,
Combining the bounds for $\Delta_{2,3,1,a}, \Delta_{2,3,1,b}, \Delta_{2,3,1,c}$, we get the desired result.
Throughout the proof, the Ising spins $\mathbf{W}=(W_i)_{i=1}^n$ are distributed according to Assumption (ref) with parameters $(\beta,h)$. For brevity, we write $\mathbb{P}$ in place of $\mathbb{P}_{\beta,h}$.
Recall
First, we will consider the effect of fluctuation of $p_i$ and $\mathbb{E}[W_i|\mathbf{W}_{-i}]$. Recall
It follows from the boundeness of $\beta \mca m_i + h$, $\mca m_i - \pi = O_{\psi_{\beta,h}}(n^{-\mathtt{r}_{\beta, h}})$ that for each $i \in [n]$,
Moreover for some $\eta_i^{\ast}$ between $M_i/N_i$ and $\pi$, using Lemma (ref) we have
Using a union bound over $i$ and an argument for the product of two terms with bounded Orlicz norm with tail control, we have
Next, we will show $n^{-\mathtt{a}_{\beta, h}} \sum_{i = 1}^n \frac{W_i - \pi}{\pi + 1} \left[ g_i \left(1, \frac{M_i}{N_i} \right) - g_i(1, \pi) - g_i^{\prime}(1,\pi) \left(\frac{M_i}{N_i} - \pi \right)\right]$ is small. Suppose $g_i(1,\cdot)$ is $p$-times continuously differentiable. Define
We will use the conditioning strategy to analyse $\delta_p$: Decompse by
with
First, we will show $\delta_{p,2}$ and $\delta_{p,3}$ are small. By Hoeffding inequality, $M_i/N_i - \mathbb{E}[W_i|\mathsf{U}_n] = O_{\psi_2}(N_i^{-1/2})$. Moreover, $\mathbb{E}[W_i|\mathsf{U}_n] - \pi = O_{\psi_{\beta,h}}(n^{-\mathtt{r}_{\beta, h}})$. Hence
For $\delta_{p,3}$, we have
where $\xi^{\ast}$ is some quantity between $\mathbb{E}[W_i|\mathsf{U}_n]$ and $\pi$. Since $x \mapsto x^{p-1}$ is either monotone or convex and none-negative, condition on $\mathbf{E}$,
Combining with boundedness of $g_i^{(p)}(1,\pi)$ and tail control of $\mathbb{E}[W_i|\mathsf{U}_n]$, we have
For $\delta_{p,1}$, we will again use the generalized version of Hanson-Wright inequality. For each $k \in [n]$,
Hence condition on $\mathbf{E}$,
Taking mixed partials w.r.p $\delta_{p,1}$ and using boundedness of $g_i^{(p)}$, we have
It follows that
It then follows from Equation (ref) and Theorem 3 in dagan2021learning that conditional on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A}$,
\paragraph*{Trade-off Between Smoothness of $g_i(1, \cdot)$ and Sparsity of Graph} Assume $g_i(1,\cdot)$ is $p+1$-times continuously differentiable. Then by the decomposition of $\Delta_{2,3,2}$, condition on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A}$,
Then by the concentration of $M_i/N_i - \pi$ given in Lemma (ref), we have
For notational simplicity, denote $\widehat{\mca p} = \frac{1}{n}\sum_{i = 1}^n T_i$ and $\mca p = \frac{1}{2}\tanh(\beta \pi + h) + \frac{1}{2} = \frac{1}{2} \pi + \frac{1}{2}$. Then
Taylor expand $x \mapsto \tanh(\beta x + h)$ at $x = \pi$, we have
where $O(\cdot)$ is up to a universal constant. Together with concentration of $\frac{1}{n}\sum_{i = 1}^n T_i Y_i$ towards $p \mathbb{E}[Y_i]$, we have
A Taylor expansion of $g_i$ and concentration of $M_i / N_i$ then implies
The conclusion then follows.
Throughout the proof, the Ising spins $\mathbf{W}=(W_i)_{i=1}^n$ are distributed according to Assumption (ref) with parameters $(\beta,h)$. For brevity, we write $\mathbb{P}$ in place of $\mathbb{P}_{\beta,h}$.
By Lemma (ref) to Lemma (ref), we show
where $R_i = \frac{g_i(1,\frac{M_i}{N_i})}{1 + \pi} + \frac{g_i(-1,\frac{M_i}{N_i})}{1 - \pi}$, and $b_i = \sum_{j \neq i} \frac{E_{ij}}{N_j} g_j^{\prime}\left(1, \pi \right)$, and $\varepsilon$ is such that condition on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A} = \{A \in \bb R^{n \times n}: \min_{i \in [n]} \sum_{j \neq i}A_{ij} \geq 32 \log n\}$,
Following the strategy as in the proof of Theorem 4 in li2022random, we will show $b_i$ is close to $R_i$: First, decompose by
By Equation (ref), condition on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A}$, $$|\sum_{j \neq i} \frac{E_{ij}}{N_j}g_j^{\prime}(1,\pi) - \sum_{j\neq i}\frac{E_{ij}}{n \mathbb{E}[G(U_i,U_j)|U_j]}g_j^{\prime}(1,\pi)| \leq C n^{-1/2}$$ with probability at least $1 - n^{-99}$. Moreover, $\frac{E_{ij}}{\mathbb{E}[G(U_i,U_j)|U_j]}g_j^{\prime}(1,\pi), j \neq i$ are i.i.d condition on $U_i$, hence $\sum_{j \neq i}\frac{E_{ij}}{n \mathbb{E}[G(U_i,U_j)|U_j]}g_j^{\prime}(1,\pi) - R_i = O_{\psi_2}((n \mathbb{E}[G(U_i,U_j)|U_j]^{-1/2}) = O_{\psi_2}(\mathbb{E}[N_j|X]^{-1/2})$. It follows that conditional on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A}$,
Again using the conditional i.i.d decomposition, Hoeffiding inequality and $\mathsf{U}_n$'s concentration for the two terms respectively,
Hence denote the term of stochastic linearization by $G_n$, i.e.
Since $R_i - \mathbb{E}[R_i] + Q_i$'s are i.i.d independent to $W_i$'s with bounded third moment, we know from Lemma (ref) that $G_n$ can be approximated by either a Gaussian or non-Gaussian law, that is order $1$, this gives
where $O(\cdot)$ does not depend on the value of $\mathbf{U}$ and
To analyse the second term, recall $\mathbb{E}[N_i|\mathbf{U}] = \rho_n \sum_{j \neq i}G(U_i,U_j)$. Hence
the last line is because with probability at least $1 - n^{-98}$, $E = \{\frac{1}{2}g(U_i) \leq \frac{1}{n}\sum_{j \neq i}G(U_i, U_j) \leq 2 g(U_i), \forall 1 \leq i \leq n\}$ happens, and by maximal inequality, $\max_i |g(U_i)|^{-1/2} = O_{\psi_2}(\sqrt{\log n})$. And on $\{A(\mathbf{U}) \in \mathcal{A}\} \cap E$, $\max_i (\frac{1}{n}\sum_{j \neq i}G(U_i,U_j))^{-1/2} \leq (32 \log n/n)^{-1/2}$, since we assume $G$ is positive. By similar argument for the last two terms in $\mathtt{r}(\mathbf{U})$, we have
Recall that $\mathcal{A} = \{A(\mathbf{U}): \min_i \sum_{j \neq i} A_{ij}(\mathbf{U})\geq 32 \log n \}$. Since $\sum_{j \neq i} A_{ij}(\mathbf{U}) \sim \operatorname{Bin}(n-1,\mathbb{E}[G(X_1,X_2)])$, we know from Chernoff bound for Binomials and union bound over $i$ that $\mathbb{P}(A(\mathbf{U}) \notin \mathcal{A}) \leq n^{-99}$. The conclusion then follows.
Our proof for Lemma (ref) to Lemma (ref) relies on the following devices:
(1) Taylor expansion of $\tanh(\cdot)$ in the inverse probability weighting for unbiased estimator, and taylor expansion of $Y_i(\ell,\cdot)$ at $\mathbb{E}[T_i]$ for $\ell \in \{0,1\}$. Then the higher order terms are in terms of $\mca m - \pi$ and $\frac{M_i}{N_i} - \pi$. In Lemma (ref) (taking $X_i \equiv 1$), we show
and in Lemma (ref), we show
where $\mathtt{K}$ is some constant that does not depend on $\beta$. This shows for the higher order terms, we always have
where the $o_\mathbb{P}(\cdot)$ terms does not depend on $\beta$.
(2) Condition i.i.d decomposition based on the de-Finetti's lemma (Lemma (ref)). Suppose $\mathsf{U}_n$ is the latent variable from Lemma (ref), we use decompositions based on $\mathsf{U}_n$: For Lemma (ref) to Lemma (ref), we break down higher order terms in the form
For the first part $F(\mathbf{W},\mathbf{E}) - \mathbb{E}[F(\mathbf{W},\mathbf{E})|\mathbf{E},\mathsf{U}_n]$, we use the conditional i.i.d of $W_i$'s given $\mathsf{U}_n$. For the second part, we use concentration from Lemma (ref) that there exists a constant $\mathtt{K}$ not depending on $\beta$ or $n$, such that $\lVert \mathsf{U}_n \rVert_{\psi_1} \leq \mathtt{K} n^{1/4}$ and the effective term $\lVert \tanh(\sqrt{\frac{\beta}{n}} \mathsf{U}_n) \rVert_{\psi_1} \leq \mathtt{K} n^{-1/4}$. In particular, the rate of concentration for conditional i.i.d Berry-Esseen and concentration of $\tanh(\sqrt{\frac{\beta}{n}}\mathsf{U}_n)$ does not depend on $\beta$.
By the same proof from Lemma (ref) to Lemma (ref), we can show in $\wh \tau_n - \tau_n$, the second and higher order terms in terms of $W_i - \pi$ can always be dominated by the first order terms, with a rate that does not depend on $\beta$.
The conclusion then follows from the two devices and the same proof logic of Lemma (ref) to Lemma (ref).
Define $g(U_j) =\mathbb{E}[G(U_i,U_j)|U_j]$, for $i \neq j$. Reordering the terms,
Hence $\tau^a_{(i)} - \overline{\tau}^a$ has the representation given by
where the second to last line is due to $-\frac{1}{n} \frac{1}{1/2} 1/2(h_i(1,0) - \mathbb{E}[h_i(1,0)]) + \frac{1}{n} \frac{1}{1 - 1/2} (1 - 1/2) (h_i(-1,0) - \mathbb{E}[h_i(-1,0)]) = - \frac{2}{n} \varepsilon_i + \frac{2}{n} \varepsilon_i = 0$.
Now we look at $b$-part. For representation purpose, we look at only the treatment part. The control part can be analysized by in the same way. Reordering the terms,
Hence $\tau_{(i)}^b - \overline{\tau}^b$ has the representation given by
The analysis follows from a Taylor expansion of $h_j(1,\cdot)$. For some $\xi_{j,i}^{\ast}$ between $\frac{M_j}{N_j}_{(i)}$ and $0$ for each $j,i$,
where we have used $\partial_2 h_j(1,\cdot) = \partial_2 [h(1,\cdot) + \varepsilon_j] = \partial_2 h(1,\cdot)$.
\paragraph*{Part 1: Linear Terms}
By a decomposition argument,
Hence
Condition on $U_j$, $(E_{lj} W_l: l \neq j)$ are i.i.d mean-zero, hence Bernstein inequality gives $\frac{1}{n} \sum_{l = 1}^n E_{lj} W_l = O_{\psi_2}(\sqrt{n^{-1}\rho_n}) + O_{\psi_1}(n^{-1})$, which implies
Putting back into Equation (ref),
Looking at contribution from the first order term in Taylor expanding $h_j(1,\cdot)$ to $\tau_{(i)}^b - \overline{\tau}^b$ in Equation (ref),
Since $(E_{ij} T_j/g(U_j): j \in [n])$ are independent condition on $U_i$, standard concentration inequality gives
Since we assumed $\partial_2 h(1,0) = \partial_2 f(1,0) + o_{\mathbb{P}}(1) = \partial_2 f_j(1,0) + o_{\mathbb{P}}(1)$ where
Together with the leading term in Equation (ref), we have
\paragraph*{Part 2: Higher Order Terms} For the second order terms, first notice that if $l \notin [n]$, then
where we have used $(M_j/N_j)_{\iota} = O_{\psi_2}((n \rho_n)^{-\frac{1}{2}})$ and $N_j^{-1} = O_{\psi_2}((n \rho_n)^{-1})$. If $l \in [n]$, then again
Hence
For the third order residual, observe that $(\frac{M_j}{N_j}_{(\iota)})^3 = O_{\psi_2}((n \rho_n)^{-3/2})$. Then
The conclusion then follows from Equations (ref), (ref) and (ref).
Define $\mathtt{r}(x) = (1,x)^\top$. Denote $\pi = \mathbb{E}[W_i] = 2 \mathbb{E}[T_i] - 1$. Then
First, consider the gram-matrix. Take $\zeta_i := \sqrt{n \rho_n} (\frac{M_i}{N_i} - \pi)$. Then $1 \lesssim \mathbb{V}[\zeta_i] \lesssim 1$. Take $b_n = \sqrt{n \rho_n} h_n$. Take
where $\mathtt{r}: \mathbb{R} \rightarrow \mathbb{R}^2$ is given by $\mathtt{r}(u) = (1,u)^{\top}$. Take $Q$ to be the probability measure of $\zeta_i$ given $\mathbf{E}$. Then
In particular, $\lambda_{\min}(\mathbf{B}) \gtrsim 1$. Now we want to show each entry of $\mathbf{B}_n$ converge to those of $\mathbf{B}$. Take
Denote $\partial_j$ to be the partial derivative w.r.p to $W_j$. Since $K$ is Lipschitz with bounded support,
Condition on $\mathbf{E}$,
Hence for all $p,q \in \{0,1\}$,
Since both $\mathbf{B}_n$ and $\mathbf{B}$ are two by two matrices, $\lVert \mathbf{B}_n - \mathbf{B} \rVert_{\operatorname{op}} \lesssim O_{\psi_2}((n b_n^4)^{-1})$. By Weyl's Theorem,
and together with $\lambda_{\min}(\mathbf{B}) \gtrsim 1$, implies $\lambda_{\min}(\mathbf{B}_n) \gtrsim 1$. Take
Hence variance can be bounded by
Next, consider the bias term. Since $f(1,\cdot) \in C^2$, whenever $|\frac{M_i}{N_i} - \pi| \leq h_n = (n \rho_n)^{-1/2} b_n$,
Hence using the fourth and third lines above respectively,
Putting together Equations (ref) and (ref),
Hence any $b_n$ such that $b_n = \Omega(n^{-1/4} + \rho_n^{1/3})$ will make $(\widehat{\gamma}_0, \widehat{\gamma}_1)$ a consistent estimator for $(\gamma_0,\gamma_1)$. For any $0 \leq \rho_n \leq 1$ such that $n \rho_n \rightarrow \infty$, such a sequence $b_n$ exists.
The order $\frac{M_i}{N_i}$ is $n^{-1/4}$ if $\liminf_{n \rightarrow \infty} n \rho_n^2 > c$ for some $c > 0$; and is $(n \rho_n)^{-1/2}$ if $n \rho_n^2 = o(1)$. We consider these two cases separately.
\paragraph*{Case 2.1: $\liminf_{n\rightarrow \infty}n \rho_n^2 > c$ for some $c > 0$} Take $\eta_i = n^{\frac{1}{4}}(\frac{M_i}{N_i} - \pi)$. Take $d_n = n^{1/4}h_n$. And with the same $\mathtt{r}$ defined in Case 1,
Under the assumption $\liminf_{n \rightarrow \infty }n \rho_n^2 \leq c$ for some $c > 0$, we have $1 \lesssim \mathbb{V}[\eta_i] \lesssim 1$. Hence $\lambda_{\min}(\mathbf{D}) \gtrsim 1$. To study the convergence between $\mathbf{D}_n$ and $\mathbf{D}$, again consider for $p, q \in \{0,1\}$,
Still let $\mathsf{U}_n$ be the latent variable from Lemma (ref), $W_i$'s are independent conditional on $\mathsf{U}_n$. Hence by similar argument as Equation (ref), we can show
Moreover, recall we denote by $\omega_i \in [k]$ the block unit $i$ belongs to, then
$p(U_l)= \mathbb{P}(W_i = 1|U_{\ell}) = \frac{1}{2}(\tanh(\sqrt{\beta_{\ell}/n}\mathsf{U}_n + h_{\ell}) + 1)$, $i \in \ca I_{\ell}$. Take the derivative term by term,
Using Lipschitz property of $x \mapsto (x/h_n)^{p+q} K(x/h_n)$,
Hence for all $\ell \in \mathscr{C}$,
Moreover, for all $\ell \in \mathscr{C}$, $\lVert U_{\ell} \rVert_{\varphi_2} \lesssim n^{1/4}$. Together, this gives
Hence if we take $d_n \gg 1$ (which implies $n d_n^4 \gg 1$), then $G_{p,q}(\mathbf{W}) = \mathbb{E}[G_{p,q}(\mathbf{W})|\mathbf{E}] + o_{\mathbb{P}}(1)$, implying $\lVert \mathbf{D}_n - \mathbf{D} \rVert_2 = o_{\mathbb{P}}(1)$ and $\lambda_{\min}(\mathbf{D}_n) - \lambda_{\min}(\mathbf{D}) = o_{\mathbb{P}}(1)$, making $\lambda_{\min}(\mathbf{D}_n) \gtrsim_{\mathbb{P}} 1$. Take
Hence variance can be bounded by
By similar argument as in Case 1, assume $d_n \gg 1$, we can show
Hence if we choose $d_n$ such that $1 \ll d_n \ll n^{1/8}$, then $(\widehat{\gamma}_0, \widehat{\gamma}_1)$ is a consistent estimator for $(\gamma_0, \gamma_1)$. The only assumption we made for the existence of such a $d_n$ is $\liminf_{n \rightarrow \infty} n \rho_n^2 \geq c$ for some $c > 0$.
\paragraph*{Case 2.2: $n \rho_n^2 = o(1)$} Take $\eta_i := \sqrt{n \rho_n}(\frac{M_i}{N_i} - \pi)$, $d_n = \sqrt{n \rho_n} h_n$. By similar decomposition based on latent variables, we can show if $n \rho_n \rightarrow \infty$ as $n \rightarrow \infty$, then there exists $h_n$ such that $(\widehat{\gamma}_0, \widehat{\gamma}_1)$ is a consistent estimator for $(\gamma_0, \gamma_1)$.
The result is a special case of Lemma (ref) in Section (ref) when $h \neq 0$.
The result follows from Lemma (ref), Lemma (ref), and the same anti-concentration argument as in the proof of Lemma (ref).
First, we consider the unbiased estimator
with $p_i = \bb P(W_i = 1| \mathbf{W}_{-i}) = \left(\exp \left(-2\beta \mca m_i\right) + 1\right)^{-1}$. Our analysis will be similar to the proofs in Section (ref), but using the concentration of $n^{-1} \sum_{i = 1}^n W_i$ conditional on $\operatorname{sgn}(\mca m)$ shown in Lemma (ref) instead of the unconditional concentration of $n^{-1} \sum_{i = 1}^n W_i$. We decompose by
For the first term, we Taylor expand the expression for $p_i^{-1}$ in terms of $\mca m_i$, and get
condition on $\operatorname{sgn}(\mca m) = \ell$, where
For the second term, we Taylor expand $g_i(1, \cdot)$ at $\pi_{\ell}$: For some $\eta_i^{\ast}$ between $\pi_{\ell}$ and $\frac{M_i}{N_i}$,
where
Term $\Delta_{2,1}^{\prime}$: Denote $\mathbf{g} = (g_i)_{1 \leq i \leq n}$. Rearranging the terms,
Term $\Delta_{2,2}^{\prime}$: Take $u_{\ell} = (\pi_{\ell} + 1)/2$ for $\ell \in \{-,+\}$. Decompose by $$\Delta_{2,2}^{\prime} = \Delta_{2,2,1}^{\prime} + \Delta_{2,2,2}^{\prime},$$ where
Since $\beta < \infty$ and $v_{-}$, $u_-$ and $u_+$ are bounded away from $0$ and $1$. Rearranging the terms, we get
where $\mathbf{H}^{\ell}$ is the $n \times n$ matrix with $H^{\ell}_{ij} = g_i^{\prime}(1, \pi_{\ell}) E_{ij} (2 u_{\ell} N_i)^{-1}$ and $\mathbf{1}$ is the $n$-dimensional vector with all entries $1$. To analyze the quadratic form, we use the same strategy as in the proof of Lemma (ref): Let $\mathsf{U}_n$ be the one defined in Lemma (ref), and we know $W_1, \cdots, W_n$ are conditional i.i.d given $\mathsf{U}_n$. Then we can decompose $\Delta_{2,2,1}^{\prime}$ into four terms based on
Conditional Berry-Esseen given $\mathsf{U}_n$, conditional concentration of $\mathsf{U}_n$, $\mca m$ and $\frac{M_i}{N_i}$ given $\operatorname{sgn}(\mca m)$ in Remark (ref), Lemma (ref) and Lemma (ref), and the same argument as in the proof for Lemma (ref) implies that condition on $\mathbf{g}, \mathbf{E}$ and $\operatorname{sgn}(\mca m)$,
Term $\Delta_{2,3}^{\prime}$: Now we proceed to $\Delta_{2,3}^{\prime}$. Decompose by $\Delta_{2,3}^{\prime} = \Delta_{2,3,1}^{\prime} + \Delta_{2,3,2}^{\prime}$, where
Define $\Delta_{2,3,1,l}^{\prime}$ to be the counterparts of $\Delta_{2,3,1,l}$ in Equation (ref) with $\pi$ by replaced by $\pi_{\ell}$ for $l \in \{a,b,c\}$, the same argument in the proof of Lemma (ref) shows
condition on $\operatorname{sgn}(\mca m) = \ell$ for $\ell = -, +$. Combining the three parts,
condition on $\operatorname{sgn}(\mca m) = \ell$ for $\ell = -, +$. Taylor expanding $p_i = (1 + \exp(- 2 \beta \mca m_i))^{-1}$ as a function of $\mca m_i$ at $\pi_{\ell}$, the same argument as in Lemma (ref) shows
Conditional concentration of $\mathsf{U}_n$, $\mca m$ and $\frac{M_i}{N_i}$ given $\operatorname{sgn}(\mca m)$ in Remark (ref), Lemma (ref) and Lemma (ref), and the same argument as in the proof for Lemma (ref) implies that condition on $\mathbf{g}, \mathbf{E}$ and $\operatorname{sgn}(\mca m)$,
Putting together. Putting together the decompositions, condition on $\mathbf{E}$ and $\operatorname{sgn}(\mca m) = \ell$,
where with $c_{i,l} = g_i(1,\pi_{\ell})(1 + \exp(2 \beta \pi_{\ell}))$, and $d_l = \frac{\beta(1 + \exp(2 \beta \pi_{\ell}))}{1 + \cosh(2 \beta \pi_{\ell})}\mathbb{E}[g_i(1, \pi_{\ell})]$,
Consider the event $\Omega_i = \{\operatorname{sgn}(\mca m) = \ell, |\sum_{j \neq i}W_j| \leq 1\}$ and $\Omega = \cup_{1 \leq i \leq n} \Omega_i$. We then have
implying $\mathbb{P}(\Omega_i) \leq C \exp(-nC)$, $1 \leq i \leq n$. Hence
Hence condition on $\mathbf{E}$ and $\operatorname{sgn}(\mca m) = \ell$,
Now, we consider the difference between the unbiased estimator and the Hajek estimator. For notational simplicity, denote $\widehat{\mca p} = \frac{1}{n}\sum_{i = 1}^n T_i$ and $\mca p_{\ell} = \frac{1}{2}\tanh(\beta \pi_{\ell} + h) + \frac{1}{2} = \frac{1}{2} \pi_{\ell} + \frac{1}{2}$. Then
Taylor expand $x \mapsto \tanh(\beta x + h)$ at $x = \pi_{\ell}$, we have
where $O(\cdot)$ is up to a universal constant. Together with the fact that condition on $\operatorname{sgn}(\mca m) = \ell$, $\frac{1}{n}\sum_{i = 1}^n T_i Y_i$ concentrates towards $\mca m \mathbb{E}[Y_i|\operatorname{sgn}(\mca m) = \ell]$, we have
condition on $\operatorname{sgn}(\mca m) = \ell$. A Taylor expansion of $g_i$ and concentration of $M_i / N_i$ then implies
where $\pi^{\ast}$ is some number between $\pi_{\ell}$ and $M_i/N_i$. The conclusion then follows.
The result follows from Lemma (ref) (3), Lemma (ref), and the same anti-concentration argument as in the proof of Lemma (ref).
As in the case of one block analyzed in Section (ref), $\wh \boldsymbol{\tau}_n$ is not an unbiased estimate of $\boldsymbol{\tau}_n$. We first consider an unbiased estimator to $\boldsymbol{\tau}_n$ and then consider the difference.
Consider $\widehat{\boldsymbol{\tau}}_{n,UB} = (\widehat{\tau}_{n,UB,1}, \cdots, \widehat{\tau}_{n,UB,K})$, where
Here $p_i = \sum_{k = 1}^K \mathbbm{1}(i \in \mathcal{C}_k) (1 + \exp(2 \beta_k \mca m_{i,k} + 2 h_k))^{-1}$ and $\mca m_{i,k} = n_k^{-1} \sum_{j \in \mathcal{C}_k, j \neq i}W_j$.
Denote $\mca m = n^{-1}\sum_{i =1}^n W_i$, $\mca m_k = n_k^{-1} \sum_{i \in \mathcal{C}_k} W_i$. For notational simplicity, we denote $\pi_{l,\operatorname{sgn}(\mca m_l)}$ by $\pi_l$ for low temperature blocks $l \in \mathscr{L}$, and omit the index by $(\mathbf{s})$ with $\mathbf{s} = \boldsymbol{sgn}$. As in the one-block case, we decompose by
where
and $\zeta_i = \frac{\sum_{k =1}^K N_{i,k} \pi_k}{\sum_{k = 1}^K N_{i,k}}$.
Condition on $\mathbf{E}, \mathbf{g} = \{g_i: i \in [n]\}$ and $\boldsymbol{sgn}$, the randomness of $\Delta_{1,k}$ only comes from $(W_i)_{i \in \mathcal{C}_k}$, that is, the Ising bits from the same block. Hence
The analysis in Lemma (ref) and Lemma (ref) with $g_i(1,\zeta_k)\mathbbm{1}(i \in \mathcal{C}_k)$ replacing $g_i(1,\pi)$ implies
condition on $\mathbf{E}, \mathbf{g}, \boldsymbol{sgn}$, where $\mathtt{c}_k = (1 + \exp(2\beta_k \pi_k + 2 h_k))/2$ and $\mathtt{d}_k = \beta_k(1 + \exp(2 \beta_k \pi_k + 2 h_k))/(1 + \cosh(2 \beta_k \pi_k + 2 h_k))$.
The linearization of $\Delta_{2,k}$ involves $M_i/N_i$, which depends all blocks even if the estimator is for block $k$. We will find its stochastic linearization in terms of units in all blocks.
By a Taylor expansion of $g_i(1,\cdot)$ at $\zeta_i$
where
with $g_i(x) = \int_0^1 (1 - t) Y_i^{(2)}(1,\zeta_i + t (x - \zeta_i))dt$. In particular, $g_i$ is $C^2$. \paragraph*{Term $\Delta_{2,1}$:} Rearranging $\Delta_{2,1}$, we get the effective term in the stochastic linearization.
\paragraph*{Term $\Delta_{2,2}$:} We want to show $\Delta_{2,2}$ is negligible. Consider the effect from each block separately. We claim that condition on $\mathbf{g}$, $\mathbf{E}$, $\boldsymbol{sgn}$,
To get the second line, notice that $p_i = (1 + \exp(2 \beta_k \mca m_{i,k} + 2 h_k))^{-1}$ is Lipschitz in $\mca m_{i,k}$, and since $(W_i: i \in \mathcal{C}_k), 1 \leq k \leq K$ form independent Ising models, we can use Lemma (ref) to get $\mca m_{i,k} - \pi_k = O_{\psi_{\beta_k,h_k}}(n_k^{-\mathtt{r}_{\beta_k, h_k}})$ for $k \in \mathscr{H} \cup \mathscr{C}$, and condition on $\operatorname{sgn}(\mca m_{k})$, $\mca m_{i,k} - \pi_k = O_{\psi_{\beta_k,h_k}}(n_k^{-\mathtt{r}_{\beta_k, h_k}})$ for $k \in \mathscr{L}$. Hence for each $k \in [K]$,
Suppose $\mathsf{U}_{n,l}$ is the latent variable underlining the distribution of $(W_i: i \in \mathcal{C}_l), l \in [K]$ as in Lemma (ref). Conditional on $\mathbf{E}$ and $\boldsymbol{sgn}$, using Hoeffiding's inequality and the concentration of $\mathsf{U}_{n,l}$, we have
From the fact that $\mathbb{P}(|Z_1 Z_2| \geq t) \leq \mathbb{P}(\sqrt{\log n} |Z_2| \geq t) + \mathbb{P}(|Z_1| \geq \sqrt{\log n})$ for any two random variables $Z_1$ and $Z_2$, and using a union bound over the summation over $i \in \mathcal{C}_k$, we get the second line for $\Delta_{2,2}$.
Now consider the first term of $\Delta_{2,2}$. With the help of the latent variables $\mathsf{U}_{n,k}, 1 \leq k \leq K$, decompose by
where
Since conditional on $\mathsf{U}_{n,k}$ and $\mathsf{U}_{n,l}$, $(W_i: i \in \mathcal{C}_k \cup \mathcal{C}_l)$ are i.i.d., we can use Hoeffding's inequality and boundedness of $g_i^{\prime}(1,\zeta_i)$ to get conditional on $\mathbf{E}$ and $\boldsymbol{sgn}$,
and
It follows that condition on $\mathbf{E}$ and $\boldsymbol{sgn}$,
For $\Gamma_{k,l,a}$, observe that with $\omega_i = \sum_{k = 1}^K k \mathbbm{1}(i \in \mathcal{C}_k)$,
Apply Hanson-Wright inequality conditional on $\mathbf{E}$, $\mathsf{U}_{n,l}$ and $\mathsf{U}_{n,k}$, we get $$\Gamma_{k,l,a} - \mathbb{E}[\Gamma_{k,l,a}|\mathbf{E}, \boldsymbol{sgn}] = O_{\psi_1}((n_k \min_i N_i)^{-\frac{1}{2}}).$$ Put together, conditional on $\mathbf{E}$ and $\boldsymbol{sgn}$,
\paragraph*{Term $\Delta_{2,3}$:} Similar to the analysis in Section (ref), we decompose $\Delta_{2,3} = \Delta_{2,3,1} + \Delta_{2,3,2}$ where
where $\Delta_{2,3,1}$ is further decomposed based on latent variables $\mathsf{U}_{n,l}, 1 \leq l \leq K$, that is, $$\Delta_{2,3,1} = \Delta_{2,3,1,a} + \Delta_{2,3,1,b} + \Delta_{2,3,1,c},$$ where
Term $\Delta_{2,3,1,a}$: Consider the $\Delta_{2,3,1,a}$ as a (random) function on $\mathbf{W}$ and $\mathsf{U}_{n,1}, \cdots, \mathsf{U}_{n,K}$. Let
Notice that conditional on $\mathsf{U}_{n,l}, 1 \leq l \leq K$, $W_j$'s are independent random variables, and we can rewrite $$\Delta_{2,3,1,a} = \frac{n}{n_k} \frac{1}{n} \sum_{i = 1}^n g_i \left(\frac{M_i}{N_i}\right) \mathbbm{1}(i \in \mathcal{C}_k)(\sum_{l =1}^K \sum_{j \in \mathcal{C}_l, j\neq i} \frac{E_{ij}}{N_i}(W_j - \mathbb{E}[W_j|\mathsf{U}_{n,l}]))^2.$$ It follows from the same concentration argument for $\Delta_{2,3,1,a}$ in the proof for Lemma (ref) that conditional on $\mathbf{E}$ and $\boldsymbol{sgn}$,
Define $p_l(u) = \mathbb{E}[W_i = 1| \mathsf{U}_{n,l} = u, i \in \mathcal{C}_l]$. Then we can write
By the same argument as in the proof for Lemma (ref),
It then follows from the concentration of $\mathsf{U}_{n,1}$ to $\mathsf{U}_{n,K}$ that condition on $\mathbf{E}$ and $\boldsymbol{sgn}$,
Moreover, since $\sum_{j \in \mathcal{C}_l, j \neq i} \frac{E_{ij}}{N_i}(W_j - \mathbb{E}[W_j|\mathsf{U}_{n,l}]) = O_{\psi_2}(N_i^{-\frac{1}{2}})$ and $\mathbb{E}[W_j|\mathsf{U}_{n,l}] - \pi_l = O_{\psi_{\beta_l,h_l}}(n^{-\mathtt{r}_{\beta_l, h_l}})$, we have conditional on $\mathbf{E}$ and $\boldsymbol{sgn}$,
Putting together, $$\Delta_{2,3,1} - \mathbb{E}[\Delta_{2,3,1}|\mathbf{E}] = O_{\psi_2, tc}\left(\sqrt{\log n} \max_i N_i^{-\frac{1}{2}} \max_{1 \leq l \leq K}n^{-\mathtt{r}_{\beta_l, h_l}} + \max_{1 \leq l \leq K} n^{-2 \mathtt{r}_{\beta_l, h_l}}\right).$$ Consider the $p$-th order term in the expansion of $\Delta_{2,3,2}$, $$\delta_p = \frac{1}{n_k} \sum_{i \in \mathcal{C}_k} \frac{T_i - \pi_k}{\pi_k} g_i^{(p)}(1, \zeta_i)\bigg(\frac{M_i}{N_i} - \zeta_i\bigg).$$ Following conditional i.i.d argument as in Lemma (ref), we can show condition on $\mathbf{E}$ and $\boldsymbol{sgn}$,
Hence assuming $g_i(1, \cdot)$ is $C^{p+1}$. Taylor expand $g_i(1,\cdot)$ up to the $p$-th order, we get
From the previous steps, condition on $\mathbf{E}$ and $\boldsymbol{sgn}$,
where $\bar{\mathtt{r}}_n = \max_{k \in [K]}n^{-2 \mathtt{r}_{\beta_k, h_k}} + \sqrt{\log n} \max_{k \in [K]} n^{-\mathtt{r}_{\beta_k, h_k}}(n \rho_n)^{-1} + \sqrt{\log n} n^{-1/2} + (n \rho_n)^{-(p+1)/2}$, and $\bar{\mathbf{S}}_{l,i}$ is the vector $(\bar{S}_{1,l,i}, \cdots, \bar{S}_{K,l,i})^{\mathbf{T}}$, where
with $\zeta_i = \frac{\sum_{\ell = 1}^k N_{i,\ell} \pi_{\ell}}{\sum_{\ell = 1}^k N_{i,\ell}}$. Condition on $U_i$, $E_{i,j}$ for all $1 \leq j \leq n$ are independent with $|E_{ij}| \leq 1$ and $\mathbb{V}[E_{ij}|U_i] \lesssim \rho_n$. Hence using Bernstein's inequality,
with $\overline{\pi} = \sum_{k = 1}^K p_k \pi_k$. The same argument as the proof for Lemma (ref) implies
with
Hence with $\mathbf{S}_{l,i} = (S_{1,l,i}, \cdots, S_{K,l,i})^{\mathbf{T}}$, where
Hence by the same analysis as Equation (ref) in the proof of Lemma (ref),
We already know $\widehat{\boldsymbol{\tau}}_{n,UB}$ is the unbiased estimator. The same argument as Equation (ref) shows that condition on $\operatorname{sgn}$,
This finishes the proof for the unbiased estimator.
The analysis will be the same as those for Lemma (ref). For simplicity, denote $\widehat{\mca p}_k = n_k^{-1} \sum_{i \in \mathcal{C}_k} W_i$ and $\mca p_k = \frac{1}{2}\tanh(\beta_k \mca m_k + h_k) + \frac{1}{2} = \frac{1}{2} \mca m_k + \frac{1}{2}$. Then
The analysis in Lemma (ref) implies
and
Hence
The conclusion then follows from step I. The Unbiased Estimator.
We want to apply Lemma (ref) to the stochastic linearizations obtained from Lemma (ref), $$n_l^{-1} \sum_{i \in \mathcal{C}_l} \mathbf{S}_{l,i,(\mathbf{s})}(W_i - \pi_{l,(\mathbf{s})}),$$ for different blocks separately. First, we need to check if $\mathbf{S}_{l,i,(\mathbf{s})}$ satisfies the covariate constraints in Lemma (ref). Recall $\mathbf{S}_{l,i,(\mathbf{s})} = (S_{1,l,i,(\mathbf{s})}, \cdots, S_{K,l,i,(\mathbf{s})})^{\mathbf{T}}$, where
Definitions of $Q_{i,(\mathbf{s})}$ and $R_{i,l,(\mathbf{s})}$ imply that $\min_{k,l \in [K]}\mathbb{E}[S_{k,l,i,(\mathbf{s})}^2] > 0$ and $\max_{k,l}|S_{k,l,i}| < \infty$ almost surely, satisfying the conditions in Lemma (ref). Hence
where
Now replacing $n_l$ by $n p_l$. The assumption that $n_l/n = p_l + O(n^{-1/2})$ and the Nazarov inequality implies
The independence between $\mathbf{S}_{l,i,(\mathbf{s})}$ for different $i$'s and the independence between Ising-spins across blocks then imply the stochastic linearization from Lemma (ref) can be approximated by summation of right hand sides of Equation (ref). Lemma (ref) and Nazarov inequality applied on the $n^{-1/2} \boldsymbol{\Sigma}^{1/2} \mathsf{Z}_K$ part then imply the conclusion.