EconBase
← Back to paper

Robust Inference for the Direct Average Treatment Effect with Treatment Assignment Interference

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

243,429 characters · 58 sections · 18 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Supplementary Appendix to “Robust Inference for the Direct Average Treatment Effect with Treatment Assignment Interference”

abstractThis supplemental appedix contains general theoretical results encompassing those discussed in the main paper, includes proofs of those general results, and discusses additional methodological and technical results.

Notations

For $n \in \mathbb{N}$, $[n] = \{1, \cdots, n\}$. For reals sequences $a_n = o(b_n)$ if $\limsup_{n\to\infty} \frac{|a_n|}{|b_n|} = 0$, $|a_n| \lesssim |b_n|$ if there exists some constant $C$ and $N > 0$ such that $n > N$ implies $|a_n| \leq C |b_n|$. For sequences of random variables $a_n = o_{\mathbb{P}}(b_n)$ if $\operatorname{plim}_{n \rightarrow \infty}\frac{|a_n|}{|b_n|} = 0$, $a_n = O_{\mathbb{P}}(b_n)$ if $\limsup_{M \rightarrow \infty} \limsup_{n \rightarrow \infty} \mathbb{P}[|\frac{a_n}{b_n}| \geq M] = 0$. For positive real sequences $a_n \ll b_n$ if $a_n = o(b_n)$. For a sequence of real-valued random variables $X_n$, we say $X_n = O_{\psi_p}(r_n)$ if there exists $N \in \mathbb{N}$ and $M > 0$ such that $\lVert X_n \rVert_{\psi_p} \leq M r_n$ for all $n \geq N$, where $\lVert \cdot \rVert_{\psi_p}$ is the Orlicz norm w.r.p $\psi_p(x) = \exp(x^p) - 1$. We say $X_n = O_{\psi_p, tc}(r_n)$, $tc$ stands for tail control, if there exists $N \in \mathbb{N}$ and $M > 0$ such that for all $n \geq N$ and $t > 0$, $\mathbb{P}(|X_n| \geq t) \leq 2 n\exp(-(t/(Mr_n))^{p}) + M n^{-1/2}$.

For a vector $\mathbf{v} \in \mathbb{R}^k$, the Euclidean norm is $\lVert \mathbf{v} \rVert = (\sum_{i = 1}^k \mathbf{v}_i^2)^{1/2}$,and the infinity norm is $\lVert \mathbf{v} \rVert_{\infty} = \max_{1 \leq i \leq k} |v_i|$. For a matrix $A = (a_{ij})_{i \in [m], j \in [n]} \in \mathbb{R}^{m \times n}$, the operator norm is $\lVert A \rVert = \lVert A \rVert_2 = \sup_{\lVert \mathbf{x} \rVert = 1} \lVert A\mathbf{x} \rVert$, the maximum absolute column sum norm is $\lVert A \rVert_1 = \sup_{1 \leq j \leq n} \sum_{i = 1}^m |a_{ij}|$, and the Frobenius norm is $\lVert A \rVert_F = \sqrt{\sum_{i = 1}^m \sum_{j = 1}^n a_{ij}^2}$. For sets $A$ and $B$, denote by $A \Delta B$ the set difference $(A \setminus B) \cup (B \setminus A)$.

$\operatorname{sgn}$ denotes the function such that $\operatorname{sgn}(x) = +$ if $x \geq 0$, and $\operatorname{sgn}(x) = -$ otherwise. $\Phi(x)$ denotes the standard Gaussian cumulative distribution function. For $\boldsymbol{\mu} \in \mathbb{R}^{k \times k}$ and $\boldsymbol{\Sigma} \in \mathbb{R}^{k \times k}$, $\mathsf{N}(\boldsymbol{\mu}, \boldsymbol{\Sigma})$ denotes the multivariate normal distribution with mean $\boldsymbol{\mu}$ and covariance matrix $\boldsymbol{\Sigma}$.

Curie-Weiss Magnetization with Independent Multipliers

For notational simplicity, we consider $$\mathbf{W} = (W_i)_{1 \leq i \leq n}, \qquad W_i = 2 T_i - 1, 1 \leq i \leq n.$$ And we consider a more general setting compare to Assumption 3 in the main paper.

assumption[Curie-Weiss] For $\beta \geq 0$ and $h \in \mathbb{R}$, suppose $\mathbf{W} = (W_i)_{1 \leq i \leq n}$ are such that for some $C_{\beta,h} \in \mathbb{R}$, \begin{align} \mathbb{P}_{\beta,h}(\mathbf{W} = \mathbf{w}) = C_{\beta,h}^{-1} \exp \bigg(\frac{\beta}{n}\sum_{1 \leq i < j \leq n} w_i w_j + h \sum_{i = 1}^n w_i\bigg), \qquad \mathbf{w} = (w_1, \cdots, w_n) \in \{-1,1\}^n, \end{align} where $C_{\beta,h}$ is a normalizing constant.

The Curie-Weiss model has a phase transition phenomena in different regimes. Let $\mca m = n^{-1}\sum_{i = 1}^n W_i$.

enumerate• High temperature or non-zero external field $\mathcal{A}_H = \{(\beta,h) \in \mathbb{R}_+ \times \mathbb{R}: h = 0, 0 \leq \beta < 1 \text{ or } h \neq 0\}$: $\mca m$ concentrates around $\pi$, where $\pi$ is the unique solution to $x = \tanh(\beta x + h)$. In particular, $\mca m = \pi + \Theta_{\mathbb{P}}(n^{-1/2})$. Moreover, $\mathcal{A}_H = \mathcal{A}_{H,1} \sqcup \mathcal{A}_{H,2}$, where $\mathcal{A}_{H,1} = \{(\beta, h) \in \mathbb{R}^+ \times \mathbb{R}: h = 0, 0 \leq \beta < 1\}$ and $\mathcal{A}_{H,2} = \{(\beta, h) \in \mathbb{R}^+ \times \mathbb{R}: h \neq 0\}$. • Critical temperature $\mathcal{A}_{C} = \{(1,0)\}$: $\mca m$ concentrates around $\pi$, where $\pi$ is the unique solution to $x = \tanh(\beta x + h)$. In particular, $\mca m = \pi + \Theta_{\mathbb{P}}(n^{-1/4})$. • Low temperature regime $\mathcal{A}_{R} = \{(\beta,h) \in \mathbb{R}_+ \times \mathbb{R}: h = 0, \beta > 1\}$: $\mca m$ concentrates on the set $\{\pi_-,\pi_+\}$, with $\pi_-$ and $\pi_+$ the unique negative and positive solutions to $x = \tanh(\beta x)$, respectively. In particular, condition on $\operatorname{sgn}(\mca m) = \ell$, $\mca m = \pi_{\ell} + \Theta_{\mathbb{P}}(n^{-1/2})$.

In the main paper, we focus on $(\beta, h)$ in $\mathscr{H}_1$. But for this section, we provide the results for all of $\mathscr{H}$, $\mathscr{C}$ and $\mathscr{L}$.

Suppose $\mathbf{X} = (X_1,\cdots,X_n)$ has i.i.d components such that $\mathbb{E} \left[|X_1|^3\right] < \infty$ independent to $\mathbf{W}$. The goal is to study the limiting distribution and the rate of convergence for

align*[align* omitted — 66 chars of source]

The magnetization $n^{-1}\sum_{i = 1}^n (W_i - \pi)$ has been studied using Stein's method eichelsbacher2010stein, chatterjee2010spin. Due to the multipliers, the Stein's method can not be directly applied for $\mca g_n$. We use a novel strategy based on the following de Finetti's lemma to show Berry Essseen results.

lemma[de Finetti's Theorem] There exists a latent variable $\mathsf{U}_n$ with density \begin{align*} f_{\mathsf{U}_n}(u) = I_{\mathsf{U}_n}^{-1}\exp \bigg(- \frac{1}{2} u^2 + n \log \cosh \bigg( \sqrt{\frac{\beta}{n}} u + h\bigg) \bigg), \end{align*} where $I_{\mathsf{U}_n} = \int_{-\infty}^{\infty} \exp (- \frac{1}{2} u^2 + n \log \cosh ( \sqrt{\frac{\beta}{n}} u + h)) d u$, such that $W_1,\cdots,W_n$ are i.i.d condition on $\mathsf{U}_n$.

The de Finetti's theorem for exchangable sequences of random variable is a classical result diaconis1988recent, diaconis1980finetti, ellis1978statistics. For completeness, we include a short proof for the Curie-Weiss model in Section (ref).

lemmaTake $\mathsf{U}_n$ to be the latent variable from Lemma (ref) and $\mathsf{W}_n = n^{-\frac{1}{4}}\mathsf{U}_n$. Then \begin{enumerate} • High-temparature case: Suppose $h \neq 0$ or $h = 0, \beta < 1$. Then $\lVert \mathsf{U}_n - \mathbb{E}[\mathsf{U}_n] \rVert_{\psi_2} \lesssim 1$. • Critical-temparature case: Suppose $h = 0$ and $\beta = 1$. Then $\lVert \mathsf{U}_n \rVert_{\psi_2} \lesssim n^{1/4}$. • Low-temparature case: Suppose $h = 0$ and $\beta > 1$. Then condition on $\mathsf{U}_n \in \mathcal{C}_l$, $\lVert \mathsf{U}_n - \mathbb{E}[\mathsf{U}_n|\operatorname{sgn}(\mathsf{U}_n) = \ell] \rVert_{\psi_2} \lesssim 1$. • Drifting sequence case: Suppose $h = 0$, $\beta = 1 - c n^{-\frac{1}{2}}, c \in \mathbb{R}^+$. Then $\lVert \mathsf{U}_n \rVert_{\psi_2} \leq \mathtt{C} n^{1/4}$ for large enough $n$ with $\mathtt{C}$ not depending on $\beta$. \end{enumerate}

Fix $\beta > 0$. We characterize the limiting distribution of $n^{-1}\sum_{i = 1}^n W_i X_i$ and the rate of convergence as $n \rightarrow \infty$ in the following lemma. In particular, we will see that the limiting distribution changes from a Gaussian distribution under high temperature, to a non-Gaussian distribution under critical temperature, to a Gaussian mixture under low temperature.

lemma[Fixed Temperature Berry-Esseen] Recall $\mca g_n = n^{-1}\sum_{i = 1}^n X_i (W_i - \pi)$. \begin{enumerate} • When $\beta < 1$ and $h = 0$ or $h \neq 0$, \begin{align*} \sup_{t\in\bb R}|\bb P_{\beta,h}(n^{\frac{1}{2}} \Big(\mathbb{E}[X_i^2](1 - \pi^2) + \mathbb{E}[X_i]^2\frac{\beta(1 - \pi^2)}{1 - \beta (1 - \pi^2)}\Big)^{-\frac{1}{2}} \mca g_{n} \leq t) -\Phi(t)| =O(n^{-\frac{1}{2}}). \end{align*} • When $\beta = 1$ and $h = 0$, denote $F_0(t) = \frac{\int_{-\infty}^t \exp(-z^4/12)d z}{\int_{-\infty}^{\infty} \exp(-z^4/12)d z}, t \in \mathbb{R}$, then \begin{align*} \sup_{t\in\bb R}|\bb P_{\beta,h}(n^{\frac{1}{4}} \mathbb{E}[X_i]^{-1} \mca g_n \leq t)-F_0(t)| =O((\log n)^3n^{-\frac{1}{2}}). \end{align*} • When $\beta > 1$ and $h = 0$, denote $\mca g_{n,\ell} = \frac{1}{n}\sum_{i = 1}^n X_i (W_i - \pi_\ell)$, then \begin{align*} \sup_{t\in\bb R}|\bb P_{\beta,h}(n^{\frac{1}{2}} \Big(\mathbb{E}[X_i^2](1 - \pi_\ell^2) + \mathbb{E}[X_i]^2\frac{\beta(1 - \pi_\ell^2)}{1 - \beta (1 - \pi_\ell^2)}\Big)^{-\frac{1}{2}} \mca g_{n,\ell} \leq t| & \operatorname{sgn}(\mca m) = \ell) -\Phi(t)| \\ & =O(n^{-\frac{1}{2}}), \quad t \in \{ -,+\}. \end{align*} \end{enumerate}
remarkLemma (ref)(3) and Lemma (ref)(3) together implies when $h = 0, \beta > 1$, condition on $\operatorname{sgn}(\mca m) = \ell$, $\lVert n^{-1/2} \mathsf{U}_n - \pi_{\ell} \rVert_{\psi_2} \lesssim n^{-1/2}$.
lemma[Size-Dependent Temperature Berry-Esseen when $h = 0$] Suppose $\mathsf{Z}$ is a standard Gaussian random variable. (1) Suppose $\beta_n = 1 + c n^{-\frac{1}{2}}$ and $h = 0$, where $c < 0$. Then \begin{align*} \sup_{t \in \mathbb{R}}\bigg|\mathbb{P}_{\beta_n, h}(n^{\frac{1}{4}} \mca g_n \leq t) - \mathbb{P} ( n^{-\frac{1}{4}}\mathbb{E}[X_i^2]^{\frac{1}{2}}\mathsf{Z} +\beta_n^{\frac{1}{2}}\mathbb{E}[X_i]\mathsf{W}_c \leq t)\bigg| = O((\log n)^{3}n^{-\frac{1}{2}}), \end{align*} where $O(\cdot)$ is up to a universal constant, and recall from Theorem 3.1 in the main paper that $\mathsf{W}_c$ is a random variable independent to $\mathsf{Z}$ with cummulative distribution function \begin{align*} \mathbb{P}[\mathsf{W}_c \leq w] = \frac{\int_{-\infty}^w \exp (-\frac{x^4}{12}-\frac{c x^2}{2})dx}{\int_{-\infty}^{\infty}\exp(-\frac{x^4}{12}-\frac{c x^2}{2})dx}, \qquad w \in \mathbb{R},\quad c \in \mathbb{R}_+. \end{align*} (2) Suppose $\beta_n = 1 + c n^{-1/2}$ and $h = 0$, where $c > 0$. Then \begin{align*} \sup_{c \in \mathbb{R}^+}\sup_{t \in \mathbb{R}}\bigg|\mathbb{P}_{1 + cn^{-1/2}, h} (n^{\frac{1}{4}} \mca g_n \leq t| \mca m \in \ca I_{c,n,\ell}) - \mathbb{P} ( n^{-\frac{1}{4}}\mathbb{E}[X_i^2]^{\frac{1}{2}}\mathsf{Z} +\beta_n^{\frac{1}{2}}\mathbb{E}[X_i]\mathsf{W}_{c,n} \leq t & |\mathsf{W}_{c,n} \in \ca I_{c,n,\ell})\bigg| \\ & = O((\log n)^{3}n^{-\frac{1}{2}}), \end{align*} where with $v_{n,+}$ and $v_{n,-}$ the positive and negative solutions to $x = \tanh(\beta_n x)$, and \begin{align*} a_{c,n} & = v_{n,+}^2 - c n^{-1/2}, \\ b_{c,n} & = 2 (1 + c n^{-1/2} - v_{n,+}^2) v_{n,+}^2, \\ c_{c,n} & = 2 (1 + c n^{-1/2} - v_{n,+}^2) (1 + c n^{-1/2} - 3 v_{n,+}^2), \end{align*} where $\mathsf{W}_{c,n}$ is a random variable taking values in $\mathbb{R}$ with density at $w \in \mathbb{R}$ proportional to $\exp(-h_{c,n}(w))$ independent to $\mathsf{Z}$, \begin{align*} & h_{c,n}(w) = \frac{\sqrt{n}a_{c,n}}{2} (w - n^{1/4} v_{n,\operatorname{sgn}(w)})^2 + \frac{n^{1/4}b_{c,n}}{6} (w - n^{1/4} v_{n, \operatorname{sgn}(w)})^3 + \frac{c_{c,n}}{24}(w - n^{1/4} v_{n,\operatorname{sgn}(w)})^4, \end{align*} and $\ca I_{c,n,-} = (-\infty,K_{c,n,-})$ and $\ca I_{c,n,+} = (K_{c,n,+},\infty)$ such that $\mathbb{E}[\mathsf{W}_{c,n}|\mathsf{W}_{c,n} \in \ca I_{c,n,\ell}] = n^{1/4} v_{c,n,\ell}$ for $\ell \in \{-,+\}$, and $O(\cdot)$ is up to a universal constant.
remarkIn (2), we consider drifting from the low temperature regime to the critical temperature regime. In Lemma (ref) we show $\mca g_n$ concentrates on the conditional means given $\operatorname{sgn}(\mca m)$ in the low temperature regime, whereas it concentrates on the unconditional mean in the critical temperature regime. The drifting region $\ca I_{c,n,\ell}$ captures this effect. $\ca I_{c,n,\ell} = (-\infty, 0)$ or $(0,\infty)$ when $c = 0$, and $\ca I_{c,n,\ell} = \mathbb{R}$ when $c = \infty$.
lemma[$\sqrt{n}$-sequence is knife-edge] Suppose $h = 0$. (1) Suppose $|\beta_n - 1| = o(n^{-\frac{1}{2}})$, then \begin{align*} \sup_{t \in \mathbb{R}}\bigg|\mathbb{P}_{\beta_n, h}(n^{\frac{1}{4}} \mca g_n \leq t) - \mathbb{P} (\mathbb{E}[X_i]\mathsf{W}_0 \leq t)\bigg| = o(1). \end{align*} (2) Suppose $1 - \beta_n \gg n^{-\frac{1}{2}}$, then \begin{align*} \sup_{t \in \mathbb{R}}\bigg|\mathbb{P}_{\beta_n, h}(\mathbb{V}[\mca g_n]^{-\frac{1}{2}} \mca g_n \leq t) - \Phi(t)\bigg| = o(1). \end{align*} (3) Suppose $\beta_n - 1 \gg n^{-\frac{1}{2}}$, then for $\ell \in \{-,+\}$, \begin{align*} \sup_{t \in \mathbb{R}} \bigg|\mathbb{P}_{\beta_n, h} \Big(\mathbb{V}[\mca g_n|\mca m \in \ca I_{\ell}])^{-\frac{1}{2}}(\mca g_n - \mathbb{E}[\mca g_n|\mca m \in \ca I_{\ell}]) \leq t \Big) - \Phi(t) \bigg| = o(1), \end{align*} where $\ca I_+ = [0,\infty)$ and $\ca I_- = (-\infty,0)$.
lemma[Fixed Temperature Berry-Esseen with Multivariate Multiplier] Suppose $\mathbf{W}$ satisfies Assumption (ref), and $\mathbf{X}_1, \cdots, \mathbf{X}_n$ are i.i.d random vectors taking values in $\mathbb{R}^d$, independent to $\mathbf{W}$. Suppose there exists some constant $b > 0$ such that $\mathbb{E}[X_{ij}^2] \geq b$ for all $j = 1, \cdots, d$, and for some sequence of constants $B_n \geq 1$, $|X_{ij}| \leq B_n$ for all $i = 1, \cdots, n$ and $j = 1, \cdots, d$. Let $\mathcal{R}$ be the collection of all hyperrectangles in $\mathbb{R}^d$. \begin{enumerate} • When $\beta < 1$ and $h = 0$ or $h \neq 0$, \begin{align*} \sup_{A \in \mathcal{R}} \Big|\bb P_{\beta,h} \Big(\frac{1}{n} \sum_{i = 1}^n \mathbf{X}_i (W_i - \pi) \in A \Big) - \mathbb{P}(n^{-1/2}\boldsymbol{\Sigma}^{1/2}\mathsf{Z}_d + n^{-1/2}\boldsymbol{\eta} \mathsf{Z} \in A) \Big| = O\Big(\Big(\frac{B_n^2 \log (n)^7}{n}\Big)^{1/6}\Big), \end{align*} where $\boldsymbol{\Sigma} = (1 - \pi^2) \mathbb{E}[\mathbf{X}_i \mathbf{X}_i^{\top}]$, $\boldsymbol{\eta} = (\frac{\beta(1 - \pi^2)^2}{1 - \beta(1 - \pi^2)})^{1/2} \mathbb{E}[\mathbf{X}_i]$, and $\mathsf{Z}_d \sim \mathsf{N}(\mathbf{0}, \mathbf{I}_{d \times d})$ independent to $\mathsf{Z} \sim \mathsf{N}(0,1)$. • When $\beta = 1$ and $h = 0$, \begin{align*} \sup_{A \in \mathcal{R}} \Big|\bb P_{\beta,h} \Big(\frac{1}{n} \sum_{i = 1}^n \mathbf{X}_i W_i \in A \Big)-\mathbb{P}(n^{-1/2}\boldsymbol{\Sigma}^{1/2}\mathsf{Z}_d + n^{-1/4}\mathbb{E}[\mathbf{X}_i] \mathsf{R} \in A) \Big| = O\Big(\Big(\frac{B_n^2 \log (n)^7}{n}\Big)^{1/6}\Big). \end{align*} where $\mathsf{R}$ be a random variable with cummulative distribution function $F_0(t) = \frac{\int_{-\infty}^t \exp(-z^4/12)d z}{\int_{-\infty}^{\infty} \exp(-z^4/12)d z}$, $t \in \mathbb{R}$, independent to $\mathsf{Z}_d$. • When $\beta > 1$ and $h = 0$, for $\ell = -,+$, \begin{align*} \sup_{A \in \mathcal{R}} \Big|\bb P_{\beta,h} \Big(\frac{1}{n} \sum_{i = 1}^n \mathbf{X}_i (W_i - \pi_{\ell}) \in A \Big| \operatorname{sgn}(\mca m) = \ell \Big) - \mathbb{P}(n^{-1/2}\boldsymbol{\Sigma}^{1/2}\mathsf{Z}_d & + n^{-1/2} \boldsymbol{\eta} \mathsf{Z} \in A) \Big| \\ & = O\Big(\Big(\frac{B_n^2 \log (n)^7}{n}\Big)^{1/6}\Big), \end{align*} where $\boldsymbol{\Sigma} = (1 - \pi_{+}^2) \mathbb{E}[\mathbf{X}_i \mathbf{X}_i^{\top}]$, $\boldsymbol{\eta} = (\frac{\beta(1 - \pi_{+}^2)^2}{1 - \beta(1 - \pi_{+}^2)})^{1/2} \mathbb{E}[\mathbf{X}_i]$. \end{enumerate}

Proof Sketch of Lemma (ref)

The magnetization $n^{-1}\sum_{i = 1}^n W_i$ has been studied using Stein's method eichelsbacher2010stein,chatterjee2010spin. Due to the multipliers, the Stein's method can not be directly applied to $n^{-1}\sum_{i = 1}^n X_i W_i$. We use a proof strategy based on the de Finetti's Lemma in Lemma (ref): There exists a latent variable $\mathsf{U}_n$ such that $W_1,\cdots,W_n$ are i.i.d condition on $\mathsf{U}_n$. Moreover, the density of $\mathsf{U}_n$ satisfies $f_{\mathsf{U}_n}(u) \propto \exp (- 1/2 u^2 + n \log \cosh( \sqrt{\beta/n} u)), u \in \mathbb{R}$.

We provide a proof sketch of Lemma (ref) (1) only. Throughout, take $c_{n, \beta} = \sqrt{n}(\beta - 1)$.

center[center omitted — 56 chars of source]

$W_i$'s are i.i.d condition on $\mathsf{U}_n$ with

align*[align* omitted — 335 chars of source]

Apply Berry-Esseen Theorem conditional on $\mathsf{U}_n$, and take $\mathsf{Z} \sim \mathsf{N}(0,1)$ independent to $\mathsf{U}_n$,

align*[align* omitted — 294 chars of source]

Lemma 2 in the supplementary material shows $\lVert \mathsf{U}_n \rVert_{\psi_2} \leq \mathtt{C}n^{1/4}$, hence by concentration arguments, $$\sup_{t \in \mathbb{R}}|\mathbb{P}(n^{-1}\sum_{i = 1}^n X_i W_i \leq t) - \mathbb{P}( \sqrt{v(\mathsf{U}_n)}\mathsf{Z} + \sqrt{n}e(\mathsf{U}_n)\leq t)|\leq K n^{-1/2}.$$

center[center omitted — 93 chars of source]

Consider $\mathsf{W}_n = n^{-1/4}\mathsf{U}_n$. By a change of variable from $\mathsf{U}_n$ and Taylor expand what is inside the exponent, we show $\mathsf{W}_n$ has density satisfying

align*[align* omitted — 143 chars of source]

where $g$ is a bounded smooth function. We show based on sub-Gaussianity of $\mathsf{W}_{n}$, with an upper bound of sub-Gaussian norm not depending on $\beta$, that the sixth order term is negligible and $$\sup_{t \in \mathbb{R}}|\mathbb{P}(\mathsf{W}_n \leq t) - \mathbb{P}(\mathsf{W} \leq t)| = O(\log^3 n n^{-1/2}),$$ where $\mathsf{W}$ has density proportional to $\exp (-c_{\beta,n}/2 w^2 - \beta_n^2 w^4 / 12)$.

center[center omitted — 55 chars of source]

Since $\mathsf{Z}$ is independent to $(\mathsf{U}_n, \mathsf{W}_n)$, we use data processing inequality and the previous two steps to show $n^{-1}\sum_{i = 1}^n X_i W_i$ is close to $n^{-1/4}v(n^{1/4}\mathsf{W}_{c_{\beta,n}})^{1/2}\mathsf{Z} + n^{1/4}e(n^{1/4}\mathsf{W}_{c_{\beta,n}}))$. Lemma 2 in the supplementary appendix imply $\lVert \mathsf{W}_{c_{\beta,n}} \rVert_{\psi_2} \leq \mathtt{K}$. By Taylor expanding $e(\cdot)$ and $v(\cdot)$ at $0$, we show $n^{1/4}e(\mathsf{U}_n)$ is close to $\mathbb{E}[X_i] \mathsf{W}_{c_{\beta,n}}$ and $n^{-1/4}\sqrt{v(\mathsf{U}_n)}\mathsf{Z}$ is close to $n^{-1/4}v(n^{1/4}\mathsf{W}_{c_{\beta,n}})^{1/2}\mathsf{Z}$.

Pseudo-Likelihood Estimator for Curie-Weiss Regimes

lemma[No Consistent Variance Estimator] Suppose Assumptions 1,2,3 in the main paper hold. Then there is no consistent estimator of $n \mathbb{V}[\widehat{\tau}_n - \tau_n]$.

The pseudo-likelihood estimator for Curie-Weiss regime with $h = 0$ is given by

align*[align* omitted — 260 chars of source]
lemma[Fixed Temperature Distribution Approximation] (1) If $\beta \in [0,1)$ and $h = 0$, then \begin{align*} \wh \beta \overset{d}{\to} \max\bigg\{1 - \frac{1 - \beta}{\chi^2(1)},0\bigg\}. \end{align*} (2) If $\beta = 1$ and $h = 0$, then \begin{align*} n^{\frac{1}{2}}(1 - \wh\beta)&\overset{d}{\to} \max \bigg\{\frac{1}{\mathsf{W}_0^2} - \frac{\mathsf{W}_0^2}{3},0\bigg\}. \end{align*} (3) If $\beta > 1$ and $h = 0$, we define an unrestricted pseud-likelihood estimator, \begin{align*} \nonumber \widehat{\beta}_UR = \operatorname*{arg\,max}_{\beta \in \mathbb{R}} \log \mathbb{P}_{\beta} \left( W_i \mid \mathbf{W}_{-i} \right) = \sum_{i \in [n]} -\log \bigg( \frac{1}{2}W_i \tanh(\beta \mca m_i) + \frac{1}{2}\bigg). \end{align*} Then \begin{align*} \sup_{t \in \mathbb{R}}|\mathbb{P}(n^{1/2}(\widehat{\beta}_UR - \beta) \leq t |\mca m \in \ca I_\ell) - \mathbb{P}((\frac{1 - \beta(1 - \pi_{\ell}^2)}{1 - \pi_{\ell}^2})^{1/2}\mathsf{Z} \leq t)| = o(1). \end{align*}
lemma[Drifting Temperature Distribution Approximation] For any $\beta \in [0,1]$ and $h = 0$, define $c_{\beta,n} = \sqrt{n}(1 - \beta)$, and suppose $z_{\beta,n}$ is a random variable such that \begin{align*} \mathbb{P}(z_{\beta,n} \leq t) = \mathbb{P}(\mathsf{Z} + n^{\frac{1}{4}}\mathsf{W}_{c_{\beta,n}} \leq t), \qquad t \in \mathbb{R}. \end{align*} then \begin{align*} \sup_{\beta \in [0,1]}\sup_{t \in \mathbb{R}}|\mathbb{P} (1 - \wh \beta \leq t) - \mathbb{P} (\min\{\max\{z_{\beta,n}^{-2} - \frac{1}{3n} z_{\beta,n}^2,0\},1\} \leq t)| = o(1). \end{align*}

Stochastic Linearization

Recall $\mathbf{W} = (W_1, \cdots, W_n)$ satisfies Assumption (ref). And for notational simplicity, let $g_i$ be the function such that

align*[align* omitted — 132 chars of source]

We denote $M_i = \sum_{j \neq i} E_{ij} W_i$, $N_i = \sum_{j \neq i} E_{ij}$. Then

align*[align* omitted — 155 chars of source]

Recall our definition of regimes: High temperature regime $\mathcal{A}_H = \{(\beta,h) \in [0,\infty) \times \mathbb{R}: h \neq 0 \text{ or } h = 0, \beta < 1\}$, critical temperature regime $\mathcal{A}_C = \{(1, 0)\}$, and low temperature regime $\mathcal{A}_L = \{(\beta,h) \in [0,\infty) \times \mathbb{R}: h = 0, \beta > 1\}$. Define the following rates that will be used in the convergence analysis:

align*[align* omitted — 412 chars of source]

and

align*[align* omitted — 429 chars of source]

Throughout Section (ref), we work with $(\beta, h) \in \mathcal{A}_H \cup \mathcal{A}_C$, and let $\pi$ be the unique solution to $x = \tanh(\beta x + h)$. Then friedli2017statistical implies $\mathbb{E}[W_i] = \pi + O(n^{-1})$. Let $\mca m = n^{-1} \sum_{i= 1}^n W_i$ and $\mca m_i = n^{-1} \sum_{j \neq i} W_j$.

The Unbiased Estimator

Denote $p_i = \bb P_{\beta, h}(W_i = 1; \mathbf{W}_{-i}) = \left(\exp \left(-2\beta \mca m_i - 2 h\right) + 1\right)^{-1}$. We propose an unbiased estimator given by

align*[align* omitted — 138 chars of source]
lemma[Unbiased Estimator] $\wh\tau_{n,\text{UB}}$ is an unbiased estimator for $\tau_n$ in the sense that, \begin{align*} \mathbb{E}[\wh\tau_{n,UB}|\mathbf{E},(f_i)_{i \in [n]}] = \tau_n. \end{align*}

We will show the followings have weak limits:

align*[align* omitted — 134 chars of source]

W.l.o.g, we analyse the error for treated data, the error for control data follows in the same way. First, decompose by

gather*[gather* omitted — 544 chars of source]
lemmaSuppose Assumption (ref), and Assumptions 2, and 3 from the main paper hold with $(\beta, h) \in \mathcal{A}_H \cup \mathcal{A}_C$. Then \begin{align*} \Delta_1 - \mathbb{E}[\Delta_1|\mathbf{E},(g_i)_{i \in [n]}] = n^{-\mathtt{a}_{\beta, h}}\sum_{i = 1}^n \Big(\frac{g_i(1,\pi)}{1 + \pi} + \frac{g_i(-1,\pi)}{1 - \pi} & - \beta \mathtt{d}\Big) \left(W_i - \pi\right) \\ + O_{\psi_{2},tc}(\sqrt{\log n}n^{-\mathtt{r}_{\beta, h}}), \end{align*} where $\mathtt{d} = (1 - \pi)\mathbb{E}[g_i(1,\pi)] + (1 + \pi)\mathbb{E}[g_i(-1,\pi)]$.

Now consider $\Delta_2$. Since $\frac{T_i}{p_i} = \frac{T_i - p_i}{p_i} + 1$, we have the decomposition,

equation[equation omitted — 272 chars of source]

where

align*[align* omitted — 466 chars of source]

where $\eta_i^{\ast}$ is some random quantity between $ \frac{M_i}{N_i}$ and $\pi$. Define $b_i = \sum_{j \neq i} \frac{E_{ij}}{N_j} Y_j^{\prime}\left(1, \pi \right)$. Then by reordering the terms,

align*[align* omitted — 102 chars of source]
lemmaAssumption (ref), and Assumptions 2, and 3 from the main paper hold with $(\beta, h) \in \mathcal{A}_H \cup \mathcal{A}_C$. Then condition on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A} = \{A \in \bb R^{n \times n}: \min_{i \in [n]} \sum_{j \neq i}A_{ij} \geq 32 \log n\}$, \begin{align*} \Delta_{2,2} = O_{\psi_{2},tc}\bigg(\log n \max_{i \in [n]}\mathbb{E}[N_i|\mathbf{U}]^{-1/2}\bigg) + O_{\psi_{\beta,\gamma},tc}(\sqrt{\log n} n^{-\mathtt{r}_{\beta, h}}). \end{align*}

For the term $\Delta_{2,3}$, we further decompose it into two parts:

align*[align* omitted — 65 chars of source]

where

align*[align* omitted — 510 chars of source]
lemmaAssumption (ref), and Assumptions 2, and 3 from the main paper hold with $(\beta, h) \in \mathcal{A}_H \cup \mathcal{A}_C$. Then condition on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A} = \{A \in \bb R^{n \times n}: \min_{i \in [n]} \sum_{j \neq i}A_{ij} \geq 32 \log n\}$, \begin{align*} & \Delta_{2,3,1} - \mathbb{E}[\Delta_{2,3,1}|\mathbf{E},(f_i)_{i \in [n]}] \\ = & O_{\psi_{\mathtt{p}_{\beta,h}/2}}(n^{-\mathtt{r}_{\beta, h}}) + O_{\psi_{\beta,h},tc}(\max_i \mathbb{E}[N_i|\mathbf{U}]^{-1/2}) + O_{\psi_1, tc}(n^{-1/2}) \\ & \qquad + O_{\psi_2,tc}(n^{\frac{1}{2} - \mathtt{a}_{\beta, h}}\max \mathbb{E}[N_i|\mathbf{U}]^{-1/2}). \end{align*}
lemmaAssumption (ref), and Assumptions 2, and 3 from the main paper hold with $(\beta, h) \in \mathcal{A}_H \cup \mathcal{A}_C$. If $g_i(1, \cdot)$ and $g_i(-1,\cdot)$ are $4$-times continuously differentiable, then condition on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A}$, \begin{align*} & \Delta_{2,3,2} - \mathbb{E}[\Delta_{2,3,2}|\mathbf{E},(f_i)_{i \in [n]}] \\ = & O_{\psi_{\mathtt{p}_{\beta, h}/2},tc}((\log n)^{-1/\mathtt{p}_{\beta, h}}n^{-2\mathtt{r}_{\beta, h}}) + O_{\psi_{1},tc}((\log n)^{-1/\mathtt{p}_{\beta, h}} (\min_i\mathbb{E}[N_i|\mathbf{U}])^{-1}) \\ & + O_{\psi_1, tc} \left(n^{1/2 - \mathtt{a}_{\beta, h}} \left( \frac{\max_i \mathbb{E}[N_i|\mathbf{U}]^3}{\min_i \mathbb{E}[N_i|\mathbf{U}]^4}\right)^{1/2} \right) + O_{\psi_{2/(p+1)},tc}\left(n^{\mathtt{r}_{\beta, h}} (\min_i \mathbb{E}[N_i|\mathbf{U}]^{-(p+1)/2}) \right). \end{align*}

Hajek Estimator

lemmaAssumption (ref), and Assumptions 2, and 3 from the main paper hold with $(\beta, h) \in \mathcal{A}_H \cup \mathcal{A}_C$. Then \begin{align*} \widehat{\tau}_n - \widehat{\tau}_{n,UB} = - \bigg(\frac{\mathbb{E}[g_i(1,\pi)]}{\pi + 1} + \frac{\mathbb{E}[g_i(-1,\pi)]}{1 - \pi} \bigg)(1 - \beta(1 - \pi^2))(\mca m - \pi) + O_{\psi_1}(n^{-2\mathtt{r}_{\beta, h}}). \end{align*}

Stochastic Linearization

lemmaSuppose Assumptions (ref), and Assumptions 2, and 3 from the main paper hold with $(\beta, h) \in \mathcal{A}_H \cup \mathcal{A}_C$. Define \begin{align*} R_i = \frac{g_i(1,\pi)}{1 + \pi} + \frac{g_i(-1,\pi)}{1 - \pi}, \qquad Q_i = \mathbb{E} [\frac{G(U_i,U_j)}{\mathbb{E}[G(U_i,U_j)|U_j]}(g_j^{\prime}(1,\pi) - g_j^{\prime}(-1,\pi))|U_i]. \end{align*} Then, \begin{align*} \sup_{t \in \mathbb{R}} \big|\mathbb{P}_{\beta,h}(\wh \tau_n - \tau_n \leq t) - \mathbb{P}_{\beta,h}(\frac{1}{n}\sum_{i =1}^n (R_i - \mathbb{E}[R_i] + Q_i)(W_i - \pi) \leq t)\big| = O\Big(\frac{\log n}{\sqrt{n \rho_n}} + \mathtt{r}_{n,\beta}\Big), \end{align*} where $\mathtt{r}_{n,\beta} = \sqrt[4]{n}\sqrt{\log n}(n \rho_n)^{-\frac{p+1}{2}}$ if $\beta = 1, h = 0$; and $\sqrt{n \log n}(n \rho_n)^{-\frac{p+1}{2}}$ if $\beta < 1$ or $h \neq 0$.
lemmaSuppose Assumption (ref), and Assumptions 2, and 3 from the main paper with $h = 0$, $\beta \in [0,1]$. Define \begin{align*} R_i = g_i(1,0) + g_i(-1,0), \qquad Q_i = \mathbb{E} [\frac{G(U_i,U_j)}{\mathbb{E}[G(U_i,U_j)|U_j]}(g_j^{\prime}(1,0) - g_j^{\prime}(-1,0))|U_i]. \end{align*} Then, \begin{align*} \sup_{\beta \in [0,1]}\sup_{t \in \mathbb{R}} \big|\mathbb{P}_{\beta,h}(\wh \tau_n - \tau_n \leq t) - \mathbb{P}_{\beta,h}(\frac{1}{n}\sum_{i =1}^n (R_i - \mathbb{E}[R_i] + Q_i) W_i \leq t)\big| = o(1). \end{align*}

Jacknife-Assisted Variance Estimation

lemmaSuppose Assumptions 1,2,3,4 from the main paper hold with $h = 0$, and $n \rho_n^3 \rightarrow \infty$ as $n \rightarrow \infty$. Suppose the non-parametric learner $\wh f$ satisfies $\wh f(\ell,\cdot) \in C_2([0,1])$, and $|\wh f(\ell,\frac{1}{2}) - f(\ell,\frac{1}{2})| = o_\mathbb{P}(1)$, $|\partial_2 \wh f(\ell,\frac{1}{2}) - \partial_2 f(\ell,\frac{1}{2})| = o_\mathbb{P}(1)$, for $\ell \in \{0,1\}$, where the rate in $o_\mathbb{P}(\cdot)$ does not depend on $\beta$. Suppose $\wh K_n$ is the jacknife estimator from Algorithm 2. Then \begin{align*} \wh K_n = \mathbb{E}[(R_i - \mathbb{E}[R_i] + Q_i)^2] + o_\mathbb{P}(1), \end{align*} where the rate in $o_\mathbb{P}(1)$ also does not depend on $\beta$.

Here we give a local-polynomial based learner $\wh f$ that satisfies requirements of Lemma (ref) (hence Theorem 4 in the main paper.)

lemmaUse a local polynomial estimator to fit the potential outcome functions: Take \begin{align*} \widehat{f}(1,x) & := \widehat{\gamma}_0 + \widehat{\gamma}_1 x, \\ (\widehat{\gamma}_0, \widehat{\gamma}_1) &:= \operatorname*{arg\,min}_{\gamma_0,\gamma_1} \sum_{i = 1}^n \Big(Y_i - \gamma_0 - \gamma_1 \frac{M_i}{N_i} \Big)^2 K_h \Big(\frac{M_i}{N_i}\Big)\mathbbm{1}(T_i = 1), \end{align*} where $K_h(\cdot) = h^{-1}K(\cdot/h)$ where $K$ is a kernel function, $h$ is the optimal bandwidth. Then $\widehat{f}(1,0) = f(1,0) + o_{\mathbb{P}}(1), \partial_2\widehat{f}(1,0) = \partial_2 f(1,0) + o_{\mathbb{P}}(1)$, the same for control group. Moreover, the rate of convergence can be made not depending on $\beta$.

Additional Distributional Results

This section presents the additional distributional results in the appendix. We continue to use the notations defined at the beginning of Section (ref).

Low Temperature Treatment Assignment

Recall we consider a conditional estimand given by

align*[align* omitted — 204 chars of source]

where $\operatorname{sgn}(\mca m) = \operatorname{sgn}(2 n^{-1}\sum_{i = 1}^n T_i - 1)$. Let $\pi_*$ be the positive root of $x = \tanh(\beta x)$, and take $\pi_+ = 1/2 + \pi_*/2$, $\pi_- = 1/2 - \pi_*/2$.

lemmaSuppose Assumptions (ref), and Assumptions 2, and 3 from the main paper hold with $\beta > 1$ and $h = 0$. Define \begin{align*} R_{i,\ell} = \frac{g_i(1,\pi_\ell)}{1 + \pi_{\ell}} + \frac{g_i(-1,\pi_\ell)}{1 - \pi_{\ell}}, \quad Q_{i,\ell} = \mathbb{E} \Big[\frac{G(U_i,U_j)}{\mathbb{E}[G(U_i,U_j)|U_j]}(g_j^{\prime}(1,\pi_{\ell}) - g_j^{\prime}(-1,\pi_{\ell}))\Big|U_i\Big], \quad \ell \in \{-,+\}. \end{align*} Then, \begin{align*} \sup_{t \in \mathbb{R}} \max_{\ell \in \{-,+\}} \big|\mathbb{P}_{\beta,h}(\wh \tau_n - \tau_{n,\ell} \leq t | \operatorname{sgn}(\mca m) = \ell) - \mathbb{P}(\frac{1}{n}\sum_{i =1}^n (R_{i,\ell} & - \mathbb{E}[R_{i,\ell}] + Q_{i,\ell})( W_i - \pi_{\ell}) \leq t)\big| \\ & = O\Big(\frac{\log n}{\sqrt{n \rho_n}} + \sqrt{n \log n}(n \rho_n)^{-\frac{p+1}{2}}\Big). \end{align*}
lemmaSuppose Assumptions (ref), and Assumptions 2, and 3 from the main paper hold with $\beta > 1$ and $h = 0$. Then \begin{align*} \sup_{t \in \mathbb{R}} \max_{\ell \in \{-,+\}}\big|\mathbb{P}_{\beta,h}(\wh \tau_n - \tau_{n,\ell} \leq t|\operatorname{sgn}(\mca m) = \ell) - L_n(t;\beta,\kappa_{1,\ell},\kappa_{2,\ell})\big| = O\Big(\sqrt{\frac{n \log n}{(n \rho_n)^{p+1}}} + \frac{\log n}{\sqrt{n \rho_n}}\Big), \end{align*} where with $\mathsf{Z} \sim \mathsf{N}(0,1)$, \begin{align*} L_{n}(t;\beta,\kappa_{1,\ell},\kappa_{2,\ell}) = \mathbb{P}\Big\{n^{-1/2}\Big(\kappa_{2,\ell} (1 - \pi_*^2) + \kappa_{1,\ell}^2 \frac{\beta (1 - \pi_*^2)}{1 - \beta (1 - \pi_*^2)}\Big)^{1/2}\mathsf{Z} \leq t\Big\}, \end{align*} where $\kappa_{s,\ell} = \mathbb{E}[(R_{i,\ell} - \mathbb{E}[R_{i,\ell}] + Q_{i,\ell})^s]$ for $s = 1,2$ and $\ell = -, +$.

Asymmetric Treatment Assignment

Recall the following treatment assignment model from Section A.1: For $\beta \in [0,\infty)$ and $h \neq 0$, the treatment vector $\mathbf{T} = (T_1, \cdots, T_n)$ satisfies a distribution on $\{0,1\}^n$ such that

align*[align* omitted — 208 chars of source]

Let $\pi$ be the unique solution to $x = \tanh(\beta x + h)$.

lemmaSuppose Assumptions (ref), and Assumptions 2, and 3 from the main paper hold with $\beta > 0$ and $h \neq 0$. Define \begin{align*} R_i = \frac{g_i(1,\pi)}{1 + \pi} + \frac{g_i(-1,\pi)}{1 - \pi}, \qquad Q_i = \mathbb{E} \Big[\frac{G(U_i,U_j)}{\mathbb{E}[G(U_i,U_j)|U_j]}(g_j^{\prime}(1,\pi) - g_j^{\prime}(-1,\pi))\Big|U_i\Big]. \end{align*} Then, \begin{align*} \sup_{t \in \mathbb{R}} \big|\mathbb{P}_{\beta,h}(\wh \tau_n - \tau_n \leq t) - \mathbb{P}(\frac{1}{n}\sum_{i =1}^n (R_i - \mathbb{E}[R_i] + Q_i)(W_i - \pi) \leq t)\big| = O\Big(\frac{\log n}{\sqrt{n \rho_n}} + \sqrt{n \log n}(n \rho_n)^{-\frac{p+1}{2}}\Big). \end{align*}
lemmaSuppose Assumptions (ref), and Assumptions 2, and 3 from the main paper hold with $\beta > 0$ and $h \neq 0$. Then \begin{align*} \sup_{t \in \mathbb{R}} \big|\mathbb{P}[\wh \tau_n - \tau_n \leq t] - L_n(t;\beta,h,\kappa_1,\kappa_2)\big| = O\Big(\frac{\log n}{\sqrt{n \rho_n}} + \sqrt{n \log n}(n \rho_n)^{-(p+1)/2}\Big), \end{align*} where $L_n(\cdot;\beta,h,\kappa_1,\kappa_2)$ is as follows: \begin{align*} L_n(t;\beta,h,\kappa_1,\kappa_2) = \mathbb{P}_{\beta,h}\Big[n^{-1/2}\Big(\kappa_2 (1 - \pi^2) + \kappa_1^2 \frac{\beta (1 - \pi^2)^2}{1 - \beta (1 - \pi^2)}\Big)^{1/2}\mathsf{Z} \leq t\Big] \end{align*} with $\mathsf{Z} \thicksim \mathsf{N}(0,1)$, and $\kappa_s= \mathbb{E}[(R_i - \mathbb{E}[R_i] + Q_i)^s]$ for $s = 1, 2$.

Ising Block Treatment Assignment

Recall our notations: For block $k$ with $h_k \neq 0$ or $h_k = 0, 0 \leq \beta_k \leq 1$, $\pi_k$ denotes the unique solution to $x = \tanh(\beta_k x + h_k)$. For block $k$ with $h_k = 0, \beta_k > 1$, $\pi_{k,+}$ and $\pi_{k,-}$ denote the unique positive and negative solutions to $x = \tanh(\beta_k x + h_k)$, respectively.

Due to the potential existence of low temperature blocks, we use $\boldsymbol{sgn}$ to collect the average spins in all low temperature blocks, and fill in the positions for high and critical temperature blocks with zeros, that is,

align*[align* omitted — 168 chars of source]

And we use $\mathscr{S}$ to denote the collection of all possible configurations of $\boldsymbol{sgn}$, that is,

align*[align* omitted — 140 chars of source]

Also we denote the conditional fixed point based on $\boldsymbol{sgn} = \mathbf{s}$ by

align*[align* omitted — 252 chars of source]

We denote by $\mathscr{R}$ the collection of all hyperrectangles in $\mathbb{R}^K$.

lemmaSuppose Assumptions 2, 3, and 6 from the main paper hold. Condition on $\boldsymbol{sgn} = \mathbf{s}$, \begin{align*} & \bigg\lVert \widehat{\boldsymbol{\tau}}_{n} - \boldsymbol{\tau}_n - \frac{1}{n} \sum_{l = 1}^K \sum_{i \in \mathcal{C}_l} \mathbf{S}_{l,i,(\mathbf{s})}(W_i - \pi_{l,(\mathbf{s})}) \bigg\rVert_2 = O_{\psi_1,tc}(\mathtt{r}_n). \end{align*} where $\mathbf{S}_{l,i,(\mathbf{s})} = (S_{1,l,i,(\mathbf{s})}, \cdots, S_{K,l,i,(\mathbf{s})})^{\mathbf{T}}$, where \begin{align*} S_{k,l,i,(\mathbf{s})} = Q_{i,(\mathbf{s})} + \mathbbm{1}(k = l) p_k^{-1}(R_{i,l,(\mathbf{s})} - \mathbb{E}[R_{i,l,(\mathbf{s})}]), \qquad 1 \leq k, l \leq K, 1 \leq i \leq n, \end{align*} with $\overline{\pi}_{(\mathbf{s})} = \sum_{k = 1}^K p_k \pi_{k,(\mathbf{s})}$, \begin{align*} R_{i,l,(\mathbf{s})} & = \frac{g_i(1, \overline{\pi}_{(\mathbf{s})})}{1 + \pi_{l,(\mathbf{s})}} + \frac{g_i(-1, \overline{\pi}_{(\mathbf{s})})}{1 - \pi_{l,(\mathbf{s})}}, \\ Q_{i,(\mathbf{s})} & = \mathbb{E} \Big[\frac{G(U_i, U_j)}{\mathbb{E}[G(U_i, U_j)|U_j]} (g_j^{\prime}(1, \overline{\pi}_{(\mathbf{s})}) - g_j^{\prime}(-1, \overline{\pi}_{(\mathbf{s})}))\Big| U_i\Big]. \end{align*} and $\mathtt{r}_n = \sqrt{\log n} \max_{1 \leq k \leq K} n^{-\mathtt{r}_{\beta_k, h_k}}(n \rho_n)^{-1/2} + (n \rho_n)^{-(p+1)/2}$.
lemmaSuppose Assumptions 2, 3, and 6 from the main paper hold. Condition on $\boldsymbol{sgn} = \mathbf{s}$, we have \begin{align*} \max_{\mathbf{s} \in \mathscr{S}} \sup_{A \in \mathscr{R}} & | \mathbb{P}_{\boldsymbol{\beta}, \boldsymbol{h}}(\widehat{\boldsymbol{\tau}}_n - \boldsymbol{\tau}_n \in A| \boldsymbol{sgn} = \mathbf{s}) - \\ & \mathbb{P}(n^{-1/2} \boldsymbol{\Sigma}_{(\mathbf{s})}^{1/2} \mathsf{Z}_K + n^{-1/2} \sum_{k \in \mathscr{H} \cup \mathscr{L}} p_k \sigma_{k,(\mathbf{s})} \mathbb{E}[\mathbf{S}_{k,i,(\mathbf{s})}] \mathsf{Z}_{(k)} + n^{-1/4} \sum_{k \in \mathscr{C}} p_k \mathbb{E}[\mathbf{S}_{k,i,(\mathbf{s})}] \mathsf{R}_{(k)} \in A)| \\ & \qquad \quad = O(n^{1/2} \mathtt{r}_n + (\log n)^{7/6} n^{-1/6}), \end{align*} where $\mathsf{Z}_K \sim \mathsf{N}(\mathbf{0},\mathbf{I}_{K \times K})$, $\mathsf{Z}_{(k)} \sim \mathsf{N}(0,1)$ for $k \in \mathscr{H} \cup \mathscr{L}$, and $\mathsf{R}_{(k)}$ has cummulative distribution function $F_0(t) = \frac{\int_{-\infty}^t \exp(-z^4/12)d z}{\int_{-\infty}^{\infty} \exp(-z^4/12)d z}, t \in \mathbb{R},$ for $k \in \mathscr{C}$, with $\mathsf{Z}_K$, $\mathsf{Z}_{(k)}, k \in \mathscr{H} \cup \mathscr{L}$ and $\mathsf{R}_{(k)}, k \in \mathscr{C}$ mutually independent, and \begin{align*} \boldsymbol{\Sigma}_{(\mathbf{s})} = (\sum_{k = 1}^K \mathbb{E}[\mathbf{S}_{k,i,(\mathbf{s})} \mathbf{S}_{k,i,(\mathbf{s})}^\top] (1 - \pi_{k,(\mathbf{s})}^2) p_k^2)^{1/2}. \end{align*}
remarkIf there is no low temperature block, then $\boldsymbol{sgn} = (0, \cdots, 0)$ almost surely, and $\mathsf{S}$ is the singleton set containing $(0, \cdots, 0)$. Hence the result reduces to the unconditional distributional approximation.

Proofs: Main Paper

Proof of Theorem 3.1

The conclusion follows from the stochastic linearization result in Lemma (ref), and the Berry-Esseen result for Curie-Weiss magnetization with independent multipliers in Lemma (ref) (1) and (2).

Proof of Theorem 3.2

The conclusion for Hajek estimator follows from the stochastic linearization result in Lemma (ref), and the (uniform in $\beta$) Berry-Esseen result for Curie-Weiss magnetization with independent multipliers in Lemma (ref) (1).

The conclusion for MPLE follows from Lemma (ref).

Proof of Lemma 3.1

The conclusion follows from Lemma (ref) and Lemma (ref).

Proof of Theorem 4.1

The uniform approximation for $\sqrt{n}(\wh \beta_n - 1)$ established in Lemma (ref) implies $$\inf_{\beta}\mathbb{P}_\beta(\beta \in \ca I(\alpha_1)) \geq \inf_{\beta}\mathbb{P}_\beta(\sqrt{n}(1 - \beta) \geq \mca q) \geq 1 - \alpha_1 + o_\mathbb{P}(1).$$ where $\mca q$ is the $\alpha_1$ quantile of $\min\{\max \{\mathsf{T}_{c_{\beta,n},n}^{-2} - \mathsf{T}_{c_{\beta,n},n}^2/(3n),0\},1\}$.

Then by a Bonferroni correction argument, the second step coverage can be lower bounded by

align*[align* omitted — 277 chars of source]

Observe that the event $\tau_n \in \widehat{\ca C}(\alpha_1,\alpha_2)$ conincides with the event $\wh \tau_n - \tau_n \in [\mathtt{L}, \mathtt{U}]$, where $\mathtt{U} = \sup_{\beta \in \ca I(\alpha_1)} H_n(1 - \frac{\alpha_2}{2};K_n, K_n,c_{\beta,n})$, $\mathtt{L} = \inf_{\beta \in \ca I(\alpha_1)} H_{n}(\frac{\alpha_2}{2};K_n, K_n, c_{\beta,n})$. Hence

align*[align* omitted — 564 chars of source]

Theorem 2 shows that the quantiles of the distributions of $\wh \tau_n - \tau_n$ can be uniformly approximated by quantiles from $H_n(\cdot;\kappa_1, \kappa_2,c_{\beta,n})$, if $\kappa_1$ and $\kappa_2$ are correctly specified, and the confidence interval is conservative, if we use upper bound $\mathtt{K}_n$ for $\kappa_1$ and $\kappa_2$. The conclusion then follows.

Proof of Theorem 4.2

The conclusion follows from Theorem 4.1 and Lemma (ref).

Proof of Lemma 5.1

The conclusion follows from Lemma (ref).

Proof of Lemma 5.2

The conclusion follows from Lemma (ref).

Proof of Lemma 5.3

The conclusion follows from Lemma (ref).

Proofs: Section (ref)

Proof of Lemma (ref)

Using Gaussian integral identity $\exp(v^2/2) = \frac{1}{\sqrt{2 \pi}} \int_{-\infty}^{\infty} \exp \left( - u^2/2 + uv\right) du$,

align*[align* omitted — 300 chars of source]

Proof of Lemma (ref)

Our proof is divided according to the different temperature regimes.

center[center omitted — 65 chars of source]

We introduce the handy notation given by $F(v):=-\frac{1}{2}v^2+\log\cosh(\sqrt{\beta}v+h)$. For the high temperature regime, we note that the term in the exponential can be expanded across its global minimum $v^*$ (which satisfies the first order stationary point condition given by $v^*=\sqrt{\beta}\tanh(\sqrt{\beta}v^*+h)$) by

align*[align* omitted — 210 chars of source]

Therefore, to obtain the limit of the expectation, we note that by the Laplace method given similar to the proof of Lemma (ref) and the definition of $\mathsf{V}_n:=n^{-1/2}\mathsf{U}_n$:

align*[align* omitted — 137 chars of source]

Then, we note that for $\ell\in\bb N$, when $h=0$ and $\beta<1$ we use the Laplace method again to obtain that for all $\ell\in\bb N$,

align*[align* omitted — 367 chars of source]

Then we can obtain that for all $t\in\bb R$, we have

align*[align* omitted — 376 chars of source]

which alternatively implies that

align[align omitted — 244 chars of source]
center[center omitted — 75 chars of source]

Then we study the critical temperature regime with $\beta=1$. Note that one has $\bb E[\mathsf{U}_n]=0$ and for all $\ell\in\bb N$ we have

align*[align* omitted — 194 chars of source]

Then we can obtain that $\ell\in\bb N$,

align*[align* omitted — 452 chars of source]

And we immediately obtain that

align*[align* omitted — 502 chars of source]

which finally leads to

align[align omitted — 171 chars of source]
center[center omitted — 72 chars of source]

We shall note that at the low temperature regime the function $F(v)$ has two symmetric global minima $v_1>0>v_2$, satisfying

align*[align* omitted — 167 chars of source]

Then we can check that by the Laplace method, for all $t>0$ (following the path given by the high temperature regime) we have

align*[align* omitted — 320 chars of source]

Then we similarly obtain that $ \bb E[\exp(t(\mathsf{V}_n-\bb E[\mathsf{V}_n|\mathsf{V}_n<0]))|\mathsf{V}_n<0]=\exp\left(\frac{(1+o(1))t^2}{2n(1-\sqrt{\beta}\operatorname{sech}^2(\sqrt{\beta}v_1))}\right)$. Hence we obtain that

align[align omitted — 304 chars of source]
center[center omitted — 56 chars of source]

Then we consider the drifting case.

First consider $\beta= 1-cn^{-\frac{1}{2}}$ with $c\in\bb R^+$ and $\beta \geq 0$. We will show that for any fixed $n$, $\lVert W_n \rVert_{\psi_2}$ is increasing in $\beta$ when $\beta \in [0,1]$. This will imply that in the drifting case, $\lVert W_n \rVert_{\psi_2}$ will be no larger than its value at the critical regime.

For a comparison argument, denote $F_{\beta}(v) = - \frac{1}{2}v^2 + \log \cosh(\sqrt{\beta} v)$. Let $0 < \beta_1 < \beta_2 \leq 1$. Then

align*[align* omitted — 146 chars of source]

where

align*[align* omitted — 203 chars of source]

Hence for any $n \in \mathbb{N}$ and $t > 0$,

align*[align* omitted — 131 chars of source]

increases as $\beta \in [0,1]$ increases. This shows that $\lVert W_n \rVert_{\psi_2}$ increases as $\beta \in [0,1]$ increases. Together with Equation (ref), we have under $\beta_n = 1 - \frac{c}{\sqrt{n}}$, $0 \leq c \leq \sqrt{n}$,

align*[align* omitted — 133 chars of source]

where $o(\cdot)$ is by an absolute constant.

Then we consider $\beta=1+cn^{-\frac{1}{2}}$. We shall note that under this situation it is not hard to check that

align*[align* omitted — 308 chars of source]

Then, under this case we have by Taylor expanding $F$ at $0$ and the fact that $\sup_{v \in \mathbb{R}}|F^{(5)}(v)| < \infty$,

align*[align* omitted — 235 chars of source]

Before we start to upper bound the moments, we first use the fact that $v_+=O(n^{-1/4})$ to obtain that

align*[align* omitted — 163 chars of source]

Then we obtain that

align*[align* omitted — 1,121 chars of source]

with $\mca C_3:=\frac{(3c)^{-1/2}}{3\int_{(-v_+,+\infty)}\exp\left(-cv^2-\frac{\sqrt{3c}}{3}v^3-\frac{1}{12}v^4\right)dv}$, $\mca C_4=\frac{1}{9\int_{(-v_+,+\infty)}\exp\left(-cv^2-\frac{\sqrt{3c}}{3}v^3-\frac{1}{12}v^4\right)dv}$,\\ and $\mca C_5=\frac{2^{-3/2}}{\int_{(-v_+,+\infty)}\exp\left(-cv^2-\frac{\sqrt{3c}}{3}v^3-\frac{1}{12}v^4\right)dv}$. Therefore, we can simply use the definition of the m.g.f. to obtain that

align*[align* omitted — 1,025 chars of source]

Then we use the fact that $\bb E[\mathsf{V}_n|\mathsf{V}_n>0]=v_+$ to obtain that (here we use proposition 2.5.2 in vershynin2018high)

align*[align* omitted — 112 chars of source]

Similarly one obtains that $ \bb E[\exp(t(\mathsf{V}_n-v_-))|\mathsf{V}_n<0]\leq\exp(18e^2n^{-1/2}\sigma^2t^2)$. And hence

align*[align* omitted — 167 chars of source]

Proof for Lemma (ref) High Temperature

We will leverage the representation of $\mathbf{W}$ as a mixture of independent Bernouli random variables after conditioning on some latent variable $\mathsf{U}_n$. We take $\mathsf{U}_n$ to be a random variable with density

align[align omitted — 284 chars of source]

Using Gaussian integral identity $\exp(v^2/2) = \frac{1}{\sqrt{2 \pi}} \int_{-\infty}^{\infty} \exp \left( - u^2/2 + uv\right) du$,

align[align omitted — 327 chars of source]

Hence condition on $\mathsf{U}_n$, $\mathbf{W}_i$ are i.i.d Bernouli with $\mathbb{P}(W_i = 1|\mathsf{U}_n) = \frac{1}{2} (\tanh (\sqrt{\frac{\beta}{n}} \mathsf{U}_n + h) + 1)$, and

align[align omitted — 253 chars of source]
align*[align* omitted — 563 chars of source]

Moreover,

align*[align* omitted — 217 chars of source]

\paragraph*{Step 1: Conditional Berry-Esseen} Apply Berry-Esseen Theorem conditional on $\mathsf{U}_n$,

align*[align* omitted — 418 chars of source]

Take $Z \sim N(0,1)$ independent to $\mathbf{W}$ and $X_i$'s. $\mathsf{U}_n$ is sub-Gaussian by Equation (ref), hence

align*[align* omitted — 468 chars of source]

\paragraph*{Step 2: Stabilization of Variance} By independence between $\mathsf{U}_n$ and $Z$, we have

align*[align* omitted — 656 chars of source]

where $v^{\ast}(\mathsf{U}_n)$ is some quantity between $\mathbb{E}[v(\mathsf{U}_n)]$ and $v(\mathsf{U}_n)$, and by Equation (ref), $v^{\ast}(\mathsf{U}_n) \geq C_2 \mathbb{V}[X_i]$. It follows from boundedness of $v(\mathsf{U}_n)$ and Lipshitzness of $\tanh$ in the expression of $v(\mathsf{U}_n)$ that

align*[align* omitted — 475 chars of source]

\paragraph*{Step 3: Reduction Through TV-distance Inequality}

align*[align* omitted — 551 chars of source]

where $b_n = \sqrt{n}v_0$. The first inequality is by relation between KS- and TV-distances. For the second inequality, denote $X = \sqrt{n}e(\mathsf{U}_n)$, $Y = \sqrt{n}e(b_n + \mathsf{U})$. Denote by $f_X, f_Y, f_Z$ the Lebesgue density of $X, Y, Z$ respectively. Then using $Z \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} X$ and $Z \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} Y$, by data processing inequality,

align*[align* omitted — 101 chars of source]

Above proves inequality (2). Inequality (3) is by scale-invariance of TV distance and data processing inequality.

\paragraph*{Step 4: Gaussian Approximation for $\mathsf{U}_n$} Consider $\mathsf{V}_n = n^{-1/2} \mathsf{U}_n$. Then

align*[align* omitted — 168 chars of source]

where $\phi(v) = - \frac{1}{2} v^2 + \log \cosh (\sqrt{\beta} v + h)$. $\phi$ is maximized at $v_0$ that solves

align[align omitted — 92 chars of source]

We will approximate the integral of $f_{\mathsf{V}_n}$ by Laplace method. We will introduce constants $c_0, c_1$ and $c_2$ that only depends on $\beta$ and $h$. By Equation (5.1.21) in bleistein1975asymptotic,

align*[align* omitted — 334 chars of source]

where the $O(n^{-1})$ term only depends on $n$ and $\phi$. It follows that

align*[align* omitted — 159 chars of source]

Then by a change of variable and the fact that $O(n^{-1})$ term does not depend on $v$,

align[align omitted — 188 chars of source]

Taylor expanding $\phi$ at $v_0 = n^{-1/2}u_0$ and using $\phi^{\prime}(v_0) = 0$, we get

align[align omitted — 422 chars of source]

where $v_{\ast}$ is some quantity between $v_0$ and $n^{-1/2}u$. Now take $b_n = u_0 = \sqrt{n}v_0$ and take $\mathsf{U} \sim N(0, (1 - \beta + v_0^2)^{-1})$, we have

align*[align* omitted — 502 chars of source]

where $v^{\ast}(u)$ is some random quantity between $v_0 = n^{-1/2}u_0$ and $n^{-1/2} u$. We will show that we can restrict the analysis to the region $[u_0 - c_0\sqrt{\log n}, u_0 + c_0 \sqrt{\log n}]$, which is where the bulk of mass lies. Since $\mathsf{U} \sim N(u_0, (1 - \beta + v_0^2)^{-1})$, for some constant $c$ only depending on $\beta$ and $h$, $\mathbb{P} \left(\left|b_n + \mathsf{U} - u_0 \right| \geq c \sqrt{\log n} \right) \leq n^{-1}$. Using a change of variable and concavity of $\phi$,

align*[align* omitted — 837 chars of source]

In the third line we used the fact that $\phi(v_0 + t) - \phi(v_0) = \int_{0}^t \phi^{\prime}(v_0 + s) d s$ and the first derivative is bounded by

align[align omitted — 354 chars of source]

where $w_0$ is the solution to $\tanh(\sqrt{\beta} v_0 + h) - \tanh(w_0) = (\sqrt{\beta}v_0 + h - w_0) \operatorname{sech}^2(w_0)$. It follows that $\phi(v) - \phi(v_0) \leq - \frac{1}{2} (1 - \operatorname{sech}(w_0)^2)(v - v_0)^2$. Using boundedness of $\tanh$ and $\operatorname{sech}$ and the Lipschitzness of $\exp$ when restricted to $[-1,1]$, we have

align*[align* omitted — 737 chars of source]

\paragraph*{Step 5: Gaussian Approximation for $\sqrt{n}e(b_n + \mathsf{U})$} In this step, we will show that $\sqrt{n}e(b_n + \mathsf{U})$ can be well-approximated by $\sqrt{\beta} \operatorname{sech}^2(\sqrt{\beta} v_0 + h)\mathsf{U}$ and hence $\frac{1}{\sqrt{n}} \sum_{i = 1}^n X_i(W_i - \pi)$ can be well-approximated by a Gaussian.

align*[align* omitted — 720 chars of source]

Since $d_{\operatorname{KS}}(\mathsf{U}_n, \mathsf{U}) = O(n^{-1/2})$ and $\pi = \mathbb{E} \left[\tanh \left( \sqrt{\frac{\beta}{n}}\mathsf{U}_n + h \right)\right]$, Taylor expanding $\tanh$ at $\sqrt{\beta}v_0 + h$,

align*[align* omitted — 559 chars of source]

It follows that $\mathbb{E} \left[\left|\sqrt{n}e(b_n + \mathsf{U}) - \sqrt{\beta} (1 - \frac{v_0^2}{\beta})\mathsf{U} \right|\right] = O(n^{-1/2})$ and hence

align*[align* omitted — 235 chars of source]

Recall $\mathsf{U} \sim N(0, (1 - \beta + v_0^2)^{-1})$, hence $\mathbb{E}[X_i] \sqrt{\beta}(1 - \frac{v_0^2}{\beta})\mathsf{U} \sim N(0, \mathbb{E}[X_i]^2\frac{(\beta - v_0)^2}{\beta (1 - \beta + v_0^2)})$. Moreover,

align*[align* omitted — 340 chars of source]

where the last line is because $\mathbb{E}[W_i|\mathsf{U}_n] = \tanh(\sqrt{\beta/n}\mathsf{U}_n)$ and $\mathsf{U}_n$ is sub-Gaussian. Since $Z \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} \mathsf{U}$,

align*[align* omitted — 267 chars of source]

Combining the previous five steps, we get

align*[align* omitted — 234 chars of source]

Proof for Lemma (ref) Critical Temperature

Throughout the proof, we denote by $\mathtt{C}$ an absolute constant, and $\mathtt{K}$ a constant that only depends on the distribution of $X_i$. The proofs for the critical temperature case will have a similar structure as the proof for the high temperature case, based the same $\mathsf{U}_n$ defined in Equation (ref).

Step 1: Conditional Berry-Esseen.

The same argument as in the high-temperature case gives

align*[align* omitted — 143 chars of source]

Step 2: Approximation for $\mathsf{U}_n$.

Take $\mathsf{W}$ to be a random variable with density function

align*[align* omitted — 141 chars of source]

independent to $\mathsf{Z}$. Take $\mathsf{W}_n = n^{-1/4}\mathsf{U}_n$ and $\mathsf{V}_n = n^{-1/2}\mathsf{U}_n$. Again $f_{\mathsf{V}_n}(v) \propto \exp(-n \phi(v))$, where $\phi(v) := - \frac{1}{2} v^2 + \log \cosh (v)$. In particular, $\phi^{(v)}(0) = 0$ for all $0 \leq v \leq 3$, and $\phi^{(4)}(0) = -2 < 0$, $\phi^{(5)}(0) = 0$, $\phi^{(6)}(0) = 16 > 0$. Example 5.2.1 in bleistein1975asymptotic leads to $$f_{\mathsf{V}_n}(v) = n^{\frac{1}{4}}\frac{\sqrt{2}}{3^{\frac{1}{4}}\Gamma(\frac{1}{4})}\exp(n \phi(v) - n \phi(0))(1 + o(1)),$$ which implies $f_{\mathsf{W}_n}(w) = f_{W}(w)(1 + o(1))$. Results in bleistein1975asymptotic do not give a rate, however. We will use a more cumbersome approach to obtain a slightly sub-optimal rate.

By a change of variable, $f_{\mathsf{W}_n}(w) = \frac{h_n(w)}{\int_{-\infty}^{\infty}h_n(u)du}$, where $h_n$ can be written as

align*[align* omitted — 187 chars of source]

The last equality follows from Taylor expanding the term in $\exp(\cdot)$ at $w = 0$, and $g$ is some bounded function.

align*[align* omitted — 209 chars of source]

Moreover, $\int_{[-10\sqrt{\log n}, 10\sqrt{\log n}]^c} h_n(w) d w = O(n^{-1/2}) = I_n [1 + O(n^{-\frac{1}{2}})]$. Hence for denominator, we have $\int_{-\infty}^{\infty} h_n(w)dw =I_n[1 + O((\log n)^3 n^{-\frac{1}{2}})]$. It follows that

align*[align* omitted — 430 chars of source]

Step 3: Data Processing Inequality.

We can use data processing inequality to get

align*[align* omitted — 262 chars of source]

\paragraph*{Step 4: Non-Gaussian Approximation for $n^{\frac{1}{4}}e(n^{\frac{1}{4}}\mathsf{W})$}

align*[align* omitted — 222 chars of source]

where we have use the fact that $\tanh^{(2)}(0) = 0$. Hence there exists $C >0$ such that for $n$ large enough, for any $t > 0$,

align*[align* omitted — 310 chars of source]

We have showed that there exists $c > 0$ such that

align*[align* omitted — 86 chars of source]

in which case $\mathsf{W}^2/\sqrt{n} \leq 1$ for large enough $n$. Hence for large enough $n$ if $t/\mathbb{E}[X_i] > c \sqrt{\log n} + 1$, then

align*[align* omitted — 270 chars of source]

If $0 < t/\mathbb{E}[X_i] < c\sqrt{\log n} + 1$, then

align*[align* omitted — 456 chars of source]

Now we study $g(x; \alpha) = (1 - \sqrt{1 - 4 x \alpha})/(2x), x >0$. Then $\sup_{\alpha \leq \frac{1}{4}} \sup_{0 \leq x \leq \frac{1}{2}}|\theta^{\prime}(x; \alpha)| \leq 2$ and $g(0;\alpha) = \alpha$. Since for large enough $n$, $0 < t/\mathbb{E}[X_i] < c\sqrt{\log n} + 1 \leq \frac{1}{4}$ and $0 \leq n^{-1/2} \leq \frac{1}{2}$, we have $\frac{1 - \sqrt{1 - 4 n^{-1/2} t/\mathbb{E}[X_i]}}{2 n^{-1/2}} \leq t/\mathbb{E}[X_i] + 2 n^{-1/2}$. Hence if $0 < t/\mathbb{E}[X_i] < c\sqrt{\log n} + 1$,

align*[align* omitted — 293 chars of source]

Combining (1), (2), (3),

align*[align* omitted — 225 chars of source]

By similar argument, we can show

align*[align* omitted — 226 chars of source]

Noticing that $W$ and $-W$ have the same distribution, the above two inequalities also hold for $t \leq 0$. Hence it follows from (0) that

align*[align* omitted — 125 chars of source]

\paragraph*{Step 5: Vanishing Variance Term.}

Denote by $f_{\mathsf{W} + n^{-1/4}\mathsf{Z}}$ the density of $\mathsf{W} + n^{-1/4}\mathsf{Z}$. Then

align*[align* omitted — 224 chars of source]

We will use Laplace method to show $f_{\mathsf{W} + n^{-1/4}\mathsf{\mathsf{Z}}}$ is close to $f_{\mathsf{W}}$. However, to get uniformity over $y$, we need to work harder than in the high temperature case. Define $\varphi(x) = x^2/2$ and $g_y(t) =\exp(-(t-y)^4/12)$. Consider

align*[align* omitted — 172 chars of source]

Following Section 5.1 in bleistein1975asymptotic, take $\tau > 0$ such that $\varphi(t) = \tau $, by a change of variable,

align*[align* omitted — 288 chars of source]

To get rate of convergence uniformly in $y$, we follow the proof of Watson's Lemma but consider only up to first order term. Taylor expanding $x \mapsto \exp(-x ^4)/12$ up to first order at $y$, we have

align*[align* omitted — 185 chars of source]

where $\tau^{\ast}$ is some quantity between $0$ and $\sqrt{2\tau}$ and

align*[align* omitted — 102 chars of source]

In particular, we have $\sup_{y \in \mathbb{R}} \sup_{u \in \mathbb{R}}|h_y(u)| < C$ for some absolute constant $C$. Then

align*[align* omitted — 235 chars of source]

Evaluating the first two terms, we get

align*[align* omitted — 281 chars of source]

Similarly, for $I_{y,-}$, change of variable by taking $\tau < 0$ such that $\varphi(t) = \tau$, we have

align*[align* omitted — 280 chars of source]

Combining the two parts, we get

align*[align* omitted — 245 chars of source]

Now take $\lambda = \sqrt{n}$ and multiply both sides by $\frac{n^{1/4}}{3^{1/4}\Gamma(\frac{1}{4})\sqrt{\pi}}$, we get

align*[align* omitted — 241 chars of source]

By a truncation argument, we have

align*[align* omitted — 434 chars of source]

Together with the fact that

align*[align* omitted — 214 chars of source]

we know

align*[align* omitted — 156 chars of source]

Putting together all previous steps, we have

align*[align* omitted — 110 chars of source]

Proof for Lemma (ref) Low Temperature

Throughout the proof, we denote by $\mathtt{C}$ an absolute constant, and $\mathtt{K}$ a constant that only depends on the distribution of $X_i$. The proofs are based on essentially the same argument as in the high temperature case.

Instead of using sub-Gaussianity of $\mathsf{U}_n$, here we use $\mathsf{U}_n$ is sub-Gaussian condition on $\mathsf{U}_n \in \ca I_\ell$, $\ell \in\{-,+\}$. In particular, the previous step 2 by:

Step 2: Approximation for $\mathsf{U}_n$.

In case $\beta > 1$, $\phi(v) = \frac{1}{2}v^2 - \log(\cosh(\sqrt{\beta}v))$ has two global minimum $v_+$ and $v_-$, which are the two solutions of $v -\sqrt{\beta}\tanh(\sqrt{\beta}v) = 0$. We want to show $\phi^{(2)}(v_+) = \phi^{(2)}(v_-) = 1 - \beta + v_+^2 > 0$. It sufffices to show $v_+ > \sqrt{\beta - 1}$. Since $\phi^{\prime}(v) < 0$ for $v \in (0,v_+)$ and $\phi^{\prime}(v) > 0$ for $v \in (v_+,\infty)$, it suffices to show $\phi^{\prime}(\sqrt{\beta - 1}) < 0$. But

align*[align* omitted — 162 chars of source]

Hence $\phi^{(2)}(v_+) = \phi^{(2)}(v_-) > 0$. Observe that on $\ca I_- = (-\infty,0)$ and $\ca I_+ = (0,\infty)$ respectively, the absolute minimum of $\phi$ occurs at $v_-$ and $v_+$, and $\phi^{\prime}$ is non-zero on $\ca I_-$ and $\ca I_+$ except at $v_-$ and $v_+$. Hence we can apply Laplace method (Equation 5.1.21 in bleistein1975asymptotic) sperarately on $\ca I_-$ and $\ca I_+$ to get

align*[align* omitted — 251 chars of source]

It follows from the definition of $f_{\mathsf{V}_n}$ and a change of variable that the density of $\mathsf{U}_n = \sqrt{n} \mathsf{V}_n$ can be approximated by

align*[align* omitted — 181 chars of source]

where $u_l = \sqrt{n} v_l, l \in \{ +,- \}$. Since $\mathbb{P}(\mathsf{U}_n \in \ca I_+) = \mathbb{P}(\mathsf{U}_n \in \ca I_-) = \frac{1}{2}$, condition on $\mathsf{U}_n \in \ca I_+$,

align*[align* omitted — 160 chars of source]

It then follows from Equation (ref) that if we define $\mathsf{U}_+$ to be a random variable with density

align*[align* omitted — 122 chars of source]

then by Taylor expanding $\phi$ at $v_+ = n^{-1/2}u_+$ and a similar argument as in the proof for high temperature case,

align*[align* omitted — 107 chars of source]

The rest follows from the same argument as in the proof for high temperature case and is sub-Gaussianity of $\mathsf{U}_n$ condition on $\mathsf{U}_n \in \ca I_\ell$, $\ell \in\{-,+\}$.

Proof for Remark (ref)

Take $\mathsf{V}_n = n^{-1/2}\mathsf{U}_n$ and $\mathsf{Z}$ be a $\mathsf{N}(0,1)$ variable independent to $\mca m$ and $\mathsf{U}_n$. Based on the conditional mean and variance formulas in Equation (ref), using the conditional on $\mathsf{U}_n$ Berry-Esseen bound,

align*[align* omitted — 248 chars of source]

Using the fact that $v(\mathsf{U}_n)$ is bounded above and $\mathsf{Z}$ is Gaussian, and Taylor expanding $\tanh$, we get

align*[align* omitted — 329 chars of source]

The proof of Lemma (ref) (low temperature) shows that $ d_{\operatorname{TV}}(\mathsf{U}_n|\mathsf{U}_n \in \ca I_+, \mathsf{U}_+) = O(n^{-1/2})$ where $\mathsf{U}_+ \sim \mathsf{N}(\sqrt{n}\pi_+, (1 - \beta(1 - \pi_+^2))^{-1})$. Hence $\mathbb{P} (\mathsf{U}_n \leq \sqrt{\log n} | \mathsf{U}_n \geq 0) \lesssim \exp(-n)$. It follows that

align*[align* omitted — 84 chars of source]

By symmetry and the fact that $\mathbb{P}(\operatorname{sgn}(\mca m) = \ell) = \mathbb{P}(\operatorname{sgn}(\mathsf{U}_n) = \ell) = 1/2$ for $\ell = -, +$, we know

align*[align* omitted — 152 chars of source]

The conclusion then follows from Lemma (ref)(3).

Proof for Lemma (ref) Drifting from High Temperature

Throughout the proof, we denote by $\mathtt{C}$ an absolute constant, and $\mathtt{K}$ a constant that only depends on the distribution of $X_i$.

Let $\mathsf{U}_n(c)$, $e(\mathsf{U}_n(c))$, $v(\mathsf{U}_n(c))$ be the latent variable, conditional mean, and conditional variance as previously defined when $\beta_n = 1 + c n^{-\frac{1}{2}}$, $c < 0$. For notational simplicity, we abbreviate the $c$, and call them $\mathsf{U}_n, e(\mathsf{U}_n), v(\mathsf{U}_n)$ respectively. By Lemma (ref), $\lVert \mathsf{U}_n \rVert_{\psi_2} \leq \mathtt{C} n^{1/4}$.

Step 1: Conditional Berry-Esseen.

Apply Berry-Esseen Theorem conditional on $\mathsf{U}_n$ in the same way as in the high temperature case, we get

align*[align* omitted — 145 chars of source]

Step 2: Non-Normal Approximation for $n^{-\frac{1}{4}}\mathsf{U}_n$.

Consider $\mathsf{W}_n = n^{-1/4}\mathsf{U}_n$. Then $f_{\mathsf{W}_n}(w) = I_n(c)^{-1} h_n(w)$, with $I_n(c) = \int_{-\infty}^{\infty}h_n(w)dw$, and

align*[align* omitted — 236 chars of source]

where by smoothness of $\log(\cosh(\cdot))$, $\lVert \theta \rVert_{\infty} \leq \mathtt{K}$. Then

align[align omitted — 342 chars of source]

Moreover, by a change of variable and the fact that $\beta_n \leq 1$,

align*[align* omitted — 328 chars of source]

Since $\lVert \mathsf{W}_n(c) \rVert_{\psi_2} \leq \mathtt{C}$, $I_n(c)^{-1}\int_{(-\mathtt{C}\sqrt{\log n},\mathtt{C} \sqrt{\log n})^c}h_n(w) dw \leq \mathtt{C} n^{-1/2}$. It follows that

align[align omitted — 142 chars of source]

Combining Equation (ref) and (ref), we have $I_n(c) =I(c)[1 + O(\mathtt{C}^6(\log n)^3 n^{-1/2})]$. It follows that

align*[align* omitted — 855 chars of source]

Step 3: A Reduction through TV-distance Inequality.

Since $\mathsf{Z} \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} (\mathsf{U}_n, \mathsf{W}_n)$, we can use data processing inequality to get

align*[align* omitted — 366 chars of source]

Step 4: Non-Gaussian Approximation for $n^{\frac{1}{4}}e(n^{\frac{1}{4}}\mathsf{W})$.

This is essentially the same as the proof for step 4 from the critical temperature case in Lemma (ref).

align*[align* omitted — 146 chars of source]

Step 5: Stabilization of Variance.

Using the same argument as Step 4 in the high temperature case for Lemma (ref), and $\lVert \mathsf{W} \rVert \leq \mathtt{K}$,

align*[align* omitted — 285 chars of source]

The conclusion then follows from putting together the previous five steps.

Proof for Lemma (ref) Drifting from Low Temperature

Consider the same $\mathsf{U}_n$ defined in Equation (ref). Recall $\phi(v) = \frac{v^2}{2} - \log \cosh(\sqrt{\beta_n}v)$, $\phi^{\prime}(v) = v - \sqrt{\beta_n} \tanh(\sqrt{\beta_n} v)$, $\phi^{(2)}(v) = 1 - \beta_n \operatorname{sech}^2(\sqrt{\beta_n}v)$. And we take $v_{n,+} > 0$, $v_{n,-} < 0$ to be the two solutions of $v - \sqrt{\beta_n}\tanh(\sqrt{\beta_n}v) = 0$.

Step 2': Non-Normal Approximation for $n^{-\frac{1}{4}}\mathsf{U}_n$.

Take $\mathsf{V}_n = n^{-1/2}\mathsf{U}_n$. Then $f_{\mathsf{V}_n}(v) \propto \exp(-n \phi(v))$. Taylor expanding $\phi^{\prime}$ at $0$, we know there exists some function $g$ that is uniformly bounded such that $\phi^{\prime}(v) = (1 - \beta_n) v + \frac{1}{3}\beta_n^2 v^3 + \beta_n^3 g(v) v^5$. Hence $$v_{n,+} = \sqrt{\frac{3(\beta_n - 1)}{\beta_n^2}} + O(\beta_n - 1) = \sqrt{3 c}n^{-1/4} + O(n^{-1/2}).$$ Taylor expand $\tanh$ and $\operatorname{sech}$ at $0$,

align*[align* omitted — 770 chars of source]

Take $\mathsf{W}_n = n^{1/4}\mathsf{V}_n = n^{-1/4}\mathsf{U}_n$, $\mca w_+ = n^{1/4}v_{n,+} = \sqrt{3 c} + O(n^{-1/4})$, and $\mca w_- = n^{1/4} v_{n,-}$. Define

align*[align* omitted — 270 chars of source]

By a change of variable and Taylor expansion, the density for $\mathsf{W}_n$ satisfies

align[align omitted — 213 chars of source]

By Lemma (ref), for $\ell \in \{-,+\}$, condition on $\mathsf{W}_n \in \ca I_{c,n,\ell}$, $\mathsf{W}_n - \mca w_\ell$ is sub-Gaussian with $\psi_2$-norm bounded by $\mathtt{C}$. Let $\mathsf{W}_{c,n}$ be a random variable with density at $w$ proportional to $\exp(h_{c,n}(w))$. By similar argument as Equations (ref) and (ref), $$d_{\operatorname{KS}}(\mathsf{W}_n| \mathsf{W}_n \in \ca I_{c,n,\ell}, \mathsf{W}_{c,n}| \mathsf{W}_{c,n} \in \ca I_{c,n,\ell}) \leq \mathtt{C} (\log n)^3 n^{-1/2}).$$

The other steps, conditional Berry-Esseen, reduction through TV-distance inequality, and non-Gaussian approximation for $n^{\frac{1}{4}}e(n^{\frac{1}{4}}\mathsf{W}_{c,n})$ can be proceeded in the same way as in the proof for Lemma (ref), with $\mathsf{W}_n - \mca w_\ell$ sub-Gaussian condition on $\mathsf{W}_n \in \ca I_{c,n,\ell}$ with $\psi_2$-norm bounded by $\mathtt{C}$, and respectively for $\mathsf{W}_{c,n}$.

Proof for Lemma (ref) Knife-Edge Representation

Again we take $\mathsf{U}_n$ to be the latent variable from Lemma (ref), and $\mathsf{W}_n = n^{-1/4}\mathsf{U}_n$. From Step 2 in the proof of Lemma (ref), $f_{\mathsf{W}_n}(w) = I_n(c)^{-1} h_n(w)$, with $I_n(c) = \int_{-\infty}^{\infty}h_n(w)dw$, and

align*[align* omitted — 204 chars of source]

where by smoothness of $\log(\cosh(\cdot))$, $\lVert \theta \rVert_{\infty} \leq \mathtt{K}$.

Case 1: When $\sqrt{n}(\beta_n - 1) = o(1)$. We can apply Berry-Esseen conditional on $\mathsf{U}_n$ the same way as in the proof of Lemma (ref), and its Step 2 can also be applied here to show that if we take $\widetilde{\mathsf{W}}_c$ to be a random variable with density proportional to $\exp(-c_n^2/2 w^2 - \beta_n^2/12 w^4)$, then $d_{\operatorname{KS}}(\mathsf{W}_n,\widetilde{\mathsf{W}}_c) = O((\log n)^3 n^{-1/2})$. Moreover, $c_n = o(1)$ and $\beta_n = 1 - o(1)$. Hence $d_{\operatorname{KS}}(\mathsf{W}_n,\mathsf{W}_0) = o(1)$. The rest of the proof then follows from Step 3 to Step 5 in the proof for the critical regime case in Lemma (ref).

Case 2: When $\sqrt{n}(1 - \beta_n) \gg 1$. Again we still have $\lVert \mathsf{U}_n \rVert_{\psi_2} = O(n^{1/4})$. And we take $v_+ > 0$, $v_- < 0$ to be the two solutions of $v - \sqrt{\beta_n}\tanh(\sqrt{\beta_n}v) = 0$. Similarly as in the previous case, the first two steps in the proof of Lemma (ref) implies $d_{\operatorname{KS}}(\mathsf{W}_n,\widetilde{\mathsf{W}}_c) = o(1)$, where the density of $\mathsf{W}_c$ is proportional to $\exp(-c_n^2/2w^2 - \beta_n^2/12 w^4)$. Since $c_n \gg 1$, the first term in the exponent dominates, and we can show $d_{\operatorname{KS}}(\mathsf{W}_n, \mathsf{W}_c^{\dag}) = o(1)$, where $\mathsf{W}_c^\dag$ has density proportional to $\exp(-c_n^2/2 w^2)$. Again, we can Taylor expand to get $n^{1/4}e(n^{1/4}\mathsf{W})) =\mathbb{E}[X_i] n^{\frac{1}{4}}\tanh\left(n^{-\frac{1}{4}}\mathsf{W}\right) = \mathbb{E}[X_i] [\mathsf{W} - O (\frac{\mathsf{W}^2}{3\sqrt{n}})]$, and show $d_{\operatorname{KS}}(n^{1/4}e(n^{1/4}\mathsf{W}_c^\dag), \mathbb{E}[X_i]\mathsf{W}_c^\dag) = o(1)$. Combining with stablization of variance as in the proof of Lemma (ref) (high temperature case), we can show $$d_{\operatorname{KS}}(\mca g_n, n^{-1/4}\mathbb{E}[X_i^2]^{1/2}\mathsf{Z} + \mathbb{E}[X_i]\mathsf{W}_c^\dag) = o(1).$$ Since $\mathsf{Z}$ and $\mathsf{W}_c^\dag$ are independent Gaussian random variables, we also have $d_{\operatorname{KS}} (\mca g_n/\sqrt{\mathbb{V}[\mca g_n]}, \mathsf{Z}) = o(1)$.

Case 3: When $\sqrt{n}(\beta_n - 1) \gg 1$. By Lemma (ref) (2),

align[align omitted — 347 chars of source]

where $\mathsf{W}_{c,n}$ has density proportional to $\exp(h_{c,n}(w))$, with

align*[align* omitted — 258 chars of source]

and $\ca I_{c,n,-} = (-\infty,K_{c,n,-})$ and $\ca I_{c,n,+} = (K_{c,n,+},\infty)$ such that $\mathbb{E}[\mathsf{W}_{c,n}|\mathsf{W}_{c,n} \in \ca I_{c,n,\ell}] = w_{c,n,\ell}$ for $\ell \in \{-,+\}$. Now we calculate the order of the coefficients under $\sqrt{n}(\beta_n - 1) \gg 1$. First, suppose $\beta_n = 1 + c n^{\gamma}$ for some $\gamma \in (0,\infty)$ and $c$ not depending on $n$. Then $v_+ = \sqrt{\frac{3(\beta_n - 1)}{\beta_n^2}} + O(\beta_n - 1) = \sqrt{3 c}n^{-\gamma/2} + O(n^{-\gamma})$. Taylor expand $\tanh$ and $\operatorname{sech}$ at $0$,

align*[align* omitted — 630 chars of source]

We see when $\gamma = 1/2$, all of $\sqrt{n}\phi^{(2)}(v_+)$, $n^{1/4}\phi^{(3)}(v_+)$ and $\phi^{(4)}(v_+)$ are of order 1. And when $c_n = \sqrt{n}(\beta_n - 1) \gg 1$, we have $\sqrt{n}\phi^{(2)}(v_+) \gg n^{1/4}\phi^{(3)}(v_+) \gg \phi^{(4)}(v_+)$. Since $w_+ = n^{1/4} v_+ = \sqrt{3 c_n} \gg 1$, and similarly, $|w_-| \gg 1$, condition on $\mathsf{W}_{c,n} \in [n]$, $\mathsf{W}_{c,n} - \mathbb{E}[\mathsf{W}_{c,n}|\mathsf{W}_{c,n} \in [n]]$ is $\mathtt{C}$-sub-Gaussian, $\ell \in \{-,+\}$. By similar concentration arguments as in the proof for Step 2 in Lemma (ref) (1), we can show the second order term in $h_{c,n}$ dominates, and for $\ell \in \{-,+\}$,

align*[align* omitted — 238 chars of source]

The conclusion then follows from pluggin the (conditional) Gaussian approximation for $\mathsf{W}_{c_n,n}$ back into Equation (ref), and the fact that $\mathsf{Z}$ is independent to $\mathsf{W}_{c,n}$ and also Gaussian.

Proof of Lemma (ref)

Throughout the proof, the Ising spins $\mathbf{W}=(W_i)_{i=1}^n$ are distributed according to Assumption (ref) with parameters $(\beta,h)$. For brevity, we write $\mathbb{P}$ in place of $\mathbb{P}_{\beta,h}$.

center[center omitted — 68 chars of source]

Let $\mathsf{U}_n$ be the latent random variable from Lemma (ref). Condition on $\mathsf{U}_n$, $\mathbf{X}_i W_i$'s are i.i.d random vectors. For $u \in \mathbb{R}$, define

align*[align* omitted — 598 chars of source]

and to save notations, we denote

align*[align* omitted — 69 chars of source]

Suppose $\mathsf{Z}_d \sim \mathsf{N}(\mathbf{0}, \mathbf{I}_{d \times d})$ independent to $\mathsf{U}_n$. By chernozhukov2017central

equation[equation omitted — 401 chars of source]

From the proofs of Lemma (ref), we know the term $t(\mathsf{U}_n)$ stabilizes, $$d_{\operatorname{KS}} \Big(t(\mathsf{U}_n), \sigma \mathsf{Z}\Big) = O(n^{-1/2}), \qquad \sigma = \Big(\frac{\beta(1 - \pi^2)^2}{1 - \beta(1 - \pi^2)}\Big)^{1/2}.$$ By Lemma (ref),

align*[align* omitted — 159 chars of source]

For each $\varepsilon$, define $A_{\varepsilon}$ to be the event $\{\lVert \Sigma(\mathsf{U}_n)^{1/2} - \Sigma^{1/2})\mathsf{Z}_d \rVert \leq \varepsilon\}$. Since $d$ is fixed, we can work with each dimension to get

equation[equation omitted — 1,398 chars of source]

where in the last line, we have chosen $\varepsilon = n^{-1/2}\sqrt{\log n}$ and used Nazarov's inequality (Lemma A.1 in chernozhukov2017central). Since $\mathsf{Z}_d$ and $\mathsf{U}_n$ are independent, we can show via data processing inequality that

align*[align* omitted — 450 chars of source]

Combining the previous results,

align*[align* omitted — 244 chars of source]
center[center omitted — 47 chars of source]

We still have conditional Berry-Esseen as in Equation (ref). The proof of Lemma (ref) implies $$d_{\operatorname{KS}}(n^{-1/4}t(\mathsf{U}_n),\mathsf{R}) = O(n^{ -1/2}).$$ Hence $\lVert \Sigma(\mathsf{U}_n) - \mathbb{E}[\Sigma(\mathsf{U}_n)] \rVert_{\operatorname{max}} = O_{\psi,2}(n^{-1/4})$. By concentration of $\mathsf{U}_n$, approximation of $n^{-1/4}t(\mathsf{U}_n)$ by $\mathsf{R}$, and anti-concentration of $\mathsf{R}$, we can use similar arguments as Equation (ref) to get

align*[align* omitted — 299 chars of source]

By independence between $\mathsf{Z}_d$ and $\mathsf{U}_n$, and approximation of $n^{-1/4}t(\mathsf{U}_n)$ by $\mathsf{R}$, we can use data processing inequality to get

align*[align* omitted — 267 chars of source]
figure[figure omitted — 847 chars of source]

It follows that

align*[align* omitted — 241 chars of source]
center[center omitted — 43 chars of source]

We still have conditional Berry-Esseen as in Equation (ref). From the proof of Lemma (ref) (3) and Remark (ref), for $\ell = -,+$, $$d_{\operatorname{KS}}(t(\mathsf{U}_n) - \sqrt{n} \pi_{\ell}|\operatorname{sgn}(\mca m) = \ell, \sigma \mathsf{Z}) = O(n^{-1/2}),$$ where $\sigma^2 = \frac{\beta(1 - \pi_+^2)^2}{1 - \beta(1 - \pi_+^2)}$. The rest of the proof follows from the arguments for I. High Temperature or Nonzero External Field, using conditional concentration of $\mathsf{U}_n$ given $\operatorname{sgn}(\mca m)$.

Proofs: Section (ref)

Proof of Lemma (ref)

Our proof is constructive. We show that consistent estimate of $n\mathbb{V}[\widehat{\tau}_n]$ would imply that one can distinguish between two constructed hypotheses easily. Let $\mathcal{P}_n$ be the class of distributions of random vectors $(\mathbf{W} = (W_1, \cdots, W_n), \mathbf{Y} = (Y_1, \cdots, Y_n))$ taking values in $\mathbb{R}^{2n}$ that satisfies Assumptions 1,2,3. Consider the following two data generating processes:

align*[align* omitted — 375 chars of source]

where $0 < u < 1$, and in both cases $(\varepsilon_i: 1 \leq i \leq n)$ are i.i.d $\mathsf{N}(0,1)$ random variables, independent to $\mathbf{W}$. Denote by $\mathbb{P}_{0,n}$ and $\mathbb{P}_{1,n}$ the laws of $(\mathbf{W},\mathbf{Y})$ under $\text{DGP}_0$ and $\text{DGP}_1$. Then {

align*[align* omitted — 372 chars of source]

} the first line uses chain rule of $d_{\text{KL}}$, the second line uses $$d_{\text{KL}}(\mathbb{P}_{0,n}(\mathbf{Y}|\mathbf{W}),\mathbb{P}_{1,n}(\mathbf{Y}|\mathbf{W})) = d_{\text{KL}}(\mathbb{P}_{0,n}(\mathbf{Y}),\mathbb{P}_{1,n}(\mathbf{Y})) = 0.$$ From Theorem 2.3 (and its proof) in bhattacharya2018inference, $$M := \lim_{n \rightarrow \infty} d_{\text{KL}}(\mathbb{P}_{0,n}(\mathbf{W}),\mathbb{P}_{1,n}(\mathbf{W})) < \infty.$$ Hence for large enough $n$,

align*[align* omitted — 278 chars of source]

Le Cam's method (Section 15.2.1 in wainwright2019high) gives for large enough $n$,

align*[align* omitted — 459 chars of source]

in the last line we used Theorem 2 (1) to get $n \mathbb{V}_{\mathbb{P}_{n,0}}[\widehat \tau - \tau] - n \mathbb{V}_{\mathbb{P}_{n,1}}[\widehat \tau - \tau] = \varepsilon (1 + o(1))$.

Proof of Lemma (ref)

The following discussions will be organized according to the three different cases: (1) When $\beta<1$. (2) When $\beta\geq 1$, $\mca m$ concentrates around $0$. (3) When $\beta\geq 1$ and $\mca m$ concentrates around two symmetric locations $w_{+}>0$ and $w_{-}<0$ with $|w_{+}|=|w_{-}|$.

We have required $\wh \beta \in [0,1]$. For analysis, consider an unrestricted pseud-likelihood estimator,

align*[align* omitted — 123 chars of source]

where $l(\beta;\mathbf{W})$ is the pseudo log-likelihood given by

align*[align* omitted — 210 chars of source]

We show that $l(\beta;\mathbf{W})$ is concave.

align*[align* omitted — 359 chars of source]

and

align*[align* omitted — 193 chars of source]

Hence $l(\cdot;\mathbf{W})$ is concave everywhere in $\mathbb{R}$. This shows $\wh \beta = \min\{\max\{\widehat{\beta}_\text{UR},0\},1\}$. Now we study limiting distribution of $\widehat{\beta}_\text{UR}$

center[center omitted — 67 chars of source]

To obtain a more precise distribution for $\widehat{\beta}_\text{UR}$, we use Fermat's condition to obtain that

align*[align* omitted — 904 chars of source]

here $O(\cdot)$'s are all up to an absolute constant. By Lemma (ref) with $X_i = 1$, we can show $\mathbb{E}[|(n \mca m)^{-1}|] \leq \mathtt{C} n^{-1/2}$. By Markov inequality, $(n \mca m)^{-1} = O_\mathbb{P}(n^{-1/2})$. Taylor expanding $\tanh$, we have

align[align omitted — 432 chars of source]

where in the above equation, both $O(\cdot)$ and $O_\mathbb{P}(\cdot)$ are up to absolute constants. The rest of the results are given according to the different temperature regimes.

(1) The High Temperature Regime. Using Lemma (ref) with $X_i = 1$, our result for the high temperature regime with $\beta<1$ implies that $n^{\frac{1}{2}}\mca m\overset{d}{\to}\mathsf{N}( 0,\frac{1}{1-\beta}) \Rightarrow (1-\beta)n\mca m^2\overset{d}{\to} \chi^2(1)$. Therefore we conclude that $\frac{1-\beta}{1-\widehat{\beta}_\text{UR}}\overset{d}{\to}\chi^2(1)$. The conclusion then follows from $\wh \beta = \min\{\max\{\widehat{\beta}_\text{UR},0\},1\}$.

(2) The Critical Temperature Regime. Using Lemma (ref) with $X_i = 1$, we have $d_{\operatorname{KS}}(n^{\frac{1}{4}}\mca m, \mathsf{W}_0) = o(1)$. This implies $n^{\frac{1}{2}}(\widehat{\beta}_\text{UR}-1) \overset{d}{\to} Law(\frac{\mathsf{W}_0^2}{3}-\frac{1}{\mathsf{W}_0^2}).$ Since $\mathsf{W}_0 = O_\mathbb{P}(1)$, $\mathbb{P}(\widehat{\beta}_\text{UR} < 0) = o(1)$. The conclusion then follows from $\wh \beta = \min\{\max\{\widehat{\beta}_\text{UR},0\},1\}$.

center[center omitted — 57 chars of source]

When $\mca m$ concentrates around $\pi_+$ and $\pi_-$ we have when $\mca m>0$, use the fact that $\pi_{\ell}=\tanh(\beta\pi_{\ell})$ for $\ell\in\{+,-\}$,

align*[align* omitted — 734 chars of source]

and the similar argument gives

align*[align* omitted — 230 chars of source]

The conclusion then Lemma (ref) (3) and the convergence of $\mca m$ to $\pi_+$ or $\pi_-$.

Proof of Lemma (ref)

Again we consider the unrestricted PMLE given by

align*[align* omitted — 123 chars of source]

where $l(\beta;\mathbf{W})$ is the pseudo log-likelihood given by

align*[align* omitted — 210 chars of source]

For $\beta \in [0,1]$, that is $c_\beta = \sqrt{n}(\beta - 1) \leq 0$, Equation (ref) and the approximation of $\mca m$ by $n^{-1/2}\mathsf{Z} + n^{-1/4}\mathsf{W}_c$ from Lemma (ref) gives

align*[align* omitted — 174 chars of source]

The conclusion follows from the fact that $x \mapsto \max\{\min\{x,0\},1\}$ is $1$-Lipschitz.

Proofs: Section (ref)

Preliminary Lemmas

lemmaRecall $\mathbf{W} = (W_i)_{1 \leq i \leq n}$ takes value in $\{-1,1\}^n$ with \begin{align*} \mathbb{P} \left( \mathbf{W} = \mathbf{w} \right) = \frac{1}{Z} \exp \bigg( \frac{\beta}{n} \sum_{i < j} W_i W_j + h \sum_{i = 1}^n W_i \bigg), \quad h \neq 0 or h = 0, 0 \leq \beta \leq 1. \end{align*} Recall $\pi$ is the unique solution to $x = \tanh(\beta x + h)$. Then $\mathbb{E}[W_i] = \pi + O(n^{-1})$.
proofIf $h = 0$, then $\pi = \mathbb{E}[W_i] = 0$. If $h \neq 0$, then the concentration of $\mca m = n^{-1} \sum_{i = 1}^n W_i$ towards $\pi$ in Lemma (ref) implies, \begin{align*} \mathbb{E}[W_i] = & \mathbb{E}[\mathbb{E}[W_i|W_{-i}]] = \mathbb{E}[\tanh(\beta \mca m_i + h)] \nonumber\\ = & \mathbb{E}[\tanh(\beta \pi + h) + \operatorname{sech}^2(\beta \pi + h)(\mca m_i - \pi) - \operatorname{sech}^2(\beta m^{\ast} + h)\tanh(\beta m^{\ast} + h)(\mca m_i - \pi)^2] \nonumber\\ = & \tanh(\beta \pi + h) + O(n^{-1}) \\ = & \pi + O(n^{-1}), \end{align*} where $\mca m^{\ast}$ is a number between $\mca m$ and $\pi$, and we have used boundedness of $\operatorname{sech}$.
lemmaSuppose Assumption (ref), 2, and 3 hold with $h = 0, 0 \leq \beta \leq 1$ or $h \neq 0$: (1) \begin{align*} & \max_{i \in [n]} \left|\frac{M_i}{N_i} - \pi \right| = O_{\psi_{\beta, \gamma}}(n^{-\mathtt{r}_{\beta, h}}) + O_{\psi_2}(N_i^{-1/2}). \end{align*} (2) Define $A(\mathbf{U}) = (G(U_i,U_j))_{1 \leq i,j \leq n}$. Condition on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A} = \{A \in \bb R^{n \times n}: \min_{i \in [n]} \sum_{j \neq i}A_{ij} \geq 32 \log n\}$, for large enough $n$, for each $i \in [n]$ and $t > 0$, \begin{align*} \mathbb{P} \left( \left|\frac{M_i}{N_i} - \pi\right| \geq 4 \mathbb{E}[N_i | \mathbf{U}]^{-1/2}t^{1/2} + C_{\beta,h}n^{-\mathtt{r}_{\beta, h}} t^{\mathtt{p}_{\beta, h}}\middle| \mathbf{U} \right) \leq 2 \exp(-t) + n^{-98}, \end{align*} where $C_{\beta,h}$ is some constant that only depends on $\beta,h$. (3) When $h = 0$, and $\beta \in [0,1]$, then there exists a constant $\mathtt{K}$ that does not depend on $\beta$, such that for large enough $n$, for each $i \in [n]$ and $t > 0$, \begin{align*} \mathbb{P} \left( \left|\frac{M_i}{N_i} - \pi\right| \geq 4 \mathbb{E}[N_i | \mathbf{U}]^{-1/2}t^{1/2} + \mathtt{K} n^{-\mathtt{r}_{\beta, h}} t \middle| \mathbf{U} \right) \leq 2 \exp(-t) + n^{-98}. \end{align*}
proofTake $\mathsf{U}_n$ to be a random variable with density \begin{align*} f_{\mathsf{U}_n}(u) = \frac{\exp \left(- \frac{1}{2} u^2 + n \log \cosh \left( \sqrt{\frac{\beta}{n}} u + h\right) \right)}{\int_{-\infty}^{\infty} \exp \left(- \frac{1}{2} v^2 + n \log \cosh \left( \sqrt{\frac{\beta}{n}} v + h\right) \right) d v}. \end{align*} Condition on $\mathsf{U}_n$, $W_i$'s are i.i.d. Decompose by \begin{align*} \frac{M_i}{N_i} - \pi = \sum_{j \neq i} \frac{E_{ij}}{N_i} \left(W_j - \mathbb{E}[W_j|\mathsf{U}_n] \right) + \mathbb{E}[W_j|\mathsf{U}_n] - \pi. \end{align*} Condition on $\mathsf{U}_n$, $W_i$'s are i.i.d. Berry-Esseen theorem condition on $\mathsf{U}_n$ and $\mathbf{E}$ gives, \begin{align} \sup_{t \in \mathbb{R}} \bigg|\mathbb{P}\Big(\frac{M_i}{N_i} - \pi \leq t\Big|\mathbf{E} \Big) - \mathbb{P} \Big(\sqrt{\frac{v(\mathsf{U}_n)}{N_i}} Z + e(\mathsf{U}_n) \leq t \Big|\mathbf{E}\Big) \bigg| = O(n^{-\frac{1}{2}}), \end{align} where $e(\mathsf{U}_n) := \mathbb{E}[W_i|\mathsf{U}_n] - \pi = \tanh(\sqrt{\beta/n}\mathsf{U}_n + h) - \pi$, and $v(\mathsf{U}_n) := \mathbb{V}[W_i - \pi|\mathsf{U}_n]$. By McDiarmid's inequality, \begin{align*} \mathbb{P} \bigg( |\sum_{j \neq i} \frac{E_{ij}}{N_i} \left(W_j - \mathbb{E}[W_j|\mathsf{U}_n] \right)| \geq 2 N_i^{-1/2} t\bigg | \mathbf{E} \bigg) \leq 2 \exp(-t^2). \end{align*} Plugging into Equation (ref), we can show (1) holds. Next, we want to show condition on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A}$, $\mathbb{P}(N_i \leq \mathbb{E}[N_i|\mathbf{U}]/3|\mathbf{U}) \leq n^{-100}$: Notice that for any $\mathbf{U}$ such that $\rho_n \min_{i \in [n]} \sum_{j \neq i}A_{ij}(\mathbf{U}) \rightarrow \infty$, Condition on $A$ such that $A \in \mathcal{A}$, $E_{ij} = \rho A_{ij}\iota_{ij}$, $1 \leq i \leq j \leq n$ are i.i.d Bernouli random variables, and for each $i,j$, $\sum_{k \neq i,j}A_{ki} \geq 32 \log n - 1 \geq 31 \log n$ for $n \geq 3$. By bounded difference inequality, for all $t > 0$, \begin{align*} \mathbb{P} \bigg( \bigg|\sum_{k \neq i,j}E_{ki} - \sum_{k \neq i,j}\rho_n A_{ki}\bigg| \geq \rho_n \sqrt{\sum_{k \neq i,j}A_{i,j}^2}t \bigg) \leq 2 \exp(-2t^2). \end{align*} Hence condition on $A$, with probability at least $1 - n^{-100}$, \begin{align} \nonumber \sum_{k \neq i,j}E_{ki} & \geq \sum_{k \neq i,j}\rho_n A_{ki} - 8 \sqrt{\log n} \rho_n \sqrt{\sum_{k \neq i,j}A_{ij}^2} \geq \rho_n \sum_{k\neq i,j} A_{ki} - 8 \sqrt{\log n} \rho_n \sqrt{\sum_{k \neq i,j}A_{ki}} \\ \nonumber & \geq \rho_n \sqrt{\sum_{k \neq i,j}A_{ki}}\left(\sqrt{\sum_{k \neq i,j}A_{ki}} - 8 \sqrt{\log n}\right) \\ \nonumber & \geq \rho_n \sqrt{\sum_{k \neq i,j}A_{ki}} \left(\sqrt{\sum_{k \neq i,j}A_{ki}} - 8 \sqrt{31^{-1} \sum_{k \neq i,j}A_{ij}}\right)\\ & \geq \rho_n \sum_{k \neq i,j}A_{ij}/3 \geq \frac{31}{3}\log n, \end{align} and since $\rho_n A_{i,j} = \mathbb{E}[E_{ij}|\mathbf{U}] \in [0,1]$, $\sum_{k \neq i,j}E_{ki} + 1 \geq \mathbb{E}[N_j|\mathbf{A}]/3$. By Equation (ref), condition on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A}$, $\mathbb{P}(N_i \leq \mathbb{E}[N_i|\mathbf{U}]/3|\mathbf{U}) \leq n^{-100}$. Hence we can disintegrate over the distribution of $\mathbf{E}$ to get \begin{align*} \mathbb{P} \left( |\sum_{j \neq i} \frac{E_{ij}}{N_i} \left(W_j - \mathbb{E}[W_j|\mathsf{U}_n] \right)| \geq 4 \mathbb{E}[N_i|\mathbf{U}]^{-1/2} t \middle | \mathbf{U} \right) \leq 2 \exp(-t^2) + n^{-100}. \end{align*} By Equation (ref) and Lemma (ref), and the Lipschitzness of $\tanh$ that \begin{align*} \mathbb{E}[W_i|\mathsf{U}_n] - \pi = O_{\psi_{\beta,h}}\left(n^{-\mathtt{r}_{\beta, h}}\right). \end{align*} Plugging into Equation (ref), we can show (2) holds. Under the setting of (3), the only part that depends on $\beta$ in our proof is $\mathsf{U}_n$. Since we show in Lemma (ref) $\lVert \mathsf{U}_n \rVert_{\psi_1} \leq \mathtt{K} n^{1/4}$ for some absolute constant $\mathtt{K}$, which is essentially the $\beta = 1$ rate, the conclusion of (3) then follows.

Proof of Lemma (ref)

Since we use the conditional probability $p_i$ in the inverse probability weight, we have

align*[align* omitted — 474 chars of source]

and the conclusion follows from $\mathbb{E}[T_i|\mathbf{T}_{-i},(f_i)_{i \in [n]}, \mathbf{E}] = p_i$.

Proof of Lemma (ref)

First consider the treatment part.

align*[align* omitted — 266 chars of source]

For the second term, taylor expand $p_i^{-1}, p_i$ as follows:

equation[equation omitted — 377 chars of source]

where $\xi_i^{\ast}$ is some random quantity that lies between $4 \frac{\beta}{n} \sum_{j \neq i} W_j$ and $4 \frac{\beta}{n} \sum_{j \neq i} \pi$. Taking the parameters $c_i^+ = g_i \left(1, \pi \right) \left(1 + \exp(-2 \beta \pi - 2 h) \right)$, $d^+ = \beta(1 - \tanh(\beta \pi + h))\mathbb{E}[g_i(1,\pi)].$ Then

align*[align* omitted — 1,422 chars of source]

\paragraph*{Proof of (1):} By Lemma (ref), $\mca m - \pi = O_{\psi_{\beta, h}}(n^{-\mathtt{r}_{\beta, h}})$. The claim follows from Equation (ref) and a union bound argument. \paragraph*{Proof of (2):}

align*[align* omitted — 267 chars of source]

By Lemma (ref),

align*[align* omitted — 84 chars of source]

Taylor expand $\tanh(x)$ at $x = \beta \pi + h$, we have

align*[align* omitted — 406 chars of source]

hence

align*[align* omitted — 165 chars of source]

\paragraph*{Proof of (3):} The first line follows from a Taylor expansion of $p_i = (1 + \exp(2 \beta \mca m_i + 2 h))^{-1}$ at $\pi$, and $\mca m_i - \pi = O_{\psi_{\beta,h}}(n^{-\mathtt{r}_{\beta, h}})$, noticing that $c_i$, $\lVert \psi^{\prime \prime} \rVert_{\infty}$ are bounded. The second line follows by reordering the terms.

\paragraph*{Proof of (4):} By Lemma (ref), $\tanh(\beta \pi + h) = \pi + O(n^{-1})$. By boundedness and i.i.d of $g_i(1,\pi)$, $\frac{1}{n} \sum_{j \neq i} c_j = \overline{c} + O(n^{-1}) = \mathbb{E}[c_i] + O_{\mathbb{P}}(n^{-1/2}) + O(n^{-1})$. Similarly, for the control part, taking the parameters $c_i^- = g_i \left(-1, \pi \right) \left(1 + \exp(2 \beta \pi + 2 h) \right)$, $d^- = \beta(1 - \tanh(-\beta \pi - h))\mathbb{E}[g_i(-1,\pi)].$

align*[align* omitted — 306 chars of source]

Using Lemma (ref) again, we can show $(1 +\exp(- 2 \beta \pi - 2 h))/2 = 1/\pi + O(n^{-1})$ and $(1 + \exp(2 \beta \pi + 2 h))/2 = 1/(1 - \pi) + O(n^{-1})$, $\tanh(- \beta \pi - h) = - \pi + O(n^{-1})$. The result then follows from replacing these quantities in $c_i^{+}, c_i^{-}, d^+, d^-$ by corresponding ones using $\pi$.

Proof of Lemma (ref)

We decompose by $\Delta_{2,2} = \Delta_{2,2,1} + \Delta_{2,2,2}$, where

align*[align* omitted — 352 chars of source]

Notice that the first term is a quadractic form. Define $\mathbf{H}$ such that $H_{ij} = \frac{g_i^{\prime}(1,\pi) E_{ij}}{2 \mathbb{E}[p_i] N_i}$. Then $\Delta_{2,2,1} = n^{-\mathtt{a}_{\beta, h}} (\mathbf{W} - \pi)^{\operatorname{T}} \mathbf{H} (\mathbf{W} - \pi)$. Take $\mathsf{U}_n$ to be the latent variable from Lemma (ref). Then we decompose

align*[align* omitted — 109 chars of source]

where

align*[align* omitted — 734 chars of source]

Since $\lVert \mathbf{H} \rVert_2 \leq \lVert \mathbf{H} \rVert_F \leq \frac{B}{2 \pi} \sqrt{n}(\min_i N_i)^{-1/2}$, we can apply Hanson-Wright inequality conditional on $\mathsf{U}_n, \mathbf{E}$,

align*[align* omitted — 108 chars of source]

Since $g_i^{\prime}(1,\pi)$'s are independent to $W_i$, by Lemma (ref), $$n^{-\mathtt{a}_{\beta, h}}\sum_{i = 1}^n (W_i - \pi) g_i^{\prime}(1,\pi) = O_{\psi_{\beta,h},tc}(1).$$ By Equation (ref), Lipschitzness of $\tanh$ and Lemma (ref), $\mathbb{E}[W_i|\mathsf{U}_n] - \pi = O_{\psi_{\beta,h}}(n^{-\mathtt{r}_{\beta, h}})$, hence

align*[align* omitted — 282 chars of source]

Then by concentration of $\frac{M_i}{N_i}$ from Lemma (ref), we have

align*[align* omitted — 538 chars of source]

The bound for $\Delta_{2,2,1,d}$ follows from the definition of $\mathbf{H}$ and $\mathsf{U}_n$,

align*[align* omitted — 278 chars of source]

For $\Delta_{2,2,2}$, a Taylor expansion of $p_i$ in terms of $\mca m_i$, and the concentration of $\frac{M_i}{N_i}$ in Lemma (ref) implies that

align*[align* omitted — 602 chars of source]

Proof of Lemma (ref)

Throughout the proof, the Ising spins $\mathbf{W}=(W_i)_{i=1}^n$ are distributed according to Assumption (ref) with parameters $(\beta,h)$. For brevity, we write $\mathbb{P}$ in place of $\mathbb{P}_{\beta,h}$.

Take $\mathsf{U}_n$ to be the latent variable given in Lemma (ref). We further decompose by

align*[align* omitted — 242 chars of source]

where $\eta_i^{\ast}$ is some value between $\pi$ and $M_i/N_i$, and

align[align omitted — 660 chars of source]
center[center omitted — 54 chars of source]

Since $\mathbb{E}[W_i|\mathsf{U}_n,\mathbf{U}] = \tanh\left(\sqrt{\frac{\beta}{n}}\mathsf{U}_n + h \right)$, we have $\mathbb{E}[W_i|\mathsf{U}_n] - \pi = O_{\psi_{\beta,h}}(n^{-\mathtt{r}_{\beta, h}})$ and $(\mathbb{E}[W_i|\mathsf{U}_n] - \pi)^2 = O_{\psi_{\mathtt{p}_{\beta, h}/2}}(n^{-2\mathtt{r}_{\beta, h}})$. It then follows from boundness of $g_i^{(2)}(1, \eta_i^{\ast})$ that

align*[align* omitted — 100 chars of source]
center[center omitted — 55 chars of source]

Condition on $\mathsf{U}_n$, $W_i$'s are i.i.d. By Mc-Diarmid inequality conditional on $\mathsf{U}_n$ for each $\sum_{j \neq i}\frac{E_{ij}}{N_i}(W_j - \mathbb{E}[W_j|\mathsf{U}_n])$ and using a union bound over $i \in [n]$, for all $i \in [n]$, for all $t > 0$,

align*[align* omitted — 222 chars of source]

The tails for $n^{\mathtt{r}_{\beta, h}}(\mathbb{E}[W_j|\mathsf{U}_n] - \pi)$ are also controlled,

align*[align* omitted — 183 chars of source]

Integrate over the distribution of $\mathsf{U}_n$ and using a union bound, for large $n$, for all $t > 0$,

align*[align* omitted — 189 chars of source]

By Equation (ref), condition on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A}$, $\mathbb{P}(N_i \leq \mathbb{E}[N_i|\mathbf{U}]/3|\mathbf{U}) \leq n^{-100}$. Hence for such $\mathbf{U}$,

align*[align* omitted — 209 chars of source]

In other words, conditional on $\mathbf{U}$ s.t. $A(\mathbf{U}) \in \mathcal{A}$,

align*[align* omitted — 101 chars of source]
center[center omitted — 56 chars of source]

For notational simplicity, we will denote

align*[align* omitted — 328 chars of source]

and since we assume $g_i(\ell,\cdot)$ is $C^4$ for $\ell \in \{-1,1\}$, we know $\theta(\ell,\cdot)$ is $C^2$ for $\ell \in \{-1,1\}$. Then we can decompose $\Delta_{2,3,1,a} - \mathbb{E}[\Delta_{2,3,1,a}|\mathbf{E}]$ as

align*[align* omitted — 309 chars of source]

where $F$ is a function that possibly depends on $\beta(\mathbf{U})$ and $\mathbf{E}$.

First part of $\Delta_{2,3,1,a}$: The first two terms have a quadratic form in $W_j - \mathbb{E}[W_j|\mathsf{U}_n]$, except for the term $\theta(M_i/N_i)$. We will handle it via a generalized version of Hanson-Wright inequality. Fix $\mathsf{U}_n$ and $\mathbf{E}$, consider

align*[align* omitted — 193 chars of source]

Denoting by $D_k H$ the partial derivative of $H$ w.r.p to $W_k$ and $D_{k,l}$ the mixed partials, then

align*[align* omitted — 410 chars of source]

Since we have assumed $f$ is at least $4$-times continuously differentiable, we can apply standard concentration inequalities for $\sum_{j \neq i}\frac{E_{ij}}{N_i}(W_j - \mathbb{E}[W_j|\mathsf{U}_n])$ to get

align*[align* omitted — 125 chars of source]

Hence the gradient of $H$ is bounded by

align*[align* omitted — 400 chars of source]

Moreover, the mix partials are

align*[align* omitted — 537 chars of source]

Hence $\lVert D_{k,l}H(\mathbf{W}) \rVert_{\infty} \lesssim n^{-1/2}\sum_{i =1}^n \frac{E_{ik} E_{il}}{N_i^2}$. Hence

align*[align* omitted — 399 chars of source]

Moreover, since $H F$ is symmetric,

align*[align* omitted — 239 chars of source]

Hence by Theorem 3 from dagan2021learning, for all $t > 0$,

align*[align* omitted — 355 chars of source]

By Equation (ref) and a similar argument for upper bound, for each $i \in [n]$, conditional on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A}$, with probability at least $1 - n^{-100}$, $\mathbb{E}[N_i|\mathbf{U}]/2 \leq N_i \leq 2 \mathbb{E}[N_i|\mathbf{U}]$. Hence for each $t > 0$,

align*[align* omitted — 274 chars of source]

that is

align[align omitted — 282 chars of source]

Second part of $\Delta_{2,3,1,a}$: Next, we will show $n^{1 - \mathtt{a}_{\beta, h}} \left(\mathbb{E} \left[B_i \middle| \mathsf{U}_n, \mathbf{U}, \mathbf{E} \right] - \mathbb{E} \left[B_i | \mathbf{E} \right]\right)$, is small. There exists a function $F$ that possibly depends on $\beta$ and $\mathbf{E}$ such that

align*[align* omitted — 182 chars of source]

Define $p(u) = \mathbb{P}(W_j = 1|\mathsf{U}_n,\mathbf{U})$. Then

align*[align* omitted — 228 chars of source]

Using chain rule and product rule for derivatives,

align*[align* omitted — 922 chars of source]

where in the last line, we have used

align*[align* omitted — 534 chars of source]

and that fact that $\lVert p^{\prime} \rVert_{\infty} = O((2 \beta/n)^{0.5})$ and Hoeffiding's inequality for $\sum_{j \neq i}\frac{E_{ij}}{N_i}(W_j - \mathbb{E}[W_j|\mathsf{U}_n])$,

align[align omitted — 312 chars of source]

Since $\mathsf{U}_n = O_{\psi_{\beta, h}}(n^{\mathtt{a}_{\beta, h} - 1/2})$, we have

align[align omitted — 353 chars of source]

Combining Equations (ref) and (ref), conditional on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A}$,

align*[align* omitted — 337 chars of source]

Combining the bounds for $\Delta_{2,3,1,a}, \Delta_{2,3,1,b}, \Delta_{2,3,1,c}$, we get the desired result.

Proof of Lemma (ref)

Throughout the proof, the Ising spins $\mathbf{W}=(W_i)_{i=1}^n$ are distributed according to Assumption (ref) with parameters $(\beta,h)$. For brevity, we write $\mathbb{P}$ in place of $\mathbb{P}_{\beta,h}$.

Recall

align*[align* omitted — 291 chars of source]

First, we will consider the effect of fluctuation of $p_i$ and $\mathbb{E}[W_i|\mathbf{W}_{-i}]$. Recall

align*[align* omitted — 176 chars of source]

It follows from the boundeness of $\beta \mca m_i + h$, $\mca m_i - \pi = O_{\psi_{\beta,h}}(n^{-\mathtt{r}_{\beta, h}})$ that for each $i \in [n]$,

align*[align* omitted — 147 chars of source]

Moreover for some $\eta_i^{\ast}$ between $M_i/N_i$ and $\pi$, using Lemma (ref) we have

align*[align* omitted — 348 chars of source]

Using a union bound over $i$ and an argument for the product of two terms with bounded Orlicz norm with tail control, we have

align*[align* omitted — 407 chars of source]

Next, we will show $n^{-\mathtt{a}_{\beta, h}} \sum_{i = 1}^n \frac{W_i - \pi}{\pi + 1} \left[ g_i \left(1, \frac{M_i}{N_i} \right) - g_i(1, \pi) - g_i^{\prime}(1,\pi) \left(\frac{M_i}{N_i} - \pi \right)\right]$ is small. Suppose $g_i(1,\cdot)$ is $p$-times continuously differentiable. Define

align*[align* omitted — 164 chars of source]

We will use the conditioning strategy to analyse $\delta_p$: Decompse by

align*[align* omitted — 72 chars of source]

with

align*[align* omitted — 654 chars of source]

First, we will show $\delta_{p,2}$ and $\delta_{p,3}$ are small. By Hoeffding inequality, $M_i/N_i - \mathbb{E}[W_i|\mathsf{U}_n] = O_{\psi_2}(N_i^{-1/2})$. Moreover, $\mathbb{E}[W_i|\mathsf{U}_n] - \pi = O_{\psi_{\beta,h}}(n^{-\mathtt{r}_{\beta, h}})$. Hence

align*[align* omitted — 74 chars of source]

For $\delta_{p,3}$, we have

align*[align* omitted — 227 chars of source]

where $\xi^{\ast}$ is some quantity between $\mathbb{E}[W_i|\mathsf{U}_n]$ and $\pi$. Since $x \mapsto x^{p-1}$ is either monotone or convex and none-negative, condition on $\mathbf{E}$,

align*[align* omitted — 340 chars of source]

Combining with boundedness of $g_i^{(p)}(1,\pi)$ and tail control of $\mathbb{E}[W_i|\mathsf{U}_n]$, we have

align*[align* omitted — 269 chars of source]

For $\delta_{p,1}$, we will again use the generalized version of Hanson-Wright inequality. For each $k \in [n]$,

align*[align* omitted — 369 chars of source]

Hence condition on $\mathbf{E}$,

align*[align* omitted — 141 chars of source]

Taking mixed partials w.r.p $\delta_{p,1}$ and using boundedness of $g_i^{(p)}$, we have

align*[align* omitted — 251 chars of source]

It follows that

align*[align* omitted — 268 chars of source]

It then follows from Equation (ref) and Theorem 3 in dagan2021learning that conditional on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A}$,

align*[align* omitted — 232 chars of source]

\paragraph*{Trade-off Between Smoothness of $g_i(1, \cdot)$ and Sparsity of Graph} Assume $g_i(1,\cdot)$ is $p+1$-times continuously differentiable. Then by the decomposition of $\Delta_{2,3,2}$, condition on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A}$,

align*[align* omitted — 619 chars of source]

Then by the concentration of $M_i/N_i - \pi$ given in Lemma (ref), we have

align*[align* omitted — 566 chars of source]

Proof of Lemma (ref)

For notational simplicity, denote $\widehat{\mca p} = \frac{1}{n}\sum_{i = 1}^n T_i$ and $\mca p = \frac{1}{2}\tanh(\beta \pi + h) + \frac{1}{2} = \frac{1}{2} \pi + \frac{1}{2}$. Then

align*[align* omitted — 235 chars of source]

Taylor expand $x \mapsto \tanh(\beta x + h)$ at $x = \pi$, we have

align*[align* omitted — 314 chars of source]

where $O(\cdot)$ is up to a universal constant. Together with concentration of $\frac{1}{n}\sum_{i = 1}^n T_i Y_i$ towards $p \mathbb{E}[Y_i]$, we have

align*[align* omitted — 250 chars of source]

A Taylor expansion of $g_i$ and concentration of $M_i / N_i$ then implies

align*[align* omitted — 243 chars of source]

The conclusion then follows.

Proof of Lemma (ref)

Throughout the proof, the Ising spins $\mathbf{W}=(W_i)_{i=1}^n$ are distributed according to Assumption (ref) with parameters $(\beta,h)$. For brevity, we write $\mathbb{P}$ in place of $\mathbb{P}_{\beta,h}$.

By Lemma (ref) to Lemma (ref), we show

align[align omitted — 186 chars of source]

where $R_i = \frac{g_i(1,\frac{M_i}{N_i})}{1 + \pi} + \frac{g_i(-1,\frac{M_i}{N_i})}{1 - \pi}$, and $b_i = \sum_{j \neq i} \frac{E_{ij}}{N_j} g_j^{\prime}\left(1, \pi \right)$, and $\varepsilon$ is such that condition on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A} = \{A \in \bb R^{n \times n}: \min_{i \in [n]} \sum_{j \neq i}A_{ij} \geq 32 \log n\}$,

align[align omitted — 492 chars of source]

Following the strategy as in the proof of Theorem 4 in li2022random, we will show $b_i$ is close to $R_i$: First, decompose by

align*[align* omitted — 307 chars of source]

By Equation (ref), condition on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A}$, $$|\sum_{j \neq i} \frac{E_{ij}}{N_j}g_j^{\prime}(1,\pi) - \sum_{j\neq i}\frac{E_{ij}}{n \mathbb{E}[G(U_i,U_j)|U_j]}g_j^{\prime}(1,\pi)| \leq C n^{-1/2}$$ with probability at least $1 - n^{-99}$. Moreover, $\frac{E_{ij}}{\mathbb{E}[G(U_i,U_j)|U_j]}g_j^{\prime}(1,\pi), j \neq i$ are i.i.d condition on $U_i$, hence $\sum_{j \neq i}\frac{E_{ij}}{n \mathbb{E}[G(U_i,U_j)|U_j]}g_j^{\prime}(1,\pi) - R_i = O_{\psi_2}((n \mathbb{E}[G(U_i,U_j)|U_j]^{-1/2}) = O_{\psi_2}(\mathbb{E}[N_j|X]^{-1/2})$. It follows that conditional on $\mathbf{U}$ such that $A(\mathbf{U}) \in \mathcal{A}$,

align[align omitted — 155 chars of source]

Again using the conditional i.i.d decomposition, Hoeffiding inequality and $\mathsf{U}_n$'s concentration for the two terms respectively,

align[align omitted — 720 chars of source]

Hence denote the term of stochastic linearization by $G_n$, i.e.

align*[align* omitted — 105 chars of source]

Since $R_i - \mathbb{E}[R_i] + Q_i$'s are i.i.d independent to $W_i$'s with bounded third moment, we know from Lemma (ref) that $G_n$ can be approximated by either a Gaussian or non-Gaussian law, that is order $1$, this gives

align*[align* omitted — 720 chars of source]

where $O(\cdot)$ does not depend on the value of $\mathbf{U}$ and

align*[align* omitted — 326 chars of source]

To analyse the second term, recall $\mathbb{E}[N_i|\mathbf{U}] = \rho_n \sum_{j \neq i}G(U_i,U_j)$. Hence

align*[align* omitted — 345 chars of source]

the last line is because with probability at least $1 - n^{-98}$, $E = \{\frac{1}{2}g(U_i) \leq \frac{1}{n}\sum_{j \neq i}G(U_i, U_j) \leq 2 g(U_i), \forall 1 \leq i \leq n\}$ happens, and by maximal inequality, $\max_i |g(U_i)|^{-1/2} = O_{\psi_2}(\sqrt{\log n})$. And on $\{A(\mathbf{U}) \in \mathcal{A}\} \cap E$, $\max_i (\frac{1}{n}\sum_{j \neq i}G(U_i,U_j))^{-1/2} \leq (32 \log n/n)^{-1/2}$, since we assume $G$ is positive. By similar argument for the last two terms in $\mathtt{r}(\mathbf{U})$, we have

align*[align* omitted — 230 chars of source]

Recall that $\mathcal{A} = \{A(\mathbf{U}): \min_i \sum_{j \neq i} A_{ij}(\mathbf{U})\geq 32 \log n \}$. Since $\sum_{j \neq i} A_{ij}(\mathbf{U}) \sim \operatorname{Bin}(n-1,\mathbb{E}[G(X_1,X_2)])$, we know from Chernoff bound for Binomials and union bound over $i$ that $\mathbb{P}(A(\mathbf{U}) \notin \mathcal{A}) \leq n^{-99}$. The conclusion then follows.

Proof of Lemma (ref)

Our proof for Lemma (ref) to Lemma (ref) relies on the following devices:

(1) Taylor expansion of $\tanh(\cdot)$ in the inverse probability weighting for unbiased estimator, and taylor expansion of $Y_i(\ell,\cdot)$ at $\mathbb{E}[T_i]$ for $\ell \in \{0,1\}$. Then the higher order terms are in terms of $\mca m - \pi$ and $\frac{M_i}{N_i} - \pi$. In Lemma (ref) (taking $X_i \equiv 1$), we show

align*[align* omitted — 73 chars of source]

and in Lemma (ref), we show

align*[align* omitted — 113 chars of source]

where $\mathtt{K}$ is some constant that does not depend on $\beta$. This shows for the higher order terms, we always have

align*[align* omitted — 111 chars of source]

where the $o_\mathbb{P}(\cdot)$ terms does not depend on $\beta$.

(2) Condition i.i.d decomposition based on the de-Finetti's lemma (Lemma (ref)). Suppose $\mathsf{U}_n$ is the latent variable from Lemma (ref), we use decompositions based on $\mathsf{U}_n$: For Lemma (ref) to Lemma (ref), we break down higher order terms in the form

align*[align* omitted — 307 chars of source]

For the first part $F(\mathbf{W},\mathbf{E}) - \mathbb{E}[F(\mathbf{W},\mathbf{E})|\mathbf{E},\mathsf{U}_n]$, we use the conditional i.i.d of $W_i$'s given $\mathsf{U}_n$. For the second part, we use concentration from Lemma (ref) that there exists a constant $\mathtt{K}$ not depending on $\beta$ or $n$, such that $\lVert \mathsf{U}_n \rVert_{\psi_1} \leq \mathtt{K} n^{1/4}$ and the effective term $\lVert \tanh(\sqrt{\frac{\beta}{n}} \mathsf{U}_n) \rVert_{\psi_1} \leq \mathtt{K} n^{-1/4}$. In particular, the rate of concentration for conditional i.i.d Berry-Esseen and concentration of $\tanh(\sqrt{\frac{\beta}{n}}\mathsf{U}_n)$ does not depend on $\beta$.

By the same proof from Lemma (ref) to Lemma (ref), we can show in $\wh \tau_n - \tau_n$, the second and higher order terms in terms of $W_i - \pi$ can always be dominated by the first order terms, with a rate that does not depend on $\beta$.

The conclusion then follows from the two devices and the same proof logic of Lemma (ref) to Lemma (ref).

Proofs: Section (ref)

Proof of Lemma (ref)

Define $g(U_j) =\mathbb{E}[G(U_i,U_j)|U_j]$, for $i \neq j$. Reordering the terms,

align*[align* omitted — 145 chars of source]

Hence $\tau^a_{(i)} - \overline{\tau}^a$ has the representation given by

align[align omitted — 996 chars of source]

where the second to last line is due to $-\frac{1}{n} \frac{1}{1/2} 1/2(h_i(1,0) - \mathbb{E}[h_i(1,0)]) + \frac{1}{n} \frac{1}{1 - 1/2} (1 - 1/2) (h_i(-1,0) - \mathbb{E}[h_i(-1,0)]) = - \frac{2}{n} \varepsilon_i + \frac{2}{n} \varepsilon_i = 0$.

Now we look at $b$-part. For representation purpose, we look at only the treatment part. The control part can be analysized by in the same way. Reordering the terms,

align*[align* omitted — 416 chars of source]

Hence $\tau_{(i)}^b - \overline{\tau}^b$ has the representation given by

align[align omitted — 239 chars of source]

The analysis follows from a Taylor expansion of $h_j(1,\cdot)$. For some $\xi_{j,i}^{\ast}$ between $\frac{M_j}{N_j}_{(i)}$ and $0$ for each $j,i$,

align[align omitted — 321 chars of source]

where we have used $\partial_2 h_j(1,\cdot) = \partial_2 [h(1,\cdot) + \varepsilon_j] = \partial_2 h(1,\cdot)$.

\paragraph*{Part 1: Linear Terms}

equation[equation omitted — 456 chars of source]

By a decomposition argument,

align*[align* omitted — 644 chars of source]

Hence

align*[align* omitted — 468 chars of source]

Condition on $U_j$, $(E_{lj} W_l: l \neq j)$ are i.i.d mean-zero, hence Bernstein inequality gives $\frac{1}{n} \sum_{l = 1}^n E_{lj} W_l = O_{\psi_2}(\sqrt{n^{-1}\rho_n}) + O_{\psi_1}(n^{-1})$, which implies

align*[align* omitted — 375 chars of source]

Putting back into Equation (ref),

align*[align* omitted — 177 chars of source]

Looking at contribution from the first order term in Taylor expanding $h_j(1,\cdot)$ to $\tau_{(i)}^b - \overline{\tau}^b$ in Equation (ref),

align*[align* omitted — 796 chars of source]

Since $(E_{ij} T_j/g(U_j): j \in [n])$ are independent condition on $U_i$, standard concentration inequality gives

align*[align* omitted — 640 chars of source]

Since we assumed $\partial_2 h(1,0) = \partial_2 f(1,0) + o_{\mathbb{P}}(1) = \partial_2 f_j(1,0) + o_{\mathbb{P}}(1)$ where

align*[align* omitted — 355 chars of source]

Together with the leading term in Equation (ref), we have

align*[align* omitted — 1,444 chars of source]

\paragraph*{Part 2: Higher Order Terms} For the second order terms, first notice that if $l \notin [n]$, then

align*[align* omitted — 452 chars of source]

where we have used $(M_j/N_j)_{\iota} = O_{\psi_2}((n \rho_n)^{-\frac{1}{2}})$ and $N_j^{-1} = O_{\psi_2}((n \rho_n)^{-1})$. If $l \in [n]$, then again

align*[align* omitted — 433 chars of source]

Hence

align*[align* omitted — 473 chars of source]

For the third order residual, observe that $(\frac{M_j}{N_j}_{(\iota)})^3 = O_{\psi_2}((n \rho_n)^{-3/2})$. Then

align*[align* omitted — 625 chars of source]

The conclusion then follows from Equations (ref), (ref) and (ref).

Proof of Lemma (ref)

Define $\mathtt{r}(x) = (1,x)^\top$. Denote $\pi = \mathbb{E}[W_i] = 2 \mathbb{E}[T_i] - 1$. Then

center[center omitted — 46 chars of source]

First, consider the gram-matrix. Take $\zeta_i := \sqrt{n \rho_n} (\frac{M_i}{N_i} - \pi)$. Then $1 \lesssim \mathbb{V}[\zeta_i] \lesssim 1$. Take $b_n = \sqrt{n \rho_n} h_n$. Take

align*[align* omitted — 186 chars of source]

where $\mathtt{r}: \mathbb{R} \rightarrow \mathbb{R}^2$ is given by $\mathtt{r}(u) = (1,u)^{\top}$. Take $Q$ to be the probability measure of $\zeta_i$ given $\mathbf{E}$. Then

align*[align* omitted — 423 chars of source]

In particular, $\lambda_{\min}(\mathbf{B}) \gtrsim 1$. Now we want to show each entry of $\mathbf{B}_n$ converge to those of $\mathbf{B}$. Take

align*[align* omitted — 218 chars of source]

Denote $\partial_j$ to be the partial derivative w.r.p to $W_j$. Since $K$ is Lipschitz with bounded support,

align[align omitted — 248 chars of source]

Condition on $\mathbf{E}$,

align*[align* omitted — 320 chars of source]

Hence for all $p,q \in \{0,1\}$,

align*[align* omitted — 139 chars of source]

Since both $\mathbf{B}_n$ and $\mathbf{B}$ are two by two matrices, $\lVert \mathbf{B}_n - \mathbf{B} \rVert_{\operatorname{op}} \lesssim O_{\psi_2}((n b_n^4)^{-1})$. By Weyl's Theorem,

align[align omitted — 189 chars of source]

and together with $\lambda_{\min}(\mathbf{B}) \gtrsim 1$, implies $\lambda_{\min}(\mathbf{B}_n) \gtrsim 1$. Take

align*[align* omitted — 222 chars of source]

Hence variance can be bounded by

align[align omitted — 424 chars of source]

Next, consider the bias term. Since $f(1,\cdot) \in C^2$, whenever $|\frac{M_i}{N_i} - \pi| \leq h_n = (n \rho_n)^{-1/2} b_n$,

align*[align* omitted — 251 chars of source]

Hence using the fourth and third lines above respectively,

equation[equation omitted — 1,356 chars of source]

Putting together Equations (ref) and (ref),

align*[align* omitted — 220 chars of source]

Hence any $b_n$ such that $b_n = \Omega(n^{-1/4} + \rho_n^{1/3})$ will make $(\widehat{\gamma}_0, \widehat{\gamma}_1)$ a consistent estimator for $(\gamma_0,\gamma_1)$. For any $0 \leq \rho_n \leq 1$ such that $n \rho_n \rightarrow \infty$, such a sequence $b_n$ exists.

center[center omitted — 46 chars of source]

The order $\frac{M_i}{N_i}$ is $n^{-1/4}$ if $\liminf_{n \rightarrow \infty} n \rho_n^2 > c$ for some $c > 0$; and is $(n \rho_n)^{-1/2}$ if $n \rho_n^2 = o(1)$. We consider these two cases separately.

\paragraph*{Case 2.1: $\liminf_{n\rightarrow \infty}n \rho_n^2 > c$ for some $c > 0$} Take $\eta_i = n^{\frac{1}{4}}(\frac{M_i}{N_i} - \pi)$. Take $d_n = n^{1/4}h_n$. And with the same $\mathtt{r}$ defined in Case 1,

align*[align* omitted — 229 chars of source]

Under the assumption $\liminf_{n \rightarrow \infty }n \rho_n^2 \leq c$ for some $c > 0$, we have $1 \lesssim \mathbb{V}[\eta_i] \lesssim 1$. Hence $\lambda_{\min}(\mathbf{D}) \gtrsim 1$. To study the convergence between $\mathbf{D}_n$ and $\mathbf{D}$, again consider for $p, q \in \{0,1\}$,

align*[align* omitted — 323 chars of source]

Still let $\mathsf{U}_n$ be the latent variable from Lemma (ref), $W_i$'s are independent conditional on $\mathsf{U}_n$. Hence by similar argument as Equation (ref), we can show

align*[align* omitted — 125 chars of source]

Moreover, recall we denote by $\omega_i \in [k]$ the block unit $i$ belongs to, then

align*[align* omitted — 228 chars of source]

$p(U_l)= \mathbb{P}(W_i = 1|U_{\ell}) = \frac{1}{2}(\tanh(\sqrt{\beta_{\ell}/n}\mathsf{U}_n + h_{\ell}) + 1)$, $i \in \ca I_{\ell}$. Take the derivative term by term,

align*[align* omitted — 229 chars of source]

Using Lipschitz property of $x \mapsto (x/h_n)^{p+q} K(x/h_n)$,

align*[align* omitted — 157 chars of source]

Hence for all $\ell \in \mathscr{C}$,

align*[align* omitted — 274 chars of source]

Moreover, for all $\ell \in \mathscr{C}$, $\lVert U_{\ell} \rVert_{\varphi_2} \lesssim n^{1/4}$. Together, this gives

align*[align* omitted — 185 chars of source]

Hence if we take $d_n \gg 1$ (which implies $n d_n^4 \gg 1$), then $G_{p,q}(\mathbf{W}) = \mathbb{E}[G_{p,q}(\mathbf{W})|\mathbf{E}] + o_{\mathbb{P}}(1)$, implying $\lVert \mathbf{D}_n - \mathbf{D} \rVert_2 = o_{\mathbb{P}}(1)$ and $\lambda_{\min}(\mathbf{D}_n) - \lambda_{\min}(\mathbf{D}) = o_{\mathbb{P}}(1)$, making $\lambda_{\min}(\mathbf{D}_n) \gtrsim_{\mathbb{P}} 1$. Take

align*[align* omitted — 220 chars of source]

Hence variance can be bounded by

align[align omitted — 428 chars of source]

By similar argument as in Case 1, assume $d_n \gg 1$, we can show

align*[align* omitted — 181 chars of source]

Hence if we choose $d_n$ such that $1 \ll d_n \ll n^{1/8}$, then $(\widehat{\gamma}_0, \widehat{\gamma}_1)$ is a consistent estimator for $(\gamma_0, \gamma_1)$. The only assumption we made for the existence of such a $d_n$ is $\liminf_{n \rightarrow \infty} n \rho_n^2 \geq c$ for some $c > 0$.

\paragraph*{Case 2.2: $n \rho_n^2 = o(1)$} Take $\eta_i := \sqrt{n \rho_n}(\frac{M_i}{N_i} - \pi)$, $d_n = \sqrt{n \rho_n} h_n$. By similar decomposition based on latent variables, we can show if $n \rho_n \rightarrow \infty$ as $n \rightarrow \infty$, then there exists $h_n$ such that $(\widehat{\gamma}_0, \widehat{\gamma}_1)$ is a consistent estimator for $(\gamma_0, \gamma_1)$.

Proofs: Section (ref)

Preliminary Lemmas

lemmaRecall $\mathbf{W} = (W_i)_{1 \leq i \leq n}$ takes value in $\{-1,1\}^n$ with \begin{align*} \mathbb{P} \left( \mathbf{W} = \mathbf{w} \right) = \frac{1}{Z} \exp \bigg( \frac{\beta}{n} \sum_{i < j} W_i W_j\bigg), \quad \beta > 1. \end{align*} Recall $\pi_+$ and $\pi_-$ are the positive and negative solutions to $x = \tanh(\beta x + h)$, respectively, and $\mca m = n^{-1} \sum_{i = 1}^n W_i$. Then $\mathbb{E}[W_i|\operatorname{sgn}(\mca m) = \ell] = \pi_{\ell} + O(n^{-1})$ for $\ell = -, +$.
proofThe conditional concentration of $\mca m = n^{-1} \sum_{i = 1}^n W_i$ towards $\pi_{\ell}$ in Lemma (ref) implies, \begin{align*} & \mathbb{E}[W_i|\operatorname{sgn}(\mca m) = \ell] \\ = & \mathbb{E}[\mathbb{E}[W_i|W_{-i}, \operatorname{sgn}(\mca m) = \ell]| \operatorname{sgn}(\mca m) = \ell] \\ = & \mathbb{E}[\tanh(\beta \mca m_i + h) \mathbbm{1}(\operatorname{sgn}(\mca m_i) = \ell)|\operatorname{sgn}(\mca m) = \ell] + O(n^{-1})\\ = & \mathbb{E}[\tanh(\beta \pi + h) + \operatorname{sech}^2(\beta \pi + h)(\mca m_i - \pi) - \operatorname{sech}^2(\beta m^{\ast} + h)\tanh(\beta m^{\ast} + h)(\mca m_i - \pi)^2] + O(n^{-1})\\ = & \tanh(\beta \pi + h) + O(n^{-1}) \\ = & \pi + O(n^{-1}), \end{align*} where $\mca m^{\ast}$ is a number between $\mca m$ and $\pi$, and we have used boundedness of $\operatorname{sech}$.
lemmaSuppose Assumption (ref), and Assumption 2, 3 hold with $h = 0$ and $\beta > 1$. Then for $\ell = _, +$: (1) Condition on $\operatorname{sgn}(\mca m) = \ell$, \begin{align*} & \max_{i \in [n]} \left|\frac{M_i}{N_i} - \pi_{\ell} \right| = O_{\psi_1}(n^{-1/2}) + O_{\psi_2}(N_i^{-1/2}). \end{align*} (2) Define $A(\mathbf{U}) = (G(U_i,U_j))_{1 \leq i,j \leq n}$. Condition on $\mathbf{U}$ such that $$A(\mathbf{U}) \in \mathcal{A} = \{A \in \bb R^{n \times n}: \min_{i \in [n]} \sum_{j \neq i}A_{ij} \geq 32 \log n\}$$ for large enough $n$, for each $i \in [n]$ and $t > 0$, \begin{align*} \mathbb{P}_{\beta,h}\left( \left|\frac{M_i}{N_i} - \pi_{\ell}\right| \geq 4 \mathbb{E}[N_i | \mathbf{U}]^{-1/2}t^{1/2} + C n^{-1/2} t^{1/2}\middle| \mathbf{U}, \operatorname{sgn}(\mca m) = \ell \right) \leq 2 \exp(-t) + n^{-98}, \end{align*} where $C$ is some absolute constant.
proofThroughout the proof, the Ising spins $\mathbf{W}=(W_i)_{i=1}^n$ are distributed according to Assumption (ref) with parameters $(\beta,h)$. For brevity, we write $\mathbb{P}$ in place of $\mathbb{P}_{\beta,h}$. Let $\mathsf{U}_n$ to be the latent variable defined in Lemma (ref). Decompose by \begin{align*} \frac{M_i}{N_i} - \pi_{\ell} = \sum_{j \neq i} \frac{E_{ij}}{N_i} \left(W_j - \mathbb{E}[W_j|\mathsf{U}_n] \right) + \mathbb{E}[W_j|\mathsf{U}_n] - \pi_{\ell}. \end{align*} Condition on $\mathsf{U}_n$, $W_i$'s are i.i.d. Berry-Esseen theorem gives that with $\mathsf{Z} \sim \mathsf{N}(0,1)$ independent to $\mathsf{U}_n$, we have \begin{align} \sup_{t \in \mathbb{R}} \bigg|\mathbb{P}\Big(\frac{M_i}{N_i} \leq t\Big|\mathbf{E}, \mathsf{U}_n \Big) - \mathbb{P} \Big(\sqrt{\frac{v(\mathsf{U}_n)}{N_i}} \mathsf{Z} + e(\mathsf{U}_n) \leq t \Big|\mathbf{E}, \mathsf{U}_n\Big) \bigg| = O(n^{-\frac{1}{2}}), \end{align} where $e(\mathsf{U}_n) = \mathbb{E}[W_i|\mathsf{U}_n] - \pi = \tanh(\sqrt{\beta/n}\mathsf{U}_n + h) - \pi$, and $v(\mathsf{U}_n) = \mathbb{V}[W_i - \pi|\mathsf{U}_n]$. By McDiarmid's inequality, \begin{align*} \mathbb{P} \bigg( |\sum_{j \neq i} \frac{E_{ij}}{N_i} \left(W_j - \mathbb{E}[W_j|\mathsf{U}_n] \right)| \geq 2 N_i^{-1/2} t\bigg | \mathbf{E} \bigg) \leq 2 \exp(-t^2). \end{align*} Conclusion (1) then follows from the conditional concentration of $\mathsf{U}_n$ in Remark (ref). Notice that $\mathbf{W}$ and $\mathsf{U}_n$ are independent to the random graph. Conclusion (1) and the same analysis as in Lemma (ref) give conclusion (2).

Proof of Lemma (ref)

The result is a special case of Lemma (ref) in Section (ref) when $h \neq 0$.

Proof of Theorem (ref)

The result follows from Lemma (ref), Lemma (ref), and the same anti-concentration argument as in the proof of Lemma (ref).

Proof of Lemma (ref)

center[center omitted — 48 chars of source]

First, we consider the unbiased estimator

align*[align* omitted — 133 chars of source]

with $p_i = \bb P(W_i = 1| \mathbf{W}_{-i}) = \left(\exp \left(-2\beta \mca m_i\right) + 1\right)^{-1}$. Our analysis will be similar to the proofs in Section (ref), but using the concentration of $n^{-1} \sum_{i = 1}^n W_i$ conditional on $\operatorname{sgn}(\mca m)$ shown in Lemma (ref) instead of the unconditional concentration of $n^{-1} \sum_{i = 1}^n W_i$. We decompose by

align*[align* omitted — 248 chars of source]

For the first term, we Taylor expand the expression for $p_i^{-1}$ in terms of $\mca m_i$, and get

align*[align* omitted — 314 chars of source]

condition on $\operatorname{sgn}(\mca m) = \ell$, where

align*[align* omitted — 194 chars of source]

For the second term, we Taylor expand $g_i(1, \cdot)$ at $\pi_{\ell}$: For some $\eta_i^{\ast}$ between $\pi_{\ell}$ and $\frac{M_i}{N_i}$,

align*[align* omitted — 202 chars of source]

where

align*[align* omitted — 492 chars of source]

Term $\Delta_{2,1}^{\prime}$: Denote $\mathbf{g} = (g_i)_{1 \leq i \leq n}$. Rearranging the terms,

align*[align* omitted — 256 chars of source]

Term $\Delta_{2,2}^{\prime}$: Take $u_{\ell} = (\pi_{\ell} + 1)/2$ for $\ell \in \{-,+\}$. Decompose by $$\Delta_{2,2}^{\prime} = \Delta_{2,2,1}^{\prime} + \Delta_{2,2,2}^{\prime},$$ where

align*[align* omitted — 335 chars of source]

Since $\beta < \infty$ and $v_{-}$, $u_-$ and $u_+$ are bounded away from $0$ and $1$. Rearranging the terms, we get

align*[align* omitted — 154 chars of source]

where $\mathbf{H}^{\ell}$ is the $n \times n$ matrix with $H^{\ell}_{ij} = g_i^{\prime}(1, \pi_{\ell}) E_{ij} (2 u_{\ell} N_i)^{-1}$ and $\mathbf{1}$ is the $n$-dimensional vector with all entries $1$. To analyze the quadratic form, we use the same strategy as in the proof of Lemma (ref): Let $\mathsf{U}_n$ be the one defined in Lemma (ref), and we know $W_1, \cdots, W_n$ are conditional i.i.d given $\mathsf{U}_n$. Then we can decompose $\Delta_{2,2,1}^{\prime}$ into four terms based on

align*[align* omitted — 169 chars of source]

Conditional Berry-Esseen given $\mathsf{U}_n$, conditional concentration of $\mathsf{U}_n$, $\mca m$ and $\frac{M_i}{N_i}$ given $\operatorname{sgn}(\mca m)$ in Remark (ref), Lemma (ref) and Lemma (ref), and the same argument as in the proof for Lemma (ref) implies that condition on $\mathbf{g}, \mathbf{E}$ and $\operatorname{sgn}(\mca m)$,

align*[align* omitted — 225 chars of source]

Term $\Delta_{2,3}^{\prime}$: Now we proceed to $\Delta_{2,3}^{\prime}$. Decompose by $\Delta_{2,3}^{\prime} = \Delta_{2,3,1}^{\prime} + \Delta_{2,3,2}^{\prime}$, where

align*[align* omitted — 423 chars of source]

Define $\Delta_{2,3,1,l}^{\prime}$ to be the counterparts of $\Delta_{2,3,1,l}$ in Equation (ref) with $\pi$ by replaced by $\pi_{\ell}$ for $l \in \{a,b,c\}$, the same argument in the proof of Lemma (ref) shows

align*[align* omitted — 365 chars of source]

condition on $\operatorname{sgn}(\mca m) = \ell$ for $\ell = -, +$. Combining the three parts,

align*[align* omitted — 103 chars of source]

condition on $\operatorname{sgn}(\mca m) = \ell$ for $\ell = -, +$. Taylor expanding $p_i = (1 + \exp(- 2 \beta \mca m_i))^{-1}$ as a function of $\mca m_i$ at $\pi_{\ell}$, the same argument as in Lemma (ref) shows

align*[align* omitted — 291 chars of source]

Conditional concentration of $\mathsf{U}_n$, $\mca m$ and $\frac{M_i}{N_i}$ given $\operatorname{sgn}(\mca m)$ in Remark (ref), Lemma (ref) and Lemma (ref), and the same argument as in the proof for Lemma (ref) implies that condition on $\mathbf{g}, \mathbf{E}$ and $\operatorname{sgn}(\mca m)$,

align*[align* omitted — 127 chars of source]

Putting together. Putting together the decompositions, condition on $\mathbf{E}$ and $\operatorname{sgn}(\mca m) = \ell$,

align*[align* omitted — 282 chars of source]

where with $c_{i,l} = g_i(1,\pi_{\ell})(1 + \exp(2 \beta \pi_{\ell}))$, and $d_l = \frac{\beta(1 + \exp(2 \beta \pi_{\ell}))}{1 + \cosh(2 \beta \pi_{\ell})}\mathbb{E}[g_i(1, \pi_{\ell})]$,

align*[align* omitted — 163 chars of source]

Consider the event $\Omega_i = \{\operatorname{sgn}(\mca m) = \ell, |\sum_{j \neq i}W_j| \leq 1\}$ and $\Omega = \cup_{1 \leq i \leq n} \Omega_i$. We then have

align*[align* omitted — 120 chars of source]

implying $\mathbb{P}(\Omega_i) \leq C \exp(-nC)$, $1 \leq i \leq n$. Hence

align[align omitted — 1,180 chars of source]

Hence condition on $\mathbf{E}$ and $\operatorname{sgn}(\mca m) = \ell$,

align*[align* omitted — 202 chars of source]
center[center omitted — 46 chars of source]

Now, we consider the difference between the unbiased estimator and the Hajek estimator. For notational simplicity, denote $\widehat{\mca p} = \frac{1}{n}\sum_{i = 1}^n T_i$ and $\mca p_{\ell} = \frac{1}{2}\tanh(\beta \pi_{\ell} + h) + \frac{1}{2} = \frac{1}{2} \pi_{\ell} + \frac{1}{2}$. Then

align*[align* omitted — 256 chars of source]

Taylor expand $x \mapsto \tanh(\beta x + h)$ at $x = \pi_{\ell}$, we have

align*[align* omitted — 384 chars of source]

where $O(\cdot)$ is up to a universal constant. Together with the fact that condition on $\operatorname{sgn}(\mca m) = \ell$, $\frac{1}{n}\sum_{i = 1}^n T_i Y_i$ concentrates towards $\mca m \mathbb{E}[Y_i|\operatorname{sgn}(\mca m) = \ell]$, we have

align*[align* omitted — 304 chars of source]

condition on $\operatorname{sgn}(\mca m) = \ell$. A Taylor expansion of $g_i$ and concentration of $M_i / N_i$ then implies

align*[align* omitted — 434 chars of source]

where $\pi^{\ast}$ is some number between $\pi_{\ell}$ and $M_i/N_i$. The conclusion then follows.

Proof of Lemma (ref)

The result follows from Lemma (ref) (3), Lemma (ref), and the same anti-concentration argument as in the proof of Lemma (ref).

Proof of Lemma (ref)

As in the case of one block analyzed in Section (ref), $\wh \boldsymbol{\tau}_n$ is not an unbiased estimate of $\boldsymbol{\tau}_n$. We first consider an unbiased estimator to $\boldsymbol{\tau}_n$ and then consider the difference.

center[center omitted — 48 chars of source]

Consider $\widehat{\boldsymbol{\tau}}_{n,UB} = (\widehat{\tau}_{n,UB,1}, \cdots, \widehat{\tau}_{n,UB,K})$, where

align*[align* omitted — 156 chars of source]

Here $p_i = \sum_{k = 1}^K \mathbbm{1}(i \in \mathcal{C}_k) (1 + \exp(2 \beta_k \mca m_{i,k} + 2 h_k))^{-1}$ and $\mca m_{i,k} = n_k^{-1} \sum_{j \in \mathcal{C}_k, j \neq i}W_j$.

Denote $\mca m = n^{-1}\sum_{i =1}^n W_i$, $\mca m_k = n_k^{-1} \sum_{i \in \mathcal{C}_k} W_i$. For notational simplicity, we denote $\pi_{l,\operatorname{sgn}(\mca m_l)}$ by $\pi_l$ for low temperature blocks $l \in \mathscr{L}$, and omit the index by $(\mathbf{s})$ with $\mathbf{s} = \boldsymbol{sgn}$. As in the one-block case, we decompose by

align*[align* omitted — 74 chars of source]

where

align*[align* omitted — 260 chars of source]

and $\zeta_i = \frac{\sum_{k =1}^K N_{i,k} \pi_k}{\sum_{k = 1}^K N_{i,k}}$.

center[center omitted — 51 chars of source]

Condition on $\mathbf{E}, \mathbf{g} = \{g_i: i \in [n]\}$ and $\boldsymbol{sgn}$, the randomness of $\Delta_{1,k}$ only comes from $(W_i)_{i \in \mathcal{C}_k}$, that is, the Ising bits from the same block. Hence

align*[align* omitted — 327 chars of source]

The analysis in Lemma (ref) and Lemma (ref) with $g_i(1,\zeta_k)\mathbbm{1}(i \in \mathcal{C}_k)$ replacing $g_i(1,\pi)$ implies

align*[align* omitted — 434 chars of source]

condition on $\mathbf{E}, \mathbf{g}, \boldsymbol{sgn}$, where $\mathtt{c}_k = (1 + \exp(2\beta_k \pi_k + 2 h_k))/2$ and $\mathtt{d}_k = \beta_k(1 + \exp(2 \beta_k \pi_k + 2 h_k))/(1 + \cosh(2 \beta_k \pi_k + 2 h_k))$.

center[center omitted — 51 chars of source]

The linearization of $\Delta_{2,k}$ involves $M_i/N_i$, which depends all blocks even if the estimator is for block $k$. We will find its stochastic linearization in terms of units in all blocks.

By a Taylor expansion of $g_i(1,\cdot)$ at $\zeta_i$

align*[align* omitted — 289 chars of source]

where

align*[align* omitted — 443 chars of source]

with $g_i(x) = \int_0^1 (1 - t) Y_i^{(2)}(1,\zeta_i + t (x - \zeta_i))dt$. In particular, $g_i$ is $C^2$. \paragraph*{Term $\Delta_{2,1}$:} Rearranging $\Delta_{2,1}$, we get the effective term in the stochastic linearization.

align*[align* omitted — 364 chars of source]

\paragraph*{Term $\Delta_{2,2}$:} We want to show $\Delta_{2,2}$ is negligible. Consider the effect from each block separately. We claim that condition on $\mathbf{g}$, $\mathbf{E}$, $\boldsymbol{sgn}$,

align*[align* omitted — 636 chars of source]

To get the second line, notice that $p_i = (1 + \exp(2 \beta_k \mca m_{i,k} + 2 h_k))^{-1}$ is Lipschitz in $\mca m_{i,k}$, and since $(W_i: i \in \mathcal{C}_k), 1 \leq k \leq K$ form independent Ising models, we can use Lemma (ref) to get $\mca m_{i,k} - \pi_k = O_{\psi_{\beta_k,h_k}}(n_k^{-\mathtt{r}_{\beta_k, h_k}})$ for $k \in \mathscr{H} \cup \mathscr{C}$, and condition on $\operatorname{sgn}(\mca m_{k})$, $\mca m_{i,k} - \pi_k = O_{\psi_{\beta_k,h_k}}(n_k^{-\mathtt{r}_{\beta_k, h_k}})$ for $k \in \mathscr{L}$. Hence for each $k \in [K]$,

align*[align* omitted — 193 chars of source]

Suppose $\mathsf{U}_{n,l}$ is the latent variable underlining the distribution of $(W_i: i \in \mathcal{C}_l), l \in [K]$ as in Lemma (ref). Conditional on $\mathbf{E}$ and $\boldsymbol{sgn}$, using Hoeffiding's inequality and the concentration of $\mathsf{U}_{n,l}$, we have

align*[align* omitted — 382 chars of source]

From the fact that $\mathbb{P}(|Z_1 Z_2| \geq t) \leq \mathbb{P}(\sqrt{\log n} |Z_2| \geq t) + \mathbb{P}(|Z_1| \geq \sqrt{\log n})$ for any two random variables $Z_1$ and $Z_2$, and using a union bound over the summation over $i \in \mathcal{C}_k$, we get the second line for $\Delta_{2,2}$.

Now consider the first term of $\Delta_{2,2}$. With the help of the latent variables $\mathsf{U}_{n,k}, 1 \leq k \leq K$, decompose by

align*[align* omitted — 99 chars of source]

where

align*[align* omitted — 869 chars of source]

Since conditional on $\mathsf{U}_{n,k}$ and $\mathsf{U}_{n,l}$, $(W_i: i \in \mathcal{C}_k \cup \mathcal{C}_l)$ are i.i.d., we can use Hoeffding's inequality and boundedness of $g_i^{\prime}(1,\zeta_i)$ to get conditional on $\mathbf{E}$ and $\boldsymbol{sgn}$,

align*[align* omitted — 169 chars of source]

and

align*[align* omitted — 345 chars of source]

It follows that condition on $\mathbf{E}$ and $\boldsymbol{sgn}$,

align*[align* omitted — 620 chars of source]

For $\Gamma_{k,l,a}$, observe that with $\omega_i = \sum_{k = 1}^K k \mathbbm{1}(i \in \mathcal{C}_k)$,

align*[align* omitted — 389 chars of source]

Apply Hanson-Wright inequality conditional on $\mathbf{E}$, $\mathsf{U}_{n,l}$ and $\mathsf{U}_{n,k}$, we get $$\Gamma_{k,l,a} - \mathbb{E}[\Gamma_{k,l,a}|\mathbf{E}, \boldsymbol{sgn}] = O_{\psi_1}((n_k \min_i N_i)^{-\frac{1}{2}}).$$ Put together, conditional on $\mathbf{E}$ and $\boldsymbol{sgn}$,

align*[align* omitted — 262 chars of source]

\paragraph*{Term $\Delta_{2,3}$:} Similar to the analysis in Section (ref), we decompose $\Delta_{2,3} = \Delta_{2,3,1} + \Delta_{2,3,2}$ where

align*[align* omitted — 312 chars of source]

where $\Delta_{2,3,1}$ is further decomposed based on latent variables $\mathsf{U}_{n,l}, 1 \leq l \leq K$, that is, $$\Delta_{2,3,1} = \Delta_{2,3,1,a} + \Delta_{2,3,1,b} + \Delta_{2,3,1,c},$$ where

align*[align* omitted — 744 chars of source]

Term $\Delta_{2,3,1,a}$: Consider the $\Delta_{2,3,1,a}$ as a (random) function on $\mathbf{W}$ and $\mathsf{U}_{n,1}, \cdots, \mathsf{U}_{n,K}$. Let

align*[align* omitted — 236 chars of source]

Notice that conditional on $\mathsf{U}_{n,l}, 1 \leq l \leq K$, $W_j$'s are independent random variables, and we can rewrite $$\Delta_{2,3,1,a} = \frac{n}{n_k} \frac{1}{n} \sum_{i = 1}^n g_i \left(\frac{M_i}{N_i}\right) \mathbbm{1}(i \in \mathcal{C}_k)(\sum_{l =1}^K \sum_{j \in \mathcal{C}_l, j\neq i} \frac{E_{ij}}{N_i}(W_j - \mathbb{E}[W_j|\mathsf{U}_{n,l}]))^2.$$ It follows from the same concentration argument for $\Delta_{2,3,1,a}$ in the proof for Lemma (ref) that conditional on $\mathbf{E}$ and $\boldsymbol{sgn}$,

align*[align* omitted — 199 chars of source]

Define $p_l(u) = \mathbb{E}[W_i = 1| \mathsf{U}_{n,l} = u, i \in \mathcal{C}_l]$. Then we can write

align*[align* omitted — 313 chars of source]

By the same argument as in the proof for Lemma (ref),

align*[align* omitted — 201 chars of source]

It then follows from the concentration of $\mathsf{U}_{n,1}$ to $\mathsf{U}_{n,K}$ that condition on $\mathbf{E}$ and $\boldsymbol{sgn}$,

align*[align* omitted — 236 chars of source]

Moreover, since $\sum_{j \in \mathcal{C}_l, j \neq i} \frac{E_{ij}}{N_i}(W_j - \mathbb{E}[W_j|\mathsf{U}_{n,l}]) = O_{\psi_2}(N_i^{-\frac{1}{2}})$ and $\mathbb{E}[W_j|\mathsf{U}_{n,l}] - \pi_l = O_{\psi_{\beta_l,h_l}}(n^{-\mathtt{r}_{\beta_l, h_l}})$, we have conditional on $\mathbf{E}$ and $\boldsymbol{sgn}$,

align*[align* omitted — 243 chars of source]

Putting together, $$\Delta_{2,3,1} - \mathbb{E}[\Delta_{2,3,1}|\mathbf{E}] = O_{\psi_2, tc}\left(\sqrt{\log n} \max_i N_i^{-\frac{1}{2}} \max_{1 \leq l \leq K}n^{-\mathtt{r}_{\beta_l, h_l}} + \max_{1 \leq l \leq K} n^{-2 \mathtt{r}_{\beta_l, h_l}}\right).$$ Consider the $p$-th order term in the expansion of $\Delta_{2,3,2}$, $$\delta_p = \frac{1}{n_k} \sum_{i \in \mathcal{C}_k} \frac{T_i - \pi_k}{\pi_k} g_i^{(p)}(1, \zeta_i)\bigg(\frac{M_i}{N_i} - \zeta_i\bigg).$$ Following conditional i.i.d argument as in Lemma (ref), we can show condition on $\mathbf{E}$ and $\boldsymbol{sgn}$,

align*[align* omitted — 302 chars of source]

Hence assuming $g_i(1, \cdot)$ is $C^{p+1}$. Taylor expand $g_i(1,\cdot)$ up to the $p$-th order, we get

align*[align* omitted — 565 chars of source]
center[center omitted — 54 chars of source]

From the previous steps, condition on $\mathbf{E}$ and $\boldsymbol{sgn}$,

align*[align* omitted — 285 chars of source]

where $\bar{\mathtt{r}}_n = \max_{k \in [K]}n^{-2 \mathtt{r}_{\beta_k, h_k}} + \sqrt{\log n} \max_{k \in [K]} n^{-\mathtt{r}_{\beta_k, h_k}}(n \rho_n)^{-1} + \sqrt{\log n} n^{-1/2} + (n \rho_n)^{-(p+1)/2}$, and $\bar{\mathbf{S}}_{l,i}$ is the vector $(\bar{S}_{1,l,i}, \cdots, \bar{S}_{K,l,i})^{\mathbf{T}}$, where

align*[align* omitted — 292 chars of source]

with $\zeta_i = \frac{\sum_{\ell = 1}^k N_{i,\ell} \pi_{\ell}}{\sum_{\ell = 1}^k N_{i,\ell}}$. Condition on $U_i$, $E_{i,j}$ for all $1 \leq j \leq n$ are independent with $|E_{ij}| \leq 1$ and $\mathbb{V}[E_{ij}|U_i] \lesssim \rho_n$. Hence using Bernstein's inequality,

align*[align* omitted — 431 chars of source]

with $\overline{\pi} = \sum_{k = 1}^K p_k \pi_k$. The same argument as the proof for Lemma (ref) implies

align*[align* omitted — 232 chars of source]

with

align*[align* omitted — 190 chars of source]

Hence with $\mathbf{S}_{l,i} = (S_{1,l,i}, \cdots, S_{K,l,i})^{\mathbf{T}}$, where

align*[align* omitted — 135 chars of source]

Hence by the same analysis as Equation (ref) in the proof of Lemma (ref),

align*[align* omitted — 371 chars of source]

We already know $\widehat{\boldsymbol{\tau}}_{n,UB}$ is the unbiased estimator. The same argument as Equation (ref) shows that condition on $\operatorname{sgn}$,

align*[align* omitted — 156 chars of source]

This finishes the proof for the unbiased estimator.

center[center omitted — 46 chars of source]

The analysis will be the same as those for Lemma (ref). For simplicity, denote $\widehat{\mca p}_k = n_k^{-1} \sum_{i \in \mathcal{C}_k} W_i$ and $\mca p_k = \frac{1}{2}\tanh(\beta_k \mca m_k + h_k) + \frac{1}{2} = \frac{1}{2} \mca m_k + \frac{1}{2}$. Then

align*[align* omitted — 292 chars of source]

The analysis in Lemma (ref) implies

align*[align* omitted — 146 chars of source]

and

align*[align* omitted — 303 chars of source]

Hence

align*[align* omitted — 331 chars of source]

The conclusion then follows from step I. The Unbiased Estimator.

Proof of Lemma (ref)

We want to apply Lemma (ref) to the stochastic linearizations obtained from Lemma (ref), $$n_l^{-1} \sum_{i \in \mathcal{C}_l} \mathbf{S}_{l,i,(\mathbf{s})}(W_i - \pi_{l,(\mathbf{s})}),$$ for different blocks separately. First, we need to check if $\mathbf{S}_{l,i,(\mathbf{s})}$ satisfies the covariate constraints in Lemma (ref). Recall $\mathbf{S}_{l,i,(\mathbf{s})} = (S_{1,l,i,(\mathbf{s})}, \cdots, S_{K,l,i,(\mathbf{s})})^{\mathbf{T}}$, where

align*[align* omitted — 193 chars of source]

Definitions of $Q_{i,(\mathbf{s})}$ and $R_{i,l,(\mathbf{s})}$ imply that $\min_{k,l \in [K]}\mathbb{E}[S_{k,l,i,(\mathbf{s})}^2] > 0$ and $\max_{k,l}|S_{k,l,i}| < \infty$ almost surely, satisfying the conditions in Lemma (ref). Hence

align*[align* omitted — 664 chars of source]

where

align*[align* omitted — 164 chars of source]

Now replacing $n_l$ by $n p_l$. The assumption that $n_l/n = p_l + O(n^{-1/2})$ and the Nazarov inequality implies

align[align omitted — 727 chars of source]

The independence between $\mathbf{S}_{l,i,(\mathbf{s})}$ for different $i$'s and the independence between Ising-spins across blocks then imply the stochastic linearization from Lemma (ref) can be approximated by summation of right hand sides of Equation (ref). Lemma (ref) and Nazarov inequality applied on the $n^{-1/2} \boldsymbol{\Sigma}^{1/2} \mathsf{Z}_K$ part then imply the conclusion.