EconBase
← Back to paper

High-Dimensional Tail Index Regression: with An Application to Text Analyses of Viral Posts in Social Media

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

42,424 characters · 9 sections · 23 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

High-Dimensional Tail Index Regression

abstractMotivated by the empirical observation of power-law distributions in the credits (e.g., “likes”) of viral posts in social media, we introduce a high-dimensional tail index regression model and propose methods for estimation and inference of its parameters. First, we propose a regularized estimator, establish its consistency, and derive its convergence rate. Second, we debias the regularized estimator to facilitate inference and prove its asymptotic normality. Simulation studies corroborate our theoretical findings. We apply these methods to the text analysis of viral posts on X (formerly Twitter). { {\ \ \ \newline Keywords: high-dimensional data, social media, tail index, power law, text data analysis.} \newline }

Introduction

A large literature is dedicated to tail features of distributions -- see deHaan07 and Resnick07 for reviews and references. As a common assumption, a distribution $F$ regularly varies with exponent $\alpha$ if its tail is well approximated by a Pareto distribution with shape parameter $\alpha$. This regularity condition implies that common tail features of interest, such as tail probabilities, extreme quantiles, and tail conditional expectations, can be expressed in terms of $\alpha$. Estimates of these features are obtained by plugging in estimated values of $\alpha$. The literature contains numerous suggestions along these lines, some of which are reviewed in the surveys cited below.

figure[figure omitted — 289 chars of source]

Consider the distribution of credits in social media to motivate this framework in contemporary applications. Figure (ref) displays the so-called log-log plot for the distribution of the number $Y$ of “likes” in LGBTQ$+$ posts in X (formerly Twitter). If the distribution of $Y$ is Pareto with exponent $\alpha$, the log-log plot would appear linear, as in this figure, with its slope indicating $-1/\alpha$. This observation motivates us to utilize the aforementioned technology for analyzing viral posts on social media.

We emphasize that the Pareto tail is not unique to our dataset. It has also been documented and explained for numerous economic, finance, and insurance datasets, including city sizes, firm sizes, stock returns, and natural disasters. See gabaix2009,gabaix2016JEP for examples. In text analysis and linguistics, Zipf's law states that, given a large sample of words, the frequency of any word is inversely proportional to its rank in the frequency table. This empirical finding can also be characterized by the Pareto distribution with $\alpha \approx 1$ fagan2010introduction.

Suppose that the conditional distribution of $Y$ given $X$ has an approximately Pareto tail with shape parameter $\alpha(X)$ depending on covariates $X$. We are interested in the effect of $X$ on the tail features of $Y$ through $\alpha(X)$. One family of existing methods imposes a parametric structure on $\alpha(X)$. wang2009tail propose a tail index regression (TIR) method by modeling $\alpha(X) = \exp(X^{\intercal}\theta_0)$ and estimating the pseudo-parameter $\theta_0$. nicolau2023tail extend the TIR method to accommodate weakly dependent data. LiLengYou2020 consider the semiparametric setup $\alpha(X)=\alpha(X_1,X_2) = \exp(X_1^{\intercal}\theta_0 + \eta(X_2))$ for some smooth function $\eta$. By combining $\alpha(X) = \exp(X^{\intercal}\theta_0)$ with a power transformation of $Y$, WangLi2013 study the estimation of conditional extreme quantiles of $Y$ given $X$. Another family of existing methods considers fully nonparametric models and local smoothing GardesGirard2010, GardesArmelleAntoine2012, DaouiaGardesGirardLekina2010, DaouiaGardesGirard2013.

Common to all of these existing approaches is that $X$ is assumed to be of a fixed and low dimension. In the current paper, we relax this restriction by allowing the dimension of $X$ to increase with the sample size and to even exceed the sample size. Our empirical question related to Figure (ref) motivates this high-dimensional model. Specifically, let $Y_i$ denote the number of “likes” of the $i$-th post, and let $X_i$ denote a long vector of binary indicators of whether this post contains a list of keywords. Smaller values of $\alpha(X)$ imply that using the words indicated by such $X$ entails more extreme numbers of “likes.” Essentially, we are asking how to write viral posts. A high-dimensional setup is crucial since the number of keywords is huge.

To address this question, we develop a novel high-dimensional tail index regression (HDTIR) method. By modifying the TIR method wang2009tail, we propose an $\ell_1$-regularized maximum likelihood estimator (MLE). For inference, we further propose debiasing the regularized estimator and establishing its asymptotic normality. Two alternative methods are provided for debiased estimation and inference: one based on sample splitting and the other based on cross-fitting.

To the best of our knowledge, the current paper is the first study of estimation and inference theory for the high-dimensional tail index regression model, which constitutes the key contribution of the current paper. Our estimation and inference problems are related to the extensive literature on high-dimensional generalized linear models van2008high, negahban2009unified, HuangZhang2012, van2014asymptotically, zhang2014confidence, belloni2018uniformly, chernozhukov2018double, cai2023statistical. None of the aforementioned papers focuses on tail index regression.

Our work also contributes to the vast literature of text analysis. $\ell_1$-regularized estimation has been applied to high-dimensional text regressions taddy2013multinomial in economics. Nonetheless, there is no method tailored to analyzing tail features relevant to distributions of credits, such as the number of “likes” for viral posts in social media. Our proposal addresses this gap in the literature as well.

The current paper is also related to the recent literature on shrinkage methods with heavy-tailed data wong2020lasso, fan2021shrinkage, zhu2021taming, babii2023machine. These methods focus on modeling the conditional mean $\mathbb{E}[Y|X]$ using all $n$ observations. The heavy tail feature typically leads to a slower convergence rate, i.e., from $\log{p}$ to a polynomial of $p$, where $p$ denotes the dimension of $X$. The asymptotic distributions become more complicated, as does the subsequent statistical inference. In contrast, our method relies on the regular variation assumption and focuses on the conditional tail index of $Y$. This tail feature requires using only the tail $n_0 < n$ observations, but restores the conventional $\log{p}$ rate. Note that our $p$ may increase with $n_0$, and we leave its relation with $n$ unspecified.

Finally, also closely related are the literature on extremal quantile regressions chernozhukov2017extremal and extremal treatment effects d2018extremal, zhang2018extremal, deuber2024estimation. The tail index plays a crucial role in the estimation and inference of these parameters.

{\bf Organization:} Section (ref) presents the method and theory of the HDTIR. Monte Carlo simulations in Section (ref) demonstrate that the proposed HDTIR has excellent finite-sample performance. We apply the method to text analyses of viral posts in X in Section (ref). Extensions, mathematical proofs, and technical details are relegated to the appendix.

{\bf Notation:} Throughout the paper, we use the following notation. For a $p$-dimensional vector $X=(X_1,\dots,X_p)^{\intercal} \in \mathbb{R}^p$, we use $\|X\|_q=(\sum_{i=1}^{p}|X_i|^q)^{1/q}$ to denote the vector $\ell_q$ norm for $1\leq q<\infty$, and $\|X\|_{\infty}=\max_{1\leq i\leq q}|X_i|$ to denote the vector maximum norm. For a set $S\subseteq\{1,\dots,p\}$, let ${X}_{S}=\{X_{j}:j\in S\}$ and $S^{c}$ be the complement of $S$. For a $p\times q$ matrix $A=(a_{i_1i_2})\in\mathbb{R}^{p\times q}$, we use $\|A\|_{1}=\sum_{i_1=1}^{p}\sum_{i_2=1}^{q}|a_{i_1,i_2}|$, $\|A\|_{2} = \|A\|_F = \{\sum_{i_1=1}^{p}\sum_{i_2=1}^{q}(a_{i_1,i_2})^2\}^{1/2}$ and $\|A\|_{\infty}=\max_{1\leq i_1\leq p,1\leq i_2\leq q}|a_{i_1i_2}|$ to denote the element-wise $\ell_1$, $\ell_2$ and $\ell_{\infty}$ norms, respectively. Let $\|A\|_{\ell_d}=\sup_{X\in\mathbb{R}^{q},|X|_d\leq 1}|AX|_d$ denote the matrix operator norm for $1\leq d\leq \infty$. More specifically, the operator $\ell_1$, $\ell_2$ and $\ell_{\infty}$ norms are denoted by $\|A\|_{\ell_1}=\max_{1\leq i_2\leq q}\sum_{i_1=1}^{p}|a_{i_1i_2}|$, $\|A\|_{\ell_2}=\max_{1\leq i_2\leq q}\{\sum_{i_1=1}^{p}(a_{i_1i_2})^2\}^{1/2}$ and $\|A\|_{\ell_\infty}=\max_{1\leq i_1\leq p}\sum_{i_2=1}^{q}|a_{i_1i_2}|$, respectively. Let $\lambda_{\min}(A)$ and $\lambda_{\max}(A)$ denote the minimum and maximum eigenvalues of the matrix $A$, respectively. Let $I_p$ be the $p\times p$ identity matrix. For two positive sequences $\{a_n\}$ and $\{b_n\}$, $a_n\gtrsim b_n$ means that $a_n>c b_n$ for all $n$ large enough and some constant $c$, $a_n\lesssim b_n$ if $b_n \gtrsim a_n$ holds, and $a_n \asymp b_n$ means that $a_n \lesssim b_n$ and $b_n \lesssim a_n$. Moreover, $a_n \ll b_n$ means $a_n/b_n \rightarrow 0$. For any random variables $X_1,\dots,X_n$ and function $h(\cdot)$, let $\mathbb{E}_n\{h(U_i)\}=\sum_{i=1}^{n}h(U_i)/n$ be the empirical average of $\{h(U_i)\}_{i=1}^{n}$. Let $\dot{h}(\cdot)$ and $\Ddot{h}(\cdot)$ be the first and second-order derivatives of a univariate function, and let $\nabla$ denote the operator for gradient or subgradient.

The High Dimensional Tail Index Regression

Regularized Estimation

Let $\{(X_{i},Y_{i})\}_{i=1}^n$ be $n$ copies of $\{Y,X\}$, where $Y$ is a real-valued response of interest and $X$ is a $p$-dimensional random vector of explanatory factors with possibly $p = p_n \rightarrow \infty$ as the sample size $n$ diverges to $\infty$. We are interested in modeling the effect of $X$ on the tail feature of the distribution of $Y$. Without loss of generality, we focus on the right tail and collect observations such that $Y_{i}>w_n$ for some $w_n$. Let $n_0:=\sum_{i=1}^{n}1 [Y_{i}\geq w_n]$ denote the effective sample size, and rearrange the indices such that $1 [Y_i\ge w_n]=1$ for all $i \in \{1,\cdots,n_0\}$. The following assumptions describe our model.

assumption[Conditional Pareto Tail] (i) $\{Y_i,X_i\}$ is i.i.d. with its distribution satisfying \begin{equation} \mathbb{P}\left( \left. Y>y\right\vert Y>w_n, X=x\right) = \left(\frac{y}{w_n}\right)^{-\alpha \left( X\right) } \end{equation} for all $y>0$ and a sufficiently large $w_n$, where $\alpha \left( X\right):=\exp \left( X^{\intercal }\theta_{0}\right)\geq \underline{\alpha}>0$ uniformly.
assumption[Compact Support] For each $j=1,\dots,p$, $X_{i,j}$ has a compact support $\mathcal{X}_{j}$ with $\sup_{x\in\mathcal{X}_{j}}f_{X_{i,j}|Y_i>w_n}\left( x\right) \leq \bar{f}<\infty$.

Assumption (ref) imposes the restriction that $Y$ conditional on $X$ has a Pareto distribution with exponent $\alpha(X)$ in the tail region $\{y\in\mathbb{R}:y>w_n\}$. Such a Pareto tail has been documented in numerous empirical datasets as emphasized in the introductory section. See gabaix2009,gabaix2016JEP for comprehensive reviews.

This condition could be relaxed by multiplying a slowly varying function $\mathcal{L}(t)$ on the right-hand side of (ref) such that $\mathcal{L}(t)\rightarrow 1$ as $t\rightarrow \infty$ wang2009tail. With this said, there are two benefits of imposing (ref). First, assuming the exact conditional Pareto tail substantially simplifies the theory, especially given that we focus on high-dimensional $X_{i}$. Second, the empirical strategy remains the same when we select a sufficiently large $w_n$ so that the higher-order approximation bias from $\mathcal{L}(t)$ becomes asymptotically negligible. This is also commonly implemented in the literature drees1998estimate, drees1998smooth. More discussions about the effect of $w_n$ can be found in deHaan07, among many others. We could also relax the i.i.d. condition at the cost of more sophisticated theory, but we focus on this sampling assumption to explicate our main contribution concerning the high-dimensional setup.

Assumption (ref) requires each coordinate of $X_i$ to have a compact support and uniformly bounded density. This is coherent with our empirical application, in which $X_i$ is a vector of binary indicators of keywords.

We now introduce our high-dimensional tail index regression (HDTIR) estimators. Define the negative log-likelihood function $\ell_{n_0}$ of $Y$ conditional on $Y\geq w_n$ by

equation[equation omitted — 194 chars of source]

Our regularized HDTIR estimator is given by

equation[equation omitted — 131 chars of source]

We denote the sparsity level of the parameter by $s_0$, i.e., $\|\theta_{0}\|_{0}\leq s_{0}$. Let $\Sigma_{w_n}=\mathbb{E}\left[X_{i}X_{i}^{\intercal}|Y_{i}>w_n\right]$. The following theorem states the convergence rate of this regularized estimator.

theoremSuppose that Assumptions (ref) and (ref) hold, and $\|\theta_{0}\|_{2}\leq C_1$. Suppose that $C_{2}^{-1}\leq\lambda_{\min}(\Sigma_{w_n})\leq\lambda_{\max}\left(\Sigma_{w_n}\right)\leq C_{2}$ holds for constants $C_{1}>0$ and $C_{2}>1$ independent of $n$, $p$, and $w_n$. Let $\lambda_{n_{0}}=c\sqrt{(\log p)/n_{0}}$ for some constant $c>0$. If $s_{0}\lesssim n_{0}/(\log p)$, then with probability approaching one, \begin{equation*} \|\widehat{\theta}-\theta_{0}\|_{1}\lesssim\sqrt{\frac{s_{0}^{2}(\log p)}{n_{0}}}, \quad \|\widehat{\theta}-\theta_{0}\|_{2}\lesssim\sqrt{\frac{s_{0}(\log p)}{n_{0}}}, \quad and\quad \frac{1}{n_{0}}\sum_{i=1}^{n_{0}}\left[X_{i}^{\intercal}\left(\widehat{\theta}-\theta_{0}\right)\right]^{2}\lesssim\frac{s_{0}(\log p)}{n_{0}}. \end{equation*}

Theorem (ref) establishes the convergence rate for the proposed regularized estimator. The condition $C_{2}^{-1}\leq\lambda_{\min}(\Sigma_{w_n})\leq\lambda_{\max}(\Sigma_{w_n})\leq C_{2}$ ensures that the conditional covariance matrix is well-behaved. This result extends previous work on generalized linear models with canonical links negahban2009unified, GardesGirard2010 and those focusing on generalized linear models with binary outcomes van2008high, cai2023statistical.

As extensively discussed in the literature, $\widehat{\theta}$ cannot be directly used to construct a confidence interval for $\theta_0$. In the next subsection, we introduce a debiased estimator to address this issue and facilitate statistical inference.

Debiased Estimation and Inference

The preliminary regularized estimator exhibits a bias that is non-negligible relative to its stochastic variation. Hence, its asymptotic distribution cannot be used directly for inference. In high-dimensional settings, it is standard to construct a debiased (de-sparsified) estimator to remove this bias and restore valid asymptotic inference. Motivated by this, the present section develops a debiased estimation and inference procedure. We adopt a cross-fitting scheme in which each bias-correction step uses a complementary subsample of that used to obtain the initial estimate.

We first present the debiasing step. Let $K$ be a fixed integer greater than 1. Take a $K$-fold random partition $(\mathcal{D}_{k})^{K}_{k=1}$ of the indices $[n_0]=\{1,\dots,n_0\}$ so that the size of each fold $\mathcal{D}_{k}$ is $n_{k}=n_0/K$ for simplicity. For each $k=1,\dots,K$, define the set $\mathcal{D}^{c}_{k}=\{1,\dots,n_0\}\backslash \mathcal{D}_{k}$ of indices in the complement of the fold.

For each $k \in \{1,...,K\}$, we estimate $\widehat{\theta}_{k}$ via (ref) by using the subsample $\mathcal{D}^{c}_{k}$, and estimate $\widehat{u}_{j,k}$ via (ref) by using the subsample $\mathcal{D}_{k}$. Specifically,

align[align omitted — 465 chars of source]

Then, for each $j \in \{1,\dots,p\}$, let

align[align omitted — 267 chars of source]

and define the debiased estimator by taking the average across the $K$ folds:

equation[equation omitted — 115 chars of source]

We emphasize that the construction of $\widehat{u}_{j}$ takes advantage of our Pareto tail approximation. Specifically, the existing literature usually constructs $\widehat{u}_{j}$ using the Hessian, which does not involve $Y_i$. In our case, the score and Hessian of $\ell_{n_{0}}$ evaluated at $\widehat\theta$ take the forms of

align*[align* omitted — 378 chars of source]

respectively. By conditioning on $X_i$ and estimating $\widehat{\theta}$ using a different subsample, the debiasing term maintains conditional zero mean. However, our Hessian involves $\exp(X_{i}^{\intercal}\widehat{\theta}) \log (Y_{i}/w_n )X_{i} X_{i}^{\intercal}$, which contains $Y_i$. To address this issue, we note that the conditional distribution of $\log (Y_{i}/w_n) $ given $X_i$ is approximately exponential with parameter $\exp (X_i^{\intercal}\theta_0)$. It follows that \[ \mathbb{E}\left[ \exp (X_{i}^{\intercal}\widehat{\theta}) \log (Y_{i}/w_n ) \} X_{i}X_{i}^{\intercal} \right] \approx \mathbb{E}\left[ X_{i}X_{i}^{\intercal} \right], \] which motivates our construction of $\widehat{u}_{j}$.

For the debiased estimator (ref), we define the asymptotic variance estimator by

equation[equation omitted — 234 chars of source]

An additional assumption is needed to derive the limiting distribution of the debiased estimator.

assumptionFor some constants $c,c',c''>0$, (i) $\lambda_{n_0}=c\sqrt{(\log p)/n_0}=o(1)$, (ii) $\gamma_{1n_0}=c'\sqrt{(\log p)/n_0}=o(1)$, and (iii) $\gamma_{2n_0}=c''\sqrt{\log n_0}$.

We now state the asymptotic distribution for the debiased estimator $\widetilde{\theta}_{j}$ defined in (ref).

theoremSuppose that Assumptions (ref)-(ref) hold. If $s_{0}\ll\frac{\sqrt{n_{0}}}{(\log p)^{3/2}}$, then \[ \sqrt{n_0}\widehat{V}_{1j}^{-1/2}\left(\widetilde{\theta}_{j}-\theta_{0j}\right)\overset{d}{\rightarrow}\mathcal{N}(0,1) \qquad\text{ as } n_0\rightarrow \infty. \]

One could obtain a similar result in Theorem (ref) without cross fitting (or sample splitting). However, following the insight from cai2023statistical, with the cross-fitting method we propose, the debiased estimator achieves asymptotic normality without requiring the inverse matrix \[ \frac{1}{n_0} \sum_{i=1}^{n_0} \log(Y_{i}/w_n) \exp(X_{i}^{\intercal}\widehat{\theta}) X_{i}X_{i}^{\intercal} \] to be weakly sparse, which relaxes a standard assumption in the literature van2014asymptotically, javanmard2014confidence, zhang2014confidence.

Sample Splitting

Instead of cross fitting, we can alternatively split the samples so that the initial estimation and bias correction steps are conducted on independent datasets. For ease of writing, let the effective sample of size be $n_0$, which is randomly divided into two disjoint subsets $\mathcal{D}_{1}=\{(X_{i},Y_{i})\}_{i=1}^{n_{0}/2}$ and $\mathcal{D}_{2}=\{(X_{i},Y_{i})\}_{i=n_0/2+1}^{n_{0}}$. We use the subsample $\mathcal{D}_{2}$ to obtain $\widehat{\theta}$ via (ref) and use the subsample $\mathcal{D}_{1}$ for the debiasing step described below. Let

align[align omitted — 404 chars of source]

where $\{e_{j}\}_{j=1}^{p}$ denotes the canonical basis of the Euclidean space $\mathbb{R}^{p}$, and $\gamma_{1n_{0}}$ and $\gamma_{2n_{0}}$ satisfy the conditions stated in Assumption (ref).

For each coordinate $j=1,\dots,p$, the debiased estimator is defined by \[ \widetilde{\theta}_{j}:=\widehat{\theta}_{j}-\frac{\widehat{u}_{j}^{\intercal}}{n_{0}/2}\sum_{i=1}^{n_0/2}\left\{ \exp\left(X_{i}^{\intercal}\widehat{\theta}\right)\log\left(Y_{i}/w_n\right)-1\right\} X_{i}, \] where $\widehat{u}_{j}\in\mathbb{R}^{p}$ is the projection direction constructed by (ref)--(ref) using the subsample $\mathcal{D}_{1}$, while $\widehat\theta$ derives from (ref) using the subsample $\mathcal{D}_{2}$. Let \[ \widehat{V}_{2j}:=\widehat{u}_{j}^{\intercal}\left(\frac{1}{n_{0}/2}\sum_{i=1}^{n_0/2} X_iX^{\intercal}_i\right)\widehat{u}_{j}. \] The following corollary states the asymptotic distribution of this estimator.

corollarySuppose that Assumptions (ref)-(ref) hold. If $s_{0}\ll\frac{\sqrt{n_{0}/2}}{(\log p)^{3/2}}$, then \[ \sqrt{n_0/2}\widehat{V}_{2j}^{-1/2}(\widetilde{\theta}_{j}-\theta_{0j})\overset{d}{\rightarrow}\mathcal{N}(0,1) \qquad\text{ as } n_0\rightarrow \infty. \]

Extensions

While our proposed method is based on the maximum likelihood principle, an alternative approach based on least squares is also possible. We develop this alternative methodology in Appendix (ref).

Our framework can be further extended to accommodate large-scale online data, which is particularly relevant in settings such as social media applications where data are generated continuously and at scale. Appendix (ref) presents this extension, where we employ a variant of stochastic gradient descent to efficiently process streaming data, thereby ensuring that the proposed methods remain scalable and practical for real-world applications.

Simulation Studies

In this section, we use simulated data to numerically evaluate the performance of our proposed method of estimation and inference. Two designs for the $p$-dimensional parameter vector $\theta_0$ are employed:

align*[align* omitted — 217 chars of source]

A random sample of $(Y_i,X_i^\intercal)^\intercal$ is generated as follows. Three designs for the $p$-dimensional covariate vector $X_i$ are employed:

align*[align* omitted — 247 chars of source]

where $I_p$ denotes the $p \times p$ identity matrix. In turn, generate the exponents by $$ \alpha_i = \exp(X_i^\intercal \theta_0) $$ and then generate $Y_i$ by $$ Y_i = \Lambda^{-1}(U_i; \alpha_i), \qquad U_i \sim \text{Uniform}(0,1), $$ where $\Lambda(\ \cdot\ ; \alpha)$ denotes the CDF of the Pareto distribution with the unit scale and exponent $\alpha$.

In each iteration, we draw a random sample $(Y_i,X_i^\intercal)^\intercal$ of size $n=$ 10,000. Setting the cutoff $w_n$ to the 95-th empirical percentile of $\{Y_i\}_{i=1}^n$, we have the effective sample size of $n_0=$ 500 from five percent of $n$. We vary the dimension $p \in \{250,500,1000\}$ of the parameter vector $\theta_0$ across sets of simulations. While there are $p$ coordinates in $\theta_0$, we focus on the first coordinate $\theta_{01} = 1.0$ for evaluating our method of estimation and inference. Throughout, we use $K = 5$ for the number of subsamples in sample splitting. The other tuning parameters are set according to Assumption (ref) where $c=1$, $c'=1$ and $c''=100$. We run 10,000 Monte Carlo iterations for each design.

Table (ref) summarizes the simulation results.

table[table omitted — 2,016 chars of source]

The sets of results vary with the effective sample size $n_0$, the dimension $p$ of the parameter vector $\theta$, the design for the parameter vector $\theta$, and the design for the covariate vector $X$. For each row, displayed are the bias (Bias) of the debiased estimator $\widetilde\theta_1$, standard deviations (SD), root mean square errors (RMSE), and the coverage frequencies by the 95% confidence interval (95%).

For each set, the bias is much smaller than the standard deviations and hence the 95% confidence interval delivers accurate coverage frequencies. We ran many other sets of simulations with different values of $n_0$ and $p$ as well as parameter designs and data-generating designs, and confirmed that the simulation results turned out to be similar in qualitative patterns to those presented here. The additional results are omitted from the paper to avoid repetitive exposition.

Application: Text Analysis of Viral Posts on LGBTQ+

In this section, we apply our proposed method to analyze LGBTQ$+$-related posts on X (formerly Twitter). Our goal is to infer the impact of specific words on attracting “likes” for these posts. The dataset comprises tweets containing the keyword “LGBT,” scraped from Twitter between August 21 and August 26, 2022.\footnote{The data set is publicly available at https://www.kaggle.com/datasets/vencerlanz09/lgbt-tweets.}

Each observation in our study represents a single post. Our sample includes a total of $n = 32,456$ posts. The data records the number of likes, $Y_i$, that the $i$-th post has received. As we will demonstrate below, $Y_i$ follows a heavy-tailed distribution: most posts attract a small number of likes, while a few viral posts garner a large number of likes. We construct a word bank consisting of 936,556 unique words used across the $n = 32,456$ posts in our sample. The $j$-th coordinate, $X_{ij}$, of the covariate vector $X_i$ takes a value of 1 if the $j$-th word in the word bank is used in the $i$-th post and 0 otherwise. From these 936,556 unique words, we only include the 500 most frequently used words to create the binary indicators in $X_i$. Therefore, the dimension $p$ of $X_i$ is 500. This list explicitly excludes articles, auxiliary verbs, and prepositions.

Figure (ref) in the introduction presents the log-log plot of the empirical distribution $\{Y_i\}_{i=1}^n$ for posts with a positive number of likes. We focus on posts with positive likes because the logarithm of zero is undefined. The horizontal axis represents the rank of $Y$, while the vertical axis represents $\log(Y)$. The approximate linearity of this log-log plot suggests that the distribution of $Y$ follows a power law, indicating that $Y$ is characterized by a Pareto distribution.

Table (ref) displays the 30 most frequently used words in LGBTQ$+$ posts. For each word, the total number of times it appeared (Count) and the total number of posts in which it appeared (Tweets) are shown. The last value corresponds to $\sum_{i=1}^n X_{ij}$ for $j=1,\dots,30$. All characters have been converted to lowercase to ensure the counting is not case-sensitive. Notice that the most frequent word, “lgbt,” is distinct from the eleventh most frequent word, “\#lgbt.” The former is a plain word, while the latter functions as a hashtag, serving to link posts with others containing the same hashtag.

table[table omitted — 1,438 chars of source]

We apply our proposed method of estimation and inference to analyze the effects of using these and other words on the tail shape of the distribution of the number of likes. Consistent with our simulation studies, we set $w_n$ to the 95th percentile of the empirical distribution of $\{Y_i\}_{i=1}^n$, which results in an effective sample size of $n_0 = 1,623$. The rules for selecting the tuning parameters remain the same as those used in our simulation studies.

Table (ref) presents the estimates, standard errors, 95% confidence intervals, and t-statistics for $\theta_j$ for the 30 most frequently used words, listed in the same order as in Table (ref). Notably, the most frequent word, “lgbt,” has a significantly negative coefficient, while the eleventh most frequent word, “\#lgbt,” has a significantly positive coefficient. Recall that smaller values of the Pareto exponent correspond to more extreme values of $Y_i$. Therefore, this finding suggests that using the plain word “lgbt” tends to attract a substantially larger number of likes, whereas using the hashtag “\#lgbt” may have the opposite effect. Most of the other words in Table (ref) are statistically insignificant, with the exceptions of “they” and “it's,” whose positive coefficients indicate their adverse effects.

table[table omitted — 2,122 chars of source]

We then identify the 10 most effective words and the 10 least effective words from the list of $p = 500$ words. Table (ref) presents the estimates, standard errors, 95% confidence intervals, and t-statistics for $\theta_j$ for these words. The words are sorted in descending order based on the absolute value of the estimate $\widetilde\theta_j$. Again, we find that the plain word “lgbt” is the only significantly effective word. In contrast, hashtags containing this effective keyword, such as “\#lgbtqia” and “\#lgbtq,” tend to have negative contributions to attracting likes.

table[table omitted — 1,609 chars of source]

Summary

This paper introduces a novel high-dimensional tail index regression (HDTIR) model inspired by observing power-law distributions in social media posts, particularly in the distribution of “likes” on viral content. We tackle the challenges of estimating and inferring the parameters of the tail index model when the dimension of the explanatory variables increases and may exceed the sample size.

We begin by developing a regularized estimation method for the HDTIR model, demonstrating its consistency and establishing its convergence rate. To facilitate inference, we introduce a debiasing technique that corrects the bias introduced by regularization. This allows us to derive the asymptotic normality of the debiased estimator, providing a robust framework for statistical inference in high-dimensional settings.

Extensive simulation studies validate the theoretical properties of our model, showing strong performance even in finite samples. In addition, we apply the HDTIR method to a dataset of viral posts on X (formerly Twitter) related to LGBTQ$+$ topics. This empirical analysis reveals insights into how specific words influence the likelihood of a post going viral, with terms like `lgbt' playing a significant role while hashtags like `\#lgbtq' do not. The results demonstrate the practical utility of the HDTIR model in understanding and predicting the factors that drive the popularity of online content.

Extensions are also provided in the appendix. While our proposed method is based on the maximum likelihood principle, an alternative approach based on least squares is also possible. We develop this alternative methodology in Appendix (ref). Our framework can be further extended to accommodate large-scale online data, which is particularly relevant in settings such as social media applications where data are generated continuously and at scale. Appendix (ref) presents this extension, where we employ a variant of stochastic gradient descent to efficiently process streaming data, thereby ensuring that the proposed methods remain scalable and practical for real-world applications.