Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
129,553 characters · 21 sections · 53 citation commands
Single-Network Finite-Sample Inference in Strategic Network Formation Models
\global\long
The formation of social and economic networks, such as friendships, information sharing, trade relationships, R&D alliances, and interbank lending, is shaped not only by the characteristics of the agents involved but also by the agents' strategic considerations over other network links. A link between two agents might be more attractive when they share common friends or when it connects to well-positioned partners, for various reasons related to information, enforcement, and coordination. Estimating structural models that take this interdependence seriously is an important problem in the econometrics of network formation de2020econometric.
Taking that interdependence seriously, however, creates at least three layers of challenges that compound each other. The first is the equilibrium itself. We adopt pairwise stability jackson1996strategic as the solution concept, which is the leading equilibrium notion in the economics and econometrics literature on strategic network formation. The mapping from primitives to realized pairwise stable networks is intractable to characterize: the outcome space is combinatorially vast,\footnote{For a standard illustration: there are $2^{435}$ possible undirected unweighted networks on $30$ agents.} pairwise stability generically admits many equilibria at once de2020econometric, and iterative procedures are not generally guaranteed to reach one even when it exists jackson2002evolution. There is thus no tractable reduced form to invert. The second is the dependence this induces: each dyad's endogenous statistic is a function of the realized network, so every link can implicitly depend on every shock, rendering the dependence structure global and vastly complicated. On a single large network, the network dependence issue renders standard statistics tool such as Law of Large Numbers, Central Limit Theorem, bootstrap and subsampling highly nontrivial, whose validity often relies on complicated and hard-to-verify weak dependence conditions that equilibrium interdependence may destroy. The third is computation: procedures built on stability conditions must simulate shocks, search over completions of the observed network, and enumerate subnetwork configurations up to isomorphism. The endogenous covariate aggravates all three at once, since it must be recomputed at every candidate parameter and every simulated network, placing an equilibrium solver inside the inference loop rather than beside it.
This paper, to the best of our knowledge, is the first in the network econometrics literature to provide a finite-sample valid inference method on the structural parameters in a strategic network formation model under a single network setting, together with the identification analysis that underlies it. The model we consider is the canonical one, where a link between agents $i$ and $j$ forms if and only if the latent surplus exceeds a threshold: \[ Y_{ij}=\mathbf{1}\left\{ Z_{ij}'\beta_{0}+X_{ij}'\gamma_{0}\geq\varepsilon_{ij}\right\} ,\quad\text{for any }ij \] where $Z_{ij}$ collects exogenous observed dyadic covariates built from the two agents' characteristics, $\varepsilon_{ij}$ is an idiosyncratic shock independent of them, and $X_{ij}:=\phi_{ij}\left(Y,Z\right)$ collects endogenous observed network statistics (e.g. the number of common friends of $i$ and $j$, the number of 2nd-order friends), which can be determined jointly by the equilibrium network $Y$ (as well as the exogenous covariates $Z$). We assume that the observed network $Y$ is pairwise stable under transferable utilities, or in other words, is a solution to the model equation above, while we leave entirely unrestricted the equilibrium selection mechanism. The model above corresponds to the transferable utility setting of sheng2020structural and, with nonnegative externalities, of miyauchi2016structural. We note that our strategy is not confined to the transferable utility setting: Remark (ref) explains how our key result extends to non-transferable utility settings. The structural parameter is $\theta_{0}=(\beta_{0},\gamma_{0})$, with $\gamma_{0}$ being the key object of interest: it measures the strength of the strategic interdependence in link formation, and the null $\gamma_{0}=0$ is the hypothesis that there is none.\footnote{That null has itself been the object of a dedicated testing literature. grahampelican2020testing,pelican2022optimal construct exact tests of no strategic interaction by conditioning on the sufficient statistics the model admits under the null, where it collapses to an exponential-family dyadic model. Such tests are sharp against the no-interdependence null, but they do not invert into confidence sets for $\gamma_{0}$. Our confidence sets deliver the no-interdependence test as a by-product.}
The key device that enables us to handle the intractable network endogeneity is a technique we call bounding by $c$.\footnote{This technique was proposed and adopted in gao2026identification, who studies sharp identification in dynamic binary choice models. That setting is very different from the strategic network formation model considered here. gao2026identification incorporates individual-level fixed effects in a panel data setting without cross-sectional strategic interdependence, while our paper considers a one-shot strategic network formation game without fixed effects.} A formed link certifies that its idiosyncratic shock lies below its latent index, $\varepsilon_{ij}\leq Z_{ij}'\beta_{0}+X_{ij}'\gamma_{0}$, and an absent link certifies the reverse. The certificate is not directly usable, because the index contains the endogenous $X_{ij}$. However, on the observable (random) event that the index does not exceed a deterministic scanning threshold, $Z_{ij}'\beta_{0}+X_{ij}'\gamma_{0}\leq c$, transitivity of the inequalities implies $\varepsilon_{ij}\leq c$, an event involving only the exogenous shock and the number $c$. Averaging such “bounding-by-$c$” certificates within a cell of exogenous characteristics therefore sandwiches the empirical distribution of that cell's exogenous shocks between two observable and endogenous envelopes, at every $n$ and in every realization, before any expectation or limit is taken. Importantly, the realization-wise nature of the inequalities implies their validity at every possible equilibrium regardless of the equilibrium selection mechanism, with the conditional distribution of the endogenous covariates $X_{ij}$ never modeled, estimated, or simulated. Identification and inference are then two ways of using the same realization-wise \emph{bounding-by-$c$} inequalities obtained: taking limits along the realized network sequence gives identifying restrictions, while characterizing the finite-sample uncertainty of an empirical average of the exogenous shocks $\epsilon$ and exogenous covariates $Z$ yields the valid non-asymptotic inference.
Specifically, for identification, we take finite-sample averages of the bounding-by-$c$ inequalities, and consider the limit restrictions arising from a realized sequence of networks as $n\to\infty$. Since we do not impose any sparsity and dependence restrictions, joint probabilities of observable events that involve the endogenous covariates $X$ are not guaranteed to converge in the single large-network limit. A key novelty of our identification result lies in that the validity of our identifying restriction only rests on the large-$n$ convergence of an exogenous statistic involving $\varepsilon$ and $Z$ only, and thus does not require convergence of probabilities that involve the endogenous $X$. Consequently, our large-network identifying restrictions are stated in terms of pathwise limit frequencies: the limit superior of the lower envelope and the limit inferior of the upper envelope, each of which involves the endogenous $X$. This departs from the standard identification template based on (conditional) expectations or probabilities, in a way we think is of independent theoretical interest. No observable quantity is required to settle down to a deterministic population limit, and along a single network sequence the observable frequencies may fluctuate indefinitely and their limit envelopes may remain genuinely random, capturing lack of determination from equilibrium selection or other forms of strong global dependence.
For finite-sample inference, the paper's central contribution, we exploit use the same sandwich inequalities with $n$ held fixed. Again, due to the realization-wise nature of our bounding-by-$c$ inequalities, at the true $\theta_{0}$, the observable criterion is bounded, realization by realization, by a dominance statistic defined on the exogenous shocks $\varepsilon$ and the exogenous covariates $Z$ alone. Importantly, this statistic involves no endogenous covariate $X$ at all, and thus its distribution is not contaminated by the interdependence generated by the equilibrium network. We provide both semiparametric and parametric inference procedures: our key idea, in either case, is to demonstrate the finite-sample uncertainty of our constructed test statistics can be bounded by a control statistics, whose exact finite-sample distribution, and thus the corresponding critical values, can be simulated. We also show how to get sharper tests (and critical values) through constrained thresholding, studentization, and more sophisciated tail weighting schemes that dampen the effect of restrictions with low effective sample sizes and high uncertainty.
We demonstrate in a simulation study the performance of our single-network confidence sets, with network sizses ranging from $100$ to $10,000$. Our parametric confidence sets certifies the sign of the strategic coefficient $\gamma_{0}$ in every replication from $n=400$ onward. We also show how two semiparametric versions (one with and one without a symmetry assumption on $F_{\varepsilon}$) of our inference procedures still produce nontrivial sign-revealing confidence sets at reasonable sample sizes.
We also apply our inference procedure to two empirical data sets: one on high-school contact networks ($n=327$) and another on Twitch gamer networks ($n$ up to $9,498$). In both settings, we find strictly positive (projected) confidence intervals for the strategic coefficients at $95\%$-level.
\
Our paper contributes mostly directly to the econometric literature on strategic network formation, with the following two highlights:
The first highlight is pathwise identification in a single large network setting, where the formulation is as novel as the restrictions. Identification analysis normally starts from a deterministic feature of the observable distribution: a population probability, moment, or probability limit that the data can recover and the model restricts. Here no such objects need exist, and thus the valid identified set can still be random even in the limit.\footnote{The randomness at issue here is not that of a confidence set. A confidence set is random in any finite sample simply because it is a function of the data; that randomness is standard, well understood, and disappears in the usual asymptotics as the set settles down around a population object. What is unusual here is that randomness may survive the large-$n$ limit: absent weak-dependence conditions the observable envelopes need not converge along the realized network sequence, so the limiting identified set is itself a random object rather than a feature of the population distribution.} What becomes deterministic in our setting is the limit of an exogenous statistic, which encodes the exogeneity assumption in the model as the identifying leverage. We call the resulting notion pathwise identification, and we know of no precedent for it in the network formation literature.
The second highlight is finite-sample inference, which we regard as the central contribution of this paper to the network econometrics literature: to our knowledge these are the first computationally and theoretically tractable, finite-sample valid confidence sets for the structural parameters of a strategic network formation game observed as a single large network under unknown equilibrium selection, with no restrictions on density or on the dependence the game induces, and with no equilibrium simulation. The tractability is structural rather than incidental. Testing a candidate parameter requires no equilibrium solve, no search over completions of the observed network, and---because our restrictions are indexed by dyads rather than by subnetwork configurations---no graph isomorphism matching. What the criterion needs are empirical averages of indicator variables over the dyads falling in cells of exogenous characteristics, which is counting. The computational cost therefore roughly scales with the number of dyads, $n(n-1)/2=O(n^{2})$, rather than with the size $2^{n(n-1)/2}$ of the graph space. The practical consequence is that network sizes out of reach for simulation-based methods are routine here: the simulation study runs to $n=10{,}000$ and the empirical application to $n=9{,}498$, each returning a certified confidence set in manageable time.
The closest papers in the strategic network formation literature are \citet*{de2018identifying} and menzel2026strategic. \citet*{de2018identifying} study partial identification of preferences in a single large network formed under pairwise stability, with non-transferable utility and complete information. They impose two restrictions to make the payoff-relevant environment finite: utility depends on the network only out to a bounded distance, and all agents have a bounded number of links. Each agent is then classified by her network type ( defined jointly over the graph of her neighborhood and the covariates of the nodes in it), and parameters are constrained by the observed distribution of network types. Their computation is carried out through a series of quadratic programs in a continuum-of-agents formulation rather than at a finite $n$, and matching network types requires deciding joint covariate-and-graph isomorphism, which is computationally hard in large finite networks. We give up sharpness and the local combinatorics, using only the threshold-crossing structure of a single dyad. In exchange we impose no bound on degree, none on interaction depth, and no restriction on density. menzel2026strategic, who studies a similar model in a single large network, also relies on dyad-level link frequencies to deliver identification and estimation. The first difference is that, while endogenous covariates are allowed in menzel2026strategic, they must take a restrictive “node-incident” form that rules out “fundamentally dyadic” network statistics such as common friends. The methodology is also very different from ours: he derives large-network asymptotic approximations, obtaining a Poisson likelihood limiting model under a Type I upper tail condition on the idiosyncratic error and unique edge response condition as well as equilibrium selection. The results are also asymptotic approximations, where ours are exact at the observed finite $n$.
Some related work consider similar strategic network formation models, but focus on the many-network setting for identification and inference. sheng2020structural studies the same model with a complementary strategy in a many network settting, who derive identifying restrictions by checking pairwise stablility conditions on subnetwork structures and resolve multiple equilibria using the 0-1 bounds a la \citet*{tamer2003incomplete}. Her approach also requires a network isomorphism algorithm that selects matches on specific subnetwork configurations where the bounds are tractable to compute, and the inference theory is built for many independent (and typically small) networks, with asymptotics established in the number of independent networks. gualdani2021identification also works with many observed networks, in a complete-information game with directed links and a spillover in the number of agents linking to the same target. Her device is a decomposition of the formation game into local games that are in equilibrium if and only if the whole game is, which cuts the number of moment inequalities characterizing the sharp identified set and makes the integrals entering them computable.
There are also other routes that resolve the equilibrium multiplicity issue rather than bracketing it. miyauchi2016structural exploits nonnegative externalities and the resulting lattice structure to bracket moments between extremal equilibria, computed by simulation. mele2017dense and badev2021nash model equilibrium selection directly, treating the observed network as a draw from the stationary distribution of a myopic adjustment process, so that estimation on a single network satisfies mixing requirements. A separate branch weakens the informational environment instead: with private information, realized links are conditionally independent given types, which severs the global dependence and permits two-step asymptotically normal estimation on one network leung2015twostep,ridder2025twostep,comola2026estimating,wu2024twostep. Hence, the model, methodology, and results obtained are quite different from ours.
In a companion paper, we further combine bounding-by-$c$ with subnetwork differencing to accommodate strategic interdependence and unobserved individual fixed effects \citep*{gao2026tractable}, but the inclusion and differencing of fixed effects imply that no node-level covariates can be included.
\paragraph{\ }
Our paper also relates to the larger network formation literature. Complementary to the strategic network formation literature where our paper belongs, the dyadic network formation literature studies models with agent unobserved heterogeneity but no link interdependence: see, e.g., graham2017econometric,jochmans2017semiparametric,dzemski2019empirical,gao2020,zeleneev2026identification, as well as work that focuses on the nontransferable-utility case gao2023logical,li2026bagging,marshall2026utility. That conditional independence is exactly what makes those models tractable, and it is exactly what link interdependence invalidates: our whole difficulty is that $X_{ij}$ is a function of the entire realized network.
There have also been theoretical advances on inference in a single network setting, which mostly focuses on deriving valid asymptotic results under suitable conditions on network dependence: $\psi$-dependence decaying in network distance kojevnikov2021limit, approximate neighborhood interference leung2022causal, exchangeable-array structure davezies2021empirical,graham2020network, or a density regime that bounds the strength of interdependence leung2019treatment,menzel2026strategic,chandrasekhar2025network. The most recent advance that delivers a central limit theorem under a network formation model with explicit strategic interdependence is leung2026normal, which requires branching-process subcriticality of the spillovers and a sufficiently decentralized selection mechanism. We impose no counterpart to any of these and provide finite-sample inference results instead of asymptotic ones. That said, the corresponding cost is that we obtain coverage rather than an asymptotic distribution, and hence conservatism. The two approaches are therefore complementary rather than competing. Where the conditions of leung2026normal do hold, their central limit theorem could in principle be used to calibrate critical values to the actual sampling distribution of our dyad-level statistics, buying sharpness at the informative margin in exchange for the assumption-free validity we insist on here. We leave that combination to future work.
More generally, our paper contributes to the literature on incomplete models and partially identifying inequalities. Our restrictions are moment inequalities, and their population treatment is the standard one for incomplete models: bracket outcome probabilities between some- and every-equilibrium envelopes tamer2003incomplete,beresteanu2011sharp,galichon2011set or model selection directly bajari2010identification, as surveyed by aradillas2020econometrics. What differs here is that we never evaluate nor invert the equilibrium correspondence, so we make no claim to sharpness, which is plausibly intractable in a single large network.\footnote{See more discussion about sharpness after Theorem (ref)} The inference we build is not the asymptotic inequality inference of andrews2013inference and chernozhukov2013intersection, which would require exactly the dependence theory we avoid. Methodologically, the “bounding-by-$c$” device descends from the partial stationarity approach of gao2026identification for nonlinear dynamic panel models, and the aggregation of inequalities by logical and monotonicity operations relates to \citet*{gao2023logical}.
Lastly, our finite-sample inference procedure relates to finite-sample randomization and Monte Carlo tests besag1989generalized. What has kept this tradition out of strategic network formation is that the null distribution of any statistic touching $X_{ij}$ depends on the intractable equilibrium map and the unknown selection mechanism, so it is neither simulable nor known by symmetry. This is solved by the bounding-by-$c$ technique, after which the finite-sample inference validity follows with technical ease. In econometrics, the closest work is rosen2025finite, who obtains finite-sample conditional inference with simulable critical values. However, their model setup is the semiparametric binary choice model underlying the maximum score estimator manski1975maximum,manski1985semiparametric,manski1987semiparametric, which is a cross-sectionally i.i.d. single-agent model, with no strategic equilibrium, multiplicity or dependence issues. In network econometrics, exact inference mostly takes the form of sharp-null tests, e.g., grahampelican2020testing,pelican2022optimal, the two-sample test of auerbach2022testing, and design-based causal inference with network interference athey2018exact,puelz2022graph. A recent paper by kasy2026causal also obtains exact finite-sample permutation tests and conservative confidence intervals a network formation setting. However, \citet*{kasy2026causal} focuses on a causal inference framework with known randomization of the initial network, and requires observation of the network structure over a panel of at least two periods; in addition their object of interest is on causal effects induced by an experiment rather than structural network formation parameters. Also related are kaido2025universal and li2026finite, who obtain finite-sample valid, selection-robust confidence sets for incomplete models using many independent observations with parametric latent variables \citep*[see also][]{epstein2016robust}. A very recent paper from the statistics literature, yanchenko2026universal, provides finite-sample model selection test on statistical network formation models (such as random graph and stochastic block/communty models) that includes neither exogenous or endogenous covariates. Our strategic network formation setting is thus out of the scope of all these papers.
\
The remainder of the paper is organized as follows. Section (ref) presents the model setup, the key realization-wise bounding-by-$c$ inequalities, and the identification analysis. Section (ref) develops the finite-sample inference under both parametric and semiparametric settings. Section (ref) reports the simulation study and Section (ref) the empirical application. All proofs are available in Appendix (ref).
Consider a set of $n$ agents indexed by $i=1,\ldots,n$, and an unweighted network among them represented by an $n\times n$ adjacency matrix $Y$, where $Y_{ij}=1$ if a link exists between agents $i$ and $j$, and $Y_{ij}=0$ otherwise. We focus on undirected networks, i.e., $Y_{ij}=Y_{ji}$ for all $i,j$. We index distinct unordered pairs as $ij$ with $i<j$, and refer to each such pair $ij$ as a dyad. Throughout the paper every sum $\sum_{ij}$ over dyads is understood to be $\sum_{i<j}$, so that each dyad is counted exactly once.
We consider the following network formation model with link interdependence, where a link between agents $i$ and $j$ exists if and only if the latent surplus from the link is non-negative:
where:
We allow the covariate function $w_{n}$ to depend on the network size $n$, for example, through the rescaling of latent positions leung2019treatment, to accommodate sparse designs in large networks.
The structural parameters of interest are $\theta_{0}:=\left(\beta_{0},\gamma_{0}\right)\in\Theta$. Throughout this paper, we use the shorthand notation
so that the link formation equation becomes $Y_{ij}=\mathbf{1}\{\varepsilon_{ij}\leq\delta_{ij}\}$.
The exogenous dyadic covariates $Z_{ij}$ can include a constant $1$ (under a proper location normalization in the distribution of $\epsilon$), homophily effects in the form of component-wise absolute difference $|Z_{i}-Z_{j}|$ for continuous variables or $\mathbf{1}\{Z_{i}=Z_{j}\}$ for discrete variables, as well as level effects in the form of $Z_{i}+Z_{j}$.\footnote{The ability of including level effects of the form $Z_{i}+Z_{j}$ is one of the key flexibility in network formation models without unobserved individual fixed effects: in the dyadic literature with additive degree heterogeneity graham2017econometric,dzemski2019empirical,gao2020, the level component is collinear with the fixed-effect profile and is annihilated by every device that eliminates it, so only difference-type effects can be included. }
A key feature of our model is the presence of the endogenous covariates $X_{ij}$, which may arise as functions of the realized network $Y$. Specifically, we allow:
where $\phi_{ij}$ is a known function mapping the $ij$-excluded\footnote{This restriction is only required to endow model $(\ref{eq:link_model})$ with the economic interpretation of “pairwise stability”, which is formulated on the utility difference of a counterfactual comparison between linking and not linking. Note that $\phi_{ij}$ above can be defined as \[ \phi_{ij}\left(Y_{-ij},Z\right)=\tilde{\phi}_{ij}\left(\left(Y_{-ij},1\right),Z\right)-\tilde{\phi}_{ij}\left(\left(Y_{-ij},0\right),Z\right), \] where $\left(Y_{-ij},y_{ij}\right)$ denotes the two networks under the two possible own-link states $y_{ij}=1$ and $y_{ij}=0$. Here, $\tilde{\phi}_{ij}$ is a function that can depend on the whole network structure, including own link $ij$: for example, the eigenvetor centrality of $ij$. That said, even if $X_{ij}$ does depend on the realized $Y_{ij}$, all the econometric inference results in our paper still apply (under the subsequently stated assumptions), with the only caveat being that we should no longer interpret (ref) as a literal pairwise stability condition.} realized network $Y_{-ij}$ and exogenous characteristics $Z=(Z_{1},\ldots,Z_{n})$ to a vector of dyadic covariates. We allow $\phi_{ij}$ to be both locally and globally defined (subject to the subtlety in Footnote (ref)).
One canonical local network statistic is the number of common friends of $i$ and $j$, together with its normalized version,
capturing transitivity: the surplus of link $ij$ may increase with the number of shared neighbors. The normalized version is the friends-in-common statistic of sheng2020structural, and a nonnegative coefficient on it delivers the nonnegative externalities exploited by miyauchi2016structural. It is the statistic we focus on in simulation and the empirical application.
Other leading choices of local network statistics are accommodated without modification: for example, the Jaccard overlap index $J_{ij}:=\frac{\text{CF}_{ij}}{{\displaystyle \sum_{k\neq i,j}\mathbf{1}\{Y_{ik}+Y_{jk}\geq1\}}}$ (with $J_{ij}:=0$ when the denominator vanishes), which measures overlap relative to the union of the two neighborhoods and is a nonlinear, non-monotone functional of the network; second-order friends in the form of $\overline{\text{SF}}_{ij}:=\frac{1}{n-2}\sum_{k\neq i,j}\left(Y_{ik}+Y_{jk}\right)\in[0,2],$ and interactions with exogenous types, such as $\overline{\text{CF}}_{ij}\cdot\mathbf{1}\{Z_{i}=Z_{j}\}$, which let the strength of transitivity vary with observables.
Notably, our model also accommodates globally defined $X_{ij}$, i.e., those that cannot pinned down completely by the local network structure around $ij$. Leading examples include various globally defined centrality measures, such as eigenvector centrality and closeness centrality (subject to the subtlety in Footnote (ref)). Such globally defined $X_{ij}$ are usually highly nonlinear and implicit aggregation of links from the whole network $Y$, for which there exist no closed-form representations.
\
Since $X_{ij}$ may depend on the realized network $Y$, while each link in $Y$ is determined by equation (ref) simultaneously, model $(\ref{eq:link_model})$ can be interpreted as the equilibrium network in a strategic network formation game under transferable utilities with pairwise stability as the equilibrium solution concept jackson1996strategic,sheng2020structural. In general, the game may have multiple equilibria for a given realization of primitives $(Z,\varepsilon)$, where $\varepsilon=(\varepsilon_{ij})_{i<j}$ collects all the idiosyncratic shocks. We represent the realized network abstractly as:
where $g$ incorporates both the equilibrium correspondence and an (arbitrary and possibly unknown) equilibrium selection mechanism that picks a single equilibrium for each realization of the primitives, potentially depending on an unrestricted random element $\xi$ that drives equilibrium selection.
The equilibrium mapping $g$ is generally intractable to characterize. Even for simple specifications of $\phi_{ij}$, the fixed-point nature of the equilibrium creates a complex interdependence structure, which is exacerbated by the discrete combinatorial nature of the graph space. A globally defined $\phi_{ij}$ would induce yet another layer of complexity and global dependence. As will be shown below, our identification approach does not require characterization, evaluation, or computation of $g$: we extract identifying information using monotonicity restrictions that hold regardless of the complexity of the equilibrium correspondence and any equilibrium selection mechanisms.
We maintain the following assumptions throughout.
Assumptions (ref)-(ref) are standard in the strategic network formation literature. Assumption (ref) essentially assumes that a pairwise stable network exists and the observed network is pairwise stable. Assumption (ref) asserts random sampling of exogenous covariates across individuals. Assumption (ref) imposes the idiosyncrasy of link-level surplus shocks but leaves their common distribution $F_{\varepsilon}$ unrestricted for now: our key identifying strategy applies to both parametric and semiparametric settings, and we will introduce a parametric assumption on $F_{\varepsilon}$ when it is explicitly required. Assumption (ref) is an exogeneity condition on the observables $Z$ and the pairwise error $\varepsilon$. Importantly, we do not assume that $X$ and $\varepsilon$ are independent: since $X$ is a function of the equilibrium network it is endogenous and generally correlated with the shocks $\varepsilon$. This assumption also formalizes the naming convention of calling $Z$ the exogenous covariates while calling $X$ the endogenous covariates.
Assumption (ref) also stresses that the data environment we consider is a single network observed once. The alternative environment of many independent networks (villages, schools, classrooms, markets) is simpler in the sense that conditional probabilities of observable events given the agents' exogenous characteristics are population moments that are directly identified and consistently estimable across networks with no restriction on within-network dependence. Our restrictions also apply to that case, and see Remark (ref) for more discussion. That said, the single network environment is what our procedures are mainly designed for and the one we focus on in the main text.
This subsection develops the identification strategy at the level of events: every statement holds realization by realization, i.e., for every value of the primitives, every equilibrium consistent with them, and every selection among equilibria. No expectation, limit, or probability model enters. These realization-wise inequalities are the cleanest foundation for the identifying restrictions of Section (ref) as well as the finite-sample inference theory of Section (ref).
Under model (ref), each link satisfies the threshold-crossing representation $Y_{ij}=\mathbf{1}\{\varepsilon_{ij}\leq\delta_{ij}\}$ with $\delta_{ij}=\delta_{ij}(\theta_{0})=Z_{ij}'\beta_{0}+X_{ij}'\gamma_{0}$. The essential difficulty is that the threshold $\delta_{ij}$ contains the endogenous covariate $X_{ij}$, which is co-determined with the entire shock vector $\varepsilon$ through the equilibrium: the event $\{\varepsilon_{ij}\leq\delta_{ij}\}$ couples the shock to an endogenous, equilibrium-determined random threshold, and its probability depends on the intractable equilibrium mapping $g$ and on the unknown selection mechanism.
The “bounding-by-$c$” idea is to trade the endogenous threshold for a deterministic one, which we now explain. Fix any constant $c\in\mathbb{R}$ and consider the event \[ Y_{ij}=1\quad\text{and}\quad\delta_{ij}\leq c. \] Given that $\left\{ Y_{ij}=1\right\} =\{\varepsilon_{ij}\leq\delta_{ij}\}$, together with the bounding-by-c event $\left\{ \delta_{ij}\leq c\right\} $, we can deduce from simple transitivity of the two inequalities that $\varepsilon_{ij}\leq c$, an event that involves only the exogenous shock and the deterministic number $c$. Similarly, a flipped event \[ Y_{ij}=0\quad\text{and}\quad\delta_{ij}\geq c \] implies $\varepsilon_{ij}>\delta_{ij}\geq c$ and thus $\varepsilon_{ij}>c$. Formally, for any $c\in\mathbb{R}$, we have
Three features of (ref) deserve emphasis, because they carry the entire methodological weight of the paper.
First, the event inequalities (ref) are valid algebraically for every realization of the primitives $(Z,\varepsilon)$, the equilibrium network $Y$, and the induced network statistics $X$. Nothing about the equilibrium correspondence, its multiplicity, or the selection mechanism is used beyond model (ref) itself. This is why the resulting restrictions below are robust to equilibrium multiplicity and arbitrary, unknown selection.
Second, (ref) can be aggregated under any nonnegative weights or more generally, any weakly increasing transformation: averaging over the dyads of a covariate cell the cornerstone of the inference procedures of Section (ref)), summing signed certificates across the links of a configuration (Remark (ref)), and taking conditional expectations in any environment where they are identified (Remark (ref)). In every such aggregate the endogeneity of $X_{ij}$ is immaterial: no conditional distribution of $X_{ij}$ given $Z$ is ever modeled, estimated, or simulated---the endogenous covariate enters only through the observable indicator $\mathbf{1}\{\delta_{ij}(\theta)\leq c\}$.
Third, the test constant $c$ generates a one-dimensional family of restrictions. Each $c$ contributes one lower and one upper bound on $\mathbf{1}\left\{ \varepsilon_{ij}\leq c\right\} $, whose probability is exactly $F_{\varepsilon}\left(c\right)$. Hence, sweeping $c$ over $\mathbb{R}$ is effectively a way to trace out an envelope for the entire error CDF $F_{\varepsilon}\left(c\right)$. The identifying content for $\theta$ comes from the interplay between the $c$-family, the aggregation across dyads, and the conditioning on the dyad's exogenous types $Z_{D}:=(Z_{i},Z_{j})$, which we will explain in the next subsection.
\
Above we presented bounding-by-$c$ inequalities for dyad $ij$ under transferable utility, which is the case we carry through the paper for expositional simplicity. The technique is tied to neither restriction: it applies to general subnetwork configurations and to nontransferable utility as well, as the next two remarks illustrate. While the identifying restrictions of Section (ref) and the inference procedures of Section (ref) are plausibly adaptable to each setting, we leave these extensions to future work.
“Standard identification analysis” usually requires converting inequalities on random variables into deterministic population (or asymptotic) restrictions after finite-sample randomness of the model is averaged out. To do so in a single large network setting, one might expect that we would need delicate weak-dependence conditions. The main message of this subsection is that no such conditions are needed for the validity of our identifying restrictions. This is because the bounding-by-$c$ inequalities hold realization by realization on the observed network, and the only unknown they sandwich is the distribution of a purely exogenous quantity, whose empirical counterpart converges in the large-network limit by classical limit theory for exchangeable arrays, regardless of what the equilibrium correspondence and the selection mechanism do.
\
Formally, we consider the following single large-netwok asymptotics. We place all primitives on a single probability space: an infinite i.i.d.\ sequence $(Z_{i})_{i\geq1}$, an infinite array $(\varepsilon_{ij})_{ij}$ of i.i.d.\ shocks independent of $Z$, an auxiliary randomization variable $U_{\mathrm{sel}}$ independent of $(Z,\varepsilon)$ (allowing randomized equilibrium selection), and, for each $n$, a realized network $Y^{(n)}=g_{n}(Z^{(n)},\varepsilon^{(n)},U_{\mathrm{sel}};\theta_{0})$ formed from the first $n$ agents by the ($n$-dependent) equilibrium-and-selection mapping, as in (ref). We write $\mathbb{P}$ for the joint law of $\big((Z_{i})_{i\geq1},(\varepsilon_{ij})_{ij},U_{\mathrm{sel}}\big)$. The equilibrium correspondence and the selection mechanism embodied in $g_{n}$ may change arbitrarily with $n$ and may depend on all primitives.
We now explain how we derive the large-network identifying restrictions.
We start by working with finite-sample conditional average of the bounding-by-$c$ inequalities derived in Section (ref).
We first define a discrete dyad-type random variable $Z_{D,ij}$ below, which allows us to cover discrete, continuous and mixed $Z_{i}$ in a unified manner.
When $Z_{i}$ has continuous or mixed support, the above device allows us to proceed with a discretized dyad type, which will result in no loss of validity as shown below.\footnote{We work throughout with the discretized dyad type because it yields the cleanest conditional-mean formulation: the cells partition the dyads exactly, so the conditioning is exact and distinct cells hold disjoint sets of shocks. What discretization costs is informational coarsening, not validity: the restrictions hold cell by cell for any partition, so a coarser partition simply delivers fewer of them. For continuous $Z_{i}$, we could in principle use a finer conditioning device by multiplying the pointwise inequalities with a nonnegative kernel, at the cost of additional regularity conditions for the limiting statements. In finite samples, however, the trade-off often runs the other way, and coarsening can be actively desirable. Cell counts govern how tightly the exogenous middle term concentrates, so in sparse networks, where fine cells hold few dyads, a coarser partition buys sharper control of that term and hence a smaller critical value (See Section (ref) on critical values in finite-sample inference). } When $Z_{i}$ has finite support, one can always take the partition to be each support point in $\text{Supp}\left(Z_{i},Z_{j}\right)$. That said, in computation, one may still want to aggregate support points in $\text{Supp}\left(Z_{i},Z_{j}\right)$ into coarser cells even when $Z_{i}$ is discrete.
We start by fixing a candidate parameter $\theta$, a threshold $c$, and a generic dyad type $z_{D}$ in the support of $Z_{D,ij}$. Define the realized cell frequencies
together with the exogenous middle term
All three objects are defined on the event that $\sum_{ij}\mathbf{1}\{Z_{D,ij}=z_{D}\}\geq1$, which occurs for all large $n$ almost surely since $\mathbb{P}\left(Z_{D,ij}=z_{D}\right)=\mathbb{P}\left(\left(Z_{i},Z_{j}\right)\in{\cal Z}_{k}\right)>0$. We emphasize that $\delta_{ij}(\theta)=Z^{'}_{ij}\beta+X^{'}_{ij}\gamma$ is still defined with the raw exogenous covariates $Z_{ij}$, not recomputed based on the discretized $Z_{D,ij}$.\footnote{We are therefore not “discretizing the model” in any way. If $Z_{i}$ is continuous in the model, then $\delta_{ij}(\theta)$ continues to vary continuously with $Z_{ij}$, and $\beta$ remains the coefficient on the raw covariate. The discretized type $Z_{D,ij}$ enters only as the conditioning variable through which dyads are aggregated into the cell averages above.} The frequencies $\hat{p}^{\,(n)}_{L}$ and $\hat{p}^{\,(n)}_{U}$ are computable from observed data $(Y,X,Z)$,\footnote{The upper frequency (ref) scans the strict index event $\{\delta_{ij}(\theta)>c\}$. The weak version $\{\delta_{ij}(\theta)\geq c\}$ is equally valid---an unformed link certifies $\varepsilon_{ij}>\delta_{ij}(\theta_{0})$ strictly, so $\delta_{ij}(\theta_{0})\geq c$ still delivers $\varepsilon_{ij}>c$---and is weakly tighter. We maintain the strict convention throughout.} while the middle term $\hat{F}_{n}$ involves the unobserved shocks and serves purely as a theoretical bridge and is never computed.
The per-dyad inclusions (ref) hold pointwise---for every realization of the primitives, every equilibrium, and every selection mechanism. Averaging within a cell therefore yields the finite-network oracle sandwich: for every $n$ and every realization,
at every $z_{D}$ such that $\sum_{ij}\mathbf{1}\{Z_{D,ij}=z_{D}\}\geq1$. The sandwich inequality (ref) is stronger than a population moment inequality: it holds before any expectation is taken, and under arbitrary dependence in the realized network. All equilibrium complications are confined to the two observable outer terms, while the middle term contains only the independent $Z$ and $\varepsilon$.
What drives our large-network identification result is the convergence of the middle term $\hat{F}_{n}(c;z_{D})$ as $n\to\infty$.
The proof (Appendix (ref)) is classical: the numerator and denominator of $\hat{F}_{n}$ are $U$-statistics of the jointly exchangeable, dissociated array formed by the i.i.d.\ types $(Z_{i})$ and the i.i.d.\ dyadic shocks $(\varepsilon_{ij})$, so the strong law of large numbers for exchangeable arrays applies. The equilibrium and the endogenous $X_{ij}$ play no role in this convergence.
Combining the oracle sandwich with the uniform strong law fixes the profiling step. Define, for each $(z_{D},c,\theta)$, the (possibly random) limit frequencies
which always exist. Crucially, the inequality $\limsup_{n}\hat{p}^{\,(n)}_{L}\leq\liminf_{n}\hat{p}^{\,(n)}_{U}$ does not follow from (ref) alone, but it does follow once the middle term is known to converge: by Lemma (ref), almost surely,
This yields the population version of single-network validity.
We refer to $\Theta^{\mathrm{\infty}}_{I}$ (or its parametric counterpart) as the pathwise identified set: it is defined along the realized network sequence.
\
We now provide some discussion about Theorem (ref).
First, we clarify that we use the phrase “identified set” without claiming sharpness, and there is a substantive reason to doubt sharp characterizations are feasible in this class. Sharp identified sets in games with multiple equilibria are often characterized by, in effect, solving the model: for each candidate parameter and each realization of the latent primitives, enumerating the full set of equilibrium outcomes and checking whether some selection reproduces the distribution of the observables. Such sharp identification results have been formalized through random-set and optimal-transport methods by beresteanu2011sharp and galichon2011set. Here in the specific context of strategic network formation, the requirement of solving or inverting the pairwise-stability correspondence over a combinatorial outcome space is intractable in both theoretical and practical senses, which is exactly the motivation behind this paper. The value of the bounding-by-$c$ restrictions is that they extract equilibrium-robust identifying content without ever solving the game, and the finite-sample inference of Section (ref) is built on their realization-wise structure rather than on sharpness.
Second, Theorem (ref) is a validity statement rather than a characterization of a deterministic population quantity. Without weak-dependence conditions, $\hat{p}^{\,(n)}_{L}$ and $\hat{p}^{\,(n)}_{U}$ need not converge: global interdependence, or an equilibrium selection that oscillates along the sequence, can keep them fluctuating, in which case $p^{\infty}_{L}$ and $p^{\infty}_{U}$, as well as the identified set $\Theta^{\infty}_{I}$ can remain random. This thus marks a departure from the standard identification analysis, and we regard it as the right one in a single large network setting. Conventional identification analysis, point- or set-valued, rests upon a deterministic feature of the observable distribution of data: a population conditional probability or moment, or its probability limit, which the data can in principle recover and the model restricts. In our context no such object needs to exist. What is deterministically pinned down in the limit (or population) is instead the middle term: the empirical distribution of the exogenous shocks $\hat{F}_{n}$ converges to $F_{\varepsilon}$ along almost every path, and all identifying restrictions come from sandwiching this convergent middle term between two possibly random observable bounds. The situation is reminiscent of laws of large numbers without ergodicity in time series settings: for time-series processes, limits of intertemporal sample means are in general random variables in the form of conditional expectations given the invariant or exchangeable $\sigma$-field rather than constants eagleson1978limit,davezies2021empirical, and inference can still proceed conditionally on the realized time-series path. In the same spirit, identification here is articulated pathwise, along the realized network sequence, without taking any stand on whether population limits of observable frequencies exist, or whether there exists a well-defined limit network model which all finite-$n$ models converge to.
The pathwise identification in Section (ref) is a large-network result along a coupled sequence of networks as $n\to\infty$. This section develops the finite-sample counterpart, which leads to a valid inference procedure that constitutes the paper's central contribution.
We start with the semiparametric case, which turns out easier to present and helps convey the conceptual core of our inference approach.
The semiparametric identifying restriction (ref) in Theorem (ref) can be encoded by the extremum criterion \[ Q^{\infty}(\theta):=\sup_{c\in\mathbb{R}}\Big[\max_{z_{D}}p^{\infty}_{L}(z_{D},c;\theta)-\min_{z_{D}}p^{\infty}_{U}(z_{D},c;\theta)\Big]_{+}, \] which satisfies $Q^{\infty}(\theta_{0})=0$ almost surely. We define its finite-sample analogue by
where $\hat{{\cal Z}}_{D}$ is the family of cells with at least $\underline{m}$ observations, i.e., \[ \hat{{\cal Z}}_{D}:=\big\{ z_{D}\in{\cal Z}_{D}:\hat{m}(z_{D})\geq\underline{m}\big\},\qquad\hat{m}(z_{D}):=\sum_{i<j}\mathbf{1}\{Z_{D,ij}=z_{D}\}, \] for a pre-chosen $\underline{m}\geq1$.\footnote{Taking $\underline{m}=1$ is the minimal requirement for the conditional averages $\hat{p}^{\,(n)}_{L}$, $\hat{p}^{\,(n)}_{U}$ and $\hat{F}_{n}$ to be well-defined. Any larger floor is equally valid, and is often preferable in practice. A larger $\underline{m}$ keeps only cells with enough dyads to carry little finite-sample uncertainty, which lowers the critical value. The effect on the confidence set is a genuine trade-off rather than a free improvement: both $\widehat{Q}_{n}$ and the critical value are computed over the same family $\hat{{\cal Z}}_{D}$, so raising $\underline{m}$ discards restrictions and lowers the test statistic as well. In sparse networks with many thin cells the first effect typically dominates, which is why we impose a moderate floor rather than $\underline{m}=1$. The validity of using a pre-chosen floor $\underline{m}$ follows again from the realization-wise validity of the bounding-by-$c$ inequalities.} Also, even though $\sup_{c\in\mathbb{R}}$ appears to be a supremum taken over the whole real line, in finite sample it is actually computable with finitely many evaluations only. This is because, at each fixed $\theta$, the maps $c\mapsto\hat{p}^{\,(n)}_{L}(z_{D},c;\theta)$ and $c\mapsto\hat{p}^{\,(n)}_{U}(z_{D},c;\theta)$ are step functions whose jumps happen exactly at the finitely many realized index values in $\{\delta_{ij}(\theta):i<j\}$. Hence, $\widehat{Q}_{n}$ is exactly computable at each candidate $\theta$.
To control the finite-sample uncertainty in $\widehat{Q}_{n}(\theta_{0})$, we use the “control statistic”
which dominates $\widehat{Q}_{n}(\theta_{0})$ as shown in the following lemma.
At a null hypothesis $\theta_{0}$, the statistic $\widehat{Q}_{n}(\theta_{0})$ on the left is observable and may carry arbitrary network endogeneity and dependence, while the oracle statistic $R_{n}$ on the right contains no $Y$ or $X$, and only involves the exogenous $\varepsilon$ and $Z$. Any probability statement about $R_{n}$ based on the independence assumption between $\varepsilon$ and $Z$ therefore translates directly into coverage for the truth $\theta_{0}$.
\
A key, and somewhat surprising, property of the control statistic $R_{n}$ is that its conditional distribution given $Z$ can be simulated without the knowledge of $F_{\varepsilon}$.
We describe our proposed inference procedure in the following. For $b=1,\dots,B$ with $B$ sufficiently large, draw $\hat{m}(z_{D})$ i.i.d. uniform random variables on $[0,1]$ for each $z_{D}\in\hat{{\cal Z}}_{D}$, and independently across $b$ and across cells. Let $G^{(b)}_{z_{D}}$ be the empirical distribution function of cell $z_{D}$, and set \[ R^{(b)}:=\sup_{u\in[0,1]}\Big[\max_{z_{D}\in\hat{{\cal Z}}_{D}}G^{(b)}_{z_{D}}(u)-\min_{z_{D}\in\hat{{\cal Z}}_{D}}G^{(b)}_{z_{D}}(u)\Big]. \] Let $q^{\,R}_{n,1-\alpha}(Z)$ be the $(1-\alpha)$-th quantile\footnote{In fact, we can deliver exact finite-$B$ coverage guarantee by using the critical value $\hat{q}^{\,R}_{n,1-\alpha}(Z)$, defined as the $\lceil(1-\alpha)(B+1)\rceil$-th smallest value of $R^{(1)},\dots,R^{(B)}$. Such choice of $\hat{q}^{\,R}_{n,1-\alpha}(Z)$ would guarantee that randomness from the finite-$B$ simulation is also properly accounted for in the confidence set construction. That said, since $B$ can be made arbitrarily large, in the main text we abstract away from finite-$B$ sampling error.} of $R^{(b)}$, which we will use as the critical value. We then construct the semiparametric confidence set as
Note that the inference procedure uses only the cell sizes $\hat{m}(z_{D})$, without the need of drawing $\varepsilon$ from the unknown distribution $F_{\varepsilon}$ or simulating any network formation process.
The proof is in Appendix (ref), where the key idea is to exploit the probability integral transform (and its discrete analog). Conditional on $Z$, different dyad cells $z_{D}\in{\cal \hat{Z}}_{D}$ contain disjoint sets of i.i.d. dyadic shocks (Assumption (ref)). When $F_{\varepsilon}$ is continuous, setting $U_{ij}:=F_{\varepsilon}(\varepsilon_{ij})$ produces i.i.d. uniform random variables by the probability integral transform, so $\hat{F}_{n}(c;z_{D})=G_{z_{D}}(F_{\varepsilon}(c))$, and thus the conditional distribution of $R_{n}$ only depends on the per-cell sample size $\hat{m}(z_{D})$, which can be computed by the $R^{(b)}$ simulation. When $F_{\varepsilon}$ has atoms, the discretized probability integral transform continues to work in delivering validity, but would make the test more conservative than it is in the case with a continuous $F_{\varepsilon}$.
\
In fact, a closed-form critical value can be derived using an analytic bound on the tail probability of $R_{n}$:
We now consider the parametric case, in which $F_{\varepsilon}$ is known up to a finite-dimensional $\theta_{\varepsilon}$. Note that, when $F_{\varepsilon}$ is taken to be normal or logistic (as commonly adopted in the literature), it is without loss of generality to set $F_{\varepsilon}$ to be standard normal or standard logistic as location and scale normalization, so that no free parameter $\theta_{\varepsilon}$ is required.\footnote{When $F_{\varepsilon}$ is set to be standard normal or standard logistic, a constant term should be included in $Z_{ij}$, and the whole vector of $\theta_{0}$ should be treated as free-varying parameter.} That said, we keep the notation $\theta_{\varepsilon}$ for theoretical generality.
Knowing $F_{\varepsilon}$ allows us to leverage the two-sided bounding-by-$c$ bounds separately, in a similar style as the identifying restrictions in Theorem (ref)(ii). Specifically, we define the finite-sample test statistic
Correspondingly, we define the oracle statistic
which controls the uncertainty in $\widehat{Q}^{par}_{n}(\overline{\theta})$ at truth.
At each null $\theta_{\varepsilon,0}$, we can again simulate the distribution of $R^{par}_{n}\left(\theta_{\varepsilon,0}\right)$ to obtain the critical value.
Specifically, for $b=1,\ldots,B$ with $B$ sufficiently large, draw a full i.i.d.\ pair-shock array $\{\varepsilon^{(b)}_{ij}\}_{ij}$ from $F_{\varepsilon}(\cdot\,;\theta_{\varepsilon})$, and then compute $\hat{F}^{(b)}_{n}(c;z_{D})$ and \[ R^{par,(b)}_{n}(\theta_{\varepsilon}):=\max_{z_{D}\in\hat{{\cal Z}}_{D}}\sup_{c\in\mathbb{R}}\big|\hat{F}^{(b)}_{n}(c;z_{D})-F_{\varepsilon}(c;\theta_{\varepsilon})\big|. \] Let $q_{n,1-\alpha}(Z;\theta_{\varepsilon})$ be the $(1-\alpha)$-th quantile\footnote{As explained in Footnote (ref), we can use the finite-$B$ adjusted $\hat{q}_{n,1-\alpha}(Z;\theta_{\varepsilon})$, defined as the $\lceil(1-\alpha)(B+1)\rceil$-th smallest value of $R^{par,(1)}_{n},\dots,R^{par,(B)}_{n}$. } of $R^{par,(b)}_{n}(\theta_{\varepsilon})$, and construct the parametric confidence set as \[ \widehat{\mathcal{C}}^{par}_{n}(1-\alpha):=\big\{\overline{\theta}:\ \widehat{Q}^{par}_{n}(\overline{\theta})\leq q_{n,1-\alpha}(Z;\theta_{\varepsilon})\big\}. \]
The previous subsections focus on simple finite-sample inference procedures that are direct “finite-sample analogues” of the identifying restrictions in Theorem (ref). That said, they are not the only valid procedures induced by the sandwich inequalities (ref). Specifically, since (ref) is a collection of pointwise inequalities, one for each cell $z_{D}$ and threshold $c$, we can first transform each side of the inequalities under an arbitrary nondecreasing transformation that preserves the inequalities, and then construct test statistics based on the transformed inequalities. This allows us to better control and account for the varying degrees of uncertainties across $z_{D}$ and $c$ values, leading to a less conservative test (and thus tighter confidence set) than the ones constructed above.
For concreteness and simplicity, we focus on the parametric case. See Remark (ref) below on how similar type of improvements also carry over in the semiparametric case.
At a candidate $\theta$, define
Evaluated at the truth $\overline{\theta}_{0}$, the sandwich inequalities (ref) translate to \[ V^{\pm}(z_{D},c;\overline{\theta}_{0})\ \leq\ R^{\pm}(z_{D},c;\theta_{\varepsilon})\qquad\text{for every }(z_{D},c,\pm), \] which we abbreviate to \[ V\left(\overline{\theta}_{0}\right)\leq R\left(\theta_{\varepsilon,0}\right), \] with $V,R$ understood as the concatenations across all $(z_{D},c,\pm)$ combinations.
Conditional on $Z$ and at the truth $\overline{\theta}_{0}$, let $T$ be any nondecreasing\footnote{That is, $T\left(v\right)\leq T\left(v^{'}\right)$ whenever $v\leq v^{'}$ elementwisely.} aggregation function, mapping from the space of $V$ or $R$ into a scalar. Then obviously we have
For example, the statistics $\widehat{Q}^{par}_{n}$ and $R^{par}_{n}$ in Section (ref) are constructed by taking the aggregation function $T$ to be the maximum function over all $(z_{D},c,\pm)$ combinations, in which case (ref) specializes to Lemma (ref).
The above implies the following generalization of the conditional Monte Carlo procedure from Section (ref).
Conditional on $Z$ and at any candidate $\overline{\theta}$, let $T_{Z,\overline{\theta}}$ be any known nondecreasing aggregation as described above. Here we emphasize that $T_{Z,\overline{\theta}}$ can freely depend on $\left(Z,\overline{\theta}\right)$, since our test will be conditional on $Z$ with $\overline{\theta}$ hypothesized as truth under the null. As before, we simulate $B$ i.i.d. copies $R^{(1)},\dots,R^{(B)}$ based on $Z$ and $F_{\varepsilon}\left(\,\cdot\,;\theta_{\varepsilon}\right)$, compute $c^{\,T}_{n,1-\alpha}(Z;\overline{\theta})$ as the $(1-\alpha)$-th quantile of the aggregated statistics $T_{Z,\overline{\theta}}(R^{(b)})$. Then we construct the $T$-induced confidence set as \[ \widehat{\mathcal{C}}^{T}_{n}(1-\alpha):=\left\{ \overline{\theta}:\ T_{Z,\overline{\theta}}\left(V\left(\overline{\theta}\right)\right)\leq c^{\,T}_{n,1-\alpha}\left(Z;\overline{\theta}\right)\right\} , \] whose validity is established below.
This freedom in the choice of $T$ matters because $T$ controls how uncertainty is aggregated and weighted across $(z_{D},c,\pm)$, which affects the distribution of $T(R(\theta_{\varepsilon}))-T\big(V(\overline{\theta}_{0})\big)$. The maximum function used in Lemma (ref) is intuitive from the identifying restrictions, but in finite sample it suffers from the sensitivity of taking maximums: one thin cell at an uninformative threshold can become a “maximum” by chance and affects the test statistic or the critical value drastically. Hence, heuristically, we should aggregate the inequalities in a way that controls the influence of any single $(z_{D},c,\pm)$ configuration with very few data points (and thus high uncertainty).
Concretely, we consider the following three “obvious improvements”. For the rest of this subsection, we suppress $\theta_{\varepsilon}$ from $F_{\varepsilon}$ and write $\theta$ instead of $\overline{\theta}$: the results still go through at each candidate $\theta_{\varepsilon}$, but, as discussed earlier, in practice $F_{\varepsilon}$ is often taken to be standard normal or standard logistic, leaving no more free $\theta_{\varepsilon}$. Since the rest of the subsection focuses more on practical implementation, this notational simplification helps better convey the key ideas of the improvements without unnecessary clutter in the notation.
\
(i) Constrained Thresholding.
First, we show how to systematically trim away thresholds $c$ that carry no identifying information about $\theta$ and would only inflate the simulated maximum. At each $(\theta,Z)$ we take the constrained threshold set $C(z_{D};\theta,Z)$ to be the a priori range of the candidate index in cell $z_{D}$. Formally, writing ${\cal X}$ as the support of $X$, \[ C(z_{D};\theta,Z)=\Big[\min_{\tilde{z}_{D}\in z_{D}}w_{n}\left(\tilde{z}\right)^{'}\beta+\inf_{x\in{\cal X}}\gamma^{\prime}x,\ \ \max_{\tilde{z}_{D}\in z_{D}}w_{n}\left(\tilde{z}\right)^{'}\beta+\sup_{x\in{\cal X}}\gamma^{\prime}x\Big], \] where $\min_{\tilde{z}_{D}\in z_{D}}w_{n}\left(\tilde{z}\right)^{'}\beta$ is the minimum exogenous index across exogenous covariate values $\tilde{z}$ that is consistent with the discretized cell $z_{D}$.\footnote{Recall from the description after model (ref) that the function $w_{n}$ generates the exogenous dyadic covariates from individual characteristics, i.e., $Z_{ij}=w_{n}\left(Z_{i},Z_{j}\right)$. And “$\tilde{z}_{D}\in z_{D}$” means that the discretized version $\tilde{z}_{D}$ of the vector of individual-leve covariates $\tilde{z}$ according to Definition (ref) belongs to the cell $z_{D}$. } The point of restricting to $C(z_{D};\theta,Z)$ is that it discards thresholds $c$ that could not have bite.\footnote{Outside the interval $C(z_{D};\theta,Z)$, neither $\hat{p}^{\,(n)}_{L}$ nor $\hat{p}^{\,(n)}_{U}$ varies with $c$: below it every dyad in the cell has $\delta_{ij}(\theta)>c$, so $\hat{p}^{\,(n)}_{L}=0$ and $\hat{p}^{\,(n)}_{U}=\overline{Y}(z_{D})$, the link frequency of the cell; above it every dyad has $\delta_{ij}(\theta)\leq c$, so $\hat{p}^{\,(n)}_{L}=\overline{Y}(z_{D})$ and $\hat{p}^{\,(n)}_{U}=1$.} That said, for $c$ outside $C(z_{D};\theta,Z)$, the control statistic $R^{\pm}(z_{D},c;\theta_{\varepsilon})$ , which does not involves the index, would still inflate the simulated maximum and thus the associated critical value. Hence, for finite-sample performance it is strictly better to throw away $c$ outside $C(z_{D};\theta,Z)$. In many practical cases, ${\cal X}$ is a compact set, so the restriction to $C(z_{D};\theta,Z)$ can be quite substantial.\footnote{If ${\cal X}$ is unbounded but sign-restricted, the same formula can still help trim away unnecessary thresholds. To illustrate, for scalar $X_{ij}$ with ${\cal X}=[0,\infty)$, we have \[ C(z_{D};\theta,Z)=
\] }
Given the threshold set $C(z_{D};\theta,Z)\subseteq\mathbb{R}$, we can aggregate restrictions only across $\left(z_{D},c\right)$ such that $c\in C(z_{D};\theta,Z)$. To illustrate, if we aggregate by taking the maximum across $\left(z_{D},c\right)$ such that $c\in C(z_{D};\theta,Z)$, then the induced test statistic and the uncertainty control statistic will be
which correspond to the aggregation $T$ that keeps the coordinates with $c\in C(z_{D};\theta,Z)$ and discards the rest before taking the maximum; multiplying each coordinate by a $(\theta,Z)$-measurable indicator is nondecreasing, so Theorem (ref) applies. Then, let $\hat{c}^{\,\mathrm{con}}_{n,1-\alpha}(Z;\theta)$ be the $\left(1-\alpha\right)$ sample quantile of $\big(T^{\mathrm{con}}_{\theta}(R^{(b)})\big)^{B}_{b=1}$, and the validity of confidence sets constructed by test inversion is then guaranteed by Theorem (ref). That said, constrained threshold sets need not be combined with the maximum aggregation rule: it can be combined with the improvements below as well.
\
(ii) Studentization.
Dividing each inequality by its null standard deviation equalizes scale across cells of unequal size, so that in finite sample a $20$-dyad cell cannot drastically influence the critical value for a $2{,}000$-dyad one. Define \[ \sigma(z_{D},c):=\sqrt{\frac{F_{\varepsilon}\left(c\right)\left(1-F_{\varepsilon}\left(c\right)\right)}{\hat{m}(z_{D})}},\quad\mathcal{S}:=\Big\{(z_{D},c):F_{\varepsilon}\left(c\right)\left(1-F_{\varepsilon}\left(c\right)\right)\geq\frac{1}{\hat{m}(z_{D})}\Big\}, \] where $\sigma(z_{D},c)$ is the standard deviation of the cell frequency $\hat{F}_{n}(c;z_{D})$, the average of $\hat{m}(z_{D})$ i.i.d.\ Bernoulli indicators $\mathbf{1}\left\{ \varepsilon_{ij}\leq c\right\} $ with success probability $F_{\varepsilon}(c)$. The set ${\cal S}$ selects the subset of $(z_{D},c)$ whose null standard deviation is at least above a threshold driven by the granularity of the effective sample size $\hat{m}(z_{D})$ in the cell, ruling out tail thresholds $c$ (very close to $0$ or $1$) that are too noisy to “estimate” given the $\hat{m}(z_{D})$ data points. The test statistic and uncertainty control statistic can then be defined as
which are again defined by a single nondecreasing aggregation rule $T$: divide each $(z_{D},c)$ by $\sigma(z_{D},c)$, and only sup over $(z_{D},c)\in{\cal S}$.
\
(iii) Berk--Jones Weights.
Studentization reweights $(z_{D},c)$ through division by the standard deviation of the cell frequency, approximately equalizing the scales of variations across $(z_{D},c)$. An alternative way to equalize uncertainty scales across $(z_{D},c)$ is to use the Berk--Jones score berk1979goodness, which we describe below.
Specifically, fix a cell $z_{D}$ with $\hat{m}(z_{D})$ dyads and a threshold $c$. Under the null the number of dyads in the cell whose shock falls below $c$ is \[ N(z_{D},c):=\sum_{i<j:\,Z_{D,ij}=z_{D}}\mathbf{1}\{\varepsilon_{ij}\leq c\}\ \sim\ \mathrm{Binomial}\big(\hat{m}(z_{D}),\,F_{\varepsilon}(c)\big), \] by Assumption (ref), so the cell frequency is $\hat{F}_{n}(c;z_{D})=N(z_{D},c)/\hat{m}(z_{D})$ with expectation $F_{\varepsilon}(c)$. Large-deviation probabilities of the form $\mathbb{P}\big(\hat{F}_{n}(c;z_{D})\geq p\big)$ can be approximated by an exponential form \[ \mathbb{P}\big(\hat{F}_{n}(c;z_{D})\geq p\big)\ =\ \exp\Big\{-\hat{m}(z_{D})\,\mathrm{KL}\big(p\,\big\|\,F_{\varepsilon}(c)\big)+o\big(\hat{m}(z_{D})\big)\Big\}\qquad\text{for }p>F_{\varepsilon}(c), \] and symmetrically for $p<F_{\varepsilon}(c)$, where \[ \mathrm{KL}(p\|u):=p\log\frac{p}{u}+(1-p)\log\frac{1-p}{1-u} \] is the Bernoulli Kullback--Leibler divergence.\footnote{Both directions are straightforward. For the upper bound, write $N\sim\mathrm{Binomial}(m,u)$ and apply Markov's inequality to $e^{\lambda N}$: for any $\lambda>0$, \[ \mathbb{P}(N\geq mp)\ \leq\ e^{-\lambda mp}\big(1-u+ue^{\lambda}\big)^{m}\ =\ \exp\Big\{-m\big[\lambda p-\log(1-u+ue^{\lambda})\big]\Big\}, \] and the bracket is maximized at $e^{\lambda}=\tfrac{p(1-u)}{u(1-p)}$, where it equals exactly $\mathrm{KL}(p\|u)$, giving $\mathbb{P}(N\geq mp)\leq e^{-m\mathrm{KL}(p\|u)}$ for every $m$ and every $p>u$. The method of types supplies the matching lower bound $\mathbb{P}(N\geq mp)\geq(m+1)^{-1}e^{-m\mathrm{KL}(p\|u)}$ at integer $mp$, so the exponent is exact. No normal approximation enters at any point.} Weighting the inequality at $(z_{D},c)$ by $\hat{m}(z_{D})\,\mathrm{KL}\big(\cdot\,\big\|\,F_{\varepsilon}(c)\big)$ therefore reports it on a tail-probability scale, while the tail is exactly where we have the most salient concern about finite-sample uncertainty. In particular, in the tails where $F_{\varepsilon}(c)\left(1-F_{\varepsilon}(c)\right)$ is near zero the studentized ratio is unstable, while the KL divergence grows more stably, and hence there is no need to “throw away” tails based on the set ${\cal S}$ as constructed in the studentization improvement above in (ii).\footnote{A concrete illustration. Take a cell of size $m$ and let $u:=F_{\varepsilon}(c)\downarrow0$ with exactly one dyad below $c$, so that $p=1/m$. The studentized deviation is \[ \frac{p-u}{\sqrt{u(1-u)/m}}\ \sim\ \frac{1}{\sqrt{mu}}, \] which diverges at a power rate as $u\downarrow0$. In comparison, the Berk--Jones weight \[ m\,\mathrm{KL}\big(\tfrac{1}{m}\big\| u\big)=\log\frac{1}{mu}+(m-1)\log\frac{1-1/m}{1-u}\ \sim\ \log\frac{1}{mu}, \] diverges only logarithmically.}
The map $p\mapsto\mathrm{KL}\big(p\,\big\|\,F_{\varepsilon}(c)\big)$ is strictly decreasing on $[0,F_{\varepsilon}(c)]$ and strictly increasing on $[F_{\varepsilon}(c),1]$, so weighting each side of the sandwich only where that side is violated yields an aggregation that is nondecreasing in the violation magnitude $V$,\footnote{Theorem (ref) requires $T$ to be nondecreasing in each coordinate of $V$, and those coordinates are the nonnegative violations of the form $V^{\pm}$. When $\hat{p}^{\,(n)}_{L}>F_{\varepsilon}(c)$ we may write $\hat{p}^{\,(n)}_{L}=F_{\varepsilon}(c)+V^{+}$, so the assigned weight at that coordinate is $\hat{m}(z_{D})\,\mathrm{KL}\big(F_{\varepsilon}(c)+V^{+}\,\big\|\,F_{\varepsilon}(c)\big)$, which is nondecreasing in $V^{+}$ and vanishes at $V^{+}=0$. Symmetrically the weight when $\hat{p}^{\,(n)}_{U}<F_{\varepsilon}(c)$ is $\hat{m}(z_{D})\,\mathrm{KL}\big(F_{\varepsilon}(c)-V^{-}\,\big\|\,F_{\varepsilon}(c)\big)$, which is again nondecreasing in $V^{-}$. } and therefore Theorem (ref) applies. Taking the maximum over the constrained threshold sets of (i),
with the arguments $(z_{D},c;\theta)$ of $\hat{p}^{\,(n)}_{L}$ and $\hat{p}^{\,(n)}_{U}$ suppressed, and the matching control statistic\footnote{At most one of the two indicators can equal one, since $\hat{p}^{\,(n)}_{L}\leq\hat{p}^{\,(n)}_{U}$. The same holds on the oracle side, where $\hat{F}^{(b)}_{n}(c;z_{D})$ lies on one side of $F_{\varepsilon}(c)$ and the two-sided deviation therefore collapses to the single score $\hat{m}(z_{D})\,\mathrm{KL}\big(\hat{F}^{(b)}_{n}(c;z_{D})\,\big\|\,F_{\varepsilon}(c)\big)$.} is
with $\hat{c}^{\,\mathrm{BJ}}_{n,1-\alpha}(Z;\theta)$ similarly defined based on simulations of $T^{\mathrm{BJ}}_{\theta}(R^{(b)})$. This is the constrained-threshold Berk--Jones statistic used in Sections (ref) and (ref).
This section numerically evaluates the performance of the confidence sets proposed in Section (ref), across different network sizes from $100$ to $10,000$.
We consider the following model specification
with $\varepsilon_{ij}$ drawn from i.i.d.\ $\mathrm{Logistic}(0,1)$, $z_{i}$ drawn from i.i.d.\ uniform on $21$ equally spaced grid on $[-10,10]$, and $\overline{\mathrm{CF}}_{ij}(Y):=\mathrm{CF}_{ij}(Y)/(n-2)$ being the normalized common-friend count on the equilibrium network $Y$. This normalization ensures that the data generating process, and thus the equilibrium network, stays stable across $n$. Conditioning cells ($z_{D}\in{\cal Z}_{D}$) are the $231$ exact unordered type pairs. We consider network size $n\in\{100,200,400,800,1600,3200,6400,10000\}$ with $K=100$ independent repetitions at every size. We calibrate $b_{0}$ via pilot simulation runs so that the outcome networks at $n=400$ have densities around 0.2, and fix $b_{0}$ thereafter.
The choice of the strategic coefficient $\gamma_{0}=16$ seems “large”, but its magnitude should be interpreted together with the scale of $X_{ij}:=\mathrm{CF}_{ij}(Y)/(n-2)$, which is normalized by $1/(n-2)$. Table (ref) reports summary statistics about the equilibrium networks, showing that the endogenous covariate $X$ are generally small in scale with mean around 0.05 and sd around 0.06. Hence, $\gamma_{0}=16$ only translates to a very moderate amount of variation in the strategic index term. For example, at $n=200$ the strategic term contributes $\gamma_{0}\cdot\mathrm{mean}(X)\approx0.75$ logit units at the mean and $\gamma_{0}\cdot\mathrm{sd}(X)\approx1.13$ logit units of variations, against a standard logistic shock $\varepsilon_{ij}$.
We also note that, since $\overline{\mathrm{CF}}$ is nondecreasing in the network and $\gamma_{0}>0$, the network formation game is supermodular, and the equilibrium networks are computed as the least fixed points by iterations from the empty graph. While our inference procedure requires no supermodularity, no equilibrium selection, and no computation of equilibrium networks, for simulations we do need to generate the simulated equilibrium networks that our inference procedure takes as data. The focus on supermodular structure enables us to compute large equilibrium networks with $n=10,000$.
Throughout the rest of this section, we compute various confidence sets as described in Section (ref) and their performance across the $K=100$ repetitions. In the parametric case, the standard logistic CDF already fixes the location and scale, so every structural parameter in $\left(\beta_{0},\gamma_{0},b_{0}\right)$ is treated as unknown, and the test (inversion) is joint over the whole vector. We work with a discrete grid on the parameter space: $b_{0}\in[-6,6]$ with step size $0.5$, $\beta\in[-3,1]$ with step size $0.25$, and $\gamma\in[-8,60]$ with step size $0.25$. The test statistic is the Berk--Jones statistics in Section (ref), with the thresholds $c$ fixed on $361$ points spanning $[-36,36]$. The conditional critical value is simulated accordingly with $B=999$ at confidence level $1-\alpha=0.95$. In the semiparametric case, we normalize the scale of the homophily effect parameter to unity, $|\beta_{0}|=1$, with the location normalization either imposed via the intercept term $b_{0}$ or the symmetry-about-$0$ assumption.
Each of the $100$ networks at a size is one replication in which the researcher observes that network alone and computes a confidence set for the full parameter triple. Table (ref) reports statistics of the confidence sets across all three coordinates. We discuss the main findings as follows.
Our inference procedure is clearly valid and informative (though conservative as expected), in a single network setting even at the smallest size $n=100$, and median width of the CS projects shrink at every tested $n$. The true parameter is covered by the CS in every of the $K=100$ replications across all $n$.\footnote{This is anticipated since our tests are built on realization-wise inequalities that heuristically ensures the “worst-case” coverage guarantee regardless of network dependence structure and equilibrium selection mechanism. } That said, the CS are nontrivial subsets of the parameter grid that contract with $n$ and deliver sign-revealing information.
Noticeably, the sign of the strategic coefficient $\gamma$, an important question of interest in the strategic network formation literature, is certified in $75\%$ of replications at $n=100$, in $96\%$ at $n=200$, and in all $100$ replications from $n=400$ onward, while the median $\gamma$ projection width falls from about 47 ($69\%$ of the candidate grid) to about $2$ ($3.3\%$ of the grid) at $n=10,000$. Again, each computed confidence set is based on a single network, with no assumption nor information on network equilibrium ever used in the inference procedure. We are unaware of prior inference procedures in the literature that deliver valid sign-revealing confidence sets in this environment.
The confidence set also delivers quite tight bounds on the homophily effect parameter $\beta_{0}$ (as well as the intercept $b_{0}$). Homophily (negative $\beta_{0}$) is certified in $98\%$ of replications by $n=400$ and in all of them from $n=800$. The projection widths for $\beta_{0}$ and $b_{0}$ both shrink with $n$, and from $n=3200$ onwards the widths around the true values effectively shrink under the grid resolution (“0.000”). This confirms the intuition that identification and inference on parameters for the exogenous covariates tend to be “easier”. What is particularly noteworthy in our exercise is the finding that this intuition appears still true even in presence of the “hard-to-learn” strategic coefficient $\gamma_{0}$.
Now, in Table (ref), we report results on the semiparametric versions of our inference procedure, based on the conditional simulation critical values described in Section (ref) as well as the symmetry-based one described in Appendix (ref). We also include corresponding results about the parametric case from the last subsection for the ease of comparison.
These confidence sets are computed on the same simulated networks with the same $\gamma$ grid as the parametric ones in Section (ref). A key operational difference lies in the different location and scale normalization between the semiparametric case and the parametric case. The parametric CS are joint over the three-dimensional parameters $(b_{0},\beta,\gamma)$ jointly, with normalization imposed through the location and scale of the standard logistic distribution. The symmetry-based semiparametric CS (Appendix (ref)) encodes a location normalization (symmetry around $0$) but no scale normalization. Consequently, we fix the scale of the homophily effect parameter $|\beta|=1$ and the symmetry-based CS is joint over an intercept $b_{0}$, the strategic coefficient $\gamma_{0}$, as well as the sign of $\beta_{0}$. Lastly, the fully semiparametric CS (Section (ref)) leave both location and scale normalization open: we thus fix $b_{0}=0$ (effectively dropping the constant term) as location normalization and set $|\beta|=1$ as scale normalization.
The semiparametric confidence sets, while less tight than the parametric ones as anticipated, remain nontrivial and informative. Given the generality and complexity of the strategic network formation problem under consideration, it is not a priori clear whether any inference method can deliver nontrivial bounds, especially on the strategic coefficient $\gamma_{0}$. It is thus surprising that our semiparametric confidence sets turn out to be informative on the sign of $\gamma_{0}$, especially given how few assumptions were used in the semiparametric inference procedure.
It should be acknowledged that the semiparametric confidence sets still materially get censored by the grid bounds, especially at small $n$. Upper-endpoint contact is about $80\%$ at $n=400$ and $27$-$45\%$ at $n=800$, so the “width” entries at $n=400$ and $n=800$ are contaminated with grid truncation and thus should be interpreted with care. The results at $n=100$ and $n=200$ should be read as sign evidence alone. That said, for larger networks (starting $n=1600$), the semiparametric confidence sets deliver two-sided bounds in the grid interior for the majority of replications.
The “cleanest” metric to focus on is the sign rate, the share of replications that certify the (correct) sign of $\gamma_{0}$, and this measure also shows a clear ordering of the confidence set performance with respect to the strength of the imposed assumptions. The strongest parametric assumption delivers the 100% sign rate starting at $n=400$, the semiparametric symmetry-based one reaches 100% at $n=1600$, while the purely semiparametric one (without any assumption on $F_{\varepsilon}$) requires $n=6400$ to do so.
The design so far has one exogenous covariate $\left|z_{i}-z_{j}\right|$. We now repeat the exercise with two and three exogenous covariates, with true coefficient vectors $\beta_{0}=(-1,+1)$ and $\beta_{0}=(-1,+1,-.5)$. Each margin of $Z_{i}$ is independently uniform on the same $21$ points in $[-10,10]$, and the scale normalization is imposed on the first coordinate $|\beta_{1}|=1$. The $\gamma$ grid spans the same $[-8,60]$, with a larger step size to save computation. The results for the parametric and the two semiparametric confidence sets are reported in Table (ref).
We note first that, under the multidimensional $Z$ design, the effective sample sizes in the “discretized cells” $z_{D}$ often tend to be too small at $n\leq400$, especially for the 3-dimensional design. For example, there are $21^{3}$ support points in the 3-dimensional $Z$ space, so at $n=100$ the parametric CS is very wide (even after cell grouping) and the semiparametric procedure is effectively not meaningful at all, which is thus labeled by “-” in Table (ref). The situation improves gradually as with $n$, but for $n\leq400$, even the parametric CS remains wide, while the semiparametric versions are mostly vacuous. That said, this is an anticipated issue as the dimension of $Z$ grows.
With the caveat above acknowledged, the results still continue to demonstrate the informativeness of our inference procedure, especially in its ability to certify the sign of $\gamma_{0}$. In addition, the performance ordering with respect to the strength of assumptions on $F_{\varepsilon}$ remains clearly shown even under multidimensional $Z$. The parametric CS reaches sign rates $\geq90\%$ at $n=200$ in both 2-dim-$Z$ and $3$-dim-$Z$ settings, while the semiparametric versions eventually reach $\geq90\%$ rates after $n\geq1600.$
We apply the parametric confidence sets of Section (ref) to two empirical data sets: contact (and friendship) among high-school students, and friendship among Twitch streamers. In each network $Y_{ij}$ records the observed link, $X_{ij}$ is a common-friend statistic computed from that same network, and $Z_{ij}$ collects observed dyadic covariates to be explained below. We do not specify which stable network is chosen when the model admits more than one.
The Thiers13 data contain $327$ students in nine classes and four class types mastrandrea2015contact.\footnote{The data set is publicly available and accessed at \href{https://sociopatterns.org/datasets/high-school-contact-and-friendship-networks/}{https://sociopatterns.org/datasets/high-school-contact-and-friendship-networks/}.} Contact sensors record face-to-face proximity in twenty-second intervals during one school week, and we set $Y_{ij}=1$ when a pair has at least $\underline{\kappa}$ recorded intervals in total. The baseline is $\underline{\kappa}=2$, with $\underline{\kappa}=1$ and $\underline{\kappa}=3$ also reported as robustness checks. The friendship layer sets $Y_{ij}=1$ if either student nominates the other in a survey to which only about 130 of the 327 students responded (and hence the friendship network should only be interpreted with this caveat). The exogenous covariates are an indicator for different classes and an indicator for different class types, so $Z_{ij}=(1,D_{ij},T_{ij})$, and the endogenous covariate is the raw common-friend count in the same network. Selected summary statistics are reporeted in Table (ref).
Again, we use the Berk--Jones statistic with constrained thresholding in Section (ref). Since the raw numbers of common friends $CF$ are already small in scale and in the empirical applications there is just one given sample size $n$, for convenience we simply use the raw $CF$ count but work with a small-scale grid for $\gamma$ with step size $0.01$ (noticing that the results can be equivalently translated into $CF/(n-2)$ with $\gamma$ correspondingly scaled up). The remaining coefficient are based on grid with step size $0.5$, and the joint grid is chosen so that every displayed projection below lies strictly inside the grid interior.
Table (ref) reports the parametric 95%-level confidence-set projection result under the assumption that $\varepsilon_{ij}$ is standard logistic.
All four common-friend projections are positive. A positive $\gamma$ indicates the tendency for triad closure in strategic network formation: sharing a contact (or friend) raises the net surplus from linking. In the baseline $\underline{\kappa}=2$ contact network one additional common friend raises the structural link index by $0.05$ to $0.37$. At the sample mean of about $2.1$ common friends, the endogenous term contributes roughly $0.1$ to $0.8$ index units, relative to standard logistic error scale. The $\underline{\kappa}=1$ and $\underline{\kappa}=3$ rows give nearly the same range, so the positive sign of $\gamma$ does not depend on a single contact threshold.
The friendship CS projection is also positive but less precise, which is consistent with the fact that “friendship” is constructed from a survey where only 130 of the 327 students responded at all. Correspondingly there is very low variation in the “friendship” data, where $97.5\%$ of dyads have no common friend and $193$ of the $327$ students are isolated. Again, this means that the results based on the friendship network should be taken with this substantive caveat in mind, and regarded as a supplemental robustness check using this data set.
The Twitch data set contains information about networks of gamers who stream in a certain language (DE, EN, ES, FR, PT, and RU): six undirected friendship networks, one per streamer language rozemberczki2021multiscale.\footnote{The data set is publicly available and accessed at \href{https://snap.stanford.edu/data/twitch-social-networks.html}{https://snap.stanford.edu/data/twitch-social-networks.html}.} We treat each network as a separate inference problem with language-specific parameters, and each confidence sets is computed on each network separately. See Table (ref) for some key summary statistics.
We construct the exogenous covariates as follows. Let $q^{V}_{i}\in\{1,\ldots,5\}$ be streamer $i$'s within-language quintile of total views and $m_{i}$ the mature-content indicator. We construct \[ V_{ij}=|q^{V}_{i}-q^{V}_{j}|,\qquad S_{ij}=q^{V}_{i}+q^{V}_{j}-6,\qquad M_{ij}=m_{i}+m_{j}-1, \] so that $V_{ij}$ measures the popularity gap, $S_{ij}$ captures the popularity level of the pair, and $M_{ij}$ the (centered) mature-account count. The structural index in language-$\ell$ network is then \[ \delta_{ij,\ell}=b_{0\ell}+b_{V\ell}V_{ij}+b_{S\ell}S_{ij}+b_{M\ell}M_{ij}+\gamma_{\ell}\,\mathbf{1}\{\mathrm{CF}_{ij}\geq1\}, \] with $Z_{ij}=(1,V_{ij},S_{ij},M_{ij})$ and the endogenous regressor $X_{ij}=\mathbf{1}\{\mathrm{CF}_{ij}\geq1\}$ as an indicator of whether the pair has at least one shared neighbor. We work with a full fixed threshold ($c$) grid from $-48$ to $48$ in steps of $0.25$. The parameter grid uses step sizes of $.05$ on all parameters. Again the joint parameter grid is chosen so that every displayed projection below lies strictly in the grid interior.
The reason why we use the binary $X_{ij}=\mathbf{1}\{\mathrm{CF}_{ij}\geq1\}$ is to control the highly skewed and heterogeneous distribution of $\mathrm{CF}_{ij}$ within each language network and across the six language networks. Table (ref) documents how extreme that skewness is. In the German network, for instance, $62.8\%$ of dyads have no common friend at all and the 99th percentile is $9$, yet the maximum $\mathrm{CF}_{ij}$ is $1{,}432$. A raw count would therefore let a handful of neighborhoods dominate the identifying variation in $X_{ij}$, while it is equally problematic to use logs to reduce skewness given the large share of zeros in the $\mathrm{CF}_{ij}$. All the six networks are very skewed, but the magnitudes of the skewness vary substantially across the six. As a result, the indicator $\mathbf{1}\{\mathrm{CF}_{ij}\geq1\}$ seems as natural “common friends” type statistic with clear interpretation and retains stability within and across the six networks.
Table (ref) reports the parametric confidence-set projections under the assumption that $\varepsilon_{ij}$ is standard logistic. Every language produces a strictly positive CS projection for $\gamma$: the lower endpoints range from $0.55$ in Russian to $1.45$ in German. German has the largest lower endpoint, which under the logistic normalization implies more than a $4.2$-fold multiplier of the latent threshold odds associated with the first common friend. The upper endpoints are less informative in these sparse networks with a binary endogenous regressor, though they are comparable in raw magnitudes (6.95-8.45) across the six languages, indicating the stability of our inference procedure across these very different networks. We thus view the robustly positive lower endpoints and the stable upper endpoints as an empirical demonstration of the informativeness and robustness of our inference procedure with respect to the strategic coefficient $\gamma$.