EconBase
← Back to paper

Logical Differencing in Dyadic Network Formation Models with Nontransferable Utilities

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

87,149 characters · 13 sections · 23 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Logical Differencing in Dyadic Network Formation Models with Nontransferable Utilities

\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long\global\long\global\long \global\long\global\long \global\long \global\long \global\long \global\long\global\long\global\long\global\long\global\long \global\long \global\long\def\mc#1{\mathscr{#1}}

\global\long\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long\global\long \global\long\def\abs#1{\left|#1\right|}

\global\long\def\norm#1{\left\Vert #1\right\Vert }

\global\long\def\rest#1{\left.#1\right|}

\global\long\def\bracket#1#2{\left\langle #1\middle\vert#2\right\rangle }

\global\long\def\sandvich#1#2#3{\left\langle #1\middle\vert#2\middle\vert#3\right\rangle }

\global\long\def\turd#1{\frac{#1}{3}}

\global\long \global\long\def\sand#1{\left\lceil #1\right\vert }

\global\long\def\wich#1{\left\vert #1\right\rfloor }

\global\long\def\sandwich#1#2#3{\left\lceil #1\middle\vert#2\middle\vert#3\right\rfloor }

\global\long\def\abs#1{\left|#1\right|}

\global\long\def\norm#1{\left\Vert #1\right\Vert }

\global\long\def\rest#1{\left.#1\right|}

\global\long\def\inprod#1{\left\langle #1\right\rangle }

\global\long\def\ol#1{\overline{#1}}

\global\long\def\ul#1{#1}

\global\long\def\td#1{\tilde{#1}}

\global\long \global\long \global\long \global\long \global\long

abstractThis paper considers a semiparametric model of dyadic network formation under nontransferable utilities (NTU). Such dyadic links arise frequently in real-world social interactions that require bilateral consent but by their nature induce additive non-separability. In our model we show how unobserved individual heterogeneity in the network formation model can be canceled out without requiring additive separability. The approach uses a new method we call logical differencing. The key idea is to construct an observable event involving the intersection of two mutually exclusive restrictions on the fixed effects, while these restrictions are as necessary conditions of weak multivariate monotonicity. Based on this identification strategy we provide consistent estimators of the network formation model under NTU. Finite-sample performance of our method is analyzed in a simulation study, and an empirical illustration using the risk-sharing network data from Nyakatoke demonstrates that our proposed method is able to obtain economically intuitive estimates.

Keywords: dyadic network formation, semiparametric estimation, nontransferable utilities, additive nonseparability

Introduction

This paper considers a semiparametric model of dyadic network formation under nontransferable utilities (NTU), which arise naturally in the modeling of real-world social interactions that require bilateral consent. For instance, friendship is usually formed only when both individuals in question are willing to accept each other as a friend, or in other words, when both individuals derive sufficiently high utilities from establishing the friendship. It is often plausible that the two individuals may derive very different utilities from the friendship for a variety of reasons: for example, one of them may simply be more introvert than the other and derive lower utilities from the friendship. In addition, there may not be a feasible way to perfectly transfer utilities between the two individuals. Monetary payments may not be customary in many social contexts, and even in the presence of monetary or in-kind transfers, utilities may not be perfectly transferable through these feasible forms of transfers, say, when individuals have different marginal utilities with respect to these transfers.\footnote{See surveys by aumann1967survey, \citet*{hart1985nontransferable} and \citet*{mclean2002values} for discussions on the implications of NTU on link (bilateral relationship) and group formation from a micro-theoretical perspective.} Given the considerable academic and policy interest in understanding the underlying drivers of network formation,\footnote{For example, the formation of friendship among U.S. high-school students has been studied by a long line of literature, such as \citet*{moody2001race}, \citet*{currarini2009economic,currarini2010identifying}, \citet*{boucher2015structural}, \Citet{currarini2016simple}, xu2018diverse among others. } it is not only theoretically interesting but also empirically relevant to incorporate NTU in the modeling of network formation.

This paper contributes to the line of econometric literature on network formation by introducing and incorporating nontransferable utilities into dyadic network formation models. Previous work in this line of literature focuses primarily on case of transferable utilities, as represented in \citet*{graham2017econometric}, which considers a parametric model with homophily effects and individual unobserved heterogeneity of the following form:

equation[equation omitted — 142 chars of source]

where $D_{ij}$ is an observable binary variable that denotes the presence or absence of a link between individual $i$ and $j$, $w\left(X_{i},X_{j}\right)$ represents a (symmetric) vector of pairwise observable characteristics specific to $ij$ generated by a known function $w$ of the individual observable characteristics $X_{i}$ and $X_{j}$ of $i$ and $j$, while $A_{i}$ and $A_{j}$ stand for unobserved individual-specific degree heterogeneity and $\epsilon_{ij}$ is some idiosyncratic utility shock. Model (ref) essentially says that, if the (stochastic) joint surplus generated by a bilateral link $s_{ij}:=w\left(X_{i},X_{j}\right)^{'}\beta_{0}+A_{i}+A_{j}-\epsilon_{ij}$ exceeds the threshold zero, then the link between $i$ and $j$ is formed. The model implicitly assumes that the link surplus can be freely distributed among the two individuals $i$ and $j$, and that bargaining efficiency is always achieved, so that the undirected link is formed if and only if the link surplus is positive. Given this specification, \citet*{graham2017econometric} provides consistent and asymptotically normal maximum-likelihood estimates for the homophily effect parameter $\beta_{0}$, assuming that the exogenous idiosyncratic pairwise shocks $\epsilon_{ij}$ are independently and identically distributed with a logistic distribution. Recently, \citet*{Candelaria2016} and \citet*{Toth2017} provide semiparametric generalizations of \citet*{graham2017econometric}, while \citet*{gao2018nonparametric} established nonparametric identification of a class of index models that further generalize (ref).

This paper, however, generalizes \citet*{graham2017econometric} along a different direction, and seeks to incorporate the natural micro-theoretical feature of NTU into this class of network formation models. To illustrate\footnote{Starting from Section (ref), we consider a more general specification than the illustrative model (ref) introduced here.}, consider the following simple adaption of model (ref) with two threshold-crossing conditions:

equation[equation omitted — 240 chars of source]

where the unobserved individual heterogeneity $A_{i}$ and $A_{j}$ separately enter into two different threshold-crossing conditions. This formulation could be relevant to scenarios where $A_{i}$ represents individual $i$'s own intrinsic valuation of a generic friend: for a relatively shy or introvert person $i$, a lower $A_{i}$ implies that $i$ is less willing to establish a friendship link, regardless of how sociable the counterparty is. For simplicity, suppose for now that $w\left(X_{i},X_{j}\right)\equiv{\bf 0}$, $\epsilon_{ij}\sim_{iid}F$, $\epsilon_{ji}\sim_{iid}F$ and $\epsilon_{ij}\perp\epsilon_{ji}$\footnote{For our general result, we do not require $\epsilon_{ij}\perp\epsilon_{ji}$, nor the log-concavity of $F$. They are used here for illustration purpose only. See a discussion after (ref).}. Focusing completely on the effects of $A_{i}$ and $A_{j}$, it is clear that the TU model (ref) implies that only the sum of “sociability”, $A_{i}+A_{j}$, matters: the linking probability among pairs with $A_{i}=A_{j}=1$ (two moderately social persons) should be exactly the same as the linking probability among pairs with $A_{i}=2$ and $A_{j}=0$ (one very social person and one very shy person), which might not be reasonable or realistic in social scenarios. In comparison, the linking probability among pairs with $A_{i}=2$ and $A_{j}=0$ is lower than the linking probability among pairs with $A_{i}=A_{j}=1$ under the NTU model (ref) with i.i.d. $\epsilon_{ij}$ and $\epsilon_{ji}$ that follow any log-concave distribution\footnote{A distribution is log-concave if $F\left(x\right)^{\lambda}F\left(y\right)^{1-\lambda}\leq F\left(\lambda x+\left(1-\lambda\right)y\right)$. Many commonly used distributions, such as uniform, normal, exponential, logistic, chi-squared distributions, are log-concave. See bagnoli2005log for more details on log-concave distributions from a microeconomic theoretical perspective.}:

align*[align* omitted — 282 chars of source]

This is intuitive given the observation that, under bilateral consent, the party with relatively lower utility is the pivotal one in link formation. Moreover, even though we maintain strict monotonicity in the unobservable characteristics $A_{i}$ and $A_{j}$, the NTU setting can still effectively incorporate homophily effects on unobserved heterogeneity: given that $w\left(X_{i},X_{j}\right)\equiv{\bf 0}$ and $A_{i}+A_{j}=2$, the linking probability is effectively decreasing in $\left|A_{i}-A_{j}\right|$ under log-concave $F$. Hence, by explicitly modeling NTU in dyadic network formation, we can accommodate more flexible or realistic patterns of conditional linking probabilities and homophily effects that are not present under the TU setting.

However, the NTU setting immediately induces a key technical complication: as can be seen explicitly in model (ref), the observable indexes, $w\left(X_{i},X_{j}\right)^{'}\beta_{0}$ and $w\left(X_{j},X_{i}\right)^{'}\beta_{0}$, and the unobserved heterogeneity terms ($A_{i}$ and $A_{j}$) are no longer additively separable from each other. In particular, notice that, even though the utility specification for each individual inside each of the two threshold-crossing conditions in model (ref) remains completely linear and additive, the multiplication of the two (nonlinear) indicator functions directly destroys both linearity and additive separability, rendering inapplicable most previously developed econometric techniques that arithmetically “difference out” the “two-way fixed effects” $A_{i}$ and $A_{j}$ based on additive separability.\footnote{Equivalently, one could write model (ref) in an alternative form as a “single” composite threshold-crossing condition: \[ D_{ij}=\mathbf{\mathbbm1}\left\{ \min\left\{ w\left(X_{i},X_{j}\right)^{'}\beta_{0}+A_{i}-\epsilon_{ij},w\left(X_{j},X_{i}\right)^{'}\beta_{0}+A_{j}-\epsilon_{ji}\right\} \geq0\right\} , \] where additive separability is again lost in this alternative formulation.}

Given this technical challenge, this paper proposes a new identification strategy termed logical differencing, which helps cancel out the unobserved heterogeneity terms, $A_{i}$ and $A_{j}$, without requiring additive separability but leveraging the logical implications of multivariate monotonicity in model (ref). The key idea is to construct an observable event involving the intersection of two mutually exclusive restrictions on the fixed effects $A_{i}$ and $A_{j}$, which logically imply an event that can be represented without $A_{i}$ or $A_{j}$. Specifically, in the context of the illustrative model (ref) above, we start by considering the event where a given individual $\ol i$ is more popular than another individual $\ol j$ among a group of individuals $k$ with observable characteristics $X_{k}=\ol x$ while $\ol i$ is simultaneously less popular than another individual $\ol j$ among a group of individuals with a certain realization of observable characteristics $\ul x$. This is the same as the conditioning event in \citet*{Toth2017} and analogous to the tetrad comparisons made in Candelaria2016. However, instead of using arithmetic differencing to cancel out the unobserved heterogeneity $A_{\ol i}$ and $A_{\ol j}$ as in Candelaria2016 and \citet*{Toth2017}, we make the following logical deductions based on the monotonicity of the conditional popularity of $\ol i$ in $w\left(X_{\ol i},\ol x\right)^{'}\beta_{0}$ and $A_{\ol i}$. First, the event that $\ol i$ is more popular than another individual $\ol j$ among the group of individuals with $X_{k}=\ol x$ implies that either $w\left(X_{\ol i},\ol x\right)^{'}\beta_{0}>w\left(X_{\ol j},\ol x\right)^{'}\beta_{0}$ or $A_{\ol i}>A_{\ol j}$, while the event that $\ol i$ is less popular than another individual $\ol j$ among a different group of individuals with $X_{l}=\ul x$ implies that either $w\left(X_{\ol i},\ul x\right)^{'}\beta_{0}<w\left(X_{\ol j},\ul x\right)^{'}\beta_{0}$ or $A_{\ol i}<A_{\ol j}$. Second, when both events occur simultaneously, we can logically deduce that either $w\left(X_{\ol i},\ol x\right)^{'}\beta_{0}>w\left(X_{\ol j},\ol x\right)^{'}\beta_{0}$ or $w\left(X_{\ol i},\ul x\right)^{'}\beta_{0}<w\left(X_{\ol j},\ul x\right)^{'}\beta_{0}$ must have occurred, because $A_{\ol i}>A_{\ol j}$ and $A_{\ol i}<A_{\ol j}$ cannot simultaneously occur. Intuitively, the “switch” in the relative popularity of $\ol i$ and $\ol j$ among the two groups of individuals with characteristics $\ol x$ and $\ul x$ cannot be driven by individual unobserved heterogeneity $A_{\ol i}$ and $A_{\ol j}$, and hence when we indeed observe such a “switch”, we obtain a restriction on the parametric indices $w\left(X_{\ol i},\ol x\right)^{'}\beta_{0}$, $w\left(X_{\ol i},\ol x\right)^{'}\beta_{0}$, $w\left(X_{\ol i},\ul x\right)^{'}\beta_{0}$, and $w\left(X_{\ol j},\ul x\right)^{'}\beta_{0}$, which helps identify $\beta_{0}$.

Based on this identification strategy we provide sufficient conditions for point identification of the parameter $\beta_{0}$ up to scale normalization as well as a consistent estimator for $\beta_{0}$. Our estimator has a two-step structure, with the first step being a standard nonparametric estimator of conditional linking probabilities, which we use to assert the occurrence of the conditioning event, while in the second step we use the identifying restriction on $\beta_{0}$ when the conditioning event occurs. The computation of the estimator essentially follows the same method proposed in \citet*{gao2018robust}, with some adaptions to the network data setting. We plot the identified sets under various restrictions on the support of the observable characteristics $X_{i}$, analyze the finite-sample performance in a simulation study, and present an empirical illustration of our method using data from Nyakatoke on risk-sharing network collected by Joachim De Weerdt.

This paper belongs to the line of literature that studies dyadic network formation in a single large network setting, including \citet*{blitzstein2011}, \citet*{chatterjee2011}, \citet*{yan2013}, \citet*{yan2016}, \citet*{graham2017econometric}, \citet*{charbonneau2017}, \citet*{dzemski2017empirical}, \citet*{jochmans2017semiparametric}, \citet*{yan2018statistical}, \citet*{Candelaria2016}, \citet*{Toth2017} and \citet*{gao2018nonparametric}. \citet*{shi2016structural} explicitly incorporates NTU into dyadic network formation models, but \citet*{shi2016structural} considers a fully parametric model and establishes the consistency and asymptotic normality of the maximum likelihood estimators. See also the recent surveys by \citet*{de2020econometric} and \citet*{graham2020dyadic}.

This paper is also related to a line of research that utilizes dyadic link formation models in order to study structural social interaction models: for instance, arduini2015parametric, \citet*{auerbach2016identification}, \citet*{goldsmith2013social}, \citet*{hsieh2016social} and \citet*{johnsson2021estimation}. In these papers, the social interaction models are the main focus of identification and estimation, while the link formation models are used mainly as a tool (a control function) to deal with network endogeneity or unobserved heterogeneity problems in the social interaction model. Even though some of the network formation models considered in this line of literature is consistent with the NTU setting, this line of literature is usually not primarily concerned with the full identification and estimation of the network formation model itself.

It should be pointed out that in this paper we do not consider link interdependence in network formation, which is studied by the line of econometric literature on strategic network formation models. This line of literature primarily uses pairwise stability \citep*{Jackson1996} as the solution concept for network formation, and also often builds NTU into the econometric specification. See, for example, \citet*{de2018identifying}, \citet*{graham2016homophily}, \citet*{leung2015random}, \citet*{menzel2015b}, boucher2017my, \citet*{mele2017dense}, \citet*{mele2017structural} and \citet*{ridder2017estimation}. However, this type of models usually do not feature unobserved heterogeneity as in this paper. See, for example, \citet*{de2020strategic} for a more detailed survey on this line of literature.

This paper is also closely related to to \citet*{gao2018robust}, which similarly leverages multivariate monotonicity in a multi-index structure under a panel multinomial choice setting. It should be pointed out that, even though there is some structural similarity between network data and panel data, there are no direct ways in the network setting to make “intertemporal comparison” as in the panel setting, which holds the fixed effects unchanged across two observable periods of time. It is precisely this additional complication induced by the network setting that requires the technique of logical differencing proposed in this paper.

The rest of the paper is organized as follows. In Section (ref), we describe the general specifications of our dyadic network formation model. Section (ref) establishes identification of the parameter of interests in our model and provides a consistent tetrad estimator. We plot the identified sets under various restrictions on $X$ and report baseline simulation results in Section (ref). We present an empirical illustration using the risk-sharing data of Nyakatoke in Section (ref). Section (ref) concludes. Proofs and additional simulation results are available in the Appendix.

A Nonseparable Dyadic Network Formation Model

We consider the following dyadic network formation model:

equation[equation omitted — 162 chars of source]

where:

itemize$i\in\left\{ 1,...,n\right\} $ denote a generic individual in a group of $n$ individuals. • $X_{i}$ is a $\mathbb{R}^{d_{x}}$-valued vector of observable characteristics for individual $i$. This could include, for example, wealth, age, education and ethnicity of individual $i$. • $D_{ij}$ denotes a binary observable variable that indicates the presence or absence of an undirected and unweighted link between two distinct individuals $i$ and $j$: $D_{ij}=D_{ji}$ for all pairs of individuals $ij$, with $D_{ij}=1$ indicating that $ij$ are linked while $D_{ij}=0$ indicating that $ij$ are not linked. • $w:\mathbb{R}^{d_{x}}\times\mathbb{R}^{d_{x}}\to\mathbb{R}^{d}$ is a known function that is symmetric\footnote{Our method can also be adapted to the case with asymmetric $w$. See Remark (ref).} with respect to its two vector arguments. We will write $W_{ij}:=w\left(X_{i},X_{j}\right)$ for notational simplicity. • $\beta_{0}\in\mathbb{R}^{d}$ is an unknown finite-dimensional parameter of interest. Assume $\beta_{0}\neq{\bf 0}$ so that we may normalize $\norm{\beta_{0}}=1$, i.e., $\beta_{0}\in\mathbb{\mathbb{S}}^{d-1}$. • $A_{i}$ is an unobserved scalar-valued variable that represents unobserved individual heterogeneity. • $\phi:\mathbb{R}^{3}\to\mathbb{R}$ is an unknown measurable function that is symmetric with respect to its second and third arguments.

In addition, we impose the following two assumptions:

assumption[Monotonicity] $\phi$ is weakly increasing in each of its arguments.

Assumption (ref) is the key assumption on which our identification analysis is based. It requires that the conditional linking probability between individuals with characteristics $\left(X_{i},A_{i}\right)$ and $\left(X_{j},A_{j}\right)$ be monotone in a parametric index $\delta_{ij}:=W_{ij}^{'}\beta_{0}$ as well as the unobserved individual heterogeneity terms $A_{i}$ and $A_{j}$. It should be noted that, given monotonicity, increasingness is without loss of generality as $\phi$, $\beta_{0}$ and $A_{i},A_{j}$ are all unknown or unobservable. In addition, Assumption (ref) only requires that $\phi$ is monotonic in the index $W_{ij}^{'}\beta_{0}$ as a whole, not individual coordinates of $W_{ij}$. Therefore, we may include nonlinear or non-monotone functions $w\left(\cdot,\cdot\right)$ on the observable characteristics as long as Assumption (ref) is maintained.

Next, we impose a standard random sampling assumption:

assumption[Random Sampling] $\left(X_{i},A_{i}\right)$ is i.i.d. across $i\in\left\{ 1,...,n\right\} $.

In particular, Assumption (ref) allows arbitrary dependence structures between the observable characteristics $X_{i}$ and the unobservable characteristic $A_{i}$.

Model (ref) along with the specifications and the two assumptions introduced above encompass a large class of dyadic network formation models in the literature. For example, the standard dyadic network formation model (ref) studied by \citet*{graham2017econometric} can be written as \[ \mathbb{E}\left[\rest{D_{ij}}X_{i},X_{j},A_{i},A_{j}\right]=F\left(W_{ij}^{'}\beta_{0}+A_{i}+A_{j}\right) \] where $F$ is the CDF of the standard logistic distribution. For the semiparametric version considered by Candelaria2016, Toth2017, and gao2018nonparametric, we can simply take $F$ to be some unknown CDF. In either case, the monotonicity of the CDF $F$ and the additive structure of $W_{ij}^{'}\beta_{0}+A_{i}+A_{j}$ immediately imply Assumption (ref).

However, our current model specification and assumptions further incorporate a larger class of dyadic network formation models with potentially nontransferable utilities. Specifically, consider the joint requirement of two threshold-crossing conditions

align[align omitted — 249 chars of source]

where $u$ is an unknown function that is not necessarily symmetric with respect to its second and third arguments $\left(A_{i},A_{j}\right)$, and $\left(\epsilon_{ij},\epsilon_{ji}\right)$ are idiosyncratic pairwise shocks that are i.i.d. across each unordered $ij$ pair with some unknown distribution. In particular, notice that model (ref) is a special case of (ref). Suppose we further impose the following two lower-level assumptions 1a and 1b:

assumption*[1a] $\left(\epsilon_{ij},\epsilon_{ji}\right)$ are independent of $\left(X_{i},A_{i},X_{j},A_{j}\right)$.
assumption*[1b] $u$ is weakly increasing in its first three arguments.

Then, the conditional linking probability

align[align omitted — 449 chars of source]

can be represented by model (ref) with Assumption (ref) satisfied.

In particular, we do not require $\epsilon_{ij}\perp\epsilon_{ji}$. In fact, $\epsilon_{ij}\equiv\epsilon_{ji}$ is readily incorporated in our model. Under the maintained assumption that $w\left(X_{i},X_{j}\right)=w\left(X_{j},X_{i}\right)$, if $\epsilon_{ij}\equiv\epsilon_{ji}$ and $u$ is furthermore assumed to be symmetric with respect to its second and third arguments ($A_{i}$ and $A_{j}$), then our model specializes to the case of transferable utilities, \[ D_{ij}=\mathbf{\mathbbm1}\left\{ u\left(W_{ij}^{'}\beta_{0},A_{i},A_{j},\epsilon_{ij}\right)\geq0\right\} , \] where effectively only one threshold crossing condition determines the establishment of a given network link. Therefore, our NTU model (ref) includes the TU model as a special case.

rem[Symmetry of $w$] To explain the key idea of our identification strategy in a notation-economical way, we will be focusing on the case of symmetric $w$ in most of the following sections. However, it should be pointed out that our method can also be applied to the case where $w$ is allowed to be asymmetric in (ref), so that individual utilities based on observable characteristics can also be made asymmetric (nontransferable). In that case, model (ref) needs to be modified as \begin{equation} \mathbb{E}\left[\rest{D_{ij}}X_{i},X_{j},A_{i},A_{j}\right]=\phi\left(W_{ij}^{'}\beta_{0},W_{ji}^{'}\beta_{0},A_{i},A_{j}\right), \end{equation} where $W_{ij}=w\left(X_{i},X_{j}\right)$ may be different from $W_{ji}=w\left(X_{j},X_{i}\right)$, but $\phi$ is symmetric with respect to its first two arguments $W_{ij}^{'}\beta_{0},W_{ji}^{'}\beta_{0}$ whenever $A_{i}=A_{j}$ Moreover, Assumption (ref) should also be changed to be $\phi$ is monotone in all its four arguments. See Appendix (ref) for a more detailed discussion on how our identification strategy can be adapted to accommodate asymmetric $w$ under appropriate conditions.

Identification and Estimation

Identification via Logical Differencing

In this section, we explain the key idea of our identification strategy. We construct a mutually exclusive event to cancel out the unobservable heterogeneity $A_{i}$ and $A_{j}$, which leads to an identifying restriction on $\beta_{0}$. We call this technique “logical differencing”.

For each fixed individual $\ol i$, and each possible $\ol x\in\mathbb{R}^{d_{x}}$, define

equation[equation omitted — 125 chars of source]

as the linking probability of this specific individual $\ol i$ with a group of individuals, individually indexed by $k$, with the same observable characteristics $X_{k}=\ol x$ (but potentially different fixed effects $A_{k}$). Clearly, $\rho_{i}\left(\ol x\right)$ is directly identified from data in a single large network.

Suppose that individual $\ol i$ has observed characteristics $X_{\ol i}=x_{\ol i}$ and unobserved characteristics $A_{\ol i}=a_{\ol i}$. Then, by model (ref) we have

align[align omitted — 434 chars of source]

where the expectation in the second to last line is taken over $A_{k}$ conditioning on $X_{k}=\ol x$. As we allow $A_{k}$ and $X_{k}$ to be arbitrarily correlated, the $\psi_{\ol x}$ function defined in the last line of (ref) is dependent on $\ol x$. In the same time, notice that $\psi_{\ol x}$ does not depend on the identity of $\ol i$ beyond the values of $w\left(x_{\ol i},\ol x\right)^{'}\beta_{0}$ and $a_{\ol i}$. By Assumption (ref), $\psi_{\ol x}\left(w\left(x_{\ol i},\ol x\right)^{'}\beta_{0},a_{\ol i}\right)$ must be bivariate weakly increasing in the index $w\left(x_{\ol i},\ol x\right)^{'}\beta_{0}$ and the unobserved heterogeneity scalar $a_{\ol i}$. We now show how to use the bivariate monotonicity to obtain identifying restrictions on $\beta_{0}$.

Fixing two distinct individuals $\ol i$ and $\ol j$ in the population, we first consider the event that individual $\ol i$ is strictly more popular than individual $\ol j$ among the group of individuals with observed characteristics $X_{k}=\ol x$:

equation[equation omitted — 100 chars of source]

which is an event directly identifiable from observable data given (ref). Even though event (ref) is the same conditioning event as considered in \citet*{Toth2017} and analogous to the tetrad comparisons made in Candelaria2016, we now exploit the following logical deduction based on the bivariate monotonicity of the conditional popularity of $\ol i$ in $w\left(X_{\ol i},\ol x\right)^{'}\beta_{0}$ and $A_{\ol i}$ without the assumption of additivity between them. Specifically, writing $\left(x_{\ol i},a_{\ol i}\right)$ and $\left(x_{\ol j},a_{\ol j}\right)$ as the observable and unobservable characteristics of individuals $\ol i$ and $\ol j$, by (ref) we have

align[align omitted — 471 chars of source]

Note that the last line of equation (ref) is a natural necessary (but not sufficient) condition for $\rho_{\ol i}\left(\ol x\right)>\rho_{\ol j}\left(\ol x\right)$ under bivariate monotonicity.

Now, consider the event that individual $\ol i$ is strictly less popular than individual $\ol j$ among the group of individuals with observed characteristics $X_{k}=\underline{x}$, i.e.,

equation[equation omitted — 108 chars of source]

Then, by a similar argument to (ref), we deduce

equation[equation omitted — 269 chars of source]

Notice that the event $\left\{ a_{\ol i}<a_{\ol j}\right\} $ in (ref) is mutually exclusive with the event $\left\{ a_{\ol i}>a_{\ol j}\right\} $ that shows up in (ref).

Next, consider the event that the two events (ref) and (ref) described above simultaneously happen. Then, by (ref), (ref) and basic logical operations, we have

align[align omitted — 2,295 chars of source]

The derivations above exploit two simple logical properties: first, \[ \left\{ a_{\ol i}>a_{\ol j}\right\} \ \text{AND}\ \left\{ a_{\ol i}<a_{\ol j}\right\} \quad=\quad\text{FALSE}, \] and second, \[ \left\{ w\left(x_{\ol i},\ol x\right)^{'}\beta_{0}>w\left(x_{\ol j},\ol x\right)^{'}\beta_{0}\right\} \ \text{AND}\ \left\{ a_{\ol i}<a_{\ol j}\right\} \quad\Rightarrow\quad\left\{ w\left(x_{\ol i},\ol x\right)^{'}\beta_{0}>w\left(x_{\ol j},\ol x\right)^{'}\beta_{0}\right\} , \] which uses only necessary but not sufficient condition, so that we can obtain an identifying restriction (ref) on $\beta_{0}$ that does not involve $a_{\ol i}$ nor $a_{\ol j}$. These two forms of logical operations together enable us to “difference out” (or “cancel out”) the unobserved heterogeneity terms $a_{\ol i}$ and $a_{\ol j}$.

In contrast with various forms of “arithmetic differencing” techniques proposed in the econometric literature (including Candelaria2016 and Toth2017 specific to the dyadic network formation literature), our proposed technique does not rely on additive separability between the parametric index $w\left(x_{\ol i},\ol x\right)^{'}\beta_{0}$ and the unobserved heterogeneity term $a_{\ol i}$. Instead, our identification strategy is based on multivariate monotonicity and utilizes logical operations rather than standard arithmetic differencing to cancel out the unobserved heterogeneity terms. Hence, we term our method “logical differencing”.

The identifying arguments above are derived for a fixed pair of individuals $\ol i$ and $\ol j$, but clearly the arguments can be applied for any pair of individuals $\left(i,j\right)$ with observable characteristics $x_{i}$ and $x_{j}$. Writing

align*[align* omitted — 546 chars of source]

for each $\beta\in\mathbb{\mathbb{S}}^{d-1}$, we summarize the identifying arguments above by the following lemma.

lem[Identifying Restriction] Under model (ref) and Assumptions (ref) and (ref), we have \begin{equation} \tau_{ij}\left(\ol x,\ul x\right)=1\quad\Rightarrow\quad\lambda\left(\ol x,\ul x;x_{i},x_{j};\beta_{0}\right)=0. \end{equation}

A simple (but clearly not unique) way to build a criterion function based on Lemma (ref) is to define

equation[equation omitted — 169 chars of source]

where the expectation is $\mathbb{E}_{ij,kl}$ taken over random samples of ordered tetrads $\left(i,j,k,l\right)$ from the population, and $\left(X_{i},X_{j},X_{k},X_{l}\right)$ denote the random variables corresponding to the observable characteristics of $\left(i,j,k,l\right)$. According to Lemma (ref), $Q\left(\beta_{0}\right)=0$, which is always smaller than or equal to $Q\left(\beta\right)\geq0=Q\left(\beta_{0}\right)$ for any $\beta\neq\beta_{0}$ because $\tau_{ij}\geq0$ and $\lambda\geq0$ by construction.

Observing that the scale of $\beta_{0}$ is never identified, we write \[ B_{0}:=\left\{ \beta\in\mathbb{\mathbb{S}}^{d-1}:Q\left(\beta\right)=0\right\} \] to represent the normalized “identified set” relative to the criterion $Q$ defined in (ref). Lemma (ref) implies that $\beta_{0}\in B_{0}$, but in general there is no guarantee that $B_{0}$ is a singleton. The next subsection contains a set of sufficient conditions that guarantees $B_{0}=\left\{ \beta_{0}\right\} $.

remWe should point out that the identified set $B_{0}$ defined above based on logical differencing is not sharp in general, since the individual unobserved heterogeneity term $A_{i}$ is canceled out by logical differencing. In fact, one can show $A_{i}$ is also identified (up to proper normalization) under certain conditions\footnote{The identification of $A_{i}$ is conceptually analogous to the identification of individual fixed effects in a long panel setting.}, and knowledge about $A_{i}$ can help with the identification of $\beta_{0}$. In fact, when all realizations of $A_{i}$ are point identified, the point identification of $\beta_{0}$ can be established under much weaker conditions than those to be presented in Section (ref).\footnote{See \citet*{gao2018nonparametric} for a related discussion. In fact, the identification strategy in \citet*{gao2018nonparametric} can be adapted to establish identification of $A_{i}$ under the NTU setting. However, a rigorous presentation of such identification results is beyond the scope of this paper.} However, we trade sharpness for simplicity: this paper provides a method of identification (and estimation) without the need to deal with the incidental parameters $A_{i}$.

It is worth mentioning that for any one-sided sign preserving function $\gamma$ such that

equation[equation omitted — 136 chars of source]

we may define

align[align omitted — 409 chars of source]

without changing the identification set at all, since \[ \tau_{ij}\left(\ol x,\ul x\right)>0\quad\text{if and only if}\quad\tau_{ij}^{\gamma}\left(\ol x,\ul x\right)>0. \] In fact, $\tau^{\gamma}$ specializes to $\tau$ when we set $\gamma\left(t\right):=\mathbf{\mathbbm1}\left\{ t>0\right\} $. Alternatively, we may set $\gamma$ to be “smoother”, say, $\gamma\left(t\right):=\left[t\right]_{+}$ the positive part function.

Such forms of “smoothing” in the population criterion will be irrelevant to all the identification results and its proofs in this paper, so for notational simplicity, we will suppress $\gamma$ and focus on the representative $\tau_{ij}$ and $Q$ in the next subsection about point identification. However, a smooth $\gamma$ will play a role when it comes to estimation and computation, and we will revisit $\gamma$ in Section (ref).

Sufficient Conditions for Point Identification

We now present a set of sufficient conditions that guarantee point identification of $\beta_{0}$ on $\mathbb{\mathbb{S}}^{d-1}$.

Assumption 1$'$ \protected@write \@auxout {\string \newlabel {as1prime}{{$1'$}{\thepage}{$1'$}{as1prime}} } \hypertarget{as1prime}

(Strict Monotonicity of $\phi$). $\phi$ defined in model (ref) is weakly increasing in $A_{i}$ and $A_{j}$ while strictly increasing in the index $w\left(X_{i},X_{j}\right)^{'}\beta_{0}$.

Assumption (ref) strengthens Assumption (ref) by requiring that $\phi$ be strictly increasing in the parametric index $w\left(X_{i},X_{j}\right)^{'}\beta_{0}$. This is used to guarantee that differences in the parametric index can indeed lead to changes in conditional linking probabilities, so that the conditional event $\left\{ \tau_{ij}\left(\ol x,\ul x\right)=1\right\} $ in Lemma (ref) may occur with strictly positive probability.

assumption[Continuity of $\phi$ and $w$] $\phi$ and $w$ are continuous functions on their domains.
assumption[Sufficient Directional Variations] There exist distinct points $\ol x$ and $\ul x$ in $Supp\left(X_{i}\right)$, such that the vector ${\bf 0}$ lies in the interior of $Supp\left(w\left(\ol x,X_{i}\right)-w\left(\ul x,X_{i}\right)\right)$.

We also provide a lower-level condition for Assumption (ref) when $w$ is the coordinate-wise Euclidean distance function.

Assumption 4$'$ \protected@write \@auxout {\string \newlabel {as3prime}{{$4'$}{\thepage}{$4'$}{as3prime}} } \hypertarget{as3prime} Suppose that (i) $w_{h}\left(\ol x,\ul x\right):=\left|\ol x_{h}-\ul x_{h}\right|$ for every coordinate $h$, and (ii) $Supp\left(X_{i}\right)$ has nonempty interior.

Essentially, since our criterion function is based on indicator functions of halfspaces in the form of $\left(w\left(\ol x,\tilde{x}\right)-w\left(\ul x,\tilde{x}\right)\right)^{'}\beta\lesseqqgtr0$, we will need the distribution of these indicators to take both values $0$ and $1$ with strictly positive probabilities under any $\beta$, so that every possible $\beta$ different from $\beta_{0}$ can be differentiated from $\beta_{0}$ by the criterion function (up to scale normalization). Hence, we need sufficient variations in the observable covariates.

Assumption (ref), though apparently not very transparent on its own, is actually implied by Assumption (ref).\footnote{Proof: Suppose that $Supp\left(X_{i}\right)$ has nonempty interior. Then there exist two distinct points $\ol x$ and $\ul x$ in the interior of $Supp\left(X_{i}\right)$ such that $\tilde{x}:=\frac{1}{2}\left(\ol x+\ul x\right)$ is also in the interior of $Supp\left(X_{i}\right)$. Clearly, $w\left(\ol x,\tilde{x}\right)=w\left(\ul x,\tilde{x}\right)=\left(\frac{1}{2}\left|\ol x_{1}-\ul x_{1}\right|,...,\frac{1}{2}\left|\ol x_{k}-\ul x_{k}\right|\right)^{'}$. Since $\ol x,\ul x$ are all interior points of $Supp\left(X_{i}\right)$ and $w$ is continuous, the vector ${\bf 0}$ must be an interior point of $Supp\left(w\left(\ol x,X_{i}\right)-w\left(\ul x,X_{i}\right)\right)$. \qedsymbol} We note that the assumption of nonempty interior is a familiar one, which is often imposed for point identification in the literature, say, on maximum score estimation. Assumption (ref) allows the support of all observable covariates to be bounded, but on the other hand require all covariates to be continuously distributed.

In Appendix (ref), we present an alternative set of assumptions that allow for the presence of discrete covariates, but require the existence of a “special covariate” with large (conditional) continuous support and nonzero coefficient, a la \citet*{horowitz1992smoothed}. Since the identification result with special covariate needs to be presented under a different scale normalization, we defer the results to the appendix.

assumption[Conditional Support of $A_{i}$] $A_{i}$ is conditionally distributed on the same support given $X_{i}=x$ for any realization $x\in Supp\left(X_{i}\right)$.

Assumption (ref) together with Assumption (ref) implies that for two randomly sampled individuals $i,j$, there is a strictly positive probability of $A_{i}$ and $A_{j}$ being sufficiently close to each other, conditional on any realizations of $X_{i}$ and $X_{j}$. Together with the continuity condition in Assumption (ref), we can ensure that the parametric index based on observable covariates alone can determine whether $i$ or $j$ is relatively more popular among a certain group of individuals.

$\ $

Next, we lay out the lemma that will be used in the proof of point identification of $\beta_{0}$.

lem[Differentiating $\beta$ from $\beta_{0}$] Under model (ref), Assumptions (ref), (ref)--(ref), for each $\beta\in\text{\ensuremath{\mathbb{S}}}^{d-1}\backslash\left\{ \beta_{0}\right\} $, \begin{align} \tau_{ij}\left(X_{k},X_{l}\right) & =1,\\ \lambda\left(X_{k},X_{l};X_{i},X_{j};\beta_{0}\right) & =0,\\ \lambda\left(X_{k},X_{l};X_{i},X_{j};\beta\right) & =1, \end{align} occur simultaneously with strictly positive probability, as we randomly sample individuals $i,j,k,l$.

For point identification of $\beta_{0}$, we need the population criterion (ref) to differentiate each $\beta\in\text{\ensuremath{\mathbb{S}}}^{d-1}\backslash\left\{ \beta_{0}\right\} $ from $\beta_{0}$. Lemma (ref) ensures this by establishing that, the conditioning event (ref) in Lemma (ref) occurs with positive probability, and, when it occurs, we can obtain different values of $\lambda$ at $\beta$ from that at $\beta$ with positive probabilities. Compared with the set identification result discussed in Section (ref), we need to strengthen Assumption (ref) to strict monotonicity in the index $w\left(X_{i},X_{j}\right)^{'}\beta_{0}$. Then, Assumptions (ref) and (ref) guarantee that, when the absolute difference between $A_{i}$ and $A_{j}$ get sufficiently small, differences in the parametric indexes can lead to differences in conditional linking probabilities, so that (ref) will occur. Since $\beta_{0}$ and $\beta$ define different intersections of halfspaces for the vector $w\left(\ol x,\tilde{x}\right)-w\left(\ul x,\tilde{x}\right)$ through $\lambda$, Assumption (ref) then guarantees that there will be on-support realizations of the observable covariates that help “detect” such differences.

We are now ready to present the point identification result.

thm[Point Identification of $\beta_{0}$] Under model (ref) and Assumptions (ref), (ref)--(ref), $\beta_{0}$ is the unique minimizer of $Q\left(\beta\right)$ defined in (ref) over the unit sphere $\text{\ensuremath{\mathbb{S}}}^{d-1}$. Furthermore, for any $\epsilon>0$, there exists $\delta>0$ such that \[ \inf_{\beta\in\text{\ensuremath{\mathbb{S}}}^{d-1}\backslash B_{\epsilon}\left(\beta_{0}\right)}Q\left(\beta\right)\geq Q\left(\beta_{0}\right)+\delta, \] where $B_{\epsilon}\left(\beta_{0}\right):=\left\{ \beta\in\mathbb{\mathbb{S}}^{d-1}:\|\beta-\beta_{0}\|\leq\epsilon\right\} $.

Theorem (ref) follows from Lemma (ref) based on the standard arguments in \citet*{newey1994asymp}.

rem*[Asymmetry of $w$, Continued] In Appendix (ref), we show how the identification arguments and assumptions above can be adapted to accommodate asymmetry of $w$. In short, the technique of logical differencing applies without changes, but the identifying restriction we obtained becomes weaker. In particular, when $w$ is antisymmetric in the sense that $w\left(\ol x,\ul x\right)+w\left(\ul x,\ol x\right)\equiv0$, the identifying restriction we obtained through logical differencing becomes trivial, and $B_{0}=\mathbb{\mathbb{S}}^{d-1}$. However, with asymmetric but not antisymmetric $w$, it is still feasible to strengthen Assumption 3 so as to obtain point identification. See more discussions in Appendix (ref).

Tetrad Estimator and Consistency

We now proceed to present a consistent estimator of $\beta_{0}$ in the framework of extremum estimation, which we construct using a two-step semiparametric estimation procedure. We clarify that we are considering the asymptotics under “a single large network” with the number of individuals $N\to\infty$. Moreover, we focus on the “dense network” asymptotics where the conditional linking probabilities are nondegenerate in the limit and can be consistently estimated.

The first step is the nonparametric estimation of \[ \rho_{i}\left(x\right):=\mathbb{E}\left[D_{ik}\big|i,X_{k}=x\right]. \] To implement this, we fix an individual $i$ in the sample, and regress $D_{ik}$, the indicator function for the link between $i$ and $k$, on the basis functions chosen by the researcher evaluated at observable characteristics $X_{k}$ for all $k\neq i$. To guarantee that a consistent nonparametric estimator of $\rho_{i}$ exists, we need to impose some regularity conditions on $\rho_{i}$. We state the following assumption as an illustrative set of such conditions, acknowledging that there may be many different versions that also work.

assumption[Regularity Conditions for $\rho_{i}$] (i) $Supp\left(X_{i}\right)$ is bounded and convex with nonempty interior; (ii) for each fixed $i$, $\rho_{i}\in{\cal C}_{M}^{d_{x}+1}\left(Supp\left(X_{i}\right)\right)$, where ${\cal C}_{M}^{d_{x}+1}\left(Supp\left(X_{i}\right)\right)$ denotes the class of functions on $Supp\left(X_{i}\right)$ whose derivatives are uniformly bounded by $M$ up to order $d_{x}+1$.

Assumption (ref) essentially requires that $\rho_{i}\left(x\right)$ is smooth enough in $x$. Given that \[ \rho_{i}\left(x\right)=\int\phi\left(w\left(\ol x,x\right)^{'}\beta_{0},a_{i},A_{k}\right)d\mathbb{P}\left(\rest{A_{k}}X_{k}=x\right) \] is an integral of $\phi$ (a strictly increasing function bounded between 0 and $1$) over the conditional distribution of $A_{k}$, Assumption (ref) is easily satisfied, say, if both $\phi$ and the conditional density of $A_{k}$ given $X_{k}=x$ have uniformly bounded derivatives up to order $d+1$, when $w$ is taken to be the coordinate-wise Euclidean distance function.

lemGiven Assumption (ref), for each $i$, there exists an estimator $\hat{\rho}_{i}\in{\cal C}_{M}^{d_{x}+1}\left({\cal X}\right)$ that is $L_{2}\left(\mathbb{P}_{X}\right)$ consistent, i.e., \[ \norm{\hat{\rho}_{i}-\rho_{i}}_{L_{2}\left(\mathbb{P}_{X}\right)}:=\sqrt{\int\left(\hat{\rho}_{i}\left(x\right)-\rho_{i}\left(x\right)\right)^{2}d\mathbb{P}_{X_{k}}\left(x\right)}=o_{p}\left(1\right). \]

Lemma (ref) follows from the large literature on many different types of consistent nonparametric estimators. See Bierens1983kernel for results on kernel estimators and chen2007sieve on sieve estimators. In our simulation and application, we use a spline-based sieve estimator.

In the second step, we use $\hat{\rho}_{i}$ to build the following sample analog of the population criterion $Q^{\gamma}\left(\beta\right)$ in (ref):

equation[equation omitted — 366 chars of source]

The two-step tetrad estimator for $\beta_{0}$ is then defined as

equation[equation omitted — 154 chars of source]

Since $\gamma$ can be arbitrarily chosen as long as it preserves strict positiveness as in Section (ref), we now impose the following continuity assumption for $\gamma$.

assumption[Continuity of $\gamma$] The one-sided sign-preserving function $\gamma$ is Lipschitz-continuous.

Assumption (ref) can be achieved by setting, say, $\gamma\left(t\right):=\left[t\right]_{+}$ or $\gamma\left(t\right):=2\times\Phi\left(\left[t\right]_{+}\right)-1$, where $\left[t\right]_{+}$ is the positive part of $t$, and $\Phi$ is the CDF of the standard normal distribution. See Section (ref) for details. The idea is that, when the difference $\rho_{i}\left(x\right)$ and $\rho_{j}\left(x\right)$ is small, the estimation of whether $\rho_{i}\left(x\right)>\rho_{j}\left(x\right)$ may be relatively imprecise, and therefore, we may wish to downweight such terms in the criterion function.

Computationally, to exploit the topological characteristics of the parameter space $\mathbb{\mathbb{S}}^{d-1}$, i.e. compactness and convexity, we develop a new bisection-style nested rectangle algorithm that recursively shrinks and refines an adaptive grid on the angle space. The key novelty of the algorithm is that instead of working with the edges of the Euclidean parameter space $\mathbb{R}^{d}$, we deterministically “cut” the angle space in each dimension of $\mathbb{\mathbb{S}}^{d-1}$ to search for the area that minimizes $\widehat{Q}_{n}^{\gamma}(\beta)$. Additional measures are taken to ensure the search algorithm is conservative. Simulation and empirical results show that our algorithm performs reasonably well with a relatively small sample size. gao2018robust provides more details regarding the implementation in a panel multinomial choice setting.

We now state the consistency result of the tetrad estimator $\widehat{\beta}_{n}$ under point identification.

thmUnder model (ref) and Assumptions (ref), (ref)--(ref), $\hat{\beta}_{n}$ is consistent for $\beta_{0}$, i.e., \[ \widehat{\beta}_{n}\overset{p}{\longrightarrow}\beta_{0}. \]
remOur estimator shares some similarity with the maximum score estimator: the parameter $\beta$ enters into the sample criterion through indicators of halfspaces about $\beta$, which creates discreteness in the sample criterion. However, our estimator features an additional complication not found in maximum score estimation: we require the first-stage nonparametric estimators $\hat{\rho}_{i}$, and we need to plug the first-stage estimators into nonlinear functions that isolate the “positive side” only: $\gamma\left(\hat{\rho}_{i}\left(X_{k}\right)-\hat{\rho}_{j}\left(X_{k}\right)\right)\gamma\left(\hat{\rho}_{j}\left(X_{l}\right)-\hat{\rho}_{i}\left(X_{l}\right)\right)$ will be nonzero only when $\hat{\rho}_{i}\left(X_{k}\right)>\hat{\rho}_{j}\left(X_{k}\right)$ and $\hat{\rho}_{j}\left(X_{l}\right)>\hat{\rho}_{i}\left(X_{l}\right)$. In fact, it is precisely the “double-threshold” feature of our model that leads to the necessity of a semiparametric two-step estimation procedure, relative to the one-step procedure of the maximum score estimator under a “single-threshold” setting. Given that the asymptotic theory for the maximum score estimator is already nonstandard (with cubic-root rate of convergence and non-normal Chernoff-type asymptotic distribution as in kim1990cube), the addition of the first-stage nonparametric estimation further complicates the asymptotic theory to a highly nontrivial extent. The first-stage nonparametric estimation of $\rho_{i}$ may further slow the rate of convergence below $n^{-1/3}$; in the meanwhile, if we take $\gamma$ to be smooth (but necessarily nonlinear), the term $\gamma\left(\hat{\rho}_{i}\left(X_{k}\right)-\hat{\rho}_{j}\left(X_{k}\right)\right)\gamma\left(\hat{\rho}_{j}\left(X_{l}\right)-\hat{\rho}_{i}\left(X_{l}\right)\right)$ may also provide some effective smoothing on the discrete $\lambda_{ij}$ term, when $\rho_{i}\left(X_{k}\right)-\rho_{j}\left(X_{k}\right)$ or $\rho_{j}\left(X_{l}\right)-\rho_{i}\left(X_{l}\right)$ is closer to zero. It is not exactly clear which effect dominates, or whether the two effects can be balanced, in our current setting. Due to such technical difficulties, we defer the investigation of such types of “two-stage maximum score estimators” in a separate paper by \citet*{gao2020two}, albeit in a simpler setting.

Simulation

In this section, we conduct a simulation study to analyze the finite-sample performance of our two-step tetrad estimator. To begin with, we calculate and plot the identified set $B_{0}$ via (ref) for various support restrictions on $X$\footnote{We thank the Editor for this suggestion.}. The graph illustrates how the size of the identified set $B_{0}$ based on our population criterion (ref) changes when the support of $X$ contains more and more discrete variables. Then, we specify the data generating process (DGP) of the Monte Carlo simulations. We show and discuss the performance of our two-step estimation method under the baseline setup with symmetric $w\left(\cdot,\cdot\right)$ function. In Appendix (ref), we vary the number of individuals $N$, the dimension of the pairwise observable characteristics $d$, and the degree of correlation between $X$ and $A$, as well as allow for asymmetric $w$ function to further examine the robustness of our method.

Identified Set $B_{0}$

In this section, we provide a graphical illustration of the identified set $B_{0}$ for various support restrictions on $X$. The analysis is based on the following network formation model:

equation[equation omitted — 233 chars of source]

where $w$ is taken to be the coordinate-wise absolute difference, i.e., $w_{h}\left(\ol x,\ul x\right):=\left|\ol x_{h}-\ul x_{h}\right|$ for $h=1,...,d$. We set $d_{x}=d=3$ and $\beta_{0}=\left(1,1,1\right)/6^{'}$\footnote{Recall that only the direction of $\beta_{0}$ is identified. Here we divide the vector $\left(1,1,1\right)^{'}$ by 6 to ensure the network is non-degenerate numerically, i.e., not all agents are connected, nor are they all disconnected.}. We incorporate correlation between $X$ and $A$ by drawing $A_{i}=\left(X_{i,1}+\xi_{i}\right)/4$, where $\xi_{i}$ is independently and uniformly distributed on $\left[-0.5,0.5\right]$ and independent of all other variables. $\epsilon_{ij}$ is the exogenous random shock independent of all other variables and uniformly distributed on $\left[0,1\right]$.

The purpose of the exercise is to show how the support restrictions on $X$ affect the size of the identified set $B_{0}$. To this end, we consider the following five support conditions on $X$: (1) all coordinates of $X$ are uniformly distributed on $[-0.5,0.5]$; (2) $X_{i,1}$ is binary with equal probability on $\left\{ 0,1\right\} $, while $X_{i,2}$ and $X_{i,3}$ are both uniformly distributed on $[-0.5,0.5]$; (3) $X_{i,1}$ is binary with equal probability on $\left\{ 0,1\right\} $, $X_{i,2}$ is discrete with equal probability on 11 points of $\left\{ -0.5,-0.4,..,0,..,0.4,0.5\right\} $, and $X_{i,3}$ is uniformly distributed on $[-0.5,0.5]$; (4) all coordinates of $X$ are discrete with equal probability on 101 points of $\left\{ -0.5,-0.49,-0.48..,0,..,0.49,0.5\right\} $; (5) all coordinates of $X$ are discrete with equal probability on 11 points of $\left\{ -0.5,-0.4,..,0,..,0.4,0.5\right\} $.

$\beta_{0}$ is point identified in (1) since the support of $X$ has a nonempty interior, thus satisfying Assumption (ref). In (2)--(5), $\beta_{0}$ is not point identified due to discreteness and boundedness of the support of $X$. Below we compare the sizes of the identified sets when the $X$ vector contains none, one, two and all discrete variables corresponding to DGP (1)--(5).

To calculate the ID set, we utilize the analytical formula of $\rho_{i}$ so that the true $\rho_{i}$ is calculated without error \footnote{We use the distribution of $\left(A,\epsilon,\xi\right)$ to obtain the analytical formula for $\rho_{i}\left(x\right)$.}. Then, we numerically approximate the population criterion by setting a large $N=1,000$ and $M=10,000$. We are able to work with a larger $N$ and $M$ than in the simulations in the next subsection because to get the identified set, we only calculate the maximizer of the population criterion once without the need to estimate $\rho_{i}\left(x\right)$. Still, one can improve the numerical results by increasing $N,M\rightarrow\infty$ and extracting more info from the data when additional computational power is available. Hence, all our results below are conservative: the true ID set must be a subset of the result shown in Figure (ref).

For clarity of illustration, we plot the results in the angle space\footnote{One can transform any $\beta=\left(\beta_{1},\beta_{2},\beta_{3}\right)^{'}$ on the unit sphere $\mathbb{\mathbb{S}}^{2}$ into the angle space by letting $\beta_{1}=\cos\theta_{1}\cos\theta_{2},\beta_{2}=\cos\theta_{1}\sin\theta_{2},\beta_{3}=\sin\theta_{1}$ for $\left(\theta_{1},\theta_{2}\right)\in\left[-\pi/2,\pi/2\right]\times\left[-\pi,\pi\right]$.}. In all cases, we maintain the correlation between $X$ and $A$. The results are summarized in Figure (ref), where we plot the identified sets $B_{0}$ on the full parameter space of $\left[-\pi/2,\pi/2\right]\times\left[-\pi,\pi\right]$ and zoom it in for more details.

center[center omitted — 323 chars of source]

In Figure (ref), the black dot represents the true $\beta_{0}$, which is guaranteed to belong to the identified set $B_{0}$. The red rectangle demonstrates the identified set when the supports of all coordinates of $X$ are continuous and bounded. The theory predicts point identification when the number of individuals $N$ and $\left(i,j\right)$ pairs $M$ go to infinity. Here the numerical result is indeed very close to point identification. The blue rectangle shows the identified set $B_{0}$ when the first dimension of $X$ is binary while all the other coordinates are continuous and bounded. The size of $B_{0}$ is larger than in case (1), which is expected due to the discreteness of $X_{i,1}$. The magenta rectangle corresponds to the case when $X_{i,1}$ is binary, $X_{i,2}$ is discrete, and $X_{i,3}$ is continuous and bounded. The size of $B_{0}$ is reasonably small, given that there are two discrete variables in $X$ and one of them is binary. The green and black rectangles illustrate $B_{0}$ when all coordinates of $X$ are discretely distributed with equal probability on 101 (fine grid) and 11 (coarse grid) points between -0.5 and 0.5, respectively. In these two scenarios, we see larger identified sets than in (1)--(3). That said, the sets still appear small relative to the full parameter space $\left[-\pi/2,\pi/2\right]\times\left[-\pi,\pi\right]$. To summarize, the size of the identified set $B_{0}$ based on logical differencing is reasonably small under each of the five settings.

Performance of the Tetrad Estimator

We maintain (ref) as our network formation model. We draw each coordinate of $X_{i}$ independently from a uniform distribution on $\left[-0.5,0.5\right]$ and calculate $w$ using coordinate-wise absolute difference. Point identification is guaranteed since Assumption (ref) is satisfied. We set $A_{i}=corr\times X_{i,1}+\left(1-corr\right)\times\xi_{i}$, and use the same distributions as in Section (ref) to generate $\xi_{_{i}}$ and $\epsilon_{ij}$. We set the true $\beta_{0}$ to be $(1,...,1)^{'}/\sqrt{d}\in\mathbb{\mathbb{S}}^{d-1}$ , and estimate the direction of $\beta_{0}$.

For the baseline result in the main text, we fix the number of individuals $N=100$, the dimension of $X$ (also $W_{ij}:=w\left(X_{i},X_{j}\right)$ and $\beta_{0}$) $d_{x}=d=3$, and $corr=0.2$. In Appendix (ref), we vary $N$, $d$ and $corr$, as well as allow for asymmetric $w$ function to investigate how robust our method is against various configurations.

To summarize, for each of the $B=100$ simulations we randomly generate data on the characteristics of and the network structure among individuals. Then, based on the observable $\left(X_{i},W_{ij},D_{ij}\right)_{i,j\in\left\{ 1,...,N\right\} }$ matrix we construct our two-step estimator $\widehat{\beta}$ for the true parameter of interest $\beta_{0}$. Specifically, we use a sieve estimator with 2nd-order spline with its knot at median for the first-stage nonparametric estimation of $\rho_{i}\left(\cdot\right)$. The spline is chosen to ensure a relatively small number of regressors in the nonparametric regression considering the small size of $N$. In the second stage, we adapt to the adaptive-gird search on the unit sphere algorithm developed in gao2018robust to calculate $\widehat{\beta}$ that minimizes the sample criterion function $\widehat{Q}\left(\beta\right)$ defined in (ref)\footnote{We suppress its dependence on $\gamma$ and $N$ for notational simplicity. In all simulations and empirical application, we use smoothing function $\gamma\left(t\right):=2\times\Phi\left(\left[t\right]_{+}\right)-1$.} over the unit sphere. It should be noted that, constrained by computational power, when calculating the sample criterion $\widehat{Q}\text{\ensuremath{\left(\beta\right)}}$ for each $\beta\in\mathbb{\mathbb{S}}^{d-1}$ we randomly draw $M=1,000$ $\left(i,j\right)$ pairs of individuals and vary across all possible $\left(k,l\right)$ pairs excluding $i$ or $j$. One can improve those results by increasing $M$ when computational constraint is not present, so again our results are conservative. Lastly, we compare our estimator $\widehat{\beta}$ with $\beta_{0}$ based on several performance metrics including root mean squared error (rMSE), mean norm deviations (MND), and maximum mean absolute deviation (MMAD).

We define for each simulation round $b$ the set estimator $\widehat{\Theta}_{b}$ as the set of points that achieve the minimum of $\widehat{Q}\text{\ensuremath{\left(\beta\right)}}$ over the unit sphere $\mathbb{\mathbb{S}}^{d-1}$. We further define for each simulation $b=1,...,B$ and each dimension $h=1,...,d$ of $\beta$ \[ \widehat{\beta}_{b,h}^{l}:=\min\widehat{\Theta}_{b,h},\quad\widehat{\beta}_{b,h}^{u}:=\max\widehat{\Theta}_{b,h},\quad\widehat{\beta}_{b,h}^{m}:=\dfrac{1}{2}\left(\widehat{\beta}_{b,h}^{l}+\widehat{\beta}_{b,h}^{u}\right), \] where $\widehat{\beta}_{b,h}^{l}$, $\widehat{\beta}_{b,h}^{u}$ and $\widehat{\beta}_{b,h}^{m}$ are the minimum, maximum, and middle point along dimension $h$ for simulation round $b$ of the identified set $\widehat{\Theta}_{b}$, respectively. One can consider $\widehat{\beta}^{m}$ as the point estimator for $\beta_{0}$. Note by construction for each simulation round $b$, the identified set $\widehat{\Theta}_{b}$ is a subset of the rectangle \[ \widehat{\Xi}_{b}:=\times_{h=1}^{d}\left[\widehat{\beta}_{b,h}^{l},\ \widehat{\beta}_{b,h}^{u}\right]. \] We calculate the baseline performance using $\widehat{\beta}^{l},\widehat{\beta}^{u},\widehat{\beta}^{m}$ respectively.

We report in Table (ref) the performance of our estimators. In the first row, we calculate the mean bias across $B=100$ simulations using $\widehat{\beta}^{m}$ along each dimension $h=1,...,d$. The result shows the estimation bias is very small across all dimensions with a magnitude between -0.0053 and 0.0052. Similar performance is observed using $\widehat{\beta}^{u}$ and $\widehat{\beta}^{l}$ as shown in row 2 and 3. We do not find any sign of persistent over/under- estimation of $\beta_{0}$ across each dimension. Row 4 measures the average width along each dimension of the rectangle $\widehat{\Xi}_{b}$ over $B$ simulations. The average size of $\widehat{\Xi}_{b}$ is small, indicating a tight area for the estimated set. In the second part of Table (ref) we report rMSE, MND, and MMAD, all of which are small in magnitude and provide evidence that our estimator works well in finite sample.

table[table omitted — 1,496 chars of source]

Empirical Illustration

As an empirical illustration, we estimate a network formation model under NTU with data of a village network called Nyakatoke in Tanzania. Nyakatoke is a small Haya community of 119 households in 2000 located in the Kagera Region of Tanzania. We are interested in how important factors, such as wealth, distance, and blood or religious ties, are relative to each other in deciding the formation of risk-sharing links among local residents. We apply our two-stage estimator to the Nyakatoke network data and obtain economically intuitive results.

Data Description

The risk-sharing data of Nyakatoke, collected by Joachim De Weerdt in 2000, cover all of the 119 households in the community. It includes the information about whether or not two households are linked in the insurance network. It also provides detailed information on total USD assets and religion of each household, as well as kinship and distance between households. See de2004risk, DeWeerdt2006, and de2011social for more details of this dataset.

To define the dependent variable link, the interviewer asks each household the following question:

$\ $

“Can you give a list of people from inside or outside of Nyakatoke, who you can personally rely on for help and/or that can rely on you for help in cash, kind or labor?”

$\ $

The data contains three answers of “bilaterally mentioned”, “unilaterally mentioned”, and “not mentioned” between each pair of households. Considering the question is about whether one can rely on the other for help, we interpret both “bilaterally mentioned” and “unilaterally mentioned” as they are connected in this undirected network, meaning that the dependent variable $D_{ij}$ link equals 1.\footnote{In the context of the village economies in our application, we think, at the time of link formation, the risk-sharing links are less likely (in comparison with the contexts of business or financial networks) to be driven by efficient arrangements of side-payment transfers, thus satisfying NTU.} We also ran a robustness check by constructing a weighted network based on the answers, i.e. “bilaterally mentioned” means link equals 2, “unilaterally mentioned” means link equals 1, and “not mentioned” means link equals 0, and obtained very similar results.

We estimate the coefficients for 3 regressors: wealth difference, distance and tie between households, with our two-step estimator. Wealth is defined as the total assets in USD owned by each household in 2000, including livestocks, durables and land. Distance measures how far away two households are located in kilometers. Tie is a discrete variable, defined to be 3 if members of one household are parents, children and/or siblings of members of the other household, 2 if nephews, nieces, aunts, cousins, grandparents and grandchildren, 1 if any other blood relation applies or if two households share the same religion, and 0 if no blood religious tie exists. Following the literature we take natural log on \textit{wealth} and \textit{distance}, and we construct the \textit{wealth difference} variable as the absolute difference in \emph{wealth}, i.e., \[ w\left(X_{i},X_{j}\right)=\left(

array[array omitted — 107 chars of source]

\right). \]

Figure (ref) illustrates the structure of the insurance network in Nyakatoke. Each node in the graph represents a household. The solid line between two nodes indicates they are connected, i.e., link equals 1. The size of each node is proportional to the USD wealth of each household. Each node is colored according to its rank in wealth: green for the top quartile, red for the second, yellow for the third and purple for the fourth quartile.

center[center omitted — 228 chars of source]

In the dataset there are 5 households that lack information on wealth and/or distance. We drop these observations, resulting in a sample size $N$ of 114. The total number of ordered household pairs is 12,882. Summary statistics for the dependent and explanatory variables used in our analysis are presented in Table (ref).

table[table omitted — 587 chars of source]

Given the data structure, we should expect close to point identification because even though the tie variable is discrete, the other two regressors wealth and distance can be considered as being continuously distributed with a large conditional support. In addition, tie also has more than one point in its support given other variables. Thus, it is straightforward to verify that Assumption (ref)\emph\textit{\emph{is satisfied, leading to point identification of $\beta_{0}$.}}

Results and Discussion

We apply our two-stage estimator proposed in Section (ref) to the Nyakatoke network data, where, similarly to the simulation studies, we use second degree splines to estimate $\rho_{i}$ and the adaptive grid search algorithm to find the minimizer of the criterion function $\widehat{Q}\left(\beta\right)$ on $\mathbb{\mathbb{S}}^{d-1}$.

table[table omitted — 855 chars of source]

Table (ref) summarizes our estimation results. The column of $\hat{\beta}^{m}$ corresponds to the center of the estimated rectangle \[ \widehat{\Xi}:=\times_{h=1}^{d}\left[\widehat{\beta}_{h}^{l},\ \widehat{\beta}_{h}^{u}\right], \] where $\widehat{\beta}_{h}^{l}$ and $\widehat{\beta}_{h}^{u}$ represent the lower and upper bound of the estimated area (the set of minimizers of sample criterion function (ref)) along each dimension $h=1,..,d$. We use $\hat{\beta}^{m}$ as the point estimator of the coefficients. While the scale of $\beta_{0}$ is unidentified, one can still compare the estimated coefficients with each other to obtain an idea about which variable affects the formation of the link more than the other.

The estimated coefficients for each variable conform well with economic intuition. The estimated set for wealth difference's coefficient is $\left[-0.1964,-0.1932\right]$, which implies the more difference in wealth between two households, the lower likelihood they are connected. The estimated set for distance is $\left[-0.8043,-0.8029\right]$. It is natural that households rely more on neighbors for help than those who live farther away, thus neighbors are more likely to be connected. The estimated coefficient for tie falls in the positive range of $\left[0.5608,0.5630\right]$, consistent with the intuition that one would depend on support from family when negative shock occurs.

It is worth mentioning the estimated set $\hat{\Xi}$ is very tight in each dimension, with the maximum width $\max_{h=1,..,d}\left(\hat{\beta}_{h}^{u}-\hat{\beta}_{h}^{l}\right)$ equal to 0.0032 for wealth difference. Usually the discreteness in tie could make the estimated set wide, but our method is able to leverage the large support in wealth difference and distance. Once again, the relative magnitude and sign of the coefficient for tie are estimated in line with expectation despite of its discreteness. In summary, the empirical results show that our proposed estimator is able to generate economically intuitive estimates under NTU.

Conclusion

This paper considers a semiparametric model of dyadic network formation under nontransferable utilities, a natural and realistic micro-theoretical feature that translates into the lack of additive separability in econometric modeling. We show how a new methodology called logical differencing can be leveraged to cancel out the two-way fixed effects, which correspond to unobserved individual heterogeneity, without relying on arithmetic additivity. The key idea is to exploit the logical implication of weak multivariate monotonicity and use the intersection of mutually exclusive events on the unobserved fixed effects. It would be interesting to explore whether and how the idea of logical differencing, or more generally the use of fundamental logical operations, can be applied to other econometric settings.

The identified sets derived by our method under various support restrictions demonstrate that the proposed method is able to achieve accurate identified set despite of discreteness and boundedness in $Supp\left(X\right)$. Simulation results show that our method performs reasonably well with a relatively small sample size, and robust to various configurations. The empirical illustration using the real network data of Nyakatoke reveals that our method is able to capture the essence of the network formation process by generating estimates that conform well with economic intuition.

This paper also reveals several further research questions regarding dyadic network formation models under the NTU setting. First, given the observation that the NTU setting can capture “homophily effects” with respect to the unobserved heterogeneity (under log-concave error distributions) while imposing monotonicity in the unobserved heterogeneity, it is interesting to investigate whether one can differentiate homophily effects generated by “intrinsic preference” from assortativity effects generated by bilateral consent, NTU and log-concave errors. Second, admittedly the identifying restriction obtained in this paper becomes uninformative when we have antisymmetric pairwise observable characteristics. However, preliminary analysis based on an adaption of gao2018nonparametric to the NTU setting suggests that individual unobserved heterogeneity can be nonparametrically identified up to location and inter-quantile range normalizations. After the identification of individual unobserved heterogeneity terms ($A_{i}$), it becomes straightforward to identify the index parameter $\beta_{0}$ based on the observable characteristics, even in the presence of antisymmetric pairwise characteristics. However, consistent estimators of $A_{i}$ and $\beta_{0}$ in a semiparametric framework based on the identification strategy in gao2018nonparametric are still being developed. We thus leave these research questions to future work.

\paragraph{Acknowledgements}

Wayne Gao is grateful to Xiaohong Chen and Peter C. B. Phillips for their advice and support while working on this paper during his PhD study at Yale University. Ming Li is grateful to Donald W. K. Andrews and Yuichi Kitamura for their advice and support while working on this paper during his PhD study at Yale University. We also thank Isaiah Andrews, Yann Bramoull\'{e}, Benjamin Connault, and Paul Goldsmith-Pinkham, as well as seminar and conference participants at Yale University, University of Arizona, University of Essex, the International Conference for Game Theory at Stony Brook University (2019), the Young Economist Symposium at Columbia University (2019), and the Latin American Meeting of the Econometric Society at Benem\'{e}rita Universidad Aut\'{o}noma de Puebla (2019) for their comments. All errors are our own.