The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
87,145 characters
Logical Differencing in Dyadic Network Formation Models with Nontransferable Utilities
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long\global\long\global\long
\global\long\global\long
\global\long
\global\long
\global\long
\global\long\global\long\global\long\global\long\global\long
\global\long
\global\long\def\mc#1{\mathscr{#1}}
\global\long\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long\global\long
\global\long\def\abs#1{\left|#1\right|}
\global\long\def\norm#1{\left\Vert #1\right\Vert }
\global\long\def\rest#1{\left.#1\right|}
\global\long\def\bracket#1#2{\left\langle #1\middle\vert#2\right\rangle }
\global\long\def\sandvich#1#2#3{\left\langle #1\middle\vert#2\middle\vert#3\right\rangle }
\global\long\def\turd#1{\frac{#1}{3}}
\global\long
\global\long\def\sand#1{\left\lceil #1\right\vert }
\global\long\def\wich#1{\left\vert #1\right\rfloor }
\global\long\def\sandwich#1#2#3{\left\lceil #1\middle\vert#2\middle\vert#3\right\rfloor }
\global\long\def\abs#1{\left|#1\right|}
\global\long\def\norm#1{\left\Vert #1\right\Vert }
\global\long\def\rest#1{\left.#1\right|}
\global\long\def\inprod#1{\left\langle #1\right\rangle }
\global\long\def\ol#1{\overline{#1}}
\global\long\def\ul#1{\underline{#1}}
\global\long\def\td#1{\tilde{#1}}
\global\long
\global\long
\global\long
\global\long
\global\long
\title{Logical Differencing in Dyadic Network Formation Models with Nontransferable
Utilities}
\author{Wayne Yuan Gao\thanks{Gao (corresponding author): Department of Economics, University of
Pennsylvania, 133 S 36th St., Philadelphia, PA 19104, USA, [email removed].}, Ming Li\thanks{Li: Department of Economics, National University of Singapore, 1 Arts
Link AS2, Singapore 117570, [email removed].}, and Sheng Xu\thanks{Xu: Department of Statistics and Data Science, Yale University, 24
Hillhouse Ave., New Haven, CT 06511, USA, [email removed].}}
\maketitle
\begin{abstract}
\noindent This paper considers a semiparametric model of dyadic network
formation under nontransferable utilities (NTU). Such dyadic links
arise frequently in real-world social interactions that require bilateral
consent but by their nature induce additive non-separability. In our
model we show how unobserved individual heterogeneity in the network
formation model can be canceled out without requiring additive separability.
The approach uses a new method we call \emph{logical differencing}.
The key idea is to construct an observable event involving the intersection
of two mutually exclusive restrictions on the fixed effects, while
these restrictions are as necessary conditions of weak multivariate
monotonicity. Based on this identification strategy we provide consistent
estimators of the network formation model under NTU. Finite-sample
performance of our method is analyzed in a simulation study, and an
empirical illustration using the risk-sharing network data from Nyakatoke
demonstrates that our proposed method is able to obtain economically
intuitive estimates.
\end{abstract}
\textbf{Keywords:} dyadic network formation, semiparametric estimation,
nontransferable utilities, additive nonseparability
\newpage{}
\section{\label{sec:C5_Intro}Introduction}
This paper considers a semiparametric model of dyadic network formation
under \emph{nontransferable utilities} (NTU), which arise naturally
in the modeling of real-world social interactions that require bilateral
consent. For instance, friendship is usually formed only when both
individuals in question are willing to accept each other as a friend,
or in other words, when both individuals derive sufficiently high
utilities from establishing the friendship. It is often plausible
that the two individuals may derive very different utilities from
the friendship for a variety of reasons: for example, one of them
may simply be more introvert than the other and derive lower utilities
from the friendship. In addition, there may not be a feasible way
to perfectly transfer utilities between the two individuals. Monetary
payments may not be customary in many social contexts, and even in
the presence of monetary or in-kind transfers, \emph{utilities} may
not be perfectly transferable through these feasible forms of transfers,
say, when individuals have different marginal utilities with respect
to these transfers.\footnote{See surveys by \citet{aumann1967survey}, \citet*{hart1985nontransferable}
and \citet*{mclean2002values} for discussions on the implications
of NTU on link (bilateral relationship) and group formation from a
micro-theoretical perspective.} Given the considerable academic and policy interest in understanding
the underlying drivers of network formation,\footnote{For example, the formation of friendship among U.S. high-school students
has been studied by a long line of literature, such as \citet*{moody2001race},
\citet*{currarini2009economic,currarini2010identifying}, \citet*{boucher2015structural},
\Citet{currarini2016simple}, \citet{xu2018diverse} among others. } it is not only theoretically interesting but also empirically relevant
to incorporate NTU in the modeling of network formation.
This paper contributes to the line of econometric literature on network
formation by introducing and incorporating \emph{nontransferable utilities}
into dyadic network formation models. Previous work in this line of
literature focuses primarily on case of \emph{transferable utilities},
as represented in \citet*{graham2017econometric}, which considers
a parametric model with homophily effects and individual unobserved
heterogeneity of the following form:
\begin{equation}
D_{ij}=\mathbf{\mathbbm1}\left\{ w\left(X_{i},X_{j}\right)^{'}\beta_{0}+A_{i}+A_{j}\geq\epsilon_{ij}\right\} \label{eq:Graham}
\end{equation}
where $D_{ij}$ is an observable binary variable that denotes the
presence or absence of a link between individual $i$ and $j$, $w\left(X_{i},X_{j}\right)$
represents a (symmetric) vector of pairwise observable characteristics
specific to $ij$ generated by a known function $w$ of the individual
observable characteristics $X_{i}$ and $X_{j}$ of $i$ and $j$,
while $A_{i}$ and $A_{j}$ stand for unobserved individual-specific
degree heterogeneity and $\epsilon_{ij}$ is some idiosyncratic utility
shock. Model \eqref{eq:Graham} essentially says that, if the (stochastic)\emph{
joint surplus} generated by a bilateral link $s_{ij}:=w\left(X_{i},X_{j}\right)^{'}\beta_{0}+A_{i}+A_{j}-\epsilon_{ij}$
exceeds the threshold zero, then the link between $i$ and $j$ is
formed. The model implicitly assumes that the link surplus can be
freely distributed among the two individuals $i$ and $j$, and that
bargaining efficiency is always achieved, so that the undirected link
is formed if and only if the link surplus is positive. Given this
specification, \citet*{graham2017econometric} provides consistent
and asymptotically normal maximum-likelihood estimates for the homophily
effect parameter $\beta_{0}$, assuming that the exogenous idiosyncratic
pairwise shocks $\epsilon_{ij}$ are independently and identically
distributed with a logistic distribution. Recently, \citet*{Candelaria2016}
and \citet*{Toth2017} provide semiparametric generalizations of \citet*{graham2017econometric},
while \citet*{gao2018nonparametric} established nonparametric identification
of a class of index models that further generalize \eqref{eq:Graham}.
This paper, however, generalizes \citet*{graham2017econometric} along
a different direction, and seeks to incorporate the natural micro-theoretical
feature of NTU into this class of network formation models. To illustrate\footnote{Starting from Section \ref{sec:NetForm}, we consider a more general
specification than the illustrative model \eqref{eq:Graham} introduced
here.}, consider the following simple adaption of model \eqref{eq:Graham}
with two threshold-crossing conditions:
\begin{equation}
D_{ij}=\mathbf{\mathbbm1}\left\{ w\left(X_{i},X_{j}\right)^{'}\beta_{0}+A_{i}\geq\epsilon_{ij}\right\} \cdot\mathbf{\mathbbm1}\left\{ w\left(X_{i},X_{j}\right)^{'}\beta_{0}+A_{j}\geq\epsilon_{ji}\right\} ,\label{eq:C4_ExNTU}
\end{equation}
where the unobserved individual heterogeneity $A_{i}$ and $A_{j}$
\emph{separately }enter into two different threshold-crossing conditions.
This formulation could be relevant to scenarios where $A_{i}$ represents
individual $i$'s own intrinsic valuation of a generic friend: for
a relatively shy or introvert person $i$, a lower $A_{i}$ implies
that $i$ is less willing to establish a friendship link, regardless
of how sociable the counterparty is. For simplicity, suppose for now
that $w\left(X_{i},X_{j}\right)\equiv{\bf 0}$, $\epsilon_{ij}\sim_{iid}F$,
$\epsilon_{ji}\sim_{iid}F$ and $\epsilon_{ij}\perp\epsilon_{ji}$\footnote{For our general result, we do not require $\epsilon_{ij}\perp\epsilon_{ji}$,
nor the log-concavity of $F$. They are used here for illustration
purpose only. See a discussion after \eqref{eq:example ntu link}.}. Focusing completely on the effects of $A_{i}$ and $A_{j}$, it
is clear that the TU model \eqref{eq:Graham} implies that only the
sum of ``sociability'', $A_{i}+A_{j}$, matters: the linking probability
among pairs with $A_{i}=A_{j}=1$ (two moderately social persons)
should be exactly the same as the linking probability among pairs
with $A_{i}=2$ and $A_{j}=0$ (one very social person and one very
shy person), which might not be reasonable or realistic in social
scenarios. In comparison, the linking probability among pairs with
$A_{i}=2$ and $A_{j}=0$ is lower than the linking probability among
pairs with $A_{i}=A_{j}=1$ under the NTU model \eqref{eq:C4_ExNTU}
with i.i.d. $\epsilon_{ij}$ and $\epsilon_{ji}$ that follow any
log-concave distribution\footnote{A distribution is log-concave if $F\left(x\right)^{\lambda}F\left(y\right)^{1-\lambda}\leq F\left(\lambda x+\left(1-\lambda\right)y\right)$.
Many commonly used distributions, such as uniform, normal, exponential,
logistic, chi-squared distributions, are log-concave. See \citet{bagnoli2005log}
for more details on log-concave distributions from a microeconomic
theoretical perspective.}:
\begin{align*}
& \mathbb{E}\left[\rest{D_{ij}}w\left(X_{i},X_{j}\right)\equiv{\bf 0},A_{i}=2,A_{j}=0\right]\\
=\ & F\left(0\right)F\left(2\right)\\
<\ & F\left(1\right)F\left(1\right)\\
=\ & \mathbb{E}\left[\rest{D_{ij}}w\left(X_{i},X_{j}\right)\equiv{\bf 0},A_{i}=A_{j}=1\right]
\end{align*}
This is intuitive given the observation that, under bilateral consent,
the party with relatively lower utility is the pivotal one in link
formation. Moreover, even though we maintain strict monotonicity in
the unobservable characteristics $A_{i}$ and $A_{j}$, the NTU setting
can still effectively incorporate homophily effects on unobserved
heterogeneity: given that $w\left(X_{i},X_{j}\right)\equiv{\bf 0}$
and $A_{i}+A_{j}=2$, the linking probability is effectively decreasing
in $\left|A_{i}-A_{j}\right|$ under log-concave $F$. Hence, by explicitly
modeling NTU in dyadic network formation, we can accommodate more
flexible or realistic patterns of conditional linking probabilities
and homophily effects that are not present under the TU setting.
However, the NTU setting immediately induces a key technical complication:
as can be seen explicitly in model \eqref{eq:C4_ExNTU}, the observable
indexes, $w\left(X_{i},X_{j}\right)^{'}\beta_{0}$ and $w\left(X_{j},X_{i}\right)^{'}\beta_{0}$,
and the unobserved heterogeneity terms ($A_{i}$ and $A_{j}$) are
no longer additively separable from each other. In particular, notice
that, even though the utility specification for each individual inside
each of the two threshold-crossing conditions in model \eqref{eq:C4_ExNTU}
remains completely linear and additive, the multiplication of the
two (nonlinear) indicator functions directly destroys both linearity
and additive separability, rendering inapplicable most previously
developed econometric techniques that arithmetically ``difference
out'' the ``two-way fixed effects'' $A_{i}$ and $A_{j}$ based
on additive separability.\footnote{Equivalently, one could write model \eqref{eq:C4_ExNTU} in an alternative
form as a ``single'' \emph{composite} threshold-crossing condition:
\[
D_{ij}=\mathbf{\mathbbm1}\left\{ \min\left\{ w\left(X_{i},X_{j}\right)^{'}\beta_{0}+A_{i}-\epsilon_{ij},w\left(X_{j},X_{i}\right)^{'}\beta_{0}+A_{j}-\epsilon_{ji}\right\} \geq0\right\} ,
\]
where additive separability is again lost in this alternative formulation.}
Given this technical challenge, this paper proposes a new identification
strategy termed \emph{logical differencing}, which helps cancel out
the unobserved heterogeneity terms, $A_{i}$ and $A_{j}$, without
requiring additive separability but leveraging the logical implications
of \emph{multivariate monotonicity} in model \eqref{eq:C4_ExNTU}.
The key idea is to construct an observable event involving the intersection
of two mutually exclusive restrictions on the fixed effects $A_{i}$
and $A_{j}$, which logically imply an event that can be represented
without $A_{i}$ or $A_{j}$. Specifically, in the context of the
illustrative model \eqref{eq:C4_ExNTU} above, we start by considering
the event where a given individual $\ol i$ is \emph{more popular}
than another individual $\ol j$ among a group of individuals $k$
with observable characteristics $X_{k}=\ol x$ while $\ol i$ is simultaneously
\emph{less popular }than another individual $\ol j$ among a group
of individuals with a certain realization of observable characteristics
$\ul x$. This is the same as the conditioning event in \citet*{Toth2017}
and analogous to the tetrad comparisons made in \citet{Candelaria2016}.
However, instead of using arithmetic differencing to cancel out the
unobserved heterogeneity $A_{\ol i}$ and $A_{\ol j}$ as in \citet{Candelaria2016}
and \citet*{Toth2017}, we make the following logical deductions based
on the monotonicity of the conditional popularity of $\ol i$ in $w\left(X_{\ol i},\ol x\right)^{'}\beta_{0}$
and $A_{\ol i}$. First, the event that $\ol i$ is \emph{more popular}
than another individual $\ol j$ among the group of individuals with
$X_{k}=\ol x$ implies that either $w\left(X_{\ol i},\ol x\right)^{'}\beta_{0}>w\left(X_{\ol j},\ol x\right)^{'}\beta_{0}$
or $A_{\ol i}>A_{\ol j}$, while the event that $\ol i$ is \emph{less
popular} than another individual $\ol j$ among a different group
of individuals with $X_{l}=\ul x$ implies that either $w\left(X_{\ol i},\ul x\right)^{'}\beta_{0}<w\left(X_{\ol j},\ul x\right)^{'}\beta_{0}$
or $A_{\ol i}<A_{\ol j}$. Second, when both events occur simultaneously,
we can logically deduce that either $w\left(X_{\ol i},\ol x\right)^{'}\beta_{0}>w\left(X_{\ol j},\ol x\right)^{'}\beta_{0}$
or $w\left(X_{\ol i},\ul x\right)^{'}\beta_{0}<w\left(X_{\ol j},\ul x\right)^{'}\beta_{0}$
must have occurred, because $A_{\ol i}>A_{\ol j}$ and $A_{\ol i}<A_{\ol j}$
cannot simultaneously occur. Intuitively, the ``switch'' in the
relative popularity of $\ol i$ and $\ol j$ among the two groups
of individuals with characteristics $\ol x$ and $\ul x$ cannot be
driven by individual unobserved heterogeneity $A_{\ol i}$ and $A_{\ol j}$,
and hence when we indeed observe such a ``switch'', we obtain a
restriction on the parametric indices $w\left(X_{\ol i},\ol x\right)^{'}\beta_{0}$,
$w\left(X_{\ol i},\ol x\right)^{'}\beta_{0}$, $w\left(X_{\ol i},\ul x\right)^{'}\beta_{0}$,
and $w\left(X_{\ol j},\ul x\right)^{'}\beta_{0}$, which helps identify
$\beta_{0}$.
Based on this identification strategy we provide sufficient conditions
for point identification of the parameter $\beta_{0}$ up to scale normalization
as well as a consistent estimator for $\beta_{0}$. Our estimator has
a two-step structure, with the first step being a standard nonparametric
estimator of conditional linking probabilities, which we use to assert
the occurrence of the conditioning event, while in the second step
we use the identifying restriction on $\beta_{0}$ when the conditioning
event occurs. The computation of the estimator essentially follows
the same method proposed in \citet*{gao2018robust}, with some adaptions
to the network data setting. We plot the identified sets under various
restrictions on the support of the observable characteristics $X_{i}$,
analyze the finite-sample performance in a simulation study, and present
an empirical illustration of our method using data from Nyakatoke
on risk-sharing network collected by Joachim De Weerdt.
~
This paper belongs to the line of literature that studies dyadic network
formation in a single large network setting, including \citet*{blitzstein2011},
\citet*{chatterjee2011}, \citet*{yan2013}, \citet*{yan2016}, \citet*{graham2017econometric},
\citet*{charbonneau2017}, \citet*{dzemski2017empirical}, \citet*{jochmans2017semiparametric},
\citet*{yan2018statistical}, \citet*{Candelaria2016}, \citet*{Toth2017}
and \citet*{gao2018nonparametric}. \citet*{shi2016structural} explicitly
incorporates NTU into dyadic network formation models, but \citet*{shi2016structural}
considers a fully parametric model and establishes the consistency
and asymptotic normality of the maximum likelihood estimators. See
also the recent surveys by \citet*{de2020econometric} and \citet*{graham2020dyadic}.
This paper is also related to a line of research that utilizes dyadic
link formation models in order to study structural social interaction
models: for instance, \citet{arduini2015parametric}, \citet*{auerbach2016identification},
\citet*{goldsmith2013social}, \citet*{hsieh2016social} and \citet*{johnsson2021estimation}.
In these papers, the social interaction models are the main focus
of identification and estimation, while the link formation models
are used mainly as a tool (a control function) to deal with network
endogeneity or unobserved heterogeneity problems in the social interaction
model. Even though some of the network formation models considered
in this line of literature is consistent with the NTU setting, this
line of literature is usually not primarily concerned with the full
identification and estimation of the network formation model itself.
It should be pointed out that in this paper we do not consider link
interdependence in network formation, which is studied by the line
of econometric literature on strategic network formation models. This
line of literature primarily uses pairwise stability \citep*{Jackson1996}
as the solution concept for network formation, and also often builds
NTU into the econometric specification. See, for example, \citet*{de2018identifying},
\citet*{graham2016homophily}, \citet*{leung2015random}, \citet*{menzel2015b},
\citet{boucher2017my}, \citet*{mele2017dense}, \citet*{mele2017structural}
and \citet*{ridder2017estimation}. However, this type of models usually
do not feature unobserved heterogeneity as in this paper. See, for
example, \citet*{de2020strategic} for a more detailed survey on this
line of literature.
This paper is also closely related to to \citet*{gao2018robust},
which similarly leverages multivariate monotonicity in a multi-index
structure under a panel multinomial choice setting. It should be pointed
out that, even though there is some structural similarity between
network data and panel data, there are no direct ways in the network
setting to make ``intertemporal comparison'' as in the panel setting,
which holds the fixed effects unchanged across two observable periods
of time. It is precisely this additional complication induced by the
network setting that requires the technique of logical differencing
proposed in this paper.
~
The rest of the paper is organized as follows. In Section \ref{sec:NetForm},
we describe the general specifications of our dyadic network formation
model. Section \ref{sec:ID_Est} establishes identification of the
parameter of interests in our model and provides a consistent tetrad
estimator. We plot the identified sets under various restrictions
on $X$ and report baseline simulation results in Section \ref{sec:C5_simu}.
We present an empirical illustration using the risk-sharing data of
Nyakatoke in Section \ref{sec:Emp}. Section \ref{sec:C5_conc} concludes.
Proofs and additional simulation results are available in the Appendix.
\section{\label{sec:NetForm}A Nonseparable Dyadic Network Formation Model}
We consider the following dyadic network formation model:
\begin{equation}
\mathbb{E}\left[\rest{D_{ij}}X_{i},X_{j},A_{i},A_{j}\right]=\phi\left(w\left(X_{i},X_{j}\right)^{'}\beta_{0},A_{i},A_{j}\right)\label{eq:Model_NF}
\end{equation}
where:
\begin{itemize}
\item $i\in\left\{ 1,...,n\right\} $ denote a generic individual in a group
of $n$ individuals.
\item $X_{i}$ is a $\mathbb{R}^{d_{x}}$-valued vector of observable characteristics
for individual $i$. This could include, for example, wealth, age,
education and ethnicity of individual $i$.
\item $D_{ij}$ denotes a binary observable variable that indicates the
presence or absence of an undirected and unweighted link between two
distinct individuals $i$ and $j$: $D_{ij}=D_{ji}$ for all pairs
of individuals $ij$, with $D_{ij}=1$ indicating that $ij$ are linked
while $D_{ij}=0$ indicating that $ij$ are not linked.
\item $w:\mathbb{R}^{d_{x}}\times\mathbb{R}^{d_{x}}\to\mathbb{R}^{d}$ is a known function that
is \emph{symmetric}\footnote{Our method can also be adapted to the case with \emph{asymmetric}
$w$. See Remark \ref{rem:AymW}.} with respect to its two vector arguments. We will write $W_{ij}:=w\left(X_{i},X_{j}\right)$
for notational simplicity.
\item $\beta_{0}\in\mathbb{R}^{d}$ is an unknown finite-dimensional parameter of interest.
Assume $\beta_{0}\neq{\bf 0}$ so that we may normalize $\norm{\beta_{0}}=1$,
i.e., $\beta_{0}\in\mathbb{\mathbb{S}}^{d-1}$.
\item $A_{i}$ is an unobserved scalar-valued variable that represents unobserved
individual heterogeneity.
\item $\phi:\mathbb{R}^{3}\to\mathbb{R}$ is an unknown measurable function that is symmetric
with respect to its second and third arguments.
\end{itemize}
In addition, we impose the following two assumptions:
\begin{assumption}[Monotonicity]
\label{assu:C5_mon} $\phi$ is weakly increasing in each of its
arguments.
\end{assumption}
Assumption \ref{assu:C5_mon} is the key assumption on which our identification
analysis is based. It requires that the conditional linking probability
between individuals with characteristics $\left(X_{i},A_{i}\right)$
and $\left(X_{j},A_{j}\right)$ be monotone in a parametric index
$\delta_{ij}:=W_{ij}^{'}\beta_{0}$ as well as the unobserved individual
heterogeneity terms $A_{i}$ and $A_{j}$. It should be noted that,
given monotonicity, increasingness is without loss of generality as
$\phi$, $\beta_{0}$ and $A_{i},A_{j}$ are all unknown or unobservable.
In addition, Assumption \ref{assu:C5_mon} only requires that $\phi$
is monotonic in the index $W_{ij}^{'}\beta_{0}$ as a whole, not individual
coordinates of $W_{ij}$. Therefore, we may include nonlinear or non-monotone
functions $w\left(\cdot,\cdot\right)$ on the observable characteristics
as long as Assumption \ref{assu:C5_mon} is maintained.
Next, we impose a standard random sampling assumption:
\begin{assumption}[Random Sampling]
\label{assu:DNF-rs} $\left(X_{i},A_{i}\right)$ is i.i.d. across
$i\in\left\{ 1,...,n\right\} $.
\end{assumption}
In particular, Assumption \ref{assu:DNF-rs} allows arbitrary dependence
structures between the observable characteristics $X_{i}$ and the
unobservable characteristic $A_{i}$.
~
Model \eqref{eq:Model_NF} along with the specifications and the two
assumptions introduced above encompass a large class of dyadic network
formation models in the literature. For example, the standard dyadic
network formation model \eqref{eq:Graham} studied by \citet*{graham2017econometric}
can be written as
\[
\mathbb{E}\left[\rest{D_{ij}}X_{i},X_{j},A_{i},A_{j}\right]=F\left(W_{ij}^{'}\beta_{0}+A_{i}+A_{j}\right)
\]
where $F$ is the CDF of the standard logistic distribution. For the
semiparametric version considered by \citet{Candelaria2016}, \citet{Toth2017},
and \citet{gao2018nonparametric}, we can simply take $F$ to be some
unknown CDF. In either case, the monotonicity of the CDF $F$ and
the additive structure of $W_{ij}^{'}\beta_{0}+A_{i}+A_{j}$ immediately
imply Assumption \ref{assu:C5_mon}.
However, our current model specification and assumptions further incorporate
a larger class of dyadic network formation models with potentially
nontransferable utilities. Specifically, consider the joint requirement
of two threshold-crossing conditions
\begin{align}
D_{ij} & =\mathbf{\mathbbm1}\left\{ u\left(W_{ij}^{'}\beta_{0},A_{i},A_{j},\epsilon_{ij}\right)\geq0\right\} \cdot\mathbf{\mathbbm1}\left\{ u\left(W_{ji}^{'}\beta_{0},A_{j},A_{i},\epsilon_{ji}\right)\geq0\right\} ,\label{eq:Model_NFNTU}
\end{align}
where $u$ is an unknown function that is not necessarily symmetric
with respect to its second and third arguments $\left(A_{i},A_{j}\right)$,
and $\left(\epsilon_{ij},\epsilon_{ji}\right)$ are idiosyncratic pairwise shocks
that are i.i.d. across each unordered $ij$ pair with some unknown
distribution. In particular, notice that model \eqref{eq:C4_ExNTU}
is a special case of \eqref{eq:Model_NFNTU}. Suppose we further impose
the following two lower-level assumptions 1a and 1b:
\begin{assumption*}[\textbf{1a}]
$\left(\epsilon_{ij},\epsilon_{ji}\right)$ are independent of $\left(X_{i},A_{i},X_{j},A_{j}\right)$.
\end{assumption*}
\begin{assumption*}[\textbf{1b}]
$u$ is weakly increasing in its first three arguments.
\end{assumption*}
Then, the conditional linking probability
\begin{align}
& \ \mathbb{E}\left[\rest{D_{ij}}X_{i},X_{j},A_{i},A_{j}\right]\nonumber \\
= & \ \int\mathbf{\mathbbm1}\left\{ u\left(W_{ij}^{'}\beta_{0},A_{i},A_{j},\epsilon_{ij}\right)\geq0\right\} \cdot\mathbf{\mathbbm1}\left\{ u\left(W_{ji}^{'}\beta_{0},A_{j},A_{i},\epsilon_{ji}\right)\geq0\right\} d\mathbb{P}\left(\epsilon_{ij},\epsilon_{ji}\right)\nonumber \\
=: & \ \phi\left(W_{ij}^{'}\beta_{0},A_{i},A_{j}\right)\label{eq:example ntu link}
\end{align}
can be represented by model \eqref{eq:Model_NF} with Assumption \ref{assu:C5_mon}
satisfied.
In particular, we do not require $\epsilon_{ij}\perp\epsilon_{ji}$. In fact,
$\epsilon_{ij}\equiv\epsilon_{ji}$ is readily incorporated in our model. Under
the maintained assumption that $w\left(X_{i},X_{j}\right)=w\left(X_{j},X_{i}\right)$,
if $\epsilon_{ij}\equiv\epsilon_{ji}$ and $u$ is furthermore assumed to be symmetric
with respect to its second and third arguments ($A_{i}$ and $A_{j}$),
then our model specializes to the case of transferable utilities,
\[
D_{ij}=\mathbf{\mathbbm1}\left\{ u\left(W_{ij}^{'}\beta_{0},A_{i},A_{j},\epsilon_{ij}\right)\geq0\right\} ,
\]
where effectively only one threshold crossing condition determines
the establishment of a given network link. Therefore, our NTU model
\eqref{eq:Model_NF} includes the TU model as a special case.
\begin{rem}[Symmetry of $w$]
\label{rem:AymW} To explain the key idea of our identification strategy
in a notation-economical way, we will be focusing on the case of symmetric
$w$ in most of the following sections. However, it should be pointed
out that our method can also be applied to the case where $w$ is
allowed to be asymmetric in \eqref{eq:Model_NFNTU}, so that individual
utilities based on observable characteristics can also be made asymmetric
(nontransferable). In that case, model \eqref{eq:Model_NFNTU} needs
to be modified as
\begin{equation}
\mathbb{E}\left[\rest{D_{ij}}X_{i},X_{j},A_{i},A_{j}\right]=\phi\left(W_{ij}^{'}\beta_{0},W_{ji}^{'}\beta_{0},A_{i},A_{j}\right),\label{eq:Model_AsymW}
\end{equation}
where $W_{ij}=w\left(X_{i},X_{j}\right)$ may be different from $W_{ji}=w\left(X_{j},X_{i}\right)$,
but $\phi$ is symmetric with respect to its first two arguments $W_{ij}^{'}\beta_{0},W_{ji}^{'}\beta_{0}$\emph{
whenever} $A_{i}=A_{j}$ Moreover, Assumption \ref{assu:C5_mon} should
also be changed to be $\phi$ is monotone in all its \emph{four} arguments.
See Appendix \ref{subsec:ID_asym} for a more detailed discussion
on how our identification strategy can be adapted to accommodate asymmetric
$w$ under appropriate conditions.
\end{rem}
\section{\label{sec:ID_Est}Identification and Estimation}
\subsection{\label{subsec:DNF_ID}Identification via Logical Differencing}
In this section, we explain the key idea of our identification strategy.
We construct a mutually exclusive event to cancel out the unobservable
heterogeneity $A_{i}$ and $A_{j}$, which leads to an identifying
restriction on $\beta_{0}$. We call this technique ``\textit{logical
differencing}''.
For each fixed individual $\ol i$, and each possible $\ol x\in\mathbb{R}^{d_{x}}$,
define
\begin{equation}
\rho_{\ol i}\left(\ol x\right):=\mathbb{E}\left[\rest{D_{\ol ik}}X_{k}=\ol x\right]\label{eq:define rho_i(x)}
\end{equation}
as the linking probability of this specific individual $\ol i$ with
a group of individuals, individually indexed by $k$, with the same
observable characteristics $X_{k}=\ol x$ (but potentially different
fixed effects $A_{k}$). Clearly, $\rho_{i}\left(\ol x\right)$ is
directly identified from data in a single large network.
Suppose that individual $\ol i$ has observed characteristics $X_{\ol i}=x_{\ol i}$
and unobserved characteristics $A_{\ol i}=a_{\ol i}$. Then, by model
\eqref{eq:Model_NF} we have
\begin{align}
\rho_{\ol i}\left(\ol x\right) & =\mathbb{E}\left[\rest{\mathbb{E}\left[\rest{D_{\ol ik}}X_{k}=\ol x,A_{k},X_{\ol i}=x_{\ol i},A_{\ol i}=a_{\ol i}\right]}X_{k}=\ol x\right]\nonumber \\
& =\mathbb{E}\left[\rest{\phi\left(w\left(x_{\ol i},\ol x\right)^{'}\beta_{0},a_{\ol i},A_{k}\right)}X_{k}=\ol x\right]\nonumber \\
& =:\psi_{\ol x}\left(w\left(x_{\ol i},\ol x\right)^{'}\beta_{0},a_{\ol i}\right),\label{eq:C5_rhorho}
\end{align}
where the expectation in the second to last line is taken over $A_{k}$
conditioning on $X_{k}=\ol x$. As we allow $A_{k}$ and $X_{k}$
to be arbitrarily correlated, the $\psi_{\ol x}$ function defined
in the last line of \eqref{eq:C5_rhorho} is dependent on $\ol x$.
In the same time, notice that $\psi_{\ol x}$ does not depend on the
identity of $\ol i$ beyond the values of $w\left(x_{\ol i},\ol x\right)^{'}\beta_{0}$
and $a_{\ol i}$. By Assumption \ref{assu:C5_mon}, $\psi_{\ol x}\left(w\left(x_{\ol i},\ol x\right)^{'}\beta_{0},a_{\ol i}\right)$
must be bivariate weakly increasing in the index $w\left(x_{\ol i},\ol x\right)^{'}\beta_{0}$
and the unobserved heterogeneity scalar $a_{\ol i}$. We now show
how to use the bivariate monotonicity to obtain identifying restrictions
on $\beta_{0}$.
~
Fixing two distinct individuals $\ol i$ and $\ol j$ in the population,
we first consider the event that \textit{individual} $\ol i$ is \emph{strictly
more popular than} \textit{individual} $\ol j$ among the \textit{group}
of individuals with observed characteristics $X_{k}=\ol x$:
\begin{equation}
\rho_{\ol i}\left(\ol x\right)>\rho_{\ol j}\left(\ol x\right),\label{eq:rho(ik) >jk}
\end{equation}
which is an event directly identifiable from observable data given
\eqref{eq:define rho_i(x)}. Even though event \eqref{eq:rho(ik) >jk}
is the same conditioning event as considered in \citet*{Toth2017}
and analogous to the tetrad comparisons made in \citet{Candelaria2016},
we now exploit the following logical deduction based on the bivariate
monotonicity of the conditional popularity of $\ol i$ in $w\left(X_{\ol i},\ol x\right)^{'}\beta_{0}$
and $A_{\ol i}$ without the assumption of additivity between them.
Specifically, writing $\left(x_{\ol i},a_{\ol i}\right)$ and $\left(x_{\ol j},a_{\ol j}\right)$
as the observable and unobservable characteristics of individuals
$\ol i$ and $\ol j$, by \eqref{eq:C5_rhorho} we have
\begin{align}
& \rho_{\ol i}\left(\ol x\right)>\rho_{\ol j}\left(\ol x\right).\nonumber \\
\Leftrightarrow\quad & \psi_{\ol x}\left(w\left(x_{\ol i},\ol x\right)^{'}\beta_{0},a_{\ol i}\right)>\psi_{\ol x}\left(w\left(x_{\ol j},\ol x\right)^{'}\beta_{0},a_{\ol j}\right)\nonumber \\
\Rightarrow\quad & \left\{ w\left(x_{\ol i},\ol x\right)^{'}\beta_{0}>w\left(x_{\ol j},\ol x\right)^{'}\beta_{0}\right\} \text{ OR }\left\{ a_{\ol i}>a_{\ol j}\right\} ,\label{eq:C5_contra}
\end{align}
Note that the last line of equation \eqref{eq:C5_contra} is a natural
necessary (but not sufficient) condition for $\rho_{\ol i}\left(\ol x\right)>\rho_{\ol j}\left(\ol x\right)$
under bivariate monotonicity.
~
Now, consider the event that \textit{individual} $\ol i$ is \emph{strictly
less popular than} \textit{individual} $\ol j$ among the \textit{group}
of individuals with observed characteristics $X_{k}=\underline{x}$,
i.e.,
\begin{equation}
\rho_{\ol i}\left(\ul x\right)<\rho_{\ol j}\left(\ul x\right).\label{eq:rho_i(l) < rho_j(l)}
\end{equation}
Then, by a similar argument to \eqref{eq:C5_contra}, we deduce
\begin{equation}
\rho_{\ol i}\left(\ul x\right)<\rho_{\ol j}\left(\ul x\right)\quad\Rightarrow\quad\left\{ w\left(x_{\ol i},\ul x\right)^{'}\beta_{0}<w\left(x_{\ol j},\ul x\right)^{'}\beta_{0}\right\} \text{ OR }\left\{ a_{\ol i}<a_{\ol j}\right\} .\label{eq:C5_contra2}
\end{equation}
Notice that the event $\left\{ a_{\ol i}<a_{\ol j}\right\} $ in \eqref{eq:C5_contra2}
is mutually exclusive with the event $\left\{ a_{\ol i}>a_{\ol j}\right\} $
that shows up in \eqref{eq:C5_contra}.
~
Next, consider the event that the two events \eqref{eq:rho(ik) >jk}
and \eqref{eq:rho_i(l) < rho_j(l)} described above\emph{ simultaneously
happen}. Then, by \eqref{eq:C5_contra}, \eqref{eq:C5_contra2} and
basic logical operations, we have
\begin{align}
& \text{\ensuremath{\phantom{\text{ OR }}} }\left\{ \rho_{\ol i}\left(\ol x\right)>\rho_{\ol j}\left(\ol x\right)\right\} \text{ AND }\left\{ \rho_{\ol i}\left(\ul x\right)<\rho_{\ol j}\left(\ul x\right)\right\} \nonumber \\
\Rightarrow\quad & \text{\ensuremath{\phantom{\text{ OR }}} }\left(\left\{ w\left(x_{\ol i},\ol x\right)^{'}\beta_{0}>w\left(x_{\ol j},\ol x\right)^{'}\beta_{0}\right\} \text{ OR }\left\{ a_{\ol i}>a_{\ol j}\right\} \right)\nonumber \\
& \ \text{AND\ \ensuremath{\left(\left\{ w\left(x_{\ol i},\ul x\right)^{'}\beta_{0}<w\left(x_{\ol j},\ul x\right)^{'}\beta_{0}\right\} \text{ OR }\left\{ a_{\ol i}<a_{\ol j}\right\} \right)}}\nonumber \\
\Leftrightarrow\quad & \phantom{\text{ OR }}\left(\left\{ w\left(x_{\ol i},\ol x\right)^{'}\beta_{0}>w\left(x_{\ol j},\ol x\right)^{'}\beta_{0}\right\} \ \text{AND}\ \left\{ w\left(x_{\ol i},\ul x\right)^{'}\beta_{0}<w\left(x_{\ol j},\ul x\right)^{'}\beta_{0}\right\} \right)\nonumber \\
& \text{ OR }\left(\left\{ w\left(x_{\ol i},\ol x\right)^{'}\beta_{0}>w\left(x_{\ol j},\ol x\right)^{'}\beta_{0}\right\} \ \text{AND}\ \left\{ a_{\ol i}<a_{\ol j}\right\} \right)\nonumber \\
& \text{ OR }\left(\left\{ a_{\ol i}>a_{\ol j}\right\} \ \text{AND}\ \left\{ w\left(x_{\ol i},\ul x\right)^{'}\beta_{0}<w\left(x_{\ol j},\ul x\right)^{'}\beta_{0}\right\} \right)\nonumber \\
& \text{ OR }\left(\left\{ a_{\ol i}>a_{\ol j}\right\} \ \text{AND}\ \left\{ a_{\ol i}<a_{\ol j}\right\} \right)\nonumber \\
\Rightarrow\quad & \phantom{\text{ OR }}\left(\left\{ w\left(x_{\ol i},\ol x\right)^{'}\beta_{0}>w\left(x_{\ol j},\ol x\right)^{'}\beta_{0}\right\} \ \text{AND}\ \left\{ w\left(x_{\ol i},\ul x\right)^{'}\beta_{0}<w\left(x_{\ol j},\ul x\right)^{'}\beta_{0}\right\} \right)\nonumber \\
& \text{ OR }\left\{ w\left(x_{\ol i},\ol x\right)^{'}\beta_{0}>w\left(x_{\ol j},\ol x\right)^{'}\beta_{0}\right\} \nonumber \\
& \text{ OR }\left\{ w\left(x_{\ol i},\ul x\right)^{'}\beta_{0}<w\left(x_{\ol j},\ul x\right)^{'}\beta_{0}\right\} \nonumber \\
\Leftrightarrow\quad & \left\{ \left(w\left(x_{\ol i},\ol x\right)-w\left(x_{\ol j},\ol x\right)\right)^{'}\beta_{0}>0\right\} \ \text{OR}\ \left\{ \left(w\left(x_{\ol i},\ul x\right)-w\left(x_{\ol j},\ul x\right)\right)^{'}\beta_{0}<0\right\} ,\label{eq:C5_logicaldiff}
\end{align}
The derivations above exploit two simple logical properties: first,
\[
\left\{ a_{\ol i}>a_{\ol j}\right\} \ \text{AND}\ \left\{ a_{\ol i}<a_{\ol j}\right\} \quad=\quad\text{FALSE},
\]
and second,
\[
\left\{ w\left(x_{\ol i},\ol x\right)^{'}\beta_{0}>w\left(x_{\ol j},\ol x\right)^{'}\beta_{0}\right\} \ \text{AND}\ \left\{ a_{\ol i}<a_{\ol j}\right\} \quad\Rightarrow\quad\left\{ w\left(x_{\ol i},\ol x\right)^{'}\beta_{0}>w\left(x_{\ol j},\ol x\right)^{'}\beta_{0}\right\} ,
\]
which uses only necessary but not sufficient condition, so that we
can obtain an identifying restriction \eqref{eq:C5_logicaldiff} on
$\beta_{0}$ that does not involve $a_{\ol i}$ nor $a_{\ol j}$. These
two forms of logical operations together enable us to ``difference
out'' (or ``cancel out'') the unobserved heterogeneity terms $a_{\ol i}$
and $a_{\ol j}$.
In contrast with various forms of ``\emph{arithmetic differencing}''
techniques proposed in the econometric literature (including \citealp{Candelaria2016}
and \citealp{Toth2017} specific to the dyadic network formation literature),
our proposed technique does \emph{not} rely on additive separability
between the parametric index $w\left(x_{\ol i},\ol x\right)^{'}\beta_{0}$
and the unobserved heterogeneity term $a_{\ol i}$. Instead, our identification
strategy is based on multivariate monotonicity and utilizes logical
operations rather than standard arithmetic differencing to cancel
out the unobserved heterogeneity terms. Hence, we term our method
``\emph{logical differencing}''.
~
The identifying arguments above are derived for a fixed pair of individuals
$\ol i$ and $\ol j$, but clearly the arguments can be applied for
any pair of individuals $\left(i,j\right)$ with observable characteristics
$x_{i}$ and $x_{j}$. Writing
\begin{align*}
\tau_{ij}\left(\ol x,\ul x\right) & :=\mathbf{\mathbbm1}\left\{ \rho_{i}\left(\ol x\right)>\rho_{j}\left(\ol x\right)\right\} \cdot\mathbf{\mathbbm1}\left\{ \rho_{i}\left(\ul x\right)<\rho_{j}\left(\ul x\right)\right\} \text{ }\text{and}\\
\lambda\left(\ol x,\ul x;x_{i},x_{j};\beta\right) & :=\mathbf{\mathbbm1}\left\{ \left(w\left(x_{i},\ol x\right)-w\left(x_{j},\ol x\right)\right)^{'}\beta_{0}\leq0\right\} \cdot\mathbf{\mathbbm1}\left\{ \left(w\left(x_{i},\ul x\right)-w\left(x_{j},\ul x\right)\right)^{'}\beta_{0}\geq0\right\}
\end{align*}
for each $\beta\in\mathbb{\mathbb{S}}^{d-1}$, we summarize the identifying arguments
above by the following lemma.
\begin{lem}[Identifying Restriction]
\label{lem:DNF1}Under model \eqref{eq:Model_NF} and Assumptions
\ref{assu:C5_mon} and \ref{assu:DNF-rs}, we have
\begin{equation}
\tau_{ij}\left(\ol x,\ul x\right)=1\quad\Rightarrow\quad\lambda\left(\ol x,\ul x;x_{i},x_{j};\beta_{0}\right)=0.\label{eq:ID_Rest}
\end{equation}
\end{lem}
A simple (but clearly not unique) way to build a criterion function
based on Lemma \ref{lem:DNF1} is to define
\begin{equation}
Q\left(\beta\right):=\mathbb{E}_{ij,kl}\left[\tau_{ij}\left(X_{k},X_{l}\right)\lambda\left(X_{k},X_{l};X_{i},X_{j};\beta\right)\right],\label{eq:pop_cri}
\end{equation}
where the expectation is $\mathbb{E}_{ij,kl}$ taken over random samples of
ordered tetrads $\left(i,j,k,l\right)$ from the population, and $\left(X_{i},X_{j},X_{k},X_{l}\right)$
denote the random variables corresponding to the observable characteristics
of $\left(i,j,k,l\right)$. According to Lemma \ref{lem:DNF1}, $Q\left(\beta_{0}\right)=0$,
which is always smaller than or equal to $Q\left(\beta\right)\geq0=Q\left(\beta_{0}\right)$
for any $\beta\neq\beta_{0}$ because $\tau_{ij}\geq0$ and $\lambda\geq0$
by construction.
Observing that the scale of $\beta_{0}$ is never identified, we write
\[
B_{0}:=\left\{ \beta\in\mathbb{\mathbb{S}}^{d-1}:Q\left(\beta\right)=0\right\}
\]
to represent the normalized ``identified set'' relative to the criterion
$Q$ defined in \eqref{eq:pop_cri}. Lemma \ref{lem:DNF1} implies
that $\beta_{0}\in B_{0}$, but in general there is no guarantee that
$B_{0}$ is a singleton. The next subsection contains a set of sufficient
conditions that guarantees $B_{0}=\left\{ \beta_{0}\right\} $.
\begin{rem}
We should point out that the identified set $B_{0}$ defined above
based on logical differencing is not \emph{sharp} in general, since
the individual unobserved heterogeneity term $A_{i}$ is \emph{canceled
out by logical differencing. }In fact, one can show $A_{i}$ is also
identified (up to proper normalization) under certain conditions\footnote{The identification of $A_{i}$ is conceptually analogous to the identification
of individual fixed effects in a long panel setting.}, and knowledge about $A_{i}$ can help with the identification of
$\beta_{0}$. In fact, when all realizations of $A_{i}$ are point identified,
the point identification of $\beta_{0}$ can be established under much
weaker conditions than those to be presented in Section \ref{subsec:PointID}.\footnote{See \citet*{gao2018nonparametric} for a related discussion. In fact,
the identification strategy in \citet*{gao2018nonparametric} can
be adapted to establish identification of $A_{i}$ under the NTU setting.
However, a rigorous presentation of such identification results is
beyond the scope of this paper.} However, we trade sharpness for simplicity: this paper provides a
method of identification (and estimation) without the need to deal
with the incidental parameters $A_{i}$.
\end{rem}
~
It is worth mentioning that for any one-sided sign preserving function
$\gamma$ such that
\begin{equation}
\gamma\left(t\right)\ \begin{cases}
\geq0, & \text{for }t>0,\\
=0, & \text{for }t\leq0.
\end{cases}\label{eq:gamma_sign}
\end{equation}
we may define
\begin{align}
\tau_{ij}^{\gamma}\left(\ol x,\ul x\right) & :=\gamma\left(\rho_{i}\left(\ol x\right)-\rho_{j}\left(\ol x\right)\right)\cdot\gamma\left(\rho_{j}\left(\ul x\right)-\rho_{i}\left(\ul x\right)\right),\label{eq:smooth tau}\\
Q^{\gamma}\left(\beta\right) & :=\mathbb{E}_{ij,kl}\left[\tau_{ij}^{\gamma}\left(X_{k},X_{l}\right)\lambda\left(X_{k},X_{l};X_{i},X_{j};\beta\right)\right],\label{eq:smooth_Q}
\end{align}
without changing the identification set at all, since
\[
\tau_{ij}\left(\ol x,\ul x\right)>0\quad\text{if and only if}\quad\tau_{ij}^{\gamma}\left(\ol x,\ul x\right)>0.
\]
In fact, $\tau^{\gamma}$ specializes to $\tau$ when we set $\gamma\left(t\right):=\mathbf{\mathbbm1}\left\{ t>0\right\} $.
Alternatively, we may set $\gamma$ to be ``smoother'', say, $\gamma\left(t\right):=\left[t\right]_{+}$
the positive part function.
Such forms of ``smoothing'' in the population criterion will be
irrelevant to all the identification results and its proofs in this
paper, so for notational simplicity, we will suppress $\gamma$ and focus
on the representative $\tau_{ij}$ and $Q$ in the next subsection
about point identification. However, a smooth $\gamma$ will play a role
when it comes to estimation and computation, and we will revisit $\gamma$
in Section \ref{subsec:DNF_est}.
\subsection{\label{subsec:PointID}Sufficient Conditions for Point Identification}
We now present a set of sufficient conditions that guarantee point
identification of $\beta_{0}$ on $\mathbb{\mathbb{S}}^{d-1}$.
\medskip{}
\noindent \textbf{Assumption 1$'$}
\protected@write \@auxout {}{\string \newlabel {as1prime}{{$1'$}{\thepage}{$1'$}{as1prime}{}} }
\hypertarget{as1prime}{}
(Strict Monotonicity of $\phi$)\textbf{.} $\phi$ defined in model
\eqref{eq:Model_NF} is weakly increasing in $A_{i}$ and $A_{j}$
while strictly increasing in the index $w\left(X_{i},X_{j}\right)^{'}\beta_{0}$.
\medskip{}
Assumption \ref{as1prime} strengthens Assumption \ref{assu:C5_mon}
by requiring that $\phi$ be strictly increasing in the parametric
index $w\left(X_{i},X_{j}\right)^{'}\beta_{0}$. This is used to guarantee
that differences in the parametric index can indeed lead to changes
in conditional linking probabilities, so that the conditional event
$\left\{ \tau_{ij}\left(\ol x,\ul x\right)=1\right\} $ in Lemma \ref{lem:DNF1}
may occur with strictly positive probability.
\begin{assumption}[Continuity of $\phi$ and $w$]
\label{assu:DNF-con} $\phi$ and $w$ are continuous functions on
their domains.
\end{assumption}
\begin{assumption}[Sufficient Directional Variations]
\label{assu:DNF-supportX} There exist distinct points $\ol x$ and
$\ul x$ in $Supp\left(X_{i}\right)$, such that the vector ${\bf 0}$
lies in the interior of $Supp\left(w\left(\ol x,X_{i}\right)-w\left(\ul x,X_{i}\right)\right)$.
\end{assumption}
We also provide a lower-level condition for Assumption \ref{assu:DNF-supportX}
when $w$ is the coordinate-wise Euclidean distance function.\medskip{}
\noindent \textbf{Assumption 4$'$}
\protected@write \@auxout {}{\string \newlabel {as3prime}{{$4'$}{\thepage}{$4'$}{as3prime}{}} }
\hypertarget{as3prime}{}
\emph{Suppose
that (i) $w_{h}\left(\ol x,\ul x\right):=\left|\ol x_{h}-\ul x_{h}\right|$
for every coordinate $h$, and (ii) $Supp\left(X_{i}\right)$ has
nonempty interior.}\medskip{}
Essentially, since our criterion function is based on indicator functions
of halfspaces in the form of $\left(w\left(\ol x,\tilde{x}\right)-w\left(\ul x,\tilde{x}\right)\right)^{'}\beta\lesseqqgtr0$,
we will need the distribution of these indicators to take both values
$0$ and $1$ with strictly positive probabilities under any $\beta$,
so that every possible $\beta$ different from $\beta_{0}$ can be differentiated
from $\beta_{0}$ by the criterion function (up to scale normalization).
Hence, we need sufficient variations in the observable covariates.
Assumption \ref{assu:DNF-supportX}, though apparently not very transparent
on its own, is actually implied by Assumption \ref{as3prime}.\footnote{\emph{Proof:} Suppose that $Supp\left(X_{i}\right)$ has nonempty
interior. Then there exist two distinct points $\ol x$ and $\ul x$
in the interior of $Supp\left(X_{i}\right)$ such that $\tilde{x}:=\frac{1}{2}\left(\ol x+\ul x\right)$
is also in the interior of $Supp\left(X_{i}\right)$. Clearly, $w\left(\ol x,\tilde{x}\right)=w\left(\ul x,\tilde{x}\right)=\left(\frac{1}{2}\left|\ol x_{1}-\ul x_{1}\right|,...,\frac{1}{2}\left|\ol x_{k}-\ul x_{k}\right|\right)^{'}$.
Since $\ol x,\ul x$ are all interior points of $Supp\left(X_{i}\right)$
and $w$ is continuous, the vector ${\bf 0}$ must be an interior
point of $Supp\left(w\left(\ol x,X_{i}\right)-w\left(\ul x,X_{i}\right)\right)$.
\qedsymbol} We note that the assumption of nonempty interior is a familiar one,
which is often imposed for point identification in the literature,
say, on maximum score estimation. Assumption \ref{as3prime} allows
the support of all observable covariates to be bounded, but on the
other hand require all covariates to be continuously distributed.
In Appendix \ref{subsec:PID_Discrete}, we present an alternative
set of assumptions that allow for the presence of discrete covariates,
but require the existence of a ``special covariate'' with large
(conditional) continuous support and nonzero coefficient, a la \citet*{horowitz1992smoothed}.
Since the identification result with special covariate needs to be
presented under a different scale normalization, we defer the results
to the appendix.
\begin{assumption}[Conditional Support of $A_{i}$]
\label{assu:DNF-supportA}$A_{i}$ is conditionally distributed on
the same support given $X_{i}=x$ for any realization $x\in Supp\left(X_{i}\right)$.
\end{assumption}
Assumption \ref{assu:DNF-supportA} together with Assumption \ref{assu:DNF-rs}
implies that for two randomly sampled individuals $i,j$, there is
a strictly positive probability of $A_{i}$ and $A_{j}$ being sufficiently
close to each other, conditional on any realizations of $X_{i}$ and
$X_{j}$. Together with the continuity condition in Assumption \ref{assu:DNF-con},
we can ensure that the parametric index based on observable covariates
alone can determine whether $i$ or $j$ is relatively more popular
among a certain group of individuals.
$\ $
Next, we lay out the lemma that will be used in the proof of point
identification of $\beta_{0}$.
\begin{lem}[Differentiating $\beta$ from $\beta_{0}$]
\label{lem:DNF2} Under model \eqref{eq:Model_NF}, Assumptions \ref{as1prime},
\ref{assu:DNF-rs}--\ref{assu:DNF-supportA}, for each $\beta\in\text{\ensuremath{\mathbb{S}}}^{d-1}\backslash\left\{ \beta_{0}\right\} $,
\begin{align}
\tau_{ij}\left(X_{k},X_{l}\right) & =1,\label{eq:DNF_ID1}\\
\lambda\left(X_{k},X_{l};X_{i},X_{j};\beta_{0}\right) & =0,\label{eq:DNF_ID2}\\
\lambda\left(X_{k},X_{l};X_{i},X_{j};\beta\right) & =1,\label{eq:DNF_ID3}
\end{align}
occur simultaneously with strictly positive probability, as we randomly
sample individuals $i,j,k,l$.
\end{lem}
For point identification of $\beta_{0}$, we need the population criterion
\eqref{eq:pop_cri} to differentiate each $\beta\in\text{\ensuremath{\mathbb{S}}}^{d-1}\backslash\left\{ \beta_{0}\right\} $
from $\beta_{0}$. Lemma \ref{lem:DNF2} ensures this by establishing
that, the conditioning event \eqref{eq:DNF_ID1} in Lemma \ref{lem:DNF1}
occurs with positive probability, and, when it occurs, we can obtain
different values of $\lambda$ at $\beta$ from that at $\beta$ with positive
probabilities. Compared with the set identification result discussed
in Section \ref{subsec:DNF_ID}, we need to strengthen Assumption
\ref{assu:C5_mon} to \emph{strict} monotonicity in the index $w\left(X_{i},X_{j}\right)^{'}\beta_{0}$.
Then, Assumptions \ref{assu:DNF-con} and \ref{assu:DNF-supportA}
guarantee that, when the absolute difference between $A_{i}$ and
$A_{j}$ get sufficiently small, differences in the parametric indexes
can lead to differences in conditional linking probabilities, so that
\eqref{eq:DNF_ID1} will occur. Since $\beta_{0}$ and $\beta$ define different
intersections of halfspaces for the vector $w\left(\ol x,\tilde{x}\right)-w\left(\ul x,\tilde{x}\right)$
through $\lambda$, Assumption \ref{assu:DNF-supportX} then guarantees
that there will be on-support realizations of the observable covariates
that help ``detect'' such differences.
We are now ready to present the point identification result.
\begin{thm}[Point Identification of $\beta_{0}$]
\label{thm:DNF-ID}Under model \eqref{eq:Model_NF} and Assumptions
\ref{as1prime}, \ref{assu:DNF-rs}--\ref{assu:DNF-supportA}, $\beta_{0}$
is the unique minimizer of $Q\left(\beta\right)$ defined in \eqref{eq:pop_cri}
over the unit sphere $\text{\ensuremath{\mathbb{S}}}^{d-1}$. Furthermore,
for any $\epsilon>0$, there exists $\delta>0$ such that
\[
\inf_{\beta\in\text{\ensuremath{\mathbb{S}}}^{d-1}\backslash B_{\epsilon}\left(\beta_{0}\right)}Q\left(\beta\right)\geq Q\left(\beta_{0}\right)+\delta,
\]
where $B_{\epsilon}\left(\beta_{0}\right):=\left\{ \beta\in\mathbb{\mathbb{S}}^{d-1}:\|\beta-\beta_{0}\|\leq\epsilon\right\} $.
\end{thm}
Theorem \ref{thm:DNF-ID} follows from Lemma \ref{lem:DNF2} based
on the standard arguments in \citet*{newey1994asymp}.
\begin{rem*}[Asymmetry of $w$, Continued]
In Appendix \ref{subsec:ID_asym}, we show how the identification
arguments and assumptions above can be adapted to accommodate asymmetry
of $w$. In short, the technique of logical differencing applies without
changes, but the identifying restriction we obtained becomes weaker.
In particular, when $w$ is \emph{antisymmetric} in the sense that
$w\left(\ol x,\ul x\right)+w\left(\ul x,\ol x\right)\equiv0$, the
identifying restriction we obtained through logical differencing becomes
trivial, and $B_{0}=\mathbb{\mathbb{S}}^{d-1}$. However, with asymmetric but not antisymmetric
$w$, it is still feasible to strengthen Assumption 3 so as to obtain
point identification. See more discussions in Appendix \ref{subsec:ID_asym}.
\end{rem*}
\subsection{\label{subsec:DNF_est}Tetrad Estimator and Consistency}
We now proceed to present a consistent estimator of $\beta_{0}$ in the
framework of extremum estimation, which we construct using a two-step
semiparametric estimation procedure. We clarify that we are considering
the asymptotics under ``a single large network'' with the number
of individuals $N\to\infty$. Moreover, we focus on the ``dense network''
asymptotics where the conditional linking probabilities are nondegenerate
in the limit and can be consistently estimated.
The first step is the nonparametric estimation of
\[
\rho_{i}\left(x\right):=\mathbb{E}\left[D_{ik}\big|i,X_{k}=x\right].
\]
To implement this, we fix an individual $i$ in the sample, and regress
$D_{ik}$, the indicator function for the link between $i$ and $k$,
on the basis functions chosen by the researcher evaluated at observable
characteristics $X_{k}$ for all $k\neq i$. To guarantee that a consistent
nonparametric estimator of $\rho_{i}$ exists, we need to impose some
regularity conditions on $\rho_{i}$. We state the following assumption
as an illustrative set of such conditions, acknowledging that there
may be many different versions that also work.
\begin{assumption}[Regularity Conditions for $\rho_{i}$]
\label{assu:uniform-reg} (i) $Supp\left(X_{i}\right)$ is bounded
and convex with nonempty interior; (ii) for each fixed $i$, $\rho_{i}\in{\cal C}_{M}^{d_{x}+1}\left(Supp\left(X_{i}\right)\right)$,
where ${\cal C}_{M}^{d_{x}+1}\left(Supp\left(X_{i}\right)\right)$
denotes the class of functions on $Supp\left(X_{i}\right)$ whose
derivatives are uniformly bounded by $M$ up to order $d_{x}+1$.
\end{assumption}
Assumption \ref{assu:uniform-reg} essentially requires that $\rho_{i}\left(x\right)$
is smooth enough in $x$. Given that
\[
\rho_{i}\left(x\right)=\int\phi\left(w\left(\ol x,x\right)^{'}\beta_{0},a_{i},A_{k}\right)d\mathbb{P}\left(\rest{A_{k}}X_{k}=x\right)
\]
is an integral of $\phi$ (a strictly increasing function bounded
between 0 and $1$) over the conditional distribution of $A_{k}$,
Assumption \ref{assu:uniform-reg} is easily satisfied, say, if both
$\phi$ and the conditional density of $A_{k}$ given $X_{k}=x$ have
uniformly bounded derivatives up to order $d+1$, when $w$ is taken
to be the coordinate-wise Euclidean distance function.
\begin{lem}
\label{lem:est_1step}Given Assumption \ref{assu:uniform-reg}, for
each $i$, there exists an estimator $\hat{\rho}_{i}\in{\cal C}_{M}^{d_{x}+1}\left({\cal X}\right)$
that is $L_{2}\left(\mathbb{P}_{X}\right)$ consistent, i.e.,
\[
\norm{\hat{\rho}_{i}-\rho_{i}}_{L_{2}\left(\mathbb{P}_{X}\right)}:=\sqrt{\int\left(\hat{\rho}_{i}\left(x\right)-\rho_{i}\left(x\right)\right)^{2}d\mathbb{P}_{X_{k}}\left(x\right)}=o_{p}\left(1\right).
\]
\end{lem}
Lemma \ref{lem:est_1step} follows from the large literature on many
different types of consistent nonparametric estimators. See \citet{Bierens1983kernel}
for results on kernel estimators and \citet{chen2007sieve} on sieve
estimators. In our simulation and application, we use a spline-based
sieve estimator.\medskip{}
In the second step, we use $\hat{\rho}_{i}$ to build the following
sample analog of the population criterion $Q^{\gamma}\left(\beta\right)$
in \eqref{eq:pop_cri}:
\begin{equation}
\begin{aligned}\widehat{Q}_{n}^{\gamma}\left(\beta\right):=\frac{\left(n-4\right)!}{n!}\sum_{i,j,k,l}\gamma\left(\hat{\rho}_{i}\left(X_{k}\right)-\hat{\rho}_{j}\left(X_{k}\right)\right)\gamma\left(\hat{\rho}_{j}\left(X_{l}\right)-\hat{\rho}_{i}\left(X_{l}\right)\right)\lambda\left(X_{k},X_{l};X_{i},X_{j};\beta\right).\end{aligned}
\label{eq:sampeQ}
\end{equation}
The two-step tetrad estimator for $\beta_{0}$ is then defined as
\begin{equation}
\widehat{\beta}_{n}:=\arg\min_{\beta\in\text{\ensuremath{\mathbb{S}}}^{d-1}}\widehat{Q}_{n}^{\gamma}\left(\beta\right).\label{eq:beta^hat}
\end{equation}
Since $\gamma$ can be arbitrarily chosen as long as it preserves strict
positiveness as in Section \ref{subsec:DNF_est}, we now impose the
following continuity assumption for $\gamma$.
\begin{assumption}[Continuity of $\gamma$]
\label{assu:as_gamma} The one-sided sign-preserving function $\gamma$
is Lipschitz-continuous.
\end{assumption}
Assumption \ref{assu:as_gamma} can be achieved by setting, say, $\gamma\left(t\right):=\left[t\right]_{+}$
or $\gamma\left(t\right):=2\times\Phi\left(\left[t\right]_{+}\right)-1$,
where $\left[t\right]_{+}$ is the positive part of $t$, and $\Phi$
is the CDF of the standard normal distribution. See Section \ref{sec:C5_simu}
for details. The idea is that, when the difference $\rho_{i}\left(x\right)$
and $\rho_{j}\left(x\right)$ is small, the estimation of whether
$\rho_{i}\left(x\right)>\rho_{j}\left(x\right)$ may be relatively
imprecise, and therefore, we may wish to downweight such terms in
the criterion function.
Computationally, to exploit the topological characteristics of the
parameter space $\mathbb{\mathbb{S}}^{d-1}$, i.e. compactness and convexity, we develop
a new bisection-style nested rectangle algorithm that recursively
shrinks and refines an adaptive grid on the angle space. The key novelty
of the algorithm is that instead of working with the edges of the
Euclidean parameter space $\mathbb{R}^{d}$, we deterministically ``cut''
the angle space in each dimension of $\mathbb{\mathbb{S}}^{d-1}$ to search for the
area that minimizes $\widehat{Q}_{n}^{\gamma}(\beta)$. Additional measures
are taken to ensure the search algorithm is conservative. Simulation
and empirical results show that our algorithm performs reasonably
well with a relatively small sample size. \citet{gao2018robust} provides
more details regarding the implementation in a panel multinomial choice
setting.
We now state the consistency result of the tetrad estimator $\widehat{\beta}_{n}$
under point identification.
\begin{thm}
\label{thm:DNF-consistent}Under model \eqref{eq:Model_NF} and Assumptions
\ref{as1prime}, \ref{assu:DNF-rs}--\ref{assu:as_gamma}, $\hat{\beta}_{n}$
is consistent for $\beta_{0}$, i.e.,
\[
\widehat{\beta}_{n}\overset{p}{\longrightarrow}\beta_{0}.
\]
\end{thm}
\begin{rem}
Our estimator shares some similarity with the maximum score estimator:
the parameter $\beta$ enters into the sample criterion through indicators
of halfspaces about $\beta$, which creates discreteness in the sample
criterion. However, our estimator features an additional complication
not found in maximum score estimation: we require the first-stage
nonparametric estimators $\hat{\rho}_{i}$, and we need to plug the
first-stage estimators into nonlinear functions that isolate the ``positive
side'' only: $\gamma\left(\hat{\rho}_{i}\left(X_{k}\right)-\hat{\rho}_{j}\left(X_{k}\right)\right)\gamma\left(\hat{\rho}_{j}\left(X_{l}\right)-\hat{\rho}_{i}\left(X_{l}\right)\right)$
will be nonzero only when $\hat{\rho}_{i}\left(X_{k}\right)>\hat{\rho}_{j}\left(X_{k}\right)$
and $\hat{\rho}_{j}\left(X_{l}\right)>\hat{\rho}_{i}\left(X_{l}\right)$.
In fact, it is precisely the ``double-threshold'' feature of our
model that leads to the necessity of a semiparametric two-step estimation
procedure, relative to the one-step procedure of the maximum score
estimator under a ``single-threshold'' setting. Given that the asymptotic
theory for the maximum score estimator is already nonstandard (with
cubic-root rate of convergence and non-normal Chernoff-type asymptotic
distribution as in \citealp{kim1990cube}), the addition of the first-stage
nonparametric estimation further complicates the asymptotic theory
to a highly nontrivial extent. The first-stage nonparametric estimation
of $\rho_{i}$ may further slow the rate of convergence below $n^{-1/3}$;
in the meanwhile, if we take $\gamma$ to be smooth (but necessarily nonlinear),
the term $\gamma\left(\hat{\rho}_{i}\left(X_{k}\right)-\hat{\rho}_{j}\left(X_{k}\right)\right)\gamma\left(\hat{\rho}_{j}\left(X_{l}\right)-\hat{\rho}_{i}\left(X_{l}\right)\right)$
may also provide some effective smoothing on the discrete $\lambda_{ij}$
term, when $\rho_{i}\left(X_{k}\right)-\rho_{j}\left(X_{k}\right)$
or $\rho_{j}\left(X_{l}\right)-\rho_{i}\left(X_{l}\right)$ is closer
to zero. It is not exactly clear which effect dominates, or whether
the two effects can be balanced, in our current setting. Due to such
technical difficulties, we defer the investigation of such types of
``two-stage maximum score estimators'' in a separate paper by \citet*{gao2020two},
albeit in a simpler setting.
\end{rem}
\section{\label{sec:C5_simu}Simulation}
In this section, we conduct a simulation study to analyze the finite-sample
performance of our two-step tetrad estimator. To begin with, we calculate
and plot the identified set $B_{0}$ via \eqref{eq:pop_cri} for various
support restrictions on $X$\footnote{We thank the Editor for this suggestion.}.
The graph illustrates how the size of the identified set $B_{0}$
based on our population criterion \eqref{eq:pop_cri} changes when
the support of $X$ contains more and more discrete variables. Then,
we specify the data generating process (DGP) of the Monte Carlo simulations.
We show and discuss the performance of our two-step estimation method
under the baseline setup with symmetric $w\left(\cdot,\cdot\right)$ function.
In Appendix \ref{subsec:Sim_Robustness}, we vary the number of individuals
$N$, the dimension of the pairwise observable characteristics $d$,
and the degree of correlation between $X$ and $A$, as well as allow
for asymmetric $w$ function to further examine the robustness of
our method.
\subsection{Identified Set $B_{0}$\label{subsec:Identified-Set}}
In this section, we provide a graphical illustration of the identified
set $B_{0}$ for various support restrictions on $X$. The analysis
is based on the following network formation model:
\begin{equation}
D_{ij}=\mathbf{\mathbbm1}\left\{ w\left(X_{i},X_{j}\right)^{'}\beta_{0}+A_{i}>\epsilon_{ij}\right\} \cdot\mathbf{\mathbbm1}\left\{ w\left(X_{j},X_{i}\right)^{'}\beta_{0}+A_{j}>\epsilon_{ji}\right\} ,\label{eq:simrule}
\end{equation}
where $w$ is taken to be the coordinate-wise absolute difference,
i.e., $w_{h}\left(\ol x,\ul x\right):=\left|\ol x_{h}-\ul x_{h}\right|$
for $h=1,...,d$. We set $d_{x}=d=3$ and $\beta_{0}=\left(1,1,1\right)/6^{'}$\footnote{Recall that only the direction of $\beta_{0}$ is identified. Here we
divide the vector $\left(1,1,1\right)^{'}$ by 6 to ensure the network
is non-degenerate numerically, i.e., not all agents are connected,
nor are they all disconnected.}. We incorporate correlation between $X$ and $A$ by drawing $A_{i}=\left(X_{i,1}+\xi_{i}\right)/4$,
where $\xi_{i}$ is independently and uniformly distributed on $\left[-0.5,0.5\right]$
and independent of all other variables. $\epsilon_{ij}$ is the exogenous
random shock independent of all other variables and uniformly distributed
on $\left[0,1\right]$.
The purpose of the exercise is to show how the support restrictions
on $X$ affect the size of the identified set $B_{0}$. To this end,
we consider the following five support conditions on $X$: (1) all
coordinates of $X$ are uniformly distributed on $[-0.5,0.5]$; (2)
$X_{i,1}$ is binary with equal probability on $\left\{ 0,1\right\} $,
while $X_{i,2}$ and $X_{i,3}$ are both uniformly distributed on
$[-0.5,0.5]$; (3) $X_{i,1}$ is binary with equal probability on
$\left\{ 0,1\right\} $, $X_{i,2}$ is discrete with equal probability
on 11 points of $\left\{ -0.5,-0.4,..,0,..,0.4,0.5\right\} $, and
$X_{i,3}$ is uniformly distributed on $[-0.5,0.5]$; (4) all coordinates
of $X$ are discrete with equal probability on 101 points of $\left\{ -0.5,-0.49,-0.48..,0,..,0.49,0.5\right\} $;
(5) all coordinates of $X$ are discrete with equal probability on
11 points of $\left\{ -0.5,-0.4,..,0,..,0.4,0.5\right\} $.
$\beta_{0}$ is point identified in (1) since the support of $X$ has
a nonempty interior, thus satisfying Assumption \ref{as3prime}. In
(2)--(5), $\beta_{0}$ is not point identified due to discreteness and
boundedness of the support of $X$. Below we compare the sizes of
the identified sets when the $X$ vector contains none, one, two and
all discrete variables corresponding to DGP (1)--(5).
To calculate the ID set, we utilize the analytical formula of $\rho_{i}$
so that the true $\rho_{i}$ is calculated without error \footnote{We use the distribution of $\left(A,\epsilon,\xi\right)$ to obtain
the analytical formula for $\rho_{i}\left(x\right)$.}. Then, we numerically approximate the population criterion by setting
a large $N=1,000$ and $M=10,000$. We are able to work with a larger
$N$ and $M$ than in the simulations in the next subsection because
to get the identified set, we only calculate the maximizer of the
population criterion once without the need to estimate $\rho_{i}\left(x\right)$.
Still, one can improve the numerical results by increasing $N,M\rightarrow\infty$
and extracting more info from the data when additional computational
power is available. Hence, all our results below are conservative:
the true ID set must be a subset of the result shown in Figure \ref{fig:id_set_plot}.
For clarity of illustration, we plot the results in the angle space\footnote{One can transform any $\beta=\left(\beta_{1},\beta_{2},\beta_{3}\right)^{'}$
on the unit sphere $\mathbb{\mathbb{S}}^{2}$ into the angle space by letting $\beta_{1}=\cos\theta_{1}\cos\theta_{2},\beta_{2}=\cos\theta_{1}\sin\theta_{2},\beta_{3}=\sin\theta_{1}$
for $\left(\theta_{1},\theta_{2}\right)\in\left[-\pi/2,\pi/2\right]\times\left[-\pi,\pi\right]$.}. In all cases, we maintain the correlation between $X$ and $A$.
The results are summarized in Figure \ref{fig:id_set_plot}, where
we plot the identified sets $B_{0}$ on the full parameter space of
$\left[-\pi/2,\pi/2\right]\times\left[-\pi,\pi\right]$ and zoom it
in for more details.
\noindent \begin{center}
\begin{figure}
\noindent \begin{centering}
\includegraphics[scale=0.75]{IDSet_0723_fullspace_5graph}
\par\end{centering}
\noindent \centering{}\includegraphics[scale=0.75]{IDSet_0723_zoomin_5graph}\caption{Identified Sets under Different Support Conditions for $X_{i}$\label{fig:id_set_plot}}
\end{figure}
\par\end{center}
In Figure \ref{fig:id_set_plot}, the black dot represents the true
$\beta_{0}$, which is guaranteed to belong to the identified set $B_{0}$.
The red rectangle demonstrates the identified set when the supports
of all coordinates of $X$ are continuous and bounded. The theory
predicts point identification when the number of individuals $N$
and $\left(i,j\right)$ pairs $M$ go to infinity. Here the numerical
result is indeed very close to point identification. The blue rectangle
shows the identified set $B_{0}$ when the first dimension of $X$
is binary while all the other coordinates are continuous and bounded.
The size of $B_{0}$ is larger than in case (1), which is expected
due to the discreteness of $X_{i,1}$. The magenta rectangle corresponds
to the case when $X_{i,1}$ is binary, $X_{i,2}$ is discrete, and
$X_{i,3}$ is continuous and bounded. The size of $B_{0}$ is reasonably
small, given that there are two discrete variables in $X$ and one
of them is binary. The green and black rectangles illustrate $B_{0}$
when all coordinates of $X$ are discretely distributed with equal
probability on 101 (fine grid) and 11 (coarse grid) points between
-0.5 and 0.5, respectively. In these two scenarios, we see larger
identified sets than in (1)--(3). That said, the sets still appear
small relative to the full parameter space $\left[-\pi/2,\pi/2\right]\times\left[-\pi,\pi\right]$.
To summarize, the size of the identified set $B_{0}$ based on logical
differencing is reasonably small under each of the five settings.
\subsection{\label{subsec:Setup-of-Simulation} Performance of the Tetrad Estimator}
We maintain \eqref{eq:simrule} as our network formation model. We
draw each coordinate of $X_{i}$ independently from a uniform distribution
on $\left[-0.5,0.5\right]$ and calculate $w$ using coordinate-wise
absolute difference. Point identification is guaranteed since Assumption
\ref{as3prime} is satisfied. We set $A_{i}=corr\times X_{i,1}+\left(1-corr\right)\times\xi_{i}$,
and use the same distributions as in Section \ref{subsec:Identified-Set}
to generate $\xi_{_{i}}$ and $\epsilon_{ij}$. We set the true $\beta_{0}$
to be $(1,...,1)^{'}/\sqrt{d}\in\mathbb{\mathbb{S}}^{d-1}$ , and estimate the direction
of $\beta_{0}$.
For the baseline result in the main text, we fix the number of individuals
$N=100$, the dimension of $X$ (also $W_{ij}:=w\left(X_{i},X_{j}\right)$
and $\beta_{0}$) $d_{x}=d=3$, and $corr=0.2$. In Appendix \ref{subsec:Sim_Robustness},
we vary $N$, $d$ and $corr$, as well as allow for asymmetric $w$
function to investigate how robust our method is against various configurations.
To summarize, for each of the $B=100$ simulations we randomly generate
data on the characteristics of and the network structure among individuals.
Then, based on the observable $\left(X_{i},W_{ij},D_{ij}\right)_{i,j\in\left\{ 1,...,N\right\} }$
matrix we construct our two-step estimator $\widehat{\beta}$ for the
true parameter of interest $\beta_{0}$. Specifically, we use a sieve
estimator with 2nd-order spline with its knot at median for the first-stage
nonparametric estimation of $\rho_{i}\left(\cdot\right)$. The spline
is chosen to ensure a relatively small number of regressors in the
nonparametric regression considering the small size of $N$. In the
second stage, we adapt to the adaptive-gird search on the unit sphere
algorithm developed in \citet{gao2018robust} to calculate $\widehat{\beta}$
that minimizes the sample criterion function $\widehat{Q}\left(\beta\right)$
defined in \eqref{eq:sampeQ}\footnote{We suppress its dependence on $\gamma$ and $N$ for notational simplicity.
In all simulations and empirical application, we use smoothing function
$\gamma\left(t\right):=2\times\Phi\left(\left[t\right]_{+}\right)-1$.} over the unit sphere. It should be noted that, constrained by computational
power, when calculating the sample criterion $\widehat{Q}\text{\ensuremath{\left(\beta\right)}}$
for each $\beta\in\mathbb{\mathbb{S}}^{d-1}$ we randomly draw $M=1,000$ $\left(i,j\right)$
pairs of individuals and vary across all possible $\left(k,l\right)$
pairs excluding $i$ or $j$. One can improve those results by increasing
$M$ when computational constraint is not present, so again our results
are conservative. Lastly, we compare our estimator $\widehat{\beta}$
with $\beta_{0}$ based on several performance metrics including root
mean squared error (rMSE), mean norm deviations (MND), and maximum
mean absolute deviation (MMAD).
We define for each simulation round $b$ the set estimator $\widehat{\Theta}_{b}$
as the set of points that achieve the minimum of $\widehat{Q}\text{\ensuremath{\left(\beta\right)}}$
over the unit sphere $\mathbb{\mathbb{S}}^{d-1}$. We further define for each simulation
$b=1,...,B$ and each dimension $h=1,...,d$ of $\beta$
\[
\widehat{\beta}_{b,h}^{l}:=\min\widehat{\Theta}_{b,h},\quad\widehat{\beta}_{b,h}^{u}:=\max\widehat{\Theta}_{b,h},\quad\widehat{\beta}_{b,h}^{m}:=\dfrac{1}{2}\left(\widehat{\beta}_{b,h}^{l}+\widehat{\beta}_{b,h}^{u}\right),
\]
where $\widehat{\beta}_{b,h}^{l}$, $\widehat{\beta}_{b,h}^{u}$ and
$\widehat{\beta}_{b,h}^{m}$ are the minimum, maximum, and middle point
along dimension $h$ for simulation round $b$ of the identified set
$\widehat{\Theta}_{b}$, respectively. One can consider $\widehat{\beta}^{m}$
as the point estimator for $\beta_{0}$. Note by construction for
each simulation round $b$, the identified set $\widehat{\Theta}_{b}$
is a subset of the rectangle
\[
\widehat{\Xi}_{b}:=\times_{h=1}^{d}\left[\widehat{\beta}_{b,h}^{l},\ \widehat{\beta}_{b,h}^{u}\right].
\]
We calculate the baseline performance using $\widehat{\beta}^{l},\widehat{\beta}^{u},\widehat{\beta}^{m}$
respectively.
We report in Table \ref{tab:C5_Sim_sum} the performance of our estimators.
In the first row, we calculate the mean bias across $B=100$ simulations
using $\widehat{\beta}^{m}$ along each dimension $h=1,...,d$. The result
shows the estimation bias is very small across all dimensions with
a magnitude between -0.0053 and 0.0052. Similar performance is observed
using $\widehat{\beta}^{u}$ and $\widehat{\beta}^{l}$ as shown in row
2 and 3. We do not find any sign of persistent over/under- estimation
of $\beta_{0}$ across each dimension. Row 4 measures the average
width along each dimension of the rectangle $\widehat{\Xi}_{b}$ over
$B$ simulations. The average size of $\widehat{\Xi}_{b}$ is small,
indicating a tight area for the estimated set. In the second part
of Table \ref{tab:C5_Sim_sum} we report rMSE, MND, and MMAD, all
of which are small in magnitude and provide evidence that our estimator
works well in finite sample.
\begin{table}
\caption{\label{tab:C5_Sim_sum}Baseline Performance}
\bigskip{}
\noindent \centering{}
\begin{tabular}{ccccc}
\hline
$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$ & & $\beta_{1}$ & $\beta_{2}$ & $\beta_{3}$\tabularnewline
\hline
$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$bias & $\frac{1}{B}\sum_{b}\left(\hat{\beta}_{b,h}^{m}-\beta_{0,h}\right)$ & -0.0021 & 0.0052 & -0.0053\tabularnewline
$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$upper bias & $\frac{1}{B}\sum_{b}\left(\hat{\beta}_{b,h}^{u}-\beta_{0,h}\right)$ & 0.0048 & 0.0118 & -0.0002\tabularnewline
$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$lower bias & $\frac{1}{B}\sum_{b}\left(\hat{\beta}_{b,h}^{l}-\beta_{0,h}\right)$ & -0.0091 & -0.0015 & -0.0105\tabularnewline
$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$mean$(u-l)$ & $\frac{1}{B}\sum_{b}\left(\hat{\beta}_{b,h}^{u}-\hat{\beta}_{b,h}^{l}\right)$ & 0.0138 & 0.0132 & 0.0103\tabularnewline
\hline
& & & & \tabularnewline
\hline
$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$rMSE & $\sqrt{\frac{1}{B}\sum_{b}\norm{\hat{\beta}_{b}^{m}-\beta_{0}}^{2}}$ & \multicolumn{3}{c}{0.0488}\tabularnewline
$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$MND & $\frac{1}{B}\sum_{b}\norm{\hat{\beta}_{b}^{m}-\beta_{0}}$ & \multicolumn{3}{c}{0.0417}\tabularnewline
$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$MMAD & $\max_{h\in\left\{ 1,..,d\right\} }\abs{\frac{1}{B}\sum_{b}\left(\hat{\beta}_{b,h}^{m}-\beta_{0,h}\right)}$ & \multicolumn{3}{c}{0.0053}\tabularnewline
\hline
\end{tabular}
\end{table}
\section{\label{sec:Emp}Empirical Illustration}
As an empirical illustration, we estimate a network formation model
under NTU with data of a village network called Nyakatoke in Tanzania.
Nyakatoke is a small Haya community of 119 households in 2000 located
in the Kagera Region of Tanzania. We are interested in how important
factors, such as wealth, distance, and blood or religious ties, are
relative to each other in deciding the formation of risk-sharing links
among local residents. We apply our two-stage estimator to the Nyakatoke
network data and obtain economically intuitive results.
\subsection{\label{subsec:Data}Data Description}
The risk-sharing data of Nyakatoke, collected by Joachim De Weerdt
in 2000, cover all of the 119 households in the community. It includes
the information about whether or not two households are linked in
the insurance network. It also provides detailed information on total
USD assets and religion of each household, as well as kinship and
distance between households. See \citet{de2004risk}, \citet{DeWeerdt2006},
and \citet{de2011social} for more details of this dataset.
To define the dependent variable \textit{link}, the interviewer asks
each household the following question:
$\ $
``\textit{Can you give a list of people from inside or outside of
Nyakatoke, who you can personally rely on for help and/or that can
rely on you for help in cash, kind or labor?}''
$\ $
The data contains three answers of ``bilaterally mentioned'', ``unilaterally
mentioned'', and ``not mentioned'' between each pair of households.
Considering the question is about whether one can rely on the other
for help, we interpret both ``bilaterally mentioned'' and ``unilaterally
mentioned'' as they are connected in this undirected network, meaning
that the dependent variable $D_{ij}$ \textit{link} equals 1.\footnote{In the context of the village economies in our application, we think,
at the time of link formation, the risk-sharing links are less likely
(in comparison with the contexts of business or financial networks)
to be driven by efficient arrangements of side-payment transfers,
thus satisfying NTU.} We also ran a robustness check by constructing a weighted network
based on the answers, i.e. ``bilaterally mentioned'' means \textit{link}
equals 2, ``unilaterally mentioned'' means \textit{link} equals
1, and ``not mentioned'' means \textit{link} equals 0, and obtained
very similar results.
We estimate the coefficients for 3 regressors: \emph{wealth difference},
\emph{distance} and \emph{tie} between households, with our two-step
estimator. \textit{Wealth} is defined as the total assets in USD owned
by each household in 2000, including livestocks, durables and land.
\textit{Distance} measures how far away two households are located
in kilometers. \textit{Tie} is a discrete variable, defined to be
3 if members of one household are parents, children and/or siblings
of members of the other household, 2 if nephews, nieces, aunts, cousins,
grandparents and grandchildren, 1 if any other blood relation applies
or if two households share the same religion, and 0 if no blood religious
tie exists. Following the literature we take natural log on \textit{wealth}
and \textit{distance}, and we construct the \textit{wealth difference}
variable as the absolute difference in \emph{wealth}, i.e.,
\[
w\left(X_{i},X_{j}\right)=\left(\begin{array}{c}
\abs{\ln\text{wealth}_{i}-\ln\text{wealth}_{j}}\\
\ln\text{distance}_{ij}\\
\text{tie}_{ij}
\end{array}\right).
\]
Figure \ref{fig:NetworkGraph} illustrates the structure of the insurance
network in Nyakatoke. Each node in the graph represents a household.
The solid line between two nodes indicates they are connected, i.e.,
\textit{link} equals 1. The size of each node is proportional to the
USD wealth of each household. Each node is colored according to its
rank in \emph{wealth}: green for the top quartile, red for the second,
yellow for the third and purple for the fourth quartile.
\begin{center}
\begin{figure}[h]
\begin{centering}
\includegraphics[scale=0.45]{Network_Plot_1}
\par\end{centering}
\caption{A Graphical Illustration of the Insurance Network of Nyakatoke\label{fig:NetworkGraph}}
\end{figure}
\par\end{center}
In the dataset there are 5 households that lack information on \textit{wealth}
and/or \textit{distance}. We drop these observations, resulting in
a sample size $N$ of 114. The total number of ordered household pairs
is 12,882. Summary statistics for the dependent and explanatory variables
used in our analysis are presented in Table \ref{tab:Empirical-Application:-Summary}.
\begin{table}
\caption{Empirical Application: Summary Statistics\label{tab:Empirical-Application:-Summary}}
\bigskip{}
\noindent \centering{}
\begin{tabular}{cccccc}
\hline
Variable & Obs & Mean & Std. Dev. & Min & Max\tabularnewline
\hline
& & & & & \tabularnewline
link & 12,882 & 0.0732 & 0.2606 & 0 & 1\tabularnewline
$\abs{\text{(ln) wealth difference}}$ & 12,882 & 1.0365 & 0.8228 & 0.0004 & 5.8898\tabularnewline
($\ln$) distance & 12,882 & 6.0553 & 0.7092 & 2.6672 & 7.4603\tabularnewline
tie & 12,882 & 0.4260 & 0.6123 & 0.0000 & 3.0000\tabularnewline
\hline
\end{tabular}
\end{table}
Given the data structure, we should expect close to point identification
because even though the \textit{tie} variable is discrete, the other
two regressors \textit{wealth} and \textit{distance} can be considered
as being continuously distributed with a large conditional support.
In addition, \textit{tie} also has more than one point in its support
given other variables. Thus, it is straightforward to verify that
Assumption \textit{\emph{\ref{as4prime2}}}\emph{ }\textit{\emph{is
satisfied, leading to point identification of $\beta_{0}$.}}
\subsection{\label{subsec: emp Results-and-Discussion}Results and Discussion}
We apply our two-stage estimator proposed in Section \ref{subsec:DNF_est}
to the Nyakatoke network data, where, similarly to the simulation
studies, we use second degree splines to estimate $\rho_{i}$ and
the adaptive grid search algorithm to find the minimizer of the criterion
function $\widehat{Q}\left(\beta\right)$ on $\mathbb{\mathbb{S}}^{d-1}$.
\begin{table}[h]
\caption{Empirical Application: Estimation Results\label{tab:EmpResults}}
\bigskip{}
\noindent \centering{}
\begin{tabular}{ccc}
\hline
\multirow{2}{*}{\textcolor{black}{$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$}Variable} & \multirow{2}{*}{$\hat{\beta}^{m}$} & \multirow{2}{*}{$\left[\hat{\beta}^{l},\ \hat{\beta}^{u}\right]$}\tabularnewline
& & \tabularnewline
\hline
& & \tabularnewline
\textcolor{black}{$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$}$\abs{\text{(ln) wealth difference}}$ & -0.1948 & $\left[-0.1964,\ -0.1932\right]$\tabularnewline
\textcolor{black}{$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$}($\ln$)
distance & -0.8036 & $\left[-0.8043,\ -0.8029\right]$\tabularnewline
\textcolor{black}{$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$}tie & 0.5619 & $\left[0.5608,\ 0.5630\right]$\tabularnewline
\hline
\end{tabular}
\end{table}
Table \ref{tab:EmpResults} summarizes our estimation results. The
column of $\hat{\beta}^{m}$ corresponds to the center of the estimated
rectangle
\[
\widehat{\Xi}:=\times_{h=1}^{d}\left[\widehat{\beta}_{h}^{l},\ \widehat{\beta}_{h}^{u}\right],
\]
where $\widehat{\beta}_{h}^{l}$ and $\widehat{\beta}_{h}^{u}$ represent
the lower and upper bound of the estimated area (the set of minimizers
of sample criterion function \eqref{eq:sampeQ}) along each dimension
$h=1,..,d$. We use $\hat{\beta}^{m}$ as the point estimator of the
coefficients. While the scale of $\beta_{0}$ is unidentified, one can
still compare the estimated coefficients with each other to obtain
an idea about which variable affects the formation of the link more
than the other.
The estimated coefficients for each variable conform well with economic
intuition. The estimated set for\textit{ wealth difference}'s coefficient
is $\left[-0.1964,-0.1932\right]$, which implies the more difference
in wealth between two households, the lower likelihood they are connected.
The estimated set for \textit{distance} is $\left[-0.8043,-0.8029\right]$.
It is natural that households rely more on neighbors for help than
those who live farther away, thus neighbors are more likely to be
connected. The estimated coefficient for \textit{tie} falls in the
positive range of $\left[0.5608,0.5630\right]$, consistent with the
intuition that one would depend on support from family when negative
shock occurs.
It is worth mentioning the estimated set $\hat{\Xi}$ is very tight
in each dimension, with the maximum width $\max_{h=1,..,d}\left(\hat{\beta}_{h}^{u}-\hat{\beta}_{h}^{l}\right)$
equal to 0.0032 for \textit{wealth difference.} Usually the discreteness
in \textit{tie} could make the estimated set wide, but our method
is able to leverage the large support in \textit{wealth difference}
and \textit{distance}. Once again, the relative magnitude and sign
of the coefficient for \textit{tie} are estimated in line with expectation
despite of its discreteness. In summary, the empirical results show
that our proposed estimator is able to generate economically intuitive
estimates under NTU.
\section{\label{sec:C5_conc} Conclusion}
This paper considers a semiparametric model of dyadic network formation
under nontransferable utilities, a natural and realistic micro-theoretical
feature that translates into the lack of additive separability in
econometric modeling. We show how a new methodology called \emph{logical
differencing} can be leveraged to cancel out the two-way fixed effects,
which correspond to unobserved individual heterogeneity, without relying
on arithmetic additivity. The key idea is to exploit the logical implication
of weak multivariate monotonicity and use the intersection of mutually
exclusive events on the unobserved fixed effects. It would be interesting
to explore whether and how the idea of \emph{logical differencing},
or more generally the use of fundamental logical operations, can be
applied to other econometric settings.
The identified sets derived by our method under various support restrictions
demonstrate that the proposed method is able to achieve accurate identified
set despite of discreteness and boundedness in $Supp\left(X\right)$.
Simulation results show that our method performs reasonably well with
a relatively small sample size, and robust to various configurations.
The empirical illustration using the real network data of Nyakatoke
reveals that our method is able to capture the essence of the network
formation process by generating estimates that conform well with economic
intuition.
This paper also reveals several further research questions regarding
dyadic network formation models under the NTU setting. First, given
the observation that the NTU setting can capture ``homophily effects''
with respect to the unobserved heterogeneity (under log-concave error
distributions) while imposing monotonicity in the unobserved heterogeneity,
it is interesting to investigate whether one can differentiate homophily
effects generated by ``intrinsic preference'' from assortativity
effects generated by bilateral consent, NTU and log-concave errors.
Second, admittedly the identifying restriction obtained in this paper
becomes uninformative when we have \emph{antisymmetric} pairwise observable
characteristics. However, preliminary analysis based on an adaption
of \citet{gao2018nonparametric} to the NTU setting suggests that
individual unobserved heterogeneity can be nonparametrically identified
up to location and inter-quantile range normalizations. After the
identification of individual unobserved heterogeneity terms ($A_{i}$),
it becomes straightforward to identify the index parameter $\beta_{0}$
based on the observable characteristics, even in the presence of antisymmetric
pairwise characteristics. However, consistent estimators of $A_{i}$
and $\beta_{0}$ in a semiparametric framework based on the identification
strategy in \citet{gao2018nonparametric} are still being developed.
We thus leave these research questions to future work.
\paragraph{Acknowledgements}
Wayne Gao is grateful to Xiaohong Chen and Peter C. B. Phillips for
their advice and support while working on this paper during his PhD
study at Yale University. Ming Li is grateful to Donald W. K. Andrews
and Yuichi Kitamura for their advice and support while working on
this paper during his PhD study at Yale University. We also thank
Isaiah Andrews, Yann Bramoull\'{e}, Benjamin Connault, and Paul Goldsmith-Pinkham,
as well as seminar and conference participants at Yale University,
University of Arizona, University of Essex, the International Conference
for Game Theory at Stony Brook University (2019), the Young Economist
Symposium at Columbia University (2019), and the Latin American Meeting
of the Econometric Society at Benem\'{e}rita Universidad Aut\'{o}noma
de Puebla (2019) for their comments. All errors are our own.
\bibliographystyle{ecta}
\bibliography{NetForm}
\newpage{}