EconBase
← Back to paper

Functional Differencing in Networks

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

65,301 characters · 18 sections · 103 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Functional Differencing in Networks

\vskip 3cm

abstractEconomic interactions often occur in networks where heterogeneous agents (such as workers or firms) sort and produce. However, most existing estimation approaches either require the network to be dense, which is at odds with many empirical networks, or they require restricting the form of heterogeneity and the network formation process. We show how the functional differencing approach introduced by bonhomme2012functional in the context of panel data, can be applied in network settings to derive moment restrictions on model parameters and average effects. Those restrictions are valid irrespective of the form of heterogeneity, and they hold in both dense and sparse networks. We illustrate the analysis with linear and nonlinear models of matched employer-employee data, in the spirit of the model introduced by abowd1999high. Keywords:\ Econometric models of networks, matching, sorting, heterogeneity, functional differencing.

\baselineskip21pt

\setcounter{page}{0}\thispagestyle{empty}

Introduction

Network data is increasingly prevalent in applied economics. In this paper we focus on models where agents (e.g., workers and firms) sort and interact on a network. In such settings, accounting for unobserved heterogeneity is empirically key. However, existing approaches to estimation in the presence of flexible heterogeneity are imperfect.

A first approach consists in treating the heterogeneity as “fixed effects” parameters to be estimated. Bias reduction methods, initially developed for single-agent panel data (hahn2004jackknife, dhaene2015split, fernandez2016individual), have been recently extended to networks (e.g., graham2017econometric, hughes2022estimating). The fixed-effects approach is appealing since it does not require modeling the distribution of heterogeneity and how it correlates with conditioning variables. Additionally, in models where agents interact on an exogenous network, the fixed-effects approach does not require specifying a model of network formation.

However, the performance of bias reduction methods in fixed-effects models hinges crucially on the network being sufficiently dense. This requirement is at odds with the nature of several empirical networks. For instance, in applications to wage determination in the presence of worker and firm heterogeneity (abowd1999high), a dense network approximation is typically inappropriate, and fixed-effects estimates suffer from a “limited mobility bias” that may be substantial (bonhomme2020much). More generally, sparsity is a feature of many empirical networks (graham2020sparse).

A second approach consists in postulating a “random effects” model for the unobserved heterogeneity. For example, bonhomme2019distributional and lentz2022anatomy propose and estimate random-effects models of worker heterogeneity in the presence of firm heterogeneity to account for sorting and complementarity on the labor market. Studying a different setting, bonhomme2021teams develops random effects models of agent heterogeneity in team production networks. Random-effects methods enjoy theoretical guarantees in sparse networks, under the assumption that the model is correctly specified.

However, modeling the full distribution of heterogeneity given conditioning variables can be challenging. In models of wage determination in the presence of worker and firm heterogeneity, this requires modeling the heterogeneity conditional on the entire network of employment relationships and job transitions, as proposed, for example, by woodcock2008wage and bonhomme2020much. More generally, the random-effects approach effectively requires modeling the network formation process, which can be a difficult task due to dimensionality and equilibrium multiplicity challenges.

In this paper our aim is to achieve the best of these two approaches, in the sense that we seek estimators that are fully robust to the form of heterogeneity, and that behave well in denser and sparser networks. Such estimators currently exist only in very special cases. Notably, andrews2008high and kline2020leave propose exact bias corrections for fixed-effects estimators of variance components in linear regressions on networks, while graham2017econometric proposes a “tetrad logit” estimator in a logistic model of network formation. These strategies mimic panel data methods that are consistent in fixed-length panels, such as conditional logit estimators (rasch1960studies, andersen1970asymptotic) or estimators of variance components (arellano2012identifying).

In panel data, the functional differencing approach (bonhomme2012functional) provides a general methodology to find moment restrictions on parameters that are robust to any distribution of heterogeneity and correlation with conditioning variables, and hold in fixed-length panels. Recent applications of the approach include the derivation of new moment restrictions in binary and discrete choice models, both static and dynamic (honore2020moment, honore2021dynamic, dano2023transition). In certain models, no exact moment restrictions exist. However, dhaene2023approximate show how to regularize the functional differencing moments to provide restrictions that are satisfied up to a vanishing approximation error as the length of the panel tends to infinity.

The starting point of this paper is the observation that the scope of functional differencing is not limited to panel data, and the approach can be applied to any setting with a parametric conditional distribution involving latent variables. Our main goal is to apply the functional differencing approach to derive moment restrictions on parameters in some network settings. Specifically, we consider linear and binary choice logit models on networks, including a novel “AKM logit model” that provides a counterpart to the AKM estimator of abowd1999high for binary outcomes. In those models, we characterize the available moment restrictions on parameters.

In addition, we study average effects that depend on the joint distribution of unobserved heterogeneity and observed covariates. In panel binary choice models, average effects have been studied by various authors (see, e.g., chernozhukov2013average, davezies2021identification, dobronyi2021identification, aguirregabiria2021identification, and pakel2021bounds). The functional differencing approach can be applied to average effects, as initially shown in the working paper version of bonhomme2012functional. This approach applies to general panel data models where the outcome distribution is parametrically specified and the distribution of heterogeneity and covariates is unrestricted. Here we use it to derive moment conditions on average effects in network settings.

Lastly, as in its panel data applications, in the settings that we consider in this paper the functional differencing approach delivers moment restrictions on parameters, yet it does not guarantee identification or consistent estimation of the parameters given those restrictions. Although, in several examples that we study, identification can be verified directly and analog estimators can be constructed, applying the approach to other models will generally require careful analysis of statistical properties. We only briefly touch on estimation at the end of the paper, and leave a deeper analysis of identification and estimation to future work. At the same time, we see the functional differencing approach as a promising building block for researchers to discover novel moment restrictions and estimators in network settings in the future.

The outline of the paper is as follows. In Section (ref) we present a class of models with multi-sided heterogeneity in networks. In Section (ref) we describe the functional differencing approach. In Sections (ref), (ref), (ref), and (ref) we illustrate the approach with various examples. Finally, in Section (ref) we briefly sketch how to construct estimators based on functional differencing moment restrictions.

Models with multi-sided heterogeneity in networks

In this section we introduce and describe a framework for heterogeneous agents interacting on a network. The subsequent sections show how to derive moment restrictions on parameters and average effects in this framework.

Description and examples

The model consists of two layers: a model of the network, and a model of agents' outcomes on the network.

The network is represented by a graph, or more generally a hypergraph, featuring nodes and edges. Nodes in the network correspond to economic agents, and edges represent their links or collaborations. The network can be static or dynamic, and we will see that, under the assumption that the network is exogenous, a specification of the network formation model is not needed.

A key feature of the framework is the presence of agent-specific types, which govern sorting patterns and affect outcomes, yet are latent to the econometrician.

Throughout the paper, we will refer to economic agents as workers and firms as a leading example. In settings with workers and firms, we model the network as bipartite, links are employment relationships, and both workers and firms are heterogeneous. We show an example in Figure (ref). Employment, job mobility and wages all depend on the workers' and firms' latent types. Prominent models of this kind were proposed in becker1973theory, shimer2000assortative, postel2002equilibrium, and lentz2022anatomy, among many others.

figure[figure omitted — 1,208 chars of source]

Outcomes in the network depend on the agents' latent types, and on covariates that are observed by the econometrician. Outcomes also depend on idiosyncratic errors, or “shocks”.

A key assumption that we will maintain in this paper is that the network is exogenous, in the sense that network links are independent of the shocks conditional on agents' types and covariates. In models with workers and firms, bonhomme2019distributional show that this assumption is satisfied in the models of shimer2000assortative and lentz2022anatomy but that it fails in the model with sequential bargaining of postel2002equilibrium, for example.

Our main focus will be on estimating parameters governing the model of outcomes while conditioning on the network. This approach only restricts the network insofar as exogeneity with respect to idiosyncratic shocks is required. However, it allows for general forms of matching and sorting patterns. Beyond the example of workers and firms, this setup is relevant to other bipartite network settings appearing in the economics of education, empirical finance, and trade (bonhomme2020econometric). Non-bipartite settings, such as models of team production, are also encompassed by this framework (ahmadpoor2019decoding, bonhomme2021teams).

In applications, the network formation model may also be of interest. Since the framework accounts for unobserved heterogeneity, it can be used to study network formation models with additive or non-additive heterogeneity and independent link-specific shocks (graham2017econometric, bickel2009nonparametric). However, since our approach relies on a parametric likelihood for the distribution of outcomes, it is not well-suited to study the determinants of link formation in models with strategic interactions (de2018identifying, gualdani2021econometric, sheng2020structural).

Probabilistic framework

Let $Y$ denote a vector of outcomes, one for every link (or edge) in the network. Let $A$ denote the vector of latent types, one for every agent (or node). Finally, let $X$ denote a matrix indicating which agents are linked together. In models with covariates, which may be agent-specific or link-specific, we include those in $X$.

We postulate a parametric model for outcomes conditional on the latent types and the network links (and possibly covariates),

equation[equation omitted — 106 chars of source]

where $f_{\theta}$ is a parametric distribution indexed by a finite-dimensional parameter vector $\theta$.

We leave the joint distribution of latent types and network links (and possibly covariates) fully unrestricted. Formally,

equation[equation omitted — 60 chars of source]

where $\pi$ is an unknown (i.e., nonparametric) distribution.

The combination of a parametric outcome distribution and a nonparametric distribution of heterogeneity and covariates is common in the panel data literature. Indeed, ((ref))-((ref)) nests the standard panel data case, where $X$ simply indicates which observations correspond to the same worker. Nevertheless, the current framework is not limited to panel data.

A key assumption implied by the specification of a likelihood conditional on $X$ (and $A$) is that the network is assumed exogenous. In panel data models, this corresponds to the assumption of strict exogeneity of covariates. This assumption is commonly relaxed in linear panel models, for example using the sequential moment restrictions introduced by arellano1991some. A subsequent literature allows for sequentially exogenous network links in linear models (e.g., kuersteiner2020dynamic). Allowing for sequential exogeneity of covariates in nonlinear panel data models is still a frontier research area (see bonhomme2023identification).

In a setting with workers and firms, a leading example of ((ref)) is a model of wage determination. Following the pioneering approach of abowd1999high, a linear specification for log-wages is

equation[equation omitted — 77 chars of source]

where $i\in\{1,...,N\}$ are workers, $t\in\{1,...,T\}$ are time periods, $j(i,t)\in\{1,...,J\}$ denotes the firm where $i$ is employed at $t$, $\varepsilon_{it}$ is an idiosyncratic shock, and we are abstracting from exogenous covariates for simplicity. In this model, network exogeneity is often referred to as “exogenous mobility”, meaning that job transitions (represented by the firm indicators $j(i,t)$) are assumed independent of the $\varepsilon_{it}$'s conditional on the $\alpha_i$'s and the $\psi_j$'s.

In order to allow for complementarity patterns, one may be interested in a nonlinear extension of ((ref)), such as a Constant Elasticity of Substitution (CES) specification in logarithms,

equation[equation omitted — 157 chars of source]

Nonlinear log wage specifications have been proposed and estimated by bonhomme2019distributional and lentz2022anatomy. However, a difference between their approaches and the one we propose here is that those authors restrict the distribution of heterogeneity and how it relates to employment relationships and job transitions. That is, in both papers, $\pi$ in ((ref)) is restricted, while it is fully unrestricted in the current framework. In addition, both bonhomme2019distributional and lentz2022anatomy rely on large firms to recover firm heterogeneity, whereas our approach will yield valid moment restrictions irrespective of the degree of sparsity of the network.

In models ((ref)) and ((ref)) one can form an $NT\times (N+J)$ matrix $X$, by having each row denoting an observation (or edge) $(i,t)$, and having both sets of agents (or nodes) $i$ and $j$ in the columns of $X$. In this case, one can equivalently write ((ref)) and ((ref)) as $$(3')\,\,\, Y=XA+\varepsilon,\quad \text{ and }\quad (4')\,\,\, Y=g\left(X,A,\lambda,\rho,\gamma\right)+\varepsilon,$$ where the vector $A$ contains the $\alpha$'s and $\psi$'s, and the function $g$ takes a known (CES) form. Since we assume that the model is parametric, we specify the distribution of $\varepsilon$ conditional on $X$ and $A$, for example as a normal distribution with zero mean and diagonal covariance matrix with variance $\sigma^2$. In this case, $\theta=\sigma^2$ in the linear model, and $\theta=(\lambda,\rho,\gamma,\sigma^2)$ in the CES model.

The framework in ((ref))-((ref)) also nests certain models of network formation. As an example, consider the model of link formation in graham2017econometric. Undirected binary links $Y_{ij}\in\{0,1\}$ between agents $i$ and $j$ are determined based on covariates $X_{ij}$ and agent-specific types $A_i$ and $A_j$ as

equation[equation omitted — 115 chars of source]

where the $\varepsilon_{ij}$ are standard logistic, independent of $A$ and $X$, and i.i.d. across pairs of agents $(i,j)\in\{1,...,N\}^2$, $i\neq j$. graham2017econometric leaves the joint distribution of $X$ and $A$ unrestricted, and his model is thus a special case of the framework in ((ref))-((ref)).

Functional differencing: A general presentation

In the framework ((ref))-((ref)) we ask two questions. First, how can one derive moment restrictions on $\theta$? Second, how can one find moment restrictions on quantities depending on $\theta$ and $(A,X)$, such as average effects? Both questions can be answered using the functional differencing approach, and we address them in turn.

Parameter $\theta$

To answer the first question, the following proposition shows how to characterize moment restrictions on $\theta$ that are valid irrespective of the heterogeneity distribution.

propositionLet $\phi_{\theta}(y,x)$ be a function of outcomes and network links (and/or covariates). Suppose that $\mathbb{E}\left[\phi_{\theta}(Y,X)\,|\, A,X\right]$ is bounded. The following three statements are equivalent: (i) For any joint distribution of $(A,X)$ we have $$\mathbb{E}\left[\phi_{\theta}(Y,X)\right]=0.$$ (ii) Almost surely in $A,X$, $$\mathbb{E}\left[\phi_{\theta}(Y,X)\,|\, A,X\right]=0.$$ (iii) For almost all $a$ and $x$, \begin{equation}\int \phi_{\theta}(y,x)f_{\theta}(y\,|\, x,a)dy=0.\end{equation}

It is immediate that (ii) implies (i), and that (ii) and (iii) are equivalent. The implication from (i) to (ii) comes from the fact that we require the moment restriction to hold irrespective of the distribution of $A$ and $X$. A formal argument is provided in Appendix (ref). Note also that (ii) implies moment restrictions conditional on $X$, $\mathbb{E}\left[\phi_{\theta}(Y,X)\,|\, X\right]=0$. Proposition (ref) implies that, to look for moment restrictions on $\theta$, it is necessary and sufficient to find solutions to the linear functional equation in ((ref)).

Proposition (ref) is the main insight of the functional differencing approach (bonhomme2012functional). Indeed, finding a $\phi$ satisfying ((ref)) amounts to finding an element in the null space of the conditional expectation operator associated with the parametric conditional model of $Y$ given $(A,X)$. To this end, bonhomme2012functional proposed numerical projection methods, while honore2020moment relied on symbolic computing to find analytical $\phi$ functions in dynamic discrete choice settings. In some models, however, one can show that no non-trivial solution $\phi$ exists, which implies the absence of informative moment equality restrictions on $\theta$. To deal with such cases, dhaene2023approximate propose an approximate functional differencing approach. In dynamic panel data logit models, dobronyi2021identification derive the identified set on the parameters, which combines moment equality and inequality restrictions.

While most applications of the functional differencing approach so far are confined to panel data settings without interactions between agents, Proposition (ref) makes clear that the scope of the approach is not limited to those settings. A conventional panel data model consists of a collection of individual-specific submodels indexed by the individual heterogeneity $A_i$, the vector of individual covariates $(X_{i1}',...,X_{iT}')'$, and the parameter $\theta$ that is common across individuals. In contrast, in network settings all units may potentially be related, and both the network matrix $X$ and the vector of heterogeneity $A$ may affect all outcomes in the network. As Proposition (ref) illustrates, this difference between panel data and network settings is immaterial from the point of view of the applicability of functional differencing. In the next sections we will provide examples to illustrate the usefulness of the functional differencing approach when applied to networks.

Average effects

We now turn our attention to linear functionals of the distribution of $A$ and $X$, and answer the second question. Using functional differencing, we show how to obtain, for given $\theta$, moment restrictions on an average effect of the form $$\mu=\mathbb{E}[m_{\theta}(A,X)],$$ where $m_{\theta}(\cdot)$ is known given $\theta$, and the expectation is taken with respect to the joint distribution $\pi$ of $(A,X)$. As an example, in the CES model of wage determination ((ref)), one may be interested in estimating average marginal effects of worker or firm heterogeneity on log wages, while accounting for the presence of complementarity between worker and firm effects. The following proposition provides a counterpart to Proposition (ref) for such target parameters.

propositionLet $\psi_{\theta}(y,x)$ be a function of outcomes and network links (and/or covariates). Suppose that $\mathbb{E}\left[\psi_{\theta}(Y,X)\,|\, A,X\right]-m_{\theta}(A,X)$ is bounded. The following three statements are equivalent: (i) For any joint distribution of $(A,X)$ we have $$\mathbb{E}\left[\psi_{\theta}(Y,X)\right]=\mathbb{E}[m_{\theta}(A,X)].$$ (ii) Almost surely in $A,X$, $$\mathbb{E}\left[\psi_{\theta}(Y,X)\,|\, A,X\right]=m_{\theta}(A,X).$$ (iii) For almost all $a$ and $x$, \begin{equation}\int \psi_{\theta}(y,x)f_{\theta}(y\,|\, x,a)dy=m_{\theta}(a,x).\end{equation}

The equivalence between the three parts is again easy to see (see Appendix (ref)), and the usefulness of the result comes from the fact that ((ref)) is a linear functional equation. In this case as well, the linear operator in ((ref)) is known given $\theta$. Proposition (ref) shows that the functional differencing approach initially applied in bonhomme2012functional to obtain restrictions on model parameters in panel data models can be applied to derive restrictions on average effects in other models with latent variables, including network settings.

Finding a moment representation for $\mu$ amounts to solving a linear functional system, which is a Fredholm integral equation of the first kind. Numerical and analytical methods can be used to construct $\psi$ functions that satisfy ((ref)). In dynamic panel logit models, aguirregabiria2021identification, dobronyi2021identification and dano2023transition show that the computation of average marginal effects is analytically straightforward. However, in models with continuous outcomes the inverse problem in ((ref)) is generally ill-posed (carrasco2007linear, engl1996regularization). Given a solution $\psi$ to the functional system, using it for estimation in a finite sample thus typically requires regularization. In this paper we focus on finding moment functions $\phi$ and $\psi$, and leave a detailed study of estimators and their properties to future work. See Section (ref) for further discussion.

Linear network models

In this section and the next three we illustrate Propositions (ref) and (ref) through various examples. Consider first a linear model with an $X$ matrix that consists of two parts, $X=(X_1,X_2)$, where $X_1$ is a matrix of network links and $X_2$ is a matrix of covariates. An example is the AKM model of abowd1999high, given by an augmented version of ((ref)) that includes covariates. We specify

equation[equation omitted — 122 chars of source]

where $n$ is the number of observations, $I_n$ denotes the $n\times n$ identity matrix, and we denote $\theta=(\beta',\sigma^2)'$.

In this model we will focus on the parameters $\beta$ and $\sigma^2$, and on quadratic forms $\mu=\mathbb{E}[A'QA]$ for a symmetric matrix $Q$. Variance components, which can be written as quadratic forms, are of interest for, e.g., decomposing the variance of log wages into components reflecting worker heterogeneity, firm heterogeneity, and sorting patterns between heterogeneous workers and firms (abowd1999high, card2013workplace, song2019firming).

Parameters $\beta$ and $\sigma^2$

In this subsection we derive moment restrictions on $\beta$ and $\sigma^2$. For this purpose we rely on Proposition (ref). We start by noting that ((ref)) can be equivalently written as

align*[align* omitted — 130 chars of source]

It is useful to introduce the Moore-Penrose pseudo-inverse $x_1^{\dagger}$ of $x_1$, and to write the two orthogonal projectors associated with $x_1$ as $x_1x_1^{\dagger}=u_1u_1'$ and $I_n-x_1x_1^{\dagger}=u_2u_2'$, where $u=(u_1,u_2)$ is orthogonal. Let $v=y-x_2\beta$, $v_1=u_1'v$, and $v_2=u_2'v$.

We then note that ((ref)) can equivalently be written as

align*[align* omitted — 256 chars of source]

Since $u=(u_1,u_2)$ is orthogonal, we equivalently obtain

align*[align* omitted — 242 chars of source]

In Proposition (ref) we are looking for restrictions holding for all real vectors $a$. Since the rows of $u_1'x_1$ are linearly independent, we equivalently search for the following equation being satisfied for all real vectors $b$:\footnote{$b$ has the same dimension as $v_1$.}

align*[align* omitted — 230 chars of source]

This convolution equation has the unique solution:

align*[align* omitted — 113 chars of source]

We have thus shown the following.

propositionIn model ((ref)), the following two statements are equivalent: (i) $\int \phi_{\theta}(y,x)f_{\theta}(y\,|\, x,a)dy=0$. (ii) $\phi_{\theta}(y,x)=\varphi_{\theta}(u_1'(y-x_2\beta),u_2'(y-x_2\beta),x)$, where the function $\varphi_{\theta}$ is such that \begin{align*} &\int \varphi_{\theta}(v_1,v_2,x)\exp\left(-\frac{1}{2\sigma^2}v_2'v_2\right)dv_2=0. \end{align*}

By Proposition (ref), Proposition (ref) characterizes all the available moment restrictions on the parameter $\theta=(\beta',\sigma^2)$ in model ((ref)). Now, there are many possible choices for $\varphi_{\theta}$, leading to many choices of moment functions in this model.

As a first example, let us take $$\varphi_{\theta}(v_1,v_2,x)=u_2v_2.$$ We obtain

eqnarray*[eqnarray* omitted — 163 chars of source]

This implies the following conditional moment restrictions on $\beta$:

equation[equation omitted — 110 chars of source]

In a panel data setting, chamberlain1992efficiency shows that the efficiency bound for $\beta$ based on the quasi-differencing restrictions ((ref)) coincides with the bound based on a semiparametric model where $\varepsilon$ has mean zero but is otherwise not restricted. In our setup, ((ref)) remains valid in general non-Gaussian linear network regression models where $\mathbb{E}[\varepsilon\,|\, X]=0$. Moreover, the Gaussian assumption on $\varepsilon$ provides additional restrictions that are fully characterized by Proposition (ref). Some of those additional restrictions arise from the variance matrix of outcomes.

As a second example, let $n_2=\dim v_2=\mbox{Trace}(I_n-x_1x_1^{\dagger})$, and take $$\varphi_{\theta}(v_1,v_2,x)=v_2'v_2-n_2\sigma^2 .$$ We obtain $$\phi_{\theta}(y,x)=(y-x_2\beta)'[I_n-x_1x_1^{\dagger}](y-x_2\beta)-n_2\sigma^2.$$ This implies the following conditional moment restrictions on $\beta$ and $\sigma^2$:

equation[equation omitted — 138 chars of source]

Note that ((ref)) remains valid under non-Gaussianity, provided $\mathbb{E}[\varepsilon\varepsilon'\,|\, X]=\sigma^2 I_n$. Restrictions akin to ((ref)) were considered in arellano2012identifying in a panel data setting. In a network context, andrews2008high derived an unconditional version of ((ref)), and applied it to the decomposition of the variance of log wages.\footnote{The assumption that the elements of $\varepsilon$ be mutually independent may be empirically restrictive. In applications of AKM, it is common to only rely on between-job-spell variation in log wages in estimation, in order not to restrict the within-job-spell correlation in $\varepsilon$ (kline2020leave, bonhomme2020much).}

Quadratic forms

In this subsection we derive moment restrictions on a quadratic form $\mu=\mathbb{E}[m_{\theta}(A,X)]$, where $m_{\theta}(a,x)=a'Qa$ for an $m\times m$ symmetric matrix $Q$, for $m$ the dimension of $A$. For this purpose we rely on Proposition (ref). We start by noting that ((ref)) is equivalent to

align*[align* omitted — 156 chars of source]

Using the above reparameterization in terms of $(v_1,v_2)$, we equivalently have

align[align omitted — 270 chars of source]

We then have the following result, shown in Appendix (ref).

propositionConsider model ((ref)), and suppose that $X_1$ has full column rank almost surely. Then the following two statements are equivalent: (i) $\int \psi_{\theta}(y,x)f_{\theta}(y\,|\, x,a)dy=a'Qa$. (ii) $\frac{1}{(2\pi \sigma^2)^{\frac{n_2}{2}}}\int \psi_{\theta}(x_2\beta+u_1v_1+u_2v_2,x)\exp\left(-\frac{1}{2\sigma^2}v_2'v_2\right)dv_2=v'(x_1^{\dagger})'Qx_1^{\dagger}v-\sigma^2\mbox{Trace}((x_1^{\dagger})'Qx_1^{\dagger})$.

By Proposition (ref), Proposition (ref) characterizes all the available moment restrictions on $\mu=\mathbb{E}[A'QA]$.\footnote{When $X_1$ is a network matrix, ensuring the assumption that $X_1$ has full column rank often requires to restrict the sample to a connected subnetwork. See abowd1999high and abowd2002computing for methods to compute connected subnetworks in settings with workers and firms.} The proof relies on Fourier transforms. A special case of Proposition (ref) is obtained when $\psi_{\theta}(y,x)$ is a function of $u_1'(y-x_2\beta)$ and $x$ only, which implies

equation[equation omitted — 160 chars of source]

The trace correction in ((ref)) is a well-known formula to obtain unbiased estimators of quadratic forms. andrews2008high, and kline2020leave in a heteroskedastic context, apply such corrections to estimate variance components in log wage variance decompositions.

Moreover, given the particular solution ((ref)), any other solution is of the form $$\widetilde{\psi}_{\theta}(y,x)=\psi_{\theta}(y,x)+\varphi_{\theta}(u_1'(y-x_2\beta),u_2'(y-x_2\beta),x),$$ where, as in Proposition (ref), the function $\varphi_{\theta}$ satisfies

align*[align* omitted — 99 chars of source]
remarkProposition (ref) can be applied to find estimators of other average effects beyond quadratic forms. Consider the quantity $\mu=\mathbb{E}[m_{\theta}(A,X)]$, for some function $m_{\theta}(\cdot)$. Examples are higher-order moments of $A$, its distribution function, or some nonlinear moments of $A$. Using similar arguments to the ones leading to Proposition (ref), one can derive a formula for all moment restrictions on $\mu$, \begin{equation}\mathbb{E}[\psi_{\theta}(Y,X)]=\mu.\end{equation} However, the expression of $\psi_{\theta}$, which we derive in Appendix (ref), depends on the inverse of the Fourier transform operator, and it does not generally admit a closed-form expression when $m_{\theta}(a,x)$ is not polynomial in $a$. In addition, estimating $\mu$ based on ((ref)) generally requires regularizing $\psi_{\theta}$ (as shown in a panel data setting in the working paper version of bonhomme2012functional).

Logit network models

In this section and the next two we study logit models for network data. Letting $X=(X_1,X_2)$, we assume

equation[equation omitted — 150 chars of source]

Model ((ref)) contains several popular binary choice models as special cases. A first example is a static panel data logit model, which obtains when the columns of $X_1$ are individual indicators. Another example is the logistic network formation model of graham2017econometric, see ((ref)), which obtains when $Y$ are link outcomes and the elements of $X_1 A$ are the individual sums $A_i+A_j$, for $i\neq j$.

Equation ((ref)) also covers models of binary outcomes on a network. To illustrate, we will consider the following binary choice counterpart to the AKM model of abowd1999high:

equation[equation omitted — 194 chars of source]

where, as in the linear case, $i$ are workers, $t$ are time periods, and $j(i,t)$ denotes the firm where $i$ is employed at $t$. It is easy to see that ((ref)) can be written in the form ((ref)) for a suitable definition of $X_1$. In this setting, the network is a bipartite multigraph: there may be multiple edges pointing from a worker $i$ to a firm $j$ indicating that worker $i$ was in an employment relationship with firm $j$ over multiple periods.

While, in many applications of AKM, outcomes $Y_{it}$ are log earnings or log wages, it may be of interest to account for worker and firm heterogeneity when studying other labor market outcomes. For example, lachowska2023work apply AKM to the analysis of log hours worked. In this setting, the AKM logit specification ((ref)) could be employed to analyze the determinants of part-time and full-time work using a binary measure of working time. In other applications, one may be interested in applying AKM to study determinants of worker promotions (e.g., benson2019promotions) or of the type of labor contract such as fixed-term or permanent contract (e.g., guell2007binding). The logit specification ((ref)) can also be useful in applications of AKM to other fields (including education, innovation, urban economics, trade, and empirical finance) where binary outcomes are common.

In the next section we will first focus on the covariates' coefficients $\theta$ in model ((ref)). For example, margolis1996cohort studies the earnings returns to seniority in France while accounting for worker and firm heterogeneity. A binary specification such as ((ref)) allows one to document the effects of seniority on binary labor market outcomes, such as working part-time or full-time, being awarded a promotion, or working under a permanent or temporary contract. In Section (ref) we will study average effects, which are functions of worker and firm heterogeneity. We will see that deriving non-trivial moment restrictions on average effects seems more challenging than obtaining informative restrictions about the $\theta$ parameter.

Model parameters in logit network models

In this section we first provide a characterization of all moment restrictions available on $\theta$ in model ((ref)), and then discuss several examples.

Characterization

We start by noting that, in model ((ref)), ((ref)) can be equivalently written as

align*[align* omitted — 217 chars of source]

where we have denoted as $y_i$ the $i$th element of $y$, $x_{i1}'$ the $i$th row of $x_1$, and $x_{i2}'$ the $i$th row of $x_2$, for $i\in\{1,...,n\}$.

Equivalently, we have

align*[align* omitted — 105 chars of source]

that is,

align*[align* omitted — 108 chars of source]

Letting, for all $x_1$, ${\cal{S}}_{x_1}=\{x_1'y\, : \, y\in\{0,1\}^n\}$ denote the set of possible values of $x_1'y$, this implies that ((ref)) can be equivalently written as

align*[align* omitted — 167 chars of source]

Finally, since $\exp(s'a)$, for $s\in{\cal{S}}_{x_1}$, are linearly independent functions of $a$, we obtain the following characterization.

propositionIn model ((ref)), the following two statements are equivalent: (i) $\int \phi_{\theta}(y,x)f_{\theta}(y\,|\, x,a)dy=0$. (ii) $\sum_{y\in\{0,1\}^n}\boldsymbol{1}\left\{x_1'y=s\right\}\phi_{\theta}(y,x)\exp\left(y'x_2\theta\right)=0$, for all $s\in{\cal{S}}_{x_1}$.

Proposition (ref) provides an exhaustive characterization of available moment restrictions in the logit model ((ref)). For a non-zero $\phi$ to exist, it is necessary that, for some $s\in{\cal{S}}_{x_1}$, $y'x_2$ varies given that $x_1'y=s$. In this model, $S=X_1'Y$ is sufficient for $A$. Below we will illustrate Proposition (ref) using several examples: the static panel data logit model, the logistic network formation model, and the AKM logit model.

Panel data: conditional logit

Consider first the panel data model

equation[equation omitted — 196 chars of source]

Let $i\in\{1,...,N\}$, $y_i=(y_{i1},...,y_{iT})'$ and $x_{i}=(x_{i1}',...,x_{iT}')'$. For simplicity we search for functions $\phi_{\theta}(y_i,x_{i})$ that only depend on $(y,x)$ through $(y_i,x_i)$. We thus look for $\phi_{\theta}(y_i,x_{i})$ such that

equation[equation omitted — 201 chars of source]

Consider the $T=2$ case, and take $s=1$. We obtain $$\phi_{\theta}(1,0,x_i)\exp\left(x_{i1}'\theta\right)+\phi_{\theta}(0,1,x_{i})\exp\left(x_{i2}'\theta\right)=0,$$ which coincides with the moment function of conditional logit (rasch1960studies, andersen1970asymptotic). When $T>2$ we recover additional moment restrictions, as in davezies2020fixed.

Network formation: tetrad logit

Consider next the logistic network formation model ((ref)) introduced in graham2017econometric. Links are undirected,\footnote{Logit models with directed links (charbonneau2017multiple) have a similar structure.} and there are $n=N(N-1)/2$ observations (one for each dyad), where $N$ is the number of agents. In this model, the sufficient statistic $s=x_1'y$ in Proposition (ref) is the vector of degrees, i.e., the degree sequence of the network.

For simplicity we focus on functions of tetrads formed by four agents $(i,j,k,\ell)$. We thus look for $\phi_{\theta}(y_{ij},y_{ik},y_{i\ell},y_{jk},y_{j\ell},y_{k\ell},x)$ that satisfy

align*[align* omitted — 468 chars of source]

Taking first $s_1=s_2=s_3=s_4=1$, we obtain

align[align omitted — 309 chars of source]

Considering next $s_1=s_2=s_3=s_4=2$, we obtain

align[align omitted — 370 chars of source]

Finally, for $s_1=s_2=2,s_3=s_4=1$, we obtain

align[align omitted — 241 chars of source]

and there will be analogous restrictions associated with permutations of the degree sequence (2,2,1,1). All other possible degree sequences have no identifying content for $\theta$.

Together, taking $\phi$ as in ((ref)), ((ref)), and ((ref)) (alongside its permutations), implies the moment restrictions underpinning the “tetrad logit” estimator in graham2017econometric. However, Proposition (ref) clarifies that these restrictions may not be unique, and it provides all available moment restrictions in the logistic network formation model.

Binary choice on a network: AKM logit

In this subsection we derive moment restrictions on $\theta$ in model ((ref)). We focus on the case where $X_{it}$ does not vary within job spells. For example, when controlling for the worker's age, job seniority only varies between spells.\footnote{If $X_{it}$ does vary within spells, then the conditional logit estimator can be used for consistent estimation of $\theta$.} We focus the analysis on the $T=2$ case, and we consider several subnetwork configurations of the data displayed in Figure (ref). Given a subnetwork configuration, we verify if moment conditions on $\theta$ exist, and what form they take.

figure[figure omitted — 4,448 chars of source]

\paragraph{Configuration A: one worker in the same firm.}

Suppose worker $i$ stays in the same firm $j$ in both periods. We look for $\phi_{\theta}(y_{it},y_{i,t+1},x)$ such that, for $s\in\{0,1,2\}$,

align*[align* omitted — 182 chars of source]

Given that $x_{2it}=x_{2i,t+1}$ (since the covariate does not vary within spell), this implies

align*[align* omitted — 168 chars of source]

Hence $\theta$ drops out from the equation, and there is no information to estimate $\theta$ in this configuration.

\paragraph{Configuration B: one worker moving between two firms.}

Suppose worker $i$ moves between firms $j$ and $j'$. We look for $\phi_{\theta}(y_{it},y_{i,t+1},x)$ such that

align*[align* omitted — 195 chars of source]

However, in this case, $(s_1,s_2)$ fully determines $(y_{it},y_{i,t+1})$. Hence, for each $(s_1,s_2)$ we obtain $\phi_{\theta}(y_{it},y_{i,t+1},x)=0$, which shows there is no information about $\theta$ in this configuration.

\paragraph{Configuration C: two workers moving between the same two firms.} Suppose workers $i$ and $i'$ both move between the same firms $j$ and $j'$. We look for a function $\phi_{\theta}(y_{it},y_{i,t+1},y_{i't},y_{i',t+1},x)$ such that

align*[align* omitted — 344 chars of source]

It turns out that, in this subnetwork configuration, there exist non-trivial moment restrictions on $\theta$. To see this, take $s_1=s_2=s_3=s_4=1$. We obtain

align*[align* omitted — 159 chars of source]

This implies the conditional moment restriction

align*[align* omitted — 242 chars of source]

\paragraph{Configuration D: two workers moving between different firms.}

Suppose worker $i$ moves between $j$ and $j'$, and worker $i'$ moves between different firms $j''$ and $j'''$. We look for $\phi_{\theta}(y_{it},y_{i,t+1},y_{i't},y_{i',t+1},x)$ such that

align*[align* omitted — 369 chars of source]

It is easy to see there is no non-trivial $\phi$ function in this case. Intuitively, since workers never share a firm, it is not possible to “difference out” the firm component of heterogeneity.

\paragraph{Configuration E: two workers moving to different firms from the same firm.}

Suppose worker $i$ moves between $j$ and $j'$, and worker $i'$ moves from the same firm $j$ to a different firm $j''$. We look for $\phi_{\theta}(y_{it},y_{i,t+1},y_{i't},y_{i',t+1},x)$ such that

align*[align* omitted — 363 chars of source]

It is easy to see there is no information about $\theta$ in this configuration.

\paragraph{Configuration F: three workers in a loop.}

There are many other subnetwork configurations providing information beyond configuration C. Indeed, consider three workers who move as follows: $i$ moves between firms $j$ and $j'$, $i'$ moves between $j'$ and $j''$, and $i''$ moves between $j''$ and $j$. We look for $\phi_{\theta}(y_{it},y_{i,t+1},y_{i't},y_{i',t+1},y_{i''t},y_{i'',t+1},x)$ such that

align*[align* omitted — 510 chars of source]

Taking $s_1=s_2=s_3=s_4=s_5=s_6=1$, one obtains

align*[align* omitted — 212 chars of source]

This implies the conditional moment restriction

align*[align* omitted — 300 chars of source]

Average effects in logit network models

In this section we again consider model ((ref)), and we study average effects of the form $$\mu=\mathbb{E}[m_{\theta}(A,X)],$$ for some known function $m_{\theta}$. In this case, ((ref)) can be equivalently written as

align[align omitted — 220 chars of source]

This equation characterizes the set of available moment restrictions on $\mu$, i.e., the set of $\psi$ functions such that ((ref)) holds.

As a simple example, consider the case where $T=2$ in the static panel logit model ((ref)), with a binary covariate $X_{it}$, and consider

align*[align* omitted — 181 chars of source]

so that $\mu$ is an average partial effect. We show in Appendix (ref) that no function $\psi$ satisfies ((ref)). Intuitively, this comes from the fact that the distribution of $A$ given $X_{i1}=X_{i2}$ (i.e., for “stayers”) is unidentified.

In contrast, as we also show in Appendix (ref), the average partial effect of “movers”, corresponding to

align*[align* omitted — 140 chars of source]

admits a characterization as in ((ref)), whenever $\psi$ satisfies

align[align omitted — 248 chars of source]

A simple example satisfying those conditions is $$\psi_{\theta}(y_1,y_2,x)=(x_2-x_1)(y_2-y_1),$$ as pointed out (in a more general nonparametric model) by chernozhukov2013average. However, ((ref)) and ((ref)) imply additional moment restrictions. For example, one can take

align*[align* omitted — 154 chars of source]

which provides an additional moment restriction on $\mu$ under the logit model's assumptions.

It appears difficult to obtain moment equality restrictions on average partial effects in logit models on networks outside of the panel data case. As an example, consider the subnetwork configuration C in Figure (ref). In this case we have seen in the previous section how to obtain moment restrictions on $\theta$. However, we show in Appendix (ref) that no function $\psi$ satisfies ((ref)) for the average partial effect corresponding to

align*[align* omitted — 122 chars of source]

where, in this model, $a_1$ is worker $i$'s fixed effect and $a_3$ is firm $j$'s fixed effect.

In models where no functional differencing restrictions are available, one may still be able to construct bounds on the average effect of interest. In panel data settings, this strategy was pursued by chernozhukov2013average, davezies2021identification, and dobronyi2021identification, among others. However, implementing bounds approaches often requires estimating conditional moments given $X$. When $X$ represents a network matrix, conditional moment estimation may be especially challenging. In a panel data setting, pakel2021bounds propose a bounding strategy that avoids the curse of dimensionality associated with conditioning covariates. Extending their approach to network settings is an interesting question for future work.

Lastly, in this section we have focused on binary choice models. The situation may be more favorable, in the sense of there existing informative functions $\phi$ and $\psi$, in models with continuous outcomes such as the CES specification ((ref)).

Remarks on estimation

To close our discussion, we briefly outline some possibilities for estimation of parameters and average effects, without providing details.

Given a moment function $\phi$ as in Proposition (ref), and a realization $(y,x)$ from the joint distribution of $(Y,X)$, one can estimate $\theta$ based on $$\widehat{\theta}=\underset{\theta}{\mbox{argmin}}\, \left\|\phi_{\theta}(y,x)\right\|,$$ for some norm $\|\cdot\|$. In some models, this approach will deliver familiar estimators. For example, in the linear model ((ref)), an estimator of $\beta$ based on ((ref)) is the “quasi-differencing” estimator

equation[equation omitted — 120 chars of source]

and an estimator of $\sigma^2$ based on ((ref)) is the “degree-of-freedom-corrected” estimator

equation[equation omitted — 161 chars of source]

When constructing a function $\phi$ using the entire data is impractical, one can construct a set of functions $\phi_{\theta}^{(k)}(y,x)$ that depend on $y$ and $x$ only though a subset of the data. An estimator of $\theta$ is then $$\widehat{\theta}=\underset{\theta}{\mbox{argmin}}\, \left\|\sum_{k=1}^K\phi_{\theta}^{(k)}(y,x)\right\|.$$ In the logit network formation model ((ref)), taking $\phi$ as in ((ref)), ((ref)), and ((ref)) (alongside its permutations), leads to estimators in the spirit of the tetrad logit estimator of graham2017econometric.

When focusing on average effects, a possible estimation approach based on Proposition (ref) consists in setting $$\widehat{\mu}=\psi_{\widehat{\theta}}(y,x),$$ for some estimator $\widehat{\theta}$. For example, in the linear model ((ref)), an estimator of the quadratic form $\mu=\mathbb{E}[A'QA]$ based on ((ref)) is

equation*[equation* omitted — 168 chars of source]

where $\widehat{\beta}$ and $\widehat{\sigma}^2$ are given by ((ref)) and ((ref)), respectively. This corresponds to the bias-corrected estimator of andrews2008high. For other average effects, regularization is typically needed for reliable estimation.

For all these estimators, there are important questions that remain to be addressed. What are their asymptotic properties (under suitable assumptions on how the network grows with the sample size)? How to conduct feasible inference on the population parameters? And, out of the available functions $\phi$ and $\psi$, how to choose a small subset of those (for tractability) without sacrificing too much precision (for efficiency)? Answering these questions will be an important task for future work.