Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
108,488 characters · 18 sections · 0 citation commands
Teams: Heterogeneity, Sorting, and Complementarity
\baselineskip21pt
\setcounter{page}{0}\thispagestyle{empty}
How to infer individual contributions when only team output is observed? This question is central to a number of areas, ranging from labor economics and the economics of innovation to sports and music. My goal in this chapter is twofold: I propose an econometric framework to estimate how team performance depends on its workers, and I apply it to study the production of economic research and the impact of individual inventors on the quality of their patents.
I focus on three aspects of the relationship between workers and teams. The first one is worker heterogeneity, which shapes team performance. The approach I propose relies on a simple intuition: when individuals work for different teams over time, the variation in team composition will be informative about which workers contribute the most to their teams and which ones contribute the least. Hence, the network structure of the data will be crucial in order to infer worker heterogeneity.
The second aspect is the sorting of workers into teams. Both heterogeneity and sorting contribute to the variation in output. I show in a model where production in additive in worker inputs that the two can be separately identified. The approach I rely on generalizes the “plus-minus” metric in US team sports (see Hvattum, 2019): intuitively, when a worker $i'$ joins a team and a worker $i$ leaves it, the difference in output identifies the difference between the types of $i'$ and $i$. More generally, I show how, in collaboration networks, worker types can be identified by using their participation in different teams, including when they work on their own.
The third aspect is the complementarity between workers. The presence of complementarity is important for policies that influence the reallocation of workers across teams. Complementarities are also central to theories of sorting (e.g., Becker, 1973, Garicano and Rossi-Hansberg, 2006, Eeckhout and Kircher, 2018). To identify complementarities empirically, I build and estimate nonlinear models of team production, which rely on the assumption that worker types are discrete. This framework allows me to jointly document the role of heterogeneity, sorting and complementarity.
In the analysis I will draw connections with methods for matched employer-employee data (e.g., Abowd and Kramarz, 1999). In that literature, both firm types and worker types contribute to the wage outcome, which is observed at the worker level, and workers do not interact with one another inside the firm.\footnote{With some exceptions, such as the recent work by Herkenhoff et al. (2018).} The worker-firm structure is thus a bipartite graph (Bonhomme, 2017). In contrast, in the collaboration networks I study here, workers interact with one another in the team. In addition, while I assume away team effects, so the team in my framework is simply a collection of workers, the outcome is only observed at the team level. Collaboration networks follow a hypergraph structure where teams vary in size (e.g., Turnbull et al., 2019, Lerner et al., 2019). Despite these differences, my framework will leverage insights from the matched employer-employee data literature. As an example, when production is additive in worker inputs, I will estimate worker types using linear regression, in the spirit of Abowd et al. (1999).
There is a fourth important aspect in the relationship between workers and teams: the presence of other factors, beyond worker types, that affect output. These other factors are lumped into the error term of the model. Their presence prevents one from recovering all worker types precisely, especially in data sets where the number of productions per worker is small and collaborations are sparse. As a result, regression-based estimates of sorting and heterogeneity may be biased. To correct the bias in additive models, I implement the method of Andrews et al. (2008) developed for matched employer-employee data.\footnote{Kline et al. (2020) generalize this approach to allow for unspecified heteroskedasticity. The nature of the bias is formally studied in Jochmans and Weidner (2020).} Bonhomme et al. (2020) find using samples from various countries that the bias can be substantial. Similarly, when estimating additive models in networks of collaborations, I find that bias correction is important.
A main goal of this chapter is to propose methods that account for complementarity between workers. Bonhomme et al. (2019, BLM) introduce a framework to allow for complementarity between workers and firms in matched employer-employee data. Here I build on this work to allow for complementarity between workers within a team. Like in BLM, I assume that heterogeneity takes a finite number of values. However, while BLM model firm heterogeneity as discrete and estimate firm groups using the kmeans algorithm, in my setting estimates of worker types are not sufficiently precise to follow a similar grouped fixed-effects approach. For this reason, I model the distribution of the discrete worker types using a random-effects approach.
The presence of unobserved worker heterogeneity in a network of teams makes estimation of nonlinear random-effects models challenging. Indeed, the likelihood function is non-separable, since the same worker may participate in multiple teams, and teams may contain multiple workers. To lower the computational cost, I use a mean-field variational method, which relies on a fully factored approximation. Variational estimators are widely used in networks and other complex data settings (Bishop, 2006, Blei et al., 2017). In stochastic blockmodels, which are closely related to the model I study, Bickel et al. (2013) provide conditions for consistency. In my model I perform Monte Carlo simulations to probe the accuracy of the variational method in finite samples.\footnote{Codes to implement the methods are available on my \href{https://sites.google.com/site/stephanebonhommeresearch/}{\color{blue}{webpage}}.}
The framework can be extended in different ways. In particular, I consider adding a model of team formation on top of the model of team production. To do so, I implement a joint random-effects approach using a stochastic blockmodel as a simple statistical model of team formation. An interesting avenue for future work will be to combine the framework with economic models of team formation. It will also be important to understand under which conditions the methods I propose are robust to the team formation model being misspecified.
I apply these methods to two empirical settings: economic research, and patents and innovation. Collaboration in scientific research is an extensively studied topic (e.g., Furman and Gaule, 2013). Aspects of the network of collaborations between economists have been studied by Goyal et al. (2006), Fafchamps et al. (2010), Hsieh et al. (2018), and Anderson and Richards-Shubik (2019), among others. Here I focus on the impact of individual researchers and the shape of the production function. In the literature on patents and innovation, there is increasing interest in the role of individual inventors in innovation as an engine of endogenous growth (e.g., Akcigit et al., 2017, Bell et al., 2019, Pearce, 2019). Relative to this literature, I propose different measures of the type of an inventor -- that is, a measure of her “quality” -- which I infer from the network of patent collaborations using econometric techniques.
I estimate the models using a sample of economists constructed by Ductor et al. (2014), and a sample of inventors and their patents constructed by Akcigit et al. (2016). In both samples, I find that an additive model of team production is useful to highlight the presence of heterogeneity and sorting. However, the additive model is only an imperfect tool to analyze these data. Indeed, for both economists and inventors, I find evidence of complementarity between high-type workers who collaborate in a team. These findings suggest that future research will need to account for complementarity and sorting when measuring individual contributions of researchers or inventors to their teams.
Related approaches have been proposed in the literature. Relative to the analysis of peer effects and spillovers (in particular, Arcidiacono et al., 2012, and Arcidiacono et al., 2017), a key difference in my context is that I only observe team output. In recent work, Devereux (2018) uses fixed-effects methods to estimate the link between individual productivity and team performance in tennis games. Most closely related to this chapter, Ahmadpoor and Jones (2019) use fixed-effects methods in nonlinear models to document heterogeneity and complementarity patterns with a focus on research, innovation and patenting.\footnote{Related methods have been used to estimate manager fixed-effects (e.g., Bertrand and Schoar, 2003).} While I build on this work, a specificity of my approach is that it explicitly accounts for the presence of other unobserved inputs in addition to the worker types, which motivates the use of new econometric methods. Moreover, estimates of nonlinear models with discrete types uncover novel forms of complementarity between workers. Another recent related article is Weidmann and Deming (2020), who design and analyze experiments to identify individual contributions to teamwork.
The outline of the chapter is as follows. I describe the framework in Section (ref), and study additive and nonlinear models in Sections (ref) and (ref), respectively. I apply the methods to study economic researchers in Section (ref), and inventors and their patents in Section (ref). Finally, I conclude in Section (ref).
Consider a set of nodes $1,...,N$, which represent workers who collaborate and produce output in teams. Nodes are linked to each other by hyperedges that represent the collaborations between workers. In a standard graph, an edge links two nodes $i$ and $i'$. In a hypergraph of collaborations, hyperedges can link a worker $i$ to herself when $i$ produces on her own, two workers $i$ and $i'$ who produce together, three workers $i,i',i''$ who produce together, and so on. I will refer to hyperedges as “teams”, and denote the workers in a team $j$ of size $n_j=n$ as $(i_1(j),...,i_n(j))$. I denote the number of teams as $J$.
As an example, Figure (ref) represents a hypergraph of collaborations with five teams and five workers. Workers 3 and 5 produce on their own. In addition, worker 3 produces jointly with worker 4, and workers 4 and 5 produce in a team together with worker 2. Lastly, worker 1 produces with worker 2 in a team.
Worker $i$ contributes to the team a quantity $\alpha_i$, which is unobserved to the econometrician. I assume $\alpha_i$ is constant across collaborations, and I refer to $\alpha_i$ as the type of worker $i$. Constancy of worker types over time is a key feature of my framework. However, assuming that individual types remain constant may be restrictive.\footnote{To relax this assumption, one could rely on hidden Markov models, which allow types to evolve as Markov processes, or one could alternatively rely on mixed membership models (e.g., Airoldi et al., 2008), which allow workers' types to depend on the team they participate in. I do not study these possibilities here.} For this reason, in the empirical analysis I will use observations from at most five consecutive years. Note that, while worker types are constant, type composition varies across teams.
The output of a team $j$ with $n$ workers is denoted as $Y_{nj}$, and it is given by
where $\phi_n$ is the production function of an $n$-worker team, and $\varepsilon_{nj}$ represent other factors, or shocks, unobserved to the econometrician, that affect team output beyond workers' inputs $\alpha_i$.\footnote{Here and in the following, $\alpha_{i_1(j)}=\sum_{i=1}^N\boldsymbol{1}\{i_1(j)=i\}\alpha_i$, and similarly for the other workers in team $j$.}$^{,}$\footnote{Since $n=n_j$ is the size of team $j$, an alternative notation for ((ref)) could be $$\widetilde{Y}_{j}=\phi_{n_j}(\alpha_{i_1(j)},...,\alpha_{i_n(j)},\widetilde{\varepsilon}_{j}),$$ where $\widetilde{Y}_{j}=Y_{n_j,j}$ and $\widetilde{\varepsilon}_{j}=\varepsilon_{n_j,j}$.} I assume that the positions of workers in the team do not matter, so $\phi_n$ is symmetric with respect to its first $n$ arguments. Hence, the framework abstracts from within-team hierarchy and organization.
I now state the first key assumption. Throughout, $\{A_k\,:\, k\}$ denotes the set of $A_k$ for all $k$.
Assumption (ref) restricts the team formation process. It states that team formation is independent of the team-specific shocks $\varepsilon_{nj}$, conditional on the worker types $\alpha_1,...,\alpha_N$. Hence, while the probability of joining a team may depend unrestrictedly on worker types and other factors that are independent of the $\varepsilon_{nj}$'s, it cannot depend on the $\varepsilon_{nj}$'s themselves. Assumption (ref) will thus fail if, before forming a team or joining one, workers have advance information about team-specific shocks. This assumption is important for the tractability of the approach. Yet it is restrictive, and in future work it will be important to relax it.\footnote{A possibility to relax Assumption (ref) is a joint random-effects approach, where production and team formation are estimated together (see Section (ref)).}
I now state the second key assumption.
Assumption (ref) rules out dependence among team-specific shocks. As an implication, shocks to a team of workers who collaborate repeatedly over time are assumed serially independent. An alternative approach, which I do not study in this chapter, would be to model the dependence, for example using a parametric process for $\varepsilon_{nj}$. More generally, the framework does not allow for dynamics in team production, such as state dependence effects.
Another implication of Assumption (ref) is the absence of “team effects” in the model. In the framework, a team $j$ is only a collection of workers, and there is no effect of $j$ per se except through the independent shocks $\varepsilon_{nj}$. In particular, the model does not allow for team-specific capital (Jaravel et al., 2018). In applications where firms (e.g., in R&D) or team identities (e.g., in sports) play an important role, it will be important to extend the framework to allow for team types and other team-specific factors, in addition to worker types.
Lastly, another assumption implicit in ((ref)) is the absence of covariates, except for team size $n$. Covariates can be incorporated as additional inputs to production. In the empirical applications, I will account for time effects, and for age effects in robustness checks.
To specify and estimate the model, I will use two approaches. In the first one (“fixed-effects”, FE) I will treat the $\alpha_i$'s as parameters and estimate them. The FE approach has the advantage of not requiring to specify a model of team formation. Moreover, when team output is additive in worker types and Assumption (ref) holds, the FE approach is tractable. However, in the empirical collaboration networks that I will analyze, $\alpha_i$ estimates will often be very noisy. The prevalence of sample noise will require using methods for bias reduction.
In the second approach (“random-effects”, RE) I will model the joint distribution of the $\alpha_i$'s. I will use this approach to estimate nonlinear models with complementarity between worker types. The RE approach accounts naturally for sample noise. However, it requires specifying how the worker types depend on the collaborations in the network. In other words, the RE approach requires making assumptions about team formation beyond Assumption (ref). I will compare several approaches based on simple statistical models.
In this section and the next, I present methods for additive production and nonlinear production, in turn.
Suppose that the production function in ((ref)) takes the following additive form
where $\lambda_n$ is a team-size scaling factor. I adopt the normalization $\lambda_1=1$. In ((ref)), output depends on the sum of worker types, or equivalently on the mean worker type in the team (given the presence of $\lambda_n$).\footnote{In model ((ref)), types $\alpha_i$ are scalar, and their effects on output in teams of varying sizes are proportional to each other. The nonlinear model in Section (ref) will allow for different patterns, possibly non-monotonic across teams of different sizes with varying worker composition.}
It is useful to write ((ref)) in vector form as $Y_n=\lambda_nA_n+\varepsilon_n$; that is, equivalently, stacking all teams and team sizes together,
where $D_{\lambda}$ is a diagonal matrix, and $A$ is a matrix of zeros and ones where a one indicates that a worker (i.e., a column) participates in a team (i.e., a row). Network exogeneity then takes the form
where $\mu_n$ is a team-size-specific constant. In the empirical analysis, I will estimate ((ref)) under the assumption that $\mu_n=0$. In addition, in a robustness check, I will report results from the following alternative specification in logarithms,
To analyze identification in this setting, I treat $\alpha$, $\lambda$, $\mu$, and $A$ as non-stochastic quantities.\footnote{Hence the conditioning on $(A,\alpha)$ in ((ref)) could be removed from the notation.} Suppose, to start with, that the $\lambda_n$'s are known, and that $\mu_n=0$ in ((ref)). Since model ((ref))-((ref)) is a linear regression, the identification of the worker effects $\alpha_1,...,\alpha_N$ is determined by the rank properties of $A$.
As an example, consider the matrix $A$ corresponding to Figure (ref): $$A=\left(
\right).$$ Since this matrix is non-singular, all $\alpha_1,...,\alpha_5$ are identified. An intuitive algorithm to show identification in this case is as follows. Let ${\cal{S}}$ denote the set of workers $i$ for which $\alpha_i$ is identified. First, focus on workers who produce on their own: since workers $3$ and $5$ work on their own, include them in ${\cal{S}}$. Then, focus on teams where the workers in ${\cal{S}}$ collaborate with others: since workers $3$ and $4$ work together, include worker $4$ in ${\cal{S}}$. Repeat this operation: since workers $2$, $4$ and $5$ work together, include worker $2$ in ${\cal{S}}$, and since workers $1$ and $2$ work together, include worker $1$ in ${\cal{S}}$. Hence ${\cal{S}}=\{1,2,3,4,5\}$.
To generalize the approach beyond this simple example, let $B=D_{\lambda}A$, continuing with the case where $D_{\lambda}$ is known. A worker type $\alpha_i$, for $i\in\{1,...,N\}$, is identified if and only if $$\exists v_i\, :\, B'v_i=e_i,$$ where $e_i$ denotes the $i$-th canonical vector in $\mathbb{R}^N$. Letting $(B')^{\dagger}$ denote the Moore-Penrose generalized inverse of $B'$, and $I$ be a conformable identity matrix, this is equivalent to
All worker types $\alpha_i$ for which ((ref)) holds are identified. In general, however, not all $\alpha_i$'s are identified. For example, when two workers always work together, with or without other co-workers, their respective $\alpha_i$'s are not separately identified.
In practice, it is convenient to further restrict the sets of workers and teams such that the types of all workers in those teams are identified. To do this, I implement the following simple iterative algorithm.
To compute Moore-Penrose generalized inverses, I rely on sparse LU matrix decompositions. With some abuse of notation, I will still denote the resulting restricted vectors and matrices as $Y$, $D_{\lambda}$, $A$, and $\varepsilon$.\footnote{Here my goal is to point-identify the $\alpha_i$ values of a set of workers, and the corresponding contributions of these workers to team output. For this reason, in the identification strategy I rely on workers who produce on their own. In other settings (e.g., team sports) this information may not be available. In such cases, it may still be possible to identify differences between $\alpha_i$'s. As an example, consider a team of five players, where player $i$ gets replaced by player $i'$. In this case the difference in the team's output identifies $\alpha_i-\alpha_{i'}$, though not $\alpha_i$ and $\alpha_{i'}$ separately. Another example is doubles tennis, as studied in Devereux (2018). In the empirical analysis, in robustness checks I will report estimates of nonlinear models that are solely based on 2-worker teams.}
Consider now the team-size parameters $\lambda_n$, under the maintained assumption that $\mu_n=0$. Since $$D_{\lambda}^{-1}Y=A\alpha+D_{\lambda}^{-1}\varepsilon,$$ it follows that $$(I-AA^\dagger)D_{\lambda}^{-1}Y=(I-AA^\dagger)D_{\lambda}^{-1}\varepsilon.$$ Hence, ((ref)) implies
where $Z_n$ is a $J\times 1$ vector with ones where $n_j=n$ and zeros elsewhere. System ((ref)) is linear in the parameters $\lambda_n^{-1}$, so the conditions for identification are standard.\footnote{This quasi-differencing approach is related to Holtz-Eakin et al. (1988) and Chamberlain (1992).} In particular, $\lambda_1,\lambda_2,\lambda_3$ are not identified in the stylized example of Figure (ref).
Lastly, when relaxing the assumption that $\mu_n=0$ in ((ref)), one can use a similar strategy. To see this, let $\mu$ be the vector of $\mu_n$'s. We have
where instruments $Z$ can be constructed as functions of collaborations in the network. As an example, the $j$-th row of $Z$ may contain some functions of the team sizes of the co-workers in $j$, in addition to the size $n_j$ of team $j$. Moreover, under assumptions on the covariance matrix of $\varepsilon$, one can add covariance restrictions on $\lambda$ and $\mu$ to the mean restrictions in ((ref)). I will not implement these approaches empirically in Sections (ref) and (ref), and proceed under the assumption that $\mu_n=0$ for all $n$.
I will focus on the following decomposition of output variance, for every team size $n$:
In this decomposition, the component labeled “heterogeneity” reflects the variation in worker effects on output, keeping team composition constant. Equivalently, it represents the effect of worker heterogeneity if the allocation of workers to teams were random, in which case covariances would be zero. In turn, the component labeled “sorting” reflects the variance contribution due to team composition not being random.\footnote{It is common in the literature on matched employer-employee data to apply a related decomposition to the variance of log-wages, as opposed to (log-) output. This raises issues for the interpretation of variance components estimates in terms of heterogeneity and sorting (Eeckhout and Kircher, 2011). In contrast, here heterogeneity and sorting can be directly interpreted as contributions to production.}
To estimate the variance components in ((ref)), I first estimate the team-size effects $\lambda_n$ using an empirical counterpart to the just-identified system ((ref)).\footnote{The additive specification ((ref)) in logs can be estimated similarly to the one in levels, the only difference being a slightly different approach to estimate $\mu_n$.} Next, I estimate $\alpha$ using OLS as $$\widehat{\alpha}=(B'B)^{-1}B'Y,$$ for $B=D_{\lambda}A$. Hence, for any variance component $V_Q=\alpha'Q\alpha$, where $Q$ is an $N\times N$ matrix, I can construct the fixed-effects (FE) estimator
However, the FE estimator $\widehat{V}_Q$ is biased. To see this, note that
Following Andrews et al. (2008), I subtract an unbiased estimate of the bias to $\widehat{V}_Q$. To construct such an estimate, note that $$\mbox{Bias}=\mbox{Trace}( (B'B)^{-1}B'QB(B'B)^{-1}\Omega),$$ where $\Omega$ is the covariance matrix of $\varepsilon$. To obtain an empirical counterpart of $\Omega$, I assume that all the $\varepsilon_{nj}$'s are independent (see Assumption (ref)), and in addition that $\mbox{Var}(\varepsilon_n|A,\alpha)=\sigma_{n}^2$ varies only with team size. I estimate $\sigma_n^2$ as
The bias correction relies on the $\varepsilon_{nj}$'s being independent, and on their variance being constant within team size. While homoskedasticity could be relaxed by following Kline et al. (2020) and suitably restricting the sets of workers and teams, independence may be empirically strong in the applications. Under independence, Kline et al. (2020) derive consistency results and limiting distributions for bias-corrected variance components estimates in environments where the numbers of rows (teams) and columns (workers) in $A$ both tend to infinity.
To implement the bias correction in large data sets, I rely on a sparse LU decomposition to compute the denominator in ((ref)), and I approximate the trace in the bias formula using Hutchinson's approximation, as in Gaure (2014) and Kline et al. (2020). I use 1000 draws in the approximation.
I now describe a nonlinear model that allows for complementarity between worker types. To keep tractability, I model types $\alpha_i$ as discrete, with $K$ points of support. I will vary $K$ in robustness checks.
I adopt a random-effects (RE) strategy that consists in modeling the distribution of types. Specifically, I suppose that $\alpha_1,...,\alpha_N$ are drawn from the joint distribution $\prod_{i=1}^N \pi(\alpha_i)$, where $\pi(\alpha_i)$ are type probabilities. I will first describe an independent RE method where $\{\alpha_i\, :\, i\}$ are independent of collaborations $\{(i_1(j),...,i_n(j))\, :\, (n,j)\}$. I will interpret $\pi$ as a prior on individual types, as in Arellano and Bonhomme (2009). While, by construction, the prior imposes that team formation is independent of types -- and thus rules out sorting -- the posterior estimates that I will report may still feature sorting (and empirically they will).\footnote{In a similar spirit, in matched employer-employee data, Woodcock (2008) points out that posterior RE estimates can still indicate the presence of sorting even if the prior on worker and firm types is independent of mobility patterns.}
However, a natural concern is that the prior $\pi$ may be empirically informative. For example, a prior that is independent of collaborations will tend to shrink sorting patterns towards zero. Unlike the FE methods of Section (ref), RE methods rely implicitly or explicitly on a model of team formation, and can be sensitive to that model being misspecified. For this reason, I will also present estimates where I will model the dependence between the types $\alpha_i$ and the collaborations $(i_1(j),...,i_n(j))$.
With discrete types, a different approach would be to treat type membership indicators as discrete parameters, and use kmeans clustering methods for estimation (Bonhomme and Manresa, 2015, Bonhomme et al., 2021). Bonhomme et al. (2019) use this grouped fixed-effects approach to incorporate discrete firm heterogeneity in a nonlinear model of wages. I will use a related approach to provide preliminary evidence of sorting and complementarity. However, accurately estimating worker types using grouped fixed-effects requires a relatively large number of observations per worker. In the data sets that I analyze in this chapter, many workers produce only a handful of times. For this reason, I instead rely on a RE approach.
Nonparametric identification of finite mixture models under independence assumptions akin to Assumption (ref) has been extensively studied. When all teams consist of a single worker, the model has independent measures and finite types, and its identification has been studied in Hall and Zhou (2003), Hu (2008), and Allman et al. (2009), among others. When teams have multiple workers, the model has a network structure that is related to the settings studied in Allman et al. (2011).
Here I outline an identification argument in the case where $\{\alpha_i\, :\, i\}$ are independent of collaborations $\{(i_1(j),...,i_n(j))\, :\, (n,j)\}$, output is i.i.d. across teams, and team size is $n\in\{1,2\}$. Formally, I treat worker types $\alpha_i$ as random, and collaborations $(i_1(j),...,i_n(j))$ as non-stochastic, and assume that the type distribution $\pi$ is common across workers, independent of the collaborations. The goal is to identify the type distribution $\pi$, as well as the type-specific conditional distributions of output.
Focusing on workers who produce at least three times on their own gives, for $\{j_1,j_2,j_3\}$ triplets such that $i_1(j_1)=i_1(j_2)=i_1(j_3)$,
where I have used Assumptions (ref) and (ref), and that $\pi$ does not depend on $\{(i_1(j),...,i_n(j))\, :\, (n,j)\}$. System ((ref)) identifies $\pi(\alpha)$ and $\Pr\left[Y_{1j}\leq y\,|\, \alpha_{i_1(j)}=\alpha\right]$ up to labeling of the latent types, under suitable rank conditions (Allman et al., 2009).
Next, focusing now on pairs of workers who produce once on their own and once together, we similarly have, for $\{j_1,j_2,j_3\}$ triplets such that $\{i_1(j_1),i_1(j_2)\}=\{i_1(j_3),i_2(j_3)\}$,
Given that $\pi(\alpha)$ and $\Pr\left[Y_{1j}\leq y\,|\, \alpha_{i_1(j)}=\alpha\right]$ are identified up to labeling, ((ref)) identifies $\Pr\left[Y_{2j}\leq y\,|\, \alpha_{i_1(j)}=\alpha,\alpha_{i_2(j)}=\alpha'\right]$ up to the same labeling, under a suitable rank condition.\footnote{Specifically, it suffices that there exist a set $\{y_r\}$ of values in the support of $Y_{1j}$ such that the matrix with $\left((y_r,y_{r'}),(\alpha,\alpha')\right)$-element $\Pr\left[Y_{1j}\leq y_r\,|\, \alpha_{i_1(j)}=\alpha\right]\Pr\left[Y_{1j}\leq y_{r'}\,|\, \alpha_{i_1(j)}=\alpha'\right]$ has full column rank.} The argument requires no monotonicity assumptions about how output depends on worker types. In addition, identification can be shown using a related argument when $\pi$ depends on collaborations and worker types are not independent of each other within teams, provided none of the joint probabilities $\pi(\alpha,\alpha')$, for pairs of workers who produce once on their own and once together, are positive.
While the previous argument relies on workers producing on their own, identification can also be shown in settings without 1-worker teams; see Allman et al. (2011) for identification results in certain random graph models.\footnote{As an example, for two workers who collaborate together at least three times, we have, for $\{j_1,j_2,j_3\}$ triplets such that $\{i_1(j_1),i_2(j_1)\}=\{i_1(j_2),i_2(j_2)\}=\{i_1(j_3),i_2(j_3)\}$,
which identifies $\pi(\alpha)$ and $\Pr\left[Y_{2j}\leq y\,|\, \alpha_{i_1(j)}=\alpha,\alpha_{i_2(j)}=\alpha'\right]$ up to labeling of the components, under suitable rank conditions.} As robustness checks, I will report empirical results based on 2-worker teams only.
Estimating the random-effects model of team production is challenging. To see this, consider a setup where the density $f^n_{\alpha_{i_1(j)},...,\alpha_{i_n(j)}}(y)$ of $Y_{nj}$ conditional on $\alpha_{i_1(j)},...,\alpha_{i_n(j)}$ is parametric, indexed by a finite-dimensional vector $\theta$. The RE likelihood $${\cal{L}}(\theta,\pi)=\sum_{\alpha_1}...\sum_{\alpha_N}\prod_i \pi(\alpha_i)\prod_n \prod_j f^n_{\alpha_{i_1(j)},...,\alpha_{i_n(j)}}\left(Y_{nj};\theta\right)$$ involves an intractable $N$-dimensional sum over all possible worker type realizations. The likelihood does not factor in simple ways, except in the special case where all teams consist of one worker.\footnote{In 1-worker teams, the identification strategy leads to practical nonparametric estimators; see Bonhomme et al. (2016) for example. This could be used to nonparametrically estimate models of team production, using ((ref)) and its extensions for $n>2$.}
To reduce computational complexity, I follow a mean-field variational approach that is increasingly popular in related settings in machine learning and statistics. The idea is to introduce an auxiliary distribution $\prod_{i=1}^N q_i(\alpha_i)$, and set it to be as close as possible to the posterior density of $\alpha_1,...,\alpha_N$. Unlike the posterior density, its variational approximation factors across $i$. This makes estimation feasible even in large data sets.
Formally, the variational objective function is
where $p(\alpha_1,...,\alpha_N;\theta,\pi)$ is the (computationally intractable) posterior density of worker types given $\theta$ and $\pi$ values. The second term on the right-hand side of ((ref)) is the Kullback-Leibler (KL) divergence between the posterior density and its variational approximation. Hence, the variational objective and the RE likelihood are closer when the variational approximation is more accurate. The variational objective is a lower bound on the RE likelihood, and it is often referred to as the “evidence lower bound” (ELBO).
To see why adding the KL term can improve tractability, note that we equivalently have
Since this expression no longer involves an $N$-dimensional sum, it can be evaluated easily, at least when team size $n$ is relatively low. To estimate the nonlinear models, I will maximize the evidence lower bound using the EM algorithm; see Bishop (2006) and Mariadassou et al. (2010). Specifically, the algorithm alternates between updates of the $q_i(\alpha)$'s given $(\theta,\pi)$ via Newton steps, and updates of $(\theta,\pi)$ given the $q_i(\alpha)$'s via weighted maximum likelihood estimation.
Computation of the evidence lower bound in ((ref)) becomes more complex as team size increases. In the applications, I will restrict the analysis to 1-worker and 2-worker teams. It will be important to devise computational strategies to efficiently handle larger teams.\footnote{Moreover, to estimate models with larger teams, one will need to impose additional structure on how production depends on types. A possibility, in the spirit of Ahmadpoor and Jones (2019), would be to model type-specific mean output using a CES function.}
\paragraph{Statistical properties of variational estimates.}
An obvious issue with replacing the RE likelihood by the evidence lower bound is that this leads to optimizing a different objective function. Indeed, in the present setting as well as in most network applications, the posterior density $p$ does not factor across workers, whereas the variational density $q$ does. This discrepancy can make the variational estimates inconsistent even when the assumptions of the RE model hold.
The analysis in Bickel et al. (2013) suggests that variational estimates can be theoretically justified in network models with discrete types (see also Celisse et al., 2012). Bickel et al. (2013) focus on a standard stochastic blockmodel (e.g., Snijders and Nowicki, 1997) with binary outcomes, where choice probabilities are functions of both agents' types. In their asymptotic environment, expected degree in the network grows at a rate faster than $\ln N$. They first show that the RE maximum likelihood estimator (MLE) is asymptotically equivalent to an infeasible MLE where the agents' types are observed.\footnote{Related perfect classification results have been derived in panel data models under discrete heterogeneity (e.g., Hahn and Moon, 2010, Bonhomme and Manresa, 2015).} As a result, the RE estimator is consistent and asymptotically normal. They then show that the mean-field variational estimator (Daudin et al., 2008) is asymptotically equivalent to both the infeasible MLE and the RE MLE, hence that it is consistent with the same asymptotically normal distribution.
Stochastic blockmodels with binary outcomes are closely related to the model I focus on. However, there are also important differences, since in my model workers both form teams and produce, and output is continuous. In addition, the conditions in Bickel et al. (2013) rely on the network becoming sufficiently dense in large samples. It is unclear whether conditions akin to the ones they assume provide a good approximation to the settings I study. For this reason, in the appendix I probe the accuracy of the variational estimates using Monte Carlo simulations. In simulated samples that mimic the data on economic researchers, I find that variational estimates recover the true parameter values well when the numbers of productions and collaborations are sufficiently large.
The assumption that the distribution of worker types $\alpha_i$ does not depend on collaborations $\{(i_1(j),...,i_n(j))\, :\, (n,j)\}$ may be restrictive. To give intuition, consider a simple assignment model of team formation, along the lines of the roommate matching problem studied by Chiappori et al. (2019). In the model, workers are assigned to teams of size $n\in\{1,2\}$ in a way that maximizes expected output. Let $\mu_1(\alpha)=\mathbb{E}(Y_{1j}\,|\, \alpha_{i_1(j)}=\alpha)$ and $\mu_2(\alpha,\alpha')=\mathbb{E}(Y_{2j}\,|\, \alpha_{i_1(j)}=\alpha,\alpha_{i_2(j)}=\alpha')$. Let $T_{1\alpha},T_{2\alpha}\in\mathbb{N}$ be exogenous “budgets” for type-$\alpha$ workers, indicating the maximum numbers of teams of size $1$ and $2$ in which they can participate. Consider an allocation solving
where $\tau_{\alpha}$ denotes the number of teams with one worker of type $\alpha$, and $\tau_{\alpha\alpha'}$ denotes the number of teams with one worker of type $\alpha$ and another one of type $\alpha'$.
The optimization in ((ref)) is related to the planner's objective in the Becker (1973) two-sided marriage model.\footnote{In the Becker model, the optimal allocation can be represented as a stable equilibrium in an economy with transfers. In one-sided problems, existence of a decentralized solution is not guaranteed in general (e.g., Talman and Yang, 2011). Chiappori et al. (2019) show existence when types are discrete and budget sizes are even. Unlike in their setup, here there are two budgets, for 1-worker and 2-worker teams, respectively. Determining optimal allocations of workers across teams of different sizes would require measuring the costs of teamwork -- e.g., effort costs -- which I do not consider here.} As in the Becker model, in ((ref)) the assignment of individuals to teams will be driven by complementarity patterns in the expected payoff function $\mu_2(\alpha,\alpha')$. Intuitively, the higher the degree of complementarity between workers, the higher the assortativeness of the optimal allocation. I will illustrate the link between complementarity and assortativeness by computing optimal allocations of economists and inventors in Sections (ref) and (ref). In particular, this simple team assignment model implies that types will generally not be uniformly distributed across teams.
To relax the independence assumption between types and collaborations, I will rely on two approaches. In the first approach, I will maintain prior independence between the $\alpha_i$'s, however I will let $\pi_i(\alpha_i)$ depend on worker $i$'s characteristics. This correlated random-effects approach (Chamberlain, 1984) is widely used in traditional panel data applications. However, in a network of teams, it is not a priori obvious which characteristics $\alpha_i$ should depend on. I will experiment with features of the degree distributions in the collaboration hypergraph. In addition, note that this approach imposes that worker types be conditionally independent of one another given collaborations, which is unlikely when the assignment of workers to teams is given by a team formation model such as ((ref)).
In the second approach, I will jointly model output production and team formation. This adds another estimation step, since one needs to estimate the team formation process as well. Identification can be analyzed using related techniques,\footnote{In this case there are two kinds of dependent variables: the output variables and the team membership indicators. See Bonhomme (2017) for an example.} and the mean-field approximation can be implemented similarly. A specific feature of the joint model of team formation and production is that worker types influence both the quality and the quantity of output. For example, in the case of researchers in economics, the type of a worker will influence both how many articles she publishes (i.e., quantity) and in which journals they appear (i.e., quality). To implement this approach, I will specify a simple stochastic blockmodel of team formation. Though of great interest, using an economic team formation model instead would raise several challenges, and I do not consider this possibility in this chapter.
In this section and the next I present two applications. I start by applying the method to study economists and their academic production.
I use data from Ductor et al. (2014). These data, initially constructed by Goyal et al. (2006), are drawn from the EconLit database, a bibliography of journals compiled by the Journal of Economic Literature. I restrict the sample to articles published between 1995 and 1999, with at most three co-authors.
I only include authors who produced at least five articles during the period. While a natural selection when aiming to infer individual contributions, restricting the number of publications per author in this way is restrictive. Indeed, starting from 62615 authors and 89834 articles in the original 1995-1999 data, the selection restricts the sample to 6509 authors and 41150 articles. To assess how representative the results on this sample are, I will run checks using larger samples.
As measure of academic output, I follow Ductor et al. (2014) and use the journal quality variable from Kodrzycki and Yu (2006), as extended by Ductor and co-authors. From the output measure I net out multiplicative year fixed-effects, using 1999 as reference year. I show descriptive statistics in panel (a) of Table (ref) for the main sample. Research output, as measured by the journal quality variable, is right-skewed. The number of collaborations per author is also skewed, with 50% of authors producing at most 8 articles, and 1% producing at least 39. The subsample in panel (b), which I will use to estimate nonlinear models of 1-author and 2-author teams, shows similar patterns.
I start by estimating the additive model of academic production ((ref)). I first restrict the set of authors and articles in order to ensure identification. This gives 6479 authors and 41049 articles. Hence, relative to the sample in panel (a) of Table (ref), only 30 authors and 101 articles drop out.
Using this sample, I first estimate the team-size effects $\lambda_n$. I find $\lambda_2=0.67$ and $\lambda_3=0.48$. This suggests that, keeping author type(s) constant, collaborations between two co-authors increase output by $2\times 0.67-1=34\%$, and collaborations between three co-authors increase output by $3\times 0.48-1=44\%$. Next, I estimate the variance components in the decomposition ((ref)). I report uncorrected estimates, which are obtained using fixed-effects, and bias-corrected estimates, which I obtain using the method described in Section (ref).
In Table (ref) I show the results of the variance decomposition ((ref)), for the three team sizes: $n=1,2,3$. The results suggest that author heterogeneity matters substantially, while playing a smaller role in larger teams. Indeed, the variance share explained by author heterogeneity is 33% for sole-authored articles, 19% in articles with two co-authors, and 14% in articles with three co-authors. Sorting also contributes a substantial share: 17% in articles with two co-authors, and 20% in articles with three co-authors. Hence, according to this analysis, both heterogeneity and sorting contribute to the variation in academic output.
Another finding in Table (ref) is the magnitude of the component due to other factors, $\sigma_n^2$. This component explains 67% of the output variance in articles with one author, 57% in articles with two co-authors, and 62% in articles with three co-authors. The large residual variance contributes to the fixed-effects estimates of variance components being substantially biased. For example, the fixed-effects estimate of the variance contribution of heterogeneity is biased upward by 33% with one author, 69% with two co-authors, and 91% with three co-authors. This suggests that, as has been documented in matched employer-employee settings (e.g., Bonhomme et al., 2020), bias correction is needed in order to obtain reliable estimates of variance components in collaboration networks.
\paragraph{Robustness analysis.}
The adequacy of the additive model ((ref)) depends on how one measures the output variable. To explore the sensitivity to this choice, I use two alternative measures of output. First, I estimate the model in logs, accounting for additive team-size effects as in ((ref)). The results are shown in Appendix Table (ref), panel (a). Second, I use year-specific ranks of journal quality as dependent variables, instead of journal quality itself. The results are shown in Appendix Table (ref), panel (b). While it is reassuring that the variance shares in the two specifications are broadly comparable to the baseline, this exercise does not fully address the issue of the sensitivity to the choice of output measure. The nonlinear model that I estimate in the next section will be less sensitive to this issue.\footnote{In this application, it would be interesting to link the journal quality variable to economic payoffs and career outcomes.}
The estimation sample represents a relatively small share of authors and articles during the period. To probe the robustness of the findings to sample definition, I consider two less restrictive rules for inclusion in the sample, requiring every author to produce at least one or two articles, respectively, as opposed to at least five in the baseline sample. The results are shown in Appendix Table (ref), panels (c) and (d). When requiring at least two articles per author, the sample contains $24555$ authors and $70550$ articles, and $23498$ authors and $69331$ articles in the sample where all author types are identified. Panel (c) of the table shows that bias-corrected variance shares are broadly similar to the baseline, and that biases are larger, so bias correction matters more. When removing all restrictions and including all authors with at least one article, the sample where all author types are identified contains $44852$ authors and $79659$ articles. Panel (d) of Appendix Table (ref) shows larger differences compared to the baseline subsample, and a substantially higher amount of bias.
Before presenting the results of the nonlinear model, I provide some preliminary evidence of complementarity, as well as sorting and heterogeneity, in the data. To do so, I select authors who produce at least five articles on their own. This gives a set of 3263 authors. Then, for every author I compute the average output (i.e., average journal quality) among their sole-authored articles, and divide this measure into four quartiles. I call the resulting quartile the “type proxy” of the author. The reason for focusing on individuals with at least five sole-authored articles is to reduce noise in the proxy. At the same time, this does not fully remove the noise and requires a selected sample.
Given this proxy of type, I focus on 2-author teams and report two sets of results, which I interpret as reflecting sorting and complementarity, respectively. In panel (a) of Figure (ref), I show the proportions of every type proxy pairs in the sample. The percentages suggest that higher-type authors, who produce higher-quality articles on their own, tend to work together, implying the presence of sorting between authors of similar productivity. In panel (b) of Figure (ref), I compute the average output for every type proxy pairs. Authors who produce higher-quality articles on their own tend to produce higher-quality output as well when they work in teams, which is consistent with the presence of author heterogeneity. In addition, joint output is highest when authors with the highest type proxies (i.e., in the top quartile of sole-authored production) work together, and the figure is suggestive of the presence of complementarity. Since the additive model ((ref)) rules those out, this motivates estimating a nonlinear model of academic production.
\paragraph{Estimates of the nonlinear model, baseline specification.}
I now report estimates of a finite mixture model with $K$ types, varying $K$ between $2$ and $6$. In this analysis I only consider 1-author and 2-author productions; see panel (b) of Table (ref) for summary statistics on the sample. I model the output distribution as a log-normal, with mean and variance that depend on the types of authors in the team: for 1-author teams mean and variance are functions of the type, and for 2-author teams mean and variance are symmetric functions of the two types. The log-normal specification is restrictive, and in future work it will be important to check how robust the results are to removing this functional form assumption. I use mean-field variational EM for estimation.\footnote{I declare convergence when the increment in evidence lower bound is less than $10^{-3}$. To alleviate local maxima issues, I started the variational EM algorithm from different sets of parameters. The values that I obtained upon reaching the tolerance threshold differed somewhat, however they had similar implications for heterogeneity, sorting and complementarity. }
For conciseness, I only comment in detail on the results for $K=4$. In this case, the type proportions are $15\%$, $20\%$, $36\%$, and $29\%$. The means of the log-normal output distribution are, in 1-author teams: $$\left(
\right),$$ and in 2-author teams: $$\left(
\right).$$ This suggests that authors have heterogeneous productivity levels, and these differences affect both 1-author and 2-author productions.
To visualize the implications of the estimated model for heterogeneity, sorting and complementarity, in Figure (ref) I report similar quantities as in Figure (ref), except that I now use the types estimated under the model instead of the type proxies. Panel (a) of Figure (ref) shows that high-type authors have a stronger propensity to work together, although this evidence of sorting is less pronounced than in Figure (ref).
Panel (b) of Figure (ref) suggests the presence of complementarity, since the return of two high types working together is 1.5 standard deviations above the return of two low or middle types working together. This suggests gains from complementarity in economic research, which are reflected in stronger sorting at the top. In addition, the results show that sole-authored productivity and productivity in teams are not one-to-one: while the lowest type produces slightly lower quality output than the second-lower type on her own, she benefits substantially more from working with others, in particular with a high-type co-author.
In addition to documenting the patterns of heterogeneity, sorting and complementarity, the nonlinear model can be used to refine the variance decomposition in ((ref)). Indeed, the nonlinear model accounts for interaction effects between co-authors, over and beyond the additive effects of author types. This gives a fourth variance component, which I will refer to as “nonlinearities”, to report in the decomposition, in addition to the contributions of heterogeneity, sorting, and other factors. This component can be computed as the difference between the variance of the mean of $Y_{nj}$ given the worker types -- which depends on both additive and interactive terms in general -- and the variance of its best linear approximation.
I report the results of the variance decomposition in Table (ref), for $K=2,4,5,6$. While $K=2$ seems to allow for too little heterogeneity, the results for $K=4,5,6$ are broadly similar to one another. They show that author heterogeneity explains approximately 30% of output variance, and sorting explains between 10% and 12%. Despite different type modeling and form of the production function, these variance shares are not too different from those based on the additive model (see Table (ref)). In addition, nonlinearities explain between 1.5% and 2.5% of the variance. While this share may seem small, the presence of complementarity at the top of the productivity distribution, as shown by panel (b) of Figure (ref), is likely to be important to study policies that re-allocate researchers across teams.
To assess how complementarities might affect worker allocations, I compute an allocation that maximizes total output by solving ((ref)). I compute expected payoffs and “budgets” based on the estimates of the nonlinear model with $K=4$. In this case, when budget sizes are even, an optimal integer-valued allocation can be computed by solving a linear program that does not impose the integral constraints. The resulting allocation in Figure (ref) shows sorting at the top of the type distribution, but not at the bottom.
\paragraph{Other specifications.}
I now present the results of three exercises based on other specifications of the nonlinear model. In the baseline specification, the distribution of author types does not depend on collaborations. To assess the impact of this assumption, I estimate two alternative specifications. In the first specification, I model the type distribution as a function of author characteristics, using a multinomial logit specification. As characteristics, I use the numbers of teams in which author $i$ participates, separately for 1-author and 2-author teams, as well as the indicators that these numbers are zero. In panels (1a) and (1b) of Figure (ref), I show type-pair proportions and mean output by pairs of types, and in panel (1c) I show the optimal allocation obtained by solving ((ref)). In Appendix Table (ref) I show the variance decomposition estimates. The results are similar to the ones from the baseline specification.
In the second specification, I augment the team production model with a model of team formation, and I jointly estimate the parameters using a variational EM algorithm. I specify team formation using a stochastic blockmodel for hypergraphs, where 1-author and 2-author collaborations follow independent Poisson distributions, with a parameter that is type-specific in the first case and pair-of-types-specific in the second case. In this model, as in most models of team formation, the conditional distribution of $\{\alpha_i\, :\, i\}$ given $\{(i_1(j),...,i_n(j))\, :\, (n,j)\}$ does not factor across $i$. In addition, in such a model, author types do not only influence heterogeneity in research quality, but also in quantity. The results show some differences with the baseline. In particular, panel (2a) of Figure (ref) shows stronger evidence of sorting. In panel (2b), I show mean output in 2-author teams, by type pairs. The estimates, and the implied optimal allocation in panel (2c), are again consistent with complementarity between high-type authors.\footnote{However, note that the high degree of sorting prevents one from assessing with confidence the gains from collaboration between highest and lowest types (i.e., between type-1 and type-4 authors): while the point-estimates in panel (2b) are large, panel (2a) shows that such collaborations are virtually non-existent. This is likely to affect the optimal allocation numbers shown in panel (2c) of Figure (ref).} In Appendix Table (ref) I report variance decomposition estimates. In 2-author teams, heterogeneity accounts for a lower share of variance compared to the baseline (22% versus 29%), sorting accounts for a larger share (19% versus 10%), and the variance share accounted for by nonlinearities is smaller than in the baseline.
In a last exercise, I report estimates based on 2-author teams only. Hence, sole-authored productions are discarded. Panel (3a) of Figure (ref) shows a higher degree of sorting in the middle of the type distribution, rather than at the top. Panel (3b) again shows evidence of complementarity, although the patterns differ somewhat from the baseline, as shown by the implied optimal allocation in panel (3c). Lastly, the variance decomposition results in Appendix Table (ref) show that, in this case, heterogeneity accounts for 32% of variance, and sorting accounts for 5% (as opposed to 29% and 10% in the baseline).
In this section I apply the method to study patents and inventors.
I use data from Akcigit et al. (2016). Their main source is the disambiguated inventor data of Li et al. (2014), which identifies unique inventors in the USPTO data. I restrict the analysis to US patents, which were granted between 1995 and 1999. Throughout, I focus on patents in the technology class “Computers and Communications”.\footnote{Among all inventors who patent in this class, 80% patent only in this class during the period.} For the output measure, I follow Akcigit et al. (2016) and use Hall et al.'s (2001) measure of patent quality, which is a truncation-adjusted measure of forward citations of the patent. From this measure I net out multiplicative year fixed-effects, using 1999 as the reference year. In this illustration, I do not use information about the firms where inventors work, although it would be interesting to incorporate this information in future research.
In panel (a) of Table (ref), I provide summary statistics about the sample, where I restrict inventors to participate in at least five patents during the period. In this setting also, output and the number of collaborations are right-skewed. The original sample of US patents in the class “Computers and Communications” that were granted between 1995 and 1999 contains 65848 inventors and 62927 patents. Hence, imposing that all inventors be on at least 5 patents restricts the sample size substantially. In the sample, team size exhibits a range of variation but teams tend to be small: 57% of teams are 1-inventor teams, 26% of teams have 2 members, and less than 1% of teams have more than 6 members. In panel (b) of Table (ref), I show summary statistics for the subsample of 1-inventor and 2-inventor teams that I will use to estimate the nonlinear models.
I first estimate the additive model ((ref)). Restricting the set of inventors and patents to ensure identification gives 5547 inventors and 29101 patents. I estimate the model by allowing for four team sizes $n$ in $\lambda_n$, where $n=4$ corresponds to all teams with at least four inventors. I find $\lambda_2=0.54$, $\lambda_3=0.39$, and $\lambda_{4}=0.29$. Hence, keeping inventor type(s) constant, patents with two inventors are cited 8% more than patents with one inventor, 3-inventor patents are cited 17% more, and patents with at least 4 inventors are cited 16% more.
In Table (ref), I show the estimates of the variance components in ((ref)), for $n=1,2,3$. The variance share explained by inventor heterogeneity is 38% in patents with one inventor, 32% in 2-inventor patents, and 19% in 3-inventor patents. Sorting contributes 6% of variance in 2-inventor patents, and 12% in 3-inventor patents. Here also, the other factors $\varepsilon_{nj}$ account for the main share of the output variance. Overall, the variance decomposition estimates in the patent sample are thus not very different from those in the sample of economists, with a somewhat larger contribution of worker heterogeneity and a smaller contribution of sorting.
In Appendix Table (ref), I augment the sample and require every inventor to produce at least one or two patents, as opposed to at least five in the baseline sample. In addition, as in all the other robustness checks in this application, using a Poisson regression I net out from the output the effects of year and “inventor age”, as measured by the difference between the year of observation and the first year where the inventor produced a patent. Compared to the baseline estimates, the results in Appendix Table (ref) show that, in these larger samples, inventor fixed-effects are estimated with more noise, and uncorrected variance components are very large in magnitude. The bias-corrected estimates indicate a larger role of inventor heterogeneity than in the baseline, and a negative sorting contribution, although, given the amount of noise, these findings should be interpreted with caution.
To specify the nonlinear model,\footnote{In Appendix Figure (ref), I construct type proxies as in Figure (ref). I group inventors according to the quartiles of the citations of their sole-authored patents, and restrict the sample to inventors with at least 5 patents on their own, so the sample is small, with 1554 inventors. The figure suggests the presence of heterogeneity and sorting, although complementarity is less salient than for economists, and sorting seems more evenly spread out along the diagonal (compare with Figure (ref)).} I model the output distribution as a negative binomial with mean and variance that depend on the types of inventors in the team. This parametric form allows for a convenient treatment of the zeros in the dependent variable. I now comment in detail on the results for $K=4$. In this case, the type proportions are $6\%$, $25\%$, $52\%$, and $17\%$. The means of the negative binomial output distribution are, in 1-inventor teams: $$\left(
\right),$$ and in 2-inventor teams: $$\left(
\right).$$
In Figure (ref), I report similar quantities as in Figure (ref) to illustrate sorting, heterogeneity, and complementarity in the sample of inventors. Panel (a) shows some evidence of sorting towards the top, however the pattern is less concentrated than in the sample of economists (compare with panel (a) of Figure (ref)). Panel (b) of Figure (ref) suggests the presence of complementarity, since the return to two high-type inventors working together is 2 standard deviations above the return to two low or middle types working together. This suggests gains from complementarity in patent production, with somewhat less sorting than in the case of economists.
In Table (ref), I show variance decomposition results based on the nonlinear model, for $K=2,4,5,6$. Focusing on the results for $K=4,5,6$, which are broadly comparable, inventor heterogeneity explains approximately 30% of output variance, and sorting explains between 6% and 9%. Nonlinearities explain a larger share than in the economists' sample, between 5% and 8% of the variance.
In addition, in Figure (ref) I report the allocation that maximizes total output according to ((ref)). As in the case of economists, the allocation has perfect assortative matching for the top type but not for the bottom types. This optimal allocation differs quite substantially from the estimated allocation shown in panel (a) of Figure (ref).
\paragraph{Other specifications.}
In Figure (ref), I show estimates based on three alternative specifications: a correlated random-effects estimator, a joint random-effects estimator where 1-inventor and 2-inventor collaborations follow independent Poisson distributions with type-specific parameters, and an estimator that is solely based on 2-inventor collaborations. In these specifications I net out multiplicative inventor-age effects from the output, in addition to year effects. I show the variance decomposition results in Appendix Table (ref). As in the case of economists, I find that the correlated RE estimates do not substantially differ from the baseline. The optimal allocation in panel (1c) of Figure (ref) does differ somewhat from the baseline (compare with Figure (ref)), yet both allocations exhibit perfect assortative matching at the top.
In addition, the estimates based on 2-inventor teams are also broadly similar to the ones based on both 1-inventor and 2-inventor teams in this sample. One difference is that the group shares are more unbalanced when using only 2-inventor collaborations. Another difference is that the implied optimal allocation in panel (3c) of Figure (ref) exhibits perfect assortative matching along the entire distribution, and not only at the top.
However, the joint RE estimates show important differences compared to the baseline. Note that, as in the corresponding specification in the economists' sample, here inventor types do not only reflect heterogeneity in the quality of innovation through patent citations, but also in the quantity of patents produced by an inventor. In 2-inventor teams, heterogeneity, sorting and nonlinearity account for 9%, 8%, and less than 1% of the variance of output, respectively, compared to 30%, 7%, and 5% in the baseline (see Appendix Table (ref)). While the estimates in panels (2a) and (2b) of Figure (ref) show a high degree of sorting, they exhibit less evidence of complementarity compared to the other specifications. The implied optimal allocation in panel (2c) is also quite different in this case. This suggests that the modeling of inventor types and how types affect team formation is important to accurately assess the contributions of heterogeneity, sorting and complementarity in this sample.
\paragraph{How do inventor types perform out of sample?} In a last exercise, I study the predictive performance of inventor types out of sample. To do so, I use the parameters of the baseline random-effects model with $K=4$ types, estimated on the 1995-1999 period. Using the estimated variational posterior type probabilities, I then calculate average patent output produced by inventors of particular types between 2000 and 2005, and the type proportions in those collaborations. Figure (ref) shows that sorting patterns and productivity differences between types persist out of sample, and that the higher returns to collaborations between high inventor types relative to other type combinations persist as well. However, the out-of-sample estimates show less inventor heterogeneity than in the baseline (compare with panel (b) of Figure (ref)). There may be several explanations for this: posterior type probabilities are only based on five years of observations, truncation issues with patent citations may make the outcome less informative towards the end of the sample (Akcigit et al., 2016), and over a 10-year period inventor quality may change.
In this chapter I have outlined a measurement framework to assess the contributions of individuals to team output. While an additive production specification leads to a tractable estimator, it rules out complementarity that appears to be a feature of the samples of economists and inventors that I study. In the applications, a natural next step will be to relate the latent types to characteristics of authors and inventors, such as their national origin or education background.
The discrete-type approach that I have implemented is promising, yet it needs to be studied more. In particular, it will be important to provide formal consistency arguments and to derive asymptotic distributions for variational estimators in this setting. Spectral clustering methods (e.g., Lei and Rinaldo, 2015) could be possible alternatives to the nonlinear random-effects methods that I have described. It will also be important to augment the framework to account for the productive effects of team-specific factors, both observed and latent.
The literature on matching and sorting has made substantial progress in one-to-one settings (Chiappori and Salani\'e, 2016, Chade et al., 2017). However, less is known about many-to-one and many-to-many sorting environments. Eeckhout and Kircher (2018) model sorting in large firms, while abstracting from complementarities between workers. Chade and Eeckhout (2018) propose a model of information aggregation in teams that has tight implications for sorting.\footnote{A related strand of the literature proposes models of co-authorship networks, see among others Goyal et al. (2004) and Gans and Murray (2014), and the co-author model in Jackson and Wolinsky (1996).} Moreover, it is well-known that models with general complementarities may not have equilibria (Kelso and Crawford, 1982). Addressing these challenges, and combining the quantitative framework of team production that I have introduced, with economic models of team formation and effort allocation (such as those recently proposed by Hsieh et al., 2018, and Anderson and Richards-Shubik, 2019), is an interesting avenue for future work.