Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
140,890 characters · 33 sections · 68 citation commands
Identification, Estimation, and Inference in Two-Sided Interaction Models
\relax \hypersetup{pageanchor=false} \hypersetup{pageanchor=true}
\thispagestyle{empty}
\thispagestyle{empty}
\hypersetup{pageanchor=true} \setcounter{page}{1}
In many economic settings, outcomes result from the interaction of two distinct types of agents and depend on the latent characteristics each side brings to the match. For example, wages reflect both worker skills and firm attributes abowd1999high; corporate performance depends on managerial ability together with company-specific features bertrand2003managing; and the productivity of public offices is shaped by local conditions and the capacity of the bureaucrats in charge fenizia2022managers.
Analyzing such settings requires using observable outcomes to (i) disentangle the contributions of the agents involved in the interaction, quantifying their latent characteristics, and (ii) understand how these characteristics interact to generate outcomes. Doing so provides the foundation for addressing a wide range of questions. In the labor market, for instance: are high-productivity workers more likely to match with high-productivity firms? How much of wage variation is due to worker heterogeneity versus firm heterogeneity? Which observable characteristics correlate with latent productivities? What is the role of complementarities in this setting? If complementarities matter and wages proxy for output, can counterfactual reallocation of matches raise aggregate productivity?
To study these questions, this paper introduces the Bipartite Interaction (BI) framework, which flexibly models two-sided interactions. Any model in this framework has two components: (i) a matching network, where nodes represent agents and edges capture which pairs are observed, and (ii) an interaction function, which maps the latent characteristics of matched agents into observed outcomes. The matching network summarizes the pattern of available data, while restrictions on the interaction function reflect assumptions about complementarities and generate distinct models within the framework. A well-known example is the two-way fixed effects (TWFE) model, which assumes an additively separable interaction function and therefore rules out complementarities. This model has become widely used by applied economists, particularly following its influential application by abowd1999high.
Building on this framework, I introduce the Tukey model, named after John Tukey, who proposed an analogous functional form in the context of nonlinear ANOVA tukey1949one. The Tukey model enriches the TWFE specification by allowing for complementarities in a simple and interpretable way: it relies on the assumption that the cross-partial derivative of the interaction function is determined by a single scalar parameter. This interaction parameter entirely captures the presence, strength, and direction of complementarities; when it equals zero, complementarities vanish, and the model coincides with TWFE.
Next, I study two extensions of the Tukey model that accommodate richer complementarity structures by relaxing the restriction of a constant cross-partial derivative. The first allows a firm-specific interaction parameter, permitting heterogeneity in complementarity across firms. This specification uses the functional form in bonhomme2019distributional but drops their grouping assumption, treating each agent individually without clustering. The second extension imposes no parametric restriction on the cross-partial derivative. It serves as a fully flexible nonparametric benchmark, requiring only that the interaction function be monotone in both arguments.
Taken together, the BI framework, the Tukey model, and these extensions lead to three main contributions. First, I show that identification in the BI framework reveals a trade-off between the flexibility of the interaction function and the structure of the matching network: weaker assumptions on the interaction function require richer structure in the matching network. For example, in the TWFE model, point identification requires only that the matching network be connected, meaning any pair of nodes can be joined by a finite path of edges. This condition is no longer sufficient in general. In the nonparametric specification, for instance, point identification requires that any two nodes of the same type be linked by a path of length at most two. Importantly, the trade-off is asymmetric: while the interaction function is unobserved and must be restricted by assumption, the matching network is directly observed, and restrictions on it are straightforward to check.
In empirical applications, researchers typically begin by choosing the restrictions to impose on the interaction function. The results in this paper can then be used to verify whether the observed matching network satisfies the corresponding conditions for parameter identification. For the Tukey model, point identification requires the same connectedness condition as the TWFE model, plus one additional requirement: the presence of at least one informative cycle of length four in the matching network (a closed path involving four distinct edges and nodes). This condition is often satisfied in practice, making the Tukey model a flexible and empirically applicable alternative to the TWFE specification. By contrast, the requirements for point identification in the two extensions are much more stringent and rarely met in empirical settings where two-sided interactions are typically studied.
Second, I propose a new estimator in the Tukey model for the interaction parameter corresponding to the cross-partial derivative of the interaction function, and study its properties in a large-graph asymptotic analysis that does not require any agent to be observed many times. As identification relies on the presence of a cycle in the matching network, the estimator uses multiple such cycles to average out noise and consistently estimate the parameter. Unlike alternative methods in similar settings, the estimator does not require estimating latent characteristics. It is consistent under the mild requirement that the number of cycles scales with graph size; in employer–employee matched data, I show that such cycles are typically abundant. The estimator is asymptotically normal, with its variance shaped by three elements: (i) the variance of the outcome error terms, (ii) heterogeneity in latent characteristics within cycles, and (iii) the ordering of agents in each cycle. Since the ordering depends on labels chosen by the researcher, I propose an instrument-based procedure that assigns them using observable characteristics correlated with latent productivities, ensuring consistency and asymptotic normality.
Third, I develop a formal test for the absence of complementarities in the BI framework. The test exploits the nesting of the TWFE model within the Tukey model: since the TWFE specification corresponds to the Tukey model with the interaction parameter equal to zero, testing this null provides a direct test for the absence of complementarities. The assumption of no complementarities implied by the TWFE model is often discussed in empirical work, and several informal diagnostics are in use, but a formal test has not, to my knowledge, been studied. The cycle-based estimator for the interaction parameter does not require estimating any individual latent characteristics, which allows construction of a test whose asymptotic properties can be studied by an asymptotic analysis aligned with common data patterns, including cases where nodes in the matching network have no more than two links.
To illustrate how the Tukey model can provide richer insights into two-sided interactions, I revisit the application in limodio2021bureaucrat on the interaction between public managers and tasks, focusing on the implementation of World Bank projects. The success of a project is modeled as the outcome of the interaction between the ability of the manager in charge and the characteristics of the country where they operate. Original estimates using the TWFE model found negative sorting, with high-performing managers more likely to be matched with low-performing countries. Estimates from the Tukey model add further insight: the interaction parameter is negative and statistically different from zero, indicating negative complementarities, where high-performing managers have larger value added when matched with low-performing countries. Given this interaction function, negative sorting is optimal for maximizing average project success, which may explain the allocation pattern documented in limodio2021bureaucrat as a rational response to the structure of complementarities.
The BI framework proposed in this paper is related to a wide range of earlier approaches to two-sided interactions, many of which are reviewed in bonhomme2020econometric. The benchmark in this class is the TWFE, which assumes a modular (additively separable) interaction function. While most empirical applications focus on worker-firm interactions abowd1999high, card2013workplace, kline2024firm, TWFE-style methods have also been applied to other two-sided settings, including managers and firms bertrand2003managing, teachers and students jackson2014teacher, chetty2014measuring, chetty2014measuring2, patients and healthcare providers finkelstein2016sources, and bureaucrats and geographic postings fenizia2022managers, limodio2021bureaucrat. Despite its versatility, the model relies on the strong modularity assumption, which rules out complementarities: the marginal productivity of each agent is assumed constant and independent of the other side of the match. In many contexts, this contrasts with the emphasis placed on matching, where significant resources are devoted to finding “the right match”.
To relax modularity, bonhomme2019distributional propose a model in which the interaction function allows for complementarities across the latent characteristics of agents. Their approach assumes that agents can be partitioned into a finite number of groups, with all members of a group sharing the same latent characteristics. In practice, this requires researchers to assign agents to groups, typically via a clustering procedure, and then estimate group-specific productivities and intergroup complementarities. A similar approach is adopted in lei2023estimating. These models introduce greater flexibility in the interaction function relative to TWFE and have delivered valuable empirical insights weigel2024super, mourot2025surgeons, but rely on the grouped heterogeneity assumption. By contrast, the approach developed in this paper offers a feasible way to introduce complementarities into two-sided interactions without relying on grouping.
Concerns about complementarities have also surfaced in many TWFE applications. A common strategy to justify their exclusion has been to estimate a saturated version of the model, compare $R^2$ values, and conclude, based on the typically small changes observed, that complementarities are not quantitatively important card2013workplace, song2019firming, fenizia2022managers, adhvaryu2024no. However, kline2024firm shows that these $R^2$ comparisons can be misleading and argues that complementarity patterns are better detected by focusing on cycles in the matching network. Building on this insight, this paper develops a procedure to formally test the absence of complementarities, providing a formal way to assess the assumption at the core of the TWFE model.
Worker-firm interactions are the most common application for this TWFE model abowd1999high, card2013workplace, kline2024firm, yet modeling labor markets in this “reduced form” way has faced criticism eeckhout2011identifying, hagedorn2017identifying, lopes2018firm, eeckhout2018sorting. These concerns extend to the BI framework as well. Still, the framework remains valuable: both in other domains where its structure may be more credible (as in the empirical illustration) and as a foundation for re-examining labor markets under alternative, potentially more realistic, restrictions than those implied by the TWFE model.
Related developments in balanced unit-time panel settings show the value of modeling interactions beyond modularity. The model in tukey1949one, for example, already incorporated complementarities more richly, and many recent contributions extend this idea using factor models or nonparametric interaction structures bai2009panel, freyberger2018non, freeman2023linear, sbaisassi2024dyadic,armstrong2025robust. Two differences distinguish the BI framework from classical panels. First, panel models typically focus on estimating coefficients on observed regressors, whereas the BI framework aims to recover the interaction function and the fixed effects themselves. Second, panel analysis usually relies on observing all unit-time combinations, while the BI framework is defined over incomplete matching networks in which only a subset of possible matches is observed. As a result, the structure of the matching network is central to both identification and inference in the BI framework, in contrast to panel settings where it is fixed by design and excluded from the modeling analysis.
Because of this crucial role of the matching network, this paper also contributes to the literature on network econometrics. For identification, as in bramoulle2009identification, graham2017econometric, and de2018identifying, I impose assumptions on the observed graph to ensure that parameters can be identified, using a strategy that eliminates nuisance parameters to identify the interaction coefficient, paralleling the approach of jochmans2017two for multiplicative models. For inference, I consider an asymptotic regime in which the graph grows without requiring the number of realized links to increase proportionally with the number of potential links, and allowing the number of edges per node to remain bounded. Most existing inference results for sparse networks assume degrees (numbers of edges per node) that may diverge with network size jochmans2019fixed, cai2022linear. Exceptions that accommodate bounded-degree graphs, such as verdier2020estimation and auerbach2021local, do not focus on parameters governing interactions. By explicitly allowing for bounded degree, the present analysis better reflects the structure of many real-world datasets on two-sided interactions and clarifies how the matching network affects the precision of BI estimators.
The remainder of the paper is organized as follows. Section (ref) introduces the BI framework, the Tukey model, and its extensions studied in this paper. Section (ref) derives necessary and sufficient conditions for parameter identification in the Tukey model and its extensions. Section (ref) turns to estimation in the Tukey model, mostly focusing on the estimator for the interaction parameter and its asymptotic properties. Section (ref) presents an empirical illustration, showing how the Tukey model can be used in practice and how it provides additional insights into the interaction function. Section (ref) reports Monte Carlo simulations assessing the finite-sample performance of the estimator. Section (ref) concludes.
Throughout the paper, lowercase letters denote scalar parameters or quantities specific to an individual agent, typically scalars but possibly vectors when productivities are multidimensional (e.g., $\alpha_i$ represents the latent productivity of worker $i$). Uppercase letters denote sets (e.g., $\alpha_i \in A$, where $A$ is typically a compact metric space). Bold symbols denote tuples: for example, $\boldsymbol{\alpha}$ denotes the collection $(\alpha_1, \dots, \alpha_I)$.
I present the Bipartite Interaction (BI) framework in the context of the labor market, where interactions involve workers and firms. This setting serves as a concrete example for exposition, but the framework itself is general and applies to any two-sided environment.
Let $I \in \mathbb{N}$ and $J \in \mathbb{N}$ denote the number of workers and firms, respectively. Each worker $i$ has a latent deterministic productivity $\alpha_i \in A$, and each firm $j$ has a latent deterministic productivity $\psi_j \in \Psi$, where $A$ and $\Psi$ are compact subsets of $\mathbb{R}^d$. Let $\boldsymbol{\alpha} \coloneqq (\alpha_1,\dots,\alpha_I)$ and $\boldsymbol{\psi} \coloneqq (\psi_1,\dots,\psi_J)$ denote the collections of these productivities.
The potential outcome of the interaction between worker $i$ and firm $j$, for example the wage worker $i$ would receive if employed by firm $j$, is
where $f \colon A \times \Psi \to \mathbb{R}$ is the interaction function and $\eta_{ij}$ is a mean-zero random term. Define $\theta_{ij} \coloneqq f(\alpha_i, \psi_j)$ as the deterministic component of the outcome, equal to $\mathbb{E}[y_{ij}]$, determined by the interaction function and the productivities of the matched agents. To focus on $f$, I omit covariates. Incorporating covariates in the BI framework is a valuable extension not considered in this paper.
To study the properties of $f$, it is useful to consider the role of its cross-partial derivative, which leads to the notions of modularity and complementarities. I formally define modularity in Appendix (ref). Intuitively, when $\alpha_i$ and $\psi_j$ are scalars and $f$ is differentiable, the interaction function is modular when the cross-partial derivative of $f$ is zero, supermodular when it is nonnegative, and submodular when it is nonpositive. This is closely related to complementarities. I say that an interaction exhibits complementarities when the cross-partial derivative is nonzero, which I call positive or negative when this derivative is positive or negative, respectively.
The potential outcome $y_{ij}$ is defined for all worker-firm pairs, representing the outcome that would be realized if match $(i,j)$ occurred, whether or not it is observed in the data. In practice, only a subset of matches is realized, and $y_{ij}$ is observed only for some pairs. Define
as the indicator of observed matches, treated as fixed in the analysis, and collect them in $\mathcal{O}_{IJ} = \{ (i,j) : D_{ij} = 1 \}$.
Let $G_{IJ} = ([I], [J], \mathcal{O}_{IJ})$ be the bipartite network linking worker $i$ to firm $j$ whenever $D_{ij} = 1$. $G_{IJ}$ is bipartite because its nodes are partitioned into two disjoint sets, workers and firms, with no edges within the same set.
Figure (ref) shows an example with 5 workers (purple), 3 firms (green), and 6 edges, each representing an observed outcome.
The fact that worker $i_2$, for example, is linked to two firms ($j_1$ and $j_2$) does not imply that the matches occur simultaneously. The matching network is constructed over a chosen time window, which may span several years, and each edge can correspond to a different period: $i_2$ may be employed by $j_1$ in one year and by $j_2$ in another. The BI framework is static, abstracting from the timing of moves, so it does not matter whether the match with $j_1$ occurs before or after the one with $j_2$.
The matching network is deterministic: the BI framework does not model its formation, and the structure of $G_{IJ}$ is taken as given. The only source of randomness in the model are the terms $\{\eta_{ij}\}_{(i,j)\in\mathcal{O}_{IJ}}$, one for each observed outcome. These terms are mutually independent, mean-zero, and may have pair-specific distributions, allowing for heteroskedasticity. Since $G_{IJ}$ is non-random, whether the match $(i,j)$ is observed does not depend on $\eta_{ij}$, making the random term exogenous.
For any realized match ($D_{ij} = 1$), the model considers only a single outcome $y_{ij}$. If repeated observations of the same worker-firm pair are available, they can be averaged to yield a single $y_{ij}$, leaving the analysis unchanged.
\paragraph{Parameters of interest.}
The primitive parameters of the model are the interaction function $f$ and the productivity vectors $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$. Recovering $f$ allows researchers to characterize how worker and firm productivities map into the outcome. A central question is whether an agent's marginal productivity depends on the characteristics of the other agent in the match, and, when this occurs, to describe the nature and pattern of such complementarities. The productivity vectors $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$, once linked to additional information on workers and firms, shed light on the determinants of productivity differences. For example, regressing $\boldsymbol{\alpha}$ on worker demographics reveals which observable traits drive variation in worker productivity.
Together, $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$ permit the study of assortative matching, the tendency of workers and firms with similar relative productivities to pair together. A common measure of sorting is the correlation between $\alpha_i$ and $\psi_j$ across observed matches: one constructs a vector containing $\alpha_i$ for each match $(i,j)$, a second vector with the corresponding $\psi_j$, and then computes their correlation. A positive value indicates positive sorting, whereas a value near zero or negative suggests weak or reverse sorting.
Joint knowledge of $f$, $\boldsymbol{\alpha}$, and $\boldsymbol{\psi}$ also enables decomposing outcome variance, quantifying the shares attributable to worker heterogeneity, firm heterogeneity, and their interaction. These quantities further allow measuring factor misallocation by comparing observed matches to the efficient allocation, and support counterfactual analyses in which workers are reassigned across firms.
These examples illustrate the wide range of parameters of interest that can be derived from $f$, $\boldsymbol{\alpha}$, and $\boldsymbol{\psi}$ as primitive inputs. A comprehensive treatment of all such parameters is beyond the scope of this paper. The focus here will be on primitive $f$, $\boldsymbol{\alpha}$, and $\boldsymbol{\psi}$ under different assumptions on $f$ and $G_{IJ}$, occasionally using selected derived quantities to highlight key features of the results.
\paragraph{Functional specifications.} Imposing restrictions on the interaction function within the BI framework yields distinct models. I focus on three specifications: the Tukey model, the BLM model, and the seriation model. The next sections describe each of these models and discuss their economic interpretation. Appendix (ref) formalizes their connection to shape restrictions on $f$, showing how their functional forms emerge from broader assumptions.
The cross-partial derivative of $f$ captures the complementarity structure of the interaction function in the BI framework. A natural way to introduce flexibility is to assume that this cross-partial derivative is constant, summarized by a single parameter. This yields the specification
with $\beta_0 \in B \subset \mathbb{R}$ compact. I refer to this as the Tukey model, after the statistician John Tukey, who studied an analogous functional form in two-way ANOVA to test whether two categorical factors affect the response additively tukey1949one, ward1952non, vsimevcek2013modification. In that setting, the focus is on testing the null hypothesis $\beta_0 = 0$ under restrictive assumptions (homoskedastic, normally distributed errors $\eta_{ij}$ and complete matching network $G_{IJ}$), and no attention is given to the estimation of the parameter itself.
Conversely, in the BI framework, the interaction parameter $\beta_0$ in the Tukey model acquires a direct economic meaning. It is the constant cross-partial $\partial^2 f / \partial\alpha \partial\psi$, capturing all departures from modularity and governing the complementarity pattern between the two productivities.
Despite its parsimony, where the entire complementarity structure is summarized by a single parameter, the Tukey model is flexible enough to encompass supermodular ($\beta_0 \geq 0$), submodular ($\beta_0 \leq 0$), and modular ($\beta_0 = 0$) interaction functions. The sign of $\beta_0$ determines the direction of complementarities, while its magnitude reflects the relative weight of the multiplicative component compared to the additive one: as $|\beta_0|$ grows, the role of the match itself becomes more important.
When $\beta_0 = 0$, the Tukey model reduces to the widely used TWFE specification, in its baseline form without covariates:
In the TWFE model, the cross-partial derivative $\partial^2 f / \partial\alpha \partial\psi$ is zero, so the interaction function is modular and complementarities are ruled out by assumption. The marginal contribution of a worker (firm) is independent of the firm (worker) they are matched with and remains constant across all matches. Under modularity, for a fixed set of matched agents, total output depends only on individual productivities and is invariant to the assignment of matches: any allocation is efficient.
While the assumption of no complementarities is restrictive, especially given the emphasis in economic theory on complementarities as a driver of sorting patterns such as positive assortative matching becker1973theory, shimer2000assortative, its appropriateness depends on the empirical setting and the research question. In practice, the TWFE model is often viewed less as a literal description of interactions and more as a tractable approximation to richer structures abowd1999high, card2013workplace. The nesting of TWFE within the Tukey model makes it possible to assess when this approximation is likely to be informative and, conversely, when ignoring complementarities may lead to misleading conclusions; see Remark (ref) for details.
The Tukey model introduces complementarities in a simple and interpretable way. To capture richer complementarity patterns, one can relax its restrictions and consider more flexible interaction functions: the BLM model allows the interaction parameter to be firm-specific, accommodating heterogeneity in complementarities across firms; the seriation model provides a fully nonparametric alternative, imposing only monotonicity and leaving the cross-partial derivative unrestricted.
Rather than modeling firm productivity as a scalar, suppose instead that each firm is characterized by a productivity vector $\psi_j = (b_j, a_j)$. This leads to the specification
where each firm is endowed with both an intercept and a slope, each varying across firms.
This interaction function is the one studied by bonhomme2019distributional, which motivates the label BLM model. In their approach, the functional form is embedded in a grouped setting, where all firms and workers within a group share the same latent characteristics. In the BI framework, by contrast, the BLM model refers to the case in which each group consists of a single worker or a single firm.
The BLM model nests the Tukey model as a special case: when $b_j = 1 + \beta_0 a_j$ for all $j$, or equivalently $\beta_0 = \tfrac{b_j - 1}{a_j}$, the specification reduces to the Tukey one, with $\psi_j = a_j$. The BLM model can thus be viewed as a generalization of the Tukey model, where the slope parameter capturing complementarities varies across firms. This additional flexibility allows researchers to distinguish a firm's intrinsic productivity $a_j$, which enters the model in levels, from its capacity to extract value from workers and exploit complementarities, governed by $b_j$. Unlike the Tukey model, which implicitly ties these two roles together, the BLM model permits them to differ, thereby enabling empirical investigation of whether more productive firms are also those in which workers' marginal productivity is higher.
The BLM model imposes a specific parametric form on the interaction function. As a benchmark, it is useful to also consider a specification that removes all parametric restrictions on the cross-partial derivative and allows for fully flexible complementarities. The natural analogue comes from nonparametric regression: when the interest is in learning $\mathbb{E}[Y|X]$ nonparametrically from observations of continuous random variables $Y$ and $X$, some regularity assumptions are needed, typically smoothness. Similarly, in the bipartite interaction setting, some regularity condition on $f$ is necessary to ensure that the data are informative.
I retain scalar productivities and impose a monotonicity assumption on $f$, leaving the cross-partial unrestricted:
with $f_m\colon A \times \Psi \to \mathbb{R}$ increasing in both arguments. The monotonicity assumption in a bipartite network setting connects directly to seriation problems in statistics flammarion2019optimal and to matrix-completion approaches for bivariate isotonic matrices under unknown permutations mao2020towards. Outside the bipartite setting, it is also related to nonparametric latent-space models for network formation; see, for example, gao2020nonparametric and references therein.
Monotonicity is a common shape restriction in economics (see the survey in chetverikov2018econometrics). In this context, it requires that the ranking of average outcomes produced by two workers (firms), when matched with the same firm (worker), is preserved across all matches. This property holds, for example, in the TWFE model.
The seriation model imposes no parametric restriction on $f_m$: when it is twice differentiable, monotonicity ensures constant signs for the partial derivatives but imposes no constraint on the cross-partial derivatives that capture complementarities. As a result, the complementarity pattern is highly flexible: the same firm can exhibit supermodular or submodular behavior depending on the productivity of the worker it is matched with. This is not possible in the BLM model, where the cross-partial derivative is fixed for each firm and does not vary across workers.
The three models described above, together with the TWFE model, trace out a spectrum within the BI framework. At one end lies the fully modular TWFE model, which rules out complementarities altogether. At the other end is the seriation model, which only imposes monotonicity and allows complementarities to vary freely across matches. The Tukey and BLM models sit between these extremes: they restrict the complementarity structure parametrically, but retain more flexibility than the TWFE model.
A second dimension of variation arises from the matching network $G_{IJ}$, which encodes the observed links between workers and firms. In the most informative case, $G_{IJ}$ is complete, with $D_{ij}=1$ for every pair. At the opposite extreme, the graph may be sparse (with far fewer links than the maximum possible) and fragmented, with each firm linked to only a handful of workers, and vice versa.
Together, the interaction function $f$ and the network $G_{IJ}$ determine what can be learned from the data. The function $f$ and the dimensions $I$ and $J$ dictate how many parameters must be recovered, while the structure of $G_{IJ}$ governs how much information is available. This creates a trade-off: richer flexibility in $f$ requires stronger assumptions on $G_{IJ}$, while sparser networks can only be informative under more restrictive assumptions on $f$.
This trade-off is asymmetric. The function $f$ is unobserved and summarizes the latent interaction between workers and firms; $G_{IJ}$ is observed and can be inspected directly. Restrictions on $G_{IJ}$ can therefore be verified in the data, while restrictions on $f$ cannot. For the TWFE model, abowd1999high and jochmans2019fixed provide conditions on $G_{IJ}$ for identifying and, under additional assumptions, estimating the productivity parameters. For the Tukey, BLM, and seriation models, there are no results that allow the matching network to be incomplete. In the next section, I derive the necessary and sufficient conditions that link the structure of $f$ with the shape of $G_{IJ}$, extending identification results to incomplete and possibly sparse graphs. These conditions involve some requirements on the matching network $G_{IJ}$. For each requirement, I will also briefly assess its plausibility in the context of employer-employee data.
In applications, the researcher can proceed in two steps. First, select the restriction on $f$ that best fits the economic environment. Then, use these results to check whether the observed $G_{IJ}$ satisfies the corresponding conditions. When the focus is on point-identification, three outcomes are possible: (i) the parameters are not identified; (ii) they are identified but not consistently estimable; (iii) they are both identified and consistently estimable. Clarifying which case applies is essential for understanding what questions the data can credibly answer. Naturally, the richer and more connected the graph, the wider the scope of questions that can be addressed.
I analyze identification in the Tukey, BLM and seriation models. Section (ref) introduces the notion of identification adopted for the BI framework. Sections (ref), (ref), and (ref) present the identification results for the three models. Finally, Section (ref) revisits the trade-off between the flexibility of the interaction function and the structure of the matching network in light of these results.
I adopt the classical notion of point identification from koopmans1949identification, which requires that the model parameters be uniquely recoverable from the distribution of observables. In the BI framework, the observables are the collection $\{ y_{ij} \}_{(i,j) \in \mathcal O_{IJ}}$. Hence, for identification purposes, I can treat $\mathbb{E}[y_{ij}] = \theta_{ij}$ as known for each observed match.
Formally, I study identification through the noiseless map
where $(f, \boldsymbol{\alpha}, \boldsymbol{\psi})$ are unknown and the graph $G_{IJ}$ is known. The parameter vector $(f, \boldsymbol{\alpha}, \boldsymbol{\psi})$ is point-identified if the map
is injective: that is, whenever parameters $(f, \boldsymbol{\alpha}, \boldsymbol{\psi})$ and $(f', \boldsymbol{\alpha}', \boldsymbol{\psi}')$ generate the same $\boldsymbol{\theta}_{\mathcal O}$, they must coincide exactly.
In practice, a parameter is identified if and only if it can be expressed as a function of $\boldsymbol{\theta}_{\mathcal O}$: this is the equivalence I will use to establish identification in the proofs.
This definition deliberately abstracts from sampling noise and treats the distribution of each $y_{ij}$ as known, even though in applications each match is typically observed only once. While such a notion does not distinguish between parameters that can or cannot be consistently estimated in the presence of error terms $\eta_{ij}$, it provides a fundamental benchmark: if a parameter is not identified in the noiseless model, then no estimator can recover it. Conversely, whenever a parameter is identified in this sense, it warrants further analysis to determine whether, and under what conditions, consistent estimation is feasible.
A location normalization is often necessary, since the absolute levels of $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$ are not uniquely determined by the observables. In the TWFE model, for example, adding a constant to all $\alpha_i$ and subtracting it from all $\psi_j$ leaves $\boldsymbol{\theta}_{\mathcal O}$ unchanged; the same invariance holds in richer specifications with complementarities, where shifts in $\boldsymbol{\alpha}$ can be offset by changes in $\boldsymbol{\psi}$ and $f$ without altering $\boldsymbol{\theta}_{\mathcal O}$. To fix the reference level and ensure point identification, a normalization such as $\sum_i \alpha_i = 0$ or $\alpha_1 = 0$ is imposed. Whenever such a normalization is required, I state it explicitly and adopt the form that yields the clearest formulas.
The Tukey model introduces complementarities in the interaction function through a single parameter, $\beta_0$, which represents the constant cross-partial derivative and fully characterizes $f$. I present the identification results in two steps: first, the identification of $\beta_0$; second, the identification of $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$.
Identifying $\beta_0$ requires additional structure on the matching network $G_{IJ}$. I begin by recalling the notion of a cycle in a bipartite graph.
The graph in Figure (ref) contains no cycles: no path starts and ends at the same node while traversing distinct edges. By contrast, adding an edge between $i_1$ and $j_2$, as in Figure (ref), creates the closed path $i_1, j_1, i_2, j_2, i_1$, which forms a cycle of length 4.
The key condition required for identification of $\beta_0$ is given in Assumption (ref).
Assumption (ref) requires the presence of a 4-cycle in which both workers and firms differ in productivity. Such heterogeneity is crucial: without it, the outcomes along the cycle would not vary, and the cycle would provide no information about $\beta_0$. When heterogeneity is present, however, the cycle reveals contrasts across matches that make the interaction parameter point-identified.
Identification of $\beta_0$ hence requires the presence of a cycle in the matching network, and if that cycle is unique, it must be of length four. The proof (Appendix (ref)) shows that outcomes from a cycle of length $2K$ generate a degree-$(K-1)$ polynomial in $\beta_0$, with coefficients given by known functions of $\boldsymbol{\theta}_\mathcal{O}$. The identification set consists of the roots of this polynomial. When the network contains at most one cycle, point identification requires $K=2$, i.e.\ a cycle of length 4. More generally, a cycle of length $2K$ yields an identification set with $K-1$ elements. If multiple longer cycles are present, the intersection of their identification sets can reduce to a singleton, achieving point identification.
The role of cycles in detecting departures from modularity was previously noted by card2013workplace and discussed by kline2024firm, though in those cases cycles were used as a diagnostic device rather than as a source of identification for an interaction parameter.
In the labor market setting, Assumption (ref) requires some degree of worker mobility across firms. Because each worker can be matched with only one firm at a time, the multiple links needed for a cycle arise only when workers change employers. The condition therefore, requires that, for some pair of firms, at least two movers exist, and that both the workers and the firms involved differ in productivity. This is not especially restrictive: labor markets are typically segmented into local or sectoral clusters, and when one worker moves between two firms, the likelihood of additional movers between the same firms is higher than under random matching. Evidence supports this: in the application of kline2024firm, for instance, about 55% of firms belong to at least one cycle.
Once $\beta_0$ is identified, the productivity vectors $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$ in the Tukey model can be identified under an additional condition on $G_{IJ}$.
The graph in Figure (ref) violates Assumption (ref), since no path exists, for example, from node $i_4$ to node $i_5$. Adding an edge between $i_4$ and $j_3$, as in Figure (ref), makes the graph connected.
The identification result for $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$ in the Tukey model is reported below.
Theorem (ref) extends the classical identification result for the TWFE model: it requires knowledge of $\beta_0$, but allows this parameter to differ from zero. The proof proceeds by first showing that, for each observed match, one can construct a function of $\boldsymbol{\theta}_\mathcal{O}$ and the identified $\beta_0$ that equals the product of known functions of the corresponding worker and firm productivities. Connectedness of $G_{IJ}$ then guarantees that all worker and firm effects can be recovered up to the normalization, completing the identification argument.
In labor market applications, the matching network is rarely fully connected, especially over short time horizons. A common practice is therefore to restrict the analysis to the largest connected component. For example, in the West German labor market studied by card2013workplace, the largest component contains over 95% of workers and 90% of firms. Similar firm coverage is reported by bonhomme2023much for Austria, Italy, Sweden, Norway, and the United States, although worker coverage is often much lower, in some cases falling below 50%.
Theorems (ref) and (ref) highlight the additional requirements introduced by allowing complementarities. Relative to the TWFE model, the Tukey model imposes only a modest cost: the same connectedness condition is needed, with the sole extra requirement that the graph contain at least one informative cycle to identify $\beta_0$. This result exploits the fact that $\beta_0$ is a global parameter: it governs the entire interaction function $f$ and does not vary across specific workers or firms. Hence, it can be identified using local information in $G_{IJ}$ and then applied across all matches to recover $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$.
The Tukey model, therefore, offers a simple and flexible way to incorporate complementarities into the BI framework, while requiring only minimal conditions on the matching network for identification. By contrast, as I show in the following sections, more general specifications such as the BLM and seriation models demand substantially stronger network requirements, which are typically stringent and rarely satisfied in applications.
The BLM model extends the BI framework by allowing each firm to have its own interaction parameter. bonhomme2019distributional study this specification under the assumption that workers and firms are partitioned into groups, with all agents in a group sharing the same productivity. Grouping substantially reduces the effective number of nodes in the matching network, since each group can be represented as a single node connected to many others. Under the assumption that the matching network is complete, bonhomme2019distributional show that $\alpha_i$, $a_j$, and $b_j$ are point-identified. Completeness, however, is stronger than necessary. In what follows, I establish a weaker connectivity condition on $G_{IJ}$ that still ensures identification of the productivity parameters.
To formalize this condition, I introduce the following property of the bipartite graph $G_{IJ}$.
Intuitively, the procedure alternates between (i) adding all workers linked to any firm already in the snowball and (ii) adding all firms connected to at least two of the workers in it. The “two-worker” condition guarantees that each newly added firm lies on a cycle, and that these cycles overlap so the snowball can propagate through the graph. In Figure (ref), the property fails because firm $j_3$ is not part of any cycle. Adding the edge $(i_5,j_1)$, as in Figure (ref), creates overlapping cycles and makes the graph Seed-and-Snowballs connected.
To verify this, run the iterative procedure with $j_1$ as the seed, so that $S^J_0 = \{ j_1 \}$. First, include the workers linked to the seed: $S^I_0 = \{ i_1, i_2, i_5 \}$. Then, add firms linked to at least two of these workers: $S^J_1 = \{ j_1, j_2 \}$. Repeating the process yields $S^I_1 = \{ i_1, i_2, i_5, i_3, i_4 \}$ and $S^J_2 = \{ j_1, j_2, j_3 \}$. Since the snowball eventually reaches every node, the graph satisfies Seed-and-Snowballs connectivity.
To the best of my knowledge, this property has not been discussed in the existing network literature. Appendix (ref) shows that it is sufficient for the form of connectivity required by kline2020leave to ensure the validity of their variance estimator in the TWFE model, where the matching network must remain connected after removing any single worker node along with its incident edges.
As in the Tukey model, Seed-and-Snowballs connectivity must be complemented by sufficient heterogeneity in the productivities appearing in the restrictions to obtain the key condition for identification.
Intuitively, since the BLM model generalizes the Tukey model by allowing a firm-specific interaction parameter, Assumption (ref) extends Assumptions (ref) and (ref). It requires that each firm lie on a cycle involving heterogeneous productivities, and that these cycles overlap at least at one node. This overlap enables information to propagate through the network and ensures identification of the parameters in the BLM model, as formalized in the next theorem.
While Assumption (ref) is necessary for identification, it demands a matching network far richer than what is typically observed in applications. In the labor market context, for example, kline2024firm finds that nearly half of the firms in their data are not part of any cycle, implying that their productivities cannot be identified. This is only a lower bound: being part of a cycle is not enough, since the cycles must also share at least one node. As a result, even under favorable conditions, the BLM model would fail to identify the productivity of a large share of firms, limiting its empirical applicability.
Theorem (ref) thus highlights that the additional flexibility of firm-specific complementarities comes at a steep cost: the matching network must satisfy a stringent requirement that is rarely met in practice. Point identification in the BLM model is therefore challenging in most empirical settings with two-sided interactions, unless one imposes additional dimension-reduction restrictions such as those in bonhomme2019distributional.
One might ask whether the restrictive graph condition found for identification with the BLM model is driven not by its richer interaction function, but by the use of multidimensional firm productivity, which introduces additional parameters.
The seriation model provides an alternative: it allows fully flexible complementarities while retaining scalar productivities. Here, the interaction function is left entirely nonparametric. As the next proposition shows, this flexibility comes with a sharp limitation: $f_m$, $\boldsymbol{\alpha}$, and $\boldsymbol{\psi}$ can be recovered only up to strictly monotonic reparameterizations.
Proposition (ref) implies that only the ordinal information (the ranking) of worker and firm productivities can be identified. Cardinal differences are not preserved under strictly monotonic transformations, so $f_m$, $\boldsymbol{\alpha}$, and $\boldsymbol{\psi}$ cannot be separately identified.
While this prevents recovery of productivity levels, the ranks of $\{ \alpha_i \}$ and $\{ \psi_j \}$, and thus the rankings of the vectors $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$, are still identified. Rank-based methods can therefore be employed to study sorting patterns and the determinants of productivity ranks without imposing cardinal structure.
To state an identification condition for the seriation model, I first introduce the notion of within-side diameters.
In the graph of Figure (ref), the within-side diameter for $I$ is $4$, as the shortest path from $i_1$ to $i_5$ spans four edges. For $J$, the diameter is $4$, corresponding to the path between $j_1$ and $j_3$. In Figure (ref), the diameter for $J$ falls to $2$ thanks to the shorter path $j_1 \to i_5 \to j_3 $, while the diameter for $I$ remains $4$.
The following condition for identification in the seriation model directly involves the within-side diameters.
Assumption (ref) requires that every pair of workers shares at least one common firm, and every pair of firms shares at least one common worker. This ensures that all nodes on the same side of the bipartition are linked through a single intermediary on the opposite side. The graph in Figure (ref) fails this property, as workers $i_3$ and $i_5$ have no firm in common. Adding the edge $(i_3,j_3)$, as in Figure (ref), creates the necessary connections and yields within-side diameters of two.
The next result shows that a within-side diameter of two is exactly the connectivity needed to recover the rankings of $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$.
From an empirical standpoint, Assumption (ref) is strong: it requires that every pair of workers have at least one firm in common and that every pair of firms share at least one worker. In labor market data, this would mean that any two workers have worked for the same firms, an unlikely occurrence outside of small or highly interconnected markets. As with the BLM model, the seriation model therefore has limited empirical applicability when the focus is on point identification.
That said, point identification is not always essential for extracting useful information from the data. Even when Assumption (ref) fails, the seriation model can still deliver partial identification of the productivity rankings. Although I do not study partial identification in this paper, it remains a promising direction to explore. Following the approach by crippa2025pairwise, for example, one could derive informative sets for the rankings of $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$.
When the identification conditions for the TWFE, Tukey, BLM, and seriation models are compared, the trade-off introduced in Section (ref) becomes clear: weaker restrictions on the interaction function require richer structure in the matching network.
Theorems (ref) and (ref) show that the conditions for point identification in the BLM and seriation models are unlikely to be satisfied in most empirical settings involving two-sided interactions. By contrast, Theorems (ref) and (ref) demonstrate that the interaction parameter $\beta_0$, together with $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$, can be identified under assumptions that are plausibly met in practice.
The Tukey model thus emerges as a more flexible yet tractable alternative to the TWFE model. It is flexible because it accommodates complementarities, through the parameter $\beta_0$. It is tractable because it imposes only mild requirements on the network, unlike its extensions and similarly to the TWFE. In the next section, I show that the additional interaction parameter can be consistently estimated under assumptions that are commonly satisfied by matching networks observed in applications.
I study estimation and inference for the parameters of the Tukey model. Section (ref) presents the asymptotic setting used to analyze estimators in the BI framework. Section (ref) introduces an estimator for $\beta_0$, establishes its consistency, derives its asymptotic distribution, and shows how it can be used to construct a test for the absence of complementarities. Section (ref) then considers the estimation of productivities $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$. While I propose an estimator, I do not study its properties; instead, I show how its analysis and the limitations it faces connect to those of the existing productivity estimators in the TWFE model.
To study the large-sample properties of the estimators introduced below, I consider an asymptotic setting in which both the number of workers $I$ and the number of firms $J$ grow large ($I \to \infty$, $J \to \infty$). Information accumulates by adding new nodes to the bipartite graph, rather than by repeatedly sampling matches along existing edges.
The analysis is hence based on an asymptotic setting in which the network expands. The sequences of matching networks $G_{IJ}$, worker productivities $\boldsymbol{\alpha}_I$, and firm productivities $\boldsymbol{\psi}_J$ are taken as deterministic, and the set of observed matches $\mathcal{O}_{IJ}$ is allowed to vary with $(I,J)$. Randomness arises solely from the error terms $\{\eta_{ij}\}_{(i,j) \in \mathcal{O}_{IJ}}$, which are assumed independent across matches.
This setup contrasts with stochastic network formation models, where the graph itself is random. Instead, the analysis reflects empirical applications in which the observed labor market network is treated as fixed, and inference concerns the role of unobserved shocks given this network structure. Equivalently, the setting can be viewed as one where the network and productivities are random, but inference is conducted conditional on their realization.
The role of the matching network in the asymptotic analysis mirrors its role in identification: conditions on the sequence of graphs $G_{IJ}$ ensure the validity of the asymptotic results. These conditions cannot be verified from a single observed graph. Instead, what matters is how the observed structure relates to the properties required of the asymptotic sequence, to assess whether the asymptotic results provide a good approximation to the finite-sample behavior.
In labor market applications, the large-$I$, large-$J$ asymptotic setting is well-suited, as available datasets often include millions of workers and firms. At the same time, the data are sparse: each worker is observed with only a handful of firms, and each firm with only a modest number of workers. For example, in the U.S. labor market over a one-year horizon, a typical worker is employed by one or two firms, while a typical firm hires only a few dozen workers. To capture this structure, the asymptotic setting allows node degrees to remain bounded as $I$ and $J$ grow, rather than requiring any worker to be linked with many firms or any firm with many workers.
Recall the Tukey model:
with $\mathbb{E}[\eta_{ij}] = 0$, and $\{\eta_{ij}\}_{(i,j) \in \mathcal{O}_{IJ}}$ independent.
This section introduces an estimator for $\beta_0$, and analyzes its properties in the large sample setting described above. Note that, despite tukey1949one and the following literature on non-additivity in ANOVA consider this same model, they do not discuss any estimator for $\beta_0$, rather focusing on directly testing the hypothesis $\beta_0 = 0$.
Theorem (ref) shows that a single informative four-cycle suffices to identify $\beta_0$. In the presence of noise, however, estimation requires pooling information across many four-cycles present in $G_{IJ}$. It is therefore convenient to treat each four-cycle as an observational unit and impose conditions on the sequence of matching networks that guarantee the number of distinct four-cycles grows with $I$ and $J$.
Index the four-cycles in $G_{IJ}$ by $\ell = 1, \dots, L$. For clarity of exposition, I restrict attention to edge-disjoint cycles, assuming that no cycles share an edge (but they can share one or two nodes). This restriction is not essential, and information from overlapping cycles can also be aggregated, but focusing on edge-disjoint cycles keeps the notation tractable.
Each cycle $\ell$ consists of two distinct workers $\{i_\ell,i'_\ell\}$ and two distinct firms $\{j_\ell,j'_\ell\}$. For expositional purposes, suppose labels are ordered so that $\alpha_{i_\ell} \geq \alpha_{i'_\ell}$ and $\psi_{j_\ell} \geq \psi_{j'_\ell}$. Of course, these labels are unknown to the researcher: they only observe the pairs in the cycle, not the underlying productivities.
To work with the formulas below, the researcher must nonetheless assign distinct labels to workers and firms in each cycle. At this stage, I leave the rule for assigning labels unspecified; later, I propose a procedure that uniquely determines them.
Formally, the researcher assigns labels $(i_{\ell,\pi_\ell}, i'_{\ell,\pi_\ell})$ and $(j_{\ell,\pi_\ell}, j'_{\ell,\pi_\ell})$ to the worker and firm pairs $\{i_\ell,i'_\ell\}$ and $\{j_\ell,j'_\ell\}$. Given this labeling, define
where
In words, $\pi_\ell^\alpha$ ($\pi_\ell^\psi$) equals $1$ when the assigned label $i_{\ell,\pi_\ell}$ ($j_{\ell,\pi_\ell}$) corresponds to the higher-productivity worker (firm), and $-1$ otherwise. Because productivities are unobserved, $\pi_\ell$ cannot be directly chosen, but each labeling rule uniquely determines its value. For this reason, I refer to the label assignment chosen by the researcher simply as $\pi_\ell$.
For each cycle $\ell$ with labeling $\pi_\ell$, define
Intuitively, $\hat\Delta_{1, \ell, \pi_\ell}$ compares the difference in outcomes between workers across firms, while $\hat\Delta_{2, \ell, \pi_\ell}$ contrasts two cross-products, each involving all four nodes in the cycle. Both statistics depend on how workers and firms are labeled within the cycle, as shown by the following example.
With these statistics in hand, define the estimator
The estimator $\hat\beta_{L,\pi}$ does not require estimating $\boldsymbol{\alpha}$ or $\boldsymbol{\psi}$. As a result, $\beta_0$ can be estimated consistently even when individual productivities cannot be, for example when each worker is observed only a few times; see Remark (ref) for details.
The estimator depends on the labelings $(\pi_1,\dots,\pi_L)$. In a network with $L$ cycles, different labeling combinations can generate up to $2^L$ distinct estimators of $\beta_0$. Such dependence on arbitrary labeling is undesirable, since two researchers analyzing the same data could obtain different estimates solely because they chose different labelings. To remove this ambiguity, I later introduce a procedure that selects a particular combination of labelings.
To see why $\hat\beta_{L, \pi}$ is an estimator for $\beta_0$, decompose $\hat\Delta_{1, \ell, \pi_\ell}$ and $\hat \Delta_{2, \ell, \pi_\ell}$ as:
where
and
Here, the absolute values of $\Delta_{1,\ell,\pi_\ell}$ and $\Delta_{2,\ell,\pi_\ell}$ depend only on the model parameters, while their sign is determined by the labeling $\pi_\ell$. The random terms $\epsilon_{\Delta_1,\ell,\pi_\ell}$ and $\epsilon_{\Delta_2,\ell,\pi_\ell}$ are mean-zero combinations of the four errors $\eta_{ij}$, with their sign again determined solely by $\pi_\ell$. This decomposition highlights why averaging across many cycles reduces sampling variability: the error terms $\{\eta_{ij}\}$ are independent and mean zero across edges. Taking the ratio of these averages then cancels the common dependence on latent productivities $(\boldsymbol{\alpha}, \boldsymbol{\psi})$ and on labelings, leaving only the parameter $\beta_0$.
The next section formalizes this intuition and establishes that, under suitable conditions on the sequence of graphs, error terms, and labelings, $\hat\beta_{L, \pi}$ converges almost surely to $\beta_0$.
To establish the consistency of $\hat\beta_{L, \pi}$, I require the following conditions.
Assumption (ref) imposes standard regularity conditions on the error terms. Importantly, it does not require the $\eta_{ij}$ to be identically distributed and allows for heteroskedasticity: each error may follow its own distribution, provided it has mean zero, is non-degenerate, and admits $(2+\delta)$ moments. While the non-degeneracy and moment conditions are not strictly necessary for consistency, they are needed for deriving the asymptotic distribution of the estimator in subsequent results.
Assumption (ref) concerns the sequence of matching networks and requires that the number of cycles grows with the overall size of the network. The condition does not restrict node degrees and is compatible with graphs $G_{IJ}$ where degrees remain bounded: for example, no worker or firm needs to appear in more than two matches. Since $\hat\beta_{L, \pi}$ treats cycles as the fundamental units of observation, the assumption guarantees that the number of such units increases sufficiently to justify asymptotic analysis.
Assumption (ref) requires that the average, across cycles, of the products of worker and firm productivity differences remains bounded away from zero. Since this condition is formulated using ordered labels, it does not depend on the specific choice of labeling. It is closely related to Assumption (ref), which underlies identification, and guarantees the presence of systematic heterogeneity in productivities across cycles. Put differently, because $\mu_L$ is an average of nonnegative terms, requiring $\mu_L$ to stay strictly positive implies that a non-vanishing fraction of the summands must themselves be bounded away from zero. This excludes degenerate cases in which heterogeneity vanishes asymptotically, requiring enough variation across workers and firms to make the cycles informative.
Assumption (ref) restricts the sequence of labelings. It rules out labeling schemes that drive the average of $\Delta_{2,\ell,\pi_\ell}$ toward zero, since in that case both the numerator and denominator of the estimator would vanish, making $\beta_0$ possibly unrecoverable. It also excludes labelings that are systematically correlated with the error terms, as such dependence would invalidate the averaging argument underpinning the law of large numbers. The assumption does not prescribe a unique labeling rule; rather, it specifies the conditions that any labeling must satisfy to be admissible, allowing for both deterministic and stochastic labeling rules. In Section (ref), I present one such rule, based on observable instruments, which ensures the condition and provides a practically implementable estimator.
Strong consistency of $\hat\beta_{L, \pi}$ for $\beta_0$ then follows from the strong law of large numbers, as stated in the following theorem.
The assumptions on $G_{IJ}$ required for consistency of $\hat\beta_{L}$ are relatively weak. For instance, under the bipartite Erdős--Rényi random graph model, Assumption (ref) holds whenever the link probability $p_{IJ}$ satisfies $\sqrt{IJ}p_{IJ} \to \infty$ (see Appendix (ref) for the formal definition of the model and the derivation of this threshold). By contrast, ensuring that the graph is connected requires the stronger condition $\frac{\sqrt{IJ}p_{IJ}}{\log(\sqrt{IJ})} > 1$. Hence, under the Erdős--Rényi model, the condition on the matching network needed for identification of $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$ automatically guarantees the one required for consistency of $\hat\beta_{L, \pi}$.
The effective sample size for estimating $\beta_0$ is $L$, the number of cycles, rather than the number of observed matches. This parallels other settings in econometrics: in local linear regression, the effective sample size is proportional to the number of observations times the bandwidth, while in clustered data it is given by the number of clusters rather than the number of units.
In labor market applications, Assumption (ref) is satisfied whenever the number of cycles grows proportionally with the number of workers and firms, even if the graph consists of many disconnected local subgraphs. For instance, in the empirical setting of kline2024firm, the data contain roughly 750{,}000 workers, 70{,}000 firms, and 5{,}000 cycles: the number of cycles is large enough that asymptotic approximations are likely to provide a good guide to finite-sample behavior.
The next step is to characterize the asymptotic distribution of $\hat\beta_{L,\pi}$, which provides the basis for valid inference and for testing the absence of complementarities.
The estimator $\hat \beta_{L, \pi}$ is a ratio of averages, which makes its asymptotic behavior amenable to analysis via the Lyapunov Central Limit Theorem. To set up the argument, define
with $\epsilon_{\Delta_1,\ell, \pi_\ell}$ and $\epsilon_{\Delta_2,\ell, \pi_\ell}$ defined in Equations (ref) and (ref). The mean-zero composite error $u_{\ell, \pi_\ell}$ depends on $\beta_0$, the four productivity terms, the labeling, and the four error terms associated with cycle $\ell$.
With this notation in place, the asymptotic distribution of $\hat{\beta}_{L,\pi}$ can be derived.
The estimator $\hat \beta_{L,\pi}$ converges at rate $\sqrt{L}$, where $L$, the number of cycles, acts as the effective sample size. In a complete bipartite graph, this corresponds to the rate $\sqrt{IJ}$, since the edges can be partitioned into $\frac{I}{2} \times \frac{J}{2}$ edge-disjoint four-cycles. This rate is faster than the $\min\{\sqrt{I}, \sqrt{J}\}$ rate obtained by iterative procedures, reflecting the more efficient aggregation of information across the graph. An important implication is that $\hat \beta_{L,\pi}$ remains consistent even when one dimension, either $I$ or $J$, is held fixed while the other grows.
In addition to the sample size, three scaling terms appear in the asymptotic distribution. The first, $\sigma_{u,L}$, is the square root of the average variance of the composite error terms $u_{\ell,\pi}$. Although these errors depend on the chosen labeling, their variances do not; hence $\sigma_{u,L}$ is invariant to labeling. It reflects only the variability of the underlying noise terms $\eta_{ij}$: greater noise in $y_{ij}$ increases $\sigma_{u,L}$ and reduces estimator precision.
The second, $\mu_L$, summarizes the heterogeneity in worker and firm productivities across cycles. If workers and firms within cycles are too similar, $\mu_L$ is small and the estimator becomes imprecise. By contrast, stronger heterogeneity makes $\mu_L$ larger, yielding sharper estimates.
The third, $c_\pi$, isolates the effect of labeling choice. It shows explicitly how different labeling rules can affect the variance of the estimator, with some labelings increasing its precision relative to others.
\paragraph{Feasible Scaling Factor Estimation.}
To make the asymptotic normality result operational for inference, the scaling factor $\frac{\sigma_{u,L}}{\mu_L c_\pi}$ must be estimated. A consistent estimator is $\frac{\sqrt{\hat\sigma_{u,L, \pi}^2}}{-\frac{1}{L}\sum_{\ell=1}^{L}\hat \Delta_{2, \ell, \pi}}$, where $-\frac{1}{L}\sum_{\ell=1}^{L}\hat \Delta_{2, \ell, \pi}$ consistently estimates $\mu_L c_\pi$ (as established in the consistency proof), and $\hat\sigma_{u,L, \pi}^2$ estimates $\sigma_{u,L}^2$. A convenient choice for $\hat\sigma_{u,L, \pi}^2$ is the average of the squares of the estimated residuals $\hat u_{\ell, \pi}$:
whose consistency is established in the following proposition.
Theorem (ref), together with these feasible estimators for the scaling terms, yields valid asymptotic confidence intervals for $\beta_0$. The next section introduces a practical procedure for selecting labelings in each cycle, ensuring Assumption (ref) holds, uniquely defining the estimator, and guaranteeing the validity of inference.
Assumption (ref) restricts how labels can be assigned within each cycle but does not prescribe a specific rule. In this section, I propose a rank-based procedure for assigning labels $(i_{\ell,\pi_\ell}, i'_{\ell,\pi_\ell})$ and $(j_{\ell,\pi_\ell}, j'_{\ell,\pi_\ell})$ in each cycle, and show that it satisfies Assumption (ref). As a result, the estimator with this rank-based labeling is consistent and asymptotically normal, and not dependent on arbitrary labeling choices.
Suppose the researcher observes some characteristics of workers and firms, denoted by $\boldsymbol{z^\alpha} = (z^\alpha_1, \dots, z^\alpha_I)$ and $\boldsymbol{z^\psi} = (z^\psi_1, \dots, z^\psi_J)$, and treated as nonrandom. Labels in each cycle are then assigned using these instruments according to the following rule.
Under this rule, the worker and firm with larger instrument values are always labeled $i_{\ell, \pi_{\ell,z}}$ and $j_{\ell, \pi_{\ell,z}}$, respectively. This induces the signs
For expositional clarity, I exclude the possibility of ties in the instruments; when ties occur, they can be resolved at random.
Let $\hat\beta_{L,z}$ denote the Rank-Based Labeling (RBL) estimator, defined using the rank-based labeling:
The properties of $\hat\beta_{L,z}$ depend on the choice of instruments $\boldsymbol{z}^{\alpha}$ and $\boldsymbol{z}^{\psi}$.
Consider first the oracle case in which the instruments available to the researcher coincide exactly with the latent productivities $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$. In this case, the induced labelings satisfy $\pi_{\ell,z} = 1$ for all $\ell$, so that Assumption (ref) holds with $c_\pi = 1$. Such oracle instruments are, of course, infeasible in practice, but Assumption (ref) does not require $c_\pi =1$, just that it remains bounded away from zero. Intuitively, this occurs whenever the instruments correctly rank the higher-productivity worker and firm more often than not.
In applications, instruments should therefore be observable characteristics plausibly associated with latent productivities. For instance, years of schooling and firm size can serve as instruments for $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$, as they are often strongly correlated with worker and firm productivity. By contrast, outcomes themselves, though mechanically related to productivities, violate the exogeneity requirement, since they also depend on the error terms $\{\eta_{ij}\}$. This means that, as shown in Appendix (ref), using outcomes as instruments leads to biased estimates and invalid inference.
The next example provides a practical illustration of how labels are assigned.
To formalize the conditions that instruments must satisfy in addition to exogeneity, define the averages
and impose the following conditions on the sequences $\{\pi_{\ell,z}^\alpha\}$ and $\{\pi_{\ell,z}^\psi\}$.
Assumption (ref) requires that the instruments contain useful information about the latent productivities: the rankings they induce must align with the true rankings often enough. Perfect accuracy is not necessary, and occasional misorderings are allowed, as long as the instruments select the correct ordering in a sufficiently large fraction of cycles.
Assumption (ref) rules out a systematic negative association between worker-side and firm-side labelings. That is, it excludes the case in which the worker instrument tends to misorder exactly when the firm instrument orders correctly (or vice versa). The restriction applies only on average: isolated instances of such behavior are admissible provided they do not dominate.
Assumption (ref) prevents a systematic association between large productivity gaps and mislabeling. Specifically, it rules out the possibility that cycles with large differences in worker-firm productivities are disproportionately associated with incorrect rankings. This condition is mild: in practice, misclassifications are more likely when productivity gaps are small, not when they are large.
The next proposition establishes that, when the instruments satisfy Assumption (ref), the rank-based labeling $\pi_{\ell,z}$ satisfies the condition required by Assumption (ref). Consequently, the (ref) $\hat\beta_{L,z}$ achieves the asymptotic properties established in Theorems (ref) and (ref).
Proposition (ref) ensures that the rank-based labeling delivers a well-defined estimator $\hat\beta_{L,z}$ with the aforementioned asymptotic properties. This result allows $\hat\beta_{L,z}$ to serve not only as an estimator of the interaction parameter but also as the basis for the construction of a formal test for the absence of complementarities and hence for modularity of the interaction function in the BI framework. The next section develops this test and studies its properties.
The TWFE model is nested within the Tukey model as the special case $\beta_0 = 0$. Hence, the asymptotic distribution of the (ref) $\hat{\beta}_{L,z}$ can be used to test the null hypothesis $H_0 : \beta_0 = 0$. Rejecting $H_0$ not only rejects the TWFE specification, but also rejects modularity of the interaction function $f$, since any modular function can be written in the additive form assumed by the TWFE model (see Appendix (ref)). To the best of my knowledge, no formal test of modularity has previously been available in the BI framework, even though empirical discussions often revolve around whether complementarities are present.
Testing $H_0 : \beta_0 = 0$ is the central focus of tukey1949one, although no estimator for $\beta_0$ is proposed there. Relying on restrictive assumptions (homoskedastic normally distributed errors $\eta_{ij}$ and complete matching network $G_{IJ}$), Tukey's procedure first estimates the additive effects $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$ under the null through two-way fixed effects regression, and then test whether including the interaction term $\alpha_i \psi_j$ significantly improves model fit, for example through regression-based tests or changes in $R^2$.
Similar heuristic approaches have been adopted in the TWFE literature (e.g., card2013workplace; fenizia2022managers), but the sparsity of the matching network $G_{IJ}$ makes consistent estimation of $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$ infeasible, so that any procedure based on regression residuals seem to be justifiable only under strong assumptions on the error terms and the matching network. By contrast, the estimator $\hat{\beta}_{L,\pi}$ remains consistent even under heteroskedasticity and in sparse matching structures.
\paragraph{Test Statistic.} Consider the t-statistic
and define the test $\phi_{L,z}$ with size $\gamma$ that rejects the null according to
where $c_{\gamma/2}$ is the $\gamma/2$ quantile of the standard normal distribution. The asymptotic validity and consistency of this test follow directly from the asymptotic normality result in Theorem (ref), as summarized in the following corollary.
The test $\phi_{L,z}$ controls asymptotic size for any modular function, but Theorem (ref) guarantees its consistency only under correct specification of the Tukey model. In particular, a non-modular function $f$ can admit a representation with $\beta_0 = 0$, in which case the test has no power. Thus, $\phi_{L,z}$ tests a necessary but not sufficient condition for modularity. The situation is analogous to testing independence between two variables using the correlation coefficient: while a nonzero correlation implies dependence, a correlation of zero does not rule out dependence. In practical terms, rejection of the null provides strong evidence against modularity, but failure to reject should be interpreted with caution, as it does not imply that the interaction function is modular.
Theorem (ref) shows that, once $\beta_0$ is known, connectedness of the matching graph guarantees identification of $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$. Two cases arise. When $\beta_0 = 0$, the model reduces to the standard TWFE specification, and existing estimators can be applied. I briefly review the available results and highlight how they rely on strong conditions that are rarely satisfied in labor market data. When $\beta_0 \neq 0$, I propose a method that uses the consistent estimator $\hat{\beta}_{L,\pi}$ as input for estimating the productivity parameters. A full analysis of its properties is left for future work. As with the TWFE model, this approach ultimately requires denser graphs than are typically observed in applications, reflecting a shared limitation between the two cases.
When $\beta_0 = 0$, the Tukey model reduces to:
and the productivity components can be estimated using the TWFE estimator:
where the design matrix $C$ is defined as in Definition (ref).
jochmans2019fixed study the asymptotic properties of the TWFE estimator. For inference on a single productivity value, they derive its asymptotic distribution under the condition that the degree of the corresponding node diverges. More generally, inference on functionals of $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$ requires that the degrees of many nodes grow. As the authors emphasize, this setting is far from typical labor market applications, where node degrees usually remain bounded: the number of workers per firm or firms per worker does not increase proportionally with $I$ and $J$.
kline2020leave also study inference for functionals of the productivities, focusing on quadratic forms such as variances and covariances. Their asymptotic setting requires the matching network $G_{IJ}$ to grow in a uniformly connected manner, without fragmenting into weakly linked subgraphs. As their Table IV shows, however, this condition is rarely satisfied in labor market data, where matching networks often exhibit considerable fragmentation.
These results highlight that in most labor market applications, and indeed in other two-sided settings as well, the network is too sparse to justify asymptotic arguments for productivity estimation. This is not surprising: information on individual productivity depends only on the relatively few edges involving that node, unlike global parameters such as $\beta_0$, which influence all observed matches.
The estimator $\hat{\beta}_{L,\pi}$ can be used to construct least-squares estimators for $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$ when $\beta_0 \neq 0$. I outline a simple procedure for doing so, which turns out to involve the same computational problem as estimating interactive fixed effects.
Starting from the Tukey model in Equation (ref), multiply both sides by $\beta_0$ and add 1:
Now substitute $\beta_0$ with its consistent estimator $\hat{\beta}_{L,\pi}$ and define:
From these transformed variables, estimates for the original productivity terms can be recovered from estimates of $\boldsymbol{\alpha}'$ and $\boldsymbol{\psi}'$. Since
a natural approach to estimate $\boldsymbol{\alpha}'$ and $\boldsymbol{\psi}'$ is to solve the least-squares problem:
This is the interactive fixed-effects problem, and can be solved via alternating least squares; see Appendix (ref) for implementation details.
A full analysis of the statistical properties of the productivity estimators derived from $(\hat{\boldsymbol{\alpha}}', \hat{\boldsymbol{\psi}}')$ lies beyond the scope of this paper and is left for future research. For this reason, in the next section, when I illustrate how the Tukey model can be applied in practice, the focus will be on $\beta_0$, the new parameter introduced by the Tukey specification, for which I developed an estimator with formally studied asymptotic properties.
In this section, I revisit the application in limodio2021bureaucrat\footnote{The data used for this illustration exercise are publicly available on the author's website.} to illustrate how to estimate the interaction parameter in the Tukey model and how the resulting estimates can be used to draw additional insights about the two-sided interaction.
limodio2021bureaucrat studies the interaction between managers and tasks in the public sector, focusing on the allocation of World Bank bureaucrats (hereafter, managers) to development projects in low- and middle-income countries. Each manager is responsible for designing, supervising, and overseeing project implementation. Project success is measured using ratings from the World Bank's Independent Evaluation Group, which assess the extent to which key objectives were achieved. In the original analysis, project success is modeled as a function of a manager-specific ability $\alpha_i$ and a country-specific characteristic $\psi_j$, using the TWFE model. An administrative dataset records manager-country assignments over time, along with project-level characteristics and evaluations, making it possible to construct the matching network and observe the outcome corresponding to each edge.
The combination of assignment data and standardized performance evaluations provides an ideal setting to study the structure and consequences of bureaucrat-task matching in an international organization. The main result in limodio2021bureaucrat is the presence of negative sorting: high-performing managers are disproportionately allocated to low-performing countries. Possible explanations include the Bank's strategic objective of assigning stronger managers to weaker countries, internal career incentives and promotion dynamics, the demand for specialized skills in more difficult environments, and reallocations following adverse shocks such as natural disasters.
For this illustration, I focus on a version of the model without controls. This differs from the main specification in limodio2021bureaucrat, which includes year and sector fixed effects in the two-way fixed effects regression. The discussion here should therefore be viewed purely as an application of the methods developed in this paper, not as a critique of or challenge to their findings. Those results rely on additional assumptions that are not addressed in the following analysis.
The data contain 3{,}385 projects, corresponding to 1{,}876 distinct manager-country pairs. When the same match $y_{ij}$ is observed multiple times, I use the average outcome. The resulting matching network consists of 697 manager nodes, 127 country nodes, and 1{,}876 edges. Within this graph, there are 228 edge-distinct four-cycles, involving 369 managers and 114 countries. Thus, the information effectively used to estimate the Tukey interaction parameter comes from approximately half of the edges and managers, and from about 90% of the countries.
I consider the (ref) $\hat{\beta}_{L,z}$, which assigns labels in each cycle using auxiliary information on managers and countries. This requires instruments, observable characteristics of managers and countries satisfying Assumption (ref). I use the average project size as the instrument for managers and the Public Infrastructure Management Index (PIMI) as the instrument for countries. limodio2021bureaucrat documents that project size, measured by the average loan amount overseen, is predictive of managerial ability: more capable managers tend to supervise larger loans. Similarly, the PIMI developed by dabla2012investing is predictive of institutional productivity, with higher-performing countries scoring higher on this index. Assumptions (ref) and (ref) hold provided that the rankings induced by these instruments are not systematically opposed within cycles and not systematically misaligned when productivity gaps are large. Since such violations would require counterintuitive patterns, the assumptions appear plausible in practice. Using average loan size for managers and the PIMI for countries, therefore, offers a feasible way to implement Assumption (ref).
Figure (ref) reports the value of the (ref), together with the corresponding confidence interval for nominal coverage of $0.9$. The estimate is negative, indicating negative complementarities between managers and countries: the relative contribution of a high-ability manager, compared to a lower-ability one, is greater when working in a lower-productivity country. The magnitude of the estimate (0.196) can be interpreted as the ratio of the multiplicative to the additive component. This suggests that, while smaller in importance than additive effects, the multiplicative part of the interaction function still plays a meaningful role.
The confidence interval $(-0.293, -0.099)$, centered around $\hat{\beta}_{L,z}$, does not include zero. The p-value for the null $\beta_0 = 0$ equals 0.001, providing strong evidence against the absence of complementarities in this interaction.
To underscore the importance of label choice, the figure also reports estimates obtained under random labeling. These are presented only to illustrate how reliance on uninformed labels affects the estimator and should not be interpreted as informative about the underlying interaction. As expected, random labeling inflates the variance, producing confidence intervals nearly four times wider than those based on rank-based labeling. In this case, the p-value of 0.126 would fail to reject the null of no complementarities, highlighting how proper labeling is crucial for test power, a theme further explored in the Monte Carlo simulations in the next section.
Figure (ref) shows the estimate for a single random labeling, but many such estimates can be computed. Figure (ref) plots the distributions of $\hat{\beta}_{L,\pi}$ and of $\hat\mu_\pi$ (the estimator of $\mu_L c_\pi$, the scaling factor in the asymptotic distribution) across 50,000 random labels assignments, with the values obtained under rank-based labeling highlighted in red.
The distribution of $\hat{\beta}_{L,\pi}$ shows that most labelings yield negative estimates, though with some variation in magnitude. The distribution of $\hat{\mu}_{\pi}$, symmetric around zero, indicates that the instruments are informative about the underlying characteristics: the rank-based labeling produces an estimate of $\hat{\mu}_{\pi}$ larger in absolute value than about 80% of random labelings. At the same time, in 20% of cases random labeling delivers even larger values, reflecting that, as expected, the instruments capture only part of the variation in latent rankings.
Figures (ref) and (ref) thus highlight the importance of selecting labels correctly. The analogy with instrumental variables is direct: valid instruments must be informative about the latent characteristics, not mere noise, in order to deliver consistent and precise estimates.
The negative estimate, statistically different from zero, implies that the skills of high-ability managers matter most in low-productivity countries, where managerial capacity has a greater marginal impact. This explanation was already suggested by limodio2021bureaucrat, but could not be formally assessed under the modularity restriction of the TWFE model. In this view, assigning high-performing bureaucrats to low-performing countries is efficient and raises average project success.
Overall, the empirical illustration shows that the Tukey model is straightforward to implement and provides insights into the structure of interactions that the TWFE model, by construction, rules out. The exercise also underscores the practical relevance of Assumption (ref): only when instruments induce informative labelings do the resulting estimates convey reliable evidence about complementarities. In the next section, I turn to Monte Carlo simulations to evaluate how these insights carry over to controlled settings and to study the finite-sample behavior of the estimator.
In this section, I investigate the finite-sample behavior of the estimator $\hat \beta_{L,z}$. Because estimation of $\beta_0$ comes entirely from the cycles, I abstract from the rest of the network: I fix a number of cycles $L$ and simulate outcomes only on those cycles. This design lets me isolate how performance depends on the effective sample size (the number of cycles), the relevance of the instruments, and the relative magnitude of productivity variation to noise.
Since in simulations the true values of $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$ are known, I can examine the role of the instruments by directly controlling the labeling within each cycle. Concretely, I set the values of $\pi^\alpha_{\ell,z}$ and $\pi^\psi_{\ell,z}$, which is equivalent to choosing instruments that induce those labelings. This approach cleanly maps instrument relevance into the labelings that matter for $\hat \beta_{L,z}$.
The simulation design proceeds in two steps.
Step 1. Generation of productivities. For each of the $L$ cycles, the vector of productivities of two workers and two firms is drawn from a multivariate normal distribution with mean vector $\mu$ and covariance matrix $\Sigma$:
Labels are then assigned by fixing $\pi^\psi_{\ell,z} = 1$ and setting $\pi^\alpha_{\ell,z}$ equal to $1$ with probability $p \geq 0.5$ and to $-1$ with probability $1-p$. This corresponds to using different instruments for $\boldsymbol{\alpha}$: the larger the value of $p$, the stronger the instrument. Equivalently, the same exercise could be conducted using instruments for $\boldsymbol{\psi}$ or for both sides simultaneously.
Step 2. Generation of outcomes. In each Monte Carlo replication, the four error terms for every cycle are drawn independently from a normal distribution with mean zero and variance $\sigma^2_{ij}$. Combined with $\beta_0$, these errors generate the $4L$ outcomes, which form the simulated data observed in that replication.
I explore the finite-sample properties of the estimator by varying the number of cycles $L$, the error variance $\sigma_{ij}$, and the relevance of the instrument $p$:
These choices allow me to examine the role of the assumptions in Theorems (ref) and (ref), which require $L$ to be large, $\sigma_{ij}$ to remain bounded, and $p \neq 0.5$. The case $p=0.5$ corresponds to non-informative labels, where $c_\pi = 2p - 1 = 0$ and Assumption (ref) fails.
For each parameter combination, Step 1 is implemented once, while Step 2 is repeated 10,000 times. Fixing productivities, their labels, and their allocation in cycles across replications mimics the identification and inference analysis, where $\boldsymbol{\alpha}$, $\boldsymbol{\psi}$, $G_{IJ}$, $\boldsymbol{z^\alpha}$, and $\boldsymbol{z^\psi}$ are treated as deterministic.
Table (ref) reports the mean squared errors of $\hat{\beta}_{L,\pi}$ as an estimator of $\beta_0$ across simulations. Table (ref) presents the average widths of the corresponding 90% confidence intervals, while Table (ref) displays the rejection rates of the test for absence of complementarities under the null hypothesis $H_0:\beta_0=0$, for $\beta_0 \in \{0,1\}$ and nominal significance level $\gamma = 0.1$.
When $p>0.5$, and hence Assumption (ref) holds, the mean squared error, the width of the confidence intervals, and the discrepancy between nominal and empirical size all increase with the error variance and decrease with the sample size. By contrast, the power of the test rises with sample size and falls as noise grows. These patterns are consistent with theoretical predictions.
The more interesting patterns concern the role of $p$, which captures the relevance of the instrument. When Assumption (ref) is violated ($p=0.5$), the estimator becomes inconsistent and the asymptotic distribution no longer applies. In this case, the mean squared error does not shrink with sample size, the test fails to control size under the null, and it is inconsistent against the alternative, regardless of sample size or error variance.
When the assumption holds but $c_\pi$ is small (that is, when $p>0.5$ but close to $0.5$), finite-sample performance deteriorates, especially when $L$ is small and $\sigma_{ij}$ is large. The simulations confirm that a small $c_\pi$ inflates the estimator's variance, producing large mean squared errors and wide confidence intervals. This effect is most pronounced when $p$ is small relative to $L$. In such cases, the test for absence of complementarities continues to control size reasonably well, but its power drops sharply, as the inflated variance reduces the ability to reject the null when it is false. This mirrors the empirical illustration, where the null hypothesis $\beta_0 = 0$ was rejected under rank-based labeling but not under random labeling.
Overall, the simulations reinforce the theoretical results and highlight the central role of labeling within cycles: appropriate label selection is crucial not only for estimator precision, but also for ensuring that the test for the absence of complementarities retains meaningful power.
This paper introduced the Bipartite Interaction (BI) framework for modeling two-sided interactions. Within this framework, I proposed notions of identification and large-sample, and studied three models defined by different restrictions on the interaction function, deriving conditions for their identification. The analysis highlighted a fundamental trade-off between flexibility in the interaction function and the density of the matching network: richer interaction structures require stronger graph conditions to achieve point identification.
Among the models, the Tukey specification emerged as a particularly useful case. By summarizing complementarities with a single parameter, it extends the TWFE model in a parsimonious yet interpretable way, and for identification it requires only mild additional conditions beyond those of the TWFE model. I developed a novel cycle-based estimator for its interaction parameter, which avoids estimating latent productivities, is consistent under bounded-degree graphs, and is asymptotically normal. Its asymptotic distribution also provides the basis for a formal test of the absence of complementarities. An empirical illustration showed that the Tukey model can be implemented in settings where the TWFE model is standard, delivering richer insights into the interaction process that would remain hidden under additively separable models. More broadly, the results demonstrate that complementarities matter in practice and can be studied with feasible methods.
Two directions for future research remain open. First, while the analysis here focused on point identification, extending the BI framework to partial identification would be valuable. For the BLM and seriation models, even when point identification fails, the data may still deliver informative bounds. Characterizing identified sets and developing computational methods to recover them could lead to more credible conclusions without strong parametric assumptions. Second, while in this paper the main focus was on estimation of $\beta_0$ in the Tukey model, researchers are often interested also in estimating the productivity parameters $\boldsymbol{\alpha}$ and $\boldsymbol{\psi}$. I outlined some estimators, but a full analysis of their properties is needed to assess when they yield reliable measures of worker and firm productivities. The BI framework, by clarifying the distinct roles of the interaction function and the matching network, provides a natural foundation for these further investigations.