Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
17,644 characters · 6 sections · 18 citation commands
Kotlarski's lemma for dyadic models
\onehalfspacing
Identifying latent variables from observed data is a central problem in econometrics. One of the important tools for addressing this problem is a lemma of kotlarski1967characterizing and its variants, which provide conditions under which the distribution (characteristic function) of latent variables is identified.
Kotlarski's lemma has been used for identification, estimation, and inference in a variety of economic settings, such as measurement error models li1998nonparametric,li2002robust,schennach2004nonparametric,kurisu2022uniform, auctions li2000conditionally,krasnokutskaya2011identification,grundl2019identification, andreyanov2022secret, and models of earning dynamics bonhomme2010generalized,botosaru2018nonparametric,hu2019semiparametric. More recently, Kotlarski's lemma has been used for robust inference in kato2021robust. Finally, generalizations of Kotlarski's lemma exist to the cases of multiple error components or unknown factor loadings szekely2000identifiability,li2020generalization,lewbel2022kotlarski, lewbel2024identification. For a more complete overview of the variants and uses of Kotlarski's lemma, see, e.g., schennach2016recent.
The classical Kotlarski lemma used in most applications assumes repeated measurement with a common latent variable and independent errors, for example, $y_{i,\ell} = c+\alpha_i + \varepsilon_{i,\ell}$, where $\ell=a,b$ for each $i$. In this note, we show how to use the classical Kotlarski lemma to identify the standard two-way dyadic model for bipartite networks, $y_{i,\ell}=c+\alpha_i+\eta_\ell + \varepsilon_{i,\ell}$. We discuss two cases: the partially connected bipartite network and the fully connected bipartite network.
Our main theorem relies on the version of Kotlarski's lemma in evdokimov2012some, and hence does not assume that the characteristic functions (CF) of the error components have no zeros, which would rule out many distributions of interest, including all continuous distributions with compact support and many discrete distributions. Instead, the CFs are allowed to have zeros, as long as they do not overlap with zeros of their first derivatives.
In this section, we briefly discuss the classical lemma by kotlarski1967characterizing and its extension in evdokimov2012some. Suppose we observe two repeated noisy measurements $X_1,X_2$ of a variable $M$,
where $U_1$, $U_2$ are noise variables. Assume that $M,U_1,U_2$ are jointly independent and $\operatorname{\mathbb{E}}[U_1]=0$. The goal is to identify the distributions of $M,U_1,U_2$.
Kotlarski's lemma states that, if the CFs $\phi_M$, $\phi_{U_1}$, and $\phi_{U_2}$ of $M$,$U_1$, and $U_2$, respectively, are nonvanishing everywhere, then these CFs (and hence the distributions) can be recovered from the CF of the observables $X_1,X_2$.
The assumption of everywhere nonvanishing CFs rules out many interesting distributions, such as any nondegenerate distribution with compact support and any discrete distribution with finite support. evdokimov2012some provide an extension of Kotlarski's lemma that relaxes this assumption. Specifically, it only requires that the real zeros of $\phi_{U_1}$ and its derivative $\phi_{U_1}'$ are disjoint and that the zeros of $\phi_{U_2}$ form a set of isolated points, see Assumption A and Lemma 1(b) in evdokimov2012some.
Consider a bipartite network with two sets of nodes, $\{1, 2\}$ and $\{a, b\}$. Assume that the node pairs $(1, a),(1, b)$, and $(2,a)$ are linked, so that the associated variables $y_{1,a}, y_{1,b}$ and $y_{2,a}$ are observed. Notice that we do not assume that nodes $2$ and $b$ are linked so that $y_{2,b}$ may be unobserved. Consider the dyadic model
where $\alpha_1,\alpha_2,\eta_a,\eta_b$ are unobserved random effects and $\varepsilon_{1,a},\varepsilon_{1,b},$ and $\varepsilon_{2,a}$ are idiosyncratic errors. We assume that all the latent variables are jointly independent. We are interested in identifying the distributions of all the latent components $\alpha_1,\alpha_2$, $\eta_a,\eta_b$, $\varepsilon_{1,a}$, $\varepsilon_{1,b}$, and $\varepsilon_{2,a}$ from the distributions of $y_{1,a}$, $y_{1,b}$, and $y_{2,a}$.
Our identification strategy consists of two parts. First, we identify the distributions of $\alpha_1$, $\eta_a$, and $\varepsilon_{1,a}$. Then, we provide two sets of restrictions under which the remaining distributions are identified: (i) the equality of the distributions of $\varepsilon_{1,a}$, $\varepsilon_{1,b}$, and $\varepsilon_{2,a}$ (Assumption (ref)) and (ii) the equality of the distributions of $\alpha_1$ and $\alpha_2$ and those of $\eta_a$ and $\eta_b$ (Assumption (ref)).
Let us now provide the intuition on how the distributions of $\alpha_1$, $\eta_a$, and $\varepsilon_{1,a}$ can be identified by a repeated application of Kotlarski's lemma. First, for a pair $(1,a), (1,b)$, write
where $M=\alpha_1$, $U_1=\eta_a + \varepsilon_{1,a}$, and $U_2= \eta_b + \varepsilon_{1,b}$. By Kotlarski's lemma, the CF $\phi_{\alpha_1}$ is identified. Similarly, for a pair $(1,a), (2,a)$, write
where $\tilde M=\eta_a$, $\tilde U_1=\alpha_1 + \varepsilon_{1,a}$, and $\tilde U_2= \alpha_2 + \varepsilon_{2,a}$. By Kotlarski's lemma, the CF $\phi_{\eta_a}$ is identified. Joint independence of $\alpha_1,\eta_a$, and $\varepsilon_{1,a}$ implies
identifying the distribution of $\varepsilon_{1,a}$, and hence the distributions of $\varepsilon_{1,b}$ and $\varepsilon_{2,a}$. Finally, joint independence of $\alpha_2,\eta_a,$ and $\varepsilon_{2,a}$ implies
and joint independence of $\alpha_1,\eta_b,$ and $\varepsilon_{1,b}$ implies
identifying the distributions of the remaining components $\alpha_2$ and $\eta_b$.
We now state the assumptions needed to make the intuition above rigorous. For a random variable $\zeta$, denote by $\mathcal{Z}_\zeta$ the set of zeros of its CF $\phi_\zeta$ and denote by $\mathcal{Z}_\zeta'$ the set of zeros of the derivative $\phi_\zeta'$ of its CF.
In the identification strategy described above, we apply Kotlarski's lemma in the case when the error terms $U_1,U_2,\tilde U_1,\tilde U_2$ consist of two latent components. The lemma relies on assumptions about these error terms, which are not primitives of our dyadic model. Instead, we want to restrict the latent components in a way that would guarantee that the assumptions on the error terms hold. The following abstract result shows how this can be achieved.
We are now ready to state our adaptation of the main theorem of evdokimov2012some to dyadic data.
Theorem (ref) establishes the distributional identification of $\alpha_1$ and $\eta_a$ and also shows that $\phi_{\varepsilon_{1,a}}$ is identified at all points of the real line except for zeros of $\phi_{\alpha_1}$ or $\phi_{\eta_a}$. We now provide two sets of assumptions under which all the latent distributions are identified (which we call total identification). Both sets of assumptions include the following condition.
Let us show that total identification holds under the distributional homogeneity of the three error terms $\varepsilon_{1,a}, \varepsilon_{1,b}, \varepsilon_{2,a}$.
Next, we show that total identification holds under the distributional homogeneity of the random effects, $\alpha_1, \alpha_2$ and $\eta_a, \eta_b$.
When the bipartite network on the node sets $\{1,2\}$ and $\{a,b\}$ is fully connected, i.e., when all the node pairs $(1,a),(1,b),(2,a)$, and $(2,b)$ are linked, the distributions of all the latent variables $\alpha_1,\alpha_2$, $\eta_a,\eta_b$, $\varepsilon_{1,a}$, $\varepsilon_{1,b}$ , $\varepsilon_{2,a}$, and $\varepsilon_{2,b}$ are identified without any assumptions on the distribution homogeneity (cf. Assumptions (ref) and (ref)).\footnote{For this, Assumption (ref) has to be extended to include the conditions on $\varepsilon_{2,b}$.} To see that, notice that applying Kotlarski's lemma to the pair $y_{i,a}, y_{i,b}$ identifies the distribution of $\alpha_i$, $i=1,2$. Then applying the lemma to the pair $y_{1,c}, y_{2,c}$ identifies the distribution of $\eta_c$, $c=a,b$. Finally, the distribution of $\varepsilon_{i,c}$ is identified via
Formulating rigorous conditions under which this identification strategy is valid can be done along the lines of Section (ref).
We show how the classical lemma of Kotlarski can be employed to identify distributions of latent components in dyadic models for bipartite networks with two-way random effects. When the bipartite graph is partially linked, we provide two sets of assumptions under which all the latent distributions are identified. The first set of assumptions restricts the errors to be identically distributed. The second set of assumptions restricts the random effects to be identically distributed. Our identification result may be applicable to dyadic regressions with random or correlated effects, as discussed in Section 2.3 of bonhomme2020econometric. It can also be used to develop an estimation procedure for this class of models. This is beyond the scope of the current paper, and we defer this agenda to future work.