EconBase
← Back to paper

Kotlarski's lemma for dyadic models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

17,644 characters · 6 sections · 18 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Kotlarski's lemma for dyadic models

\onehalfspacing

abstractWe show how to identify the distributions of the latent components in the two-way dyadic model for bipartite networks $y_{i,\ell}= \alpha_i+\eta_{\ell}+\varepsilon_{i,\ell}$. This is achieved by a repeated application of the extension of the classical lemma of kotlarski1967characterizing in evdokimov2012some. We provide two separate sets of assumptions under which all the latent distributions are identified. Both rely on some of the latent components being identically distributed. JEL Classification: C23 Keywords: Kotlarski lemma, deconvolution, dyadic data, two-way error component, bipartite network

Introduction

Identifying latent variables from observed data is a central problem in econometrics. One of the important tools for addressing this problem is a lemma of kotlarski1967characterizing and its variants, which provide conditions under which the distribution (characteristic function) of latent variables is identified.

Kotlarski's lemma has been used for identification, estimation, and inference in a variety of economic settings, such as measurement error models li1998nonparametric,li2002robust,schennach2004nonparametric,kurisu2022uniform, auctions li2000conditionally,krasnokutskaya2011identification,grundl2019identification, andreyanov2022secret, and models of earning dynamics bonhomme2010generalized,botosaru2018nonparametric,hu2019semiparametric. More recently, Kotlarski's lemma has been used for robust inference in kato2021robust. Finally, generalizations of Kotlarski's lemma exist to the cases of multiple error components or unknown factor loadings szekely2000identifiability,li2020generalization,lewbel2022kotlarski, lewbel2024identification. For a more complete overview of the variants and uses of Kotlarski's lemma, see, e.g., schennach2016recent.

The classical Kotlarski lemma used in most applications assumes repeated measurement with a common latent variable and independent errors, for example, $y_{i,\ell} = c+\alpha_i + \varepsilon_{i,\ell}$, where $\ell=a,b$ for each $i$. In this note, we show how to use the classical Kotlarski lemma to identify the standard two-way dyadic model for bipartite networks, $y_{i,\ell}=c+\alpha_i+\eta_\ell + \varepsilon_{i,\ell}$. We discuss two cases: the partially connected bipartite network and the fully connected bipartite network.

Our main theorem relies on the version of Kotlarski's lemma in evdokimov2012some, and hence does not assume that the characteristic functions (CF) of the error components have no zeros, which would rule out many distributions of interest, including all continuous distributions with compact support and many discrete distributions. Instead, the CFs are allowed to have zeros, as long as they do not overlap with zeros of their first derivatives.

Classical lemma by Kotlarski

In this section, we briefly discuss the classical lemma by kotlarski1967characterizing and its extension in evdokimov2012some. Suppose we observe two repeated noisy measurements $X_1,X_2$ of a variable $M$,

align*[align* omitted — 54 chars of source]

where $U_1$, $U_2$ are noise variables. Assume that $M,U_1,U_2$ are jointly independent and $\operatorname{\mathbb{E}}[U_1]=0$. The goal is to identify the distributions of $M,U_1,U_2$.

Kotlarski's lemma states that, if the CFs $\phi_M$, $\phi_{U_1}$, and $\phi_{U_2}$ of $M$,$U_1$, and $U_2$, respectively, are nonvanishing everywhere, then these CFs (and hence the distributions) can be recovered from the CF of the observables $X_1,X_2$.

The assumption of everywhere nonvanishing CFs rules out many interesting distributions, such as any nondegenerate distribution with compact support and any discrete distribution with finite support. evdokimov2012some provide an extension of Kotlarski's lemma that relaxes this assumption. Specifically, it only requires that the real zeros of $\phi_{U_1}$ and its derivative $\phi_{U_1}'$ are disjoint and that the zeros of $\phi_{U_2}$ form a set of isolated points, see Assumption A and Lemma 1(b) in evdokimov2012some.

Application to dyadic models

Case of partially connected network

Consider a bipartite network with two sets of nodes, $\{1, 2\}$ and $\{a, b\}$. Assume that the node pairs $(1, a),(1, b)$, and $(2,a)$ are linked, so that the associated variables $y_{1,a}, y_{1,b}$ and $y_{2,a}$ are observed. Notice that we do not assume that nodes $2$ and $b$ are linked so that $y_{2,b}$ may be unobserved. Consider the dyadic model

align*[align* omitted — 172 chars of source]

where $\alpha_1,\alpha_2,\eta_a,\eta_b$ are unobserved random effects and $\varepsilon_{1,a},\varepsilon_{1,b},$ and $\varepsilon_{2,a}$ are idiosyncratic errors. We assume that all the latent variables are jointly independent. We are interested in identifying the distributions of all the latent components $\alpha_1,\alpha_2$, $\eta_a,\eta_b$, $\varepsilon_{1,a}$, $\varepsilon_{1,b}$, and $\varepsilon_{2,a}$ from the distributions of $y_{1,a}$, $y_{1,b}$, and $y_{2,a}$.

Our identification strategy consists of two parts. First, we identify the distributions of $\alpha_1$, $\eta_a$, and $\varepsilon_{1,a}$. Then, we provide two sets of restrictions under which the remaining distributions are identified: (i) the equality of the distributions of $\varepsilon_{1,a}$, $\varepsilon_{1,b}$, and $\varepsilon_{2,a}$ (Assumption (ref)) and (ii) the equality of the distributions of $\alpha_1$ and $\alpha_2$ and those of $\eta_a$ and $\eta_b$ (Assumption (ref)).

Let us now provide the intuition on how the distributions of $\alpha_1$, $\eta_a$, and $\varepsilon_{1,a}$ can be identified by a repeated application of Kotlarski's lemma. First, for a pair $(1,a), (1,b)$, write

align*[align* omitted — 142 chars of source]

where $M=\alpha_1$, $U_1=\eta_a + \varepsilon_{1,a}$, and $U_2= \eta_b + \varepsilon_{1,b}$. By Kotlarski's lemma, the CF $\phi_{\alpha_1}$ is identified. Similarly, for a pair $(1,a), (2,a)$, write

align*[align* omitted — 170 chars of source]

where $\tilde M=\eta_a$, $\tilde U_1=\alpha_1 + \varepsilon_{1,a}$, and $\tilde U_2= \alpha_2 + \varepsilon_{2,a}$. By Kotlarski's lemma, the CF $\phi_{\eta_a}$ is identified. Joint independence of $\alpha_1,\eta_a$, and $\varepsilon_{1,a}$ implies

align*[align* omitted — 110 chars of source]

identifying the distribution of $\varepsilon_{1,a}$, and hence the distributions of $\varepsilon_{1,b}$ and $\varepsilon_{2,a}$. Finally, joint independence of $\alpha_2,\eta_a,$ and $\varepsilon_{2,a}$ implies

align[align omitted — 132 chars of source]

and joint independence of $\alpha_1,\eta_b,$ and $\varepsilon_{1,b}$ implies

align[align omitted — 130 chars of source]

identifying the distributions of the remaining components $\alpha_2$ and $\eta_b$.

We now state the assumptions needed to make the intuition above rigorous. For a random variable $\zeta$, denote by $\mathcal{Z}_\zeta$ the set of zeros of its CF $\phi_\zeta$ and denote by $\mathcal{Z}_\zeta'$ the set of zeros of the derivative $\phi_\zeta'$ of its CF.

assumption\begin{enumerate}[label=(\roman*)] • $\alpha_1,\alpha_2,\eta_a,\eta_b,\varepsilon_{1,a},\varepsilon_{1,b},\varepsilon_{2,a}$ are integrable with zero means. • $\alpha_1,\alpha_2,\eta_a,\eta_b,\varepsilon_{1,a},\varepsilon_{1,b},\varepsilon_{2,a}$ are jointly independent. \end{enumerate}
assumption\begin{enumerate}[label=(\roman*)] • The sets $\mathcal{Z}_{\alpha_1}, \mathcal{Z}_{\eta_a}, \mathcal{Z}_{\varepsilon_{1,a}}$ are pairwise disjoint. • The sets $\mathcal{Z}_{\alpha_1}$ and $\mathcal{Z}_{\alpha_1}'$ are disjoint. • The sets $\mathcal{Z}_{\eta_a}$ and $\mathcal{Z}_{\eta_a}'$ are disjoint. • The sets $\mathcal{Z}_{\varepsilon_{1,a}}$ and $\mathcal{Z}_{\varepsilon_{1,a}}'$ are disjoint. • The sets $\mathcal{Z}_{\alpha_2}, \mathcal{Z}_{\eta_b}, \mathcal{Z}_{\varepsilon_{1,b}}, \mathcal{Z}_{\varepsilon_{2,a}}$ consist of isolated points. \end{enumerate}

In the identification strategy described above, we apply Kotlarski's lemma in the case when the error terms $U_1,U_2,\tilde U_1,\tilde U_2$ consist of two latent components. The lemma relies on assumptions about these error terms, which are not primitives of our dyadic model. Instead, we want to restrict the latent components in a way that would guarantee that the assumptions on the error terms hold. The following abstract result shows how this can be achieved.

lemLet $A=B+C$, where $B$ is independent of $C$. Assume that \begin{enumerate} • $\mathcal{Z}_B \cap \mathcal{Z}_C = \varnothing$, • $\mathcal{Z}_B \cap \mathcal{Z}_B' = \varnothing$, • $\mathcal{Z}_C \cap \mathcal{Z}_C' = \varnothing$. \end{enumerate} Then $\mathcal{Z}_A \cap \mathcal{Z}_A' = \varnothing$.
proofTake any $t \in \mathcal{Z}_A$. Then either (i) $t \in \mathcal{Z}_B$ or (ii) $t\in \mathcal{Z}_C$. In the case (i), we have $\phi_A'(t) = \phi_B'(t)\phi_C(t) + \phi_B(t) \phi_C'(t) = \phi_B'(t)\phi_C(t)$. By condition 1, $t\notin \mathcal{Z}_C$, and by condition 2, $t \notin \mathcal{Z}_B'$. Therefore, $\phi_A'(t)\neq 0$ and so $t\notin \mathcal{Z}_A'$. Case (ii) is analogous.

We are now ready to state our adaptation of the main theorem of evdokimov2012some to dyadic data.

theoremUnder Assumptions (ref) and (ref), $\phi_{\alpha_1}$ and $\phi_{\eta_a}$ are identified and \begin{align*} \phi_{\varepsilon_{1,a}}(s) = \frac{\phi_{y_{1,a}}(s)}{\phi_{\alpha_1}(s)\phi_{\eta_a}(s)}, \,\,\, s\notin \mathcal{Z}_{\alpha_1} \cup \mathcal{Z}_{\eta_a}. \end{align*}
proofWe use the notations $M,U_1,U_2,\tilde M,\tilde U_1,\tilde U_2$ from the discussion above. Assumptions (ref)(ref), (ref)(ref), (ref)(ref) and Lemma (ref) imply that the zeros of $\phi_{U_1}$ and $\phi_{U_1}'$ are disjoint. Assumption (ref)(ref) implies that the zeros of $\phi_{U_2}$ are isolated. Combining with Assumption (ref) proves Assumption A in evdokimov2012some. Invoking their Lemma 1(b) establishes identification of $\phi_{\alpha_1}$. Similarly, Assumptions (ref)(ref), (ref)(ref), (ref)(ref) and Lemma (ref) imply that the zeros of $\phi_{\tilde U_1}$ and $\phi_{\tilde U_1}'$ are disjoint. Assumption (ref)(ref) implies that the zeros of $\phi_{\tilde U_2}$ are isolated. Combining with Assumption (ref) proves Assumption A in evdokimov2012some. Invoking their Lemma 1(b) establishes identification of $\phi_{\eta_a}$. Finally, independence implies $\phi_{y_{1,a}}(t)=\phi_{\alpha_1}(t) \phi_{\eta_a}(t) \phi_{\varepsilon_{1,a}}(t)$, completing the proof.

Theorem (ref) establishes the distributional identification of $\alpha_1$ and $\eta_a$ and also shows that $\phi_{\varepsilon_{1,a}}$ is identified at all points of the real line except for zeros of $\phi_{\alpha_1}$ or $\phi_{\eta_a}$. We now provide two sets of assumptions under which all the latent distributions are identified (which we call total identification). Both sets of assumptions include the following condition.

assumptionThe sets $\mathcal{Z}_{\alpha_1}$ and $\mathcal{Z}_{\eta_a}$ consist of isolated points.

Let us show that total identification holds under the distributional homogeneity of the three error terms $\varepsilon_{1,a}, \varepsilon_{1,b}, \varepsilon_{2,a}$.

assumption$\varepsilon_{1,a}\overset{d}{=} \varepsilon_{1,b} \overset{d}{=} \varepsilon_{2,a}$.
corollaryUnder Assumptions (ref), (ref), (ref), and (ref), the distributions of all the latent variables $\alpha_1,\alpha_2$, $\eta_a,\eta_b$, $\varepsilon_{1,a}$, $\varepsilon_{1,b}$, and $\varepsilon_{2,a}$ are identified.
proofBy Assumption (ref), $\mathcal{Z}_{\alpha_1} \cup \mathcal{Z}_{\eta_a}$ consists of isolated points, and hence the formula for $\phi_{\varepsilon_{1,a}}(s)$ in Theorem (ref) can be extended to all $s\in\mathbb{R}$ by continuity. This identifies the distribution of $\phi_{\varepsilon_{1,a}}$, and, in view of Assumption (ref), the distributions of $\varepsilon_{1,b}$ and $\varepsilon_{2,a}$. The formulas (ref) and (ref) then identify the distributions of $\alpha_2$ and $\eta_b$.

Next, we show that total identification holds under the distributional homogeneity of the random effects, $\alpha_1, \alpha_2$ and $\eta_a, \eta_b$.

assumption$\alpha_1 \overset{d}{=} \alpha_2$ and $\eta_a \overset{d}{=} \eta_b$.
corollaryUnder Assumptions (ref), (ref), and (ref), the distributions of all the latent variables $\alpha_1,\alpha_2$, $\eta_a,\eta_b$, $\varepsilon_{1,a}$, $\varepsilon_{1,b}$, and $\varepsilon_{2,a}$ are identified.\footnote{Notice that Assumption (ref) is implied by Assumptions (ref)(v) and (ref).}
proofBy Theorem (ref), $\phi_{\alpha_1}$ and $\phi_{\eta_a}$ are identified everywhere. By Assumption (ref), $\phi_{\alpha_1}=\phi_{\alpha_2}:=\phi_\alpha$ and $\phi_{\eta_a}=\phi_{\eta_b}:=\phi_\eta$. Finally, joint independence yields \begin{align*} \phi_{\varepsilon_{1,b}} = \frac{\phi_{y_{1,b}}(s)}{\phi_{\alpha}(s) \phi_{\eta}(s)}, \quad s\notin \mathcal{Z}_{\alpha} \cup \mathcal{Z}_{\eta}, \\ \phi_{\varepsilon_{2,a}} = \frac{\phi_{y_{2,a}}(s)}{\phi_{\alpha}(s) \phi_{\eta}(s)}, \quad s\notin \mathcal{Z}_{\alpha} \cup \mathcal{Z}_{\eta}. \end{align*} By Assumption (ref), $\mathcal{Z}_{\alpha} \cup \mathcal{Z}_{\eta}$ consists of isolated points, and hence the formulas above can be extended by continuity. This identifies the distributions of $\varepsilon_{1,b}$ and $\varepsilon_{2,a}$.

Case of fully connected network

When the bipartite network on the node sets $\{1,2\}$ and $\{a,b\}$ is fully connected, i.e., when all the node pairs $(1,a),(1,b),(2,a)$, and $(2,b)$ are linked, the distributions of all the latent variables $\alpha_1,\alpha_2$, $\eta_a,\eta_b$, $\varepsilon_{1,a}$, $\varepsilon_{1,b}$ , $\varepsilon_{2,a}$, and $\varepsilon_{2,b}$ are identified without any assumptions on the distribution homogeneity (cf. Assumptions (ref) and (ref)).\footnote{For this, Assumption (ref) has to be extended to include the conditions on $\varepsilon_{2,b}$.} To see that, notice that applying Kotlarski's lemma to the pair $y_{i,a}, y_{i,b}$ identifies the distribution of $\alpha_i$, $i=1,2$. Then applying the lemma to the pair $y_{1,c}, y_{2,c}$ identifies the distribution of $\eta_c$, $c=a,b$. Finally, the distribution of $\varepsilon_{i,c}$ is identified via

align*[align* omitted — 134 chars of source]

Formulating rigorous conditions under which this identification strategy is valid can be done along the lines of Section (ref).

Conclusion

We show how the classical lemma of Kotlarski can be employed to identify distributions of latent components in dyadic models for bipartite networks with two-way random effects. When the bipartite graph is partially linked, we provide two sets of assumptions under which all the latent distributions are identified. The first set of assumptions restricts the errors to be identically distributed. The second set of assumptions restricts the random effects to be identically distributed. Our identification result may be applicable to dyadic regressions with random or correlated effects, as discussed in Section 2.3 of bonhomme2020econometric. It can also be used to develop an estimation procedure for this class of models. This is beyond the scope of the current paper, and we defer this agenda to future work.