EconBase
← Back to paper

Estimating Peer Effects Using Partial Network Data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

35,198 characters · 7 sections · 4 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Editor

Dear Peter (if we may),

We want to thank you for the opportunity to revise our manuscript a second time.

We have worked hard in order to address the remaining comments raised by the three referees as well as by yourself. We believe that we were able to address satisfactorily all of the comments raised by all the referees. For completeness, you can find detailed responses to all of the referees attached to this letter.

With respect to the specific points raised in your letter:

itemize• R1 (Consistency, Definition 1 and Theorem 1) We made the suggested change to Definition 1. We also rewrote the proof of Lemma 3 in order to be more transparent. In particular, Lemma 3 did not use Definition 1, as noted by R1. This is because we used the stronger Assumptions 5 and 9 (which we need for asymptotic normality). We modified the proof following the suggestion of R1, making explicit the use of Definition 1. • R1 (SGMM vs 2SLS) R1 suggested to use a linear GMM instead of our non-linear SGMM. From our understanding of their comment, we believe that they are mistaken. We have rewritten the discussion preceding Theorem 1 in order to improve clarity. We also sketched a proof of the inconsistency of the linear GMM approach suggested by R1 in our letter to them. (This is a slight modification from our Proposition 2 in the Online Appendix C.1.) • R2 (Comment 1 on graham2017econometric) We believe that there was a confusion about the estimator in graham2017econometric, which is not based on the marginal distribution (integrated over the degree heterogeneity), but is rather based on the conditional distribution (for a specific degree sequence). We have now (also following R3's comment) rewritten this part of the Online Appendix. We are more precise in our discussions of these alternative network formation processes and are explicit about how to 1) estimate $\boldsymbol{\rho}$ and 2) compute $\hat{P}(\mathbf{A}|\boldsymbol{\rho})$. In particular, the partial information $\mathcal{A}$ is formally defined in each case. This also echos R1's comment on the examples. • R2 (Comment 2 on graham2017econometric and boucher2017my) Again, there seems to be confusion about these models. Our objective is to show that these models can be estimated with partial network data. As mentioned above, we have rewritten this section and are now much more precise and we formally define $\mathcal{A}$ in each case. • R2 (Reorganizing Section 2.2) This is a good suggestion. We did as suggested.

We also restructured and updated the literature review and made sure that the revision remains in agreement with the journal formatting guidelines. We restructured part of the Online Appendix in order to make it more usable to applied researchers wanting to apply the methodology to cases not covered in the main text.

We hope this new version will be to your (as well as the referees') liking.

Sincerely,

Vincent and Aristide.

Referee 1

Dear Referee,

We would like to thank you for your detailed comments. We have worked hard to address them. We hope that the revised version of the manuscript will be to your liking.

Below, we present detailed answers to each of your comments and questions.

Sincerely,

The authors.

Comments from Referee 1

\noindentAbout the Simulated GMM

itemize• (Consistency in Definition 1.) I have two questions. First, I believe the proper definition should be convergence in probability uniformly over the support of $A_m$, rather than the sample size m, as is usually the case in two-step estimators. Could the authors please confirm this (or, explain why the supremum should be over the sample size m)? Second, is such consistency implicitly assumed in Theorem 1? It is not stated as a condition in Theorem 1; nor is it established in the proof of Theorem 1. Yet I am sure it has to play a role in showing the uniform convergence of the objective function. [Answer] You are perfectly right: the convergence in probability should be uniformly over the support of $A_m$. Thank you for catching this. In the revised version, we made the appropriate changes to the definition, so the $\sup$ is over $(\mathbf{A}_m,m)$. We also rewrote the proof of Lemma 3 in order to clarify where the assumption is used. Specifically, you are right that in the previous version, we did not refer to Definition 1 in the proof of Lemma 3. This is because Lemma 3 uses (or used) Assumption 9 (derivatives of $\hat{P}$ with respect to $\boldsymbol{\rho}$ are uniformly bounded). Applying the mean value theorem to the difference $\hat{P}(\mathbf{A}_m \mid \hat{\boldsymbol{\rho}}, \mathbf{X}_m, \kappa(\mathcal{A}_m)) - P(\mathbf{A}_m \mid \boldsymbol{\rho}_0, \mathbf{X}_m, \kappa(\mathcal{A}_m))$ yields the needed condition. This is why Definition 1 did not explicitly appear in the original proof. We now provide a proof of Lemma 3 that remains valid even if Assumption 9 fails, by relying on Definition 1. • (Examples.) These examples are helpful, as they illustrate the applicability of the method. Two suggestions here. First, please write out the precise form of $P (A_m |X_m , A_m )$ in each example. It is especially important to be clear about what this probability involve any additional parameters such as misclassification or censoring probability, in addition to the network formation probability itself. Second, explain how consistent estimation of $\rho$, the parameters in network formation, can be done given these forms of $P(A_m |X_m , A_m )$. The discussions may be brief and informal, but they should at least touch on the nature of asymptotic experiment and the probabilistic/statistic tools used for showing convergence. [Answer] \emph{Thank you very much for this comment. We have tried to address your comment as best as possible given the space constraint.} \emph{In particular, we added the computation of estimator of the distribution of the true network to Example 3 (the only one for which it was missing)}. \emph{With respect to the statistical tools needed to estimate $\boldsymbol{\rho}$, we chose to discuss this jointly for all examples, following the statement of Assumption 5. See Footnote 12. In particular, note that our Assumption 1 (many groups asymptotic framework) greatly simplifies the estimation. Indeed, LLN and CLT for independent (but not identically distributed) random variables are sufficient, along with standard regularity conditions. The substantial assumption is identification. We are now more explicit on this.} • (Theorem 1 and discussions on page 15-17.) I thank the authors for revising these parts to respond to previous comments. I do have a few followup questions. To be specific, it helps to recap the SGMM idea below. [\emph{omitted}] Given this recap, I have the following specific comments/requests, some of which are only reiterations of earlier comments 1 and 2 in the last report. \begin{itemize} • First, the identification condition, i.e., $\theta_0$ is the unique minimizer of (2), should be presented and included in the text. Assumption 11 in the appendix is related, but not quite illustrative, especially because $\boldsymbol{\rho}$ now enters $m(y, X, G; \theta)$ nonlinearly through a matrix inverse. I understand it might not be possible (nor is it necessary) to present exact mathematical conditions for point identification here; however, at the very minimum, standard arguments such as the order conditions and no immediate failure of rank conditions should be included with some degree or rigor. Relatedly, Lemma 1 in A.1, as a step for showing identification, is also not sufficient. We also need to show that these conditions do not hold for any other $\theta\neq\theta_0$. I would recommend the authors include both results in one identification lemma and present it in the text as a prelude to Theorem 1. \emph{\textbf{[Answer]}} \emph{Thank you very much for this comment. We now present the identification condition (assumption) directly the statement of the Theorem.} \emph{Specifically, we now define the matrices of simulated instruments and explanatory variables, as well as the empirical moment condition in the text, before the formal statement of the theorem. This allows us to include the identification condition in the statement of the theorem.} \emph{In particular, in order to avoid confusion, it is stated as: for any $\theta\neq \theta_0$, the (asymptotic) moment condition fails. In Footnote 19 (P. 17), we state that Lemma 1 of the Appendix A shows that the asymptotic moment condition is verified for $\boldsymbol{\theta}_0$.} \emph{We also discuss, in Footnote 20 that the number of instruments needs to be larger than the number of explanatory variables in order to avoid an immediate failure of the rank condition.} • Second, I am still puzzled as to why SGMM needs to make three independent sets of draws of $A^{(s)}$ to work, despite the authors' edits on page 15-16. It appears (in the last display on page 15) that the concern was about some “bias” when the same simulated network $G^{(s)}$ is used in construction of $Z$ and $\varepsilon(y, X, G; \theta) \equiv (I-\alpha G)^{-1} y - V \theta$. But I do not think that is a cause for concern about the failure of the moment condition at the true parameter $\theta_0$ . Recall the original moment condition in (1) holds at the true $\theta_0$ even when the same $G$ enters $Z$ and $\varepsilon(y, X, G; \theta_0)$. In fact, this moment condition holds at $\theta_0$ even if we were to further condition on $A$ in addition to $(A, X)$. Therefore, I do not see why using the same $G^{(s)}$ leads to any bias for simulated GMM. \emph{\textbf{[Answer]}} \emph{Thank you for this comment. We apologize for the confusion. The same draw for the instrument and the explanatory variables can be used, but precision is much lower. In our letter to you from the previous round, we included the sketch of a proof showing that the variance of the moment condition is larger. In simulation, the precision of the estimator is much better with different draws. We have included Footnote 16 to clarify.} \emph{We unfortunately cannot have fewer than two set of draws. Consider the simple case $y=x\beta + \alpha Gy + \varepsilon$. If one replaces $G$ with some draw $\dot{G}$, we obtain: $y=x\beta + \alpha\dot{G}y + [\varepsilon + \alpha(G-\dot{G})y]$, where the term in bracket is the new “error” term which includes $\varepsilon$ as well as the approximation error. In particular, note that $y$ in the approximation error is a function of $G$ (the true $G$). So the approximation error is $\alpha(G-\dot{G})y(G)$. Thus, we need an instrument that is orthogonal with this approximation error. In general, $\dot{G}x$ (or $\ddot{G}x$) won't do because $E[x'\dot{G}'Gy(G)|\cdot]\neq E[x'\dot{G}'\dot{G}y(G)|\cdot]$.} \emph{In this revised version, we emphasized that the bias is due to the fact that $\mathbf{y}_m$ is a function of $\mathbf{G}_m$ (see P.16). We also added Footnote 17 (P. 16) in order to direct the interested reader to Proposition 2 of the Online Appendix C.1. The proposition studies a simple linear GMM (see our answer to your comment on 2SLS vs non-linear GMM below) for the case where $\mathbf{G}_m\mathbf{X}_m$ is observed. We show that the asymptotic bias is: 1) non-zero and 2) minimized when the weighting matrix $\mathbf{W}_M=\mathbf{I}_M$.} • Third, the use of multiple sets of independent draws of $G^{(s)}$ , $G^{(r)}$ in $Z$ and $\varepsilon(y, X, G; \theta)$ in the calculation of $\bar{m}(\theta; \hat{\rho})$ in fact raises questions about the uniform convergence of $\bar{m}(·; \hat{\rho})$ to $m^* (·)$, as this practice removes one source of correlation between $Z$ and $\varepsilon(y, X, G; \theta)$ (through the actual $A$) conditional on $(A, X)$, which does exist in the actual conditional moment $Z'\varepsilon(y, X, G; \theta)$. One needs to show such imposed difference does not affect the uniform convergence of $\bar{m}(·; \hat{\rho})$. (The proof of Lemma 3 in Appendix B does not appear to explicitly take into account this subtle difference due to the use of three sets of simulated draws.) On a high level, the uniform convergence of $\bar{m}(\cdot; \hat{\rho})$ (which uses estimates of $\rho$ in network formation) to $m^*(\cdot)$ over all $\theta$ is a central piece for the theoretical foundation of the SGMM method. I would recommend the authors include a proposition in the text as a preclude to Theorem 1, and be specific with how the concept of consistency in Definition 1 is used or established in that proposition. \emph{\textbf{[Answer]}} \emph{Thank you for raising this point. The uniform convergence of $\bar{m}(\theta; \hat{\rho})$ to $m^{\ast}(\theta)$ is indeed a key component in the proof of the consistency of the SGMM estimator. However, in our proof, we establish a similar result that $\bar{m}(\theta; \hat{\rho}) - \mathbb{E}[\bar{m}(\theta; \rho_0)]$ converges uniformly to zero, which can also be used to show the consistency of the SGMM estimator. To obtain this result, we rely on the intermediary Lemmas 3a and 3b. In particular, Lemma 3a shows that $\mathbb{E}[\bar{m}(\theta; \hat{\rho})] - \mathbb{E}[\bar{m}(\theta; \rho_0)]$ converges uniformly to zero, which is the central step in the proof. If uniform convergence holds for the expected moments, then it also holds for the sample moments by the uniform law of large numbers, as established in Lemma 3b. Taking the expectation of the moment function simplifies the proof and shows that uniform convergence is not affected by the multiple draws.}\\ \emph{The first step in the proof of Lemma 3a is to establish uniform convergence with conditional expectations. Note that $\mathbb{E}[\bar{m}(\theta; \hat{\rho}) \mid \hat{\boldsymbol{\rho}}, \mathbf{X}_m, \kappa(\mathcal{A}_m)]$ can be written as:} $$\sum_{\dot A, \ddot A, \dddot A} \mathbf B(\dot A, \ddot A, \dddot A, \mathbf{X}) \hat P( \dot A\mid \hat{\boldsymbol{\rho}}, \mathbf{X}, \kappa(\mathcal{A}))\hat P(\ddot A\mid \hat{\boldsymbol{\rho}}, \mathbf{X}, \kappa(\mathcal{A}))\hat P(\dddot A \mid \hat{\boldsymbol{\rho}}, \mathbf{X}, \kappa(\mathcal{A})),$$ \emph{where the sum is taken over the support of the draws, and $\mathbf{B}(\dot{A}, \ddot{A}, {\mathop{\kern\z@A}\limits^{\makebox[0pt][c]{\vbox to-1.5\ex@{\kern-\tw@\ex@ \hbox{\normalfont\scalebox{.95}{.}\kern-0.5pt\scalebox{.95}{.}\kern-0.5pt\scalebox{.95}{.}}\vss}}}}, \mathbf{X})$ denotes the conditional expectation of the moment function—that is, the term with $\boldsymbol{\varepsilon}_m$ in the moment function disappears. Similarly, $\mathbb{E}[\bar{m}(\theta; \hat{\rho}) \mid \boldsymbol{\rho}, \mathbf{X}_m, \kappa(\mathcal{A}_m)]$ has the same expression, except that $\hat{P}(\cdot \mid \hat{\boldsymbol{\rho}}, \mathbf{X}_m, \kappa(\mathcal{A}_m))$ is replaced by the true distribution $P(\cdot \mid \boldsymbol{\rho}, \mathbf{X}_m, \kappa(\mathcal{A}_m))$.\\ Consequently, uniform convergence is guaranteed if $\hat{P}(\dot{A} \mid \hat{\boldsymbol{\rho}}, \mathbf{X}_m, \kappa(\mathcal{A}_m)) - P(\dot{A} \mid \boldsymbol{\rho}, \mathbf{X}_m, \kappa(\mathcal{A}_m))$ is sufficiently close to zero over the support of $\dot{A}$ (and similarly for $\ddot{A}$ and $ {\mathop{\kern\z@A}\limits^{\makebox[0pt][c]{\vbox to-1.5\ex@{\kern-\tw@\ex@ \hbox{\normalfont\scalebox{.95}{.}\kern-0.5pt\scalebox{.95}{.}\kern-0.5pt\scalebox{.95}{.}}\vss}}}}$), and if $\mathbf{B}(\dot{A}, \ddot{A}, {\mathop{\kern\z@A}\limits^{\makebox[0pt][c]{\vbox to-1.5\ex@{\kern-\tw@\ex@ \hbox{\normalfont\scalebox{.95}{.}\kern-0.5pt\scalebox{.95}{.}\kern-0.5pt\scalebox{.95}{.}}\vss}}}})$ is bounded, which holds under Assumptions 1 and 7. Importantly, this does not depend on the number of draws, as they are simulated from the same distribution. Moreover, the probabilities of the draws appear only as products in the expression of $\mathbb{E}[\bar{m}(\theta; \hat{\rho}) \mid \hat{\boldsymbol{\rho}}, \mathbf{X}_m, \kappa(\mathcal{A}_m)]$. Therefore, it suffices to impose the condition stated in Definition 1.} \emph{We hope that these additional details adequately address your comment. At this stage, we do not feel that the above discussion requires any modification to the proof of Lemma 3. But of course, we could add more details if you strongly feel otherwise.} • Fourth, in the previous report, I asked whether the idea of using simulated networks can be applied to a version of 2SLS as in by Bramoullé et al (2009). The authors claimed in the response that this would not work, but I can not find any details in support of this claim in the updated Section 3. Specifically, my question was why one can not use $\varepsilon(y, X, G; \theta) = y-c-X\beta-\alpha Gy -GX\gamma$ to construct the moments, instead of $\varepsilon(y, X, G; \theta) = (I -\alpha G)^{-1} y - V \theta$. As I tried to explain, the former choice would at least make the identification condition very transparent (just standard rank condition in 2SLS). I can not find answers on page 12 of the authors' response, or in the revised text. If there is no fundamental reason to rule out a SGMM based on 2SLS, I’d recommend including it in the paper for comparison, as it is likely to be easier to compute (especially in the just-identified case). \emph{\textbf{[Answer]}} \emph{Thank you for this comment. We believe it is linked to your other comment regarding the use of many draws. Our apologies if we misunderstood your comment.} \emph{Ideally, of course, we would prefer a linear GMM (simulated or not) since in particular identification conditions are much more transparent. (It is also numerically attractive.) The reason we \emph{need} a non-linear moment condition is because of the approximation error when replacing $Gy$ with $\dot{G}y$.} \emph{It is very hard (we could not) to find a valid instrument for $\ddot{\boldsymbol{\varepsilon}}_m$ (that would lead to a linear GMM moment function). Our strategy is to use another draw in order to remove the asymptotic bias. Unfortunately, this is what leads to a non-linear moment function.} \emph{As we discussed above (and unless we misunderstood your suggestion) a simple linear GMM (such as a 2SLS) would not work. This is covered by our Proposition 2 in the Online Appendix C.1. We sketch the argument below.} \emph{For simplicity, assume that $\boldsymbol{\gamma}=\mathbf{0}$ (no contextual effects) so that the observation or not of $\mathbf{GX}$ is inconsequential. Let $\ddot{\mathbf{V}}_m$ and $\dot{\mathbf{Z}}_m$ be the simulated matrices of explanatory variables and instruments obtained by simply replacing the true (unobserved) interaction matrix $\mathbf{G}_m$ by their simulated counterparts $\ddot{\mathbf{G}}_m$ and $\dot{\mathbf{G}}_m$. (Note that the argument is the same if $\dot{\mathbf{G}}_m=\ddot{\mathbf{G}}_m$.)} \emph{Consider the linear moment function $\frac{1}{M} \sum_{m = 1}^M \dot{\mathbf{Z}}'_m (\mathbf y_m - \ddot{\mathbf{V}}_m\ddot{\boldsymbol{\theta}})$ and let $\check{\boldsymbol{\theta}}$ be the associated GMM estimator of $\ddot{\boldsymbol{\theta}}$.} \emph{Finally, define the \emph{sensitivity matrix}: $${\mathbf{R}}_M=\left(\frac{\sum_m\ddot{\mathbf{V}}_m'\dot{\mathbf{Z}}_m}{M}\mathbf{W}_M\frac{\sum_m \dot{\mathbf{Z}}_m'\ddot{\mathbf{V}}_m}{M}\right)^{-1}\frac{\sum_m\ddot{\mathbf{V}}_m'\dot{\mathbf{Z}}_m}{M}\mathbf{W}_M.$$} \emph{Then, we can show that the asymptotic bias for $\check{\boldsymbol{\theta}}$ is given by $$\alpha_0 \operatorname{plim}\frac{\mathbf{R}_M\sum_m \dot{\mathbf{Z}}_m'({\mathbf{G}}_m-\ddot{\mathbf{G}}_m)\mathbf{y}_m}{M}.$$ (We can also show that the Frobenius norm of $\mathbf{R}_m$ is minimized when $\mathbf{W}_m=\mathbf{I}_m$.)} \emph{Note that because $\mathbf{y}_m$ is a function of the true $\mathbf{G}_m$, the two quantities: $\operatorname{plim}\frac{\mathbf{R}_M\sum_m \dot{\mathbf{Z}}_m'{\mathbf{G}}_m\mathbf{y}_m}{M}$ and $\operatorname{plim}\frac{\mathbf{R}_M\sum_m \dot{\mathbf{Z}}_m'\ddot{\mathbf{G}}_m\mathbf{y}_m}{M}$ are in general different so the asymptotic bias is non-zero.} \end{itemize}

\noindentAbout the Bayesian Estimator

itemize• (Response to Comment 1.) Thank you for the explanation here. It is helpful. I think one way to formalize this is to say that the posterior of $(\rho, \theta)$ is proportional to the product of $L(y|A, X, \theta, \rho)$,$ L(A|X, \rho)$, and $\pi(\theta, \rho)$, where $\rho$ enters in both conditional likelihoods. So, not only does the MCMC algorithm needs to lump both $\theta,\rho$ in the iterations, but also because there is more to learn about $\rho$ from the first conditional likelihood $L(y|A, X, \theta,\rho)$ as well. This is a source of efficiency gains. [Answer] Thank you. We fully agree. We added this precision in Footnote 31. • (Response to Comment 3.) Again, thank the authors for making edits regarding the computational complexity in the Bayesian approach; the two points they mentioned in the response also makes sense. I have a followup suggestion: It would be helpful to report how the convergence in MCMC is affected by the group size $n_m$ . For instance, report the length of “burn-in” period before MCMC reaches stationarity for different values of $n_m$ . This would give practitioners a clear sense of the computational costs involved for implementing the Bayesian method. [Answer] Thank you for this comment. We conduct a simulation study considering the case of missing links, as in our SGMM estimator (see Figure F.1 in Online Appendix F). The results indicate that the burn-in period increases with $N_m$, the number of nodes in the network, and especially with the number of missing entries in the network. However, even with $75\%$ missing entries, the MCMC converges quickly (before 2{,}000 iterations) for $N_m = 50$ and $M = 50$, and faster when $N_m = 30$ and $M = 50$.

Referee 2

Dear Referee,

We would like to thank you for this second round of detailed comments. We have worked hard to address you comments and we hope that the revised version of the manuscript will be to your liking.

Below, we present detailed answers to each of your comments and questions.

Sincerely,

The authors.

Comments from Referee 2

\noindentMajor Comments

enumerate• In my previous report (Comment 3 Network formation process), I pointed out to the authors that the network formation model considered by Graham (2017) includes unobserved degree heterogeneity. In their response, the authors wrote: “However, the degree distribution is a sufficient statistics and so his model can be used in our setting provided the degree distribution is observed.” I understand that the network formation in equation (2) of the current version is an integrated version of that in Graham (2017) with respect to the degree heterogeneity. However, I am not fully convinced by their argument. The reason is that the sufficient statistics for the degree heterogeneity is not the degree of each individual, but rather the degree sequence of all individuals (see page 1040 of Graham 2017). Therefore, to properly integrate out the degree heterogeneity, we need the joint distribution of a random vector (i.e., the degree sequence), rather than simply the degree distribution of a scalar random variable (i.e., the degree). In footnote 11, the authors claim that “the degree distribution can be obtained from survey question by simply asking individuals about the number of links they have.” However, this only gives the degree distribution, not the distribution of the degree sequence. [Answer] Thank you for this comment. We apologise for the confusion. We are not proposing to integrate out the degree heterogeneity. That would indeed require to know the distribution. What we propose instead (and as in Graham (2017) p. 1040) is to condition on the degree sequence of the true network. That is, the degree sequence is in $\mathcal{A}$. In other words, we propose to use Equation (3) in Graham (2017) in order to estimate $\boldsymbol{\rho}$. To avoid confusion, we rewrote Section 2.2 following your suggestion below. We also completely rewrote that part of the online appendix (Appendix B). It is now designed to be used by applied researchers wanting to apply our methodology to more general settings. As such, we are more precise about what information is required to estimate $\boldsymbol{\rho}$ as well as the estimate of the network formation process. In particular, we now formally define $\mathcal{A}$ in each example. • Related to comment 1, in the revised manuscript (page 10) and Appendix D, the authors provided detailed discussions of the network formation models by Graham (2017) and Boucher and Mourifi\'e (2017), as well as how to consistently estimate these two models. In my view, these two models serve as motivating examples for the network formation model in equation (2). While, the discussions (especially those in Appendix D) regarding the estimation of these two models seem to require correctly observing the entire network and might be irrelevant in the current context. Thus, the purpose of these two models is not clearly stated and can be confusing. I suggest that the authors restructure section 2.2 and remove irrelevant discussions. One possible way is to divide section 2.2 into two sections: section 2.2 introduces the true network formation model (i.e. equation 2), along with its motivating examples of Graham (2017) and Boucher and Mourifie (2017); then a new section 2.3 presents Assumption 5, the definition of consistent estimator of true network distribution, and the three examples for Assumption 5. \emph{\textbf{[Answer]}} \emph{Thank you. Again, we apologise for the confusion. Our main point is that the models in Graham (2017) and Boucher and Mourifi\'e (2017) \emph{can} be estimated with partial information. As mention above, we rewrote this part and are now explicit about which information about the network structure is required.} \emph{We also reorganized Section 2.2 as suggested. Thank you for the suggestion.} • The discussion of Definition 1 can be improved. First, the discussion immediately after Definition 1, regarding why the partial information $\mathcal{A_m}$ is used twice, is not clear until the more detailed explanations given in Examples 1--3 are provided. Second, the connection between equation (3) and its simplified solution, which disregards the information in $\mathcal{A}_m$, is not clearly established in relation to Examples 1--3. For instance, it is not clear which of the three examples utilizes equation (3) and which relies on the simplified solution. \emph{\textbf{[Answer]}}\emph{Thank you. We rewrote this discussion in order to make it clearer. In particular, for our Examples 1--3, we suggest to use Bayes rule and it is easy to compute. Also, following a comment from R1, we now formally define the estimator of the network formation process for each example.} • In the reply to my comment 6, the authors stated that “missing and misclassification rates can depend on network statistics”. In Example 3 (the case of misclassification), they claimed that the true network distribution can be consistently estimated using the method of Hausman et al (1998). However, if I understand correctly, Hausman et al (1998) assumes constant misclassification rates. In the simulation experiments in Section 3.1, the misclassification is also at random (I guess it means with constant misclassification rates). So I wonder if the authors could provide further explanations/references on this point. \emph{\textbf{[Answer]}}\emph{Our apology for the confusion. The method in Hausman et al (1998) assumes constant misclassification. This is also what we assume in Section 3.1. However, our methodology can be adapted to cases where the misclassification can depend on network statistics (but not using Hausman et al (1998)). This is the case for example of “top coding” which we implement in our application. }

\noindentOther Comments

enumerate• I suggest that the authors move the explanation regarding how to simulate $\mathbf{G}_m$ as a Bernoulli variable (the remark at the bottom of page 17) to page 15, before introducing the moment function. [Answer]Thank you. We did as suggested. • If I understand correctly, the moment function presented in Theorem 1 is obtained from the conditional moment $E[\dot{\mathbf{Z}}'_m (\ddot{\varepsilon}_m -\delta_m ) | \mathbf{A}_m, \mathcal{A}_m, \mathbf{X}_m ]$. If so, it would be helpful to add this as an intermediate step before presenting Theorem 1. [Answer]Thank you. Yes, we added the precision. • The authors claimed repeatedly that the partial network data results in a “first-order” bias in Add Health study. What is the meaning of the “first-order” here? Does it mean that the bias only appears in the moment condition (first-order), not in the second-order asymptotics? \emph{\textbf{[Answer]}}\emph{Our apology for the confusion. We used “first order” in the sense of “very important” or “large” and not the sense of a first-order Taylor expansion. We changed the manuscript and removed the use of “first-order” when not referring to Taylor approximations. } • In Section 3.1, how many groups are generated in the simulations? What are the values for R, S, and T for the simulated networks? \emph{\textbf{[Answer]}}\emph{Thank you. We added the following precision to the note to Figure 1: we simulate data for 100 groups of 30 individuals. We also set $R=100$, $S=1$, and $T=1$.} • There are still quite a few typos in the revised manuscript. Below, I point out some of them, and I suggest the authors conduct a thorough check throughout the paper. \emph{\textbf{[Answer]}}\emph{ Thank you, we made the appropriate changes.}

Referee 3

Dear Referee,

We would like to thank you for your comments. We hope that the revised version of the manuscript will be to your liking.

Below, we present detailed answers to each of your comments and questions.

Sincerely,

The authors.

Comments from Referee 3

\noindentComments

enumerate• The assumption that one can take draws from a consistent estimate of the distribution of networks is a quite strong one, which the authors make clear. As I understand it, this requires a consistent estimate of the distribution of networks, not just features of networks. Their assumption of conditional link independence does a lot of work in the actual estimation of the network, however, which is not 100% clear at the outset. [Answer] Thank you. Indeed, conditional link independence is extremely helpful, although not formally necessary. Still, most of our examples have conditionally independent links. We added the precision in the second paragraph of the introduction. • While their method is more general, the examples that they give have the form of estimators that rely on data augmentation, which generally require some missing (conditionally) at random assumption for validity. While this isn't their main point (which assumes consistent estimation of the distribution of networks), to actually implement this would require something like that. This is also related to the assumption needed in Griffith (2022) and related work. A bit more here would be helpful for applied people who want to implement their methods. [Answer] Thank you. We added Footnote 2 in order to emphasize this point. • I understand that it was cut in the interest of space, but I think both the discussion of ARD and also how edgewise sampling in Conley and Udry (2010) led to good estimation properties, were nice ideas that showed how a consistent estimate could be obtained with different types of network data (not just noisy data on which edges exist). I'd probably add a short discussion of these, or at least a more clear direction to the online appendix. \emph{\textbf{[Answer]}} \emph{We agree. In the revised version, we rewrote the Online Appendix in order to make it more suitable to applied researchers. We discuss alternative network formation process in details and added a section releated to survey design in which we briefly discuss the idea in Conely and Udry (2010) } • This is mentioned at scattered points, but I think the authors should make clear up front that they are assuming network exogeneity. It shows up in Assumption 4, but seems that it should be mentioned in the intro since it's so important and makes clear what this paper does not do. This contrasts the paper with many works, including Goldsmith-Pinkham and Imbens (2013, JBES), Hsieh and Lee (2016, J Applied Metrics), Graham (2017, Ecma), Griffith (2022, REStat), Johnsson and Moon (2018, REStat), and Auerbach (2022, Ecma). \emph{\textbf{[Answer]}} \emph{Indeed, we added the precision in the third paragraph of the introduction. } • I also think the authors should more clearly note in the intro that they need some network data, which contrasts the paper with Manresa (2016), de Paula et al. (2024), and Griffith and Peng (2023) which require no network data. \emph{\textbf{[Answer]}} \emph{We added the precision in the second paragraph of the introduction. } • The formatting of Figure 1 is confusing. I appreciate the table notes, but the “Left: ...” and “Right: ...” lines took me some time to figure out. Maybe format as sub-figures rather than “Left: ...” and “Right: ...” ? \emph{\textbf{[Answer]}} \emph{Thank you, we now created two subfigures, which are referred to independently in the text.}

\cleardoublepage