EconBase
← Back to paper

Treatment Effect Models with Strategic Interaction in Treatment Decisions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

115,702 characters · 18 sections · 2 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Treatment Effect Models with Strategic Interaction in Treatment Decisions

\onehalfspacing

abstractThis study considers treatment effect models in which others' treatment decisions can affect both one's own treatment and outcome. Focusing on the case of two-player interactions, we formulate treatment decision behavior as a complete information game with multiple equilibria. Using a latent index framework and assuming a stochastic equilibrium selection, we prove that the marginal treatment effect from one's own treatment and that from the partner are identifiable on the conditional supports of certain threshold variables determined through the game model. Based on our constructive identification results, we propose a two-step semiparametric procedure for estimating the marginal treatment effects using series approximation. We show that the proposed estimator is uniformly consistent and asymptotically normally distributed. As an empirical illustration, we investigate the impacts of risky behaviors on adolescents' academic performance. Keywords: binary games; latent index models; marginal treatment effects; series estimation; strategic interaction. JEL Classification: C14, C31, C57.

Introduction

Estimating the marginal treatment effects (MTE) is essential in treatment evaluation. The MTE can provide rich information on how treatment effects vary across economic agents in terms of their observed and unobserved characteristics. Furthermore, after estimating the MTE, researchers can identify many treatment parameters of interest, such as the average treatment effects (ATE), local ATE (LATE), and policy-relevant treatment effects (PRTE) as some weighted averages of the MTE (\citealpmain{heckman1999local, heckman2005structural}). Prior studies clearly demonstrate the usefulness of MTE methods in various empirical fields (e.g., \citealpmain{basu2007use}; \citealpmain{carneiro2011estimating}; \citealpmain{cornelissen2016late}; \citealpmain{felfe2018does}).

An important but often neglected issue in studies of treatment effects is the presence of potential interference between agents. For example, when evaluating the effect of smoking behavior on health outcomes for couples, one partner's smoking behavior would affect the health outcome of the other. That is, the stable unit treatment value assumption (SUTVA) does not hold due to the “treatment spillover”. In addition, it is natural to expect that one partner's smoking behavior interacts “strategically” with the other partner's smoking behavior, in the sense that one's action can directly affect the utility of smoking of the other, and vice versa. This example suggests the presence of two different types of interference that need to be addressed: (i) treatment spillover and (ii) strategic interaction in the treatment decisions.

In reality, the co-existence of treatment spillover and strategic interaction should be common. \phantomsection\Copy{AE-2}{ Here, we provide three empirical examples that would fit into our framework. First, many previous studies empirically demonstrate that strategic/social interaction between friends is a primary cause of delinquency. It is also often observed that delinquent activities among close friends have significant impacts on students' academic performance. Understanding such spillover effects, as well as direct effects, is important in the education literature. Second, consider two competing airlines deciding whether to introduce a direct flight between a given pair of cities. Then, it would be interesting to investigate to what extent the number of passengers using one airline's transfer flight to travel between that city pair drops when the other introduces a direct flight. Third, suppose there are two candidates in a mayoral election. The outcome of interest is the vote share of each candidate. Here, we would be interested in the effect of the candidates' political position (e.g., pro-choice or pro-life) or their main campaign strategy (e.g., online or grassroots) on their outcomes. }

This study aims to develop identification and estimation procedures for MTE models that allow both treatment spillover and strategic interaction in the treatment decisions. \phantomsection\Copy{AE-3-1}{ In particular, we focus on an empirically relevant setup in which interactions occur only within a pair of agents (e.g., couples, best friends, duopoly firms, or incumbent and challenger), and there are no treatment spillovers across the pairs. We postulate that they decide their treatment status simultaneously in a binary game of complete information. Within this framework, we formulate a set of sufficient conditions under which it is possible to point-identify the MTE parameters of interest. Note that if we consider each pair of players as a single observation unit (ignoring the strategic interaction), such data do not violate SUTVA, and we may be able to apply an existing causal inference method with multiple treatments to identify some causal parameters. However, we do not adopt such approach because doing so would make it difficult to uncover the nature of the interaction structure and, more importantly, may not allow us to perform a precise player-level treatment evaluation. }

To achieve our goal, we need to address the following two issues. The first issue is the non-applicability of {\it (unordered) monotonicity} (\citealpmain{imbens1994identification}; \citealpmain{heckman1999local, heckman2005structural}; \citealpmain{heckman2018unordered}). The monotonicity conditions require that shifts in instrumental variables (IVs) determine the direction of changes in the treatment choices uniformly for all agents. \phantomsection\Copy{AE-4-1}{ \citetmain{vytlacil2002independence} and \citetmain{heckman2018unordered} demonstrate that these monotonicity conditions are equivalent to assuming that each treatment realization is characterized by a single threshold-crossing equation in which the IVs and an unobserved error term are weakly separable. However, in our case, an agent's treatment choice may complexly depend on that of another agent through strategic interaction. As a result, each player's treatment choice cannot be expressed as a weakly separable threshold crossing model, implying the failure of the monotonicity. } The second issue is the possibility of multiple equilibria in the treatment decisions. The presence of multiple equilibria leads to an {\it incomplete} econometric model (e.g., \citealpmain{tamer2003incomplete}; \citealpmain{lewbel2007coherency}; \citealpmain{ciliberto2009market}; \citealpmain{chesher2020structural}) in the sense that the model-consistent treatment assignment is not unique. The issue of incompleteness is common in the literature on game model estimation; however, it is not yet well understood in the context of treatment evaluation.

Our identification strategy solves these two issues simultaneously by combining the local IV (LIV) method in \citemain{heckman1999local, heckman2005structural} and a stochastic equilibrium selection rule in the treatment decision game. The key idea is to use local variations of player-specific continuous IVs that alter players' treatment status, but do not directly affect their outcomes. Although this is a natural extension of the LIV method to a multi-dimensional space, it is still insufficient to point-identify the MTEs due to the presence of multiple equilibria. We overcome this issue by explicitly introducing an equilibrium selection rule.

To identify the MTE parameters, we first need to identify the parameters of the treatment decision game. We model the treatment decision game as a binary game of complete information. In particular, for the identification of MTE parameters, the model needs to have strategic interaction effects that are functions of the IV. This would be empirically plausible under many different situations. \phantomsection\Copy{R1-6-1}{ We then provide a new identification result for our game model based on a large support condition and a particular dependence property for a parametric copula function with a scalar correlation parameter (cf. \citealpmain{han2017identification}). } We also show that the large support condition can be mitigated when we introduce additional parametric model restrictions. \phantomsection\Copy{R1-6-2}{ Although treatment evaluation is our main concern, these identification results provide independent contributions to the literature; to the best of our knowledge, we would be the first to address the identification of game models where the joint distribution of the unobserved payoff components is characterized by a certain class of parametric copulas. }

Given the identification of the treatment decision game, we present the following identification results for treatment parameters. First, we can point-identify the following MTE parameters: the {\it direct} MTE, in which only one agent's treatment status switches from untreated to treated, whereas the partner's remains unchanged; the {\it indirect} MTE, in which only the partner's treatment status switches from untreated to treated, and the focal agent's status remains unchanged; and the {\it total} MTE, in which the treatment status of both players switches from untreated to treated. In contrast to the conventional MTE framework, our MTEs can reveal the treatment effect heterogeneity in terms of the pair of unobservables. We show that the regions in which the MTEs are identifiable are determined by the conditional supports of appropriately transformed threshold variables in the game model. This result extends \citemain{heckman1999local,heckman2005structural} and \citemain{carneiro2009estimating}, who prove the identification of conventional MTEs in terms of the supports of propensity scores. \phantomsection\Copy{AE-6-1}{ Second, we demonstrate that the MTE parameters are over-identified under our identification conditions, which is essentially a consequence of the stochastic equilibrium selection assumption. This over-identification result provides an opportunity to improve the efficiency of MTE estimation. } Finally, we present the identification of several other treatment parameters, including the LATE and PRTE (these results are relegated to Appendix (ref)).

Our identification is constructive in that we can estimate the MTEs directly by following the identification strategy. We propose a two-step semiparametric procedure for estimating the MTEs. In the first step, we estimate the parameters in the treatment decision game using a maximum likelihood (ML) approach based on a fully parametric model specification. Using the ML estimates, we estimate the MTEs in the second step by employing semiparametric series (sieve) technique. The proposed estimator is uniformly consistent with the optimal convergence rate and is asymptotically normally distributed. In addition, the estimator possesses an oracle property; that is, its limiting distribution is the same as that of the infeasible estimator where the parameters in the treatment decision game are known.

To illustrate our methods empirically, we investigate the impacts of the delinquency (e.g., smoking and drinking) of an opposite-gender best friend on the academic outcomes of an adolescent, i.e., their grade point average (GPA). Following the literature (e.g., \citealpmain{card2013peer}), we model the decision to participate in risky activities as a complete information game. Our method revealed the following empirical evidence: (i) the susceptibility to peer effects varies with the personality of the students, (ii) the direct treatment effect of risky behaviors on the GPA is significantly negative for both male and female students, and (iii) the total treatment effect is even larger than the direct effect for both genders. The third finding implies that the delinquency of a best friend has a negative causal impact on the academic performance of the students. These results partially overlap with those obtained in previous empirical studies, but the current study would be the first to report such results by formally integrating social interactions in risky activities and the resulting causal effects on academic performance in a single framework. We also demonstrate that if we ignore the strategic interaction between each pair of students and estimate the treatment model simply as a bivariate probit model, the resulting MTE estimates can differ significantly from those obtained in our framework. This indicates the importance of correctly addressing the interaction structure in the treatment choices, not just as correlated multiple treatments.

\phantomsection\Copy{R2-1-2}{ This study has a clear connection to both the treatment effect and the structural estimation literature. By formulating the treatment choice as a structural game model within the MTE framework, we can take advantage of both reduced-form causal inference and structural estimation. In particular, unlike a structural approach that fully specifies the outcome and treatment choice equations, our approach provides some robustness to the problem of model misspecification for the outcome equation. At the same time, compared to the conventional reduced-form analysis, we can perform a certain class of interesting counterfactual exercises through the structural estimation of treatment decision game. }

To the best of our knowledge, few studies have addressed treatment evaluation in the presence of strategic interactions modeled explicitly as games. One important exception is \citetmain{balat2020multiple}. Their approach is more general than ours in that it allows for nonparametric strategic interactions with more than two players and does not require point identification of the game parameters. To overcome the multiple equilibria issue, \citetmain{balat2020multiple} employ nonparametric shape restrictions and variations of possibly discrete IVs. However, because of such generality, the treatment parameter of interest in their study, ATE, is only partially identified in general. In contrast, we focus only on two-player games with a stochastic equilibrium selection to establish point identification for the treatment decision game and MTE. While this loses some generality, we can obtain rich information on the unobserved heterogeneity in the treatment effects that cannot be learned by estimating ATE alone. Thus, these two studies complement each other from different perspectives.

Another closely related study is \citetmain{lee2018identifying}, who showed that it is possible to identify MTE parameters by modeling the treatment selection with a set of threshold crossing rules. Their identification strategy is essentially the same as ours. They characterize a two-player binary game by combining five threshold crossing rules and identify the MTE parameters by computing the changes in the expected outcome with respect to local variations of all these threshold variables (see Appendix C of \citealpmain{lee2018identifying}). \phantomsection\Copy{AE-7-1}{ Compared to their model, each treatment realization in our model can be characterized by a smaller set of threshold crossing equations and consequently the target MTE parameters are not completely identical between these studies. This difference is crucial from a practical perspective because our MTE parameter can be estimated reasonably well even with moderate sample size (see Remark (ref) for more details). } \phantomsection\Copy{AE-8-1}{ In addition, they do not discuss in depth how to cope with the identification problem in the presence of multiple equilibria, whereas we formally prove the identification of both the game model and the treatment parameters under multiplicity. }

The rest of the paper is organized as follows. In Section (ref), we introduce our model and review the incompleteness problem for discrete game models. Section (ref) presents the identification analysis. We develop our MTE estimators and study their asymptotic properties in Section (ref). Section (ref) provides the numerical illustrations, including Monte Carlo simulations and the empirical analysis of adolescents' academic performance. Section (ref) concludes the paper. The proofs of all theorems are provided in Appendix (ref), and the other supplementary technical results in Appendix (ref). In Appendix (ref), we discuss the identification of several treatment parameters besides the MTE. Appendix (ref) presents detailed information about the Monte Carlo experiments in Section (ref). Finally, we provide supplementary material for the empirical analysis in Appendix (ref).

Model

In this section, we introduce our treatment effect model with strategic interaction. We denote a player by $j \in \{1, 2\}$ and his/her partner (or opponent) by $-j$. We aim to evaluate the effects of player $j$'s treatment $D_j \in \{0, 1\}$ and/or partner's treatment $D_{-j} \in \{0, 1\}$ on player $j$'s outcome $Y_j$ and/or the partner's outcome $Y_{-j}$. The outcomes may or may not be common to both players; that is, we allow both $Y_j = Y_{-j}$ and $Y_j \neq Y_{-j}$. For $(d_j, d_{-j}) \in \{ 0, 1 \}^2$, let $Y_j^{(d_j, d_{-j})}$ be the potential outcome when $j$'s own treatment status is $D_j = d_j$ and the partner's is $D_{-j} = d_{-j}$. Then, the observed outcome can be written as $Y_j = \sum_{d_j}\sum_{d_{-j}} I^{(d_j, d_{-j})} Y_j^{(d_j, d_{-j})}$, where $I^{(d_j, d_{-j})} \coloneqq \mathbf{1}\{(D_j, D_{-j}) = (d_j, d_{-j})\}$. \phantomsection\Copy{AE-9-1}{ Suppose that player $j$'s potential outcome can be written as

align[align omitted — 119 chars of source]

where $X_j \in \mathbb{R}^{\mathrm{dim}(X)}$ is a vector of observable covariates, $U_j^{(d_j, d_{-j})}$ is an unobserved random variable with arbitrary dimension, and $\mu_j^{(d_j, d_{-j})}$ is an unknown structural function. } The covariates $X_1$ and $X_2$ may contain common elements as well as some player-specific elements. For simplicity, we assume that the dimensions of $X_1$ and $X_2$ are both equal to $\mathrm{dim}(X)$, and the same assumption will be made for the other variables.

\phantomsection\Copy{AE-18-1}{ Note that we do not explicitly consider the stochastic process on who becomes whose partner, but we treat the pair formation as a given object. In other words, all subsequent analyses are conditioned on the existing pair formation and may not be generalizable to other pair data. While this could be a shortcoming of this study, such an approach has been adopted quite commonly in the social interaction literature to circumvent the notorious endogenous network formation problem (with a few exceptions: e.g., \citealpmain{goldsmith2013social,hsieh2016social,johnsson2021estimation}). }

Hereinafter, for a generic player-specific variable, say $A_j$, we write $A$ without a subscript to denote the union $A = (A_1, A_2)$. For example, $Y = (Y_1, Y_2)$, $X = (X_1, X_2)$, and $U^{(d_1, d_2)} = ( U_1^{(d_1, d_2)}, U_2^{(d_2, d_1)} )$. In addition, let $F_{A | E}$ denote the cumulative distribution function (CDF) of $A$ conditional on $E$. When $A$ is continuously distributed given $E$, we write its conditional probability density function (PDF) as $f_{A|E}$.

Strategic treatment decision

Suppose that each player $j$ obtains the payoff $\pi_j(D_{-j}, W_j) - \varepsilon_j$ when $D_j = 1$, and he/she obtains $0$ (for normalization) when $D_j = 0$, where $W_j \coloneqq (X_j^\top, Z_j^\top)^\top$ is a vector that includes $X_j$ and the instruments $Z_j \in \mathbb{R}^{\mathrm{dim}(Z)}$, $\varepsilon_j \in \mathbb{R}$ is an unobserved continuous variable, and $\pi_j$ is an unknown function. Then, for $j = 1,2$, the payoff function can be written as follows:

align*[align* omitted — 119 chars of source]

where $\Delta_j(W_j) \coloneqq \pi_j^1(W_j) - \pi_j^0(W_j)$ corresponds to the strategic interaction effect with $\pi_j^0(W_j) \coloneqq \pi_j(0, W_j)$ and $\pi_j^1(W_j) \coloneqq \pi_j(1, W_j)$.\footnote{ This expression clearly indicates that the additive separability imposed on the payoff function implies the interaction effect to be a function of only the observables. In our empirical setting, this requires that the impact of friends' delinquency is independent of own unobserved factors, such as his/her latent attitude towards risky activities. With few exceptions (e.g., \citealpmain{kline2015identification}), models where the interaction effects can depend on some unobservables have not been studied in detail in the literature. } Whereas the strategic interaction effect is often assumed to be constant in the game econometrics literature, we require it to be a non-constant function of the instruments for identification of the MTE parameters.

commentThe payoff matrix for this game is summarized in Table (ref). \begin{table}[!ht] \caption{Payoff matrix} \begin{tabular}{c|c|c} \hline & $D_2 = 0$ & $D_2 = 1$ \\ \hline $D_1 = 0$ & $(0, 0)$ & $(0, \pi^0_2(W_2) - \varepsilon_2)$ \\ \hline $D_1 = 1$ & $(\pi^0_1(W_1) - \varepsilon_1, 0)$ & $(\pi^1_1(W_1) - \varepsilon_1, \pi^1_2(W_2) - \varepsilon_2)$\\ \hline \end{tabular} \end{table}

Based on the payoff-maximization principle, the best response is $D_j(d_{-j}) = \operatorname*{argmax}_{d_j} u_j(d_j, d_{-j})$. Here, we assume that $(W, \varepsilon)$ is common knowledge for both players (i.e., a complete information setup). Furthermore, assume that the set of realized treatments $(D_j, D_{-j})$ is characterized as a pure strategy Nash equilibrium.\footnote{ Most of the studies in the literature focus on pure strategy Nash equilibria as the solution concept for $2 \times 2$ games, except when no Nash equilibria in pure strategies exist, for example, because of asymmetric strategic interaction (e.g., \citealpmain{bjorn1984simultaneous}; \citealpmain{tamer2003incomplete}; \citealpmain{depaula2013econometric}). For the motivations of this choice, see, for example, Section 6 of \citemain{bjorn1984simultaneous}. } Then, the players' treatment decisions follow the simultaneous binary response model:

align[align omitted — 235 chars of source]

The first line of (ref) shows that our treatment decision model can be viewed as a direct extension of the latent index model of \citetmain{heckman1999local, heckman2005structural}.\footnote{\Copy{R2-1-e2}{ We can see that the additive separability in the treatment payoff function essentially rules out unobserved preference heterogeneity in the treatment choice. One possible way to relax this restriction is to use a finite-mixture model to permit the preference heterogeneity across pairs, as in \citetmain{hoshino2022estimating}, but we leave this task for a future study. } } \phantomsection\Copy{AE-10-1}{ As we will discuss in Remark (ref), the model in (ref) does not satisfy the IV monotonicity. Studies in the MTE literature that consider a “non-monotonic” treatment model similar to ours include \citemain{klein2010heterogeneous} and \citemain{lee2018identifying}. }

commentNote that we do not restrict the dependence structure between $\varepsilon_j$ and $U_j^{(d_j, d_{-j})}$, which is the source of endogeneity. Without endogeneity, one may relatively easily identify some causal parameters by standard approaches based on an unconfoundedness assumption, even in the presence of strategic interaction.

In the following, we assume that the treatment decisions are known to be strategic complements (to be in line with our empirical application).

assumption\hfil \begin{enumerate}[(i)] • For both $j = 1, 2$, $\Delta_j(W_j) > 0$ almost surely (a.s.). • \phantomsection\Copy{AE-11-1}{ $(\varepsilon_1, \varepsilon_2)$ are conditionally independent of $Z$ given $X$ and continuously distributed with strictly increasing marginal CDFs $F_{\varepsilon_1 | X}$ and $F_{\varepsilon_2 | X}$, respectively. } \end{enumerate}

Note that Assumption (ref)(i) is empirically testable.\footnote{ For example, \citetmain{aradillas2019nonparametric} has developed nonparametric tests for the presence and direction of strategic interaction effects in $2 \times 2$ games of complete information. For another example, one may use a Vuong-type model selection test for non-nested alternatives (e.g., \citealpmain{hsu2017model}; \citealpmain{schennach2017simple}). } We also note that our analysis can be easily converted to the situation with strategic substitutes where $\Delta_j(W_j) < 0$ a.s. for both $j = 1, 2$. We introduce Assumption (ref)(ii) not only for the identification of the game model but also for that of MTE.

Define

align[align omitted — 221 chars of source]

By construction, $V_j $ is distributed as $\text{Uniform}[0,1]$ conditional on $X$. Then, we can rewrite (ref) as

align*[align* omitted — 107 chars of source]

By Assumption (ref), $P_j^1 - P_j^0 > 0$ a.s. for both $j = 1, 2$.

Incompleteness

A major difficulty in our model is the {\it incompleteness} of the treatment decision model. \phantomsection\Copy{AE-11-2}{ We have the following relationships:

equation*[equation* omitted — 340 chars of source]

}

comment\begin{equation*} \begin{array}{lcllcl} D = (1, 0) & \Longleftrightarrow & \varepsilon_1 \le \pi_1^0(W_1), \; \varepsilon_2 > \pi_2^1(W_2), & \quad D = (0, 1) & \Longleftrightarrow & \varepsilon_1 > \pi_1^1(W_1), \; \varepsilon_2 \le \pi_2^0(W_2), \\ D = (1, 1) & \Longrightarrow & \varepsilon_1 \le \pi_1^1(W_1), \; \varepsilon_2 \le \pi_2^1(W_2), & \quad D = (0, 0) & \Longrightarrow & \varepsilon_1 > \pi_1^0(W_1), \; \varepsilon_2 > \pi_2^0(W_2). \end{array} \end{equation*}
comment\begin{equation*} \begin{array}{lclcl} (D_1, D_2) = (1, 0) & \Longleftrightarrow & \varepsilon_1 \le \pi_1^0(W_1), \; \varepsilon_2 > \pi_2^1(W_2) & \Longleftrightarrow & V_1 \le P_1^0, \; V_2 > P_2^1, \\ (D_1, D_2) = (0, 1) & \Longleftrightarrow & \varepsilon_1 > \pi_1^1(W_1), \; \varepsilon_2 \le \pi_2^0(W_2) & \Longleftrightarrow & V_1 > P_1^1, \; V_2 \le P_2^0, \\ (D_1, D_2) = (1, 1) & \Longrightarrow & \varepsilon_1 \le \pi_1^1(W_1), \; \varepsilon_2 \le \pi_2^1(W_2) & \Longleftrightarrow & V_1 \le P_1^1, \; V_2 \le P_2^1, \\ (D_1, D_2) = (0, 0) & \Longrightarrow & \varepsilon_1 > \pi_1^0(W_1), \; \varepsilon_2 > \pi_2^0(W_2) & \Longleftrightarrow & V_1 > P_1^0, \; V_2 > P_2^0. \end{array} \end{equation*}

Figure (ref) visually summarizes these relationships. As shown in the figure, the space of $(V_1, V_2)$ cannot be partitioned into non-overlapping regions associated with the four alternative realizations of $D$. Both $D = (1, 1)$ and $D = (0, 0)$ can occur when $P_1^0 < V_1 \le P_1^1$ and $P_2^0 < V_2 \le P_2^1$ (i.e., multiple equilibria). This non-uniqueness of model-consistent decisions is called incompleteness and has been extensively studied in the literature on simultaneous equation models for discrete outcomes.

figure[figure omitted — 190 chars of source]

In the game econometrics literature, there are several approaches to handle the incompleteness problem to achieve point identification.\footnote{ See \citetmain{depaula2013econometric} for an excellent survey on this topic. } One major approach is to explicitly introduce a stochastic (or possibly deterministic) equilibrium selection mechanism (e.g., \citealpmain{bjorn1984simultaneous}; \citealpmain{kooreman1994estimation}; \citealpmain{soetevent2007discrete}; \citealpmain{bajari2010identification}; \citealpmain{card2013peer}). This approach allows us to predict the choice behavior in the region of multiplicity and to perform counterfactual exercises, which is essential to establish the identification of the MTE parameters, and thus is adopted in this study.

figure[figure omitted — 184 chars of source]
remark[Monotonicity] To see why the monotonicity conditions are not applicable to our situation, let $D_j(w)$ be the potential treatment status of player $j$ when $W = w$, and $D(w) = (D_1(w), D_2(w))$. The (ordered) monotonicity in \citetmain{imbens1994identification} is that, for any $w, w' \in \text{supp}[W]$ and $j = 1, 2$, either $\Pr[ D_j(w) \ge D_j(w') ] = 1$ or $\Pr[ D_j(w) \le D_j(w') ] = 1$ should hold. The unordered monotonicity in \citetmain{heckman2018unordered} is that, for any $w, w' \in \text{supp}[W]$ and $(d_1, d_2) \in \{ 0, 1 \}^2$, either $\Pr[ \mathbf{1}\{ D(w) = (d_1, d_2) \} \ge \mathbf{1}\{ D(w') = (d_1, d_2) \} ] = 1$ or $\Pr[ \mathbf{1}\{ D(w) = (d_1, d_2) \} \le \mathbf{1}\{ D(w') = (d_1, d_2) \} ] = 1$ should hold. Here, consider two groups of pairs: the one in the region of $D = (0, 0)$ and the other in the region of $D = (1, 1)$ for a given $W = w$. Suppose that shifting $W$ from $w$ to $w'$ induces both groups to the multiple equilibria. Figure (ref) illustrates such a situation where the pairs in the former group and those in the latter are respectively located in regions [A] and [B]. Here, $(p_1^0, p_1^1, p_2^0, p_2^1)$ and $(p_1^{0'}, p_1^{1'}, p_2^{0'}, p_2^{1'})$ correspond the values of $(P_1^0, P_1^1, P_2^0, P_2^1)$ when $W = w$ and $W = w'$, respectively. In this situation, some of the former pairs would switch their treatment statuses from $D = (0, 0)$ to $(1, 1)$, while some of the latter pairs would switch from $D = (1, 1)$ to $(0, 0)$. Thus, the impact of $W$ on the treatment choice may be non-monotonic.

Identification

In this section, we present the identification results for the treatment decision game and the MTE parameters. We refer to the treatment decision game (ref) as the “first stage” and the realization of the outcome (ref) as the “second stage”. Because the identification of MTE relies on the knowledge of the first-stage parameters, we first examine the identification of them in Subsection (ref), and then present the identification results for the second-stage parameters in Subsection (ref).

Identification of the first-stage parameters

We first introduce the following assumptions on the joint distribution of $(\varepsilon_1, \varepsilon_2)$.

assumption\hfil \begin{enumerate}[(i)] • The joint distribution of $(\varepsilon_1, \varepsilon_2)$ given $X = x$ is represented by the copula function $H_{\rho_x}$ such that $\Pr[\varepsilon_1 \le a_1, \varepsilon_2 \le a_2 | X = x] = H_{\rho_x}(F_{\varepsilon_1 | X = x}(a_1), F_{\varepsilon_2 | X = x}(a_2))$, where $\rho_x \in (\underline{c}, \bar c)$ is a scalar correlation parameter and $\underline{c}$ and $\bar c$ are real numbers whose values depend on the choice of the copula function. • $H_{\rho_x}(\cdot, \cdot)$ is twice differentiable in its arguments and ${\rho_x}$. • $H_{\rho_x}(\cdot, \cdot)$ is strictly more stochastically increasing in joint distribution with respect to ${\rho_x}$ (see Definition 3.3 of \citetmain{han2017identification}). \end{enumerate}

\phantomsection\Copy{AE-12-1}{ Assumption (ref)(iii) restricts the dependence ordering of the copula function in terms of stochastic monotonicity. \citemain{han2017identification} use this property to identify generalized bivariate probit models, and they demonstrate that various commonly used copula functions satisfy it. Note that this dependence property is only for the class of copulas with a scalar correlation parameter. Investigating the identification with a more general multi-parameter copula is beyond the scope of our study. } A typical example satisfying this assumption is a standard bivariate normal distribution, in which $H_{\rho_x}$ corresponds to the Gaussian copula $H_{\rho_x}(v_1, v_2) = \Phi_2(\Phi^{-1}(v_1), \Phi^{-1}(v_2); \rho_x)$, with $\Phi_2(\cdot, \cdot; \rho_x)$ and $\Phi(\cdot)$ being the standard bivariate normal CDF with correlation $\rho_x$ and the standard normal CDF, respectively. For another example, the assumption is also satisfied with the Farlie--Gumbel--Morgenstern (FGM) copula $H_{\rho_x}(v_1, v_2) = v_1v_2[1 + \rho_x(1 - v_1)(1 - v_2)]$. For other examples and further discussions on the dependence ordering properties of copula functions, see \citemain{han2017identification}.

assumptionIn the case of multiple equilibria for a given $X = x$, $D = (0, 0)$ occurs if and only if $\epsilon \le \lambda_x$, where $\lambda_x \in [0, 1]$ is a constant and $\epsilon$ is a random variable distributed as $\text{Uniform}[0, 1]$ independent of $(Z, \varepsilon, U^{(d_1, d_2)})$ given $X$ for $(d_1,d_2)\in\{0,1\}^2$.

This assumption states that, given $X = x$, $D = (0, 0)$ is observed with probability $\lambda_x$ in the multiple equilibria situation. For example, some authors assume that under multiple equilibria, one of model-consistent actions is selected uniformly at random (e.g., \citealpmain{bjorn1984simultaneous}; \citealpmain{soetevent2007discrete}; \citealpmain{card2013peer}). If we adopt the same assumption, we can set $\lambda_x = 0.5$ a priori. As another example, one may assume that the realized treatment status corresponds to the “largest” Nash equilibrium, as in \citemain{xu2015estimation}. In this case, because $u_j(1,1) > u_j(0,0)$ in the multiple equilibria region for both players, $D = (1,1)$ holds almost surely, i.e., $\lambda_x = 0$. Assumption (ref) allows for a more general equilibrium selection in that $\lambda_x$ can be unknown and can depend on $X$. \phantomsection\Copy{R2-2-1}{ However, note that the IV $Z$ should not affect equilibrium selection. Such an assumption is not necessary for identifying the game model but is required for nonparametrically identifying MTE (see also Remark (ref)). }

In addition, Assumption (ref) rules out the possibility of “endogenous” equilibrium selection such that the selection probability depends on unobservables $(\varepsilon, U^{(d_1,d_2)})$. In our empirical setting, this requires that the latent academic abilities and attitudes toward risky activities do not affect which is selected between the two equilibria. Although this requirement might be restrictive for certain empirical applications, it would be challenging to identify such a game model (cf. \citealpmain{jun2020counterfactual}).

For any given $x \in \text{supp}[X]$ and $w \in \text{supp}[W | X = x]$, we write the conditional probability of $D = (d_1, d_2)$ as $\mathcal{L}^{(d_1, d_2)}(w) \coloneqq \Pr[D = (d_1, d_2) | W = w]$ and the value of $P_j^d$ as $p_j^d = F_{\varepsilon_j | X = x}(\pi_j^d(w_j))$. Then, the above assumptions give

equation[equation omitted — 394 chars of source]

where $\mathcal{L}_{\text{mul}}(w) \coloneqq H_{\rho_x}(p_1^1, p_2^1) - H_{\rho_x}(p_1^1, p_2^0) - H_{\rho_x}(p_1^0, p_2^1) + H_{\rho_x}(p_1^0, p_2^0)$ is the probability that the pair of players is in the multiple equilibria region. There are six unknown parameters $(\mathbf{p}, \rho_x, \lambda_x)$ in the moment equations in (ref), where $\mathbf{p} = (p_1^0, p_1^1, p_2^0, p_2^1)$.

assumption\hfil \begin{enumerate}[(i)] • For both $j = 1, 2$, there exists a player-specific continuous random variable, say $W_{j,1}$, whose PDF is everywhere positive on $\mathbb{R}$ given $(W_{j,-1}, W_{-j})$, where $W_j = (W_{j, 1}, W_{j, -1})$, so that each of $\pi_j^0(W_j)$ and $\pi_j^1(W_j)$ is non-degenerate and continuously distributed conditional on $(W_{j,-1}, W_{-j})$. • For both $j = 1, 2$, $\pi_j^1(w_{j, 1}, w_{j, -1}) \to -\infty$ as $w_{j, 1} \to -\infty$ for any $w_{j, -1} \in \text{supp}[W_{j, -1} | X = x]$. \end{enumerate}

Assumption (ref)(i) requires that at least one player-specific element of $W_j$ can tend to $-\infty$ and $\infty$, and Assumption (ref)(ii) says that this player-specific variable has a significant impact on the payoff function $\pi_j^1$. We note that this assumption can be replaced by $\pi_j^1(w_{j, 1}, w_{j, -1}) \to -\infty$ as $w_{j, 1} \to \infty$. Some prior studies have utilized assumptions similar to ours (e.g., tamer2003incomplete; kline2015identification). Importantly, this assumption is empirically verifiable because almost all pairs with sufficiently small $w_{1, 1}$ and $w_{2, 1}$ should choose $D = (0, 0)$ if it is true.

The next theorem presents identification for a general parameter value $(\mathbf{p}, \rho_x, \lambda_x)$, for any given $x \in \text{supp}[X]$ and $w \in \text{supp}[W | X = x]$. To the best of our knowledge, the identification result for the correlation parameter $\rho_x$ is novel in the literature of discrete games with complete information.

theorem\phantomsection\Copy{AE-11-3}{ Fix arbitrary $x \in \text{supp}[X]$ and $w \in \text{supp}[W | X = x]$. Let $p_j^d = F_{\varepsilon_j | X = x}(\pi_j^d(w_j))$ with $d = 0, 1$ and $j = 1, 2$. \begin{enumerate}[(i)] • Suppose that Assumptions (ref), (ref)(i), and (ref) hold. Then, $p_j^0$ is identified for $j = 1, 2$. • In addition, suppose that Assumptions (ref)(ii)--(iii) hold. If $\text{supp}[P_j^1 | X = x] \times (\underline{c}, \bar c)$ is a simply connected set, then $(p_j^1, \rho_x)$ are identified for $j = 1,2$. • If Assumption (ref) additionally holds, $\lambda_x$ is identified. \end{enumerate} }

Note that $\mathbf{p} = (p_1^0, p_1^1, p_2^0, p_2^1)$ can be identified without assuming any particular form of equilibrium selection. Indeed, the equilibrium selection assumption is imposed not to simplify the identification of the game model but to achieve the identification of the MTE parameters.

remark[Identification without the large support condition] The large support condition in Assumption (ref) can be replaced by other identification conditions. For example, consider $(w_1, w_2)$, $(w_1', w_2)$, $(w_1, w_2')$, $(w_1', w_2') \in \text{supp}[W | X = x]$ such that $w_j \neq w_j'$ for $j = 1, 2$. Define $\vartheta_{w, w'} \coloneqq (\mathbf{p}, \mathbf{p}', \rho_x, \lambda_x)$ and \begin{align*} \mathcal{G}(w_1, w_2; \vartheta_{w, w'}) \coloneqq \begin{pmatrix} p_1^0 - H_{\rho_x}(p_1^0, p_2^1), \;\; p_2^0 - H_{\rho_x}(p_1^1, p_2^0), \;\; H_{\rho_x}(p_1^1, p_2^1) - \lambda_x \cdot \mathcal{L}_{mul}(w) \end{pmatrix} , \end{align*} where $\mathbf{p}' = (p_1^{0'}, p_1^{1'}, p_2^{0'}, p_2^{1'})$ with $p_j^{d'} = F_{\varepsilon_j | X = x} (\pi_j^d(w_j'))$. Further, let $\mathcal{G}(\vartheta_{w, w'}) \coloneqq [\mathcal{G}(w_1, w_2; \vartheta_{w, w'})$, $\mathcal{G}(w_1', w_2; \vartheta_{w, w'}), \mathcal{G}(w_1, w_2'; \vartheta_{w, w'}), \mathcal{G}(w_1', w_2'; \vartheta_{w, w'})]^\top$. Then, $\mathcal{G}(\vartheta_{w, w'})$ is a $12 \times 1$ vector whose value is knowable from data in view of (ref), while $\vartheta_{w, w'}$ contains 10 unknown parameters. Thus, we can locally identify $\vartheta_{w, w'}$ if and only if the rank of the Jacobian matrix $\partial \mathcal{G}(\vartheta_{w, w'}) / \partial \vartheta_{w, w'}$ is equal to 10 (Theorem 6, \citealpmain{rothenberg1971identification}). Moreover, we can achieve the global identification if we introduce additional conditions on the structure of the Jacobian matrix.\footnote{ By Lemma 4.2 of \citetmain{han2017identification}, $\vartheta_{w, w'}$ is globally identified if there exists a $10 \times 1$ sub-vector $\mathcal{G}'$ of $\mathcal{G}$ such that (i) $\mathcal{G}'$ is proper, (ii) the Jacobian of $\mathcal{G}'$ vanishes nowhere, and (iii) the range of $\mathcal{G}'$ is simply connected. } In general, proving the full-rankness of the Jacobian matrix with easy-to-check conditions is difficult unless we introduce specific parametric specifications. As such an example, we consider a particular linear index form for the payoff function and the equilibrium selection in Appendix (ref), and demonstrate that the parameters in this game model can be identified globally without the large support condition. Another approach to circumvent the large support condition is to employ the strategy of \citetmain{kline2016empirical}, which is based on the unimodality of the error distribution and the linearity of the payoff function.
comment\begin{remark}[Dependence ordering property of the copula] The dependence ordering property in Assumption (ref)(iii) may be slightly stronger than necessary. To see this, we denote the partial derivatives of copula $H_{\rho_x}(\cdot, \cdot)$ with respect to the first argument and ${\rho_x}$ as $H_{\rho_x}^{(1)}(\cdot, \cdot)$ and $H_{\rho_x}^{(\rho)}(\cdot, \cdot)$, respectively. The proof of Theorem (ref) indicates that $(p_1^1, \rho_x)$ is locally identifiable if and only if $H_{\rho_x}^{(\rho)}(p_1^1, v_2)/H_{\rho_x}^{(1)}(p_1^1, v_2)$ is a nonconstant function in $v_2$ (see Proposition 4.1 of \citealpmain{han2017identification}). In addition, according to Lemma 4.1 of \citetmain{han2017identification}, Assumption (ref)(iii) is equivalent to that $H_{\rho_x}^{(\rho)}(v_1, v_2)/H_{\rho_x}^{(1)}(v_1, v_2)$ is strictly decreasing in $v_2$ and $H_{\rho_x}^{(\rho)}(v_1, v_2)/[1 - H_{\rho_x}^{(1)}(v_1, v_2)]$ is strictly increasing in $v_2$ for all $v_1 \in (0, 1)$ and $\rho_x \in (\underline{c}, \bar c)$. These facts may imply that assuming the dependence ordering property is not strictly necessary for the result in Theorem (ref) to hold. Nonetheless, as discussed in \citetmain{han2017identification}, this assumption should be minimal in terms of interpretability. \end{remark}

Finally, we note that any counterfactual analysis requires that the structure of the game model, including the equilibrium selection mechanism, should remain invariant under the new environment. Although this requirement sounds restrictive, it cannot be omitted in principle, except for a partial identification analysis. For a more specific discussion, see \citemain{jun2020counterfactual} and Appendix (ref) of this paper.

Identification of the second-stage parameters

Throughout this subsection, we denote the conditional joint CDF and PDF of $(V_1, V_2)$ given $X = x$ as $H(\cdot, \cdot | x)$ and $h(\cdot, \cdot | x)$, respectively. Theorem (ref)(ii) ensures that these functions are identified via the copula function $H_{\rho_x}$. Including these, all the first-stage components that are identifiable through Theorem (ref) are treated as observable information in the following analysis.

assumption\hfil \begin{enumerate}[(i)] • $Z$ is excluded from the structural functions in (ref) and is independent of the unobservables $(\varepsilon, U^{(d_1, d_2)})$ given $X$ for $(d_1, d_2) \in \{0, 1\}^2$. • \phantomsection\Copy{AE-14-1}{ For all $(d_1, d_1', d_2, d_2') \in \{ 0, 1 \}^4$ such that $d_1 \neq d_1'$ and $d_2 \neq d_2'$, $(P_1^{d_1}, P_2^{d_2})$ is continuously distributed conditional on $(X, P_1^{d_1'}, P_2^{d_2'})$. } \end{enumerate}

Assumption (ref)(i) requires $Z$ not to directly affect $Y$ and to be conditionally independent of the error terms. By construction, the transformed error $V$ is also conditionally independent of $Z$ given $X$. \phantomsection\Copy{AE-15}{ Assumption (ref)(ii) can be read as the exclusion restriction assumption that is often required for nonparametric identification. That is, for this to hold in a nonparametric model, both player- and action-specific continuous IVs must exist. However, note that when considering a fully parametric game model with a linear index form (i.e., the one given in Assumption (ref)), the existence of action-specific IVs is not strictly required to maintain Assumption (ref)(ii), and the assumption can be satisfied if some IVs are player-specific and continuously distributed. }

\phantomsection\Copy{R2-1-1}{ As an example of IV $Z$ satisfying these conditions, in the empirical application of this study, with $Y$ being students' academic performance and $D$ being their delinquency, we consider employing their non-best friends' parental characteristics. For another example, in the case of an airline market with two competitors, we may estimate the effect of introducing a direct flight ($D$) on the number of total passengers for each company ($Y$) by using cost variables as $Z$ (cf. \citealpmain{ciliberto2009market}). As the last example, when we are interested in measuring how the choice of electoral campaign strategy ($D$) affects the candidate's vote share ($Y$) in a mayoral election, the amount spent in the previous election may be a good candidate for $Z$ (in a similar sense to \citealpmain{gerber1998estimating}). Then, once the MTE is obtained, we can estimate some interesting PRTEs; for example, what if there had been a cap on the election budget imposed by the government? }

We provide below a series of identification results only for player 1 (the results for player 2 are symmetric and thus omitted). Define

align[align omitted — 124 chars of source]

for a given $x \in \text{supp}[X]$ and $(p_1, p_2) \in [0, 1]^2$. We call the function $m^{(d_1, d_2)}$ the marginal treatment response (MTR) function, as in \citetmain{mogstad2018using}. From the viewpoint of player 1, the parameters of interest are the direct MTE, indirect MTE, and total MTE:

align*[align* omitted — 426 chars of source]

Notably, these MTE parameters are informative about the variation of the treatment effects in terms of the pair of unobservables $V_1$ and $V_2$, in contrast to the conventional MTE framework. The identification of the MTE parameters is straightforward once the MTR function for each $(d_1,d_2)$ is identified.

Identification of the marginal treatment response functions

We first investigate the identification of the MTRs $m^{(1, 0)}(x, p_1, p_2)$ and $m^{(0, 1)}(x, p_1, p_2)$. Since $D = (1, 0)$ and $D = (0, 1)$ are unique equilibria, the presence of multiple equilibria does not cause problems in these cases. Indeed, we can achieve point identification of MTRs with a natural extension of the conventional LIV method without using the equilibrium selection rule in Assumption (ref).\footnote{ For this study, when it is stated that MTR $m^{(d_1,d_2)}(x, p_1, p_2)$ is “(point) identified”, the value of the MTR at this specific $(p_1, p_2)$ is identified, but the whole functional form with respect to $(p_1, p_2)$ is not necessarily identified. When the latter is true, we say, for example, the MTR function $m^{(d_1,d_2)}(x, \cdot, \cdot)$ is “fully” (point) identified on $[0,1]^2$. The same expression applies to the identification of MTEs. }

For the subsequent analysis, we introduce a standard overlap condition.

assumption\phantomsection\Copy{R1-1-1}{ For all $(d_1, d_2) \in \{ 0, 1 \}^2$, $0 < \Pr[ D = (d_1, d_2) | W ] < 1$ a.s. }

\phantomsection\Copy{R1-1-2}{ Define

align*[align* omitted — 254 chars of source]

Under Assumption (ref), these quantities are identified on $\mathcal{S}_x^{(1,0)} \coloneqq \text{supp}[P_1^0, P_2^1 | X = x]$ and $\mathcal{S}_x^{(0,1)} \coloneqq \text{supp}[P_1^1, P_2^0 | X = x]$, respectively. The next theorem demonstrates that the MTR functions $m^{(1,0)}(x, \cdot, \cdot)$ and $m^{(0,1)}(x, \cdot, \cdot)$ are also identified on these supports. } Hereinafter, we use the following differential operator notations: $\partial_{p_j} \coloneqq \partial / (\partial p_j)$ and $\partial_{p_1 p_2} \coloneqq \partial^2 / (\partial p_1 \partial p_2)$.

theoremSuppose that Assumptions (ref), (ref), (ref), (ref), and (ref) hold. \phantomsection\Copy{R1-3-1}{ Under these assumptions, $\mathbf{P} = (P_1^0, P_1^1, P_2^0, P_2^1)$ and $h(\cdot, \cdot | x)$ are identified from Theorem (ref)(i)--(ii). } If $m^{(1, 0)}(x, \cdot, \cdot)$, $m^{(0,1)}(x, \cdot, \cdot)$, and $h(\cdot, \cdot | x)$ are continuous, the MTRs are identified in the following way: \begin{align*} & m^{(1, 0)}(x, p_1, p_2) = - \frac{\partial_{p_1 p_2} [\psi^{(1,0)}(x, p_1, p_2)]}{h(p_1, p_2 | x)} \quad for \;\; (p_1, p_2) \in \mathcal{S}_x^{(1,0)}, \\ & m^{(0, 1)}(x, p_1, p_2) = - \frac{\partial_{p_1 p_2} [\psi^{(0,1)}(x, p_1, p_2)]}{h(p_1, p_2 | x)} \quad for \;\; (p_1, p_2) \in \mathcal{S}_x^{(0,1)}. \end{align*}

The main idea behind Theorem (ref) can be visually understood by Figure (ref). We consider the pairs of players with $X = x$ and $\mathbf{P} = \mathbf{p}$, where $\mathbf{p} = (p_1^0, p_1^1, p_2^0, p_2^1)$. Because $(V_1, V_2)$ are continuously distributed on $[0, 1]^2$, a certain proportion of the pairs are in the neighborhood of $(p_1^0, p_2^1)$, point A, at the margin of $D = (1, 0)$, $(0, 0)$, and $(1, 1)$. Hence, for a small fluctuation in $(P_1^0, P_2^1)$ at this point, some pairs would switch their treatment status between $D = (1, 0)$ and $D = (0, 0)$ or $(1, 1)$. As $\mathbf{P}$ is exogenous under Assumption (ref)(i), such variation in the treatment status is generated exogenously, which in turn can be used to identify $m^{(1, 0)}(x, p_1^0, p_2^1)$. Meanwhile, since fluctuations of $(P_1^1, P_2^1)$ around point B, that of $(P_1^0, P_2^0)$ around C, and of $(P_1^1, P_2^0)$ around D do not generate any shift of the treatment status from/to $D = (1, 0)$, they do not possess identification power for $m^{(1, 0)}(x, p_1^0, p_2^1)$. The analogous argument applies to the identification of $m^{(0, 1)}(x, p_1^1, p_2^0)$ at point D.

figure[figure omitted — 182 chars of source]

Presenting a sketch of the proof for $m^{(1, 0)}(x, p_1, p_2)$ would be helpful. It can be shown that

align[align omitted — 287 chars of source]

The first line follows from Assumption (ref)(i) and the fact that the treatment status $D = (1, 0)$ is uniquely linked with the set of threshold crossing rules $V_1 \le P_1^0$ and $V_2 > P_2^1$. Since $(P_1^0, P_2^1)$ are jointly continuously distributed under Assumption (ref)(ii), we can take the partial derivatives of both sides with respect to $p_1$ and $p_2$, which yields the desired result. Note that Assumptions (ref) and (ref) are not used to show (ref), but they are used to recover $\mathbf{P}$ and $h$.

In view of (ref), we can see that Assumption (ref)(ii) is in fact stronger than necessary for Theorem (ref). Indeed, the partial differentiation of (ref) with respect to $(p_1, p_2)$ is well-defined as long as only $(P_1^0, P_2^1)$ are continuously distributed conditional on $X = x$.

We move on to the identification of $m^{(0, 0)}(x, p_1, p_2)$ and $m^{(1, 1)}(x, p_1, p_2)$. Due to the presence of multiple equilibria, the treatment statuses $D = (0, 0)$ and $D = (1, 1)$ are not uniquely determined by the threshold crossing rules regarding $V$ only. Consequently, in general, we cannot achieve point identification of the MTR for $(d_1, d_2) \in \{ (0, 0), (1, 1) \}$ with no assumptions on the equilibrium selection. In the following, we demonstrate that Assumption (ref) is sufficient to overcome this issue, and enables us to utilize the same identification strategy as in Theorem (ref).

\phantomsection\Copy{R1-1-3}{ To state the next theorem, define

align*[align* omitted — 128 chars of source]

which is identified on $\text{supp}[\mathbf{P} | X = x]$. }

theoremSuppose that Assumptions (ref) and (ref)--(ref) hold. \phantomsection\Copy{R1-3-2}{ Under these assumptions, $\mathbf{P} = (P_1^0, P_1^1, P_2^0, P_2^1)$, $h(\cdot, \cdot | x)$, and $\lambda_x$ are identified from Theorem (ref)(i)--(iii). } If $m^{(0, 0)}(x, \cdot, \cdot)$, $m^{(1, 1)}(x, \cdot, \cdot)$, and $h(\cdot, \cdot | x)$ are continuous and $0 < \lambda_x < 1$, the MTRs are identified in the following way: for $\mathbf{p} = (p_1^0, p_1^1, p_2^0, p_2^1) \in \text{supp}[\mathbf{P} | X = x]$, \begin{alignat*}{3} m^{(0, 0)}(x, p_1^0, p_2^0) & = \frac{\partial_{p_1^0 p_2^0} [\psi^{(0, 0)}(x, \mathbf{p})]}{\lambda_x h(p_1^0, p_2^0 | x)}, &\quad m^{(0, 0)}(x, p_1^1, p_2^1) & = - \frac{\partial_{p_1^1 p_2^1} [\psi^{(0, 0)}(x, \mathbf{p})]}{(1 -\lambda_x) h(p_1^1, p_2^1|x)},\\ m^{(0, 0)}(x, p_1^1, p_2^0) & = \frac{\partial_{p_1^1 p_2^0} [\psi^{(0, 0)}(x, \mathbf{p})]}{(1 - \lambda_x) h(p_1^1, p_2^0|x)}, &\quad m^{(0, 0)}(x, p_1^0, p_2^1) & = \frac{\partial_{p_1^0 p_2^1} [\psi^{(0, 0)}(x, \mathbf{p})]}{(1 - \lambda_x) h(p_1^0, p_2^1|x)}, \end{alignat*} and \begin{alignat*}{3} m^{(1, 1)}(x, p_1^0, p_2^0) & = - \frac{\partial_{p_1^0 p_2^0} [\psi^{(1, 1)}(x, \mathbf{p})]}{\lambda_x h(p_1^0, p_2^0 | x)}, &\quad m^{(1, 1)}(x, p_1^1, p_2^1) & = \frac{\partial_{p_1^1 p_2^1} [\psi^{(1, 1)}(x, \mathbf{p})]}{(1 - \lambda_x) h(p_1^1, p_2^1|x)},\\ m^{(1, 1)}(x, p_1^1, p_2^0) & = \frac{\partial_{p_1^1 p_2^0} [\psi^{(1, 1)}(x, \mathbf{p})]}{\lambda_x h(p_1^1, p_2^0|x)}, &\quad m^{(1, 1)}(x, p_1^0, p_2^1) & = \frac{\partial_{p_1^0 p_2^1} [\psi^{(1, 1)}(x, \mathbf{p})]}{\lambda_x h(p_1^0, p_2^1|x)}. \end{alignat*}

The proof of Theorem (ref) is straightforward from the following fact:

align[align omitted — 333 chars of source]

A similar representation can be obtained for $\psi^{(1, 1)} (x, \mathbf{p} )$. Again, these identification results can be visually understood in Figure (ref). For example, a small fluctuation of $(P_1^0, P_2^0)$ around point C generates an exogenous shock to change the treatment status of a certain portion of pairs between $D = (0, 0)$ and $D = (1,1)$, provided that $\lambda_x > 0$. This local exogenous variation in the treatment choice can be used to recover $m^{(0, 0)}(x, p_1^0, p_2^0)$. Note that, if $\lambda_x = 0$, the identification at point C fails because $D = (0, 0)$ is never chosen at this point. Similarly, if $\lambda_x < 1$, a variation in $(P_1^1, P_2^1)$ in the neighborhood of point B also causes an exogenous treatment shift between $D = (0, 0)$ and $D = (1, 1)$, which enables us to identify $m^{(0, 0)}(x, p_1^1, p_2^1)$.

Theorems (ref) and (ref) differ in two important respects. First, Assumption (ref)(ii) is necessary for Theorem (ref). This point should be obvious because the partial differentiation with respect to $(p_1^{d_1}, p_2^{d_2})$ with $(p_1^{d_1'}, p_2^{d_2'})$ being fixed ($d_1 \neq d_1'$, $d_2 \neq d_2'$) is not well-defined without Assumption (ref)(ii). Second, recall that Theorem (ref) requires only two out of $(P_1^0, P_1^1, P_2^0, P_2^1)$ as the conditioning variables because the point at which identification is achieved is well characterized as the upper-left or lower-right corner in the space of $V$. By contrast, to characterize the multiple equilibria region, we should fix the values of all four items $(P_1^0, P_1^1, P_2^0, P_2^1)$, of which two are marginal probabilities of each player's counterfactual treatment. This is the main reason why we need to identify all of the payoff functions, the equilibrium selection probability, and the correlation parameter in the first-stage game (as kindly pointed out by a referee).

\phantomsection\Copy{R1-1-4}{ The results in Theorem (ref) enable us to recover the MTR $m^{(0, 0)}(x, \cdot, \cdot)$ on the union of the supports of $(P_1^{d_1}, P_2^{d_2})$ for $(d_1, d_2) \in \{ 0, 1 \}^2$:

align*[align* omitted — 134 chars of source]

To observe this, note that, for arbitrary $(p_1^0, p_2^0) \in \text{supp}[P_1^0, P_2^0 | X = x]$, some $p_1^1$ and $p_2^1$ exist such that $(p_1^0, p_1^1, p_2^0, p_2^1) \in \text{supp}[\mathbf{P} | X = x]$. Therefore, $m^{(0, 0)}(x, p_1^0, p_2^0)$ is identified for any $(p_1^0, p_2^0) \in \text{supp}[P_1^0, P_2^0 | X = x]$ as in Theorem (ref). The same argument holds for the other cases. Consequently, $m^{(0, 0)}(x, p_1, p_2)$ is identifiable for any $(p_1, p_2) \in \overline{\mathcal{S}}_x$. Similarly, $m^{(1, 1)}(x, \cdot, \cdot)$ is identified on $\overline{\mathcal{S}}_x$. }

remark[Identifiable regions] \phantomsection\Copy{R1-3-3}{ The identification regions in Theorems (ref) and (ref), namely, $\mathcal{S}_x^{(1,0)}$, $\mathcal{S}_x^{(0,1)}$, and $\overline{\mathcal{S}}_x$, are characterized indirectly through the parameters in the first-stage treatment decision game as $P_j^d = F_{\varepsilon_j | X = x}(\pi_j^d(W_j))$, not directly from the standard propensity score, say $\text{PS}_j \coloneqq \mathbb E[D_j | W_j]$. In the canonical single agent setting, the MTR parameter for each agent $j$ can be identified on the conditional support of $\text{PS}_j$ (cf. \citealpmain{heckman1999local, heckman2005structural}). Even in the presence of treatment spillovers within each pair of players, if there is no strategic interaction in the treatment decision, it is straightforward to show that the MTR parameter can be identified on the joint support of $(\text{PS}_1, \text{PS}_2)$ (see Appendix (ref) for more details). Once a strategic interaction comes into play, such simple identification arguments no longer hold. }
remark[Role of the large support condition] \phantomsection\Copy{R1-5-1}{ In Theorems (ref) and (ref), the large support condition in Assumption (ref)(i) is used only for recovering the parameters in the treatment decision game. If the first-stage parameters can be identified under alternative assumptions as in Remark (ref), the identification of the MTR at an interior point of the support of $(P_1^{d_1}, P_2^{d_2})$ can be established without the large support condition. However, note that, if the goal is to identify other treatment parameters, such as the LATE and PRTE, we generally need the large support condition to recover the MTR fully on $[0,1]^2$. See Appendix (ref) for the identification of these parameters. }
remark[Imposing a parametric restriction on MTR] \phantomsection\Copy{R2-3}{ Theorems (ref) and (ref) are uninformative about the value of the MTRs for $(p_1, p_2)$ outside the support of $(P_1^{d_1}, P_2^{d_2})$. This is due to the nonparametric nature of the identification strategy. Alternatively, if we impose an explicit functional form on $m^{(d_1,d_2)}(x, \cdot, \cdot)$, such as polynomials, there is a possibility to identify $m^{(d_1,d_2)}(x, \cdot, \cdot)$ by interpolating from the other identified MTR values. This approach is analogous to the identification strategy discussed in \citetmain{brinch2017beyond} in the conventional MTE framework. If we adopt such an approach, even when the equilibrium selection probability $\lambda_x$ also depends on $Z = z$, we would be able to identify $m^{(0, 0)}(x, \cdot, \cdot)$ by directly solving the integral equation (ref). The same is true for $m^{(1, 1)}(x, \cdot, \cdot)$. In this sense, for identification of $m^{(0, 0)}(x, \cdot, \cdot)$ and $m^{(1, 1)}(x, \cdot, \cdot)$, there is a trade-off between their nonparametric identifiability and the exclusion restriction on the equilibrium selection. }
remark[\citemain{lee2018identifying}] A major distinction between \citemain{lee2018identifying} and our study is that they allow the strategic effect to depend on unobservable factors by assuming that the error term $V_j$ has the form of $V_j = V_j^0 + D_{-j} (V_j^1 - V_j^0)$, where $V_j^0$ denotes the unobserved payoff determinant when only player $j$ is treated, and $V_j^1 - V_j^0$ corresponds to the unobservable strategic interaction effect. To deal with the multiple equilibria, similarly to us, they assume that one of the equilibria is selected if an unobserved variable $\bar V$ does not exceed a threshold $\bar P$. All possible treatment assignments are then characterized uniquely by the set of threshold crossing rules determined by $\mathbf{V} = (V_1^0, V_1^1, V_2^0, V_2^1, \bar V)$ and the threshold variables $\mathbf{P} = (P_1^0, P_1^1, P_2^0, P_2^1, \bar P)$. Then, \citemain{lee2018identifying} showed that \begin{align*} \mathbb E [Y_1^{(d_1, d_2)} | X=x, \mathbf{V} = \mathbf{p}] = \frac{\partial_{\mathbf{p}} \mathbb E [I^{(d_1, d_2)} Y_1 | X = x , \mathbf{P} = \mathbf{p}]}{\partial_{\mathbf{p}} \Pr[D = (d_1, d_2) | X=x, \mathbf{P} = \mathbf{p}]}, \end{align*} where $\partial_{\mathbf{p}} \coloneqq \partial^5 / (\partial p_1^0 \partial p_1^1 \partial p_2^0 \partial p_2^1 \partial \bar p)$. This MTR parameter includes richer information related to unobserved heterogeneity than ours in that it can reveal unobserved heterogeneity with respect to the unobservable interaction effect, preference on equilibrium selection, and payoff determinants. To estimate this MTR, $\mathbf{P}$ must be jointly continuously distributed, which is more demanding than Assumption (ref)(ii). In addition, the estimator resulting from this identification result would suffer from the curse of dimensionality, whereas our estimator does not at the cost of less-flexible interaction structure. Indeed, \citemain{lee2018identifying} is concerned mainly with the identification of MTR and not its estimation, whereas we focus on both.

Identification of the marginal treatment effects

\phantomsection\Copy{R1-2-1}{ Given Theorems (ref) and (ref), identification of MTE is straightforward. For example, $\text{MTE}_{\text{direct}}^{(0)}(x, \cdot, \cdot) = m^{(1,0)}(x, \cdot, \cdot) - m^{(0, 0)}(x, \cdot, \cdot)$ is identifiable on the intersection of $\mathcal{S}_x^{(1,0)}$ and $\overline{\mathcal{S}}_x$ because $m^{(1, 0)}(x, \cdot, \cdot)$ and $m^{(0, 0)}(x, \cdot, \cdot)$ are identified on the respective regions, meaning that the identifiable region of $\text{MTE}_{\text{direct}}^{(0)}(x, \cdot, \cdot)$ reduces to $\mathcal{S}_x^{(1,0)}$ as $\mathcal{S}_x^{(1,0)} \subseteq \overline{\mathcal{S}}_x$. As another example, the identifiable region of $\text{MTE}_{\text{total}}(x, \cdot, \cdot) = m^{(1, 1)}(x, \cdot, \cdot) - m^{(0, 0)}(x, \cdot, \cdot)$ corresponds to $\overline{\mathcal{S}}_x$, on which both MTR functions are identified. We can characterize the identifiable regions of the other MTE parameters analogously. }

\phantomsection\Copy{R1-1-5}{ As a result of Theorem (ref), $m^{(d_1, d_2)}(x, p_1, p_2)$ for $(d_1, d_2) \in \{ (0, 0), (1, 1) \}$ can be “over-identified” if $(p_1, p_2) \in \underline{\mathcal{S}}_x$, where

align*[align* omitted — 135 chars of source]

Thus, the MTE parameters are also over-identifiable if $\underline{\mathcal{S}}_x$ is non-empty, and this over-identification result can be used to improve the estimation efficiency (see Subsection (ref)). }

remark[Partial identification of MTE without equilibrium selection] As shown above, the MTR functions for $(d_1, d_2) \in \{ (1, 0), (0, 1) \}$ can be identified without using the equilibrium selection assumption. When no assumptions are imposed on the equilibrium selection, the equilibrium selection probability can take any value on $[0,1]$. Thus, as can be seen in Theorem (ref), the identified sets of the MTRs for $(d_1, d_2) \in \{ (0, 0), (1, 1) \}$ are typically unbounded. However, because the equilibrium selection probability is non-negative, we can still identify the signs of these MTR functions. Therefore, for example, if $m^{(1,0)}(x, p_1, p_2)$ has a positive value and the sign of $m^{(0,0)}(x, p_1, p_2)$ is non-positive, we can obtain an informative lower bound of $\text{MTE}_{\text{direct}}^{(0)}(x, p_1, p_2)$ based on the inequality $\text{MTE}_{\text{direct}}^{(0)}(x, p_1, p_2) \ge m^{(1,0)}(x, p_1, p_2)$. A more promising approach would be to introduce some shape restrictions, as in \citemain{balat2020multiple}, to derive informative bounds on $m^{(0,0)}(x, p_1, p_2)$.

Estimation and Asymptotics

In this section, we propose a two-step semiparametric procedure for estimating the MTE parameters given the data $\{\{(Y_{ji}, D_{ji}, W_{ji})\}_{j=1}^2\}_{i=1}^n$ are observed. In the subsequent analysis, we consider the following parametric treatment decision model:

assumption\hfil \begin{enumerate}[(i)] • \phantomsection\Copy{R2-4-1}{ For each $j = 1, 2$, $D_{ji} = \mathbf{1}\{ W_{ji}^\top \gamma_0 + D_{-j, i} \cdot \Delta(W_{ji}^\top \gamma_1) \ge \varepsilon_{ji} \}$, where $\gamma = (\gamma_0^\top, \gamma_1^\top)^\top \in \mathbb{R}^{2\mathrm{dim}(W)}$ is a vector of unknown parameters such that $\gamma_0 \neq \gamma_1$, and $\Delta$ is a known positive function. } • $(\varepsilon_1, \varepsilon_2)$ are independent of $W$ and continuously distributed with known strictly increasing marginal CDFs $F_{\varepsilon_1}$ and $F_{\varepsilon_2}$, respectively. Their joint distribution is given by $\Pr[\varepsilon_1 \le a_1, \varepsilon_2 \le a_2] = H_\rho(F_{\varepsilon_1}(a_1), F_{\varepsilon_2}(a_2))$, where $\rho \in (\underline{c}, \bar c)$ is an unknown parameter. The copula $H_\rho$ has a density function $h_\rho$. • Assumption (ref) holds with $\lambda_x = \Lambda(\widetilde x^\top \lambda)$, where $\Lambda$ is a known strictly increasing CDF, $\widetilde X$ is a linearly independent subset of $X$, and $\lambda \in \mathbb{R}^{\mathrm{dim}(\widetilde X)}$ is a vector of unknown parameters. \end{enumerate}

In Assumption (ref)(i), we assume $\Delta > 0$ to ensure strategic complementarity. We also assume that the coefficients $\gamma$ are common to both players. \phantomsection\Copy{R2-4-2}{ Note that the assumption $\gamma_0 \neq \gamma_1$ is necessary to identify the interaction effect. If this does not hold, the value of $P_j^1$ is automatically determined once $P_j^0$ is fixed, violating Assumption (ref)(ii). } Assumption (ref)(ii) strengthens Assumptions (ref)(ii) and (ref)(i) by requiring a known marginal CDF of $\varepsilon_j$ and full independence between $\varepsilon$ and $W$. Assuming a known distribution for the unobservable is common in the literature, as nonparametrically identifying the payoff function and the error distribution simultaneously is generally impossible (cf. Section 4.1 of \citealpmain{bajari2010identification}). Furthermore, as shown by \citetmain{khan2018information}, when the marginal CDFs of the errors are unknown, the interaction effect cannot be estimated at the parametric rate in general. Estimating the game parameters at the parametric rate is important for the estimation of the MTE parameters. Assumption (ref)(iii) introduces a parametric specification on the equilibrium selection probability. Identification of this parametric treatment decision model is discussed in detail in Appendix (ref).

For the potential outcome, we assume the following linear model.

assumption\hfil \begin{enumerate}[(i)] • \phantomsection\Copy{AE-9-2}{ For each $j = 1, 2$ and $(d_j, d_{-j}) \in \{ 0, 1 \}^2$, $Y_{ji}^{(d_j, d_{-j})} = X_{ji}^\top \beta_j^{(d_j, d_{-j})} + U_{ji}^{(d_j, d_{-j})}$, where $\beta_j^{(d_j, d_{-j})} \in \mathbb{R}^{\mathrm{dim}(X)}$ is a vector of unknown parameters and $U_{ji}^{(d_j, d_{-j})} \in \mathbb{R}$ is a scalar error term. } • $(\varepsilon, U^{(d_1,d_2)})$ are independent of $W$ for $(d_1, d_2) \in \{0, 1\}^2$. \end{enumerate}

Two-step estimation

\paragraph{First step: Estimation of the treatment decision game.}

In accordance with (ref), we write $V_{ji} = F_{\varepsilon_j}(\varepsilon_{ji})$, $P_{ji}^0(\gamma) = F_{\varepsilon_j}(W_{ji}^\top \gamma_0)$, and $P_{ji}^1(\gamma) = F_{\varepsilon_j}(W_{ji}^\top \gamma_0 + \Delta(W_{ji}^\top \gamma_1))$. Further, let $\theta^* = (\gamma^{* \top}, \lambda^{* \top}, \rho^*)^\top$ be the true value of $\theta = (\gamma^\top, \lambda^\top, \rho)^\top$. For a given $\theta$, the conditional probability that the $i$-th pair of players is in the multiple equilibria region is given by

align*[align* omitted — 238 chars of source]

Then, letting $\mathcal{L}_i^{(d_1, d_2)}(\theta)$ be the probability that they choose an action $D_i = (d_1, d_2)$, we have

equation*[equation* omitted — 512 chars of source]

The ML estimator $\widehat \theta$ of $\theta^*$ is obtained by maximizing the log-likelihood $\sum_{i}\sum_{d_1, d_2} I_i^{(d_1, d_2)} \log \mathcal{L}_i^{(d_1, d_2)}(\theta)$ with respect to $\theta$.

Given this $\widehat \theta$, we can estimate $P_{ji}^0 = P_{ji}^0(\gamma^*)$ and $P_{ji}^1 = P_{ji}^1(\gamma^*)$ by $\widehat P_{ji}^0 = P_{ji}^0(\widehat \gamma)$ and $\widehat P_{ji}^1 = P_{ji}^1(\widehat \gamma)$, respectively. We denote the estimator of $\mathbf{P} = (P_1^0, P_1^1, P_2^0, P_2^1)$ by $\widehat{\mathbf{P}} = (\widehat P_1^0, \widehat P_1^1, \widehat P_2^0, \widehat P_2^1)$. Similarly, the true equilibrium selection probability $\lambda_x^* = \Lambda(\widetilde x^\top \lambda^*)$ and the true joint CDF and density of $(V_1, V_2)$, which we denote by $H = H_{\rho^*}$ and $h = h_{\rho^*}$, can be respectively estimated by $\widehat \lambda_x = \Lambda(\widetilde x^\top \widehat \lambda)$, $\widehat H = H_{ \widehat \rho}$, and $\widehat h = h_{ \widehat \rho}$. Moreover, we can estimate the true choice probability $\mathcal{L}_i^{(d_1, d_2)} = \mathcal{L}_i^{(d_1, d_2)}(\theta^*)$ by $\widehat{\mathcal{L}}_i^{(d_1, d_2)} = \mathcal{L}_i^{(d_1, d_2)}(\widehat \theta)$.

remark[Constant equilibrium selection probability] \phantomsection\Copy{R2-2-2}{ Since the equilibrium selection probability is estimated using only the observations in multiple equilibria, the estimation of $\lambda_x$ is very challenging in practice when $\text{dim}(\widetilde X)$ is not small or the strategic effects are so weak that only a small number of observations are in the multiple equilibria region. Indeed, when we tried to estimate the function $\lambda_x$ in our empirical application, we could not obtain meaningful estimate. To circumvent this issue in practical situations with moderate sample size, we suggest simplifying Assumption (ref)(iii) such that $\lambda_x = \lambda$ for a scalar parameter $\lambda \in [0, 1]$. }

\paragraph{Second step: Estimation of the MTE.}

As discussed in Subsection (ref), a variety of treatment effect parameters can be identified. We here specifically discuss the estimation of $\text{MTE}_{\text{direct}}^{(0)}(x, p_1^0, p_2^1)$ at point A in Figure (ref) (the estimation at the other points is analogous). The estimation of the total MTE is relegated to Appendix (ref).

Recall that $\text{MTE}_{\text{direct}}^{(0)}(x, p_1^0, p_2^1) = m^{(1, 0)}(x, p_1^0, p_2^1) - m^{(0, 0)}(x, p_1^0, p_2^1)$. First, we discuss how to estimate $m^{(1, 0)}(x, p_1^0, p_2^1)$. Assumption (ref) implies that

align*[align* omitted — 231 chars of source]

where $g^{(1, 0)}(p_1^0, p_2^1) \coloneqq \int_{0}^{p_1^0} \int_{p_2^1}^1 \mathbb E [U_{1}^{(1, 0)}| V_1= v_1, V_2= v_2] h(v_1, v_2) \mathrm{d}v_1 \mathrm{d}v_2$. Further, observe that

align*[align* omitted — 185 chars of source]

A similar argument to Theorem (ref) gives $\mathbb E [ U_{1}^{(1, 0)}| V_1\le p_1^0, V_2> p_2^1] = g^{(1, 0)}(p_1^0, p_2^1) / \mathcal{L}^{(1,0)}(p_1^0, p_2^1)$, where $\mathcal{L}^{(1, 0)}(p_1^0, p_2^1) = p_1^0 - H(p_1^0, p_2^1)$. Then, we have the following partially linear regression model:

align[align omitted — 136 chars of source]

where $T^{(1, 0)} \coloneqq I^{(1, 0)} / \mathcal{L}^{(1, 0)}$, and $\mathbb E [e^{(1, 0)} | I^{(1, 0)}, X, P_1^0, P_2^1] = 0$ by construction. Here, note that since $\mathcal{L}^{(1, 0)}$ may take an arbitrarily small value close to zero, the weight term $T^{(1, 0)}$ can be extremely large for some observations. To deal with this issue, we exclude such observations from the analysis. Specifically, we introduce a non-negative smoothed indicator function $\tau_\varpi(\mathcal{L}^{(1, 0)})$ such that (i) $\tau_\varpi:[0, 1] \to [0, 1]$ is a non-decreasing function and $\tau_\varpi(a) = 0$ if and only if $a < \varpi$ for a small constant $\varpi > 0$, and (ii) $\tau_\varpi$ is continuously differentiable with a bounded derivative. Multiplying both sides of (ref) by $\tau_\varpi(\mathcal{L}^{(1, 0)})$ gives

align[align omitted — 180 chars of source]

where $\widetilde I^{(1, 0)} \coloneqq \tau_\varpi(\mathcal{L}^{(1, 0)}) I^{(1, 0)}$, $\widetilde T^{(1, 0)} \coloneqq \widetilde I^{(1, 0)} / \mathcal{L}^{(1,0)}$, and $\widetilde e^{(1, 0)} \coloneqq \tau_\varpi(\mathcal{L}^{(1, 0)}) e^{(1, 0)}$. Since $\mathcal{L}^{(1, 0)}$ is a function of $P_1^0$ and $P_2^1$, $\mathbb E [\widetilde e^{(1, 0)} | I^{(1, 0)}, X, P_1^0, P_2^1] = 0$ still holds.

Based on (ref), we consider estimating $\beta_1^{(1, 0)}$ and $g^{(1,0)}$ using the series (sieve) method. Let $b_K(\cdot, \cdot) = (b_{1K}(\cdot, \cdot), \dots, b_{KK}(\cdot, \cdot))^\top$ be a $K \times 1$ vector of bivariate basis functions. We assume that $g^{(1, 0)}$ can be well-approximated by $g^{(1, 0)}(\cdot, \cdot) \approx b_K(\cdot, \cdot)^\top \alpha^{(1,0)}$ for some coefficient vector $\alpha^{(1, 0)}$ with sufficiently large $K$. Then, letting $\widehat I_i^{(1, 0)} \coloneqq \tau_\varpi(\widehat{\mathcal{L}}_i^{(1, 0)}) I_i^{(1, 0)}$ and $\widehat T_i^{(1, 0)} \coloneqq \widehat I_i^{(1, 0)} / \widehat{\mathcal{L}}_i^{(1,0)}$, we have

align*[align* omitted — 216 chars of source]

Let $(\widehat \beta_1^{(1,0)}, \widehat \alpha^{(1,0)})$ be the least squares (LS) estimator of $(\beta_1^{(1,0)}, \alpha^{(1,0)})$ obtained by regressing $\widehat I_i^{(1, 0)} Y_{1i}$ on $(\widehat I_i^{(1, 0)} X_{1i}, \widehat T_i^{(1, 0)} b_K(\widehat P_{1i}^0, \widehat P_{2i}^1) )$. Then, the estimator of $g^{(1, 0)}(p_1^0, p_2^1)$ can be obtained by $\widehat g^{(1, 0)}(p_1^0, p_2^1) \coloneqq b_K(p_1^0, p_2^1)^\top \widehat \alpha^{(1, 0)}$, and we can estimate $\mathbb E [U_1^{(1, 0)} | V_1 = p_1^0, V_2 = p_2^1]$ by

align*[align* omitted — 303 chars of source]

where $\ddot{b}_K(p_1, p_2) \coloneqq \partial_{p_1 p_2} [ b_K(p_1, p_2) ]$. Finally, we can estimate $m^{(1, 0)}(x, p_1^0, p_2^1)$ by

align[align omitted — 174 chars of source]

Next, we describe the estimation of $m^{(0, 0)}(x, p_1^0, p_2^1) = x_1^\top \beta_1^{(0, 0)} + \mathbb E [U_1^{(0, 0)} | V_1 = p_1^0, V_2 = p_2^1]$. First, observe that

align*[align* omitted — 186 chars of source]

The same argument as in Theorem (ref) gives

align[align omitted — 729 chars of source]

where $\mathcal{L}^{(0, 0)}(\mathbf{P})$ is the conditional probability of $D = (0,0)$ given $\mathbf{P}$, and $g_l^{(0,0)}$'s, $l = 1, \ldots, 4$, are bivariate real-valued functions.\footnote{ Specifically, $g_1^{(0,0)}(p_1^0, p_2^0) = \int_{p_1^0}^1 \int_{p_2^0}^1 \mathbb E [U_{1}^{(0, 0)}| V_1= v_1, V_2= v_2] h(v_1, v_2) \mathrm{d}v_1 \mathrm{d}v_2$, and $g_2^{(0,0)}(p_1^1, p_2^0)$ and $g_3^{(0,0)}(p_1^0, p_2^1)$ are obtained by replacing $(p_1^0, p_2^0)$ in the right-hand side with $(p_1^1, p_2^0)$ and $(p_1^0, p_2^1)$, respectively. Moreover, $g_4^{(0,0)}(p_1^1, p_2^1) = -\int_{p_1^1}^1 \int_{p_2^1}^1 \mathbb E [U_{1}^{(0, 0)}| V_1= v_1, V_2= v_2] h(v_1, v_2) \mathrm{d}v_1 \mathrm{d}v_2$. } Hence, similarly to (ref), we obtain the following regression model:

align[align omitted — 437 chars of source]

where $\widetilde I^{(0, 0)} \coloneqq \tau_\varpi(\mathcal{L}^{(0, 0)}) I^{(0, 0)}$, $\widetilde T^{(0, 0)} \coloneqq \widetilde I^{(0, 0)} / \mathcal{L}^{(0,0)}$, and $\mathbb E [\widetilde e^{(0, 0)} | I^{(0, 0)}, X, \mathbf{P}] = 0$. Assuming again that each $g_l^{(0,0)}$, $l = 1, \ldots , 4$, can be approximated by $g_l^{(0,0)}(\cdot, \cdot) \approx b_K(\cdot, \cdot)^\top \alpha_l^{(0,0)}$ and replacing $\widetilde I^{(0,0)}$, $\widetilde T^{(0,0)}$, $\lambda_X^*$, and $\mathbf{P}$ with their estimators $\widehat I^{(0,0)} \coloneqq \tau_\varpi(\widehat{\mathcal{L}}^{(0, 0)}) I^{(0, 0)}$, $\widehat T^{(0, 0)} \coloneqq \widehat I^{(0, 0)} / \widehat{\mathcal{L}}^{(0,0)}$, $\widehat \lambda_X$, and $\widehat{\mathbf{P}}$, respectively, $\beta_1^{(0, 0)}$ and $\alpha_l^{(0,0)}$'s can be estimated by LS regression.\footnote{ It is possible to use different orders of basis terms to approximate each component of the functions $g_l^{(0,0)}$'s for $l = 1, \ldots, 4$, but we use the same order $K$ for all, for simplicity. Also note that the “locations” of the functions $g_l^{(0,0)}$'s are not identified without further restrictions. To simplify our presentation, we postulate that an appropriate location normalization is made implicitly. } Let $\widehat \beta_1^{(0, 0)}$ and $\widehat \alpha_l^{(0,0)}$ be the resulting LS estimators. Then, each $g_l^{(0,0)}(p_1, p_2)$ can be estimated by $\widehat g_l^{(0, 0)}(p_1, p_2) \coloneqq b_K(p_1, p_2)^\top \widehat \alpha_l^{(0, 0)}$. Moreover, (ref) implies that

align*[align* omitted — 145 chars of source]

and thus we can estimate this by

align*[align* omitted — 301 chars of source]

Consequently, $m^{(0, 0)}(x, p_1^0, p_2^1)$ can be estimated by

align[align omitted — 174 chars of source]

Finally, $\text{MTE}_{\text{direct}}^{(0)}(x, p^0_1, p^1_2)$ can be estimated by

align*[align* omitted — 184 chars of source]

Asymptotics

In this subsection, we discuss the asymptotic properties of the proposed estimators. We mainly investigate the MTR estimator for $D = (1,0)$. Analogous arguments apply to the other cases.

For simplicity of presentation, all the technical assumptions are summarized in Appendix (ref). Let us comment on some of the high-level assumptions employed here. \phantomsection\Copy{R2-1-e}{ First, in Assumption (ref), we assume that $\{\{(Y_{ji}, D_{ji}, W_{ji})\}_{j=1}^2\}_{i=1}^n$ are independent and identically distributed across $i$. Although we do not restrict the dependence between the variables for $j$ and those for $-j$, the dependence across different pairs is ruled out. We conjecture that if the dependence is sufficiently weak, we can establish similar results to those given below (cf. \citealpmain{chen2015optimal, lee2016series}). In addition, if there is pair-specific unobserved heterogeneity, the assumption of identical distribution might be questionable. To mitigate this problem, in our empirical application, we include school fixed effects in the model. } Assumption (ref)(i) assumes the $\sqrt{n}$-consistency of the ML estimator. In Assumption (ref), we require that the functions $g^{(1,0)}$ and $b_K$ are sufficiently smooth so that they are at least $s$-times continuously differentiable for some $s \ge 2$ and that the derivatives of $g^{(1,0)}$ can be well approximated in the space spanned by the derivatives of the basis $b_K$. Lastly, Assumption (ref) requires that the linear projection of any bounded and continuous function onto $(\widetilde I_i^{(1,0)} X_{1i}, \widetilde T_i^{(1,0)} b_K (P_{1i}^0, P_{2i}^1))$ is “stable” in a sense that its sup-operator norm is stochastically bounded. Similar conditions can be found, for example, in \citetmain{huang2003local} and \citetmain{chen2018optimal}.

The next theorem establishes the uniform convergence rate of our MTR estimator $\widehat m^{(1, 0)}(x, p_1, p_2)$ and that of the “oracle” estimator $\widetilde m^{(1, 0)}(x, p_1, p_2)$ obtained by assuming that the true parameters in the first stage are known (see Appendix (ref) for more precise definition). Denote $\mathcal{S}^{(1,0)} \coloneqq \text{supp}[P_1^0, P_2^1]$.

theoremSuppose that Assumptions (ref), (ref), (ref), (ref), (ref)(i), and (ref)--(ref) hold. Then, for a given $x \in \text{supp}[X]$, we have \begin{align*} \begin{array}{cl} (i) & \sup_{(p_1, p_2) \in \mathcal{S}^{(1,0)}}\left| \widetilde m^{(1, 0)}(x, p_1, p_2) - m^{(1, 0)}(x, p_1, p_2) \right| = O_P(\zeta_0(K) K \sqrt{\log n / n} ) + O_P(K^{(2 - s) /2}),\\ (ii) & \sup_{(p_1, p_2) \in \mathcal{S}^{(1,0)}}\left| \widehat m^{(1, 0)}(x, p_1, p_2) - m^{(1, 0)}(x, p_1, p_2) \right| = O_P(\zeta_0(K) K \sqrt{\log n / n} ) + O_P(K^{(2 - s) /2}), \end{array} \end{align*} where $\zeta_0(K) \coloneqq \sup_{(p_1, p_2) \in [0,1]^2 }\| b_K(p_1, p_2) \|$ with $\| \cdot \|$ being the Euclidean norm.

The proof of the theorem is straightforward from Lemmas (ref) and (ref), and thus it is omitted. When $\zeta_0(K) \asymp \sqrt{K}$, by choosing $K \asymp (\log n /n)^{-1/(1 + s)}$, we can obtain \[ \sup_{(p_1, p_2) \in \mathcal{S}^{(1,0)}}\left| \widehat m^{(1, 0)}(x, p_1, p_2) - m^{(1, 0)}(x, p_1, p_2) \right| = O_P\left(\left(\log n / n\right)^{\frac{s - 2}{2s + 2}} \right). \] The above result implies that our MTR estimator can converge at the optimal uniform rate of \citetmain{stone1982optimal}.

The next theorem shows that the asymptotic distribution of the feasible estimator is also equivalent to that of the infeasible oracle estimator.

theoremSuppose that Assumptions (ref), (ref), and (ref)--(ref) hold. For given $x \in \text{supp}[X]$ and $(p_1^0, p_2^1) \in \mathcal{S}^{(1,0)}$, if in addition $K / \| \ddot{b}_K(p_1^0, p_2^1) \| \to 0$, $\zeta_0(K) \sqrt{K/n} \to 0$, and $\sqrt{n} K^{(2 - s)/2} = O(1)$ hold, then we have \begin{align*} (i) & \qquad \frac{\sqrt{n} \left( \widetilde m^{(1, 0)}(x, p_1^0, p_2^1) - m^{(1, 0)}(x, p_1^0, p_2^1) \right)}{\sigma_K^{(1,0)}(p_1^0, p_2^1)} \overset{d}{\to} N(0, 1),\\ (ii) & \qquad \frac{\sqrt{n} \left( \widehat m^{(1, 0)}(x, p_1^0, p_2^1) - m^{(1, 0)}(x, p_1^0, p_2^1) \right)}{\sigma_K^{(1,0)}(p_1^0, p_2^1)} \overset{d}{\to} N(0, 1), \end{align*} where the definition of $\sigma_K^{(1,0)}(p_1^0, p_2^1)$ can be found in (ref).

The standard deviation $\sigma_K^{(1,0)}(p_1^0, p_2^1)$ can be easily estimated by a sample analog, replacing the true values and functions with their estimates.

Let us briefly discuss the asymptotic properties of the MTR estimator for $D=(0,0)$ given in (ref). Under conditions similar to those in Theorem (ref), we can show the following asymptotic normality result:

align*[align* omitted — 173 chars of source]

for $(p_1^0, p_2^1) \in \mathcal{S}^{(1,0)}$, where the definition of $\sigma_{K}^{(0,0)}(p_1^0, p_2^1)$ can be found in (ref).

Finally, the limiting distribution of $\widehat{\text{MTE}}_{\text{direct}}^{(0)}(x, p_1^0, p_2^1)$ can be characterized as follows:

align*[align* omitted — 289 chars of source]

The limiting distributions of the other MTE estimators can be derived similarly.

An over-identified estimator

As mentioned above, we can improve the estimation efficiency based on the over-identification result. For $(p_1, p_2) \in \underline{\mathcal{S}}$, where $\underline{\mathcal{S}} \coloneqq \bigcap_{(d_1, d_2) \in \{ 0, 1 \}^2} \text{supp}[P_1^{d_1}, P_2^{d_2}]$, Theorem (ref) suggests that we can estimate $m^{(0, 0)}(x, p_1, p_2)$ in four ways: $\widehat m_l^{(0, 0)}(x, p_1, p_2) \coloneqq x_1^\top \widehat \beta_1^{(0, 0)} + \widehat \mathbb E_{ln}[ U_1^{(0, 0)} | V_1 = p_1, V_2 = p_2 ]$ for $l = 1, \dots, 4$, where $\widehat \mathbb E_{ln} [ U_1^{(0, 0)} | V_1 = p_1, V_2 = p_2 ] \coloneqq \widehat \kappa_l \ddot b_K(p_1, p_2)^\top \widehat \alpha_l^{(0, 0)}$ with $\widehat \kappa_1 = \widehat \kappa_2 = \widehat \kappa_3 \coloneqq 1 / \widehat h(p_1, p_2)$ and $\widehat \kappa_4 \coloneqq - 1 / \widehat h(p_1, p_2)$.\footnote{ \phantomsection\Copy{R2-5}{ The over-identification result essentially builds on the continuity of strategic effects, which can be examined by the first-stage estimation result. Then, if the continuity is verified, the estimated $\widehat m_l^{(0, 0)}(x, p_1, p_2)$, $l = 1, \ldots, 4$, must all take a “close” value at each $(p_1, p_2)$ in the common support. This gives us a testable implication for our model specification. Investigating the property of this testing procedure is quite intriguing, but we leave this for future work. } } Let $w = (w_1, w_2, w_3, w_4)^\top \in [0,1]^4$ be a vector of fixed weights such that $\sum_{l = 1}^4 w_l = 1$. Then, $m^{(0, 0)}(x, p_1, p_2)$ can be estimated by the weighted average of the four estimators:

align[align omitted — 152 chars of source]

In the same manner as in the proof of Theorem (ref), we can show

align*[align* omitted — 214 chars of source]

By easy calculations, we can see that the variance $[\sigma_{\text{over-id}, K}^{(0, 0)}(p_1, p_2; w)]^2$ has a quadratic form with respect to $w$, and, thus, the “optimal” weights that attain the minimum variance in the class of the weighted average estimators $\widehat m_{\text{over-id}}^{(0, 0)}(x, p_1, p_2; w)$ can be easily found by quadratic programming.\footnote{ Specifically, $[ \sigma_{\text{over-id}, K}^{(0, 0)}(p_1, p_2; w) ]^2 = w^\top \Upsilon w$, and $\Upsilon = (\Upsilon_{lm})$ is the $4 \times 4$ symmetric matrix with the $(l, m)$-th element given by $\Upsilon_{lm} \coloneqq \kappa_l \kappa_m \ddot b_K(p_1, p_2)^\top \mathbb{S}_{lK} [\Psi_K^{(0,0)}]^{-1} \Sigma_K^{(0, 0)} [\Psi_K^{(0,0)}]^{-1} \mathbb{S}_{mK}^\top \ddot b_K(p_1, p_2)$, where $\kappa_l$ is the true value of $\widehat \kappa_l$, and $\mathbb{S}_{lK}$, $\Psi_K^{(0,0)}$, and $\Sigma_K^{(0, 0)}$ are as in (ref). } Although $[\sigma_{\text{over-id}, K}^{(0, 0)}(p_1, p_2; w)]^2$ is unknown in practice, we can estimate the optimal weights by minimizing its sample analog.

\phantomsection\Copy{R1-2-2}{ It should be noted that the over-identified estimator relies on the assumption of non-empty $\underline{\mathcal{S}}$. If $\underline{\mathcal{S}}$ is empty, there is at least one estimator among the four that does not contribute to the estimation of $m^{(0, 0)}(x, p_1, p_2)$. In this situation, the over-identified estimator may entail a bias arising from an inaccurate extrapolation. Thus, it is always recommended to check the joint distribution of $(\widehat P_1^{d_1}, \widehat P_2^{d_2})$ before computing $\widehat m_{\text{over-id}}^{(0, 0)}(x, p_1, p_2; w)$ to see if $(p_1, p_2) \in \underline{\mathcal{S}}$ actually holds. }

\phantomsection\Copy{R2-7-b-1}{ Finally, as mentioned by a referee, we can consider other estimation strategies than the above-mentioned over-identified estimator, which also use the over-identification result but in a more implicit way. As far as we have checked via numerical simulations, the performances of these estimators, including ours, are similar. To save space, we discuss the constructions of these alternative estimators in Appendix (ref). }

Numerical Illustrations

Monte Carlo simulation results

In order to evaluate the finite sample performance of our MTE estimator, we conduct a set of Monte Carlo simulations based on the over-identified estimator (ref). To save space, we here briefly report the summary of the results, and detailed information about the experiments are provided in Appendix (ref).

The simulation results provide us with several practical guidelines as follows. First, when comparing bivariate power series and tensor-product B-splines for the basis functions, we find that the power-series-based estimator achieves smaller RMSE (root mean squared error) than the B-splines-based estimator, especially when the sample size is not large. For this result, note that the power series is a globally supported series, and thus the power-series estimator tends to have smaller variance at the cost of lower flexibility than the B-splines estimator. Theoretically speaking, B-splines estimators have a better approximation property than power-series estimators, and in fact the latter may not attain the optimal convergence rate. Hence, the above result may be due to the chosen sample sizes in this simulation setup. Second, in terms of RMSE, the MTE parameters can be more precisely estimated by ridge regression with a small regularization parameter. Remarkably, using the ridge regression often reduces the bias for the B-splines estimator. Since the tensor-product B-splines involves a large number of basis terms, which causes a severe multicollinearity problem for the non-regularized estimators, there is a higher risk of generating extremely outlying estimates compared with the regularized estimators. Thus, it would be desirable to employ a regularized regression in practical situations with moderate sample size. Note that adding a sufficiently small regularization factor does not alter our asymptotic theory.

Empirical application: Adolescent delinquency and academic performance

As an empirical illustration, we estimate the impacts of risky behaviors such as smoking and drinking alcohol on adolescents' academic performance. The empirical analysis is performed on the National Longitudinal Study of Adolescent Health (Add Health) data that provide nationally representative information on 7--12th graders in the U.S. The survey was conducted during the 1994--1995 school year. The survey elicits information on variables such as the social and demographic characteristics of the respondents, education levels and occupations of their parents, and friendship connections.

In this analysis, following \citemain{card2013peer}, we focus on the interactions among pairs of closest opposite-gender friends. In the survey, each respondent was asked to list up to five friends of each gender in order from best to 5th best friend. We first exclude missing nomination data (caused by non-response or because the nominee was not included in the Add Health data) and identify each respondent's closest opposite-gender friend in the available dataset. Since a student's closest friend's friends might not include the student him/herself, we consider that the pair is formed only when his/her best friend nominates him/her either in the first place or second place. This procedure results in 7,631 opposite-gender pairs of students.

\phantomsection\Copy{AE-17}{ Before proceeding, let us clarify two important caveats in interpreting the results of the following analysis. First, as previously mentioned, the results presented are all conditional on the already formed pairs. In reality, it is often observed that people with similar profiles are more likely to be friends with each other. Such homophilic preference generally poses an additional endogeneity problem, making it difficult to generalize the results beyond the current data. Second, in reality, not only the interference within best friends, but also that between different friend groups may affect the students' activities. This type of interaction cannot be accommodated in our framework. }

The treatment variable ($D$) is defined as the student's participation in risky behaviors including smoking, drinking, truancy, and fighting, and the outcome variable ($Y$) is the natural log of GPA. Table (ref) summarizes the explanatory variables used ($X$) and their definitions. For continuous IVs ($Z$) for each gender, we construct them based on the average parental characteristics of the students' fourth and fifth ranked non-best friends. For our method to work with this choice of IV, the following two identification assumptions must be met: (1) the characteristics of non-best friends affect the students' delinquency, and (2) the non-best friends' characteristics exert no direct effect on the students' academic performance. The arguments supporting these assumptions are presented below.

For assumption (1), in the econometrics literature, \citemain{lee2014binary} and \citemain{lin2014peer} present empirical evidence that the personalities of friends, including non-best friends, influence participation in risky behaviors. In the education and sociology literature, several studies reveal a close relation between student delinquency and the school environment, especially the shared perception of social norms (e.g., \citealpmain{paluck2016changing}). Because the school environment for a student is created by all school mates, not only by one's own best friends, this point is consistent with our assumption. \citemain{rees2011one} also point out that a kind of social pressure exists in youth friendships to achieve peer acceptance, which drives the students to engage in risky behavior in groups. Here, the role of non-best friend relationships particularly matters because non-best friends have less intimate and caring relationships than best friends have (p.199, \citealpmain{rees2011one}).

Next, for assumption (2), several studies indicate the existence of positive interactions in terms of academic outcomes between reciprocal best friends (e.g., \citealpmain{vaquera2008you,cherng2013along}). For example, \citemain{vaquera2008you} argue that reciprocal friendships are likely to be more emotionally supportive as well as superior information resources compared to friendships that are not reciprocal. Because our definition of non-best friends does not involve reciprocity, in view of the above described studies, it would be reasonable to assume that the non-best friend characteristics do not directly affect the student's academic performance.\footnote{ \phantomsection\Copy{AE-19-1}{ To verify whether this exclusion restriction assumption is actually supported in our data, we have performed the test proposed by \citemain{kedagni2020generalized}. The results are reported in Table (ref) in Appendix (ref), which indicate that our chosen IVs are likely to be valid. However, note that this testing procedure is not a formal diagnosis for our specific setting, and that developing such a formal test is beyond the scope of our study. } }

After excluding observations with missing values for $(D,X,Z)$, the size of the sample used in the estimation of the treatment decision model is 6,053. When estimating the MTRs, observations with missing values for $Y$ are further excluded. Table (ref) presents the empirical distribution of $(D_1, D_2)$, where male students are labeled as “player 1” and females as “player 2”. As expected, the case of $D = (0,0)$ shows the largest share of our sample. There is an interesting asymmetry between the cases $D = (1,0)$ and $D = (0,1)$; that is, the number of males who solely participate in delinquent behaviors is significantly larger than the number of such females.

\begingroup

table[table omitted — 484 chars of source]

\endgroup

To save space, the detailed estimation results of the treatment decision model are omitted here and are provided in Table (ref). Some noteworthy findings are as follows. First, approximately three-quarters of the pairs choose $D = (0,0)$ in the multiple equilibria situation. Second, the education levels of the parents (the father's level of education, in particular) have a significant impact on reducing the risky activities of the students. Third, having friends whose mothers have a professional job tends to discourage risky behavior. Fourth, white and Asian students are more likely to be influenced by their friends' behaviors than other students. Finally, students who belong to a sports club are less susceptible to their friends' behaviors. The last two findings clearly indicate the presence of heterogeneous peer effects. Although such heterogeneity is often overlooked in the literature on structural game estimation of adolescent activities (e.g., \citealpmain{soetevent2007discrete, card2013peer, lee2014binary}), it is consistent with the results in other contexts (e.g., \citealpmain{eisenberg2014peer, hsieh2018smoking}).

Figure (ref) in Appendix (ref) presents the histograms of the estimated $(\pi_1^0(W_1), \pi_2^0(W_2))$ and $(\Delta_1(W_1), \Delta_2(W_2))$, which indicate that these variables are distributed over some bounded range and, thus, that estimating the whole functional form of the MTEs would be challenging. To see this point more concretely, we provide the histograms of the estimated $(P_j^0, P_j^1)$ for $j=1$ (male) and $2$ (female) in Figure (ref). These figures suggest that the overlapping support condition can hardly hold outside the interval $[0.2, 0.7]$, implying that MTE estimates might not be reliable when $V$ is outside this interval.

Now, we present our main results of estimating the MTE parameters. Since our sample size is not very large, following the suggestion from the Monte Carlo results, we employ the over-identified estimator (ref) with a third-order bivariate power series for the basis function and use ridge regression for the parameter estimation with penalty equal to $n^{-1}$. Figure (ref) summarizes the estimated direct, total, and indirect MTEs for both genders. For comparison, we also report the MTE estimates obtained from a model where the first-stage treatment decision model is a standard bivariate probit model with no strategic interactions. The MTE estimator for this model can be constructed in exactly the same way as our two-step estimator based on the identification result in Appendix (ref). For the value of $X$, the MTEs are evaluated at the median over all observations of men and women altogether.

Figure (ref) shows that the direct impact of their own delinquent activities is significantly negative for both male and female students. The direct MTE for female students tends to have a weak negative slope with respect to the male partner's $V$. That is, when the male partner has a stronger hesitation in participating in risky activities, the direct MTE on female students' GPA becomes larger. For male students, their direct MTE is relatively flat with respect to the female partner's $V$ (partially due to regularization). Figure (ref) illustrates that the total MTE is also significantly negative and is larger than the direct MTE for both male and female students. The gap between the direct and total MTE corresponds to indirect MTE, which suggests that the students' own GPA is indirectly influenced by their peers' delinquent behavior, as shown in Figure (ref).\footnote{ Although the asymptotic distribution of the indirect MTE estimator is not formally discussed in this study, it can be easily derived by combining the results of Theorem (ref) and that in Appendix (ref). } As expected, the magnitude of the indirect MTE is clearly smaller than that of the direct MTE. Interestingly, in contrast to the direct MTE estimates, we can observe that the indirect MTE estimates have a weak increasing tendency with respect to the best friend's $V$. That is, if the best friend is a person who hesitates in engaging in risky behavior, the indirect effect would be slightly smaller.

The MTE estimates obtained from the non-strategic treatment model are significantly different from those obtained in our game-theoretic model. In particular, the direct and total MTE for female students are clearly overestimated, and the indirect MTEs for male students have opposite sign as compared to the strategic model. These results suggest that the misspecification in the first-stage model may produce certain biases that cannot be easily controlled unless the strategic interactions are actually incorporated into the model.

figure[figure omitted — 533 chars of source]

Conclusion

This study developed treatment effect models that admit both the treatment spillover and strategic interaction in the treatment decisions within a pair of agents. We first demonstrated that the interaction in the treatment decisions can be modeled as a binary game of complete information with multiple equilibria. Assuming an equilibrium selection rule, we showed that the MTE can be identified using an extended version of the LIV method. Based on our identification results, we proposed two-step semiparametric series estimation for MTE parameters. We showed that the proposed MTE estimator is uniformly consistent and asymptotically normal.

The results of this study suggest several extensions that would be promising to investigate, some of which have been mentioned above. In addition, it would be worth extending our results to the case of strategic interaction among more than two players. In the literature on game econometrics, only a few studies have addressed the point identification of game models of complete information with more than two players. The main reason for this is that the characterization of the equilibrium is extremely complicated compared to the case of two-player games; in general, the number of multiple equilibria regions increases with the number of players (see, e.g., \citealpmain{soetevent2007discrete}). Accordingly, the direct applicability of our approach to such cases is unclear. On the other hand, while fixing the number of players at two, extending our model to an ordered treatment setup, as in \citetmain{vytlacil2006ordered}, should be more tractable. As shown in \citetmain{card2013peer}, the Nash equilibrium in the two-player ordered-response game of complete information is uniquely characterized by an equilibrium selection assumption similar to ours. Hence, although computation would be much more complicated than the present case, the identification and estimation approach proposed in this study would be applicable to such models. Finally, it also may be beneficial to develop treatment evaluation techniques for treatment decision games of incomplete information. Under incomplete information, we may use the rational expectation model developed in \citetmain{lee2014binary}, for example, as a treatment decision model, even when the number of players is large. These topics are left for future research.

Acknowledgments

The authors are grateful to the coeditor (Elie Tamer), the associate editor, and two anonymous referees for their insightful comments that significantly improved the paper. We also thank to Yoichi Arai, Sukjin Han, Hiroaki Kaido, Shin Kanaya, Hiroyuki Kasahara, Toru Kitagawa, Tatsushi Oka, Ryo Okui, Yuya Sasaki, Yuta Toyama, Takuya Ura, Haiqing Xu, and the participants of conferences and seminars at various places for their valuable comments and suggestions. This work was supported by JSPS KAKENHI Grant Numbers 15K17039, 17K13715, 19H01473, and 20K01597.

This research uses data from Add Health, a program project directed by Kathleen Mullan Harris and designed by J. Richard Udry, Peter S. Bearman, and Kathleen Mullan Harris at the University of North Carolina at Chapel Hill, and funded by grant P01-HD31921 from the Eunice Kennedy Shriver National Institute of Child Health and Human Development, with cooperative funding from 23 other federal agencies and foundations. Special acknowledgment is due Ronald R. Rindfuss and Barbara Entwisle for assistance in the original design. Information on how to obtain the Add Health data files is available on the Add Health website (http://www.cpc.unc.edu/addhealth). No direct support was received from grant P01-HD31921 for this analysis.

center[center omitted — 370 chars of source]