EconBase
← Back to paper

Multiple Treatments with Strategic Interaction

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

123,375 characters · 11 sections · 71 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Multiple Treatments with Strategic Interaction

\spacing{1.5}

abstractWe develop an empirical framework to identify and estimate the effects of treatments on outcomes of interest when the treatments are the result of strategic interaction (e.g., bargaining, oligopolistic entry, peer effects). We consider a model where agents play a discrete game with complete information whose equilibrium actions (i.e., binary treatments) determine a post-game outcome in a nonseparable model with endogeneity. Due to the simultaneity in the first stage, the model as a whole is incomplete and the selection process fails to exhibit the conventional monotonicity. Without imposing parametric restrictions or large support assumptions, this poses challenges in recovering treatment parameters. To address these challenges, we first establish a monotonic pattern of the equilibria in the first-stage game in terms of the number of treatments selected. Based on this finding, we derive bounds on the average treatment effects (ATEs) under nonparametric shape restrictions and the existence of excluded exogenous variables. We show that instrument variation that compensates strategic substitution helps solve the multiple equilibria problem. We apply our method to data on airlines and air pollution in cities in the U.S. We find that (i) the causal effect of each airline on pollution is positive, and (ii) the effect is increasing in the number of firms but at a decreasing rate. JEL Numbers: C14, C31, C36, C57 Keywords: Heterogeneous treatment effects, strategic interaction, endogenous treatments, average treatment effects, multiple equilibria.

\spacing{1.5}

Introduction

We develop an empirical framework to identify and estimate the heterogeneous effects of treatments on outcomes of interest, where the treatments are the result of agents' interaction (e.g., bargaining, oligopolistic entry, decisions in the presence of peer effects or strategic effects). Treatments are determined as an equilibrium of a game and these strategic decisions of players endogenously affect common or player-specific outcomes. For example, one may be interested in the effects of entry of newspapers on local political behavior, entry of carbon-emitting companies on local air pollution and health outcomes, the presence of potential entrants in nearby markets on pricing or investment decisions of incumbents, the exit decisions of large supermarkets on local health outcomes, or the provision of limited resources when individuals make participation decisions under peer effects and their own gains from the treatment.\footnote{The entry and pollution is our leading example introduced in Section (ref); the other examples are discussed in detail in Appendix (ref).} As reflected in some of these examples, our framework allows us to study the externalities of strategic decisions, such as societal outcomes resulting from firm behavior. Ignoring strategic interaction in the treatment selection process may lead to biased, or at least less informative, conclusions about the effects of interest.

We consider a model in which agents play a discrete game of complete information, whose equilibrium actions (i.e., a profile of binary endogenous treatments) determine a post-game outcome in a nonseparable model with endogeneity. We are interested in the various treatment effects of this model. In recovering these parameters, the setting of this study poses several challenges. First, the first-stage game posits a structure in which binary dependent variables are simultaneously determined in threshold crossing models, thereby, making the model, as a whole, incomplete. This is related to the problem of multiple equilibria in the game. Second, due to this simultaneity, the selection process for each treatment in the profile does not exhibit the conventional monotonic property \`a la imbens1994identification. Furthermore, we want to remain flexible with other components of the model. That is, we make no assumptions on the joint distributions of the unobservables nor parametric restrictions on the player's payoff function and how treatments affect the outcome. In addition, we do not impose any arbitrary equilibrium selection mechanism to deal with the multiplicity of equilibria, nor require that players be symmetric. In nonparametric models with multiplicity and/or endogeneity, identification may be achieved using excluded instruments with large support. Although such a strong requirement can be met in practice, estimation and inference can still be problematic (andrews1998semiparametric, khan2010irregular). Thus, we avoid such assumptions for instruments and other exogenous variables.

The first contribution of this study is to establish that under strategic substitutability, regions that predict the equilibria of the treatment selection process in the first-stage game can present a monotonic pattern in terms of the number of treatments selected.\footnote{To estimate payoff parameters, berry1992estimation partly characterizes equilibrium regions. To calculate the bounds on these parameters, CT09 simulate their moment inequalities model that are implied by the shape of these regions, especially the regions for multiple equilibria. While their approaches are sufficient for their analyses, full analytical results are critical for the identification analysis in this current study.} The second contribution of this study is to show, after restoring the generalized monotonicity in the selection process, how the model structure and the data can provide information about treatment parameters, such as the average treatment effects (ATEs). We first establish the bounds on the ATE and other related parameters with possibly discrete instruments. We also show that tighter bounds on the ATE can be obtained by introducing (possibly discrete) exogenous variables excluded from the first-stage game. This is especially motivated when the outcome variable is affected by externalities generated by the players. We can derive sharp bounds as long as the outcome variable is binary. To deal with the multiple equilibria problem in our analysis, we assume that instruments vary sufficently to offset the effect of strategic substitutability. We provide a simple testable implication for the existence of such instrument variation in the case of mutually independent payoff unobservables. This requirement of variation is qualitatively different and substantially weaker than a typical large support assumption. A marked feature of our analyses is that for the sharp bounds on the ATE, player-specific instruments are not necessary.

Our bound analysis builds on VY07 and SV11, which consider point and partial identification in single-agent nonparametric triangular models with binary endogenous variables. Unlike them, however, we allow for multi-agent strategic interaction as a key component of the model. Some studies have extended a single-treatment model to a multiple-treatment setting (e.g., heckman2006understanding, jun2011tighter), but their models maintain monotonicity in the selection process and none of them allow simultaneity among the multiple treatments resulting from agents' interaction, as we do in this study.

In interesting recent work, heckman2015unordered, and lee2016identifying extend the monotonicity of the selection process in multi-valued treatments settings. heckman2015unordered introduce unordered monotonicity, which is a different type of treatment selection mechanisms than ours. lee2016identifying consider more general non-monotonicity and do mention entry games as one example of the treatment selection processes they allow. However, they assume known payoffs and bypass the multiplicity of equilibria by assuming a threshold-crossing equilibrium selection mechanism, both of which we do not assume in this study. In addition, lee2016identifying's focus is on the identification of marginal treatment effects with continuous instruments. In another related work, chesher2014generalized consider a class of generalized instrumental variable models in which our model may fall and propose a systematic method of characterizing sharp identified sets for admissible structures. This present study's characterization of the identified sets is analytical, which helps investigate how the identification is related to exogenous variation in the model and to the equilibrium characterization in the treatment selection. Also, calculating the bounds on the treatment parameters using their approach involves projections of identified sets that may require parametric restrictions. Lastly, Han18,Han19c consider identification of dynamic treatment effects and optimal treatment regimes in a nonparametric dynamic model, in which the dynamic relationship causes non-monotonicity in the determination of each period's outcome and treatment.

Without triangular structures, manski1997monotone, MP00 and Man13 also propose bounds on the ATE with multiple treatments under various monotonicity assumptions, including an assumption on the sign of the treatment response. We take an alternative approach that is more explicit about treatments interaction while remaining agnostic about the direction of the treatment response. Our results suggest that provided there exist exogenous variation excluded from the selection process, the bounds calculated from this approach can be more informative than those from their approach.

Identification in models for binary games with complete information has been studied in Tam03, CT09, and bajari2010identification, among others.\footnote{See also galichon2011set and beresteanu2011sharp for a more general setup that includes complete information games as an example.} This present study contributes to this literature by considering post-game outcomes that are often not of players' direct concern. As related work that considers post-game outcomes, ciliberto2015market introduce a model in which firms make simultaneous decisions of entry and pricing upon entry. Consequently, their model can be seen as a multi-agent extension of a sample selection model. On the other hand, the model considered in this study is a multi-agent extension of a model for endogenous treatments. At a more general level, our approach is an attempt to bridge the treatment effect literature and the industrial organization (IO) literature. We are interested in the evaluation of treatments that are the result of agents' strategic interaction, an aspect that is key in the IO literature. To conduct the counterfactual analysis, however, we closely follow the treatment effect literature, instead of the structural approach of the IO literature. For example, CT09 and ciliberto2015market impose economic structure and parametric assumptions to recover model primitives for policy analyses. In contrast, our parameters of interest are treatment effects as functionals of the primitives (but excluding the game parameters), and thus, allow our model to remain nonparametric. In addition, as the goal is different, we employ a different approach to partial identification under the multiplicity of equilibria than theirs.\footnote{Even if we are willing to assume a known distribution for the unobserved payoff types, their approach to multiplicity is not applicable to the particular setting of this study.}

To demonstrate the applicability of our method, we take the bounds we propose to data on airline market structure and air pollution in cities in the U.S. Aircrafts and airports land operations are a major source of emissions, and thus, quantifying the causal effect of air transport on pollution is of importance to policy makers. We explicitly allow market structure to be determined endogenously as the outcome of an entry game in which airlines behave strategically to maximize their profits and where the resulting pollution in this market is not internalized by the firms. Additionally, we do not impose any structure on how airline competition affects pollution and allow for heterogenous effects across firms. In other words, not only do we allow the effect of a different number of firms in the market on pollution to be nonlinear and not restricted, but also distinguish the identity of the firms. The latter is important if we believe that behavior post-entry differs across airlines. For example, different airlines might operate the market with a higher frequency or with different types of airplanes, hence affecting pollution in a different way. To implement our application, we combine data from two sources. The first contains airline information from the Department of Transportation, which we use to construct a dataset of airlines' presence in each market. We then merge it with air pollution data in each airport from air monitoring stations compiled by the Environmental Protection Agency. In our preferred specification, our outcome variable is a binary measure of the level of particulate matter in the air.

We consider three sets of ATE exercises to investigate different aspects of the relationship between market structure and pollution in equilibrium. The first simply quantifies the effects of each airline operating as a monopolist compared to a situation in which the market is not served by any airline. We find that the effect of each airline on pollution is positive and statistically significant. We also find evidence of heterogeneity in the effects across different airlines. The second set of exercises examines the ATEs of all potential market structures on pollution. We find that the probability of high pollution is increasing with the number of airlines in the market, but at a decreasing rate. Finally, the third set of exercises quantifies the ATE of a single airline under all potential configurations of the market in terms of its rivals. We observe that in all cases, Delta entering a market has a positive effect on pollution and this effect is decreasing with the number of rivals. The results from the last two set of exercises are consistent with the results of a Cournot-competition oligopolistic model in which incumbents accommodate new entrants by reducing the quantity they produce.

This paper is organized as follows. Section (ref) summarizes the analysis of this study using a stylized example. Section (ref) presents a general theory: Section (ref) introduces the model and the parameters of interest; Section (ref) presents the generalized monotonicity for equilibrium regions for many players; and Section (ref) delivers the partial identification results of this study. Section (ref) presents a numerical illustration and Section (ref) the empirical application on airlines and pollution. In the Appendix, Section (ref) provides more examples to which our setup can be applied. Section (ref) contains four extensions of our main results. Finally, Section (ref) collects the proofs of theorems and lemmas.

A Stylized Example

We first illustrate the main results of this study with a stylized example. Suppose we are interested in the effects of airline competition on local air quality (or health). Let $Y_{i}$ denote the binary indicator of air pollution in market $i$. For illustration, we assume there are two potential airlines. In the next section, we present a general theory with more than two players. Let $D_{1,i}$ and $D_{2,i}$ be binary variables that indicate the decisions to enter market $i$ by Delta and United, respectively. We allow the decisions $D_{1,i}$ and $D_{2,i}$ to be correlated with some unobserved characteristics of the local market that affect $Y_{i}$. Moreover, since $D_{1,i}$ and $D_{2,i}$ are equilibrium outcomes of the entry game, we allow them to be outcomes from multiple equilibria. The endogeneity and the presence of multiple equilibria are our key challenges in this study.

Let $Y_{i}(d_{1},d_{2})$ be the potential air quality had Delta and United's decisions been $(D_{1},D_{2})=(d_{1},d_{2})$; for example, $Y_{i}(1,1)$ is the potential air quality from duopoly, $Y_{i}(1,0)$ is with Delta being a monopolist, and so on. Let $X_{i}$ be a vector of market characteristics that affect $Y_{i}$. Our parameter of interest is the ATE, $E[Y_{i}(d_{1},d_{2})-Y_{i}(d_{1}',d_{2}')|X_{i}=x]$, which captures the effect of market structure on pollution. One interesting ATE is $E[Y_{i}(1,d_{2})-Y_{i}(0,d_{2})|X_{i}=x]$ for each $d_{2}$, where we can learn the interaction effects of treatments, e.g., how much the average effect of Delta's entry is affected by United's entry: $E\left[Y_{i}(1,1)-Y_{i}(0,1)\right]-E\left[Y_{i}(1,0)-Y_{i}(0,0)\right]$ (suppressing $X_{i}$). In our empirical application (Section (ref)), we consider this and other related parameters in a more realistic model, where there are more than two airlines.

We show how we overcome the problems of endogeneity and multiple equilibria and how to construct bounds on the ATE using the excluded instruments and other exogenous variables. Let $Z_{1,i}$ and $Z_{2,i}$ be cost shifters for Delta and United, respectively, which serve as instruments. As a benchmark, we first consider naive bounds analogous to manski1990nonparametric using excluded instruments which satisfy

align[align omitted — 82 chars of source]

for all $(d_{1},d_{2})$. To simplify notation, we suppress the index $i$ henceforth, let $\boldsymbol{D}\equiv(D_{1},D_{2})$ and $\boldsymbol{Z}\equiv(Z_{1},Z_{2})$, and write $E[\cdot|w]\equiv E[\cdot|W=w]$ for a generic r.v. $W$. As an illustration, we focus on calculating bounds on $E[Y(1,1)|X=x]$. Note that

align[align omitted — 347 chars of source]

where the first equality is by (ref). Manski-type bounds can be obtained by observing that the counterfactual term $E[Y(1,1)|\boldsymbol{D}=\boldsymbol{d}^{\prime},\boldsymbol{z},x]=\Pr[Y(1,1)=1|\boldsymbol{D}=\boldsymbol{d}^{\prime},\boldsymbol{z},x]$ is bounded above by one and below by zero. By further using the variation in $\boldsymbol{Z}$, which is excluded from $Y(1,1)$, the lower and upper bounds on $E[Y(1,1)\vert x]$ can be written as

align*[align* omitted — 297 chars of source]

The goal of our analysis is to derive tighter bounds than $L_{Manski}(x)$ and $U_{Manski}(x)$ by introducing further assumptions motivated by economic theory.

To illustrate, we introduce the following semi-triangular model with linear indices. In the next section, we generalize this model by introducing fully nonparametric models that allow continuous $Y$. All the assumptions and results illustrated in the current section are formally stated and proved in the next section. Consider

align[align omitted — 238 chars of source]

where $(\epsilon,U_{1},U_{2})$ are continuously distributed unobservables that can be arbitrarily correlated, $(U_{1},U_{2})$ are uniform, and assume

align[align omitted — 189 chars of source]

Note that (ref) replaces (ref), (ref) assumes strategic substitutability, and (ref) is plausible in the current example of air quality and entry. Owing to the first stage simultaneity, the model (ref)--(ref) is incomplete, i.e., the model primitives and the covariates do not uniquely predict $(Y,\boldsymbol{D})$. In this model, we are not interested in the players' payoff parameters $(\delta_{-s},\gamma_{s})$ for $s=1,2$, individual parameters $(\mu_{1},\mu_{2},\beta)$ that generate the outcome, nor distributional parameters. Instead, we are interested in the ATE as a function of $(\mu_{1},\mu_{2},\beta)$. This is in contrast to ciliberto2015market, where payoff and pricing parameters are direct parameters of interest, and thus, our identification question and strategy (especially how we deal with multiple equilibria) are different from theirs.

Typically, a standard approach that utilizes instrumental variables compares the reduced-form relationship between the outcome and treatment with the reduced-form relationship between the treatment and instrument. We apply the same idea here by changing the values of $Z_{1}$ and $Z_{2}$ and measure the change in $Y$ relative to the change in $D_{1}$ and $D_{2}$. To this end, for two realizations $\boldsymbol{z},\boldsymbol{z}'$ of $\boldsymbol{Z}$, say low and high entry cost for both airlines, we introduce reduced-form objects directly recovered from the data:

align[align omitted — 320 chars of source]

for $d\in\{(0,0),(1,0),(0,1),(1,1)\}\equiv\mathcal{D}$. We show that (ref)--(ref) deliver useful information about the outcome index function ($\mu_{1}D_{1}+\mu_{2}D_{2}+\beta X$), which in turn is helpful in constructing bounds on the ATE. Note that

align[align omitted — 680 chars of source]

where $\boldsymbol{D}=(1,0)$ and $(0,1)$ are the airlines' decisions that may arise as multiple equilibria. The increase in cost (from $\boldsymbol{z}$ to $\boldsymbol{z}'$) will make the operation of these airlines less profitable in some markets, depending on the values of the unobservables $\boldsymbol{U}=(U_{1},U_{2})$. This will result in a change in the market structure in those markets. Specifically, markets “on the margin” may experience one of the following changes in structure as cost increases: (a) from duopoly to Delta-monopoly; (b) from duopoly to United-monopoly; (c) from Delta-monopoly to no entrant; (d) from United-monopoly to no entrant; and (e) from duopoly to no entrant. These changes are depicted in Figure (ref), where each $R_{d_{1},d_{2}}(\boldsymbol{z})$ denotes the maximal region that predicts $(d_{1},d_{2})$, given $\boldsymbol{Z}=\boldsymbol{z}$.\footnote{See Section (ref) in the Appendix for a formal definition. The figure is drawn in a way that $\gamma_{1}$ and $\gamma_{2}$ are negative.}

figure*[figure* omitted — 2,778 chars of source]

These changes (a)--(e) are a consequence of the monotonic pattern of equilibrium regions, which we formally establish in a general setting of more than two players in Theorem (ref) of Section (ref).

In general, besides these five scenarios, there may be markets that used to be Delta-monopoly but become United-monopoly and vice versa, i.e., markets that exhibit non-monotonic behaviors; see Remark (ref) below for details. Owing to possible multiple equilibria, we are agnostic about these latter types of changes except in extreme cases, where one equilibrium is selected with probability one. We generally do not know the equilibrium selection mechanism in play, much less about how such mechanism changes as cost $\boldsymbol{Z}$ changes. The key idea in this study is to overcome the non-monotonicity by shifting the cost sufficiently so that there is no market that switches from one monopoly to another. We show that the shift in cost that compensates the strategic substitutability does just that, as is depicted in Figure (ref). In this figure, we assume $\delta_{2}+\gamma_{1}z_{1}>\gamma_{1}z_{1}'$ and $\delta_{1}+\gamma_{2}z_{2}>\gamma_{2}z_{2}'$. In other words, we assume

align[align omitted — 114 chars of source]

Importantly, we do not require infinite variation in $\boldsymbol{Z}$.\footnote{Of course, changing each $Z_{s}$ from $-\infty$ to $\infty$ will trivially achieve our requirement of having no market that switches from one monopoly to another.} In fact, we show that the compensating strategic substitutability (ref) is implied by the following condition, which can be tested using the data: there exist $\boldsymbol{z},\boldsymbol{z}'\in\mathcal{Z}$ such that

align[align omitted — 126 chars of source]

Suppose $\boldsymbol{z},\boldsymbol{z}'$ satisfy (ref). Then, by (ref), we can derive from (ref) that (suppressing $X=x$ for simplicity)

align[align omitted — 493 chars of source]

where $\Delta_{i}$ ($i\in\{a,...,e\}$) are disjoint and each $\Delta_{i}$ characterizes those markets on the margin described above: $\Delta_{a}$ corresponds to the set of $\boldsymbol{U}$'s that experience (a), $\Delta_{b}$ corresponds to (b), and so on. Once (ref) is derived, it is easy to see that

align[align omitted — 106 chars of source]

which is formally shown in Lemma (ref)(i). See Section (ref) in the Appendix for a proof in this specific two-player case, which simplifies the argument in the general proof. The result (ref) is helpful for our bound analysis. Again, focus on $E[Y(1,1)|x]$ and suppose $h(\boldsymbol{z},\boldsymbol{z}',x)>0$. Then, $\mu_{1}>0$ and $\mu_{2}>0$, and thus, we can derive the lower bound on, e.g., $E[Y(1,1)|\boldsymbol{D}=(1,0),\boldsymbol{z},x]$ in (ref) as

align[align omitted — 318 chars of source]

which is larger than zero, the previous naive lower bound. Similarly, we can calculate the lower bounds on all $E[Y(1,1)|\boldsymbol{D}=\boldsymbol{d},\boldsymbol{z},x]$ for $\boldsymbol{d}\neq(1,1)$. Consequently, by (ref), we have

align*[align* omitted — 57 chars of source]

i.e., the lower bound on $E[Y(1,1)|x]$ is

align*[align* omitted — 82 chars of source]

Note that $\tilde{L}(x)\ge L_{Manski}(x)$. In this case, $\tilde{U}(x)=U_{Manski}(x)$. In Section (ref), we show that $\tilde{L}(x)$ and $\tilde{U}(x)$ are sharp under (ref)--(ref).

We can further tighten the bounds if we have exogenous variables that are excluded from the entry decisions, i.e., from the $D_{1}$ and $D_{2}$ equations. The existence of such variables is not necessary but helpful in tightening the bounds, and can be motivated by the notion of externalities. That is, there can exist factors that affect $Y$ but do not enter the players' first-stage payoff functions. Modify (ref) and assume

align[align omitted — 76 chars of source]

where conditioning on other (possibly endogenous) covariates is suppressed. Here, $X$ can be the characteristics of the local market that directly affect pollution or health levels, such as weather shocks or the share of pollution-related industries in the local economy. We assume that conditional on other covariates, these factors affect the outcome but do not enter the payoff functions, since the airlines do not take them into account in their decisions.

To exploit the variation in $X$ (in addition to the variation in $\boldsymbol{Z}$), let $(x,\tilde{x},\tilde{\tilde{x}})$ be (possibly different) realizations of $X$, and define

equation[equation omitted — 310 chars of source]

Under (ref) and analogous to (ref), we can show that if

align[align omitted — 123 chars of source]

is positive (negative), then $sgn\{\mu_{1}+\beta(x-x')\}=sgn\{\mu_{2}+\beta(x-x')\}$ is positive (negative). This is formally shown in Lemma (ref)(ii).

As before, suppose $h(\boldsymbol{z},\boldsymbol{z}',x)>0$, and thus, $\mu_{1}>0$ and $\mu_{2}>0$ by (ref). Now, if $\tilde{h}(\boldsymbol{z},\boldsymbol{z}',x',x',x)<0$, then $\mu_{1}+\beta x<\beta x'$ and $\mu_{2}+\beta x<\beta x'$. Therefore, we can derive

align*[align* omitted — 277 chars of source]

where the second equality also uses (ref) and (ref)--(ref). Similarly, we have $E[Y(1,1)|\boldsymbol{D}=(0,1),\boldsymbol{z},x]\le\Pr[Y=1|\boldsymbol{D}=(0,1),\boldsymbol{z},x']$, and consequently, the upper bound on $E[Y(1,1)\vert x]$ becomes \[ U(x)\equiv\inf_{\boldsymbol{z}\in\mathcal{Z}}\left\{ \Pr[Y=1,\boldsymbol{D}=(1,1)\vert\boldsymbol{z},x]+\Pr[Y=1,\boldsymbol{D}\in\{(1,0),(0,1)\}\vert\boldsymbol{z},x']+\Pr[\boldsymbol{D}=(0,0)\vert\boldsymbol{z},x]\right\} \] by (ref), and the lower bound is $L(x)=\tilde{L}(x)$. Note that we can further take infimum over $x'$ such that $\tilde{h}(\boldsymbol{z},\boldsymbol{z}',x',x',x)<0$.

To summarize our illustration, by using two values of $\boldsymbol{Z}$ that satisfy (ref) and $h(\boldsymbol{z},\boldsymbol{z}',x)>0$ and two values of $X$ that satisfy $\tilde{h}(\boldsymbol{z},\boldsymbol{z}',x',x',x)<0$, our lower and upper bounds, $L(x)$ and $U(x)$, on $E[Y(1,1)\vert x]$ achieve

align*[align* omitted — 92 chars of source]

where the inequalities are strict if $\sum_{\boldsymbol{d}\neq(1,1)}\Pr[Y=1,\boldsymbol{D}=\boldsymbol{d}|\boldsymbol{z},x]>0$ and $\Pr[Y=0,\boldsymbol{D}\in\{(1,0),(0,1)\}\vert\boldsymbol{z},x']>0$. We discuss the sharpness of $L(x)$ and $U(x)$ in the next section. Similarly, we can derive lower and upper bounds on other $E[Y(\boldsymbol{d})\vert x]$'s for $\boldsymbol{d}\neq(1,1)$, and eventually construct bounds on any ATE. The gain from our approach is also exhibited in Figure (ref) in Section (ref), where we use the same data generating process as in this section and calculate different bounds on the ATE, $E[Y(1,1)\vert x]-E[Y(0,0)\vert x]$.

remark[Point Identification of the ATE] When there exist player-specific excluded instruments with large support, we point identify the ATEs. To invoke an identification-at-infinity argument, the following assumptions are instead needed to hold: \begin{align} & \gamma_{1} and \gamma_{2} are nonzero,\\ & Z_{1}|(X,Z_{2}) and Z_{2}|(X,Z_{1}) has an everywhere positive Lebesgue density. \end{align} These assumptions impose a player-specific exclusion restriction and large support. Under (ref)--(ref), we can easily show that the ATE in (ref) is point identified. In this case, the structure we impose, especially on the outcome function (such as the threshold-crossing structure, or more generally Assumption M in Section (ref) below) is not needed. The identification strategy is to exploit the large variation of player specific instruments based on (ref)--(ref), which simultaneously solves the multiple equilibria and the endogeneity problems. For example, to identify $E[Y(1,1)|x]$, consider \begin{align*} & E[Y|\boldsymbol{D}=(1,1),\boldsymbol{z},x]=E[Y(1,1)|\boldsymbol{D}=(1,1),\boldsymbol{z},x]\\ & =E[Y(1,1)|\delta_{2}+\gamma_{1}z_{1}\geq U_{1},\delta_{1}+\gamma_{2}z_{2}\geq U_{2},x]\rightarrow E[Y(1,1)|x], \end{align*} where the second equation is by (ref) and $Y(1,1)=\mu_{1}+\mu_{2}+\beta X$, and the convergence is by (ref)--(ref) with $z_{1}\rightarrow\infty$ and $z_{2}\rightarrow\infty$. The identification of $E[Y(0,0)|x]$, $E[Y(1,0)|x]$ and $E[Y(0,1)|x]$ can be achieved by similar reasoning. Note that $\boldsymbol{D}=(1,0)$ or $\boldsymbol{D}=(0,1)$ can be predicted as an outcome of multiple equilibria. However, when either $(z_{1},z_{2})\rightarrow(\infty,-\infty)$ or $(z_{1},z_{2})\rightarrow(-\infty,\infty)$ occurs, a unique equilibrium is guaranteed as a dominant strategy, i.e., $\boldsymbol{D}=(1,0)$ or $\boldsymbol{D}=(0,1)$, respectively.
remark[Non-Monotonicity of Treatment Selection]In the case of a single binary treatment, the standard selection equation exhibits monotonicity that facilitates various identification strategies (e.g., imbens1994identification, heckman2005structural, VY07 to name a few). Relatedly, vytlacil2002independence shows the equivalence between imposing the selection equation with threshold-crossing structure and assuming the local ATE (LATE) monotonicity. This equivalence (and thus, previous identification strategies) is inapplicable to our setting due to the simultaneity in the first stage (ref). To formally state this, let $\boldsymbol{D}(\boldsymbol{z})$ be a potential treatment vector, had $\boldsymbol{Z}=\boldsymbol{z}$ been realized. When cost $\boldsymbol{Z}=(Z_{1},Z_{2})$ increases from $\boldsymbol{z}$ to $\boldsymbol{z}'$, it may be that some markets witness Delta entering and United going out of business (i.e., $\boldsymbol{D}(\boldsymbol{z})=(0,1)$ and $\boldsymbol{D}(\boldsymbol{z}')=(1,0)$), while other markets witness the opposite (i.e., $\boldsymbol{D}(\boldsymbol{z})=(1,0)$ and $\boldsymbol{D}(\boldsymbol{z}')=(0,1)$). The direction of monotonicity is reversed in the two groups of markets, and thus, $\Pr[\boldsymbol{D}(\boldsymbol{z})\ge\boldsymbol{D}(\boldsymbol{z}')]\neq1$ and $\Pr[\boldsymbol{D}(\boldsymbol{z})\le\boldsymbol{D}(\boldsymbol{z}')]\neq1$ where the inequality for vectors is pair-wise inequalities, which violates the LATE monotonicity.\footnote{The same argument applies with a scalar multi-valued treatment $\tilde{D}\in\{1,2,3,4\}$, which has a one-to-one map with $\boldsymbol{D}\in\{(0,0),(0,1),(1,0),(1,1)\}$. Then, some markets can experience $\tilde{D}(\boldsymbol{z})=2$ and $\tilde{D}(\boldsymbol{z}')=3$ while others experience $\tilde{D}(\boldsymbol{z})=3$ and $\tilde{D}(\boldsymbol{z}')=2$, and thus, it is possible to have $\Pr[\tilde{D}(\boldsymbol{z})\ge\tilde{D}(\boldsymbol{z}')]\neq1$ and $\Pr[\tilde{D}(\boldsymbol{z})\le\tilde{D}(\boldsymbol{z}')]\neq1$.} Despite this non-monotonic pattern, Theorem (ref) below restores generalized monotonicity, i.e., monotonicity in terms of the algebra of sets. This generalized monotonicity, combined with the compensating strategic substitutability (ref), allows us to use a strategy analogous to the single-treatment case for our bound analysis. This also suggests that we can introduce a generalized version of the LATE parameter in the current framework, although we do not pursue it in this study. Related to our study, lee2016identifying introduce a framework for treatment effects with general non-monotonicity of selection, and consider the simultaneous treatment selection as one of the examples. Although they engage in a similar discussion on non-monotonicity, their approach to gain tractability for identification is different from ours. When they allow the identity of players being observed as in our setting, they show that their treatment measurability condition (Assumption 2.1) introduced to restore monotonicity is satisfied, provided they assume a threshold-crossing equilibrium selection mechanism. In contrast, we avoid making assumptions on equilibrium selection, but require compensating variation of instruments. In addition, for this particular example, they assume the first-stage is known (i.e., payoff functions are known), and focus on point identification of the MTE with continuous instruments.

General Theory

Setup

Let $\boldsymbol{D}\equiv(D_{1},...,D_{S})\in\mathcal{D}\subseteq\{0,1\}^{S}$ be an $S$-vector of observed binary treatments and $\boldsymbol{d}\equiv(d_{1},...,d_{S})$ be its realization, where $S$ is fixed. We assume that $\boldsymbol{D}$ is predicted as a pure strategy Nash equilibrium of a complete information game with $S$ players who make entry decisions or individuals who choose to receive treatments.\footnote{While this study does not consider mixed strategy equilibria, it may be possible to extend the setup to incorporate mixed strategies, following the argument in CT09.} Let $Y$ be an observed post-game outcome that results from profile $\boldsymbol{D}$ of endogenous treatments. It can be an outcome common to all players or an outcome specific to each player. Let $(X,Z_{1},...,Z_{S})$ be observed exogenous covariates. We consider a model of a semi-triangular system:

align[align omitted — 213 chars of source]

where $s$ is indices for players or interchangeably for treatments, and $\boldsymbol{D}_{-s}\equiv(D_{1},...,D_{s-1},D_{s+1},...,D_{S})$. Without loss of generality, we normalize the scalar $U_{s}$ to be distributed as $Unif(0,1)$, and $\nu^{s}:\mathbb{R}^{S-1+d_{z_{s}}}\rightarrow(0,1]$ and $\theta:\mathbb{R}^{S+d_{x}+d_{\epsilon}}\rightarrow\mathbb{R}$ are unknown functions that are nonseparable in their arguments. We allow the unobservables $(\epsilon_{\boldsymbol{D}},U_{1},...,U_{S})$ to be arbitrarily dependent on one another. Although the notation suggests that the instruments $Z_{s}$'s are player/treatment-specific, they are not necessarily required to be so for the analyses in this study; see Appendix (ref) for a discussion. The exogenous variables $X$ are variables excluded from all the equations for $D_{s}$. The existence of $X$ is not necessary but useful for the bound analysis of the ATE. There may be covariates $W$ common to all the equations for $Y$ and $D_{s}$, which is suppressed for succinctness. Implied from the complete information game, player $s$'s decision $D_{s}$ depends on the decisions of all others $\boldsymbol{D}_{-s}$ in $\mathcal{D}_{-s}$, and thus, $\boldsymbol{D}$ is determined by a simultaneous system. As before, the model (ref)--(ref) is incomplete because of the simultaneity in the first stage, and the conventional monotonicity in the sense of imbens1994identification is not exhibited in the selection process because of simultaneity. The unit of observation, a market or geographical region, is indexed by $i$ and is suppressed in all the expressions.

The potential outcome of receiving treatments $\boldsymbol{D}=\boldsymbol{d}$ can be written as

align*[align* omitted — 128 chars of source]

and $\epsilon_{\boldsymbol{D}}=\sum_{\boldsymbol{d}\in\mathcal{D}}1[\boldsymbol{D}=\boldsymbol{d}]\epsilon_{\boldsymbol{d}}$. We are interested in the ATE and related parameters. Using the average structural function (ASF) $E[Y(\boldsymbol{d})|x]$, the ATE can be written as

align[align omitted — 203 chars of source]

for $\boldsymbol{d},\boldsymbol{d}'\in\mathcal{D}$. Another parameter of interest is the average treatment effects on the treated or the untreated: $E[Y(\boldsymbol{d})-Y(\boldsymbol{d}^{\prime})|D=\boldsymbol{d}'',z,x]$ for $\boldsymbol{d}''\in\{ \boldsymbol{d},\boldsymbol{d}' \}$.\footnote{Technically, $\boldsymbol{d}''$ does not necessarily have to be equal to $\boldsymbol{d}$ or $\boldsymbol{d}'$, but can take another value.} One might also be interested in the sign of the ATE, which in this multi-treatment case is essentially establishing an ordering among the ASF's.

As an example of the ATE, we may choose $\boldsymbol{d}=(1,...,1)$ and $\boldsymbol{d}'=(0,...,0)$ to measure the cancelling-out effect or more general nonlinear effects. Another example would be choosing $\boldsymbol{d}=(1,\boldsymbol{d}_{-s})$ and $\boldsymbol{d}'=(0,\boldsymbol{d}_{-s})$ for given $\boldsymbol{d}_{-s}$, where we use the notation $\boldsymbol{d}=(d_{s},\boldsymbol{d}_{-s})$ by switching the order of the elements for convenience. Sometimes, we instead want to focus on learning about complementarity between two treatments, while averaging over the remaining $S-2$ treatments. This can be examined in a more general framework of defining the ASF and ATE by introducing a partial potential outcome; this is discussed in Appendix (ref).

In identifying these treatment parameters, suppose we attempt to recover the effect of a single treatment $D_{s}$ in model (ref)--(ref) conditional on $\boldsymbol{D}_{-s}=\boldsymbol{d}_{-s}$, and then recover the effects of multiple treatments by transitively using these effects of single treatments. This strategy is not valid since $\boldsymbol{D}_{-s}$ is a function of $D_{s}$ and also because of multiplicity. Therefore, the approaches in the literature with single-treatment, single-agent triangular models are not directly applicable and a new theory is necessary in this more general setting.

Monotonicity in Equilibria

As an important step in the analyses in this study, we establish that the equilibria of the treatment selection process in the first-stage game present a monotonic pattern when the instruments move. Specifically, we consider the regions in the space of the unobservables that predict equilibria and establish a monotonic pattern of these regions in terms of instruments. The analytical characterization of the equilibrium regions when there are more than two players ($S>2$) can generally be complicated (CT09); however, under a mild uniformity assumption (Assumption M1), our result is obtained under strategic substitutability. Let $\mathcal{Z}_{s}$ be the support of $Z_{s}$. We make the following assumptions on the first-stage nonparametric payoff function for each $s\in\{1,...,S\}$.

asSSFor every $z_{s}\in\mbox{\ensuremath{\mathcal{Z}}}_{s}$, $\nu^{s}(\boldsymbol{d}_{-s},z_{s})$ is strictly decreasing in each element of $\boldsymbol{d}_{-s}$.

\begin{asM1}For any given $z_{s},z_{s}'\in\mathcal{Z}_{s}$, either $\nu^{s}(\boldsymbol{d}_{-s},z_{s})\geq\nu^{s}(\boldsymbol{d}_{-s},z_{s}')$ $\forall\boldsymbol{d}_{-s}\in\mathcal{D}_{-s}$, or $\nu^{s}(\boldsymbol{d}_{-s},z_{s})\leq\nu^{s}(\boldsymbol{d}_{-s},z_{s}')$ $\forall\boldsymbol{d}_{-s}\in\mathcal{D}_{-s}$.\end{asM1}

Assumption SS asserts that the agents' treatment decisions are produced in a game with strategic substitutability. The strictness of the monotonicity is not important for our purpose but convenient in making statements about the equilibrium regions. In the language of CT09, we allow for heterogeneity in the fixed competitive effects (i.e., how each of other entrants affects one's payoff), as well as heterogeneity in how each player is affected by other entrants, which is ensured by the nonseparability between $\boldsymbol{d}_{-s}$ and $z_{s}$ in $\nu^{s}(\boldsymbol{d}_{-s},z_{s})$; this heterogeneity is related to the variable competitive effects. Assumption M1 is required in this multi-agent setting, and the uniformity is across $\boldsymbol{d}_{-s}$. Note that this assumption is weaker than a conventional monotonicity assumption that $\nu^{s}(\boldsymbol{d}_{-s},z_{s})$ is either non-decreasing or non-increasing in $z_{s}$ for all $\boldsymbol{d}_{-s}$. Assumption M1 is justifiable, especially when $z_{s}$ is chosen to be of the same kind for all players. For example, in an entry game, if $z_{s}$ is chosen to be each player's cost shifters, the payoffs would decrease in their costs for any given opponents.

As the first main result of this study, we establish the geometric property of the equilibrium regions. For $j=0,...,S$, let $\boldsymbol{R}_{j}(\boldsymbol{z})\subset\mathcal{U}\equiv(0,1]^{S}$ denote the region that predicts all equilibria with $j$ treatments selected or $j$ entrants, defined as a subset of the space of the entry unobservables $\boldsymbol{U}\equiv(U_{1},...,U_{S})$; see Section (ref) in the Appendix for a formal definition. Then, define the region of all equilibria with at most $j$ entrants as

align*[align* omitted — 113 chars of source]

Although this region is hard to express explicitly in general, it has a simple feature that serves our purpose. For given $j$, choose $z_{s},z_{s}'\in\mathcal{Z}_{s}$ such that

align[align omitted — 181 chars of source]

for all $s$. This condition is to merely fix $\boldsymbol{z},\boldsymbol{z}'$ that change the joint propensity score, and the direction of change is without loss of generality. Such $\boldsymbol{z},\boldsymbol{z}'$ exist by the relevance of the instruments, which is assumed below. Let $\mathcal{Z}$ be the support of $\boldsymbol{Z}\equiv(Z_{1},...,Z_{S})$.

theoremUnder Assumptions SS and M1 and for $\boldsymbol{z},\boldsymbol{z}'\in\mathcal{Z}$ that satisfy (ref), we have \begin{equation} \boldsymbol{R}^{\le j}(\boldsymbol{z})\subseteq\boldsymbol{R}^{\le j}(\boldsymbol{z}') \forall j. \end{equation}

Theorem (ref) establishes a generalized version of monotonicity in the treatment selection process. This theorem plays a crucial role in calculating the bounds on the treatment parameters and in showing the sharpness of the bounds. Relatedly, berry1992estimation derives the probability of the event that the number of entrants is less than a certain value, which can be written as $\Pr[\boldsymbol{U}\in\boldsymbol{R}^{\le j}(\boldsymbol{z})]$ using our notation. However, his result is not sufficient for our study and relies on stronger assumptions, such as restricting the payoff functions to only depend on the number of opponents.

Main Assumptions

To characterize the bounds on the treatment parameters, we make the following assumptions. Unless otherwise noted, the assumptions hold for each $s\in\{1,...,S\}$.

asIN$(X,\boldsymbol{Z})\perp(\epsilon_{\boldsymbol{d}},\boldsymbol{U})$ $\forall\boldsymbol{d}\in\mathcal{D}$.
asEThe distribution of $(\epsilon_{\boldsymbol{d}},\boldsymbol{U})$ has strictly positive density with respect to Lebesgue measure on $\mathbb{R}^{S+1}$ $\forall\boldsymbol{d}\in\mathcal{D}$.
asEXFor each $\boldsymbol{d}_{-s}\in\mathcal{D}_{-s}$, $\nu^{s}(\boldsymbol{d}_{-s},Z_{s})|X$ is nondegenerate.

Assumptions IN, EX and all the following analyses can be understood as conditional on $W$, the common covariates in $X$ and $\boldsymbol{Z}=(Z_{1},...,Z_{S})$. Assumption EX is related to the exclusion restriction and the relevance condition of the instruments $Z_{s}$.

We now impose a shape restriction on the outcome function $\theta(\boldsymbol{d},x,\epsilon_{\boldsymbol{d}})$ via restrictions on \[ \vartheta(\boldsymbol{d},x;\boldsymbol{u})\equiv E[\theta(\boldsymbol{d},x,\epsilon_{\boldsymbol{d}})|\boldsymbol{U}=\boldsymbol{u}] \] a.e. $\boldsymbol{u}$. This restriction on the conditional mean is weaker than the one directly imposed on $\theta(\boldsymbol{d},x,\epsilon_{\boldsymbol{d}})$. Let $\mathcal{X}$ be the support of $X$. Recall that we use the notation $\boldsymbol{d}=(d_{s},\boldsymbol{d}_{-s})$ by switching the order of the elements for convenience.

asMFor every $x\in\mathcal{X}$, either $\vartheta(1,\boldsymbol{d}_{-s},x;\boldsymbol{u})\geq\vartheta(0,\boldsymbol{d}_{-s},x;\boldsymbol{u})$ a.e. $\boldsymbol{u}$ $\forall\boldsymbol{d}_{-s}\in\mathcal{D}_{-s}$ $\forall s$ or $\vartheta(1,\boldsymbol{d}_{-s},x;\boldsymbol{u})\leq\vartheta(0,\boldsymbol{d}_{-s},x;\boldsymbol{u})$ a.e. $\boldsymbol{u}$ $\forall\boldsymbol{d}_{-s}\in\mathcal{D}_{-s}$ $\forall s$. Also, $Y\in[\underline{Y},\overline{Y}]$.

Assumption M holds in, but is not restricted to, the leading case of binary $Y$ with a threshold crossing model that satisfies uniformity.

\begin{asM2}(i) $\theta(\boldsymbol{d},x,\epsilon_{\boldsymbol{d}})=1[\mu(\boldsymbol{d},x)\geq\epsilon_{\boldsymbol{d}}]$ where $\epsilon_{\boldsymbol{d}}$ is scalar and $F_{\epsilon_{\boldsymbol{d}}|\boldsymbol{U}}=F_{\epsilon_{\boldsymbol{d}'}|\boldsymbol{U}}$ for any $\boldsymbol{d},\boldsymbol{d}'\in\mathcal{D}$; (ii) for every $x\in\mathcal{X}$, either $\mu(1,\boldsymbol{d}_{-s},x)\geq\mu(0,\boldsymbol{d}_{-s},x)$ $\forall\boldsymbol{d}_{-s}\in\mathcal{D}_{-s}$ $\forall s$ or $\mu(1,\boldsymbol{d}_{-s},x)\le\mu(0,\boldsymbol{d}_{-s},x)$ $\forall\boldsymbol{d}_{-s}\in\mathcal{D}_{-s}$ $\forall s$.\end{asM2}

Assumption M$^{*}$ implies Assumption M. The second statement in Assumption M is satisfied with binary $Y$.\footnote{Another example would be when $Y\in[0,1]$, as in Example (ref).} The first statement in Assumption M can be stated in two parts, corresponding to (i) and (ii) of Assumption M$^{*}$: (a) for every $x$ and $\boldsymbol{d}_{-s}$, either $\vartheta(1,\boldsymbol{d}_{-s},x;\boldsymbol{u})\geq\vartheta(0,\boldsymbol{d}_{-s},x;\boldsymbol{u})$ a.e. $\boldsymbol{u}$, or $\vartheta(1,\boldsymbol{d}_{-s},x;\boldsymbol{u})\leq\vartheta(0,\boldsymbol{d}_{-s},x;\boldsymbol{u})$ a.e. $\boldsymbol{u}$; (b) for every $x$, each inequality statement in (a) holds for all $\boldsymbol{d}_{-s}$. For an outcome function with a scalar index, $\theta(\boldsymbol{d},x,\epsilon_{\boldsymbol{d}})=\tilde{\theta}(\mu(\boldsymbol{d},x),\epsilon_{\boldsymbol{d}})$, part (a) is implied by $\epsilon_{\boldsymbol{d}}=\epsilon_{\boldsymbol{d}'}=\epsilon$ (or more generally, $F_{\epsilon_{\boldsymbol{d}}|\boldsymbol{U}}=F_{\epsilon_{\boldsymbol{d}'}|\boldsymbol{U}}$) for any $\boldsymbol{d},\boldsymbol{d}'\in\mathcal{D}$ and $E[\tilde{\theta}(t,\epsilon_{\boldsymbol{d}})|\boldsymbol{U}=\boldsymbol{u}]$ being strictly increasing (decreasing) in $t$ a.e. $\boldsymbol{u}$.\footnote{A single-treatment version of the latter assumption appears in VY07 (Assumption A-4), which is weaker than assuming $\tilde{\theta}(t,\epsilon)$ is strictly increasing (decreasing) a.e. $\epsilon$; see VY07 for related discussions.} Functions that satisfy the latter assumption include strictly monotonic functions, such as transformation models $\tilde{\theta}(t,\epsilon)=r(t+\epsilon)$ with $r(\cdot)$ being possibly unknown strictly increasing, or their special case $\tilde{\theta}(t,\epsilon)=t+\epsilon$, allowing continuous dependent variables; and functions that are not strictly monotonic, such as models for limited dependent variables, $\tilde{\theta}(t,\epsilon)=1[t\ge\epsilon]$ or $\tilde{\theta}(t,\epsilon)=1[t\ge\epsilon](t-\epsilon)$. However, there can be functions that violate the latter assumption but satisfy part (a). For example, consider a threshold crossing model with a random coefficient: $\theta(\boldsymbol{d},x,\epsilon)=1[\phi(\epsilon)\boldsymbol{d}\beta^{\top}\geq x\gamma^{\top}]$, where $\phi(\epsilon)$ is nondegenerate. When $\beta_{s}\geq0$, then $E[\theta(1,\boldsymbol{d}_{-s},x,\epsilon)-\theta(0,\boldsymbol{d}_{-s},x,\epsilon)|\boldsymbol{U}=\boldsymbol{u}]=\Pr\left[\frac{x\gamma^{\top}}{\beta_{s}+\boldsymbol{d}_{-s}\beta_{-s}^{\top}}\leq\phi(\epsilon)\leq\frac{x\gamma^{\top}}{\boldsymbol{d}_{-s}\beta_{-s}^{\top}}|\boldsymbol{U}=\boldsymbol{u}\right]$, and thus, nonnegative a.e. $\boldsymbol{u}$, and vice versa. Part (a) also does not impose any monotonicity of $\theta$ in $\epsilon_{\boldsymbol{d}}$ (e.g., $\epsilon_{\boldsymbol{d}}$ can be a vector).

Part (b) of Assumption M imposes uniformity, as we deal with more than one treatment. Uniformity is required across different values of $\boldsymbol{d}_{-s}$ and $s$. For instance, in the empirical application of this study, this assumption seems reasonable, since an airline's entry is likely to increase the expected pollution regardless of the identity or the number of existing airlines. On the other hand, in Example (ref) in the Appendix regarding media and political behavior, this assumption may rule out the “over-exposure” effect (i.e., too much media exposure diminishes the incumbent's chance of being re-elected). In any case, knowledge on the direction of the monotonicity is not necessary in this assumption, unlike manski1997monotone or Man13, where the semi-monotone treatment response is assumed for possible multiple treatments.

Lastly, we require that there exists variation in $\boldsymbol{Z}$ that offsets the effect of strategic substitutability. Similar as before, using the notation $\boldsymbol{d}_{-s}=(d_{s'},\boldsymbol{d}_{-(s,s')})$ where $\boldsymbol{d}_{-(s,s')}$ is $\boldsymbol{d}$ without $s$-th and $s'$-th elements, note that Assumption SS can be restated as $\nu^{s}(0,\boldsymbol{d}_{-(s,s')},z_{s})>\nu^{s}(1,\boldsymbol{d}_{-(s,s')},z_{s})$ for every $z_{s}$. Given this, we assume the following.

asEQThere exist $\boldsymbol{z},\boldsymbol{z}'\in\mathcal{Z}$, such that $\nu^{s}(0,\boldsymbol{d}_{-(s,s')},z_{s}')\le\nu^{s}(1,\boldsymbol{d}_{-(s,s')},z_{s})$ $\forall\boldsymbol{d}_{-(s,s')}$ $\forall s,s'$.

For example, in an entry game with $Z_{s}$ being cost shifters, Assumption EQ may hold with $z_{s}'>z_{s}$ $\forall s$. In this example, players may become less profitable with an increase in cost from government regulation. In particular, players' decreased profits cannot be overturned by the market being less competitive, as one player is absent due to unprofitability. Recall that Assumption EQ is illustrated in Figure (ref) with $\nu^{s}(0,z_{s}')=\gamma_{s}z_{s}'<\nu^{s}(1,z_{s})=\delta_{-s}+\gamma_{s}z_{s}$ for $s=1,2$. Assumption EQ is key for our analysis. To see this, let $R_{j}^{M}(\cdot)$ denote the region that predicts multiple equilibria with $j$ treatments selected or $j$ entrants. In the proof of a lemma that follows, we show that Assumption EQ holds if and only if $R_{j}^{M}(\boldsymbol{z})\cap R_{j}^{M}(\boldsymbol{z}')=\emptyset$. That is, we can at least ensure that there is no market where firms' decisions change from one realization of multiple equilibria to another realization of multiple equilibria with the same number of entrants. To the extent of our analysis, this liberates us from concerns about the regions of multiple equilibria and about a possible change in equilibrium selection when changing $\boldsymbol{Z}$.\footnote{In Section (ref), we discuss an assumption, partial conditional symmetry, which can be imposed alternative to Assumption EQ.} Assumption EQ has a simple testable sufficient condition, provided that the unobservables in the payoffs are mutually independent.

\begin{asEQ2}There exist $\boldsymbol{z},\boldsymbol{z}'\in\mathcal{Z}$, such that

align[align omitted — 156 chars of source]

for all $\boldsymbol{d}^{j}\in\mathcal{D}^{j}$, $\boldsymbol{d}^{j-2}\in\mathcal{D}^{j-2}$ and $2\le j\le S$.

\end{asEQ2}

When $S=2$, the condition is stated as $\Pr[\boldsymbol{D}=(1,1)|\boldsymbol{z}]+\Pr[\boldsymbol{D}=(0,0)|\boldsymbol{z}']>2-\sqrt{2}$. As is detailed in the proof, this essentially restricts the sum of radii of two circular isoquant curves to be less than the length of the diagonal of $\mathcal{U}$: $(1-\Pr[\boldsymbol{D}=(1,1)|\boldsymbol{z}])+(1-\Pr[\boldsymbol{D}=(0,0)|\boldsymbol{z}'])<\sqrt{2}$. This ensures the required variation in Assumption EQ.

lemmaUnder Assumptions SS, M1, and $U_{s}\perp U_{t}$ for all $s\neq t$, Assumption EQ$^{*}$ implies Assumption EQ.

The mutual independence of $U_{s}$'s (conditional on $W$) is useful in inferring the relationship between players' interaction and instruments from the observed choices of players. The intuition for the sufficiency of Assumption EQ$^{*}$ is as follows. As long as there is no dependence in unobserved types, (ref) dictates that the variation of $\boldsymbol{Z}$ is large enough to offset strategic substitutability, because otherwise, the payoffs of players cannot move in the same direction, and thus, will not result in the same decisions. The requirement of $\boldsymbol{Z}$ variation in (ref) is significantly weaker than the large support assumption invoked for an identification at infinity argument to overcome the problem of multiple equilibria.

Partial Identification of the ATE

Under the above assumptions, we now present a generalized version of the sign matching results (ref) and (ref) in Section (ref). We need to introduce additional notation. Let $\boldsymbol{d}^{j}\in\mathcal{D}^{j}$ denote an equilibrium profile with $j$ treatments selected or $j$ entrants, i.e., a vector of $j$ ones and $S-j$ zeros, where $\mathcal{D}^{j}$ is a set of all equilibrium profiles with $j$ treatments selected. For realizations $x$ of $X$ and $\boldsymbol{z},\boldsymbol{z}'$ of $\boldsymbol{Z}$, define

align[align omitted — 438 chars of source]

Since $\sum_{j=0}^{S}\sum_{\boldsymbol{d}^{j}\in\mathcal{D}^{j}}\Pr[\boldsymbol{D}=\boldsymbol{d}^{j}|\cdot]=1$, $h(\boldsymbol{z},\boldsymbol{z}',x)=\sum_{j=0}^{S}\sum_{\boldsymbol{d}^{j}}h_{\boldsymbol{d}^{j}}(\boldsymbol{z},\boldsymbol{z}',x)$. Let $\tilde{\boldsymbol{x}}=(x_{0},...,x_{S})$ be an $(S+1)$-dimensional array of (possibly different) realizations of $X$, i.e., each $x_{j}$ for $j=0,...,S$ is a realization of $X$, and define \[ \tilde{h}(\boldsymbol{z},\boldsymbol{z}',\tilde{\boldsymbol{x}})\equiv\sum_{j=0}^{S}\sum_{\boldsymbol{d}^{j}\in\mathcal{D}^{j}}h_{\boldsymbol{d}^{j}}(\boldsymbol{z},\boldsymbol{z}',x_{j}). \] For $1\le k\le j$, define a reduction of $\boldsymbol{d}^{j}=(d_{1}^{j},...,d_{S}^{j})$ as $\boldsymbol{d}^{j-k}=(d_{1}^{j-k},...,d_{S}^{j-k})$, such that $d_{s}^{j-k}\le d_{s}^{j}$ $\forall s$. Symmetrically, for $1\le k\le S-j$, define an extension of $\boldsymbol{d}^{j}$ as $\boldsymbol{d}^{j+k}=(d_{1}^{j+k},...,d_{S}^{j+k})$, such that $d_{s}^{j+k}\ge d_{s}^{j}$ $\forall s$. For example, given $\boldsymbol{d}^{2}=(1,1,0)$, a reduction $\boldsymbol{d}^{1}$ is either $(1,0,0)$ or $(0,1,0)$ but not $(0,0,1)$, a reduction $\boldsymbol{d}^{0}$ is $(0,0,0)$, and an extension $\boldsymbol{d}^{3}$ is $(1,1,1)$. Let $\mathcal{D}^{<}(\boldsymbol{d}^{j})$ and $\mathcal{D}^{>}(\boldsymbol{d}^{j})$ be the set of all reductions and extensions of $\boldsymbol{d}^{j}$, respectively, and let $\mathcal{D}^{\le}(\boldsymbol{d}^{j})\equiv\mathcal{D}^{<}(\boldsymbol{d}^{j})\cup\{\boldsymbol{d}^{j}\}$ and $\mathcal{D}^{\ge}(\boldsymbol{d}^{j})\equiv\mathcal{D}^{>}(\boldsymbol{d}^{j})\cup\{\boldsymbol{d}^{j}\}$. Recall $\vartheta(\boldsymbol{d},x;\boldsymbol{u})\equiv E[\theta(\boldsymbol{d},x,\epsilon)|\boldsymbol{U}=\boldsymbol{u}]$. Now, we state the main lemma of this section.

lemmaIn model (ref)--(ref), suppose Assumptions SS, M1, IN, E, EX, and M hold, and $h(\boldsymbol{z},\boldsymbol{z}',x)$ and $h(\boldsymbol{z},\boldsymbol{z}',\tilde{\boldsymbol{x}})$ are well-defined. For $\boldsymbol{z},\boldsymbol{z}'$ such that (ref) and Assumption EQ hold, and for $j=1,...,S$, it satisfies that\\ (i) $sgn\{h(\boldsymbol{z},\boldsymbol{z}',x)\}=sgn\left\{ \vartheta(\boldsymbol{d}^{j},x;\boldsymbol{u})-\vartheta(\boldsymbol{d}^{j-1},x;\boldsymbol{u})\right\} $ a.e. $\boldsymbol{u}$ $\forall\boldsymbol{d}^{j-1}\in\mathcal{D}^{<}(\boldsymbol{d}^{j})$;\\ (ii) for $\iota\in\{-1,0,1\}$, if $sgn\{\tilde{h}(\boldsymbol{z},\boldsymbol{z}',\tilde{\boldsymbol{x}})\}=sgn\{-\vartheta(\boldsymbol{d}^{k},x_{k};\boldsymbol{u})+\vartheta(\boldsymbol{d}^{k-1},x_{k-1};\boldsymbol{u})\}=\iota$ $\forall\boldsymbol{d}^{k-1}\in\mathcal{D}^{<}(\boldsymbol{d}^{k})$ $\forall k\neq j$ ($k\ge1$), then $sgn\{\vartheta(\boldsymbol{d}^{j},x_{j};\boldsymbol{u})-\vartheta(\boldsymbol{d}^{j-1},x_{j-1};\boldsymbol{u})\}=\iota$ a.e. $\boldsymbol{u}$ $\forall\boldsymbol{d}^{j-1}\in\mathcal{D}^{<}(\boldsymbol{d}^{j})$.

Parts (i) and (ii) parallel (ref) and (ref), respectively. Using Lemma (ref), we can learn about the ATE. First, note that the sign of the ATE is identified by Lemma (ref)(i), since $E[Y(\boldsymbol{d})|x]=E[\vartheta(\boldsymbol{d},x;\boldsymbol{U})]$. Next, we establish the bounds on $E[Y(\boldsymbol{d}^{j})|x]$ for given $\boldsymbol{d}^{j}$ for some $j=0,...,S$.

We first present the bounds using the variation in $\boldsymbol{Z}$ only, i.e., by using Lemma (ref)(i). To this end, we fix $X=x$ and suppress it in all relevant expressions. To gain efficiency we define the integrated version of $h$ as

align[align omitted — 182 chars of source]

where $\mathcal{Z}_{EQ,j}$ is the set of $(\boldsymbol{z},\boldsymbol{z}')$ that satisfy (ref) and Assumption EQ given $j$, and $h(\boldsymbol{z},\boldsymbol{z}',x)=0$ whenever it is not well-defined. We focus on the case $H(x)>0$; $H(x)<0$ is symmetric and $H(x)=0$ is straightforward. Using Lemma (ref)(i), one can readily show that $L_{\boldsymbol{d}^{j}}(x)\le E[Y(\boldsymbol{d}^{j})|x]\le U_{\boldsymbol{d}^{j}}(x)$ with

align[align omitted — 476 chars of source]

We can simplify these bounds and show that they are sharp under the following assumption.

asC(i) $\mu_{\boldsymbol{d}}(\cdot)$ and $\nu_{\boldsymbol{d}_{-s}}(\cdot)$ are continuous; (ii) $\mathcal{Z}$ is compact.

Under Assumption C, for given $\boldsymbol{d}^{j}$, there exist vectors $\bar{\boldsymbol{z}}\equiv(\bar{z}_{1},...,\bar{z}_{S})$ and $\underline{\boldsymbol{z}}\equiv(\underline{z}_{1},...,\underline{z}_{S})$ that satisfy

equation[equation omitted — 417 chars of source]

The following is the first main result of this study, which establishes the sharp bounds on $E[Y(\boldsymbol{d}^{j})|x]$, where $X=x$ is fixed in the model.

theoremGiven model (ref)--(ref) with fixed $X=x$, suppose Assumptions SS, M1, IN, E, EX, M$^{*}$, EQ and C hold. In addition, suppose $H(x)$ is well-defined and $H(x)\ge0$. Then, the bounds $U_{\boldsymbol{d}^{j}}$ and $L_{\boldsymbol{d}^{j}}$ in (ref) and (ref) simplify to \begin{align*} U_{\boldsymbol{d}^{j}}(x) & =\Pr[Y=1,\boldsymbol{D}\in\mathcal{D}^{\ge}(\boldsymbol{d}^{j})|\bar{\boldsymbol{z}},x]+\Pr[\boldsymbol{D}\in\mathcal{D}\backslash\mathcal{D}^{\ge}(\boldsymbol{d}^{j})|\bar{\boldsymbol{z}},x],\\ L_{\boldsymbol{d}^{j}}(x) & =\Pr[Y=1,\boldsymbol{D}\in\mathcal{D}^{\le}(\boldsymbol{d}^{j})|\boldsymbol{z},x], \end{align*} and these bounds are sharp.

With binary $Y$ (Assumption M$^{*}$), sharp bounds on the mean treatment parameters can be obtained, which is reminiscent of the findings of studies that consider single-treatment models. However, our analysis is substantially different from earlier studies. In a single treatment model, SV11 use the propensity score as a scalar conditioning variable, which summarizes all the exogenous variation in the selection process and is convenient for simplifying the bounds and proving sharpness. However, in the context of our paper this approach is invalid, since $\Pr[D_{s}=1|Z_{s}=z_{s},\boldsymbol{D}_{-s}=\boldsymbol{d}_{-s}]$ cannot be written in terms of a propensity score of player $s$ as $\boldsymbol{D}_{-s}$ is endogenous. We instead use the vector $\boldsymbol{Z}$ as conditioning variables and establish partial ordering for the relevant conditional probabilities (that define the lower and upper bounds) with respect to the joint propensity score $\Pr[\boldsymbol{D}=\boldsymbol{d}|\boldsymbol{Z}=\boldsymbol{z}]$ $\forall\boldsymbol{d}\in\mathcal{D}^{\ge}(\boldsymbol{d}^{j})$.

When the variation of $X$ is additionally exploited in the model, the bounds will be narrower than the bounds in Theorem (ref). We now proceed with this case, utilizing Lemma (ref) (i) and (ii). First, analogous to (ref), we define the integrated version of $\tilde{h}(\boldsymbol{z},\boldsymbol{z}',\tilde{\boldsymbol{x}})$ as

align*[align* omitted — 229 chars of source]

where $\tilde{h}(\boldsymbol{z},\boldsymbol{z}',\tilde{\boldsymbol{x}})=0$ whenever it is not well-defined. Then, we define the following sets of two consecutive elements $(x_{j},x_{j-1})$ of $\boldsymbol{x}$ that satisfy the conditions in Lemma (ref): for $j=1,...,S$, define $\mathcal{X}_{j,j-1}^{0}(\iota)\equiv\{(x_{j},x_{j-1}):sgn\{\tilde{H}(\tilde{\boldsymbol{x}})\}=\iota,x_{0}=\cdots=x_{S}\}$ and for $t\ge1$,

align*[align* omitted — 243 chars of source]

where the sets are understood to be empty whenever $\tilde{h}(\boldsymbol{z},\boldsymbol{z}',\tilde{\boldsymbol{x}})$ is not well-defined for any $p_{M^{\le j}}(\boldsymbol{z})<p_{M^{\le j}}(\boldsymbol{z}')$ $\forall j$. Note that $\mathcal{X}_{j,j-1}^{t}(\iota)\subset\mathcal{X}_{j,j-1}^{t+1}(\iota)$ for any $t$. Define $\mathcal{X}_{j,j-1}(\iota)\equiv\lim_{t\rightarrow\infty}\mathcal{X}_{j,j-1}^{t}(\iota)$.\footnote{In practice, the formula for $\mathcal{X}_{j,j-1}^{t}$ provides a natural algorithm to construct the set $\mathcal{X}_{j,j-1}$ for the computation of the bounds. The calculation of each $\mathcal{X}_{j,j-1}^{t}$ is straightforward, as it is a search over a two-dimensional space for $(x_{j},x_{j-1})$ once the set $\mathcal{X}_{j,j-1}^{t-1}$ from the previous step is obtained. Practitioners can employ truncation $t\le T$ for some $T$ and use $\mathcal{X}_{j,j-1}^{T}$ as an approximation for $\mathcal{X}_{j,j-1}$.} Then, by Lemma (ref), if $(x_{j},x_{j-1})\in\mathcal{X}_{j,j-1}(\iota)$, then

align[align omitted — 276 chars of source]

In conclusion, for bounds on the ATE $E[Y(\boldsymbol{d}^{j})|x]$, we can introduce the sets $\mathcal{X}_{\boldsymbol{d}^{j}}^{L}(x;\boldsymbol{d}')$ and $\mathcal{X}_{\boldsymbol{d}^{j}}^{U}(x;\boldsymbol{d}')$ for $\boldsymbol{d}'\neq\boldsymbol{d}^{j}$ as follows: for $\boldsymbol{d}'\in\mathcal{D}^{<}(\boldsymbol{d}^{j})\cup\mathcal{D}^{>}(\boldsymbol{d}^{j})$,

align[align omitted — 734 chars of source]

The following theorem summarizes our results:

theoremIn model (ref)--(ref), suppose the assumptions of Lemma (ref) hold. Then the sign of the ATE is identified, and the upper and lower bounds on the ASF and ATE with $\boldsymbol{d},\tilde{\boldsymbol{d}}\in\mathcal{D}$ are \begin{align*} L_{\boldsymbol{d}}(x) & \leq E[Y(\boldsymbol{d})|x]\leq U_{\boldsymbol{d}}(x) \end{align*} and $L_{\boldsymbol{d}}(x)-U_{\tilde{\boldsymbol{d}}}(x)\leq E[Y(\boldsymbol{d})-Y(\tilde{\boldsymbol{d}})|x]\leq U_{\boldsymbol{d}}(x)-L_{\tilde{\boldsymbol{d}}}(x)$, where for any given $\boldsymbol{d}^{j}\in\mathcal{D}^{j}\subset\mathcal{D}$ for some $j$, \begin{align*} U_{\boldsymbol{d}^{j}}(x) & \equiv\inf_{\boldsymbol{z}\in\mathcal{Z}}\Biggl\{ E[Y|\boldsymbol{D}=\boldsymbol{d}^{j},\boldsymbol{z},x]\Pr[\boldsymbol{D}=\boldsymbol{d}^{j}|\boldsymbol{z}]+\Pr[\boldsymbol{D}\in\mathcal{D}^{j}\backslash\{\boldsymbol{d}^{j}\}|\boldsymbol{z}]\overline{Y}\\ & +\sum_{\boldsymbol{d}^{\prime}\in\mathcal{D}^{<}(\boldsymbol{d}^{j})\cup\mathcal{D}^{>}(\boldsymbol{d}^{j})}\inf_{x'\in\mathcal{X}_{\boldsymbol{d}^{j}}^{U}(x;\boldsymbol{d}')}E[Y|\boldsymbol{D}=\boldsymbol{d}',\boldsymbol{z},x']\Pr[\boldsymbol{D}=\boldsymbol{d}'|\boldsymbol{z}]\Biggr\},\\ L_{\boldsymbol{d}^{j}}(x) & \equiv\sup_{\boldsymbol{z}\in\mathcal{Z}}\Biggl\{ E[Y|\boldsymbol{D}=\boldsymbol{d}^{j},\boldsymbol{z},x]\Pr[\boldsymbol{D}=\boldsymbol{d}^{j}|\boldsymbol{z}]+\Pr[\boldsymbol{D}\in\mathcal{D}^{j}\backslash\{\boldsymbol{d}^{j}\}|\boldsymbol{z}]Y\\ & +\sum_{\boldsymbol{d}'\in\mathcal{D}^{<}(\boldsymbol{d}^{j})\cup\mathcal{D}^{>}(\boldsymbol{d}^{j})}\sup_{x'\in\mathcal{X}_{\boldsymbol{d}^{j}}^{L}(x;\boldsymbol{d}')}E[Y|\boldsymbol{D}=\boldsymbol{d}',\boldsymbol{z},x']\Pr[\boldsymbol{D}=\boldsymbol{d}'|\boldsymbol{z}]\Biggr\}. \end{align*}

See Sections (ref) and (ref) for concrete examples of the expression of $U_{\boldsymbol{d}^{j}}(x)$ and $L_{\boldsymbol{d}^{j}}(x)$. The terms $\Pr[\boldsymbol{D}=\boldsymbol{d}^{'}|\boldsymbol{z}]\overline{Y}$ and $\Pr[\boldsymbol{D}=\boldsymbol{d}^{'}|\boldsymbol{z}]\underline{Y}$ appear in the expression of the bounds because Lemma (ref) cannot establish an order between $\vartheta(\boldsymbol{d},x;\boldsymbol{u})$'s for $\boldsymbol{d}\in\mathcal{D}^{j}$, which is related to the complication due to multiple equilibria, which occurs for $\boldsymbol{d}\in\mathcal{D}^{j}$. When the variation in $\boldsymbol{Z}$ is only used in deriving the bounds, $\mathcal{X}_{k,k-1}(\iota)$ should simply reduce to $\mathcal{X}_{k,k-1}^{0}(\iota)$ in the definition of $\mathcal{X}_{\boldsymbol{d}^{j}}^{L}(x;\boldsymbol{d}')$ and $\mathcal{X}_{\boldsymbol{d}^{j}}^{U}(x;\boldsymbol{d}')$. When $Y$ is binary with no $X$, such bounds are equivalent to (ref) and (ref). The variation in $X$ given $\boldsymbol{Z}$ yields substantially narrower bounds than the sharp bounds established in Theorem (ref) under Assumption C. However, the resulting bounds are not automatically implied to be sharp from Theorem (ref), since they are based on a different DGP and the additional exclusion restriction.

remarkMaintaining that $Y$ is binary, sharp bounds on the ATE with variation in $X$ can be derived assuming that the signs of $\vartheta(\boldsymbol{d},x;\boldsymbol{u})-\vartheta(\boldsymbol{d}',x';\boldsymbol{u})$ are identified for $\boldsymbol{d}\in\mathcal{D}$ and $\boldsymbol{d}'\in\mathcal{D}^{<}(\boldsymbol{d})$ and $x,x'\in\mathcal{X}$ via Lemma (ref). To see this, define \begin{align*} \mathcal{\tilde{X}}_{\boldsymbol{d}}^{U}(x;\boldsymbol{d}') & \equiv\left\{ x':\vartheta(\boldsymbol{d},x;\boldsymbol{u})-\vartheta(\boldsymbol{d}',x';\boldsymbol{u})\leq0 a.e. \boldsymbol{u}\right\} ,\\ \mathcal{\tilde{X}}_{\boldsymbol{d}}^{L}(x;\boldsymbol{d}') & \equiv\left\{ x':\vartheta(\boldsymbol{d},x;\boldsymbol{u})-\vartheta(\boldsymbol{d}',x';\boldsymbol{u})\geq0 a.e. \boldsymbol{u}\right\} , \end{align*} which are identified by assumption. Then, by replacing $\mathcal{X}_{\boldsymbol{d}}^{i}(x;\boldsymbol{d}')$ with $\mathcal{\tilde{X}}_{\boldsymbol{d}}^{i}(x;\boldsymbol{d}')$ (for $i\in\{U,L\}$) in Theorem (ref), we may be able to show that the resulting bounds are sharp. Since Lemma (ref) implies that $\mathcal{X}_{\boldsymbol{d}^{j}}^{i}(x;\boldsymbol{d}')\subset\mathcal{\tilde{X}}_{\boldsymbol{d}^{j}}^{i}(x;\boldsymbol{d}')$ but not necessarily $\mathcal{X}_{\boldsymbol{d}^{j}}^{i}(x;\boldsymbol{d}')\supset\mathcal{\tilde{X}}_{\boldsymbol{d}^{j}}^{i}(x;\boldsymbol{d}')$, these modified bounds and the original bounds in Theorem (ref) do not coincide. This contrasts with the result of SV11 for a single-treatment model, and the complication lies in the fact that we deal with an incomplete model with a vector treatment. When there is no $X$, Lemma (ref)(i) establishes equivalence between the two signs, and thus, $\mathcal{X}_{\boldsymbol{d}^{j}}^{i}(x;\boldsymbol{d}')=\mathcal{\tilde{X}}_{\boldsymbol{d}^{j}}^{i}(x;\boldsymbol{d}')$ for $i\in\{U,L\}$, which results in Theorem (ref). Relatedly, we can also exploit variation in $W$, namely, variables that are common to both $X$ and $\boldsymbol{Z}$ (with or without exploiting excluded variation of $X$). This is related to the analysis of Chi10 and mourifie2015sharp in a single-treatment setting. One caveat of this approach is that, similar to these papers, we need to additionally assume that $W\perp(\epsilon,\boldsymbol{U})$.
remarkWhen $X$ does not have enough variation, we can calculate the bounds on the ATE. To see this, suppose we do not use the variation in $X$ and suppose $H(x)\ge0$. Then $\vartheta(\boldsymbol{d}^{j},x;\boldsymbol{u})\ge\vartheta(\boldsymbol{d}^{j-1},x;\boldsymbol{u})$ $\forall\boldsymbol{d}^{j-1}\in\mathcal{D}^{<}(\boldsymbol{d}^{j})$ $\forall j=1,...,S$ by Lemma (ref)(i) and by transitivity, $\vartheta(\boldsymbol{d}^{'},x;\boldsymbol{u})\ge\vartheta(\boldsymbol{d},x;\boldsymbol{u})$ with $\boldsymbol{d}'$ being an extension of $\boldsymbol{d}$. Therefore, we have \begin{align} E[Y(\boldsymbol{d})|x] & \le E[Y|\boldsymbol{D}=\boldsymbol{d},\boldsymbol{z},x]\Pr[\boldsymbol{D}=\boldsymbol{d}|\boldsymbol{z}]+\sum_{\boldsymbol{d}'\in\mathcal{D}^{>}(\boldsymbol{d})}E[Y|\boldsymbol{D}=\boldsymbol{d}',\boldsymbol{z},x]\Pr[\boldsymbol{D}=\boldsymbol{d}'|\boldsymbol{z}]\nonumber \\ & +\sum_{\boldsymbol{d}'\in\mathcal{D}\backslash\mathcal{D}^{\ge}(\boldsymbol{d})}E[Y(\boldsymbol{d}^{j})|\boldsymbol{D}=\boldsymbol{d}',\boldsymbol{z},x]\Pr[\boldsymbol{D}=\boldsymbol{d}'|\boldsymbol{z}]. \end{align} Without using variation in $X$, we can bound the last term in (ref) by $Y\in[\underline{Y},\overline{Y}]$. This is done above with $\theta(\boldsymbol{d},x,\epsilon)=1[\mu(\boldsymbol{d},x)\geq\epsilon_{\boldsymbol{d}}]$ and $\vartheta(\boldsymbol{d},x;\boldsymbol{u})=F_{\epsilon|\boldsymbol{U}}(\mu(\boldsymbol{d},x)|\boldsymbol{u})$.

Numerical Study

To illustrate the main results of this study in a simulation exercise, we calculate the bounds on the ATE using the following data generating process:

align*[align* omitted — 199 chars of source]

where $(\epsilon,V_{1},V_{2})$ are drawn, independent of $(X,\boldsymbol{Z})$, from a joint normal distribution with zero means and each correlation coefficient being $0.5$. We draw $Z_{s}$ ($s=1,2$) and $X$ from a multinomial distribution, allowing $Z_{s}$ to take two values, $\mathcal{Z}_{s}=\{-1,1\}$, and $X$ to take either three values, $\mathcal{X}=\{-1,0,1\}$, or fifteen values, $\mathcal{X}=\{-1,-\frac{6}{7},-\frac{5}{7},...,\frac{5}{7},\frac{6}{7},1\}$. Being consistent with Assumption M, we choose $\tilde{\mu}_{11}>\tilde{\mu}_{10}$ and $\tilde{\mu}_{01}>\tilde{\mu}_{00}$. Let $\tilde{\mu}_{10}=\tilde{\mu}_{01}$. With Assumption SS, we choose $\delta_{1}<0$ and $\delta_{2}<0$. Without loss of generality, we choose positive values for $\gamma_{1}$, $\gamma_{2}$, and $\beta$. Specifically, $\tilde{\mu}_{11}=0.25$, $\tilde{\mu}_{10}=\tilde{\mu}_{01}=0$ and $\tilde{\mu}_{00}=-0.25$. For default values, $\delta_{1}=\delta_{2}\equiv\delta=-0.1$, $\gamma_{1}=\gamma_{2}\equiv\gamma=1$ and $\beta=0.5$.

In this exercise, we focus on the ATE $E[Y(1,1)-Y(0,0)\vert X=0]$, whose true value is $0.2$ given our choice of parameter values. For $h(\boldsymbol{z},\boldsymbol{z}',x)$, we consider $\boldsymbol{z}=(1,1)$ and $\boldsymbol{z}'=(-1,-1)$. Note that $H(x)=h(\boldsymbol{z},\boldsymbol{z}',x)$ and $\tilde{H}(x,x',x'')=\tilde{h}(\boldsymbol{z},\boldsymbol{z}',x,x',x'')$, since $Z_{s}$ is binary. Then, we can derive the sets $\mathcal{X}_{\boldsymbol{d}}^{U}(0;\boldsymbol{d}')$ and $\mathcal{X}_{\boldsymbol{d}}^{L}(0;\boldsymbol{d}')$ for each $\boldsymbol{d}\in\{(1,1),(0,0)\}$ and $\boldsymbol{d}'\neq\boldsymbol{d}$ in Theorem (ref).

Based on our design, $H(0)>0$, and thus, the bounds when we use $Z$ only are, with $x=0$, \[ \max_{\boldsymbol{z}\in\mathcal{Z}}\Pr[Y=1,\boldsymbol{D}=(0,0)\vert\boldsymbol{z},x]\le\Pr[Y(0,0)=1\vert x]\le\min_{\boldsymbol{z}\in\mathcal{Z}}\Pr[Y=1\vert\boldsymbol{z},x], \] and \[ \max_{\boldsymbol{z}\in\mathcal{Z}}\Pr[Y=1\vert\boldsymbol{z},x]\le\Pr[Y(1,1)=1\vert x]\le\min_{\boldsymbol{z}\in\mathcal{Z}}\left\{ \Pr[Y=1,\boldsymbol{D}=(1,1)\vert\boldsymbol{z},x]+1-\Pr[\boldsymbol{D}=(1,1)\vert\boldsymbol{z},x]\right\} . \] Using both $\boldsymbol{Z}$ and $X$, we obtain narrower bounds. For example, when $\left|\mathcal{X}\right|=3$, with $\tilde{H}(0,-1,-1)<0$, the lower bound on $\Pr[Y(0,0)=1\vert X=0]$ becomes \[ \max_{\boldsymbol{z}\in\mathcal{Z}}\left\{ \Pr[Y=1,\boldsymbol{D}=(0,0)\vert\boldsymbol{z},0]+\Pr[Y=1,\boldsymbol{D}\in\{(1,0),(0,1)\}\vert\boldsymbol{z},-1]\right\} . \] With $\tilde{H}(1,1,0)<0$, the upper bound on $\Pr[Y(1,1)=1\vert X=0]$ becomes \[ \min_{\boldsymbol{z}\in\mathcal{Z}}\left\{ \Pr[Y=1,\boldsymbol{D}=(1,1)\vert\boldsymbol{z},0]+\Pr[Y=1,\boldsymbol{D}\in\{(1,0),(0,1)\}\vert\boldsymbol{z},1]+\Pr[\boldsymbol{D}=(0,0)\vert\boldsymbol{z},0]\right\} . \] For comparison, we calculate the bounds in manski1990nonparametric using $\boldsymbol{Z}$. These bounds are given by

align*[align* omitted — 284 chars of source]

and

align*[align* omitted — 284 chars of source]

We also compare the estimated ATE using a standard linear IV model in which the nonlinearity of the true DGP is ignored:

align*[align* omitted — 408 chars of source]

Here, the first stage is the reduced-form representation of the linear simultaneous equations model for strategic interaction. Under this specification, the ATE becomes $E[Y(1,1)-Y(0,0)\vert X=0]=\pi_{1}+\pi_{2}$, which is estimated via two-stage least squares (TSLS).

The bounds calculated for the ATE are shown in Figures (ref)--(ref). Figure (ref) shows how the bounds on the ATE change, as the value of $\gamma$ changes from $0$ to $2.5$. The larger $\gamma$ is, the stronger the instrument $\boldsymbol{Z}$ is. The first conspicuous result is that the TSLS estimate of the ATE is biased because of the problem of misspecification. Next, as expected, Manski's bounds and our proposed bounds converge to the true value of the ATE as the instrument becomes stronger. Overall, our bounds, with or without exploiting the variation of $X$, are much narrower than Manski's bounds.\footnote{Although we do not make a rigorous comparison of the assumptions here, note that the bounds by MP00 under the semi-MTR is expected to be similar to ours. However, their bounds need to assume the direction of the monotonicity.} Notice that the sign of the ATE is identified in the whole range of $\gamma$, as predicted by the first part of Theorem (ref), in contrast to Manski's bounds. Using the additional variation in $X$ with $\left|\mathcal{X}\right|=3$ decreases the width of the bounds, particularly with the smaller upper bounds on the ATE in this simulation design. Figure (ref) depicts the bounds using $X$ with $\left|\mathcal{X}\right|=15$, which yields narrower bounds than when $\left|\mathcal{X}\right|=3$, and substantially narrower than those only using $\boldsymbol{Z}$.

Figure (ref) shows how the bounds change as the value of $\beta$ changes from $0$ to $1.5$, where a larger $\beta$ corresponds to a stronger exogenous variable $X$. The jumps in the upper bound are associated with the sudden changes in the signs of $\tilde{H}(-1,0,-1)$ and $\tilde{H}(0,1,1)$. At least in this simulation design, the strength of $X$ is not a crucial factor for obtaining narrower bounds. In fact, based on other simulation results (omitted in the paper), we conclude that the number of values $X$ can take matters more than the dispersion of $X$ (unless we pursue point identification of the ATE).

Finally, Figure (ref) shows how the width of the bounds is related to the extent to which the opponents' actions $D_{-s}$ affect one's payoff, captured by $\delta$. We vary the value of $\delta$ from $-2$ to $0$, and when $\delta=0$, the players solve a single-agent optimization problem. Thus, heuristically, the bound at this point would be similar to the ones that can be obtained when SV11 is extended to a multiple-treatment setting with no simultaneity. In the figure, as the value of $\delta$ becomes smaller, the bounds get narrower.

Empirical Application: Airline Markets and Pollution

Aircrafts are a major source of emissions, and thus, quantifying the causal effect of air transport on pollution is of importance to policy makers. Therefore, in this section, we take the bounds proposed in Section (ref) to data on airline market structure and air pollution in cities in the U.S.

In 2013, aircrafts were responsible for about 3 percent of total U.S. carbon dioxide emissions and nearly 9 percent of carbon dioxide emissions from the U.S. transportation sector, and it is one of the fastest growing sources.\footnote{See https://www.c2es.org/content/reducing-carbon-dioxide-emissions-from-aircraft/7/} Airplanes remain the single largest source of carbon dioxide emissions within the U.S. transportation sector, which is not yet subject to greenhouse gas regulations. In addition to aircrafts, airport land operations are also a big source of pollution, making airports one of the major sources of air pollution in the U.S. For example, 43 of the 50 largest airports are in ozone non-attainment areas and 12 are in particulate matter non-attainment areas.\footnote{Ozone is not emitted directly but is formed when nitrogen oxides and hydrocarbons react in the atmosphere in the presence of sunlight. In United States environmental law, a non-attainment area is an area considered to have air quality worse than the National Ambient Air Quality Standards as defined in the Clean Air Act.}

There is growing literature showing the effects of air pollution on various health outcomes (see, schlenker2015airports, cg2003, kms2011). In particular, schlenker2015airports show that the causal effect of airport pollution on the health of local residents---using a clever instrumental variable approach---is sizable. Their study focuses on the 12 major airports in California and implicitly assume that the level of competition (or market structure) is fixed. Using high-frequency data, they exploit weather shocks in the East coast---that might affect airport activity in California through network effects---to quantify the effect of airport pollution on respiratory and cardiovascular health complications. In contrast, we take the link between airport pollution and health outcomes as given and are interested in quantifying the effects of different market structures of the airline industry on air pollution.\footnote{In this section, we refer to market structure as the particular configuration of airlines present in the market. In other words, market structure not only refers to the number of firms competing in a given market but to the actual identities of the firms. Thus, we will regard a market in which, say, United and American operate as different from a market in which Southwest and Delta operate, despite both markets having two carriers.} We explicitly allow market structure to be determined endogenously, as the outcome of an entry game in which airlines behave strategically to maximize their profits and the resulting pollution in this market is not internalized by the firms. Understanding these effects can then help inform the policy discussion on pollution regulation. Given that we treat market structure as endogenous, one cannot simply run a regression of a measure of pollution on market structure (or the number of airlines present in a market) to obtain the causal effect, if there are unobserved market characteristics that affect both firm competition and pollution outcomes. For example, if at both ends of a city-pair there are firms from a high-polluting industry which engage in a lot of business travel, it would drive both pollution and the entry of airlines in the market. Therefore, we use the method presented in Section (ref).

In each market, we assume that a set of airlines chooses to be in or out as part of a simultaneous entry game of perfect information, as introduced in Section (ref). Therefore, we treat market structure as our endogenous treatment. We then model air pollution as a function of the airline market structure as in equation ((ref)), where $Y$ is a measure of air pollution at the airport level (including both aircraft and land operation pollution), the vector $\boldsymbol{D}$ represents the market structure, and $X$ includes market specific covariates that affect pollution directly (i.e., not through airline activity), such as weather shocks or the share of pollution-related activity in the local economy.\footnote{Note that our definition of market is a city-pair; hence, all of our variables are, in fact, weighted averages over the two cities.} Additionally, we allow for market-level covariates, $\boldsymbol{W}$, which affect both the participation decisions and pollution (e.g., the size of the market as measured by population or the level of economic activity). As instruments, $Z_{s}$, we consider a firm-market proxy for cost introduced in CT09. We discuss the definition and construction of the variables in detail below.

Our objective is to estimate the effect of a change in market structure on air pollution, $E[Y(\boldsymbol{d})-Y(\boldsymbol{d}')]$. For example, we might be interested in the average effect on pollution of moving from two airlines operating in the market to three, or how the pollution level changes on average when Delta is a monopolist versus a situation in which Delta faces competition from American. Following entry, firms compete by choosing their pricing, frequency, and which airplanes to operate. Different market structures will have different impacts on these variables, which in turn, affect the level of pollution. Note that the effects might be asymmetric. That is, for a given number of entrants, their identities are important to determine the pollution level. For example, when comparing the effect of a monopoly on pollution, we find that the airline operating plays a role.

To illustrate our estimation procedure, we consider three types of ATE exercises. The first examines the effects on pollution from a monopolist airline vis-a-vis a market that is not served by any airline. The second set of exercises examine the total effect of the industry on pollution under all possible market configurations. Finally, the third type of exercises examine how the (marginal) effect of a given airline changes when the firm faces different levels of competition. Notice that regardless of the exercise we run, we quantify “reduced-form” effects, in that they summarize structural effects resulting from a given market structure. The idea is that given the market structure, prices are determined, and given demand, ultimately the frequency of flights in the market is determined, which in fact, causes pollution.

In the rest of this section, we first describe our data sources, then show results for three different ATE exercises, and conclude with a brief discussion relating our results to potential policy recommendations.

Data

For our analysis, we combine data spanning the period 2000--2015 from two sources: airline data from the U.S. Department of Transportation and pollution data from the Environmental Protection Agency (EPA).

Airline Data. Our first data source contains airline information and combines publicly available data from the Department of Transportation's Origin and Destination Survey (DB1B) and Domestic Segment (T-100) database. These datasets have been used extensively in the literature to analyze the airline industry (see, e.g., borenstein89, berry1992estimation, CT09, and more recently, robertssweeting2013 and ciliberto2015market). The DB1B database is a quarterly sample of all passenger domestic itineraries. The dataset contains coupon-specific information, including origin and destination airports, number of coupons, the corresponding operating carriers, number of passengers, prorated market fare, market miles flown, and distance. The T-100 dataset is a monthly census of all domestic flights broken down by airline, and origin and destination airports.

Our time-unit of analysis is a quarter and we define a market as the market for air connection between a pair of airports (regardless of intermediate stops) in a given quarter.\footnote{In cities that operate more than one airport, we assume that flights to different airports in the same metropolitan area are in separate markets.} We restrict the sample to include the top 100 metropolitan statistical areas (MSA's), ranked by population at the beginning of our sample period. We follow berry1992estimation and CT09 and define an airline as actively serving a market in a given quarter, if we observe at least 90 passengers in the DB1B survey flying with the airline in the corresponding quarter.\footnote{This corresponds to approximately the number of passengers that would be carried on a medium-size jet operating once a week.} We exclude from our sample city pairs in which no airline operates in the whole sample period. Notice that we do include markets that are temporarily not served by any airline. This leaves us with 181,095 market-quarter observations.

In our analysis, we allow for airlines to have a heterogeneous effect on pollution, and to simplify computation, in each market we allow for six potential participants: American (AA), Delta (DL), United (UA), Southwest (WN), a medium-size airline, and a low-cost carrier.\footnote{That is, to limit the number of potential market structures, we lump together all the low cost carriers into one category, and Northwest, Continental, America West, and USAir under the medium airline type.} The latter is not a bad approximation to the data in that we rarely observe more than one medium-size or low-cost in a market but it assumes that all low-cost airlines have the same strategic behavior, and so do the medium airlines. Table (ref) shows the number of firms in each market broken down by size as measured by population. As the table shows, market size alone does not explain market structure, a point first made by CT09.

table[table omitted — 802 chars of source]

In our application, we consider two instruments for the entry decisions. The first is the airport presence of an airline proposed by berry1992estimation. For a given airline, this variable is constructed as the number of markets it serves out of an airport as a fraction of the total number of markets served by all airlines out of the airport. A hub-and-spoke network allows firms to exploit demand-side and cost-side economies, which should affect the firm's profitability. While berry1992estimation assumes that an airline's airport presence only affects its own profits (and hence, is excluded from rivals' profits), CT09 argue that this may not be the case in practice, since airport presence might be a measure of product differentiation, rendering it likely to enter the profit function of all firms through demand. While an instrument that enters all of the profit functions is fine in our context (see Appendix (ref)), we also consider the instrument proposed by CT09, which captures shocks to the fixed cost of providing a service in a market. This variable, which they call cost, is constructed as the percentage of the nonstop distance that the airline must travel in excess of the nonstop distance, if the airline uses a connecting instead of a nonstop flight.\footnote{Mechanically, the variable is constructed as the difference between the sum of the distances of a market's endpoints and the closest hub of an airline, and the nonstop distance between the endpoints, divided by the nonstop distance.} Arguably, this variable only affects its own profits and is excluded from rivals' profits.

table[table omitted — 777 chars of source]

Table (ref) presents the summary statistics of the airline related variables. Of the leading airlines, we see that American and Delta are present in about half of the markets, while United and Southwest are only present in about a quarter of the markets. American and Delta tend to dominate the airports in which they operate more than United and Southwest. From the cost variable, we see that both American and United tend to operate a hub-and-spoke network, while Southwest (and to a lesser extent Delta) operates most markets nonstop.

Pollution Data. The second component of our dataset is the air pollution data. The EPA compiles a database of outdoor concentrations of pollutants measured at more than 4,000 monitoring stations throughout the U.S., owned and operated mainly by state environmental agencies. Each monitoring station is geocoded, and hence, we are able to merge these data with the airline dataset by matching all the monitoring stations that are located within a 10km radius of each airport in our first dataset.

The principal emissions of aircraft include the greenhouse gases carbon dioxide ($\text{CO}_{2}$) and water vapor ($\text{H}_{2}\text{O}$), which have a direct impact on climate change. Aircraft jet engines also produce nitric oxide ($\text{NO}$) and nitrogen dioxide ($\text{NO}_{2}$) (which together are termed nitrogen oxides ($\text{NO}_{\text{x}}$)), carbon monoxide (CO), oxides of sulphur ($\text{SO}_{\text{x}}$), unburned or partially combusted hydrocarbons (also known as volatile organic compounds or VOC's), particulates, and other trace compounds (see, FAA2015). In addition, ozone ($\text{O}_{3}$) is formed by the reaction of VOC's and $\text{NO}_{\text{x}}$ in the presence of heat and sunlight. The set of pollutants other than $\text{CO}_{2}$ are more pernicious in that they can harm human health directly and can result in respiratory, cardiovascular, and neurological conditions. Research to date indicates that fine particulate matter (PM) is responsible for the majority of the health risks from aviation emissions, although ozone has a substantial health impact too.\footnote{See FAA2015.} Therefore, as our measure of pollution, we will consider both.

Our measure of ozone is a quarterly mean of daily maximum levels in parts per million. In terms of PM, as a general rule, the smaller the particle the further it travels in the atmosphere, the longer it remains suspended in the atmosphere, and the more risk it poses to human health. PM that measure less than 2.5 micrometer can be readily inhaled, and thus, potentially pose increased health risks. The variable PM2.5 is a quarterly average of daily averages and is measured in micrograms/cubic meter. For each airport in our sample, we take an average (weighted by distance to the airport) of the data from all air monitoring stations within a 10km radius. The top panel of Table (ref) shows the summary statistics of the pollution measures.

table[table omitted — 655 chars of source]

Other Market-Level Controls. We also include in our analysis market-level covariates that may affect both market structure and pollution levels. In particular, we construct a measure of market size by computing the (geometric) mean of the MSA populations at the market endpoints and a measure of economic activity by computing the average per capita income at the market endpoints, using data from the Regional Economic Accounts of the Bureau of Economic Analysis.

Finally, as we mentioned in Section (ref), having access to data on a variable that affects pollution but is excluded from the airline participation decisions can greatly help in calculating the bounds of the ATE. Therefore, we construct a variable that measures the economic activity of pollution related industries (manufacturing, construction, and transportation other than air transportation) in a given market (MSA) as a fraction of total economic activity in that market, again, using data from the Regional Economic Accounts of the Bureau of Economic Analysis. Our implicit assumption is as follows. The size of the market, among other things, determines whether a firm might enter it, but not the type of economic activity in the cities. The idea is that conditional on the market GDP, a market with a higher share of polluting industries will have a higher level of pollution but this share would not affect the airline market structure.

Estimation and Results

To simplify the estimation, we discretize all continuous variables into binary variables (taking a value of 0 (1) if the corresponding continuous variable is below (above) its median). Using the notation from Section (ref), let the elements of the treatment vector $\boldsymbol{d}=(d_{\text{DL}},d_{\text{AA}},d_{\text{UA}},d_{\text{WN}},d_{\text{med}},d_{\text{low}})$ be either 0 or 1, indicating whether each firm is active in the market. We compute the upper and lower bounds on the ATE using the result from Theorem (ref) and the fact that our $Y$ variable is binary. Specifically, given two treatment vectors $\boldsymbol{d}$ and $\tilde{\boldsymbol{d}}$ we can bound the ATE

align*[align* omitted — 165 chars of source]

where the upper bound can be characterized by

align*[align* omitted — 798 chars of source]

for every $\boldsymbol{z}$, $x'(\boldsymbol{d}')\in\mathcal{X}_{\boldsymbol{d}}^{U}(x;\boldsymbol{d}')$ for $\boldsymbol{d}'\neq\boldsymbol{d}$, and $x''(\boldsymbol{d}'')\in\mathcal{X}_{\tilde{\boldsymbol{d}}}^{L}(x;\boldsymbol{d}'')$ for $\boldsymbol{d}''\neq\tilde{\boldsymbol{d}}$. Similarly, the lower bound can be characterized by

align*[align* omitted — 827 chars of source]

for every $\boldsymbol{z}$, $x'(\boldsymbol{d}')\in\mathcal{X}_{\boldsymbol{d}}^{L}(x;\boldsymbol{d}')$ for $\boldsymbol{d}'\neq\boldsymbol{d}$, and $x''(\boldsymbol{d}'')\in\mathcal{X}_{\tilde{\boldsymbol{d}}}^{U}(x;\boldsymbol{d}'')$ for $\boldsymbol{d}''\neq\tilde{\boldsymbol{d}}$. We estimate the population objects above using their sample counterparts. We experimented with both measures of pollution discussed earlier and obtain qualitatively and quantitatively similar results in all cases, which is not surprising given that the two pollution measures are highly correlated. In order to save space, we only show results using PM2.5 as our outcome variable. We also experimented with several specifications of the covariates, $\boldsymbol{X}$ and $\boldsymbol{W}$, and instruments, $\boldsymbol{Z}$. In particular, we tried different discretizations of each variable (including allowing for more than two points in their supports and different cutoffs). Clearly, there is a limit to how finely we can cut the data even with a large sample size such as ours. The coarser discretization occurs when each covariate (and instrument) is binary and it seems to produce reasonable results; hence, we stick with this discretization in all of our exercises. Again, aiming at the most parsimonious model, and after some experimentation, we obtained reasonable results when both $\boldsymbol{X}$ and $\boldsymbol{W}$ are scalars (share of pollution related industries in the market and total GDP in the market, respectively).

We also compute confidence sets by deriving unconditional moment inequalities from our conditional moment inequalities and implementing the Generalized Moment Selection test proposed by as2010. The confidence sets are obtained by inverting the test.\footnote{For details of this procedure, see dm2016.}

figure*[figure* omitted — 462 chars of source]

Monopoly Effects. Here we examine a very simple ATE of a change in market structure from no airline serving a market to a monopolist serving it. Intuitively, we want to understand the change in the probability of being a high-pollution market when an airline starts operating on it. Recall that we allow each firm to have different effects on pollution; hence, we estimate the effects of each one of the six firms in our data becoming a monopolist. Thus, we are interested in the ATEs of the form \[ E[Y(\boldsymbol{d}_{\text{monop}})-Y(\boldsymbol{d}_{\text{noserv}})|X,W] \] where $\boldsymbol{d}_{\text{monop}}$ is one of the six vectors in which only one element is 1 and the rest are 0's, and $\boldsymbol{d}_{\text{noserv}}$ is a vector of all 0's. The results are shown in Figure (ref), where the solid black intervals are our estimates of the identified sets and the thin red lines are the 95% confidence sets. We see that all ATEs are positive and statistically significant different from 0, except for the medium-size carriers. While there no major differences on the effects of the leader carriers, with the exception of Delta which seems to induce a higher increase in the probability of high pollution, the medium and low-cost carriers induce a smaller effect.

figure*[figure* omitted — 525 chars of source]

Total Market Structure Effect. We now turn to our second set of exercises. Here, we quantify the effect of the airline industry on the likelihood of a market having high levels of pollution. To do so, we estimate ATEs of the form \[ E[Y(\boldsymbol{d})-Y(\boldsymbol{d}_{\text{noserv}})|X,W] \] for all potential market configurations $\boldsymbol{d}$, and where, as before, $\boldsymbol{d}_{\text{noserv}}$ is a vector of all 0's. Figure (ref) depicts the results. The left-most set of intervals corresponds to the 6 different monopolistic market structures, and by construction, coincide with those from Figure (ref). The next set corresponds to all possible duopolistic structures, which has 15 possibilities, and so on. Not surprisingly, we observe that the effect on the probability of being a high-pollution market is increasing in the number of firms operating in the market. More interesting is the non-linearity of the effect: the effect increases at a decreasing rate. This would be consistent with a model in which firms accommodate new entrants by decreasing their frequency, which is analogous to the prediction of a Cournot competition model, as we increase the number of firms. To further investigate this point, in the next set of exercises, we examine the effect of one firm as we change the competition it faces.

figure*[figure* omitted — 567 chars of source]

Marginal Carrier Effect. In our last set of exercises, we are interested in investigating how the marginal effect (i.e., the effect of introducing one more firm into the market) changes under different configurations of the market structure. Say we are interested in the effect of Delta entering the market on pollution, given that the current market structure (excluding Delta) is $\boldsymbol{d}_{\text{--DL}}=(d_{\text{AA}},d_{\text{UA}},d_{\text{WN}},d_{\text{med}},d_{\text{low}})$. Then, we want to estimate \[ E[Y(1,\boldsymbol{d}_{\text{--DL}})-Y(0,\boldsymbol{d}_{\text{--DL}})|X,W]. \]

Figure (ref) shows the identified sets and confidence sets of the marginal effect of Delta on the probability of high pollution under all possible market configuration for Delta's rivals. We obtain qualitatively similar results when estimating the marginal effects of the other five carriers, and hence, we omit the graphs to save space. In the Figure, the left-most exercise is the effect of Delta as a monopolist, and coincides, by construction, with the left-most exercise in Figure (ref). The second exercise (from the left) corresponds to the additional effect of Delta on pollution when there is already one firm operating in the market, which yields five different possibilities. The next exercise shows the effect of Delta when there are two firms already operating in the market yielding 10 possibilities, and so on. Again, the estimated marginal ATEs in all cases are positive and statistically significant. Interestingly, although we cannot entirely reject the null hypothesis that all the effects are the same, it seems that the marginal effect of Delta is decreasing in the number of rivals it faces. Intuitively, this suggests a situation in which Delta enters the market and operates with a frequency that is decreasing with the number of rivals (again, as we would expect in a Cournot competition model) and is consistent with the findings in our previous set of exercises.

The conclusions from the total market and marginal ATEs are also interesting from a policy perspective. For example, a merger of two airlines in which duplicate routes are eliminated would imply a decrease in total pollution in the affected markets, but by less than what one would have naively anticipated from removing one airline while keeping everything else constant.\footnote{Note that, however, to the extent that a merger alters the way the merged firm behaves post entry, the treatment effects we estimate will not be informative. In other words, our model can only speak to the effects of a merger on pollution that only affects behavior through entry.} In other words, there are two effects of removing an airline from a market. The first is a direct affect: pollution decreases by the amount of pollution by the carrier that is no longer present in the market. However, the remaining firms in the market will react strategically to the new market structure. In our exercises, we find that this indirect effect implies an increase in pollution. The overall effect is a net decrease in pollution. Moreover, given the non-linearities of the ATEs we estimate it looks like the overall effect, while negative, might be negligible in markets with four or more competitors. While it is unclear that merger analysis, which is typically concerned with price increases post-merge or cost savings of the merging firms, should also consider externalities such as pollution, (social) welfare analysis should. Hence, our findings may serve as guidance to policy discussion on air traffic regulation.