EconBase
← Back to paper

Demand and Welfare Analysis in Discrete Choice Models with Social Interactions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

149,170 characters · 16 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Demand and Welfare Analysis in Discrete Choice Models with Social Interactions

abstractMany real-life settings of individual choice involve social interactions, causing targeted policies to have spillover effects. This paper develops novel empirical tools for analyzing demand and welfare effects of policy interventions in binary choice settings with social interactions. Examples include subsidies for health product adoption and vouchers for attending a high-achieving school. We show that even with fully parametric specifications and unique equilibrium, choice data, that are sufficient for counterfactual demand prediction under interactions, are insufficient for welfare calculations. This is because distinct underlying mechanisms producing the same interaction coefficient can imply different welfare effects and deadweight-loss from a policy intervention. Standard index restrictions imply distribution-free bounds on welfare. We propose ways to identify and consistently estimate the structural parameters and welfare bounds allowing for unobserved group effects that are potentially correlated with observables and are possibly unbounded. We illustrate our results using experimental data on mosquito-net adoption in rural Kenya.

INTRODUCTION

Social interaction models -- where an individual's payoff from an action depends on aggregate choice -- feature prominently in economic and sociological research. In this paper, we address a substantively important issue that has received limited attention within these literatures, viz., how to conduct welfare analysis of economic policy intervention in such settings. Examples include subsidies for adopting a health product and merit-based vouchers for attending a high-achieving school, where the welfare gain of beneficiaries may be accompanied by spillover-led welfare effects on those unable to adopt or move, respectively. Ex-ante welfare analysis of policies is ubiquitous in economic applications, and informs the practical decision of whether to implement the policy in question. Furthermore, common public interventions such as taxes and subsidies are often motivated by efficiency losses resulting from externalities. Therefore, it is important to develop empirical methods for welfare analysis in presence of such externalities, which cannot be done using available tools in the literature. Developing such methods and making them practically relevant also requires one to clarify and extend some aspects of existing empirical models of social interaction.

\noindentLiterature Review and Contributions: Seminal contributions to the econometrics of social interactions include Manski (1993) for continuous outcomes, and Brock and Durlauf (2001) (henceforth, BD01) for binary outcomes. Bisin, Moro and Topa (2011) discuss some issues related to identification and estimation of structural parameters in choice models with social interaction and multiple equilibria but do not cover welfare analysis. More recently, there has been a surge of research on the related theme of network models, cf. de Paula (2017) who provides a comprehensive review of the relevant literature. On the other hand, the econometric analysis of welfare in standard discrete choice settings, i.e., with heterogeneous consumers but without social spillover, started with Domencich and McFadden (1977), with later contributions by Daly and Zachary (1978), Small and Rosen (1981), and Bhattacharya (2018). The present paper builds on these two separate literatures to examine how social interactions influence welfare effects of policy interventions and the identifiability of such welfare effects from standard choice data. In the context of a logit binary choice model with social interactions, BD01 Sec 3.3 equations (16) and (17) discussed how to infer the sign of the differences between expected ex ante (indirect) utility at each possible equilibria resulting from the policy intervention being studied. This differs from the average of individual compensating variations that restore realized individual utilities to their pre-intervention level which is a money metric, unlike the BD01 measure, and hence can be directly compared with the cost of the intervention, yielding a theoretically justified measure of deadweight loss. Consequently, this measure has received the most attention in the recent literature on applied welfare analysis, cf. Hausman and Newey (2016), Bhattacharya (2015), McFadden and Train (2019). However, in settings involving spillover, we cannot use the methods of the above papers, as they do not allow for individual utilities to be affected by aggregate choices -- a feature that has fundamental implications for welfare analysis. Therefore, new methods are required for welfare calculations under spillover, which we develop in the present paper.

Our starting point is a theoretically coherent empirical model where many individuals with some observed and some unobserved attributes interact with each other to produce the aggregate choice in equilibrium before and after the policy intervention. Individual choice data can be used to estimate identifiable parameters of this model in BD01, which can then be used to predict counterfactual demand, i.e. equilibrium choice probabilities resulting from a hypothetical price intervention, e.g., a price subsidy. However, we show that unlike counterfactual demand estimation, welfare effects are generically not identified from standard choice data under interactions, even when utilities and the distribution of unobserved heterogeneity are parametrically specified, equilibrium is unique, and there are no endogeneity concerns. To understand the heuristics behind under-identification, consider the empirical example of evaluating the welfare effect of subsidizing an anti-malarial, insecticide-treated mosquito net. Suppose, under suitable restrictions, we can model choice behavior in this setting via a Brock-Durlauf type social interaction model, and the data can identify the coefficient on the social interaction term. However, this coefficient may reflect an aggregate effect of several distinct mechanisms, viz. (a) a social preference for conforming, (b) learning from others' experiences, (c) a health-concern led desire to protect oneself from mosquitoes deflected from neighbors who adopt a bednet, and (d) desire to free-ride on other users who increase herd-immunity by protecting themselves and/or protect neighbors via the insecticide effects. These distinct mechanisms, with different magnitudes in general, can make the social interaction coefficient positive, but are not separately identifiable from choice data (only their sum is). But they have different implications for welfare if, say, a subsidy is introduced. In particular, if spillovers are all due to preference for social conforming or learning and there is no (perceived) health externality, then as more neighbors buy, a household's perceived utility from buying will increase over and above the gain due to price reduction. At the other extreme, if spillovers are solely due to perceived negative health externality of buyers on non-buyers, then increased purchase by neighbors would lower the utility of a household upon not buying via the health-route, but not affect it upon buying since the household is then protected anyway. These different aggregate welfare effects are both consistent with the same positive aggregate social interaction coefficient. This conclusion continues to hold even if eligibility for the subsidy is universal and there are no income effects or endogeneity concerns.

This feature is present in many other choice situations that economists routinely study. For example, merit-based school vouchers for attending a high-achieving school can potentially have a range of possible welfare effects. Aggregate welfare change could be negative if, for example, with high-ability children moving with the voucher the academic quality declines in the resource-poor schools more than the improvement in the selective school via peer effects. In the absence of such negative externalities, aggregate welfare could be positive due to the subsidy-led price decline for voucher users and any positive conforming effects that raise the utility of attending the high-achieving school when more high-ability children also do so. These contradictory welfare implications are compatible with the same positive coefficient on the social interaction term in an individual school choice model.

For standard discrete choice without spillover, Bhattacharya (2015) showed that the choice probability function itself contains all the information required for exact welfare analysis. For the special case of quasilinear random utility models with extreme value errors, the popular `logsum' formula of Small and Rosen (1981) yields average welfare of policy interventions. These results fail to hold in a setting with spillovers because here one cannot set\ the utility from the outside option to zero -- an innocuous normalization in standard discrete choice models -- since this utility changes as the equilibrium choice-rate changes with the policy intervention. This is in contrast to binary choice without spillover, where utility from the outside option, i.e., non-purchase, does not change due to a price change of the inside good.

Nonetheless, under a standard, linear-index specification of utilities, one can calculate distribution-free bounds on average welfare, based solely on choice probability functions. The width of the bounds increases with (i) the extent of net social spillover, i.e. how much the (belief about) average neighborhood choice affects individual choice probabilities, and (ii) the difference in average peer-choice corresponding to realized equilibria before and after the price change. The index structure, which has been universal in the empirical literature on social interactions, leads to dimension reduction that helps identify spillovers effects. We therefore continue to use the index structure as it simplifies our expressions, and comes \textquotedblleft for free\textquotedblright, because social spillovers cannot in general be identified without such structure anyway. Under stronger and untestable restrictions on the nature of spillover, our bounds can shrink to a singleton, implying point-identification of welfare. Two such restrictions are (a) the effects of an increase in average peer-choice on individual utilities from buying and not buying are exactly equal in magnitude and opposite in sign, or (b) the effect of aggregate choice on either the purchase utility or the non-purchase utility is zero.

A separate identification problem arises when there are, in addition to social interaction, unobserved group-effects that are potentially correlated with observed individual covariates. We address this problem through a novel latent factor structure on the relevant variables, and developing a method of asymptotic analysis where the dimension of parameters, i.e. the group-effects whose magnitude may be unbounded, increases as the number of groups increase.

\noindentEmpirical Illustration: We illustrate our theoretical results with an empirical example of a hypothetical, targeted public subsidy scheme for anti-malarial bednets. In particular, we use micro-data from a pricing experiment in rural Kenya (Dupas, 2014) to estimate an econometric model of demand for bednets, where spillovers can arise via different channels, including a preference for conformity and perceived negative externality arising from neighbors' use of a bednet. In this setting, we calculate predicted effects of hypothetical income-contingent subsidies on bednet demand and welfare. We perform these calculations by first accounting for social interactions, and then compare these results with what would be obtained if one had ignored these interactions. We find that allowing for (positive) interaction leads to a prediction of lower demand when means-tested eligibility is restricted to fewer households and higher demand when the eligibility criterion is more lenient, relative to ignoring interactions. To illustrate, consider a policymaker debating whether to expand eligibility for the subsidy from 40% to 60% of the population. In the presence of conforming effects, the increase in eligibility will spur more non-eligible to adopt, such that the total demand with spillovers are larger than without spillovers. Conversely, if eligibility was cut from 40% to 20%, the drop in adoption would be magnified by conforming effects, such that the total demand with spillovers are lower than without spillovers.\footnote{The intuition can be understood via a simple example. Suppose the true regression model is $y=\beta_{0}+\beta_{1}x+u$, where $\beta_{1}>0$. Suppose, to predict $y$ at a value $x_{0}$ of $x$, we ignore the covariate and simply use $\bar {y}$ as the prediction. If $x_{0}<\bar{x}$, then the naive prediction $\bar {y}=$ $\beta_{0}+\beta_{1}\bar{x}$ will be larger than the true value $\beta_{0}+\beta_{1}x_{0}$, whereas if $x_{0}>\bar{x}$, then the naive prediction $\bar{y}$ will be smaller than the true value $\beta_{0}+\beta _{1}x_{0}$.} As for welfare, allowing for social interactions may lead to a welfare loss for ineligible households, in turn implying higher deadweight loss from the subsidy scheme, relative to estimates obtained ignoring social spillovers where welfare effects for ineligibles are zero by definition. The resulting net welfare effect, aggregated over both eligibles and ineligibles, admits a large range of possible values including both positive and negative ones, with associated large variation in the implied deadweight loss estimates, all of which are consistent with the same coefficient on the social interaction term in the choice probability function.

An implication of these results for applied work is that welfare analysis under spillovers effects requires knowledge of the different channels of spillovers separately, possibly via conducting a `belief elicitation' survey where subjects are asked the reasons for their actions; knowledge of only the choice probability functions, inclusive of a social interaction term, is insufficient.

\noindentPlan of the Paper: The rest of the paper is organized as follows. Section (ref) describes the set-up, Section (ref) develops the tools for empirical welfare analysis of a price intervention in such models, and associated deadweight loss calculations. Section (ref) specifies the stochastic environment and derives the convergence of equilibrium beliefs under I.I.D. unobservables. Section (ref) establishes consistency of our estimator. Section (ref), describes the context of our empirical application and the data; Section (ref) describes the empirical results; Section (ref) summarizes and concludes. Technical derivations, formal proofs and additional results are collected in an Appendix.

SET-UP

Consider a population of villages indexed by $v\in\left\{ 1,\dots,\bar {v}\right\} $ and resident households in village $v$ indexed by $\left( v,h\right) $, with $h\in\left\{ 1,\dots,N_{v}\right\} $. For the purpose of inference discussed later, we will think of these households as a random sample drawn from an infinite superpopulation. The total number of households we observe is $N=\sum_{v=1}^{\bar{v}}N_{v}$. Each household faces a binary choice between buying one unit of an indivisible good (alternative $1$) or not buying it (alternative $0$). Its utilities from the two choices are given by $U_{1}(Y_{vh}-P_{vh},\Pi_{vh},\boldsymbol{\eta}_{vh})$ and $U_{0}(Y_{vh} ,\Pi_{vh},\boldsymbol{\eta}_{vh})$ where the variables $Y_{vh}$, $P_{vh}$, and $\boldsymbol{\eta}_{vh}$ denote respectively the income, price, and heterogeneity of household $(v,h)$, and $\Pi_{vh}$ is household $\left( v,h\right) $'s subjective belief of what fraction of households in her village would choose to buy. The variable $\boldsymbol{\eta}_{vh}$ is privately observed by household $\left( v,h\right) $ but is unobserved by the econometrician and other households. The dependence of utilities on $\Pi_{vh}$ captures social interactions. Below, we will specify how $\Pi_{vh}$ is formed. Household $(v,h)$'s choice is described by

equation[equation omitted — 205 chars of source]

where $1\left\{ \cdot\right\} $ denotes the indicator function. In the mosquito-net example of our application, one can interpret $U_{1}$ and $U_{0}$ as expected utilities resulting from differential probabilities of contracting malaria from using and not using the net, respectively.

The utilities, $U_{1}$ and $U_{0}$, may also depend on other covariates of $(v,h)$. For notational simplicity, we will occasionally write $W_{vh} =(Y_{vh},P_{vh})$\footnote{All vectors are defined as row vectors.}, and suppress other covariates for now; additional covariates are used in our empirical implementation in Section (ref).

WELFARE\ ANALYSIS

We now lay out the empirical framework for welfare analysis of policy interventions under spillovers. We will assume spillovers are restricted to the village where households reside, hence welfare effects of a policy intervention can be analyzed village by village; so for economy of notation, we drop the $\left( v,h\right) $ subscripts except when we account explicitly for village-effects during estimation. Also, we use the same notation $\pi$ to denote both individual beliefs $\Pi_{vh}$ entering individual utilities, and the unique equilibrium belief about village take-up rate entering the average demand function. The assumption of a constant (within village) $\pi$ is justified via Proposition 1 and Proposition 2 in Section (ref).

In the welfare results derived below, all probabilities and expectations -- e.g., mean welfare loss -- are calculated with respect to the marginal distribution of aggregate unobservables, denoted by $\boldsymbol{\eta}$. In this sense, they are analogous to `average structural functions' (ASF), introduced by Blundell and Powell (2004). Later, when discussing estimation of the ASF, together with the implied pre- and post-intervention aggregate choice probabilities and average welfare in Section (ref), we will allude to village-effects explicitly, and show how they are estimated and incorporated in demand and welfare predictions.

Define $q_{1}\left( p,y,\pi\right) $ to be the structural probability (i.e. average structural function; ASF) of a household choosing option $1$ (e.g., buying mosquito-net) when it faces a price of $p$, has income $y$ and belief $\pi$:

equation[equation omitted — 228 chars of source]

where $F_{\boldsymbol{\eta}}$ is the cumulative distribution function (CDF) of $\boldsymbol{\eta}$. This probability can be estimated via the conditional probability of purchase given covariates when household level unobservables are uncorrelated with the covariates, as will be assumed in our application. The reason for focusing on the ASF, rather than the purchase probability conditional on covariates is that ultimately, we will be interested in the marginal distribution of welfare in a village resulting from a potential price-intervention, e.g. a means-tested subsidy, which is counterfactual.\footnote{Expressing our results in terms of ASFs help clarify that in general, the object of interest is one, whose identification and consistent estimation may require non-experimental methods if the data at hand are observational. Also, later in the paper, we will allude to village level unobservables i.e. $\eta_{vh}=\xi_{v}+u_{vh}$, where $\xi_{v}$ is a village specific unobservable variable (introduced in Section (ref)). In that context, the object of interest will be the marginal distribution of welfare in each village; thus the relevant distribution to used to compute the ASF will be that of $u_{vh}$ given village specific variables.}

\noindentLinear Index Structure: We now specify the forms of the utility functions. Given a moderate/small number of large peer groups (e.g., there are eleven large villages in our application dataset), it is not easy to consistently estimate the impact of the belief $\Pi_{vh}$ on the choice probability function nonparametrically holding other regressors constant.\footnote{This is because $\Pi_{vh}$ is a constant within a village (as discussed in Section (ref)). In particular, the fixed point constraint, which is a notable feature of the social interaction model, does not help because of dimensionality problems. Indeed, in the fixed point condition: $\pi=\int q_{1}\left( p,y,\pi\right) dF_{P,Y}\left( p,y\right) $, where the joint CDF\ $F_{P,Y}\left( p,y\right) $ of $\left( P,Y\right) $ is identified, the unknown function $q_{1}\left( p,y,\pi\right) $ has more arguments than the identified $F_{P,Y}\left( p,y\right) $.} Accordingly, following Manski (1993), and Brock and Durlauf (2001, 2007), we assume a linear index structure with $\boldsymbol{\eta}=(\eta^{0},\eta^{1})$ viz. the utilities are given by

equation[equation omitted — 279 chars of source]

where we assume that $\beta_{0}>0$, $\beta_{1}>0$, i.e., non-satiation in numeraire, and $\beta_{1}$ need not equal $\beta_{0}$, i.e. income effects can be present.\footnote{We can also allow for concave income effects by specifying, say,

align*[align* omitted — 223 chars of source]

but we wish to keep the utility formulation as simple as possible to highlight the complications in welfare calculations even in the simplest linear utility specification.}

In our empirical setting of anti-malarial bednet (Insecticide-Treated Net; ITN, henceforth) adoption, there are multiple potential sources of interactions (i.e. $\alpha_{1},\alpha_{0}\neq0$). The first is a pure preference for conforming; the second is increased awareness of the benefits of a bednet when more villagers use it; the third is the perceived health externality. The medical literature suggests that the technological health externality is positive, i.e. as more people are protected, the lower is the malaria burden, but the perceived health externality can be negative if households believe that other households' bednet use deflects mosquitoes to unprotected households, but ignore the fact that those deflected mosquitoes are less likely to carry the parasite. Indeed, the implications for adoption are different: under the positive health externality, one would expect free-riding, hence a negative effect of others' adoption on own adoption; under the negative health externality, the correlation would be positive.

In particular, let $\gamma_{p}$ denote the conforming plus learning effect, and $\gamma_{H}$ denote the health externality. Then it is reasonable to assume that $\alpha_{1}\equiv\gamma_{p}\geq0$, while $\alpha_{0}\equiv \gamma_{H}-\gamma_{p}$ could be either negative or positive. It is natural that the conforming/learning/peer effect $\gamma_{p}$ affects utilities from buying and non-buying symmetrically, i.e. if $\Pi$ changes from $0$ to $\pi$ the resulting change in the utility (relative to when $\Pi$ was $0$) from buying and the one from not buying are symmetric and of opposite sign, as is also assumed in BD01, BD07. Further, if a household uses an ITN, then there is no health externality from the neighborhood adoption rate since the household is protected anyway,\footnote{We can allow for a smaller health externality, say $\gamma_{h}<\gamma_{H}$ when one adopts the bednet. But this does not change the fundamental point about the asymmetric effect of $\pi$ on the utility from buying and from not buying. So we avoided adding this to save on notation.} but if it does not adopt, then there is a net health externality effect $\gamma_{H}$ from neighborhood use, which makes the overall effect $\alpha_{0}=\gamma_{H}-\gamma_{p}$ and in general, there is no exact relationship between $\alpha_{0}=$ $\gamma_{H}-\gamma_{p}$ \ and $\alpha _{1}=\gamma_{p}$.\footnote{An analogous asymmetry is also likely in the school voucher example mentioned in the introduction if the voucher-led `brain-drain' leads to utility gains and losses of different amounts, e.g., if better teaching resources in the high-achieving school substitute for -- or complement -- peer effects in a way that is not possible in the resource-poor local school.} Accordingly, we first assume that the perceived net health externality is non-positive, and thus $\alpha_{1}\geq 0\geq\alpha_{0}$, and derive welfare results. In the next subsection, we present the results under the case $\alpha_{1}\geq\alpha_{0}\geq0$. Note that the sign of $\alpha=\alpha_{1}-\alpha_{0}$ is identified, and is positive in our data, which rules out $\alpha_{0}>\alpha_{1}\geq0$. In the application, we present the bounds separately for $\alpha_{1}\geq0\geq\alpha_{0}$ and $\alpha_{1}\geq\alpha_{0}\geq0$, and then the union of these.

Given the linear index specification, the structural choice probability of buying at $\left( p,y,\pi\right) $ is given by

equation[equation omitted — 275 chars of source]

where $F\left( \cdot\right) $ denotes the marginal distribution function of $\eta^{0}-\eta^{1}$. It is known from Brock and Durlauf (2007) that the structural choice probabilities $F\left( c_{0}+c_{1}p+c_{2}y+\alpha \pi\right) $ identify $c_{0},c_{1},c_{2}$ and $\alpha$, i.e. $\left( \delta_{1}-\delta_{0}\right) $, $\beta_{0}$, $\beta_{1}$ and $\left( \alpha_{1}-\alpha_{0}\right) =2\gamma_{p}-\gamma_{H}$, up to scale even without knowledge of the probability distribution of $\eta^{0}-\eta^{1}$. In the application, we will consider two different estimates of the choice probabilities, first ignoring village-specific unobservables and using a standard probit, and then allowing for village-fixed unobservables using a variant of correlated random effects.

The distinct presence of $\alpha_{1},\alpha_{0}$ makes the model different from standard demand models for binary choice. In the standard case, for the so-called \textquotedblleft outside option\textquotedblright, i.e. not buying, the utility is normalized to zero. In a social spillovers setting, this cannot be done because that utility depends on the aggregate purchase rate $\pi$. As we will see below, in welfare evaluations of a subsidy, $\alpha_{1}$ and $\alpha_{0}$ appear separately in the expressions for welfare-distributions, but cannot be separately identified from demand data, which can only identify $\alpha\equiv\alpha_{1}-\alpha_{0}$. As a result, point-identification of welfare will in general not be possible. Below, we will consider some untestable special cases, under which one obtains point-identification, e.g., with $\alpha_{1}\geq0\geq\alpha_{0}$, the interesting special cases are (i) $\alpha_{1}=\alpha/2=-\alpha_{0}$ (i.e. $\gamma_{H}=0$: no health externality and symmetric spillover), which is considered in BD01 for social welfare analysis, (ii) $\alpha_{1}=\alpha,$ $\alpha_{0}=0$ (i.e. $\gamma _{H}=\gamma_{p}$: technological health externality dominates deflection channel and net health externality exactly offsets conforming effect), and (iii) $\alpha_{1}=0$, $\alpha_{0}=-\alpha$ ($\gamma_{p}=0$ and $\gamma _{H}=-\alpha$: no conforming effect and deflection channel dominates). Cases (ii) and (iii) will yield respectively the upper and lower bounds on welfare gain for the case $\alpha_{1}\geq0\geq\alpha_{0}$. Analogously for the case $\alpha_{1}\geq\alpha_{0}\geq0$.

\noindentPolicy Intervention and Welfare Expressions: We start with a situation where the price of the product is $p_{0}$ and the value of $\pi$ is $\pi_{0}$. Now suppose a price subsidy is introduced such that individuals with income less than a threshold $\tau$ become eligible to buy the product at price $p_{1}<p_{0}$. This policy will alter the equilibrium adoption rate; suppose the new equilibrium adoption rate changes to $\pi_{1}$, where $\pi _{0}$ and $\pi_{1}$, solve the fixed point conditions:

align[align omitted — 370 chars of source]

where $F_{Y}$ is the CDF of $Y_{vh}$, and $F$ is defined in ((ref)). Since the price coefficient $c_{1}<0$ and $p_{1}>p_{0}$, therefore if $\alpha>0$, we have that

align*[align* omitted — 252 chars of source]

for each $\pi$, and therefore, the integrand of ((ref)) is smaller for every $\pi$ than the integrand of ((ref)). If the solutions to ((ref)) and ((ref)) are unique, then the value of $\pi$ at which ((ref)) holds must be smaller than the value of $\pi$ where ((ref)) holds. So we shall get $\pi_{1}>\pi_{0}$. This is borne out in our application where sufficient conditions on $\alpha$ for a contraction are satisfied. Under multiple solutions, we can at least say that if $p_{1}<p_{0}$, the smallest solution $\pi_{1}$\ to ((ref))\ is greater than the smallest solution $\pi_{0}$\ to ((ref)).

For given values of $\pi_{0}$ and $\pi_{1}$, we now derive expressions for welfare resulting from the intervention. By \textquotedblleft welfare\textquotedblright\ we mean the compensating variation (CV), viz. what hypothetical income compensation would restore the post-change indirect utility for an individual to its pre-change level. For a subsidy-eligible individual, for any potential value of $\pi_{1}$ corresponding to the new equilibrium, the individual compensating variation is the solution $S$ to the equation

equation[equation omitted — 294 chars of source]

whereas for a subsidy-ineligible individual, it is the solution $S$ to

equation[equation omitted — 294 chars of source]

Thus we interpret the CV as measuring utility changes via the value of hypothetical income compensation that would restore utilities to their initial level.\footnote{Note that we do not take account of peer effects of this hypothetical income compensation, which might be an alternative way to define the CV.} Now, since $S$ depends on the unobservables $\boldsymbol{\eta}$, the same price change will produce a distribution of welfare effects across individuals; we are interested in calculating that distribution and its functionals such as mean welfare.

The welfare effect of the subsidy can be calculated as described below.

Welfare Calculation under $\alpha_{1}\geq0>\alpha_{0}$

Recall that $\alpha_{1}=\gamma_{p}\geq0$, $\alpha_{0}=\gamma_{H}-\gamma_{p}$; thus $\alpha_{1}\geq0>\alpha_{0}$ corresponds to the case where either $\gamma_{H}<0$, i.e., deflection effect dominates positive health effect in perception, or is positive but smaller than conforming/learning effect.

\noindentWelfare for Eligibles ($\alpha_{1}\geq0>\alpha_{0} $): The CV for a subsidy-eligible household is given by the solution $S$ to

align[align omitted — 352 chars of source]

The resulting solution $S$ depends on the unobservable heterogeneity $\eta ^{0}$ and $\eta^{1}$ and hence we are interested in deriving its distribution and functionals thereof such as mean welfare. Calculating the welfare distribution requires us to compute the CDF of $S$, i.e. $\Pr\left( S\leq a\right) $ for various values of $a$ (for given $(p_{0},\pi_{0},p_{1},\pi _{1})$). Let $f_{\eta^{0}-\eta^{1}}\left( \cdot\right) $ denote the marginal density function of $\eta^{0}-\eta^{1}$. Then the expression for the CDF of welfare is as follows:

theoremSuppose the linear index structure described above holds with $\beta_{0}>0$, $\beta_{1}>0$ and $\alpha_{1}\geq0\geq\alpha_{0}$ with $\alpha=\alpha_{1}-\alpha_{0}$ satisfying $\left\vert \alpha\right\vert \sup_{e\in\mathbb{R}}f_{\eta^{0}-\eta^{1}}\left( e\right) <1$. Then $\pi _{1}>\pi_{0}$, and the distribution of compensating variation for the eligibles, $S=S^{\mathrm{Elig}}$, is given by \begin{align} & \Pr\left( S^{\mathrm{Elig}}\leq a\right) \nonumber\\ & =\left\{ \begin{array} [c]{ll} 0, & if a<p_{1}-p_{0}-\frac{\alpha_{1}}{\beta_{1}}\left( \pi_{1}-\pi_{0}\right) ,\\ q_{1}\left( p_{1}-a,y,\pi_{0}+\frac{\alpha_{1}}{\alpha}\left( \pi_{1} -\pi_{0}\right) \right) , & if p_{1}-p_{0}-\frac{\alpha_{1} }{\beta_{1}}\left( \pi_{1}-\pi_{0}\right) \leq a<\frac{\alpha-\alpha_{1} }{\beta_{0}}\left( \pi_{1}-\pi_{0}\right) ,\\ 1\text{,} & \text{if }a\geq\frac{\alpha-\alpha_{1}}{\beta_{0}}\left( \pi _{1}-\pi_{0}\right) \text{.} \end{array} \right. \end{align}

The proof is provided in Appendix (ref). The condition $\left\vert \alpha\right\vert \sup_{e\in\mathbb{R}}f_{\eta^{0}-\eta^{1} }\left( e\right) <1$ essentially says that the social interaction parameter is not too large in magnitude, so that ((ref)) and ((ref)) have unique solutions in $\pi_{0}$ and $\pi_{1}$ respectively, whence by the argument following ((ref)) and ((ref)), we have that $\pi_{1}>\pi_{0}$.

Now, note that in the intermediate case in ((ref)), where $a\in\lbrack p_{1}-p_{0}-\frac{\alpha_{1}}{\beta_{1}}\left( \pi_{1}-\pi_{0}\right) ,$ $\frac{\alpha_{0}}{\beta_{0}}\left( \pi_{0}-\pi_{1}\right) ]$, $\Pr\left( S\leq a\right) $ equals

equation[equation omitted — 235 chars of source]

In ((ref)), the intercept $c_{0}=\delta_{1}-\delta_{0}$, the slopes $c_{1}=-\beta_{1},c_{2}=\beta_{1}-\beta_{0}$ and $\alpha=\alpha_{1}-\alpha _{0}$ are all identified from conditional choice probabilities; however $\alpha_{1}$ is not identified and therefore ((ref)) is not point-identified from the structural choice probabilities. However, since $\alpha_{1}\in\left[ 0,\alpha\right] $, for each feasible value of $\alpha_{1}\in\left[ 0,\alpha\right] $, we can compute a corresponding value of ((ref)), giving us bounds on the welfare distribution.

Note also that the thresholds of $a$ at which the CDF expression changes are also not point-identified for the same reason. However, since $\pi_{1}-\pi _{0}>0$ and $\beta_{0}>0$, $\beta_{1}>0$, the interval \[ p_{1}-p_{0}-\frac{\alpha_{1}}{\beta_{1}}\left( \pi_{1}-\pi_{0}\right) \leq a<\frac{\alpha_{0}}{\beta_{0}}\left( \pi_{0}-\pi_{1}\right) \] will translate to the left as $\alpha_{1}$ varies from $0$ to $\alpha $.

remarkNote that the above theorem continues to hold even if the subsidy is universal; we have not used the means-tested nature of the subsidy to derive the result.
corollary[Mean Welfare]From ((ref)), the mean welfare for the eligible is given by \begin{align} E[S^{\mathrm{Elig}}] & =\underset{welfare gain}{\underbrace{-\int _{p_{1}-p_{0}-\frac{\alpha_{1}}{\beta_{1}}\left( \pi_{1}-\pi_{0}\right) }^{0}q_{1}\left( p_{1}-a,y,\pi_{0}+\frac{\alpha_{1}}{\alpha}\left( \pi _{1}-\pi_{0}\right) \right) da}}\nonumber\\ & +\underset{welfare loss}{\underbrace{\int_{0}^{\frac{\alpha -\alpha_{1}}{\beta_{0}}\left( \pi_{1}-\pi_{0}\right) }\left[ 1-q_{1}\left( p_{1}-a,y,\pi_{0}+\frac{\alpha_{1}}{\alpha}\left( \pi_{1}-\pi_{0}\right) \right) \right] da}}, \end{align} where the following formula for a random variable $X$ that has finite mean and the CDF, $F_{X}$, is used: $E\left[ X\right] =\int_{0}^{\infty} [1-F_{X}\left( x\right) ]dx-\int_{-\infty}^{0}F_{X}\left( x\right) dx$.

The width of the bounds on ((ref)) and ((ref)), obtained by varying $\alpha_{1}$ over $\left[ 0,\alpha\right] $, depends on the extent to which $q_{1}\left( \cdot,\cdot,\pi\right) $ is affected by $\pi$, i.e. the extent of social spillover, and also the difference in the realized values $\pi_{1}$ and $\pi_{0}$. For our single-index model, the fixed point restrictions imply that these counterfactual $\pi_{1}$ and $\pi_{0}$ depend on $\alpha_{1}$ and $\alpha_{0}$ only via $\alpha=\alpha_{1}-\alpha_{0}$ (cf. ((ref)) and ((ref)) above) which is point-identified; thus every potential value of counterfactual demand is point-identified. But given any feasible value of $\pi_{1}$ and $\pi_{0}$, the welfare ((ref)) is not point-identified in general, since $\alpha_{1}$ is unknown.

However, given $\alpha$, the welfare gain in expression ((ref)) is increasing in $\alpha_{1}$; i.e., the welfare gain is largest in absolute value when $\alpha_{1}=\alpha$ and $\alpha_{0}=0$, and the smallest when $\alpha_{1}=0$ and $\alpha_{0}=-\alpha$; conversely for welfare loss. Intuitively, if there is no negative externality from increased $\pi$ on non-purchasers, then they do not suffer any welfare loss, but purchasers have a welfare gain from both lower price and higher $\pi$. Conversely, if all the spillovers are negative, then purchasers still receive a welfare gain via price reduction, but non-purchasers suffer welfare loss due to increased $\pi $. Also, note that under quasilinear utilities (i.e., utilities with $\beta_{0}=\beta_{1}$), where income effects are absent, the $y$ drops out of the above expressions, but the same identification problem remains, since $\alpha_{1}$ does not disappear. Changing variables $p=p_{1}-a$, one can rewrite ((ref)) as

align[align omitted — 541 chars of source]

Note that if $\alpha_{1}=0$, then the first term is the usual consumer surplus capturing the effect of price reduction on consumer welfare; for a positive $\alpha_{1}$, the term $\frac{\alpha_{1}}{\beta_{1}}\left( \pi_{1}-\pi _{0}\right) $ yields the additional effect arising via the conforming channel. Also, if $\alpha_{1}=0$, then the second term, i.e. the welfare loss from not buying, is the largest (given $\alpha$): this corresponds to the case where all of $\alpha$ is due to the negative externality.

The second term in ((ref)), which represents welfare change caused solely via spillovers and no price change, is still expressed as an integral with respect to price. This is a consequence of the index structure which enables us to express this welfare loss in terms of foregone utility from an equivalent price change.

\noindentSpecial Cases: For the special case of symmetric interactions considered in BD01, where their social welfare is calculated with $\alpha_{1}=-\alpha_{0}$ in ((ref)) (e.g., if $\gamma_{H}=0$, i.e. there is no health externality in the health-good example), we have $\dfrac{\alpha_{1} }{\alpha}=\dfrac{-\alpha_{0}}{-2\alpha_{0}}=\dfrac{1}{2}$, and from ((ref)) mean welfare equals:

equation[equation omitted — 446 chars of source]

If $\alpha_{0}=0$, and $\alpha=\alpha_{1}$, i.e. all spillovers are via conforming, mean welfare is given by

equation[equation omitted — 193 chars of source]

and if, on the other hand, any spillovers are due to perceived health risk, i.e. $\alpha=-\alpha_{0}$ and $\alpha_{1}=0$, then mean welfare is given by

equation[equation omitted — 314 chars of source]

Expressions ((ref)) and ((ref)) correspond to the upper and lower bounds, respectively, of the overall welfare gain for eligibles.\footnote{Gautam (2018) obtained point-identified estimates of welfare in parametric discrete choice models with social interactions, purportedly using Dagsvik and Karlstrom's (2005) results for the setting without spillover. Gautam's paper contains no explicit expression for average welfare, but we conjecture that her derivation had implicitly assumed one of the normalizations ((ref)), ((ref)) or ((ref)) under which average welfare is point-identified.}

\noindentWelfare for Ineligibles ($\alpha_{1}\geq0>\alpha_{0} $): Welfare change for ineligibles is measured by the CV defined as the solution $S$ to the equation:

align[align omitted — 365 chars of source]

which is simply ((ref)) with $p_{1}$ replaced by $p_{0}$. Therefore, the mean CV is simply ((ref)) with $p_{1}$ replaced by $p_{0}$.

corollarySuppose the linear index structure described above holds with $\beta_{0}>0$, $\beta_{1}>0$, and $\alpha_{1}\geq0\geq\alpha_{0}$. Then for each $\alpha_{1}\in\left[ 0,\alpha\right] $, the mean welfare for the ineligible, $S=S^{\mathrm{Inelig}}$, is given by \begin{align} E[S^{\mathrm{Inelig}}] & =-\int_{p_{0}}^{p_{0}+\frac{\alpha_{1}}{\beta_{1} }\left( \pi_{1}-\pi_{0}\right) }q_{1}\left( p,y,\pi_{0}+\frac{\alpha_{1} }{\alpha}\left( \pi_{1}-\pi_{0}\right) \right) dp\nonumber\\ & +\int_{p_{0}+\frac{\alpha_{1}-\alpha}{\beta_{0}}\left( \pi_{1}-\pi _{0}\right) }^{p_{0}}\left[ 1-q_{1}\left( p,y,\pi_{0}+\frac{\alpha_{1} }{\alpha}\left( \pi_{1}-\pi_{0}\right) \right) \right] dp. \end{align}

For ineligibles, all of the welfare effects come from spillovers, since they experience no price change. In particular, for ineligibles who buy, there is a welfare gain from positive spillovers due to a higher $\pi$. For ineligibles who do not buy, there is, however, a potential welfare loss due to increased $\pi$. This is why the CV distribution has the support that includes both positive and negative values. The first term in ((ref)) captures the welfare gain resulting from a positive $\alpha_{1}$ and higher $\pi$; this term would be zero if $\alpha_{1}=0$. The second term in ((ref)) captures the welfare loss also resulting from higher $\pi$; this loss would be zero if there are no negative impacts, i.e. $\alpha_{0}=0$. Of course, both would be zero if $\alpha =0=\alpha_{1}=\alpha_{0}$, reflecting the fact that welfare effect on ineligibles would be zero if there is no spillover.

In the three special cases where we have point-identification, viz. (i) $\alpha_{1}=-\alpha_{0}=\frac{\alpha}{2}$; (ii) $\alpha=\alpha_{1}$, $\alpha_{0}=0$; and (iii) $\alpha=-\alpha_{0}$, $\alpha_{1}=0$, mean CV ((ref)) reduces respectively to:

align[align omitted — 863 chars of source]

Expressions ((ref)) and ((ref)) correspond to the upper and lower bounds, respectively, of the overall welfare gain for ineligibles, and therefore, the overall bounds generically contain both positive and negative values, since $\alpha\neq0$.

\noindentDeadweight Loss ($\alpha_{1}\geq0>\alpha_{0}$): The mean deadweight loss (DWL) can be calculated as the expected subsidy spending less the net welfare gain:

align*[align* omitted — 1,310 chars of source]

The bounds on $\alpha_{1}$ then translate into bounds for mean DWL. In particular, if $\alpha_{0}=0$ (so that $\alpha=\alpha_{1}$), then

align*[align* omitted — 429 chars of source]

Therefore, if $\frac{\alpha}{\beta_{1}}\left( \pi_{1}-\pi_{0}\right) $ is sufficiently large, then the mean DWL will be negative, i.e. the subsidy will increase economic efficiency under positive spillover, as in the standard textbook case. This happens because there is no subsidy expenditure on ineligibles, and yet those ineligibles who buy enjoy a subsidy-induced welfare gain due to positive spillover. Subsidy-eligibles receive an additional welfare gain via positive spillover, over and above the welfare-gain due to reduced price, and it is only the latter that is financed by the subsidy expenditure. In general, the deadweight loss will be lower (more negative) when (i) the positive spillovers ($\alpha_{1}$) is larger, (ii) the change in equilibrium adoption ($\pi_{1}-\pi_{0}$) due to the subsidy is greater, and (iii) the price elasticity of demand ($-\beta_{1}$) is lower -- the last effect lowers deadweight loss simply by reducing the substitution effect, even in absence of spillover.

Mean Welfare under $\alpha_{1}\geq\alpha_{0}\geq0$

Recall that $\alpha_{1}=\gamma_{p}\geq0$, $\alpha_{0}=\gamma_{H}-\gamma_{p}$; in our application, it holds that $\alpha=\alpha_{1}-\alpha_{0}>0$; thus $\alpha_{1}>\alpha_{0}\geq0$ corresponds to the case where $\gamma_{H}>0$ i.e. insecticide effect dominates deflection effect, is also larger than conforming/learning but less than twice the conforming effect. Note that under this assumption, we must also have $\alpha\leq\alpha_{1}$.

\noindentWelfare for Eligibles ($\alpha_{1}\geq\alpha_{0}\geq 0$): For subsidy-eligibles, the mean welfare (for given $(p_{0} ,\pi_{0},p_{1},\pi_{1})$) is presented in the following theorem:

theoremSuppose the linear index structure described above holds with $\beta_{1} \geq\beta_{0}>0$, and $\alpha_{1}\geq\alpha_{0}\geq0$. Let $\beta=\beta _{1}-\beta_{0}$ and $\alpha=\alpha_{1}-\alpha_{0}$, which are estimable from the choice probability function, and define \begin{align*} C_{1}\left( \alpha_{1}\right) & :=- {\int_{p_{1}-\frac{\alpha-\alpha_{1}}{\beta_{0}}\left( \pi _{1}-\pi_{0}\right) }^{p_{0}+\frac{\alpha_{1}}{\beta_{1}}\left( \pi_{1} -\pi_{0}\right) }} q_{1}\left( p,y,\pi_{0}+\frac{\alpha_{1}}{\alpha}\left( \pi_{1}-\pi _{0}\right) \right) dp,\\ C_{2}\left( \alpha_{1}\right) & :=- {\int_{p_{0}+\frac{\alpha-\alpha_{1}}{\beta_{0}}\left( \pi _{1}-\pi_{0}\right) }^{p_{1}-\frac{\alpha_{1}}{\beta_{1}}\left( \pi_{1} -\pi_{0}\right) }} \left[ 1-q_{1}\left( p,y+p-p_{0},\pi_{1}-\frac{\alpha_{1}}{\alpha}\left( \pi_{1}-\pi_{0}\right) \right) \right] dp. \end{align*} Then mean welfare for the eligible is given by \[ E\left[ S^{\mathrm{Elig}}\right] =\left\{ \begin{array} [c]{ll} C_{1}\left( \alpha_{1}\right) \text{,} & \text{if }\alpha\leq\alpha_{1} \leq\frac{\beta_{1}}{\beta}\left( \beta_{1}-\beta\right) \frac{p_{0}-p_{1} }{\pi_{1}-\pi_{0}}+\frac{\beta_{1}}{\beta}\alpha\text{,}\\ C_{2}\left( \alpha_{1}\right) \text{,} & \text{if }\alpha_{1}>\frac {\beta_{1}}{\beta}\left( \beta_{1}-\beta\right) \frac{p_{0}-p_{1}}{\pi _{1}-\pi_{0}}+\frac{\beta_{1}}{\beta}\alpha\text{.} \end{array} \right. \]

The proof is provided in Appendix (ref). Given that $\alpha_{1}$ is unknown, this result implies the lower and upper bounds of the mean welfare:

align[align omitted — 1,001 chars of source]

Thus, allowing for both $\alpha_{1}\geq\alpha_{0}\geq0$ and $\alpha_{1} \geq0\geq\alpha_{0}$ yields the wider bounds on mean welfare for the eligible to:

align[align omitted — 369 chars of source]

where $LB_{\alpha_{1}\geq0\geq\alpha_{0}}^{\mathrm{Elig}}$ and $UB_{\alpha _{1}\geq0\geq\alpha_{0}}^{\mathrm{Elig}}$ are defined as expressions ((ref)) and ((ref)), respectively. Since we expect $\beta_{1}>\beta_{0}$ (also borne out by the empirical results), $C_{2}\left( \alpha_{1}\right) $ will tend to $-\infty$ as $\alpha_{1},\alpha_{0}\rightarrow\infty$. Therefore, as $\alpha_{1},\alpha_{0}\rightarrow\infty$, the integrand in $C_{2}\left( \alpha_{1}\right) $ will tend to 1 and $LB_{\alpha_{1}\geq\alpha_{0}\geq 0}^{\mathrm{Elig}}$ in ((ref)) will tend to $-\infty$ whereas the $UB_{\alpha_{1}\geq\alpha_{0}\geq0}^{\mathrm{Elig}}$ in ((ref)) will remain bounded. Therefore, the lower bound on welfare gain will be finite and the upper bound infinite under $\alpha_{1}\geq\alpha_{0}\geq0$.

\noindentWelfare for Ineligibles ($\alpha_{1}\geq\alpha_{0}\geq 0$): For subsidy-ineligibles, the mean welfare is obtained simply by replacing $p_{1}$ by $p_{0}$ in the expressions for eligibles:

corollarySuppose the linear index structure described above holds with $\beta_{1} \geq\beta_{0}>0$, and $\alpha_{1}\geq\alpha_{0}\geq0$. Define \begin{align*} D_{1}\left( a_{1}\right) & :=-\int_{p_{0}+\frac{\alpha_{1}-\alpha} {\beta_{0}}\left( \pi_{1}-\pi_{0}\right) }^{p_{0}+\frac{\alpha_{1}} {\beta_{1}}\left( \pi_{1}-\pi_{0}\right) }q_{1}\left( p,y,\pi_{0} +\frac{\alpha_{1}}{\alpha}\left( \pi_{1}-\pi_{0}\right) \right) dp,\\ D_{2}\left( \alpha_{1}\right) & :=-\int_{p_{0}+\frac{\alpha-\alpha_{1} }{\beta_{0}}\left( \pi_{1}-\pi_{0}\right) }^{p_{0}-\frac{\alpha_{1}} {\beta_{1}}\left( \pi_{1}-\pi_{0}\right) }\left[ 1-q_{1}\left( p,y+p-p_{0},\pi_{1}-\frac{\alpha_{1}}{\alpha}\left( \pi_{1}-\pi_{0}\right) \right) \right] dp. \end{align*} Then, the mean welfare for the ineligible is given by \[ E\left[ S^{\mathrm{Inelig}}\right] =\left\{ \begin{array} [c]{cc} D_{1}\left( \alpha_{1}\right) \text{, } & \text{if }\alpha\leq\alpha_{1} \leq\frac{\beta_{1}}{\beta_{1}-\beta_{0}}\alpha\text{,}\\ D_{2}\left( \alpha_{1}\right) \text{,} & \text{if }\frac{\beta_{1}} {\beta_{1}-\beta_{0}}\alpha<\alpha_{1}<\infty\text{.} \end{array} \right. \]

From the above results, it follows that allowing for $\alpha_{1}>\alpha _{0}\geq0$ in addition to the possibility $\alpha_{1}\geq\alpha_{0}\geq0$ widens the overall bounds for mean welfare of the ineligible from ((ref)) and ((ref)) to

align[align omitted — 836 chars of source]

The deadweight loss expressions are analogous to those for the case with $\alpha_{1}\geq0\geq\alpha_{0}$ and not repeated here.

STOCHASTIC ENVIRONMENT AND\ EQUILIBRIUM BELIEFS

\noindentIncomplete-Information Setting: In this section, we formulate interactions of households as an incomplete-information Bayesian game, whose stochastic structure will be laid out below. In each village $v$, each of the $N_{v}$ households is provided the opportunity to buy the product at a researcher-specified price $P_{vh}$ randomly varied across households. They have incomplete information in that each household $(v,h)$ knows her own variables $(A_{vh},W_{vh},\boldsymbol{\eta}_{vh})$ but does not know the values of all the variables $W_{\tilde{v}k},\boldsymbol{\eta}_{\tilde{v} k},A_{\tilde{v}k}$ for every household $k\neq h$ selected in the experiment.

We assume households have `consistent beliefs' in accordance with the standard Bayes-Nash setting, i.e., each $(v,h)$'s belief is formed as

equation[equation omitted — 163 chars of source]

where $A_{vk}$ is given in ((ref)) and $E\left[ \cdot\text{ }|\mathcal{I}_{vh}\right] $ is the conditional expectation computed through the probability law that governs all the relevant variables given $(v,h)$'s information set $\mathcal{I}_{vh}$ that includes $(W_{vh},\boldsymbol{\eta}_{vh})$. The explicit form of ((ref)) in equilibrium is investigated in the next subsection.

Each household $(v,h)$ is solely concerned with behavior of other households in the same village $v$. Thus the econometrician observes $\bar{v}$ games ($\bar{v}=11$ in our application), each with `many' households. To formalize our model as a Bayesian game, given the form of ((ref)), $U_{1}$ and $U_{0}$ are to be interpreted as expected utilities. This is possible when the underlying von Neumann-Morgenstern utility indices $u_{1}$ and $u_{0}$ satisfy \[ U_{1}\left( Y_{vh}-P_{vh},\Pi_{vh},\boldsymbol{\eta}_{vh}\right) =E[u_{1}(Y_{vh}-P_{vh},\tfrac{1}{N_{v}-1} {\displaystyle\sum\nolimits_{1\leq k\leq N_{v}\text{; }k\neq h}} A_{vk},\boldsymbol{\eta}_{vh})|\mathcal{I}_{vh}]\text{,} \] i.e., $u_{1}$ is linear in the second argument; $U_{0}$ and $u_{0}$ satisfy an analogous relationship. This will hold in particular when utilities have a linear index structure as in Manski (1993) and Brock and Durlauf (2001, 2007). We have already presented our linear specifications of $U_{1}$ and $U_{0}$ in ((ref)), but these are further elaborated below in this section and in Section (ref).

\noindentUnobserved Heterogeneity: We assume that unobserved heterogeneity $\{\boldsymbol{\eta}_{vh}\}_{v=1}^{N_{v}}$ ($v=1,\dots\bar{v}$) takes the following form:

equation[equation omitted — 112 chars of source]

where $\boldsymbol{\xi}_{v}$ stands for a village-specific vector of variables that are common to all members in the $v$th village and $\boldsymbol{u}_{vh}$ represents an individual specific variable. Let $\left( d_{v},e_{v}\right) $ be an underlying vector of village-specific variables such that $d_{v}$ is a common factor affecting both the unobservable $\boldsymbol{\xi}_{v}$ and the observable covariates $W_{vh}$, and $e_{v}$ affects only $\boldsymbol{\xi} _{v}$ with $\boldsymbol{\xi}_{v}$ fully determined by $\left( d_{v} ,e_{v}\right) $, i.e., $\boldsymbol{\xi}_{v}=\boldsymbol{\xi}\left( d_{v},e_{v}\right) $.\footnote{The need to separate $d_{v}$ and $e_{v}$ will become clear below in the context of identification of model parameters in presence of unobserved group-effects.} Each household in village $v$ is assumed to know $\left( d_{v},e_{v}\right) $, the functional form $\boldsymbol{\xi}\left( \cdot\right) $, and thus $\boldsymbol{\xi}_{v}$, while $\boldsymbol{u}_{vh}$ is a purely private variable known only to individual $(v,h)$. None of $\left\{ \left( d_{v},e_{v}\right) \right\} $, $\left\{ \boldsymbol{\xi}_{v}\right\} $, and $\left\{ \boldsymbol{u} _{vh}\right\} $ is observable to the econometrician. Denote household $(v,h)$'s information set by

equation[equation omitted — 102 chars of source]

We now impose the following conditions on the probabilistic law for these variables:

description$\{(W_{vh},\boldsymbol{u}_{vh},d_{v},e_{v})\}_{h=1}^{N_{v}}$, $v=1,\dots,\bar{v}$, are independent across $v$.

Assumption C1 says that variables in village $v$ are independent of those in village $\tilde{v}(\neq v)$.

description• (i)\ For each $v$, the sequence $\left\{ (W_{vh},\boldsymbol{u} _{vh})\right\} _{h=1}^{N_{v}}$ is I.I.D. conditionally on $\left( d_{v},e_{v}\right) $. (ii) $\left\{ \boldsymbol{u}_{vh}\right\} _{h=1}^{N_{v}}$ is independent of $\left\{ W_{vh}\right\} _{h=1}^{N_{v}}$ conditionally on $\left( d_{v},e_{v}\right) $.

The conditional I.I.D.-ness imposed in C2 (i) leads to equi-dependence within each village, i.e., $\mathrm{Cov}\left[ \boldsymbol{\eta}_{vh},\boldsymbol{\eta}_{vk}\right] =\mathrm{Cov}\left[ \boldsymbol{\eta}_{v\tilde{h}},\boldsymbol{\eta}_{v\tilde{k}}\right] (\neq0)$ for any $h\neq k$ and $\tilde{h}\neq\tilde{k}$. Further, each household $(v,h)$'s unobservable $\boldsymbol{u}_{vh}$ is not useful for predicting another household $(v,k)$'s variables and behavior, and therefore her belief $\Pi_{vh}$ ((ref)) is reduced to the average of the unconditional expectations (as formally shown in Proposition (ref)) below. This condition rules out spatial correlation in unobservables which, if present, would complicate the analysis in a non-trivial way by making a household's belief a function of its privately known variables.

C2 (ii) is the exogeneity condition. This allows for identification and consistent estimation of model parameters. In the context of the field experiment in our empirical exercise, this exogeneity condition can be interpreted as saying that realization of unobserved heterogeneity is independent of how researchers have selected the sample. Note that the exogeneity condition is conditional on $\left( d_{v},e_{v}\right) $, and it does not exclude correlation of $\boldsymbol{u}_{vh}$ and $W_{vh}=\left( P_{vh},Y_{vh}\right) $ in the unconditional sense. In our application, prices $P_{vh}$ are randomly assigned to individuals by researchers and thus $P_{vh}$ and $\boldsymbol{u}_{vh}$ are independent both unconditionally and conditionally.

Note that under the ((ref)) introduced later, we compute the ASF ((ref)) using the marginal distribution of $\boldsymbol{u}_{vh}$ conditionally on $\boldsymbol{\xi}_{v}$ in later sections (see also Footnote 3).

Equilibrium Beliefs

We now investigate the forms of households' beliefs defined in ((ref)). We show that under C2, the high-level assumption in BD01 that beliefs, corresponding to our $\Pi_{vh}$, are constant and symmetric across all households in the same village can be formalized in our incomplete-information game setting via the specification of a Bayes-Nash equilibrium.

propositionSuppose that Conditions C1 and C2 are common knowledge in the Bayesian game described above. Then, for any $k\neq h$ in village $v$ with $\left( d_{v},e_{v}\right) $, \[ E[A_{vk}|\mathcal{I}_{vh}]=E[A_{vk}|d_{v},e_{v}], \] where the information set $\mathcal{I}_{vh}$ is defined in ((ref)).

The proof of Proposition (ref) is provided in Appendix (ref). Note that this proposition does not utilize any equilibrium condition. It simply confirms, formally, the intuitive statement that $(v,h)$'s own variables are not useful to predict other $(v,k)$'s behavior $A_{vk}$. Given this result, we can write the belief $\Pi_{vh}$ (defined in ((ref))) as

equation[equation omitted — 74 chars of source]

where \[ \bar{\Pi}_{vh}=\bar{\Pi}_{vh}(d_{v},e_{v}):=\tfrac{1}{N_{v}-1} {\displaystyle\sum\nolimits_{1\leq k\leq N_{v}\text{; }k\neq h}} E[A_{vk}|d_{v},e_{v}], \] and $\bar{\Pi}_{vh}$ is a function of $\left( d_{v},e_{v}\right) $ and independent of $(v,h)$-specific variables, $(W_{vh},\boldsymbol{u}_{vh})$; for notational simplicity, we suppress the dependence of $\bar{\Pi}_{vh}$ on $\left( d_{v},e_{v}\right) $ from now on.

Beliefs in equilibrium solve the system of $N_{v}$ equations:

equation[equation omitted — 360 chars of source]

where $E_{v}\left[ \cdot\right] $ denotes the conditional expectation operator given $\left( d_{v},e_{v}\right) $ (i.e., $E\left[ \cdot |d_{v},e_{v}\right] $). BD01 focus on equilibria with constant and symmetric beliefs. Using our notation above, we say that (constant) beliefs are symmetric when $\bar{\Pi}_{vh}=\bar{\Pi}_{vk}$ for any $h,k\in\{1,\dots ,N_{v}\}$ (for each $v$). When Brock and Durlauf's framework is interpreted as a Bayesian game, one can justify their focus on constant and symmetric beliefs under conditions laid out in Proposition (ref) below.

To establish this proposition, define for each $v$, given $\left( d_{v} ,e_{v}\right) $, a function $m_{v}:\left[ 0,1\right] \rightarrow \lbrack0,1]$ as

equation[equation omitted — 248 chars of source]

note that $m_{v}\left( r\right) $ is independent of individual index $h$ under the conditional I.I.D. assumption given $\left( d_{v},e_{v}\right) $. Then the following characterization of beliefs holds:

propositionSuppose that the same conditions hold as in Proposition (ref) and the function $m_{v}\left( \cdot\right) $ defined in ((ref)) is a contraction, i.e., for some $\rho\in\left( 0,1\right) $, \begin{equation} |m_{v}\left( r\right) -m_{v}\left( \tilde{r}\right) |\leq\rho|r-\tilde {r}| \ for any r,\tilde{r}\in\left[ 0,1\right] . \end{equation} Then, a solution $(\bar{\Pi}_{v1},\dots,\bar{\Pi}_{vN_{v}})$ of the system of $N_{v}$ equations in ((ref)) uniquely exists and is given by symmetric beliefs, i.e., \[ \bar{\Pi}_{vh}=\bar{\Pi}_{vk}\text{ \ for any }h,k\in\{1,\dots,N_{v}\}. \]

The proof is given in Appendix (ref). Propositions (ref)-(ref) show that, given the (conditional) I.I.D. and contraction conditions, the equilibrium is characterized through \[ \Pi_{vh}=\bar{\pi}_{v}\text{ \ for any }h=1,\dots,N_{v}, \] for some constant $\bar{\pi}_{v}:=\bar{\pi}_{v}(d_{v},e_{v})\in\left[ 0,1\right] $ within each village (given $\left( d_{v},e_{v}\right) $). This implies that the beliefs can be consistently estimated by the sample average of $A_{vk}$ over village $v$, which is exploited in our empirical study.

The contraction condition ((ref)) holds when the social interactions coefficient $\alpha$ is not large (in our linear index specification). In Section (ref) below, we will provide sufficient conditions for the contraction and equilibrium uniqueness, as well as explain additional procedures that are needed for estimation and counterfactual analysis when multiplicity of equilibria may arise.

ECONOMETRIC SPECIFICATION, IDENTIFICATION AND ESTIMATION

Taking the belief variable $\Pi_{vh}$ in the linear-index choice probability function $q_{1}\left( \cdot\right) $ to be the (limit of) observed fraction of usage in each village i.e. $\bar{\pi}_{v}=E_{v}\left[ A_{vh}\right] (=\lim_{N_{v}\rightarrow\infty}\sum_{h=1}^{N_{v}}A_{vh}/N_{v})$ as justified in Propositions (ref)-(ref), the index coefficients can be estimated semiparametrically using say, Bhattacharya (2008). However, unobserved village-effects may confound the consistency of these estimates; we overcome this by using a correlated random effects (CRE, henceforth) probit approach to estimate $q_{1}\left( \cdot\right) $, which is derived from a factor structure on the covariates and the village-effects, as follows.

Village Effects Specification

Our data for the application come from eleven different villages with an average of $195$ households per village. It is plausible that utilities from using and from not using an ITN\ are affected by village-specific unobservable characteristics (i.e., $\xi_{v}=\xi_{v}^{1}-\xi_{v}^{0}$ introduced in ((ref))), such as the chance of contracting malaria when not using an ITN. Recall the linear utility structure ((ref)) from Section (ref). Given this, together with the unobserved heterogeneity specification in ((ref)), $\boldsymbol{\eta}$ $_{vh}=\boldsymbol{\xi}_{vh}+\boldsymbol{u}_{vh}$, we model

equation[equation omitted — 444 chars of source]

where $\boldsymbol{\xi}_{v}=\left( \xi_{v}^{0},\xi_{v}^{1}\right) $ and $\boldsymbol{u}_{vh}=\left( u_{vh}^{0},u_{vh}^{1}\right) $ denote village and individual specific characteristics, respectively, both of which are unobservable. Therefore,

align[align omitted — 569 chars of source]

where $\varepsilon_{vh}$ is assumed to have zero mean and unit variance for scale and location normalization.

\noindentNon-identification of the village effects $\xi_{v} $: Brock and Durlauf (2007) discussed difficulties of estimating social interactions models in presence of group-specific unobservables and presented a non-identification result (their Proposition 2). To see this in our context, consider constant beliefs, $\Pi_{vh}=\bar{\pi}_{v}$ (justified in Propositions (ref)-(ref)). Since $\xi_{v}$ is village specific and many observations per village are available, we can estimate village specific intercepts $\gamma_{v}$ by regression of take-up $A_{vh}$ on price and income $W_{vh}=\left( P_{vh},Y_{vh}\right) $ that vary across households $h$ within village $v$, together with village dummies, i.e.,

equation[equation omitted — 203 chars of source]

where the left-hand side (LHS) is computed under the conditional law given $\left( d_{v},e_{v}\right) $, and $F_{\varepsilon}\left( \cdot\right) $ is the CDF of $-\varepsilon_{vh}$.\footnote{Recall that $\varepsilon_{vh}\left( =u_{vh}^{1}-u_{vh}^{0}\right) $ is assumed to be independent of $W_{vh}$ conditionally on $\left( d_{v},e_{v}\right) $ in C2; and $\xi_{v}$ is determined by $\left( d_{v},e_{v}\right) $. Below, it is further assumed that $\varepsilon_{vh}$ is jointly independent of $W_{vh}$ and $\left( d_{v},e_{v}\right) $.}

The realized $\xi_{v}$ is a constant within each village; thus, $\xi_{1} ,\dots,\xi_{\bar{v}}$ and the universal constant $c_{0}$ cannot be separately identified and we reparametrize $\bar{\xi}_{v}:=c_{0}+\xi_{v}$. For each $v$ and each realized $\xi_{v}$, the LHS of ((ref)) is identifiable as a function of $w$; thus, under a parametric specification of $F_{\varepsilon }\left( \cdot\right) $ together with the exogeneity condition C2 (ii) and a rank condition for covariates (stated below), $\left( \boldsymbol{c},\gamma_{1},\dots,r_{\bar{v}}\right) $ is also identified. The identified coefficients $\gamma_{1},\dots,\gamma_{\bar{v}}$ on the village dummies therefore satisfy the equations $\gamma_{v}=\alpha\pi_{v}+\bar{\xi }_{v}\equiv c_{0}+\xi_{v}$ ($v=1,\dots,\bar{v}$).\footnote{In the application, we a run a probit of $A_{vh}$ on covariates $W_{vh}$ ($P_{vh}$, $Y_{vh}$, and other variables) and a dummy for each village which corresponds to converting these conditional moments to a set of unconditional ones.} However, even in the reparametrized equations, there are as many $\bar{\xi}_{v}$ as there are $\gamma_{v}$, so that we have $\bar{v}$ equations with $\bar{v}+1$ unknowns $\bar{\xi}_{1},\dots,\bar{\xi}_{\bar{v}}$, and $\alpha$, which are needed for policy and counterfactual analysis but cannot be separately identified.

Factor Structure and Correlated Random Effects Modelling

We surmount non-identification of $\xi_{v}$ by an approximate version of the Mundlak-Chamberlain correlated random effects (CRE) structure, cf. Section 15.8.2 of Wooldridge (2010), which is routinely used as a reasonable middle ground between fixed and random effects in the panel econometrics literature. While the CRE device is typically intended for short panels, our setting here may be seen like a \textquotedblleft long panel\textquotedblright\ in that each village is supposed to have its own effect that is shared by a large number of households (note that our dataset does not have a panel structure but consists of several cross sectional datasets). To have our specification consistent with the long-panel-like\ setting and Section (ref)\ (in particular, C2), we consider the following factor structure for the observable covariate $W_{vh}$ and the village specific variable $\xi_{v}$,

equation[equation omitted — 129 chars of source]

where $d_{v}$ is a vector of \textquotedblleft factor\textquotedblright \ variables (with the same dimension as $W_{vh}$) that are common in $W_{vh}$ and $\xi_{v}$, $\tau_{vh}$ is the covariate specific, idiosyncratic component that is assumed to have zero mean (for location normalization) and is defined through $\tau_{vh}:=W_{vh}-d_{v}$, $\boldsymbol{\delta}$ is a (row) vector of constant coefficients on the factor, and $e_{v}$ is a village specific variable that affects only $\xi_{v}$.\footnote{Note that one component $W_{vh}$ is the price $P_{vh}$ faced by the household, which is randomized across households. The corresponding component of $d_{v}$ in ((ref)) is the average price within the village and its coefficient in $\boldsymbol{\delta}$\ is set as zero in our application (as the randomized price does not capture village specific features).}

We assume each household in village $v$ knows the realization of $\left( d_{v},e_{v}\right) $, while all the right-hand-side components in ((ref)) are unobservable to researchers. Let $\bar{W}_{v}:=\left( 1/N_{v}\right) \sum_{v=1}^{N_{v}}W_{vh}$. Then, we can write $d_{v}=\bar {W}_{v}-\left( 1/N_{v}\right) \sum\nolimits_{v=1}^{N_{v}}\tau_{vh}$. Plugging this into the second equation in ((ref)), we can write

equation[equation omitted — 118 chars of source]

for each $N_{v}$, which follows from $\left( 1/N_{v}\right) \sum \nolimits_{v=1}^{N_{v}}\tau_{vh}=O_{p}(1/\sqrt{N_{v}})=o_{p}(1)$ by a standard central limit theorem. We note that ((ref)) is a reduced-form representation for each (sufficiently large) $N_{v}$ derived from the structural assumption ((ref)). We further assume that the error term satisfies

equation[equation omitted — 193 chars of source]

for each $v$, where we note that $\left\{ e_{v}\right\} _{v=1}^{\bar{v}}$ is I.I.D. under C1 and ((ref)) (we denote by $\sigma_{e}^{\ast}$ the true standard deviation parameter; and subsequently, $\ast$ is often used to denote true parameters).\footnote{Brock and Durlauf (2007) have also considered a (linear) restriction on the group effects similar to ((ref)) (see their Section 4.1.2 and Assumption L.1) and argue that it may help partial identification.} In standard short-panel cases, a distributional assumption is directly imposed on group effects, say, $\xi_{v}|\left\{ W_{vh}\right\} _{h=1}^{N_{v}}\sim N(\bar {W}_{v}\boldsymbol{\delta}^{\prime},(\sigma_{e}^{\ast})^{2})$ (see Wooldridge, 2010, p. 615); in our setting, this conditional normality of $\xi_{v}$ holds in an approximate sense with a small order $o_{p}\left( 1\right) $ term in ((ref)).\footnote{The original CRE model, the so-called Mundlak-Chamberlain device, is not derived from a factor structure as in ((ref)); we do not know of any other paper that considers a CRE model as a reduced form derived from some factor structure, which can be thought of as a separate contribution of the present paper. Our derivation of the approximate CRE model makes the households' information structure transparent which is required for constructing an econometric framework consistent with the game structure and C2. If we directly imposed ((ref)) as is done in standard short panel contexts, it would be difficult to see which parts of $\xi_{v}$ should be known to households and to interpret the conditional i.i.d.-ness in C2 given the village specific variables.} We further assume that

equation[equation omitted — 152 chars of source]

which is analogous to specifications in Chamberlain (1980) and Wooldridge (2010). Putting all of this together, we can write \[ A_{vh}=1\{W_{vh}\boldsymbol{c}^{\prime}+c_{0}+\alpha\bar{\pi}_{v} +d_{v}\boldsymbol{\delta}^{\prime}+e_{v}+\varepsilon_{vh}\geq0\}\text{ \ for each }(v,h) \] and compute the conditional probability as

align[align omitted — 544 chars of source]

where the probability on the LHS is computed under the conditional law given $d_{v}$ (i.e., it is with respect to the distribution of $-\left( \varepsilon_{vh}+e_{v}\right) $); $F_{\varepsilon+e}$ is the CDF of $-(\varepsilon_{vh}+e_{v})\sim N(0,1+\sigma_{e}^{2})$, and last equality holds since $d_{v}=\bar{W}_{v}+o_{p}\left( 1\right) $; and ((ref)) can be shown to hold uniformly over $\left( v,h\right) $, $w$, and $\left( \boldsymbol{c},c_{0},\alpha,\boldsymbol{\delta}\right) $, under compactness of the parameter space and Condition CR2 (ii), imposed in the next subsection. Denote by $\Phi$ the CDF of $N\left( 0,1\right) $. Then, the leading term on the right-hand side (RHS) of ((ref)) can be written as

equation[equation omitted — 323 chars of source]

For calculating the LHS of ((ref)), $e_{v}$ is treated as a part of the parameter $\gamma_{v}$; in contrast, the LHS of ((ref)) is calculated with respect to the distribution of the unobservable $\varepsilon_{vh}+e_{v}$ across all households over all villages. Both the probabilities in ((ref)) and ((ref)) concern the same outcome variable $A_{vh}$ but they differ in conditioning variables. The former probability can be consistently estimated within each village as $N_{v}\rightarrow\infty$ for each $v$, while consistent estimation \ of the latter requires $\bar {v}\rightarrow\infty$ (in addition to $N_{v}\rightarrow\infty$) since village specific effects $e_{v}$ have to be averaged out to match the probability ((ref)) computed as the integral of $e_{v}$ via its approximation ((ref)).

Putting all this together, our estimation steps are as follows:

enumerate• First run a probit of $A_{vh}$ on $W_{vh},$ $\bar{W}_{v}$ and $\hat{\pi }_{v}\equiv\left( 1/N_{v}\right) \sum_{v=1}^{N_{v}}A_{vh}$ corresponding to ((ref)) to obtain estimates of $\boldsymbol{\bar{c}},\bar{c} _{0},\bar{\alpha},\boldsymbol{\bar{\delta}}$ ; • Then run a probit of $A_{vh}$ on $W_{vh}$ and village dummies corresponding to ((ref)) and obtain estimates $\boldsymbol{c}^{\prime },\gamma_{1},\gamma_{2},...\gamma_{\bar{v}}$; • Estimate $\sigma_{e}^{\ast}$ by the ratio of the price coefficient in the former to that in the latter probit; • Estimate $c_{0}$ via $\bar{c}_{0}\times\sqrt{1+(\sigma_{e}^{\ast})^{2}}$ and $\alpha$ via $\bar{\alpha}\times\sqrt{1+(\sigma_{e}^{\ast})^{2}}$; • From ((ref)), estimate $\xi_{v}=\gamma_{v}-$ $c_{0}-\alpha \hat{\pi}_{v}$.

These are all the quantities we need for empirical calculation of welfare expressions outlined in Section 3. In the empirical application below, the parameters $\boldsymbol{\bar{c}},\bar{c}_{0},\bar{\alpha},\boldsymbol{\bar {\delta}}$ are estimated via pseudo-MLE by running an ordinary probit regression of $A_{vh}$ on $W_{vh},\bar{\pi}_{v}$ and $\bar{W}_{v}$.

Thus to summarize, it follows from Brock and Durlauf's (2007) arguments, outlined above, that identification of village specific parameters, $\xi_{v}$, is in general impossible in the presence of social interaction effects. We overcome this through our CRE condition ((ref)) which imposes more structure on $\xi_{v}$ and letting the number of groups, i.e. $\bar {v}\rightarrow\infty$, as formally stated in the next subsection and the proof of consistency in Appendix (ref). As such, this is a new finding for social-interactions models. Note that if $e_{v}$ is non-stochastic (i.e., $\sigma_{e}^{\ast}=0$ and $\xi_{v}=\bar{W} _{v}\boldsymbol{\delta}^{\prime}+o_{p}\left( 1\right) $, instead of ((ref))), the above scheme using two probit regressions leads to identification and consistent estimation without the many-village assumption of $\bar{v}\rightarrow\infty$.\footnote{This identification/estimation scheme of CRE models using two probit regressions appears new, which allows us to recover the standard deviation $\sigma_{e}^{\ast}$ (which is not typically identified in standard short-panel cases; see e.g. p. 617 of Wooldridge, 2010) and further all the realized values of $e_{1},\dots,e_{\bar{v}}$.}

Estimation and Consistency

Now we discuss consistency of the estimation procedure outlined in the previous subsection. We focus on the consistency of the first probit ((ref)), the setting of which is non-standard under the CRE structure and the many-village asymptotics $\bar{v}\rightarrow\infty$; in contrast, the setting of the second probit ((ref)) or ((ref)) can be analyzed in the same way as in Hahn and Kuersteiner (2011), and a detailed discussion of its consistency is omitted.\footnote{Our second probit setting is even simpler than Hahn and Kuersteiner's in that the number of parameters do not increase $N_{v}\rightarrow\infty$ or $\bar{v}\rightarrow\infty$. A notable difference is that the objective function $\hat{R}$ incurs some approximation error $o_{p}\left( 1\right) $ by using ((ref)) instead of ((ref)); but given the uniformity of the $o_{p}\left( 1\right) $ as stated, this error can be negligible for the consistency discussion.}

For verification of consistency, we assume that the number of households in each village can be written as

equation[equation omitted — 59 chars of source]

where $r_{v}\in\left( \underline{r},\bar{r}\right) $ is a constant that is independent of $N_{0}$ and $\bar{v}$ with $0<\underline{r}\leq\bar{r}<\infty$ (i.e., $r_{v}$ is uniformly bounded from below and above), and let $N=\sum_{v=1}^{\bar{v}}N_{v}$ is the total number of households in all villages combined. This assumption means that all $N_{1},\dots,N_{\bar{v}}$ grow at the same rate, so that none of villages is asymptotically negligible.

Comparing the two probabilities in ((ref)), we can identify/estimate all the parameters $\gamma_{v}^{\ast}$, $\left( \boldsymbol{c}^{\ast},c_{0}^{\ast},\alpha^{\ast},\boldsymbol{\delta}^{\ast }\right) $, and $\sigma_{e}^{\ast}$, which allows us to obtain estimates of $\xi_{1},\dots\xi_{\bar{v}}$. Consistent estimation of these parameters can be achieved through the following two probit regressions.\footnote{Note that our practical estimation procedure exploits the fixed point restriction by an iteration process (discussed in Section (ref)). It is slightly more complicated than the procedure outlined here; but the substance of our identification arguments does not change between the two procedures; our exposition here is based on the simpler procedure.} First, a probit of $A_{vh}$ \ on $W_{vh}$ and village dummies allows us to obtain estimates \[ \left( \boldsymbol{\hat{c}},\hat{\gamma}_{1},\dots.,\hat{\gamma}_{\bar{v} }\right) =\underset{\boldsymbol{c}\in \Upsilon_{1};\text{ }\left( \gamma _{1},\dots,\gamma_{\bar{v}}\right) \in \Upsilon\left( \bar{v}\right) \times\dots\times \Upsilon\left( \bar{v}\right) }{\operatorname{argmax}} \hat{Q}\left( \boldsymbol{c},\gamma_{1},\dots.,\gamma_{\bar{v}}\right) , \] where the objective function $\hat{Q}$ is

equation[equation omitted — 267 chars of source]
equation[equation omitted — 276 chars of source]

$\Upsilon_{1}$ is a compact set in $\mathbb{R}^{d_{W}}$, and $\Upsilon\left( \bar{v}\right) $ is a compact interval on $\mathbb{R}$ that may grow as $\bar{v}\rightarrow\infty$ (specified in ((ref)) in Appendix (ref)). Second, via a second probit of $A_{vh}$ on $(W_{vh},1,\hat{\pi}_{v},\bar{W}_{v})$, we can estimate the coefficients $(\boldsymbol{\bar{c}}^{\ast},\bar{c}_{0}^{\ast},\bar{\alpha}^{\ast },\boldsymbol{\bar{\delta}}^{\ast})$ through \[ (\widehat{\boldsymbol{\bar{c}}},\widehat{\bar{c}}_{0},\widehat{\bar{\alpha} },\widehat{\boldsymbol{\bar{\delta}}})=\underset{(\boldsymbol{\bar{c}},\bar {c}_{0},\bar{\alpha},\boldsymbol{\bar{\delta}})\in \Upsilon_{2} }{\operatorname{argmax}}\hat{R}(\boldsymbol{\bar{c}},\bar{c}_{0},\bar{\alpha },\boldsymbol{\bar{\delta}}), \] where the objective function $\hat{R}$ is defined as

align*[align* omitted — 507 chars of source]

and $\Upsilon_{2}$ is a compact set in $\mathbb{R}^{2d_{W}+1}$.\footnote{By the results in Section (ref), we have $\bar{\pi}_{v}=E\left[ A_{vh}|d_{v},e_{v}\right] $ in the equilibrium, which can be consistently estimated by an average within each village, $\hat{\pi}_{v}=\frac{1}{N_{v} }\sum_{h=1}^{N_{v}}A_{vh}$.} Then, we can recover an estimate of $\hat{\sigma }_{e}^{2}$ through a ratio of the first (or any other) components of $\boldsymbol{\hat{c}}$ and $\widehat{\boldsymbol{\bar{c}}}$, which yields $\sqrt{1+\hat{\sigma}_{e}^{2}}$ and further $(\hat{c}_{0},\hat{\alpha },\boldsymbol{\hat{\delta}})=(\widehat{\bar{c}}_{0},\widehat{\bar{\alpha} },\widehat{\boldsymbol{\bar{\delta}}})\times\sqrt{1+\hat{\sigma}_{e}^{2}}$. Finally, given these estimates, we can compute \[ \hat{\xi}_{v}=\hat{\gamma}_{v}-\hat{c}_{0}-\hat{\alpha}\hat{\pi}_{v} \] for each $v$, which then allows us to calculates the welfare estimates of Section (ref).

The proof of consistency for the first probit is involved due to the CRE structure and the formal steps are provided in Appendix (ref). The key substantive assumption delivering consistency is as follows:

description• (i) For each $v$, let $\lambda_{v}^{\min}$ be the minimum of the eigenvalues of $E_{\boldsymbol{\omega}_{v}}[(W_{vh},1)^{\prime}(W_{vh},1)]$, which is a square (real symmetric) matrix of order $d_{W}+1$, where $d_{W}$ is the dimension of $W_{vh}$. Then, $\inf_{v\geq1}$ $\lambda_{v}^{\min}>0$. (ii) The covariates and unobservables $(W_{vh},\xi_{v},\varepsilon_{vh})\ $satisfy ((ref)), ((ref)), and ((ref)). (iii) Let $\bar{W}_{v}^{\ast}:=\operatorname*{plim}\limits_{N_{v} \rightarrow\infty}\frac{1}{N_{v}}\sum_{h=1}^{N_{v}}W_{vh} (=E_{\boldsymbol{\omega}_{v}}\left[ W_{vh}\right] )$ for each $v$ (the existence of $\bar{W}_{v}^{\ast}$ is supposed), and \[ \mathbb{\bar{W}}:=\left[ \begin{array} [c]{ccc} 1 & \bar{\pi}_{1} & \bar{W}_{1}^{\ast}\\ \vdots & & \vdots\\ 1 & \bar{\pi}_{\bar{v}} & \bar{W}_{\bar{v}}^{\ast} \end{array} \right] , \] which is a $\bar{v}\times\left( 2+d_{W}\right) $ matrix. Then, suppose that $\mathbb{\bar{W}}$ is of rank $2+d_{W}$.

Condition (i) of CR1 allows us to identify $(\boldsymbol{c}^{\ast },\gamma_{v}^{\ast})$ for each $v$ as the maximizer of $Q_{v}\left( \boldsymbol{c},\gamma_{v}\right) =E_{\boldsymbol{\omega}_{v}}\left[ \mathcal{L}_{vh}\left( \boldsymbol{c},\gamma_{v}\right) \right] $ whose empirical analogue $\hat{Q}_{v}$ defined in ((ref)) is a constituent of the objective $\hat{Q}$ defined in ((ref)); see the expression ((ref)) in Appendix (ref). The condition on the uniform lower bound of $\lambda_{v}^{\min}$ together with ((ref)) guarantees the identification/consistency of all $\gamma_{v}^{\ast}$ when $\bar{v}\rightarrow\infty$. (iii) of CR1 is used to verify identification of the parameters in the second probit, $(\boldsymbol{\bar{c}}^{\ast},\bar{c}_{0}^{\ast},\bar{\alpha}^{\ast },\boldsymbol{\bar{\delta}}^{\ast})$, where we note that by the definition in ((ref)) and the reduced form expression of $\xi_{v}$ in ((ref)), we can write the true village-specific effect $\gamma _{v}^{\ast}=c_{0}^{\ast}+\alpha^{\ast}\bar{\pi}_{v}+\bar{W}_{v}^{\ast }\boldsymbol{\delta}^{\prime}+e_{v}$.

The formal consistency statement is expressed in the following proposition:

propositionSuppose that Conditions C1, C2, CR1 (i)-(ii), and the technical condition CR2 (stated in Appendix (ref)), Specifications ((ref)) and ((ref)) hold and that $\bar {v}$ and $N_{0}$ satisfy Assumption ((ref)) with \begin{equation} \left. \bar{v}^{4(\sigma_{e}^{\ast})^{2}}\left( \log\bar{v}\right) ^{3+4(\sigma_{e}^{\ast})^{2}}\left( \log N_{0}\right) \right/ N_{0}\rightarrow0 \ (as N_{0}\rightarrow\infty). \end{equation} Then, as $N_{0}\rightarrow\infty$ and $\bar{v}\rightarrow\infty$, \[ \left\Vert \boldsymbol{\hat{c}}-\boldsymbol{c}^{\ast}\right\Vert \overset{p}{\rightarrow}0\text{ \ and \ }\max_{v\in\left\{ 1,\dots,\bar {v}\right\} }\left\vert \hat{\gamma}_{v}-\gamma_{v}^{\ast}\right\vert \overset{p}{\rightarrow}0. \]

The proof is provided in Appendix (ref).

Verification of this proposition for the first probit is not trivial. This is because (I) given the asymptotic assumption $\bar{v}\rightarrow\infty$, required for the consistency in the second probit, the number of parameters tends to infinity; and (II) each parameter $\gamma_{v}^{\ast}=c_{0}^{\ast }+\alpha^{\ast}\bar{\pi}_{v}+\bar{W}_{v}^{\ast}\boldsymbol{\delta}^{\ast \prime}+e_{v}$ includes a realization of $e_{v}\sim N\left( 0,\sigma_{e} ^{2}\right) $ and thus the maximum of realized $\left\vert \gamma_{1}^{\ast }\right\vert ,\dots,\left\vert \gamma_{\bar{v}}^{\ast}\right\vert $ grows with positive probability as $\bar{v}\rightarrow\infty$ since $e_{v}$ has unbounded support $(-\infty,\infty)$.

Several previous papers on panel models have considered a setting like (I), such as Hahn and Kuersteiner (2011) and Fern\'{a}ndez-Val and Weidner (2016). However, in these papers, the \textquotedblleft growing magnitude of parameters\textquotedblright\ as (II) is not allowed for, i.e., typically, all parameters are supposed to be in a fixed compact set.\footnote{This is explicitly assumed in Hahn and Kuersteiner's Condition 4, while cases like (II) have to be typically excluded by Fern\'{a}ndez-Val and Weidner's Assumption 4.1 (v) (the presence of uniform bounds $b_{\min}$ and $b_{\max}$ for the derivative of their objective function).}

These problems, in particular (II), make it hard to establish the identification of $\left( \boldsymbol{c}^{\ast},\gamma_{v}^{\ast}\right) $. However, we can overcome this by showing that given the I.I.D. $\left\{ e_{v}\right\} _{v\geq1}$, the maximum of $\left\vert e_{1}\right\vert ,\dots,\left\vert e_{\bar{v}}\right\vert $ is bounded by $\sqrt{4(\sigma _{e}^{\ast})^{2}\log[\bar{v}\left( \log\bar{v}\right) ^{t}]}$ almost surely for any $t>1/2$ (Lemma (ref)). This result allows us to restrict possible support of each $\gamma_{v}^{\ast}$ as a compact set that grows (as $\bar{v}\rightarrow\infty$); and within this support, if $||\left( \boldsymbol{c},\gamma_{v}\right) -\left( \boldsymbol{c}^{\ast},\gamma _{v}^{\ast}\right) ||>\epsilon_{1}$, we can always find some constant $C_{Q}>0$ such that $Q_{v}\left( \boldsymbol{c}^{\ast},\gamma_{v}^{\ast }\right) -Q_{v}\left( \boldsymbol{c},\gamma_{v}\right) \geq C_{Q}[\bar {v}(\log\bar{v})]^{-2(\sigma_{e}^{\ast})^{2}}$ almost surely (shown in Lemma (ref)), which means the identification of $\left( \boldsymbol{c}^{\ast},\gamma_{v}^{\ast}\right) $ as the unique maximizer of $Q_{v}$. If the support were not restricted, the lower bound of the difference between $Q_{v}\left( \boldsymbol{c}^{\ast},\gamma_{v}^{\ast}\right) $ and $Q_{v}\left( \boldsymbol{c},\gamma_{v}\right) $ would be zero as $\bar {v}\rightarrow\infty$, which would complicate identification and consistency.\footnote{To see this, let $||\boldsymbol{c}-\boldsymbol{c}^{\ast }||>\epsilon_{1}$ and $\gamma_{v}=\gamma_{v}^{\ast}$, for example. Then, for any $\boldsymbol{c}\neq\boldsymbol{c}^{\ast}$, $[Q_{v}\left( \boldsymbol{c} ^{\ast},\gamma_{v}^{\ast}\right) -Q_{v}\left( \boldsymbol{c},\gamma _{v}^{\ast}\right) ]\rightarrow0$ as $\left\vert \gamma_{v}^{\ast}\right\vert \rightarrow\infty$, which holds for some $v$ as $\bar{v}\rightarrow\infty$ since $\gamma_{v}^{\ast}=c_{0}^{\ast}+\alpha^{\ast}\bar{\pi}_{v}+\bar{W} _{v}^{\ast}\boldsymbol{\delta}^{\ast}{}^{\prime}+e_{v}$ includes the normally distributed variable $e_{v}$. Note that for large $\left\vert \gamma_{v} ^{\ast}\right\vert $, both $\Phi\left( W_{vh}(\boldsymbol{c}^{\ast})^{\prime }+\gamma_{v}^{\ast}\right) $ and $\Phi\left( W_{vh}\boldsymbol{c}^{\prime }+\gamma_{v}^{\ast}\right) $ (the normal CDF's) are very close to $1$ or $0$ regardless of $\boldsymbol{c}\neq\boldsymbol{c}^{\ast}$ (i.e., variation of $\Phi\left( \cdot\right) $ is tiny in the tail region); thus, the difference between $Q_{v}\left( \boldsymbol{c}^{\ast},\gamma_{v}^{\ast}\right) $ and $Q_{v}\left( \boldsymbol{c},\gamma_{v}^{\ast}\right) $ is very small, which are computed through these CDF's.}

The rate condition ((ref)) requires $\bar{v}$ to grow slower than $N_{0}$ in particular when the variance of $e_{v}$ is large, which is reasonable in the context of our empirical application, where $\bar{v}=11$ may be regarded as small relative to $N_{v}=r_{v}N_{0}$ which is $195$ on average.\footnote{Note that regardless of the rate condition ((ref)), for the first probit, the magnitude of $\bar{v}$ does not directly affect estimation precision of $\left( \boldsymbol{\hat{c}},\hat{\gamma}_{v}\right) $ (up to first order), whose convergence rate is $1/\sqrt{N_{v}}$. In contrast, the rate condition matters for the second probit, for which the integration with respect to $e_{v}$ has to be approximated by the sum over $e_{1},\dots,e_{\bar{v}}$.} It guarantees that the difference between $Q_{v}\left( \boldsymbol{c}^{\ast},\gamma_{v}^{\ast}\right) $ and $Q_{v}\left( \boldsymbol{c},\gamma_{v}\right) $ is larger than that between $\hat{Q}_{v}$ and $Q_{v}$, justifying the maximizer of $\hat{Q}_{v}$ as an estimator, implying the consistency.

To see how our two-step probit performs in finite samples, we implemented a small Monte Carlo exercise that is reported in Appendix (ref) \ and shows reassuring results for magnitudes of sample size resembling ours.

Calculation of Predicted Demand and Welfare

In order to calculate our welfare-related quantities, we need to estimate the structural choice probabilities $q_{1}\left( p,y,\pi\right) $ and the equilibrium values of the choice probabilities, $\pi_{0}$ and $\pi_{1}$, in the pre and post intervention situations. To do this we will assume that the unobservables $\varepsilon_{vh}\left( =u_{vh}^{1}-u_{vh}^{0}\right) $ are independent of price and income, conditional on unobserved village-effects.\footnote{$q_{1}\left( p,y,\pi\right) $ defined in ((ref)) as the probability computed with respect to the distribution of $\eta^{0}-\eta^{1}\left( =\eta_{vh}^{0}-\eta_{vh}^{1}\right) $. But given the specification of $\boldsymbol{\eta}_{vh}=\left( \eta_{vh}^{0},\eta _{vh}^{1}\right) $ in Sections (ref)-(ref), it should be now interpreted as the one with respect to the distribution of $\varepsilon_{vh}$ (conditionally on the village-fixed effects $\left( d_{v},e_{v}\right) $ or $\xi_{v}$), i.e., the probability ((ref)) as a function of $w=\left( p,y\right) $ and $\bar{\pi}_{v}=\pi$.} Note that prices in our data are randomly assigned, so the endogeneity concern is solely regarding income. Under income endogeneity, Bhattacharya (2018) had discussed interpretation of welfare distributions as conditional on income (see our discussion at the end of Section (ref) ).

\noindentWelfare\ Calculations: Once we have estimates of the structural choice probabilities from the parametric model above, we can proceed with welfare calculation in presence of social spillovers and unobserved group-effects, as follows. Consider an initial situation where everyone faces the unsubsidized price $p_{0}$, so that the predicted take-up rate $\pi_{0v}$ in village $v$ solves:

equation[equation omitted — 143 chars of source]

where $F_{Y}^{v}\left( y\right) $ is the CDF of income $Y_{vh}$ in village $v$. Section 5.2 above outlines the calculations of all parameters including the $\xi_{v}$'s appearing in ((ref)).

Now consider a policy induced price regime $p_{0}$ for ineligibles (with wealth larger than $a$) and $p_{1}$ for the eligible (with wealth less than or equal to $a$). Then the resulting usage $\pi_{1}=\pi_{1v}$ in village $v$ is obtained via solving the fixed point $\pi_{1v}$ in the equation:

equation[equation omitted — 317 chars of source]

For fixed $\left( p_{0},p_{1}\right) $, the right-hand sides of the above fixed point equations ((ref)) and ((ref)), viewed as functions of $\pi_{0v}$ and $\pi_{1v}$ respectively, are a map from $\left[ 0,1\right] $ to $\left[ 0,1\right] $ ($\pi_{0v}$ and $\pi_{1v}$ are probabilities taking values in $\left[ 0,1\right] $). By the continuity of $\Phi$ and Brouwer's fixed point theorem, there is at least one solution in $\pi_{0v}$ and $\pi_{1v}$, respectively, implying \textquotedblleft coherence\textquotedblright. However, there may be multiple solutions, and then our welfare expressions would have to be applied separately for each feasible pair of values $\left( \pi_{0v},\pi_{1v}\right) $ (see our discussion on the equilibrium multiplicity in Section (ref) ). Note that even if the solutions to ((ref)) and ((ref)) are unique, our expressions in Theorem (ref) and Corollary (ref) imply that welfare distributions are still not point-identified.

Finally, mean welfare effect of the policy change in village $v$ can be calculated as

equation[equation omitted — 254 chars of source]

where $\mathcal{W}_{v}^{\mathrm{Elig}}\left( y\right) $ and $\mathcal{W} _{v}^{\mathrm{Inelig}}(y)$ are mean welfare at income $y$ in village $v$, calculated from ((ref)) for the eligible and ((ref)) for the ineligible, respectively, using $\pi_{0v}$ and $\pi_{1v}$ as the predicted take-up probabilities in village $v$ (analogous to $\pi_{0}$ and $\pi_{1}$ in ((ref)) and ((ref))), $\alpha_{1}\in\left[ 0,\alpha\right] $ as above).

Equilibrium Existence and Uniqueness

In this section, we present sufficient conditions of unique equilibrium and then discuss multiplicity of equilibria as well as its implication for our demand and welfare estimation. Note that given the parametric model, our equilibrium condition (or fixed point restriction) takes the form:

equation[equation omitted — 196 chars of source]

where $\tilde{F}_{W}\left( \cdot\right) $ is a CDF of $W_{vh}=\left( P_{vh},Y_{vh}\right) $. Some different $\tilde{F}_{W}$ has to be used, depending on the context. For example, on the RHS of ((ref)), $\tilde{F} _{W}$ corresponds to the distribution that gives a point mass for $P_{vh}=p_{0}$ (when considering a counterfactual analysis with this $p_{0}$) and $Y_{vh}\sim F_{Y}^{v}\left( y\right) $ (the marginal distribution of the observable variable $Y_{vh}$); and on the RHS of ((ref)), a different $\tilde{F}_{W}$ representing the new subsidy scheme is used.

As stated in the previous subsection, existence of a solution to ((ref)) follows from Brouwer's fixed point theorem. It is also clear that if $\alpha\leq0$, the solution is unique. On the other hand, if $\alpha>0$, a contraction condition is sufficient for uniqueness. The contraction condition ((ref)) in Proposition (ref) can be verified on a case by case basis. In particular, for the linear index model, it is easy to see that the condition for contraction is \[ \left\vert \alpha\right\vert \sup_{e\in\mathbb{R}}f_{\varepsilon}\left( e\right) <1, \] where $\alpha$ denotes the social interaction term, and $f_{\varepsilon }\left( \cdot\right) $ denotes the probability density of $-\varepsilon _{vh}$. In a probit specification in which $\varepsilon_{vh}$ is the standard normal variable, $\sup_{e\in\mathbb{R}}f_{\varepsilon}\left( e\right) =1/\sqrt{2\pi}$ and thus we require $\left\vert \alpha\right\vert <\sqrt{2\pi }(\simeq2.506)$ and for a logit specification, $\sup_{e\in\mathbb{R} }f_{\varepsilon}\left( e\right) =1/4$, and thus $\left\vert \alpha \right\vert <4$. We check that our probit estimate satisfies this condition in our application.

Note that the contraction condition ((ref)) is not necessary for uniqueness. That is, if a solution $(\bar{\Pi}_{v1},\dots ,\bar{\Pi}_{vN_{v}})$ to the system of equations ((ref) ) is unique and $m_{v}\left( \cdot\right) $ (defined in ((ref))), which also depends on the distribution of covariates, has a unique fixed point (i.e., a solution to $r=m_{v}\left( r\right) $ is unique), the uniqueness for the equilibrium solution holds. We have imposed ((ref)) as it is a convenient condition that guarantees uniqueness equilibrium solution; it is also typically easy to verify in applications.

\noindentThe Maximum Number of Equilibrium Solutions: The variable $w$ in ((ref)) is multivariate but its RHS can be written as $\int\Phi\left( q+c_{0}+\alpha\pi_{v}+\xi_{v}\right) dF_{W\boldsymbol{c} ^{\prime}}^{v}\left( q\right) $ in terms of the integral with respect to the univariate variable $W_{vh}\boldsymbol{c}^{\prime}$, using its CDF $F_{W\boldsymbol{c}^{\prime}}^{v}$ and support $[\underline{q}_{v},\bar{q} _{v}]$, where existence of the finite endpoints of the support of $W_{vh}\boldsymbol{c}^{\prime}$ is guaranteed under Condition CR2 (ii) (provided in Appendix A.3). Then, applying the mean value theorem for Stieltjes integrals, we can find some $q_{v}\in\lbrack\underline{q}_{v} ,\bar{q}_{v}]$ such that

align*[align* omitted — 370 chars of source]

Therefore, for each $v$, the fixed point restriction can be re-written as

equation[equation omitted — 132 chars of source]

Here, by shape properties of the standard normal CDF $\Phi\left( x\right) $ (e.g., its derivative is the normal density $\phi\left( x\right) $ with the two inflection points, $-1$ and $1$), we can see that ((ref) ) has at most three solutions ($\pi_{v}=0$ or $1$ cannot be a solution since each value of $\Phi$ is on $\left( 0,1\right) $). In particular, a continuum of solutions cannot exist since $\Phi\left( x\right) $ does not have a linear part on any interval in the real line. This is summarized via the following proposition whose proof is also evident from the above discussion.

propositionFor each $v$, the maximum number of (equilibrium) solutions to ((ref)) is three.

This is analogous to Proposition 2 of BD01 for the logit distribution case without covariates $W_{vh}$. The number of equilibria is determined by the value of $\alpha$ as well as those of $q_{v}$, $c_{0}$, and (in particular) the unobserved group effects $\xi_{v}$. We now discuss implications of equilibrium multiplicity in estimation.

\noindentPreference and Demand Estimation under Multiple Equilibria: Our estimation involves maximum likelihood in nonlinear models with strategic interactions; thus, it is useful to recall Hahn and Moon (2010) who consider estimation of game theoretic models possibly with multiple equilibria under a panel setting, i.e., observations from many markets are obtained repeatedly over several time periods. These authors interpret unobservables affecting equilibrium selection as an unobserved fixed effect, assuming that which equilibrium is selected in one group is fully characterized by each unobserved fixed effect, which may be correlated with observed characteristics. Then they show that equilibrium multiplicity is unlikely to be a problem in panel settings when the number of equilibria is finite which, in the panel terminology, is equivalent to the fixed effect having finitely many support points. However, this result requires that the number of equilibria be constant across parameters and covariates, which is a strong restriction.

In contrast, in our setting, each village/group effect can be interpreted as a part of players' preference parameters and it does not fully determine which equilibrium arises. That is, under the same value of group/village effect $\xi_{v}$, we may observe different equilibria in village $v$. Given the knowledge of all the preference parameters and the distribution of all the covariates and error variables, one can determine the number of possible equilibria and possible values of beliefs by investigating all solutions of the fixed point equation ((ref)), but cannot in advance see which equilibrium would be realized. We note that our model setting is not equipped with any equilibrium selection mechanism (just like thus in Brock and Durlauf's). In our setup, given $\bar{\pi}_{v}$ and $\xi_{v}$, we do not need to solve the equilibrium system to predict each player's behavior, which is determined by $A_{vh}=1\{W_{vh}\boldsymbol{c}^{\prime}+c_{0}+\alpha\bar{\pi }_{v}+\xi_{v}+\varepsilon_{vh}\geq0\}$, and preference parameters can be estimated without exploiting the equilibrium fixed point condition.\footnote{This is quite different from the so-called two step estimation approach (typically used in the empirical industrial organization game literature) as in Hotz and Miller (1993) and Pesendorfer and Schmidt-Dengler (2008), in which equilibrium conditions provide the basis of identification.} This is possible since (A)\ the only objects that are endogenously determined in equilibrium are $\bar{\pi}_{v}$ ($v=1,\dots,\bar {v}$), which can be identifiable as $E_{v}\left[ A_{vh}\right] $ and thus consistently estimated in our \textquotedblleft large market\textquotedblright \ setting with a large number of players in each village (as discussed in Section (ref)); and (B) the group effects parameters $\xi _{v}$ can also be consistently estimated under the correlated random effects structure.

These features of our (and Brock and Durlauf's) modeling allow us to avoid intrinsically difficult problems caused by the equilibrium multiplicity. In particular, in any equilibrium realization, the same preference parameters (that are invariant under different equilibria) can be identified and thus consistently estimated. As a further illustration, consider a case in which there are three equilibria, i.e., the fixed point equation ((ref)) has three solutions, $\bar{\pi}_{v}^{H}$, $\bar{\pi}_{v}^{M}$, and $\bar{\pi}_{v}^{L}$, where we let $\bar{\pi}_{v}^{H}>\bar{\pi}_{v}^{M}>\bar{\pi}_{v}^{L}$ and call each of equilibria as $H$, $M$, or $L$. Then, depending on $t\in\{H,M,L\}$, we have a different discrete choice model: \[ A_{vh}=1\{W_{vh}\boldsymbol{c}^{\prime}+c_{0}+\alpha\pi_{v}^{t}+\xi _{v}+\varepsilon_{vh}\geq0\}. \] Note that the outcome variable changes depending on which equilibrium arises (i.e., one can write $A_{vh}=A_{vh}^{t}$); and thus, for each equilibrium $t$, $\pi_{v}^{t}$ can be consistently estimated by $\frac{1}{N_{v}}\sum _{h=1}^{N_{v}}A_{vh}$. By plugging in the estimated version of $\pi_{v} =\pi_{v}^{t}$, we can construct objective functions (i.e., likelihood functions) to be maximized, based on which we can consistently estimate the preference parameters regardless of the realized equilibrium $t\in\{H,M,L\}$. As for consistency, the preference parameters can be identified as the unique maximizers of the limits of the objective functions. In particular, for our first probit regression (cf. Section (ref)), the choice probability is \[ \Pr\left( A_{vh}^{t}=1|W_{vh}=w;d_{v},e_{v}\right) =\Phi(w(\boldsymbol{c} ^{\ast})^{\prime}+\gamma_{v}^{\ast t}), \] under equilibrium $t$, where $\gamma_{v}^{\ast t}(=c_{0}^{\ast}+\alpha^{\ast }\bar{\pi}_{v}^{t}+\xi_{v}$) depends on which equilibrium has occurred, and the (limit) objective function (under equilibrium $t$) is \[ Q_{v}\left( \boldsymbol{c},\gamma_{v}\right) =E_{v}\left[ A_{vh}^{t} \log\Phi\left( W_{vh}\boldsymbol{c}^{\prime}+\gamma_{v}\right) +\left( 1-A_{vh}^{t}\right) \log\left( 1-\Phi\left( W_{vh}\boldsymbol{c}^{\prime }+\gamma_{v}\right) \right) \right] , \] where $E_{v}\left[ \cdot\right] =E\left[ \cdot|d_{v},e_{v}\right] $. Then, given the true parameter $(\boldsymbol{c}^{\ast},\gamma_{v}^{t\ast} )(\neq\left( \boldsymbol{c},\gamma_{v}\right) )$

align*[align* omitted — 851 chars of source]

where the equality has used the law of iterated expectation and correct specification assumption (i.e., $E_{v}\left[ A_{vh}^{t}|W_{vh}\right] =\Phi(W_{vh}(\boldsymbol{c}^{\ast})^{\prime}+\gamma_{v}^{t\ast})$), and the strict inequality follows from Jensen's inequality, the strict convexity of $-\log\left( \cdot\right) $, and the rank condition on $(W_{vh},1)$ (CR1 (ii)).\footnote{This identification argument is standard (as in Newey and McFadden, 1994, Example 1.2 on page 2125). \newline While we believe that this inequality for each $v$ is useful for illustrating identification under the equilibrium multiplicity, it is not sufficient for consistency when $\bar{v}\rightarrow\infty$ i.e., Proposition (ref). For verification of the proposition, we have derived uniform lower bound of $Q_{v}\left( \boldsymbol{c}^{\ast},\gamma_{v}^{\ast t}\right) -Q_{v}\left( \boldsymbol{c},\gamma_{v}\right) $ (Lemma (ref) in the Appendix).} That is, we have \[ Q_{v}\left( \boldsymbol{c}^{\ast},\gamma_{v}^{t\ast}\right) >Q_{v}\left( \boldsymbol{c},\gamma_{v}\right) \text{ for any }\left( \boldsymbol{c} ,\gamma_{v}\right) \neq\left( \boldsymbol{c}^{\ast},\gamma_{v}^{t\ast }\right) . \] Thus, $\left( \boldsymbol{c}^{\ast},\gamma_{v}^{t\ast}\right) $ is identified as the unique maximizer of $Q_{v}\left( \cdot,\cdot\right) $; in particular, the same $\boldsymbol{c}^{\ast}$ is always identified under any of equilibrium $t$, while the identified $\gamma_{v}^{t\ast}$ depends on $t$, which corresponds to our specification of $\gamma_{v}^{t\ast}=c_{0}^{\ast }+\alpha^{\ast}\bar{\pi}_{v}^{t}+\xi_{v}$ (including the equilibrium object $\bar{\pi}_{v}^{t}$). The identification argument for the second probit under the CRE structure is analogous (so details are omitted here): under any equilibrium $t$, the same $(\boldsymbol{\bar{c}}^{\ast},\bar{c}^{\ast} ,\bar{\alpha}^{\ast}\boldsymbol{\bar{\delta}}^{\ast})$ is obtained as the unique maximizer of the limit of $\hat{R}(\boldsymbol{\bar{c}},\bar{c} _{0},\bar{\alpha},\boldsymbol{\bar{\delta}})$. That is, given the CRE structure, the same group effects $\xi_{v}$ can be identified in any equilibrium $t$ through the procedure outlined in Section (ref). Thus our estimation procedure need not to use the equilibrium fixed point restriction and is robust to equilibrium multiplicity.

In our empirical application, we use an iterative estimator that exploits the equilibrium fixed point restriction as in Pastorello, Patilea, and Renault (2003), and Dominitz and Sherman (2005). This estimator is more efficient under correct specification than the estimator that does not use ((ref)). For this iterative estimation, the contraction property of the fixed point mapping (implying unique equilibrium) is key, the sufficient condition for which is a \textquotedblleft small $\alpha $\textquotedblright. Through preliminary investigation i.e. checking estimates obtained without exploiting the fixed-point condition ((ref)), we have confirmed that the estimate of $\alpha$ is small, so that the contraction condition is satisfied.

\noindentCounterfactual Welfare Estimation under Equilibrium Multiplicity: As discussed above, we do not need to solve the equilibrium condition ((ref)) for estimation of preference parameters and the $\xi_{v}$'s (we do use the equilibrium conditions to predict the counterfactual $\pi$ resulting from the policy experiment), and are therefore not affected by the multiplicity. However, when predicting counterfactual outcomes, we need to solve the equilibrium fixed point condition, and find solutions $\pi_{v}$'s in the counterfactual scenario, e.g., a hypothetical subsidy rule to buy an ITN. Given the already estimated structural parameter values, the solutions $\pi_{v}$ of the fixed point equation ((ref)) in the counterfactual scenario can be computed for each $v$. If equilibrium multiplicity is anticipated, e.g., when the estimated $\alpha$ is larger than the threshold for contraction, we can compute these multiple solutions for each $v\in\left\{ 1,\dots,\bar{v}\right\} $ through eye-balling\ since the number of solutions is at most three. This can be done by drawing a graph of the LHS of ((ref)) and checking points on the graph that intersect the 45 degree line (or the graph for a least squares objective function as in ((ref)), presented in Figure 2). The number of villages is eleven in our application and eye-balling is not difficult. Then, given multiple solutions, a bound for average welfare can be computed for each solution, and one can report multiple (at most three) bounds of them or a single union of the multiple intervals.

EMPIRICAL\ CONTEXT AND\ DATA

Our empirical application concerns the provision of anti-malarial bednets. Malaria is a life-threatening parasitic disease transmitted from human to human through mosquitoes. In 2019, an estimated 229 million cases of malaria occurred worldwide, with 90% of the cases in sub-Saharan Africa (WHO, 2017). The main tool for malaria control in sub-Saharan Africa is the use of insecticide treated bednets. Regular use of a bednet reduces overall child mortality by around 18 percent and reduces morbidity for the entire population (Lengeler, 2004). However, at \$6 or more a piece, bednets are unaffordable for many households, and to palliate the very low coverage levels observed in the mid-2000s, public subsidy schemes were introduced in numerous countries in the last 15 years. Our empirical exercise is designed to evaluate such subsidy schemes not just in terms of their effectiveness in promoting bednet adoption, but also their impact on individual welfare and deadweight loss. Based on our discussion in Section 4, we focus on two main sources of spillover, viz. (a) a preference for conformity, and (b) a concern that mosquitoes will be deflected to oneself when neighbors protect themselves. Both will generate a positive effect of the aggregate adoption rate on one's own adoption decision, but they have different implications for the welfare impact of a price subsidy policy.

\noindentExperimental Design: We exploit data from a 2007 randomized bednet\ pricing experiment conducted in eleven villages of Western Kenya, where malaria is transmitted year-round. In each village, a list of $150$ to $200$ households was compiled from school registers, and households on the list were randomly assigned to a price at which they could purchase a long-lasting ITN, a new, highly effective type of antimalarial bednet. After the random assignment had been performed in office, trained enumerators visited each sampled household to administer a baseline survey. At the end of the interview, the household was given a voucher for one long-lasting ITN at the randomly assigned price level. The amount of subsidy (for those who received any) varied from $40$% to $100$% of the market price in two villages, and from $40$% to $90$% in the remaining $9$ villages; there were $22$ corresponding final prices faced by households, ranging from $0$ to $300$ Ksh (US \$$5.50$), where $300$ Ksh would be the non-subsidized sale price. Vouchers could be redeemed within three months at participating local retailers.

\noindentData: We use data on bednet adoption as observed from coupon redemption and verified acquisition through a follow-up survey. We also use data on baseline household characteristics measured during the baseline survey. The three main baseline characteristics we consider are wealth (the combined value of all durable and animal assets owned by the household); the number of children under ten years old; and the education level of the female head of household.

\noindentNonparticipating Households: While all households in a given village were potentially interacting, our sample does not cover all village members. This can potentially cause a problem since selected households might interact with non-selected ones. However, at the time of the experiment, non-selected households did not have the opportunity to buy a long-lasting ITN, so the outcome variable $A$ for such households is zero, whose conditional expectations are zero as well. Thus, in our specification, even if we allow for interactions among all the village members, it is easy to do the necessary adjustments in the empirics, viz. replace $\Pi_{vh}$ in ((ref)) by

equation[equation omitted — 246 chars of source]

where $\check{N}_{v}$ equals the total number of households in the village, and $N_{v}$ those participating in the game. In our empirical setting, this ratio is about $0.8$ for each village, and we apply this adjustment throughout the empirical analysis.

EMPIRICAL SPECIFICATION AND RESULTS

We work with the linear index structure ((ref)), where $Y_{vh}$ is taken to be the household wealth, $P_{vh}$ is the experimentally set price faced by the household, $\Pi_{vh}$ is the observed average adoption in the village. We also use additional controls, denoted by $Z_{vh}$ below, that can potentially affect preferences and therefore the take-up of bednet, viz. presence of children under the age of ten, and years of education of the oldest female member of the household. A village-specific variable that could affect adoption is the extent of malaria exposure risk in the village. We measure this in our data from the response to the question: “Did anyone in your household have malaria in the past month?”. Summary statistics are reported in Table 1, and their village averages are shown in Table 2, for each of the eleven villages in the data.

Our first set of results correspond to taking $F\left( \cdot\right) $ to be the probit CDF of $\eta_{vh}=\eta_{vh}^{0}-\eta_{vh}^{1}$ (as in ((ref))), i.e. with no village-effects), and then our main results use the correlated random effects probit model that accounts for village-level unobservables. The marginal effects at mean are presented in Table 3, corresponding to both the probit model without village-effects and the CRE probit model that accounts for village-level unobservables. The fact that the price elasticity is very similar in the two specifications is expected, since price was exogenously assigned in the experiment, so accounting for village-fixed unobservables has no impact on the marginal effect.

It is evident from the table that the demand coefficient is negative and significantly different from zero (the averaged price elasticity is $-0.12)$, and that bednet adoption in the village has a significant positive association with private adoption, conditional on price and other household characteristics, i.e. $\alpha>0$ in our notation above. The social interaction coefficient $\alpha$ is $2.2$ for the probit, which is less than the upper bound for the fixed point map to be a contraction (see discussion in Section (ref)). The effect of children is negative, likely reflecting that households with children had already invested in other anti-malarial steps, e.g., had bought a less effective traditional bednet prior to the experiment.

Next, we consider a hypothetical subsidy rule, where those with wealth less than $\tau$ are eligible to get the bednet for $50$ KSh ($83$% subsidy), whereas those with wealth larger than $\tau$ get it for the price of $250$ KSh ($17$% subsidy). Based on our preferred CRE probit model (which is used for all subsequent results, unless mentioned otherwise), we plot the predicted aggregate take-up of bednets corresponding to different income thresholds $\tau$. In Figure 1, for each threshold $\tau$, we plot the fraction of households eligible for a subsidy on the horizontal axis, and the predicted fraction choosing the bednet on the vertical axis, based on coefficients obtained by including (solid) and excluding (small dash) the spillover effect. The 45 degree line (large dash), showing the fraction eligible for the subsidy, is also plotted in the same figure for comparison.

It is evident from Figure 1 that ignoring spillovers leads to over-estimation of adoption at lower thresholds and underestimation at higher thresholds of eligibility. This happens because ignoring a covariate (here $\pi$) with positive impact on the outcome in prediction amounts to \textquotedblleft smoothing\textquotedblright\ over values of $\pi$.

Having obtained these (uncompensated) effects, we now turn to calculating the demand and the mean compensating variation for a hypothetical subsidy scheme. We consider an initial situation where everyone faces a price of $250$ KSh for the bednet, and a final situation where a bednet is offered for $50$ KSh to households with wealth less than $\tau=8000$ KSh (about the $27$th percentile of the wealth distribution), and for the price of $250$ KSh to those with wealth above that. The demand results are reported in Table 4, and the welfare results in Table 5. We perform these calculations village-by-village, and then aggregate across villages. To calculate these numbers, we first predict the bednet adoption when everyone is facing a price of $250$ KSh, and then when eligibles face a price of $50$ KSh and the rest stay at $250$ KSh, giving us the equilibrium values of $\pi_{0}$ and $\pi_{1}$, respectively. In all such calculations with our data, we always detected a single solution to the fixed point $\pi$ (i.e. a unique equilibrium) as can be seen from Figure 2, where we plot the squared difference between the RHS and the LHS of an empirical version of the fixed point equation ((ref)) (with the additional covariate $Z_{vh}$), i.e.

equation[equation omitted — 266 chars of source]

on the vertical axis, and $\pi_{1}$ on the horizontal axis, separately for each of the eleven villages, where $\hat{q}_{1}\left( p,y,z,\pi\right) $ is the predicted demand (choice probability) function at $\left( p,y,z,\pi \right) $. The globally convex nature of each objective function is evident from Figure 1. The minima are relatively close to each other around $0.15$, except village 7 and 10, where it is larger. As for $\pi_{0}$, which minimizes $\left[ \pi_{0}-\int\hat{q}_{1}\left( p_{1},y,z,\pi_{0}\right) d\hat {F}_{Y,Z}\left( y,z\right) \right] ^{2}$, the objective function is also convex with a minimizing value close to zero in every village, reflecting that very few households would buy at this high price. These predicted values of $\pi_{0}$ and $\pi_{1}$ are used as inputs into the prediction of demand as the structural choice probability ((ref)) and welfare as per Theorem (ref) and Corollary (ref).

The first row of Table 4 shows the pre-subsidy predicted demand by subsidy eligibility. In the second row, we calculate the predicted effect of the subsidy on demand, and break that up by the own price effect (Row 2) and the spillover effect (Row 3). The own effect is obtained by changing the price in accordance with the subsidy but keeping the village demand equal to the pre-subsidy value; the spillover effect is the difference between the overall effect and the own effect. It is clear that spillover effects on both eligibles and ineligibles are large in magnitude. In particular, the spillovers effect raises demand for ineligibles by an amount that nearly equals its pre-subsidy level.

In Table 5, we report welfare calculations with standard errors computed via the simple nonparametric bootstrap where households were resampled within each village in each bootstrap replication. In the first row, we report the welfare gain of the subsidy rule for eligibles, first assuming no spillovers and using a probit model without village-effects. In this case, we simply use the results of Bhattacharya (2015) to calculate the (point-identified) CV for eligibles as the price changes from $250$ KSh to $50$ KSh. This yields the value of welfare gain to be $52.589$ KSh. As there is no spillover, the welfare change of ineligibles is zero by definition, and therefore the net welfare gain is simply the fraction eligible ($0.27$) times the CV for eligibles. This is reported in the third column of Table 5. The case with spillovers under probit and assuming $\alpha_{1}\geq0\geq\alpha_{0}$ are reported in the second panel of Row 1 using (the negatives of) ((ref)), ((ref)) and ((ref)) for eligibles, and using ((ref)), ((ref)) and ((ref)) for ineligibles.

The 2nd-4th row present analogous results using CRE probit to control for fixed effects; the 2nd row does this for $\alpha_{1}\geq0\geq\alpha_{0}$ ; the third row for $\alpha_{1}\geq\alpha_{0}\geq0$, using a large upper limit of $\alpha_{1}$ (and concurrently $\alpha_{0}=\alpha_{1}-\alpha>0$) to proxy $\alpha_{1}\nearrow\infty$, (cf. ((ref)) and ((ref)) above). Finally, the 4th row presents the overall bounds by taking union of the previous two cases.

Under $\alpha_{1}\geq0\geq\alpha_{0}$, both specifications imply that ineligibles can suffer a large welfare loss due to the subsidy. This is because the subsidy facilitates usage for solely the eligibles, raising the equilibrium usage $\pi$ in the village, but the ineligibles keep facing the high price, and thus a lower utility from not buying because $\pi$ is now higher and $\alpha_{0}\leq0$. However, the few ineligibles who buy, despite the high price, get some welfare increase from a rise in the adoption rate, that explains the small upper bound corresponding to the case $\alpha_{0}=0$. The overall welfare gain aggregated over eligibles and ineligibles is reported in the column headed \textquotedblleft Net Welfare Gain\textquotedblright.

\noindentDeadweight Loss: To compute the deadweight loss, we subtract the net welfare from the predicted subsidy expenditure. The latter equals the amount of subsidy ($200$ KSh) times the demand at the subsidized price 50 KSh of the eligibles. Thus the expression for DWL is given by \[ D=\int\left[

array[array omitted — 263 chars of source]

\right] dF\left( y,z\right) , \] where $y$ denotes wealth, $z$ denotes other covariates, $q_{1}\left( 50,y,z,\pi_{1}\right) $ denotes predicted demand at price $50$ KSh including the effect of spillover, and $\mu^{\mathrm{Elig}}$ and $\mu^{\mathrm{Inelig}}$ refer to welfare gain for eligibles and ineligibles, respectively. Ignoring spillovers leads to the point-identified deadweight loss \[ D=\int\left[ 200\times1\left\{ y\leq\tau\right\} \times q_{1} ^{\mathrm{No}\text{\textrm{-}}\mathrm{spillover}}\left( 50,y,z\right) -1\left\{ y\leq\tau\right\} \mu^{\mathrm{No}\text{\textrm{-}} \mathrm{spillover}}\left( y,z\right) \right] dF\left( y,z\right) \text{.} \]

For the case $\alpha_{1}\geq\alpha_{0}\geq0$, in the last-but-one row of Table 5, there is no welfare loss for anyone, since all spillover is positive, which explains the negative deadweight loss lower bound, i.e. an efficiency from subsidizing a positive externality.

These welfare and DWL numbers support the overall conclusion that accounting for spillovers can lead to much lower estimates of net welfare gain from the subsidy program and higher deadweight loss. Some of this difference arises from potential welfare loss suffered by ineligibles that is missed upon assuming no spillover, and some from the impact of including spillovers terms on the prediction of counterfactual purchase-rates (cf. Fig 1). Furthermore, the two cases $\alpha_{1}\geq0\geq\alpha_{0}$ and $\alpha_{1}\geq\alpha _{0}\geq0$, which are both consistent with the observed $\alpha=\alpha _{1}-\alpha_{0}>0$, yield vastly different bounds on welfare, resulting in wide overall bounds on net welfare gain and deadweight loss that include zero (cf. last row of Table 5), which is the key substantive point of this paper.

\noindentEndogeneity: Price variation is exogenous in our application, since price was varied randomly by the experimenter. Indeed, it is still possible that wealth $Y$ is correlated with $\boldsymbol{\eta}$, the unobserved determinants of bednet purchase (even conditionally on the village specific effects). However, experimental variation in price $P$ implies also that $P$ is independent of $\boldsymbol{\eta}$, given $Y$. Consequently, one can invoke the argument presented in Bhattacharya (2018, Section 3.1), and interpret the estimated choice probabilities and the corresponding welfare numbers as conditional on $y$, and then integrating with respect to the marginal distribution of $y$. This overcomes the problem posed by potentially endogenous income.

CONCLUSION

This paper develops tools for economic demand and welfare analysis in binary choice models with social interactions. The key finding is that under interactions, welfare distributions resulting from policy changes such as a price subsidy are generically not point-identified for given values of counterfactual aggregate demand, unlike the case without spillovers. This is true even when utility functions and distribution of unobserved heterogeneity are fully parametrized and there is a unique equilibrium. Non-identification results from the inability of standard choice data to distinguish between different underlying latent mechanisms, e.g. conforming motives, consumer learning, negative externalities etc., which produce the same aggregate social interaction coefficient, but have different welfare implications depending on which mechanism dominates. This feature is endemic to many practical settings that economists study, including the health-product adoption case examined here. Another prominent example is school-choice, where merit-based vouchers to attend a fee-paying selective school can create negative externalities by lowering the academic quality of the free local school via increased departure of high-achieving students. The resulting welfare implications cannot be calculated based solely on a Brock-Durlauf style empirical model of individual school-choice inclusive of a social interaction term. This is in contrast to models without social interaction, where choice probability functions have been shown to contain all the information required for welfare analysis. Nonetheless, we show that under standard linear index restrictions, welfare distributions can be bounded. Under some special and empirically untestable cases e.g. exactly symmetric spillovers effects or absence of negative externalities, these bounds shrink to point-identified values. Next, we develop methods of identification and consistent estimation for the structural utility parameters, required for prediction of counterfactual outcome and welfare bounds, when there is unobserved group-level heterogeneity possibly correlated with observable covariates. This is achieved via a novel latent factor modelling of unobserved group-effects and observed covariates, and developing a method of asymptotic analysis where the dimension of nuisance parameters, i.e. the group-effects whose magnitude may be unbounded, increases as the number of groups increase.

We apply our methods to an empirical setting of adoption of anti-malarial bednets, using data from a pricing experiment by Dupas (2014) in rural Kenya. We find that accounting for spillovers provides different predictions for demand and welfare resulting from hypothetical, means-tested subsidy rules. In particular, with positive interaction effects, predicted demand when including spillovers are lower for less generous eligibility criteria, compared to demand predicted by ignoring spillovers. At more generous eligibility thresholds, the conclusion reverses. As for welfare, if negative health externalities are present, then subsidy-ineligibles can suffer welfare loss due to increased use by subsidized buyers in the neighborhood; if solely conforming effects are present and there is no health-related externality, then welfare can improve.

The implication of these results for applied work is that under social interactions, welfare analysis of potential interventions requires more information regarding individual channels of spillovers than knowledge of solely the choice probability functions inclusive of a social interaction term. Belief-eliciting surveys, recording the reasons behind the subjects' actions, can provide a potential solution.

thebibliography{99} \bibitem BHATTACHARYA, D. (2008), A Permutation-based estimator for monotone index models. Econometric Theory 24, 795-807. \bibitem BHATTACHARYA, D. (2015), Nonparametric welfare analysis for discrete choice. Econometrica 83, 617-649. \bibitem BHATTACHARYA, D. (2018), Empirical welfare analysis for discrete choice: Some general results. Quantitative Economics 9, 571-615. \bibitem BISIN, A., MORO, A. and TOPA, G., (2011), The empirical content of models with multiple equilibria in economies with social interactions (No. w17196). National Bureau of Economic Research. \bibitem BLUNDELL, R. & POWELL, J. (2004), Endogeneity in nonparametric and semiparametric regression Models, in Advances in Economics and Econometrics, Cambridge University Press, Cambridge, UK. \bibitem BROCK, W.A. & DURLAUF, S.N. (2001), Discrete choice with social spillover. The Review of Economic Studies 68, 235-60. \bibitem BROCK, W.A. & DURLAUF, S.N. (2007), Identification of binary choice models with social interactions. Journal of Econometrics 140, 52-75. \bibitem CHAMBERLAIN, G. (1980), Analysis of covariance with qualitative data. The Review of Economic Studies 47, 225-238. \bibitem DALY, A. & ZACHARY, S. (1978), Improved multiple choice models, in Determinants of travel choice, 335-357. \bibitem DOMINITZ, J. & SHERMAN R.P. (2005), Some convergence theory for iterative estimation procedures with an application to semiparametric estimation. Econometric Theory 21, 838-863. \bibitem DUPAS, P. (2014), Short-run subsidies and long-run adoption of new health products: Evidence from a field experiment. Econometrica 82, 197-228. \bibitem FERNANDEZ-VAL, I. & WEIDNER, M. (2016), Individual and time effects in nonlinear panel models with large $N,T$, Journal of Econometrics 192, 291-312. \bibitem GAUTAM, S. (2018), Quantifying welfare effects in the presence of externalities: An ex-ante evaluation of a sanitation intervention, Mimeo. \bibitem HAUSMAN, J.A, & NEWEY W. (2016), Individual heterogeneity and welfare. Econometrica 84, 1225-48. \bibitem HAHN, J. & KUERSTEINER, G. (2011), Bias reduction for dynamic nonlinear panel models with fixed effects, Econometric Theory 27, 1152-1191. \bibitem HAHN, J. & MOON, H.R. (2010), Panel data models with finite number of multiple equilibria, Econometric Theory 26, 863-881. \bibitem HOTZ, V.J. & MILLIER, R.A. (1993), Conditional choice probabilities and the estimation of dynamic models. The Review of Economic Studies 60, 497-529. \bibitem LENGELER, C. (2004), Insecticide-treated bed nets and curtains for preventing malaria. The Cochrane Library. \bibitem MANSKI, C.F. (1993), Identification of endogenous social effects: The reflection problem. The Review of Economic Studies 60, 531-542. \bibitem McFADDEN, D. & TRAIN, K., (2019), Welfare economics in product markets. Working paper, University of California, Berkeley. \bibitem NEWEY, W.K. & McFADDEN, D. (1994), Large sample estimation and hypothesis testing, Ch. 36 in Handbook of Econometrics, Volume IV, (Ed. by R.F. Engle & D.L. McFadden) \bibitem PESENDORFER, M. and SCHMIDT-DENGLER, P. (2008), Asymptotic least squares estimator for dynamic games. The Reviews of Economics Studies 75, 901-928. \bibitem PASTORELLO, S., PATILEA, V. & RENAULT, E. (2003), Iterative and recursive estimation in structural nonadaptive models. Journal of Business & Economic Statistics 21, 449-509. \bibitem RUST, J. (1987), Optimal replacement of GMC bus engines: An Empirical model of Harold Zurcher. Econometrica 55, 999-1033. \bibitem DE PAULA, A. (2017), Econometrics of network models. In Advances in economics and econometrics: Theory and applications, eleventh world congress (pp. 268-323), Cambridge University Press. \bibitem SMALL, K. & ROSEN, H. (1981). Applied welfare economics with discrete choice models. Econometrica 49, 105-130. \bibitem WHO (2020), World malaria report 2020, Geneva, World Health Organization, Licence: CCBY-NC-SA 3.0 IGO. \bibitem WOOLDRIDGE, J.M. (2010), Econometric analysis of cross section and panel data. MIT press.