Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
51,821 characters · 7 sections · 3 citation commands
A Trimming Estimator for the Latent-Diffusion-Observed-Adoption Model
Diffusion in networks has been vastly studied in biology and Computer Science. Yet recently it has increasingly been acknowledged that also economic agents exchange goods, services or assets in social networks and that it is highly important and interesting to estimate at which rate they do so. A key challenge here is that economic interaction in social networks is oftentimes unobserved. \\ We can tackle such a situation assuming that economic events which we spot are but a signal of the process that jointly determines an observed outcome and an unobserved state variable. Yet, fitting such a model to the data is challenging due to the high dimensionality: at each point in time, there exists a system of equations relating all the unobserved individual state variables to one another, and the dynamics of this system are governed by the network. With social networks typically being dense and large, we thus have a gigantic number of scenarios that could have given rise to the data at hand. This is problematic in particular if the time horizon we wish to consider is large. \\ Due to the nonlinear nature of network diffusion models, Maximum Likelihood (ML) estimation would be the preferred estimation method. However, setting up the log likelihood function is computationally infeasible for larger time horizons. The gigantesque number of possible realizations of the latent diffusion process that could have generated the observed adoption pattern makes it impossible to establish the probability to observe the given data by applying the law of total probability (i.e.\ by integrating out the latent process). To ensure tractability one needs to either restrict the model time horizon or to use a dimension reduction strategy. In this paper, I conduct a Monte Carlo study in order to compare these two alternatives. This shows that the “Trimming estimator", derived in the present chapter, outperforms the previously used estimator that restrict the number of rounds of network interaction that can take place in most cases (precisely, whenever the sample includes a sufficient number of villages for which trimming impacts no more than one-third of the individuals that are eligible to be trimmed). The “Trimming estimator" is thereafter applied to four real villages from the BCDJ study. \\ I propose a simple trimming rule that imposes restrictions on the number of permissible results of the unobservable network interaction and compare it to the estimate that results from restricting the time horizon modeled. Fast and flexible, the algorithm can be used to estimate a class of network models that comprises an observed outcome and an unobserved state variable. The strategy is easily implemented and oftentimes succeeds in selecting the interaction outcomes with the largest probability mass. I also investigate how features of the network affect performance of the algorithm.\\ I demonstrate, that the MLE obtained from maximizing the approximate log likelihood function quickly converges to the exact MLE as the trimming value increases. I proceed by establishing the (approximate) log likelihood function and finding its maximum by means of grid search. This entails certain advantages. As a byproduct we obtain the log likelihood function which, plotted over the grid, can provide us with additional insights (e.g.\ I can detect local maxima, areas of weak identification or illustrate the covariance of the parameters). Confidence sets can be constructed straightforwardly using the Likelihood Ratio (LR) test and corner solutions can be dealt with. \\ The application considered stems from the seminal study BCDJ, in which the authors investigate the diffusion of micro-finance in Indian village networks. Linked agents repeatedly exchange new information, the information status being unobserved. Upon being newly informed, they can make an observable borrowing choice. We use eleven of the village networks from this data set in our simulation study. Nonetheless, it is important to emphasize that the applicability of our algorithm goes well beyond this particular model. A model in which an observed outcome depends on an unobserved status lends itself to various applications in economics. Instead of information, one could model technology spill-over, where we observe the adoption of new technologies, but for those that abstain from it we do not know whether this is by deliberate choice or by lack of knowledge. Alternatively, one can estimate a model in which the probability of (observed) micro-business default depends on (unobserved) mutual lending among friends. \\ In recent years, the amount of data on (action taking place through) social networks that is available has increased drastically, providing new exciting research possibilities, but also serious challenges for econometricians. What shies away many applied researchers from fully embracing this opportunity and using MLE in settings where it provides clear advantages is often the lack of tractable, easily implementable algorithms. While established software packages exist for common estimation procedures, these typically work very well for specific tasks, but lack the flexibility required by network models. They are not fully efficient in terms of speed and provide little opportunity to deal with large dimensionality of the data. Research into methodological approaches facilitating ML estimation in such models thus has the potential to be beneficial. In the current paper, I make a contribution to the this new strand of research.\\ The paper is organized as follows: Section 2 provides an overview of the model. In section 3 we highlight the challenges of ML estimation in this model and introduce the basic trimming procedure. Section 4 outlines the Monte Carlo study. Section 5 provides and analyses the Monte Carlo results. Section 6 contains the results from applying the trimming estimator to four real villages using actual adoption data. Section 7 concludes. The Appendices highlight how features of the network impact the algorithm's performance and the shape of the error curve.
In this and the following sections, we use the terminology of the concrete application despite our algorithm being applicable to many research questions. \\ We dispose over a sample of villages, indexed by $v=1,...,V$. \\ $N_v$ individuals in village $v$ are linked in a network, represented by the network adjacency matrix $G_v$. While villagers $i=1,...,N_v$ entertain relationships among themselves, they are supposed not to interfere with agents living in other villages. Consequently, to ease notation, in the following we drop the subscript $v$ and we outline the model with an exemplary village. \\ The symmetric matrix $G$ comprises binary variables indicating the presence (absence) of a link. If the entry in row $i$, column $j$ takes the value of one (i.e. $g_{ij}=g_{ji}=1$), we denominate $i$ and $j$ as “neighbors" or “friends". An organization enters the community and starts providing a new technology to its members. In period $t=0$, they advertise their innovation to a subset of them (referred to as the “information injection points" or “IPs") and thereafter rely on word-of-mouth marketing. \\ In each subsequent period ($t=1,...,4$), two processes take place: first, newly informed individuals face the choice of whether or not to adopt the technology, second, informed individuals can instruct their neighbors about the novel opportunity. \\ The process is modeled using the random matrices $Y$ and $S$, containing dummy variables for respectively the participation and information statuses of each of the villagers at each point in time. Using the first subscript for individuals and the second subscript for time periods, thus $Y_{it}=1$ ($S_{it}=1$) (i.e.\ the row $i$, column $t$ element of $Y$ ($S$) displays the value one) implies that $i$ participates (is informed) in period $t$. The distributions of the random matrices $Y$ and $S$ depend on the village network ($G$) and the information injection $s_{0}$). \\ For each village, we observe the network ($G$), the information initialization $s_{0}$ and the individual participation decisions over time (i.e. one realization of $Y$). The information statuses of all inhabitants but the IPs are generally unobserved. In particular, there usually exist various realizations of $S$ that are in accordance with the data at hand. \\ The following assumptions are made: \\
Assumption 1: Independence across villages \\ Villages can be treated as independent entities. In particular, there is no link between any two agents that reside in different communities. \\ \\ Assumption 2: Exogenous network \\ The network is exogenous and observed. Measurement error is negligible. \\ \\ Assumption 3: Timing \\ Each period consists of two processes: First, newly informed individuals decide upon participation, second, information is exchanged. This implies that $Y$ is of dimension $N \times 4$ while $S$ is of dimension $N \times 3$: since information is exchanged after the participation decisions, modeling the fourth period's information exchange is redundant. \\ \\ Assumption 4: Information is a pre-condition for participation \\ \\ Assumption 5: Participation is a one-time opportunity \\ Each period only newly informed individuals face the participation decision. Having opted in (out), the respective individual will thereafter stay in the set of participants (non-participants) forever. \\ \\ \textbf{Assumption 6:} Distributional assumption \\ Conditional on being newly informed, the random variables $ Y_{it} \hspace{0,15cm} \forall i=1,...,N; \hspace{0,15cm} t=1,...,4$ are i.i.d. and follow a Bernoulli distribution. The probability to participate conditional on being newly informed is $p$. \\ \\ \textbf{Assumption 7:} Information is never forgotten \\ Individuals can only switch their information status once (from being uninformed ($S_{i(t-1)}=0$) to being informed ($S_{it}=1$)). Any informed individual remains potentially informing her neighbors every period. \\ \\ \textbf{Assumption 8:} Information exchange \\ In each period, informed individuals may transmit the information to their neighbors. On average, they do so with probability $q$. Transmitting is independent across all links. \\ Let the random variable $I_{(ij)t}$ denote the indicator that shows that individual $i$ has sent the information to individual $j$ in period $t$. Conditional on $i$ being a sender, the variables $I_{(ij)t} \hspace{0,1cm} \forall i\neq j; t=1,...,3 $ are i.i.d. and follow a Bernoulli distribution with $E[I_{(ij)t}]=q$. \\ \\ The aim is to estimate \\ (i) the individual's probability to participate if informed ($p$) and \\ (ii) individual's probability to share information with acquaintances ($q$)
Using the notation introduced above,
is the probability to observe the actual data at hand, given the (nonrandom) village network, the parameters and the information initiation and
is the village log likelihood function. \\ However, the probability to observe any particular outcome pattern depends on the (unobserved) information propagation.\footnote{Note that knowing the probability of (not) observing somebody as a participant depends on her chances to receive the information which in turn depends on which of her neighbors are informed} In the following, we use “information scenario" to denote a unique sequence $S_1=s_1,...,S_{(T-1)}=s_{(T-1)}$. The set $\mathbb{S}$ comprises all possible information scenarios. Then the log likelihood function is (by the law of total probability)
The joint probability of the outcome matrix and the information status matrix can, by iterated conditioning, conveniently be factorised into four time period specific factors, which in turn are a product of $N$ individual specific factors. Denoting by $l_t$ the likelihood up to and including period $t$ and by $\Delta \ell_t$ the change in the likelihood function as we move from period $t-1$ to period $t$, we can rewrite (ref) as
As we shall see, for any particular information scenario, the factors appearing in (ref) are straightforward to evaluate and thus the joint probability of the outcome and that particular scenario are easily computed. Applying the law of total probability then amounts to summing up these joint probabilities over all possible information scenarios. The challenge stems from the fact that the cardinality of $\mathbb{S}$ quickly becomes prohibitively large.
Individual $i$'s contribution to the village log likelihood function depends on any other individual only insofar as that $i$ may newly receive the information from her. Let
denote $i$'s probability to newly receive the information (from any of her neighbors) in period $t$. Using $r_{i(t-1)}$, we can categorize villagers at any point in time into four groups according to what they contribute to $\ell$ at this point in time (that is, according to $\Delta \ell_{it}$):
We can thus see that the PIIs are the origin of the dimensionality problem: without them, there would be only one information status matrix that could have generated the data. PIIs introduce a twofold challenge. First, their own information status cannot be derived from the data and hence their own contribution to the likelihood function is scenario specific. Second, whether or not they enter any particular period informed impacts all their neighbours' (and later on their neighbours' neighbours') information reception probabilities, implying that these individuals' likelihood contributions become scenario specific as well. As the cardinality of $\mathbb{S}$ increases exponentially in the number of PIIs, hence (ref) cannot be computed. \\ Due to the interdependence described above, likelihood contributions being scenario specific implies that $i's$ contribution depends on the previous information statuses of all her neighbours as well as on her own current status. \\ As claimed above, for any particular information scenario, the individual likelihood contributions can be calculated straightforwardly: first, the individual information reception probabilities $r_{it}$ are computed, then these are plugged into the formulas above, where for PIIs we choose formula A, B or C, depending on her own past and current status. Since $i$ is informed as long as {\it any} of her neighbors sends the information, we can use the counter-factual probability (nobody informs $i$) to calculate
From period two onward, different information status vectors from the last period lead to different PIIs and different probabilities of any individual to receive the information in the current period, thus $r_{it}$ needs to be computed for each $S_{(t-1)} \in \mathbb{S}$. \\ The origin of the dimensionality problem will also be the point of attack for our algorithm since a {\bf restriction of the number of PIIs} directly translates into a pronounced reduction of information scenarios. \\ For any PII who enters a time period uninformed, the scenarios A: “newly informed and opted out" and B: “uninformed" are observationally equivalent. The basic idea of our trimming procedure is to, for a certain number of PIIs, restrict the log likelihood function to include only one of these two cases. \\ We illustrate the dimension reduction strategy using two toy villages. The first village consists of six individuals one of which is the IP. Even in this very small and rather sparse network, three rounds of information exchanges result in 92 information propagation possibilities.
Consider a slightly larger and denser network.
What has changed from one graph to the other is a substantial increase in the number of intermediate PIIs. A large number of PIIs implies that the number of possible information propagation possibilities grows very large. What can happen in any later information exchange always depends on what has happened in previous ones. Consequently, for each information scenario, we would, in each intermediate time period, need to store the information who has been informed already in order to know who can be informed thereafter and to calculate the scenario specific individual information reception probabilities ($r_{it}$). If we wanted to evaluate the exact log likelihood function that would imply storing $2^4$ period one information status vectors, for each of them computing and storing all possible period two information status vectors, in order to finally know - for each realization of $S_{1:2}$ - the final (scenario specific) information reception probabilities of all agents. If the network is large and dense, it is our inability to store so many intermediate information status vectors and the time needed to process files that prevent us from establishing the log likelihood function. \\ If we have not eight, but roughly sixty to a hundred intermediate PIIs, then there would be not a billion or trillion, but a quintillion or nonillion of possible realizations of $S_{1:2}$. \\ To tackle the problem, we apply a trimming strategy. We will only allow a fixed number of $d$ PIIs to have two possible information status variables in the next period. For all other PIIs, the next period's information status will be set to a default (either one, in which case we assume the agent to be informed and opted out, or zero in which case we presume her to be uninformed). In particular, the individual will be trimmed to whichever of the two scenarios is more likely in the given period. If in each of the first two exchanges, the maximal number of PIIs is set to a trimming value of $d$, then we end up with a number of $s_{1:2}$ scenarios that cannot exceed $2^{2d}$. Storing information scenarios is only necessary to calculate the probabilities of what can follow. Since nothing follows after the third information exchange, a large number of final PIIs is absolutely un-problematic and there is no need to restrict the last wave of PIIs. \\ Accordingly, we need to select $d$ PIIs for which the two scenarios will be considered. Recapitulating that each PII can, in each exchange in which she can reached be either newly informed and uninterested or uninformed, we can see that there exists a threshold $r_{it}$ such that both are equally likely. \[ \underbrace{(1-p) r_i}_{ \substack{ \mbox{probability of scenario A:}\\ \mbox{$i$ is informed and opts out} } } = \underbrace{(1-r_i)}_{ \substack{ \mbox{probability of scenario B:}\\ \mbox{$i$ is uninformed} } } \]
The simple trimming rule identifies agents that are the furthest away from this threshold and trims their information status to either zero (if the agent's information reception probability falls below the threshold) or one (if it exceeds the threshold). This is repeated until only $d$ PIIs are left unrestricted. The basic idea of the proceeding is that if an agent is far away from the threshold, the probability of one scenario will be large while the other will be small and as such, neglecting the less likely scenario for these agents is the optimal choice from the current perspective (i.e.\ not taking into account what can happen thereafter) in the sense that the currently most likely scenarios are chosen. It may be the case, however, that a trimming choice reveals itself as sub-optimal in one of the following periods. Most dramatically, imagine a period-3 participant that is two links away from the nearest IP. The algorithm may set {\it} all neighbors of her neighbors to the default status “uninformed" in the first information exchange, which obviously has probability zero given that {\it someone} must have informed the period-3 participant. In the Appendix, we discuss under which circumstances such erroneous choices occur and how they depend on the network topology. \\ We demonstrate the trimming with the second graph above. Assume we are at a grid-point $p,q$ for which $2-\frac{1}{q}>p$ implying that scenario A is more likely for any PII and that we pick a $d$ value of 2. All PIIs featuring the same number of IP links in the first information exchange, the algorithm arbitrarily picks individual 2 and 3 to be trimmed. With 2 and 3 being newly informed and opting out, there are 4 scenarios resulting from the first exchange. Depending on which scenario was realized, there are four or five period-two PIIs. In any case, however, 6 and 7 will always be the ones with the largest number of links to the information and hence the algorithm trims them to scenario A. In the last exchange, there will be three or four PIIs. No individual will have more links to the information than 10 and as such, she will be subject to trimming. If there are four PIIs, then another PII to be trimmed will be chosen from the remaining three ones. \\ Assume next that the grid-point is such that $2-\frac{1}{(1-(1-q)^4)} < p$ implying that scenario B is more likely for any PIIs and that we again pick a $d$ value of 2. This time, 2 and 3 will be trimmed to being uninformed in the first exchange. As such, the second exchange features four to six PIIs. No PII can have less links to the information than 2 and 3 and thus they will be trimmed again and not receive the information. Whenever there are six PIIs, two additional PIIs that have one link to the information will be trimmed. The last exchange again trimms 2 and 3 and if necessary up to three other PIIs.\footnote{Trimming an individual to zero implies that she temporarily looses her right to receive the information or that the link is temporarily deactivated. She can, however, enter the PIIs in the next period. Whether or not she will be cut again depends on her and everybody else's probability to receive the information in that period. } \\ For each link portfolio an agent can have, there exist a specific $pq$ relationship such that the two scenarios are equally likely. Figure (ref) illustrates this with a link portfolio of one, two, three and four links.
The trimming strategy works better when agents are furthest away from the threshold rate. As a direct consequence, the best performance can be expected whenever either $p$ or $q$ is small and the respective other large. \\ Worth mentioning, note that the number of scenarios considered is endogenous. Each of the $2^d$ period-1 information status vectors entails a different number of scenarios that are subsequently possible as the umber of agents in reach of the information and the number of agents already informed vary. As a consequence, each trimming choice induces a different number of scenarios taken into account when establishing the log likelihood function. The rates at which trimming occurs are endogenous too. If the aim is to always trim to a certain number of PIIs, thus the lower and upper trimming rates vary on the grid.
The aim of the study is to evaluate the performance of the simple trimming technique as the trimming value ($d$) varies.
We use eleven network adjacency matrices from BCDJ. These were surveyed by asking villagers in India with whom they engage in regular activities. We here use the union of the activity-specific networks, thus assuming that any joint activity indicates the presence of a link. The alternative would be to simulate the network data as well. However, this would require us to specify a network formation model and to choose its parameters and these choices would impact the study results. \\ The eleven villages are large and dense. In the Monte Carlo study, we need to be able to evaluate the exact log likelihood function. To achieve tractability, we pick a sub-matrix consisting of $N$ individuals. By the same token, to limit the number of PIIs, only one individual will obtain the information initially. This individual is randomly drawn from the $N$ villagers. The sub-matrix used in each replication is different.\footnote{The algorithm starts at the element row seed S, column seed S situated on the diagonal and reads out a square 20 times 20 matrix} \\ Each repetition consists of simulation and estimation. While $N$, $p_0$, $q_0$ and $V$ are held fixed, each repetition distinguishes itself by its unique combination of the seeds used to initialize the random number generators. We consider the settings: case 1: $p_0=0.5, q_0=0.5$, case 2: $p_0=0.1, q_0=0.9$ and case 3: $p_0=0.9, q_0=0.1$ with $N=20$ and $V=11$. \\ \\ Simulation: The participation data is simulated according to the model assumptions using parameters $p_0$ (participation probability) and $q_0$ (information transmission probability) and the network sub-matrices. Each replication chooses a different sub-matrix out of the network matrix. Following the random draw of the initially informed individual and the random determination of her participation status, there are three rounds of information diffusion and participation. Each round, for each individual linked to the information, we draw a number from the uniform distribution in the interval $[0,1]$ and specify that the information is passed on if the number drawn is smaller or equal to $q_0$. Equivalently, for each newly informed individual, the number drawn from the standard uniform distribution needs to be smaller or equal to $p_0$ for her to participate. \\ This leads to a participation data matrix for each village, i.e.\ one sample. \\ \\ Estimation: The log likelihood function is established as the sum over information scenarios and each scenarios has a different number of potentially informed individuals (PIIs). Our trimming strategy sets a maximum to this number and sets PIIs exceeding this number to their more likely scenario in the first two exchanges (trimming is computationally unnecessary in the last exchange). In the simulation study, we wish to let the trimming value ($d$) vary and evaluate the consequences. Therefore, for each village, we find the maximal number of PIIs ($\bar{d}_v$). We then compute the approximate log likelihood function for $d=0,...,\bar{d}_v-1$ (denoted $\mathcal{L}_{v,d}$) and the correct log likelihood function, i.e.\ $d=\bar{d}_v$ which implies no trimming (denoted $\mathcal{L}_{v,\bar{d}_v }$). Each time, the number of PIIs is trimmed down to size $d$ in the first two information exchanges. Then the approximate sample log likelihood function for each trimming value is computed by aggregating all villages, i.e. \[ \mathcal{L}_d= \sum_{v=1}^V \mathcal{L}_{v,max(d, \bar{d}_v)} \] (i.e. if the trimming value exceeds the maximal number of PIIs in the village, the village is not subject to trimming). Changes of the trimming value when it is low hence impact practically all villages and in the villages practically all scenarios but as the trimming value grows towards its maximum, only few villages and scenarios are actually subject to trimming. Then the peak of the log likelihood function for each $d$ is found by grid search and confidence sets are established using the LR test. Each repetition hence delivers a sequence of estimates $\hat{p}_d,\hat{q}_d$ together with their confidence sets for $d=0,...,max(\bar{d}_v)$. We compare the estimates at different trimming values with the estimates that result from applying ML to a shortened time horizon of two periods.
For each case, the proceeding described above leads to ninety series of estimates. Remembering that trimming takes place whenever the number of PIIs exceed the desired trimming value, thus the lengths of this series is pinned down by the largest number of PIIs in any village and the last estimate of this series is the four-period MLE. Additionally, we dispose over the MLE that results from restricting the time horizon to two periods (i.e. the “Two-period Estimator" from chapter two).\\ In Figures (ref)-(ref), we plot the mean estimates and their empirical standard errors, the estimated density and the histograms of the empirical probability mass function. \\ For each case the first two plots depicts the mean estimate ($\Bar{\hat{p}}$ and $\Bar{\hat{q}}$) as a function of the trimming value, the mean being taken over the 90 estimates (Figures (ref), (ref) and (ref)). The vertical bars have the length of the empirical standard error in each direction, the latter being simply computed as the square root of average squared deviation of the individual estimate from the mean estimate. \[ \left( \frac{1}{89} \sum_{r=1}^{90} (\hat{p}_r- \Bar{\hat{p}})^2 \right)^{1/2} \hspace{2cm} \left( \frac{1}{89} \sum_{r=1}^{90} (\hat{q}_r- \Bar{\hat{q}})^2 \right)^{1/2} \] The trimmed estimates converge to the MLE rather quickly. However, it needs to be kept in mind that we have restricted the village sizes to twenty and as a result, a substantial fraction of villages simply feature very few PIIs, in particular when participation rates are not too small. Then, these villages are already at the MLE, even for minor trimming values. The results are nonetheless encouraging in the sense that trimming substantially increases computational speed and does not hurt much (compared to the MLE) as long as the sample includes enough villages for which no more than one third of the PIIs are trimmed. The efficiency gains compared to the “Two-period Estimator" are substantial, in particular for the diffusion rate. \\ Thereafter, for each case, we depict depict the kernel density estimates together with the empirical quantiles (first and third)(Figures (ref), (ref) and (ref)). The density of the “Two-period Estimator", depicted in orange, is substantially flatter and more often exhibit multiple local maxima. \\ The histograms in the third plot series (Figures (ref), (ref), (ref) and (ref), (ref), (ref)) reconfirm that the “Two-period estimates" are much more dispersed, especially for the diffusion rates. As the Trimming estimates converge, the histograms for larger trimming values become indistinguishable from the MLE. \\ \\ The following case-specific observations can be made. \\ {\bf Case 1:} \\ The densities for “Trimming estimates" exhibit a distinct peak in the probability mass function, which is not the case for the “Two-period estimates". Observe that for the latter, the estimated diffusion rate for a substantial fraction of the sample is at 0.99. Also, the estimates are much less dispersed for the “Trimming-Estimator" than the “Two-period Estimator". \\ {\bf Case 2:} \\ The gains from increasing the trimming value are made exclusively for the diffusion rate estimates, while the participation rate estimates are very similar. \\ The standard deviations are smaller and the log likelihood surfaces are more peaked for small trimming values. \\ {\bf Case 3:} \\ Both $p$ and $q$ are well identified whenever adoption rate is high (except for the extreme case of $d=0$, which would imply trimming all PIIs). Unsurprisingly, in this case also the performance difference as compared to the “Two-Period estimator" is smallest.
This section presents the results of an application of the “Trimming-estimator" to four selected real and full-size villages. This estimation was run on the powerful Amazon-Web-Service server cluster and I am grateful that funds have been made available for this estimation, which also provided me with the opportunity to gain valuable experience with respect to the usage of this server cluster.\\ The time horizon is set to 3, the chosen villages are 1, 12, 31 and 67. The choice was guided by the fact that in particular these four villages exhibit a steady increase in the observed real adoption rates, whereas for many other villages there is a distinct jump. With jumps being potentially associated to measurement error, I hope to have chosen the villages with the most informative observed adoption patterns. These villages have, respectively, 182, 175, 153 and 193 inhabitants. \\ Due to the computational cost, it was infeasible to calculate the objective function over the entire grid. However, this was also not necessary. Extensive testing indicated that the convergence was relatively slow (approximately 0.01 when the trimming value increases by two for $q$ and practically nonexistent for $p$). A relatively low trimming value of eighteen was chosen to define the subgrid to be evaluated. The subgrid contains 48 point with the diffusion rate varying from 0.58 to 0.73 and the adoption rate from 0.23 to 0.25. This range was chosen according to a prognostic calculation such that the resulting grid would cover approximately most of the five present confidence set of the final largest feasible trimming value, which was 32. \\ For a trimming estimation with three time periods that starts from a trimming value that is low compared to the total number of villagers, the computation time roughly doubles when the trimming value is increased by one. Parallelization was done over both villages and gridpoints. The final computation times were 651.2 hours (village 1), 609.5 hours (village 12), 480.7 hours (village 31) and 751.8 hours (village 67).
The fact that the estimate is situated at the border of the grid is not problematic given that testing indicated that convergence of $p$ was achieved much faster than convergence of $q$ and as such, not much change can be expected in the adoption parameter when the trimming value further increases. The confidence sets are enlarging when the trimming value increases. This is intuitive: the subgrid around the peak is situated in an area in which all agents are trimmed to being “informed and opted out” (the small $p$, high $q$ area). A change in $q$ will thus have the same effect on all trimmed individuals leading to a more pronounced peak, while in the full likelihood, the impact of a change in $q$ would be different for different scenarios, thus resulting overall in a flatter log likelihood function. This hints towards the finding that indeed the trimming estimate may, in cases of relatively low trimming values, lead to an over-pronounced peak such that it may be necessary to adjust critical values accordingly. This same finding is also in the contour and surface plots. \\ The “Two-period Estimate" estimate is outside the interval considered. Given the test results, it is expected that the full MLE with three time periods will converge to a value that is somewhere in between the trimming and the “Two-period estimate". \\ Estimation using only four villages is highly challenging as we can deduct from the sizes of the “Two-period estimate"’s confidence sets. Note that when testing on the basis of the “Two-period estimate", the hypothesis that it equals the trimming estimate cannot be rejected even at 30%. \\ Since practically all points in the sub-grid are in the 5% confidence set, hence I compute the confidence sets for higher confidence values. As compared to the “Two-period estimate", the “Trimming estimate"’s confidence set is extremely narrow. This is in line with the observation that the confidence sets widen when the trimming value is increased.
The “Trimming estimator" provides an attractive alternative when exact MLE estimation is infeasible. Substantial increases in computational speed can be achieved whenever the sample comprises a sufficiently large fraction of villages for which no more than one third of the PIIs are trimmed. The efficiency gains are substantial, in particular for the diffusion rate estimate. This is in particular the case when the adoption rate is low.