Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
41,893 characters · 7 sections · 31 citation commands
Difference-in-Differences with Geocoded Microdata
The rise of microdata with precisely geocoded locations has allowed researchers to begin answering questions about the effects of spatially-targeted treatments at a very granular level. What are the effects of local pollutants on child health?\oldfootnote{\ See, e.g., Currie_Davis_Greenstone_Walker_2015 and Marcus_2021.} Does living within walking distance to a new bus stop improve labor market outcomes?\oldfootnote{\ See, e.g., Gibbons_Machin_2005 and Billings_2011.} How far do neighborhood shocks, such as foreclosures or new construction spread?\oldfootnote{\ See, e.g., Asquith_Mast_Reed_2021,Cui_Walsh_2015,Gerardi_Rosenblatt_Willen_Yao_2015 and Campbell_Giglio_Pathak_2011.} When treatment is located at a specific point in space, a standard method of evaluating the effects of the treatment is to compare units that are close to treatment to those slightly further away -- what I will label the `ring method'. This paper formalizes the assumptions required for identification in the ring method, highlighting potential pitfalls of the currently used estimator, and proposes an improved estimator which relaxes these assumptions.
The ring method is illustrated in (ref). The center of the figure is marked with a triangle which represents the location of treatment, e.g. a foreclosed home. Units within the inner circle, marked by dots, are considered treated due to their proximity to the treatment location; units between the inner and outer circles, marked in triangles, are considered control units; and then the remaining units are removed from the sample. The appeal of this identification strategy is that since the treated and control units are all very close in physical location, e.g. having access to the same labor market and consumptive amenities, the counterfactual untreated outcomes will approximately be equal for units within each ring. The ring estimate for the treatment effect compares average changes in outcomes between units in the inner `treated' ring and the outer `control' ring to form an estimate for the treatment effect, i.e. a difference-in-differences estimator.
My first contribution is to fill a gap in the econometrics literature by formalizing the necessary assumptions for unbiased estimates of the average treatment effect on the affected units.\oldfootnote{\ This generalizes the treatment effect on the treated in the case where treatment isn't assigned to specific units.} The first assumption is the well understood parallel trends assumption for the treated and control units. This requires that the average change in (counterfactual) untreated outcomes in the treated ring is equal to the average change in the control ring. This allows the control units to estimate the counterfactual trend for the treated units.
The second assumption requires the researcher to correctly identify how far treatment effects are experienced (the inner ring). This is a very strict assumption that when not satisfied, causes biased estimates of the treatment effect. If the treated ring is too narrow, then units in the control ring experience effects of treatment and the change among `control' units would no longer identify the counterfactual trend. On the other hand, if the treated ring is too wide, then the zero treatment effect of some unaffected units are averaged into the change among `treated' units. Therefore, results will be attenuated towards zero.
Since researchers often do not know how far treatment effects extend in most circumstances, I propose an estimator that replaces the second assumption with a less strict assumption by using a nonparametric, partitioning-based, least square estimator Cattaneo_Crump_Farrell_Feng_2019,Cattaneo_Farrell_Feng_2019. My proposed methodology estimates the treatment effect curve as a function of distance by using many rings rather than trying to estimate the average treatment effect with one inner ring. This method requires that treatment effects become zero somewhere between the distance of 0 and the control ring without the need to specify the exact distance. It however requires a stronger assumption that the counterfactual trend is constant across distance.\oldfootnote{\ Note that the original requirement is that the average change in each ring is equal but allows variation across distance.} This new assumption is more strict in that the standard method only requires that parallel trends holds on average in each ring. However, researchers motivate the identification strategy by saying within a small distance from treatment that units are subject to a common set of shocks which implies the more strict assumption. While this assumption is not directly testable, the estimator creates a set of point estimates of treatment effects that can be used to visually inspect the plausability of the assumption. If after some distance, treatment effects become centered at zero, this suggests that common trends hold, akin to the pre-trends test in event study regressions.
The nonparametric approach allows the researcher to get a more complete picture of how the intervention affects units at various distances rather than estimating an “overall effect”. For example, the construction of a new bus-stop potentially creates net costs to immediate neighbors while providing net benefits for homes slightly further away. Estimation of the treatment effect curve can illustrate these different effects that the “overall effect” would mask. In this case, the average effect could be zero even though most units experience non-zero effects.
This paper relates to a few papers that address difficulties with using the rings method for causal effect estimation. In the online appendix, Gerardi_Rosenblatt_Willen_Yao_2015 discuss the problem that if the treated ring is defined too narrowly, then control units will be affected by treatment causing a biased estimate of the counterfactual trend. Sullivan_2017 discusses the problem more formally and derives that the bias will be the difference in treatment effects experienced by the `treated' ring and the `control' ring. My paper expands on the results of Sullivan_2017 by including the additional source of bias that can result from a violation of parallel trends. Other researchers have recognized that estimating a single average treatment effect is less informative than a treatment effect curve. They solve this by using multiple rings to estimate treatment effects at different distances (e.g. Alexander_Currie_Schnell_2019,Casey_Schiman_Wachala_2018,Di_Tella_Schargrodsky_2004). However, this approach selects multiple rings in an ad-hoc manner, still requires treatment effects to become zero after the outer-most treatment ring, and is prone to problems of specification searching. The current study's proposed estimator selects the number and location of rings in a data-driven way and does not require correct specification of where treatment effects become zero.
Diamond_McQuade_2019 propose a nonparametric estimator aimed at estimating a treatment effect surface. They use two-dimensions (latitude/longitude) to better approximate a smooth change in counterfactual outcomes (e.g. north-west and south-east from treatment might have different treatment effects). My method, instead, uses a singular measure of distance which pools units at similar distances but different directions from treatment and therefore delivers more precise estimates. However, the treatment effect estimate may mask heterogeneity of effects at different directions. If a researcher has a reason to suspect significant heterogeneity, then their proposed estimator will make a better fit.
This paper also contributes to a small literature on difference-in-differences estimators from a spatial lens Butts_2021,Clarke_2017,Berg_Streitz_2019,Verbitsky-Savitz_Raudenbush_2012,Delgado_Florax_2015. These papers address instances where treatment is well defined by administrative boundaries but spillovers cause problems of defining who is `treated' and at what level of exposure. Butts_2021 and Clarke_2017 both recommend a method of using many rings to estimate treatment effects similar to the proposed nonparametric estimator. Clarke_2017 does specify a cross-validation approach for selecting rings, but does not specify that this result requires common local trends. Since this paper focuses on local shocks where constant parallel trends are plausible, I am able to provide a data-driven approach to choosing rings.
Last, there is a growing literature around design-based estimation of treatment effects in the presence of spillover effects Sävje_Aronow_Hudgens_2019,Aronow_Eckles_Samii_Zonszein_2020. Aronow_Samii_Wang_2021 specifically discuss estimation of what this paper calls the treatment effect curve, or the average treatment effect at a certain distance away from treatment. This paper compliments this literature by introducing model-based assumptions for cases where treatment is not assigned following an experimental design.
To illustrate the methodological difficulties in this method, I present an illustriative example. Suppose that an overgrown empty lot in a high-poverty neighborhood is cleaned up by the city and the outcome of interest is home prices. The researcher observes a panel of home sales before and after the lot is cleaned. Cleaning up the lot causes home values to go up directly nearby and as you move away from the lot, the positive treatment effect will decay to zero effect at, say, 3/4 of a mile. Since treatment is targeted to the high poverty neighborhood, comparisons with other neighborhoods in the cities could be biased if the neighborhood home prices are on different trends. Hence, the researcher wants to look only at the homes in the immediate neighborhood.
(ref) shows a plot of simulated data from this example. The black line is treatment effect at different distances from the empty lot and the grey line is the underlying (constant) counterfactual change in home prices, normalized to 0. Panel (a) of (ref) shows the best-case scenario where the treated ring is correctly specified. The two horizontal lines show the average change in outcome in the treated ring and the control ring. The treatment effect estimate, $\hat{\tau}$, is the difference between these two averages. However, this singular number masks over a large amount of treatment effect heterogeneity with units very close to treatment having a treatment effect double that of $\hat{\tau}$ and units near 3/4 miles experience a treatment efffect half as large as $\hat{\tau}$. For this reason, even if a researcher identifies the correct average treatment effect, they are masking a lot of heterogeneity that is potentially interesting. Therefore, later in this paper I recommend nonparametrically estimating the treatment effect curve as a function of distance rather than using average effect.
However, the researcher does not typically know the distance at which treatment effects stop. Panels (b) and (c) highlights how treatment effect estimates change with a change in ring distances. Panel (b) shows when the `treatment' ring is too wide. In this case, some of the units in the treatment ring receive no effect from treatment and therefore makes the average treament effect among units in the treatment ring smaller. Therefore when the treated ring is too large, the estimated treatment effect is too small. Panel (c) of (ref) shows the opposite case, where the treated ring is too narrow. In this case, there are some units in the `control' ring that experience treatment effects. Hence, the average change in outcome among the control unit is too large. This does not, though, decrease the treatment effect as one may suspect. Since the treatment effect decays with distance, the average change in outcome among the more narrow `treatment' ring is larger than the correct specification. The estimated treatment effect in this case grows, but it is not clear more generally whether the treatment effect will increase or decrease.\oldfootnote{\ This primarily depends on the curvature of the treatment effect curve.} From these three examples, it's clear that the estimation strategy requires researchers to know the exact distance at which treatment effects become zero. Since this is a very demanding assumption, I propose an improved estimator in (ref) that relaxes this assumption.
Often times, researchers try multiple sets of rings and if the estimated effect remains similar across specifications, they assume the results are `robust'. Panel (d) of (ref) shows an example of why this a problem. If Panel (c) was the researchers' original specification and Panel (d) was run as a robustness check, then the researcher would be quite confident in their results even though the estimate is too large in both cases. Now, I turn to econometric theory in order to formalize the intuition developed in this section.
Now, I develop econometric theory to formalize the intuition developed in the previous section. A researcher observes panel data of a random sample of units $i$ at times $t = 0, 1$ located in space at point $\theta_i = (x_i, y_i)$. Treatment occurs at a location $\bar{\theta} = (\bar{x}, \bar{y})$ between periods. Therefore, units differ in their distance to treatment, defined by $\text{Dist}_i \equiv d(\theta_i, \bar{\theta})$ for some distance metric $d$ (e.g. Euclidean distance) with a distribution function $F$. Outcomes are given by
where $\mu_i$ is unit-specific time-invariant factors, $\lambda_i$ is the change in outcomes due to non-treatment shocks in period 1, $\tau_i$ is unit $i$'s treatment effect. Both $\lambda$ and $\tau$ can be split into a systematic function of distance $z(\text{Dist}_i)$ and an idiosyncratic term $\tilde{z}_i \equiv z_i - z(\text{Dist}_i)$ with $z$ being $\tau$ and $\lambda$. $\tau(d)$ is the average effect of treatment at a given distance and $\lambda(d)$ summarizes how covariates and shocks change over distance. Therefore, we could rewrite our model as
where $\varepsilon = u_{it} + \tilde{\tau}_i + \tilde{\lambda}_i$ which is uncorrelated with distance to treatment. Researchers are trying to identify the average treatment effect on units experiencing treatment effects, i.e. $\bar{\tau} = \mathbb{E}\left[\tau_i \ \vert \ \tau(\text{Dist}_i) > 0\right]$.
Taking first-differences of our model, we have $\Delta Y_{it} = \tau(\text{Dist}_i) + \lambda(\text{Dist}_i) + \Delta \varepsilon_{it}$. It is clear that $\tau(\text{Dist}_i)$ and $\lambda(\text{Dist}_i)$ are not seperately identified unless additional assumptions are imposed. The central identifying assumption that researchers claim when using the ring method is that counterfactual trends likely evolve smoothly over distance, so that $\lambda(\text{Dist}_i)$ is approximately constant within a small distance from treatment. This is formalized in the context of our outcome model by the following assumption.
This assumption requires that, in the absence of treatment, outcomes would evolve the same at every distance from treatment within a certain maximum distance, $\bar{d}$. To clarify the assumption, it is helpful to think of ways that it can fail. First, if treatment location is targeted based on trends within a small-area/neighborhood, then trends would not be constant within the control ring. Second, if units sort either towards or away from treatment in a way that is systematically correlated with the outcome variable, then the compositional change can cause a violation in trends over time. Note that \nameref{assum:parallel} implies the standard assumption that parallel trends holds on average between the treated and control rings:
If \nameref{assum:parallel} holds for some $d_c$, then our first-difference equation can be simplified to $\Delta Y_{it} = \tau(\text{Dist}_i) + \lambda + \Delta \varepsilon_{it}$ where $\lambda$ is some constant for units in the subsample $\mathcal{D} \equiv \{i \ : \ \text{Dist}_i \leq d_c \} $. Therefore, the treatment effect curve $\tau(\text{Dist}_i)$ is identifiable up to a constant under Assumption (ref). To identify $\tau(\text{Dist}_i)$ seperately from the constant, researchers will often claim that treatment effects stop occuring before some distance $d_t < d_c$. This is formalized in the following assumption.
With this assumption, the first difference equation simplifies to $\Delta Y_{it} = \lambda + \Delta \varepsilon_{it}$ for units with $d_t < \text{Dist}_i < d_c$. These units therefore identify $\lambda$. The `ring method' is the following procedure. Researchers select a pair of distances $d_t < d_c$ which define the “treated” and “control” groups. These groups are defined by $\mathcal{D}_t \equiv \{ i : 0 \leq \text{Dist}_i \leq d_t \}$ and $\mathcal{D}_c \equiv \{ i : d_t < \text{Dist}_i \leq d_c \}$. On the subsample of observations defined by $\mathcal{D} \equiv \mathcal{D}_t \cup \mathcal{D}_c$, they estimate the following regression:
From standard results for regressions involving only indicators, $\hat{\beta}_1$ is the difference-in-differences estimator with the following expectation: \[ \expec{\hat{\beta}_1} = \mathbb{E}\left[\Delta Y_{it} \ \vert \ \mathcal{D}_t\right] - \mathbb{E}\left[\Delta Y_{it} \ \vert \ \mathcal{D}_c\right]. \] This estimate is decomposed in the following proposition.\oldfootnote{\ A similar derivation of part (i) is found in Sullivan_2017 but does not include difference in parallel trends.}
Part (i) of this proposition shows that the estimate is the sum of two differences. The first difference is the difference in average treatment effect among units in the treated ring and units in the control ring. The second difference is the difference in counterfactual trends between the treated and control rings. This presents two possible problems. If some units in the control group experience effects from treatment, the average of these effects will be subtracted from the estimate. Second, since treatment can be targeted, the treated ring could be on a different trend than units further away and hence control units do not serve as a good counterfactual for treated units.
Part (ii) says that if $d_c$ satisfies \nameref{assum:parallel}, then the difference in trends from part (i) is equal to 0. As discussed above, the decomposition in part (ii) of Proposition (ref) is not necessarily unbiased estimate for $\bar{\tau}$. First, if $d_t$ is too wide, then $\mathcal{D}_t$ contain units that are not affected by treatment. In this case, $\hat{\beta}_1$ will be biased towards zero from the inclusion of unaffected units from $d_t$ being too wide. Second, if $d_t$ is too narrow then the $\mathcal{D}_c$ will contain units that experience treatment effects. It is not clear in this case, though, whether $\hat{\beta}_1$ will grow or shrink without knowledge of the $\tau(\text{Dist})$ curve, but typically $\hat{\beta}_1$ will not be an unbiased estimate for $\bar{\tau}$. See the previous section for an example.
Part (iii) of Proposition (ref) shows that if $d_t$ is correctly specified as the maximum distance that receives treatment effect, then $\hat{\beta}_1$ will be an unbiased estimate for the average treatment effect among the units affected by treatment. However, Assumption (ref) is a very demanding assumption and unlikely to be known by the researcher unless there are a priori theory dictating $d_t$.\oldfootnote{\ As an example, Currie_Davis_Greenstone_Walker_2015 uses results from scientific research on the maximum spread of local pollutants and Marcus_2021 use the plume length of petroleum smoke.} The following section will improve estimation by allowing consistent nonparametric estimation of the entire $\tau(\text{Dist})$ function. An estimate of $\tau(\text{Dist})$ can then be numerically integrated to for an estimate of $\bar{\tau}$.
In this section, I propose an estimation strategy that nonparametrically identifies the treatment effect curve $\tau(\text{Dist}_i)$ using partitioning-based least squares estimation and inference methods developed in Cattaneo_Crump_Farrell_Feng_2019, Cattaneo_Farrell_Feng_2019. Partition-based estimators seperate the support of a covariate, $\text{Dist}_i$, into a set of quantile-spaced intervals (e.g. 0-25th percentiles of $\text{Dist}_i$, 25-50th, 50-75th, and 75-100th). Then the conditional $\mathbb{E}\left[Y_i \ \vert \ \text{Dist}_i\right]$ is estimated seperately within each interval as a $k$-degree polynomial of the covariate $x_i$.
For a given $d_c$, we will form a partition of our sample $\mathcal{D} = \{ i : \text{Dist}_i \leq d_c \}$ into $L$ intervals based on quantiles of the distance variable. Denote a given quantile as $\mathcal{D}_j \equiv\{ i : F_n^{-1}(\frac{j-1}{L}) \leq \text{Dist}_i < F_n^{-1}(\frac{j}{L}) \}$ where $F_n$ is the empirical distribution of $\text{Dist}$. Let $\{ \mathcal{D}_1, \dots, \mathcal{D}_L \}$ be the collection of the $L$ intervals. This paper will impose $k = 0$ which will predict $\Delta Y_{it}$ with a constant within each interval.\oldfootnote{\ Approximation can be made arbitrarily close to the true conditional expectation function by either increasing the number of intervals or by increasing the polynomial order to infinity, so setting $k = 0$ does not impose any cost.}
These averages are defined as \[ \overline{\Delta Y}_j \equiv \frac{1}{n_j} \sum_{i \in \mathcal{D}_j} \Delta Y_{it}, \] where the number of units in bin $\mathcal{D}_j$ is $n_j \approx n/L$. Our estimator for $\mathbb{E}\left[\Delta Y_{it} \ \vert \ \text{Dist}_i\right]$ is then given by \[ \widehat{\Delta Y_{it}} = \sum_{j = 1}^{L} \mathbf{1}_{i \in \mathcal{D}_j} \overline{\Delta Y}_j \] As the number of intervals approach infinity, this estimate will approach $\mathbb{E}\left[\Delta Y_{it} \ \vert \ \text{Dist} = d\right]$ in a mean-squared error sense. Under \nameref{assum:parallel}, $\mathbb{E}\left[\Delta Y_{it} \ \vert \ \text{Dist} = d\right] \equiv \mathbb{E}\left[\tau(\text{Dist}) \ \vert \ \text{Dist} = d\right] + \lambda$. To remove $\lambda$, we require a less-strict version of assumption (ref).
If a distance $d_c$ satisfies \nameref{assum:parallel} and ((ref)), the mean within the last ring $\mathcal{D}_k$ will estimate $\lambda$ as the number of bins $L \to \infty$. The reason for this is simple, as $L \to \infty$, the last bin will have the left end-point $> d_t$ and therefore $\tau(\text{Dist}) = 0$ in $\mathcal{D}_L$. Under local parallel trends, the last ring will therefore estimate $\lambda$. Therefore, estimates of $\tau(\text{Dist}_i)$ can be formed for each interval as $\hat{\tau}_j \equiv \overline{\Delta Y}_j - \overline{\Delta Y}_L$.
As discussed in Section (ref), specifying $d_t$ correctly is important to identify the average treatment effect among the affected in the parametric estimator. The nonparametric estimator only requires that treatment effects become zero before $d_c$, i.e. that such a $d_t$ exists. However, the estimator would no longer identify the treatment effect curve under the milder \nameref{assum:parallel_weak} assumption. Therefore, a researcher should justify explicity the assumption that, within the $d_c$ ring, every unit is subject to the same trend. This is most likely to be satisfied on a very local level and not very plausible in the case of larger units, e.g. counties.
The nonparametric approach allows estimation of the treatment effect curve whereas the indicator approach, at best, can only estimate an average effect among units experiencing effects. The treatment effect curve allows researcher to understand differences in treatment effect across distance. For example, typically one would assume treatment effects shrink over distance and evidence of this from the nonparametric approach can strengthen a causal claim. In some cases, such as a negative hyper-local shock and a postivie local shock (e.g. a local bus-stop), the treatment effect can even change sign across distances. In this case, the average effect could be near zero even though there are significant effects occuring.
Plotting estimates $\hat{\tau}_j$ can provide visual evidence for the underlying \nameref{assum:parallel} assumption. Typically, treatment effect will stop being experienced far enough away from $d_c$ that some estimates of $\hat{\tau}_j$ with $j$ `close to' $L$ will provide informal tests for parallel trends holding. (ref) provide an example where plotting of $\hat{\tau}_j$ provide strong evidence in support of local parallel trends as it appears that after some distance, average effects are consistetly centered around zero. This is not a formal test as it could be the case that the true treatment effect curve, $\tau(\text{Dist})$ is perfectly cancelling out with the counterfactual trends curve $\lambda(\text{Dist})$ producing near zero estimates, but this is a knive's edge case.
The above proposition shows that the series estimator will consistenly estimate the treatment effect curve, $\tau(\text{Dist})$ as the number of bins $L$ and the number of observations $n$ both go to infinity. In finite-samples though, we will have a fixed $L$ and hence a fixed set of treatment effect estimates $\{ \tau_1, \dots, \tau_{L}\}$ with $\tau_L \equiv 0$ by definition. The estimates $\hat{\tau}_j$ are approximately equal to $\mathbb{E}\left[\tau(\text{Dist}) \ \vert \ \text{Dist} \in \mathcal{D}_j\right]$ or the average treatment effect within the interval $\mathcal{D}_j$.
The choice of $L$ in finite samples is not entirely clear. Cattaneo_Crump_Farrell_Feng_2019 derive the IMSE-optimal choice of $L$ which is a completely data-driven choice. The optimal $L$ is driven by two competing terms in the IMSE formula. On the one hand, as $L$ increases, the conditional expectation function is allowed to vary more across values of $\text{Dist}$ and hence bias of the estimator decreases. However, larger values of $L$ increase the variance of the estimator. Balancing this trade-off depends on the shape and curvature of $\tau(\text{Dist})$. The resulting choice of $L^*$ and the use of quantiles of the data allows a completely data-driven choice of the number of rings and their endpoints which allows for estimation in a principled and objective way. This principled estimator removes researcher-incentives to search across choices of rings to provide the best evidence.
For a given $L^*$, Cattaneo_Crump_Farrell_Feng_2019 show the large-sample asymptotics of the estimates $\overline{\Delta Y}_j$ and provide robust standard errors for the conditional means that account for the additional randomness due to quantile estimation. Since our estimator is a difference in means, standard errors on our estimate $\hat{\tau}_j$ are given by $\sqrt{\sigma^2_j + \sigma^2_L}$, where $\sigma_j$ is the standard error recommended by Cattaneo_Crump_Farrell_Feng_2019. These standard errors are produced by the Stata/R package binsreg. Inference can be done by using the estimated t-stat with the standard normal distribution. There may be concerned that the standard errors need to adjust for spatial correlation. However, this is not the case under assumption ((ref)) as this implies the error term is uncorrelated with distance.
To highlight the advantages of my proposed estimator, I revisit the analysis of Linden_Rockoff_2008. This paper analyzes the effect of a sex offender moving to a neighborhood on home prices. This paper uses the ring method with treated homes being defined as being within $1/10^{th}$ of the sex offender's home and the control units being between $1/10^{th}$ and $1/3^{rd}$ of a mile from the home. The authors make a case for the ring method by arguing that within a neighborhood, \nameref{assum:parallel} holds since they are looking in such a narrow area and purchasing a home is difficult to be precisely located with concurrent hyper-local shocks.
As for the choice of the treatment ring, there is little a priori reasons to know how far the effects of sex offender arrival will extend in the neighborhood. The authors provide graphical evidence of nonparametric estimates of the conditional mean home price at different distances in the year before and the year after the arrival of a sex offender. The published plot can be seen in Panel (b) of (ref). They `eyeball' the point at which the two estimates are approximately equal to decide how far treatment effects extend. However, this approach is less precise than it may seem. Panels (a) and (c) show that changing the bandwidth for the kernel density estimator will produce very different guesses at how far treatment effects extend. My proposed estimator works in a data-driven way that does not require these ad-hoc decisions.
The standard rings approach is equivalent to my proposed method with two rings: $\mathcal{D}_1$ being the treated homes between 0 and 0.1 miles away and $\mathcal{D}_2$ being the control homes between 0.1 and 0.3 miles away. The average change among $\mathcal{D}_2$ estimates the counterfactual trend and the average change among $\mathcal{D}_1$ minus the estimated counterfactual trend serves as the treatment effect. Panel (a) of (ref) shows the basic results of their difference-in-differences analysis which plots estimates $\hat{\tau}_j$ for $j = 1,2$. On average, homes between 0 and 0.1 miles decline in value by about 7.5% after the arrival of a sex offender. As an assumption of the rings method, homes between 0.1 and 0.3 miles away are not affected by a sex offender arrival. The choice of 0.1 miles is an untestable assumption and as seen above the evidence provided is highly dependent on the choice of bandwidth parameter. My proposed estimator does not require a specific choice for a `treated' area.
Linden_Rockoff_2008 only have access to a non-panel sample of home sales, so identification requires another assumption for identification, namely that the composition of homes at a given distance does not change over time. Further, since we can no longer form first-differences of the outcome variable, seperate nonparametric estimators must be estimated before and after treatment and subtracted from one another. Details of this theory are in the Online Appendix
Panel (b) of (ref) applies the nonparametric approach described in Section (ref). Two differences in results occur. First, homes in the two closest rings i.e. within a few hundred feet, are most affected by sex-offender arrival with an estimated decline of home value of around 20%. homes a bit further away but still within in Linden and Rockoff's `treated' sample do not experience statistically significant treatment effects. As discussed above, Linden and Rockoff's estimate of $\bar{\tau}$ is attenuated towards zero because of the inclusion of homes with little to no treatment effects, leading them to understate the effect of arrival on home prices. The nonparametric approach improves on answering this question by providing a more complete picture of the treatment effect curve. The magnitude of treatment effects decrease over distance, providing additional evidence that the arrival causes a drop in home prices.\oldfootnote{\ This is similar to estimating a dose-response function as evidence supporting a causal mechanism.}
The second advantage of this approach is that the produced figure provides an informal test of the local parallel trends assumption. After 0.1 miles, the estimated treatment effect curve becomes centered at zero consistently. This implies that units within each ring have the same estimated trend as the outer most ring, providing suggestive evidence that homes in this neighborhood are subject to the same trends.
This article formalizes a common applied identification strategy that has a strong intuitive appeal. When treatment effects of shocks are experienced in only part of an area that would otherwise be on a common neighborhood-trend, difference-in-differences comparisons within a neighborhood can identify treatment effects. However, this paper shows that the typical estimator for treatment effects requires a very strong assumption and returns only an average treatment effect among affected units when this assumption holds.
This article then proposes an improved estimator that relies on nonparametric series estimators. The nonparametric estimator allows for estimation of the treatment effect at different distances from treatment, similar to a dose-response function, which can allow better understanding of who is experiencing effects and how this changes across `exposure' to a shock. More, in some cases it can provide explanation for null results. For example, if a bus station creates negative externalities for apartments that border the station but positive externalities for apartments within walking distance, the average effect could be zero. However, nonparametric estimation would reveal the two effects seperately.