Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
113,883 characters · 32 sections · 91 citation commands
Causal Inference in Financial Event Studies
\pagenumbering{Alph}
\pagenumbering{arabic} \setstretch{1.25}
Financial economists were practicing causal inference well before the credibility revolution angrist2010credibility. By examining how asset prices respond to information events---such as merger announcements, earnings releases, or regulatory changes---financial event studies compare the returns of treated assets to benchmark comparison asset returns. The approach remains central: between 2010 and 2025, 305 articles in the Journal of Finance and the Review of Financial Studies reference event‐study methods ((ref)).
While financial event studies target causal effects as their estimands, the suite of estimators used in financial event studies are antiquated relative to the many tools available. The textbook approach, starting as early as fama1969adjustment and canonized in the campbell1997econometrics textbook, relies on linear factor models with known factors to construct counterfactual returns, i.e. what a security's return would have been absent the event. Researchers typically estimate a security's exposure to market factors during a pre-event window, then use these estimated loadings to predict what returns during the event window. The difference between actual and predicted returns constitutes the abnormal return.
This paper demonstrates that this standard approach faces a fundamental identification challenge. We show analytically that when the factor model is misspecified---which is almost certainly the case given the ongoing debates about the appropriate asset pricing model---abnormal return estimators are generally inconsistent estimators for causal effects. The problem is particularly severe in three empirically relevant scenarios. First, when events occur during periods of extreme market volatility, even small misspecification in factor loadings gets amplified by large factor realizations, potentially generating economically significant bias. Second, in long-horizon event studies that examine returns over months or years, misspecification bias accumulates over time, making the resulting estimates potentially more reflective of model error than treatment effects. Third, when event timing coincides with particular market conditions—for instance, if mergers cluster during market downturns—the standard estimators conflate selection effects with treatment effects.
We provide precise conditions under which traditional event study methods identify causal effects. Identification of the average treatment effect on the treated requires either correct specification of the factor model (unlikely given decades of asset pricing research showing the difficulty of this task), or random assignment of treatment across securities. When these conditions fail, we derive expressions for the asymptotic bias that provide guidance on when concerns should be most acute.
Our results stand in contrast to folk wisdom that the structure of the factor model does not have significant impacts on the size of the estimated effects. For example, in footnote 5, shleifer1986demand states “The [index inclusion] results were not materially different when returns were not corrected for market movements.” We show that this irrelevance is due to two key features: (1) random timing of many events over time (in the shleifer1986demand case, index inclusions) and (2) very short-run estimates such that the treatment effect dominates any omitted risk premium. If these two cases do not hold, this irrelevance will disappear.
We contrast three types of estimators that can be used for financial event studies, and compare their properties: (1) classic abnormal return estimators, based on specified factors, (2) difference-in-mean estimators, which construct control groups through decisions of the econometrician, and (3) synthetic estimators abadie2003economic,abadie2010synthetic, which use historical prices from control assets to construct either a replicating portfolio (synthetic control) or to construct a set of factors using PCA xu2017generalized. The key insight is that rather than imposing a specific factor structure ex ante, the synthetic methods construct a portfolio of control securities that best match the pre-event return path of treated securities. If such a replicating portfolio exists, it should provide valid counterfactual returns in the post-event period without requiring correct specification of the underlying factor model.
Our theoretical results hinge on the assumption that the expected return for a cohort of treated stocks follows an unknown time-invariant linear factor model. This assumption is not innocuous, and likely not true for all time periods. But, it is also a weaker assumption than the traditional abnormal return estimators. Any approach that uses a model to infer the counterfactual outcomes for the treated stocks will require some kind of model stability assumption (without additional structure like kelly2019characteristics). We view it as valuable future work to see if other more robust asset pricing models can be used to generate counterfactual returns, such as giglio2025test and kelly2019characteristics.
One key benefit of focusing carefully on the estimand of interest is that we are able to show that buy-and-hold abnormal return estimates are particularly challenging to estimate because they require the matching portfolio to not just match on expected returns, but also on volatility. If the control group's returns have different variance, then the differing volatility drag will lead to very different results. To make this concrete: imagine that there is no treatment effect, but a diversified portfolio is used as a control group for a stock, both with equal expected returns. The lower variance for the diversified portfolio will lead to a negative treatment effect from a buy-and-hold perspective, despite no actual treatment effect. This implies that doing buy-and-hold abnormal returns with an index can be seriously flawed.
We revisit four empirical settings that span the range of typical applications. First, we reexamine the acemoglu2016value study of political connections during the 2008 financial crisis, where the Treasury Secretary announcement coincided with extreme market volatility—daily returns exceeded 6% on multiple event days. The original estimates using simple averaging suggest economically large effects of political connections. Even abnormal return models using the Fama-French 3 factor model suggest economically meaningful effects. However, the estimates disappear when using our proposed synthetic methods, suggesting that model misspecification with a single event can create spurious results when events coincide with volatile market conditions.
Next, we analyze S&P 500 index inclusions, and show that since the index inclusion events appear random across time, the effect of short-run model misspecification is non-existent, echoing the folk wisdom above. However, we show that the substantial pre-announcement drift, often pointed to as a source of possible index inclusion front-running, disappears once we properly account for the unobserved factor exposures of included firms. This finding suggests that what appears to be anticipation or momentum may actually reflect model misspecification.
Third, we examine the effect of acquistions in merger deals on acquiring firms with some studies finding large negative abnormal returns over several years.loughran1997long, rau1998glamour We demonstrate that these long-run patterns are highly sensitive to model specification, consistent with our theoretical prediction that misspecification bias accumulates over longer horizons.
Our last empirical result applies a version of lalonde1986evaluating to our analysis by using quasi-experimental variation to provide a benchmark for our model-based approaches. The treatment and control groups for the baseline are found in close merger contests where multiple firms bid for the same target. Following malmendier2018winning, contest losers provide a natural counterfactual for winners since they are ex ante similar firms competing for identical targets. The results are not supportive of abnormal return models at all, but only weakly support synthetic methods. These results suggest that for long-run analyses, it is far better to construct counterfactuals based on quasi-experimental variation than using model-based approaches.
These empirical findings have important implications for the interpretation of the vast event study literature in finance. Many influential results---particularly those involving long horizons, volatile periods, or systematic event timing---may reflect factor model misspecification rather than true treatment effects. However, we emphasize that our results do not invalidate the entire enterprise. For short-horizon studies with plausibly random event timing, traditional methods remain reliable and our empirical work confirms they produce similar estimates to more sophisticated approaches. The key insight is recognizing when standard methods are likely to fail and having appropriate alternatives available.
Our work connects several distinct literatures. Methodologically, we build on the econometrics of event studies in finance mackinlay1997event, kothari2007econometrics while incorporating insights from the modern causal inference literature imbens2015causal,abadie2021introduction. We also contribute to the older debate about long-run event studies mitchell2000managerial, barber1997detecting by providing a formal framework for understanding when and why these studies are problematic.
This section formalizes the setup of financial events on stock market returns in the language of potential outcomes. We begin by introducing the basic notation (Section (ref)), defining potential returns and treatment indicators for each security over time. We then specify the causal estimands of interest, clarifying what it means to identify a treatment effect in an event study context. Finally, we discuss how these causal quantities relate to traditional event study methods based on “abnormal returns” and factor model adjustments.
We study the causal effects of corporate events on security returns using a potential outcomes framework. Consider a panel of $N$ securities indexed by $i = 1, 2, \ldots, N$ observed over $T$ time periods indexed by $t = 1, 2, \ldots, T$.
For each security $i$, let $T_i$ denote the time when an event occurs:
We denote the set of event times as $\mathcal{S} \subseteq \{1, \ldots, T\}$ and the set of never-treated (control) securities as $\mathcal{C} = \{i: T_i = \infty\}$. Following standard practice in event studies, we assume events are irreversible---once an event occurs (e.g., a merger announcement or earnings release), it cannot be undone.
Now, we define the potential outcomes framework for our returns. Let \(R_{i,t}(s)\) be the potential return for security \(i\) at time \(t\) if it has the event occur in period $s$, and \(R_{i,t}(\infty)\) the potential return in the absence of any event. Because a security cannot be both treated and untreated, we only observe one of the potential returns for each \((i,t)\):
We postulate that financial event studies are focused on identifying the difference between the realized returns for a treated firm ($R_{it}(s)$) versus the returns in the absence of the event. We define the difference in returns due to the event in period $s$ for firm $i$ in period $t$ as the individual treatment or equivalently, the abnormal firm return:
For a firm that has the event occur in period $s$, $R_{it}(s)$ is observed, and hence is identified. But, $R_{it}(\infty)$ is not. Indeed, in asset pricing, the challenge of modeling the exact return for an individual asset is viewed as an near-impossible task, even with a structural model. Instead, a large number of asset pricing papers focus on the challenge of estimating the average return for firms given a set of characteristics and/or risk factors chamberlain1983arbitrage,CONNOR198413,fama1993common, ross2013arbitrage, kelly2019characteristics, bryzgalova2025forest.
This focus on expected returns makes causal inference and asset pricing models happy bedfellows. The inability to known the exact counterfactual return is known as the fundamental problem of causal inference and leads to a focus on other alternative estimators, often constructing average counterfactual returns for a group of treated units.
A significant body of empirical work and legal scholarship focuses on identifying the effect of events on single firms' valuations, since these valuations are used in litigation to estimate damages baker2020machine. But our view is that in academic research studying financial event studies, a much more natural estimand to target is the average treatment effect on the treated (ATT), using many treated firms to estimate an overall average effect, rather than the effect on a single firm. We view the estimated abnormal returns for single firm events as case studies of a much more stable design that focuses on the average effect.
This cohort-period ATT describes the effect of a treatment happening in period $s$ during period $t$ for those firms who are experience the period $s$ event. If these firms are special in some way, then this may not be the same effect for other firms (for example, if these firms are riskier, and the effect differs by risk profile).
These cohort-period ATTs can be combined in a number of ways. Most crucially for our results, combining event cohorts to study effects relative to an event time will average across different event timings. The average treatment effect $\kappa$ periods after an event is:
where $w_s$ represents the weight on event cohort $s$. A natural choice is $w_s = N_s/\sum_{s'} N_{s'}$, where $N_s$ is the number of securities with $T_i = s$.
Many empirical papers studying these announcements are interested in cumulating the effects. The Cumulative Average Treatment Effect (CATT), analogous to cumulative abnormal returns (CAR), from event time 0 to $H$ is:
In this paper, we focus on these linear transformations of the ATT because they are well-behaved econometrically. However, an alternative approach to cumulative arithmetric returns is the buy-and-hold abnormal return, which we discuss briefly here to highlight its econometric challenges.
Announcement effects are often cumulated using buy-and-hold returns, which correspond to geometric returns. The usual approach for defining abnormal buy-and-holds returns in the literature differences out the buy and hold return of a counterfactual portfolio or stock savor2009stock, barber1997detecting from a stock's buy-and-hold return. In our setting, this is analogous to
This object is challenge to analyze analytically, and has many challenging statistical properties barber1997detecting, mitchell2000managerial.
In our notation, this corresponds to the following geometric estimands. Let the cohort-horizon geometric ATT for cohort $s$ at horizon $H$ as
As with the arithmetic ATT, this can be averaged over the event timings:\footnote{Note that in the case of the geometric cumulative return, we first cumulate over the holding period, and then average across periods, since the non-linear structure makes the two non-interchangeable. }
Note that this estimand effectively studies the percentage difference in gross cumulative returns, rather than level difference in gross cumulative returns. For researchers interested in sign tests (e.g. positive or negative long-run returns), both objects work equally well.
We now present a result tying the arithmetic (abnormal return) and geometric (buy and hold) ATT together.
This result shows that buy-and-hold returns incorporate both volatility drag and the interaction between base returns and treatment effects, making them more complex to analyze than arithmetic returns. An important implication of this is if the counterfactual return $\hat{R}$ chosen for $R_{it}(s)$ identifies $E(R_{it}(\infty) | T_{i} = s)$, it may be a bad counterfactual for buy-and-hold returns because it does not match on volatility. For example, a portfolio with identical returns to $E(R_{it}(\infty)|T_{i}=s)$ may have much lower variance (due to diversification). As a result, the volatility drag from the treated observed units will bring down the geometric returns, even in the absence of any true effect!
As a result of Lemma (ref), we focus on estimating the arithmetic ATT, rather than approximating the buy-and-hold return.\footnote{The issues raised here are analogous to problems in difference-in-difference for log vs. level outcomes. If the parallel trends assumption holds for a level outcome, than it almost surely cannot equivalently hold for a log outcome, unless the treatment is randomly assigned. roth2023parallel} Geometric returns would require a counterfactual return portfolio that matches on both level and variance, and since the variance of a portfolio does not have the same theoretical guidance for a model as expected returns, finding this counterfactual portfolio is quite hard. It also suggests that papers that use buy-and-hold abnormal returns may contaminate their results as a function of how many firms are included in the counterfactual return portfolio due to diversification differences.
We now operationalize our model for $E(R_{it}(\infty) | T_{i} = s,)$, based on a long literature in asset pricing chamberlain1983arbitrage, CONNOR198413.
Note that the linear factor assumption is quite strong. For example, it does not allow for changing factor loadings barberis2005comovement. It also does not allow for the market to anticipate an event (rationally) in the future if the event does not eventually occur.\footnote{This issue is considered in a series of papers in the finance literature, e.g. prabhala1997conditional, that consider conditional events.} However, it nests generally almost all financial event study methods, such as using the market model, CAPM, or Fama-French factors to construct the counterfactual return campbell1997econometrics. It is also possible that this model could only hold for a short period of time, allowing for varying loadings over a longer period of time (as in kelly2019characteristics).
Two important special cases are:
These assumptions formalize when simple estimators will be unbiased and when more sophisticated methods are needed.
This assumption implies that the treated group cannot have an impact from the announcement for a sufficient window prior to the date of the release. There is obvious evidence in the finance literature of hidden information leaking out, with prices responding beforehand (e.g. schwert1996markup). Indeed, this is often pointed to evidence for the strong version of the efficient markets hypothesis. Hence, limited anticipation will be necessary to set a benchmark for when leakage has not yet occurred. This will allow the researcher to identify the periods in which we can estimate the counterfactual returns. This is the assumption necessary to use the pre-event estimation window commonly used in financial event studies campbell1997econometrics,kothari2007econometrics.
However, it is important to distinguish between selection into the treatment (e.g. $\{R_{it}(s)\}_{s\in\mathcal{S}}$ being correlated $T_{i}$) and anticipation of the treatment. The former is quite plausible, as we see in our analysis of the S&P 500 index inclusion effect in (ref) -- firms that are growing and having a large market cap are more likely to be selected into the S&P. The latter will bias our estimates of the true treatment effect, and can be caused by market participants anticipating the event.
We now present four sets of estimators and characterize the conditions under which they identify the ATT. In all cases, we assume returns are already adjusted for the risk-free rate.
Consider first the canonical abnormal returns model used in finance research campbell1997econometrics, BROWN19853. The researcher begins by selecting a set of observable factors $F^o_t$ and estimates factor loadings $(\hat{\alpha}_i, \hat{\beta}_i)$ using ordinary least squares on data prior to $T_i - \delta$:
These estimates $\hat{\alpha}_i$ and $\hat{\beta}_i$ minimize squared prediction errors for stock $i$'s returns using the observed factors. The factors $F^o_t$ may include no factors, a single factor (the market return), or multiple factors (e.g., Fama-French factors).
This approach attempts to remove the component of returns attributable to systematic factor exposure, leaving only the “abnormal” component. Under correct specification of the factor structure ($F^o_t = F_t$ for all relevant factors), this abnormal component should isolate the treatment effect. However, when factors are omitted or mismeasured, the estimated loadings $\hat{\beta}_i$ may fail to capture the true exposures $\beta_i$, leading to bias.
We compare the abnormal returns approach to three alternatives estimation approaches.
When the control group consists of all securities weighted by market capitalization, this estimator corresponds to the “market-adjusted-return model” of campbell1997econometrics and BROWN19853. Alternatively, the control group might consist of matched firms selected based on observable characteristics, as in barber1997detecting and loughran1997long.
Second, we consider a synthetic control estimator abadie2021introduction that uses the pre-event data to construct a synthetic control group:
The synthetic control method originated in abadie2003economic, abadie2010synthetic and has expanded and grown as a method over the last decade. Synthetic control directly constructs a counterfactual by matching the pre-event dynamics of treated securities using a portfolio of controls. This synthetic control is then used as a counterfactual return following the event.
The key distinction from abnormal returns is that synthetic control does not require the researcher to specify or estimate the underlying factor structure. Instead, it searches for portfolio weights that replicate the treated group's returns in the pre-period, effectively letting the data determine the appropriate factor exposures. If such a replicating portfolio exists, it should continue to provide valid counterfactual returns in the post-period (absent the treatment).
We could depart from the original synthetic control applications by allowing negative weights. Traditionally, synthetic control methods restrict $\omega_j \geq 0$ to ensure the counterfactual represents a convex combination of control units. However, this restriction is unnecessarily limiting in financial applications. Allowing negative weights permits short positions and significantly expands the set of achievable factor loadings, making it more likely that a replicating portfolio exists. This flexibility is natural in financial markets and consistent with standard long-short portfolio construction. However, absent this restriction, we are not able to prove our results on unbiasedness using results from ferman2021properties.
We focus on constructing a single synthetic control for the portfolio of treated securities ($R_{s,t}$) rather than constructing separate synthetic controls for each individual security. This choice reflects both practical and theoretical considerations. Empirically, individual stock returns contain substantial idiosyncratic noise that would make firm-by-firm matching challenging. Theoretically, our estimands target average treatment effects for groups of securities, not individual effects, making portfolio-level analysis natural. This approach follows very naturally the approach advocated in ben2022synthetic for staggered synthetic control.
In practice, perfect pre-period fit may not be achievable. Extensions by abadie2021penalized, ben2021augmented, ben2022synthetic allow for approximate rather than exact matching, trading off pre-period fit against overfitting concerns. However, most importantly for our analysis in financial event studies, ferman2021properties shows that if the data follows a linear factor structure, then with sufficient pre-event time periods and control units, the estimator is consistent. This is consistent with a wide-range of asset pricing work highlighting the importance of having assets that span risk factors giglio2021asset, giglio2025test.There is also a close connection to the mimicking-portfolio approach huberman1987mimicking.
We also consider a third estimator, following xu2017generalized, which uses PCA regression with cross-validation to estimate a factor structure with unknown factors:
The Gsynth estimator more directly leans on the linear factor structure, but does not require knowing the true factors, and uses the set of control firms to construct the set of counterfactual returns.
We focus on these two alternative estimators, but other alternative methods, such as IPCA kelly2019characteristics or the three-pass method in giglio2021asset may work as well or better. We leave it to future work to consider what approaches may work best.
We now establish conditions under which these estimators identify the ATT. For this proposition, it is convenient to see how these estimators differ from the target single event-period estimand:
where $\alpha_s = \mathbb{E}(\alpha_i | T_i = s)$, $\beta_s = \mathbb{E}(\beta_i | T_i = s)$ are the average intercept and factor loadings for treated securities, $\alpha_\infty$ and $\beta_\infty$ are corresponding quantities for the control group, $\hat{\alpha}_s$ and $\hat{\beta}_s$ are the estimated loadings from the abnormal returns approach, and $\hat{\alpha}^{alt}_s$ and $\hat{\beta}^{alt}_s$ are the implied loadings from either the synthetic control or gsynth estimator. $\varepsilon_{st} = n^{-1}_{s}\sum \varepsilon_{it}$ is the average idiosyncratic noise for the $i$ cohort, and $\varepsilon_{\infty,t} = n^{-1}_{i \in \mathcal{C}}\sum varepsilon_{it}$ is the average noise for the control group.
All proofs are in (ref).
The most complex part of this proof, proof of asymptotic unbiasedness of the synthetic control estimator, follows directly from ferman2021properties, who show that the synthetic control estimator is asymptotically unbiased under the assumption of an unknown linear factor model and many control units. The results for Gsynth also follow directly from xu2017generalized. The other two results follow from the assumptions and the definition of the estimators.
Of course, intuitively, in many applications the factor loadings are often not too large, and the underlying risk premia are, on average, typically small relative to $\tau^{ATT}(s,t)$. For example, the one-day index inclusion effect is estimated to be somewhere between 1-4%, depending on the time period. By comparison, the market return is, on average, 0.05%, two orders of magnitude smaller than the treatment effects.
However, there are many periods when the market return can be far larger, such as during periods of market volatility. There is substantial variation in the size of these factors, with an interquartile range of 1% and very large fat tails. Hence, the correlation of the factors with the timing of the event is very important. This will be apparent in our first empirical example of acemoglu2016value. As a result, this bias can be quite large. Formally, we can write the following from (ref):
We next consider how these results change if there are multiple event periods.
An implication of this is that the abnormal returns estimator is can be quite close to the true treatment effect, even when the factor model is misspecified. Moreover, this bias could be small even for a model that ignores factors, consistent with the simulation evidence in BROWN19853 that the form of the abnormal return estimator has limited effects on the estimates.\footnote{The simulations in BROWN19853 are such that the event days are exactly randomly assigned across time: “Each time a security is selected, a hypothetical event day is generated. Events are selected with replacement, and are assumed to occur with equal probability on each trading day from July 2, 1962, through December 31, 1979.”} In fact, a common phrase described in event studies is that the structure of the model used in $\tau^{AR}$ does not have significant impacts on the estimated effects. For example, in footnote 5, shleifer1986demand states “The [index inclusion] results were not materially different when returns were not corrected for market movements. Similarly, combining the before and after estimation periods did not make much difference.” Or in edmans2012link “I use the standard short event-study window so that the calculation of abnormal returns is relatively insensitive to the benchmark asset pricing model used.”
Researchers are often interested in the trends or cumulative impact of events on returns, as measured by cumulative abnormal returns or buy-and-hold abnormal returns. This gets mapped to different economic and behavioral theories about how the market processes information (e.g. daniel1998investor is a theory to explain these effects from a behavioral perspective; kwon2022extreme consider 90 day post-announcement effects relative to announcement day effects).
Some papers have pointed to flaws in studying these types of long-run perspectives -- for example, Mitchell and Stafford (2000) highlight the flaws in the inference around long-run abnormal return studies of firm activity. As we show in (ref), the buy-and-hold abnormal return has additional challenges caused by variance considerations in the counterfactual portfolio. We now use our results in (ref) and (ref) to show that even estimating arithmetic cumulative abnormal returns in the long-run amplifies the misspecification bias.
Following the analogy principle for the CATT, (ref) tells us that the bias in for the abnormal return and difference-in-mean estimators is intimately related to the cumulative sum of the factor premium. Under random timing and many events, the bias for the CATT at horizon $H$ is $H E(\beta_{i} \mid T_{i} \in \mathcal{S})E(F_{t})$. The factors have a positive mean (since the risk of the factors leads to positive expected return), and thus the bias in the estimator will drift proportional to the expected value of the factors during the time period, scaled by the relative estimation error in $E(\beta_{i} \mid T_{i} \in \mathcal{S})$.
Consider estimating the long-run impact of a merger on stock market prices. RAGHAVENDRARAU1998223 find a three-year long run effect of -4% for all mergers, while savor2009stock find a three-year long-run effect of -13.1% for stock-financed mergers and 1.6% for cash financed mergers. These results are well-motivated by shleifer2003stock, but their magnitude may reflect bias due to the errors in $\hat{\beta}$:
If $K=1$, for example, and was equal to the market, then our expected excess return is 6%. If $\beta_{sk} - \hat{\beta}_{sk}$ was $-0.1$, then at the three year level, we might expect a bias of -1.8%. This is of course an empirical question of which way the biases would go; is the constructed portfolio of firms too heavily loaded on risk factors?
Note that these issues are not solved by using multiple event timings. This bias in factors cannot average out to zero, and so the only source by which we can achieve zero bias is through mean zero differences in the loadings.
It is also worth remarking how the results from Mitchell and Stafford (2000) can be seen analytically in our statistical terms. While the misspecification term $ (\beta_{s}-\hat{\beta}_{s})\sum_{\kappa=0}^{K^{0}}F_{s+\kappa}$ creates bias, it also creates cross-correlation in errors for every event-timing.\footnote{As they state: “[M]ajor corporate events cluster through time by industry. This leads to positive cross-correlation of abnormal returns, making test statistics that assume independence severely overstated.” }
Key takeaways for practitioners are four-fold:
We briefly discuss the case of a single firm being treated. To analyze this case, we need to allow for slightly more flexibility in our notation.
Then, consider the case of a single firm estimated in each estimator:
Statistically, there are now three objects with randomness to worry about: the estimated parameters, the aggregate factors, and the idiosyncratic variance for the individual firm. Note that with several treated units, this last term disappears, but with a single unit, we have insurmountable noise. This is a common problem flagged in the event studies literature looking at securities litatigation baker2020machine.
However, consider an approach that estimates many individual treatment effects in this manner (such as kogan2017technological). On, average, these estimates will be subject to the same results outlined above, but each one is quite noisy. This is equivalent to problems associated with estimating many treatment effects. One approach is to consider shrinkage estimators. Another would be to pool the firms based on characteristics of interest, and construct portfolios this way. This would remove $\varepsilon$.
We highlight how the non-random timing and assignment, together with a misspecified factor model, could affect the bias with different estimators of treatment effects, using a simple simulation exercise. In the simulation, the returns follow a two-factor structure, with the second factor omitted in the estimation of abnormal returns. We compare the expected bias, root mean square error, and coverage with random vs. nonrandom assignment and timing.
We simulate a panel of stock returns with a linear factor structure:
where the return for each stock equals to the risk-free rate, plus the exposure times risk premium of a market factor and a size factor (small-minus-big), and a stock-level idiosyncratic component.
We assume that both factor loadings follow independent normal distributions: $\beta_{i,mkt},\beta_{i,smb}\sim \mathcal{N}(1,0.3^{2})$. We further assume that the idiosyncratic component of each stock is drawn i.i.d. from a Normal distribution: $\varepsilon_{i,t}\sim \mathcal{N}(0,0.1^{2})$. We choose a standard deviation of around 0.1 so that the residual variance constitutes approximately half of the total variance.
We simulate returns for 500 firms, with pre-treatment period of 239 days, 1 event day, and 10 post-treatment periods. Roughly 10% of firms are treated, following one of two treatment assignment processes, discussed below. Treated firms get a true effect of 3% on the treatment day, and nothing afterwards. The factor returns and the risk-free rate are randomly sampled from daily Fama-French returns from July 1926 to 2022 with block sampling to preserve the correlation structure between factors.
\paragraph{Treatment assignment process} We compare expected bias with different treatment assignment selection and timing selection. For firm assignment, we either completely randomly assign the treatment to 10% of firms, or to instead relax this assumption, we model that the probability of a firm getting treated follows a logit function of the beta on the SMB factor
where $\delta=\frac{\log(0.1)}{E(\beta_{i,smb})}<0$ to achieve an average probability of 10%. The lower the simulated SMB factor loading of the firm, the more likely to be treated.
For treatment period selection, we similarly use two different assignment mechanisms. The first is to randomly sample the 250 data periods, and always set the treatment period equal to $t=240$. This effectively makes the treatment period's factor draw uncorrelated with the treated firms' factor loadings. The secon approach with timing selection works as follows. First, we rank the SMB factor in 250 candidate treatment periods. We then use the rank of SMB returns as inputs to the selection function.\footnote{Raw factors returns have positive and negative values with mean close to 0, which will make the logit function highly sensitive.} The probability of any one of the candidate period being the treatment period is
where $\delta=\frac{\log(1/250)}{E(Rank_t)}$. We then draw indicator variables for each candidate period from binomial distributions with respective treatment probability in each period. If multiple periods are drawn to be the event period, we use the one with the highest factor realization. Thus, if a period has a high factor realization of the omitted factor, it is more likely to become the treatment period.
In (ref), we compare the performance of four different estimators across 50 simulations: mean difference between treated and control firms, average abnormal returns using the market factor (estimating the factor loading for each treated firm in the pre-period), average abnormal returns using the both factors (estimating the factor loadings for each treated firm in the pre-period), and average treatment effects from the generalized synthetic control method (Gsynth). Estimated bias is reported in percentage points. We also report the root mean square error (RMSE) and coverage of 95% confidence intervals.
First, in Panel A, we see that the average bias is small even with the wrong factor structure, if the treatment is randomly assigned. Similarly, in Panel B, if we only have non-random assignment selection, the expected bias is also insignificant on average. However, this masks the variation across simulations - if a time period has a larger factor draw on the treatment day, that leads to much larger bias.
In Panel C, we consider random assignment of treatment to units, but non-random event timing. As in Panel A, the difference in means is unbiased thanks to the results in (ref). Since treatment is uncorrelated with factor loadings, there is no endogeneity and the simple means estimator is an unbiased estimator of the treatment effect. However, with non-random timing, the CAPM model is biased, because the abnormal return (as discussed in Section 2.2) will be the average $\beta$ for the omitted factor multiplied by the largest possible factor draw. In contrast, the difference in means is unbiased because while both treated and untreated firms are exposed to the high factor draw, they have identical factor exposures, which cancels out. For the correctly specified model, the estimated model correctly specifies the counterfactual, and so there is no bias. Finally, the Gsynth estimator is able to identify the correct underlying factor structure, and has limited bias as well.
Once we have both types of selection in treatment in Panel D, we see that the simple difference in means is now biased. However, it is still less biased in absolute value than the misspecified CAPM model. This is because the gap in the treatment and control factor loadings for the simple mean difference is still smaller than the level misspecification in the factor loadings in the CAPM estimation. Again, the Gsynth approach does quite well, with similar performance to the correctly specified factor model.
We now turn to our first empirical example, examining the period when the announcement of Timothy Geithner as Treasury Secretary was leaked, following the setup of acemoglu2016value. This example highlights the results of Proposition (ref) in a simultaneous treatment setting. We demonstrate that the bias from an incorrect factor structure can be substantial in this setting, and that synthetic control methods help alleviate this bias. We argue that the bias arises from two sources: first, the event window coincides with turbulent market conditions characterized by large daily factor realizations; second, the counterfactual returns are constructed from control firms with substantially different factor exposures. We show that synthetic methods, which greatly reduce these biases, also match the factor loadings of treated firms for known factors such as size and value.
\paragraph{Empirical setup.} We examine the announcement of Timothy Geithner as nominee for Treasury Secretary on November 21, 2008. Following acemoglu2016value, we estimate average treatment effects over the 11-day window encompassing and following the announcement date, from November 21, 2008 (day 0) through December 8, 2008 (day 10).\footnote{November 24, 2008 corresponds to day 1 due to the weekend.} For treated and control bank returns, we use the data provided by the authors, who collected daily returns from Datastream.\footnote{We thank Amir Kermani for providing the replication code and data on his website.} For all trading days before and after the event, returns represent full trading day returns during regular trading hours. For the event day, returns are calculated from 3:00 p.m. (when the news leaked) until market close at 4:00 p.m.
We consider two sets of control firms. First, we use the same set of financial firms listed on the NYSE or NASDAQ that are not connected to Geithner, as in acemoglu2016value. Second, we expand the control group to include all NYSE, AMEX, and NASDAQ (exchange codes 1--3) common stocks (share codes 10 or 11).
\paragraph{Non-connected banks as controls.} We first use public financial institutions without connections to Geithner as control firms. Panel A of Table (ref) reports the average treatment effects over the 11-day post-event window. Column 1 presents the difference in average returns between treated and control firms, implementing the counterfactual as a simple average of returns from non-connected firms—the same approach used in Table 2 of acemoglu2016value. Column 2 reports difference-in-differences estimates. Columns 3--5 present traditional factor model adjustments: the market model (Column 3), CAPM (Column 4), and Fama-French three-factor model (Column 5). Columns 6--8 employ synthetic control methods: standard synthetic control abadie2010synthetic, synthetic difference-in-differences arkhangelsky2021synthetic, and generalized synthetic control (Gsynth) from xu2017generalized. For all models requiring pre-event estimation, we use days $-256$ to $-31$, slightly shorter than the $-280$ to $-31$ window in the original paper to maintain a balanced panel. We report a graphical version of (ref) in (ref).
The results reveal a clear pattern. Simple averaging and difference-in-differences (Columns 1--2) show that firms with schedule connections experience 2.6--2.7% higher cumulative returns, those with personal connections show 2.9--3.0% higher returns, and firms with New York connections exhibit 1.9--2.0% higher returns. The market model adjustment (Column 3) produces minimal changes. However, risk-adjusted returns using CAPM and Fama-French models (Columns 4--5) reduce these estimates by approximately 40--50%, suggesting that connected firms have higher market betas. The synthetic control methods (Columns 6--8) produce even larger reductions, with standard synthetic control reducing schedule connection effects by 38% and personal connection effects becoming statistically insignificant.\footnote{Our results contrast with acemoglu2016value, who employ synthetic control methods as robustness checks. Their approach was necessarily ad hoc given the limited literature at the time on handling multiple treated units in synthetic control settings.}
\paragraph{All public firms as controls.} We next expand the control group to include all common shares traded on NYSE, AMEX, and NASDAQ, with results reported in Panel B of Table (ref). This expansion is motivated by the integration of equity markets: systematic factors should be well-identified using the universe of traded stocks. Restricting controls to financial firms alone may be suboptimal unless banking-specific factors exist that cannot be spanned by the broader market.
The expanded control group dramatically changes the results from synthetic control methods while leaving traditional methods largely unaffected. Simple averaging and difference-in-differences (Columns 1--2) continue to show significant effects of approximately 2% for all connection types. The factor model adjustments (Columns 3--5) show a similar pattern to Panel A, with the market model producing minimal changes while CAPM and Fama-French adjustments reduce estimates by 30--50%.
Strikingly, the synthetic control methods now produce near-zero and statistically insignificant estimates. Standard synthetic control (Column 6) yields point estimates of 0.4% for schedule connections and $-0.3\%$ for personal connections. Gsynth (Column 8) estimates are particularly close to zero: 0.1% for schedule connections, 0.3% for personal connections, and 0.1% for New York connections—all statistically insignificant. This dramatic difference suggests that the broader control group allows synthetic methods to better match the factor exposures of treated firms, effectively eliminating the estimated treatment effects. The contrast between traditional factor adjustments (which still show significant effects) and synthetic methods (which do not) highlights the importance of allowing flexible, data-driven matching of factor exposures rather than imposing a specific factor structure.
We now investigate the sources of bias in the original estimates. First, we examine the distribution of market returns during the event window. Figure (ref) displays the kernel density of daily S&P 500 returns from 1962--2023, overlaid with the realized returns during the 11-day event window. The event period coincides with extraordinary market volatility, with returns falling in the extreme tails of the historical distribution. The market surged 6.6% on November 21 (day 0) and 6.5% on November 24 (day 1), while the largest decline of $-8.4\%$ occurred on December 1 (day 5).
These extreme factor realizations have important implications for identification. As demonstrated in Proposition (ref), when treatment occurs simultaneously for all units, abnormal return estimators are particularly sensitive to factor model misspecification. The bias is proportional to both the magnitude of factor realizations and the difference in factor loadings between treated and control firms: $(\beta_s - \hat{\beta}_s)F_t$. Large factor realizations during the event window amplify any misspecification bias arising from imperfect matching of factor exposures.
This mechanism explains the substantial reduction in estimated treatment effects when using synthetic control methods rather than simple averaging. Proposition (ref) shows that both synthetic control and gsynth estimators are asymptotically unbiased even with omitted factors, as they construct control portfolios that match the pre-event factor structure of treated firms without requiring explicit factor model specification. The extreme market conditions during the Geithner announcement thus reveal the importance of proper counterfactual construction in volatile periods.
We now provide direct evidence on the factor exposure differences between treated and control firms. We estimate market betas using daily returns from day $-280$ to day $-31$ before the event, running firm-level time-series regressions on the S&P 500 index return for CAPM betas and on the Fama-French three factors for multifactor betas.\footnote{Fama-French factor returns are obtained from Kenneth French's data library: \url{https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html}.}
Table (ref) reports the weighted average betas for treated and control portfolios. Panel A presents equal-weighted averages for treated firms and two control groups: financial institutions only and all public firms. The results reveal substantial factor exposure mismatches. Treated firms have an average CAPM beta of 1.43, compared to 0.83 for financial controls—a difference of 0.60. The Fama-French three-factor model confirms this pattern: treated firms exhibit a market beta of 1.28 versus 0.66 for controls, with similar disparities in SMB (0.23 vs.\ 0.75) and HML (0.61 vs.\ 0.72) exposures.
These factor loading differences, combined with the extreme market realizations documented in Section (ref), generate substantial bias in simple difference estimators. During the event window, the 0.60 difference in market beta translates to a bias of approximately $0.60 \times 6.9\% = 4.1\%$ on November 21 alone. This mechanical bias explains much of the estimated effect found using naive averaging methods.
We report the (weighted) average of betas of treated and control firms in Table (ref). First, in Panel A, we first show the average CAPM and Fama-French three-factor betas of the treated firms and equal-weighted averages of financial firm controls and all public firm controls. The average CAPM beta of the treated firms is 1.43, much higher than 0.83 from the control firms. Expanding to a three-factor model, we still see a higher market beta in treated firms. Given these mismatches of treated and control betas, together with turmoil market returns, as shown in Section (ref), could lead to large biases in average treatment effects by comparing treated versus control firms.
In Panel B, we compute the weighted average betas of control firms using synthetic control weights, with both standard synthetic control and synthetic difference in differences. First, we see that synthetic methods match the beta in the treated firms well. For example, the synthetic control gives a weighted average beta of 1.33, much closer to the treated beta of 1.43 than the equal-weighted average. Fama-French three-factor betas of the treated firms are 1.28 on the market, 0.23 on SMB, and 0.61 on HML, and synthetic control weights give a market beta of 1.15, SMB beta of 0.48, 0.75 ( closer than 0.66, 0.75, and 0.72 with simple average). Second, if we extend the set of possible control firms from financial firms in acemoglu2016value to all public firms in CRSP, we obtain better matches across all synthetic methods. For synthetic control specifically, controlled firms give an average beta of 1.38, closer to 1.43 in the treated firm. There is also a significant improvement in matching the Fama-French three-factor betas, synthetic control betas are 1.22, 0.38, and 0.67 (compared to treated betas of 1.28, 0.23, and 0.61). Finally, standard synthetic control methods give slightly better weights than synthetic difference-in-differences, who is more directly related to a mimicking portfolio approach.
Overall, synthetic methods matches the beta of treated firms well, which results in a lower bias in the average treatment effects.
We next examine S&P 500 index inclusion announcements, analyzing both immediate announcement returns and pre-announcement price dynamics to test our theoretical predictions regarding identification in staggered event settings.
We first demonstrate that in staggered event settings, announcement-day bias is negligible because factor returns on event days average close to zero, particularly when compared to the large treatment effects of 3--4%. However, consistent with Greenwood2025Index, we document substantial pre-announcement drift. Synthetic control methods that match on pre-event returns nearly eliminate this drift. This pattern is consistent with selection on unobserved factors: firms added to the index differ systematically from control firms along dimensions not captured by observable factors, generating apparent pre-event "drift" that actually reflects factor model misspecification.
\paragraph{Empirical setting.} Following Greenwood2025Index, we obtain index inclusion dates from Siblis Research and match tickers to CRSP PERMNOs using header information. Siblis provides announcement dates for S&P 500 additions. For the period September 1976 through September 1989, when announcement dates are missing, we exploit the institutional detail that index changes were announced after Wednesday market close and became effective the following day, allowing us to infer announcement dates.\footnote{During this period, S&P followed a predictable schedule of announcing changes after Wednesday close for Thursday implementation.} We measure returns on the announcement date when it falls on a trading day; otherwise, we use the most recent prior trading day.
To assess whether event timing can be treated as random, we examine the distribution of factor returns on announcement days. Appendix (ref), Panel A shows that the distribution of daily market returns on S&P 500 index inclusion announcement days is virtually indistinguishable from the distribution on non-announcement days. This pattern holds consistently across our entire sample period, from 1980--1989 through 2010--2020. The small-minus-big (SMB) factor exhibits similar distributional stability (Panel B). These results support treating announcement timing as conditionally random with respect to factor realizations, satisfying a key identification assumption for our short-horizon analysis.
Table (ref) reports CAPM and Fama-French three-factor betas for firms added to the S&P 500 index, estimated using daily returns from days $-250$ to $-100$ relative to announcement.\footnote{We exclude the immediate pre-announcement period to avoid contamination from potential information leakage.} We present results separately by decade from 1980 through 2020 to examine temporal variation in the characteristics of included firms.
Across all decades, the average market beta of included firms is approximately one. When treated firms have market betas near unity, the simple market-adjusted return (which implicitly assumes $\beta = 1$) yields similar results to the more sophisticated CAPM adjustment that estimates firm-specific betas. This convergence occurs because the bias term $(1 - \beta_i) \times r_{m,t}$ approaches zero when $\beta_i \approx 1$, consistent with the theoretical predictions in (ref).
The combination of two empirical regularities—random event timing with respect to factor realizations and limited selection on factor loadings—suggests that short-horizon abnormal return estimates should exhibit minimal bias regardless of the specific factor model employed. This prediction from Theorem (ref) finds strong empirical support in Table (ref), where announcement-day treatment effects are remarkably stable across estimation methods. The difference between simple market adjustment and sophisticated synthetic control methods is less than 0.2 percentage points in most decades, confirming that model specification has negligible impact on short-horizon estimates when the conditions of Theorem (ref) are satisfied.
While (ref) predicts negligible bias in short-horizon studies, it also implies that long-horizon estimates may suffer from substantial bias unless factor exposures are correctly specified. We now examine the "pre-announcement drift" documented by Greenwood2025Index, analyzing it decade by decade as a manifestation of potential long-horizon bias.
Interpreting pre-announcement price movements requires careful consideration of (ref), our limited anticipation assumption. This assumption is particularly tenuous in the index inclusion setting for two reasons. First, market participants have incentives to anticipate market index changes. Second, as Greenwood2025Index document, inclusion is partially predictable: firms with market capitalizations just below the S&P 500 cutoff face substantially higher inclusion probabilities than other firms. This predictability complicates the identification of treatment effects, as observed pre-announcement returns may reflect either genuine anticipation (violating (ref)) or selection on unobserved characteristics that drive both inclusion probability and returns.
To disentangle these effects, we pursue a two-pronged empirical strategy. First, we implement propensity score matching based on observable firm characteristics to account for selection on observables. Second, we employ synthetic control methods that match on pre-event returns, effectively controlling for unobserved factors that drive both selection and returns. The difference between these two approaches helps identify whether pre-announcement drift reflects anticipation or factor model misspecification.
Index inclusion predictability operates along two dimensions: the timing of additions (when inclusions occur) and the cross-section of selections (which firms are added). While ideally we would model both, we focus on cross-sectional predictability by estimating inclusion propensities based on observable firm characteristics. Specifically, we estimate annual logistic regressions:
where $\text{MktCapRank}_{i,y,m-1}$ is firm $i$'s market capitalization rank at the end of month $m-1$, and inclusion occurs in month $m$ of year $y$. Consistent with Greenwood2025Index, we find increasing predictability over time, with recent decades showing stronger relationships between lagged size and inclusion probability.
Using these propensity scores, we construct matched control groups via nearest-neighbor matching, creating portfolios of "pseudo-included" firms with similar inclusion probabilities but no actual inclusion. Under the assumption that selection between observationally equivalent firms is quasi-random, differences between included and pseudo-included firms should primarily reflect the causal effect of inclusion rather than selection bias.
To address selection on unobservables, we additionally implement the generalized synthetic control method of xu2022gsynth. For each announcement date, we estimate factor loadings using returns from days $-250$ to $-101$, deliberately excluding the immediate pre-announcement period where anticipation effects may contaminate estimation. We then construct synthetic control portfolios that match the pre-event return dynamics of included firms, examining the period from day $-100$ to $-15$.
This dual approach yields three distinct counterfactuals for cumulative abnormal returns (CARs): (i) simple market adjustment as in Greenwood2025Index (we also do CAPM and FF3F adjustments for completeness, but do not subtract $\alpha$ for reasons that will be clear shortly) (ii) propensity score-matched pseudo-included firms that control for selection on observables, and (iii) synthetic controls that account for selection on unobserved factors. Comparing these counterfactuals from day $-100$ through the announcement date allows us to decompose pre-announcement drift into components attributable to observable characteristics versus unobserved factor exposures. If drift persists after propensity score matching but disappears with synthetic controls, this would suggest that unobserved factors—rather than anticipation based on observables—drive the pre-announcement returns.
First, we find the pre-announcement drift as estimated by either the propensity score matched difference, or by the Gsynth approach drops significantly when compared to the market adjusted method. The effectiveness of Gsynth is quite striking in this setting, and suggests that longer-run cumulative effects can be substantially biased. What can explain the differences identified between these estimated methods? In Appendix (ref), we show that there is a substantial drift in our known factors across most decades. Considering the positive loadings in (ref), this suggests that the counterfactual return needs to sufficiently account for any and all potential unobserved factors driving the expected returns to avoid this bias highlighted in (ref).
Capturing all potential unobserved factors is not easy, however. In (ref), Panel A, we plot the average return in event time for the treated group, and then our five counterfactuals. We report the average daily return for each group, and note a 0.13 p.p. daily return for the included firms (prior to inclusion), an unusually high daily return. In contrast, the S&P500 has a daily return of only 0.03 p.p. during this period. This suggests that the included firms are quite unusual. Recall that if our factor choice in a linear model sufficiently spans the risk factors, we should estimate an alpha of zero, even in the presence of positive average returns. Our counterfactual models do an excellent job of matching the average return in the training and pre-periods. However, for the CAPM and FF3F, more than half of the average predicted return comes from just the intercept, $\alpha$. For gsynth, the estimated $\alpha$ is less than half, but still 0.06 p.p. Strikingly, the synthetic control counterfactual, which only takes a positive weighted average of control firms and does not include a constant, matches the pre-period return closely.
How should we interpret the estimated alpha in this linear factor models? There is presumably two components in an estimated factor model's alpha, true alpha, and model error:
where the model misspecification captures the return premium over this period that is not included in our model. In this setting, we view true alpha as zero, especially 280 days prior to the inclusion event. As a result, the positive alpha likely suggests model misspecification. The implications of this misspecification depend on the stability of this misspecification term. In Panel B of (ref), we see that the inclusion of alpha ensures that the various linear factor models do as well the synthetic control method in removing almost all pre-inclusion drift. This suggests that the trend beforehand is not due to front-running, but instead differential return profiles for included stocks. However, failing to include alpha for the CAPM, FF3F and gsynth fail to remove the pre-inclusion drift.
We examine acquirer returns around merger announcements using deal data from SDC Platinum. Following malmendier2018behavioral and savor2009stock, we implement several sample restrictions to ensure clean identification. We require targets to be classified as “Public,” “Private,” or “Subsidiary” and restrict to completed deals with all-cash or all-stock payment structures, as mixed consideration complicates the interpretation of market-timing effects. To ensure economic materiality, we require the target's pre-announcement market value to exceed 5% of the acquirer's market capitalization. We exclude repurchases, self-tenders, and minority stake purchases by requiring deal types to be “Disclosed Dollar Value” or “Undisclosed Dollar Value,” and mandating that acquirers hold less than 50% of the target six months before announcement.
We match acquirers to CRSP using six-digit CUSIPs, restricting to U.S. common shares (share codes 10 or 11) traded on NYSE, NASDAQ, or AMEX. Our event window spans days $-280$ to $+250$ relative to announcement. For the 20% of deals announced on non-trading days, we define $t = 0$ as the next trading day. Control firms comprise all CRSP-listed firms without contemporaneous merger announcements that have complete returns data over the event window. Our final sample contains 14,847 merger events across 6,625 unique dates, providing substantial variation in event timing for identification.
To assess whether event timing can be treated as random, we examine the distribution of factor returns on announcement days. Appendix (ref) shows that the distribution of daily market returns on S&P 500 index inclusion announcement days is virtually indistinguishable from the distribution on non-announcement days. The small-minus-big (SMB) factor exhibits similar distributional stability (unreported) These results support treating announcement timing as conditionally random with respect to factor realizations, satisfying a key identification assumption for our short-horizon analysis.
Table (ref) reports CAPM and Fama-French three-factor betas for firms with merger announcements, estimated using daily returns from days $-250$ to $-100$ relative to announcement.\footnote{We exclude the immediate pre-announcement period to avoid contamination from potential information leakage.} We also examine how the betas change following the announcement as well, and show that there are statistically significant changes after announcement, but they are small economically.
We first examine three-day announcement returns $[-1, +1]$ to test whether short-horizon estimates are robust to model specification, as predicted by (ref). We compare two approaches: market-adjusted returns using the CRSP value-weighted index (including distributions) following malmendier2018behavioral, and gsynth estimates using the generalized synthetic control method of xu2017generalized.
For the synthetic control approach, we estimate a separate model for each of the 6,625 event dates, treating all firms announcing mergers on that date (typically one or two firms) as the treatment group. We construct factor loadings using returns from days $-280$ to $-31$, excluding the immediate pre-announcement period to avoid contamination. Control firms consist of all CRSP securities without merger announcements that satisfy our data requirements.
(ref) reports cumulative abnormal returns by target type and payment method. Consistent with our theoretical predictions, the difference between market-adjusted and synthetic control estimates is economically negligible—less than 10 basis points in most specifications. This robustness to model choice confirms that short-horizon merger announcement effects are identified regardless of the specific factor adjustment employed.
In (ref), we plot the daily returns for the treated firms and the various control returns. Similar to the preannoucnment drift for the S&P index inclusion, the treated firm has significant daily returns prior in the year prior to the announcement, with 0.14 p.p. daily average returns (contrast with the market having 0.05 p.p. daily average returns). Again, the CAPM and FF3F models have significant alpha (0.08 p.p.), while gsynth has remarkably small alpha (0.02 p.p.). However, both gsynth and synthetic control do a poorer job matching the overall pre-period return, with 0.12 p.p. and 0.19p.p. returns respectively, relative to 0.14 for the treated group.
The path of the daily return line for the treated group spikes significantly on the event date, and then declines precipitously to a new steady state. It is worth remarking that within 30 days the event announcement, the treated firms' returns appear to line up almost exactly with the market returns. This is suggestive that there is a structural shift in the underlying return performance to these acquiring firms, perhaps due to change in true alpha, or perhaps due to factor loadings.
In (ref), we see the long-run implications of these alphas in in that the difference counterfactual predictions have wildly different long-run cumulative ATT a year after the event. In the literature, the presence of a negative post-acquisition event is often pointed to as evidence in favor of shleifer2003stock, but the modeling assumptions to make these types of assessments seem quite strong. This type of analysis also has implications for papers studying over- and under-reaction in the stock market.
What can we do to deal with this specification error? A crucial alternative is to examine the effect in a setting where treatment is as-if randomly assigned.
Our previous empirical examples rely on model-based identification strategies that require correct specification of the factor structure. To validate these approaches, we now examine a setting with quasi-experimental variation: close merger contests where assignment to treatment (winning the contest) is plausibly random conditional on observables. malmendier2018winning show that in protracted bidding contests, winners and losers are ex ante similar firms competing for the same target, making losers natural counterfactuals for winners.
This setting provides a unique opportunity to assess the performance of different estimators. Since losing bidders offer a design-based counterfactual, we can compare our model-based estimates (market adjustment, factor models, synthetic controls) against this quasi-experimental benchmark. Agreement between model-based and design-based estimates would validate our econometric approaches; divergence would suggest specification problems in the model-based methods.
Following malmendier2018winning, we analyze close merger contests defined as those with above-median duration.\footnote{We thank the authors for providing data on winning and losing bidders, announcement and completion dates, and contest duration.} Protracted contests involving multiple rounds of bids and counterbids suggest that participants had similar ex ante winning probabilities, supporting the identifying assumption that contest outcomes are quasi-random.
We construct event-time at the monthly frequency, with $t = 0$ marking the month-end before the initial bid announcement. The pre-contest period spans months $t = -35$ to $t = 0$. The contest period ($t = 1$) encompasses all months from initial bid through completion, averaging 361 days in our sample. The post-merger period runs from $t = 2$ to $t = 36$. This structure accommodates contests of varying duration while maintaining a consistent event-time framework.
We match contest participants to CRSP monthly returns, filling missing observations with market returns following malmendier2018winning. For synthetic control estimation, we augment the sample with all CRSP common shares (share codes 10 or 11) traded on NYSE, NASDAQ, or AMEX that have complete returns over the event window. This expanded control group allows the synthetic control algorithm to construct appropriate counterfactuals even when losing bidders may themselves be poor matches due to contest-specific shocks affecting all participants.
To check how well the as-if random counterfactual losing bidders match to winners, In (ref), we compare the risk exposures on different factors of winners and losers. We see that these two groups are relatively similar, although the losers have slightly higher HML beta than winners.
In (ref), we plot the cumulative ATTs for the winners relative to our various controls. Our benchmark is the “Loser” control, in solid blue. We see that in the pre-period, there is reasonable balance between the two groups, and then a small but significant decline following the announcement, suggesting a negative effect. Notably, this counterfactual has the smallest and least trending of the different counterfactuals. The only alternative model-based portfolio that is meaningfully close is the Gsynth control.
It is worth noting that the synthetic control methods have much larger declines during the merger battle, and afterwards as well. The abnormal return models fit well in the pre-period, but then predict significant and continuing declines in cumulative returns. Overall, this evidence suggests that most of the models (especially abnormal return) do poorly in the longer run, although gsynth is the exception.
This paper brings modern causal inference techniques to financial event studies, highlighting important limitations in standard approaches while providing constructive solutions. We demonstrate that traditional abnormal return estimators face inconsistency problems due to factor model misspecification -- a concern that becomes particularly severe in long-horizon analyses where small daily biases accumulate substantially over time.
While staggered event timing helps mitigate these issues in short-horizon studies by averaging out factor realizations, this solution proves inadequate for long-horizon analyses. The key insight is that misspecification bias compounds over longer horizons, regardless of how events are distributed across time.
Synthetic control methods offer a promising alternative by directly modeling counterfactual security paths without requiring correct specification of the underlying factor structure. Our empirical applications to political connections during market turbulence and S&P 500 index inclusions convincingly demonstrate the practical value of these methods.
Our findings suggest that many influential results based on long-horizon event studies may reflect factor model misspecification rather than genuine causal effects. We recommend that researchers employ synthetic control methods as a robust complement to traditional approaches, particularly when studying extended price responses or when events occur during periods of high market volatility.
\printbibliography