EconBase
← Back to paper

Comparing Experimental and Nonexperimental Methods: What Lessons Have We Learned Four Decades After LaLonde (1986)?

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

76,236 characters · 20 sections · 105 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

\setcounter{page}{1} \thispagestyle{empty}

flushleft\doublespacing { Comparing Experimental and Nonexperimental Methods: What Lessons Have We Learned Four Decades After LaLonde (1986)?}\\ { Guido W. Imbens and Yiqing Xu \footnote{Guido W. Imbens is an Applied Econometrics Professor and a Professor of Economics at the Graduate School of Business and the Department of Economics, Stanford University, Stanford, California. He is also a Research Associate, National Bureau of Economic Research, Cambridge, Massachusetts. Yiqing Xu is an Assistant Professor in Political Science at the Department of Political Science and a W. Glenn Campbell and Rita Ricardo-Campbell National Fellow at the Hoover Institution at Stanford University, Stanford, California. Their email addresses are [email removed] and [email removed].\\ For supplementary materials such as appendices, datasets, tutorials, and author disclosure statements, see the article page at https://doi.org/XXXX} }

\setcounter{footnote}{1}

\lettrine[lines=3]{I}{n} 1986, Robert LaLonde published a paper based on part of his PhD thesis LaLonde, which has had profound impact on both the methodological and empirical literatures on estimating causal effects. As of May 2025, this paper has been cited over 3,000 times, a number that only partially reflects its tremendous influence on the field of causal inference and the credibility revolution angrist2010.

The context for LaLonde's (1986) paper was the National Supported Work demonstration program. This program targeted individuals with extremely poor employment prospects: for females, recipients of Aid to Families with Dependent Children; for males, ex-drug addicts, ex-criminal offenders, and high school dropouts. Although about 10,000 people participated across 15 cities, the evaluation focused on approximately 6,600 participants in 10 cities. Approximately half were randomly assigned to a control group, while the other half received up to 12 months of training in small groups, with a supervisor and a counselor, followed by efforts to place them with outside employers. According to the Manpower Development Research Corporation (MDRC), which operated the program and conducted the evaluation, the National Supported Work program substantially increased 1978 earnings for female participants, particularly those without prior job experience; in contrast, the effects for male participants were smaller and, for some subgroups, essentially nonexistent (MDRC, mdrc1980).

LaLonde used this experimental evaluation to address a broader methodological question: whether the state-of-the-art nonexperimental evaluation methods of that time could replicate experimental benchmarks---that is, estimates obtained from randomized controlled trials. By nonexperimental methods, we mean statistical approaches that estimate a program’s impact using naturally occurring variation in treatment assignment, without randomization or experimental control. LaLonde's conclusion was sharply negative regarding the credibility of the nonexperimental methods he examined. He wrote LaLonde:

quoteThis comparison shows that many of the econometric procedures do not replicate the experimentally determined results, and it suggests that researchers should be aware of the potential for specification errors in other nonexperimental evaluations.

and concluded LaLonde:

quote[P]olicymakers should be aware that the available nonexperimental evaluations of employment and training programs may contain large and unknown biases resulting from specification errors.

Around the time LaLonde’s paper was published, several other studies made similar points about the credibility of nonexperimental methods, including work by LaLonde’s thesis advisers ashenfelter1985using, as well as fraker1987adequacy and heckman1987we.

One reason LaLonde’s study became so influential is that the full data underlying the original analysis later became publicly available. As part of a research project initiated in a 1996 graduate class taught by Imbens and Rubin at Harvard, dehejiawahba, dehejia2002propensity retrieved LaLonde’s original male subsample data—stored on tapes—by locating a tape reader capable of reading the files. The data are now publicly available on Dehejia’s website (\url{https://users.nber.org/ rdehejia/data/.nswdata2.html}). We refer to a subsample of these data that is most commonly used in the literature as the LaLonde-Dehejia-Wahba (LDW) data. More recently, calonico2017women also reconstructed the female samples. These public datasets are highly valuable for both teaching and future research and are provided alongside this paper, with code to replicate the estimates reported here.

The methodological literature on nonexperimental evaluation has advanced significantly. Nearly four decades after LaLonde’s paper, the answer to his original question—whether nonexperimental methods can successfully replicate experimental benchmarks—is more nuanced than his initial conclusion: sometimes they can, and we now have better tools to help assess when they are likely to succeed.

Here are five key lessons that emerge from this literature since the publication of LaLonde. First, the unconfoundedness assumption has become central to modern methods for nonexperimental data. Unconfoundedness means that, conditional on pre-treatment variables or covariates, treatment is as if randomly assigned. This assumption elegantly separates the underlying identification assumptions from functional form considerations. It emphasizes design—how treatment is assigned—over full specification of the data-generating process.

Second, inspecting and ensuring overlap is critical in the process of estimating treatment effects. Overlap means that the covariate distributions in treatment and control groups have common support. Improved overlap reduces sensitivity to the choice of estimation strategy and improves the robustness of estimates.

Third, the use of propensity scores, defined as the probability of receiving treatment given observed covariates, has become widespread. The propensity score, introduced in rosenbaum1983central, had only just entered the economics literature when LaLonde was writing his thesis. Since then, propensity scores have become a core component of many analyses that rely on unconfoundedness, including for assessing overlap, but also for estimation.

Fourth, researchers have increasingly taken heterogeneity in treatment effects seriously. This includes both methods for estimating average treatment effects in the presence of heterogeneity, but also going beyond average treatment effects to examine estimates of treatment effect heterogeneity, such as conditional treatment effects and quantile treatment effects, {\it e.g.}, wager2018estimation. These approaches provide a deeper understanding of how the treatment impacts different individuals or segments of the outcome distribution.

Finally, validation exercises—particularly placebo tests—are essential for assessing key assumptions and evaluating the credibility of causal claims. In a placebo test, researchers use an outcome that should not be affected by the treatment. For example, LaLonde used 1975 earnings as a placebo outcome, since it predates the training program. A non-zero estimated effect on the placebo outcome indicates potential unobserved confounding. Placebo tests are important diagnostic tools in modern nonexperimental research since the credibility revolution angrist2010. \footnote{We focus on five issues specific to the setting studied in LaLonde: the evaluation of an individual-level intervention using detailed background information on participants. While the broader literature on causal inference has expanded substantially over the past four decades currie2020technology, we do not attempt to cover that broader landscape here. For more comprehensive overviews, see surveys such as imbens2009recent and abadie2018econometric, as well as a number of textbooks angrist2008mostly, imbens2015causal, cunningham2018causal, huntington2021effect, huber2023causal, ding2024first, wager2024, chernozhukov2024applied.}

To illustrate these lessons, we reexamine the original and reconstructed LaLonde datasets. We show that, once overlap is ensured, various modern methods yield similar estimates. However, these estimates lack a causal interpretation if unconfoundedness is violated. Validation exercises such as placebo tests are critical for assessing the credibility of unconfoundedness. For the LaLonde datasets, placebo estimates generally do not support the unconfoundedness assumption.

LaLonde's Findings

We begin by describing the data used in LaLonde and outlining the main econometric approaches and results presented in the paper. We then introduce the LaLonde-Dehejia-Wahba data and examine some of the immediate methodological responses to LaLonde’s findings.

LaLonde's Data

Although the MDRC evaluation included 6,600 individuals, the experimental samples LaLonde used to evaluate nonexperimental methods were much smaller. For males, the sample included 722 individuals (297 treated and 425 controls), and for females, 1,158 individuals (600 treated and 585 controls). This substantial reduction in sample size stemmed from three main factors. First, many participants failed to complete follow-up interviews needed to collect post-program earnings data. Second, due to budget constraints, the MDRC team randomly selected subsamples for the 27- and 36-month follow-up interviews. Third, LaLonde excluded male participants who entered the program before January 1976 or were still enrolled in January 1978.

To conduct a nonexperimental analysis of the effects of the training program, LaLonde combined the experimental treatment group with comparison groups drawn from external, nonexperimental datasets used as controls. For both female and male participants, he used two main sources. The first, CPS-SSA-1, is based on Westat's Matched Current Population Survey–Social Security Administration File and includes all females or males under age 55 who met Westat's eligibility criteria. The second, PSID-1, comes from the Panel Study of Income Dynamics and includes all female or male household heads under age 55 from 1975 to 1978 for males and 1979 for females, excluding those who identified as retired in 1975. The age cutoff was intended to improve comparability between the comparison group and the experimental sample.

table[table omitted — 536 chars of source]

Table (ref) columns 1–4 present summary statistics for the male samples in the experimental data and the comparison groups used in LaLonde. Respondents in both the CPS and PSID comparison groups were, on average, older, more educated, substantially less likely to be high school dropouts, more likely to be married, and less likely to be Black or Hispanic than participants in the National Supported Work experiment. They also had much higher earnings in the years before the program. For example, average 1975 earnings were \$3,000 in the experimental data, compared to \$14,000 in the CPS-SSA-1 sample and \$19,000 in the PSID-1 sample. Notably, the comparison groups are much larger: the male CPS-SSA-1 sample includes 15,922 observations, and the male PSID-1 sample contains 2,490 observations. To improve comparability with the experimental sample, LaLonde dropped from these datasets individuals based on employment status, time of survey, and poverty status, creating four additional, smaller comparison groups.

Columns 5 and 6 of Table (ref) present summary statistics for the experimental treated and control units in the LaLonde-Dehejia-Wahba data (discussed further below), a subset of the original LaLonde male sample selected for containing information on 1974 earnings. Our reanalysis focuses on this dataset because it is widely used in the post-LaLonde methodological literature and includes two pretreatment outcomes—earnings in 1974 and 1975—which enable adjustment for longer earnings histories and placebo analyses. Results for the original LaLonde male samples and reconstructed female samples are reported in the online appendix.

Econometric Approaches in LaLonde (1986)

To estimate the causal effect of the National Supported Work program on 1978 earnings using both experimental and nonexperimental data, LaLonde employed a variety of models that fall into two broad categories: regression methods, in which earnings serve as the outcome (referred to as the “earnings equation”), and selection models that also include a “participation equation,” where the outcome is program participation. In LaLonde's paper, Tables 4 and 5 present the regression results, while Table 6 reports the selection model results. These tables are reproduced in the appendix.

LaLonde reported training effect estimates for both female and male participants using seven estimators across six comparison groups: two large groups described in Table (ref) and the aforementioned four subgroups defined through additional selection criteria to improve overlap with the National Supported Work demonstration sample. All regression models were linear and implicitly assumed constant treatment effects. Of the seven models, two are simple regressions estimated without or with controls for age, education, and race (but notably excluding 1975 earnings). Two models use a difference-in-differences estimator, replacing the outcome with the change in earnings between 1975 and 1978, estimated with and without controlling for age. Two more use a quasi-difference-in-differences specification, including 1975 earnings on the right-hand side to account for transitory shocks—known as the “Ashenfelter dip” ashenfelter1978estimating—again estimated with and without controlling for age. The final specification includes all available pretreatment covariates, including 1975 earnings, 1975 unemployment status, and marital status.

In an additional analysis in the spirit of the modern causal literature, LaLonde also conducted placebo tests using 1975 earnings as the outcome. Because the training occurred after 1975, the true treatment effect on this placebo outcome is zero. And indeed, the placebo estimates from the linear regression approach on the experimental sample were close to zero.

LaLonde emphasized several key findings from the linear regression results. First, using the experimental data, all seven estimators produced similar estimates: around \$851 for female participants and \$886 for male participants. Second, when using nonexperimental comparison groups, the estimates diverged sharply from these benchmarks, often yielding large, negative values in both female and male samples, with modest standard errors. Third, the estimates from nonexperimental data varied widely across specifications, and goodness-of-fit tests offered little guidance for selecting models that aligned with the experimental results. Taken together, these findings led LaLonde to conclude that the regression adjustment methods commonly used at the time were not credible when applied to nonexperimental data.

In addition to the regression models, which rely on exogeneity of the treatment indicator, LaLonde also presented results from selection models that allow for potential endogeneity. These models use the two-step estimator proposed by heckman1978, which permits correlation between the error terms in the earnings and participation equations. Identification in this framework relies either on exclusion restrictions—variables that appear in the participation equation but not in the earnings equation—or on functional form and distributional assumptions (linearity and joint normality of the error terms). For both female and male samples, LaLonde used three comparison groups (the experimental controls, CPS-SSA-1, and PSID-1) and estimated four specifications for females and three for males. Each specification used a different set of excluded variables. {\it A priori}, no single specification was clearly more defensible based on economic or econometric reasoning.

LaLonde found that the selection model estimates using the experimental data remained close to \$851 for females and \$886 for males. However, those using nonexperimental data again varied substantially and deviated from the experimental benchmarks. He concluded that, although the two-step procedure brought the estimates somewhat closer to the benchmarks, it still yielded a “considerable range of imprecise estimates” LaLonde.

The LaLonde-Dehejia-Wahba Data

The data reconstructed by dehejiawahba focused exclusively on male participants, noting that “estimates for this group were the most sensitive to functional-form specification” dehejiawahba. They constructed a subsample from LaLonde’s original data consisting of individuals with available information on 1974 earnings and unemployment status. They argued that this subsample remained a valid experimental sample because it was constructed solely using pretreatment information—such as month of assignment and employment history—thereby preserving the orthogonality of treatment assignment with respect to observed and unobserved characteristics. Notably, this subsample includes only 62 percent of the original treatment group in LaLonde’s study. For nonexperimental controls, they used subsets of the same datasets LaLonde employed, restricted to units with 1974 earnings and unemployment data. This collection, now widely known as the LaLonde-Dehejia-Wahba (LDW) data, has become a standard benchmark in the causal inference literature. In particular, most methodological studies focus on the version combining the experimental treated units with CPS-SSA-1 controls. For brevity, we refer to the LaLonde-Dehejia-Wahba experimental sample as LDW-Experimental, and the samples combining these treated units with CPS-SSA-1 and PSID-1 controls as LDW-CPS and LDW-PSID, respectively.

As noted earlier, columns 5 and 6 of Table (ref) report summary statistics for the LDW-Experimental sample. Compared to LaLonde's original male sample, participants in this subsample had higher unemployment rates and lower average earnings in 1975. The inclusion of 1974 earnings---available only in the LaLonde-Dehejia-Wahba data---further suggests that many participants faced long-term unemployment. These differences may help explain why the estimated training effect in this sample (\$1,794) is more than twice that of LaLonde's original male sample (\$886).

Subsequent Literature

The publication of LaLonde sparked a debate in the applied econometrics literature. For example, heckman1989choosing responded to LaLonde’s critique of nonexperimental evaluation methods by advocating the use of specification tests to rule out particularly poor estimators. However, this approach did not offer a clear way to distinguish among the many estimators that fit the data reasonably well but rely on different identification assumptions. As a result, it gained limited traction in subsequent work.

In contrast, dehejiawahba proposed an alternative, more flexible estimators to address LaLonde’s challenge. Their proposals included approaches based on propensity score stratification and matching. Their estimates closely matched the experimental benchmark, leading them to conclude, in sharp contrast to LaLonde:

quote“[T]he estimates of the training effect for LaLonde's ... dataset are close to the benchmark experimental estimates and are robust to the specification of the comparison group and to the functional form used to estimate the propensity score. ... our methods succeed for a transparent reason: They use only the subset of the comparison group that is comparable to the treatment group, and discard the complement.” dehejiawahba.

The contrast between the conclusions of dehejiawahba and LaLonde led to a wave of methodological research aimed at probing and reconciling these findings. We turn to these developments in the next section.

Methodological Improvements since LaLonde

To facilitate discussion of methodological advances since LaLonde, we begin with the “potential outcomes framework,” then introduce the main causal estimand and a closely related statistical estimand. We then outline the key assumptions required for nonexperimental data to identify the causal estimand.

Causal and Statistical Estimands

Over the past four decades, the applied econometrics literature has made significant progress—most notably in clarifying estimands (what is being estimated) and the assumptions required for identification (how to use data to recover them). Once these assumptions are clearly stated, researchers can assess how plausible they are.

In our discussion, we adopt the potential outcome framework, originally developed by Jerzy Neyman in the context of randomized experiments neyman1923 and extended to nonexperimental settings by Donald Rubin rubin1974estimating, rubin2006matched, imbens2015causal. For each individual $i$, two potential outcomes are postulated: $Y_i(0)$ and $Y_i(1)$. In the LaLonde setting, they represent the individual's 1978 earnings had they not participated in the program and had they participated, respectively. The causal effect for individual $i$ is the difference $Y_i(1) - Y_i(0)$. Let $W_i$ denote the binary treatment indicator, equal to 1 if the individual participated in the program and 0 otherwise. The observed outcome is linked to the potential outcomes by the identity $Y_i = (1 - W_i) Y_i(0) + W_i Y_i(1)$. We also observe pretreatment characteristics or covariates $X_i$. In the LaLonde study, these include age, years of schooling, high school dropout status, marital status, and indicators for Black and Hispanic backgrounds.\footnote{The modern methodological literature has paid particular attention to settings where the vector of pretreatment variables is high-dimensional.} Following dehejiawahba, we may augment this vector to include 1974 and 1975 earnings and indicators for zero earnings (unemployment) in each year.

Our primary goal is to estimate the average treatment effect on the treated (ATT), \[ ATT = \mathbb{E}[Y_i(1) - Y_i(0) \mid W_i = 1]\ = \ \mathbb{E}[Y_i \mid W_i = 1] - \mathbb{E}[Y_i(0) \mid W_i = 1]. \] ATT represents the individual-level causal effect averaged over for those who actually received the treatment.\footnote{For simplicity, here we do not distinguish between a finite study population and a super-population.} In LaLonde's context, it means the average treatment effect of the job training program on individuals who actually participated. Since $Y_i(1) = Y_i$ for treated individuals (individuals with $W_i = 1$), the second equality follows. The first term on the right-hand side, $\mathbb{E}[Y_i \mid W_i = 1]$, can be directly estimated from the observed data. The second term, $\mathbb{E}[Y_i(0) \mid W_i = 1]$, cannot, because it involves counterfactual outcomes—in this case, what participants' earnings would have been, had they not participated in the program.

Most analyses of the LaLonde data that explicitly allow for heterogeneous treatment effects focus on the ATT, as it makes little sense to estimate or even contemplate the effect of the program for nonparticipants with stable jobs and high earnings. In other contexts, researchers may be interested in the average treatment effect (ATE), which captures the average effect across the entire population while the ATT focuses specifically on the treated group. LaLonde did not explicitly distinguish between different estimands, as his analysis did not consider treatment effect heterogeneity.

Because we cannot observe the counterfactual outcomes, we cannot directly estimate the ATT. This is the core of what holland1986statistics termed the “fundamental problem of causal inference.” To make progress, we can instead estimate the covariate-adjusted difference in average outcomes between treated and control groups: \[ \text{(Covariate-adjusted difference)}\quad \mathbb{E}[Y_i \mid W_i = 1] - \mathbb{E}\left\{\mathbb{E}[Y_i|W_i=0,X_i]\mid W_i=1\right\}. \] We refer to this as a statistical estimand, as opposed to a causal estimand, because it can be estimated from observed data when overlap holds, an assumption we return to below. In LaLonde's context, the statistical estimand represents the difference between the average 1978 earnings of individuals who participated in the program and the average 1978 earnings of nonexperimental individuals with similar observed characteristics.

Unlike a statistical estimand, the ATT is a causal estimand because it compares potential outcomes for the same individuals—specifically, those who participated in the program. These two estimands are not necessarily equal. The counterfactual mean for the treated individuals, $\mathbb{E}[Y_i(0) \mid W_i = 1]$, may differ from $\mathbb{E}\left\{\mathbb{E}[Y_i \mid W_i = 0, X_i] \mid W_i = 1\right\}$, a weighted average of observed outcomes for individuals in the nonexperimental control group with similar characteristics. For the statistical estimand to recover the ATT, a key assumption—unconfoundedness—must hold. That is, if unconfoundedness holds, we can approximate the average untreated outcome for the treated individuals using outcomes from control individuals of similar observed characteristics.

A substantial body of subsequent research has focused on improving methods for estimating the covariate-adjusted difference, particularly in high-dimensional settings. Many recent approaches draw on machine learning techniques to flexibly estimate the relationships between covariates, treatment, and outcomes. Formal results often rely on additional regularity conditions, such as the smoothness of conditional means and propensity scores. For formal treatments, see the references in imbens2009recent and abadie2018econometric. The methodological advances in this literature apply well beyond the LaLonde setting.

Our focus here is on the plausibility of the key assumptions. In particular, if unconfoundedness does not hold, we may still be able to robustly and precisely estimate the covariate-adjusted difference, but it cannot be interpreted as a causal effect and may have little substantive relevance.

Unconfoundedness

The unconfoundedness assumption was first explicitly introduced by rubin1978bayesian as part of the “ignorable treatment assignment” concept, see also rosenbaum1983central. If unconfoundedness holds, then treatment assignment can be viewed as effectively random once differences in observed covariates between the treatment and control groups are adjusted for. This assumption plays a central role in identifying causal effects from nonexperimental data. In the context of estimating the ATT, unconfoundedness is formally stated as: \[ \text{(Unconfoundedness)}\qquad W_i\ \perp\!\!\!\perp \ Y_i(0)\ |\ X_i\ , \] which means that treatment status is conditionally independent of the control potential outcome given the covariates. In LaLonde’s context, unconfoundedness means that once we account for observable characteristics (including age, education, race, marital status, and prior earnings), whether someone was in the experimental treatment group or nonexperimental comparison group provides no additional information about what their 1978 earnings would have been had they not participated in the program.

Unconfoundedness is also referred to as conditional independence lechner1999earnings, lechner2002program, or informally as {\it exogeneity} imbens2004, or {selection on observables} barnow1980issues. However, its definition departs from traditional econometric definitions of exogeneity which are expressed in terms of residuals and specific functional forms. By contrast, the formal statement of unconfoundedness avoids functional form assumptions, focusing instead on the treatment assignment mechanism. It allows researchers to separate the essence of the identification assumptions from functional form considerations, emphasizing the role of design rather than the full specification of the data-generating process.

A key result from rosenbaum1983central shows that unconfoundedness implies that conditioning on the scalar propensity score is sufficient to remove bias from covariate imbalance. This dimensionality reduction—from the full vector $X_i$ to a single index—has made the propensity score a cornerstone of many modern estimators.

When the parametric model for the conditional expectation of the outcome—such as the earnings equation used in LaLonde—is correctly specified, unconfoundedness implies a zero conditional mean for the error term. Thus, the results reported in Tables 4 and 5 of LaLonde can be viewed as relying on a combination of unconfoundedness and correct functional form assumptions. At the time, however, the nonparametric framing of the identifying assumptions and the emphasis on assignment mechanisms were not yet standard in applied work.

In practice, unconfoundedness is a strong assumption, and its plausibility depends heavily on the context. When the treatment assignment mechanism is poorly understood, this assumption may not be credible. Nonetheless, often researchers can assess its plausibility through supplementary analyses. Tools such as placebo tests and sensitivity analyses can help probe the credibility and robustness of causal claims that rely on unconfoundedness. For a general discussion, see rosenbaum1983central and imbens2004.

In the next section, we illustrate the use of placebo tests to evaluate unconfoundedness. In the LaLonde setting, the ten covariates are clearly pretreatment variables and should be included in any adjustment strategy. In other settings, however, whether a given covariate should be adjusted for is less obvious. rosenbaum1984consequences cautions against adjusting for post-treatment variables, and cinelli2022crash offer guidance on selecting from among valid pretreatment covariates for causal inference.

Overlap and Balance

To identify the ATT—and to ensure that the statistical estimand is properly defined—we require an overlap assumption, which states that the propensity score is strictly less than one: \[ \text{(Overlap)}\qquad\Pr(W_i = 1 \mid X_i) < 1. \] In LaLonde's context, this means that for any individual $i$ assigned to the treatment group with a covariate profile $X_i = x_0$, there must also be individuals in the nonexperimental comparison group with the same profile; otherwise, $\Pr(W_i = 1 \mid X_i = x_0) = 1$, violating the overlap assumption. Overlap ensures that the weighted average $\mathbb{E}\left\{\mathbb{E}[Y_i \mid W_i = 0, X_i] \mid W_i = 1\right\}$, the second term in the statistical estimand, is well-defined. When overlap is violated, it can be restored by trimming the treatment group—that is, by removing treated units whose covariate profiles are not represented in the control group.\footnote{If the target parameter is the ATE, a stronger version of the overlap assumption is needed: the propensity score must lie strictly between zero and one for all units.}

Overlap implies that treated and control units must share common support in their covariate distributions. Without overlap—for instance, when some covariate values appear in the treatment group but not in the control group—it becomes difficult to make credible comparisons because estimates rely on extrapolation. Overlap is conceptually distinct from balance, which refers to the similarity in covariate distributions across groups. In a randomized control trials overlap holds by design, and balance is achieved in expectation. In that case, balance can be further improved through stratification or post-stratification. In nonexperimental settings, ensuring overlap is a key condition for credible estimation of causal effects.

Overlap is especially important when researchers do not wish to impose strong functional form assumptions on the conditional means of potential outcomes or on the structure of treatment effect heterogeneity. When the number of covariates is small, overlap can be assessed by examining marginal or joint covariate distributions across treatment groups. But this becomes impractical in high-dimensional settings. In such cases, it is more effective to inspect the distribution of estimated propensity scores across treated and control groups. Lack of overlap in covariate distributions implies, and is implied by, a lack of overlap in propensity score distributions.

LaLonde did not explicitly discuss overlap, nor did he assess it beyond reporting covariate means by treatment status. Both the regression and selection models he used rely on correct functional form assumptions, which permit interpolation or extrapolation of treatment effects even when treated and control units differ substantially in their covariates—thus bypassing the need for overlap. Still, LaLonde clearly recognized the potential problem: to improve comparability, he trimmed the comparison groups based on “characteristics [that] are consistent with some of the eligibility criteria used to admit applicants into the NSW program” LaLonde. However, by modern standards, his trimming procedures—such as removing all men working in March 1976 in one subset (CPS-SSA-2) or further excluding unemployed respondents with 1975 incomes above the poverty line in another (CPS-SSA-3)—are {\it ad hoc} and certainly do not guarantee overlap on all relevant covariates.

Over the past four decades, researchers have developed more systematic approaches for improving overlap, often using the propensity score. These methods vary depending on whether the goal is simply to ensure overlap or to further improve covariate balance.

Improving overlap or balance typically requires dropping some units from the sample. Although this reduces the sample size, by ensuring better balance it may improve precision for estimates of the average treatment effect. In addition, the gain in robustness and reduction in bias often outweigh any loss in precision. In practice, any increase in variance, even from substantial trimming is typically modest.\footnote{For example, suppose we have a sample with $N_{tr}$ treated units and $N_{co}$ control units. Under homoskedasticity and random assignment, the variance of the difference-in-means estimator is $\sigma^2(1/N_{tr} + 1/N_{co})$. In the LDW-CPS sample, starting with $N_{tr} = 185$ and $N_{co} = 15,922$, dropping 15,737 control units—a 99% reduction—raises the standard error by only about 30% in the “best-case scenario,” which assumes no bias from including the additional controls.}

Several specific trimming strategies have been proposed. For instance, focusing on overlap for ATT estimation, dehejiawahba dropped control units with estimated propensity scores below the minimum observed in the treated group. crump2009dealing proposed a more aggressive approach, selecting subsamples that minimize the variance of ATE estimates and recommending trimming units with propensity scores outside the $[.1, .9]$ interval. crump2006moving and li2018balancing proposed improving balance through propensity score weighting, introducing “overlap weights” proportional to the product of the propensity score and one minus the propensity score. Another effective strategy—especially when targeting the ATT—is to match each treated unit to a control unit with a similar estimated propensity score. This approach not only ensures overlap but also tends to improve balance across the covariate distributions.

In practice, overlap—like unconfoundedness—is essential for credible estimation. In settings with poor overlap, such as the LaLonde nonexperimental samples, trimming the sample to ensure overlap is often more important than the choice of estimation method.

Estimation Given Unconfoundedness and Overlap

All estimators in LaLonde are linear in the covariates. Since then, a wide range of more flexible methods have been proposed to estimate average causal effects under unconfoundedness and overlap. These approaches can be broadly categorized into three groups: (i) outcome modeling, including linear regressions, (ii) methods that directly adjust for covariate imbalance, including those based on propensity scores, and (iii) doubly robust methods.

First, outcome modeling remains the most widely used approach among applied researchers. It typically involves regressing the outcome on the treatment indicator and covariates (usually, the level terms)—what LaLonde called the earnings equation. This method assumes linearity in covariates and constant treatment effects. A modest relaxation of this approach involves estimating two separate linear regressions for the treated and control groups, sometimes referred to as the Oaxaca-Blinder estimator kline2011oaxaca. More flexible alternatives include semiparametric and nonparametric methods to model the conditional means of the potential outcomes heckman1997matching, athey2019generalized.

The second group of methods focuses on directly adjusting for covariate imbalance between the treatment and control groups. This includes blocking on covariates ({\it i.e.}, grouping units with similar characteristics and comparing outcomes within groups, and then aggregating), covariate matching abadie2006, abadie2008failure, abadie2011bias, abadie2016matching, diamond2013genetic, rubin2006matched, imbens2015, and weighting methods to achieve covariate balance hirano2003efficient, hainmueller, zubizarreta2023handbook, zubizarreta2015stable.

Covariate matching is a nonparametric method that avoids imposing modeling assumptions. However, it suffers from the curse of dimensionality when many covariates are present, making it prone to large biases abadie2006. In such settings, adjusting for the estimated propensity score is often more practical. This can be implemented through blocking and matching dehejiawahba, abadie2011bias or through inverse propensity weighting (IPW). IPW reweights observations by the inverse of their estimated propensity score, creating a pseudo-population in which treatment assignment is uncorrelated to observed covariates. hirano2003efficient show that the Hájek variant of the IPW estimator can achieve the semiparametric efficiency bound even when the propensity score is estimated nonparametrically. In the LaLonde setting, the IPW estimator reweights the nonexperimental control group based on estimated propensity scores so that its covariate distribution closely approximates that of the experimental treated group.

Since covariate imbalance is the sole source of bias under unconfoundedness, scholars have developed methods to improve balance either by refining propensity score estimation or by bypassing it entirely. For instance, imai2014covariate propose estimating a covariate-balancing propensity score using the generalized method of moments, whereas hainmueller introduces entropy balancing to directly achieve balance in specified covariate moments. Entropy balancing can be viewed as an IPW estimator that implicitly relies on a correctly specified propensity score model, where the link function is logistic and the log-odds are linear in covariates zhao2017entropy.

However, neither outcome modeling nor matching or weighting methods on their own are currently most favored in the methodological literature. Instead, hybrid approaches combine outcome modeling, such as regression, with techniques that address covariate imbalance, such as propensity score weighting, to leverage the strengths of both. Examples include regression within propensity score blocks rosenbaumrubin1983assessing, imbens2015, matching followed by regression adjustment rubin1973use, abadie2011bias, and methods that integrate weighting with regression robins1994estimation, robins1995semiparametric. These approaches are motivated by the fact that even if balancing or propensity score methods are consistent and efficient in large samples, combining them with outcome modeling can reduce small-sample bias and improve precision. For instance, in high-dimensional settings, the bias from a matching estimator due to remaining covariate imbalance may dominate the variance, and regression adjustment that accounts for this imbalance can help reduce the bias abadie2011bias.

robins2001comment introduced the term double robustness, a key concept for hybrid methods. They show that if either the propensity score or the outcome model is correctly specified, the augmented inverse propensity weighting (AIPW) estimator—which combines propensity score weighting and regression—is consistent. AIPW can be viewed as an outcome model augmented by a correction term: an IPW estimator applied to the residuals from the outcome model, rather than the raw outcomes. The double robustness property arises because, if the outcome model is correctly specified, the correction term has mean zero even if the propensity score model is misspecified; if the outcome model is misspecified, a correctly specified IPW component can debias the regression model. Beyond robustness, AIPW is also desirable for its efficiency: when both models are correctly specified, it achieves the semiparametric efficiency bound bang2005doubly.

More recently, machine learning methods have become increasingly popular in applied causal inference van2011targeted, wager2017estimation, chernozhukov2017double, athey2018approximate, athey2019generalized; see athey2017state, athey2019machine for reviews. These methods are particularly useful for estimating “nuisance parameters”---such as propensity scores or conditional outcome means---that are not of direct interest but are are essential for identifying causal effects. In the LaLonde setting, the goal is to estimate the ATT on post-training wages, which requires modeling the propensity score and/or the conditional means of potential outcomes. Many modern estimators satisfy the “Neyman orthogonality” condition chernozhukov2017double, which reduces the effect of small estimation errors in the nuisance parameters on estimates of the target parameter. In binary treatment settings like LaLonde’s, these estimators resemble the AIPW estimator introduced earlier, but with the nuisance parameters estimated via flexible machine learning methods instead of parametric models. chernozhukov2017double, chernozhukov2018double emphasize that this property ensures valid inference even when using machine learning algorithms that converge more slowly than is required for estimators based on only estimating the conditional outcome distributions or the propensity score.

Alternative Estimands and Heterogeneous Treatment Effects

Much of the methodological and applied research has focused on estimating average causal effects, such as the ATT. However, other quantities may also be of interest. For example, researchers often seek to understand treatment effect heterogeneity among treated units, which can shed light on mechanisms, improve evaluations of effectiveness, and guide personalized policy design. For example, the MDRC team reported that the National Supported Work program had a large and positive effect on female participants, a significant impact on ex-addict male participants, a small and highly variable impact on ex-criminal male participants, and almost no effect on youth participants (MDRC, mdrc1980). Econometrically, exploring treatment effect heterogeneity given observed characteristics involves estimating the conditional average treatment effects on the treated (CATT). Machine learning methods have been proposed to estimate CATT either nonparametrically or using low-dimensional representations, such as causal forests, while still permitting valid inference or error bounds athey2016recursive, wagerathey, athey2019generalized.

Another important, though less frequently used, estimand is the quantile treatment effects—defined as the difference between quantiles of the treated and untreated potential outcome distributions, either for the population or for the treated group. Under unconfoundedness and overlap, the full marginal distributions of potential outcomes are identified, enabling identification of quantile treatment effects. firpo2007efficient proposes a semiparametrically efficient inverse propensity weighting estimator for these quantities.

Validation through Placebo Analyses

While researchers can assess overlap using observed data, the unconfoundedness assumption is not directly testable. To evaluate the credibility of treatment effect estimates, the literature has developed two main strategies: placebo analyses and sensitivity analyses. Here, we focus on the former and relegate discussion of the latter to the online appendix.

Placebo analyses offer an indirect way to probe the plausibility of unconfoundedness. These analyses typically involve estimating a model similar to the main specification, but replacing the outcome variable with a pseudo-outcome—usually a variable known to be unaffected by the treatment. A common approach is to test for a treatment effect on a pretreatment variable, such as a lagged outcome, which should not be influenced by the treatment but may still correlate with unobserved confounders. Another variant tests the effect of a pseudo-treatment—often a proxy for the treatment—on the actual outcome. This strategy is often implemented using multiple control groups. For further discussion, see rosenbaum1987role, imbens2015causal, and imbens2015.

Although the term “placebo test” was not yet in use in the mid-1980s, LaLonde conducted such an analysis. He regressed 1975 earnings, a pretreatment variable, on the treatment indicator and covariates (reported in columns 2 and 3 of his Tables 4 and 5). Using nonexperimental data, he found that many of the estimated effects were large, negative, and statistically significant, suggesting a violation of unconfoundedness. One limitation of the LaLonde data is that it contains only a single pretreatment outcome. In contrast, the LaLonde-Dehejia-Wahba subset of the LaLonde data allow for placebo analyses that condition on an earlier pretreatment variable, 1974 earnings, when testing for associations between treatment and 1975 earnings. In general, when multiple pretreatment periods are available, placebo tests can be both statistically more powerful and substantively more credible.

Formally, a placebo test assesses a conditional independence restriction, which implies that the placebo outcomes are unrelated to treatment assignment once we account for the remaining observed covariates. A limitation of LaLonde’s original test is that it checks only one implication of the full conditional independence assumption: whether the average outcomes, after adjusting for observed covariates, are the same across groups. imbens2015 discusses other testable implications of this assumption.

Reanalyzing the LaLonde Data

To demonstrate the methodological advances since LaLonde in practice, we revisit the LaLonde data, including the LaLonde-Dehejia-Wahba data, the original LaLonde male samples, and the Lalonde-Cal{\'o}nico-Smith female samples. Our primary focus is on the LaLonde-Dehejia-Wahba data, while results for the other two datasets are reported in the online appendix. For all three datasets, overlap is a central concern. The ATT and CATT are our primary causal estimands.

Propensity Scores and Overlap

We focus on the LaLonde-Dehejia-Wahba (LDW) data because it includes earnings and employment information from 1974. Our analysis uses three LDW datasets: (1) LDW-Experimental, consisting of 185 treated and 280 control individuals from the experimental sample; (2) LDW-CPS, which includes the same treated individuals and 15,922 controls from CPS-SSA-1; and (3) LDW-PSID, comprising the same treated individuals and 2,490 controls from PSID-1. As noted earlier, LaLonde constructed smaller subsamples to improve comparability, but we instead rely on more modern, data-driven methods to assess and address overlap issues.

First, we estimate propensity scores using generalized random forest athey2019generalized, a machine learning method that flexibly models conditional probabilities. Unlike simple logistic regression, GRF can capture complex nonlinearities and higher-order interactions in the covariates, potentially yielding more accurate estimates of the propensity scores.

Figure (ref) (A)–(C) use these estimated scores to assess overlap between treated and control units in each sample. Each panel plots histograms of the log-odds of the estimated propensity scores, defined as $\log\left(\hat{e}/(1-\hat{e})\right)$, where $\hat{e}$ is the estimated propensity score. We use the log-odds scale because it more effectively distinguishes differences at the tails of the distribution. For reference, a log-odds of –3 corresponds roughly to a probability of 0.05, and by symmetry, a log-odds of 3 corresponds to a probability of about 0.95.

figure[figure omitted — 1,902 chars of source]

In panel (A), the LDW-Experimental sample shows near-perfect overlap: the treated and control groups have closely aligned propensity score distributions. In contrast, panels (B) and (C) reveal severe overlap problems in the nonexperimental samples, with many treated units having propensity scores that fall outside the support of the control group, and large segments of the control group exhibiting extremely low log-odds. Similar patterns are observed in the original LaLonde male samples, as shown in the online appendix.

To address these issues, we construct trimmed versions of the LDW-CPS and LDW-PSID samples. Following the earlier discussion, trimming can improve robustness with only modest loss of precision. We begin by augmenting each nonexperimental sample with the experimental controls and estimating each unit’s probability of being in the experimental data using GRF. We then trim based on preset thresholds, which may exclude some treated units crump2009dealing. Next, we re-estimate the propensity scores in the trimmed sample and perform 1:1 matching to refine the control group. This yields two sets of trimmed samples: one with experimental treated units and matched nonexperimental controls, and a second with treated and control units from the experiment, which serves as a benchmark. This two-step trimming and matching procedure improves overlap while preserving a comparison to an experimental benchmark.\footnote{Details of the trimming procedure are provided in the online appendix.} As shown in Figure (ref)(D)–(E), overlap improves substantially in both trimmed samples, albeit at the cost of smaller sample sizes.

Estimating the ATT

We now estimate the ATT using both the original LaLonde-Dehejia-Wahba nonexperimental samples and the newly constructed trimmed samples. We apply a range of estimators, some likely familiar to most readers and others perhaps more novel. Our primary goal is to compare these methods; full computational details are provided in the online appendix.

Figure (ref) presents the results. The top row reports experimental benchmarks using both the untrimmed and the trimmed LaLonde-Dehejia-Wahba samples. The remaining rows combine the experimental treated units with a nonexperimental control group and apply the following methods: difference-in-means, simple regression; regression with interactions; generalized random forest (GRF) for outcome modeling; nearest neighbor matching with bias correction (matching each treated unit with five control units based on covariates); inverse propensity weighting with GRF-estimated propensity scores; covariate balancing propensity score; entropy balancing; double/debiased machine learning using elastic net (DML-ElasticNet); and augmented inverse propensity weighting via GRF (AIPW-GRF). All estimators use the same ten covariates used in prior analyses.\footnote{The difference-in-means estimates are shown in the online appendix but omitted here due to their extreme values in the LDW-CPS and LDW-PSID samples (\$–8,497 and \$–15,204, respectively). In the trimmed samples, they are closer to the other estimates (\$1,483 and \$–1,505).}

figure[figure omitted — 1,317 chars of source]

The left panel of Figure (ref) shows ATT estimates and 95% confidence intervals from the LDW-CPS sample; the right panel shows results from the LDW-PSID sample. For each method, estimates from the full sample are shown in black, and those from the trimmed data are in red. Solid circles mark point estimates; lines represent 95% confidence intervals.

Using the LDW-CPS sample (left panel), all estimators yield positive ATT estimates, though they vary in magnitude. Nearest neighbor matching aligns most closely with the experimental benchmark of \$1,794; covariate balancing propensity score, entropy balancing, and AIPW-GRF also produce estimates near this benchmark. Despite numerical differences, the estimates are not statistically distinguishable from one another. In the trimmed LDW-CPS sample, the estimates show less variation across methods, although confidence intervals widen slightly. All estimates in the trimmed sample center around the experimental benchmark of \$1,911. This stability of estimates after trimming is consistent with the simulation results in athey2021using.

Using the LDW-PSID sample (right panel), estimates from the full data show more dispersion, ranging from \$4 to \$2,420. AIPW-GRF yields an estimate closest to the experimental benchmark. In the trimmed LDW-PSID sample, the experimental benchmark is \$306, which is not statistically different from zero at the 5% level. Although the estimates from the trimmed sample are all negative and more similar to each other, they seem qualitatively different from the experimental benchmark. However, due to large standard errors, these differences are not statistically significant.

Overall, the results suggest that improving overlap based on observed covariates reduces model dependence and estimate variability, yielding more robust estimates of the statistical estimand. However, the fact that many methods produce ATT estimates close to the experimental benchmark using LDW-CPS may have given researchers a false sense of confidence that modern estimators can recover causal effects---even in settings where the unconfoundedness assumption is likely violated.

Treatment Effect Heterogeneity.

We explore treatment effect heterogeneity by estimating the CATT, comparing estimates from the experimental and nonexperimental samples in the LaLonde-Dehejia-Wahba data. For simplicity, we focus on the LDW-CPS data—both full and trimmed—for the nonexperimental analyses, using the corresponding experimental samples for benchmarking. In the trimmed sample, 21 treated units lacking comparable controls are removed, and the nonexperimental control group is trimmed using one-to-one matching on the estimated propensity score with the remaining 164 treated units. Details on the trimming procedure are discussed earlier in the paper and in the online appendix. We estimate CATT using a causal forest estimator athey2017generalized, omitting technical details.

figure[figure omitted — 1,361 chars of source]

Figure (ref) plots the estimated CATT at the covariate values of each treated unit, with experimental data on the x-axis and nonexperimental data on the y-axis, using both the full sample and the trimmed sample. Each dot represents a pair of CATT estimates for an individual with a given covariate profile. The red cross marks the pair of ATT estimates obtained using the AIPW-GRF estimator.

Figure (ref) reveals several interesting patterns. First, there is substantial heterogeneity in the estimated treatment effects: CATT estimates based on the experimental data range from \$–236 to \$3,817 in the full sample and from \$–218 to \$4,324 in the trimmed sample. Second, in the full sample, although the AIPW-GRF estimator yields ATT estimates that closely match the experimental benchmark, the CATT estimates from the experimental and nonexperimental data diverge sharply from the 45-degree line. In particular, the nonexperimental CATT estimates span a much wider range—from \$–5,667 to \$7,102—far exceeding that of the experimental estimates, with over one-fourth of treated units obtaining negative CATT estimates. Third, improving overlap substantially enhances the robustness of the CATT estimates: in the trimmed sample, CATT estimates from the experimental and nonexperimental data roughly align along the 45-degree line, and the range of the nonexperimental estimates—\$–2,877 to \$5,758—is much narrower than in the untrimmed sample.

Applying the same procedure to the LDW-PSID data yields CATT estimates ranging from \$–8,422 to \$4,870 in the full sample and from \$–5,088 to \$1,571 in the trimmed sample—showing greater deviation from the experimental benchmark than in the LDW-CPS case. Analysis of the LaLonde male sample and the reconstructed female data further demonstrates that recovering the CATT is substantially more difficult than estimating the ATT. The combination of the experimental subsample constructed by dehejiawahba and the CPS nonexperimental comparison group appears to be an outlier in its ability to recover experimental benchmarks. These additional results are reported in the online appendix.

Validation Through Placebo Analyses

While modern nonexperimental methods may be effective in estimating the statistical estimand---the covariate-adjusted difference in average outcomes between treated and control groups---this does {\it not} imply that the estimate approximates the causal estimand, such as the ATT. Identifying the ATT requires the unconfoundedness assumption, which is fundamentally untestable. However, we can assess its plausibility through placebo analyses.

In the LaLonde setting, we use 1975 earnings as a placebo outcome and exclude both 1975 earnings and employment status from the set of conditioning variables. By construction, 1975 earnings could not have been affected by the treatment, which occurred afterward. We also construct two new trimmed samples, omitting these variables during the trimming process. We then estimate the ATT for the placebo outcome, adjusting for the remaining covariates using a range of estimators.

If the ATT estimates for the placebo outcome are close to zero, this lends support to the unconfoundedness assumption, as it suggests that even without conditioning on 1975 earnings or employment status, treatment assignment is likely independent of the untreated potential outcome $Y_i(0)$—that is, 1978 earnings had the individual not participated in the program. Including these variables in the main analysis would then make unconfoundedness even more credible. However, if the placebo ATT estimates differ significantly from zero, this may indicate either that 1975 earnings and employment status are key confounders, or that unmeasured factors such as perseverance influence both treatment assignment and outcomes. In that case, the placebo analysis fails to bolster unconfoundedness in the main analysis.

Figure (ref) presents the results. As expected, the experimental benchmarks are close to zero and statistically insignificant. In contrast, all estimators using nonexperimental data yield large, negative estimates. While trimming improves the stability of these estimates, they remain statistically different from zero.\footnote{The difference-in-means estimates are not shown here; they are \$–12,118, \$–17,531, \$–14,056, and \$–4,670 across the four panels, respectively. Details are available in the online appendix.} Moreover, as shown in the online appendix, CATT estimates for a similar placebo test based on experimental data cluster around zero, while their nonexperimental counterparts are consistently negative and sizable, indicating substantial bias. These patterns further illustrates the point we aim to emphasize: With sufficient overlap, modern methods can robustly estimate the statistical estimands, but not necessarily the causal estimands if unconfoundedness is violated.

figure[figure omitted — 784 chars of source]

Alternative Samples

For comparison, we also revisit the original male samples used in LaLonde. These datasets lack information on 1974 earnings and employment status. We find that with sufficient overlap, most modern estimators yield estimates within relatively narrow ranges when using CPS-SSA-1 or PSID-1 as control groups. However, these estimates—most of which are negative—do not align with the experimental ATT benchmarks. Similar patterns are reported by smith2001reconciling, smith2005does. Using these nonexperimental data, modern methods also fail to recover the experimental CATT benchmarks, as noted earlier.

Using the LaLonde-Cal{\'o}nico-Smith female sample with PSID data as nonexperimental controls, we find that many modern methods produce estimates close to the experimental benchmarks, though standard errors are often quite large. As noted by calonico2017women, selection appears to be less severe for female participants than for male participants in the National Supported Work program. However, overlap remains a significant challenge. Moreover, a placebo test using the number of children in 1975—a variable not included in LaLonde’s original analysis—does not support the unconfoundedness assumption. CATT estimates also fail to recover the experimental benchmarks.

As an example of a dataset that readily passes a placebo test, we include in the appendix a reanalysis of data from imbensrubinsacerdote, who conducted an original survey to study the impact of lottery prize size in Massachusetts during the mid-1980s on the economic behavior of lottery players. The primary outcome is post-winning labor earnings, with data available for earnings in the six preceding years. While lottery outcomes may seem random, there are systematic differences in pre-treatment variables, likely related to the number of tickets purchased or survey response rates. However, in this case, placebo tests provide strong evidence supporting the unconfoundedness assumption, bolstering the credibility of the causal estimates. The availability of six pretreatment earnings measures proves particularly valuable: they likely capture both selection and outcome-relevant factors and serve as strong placebo outcomes due to their comparability to the post-treatment outcome. We report the details in the online appendix.

Summary

Our reexamination of the LaLonde data shows that when overlap is ensured, the choice of estimation method becomes less consequential, as most methods yield similar results. However, these estimates may not represent the causal quantity of interest if the unconfoundedness assumption is violated. Supplementary analyses, such as placebo tests, can help assess how plausible this assumption is. Although in both the LDW-CPS and LaLonde-Cal{\'o}nico-Smith datasets, many modern methods appear to recover the experimental benchmark for the ATT, placebo tests fail to support unconfoundedness. In typical research settings where experimental benchmarks are unavailable, we cannot know whether such estimates credibly recover the causal estimand. Indeed, with the original LaLonde male sample or the LDW-PSID dataset, modern estimators fail to recover the ATT.

The answer to whether modern nonexperimental methods can credibly estimate causal effects, forty years after LaLonde's critique, is therefore nuanced. Researchers can now use data-driven approaches to improve overlap and apply modern methods to reliably estimate the statistical estimand. However, credible causal inference still depends on the validity of the unconfoundedness assumption. If supplementary analyses—such as placebo tests—support this assumption, modern methods can yield highly credible estimates. Otherwise, they cannot.

Lessons Learned

What lessons has the methodological literature since LaLonde taught us? And what specific analyses should researchers conduct today in similar nonexperimental studies? We return to the five lessons and offer practical recommendations.

First, any analysis of causal effects using nonexperimental data should begin with a careful investigation of the treatment assignment mechanism. A clear understanding of the “design” is essential for evaluating the plausibility of the unconfoundedness assumption. The case for relying on unconfoundedness is strongest when researchers believe that selection into treatment is driven by factors that are well understood, observed, and measured. When that condition holds, flexibly adjusting for observed pretreatment covariates can reduce reliance on strong modeling assumptions. In cases where important confounders are unobserved, researchers may turn to panel data methods that account for time-invariant unobserved confounding; for recent overviews, see xu2023causal and arkhangelsky2023causal.

Second, the literature has underscored the importance of assessing and improving overlap in covariate distributions. The comparison groups LaLonde used differed substantially from the experimental sample, prompting him to discard some units based on age, employment status, and earnings. Since then, more systematic and data-driven approaches have been developed to diagnose and address lack of overlap, including propensity score-based trimming and weighting strategies. The loss of efficiency is often modest and a worthwhile cost.

Third, relatedly, the propensity score has become central to both diagnosing overlap and estimating treatment effects. Researchers now routinely estimate the propensity score using flexible methods and evaluate overlap by comparing the distribution of propensity scores across treated and control groups. When necessary, samples can be trimmed to improve comparability. While the concept of the propensity score was introduced just prior to LaLonde’s work in rosenbaum1983central, it has since become foundational. Doubly robust estimators, particularly those that combine outcome modeling with inverse propensity score weighting, yield consistent estimates when either the outcome model or the propensity score model is correctly specified. The integration of machine learning techniques into causal inference has further strengthened these methods by reducing dependence on ad hoc specification choices. We expect these methods to see broader adoption among economists and social scientists.

Fourth, attention has shifted from solely estimating average treatment effects to understanding treatment effect heterogeneity. Policymakers increasingly ask not only “Does it work?” but “For whom does it work?” and “Where does it cause harm?” The literature has developed tools to estimate conditional average treatment effects and quantile treatment effects. These estimands help unpack heterogeneous impacts and can inform personalized policy decisions. Large datasets and algorithmic advances have made it easier to estimate these effects flexibly and at scale wager2018estimation. We encourage the estimation and visualization of these additional quantities of interest.

Finally, the credibility of causal estimates increasingly hinges on validation exercises—especially placebo tests. While LaLonde included some placebo analyses using lagged earnings, his main focus was on comparing nonexperimental estimates to experimental benchmarks. In contrast, modern practice places greater emphasis on formal diagnostic checks. Placebo tests, such as estimating effects on outcomes known to be unaffected by the treatment, offer informative checks on key identifying assumptions like unconfoundedness and should be more routinely used in empirical analyses.

To help researchers apply these lessons in practice, we provide a detailed online tutorial—including R code and the data used in this paper and the online appendix—available at \url{https://yiqingxu.org/tutorials/lalonde/}. The tutorial replicates our analyses and can be easily adapted to other datasets.

{\it $\blacksquare\quad$ We thank the Office of Naval Research for support under grant numbers N00014-17-1-2131 and N00014-19-1-2468, and Amazon for a gift. We are grateful to Susan Athey, Scott Cunningham, Alexis Diamond, Peng Ding, Dean Eckles, and Xiang Zhou for helpful comments and feedback, and to Timothy Taylor, Jonathan Parker, and Heidi Williams for detailed editorial suggestions. We also thank Zihan Xie and Jinwen Wu for excellent research assistance.}