Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
76,072 characters · 19 sections · 21 citation commands
Fixed Effects and the Generalized Mundlak Estimator
\thispagestyle{empty}
Keywords: fixed effects, cross-section data, groups, causal effects, treatment effects, unconfoundedness.
\baselineskip=20pt \setcounter{page}{1}
A common specification for regression functions when estimating causal effects with grouped data is a fixed effect model:
where the index $g(i)$ indicates the group a unit $i$ belongs to, the $\alpha_g$ are the group fixed effects, and $\varepsilon_i$ is an error term, independent of the regressors $W_i$ and $X_i$ (see \citet*{ arellano2003panel, angrist2008mostly,wooldridge2010econometric} for textbook discussions). In empirical work the groups (also referred to as strata, subpopulations, or clusters) may correspond to states, cities, MSAs, classrooms, birth cohorts, families, siblings, or other geographic or demographic groups. The model is typically estimated by least squares. We refer to the corresponding estimator as the fixed effect estimator. The parameter $\tau$, which we would like to interpret as an average causal effect of the treatment $W_i$ (binary through most of this paper), is the object of interest. The fixed effects $\alpha_g$ are intended to capture unobserved differences between the groups. The motivation for including these fixed effects in the specification is that their presence improves the credibility of a causal interpretation of the fixed effect estimator for $\tau$ ({\it e.g.,} \citet*{arellano2003panel, angrist2008mostly}). Inference for $\tau$ is typically based on asymptotic approximations with a growing number of groups and a fixed number of units per group. In that setting one cannot rely on consistent estimation of the fixed effects because of the incidental parameter problem ({\it e.g.,} \citet*{neyman1948consistent, lancaster2000incidental}).
In this paper we make two sets of contributions. First, we unpack the assumptions underlying the fixed effects estimator based on specification ((ref)), and second, we re-interpret the fixed effect estimator and use that re-interpretation to propose a new class of estimators.
The popularity of the fixed effect estimator may be due partly because it is viewed as accounting for differences between groups in a flexible way. However, the specification in ((ref)) embodies multiple strong assumptions. The current paper clarifies these assumptions by separating them into three parts. First, the specification in ((ref)) assumes a constant treatment effect. In practice, it is likely the effects of the treatment are heterogeneous. We show that in general the average effect of the treatment is {\em not} point-identified. We then describe the estimand corresponding to the fixed effect estimator under general treatment-effect heterogeneity and demonstrate that it estimates, under some conditions, a weighted average treatment effect, with the weights depending in an unusual way on the sampling scheme. For example, we show that in the presence of treatment effect heterogeneity the interpretation of the fixed effect estimator changes if we double the number of units we sample per cluster. Second, a key assumption motivating ((ref)), what we call group unconfoundedness (Assumption (ref) below), validates all within-group comparisons of treated and control units with the same covariates. However, the additional assumptions underlying ((ref)) also implicitly validate some, but not all, between-group comparisons of treated and control units with the same covariates. We describe which between-group comparisons are validated by the fixed effect specification. Third, the fixed effect specification assumes linearity and additivity in the fixed effects, covariates, and treatments.
In the second set of contributions, we propose a new class of estimators. To motivate these new estimators we start with the re-interpretation of the fixed effect estimator due to mundlak1978pooling. Instead of thinking of the fixed effect estimator for $\tau$ as corresponding to the specification in ((ref)), we can also think of it as the least squares estimator for $\tau$ corresponding to the specification
where, throughout the paper, we use $\overline{W}_g$ and $\overline{X}_g$ to denote the average value of $W_i$ and $X_i$ for units in group $g$, that is, for units with $g(i)=g$. Compared to ((ref)), this specification does not have the fixed effects $\alpha_g$. Instead, it has two additional regressors, the group averages of $W_i$ and $X_i$. The least squares estimators for $\tau$ based on ((ref)) and ((ref)) are identical, as first noted by mundlak1978pooling (see also wooldridge2021two). We can therefore interpret the fixed effect estimator as controlling for group differences by simply including group averages of the covariates as additional regressors in a linear regression. This interpretation motivates our exploration of new estimators, which we refer to as Generalized Mundlak Estimators (GMEs). These estimators, for suitable defined average treatment effects as made precise in Section (ref), maintain group unconfoundedness (Assumption (ref)) which justifies within-group comparisons of treated and control units. We then formulate additional assumptions that allow for some, but not all, cross-group comparisons of treated and control units. The estimators can be thought of as starting with a parsimonious specification as in ((ref)), and, like ((ref)), are based on adjusting for group characteristics to justify cross-group comparisons. They differ from ((ref)) in that they free up the specification or change the estimation procedure in three distinct ways, inspired by the modern literature on estimating average treatment effects under unconfoundedness ({\it e.g.,} imbens2015causal, abadie2018econometric). First, we can adjust for differences in the group averages in a more flexible, nonlinear way, for example, by including nonlinear functions of the averages $\overline{W}_{g}$ and $\overline{X}_{g}$ in the regression function in ({(ref))}. Second, we can make the estimator more robust by estimating the propensity score and using inverse propensity score weighting in combination with regression to create doubly robust estimators (\citet*{robins1995semiparametric, zubizarreta2015stable, chernozhukov2017double, athey2018approximate}). Third, we can account for more general differences between groups than accounted for by the fixed effect estimator by adjusting for differences in averages of other functions of the covariates $X_i$ and the treatment indicators $W_i$, for example, the average of the product, $\overline{WX}_{g}$. We show how the choice of group characteristics to adjust for can be motivated by structural choice models in Section (ref).
We can capture the three modifications in the proposed Generalized Mundlak Estimator relative to the fixed effect estimators by an unconfoundedness assumption that strengthens group unconfoundedness. Instead of conditioning on group indicators, it assumes it is sufficient to condition on a set of group characteristics: \[W_i\ \perp\!\!\!\perp\ \Bigl(Y_i(0),Y_i(1)\Bigr)\ \Bigl|\ X_i, \overline{H}_{g(i)},\] where $Y_i(0)$ and $Y_i(1)$ are potential outcomes for unit $i$ (\citet*{neyman1923,rubin1974estimating, imbens2015causal}), and $\overline{H}_{g(i)}$ is the group average of some pre-specified function $H(W_i,X_i)$ of the treatment indicators and the covariates. We refer to $\overline{H}_{g(i)}$ as a balancing score, following the terminology in rosenbaum1983central. By the unconfoundedness condition, treated and control units with the same values for covariates and balancing statistics can be compared even if they belong to different groups. If $H(W_i,X_i)=(W_i,X_i)$ and the adjustment is only linear, this gets us back to the fixed effect estimator, but in general the proposed Generalized Mundlak Estimator differs in two ways. First, it adjusts more flexibly for the components of $\overline{H}_{g(i)}$, and second, $H(\cdot)$ can contain more components than just $(W_i,X_i)$.
Our formal asymptotic results focus on the case with a fixed number of units per group. In particular, we discuss settings where the group size is not large enough to carry out a two-stage procedure where we first estimate the effects entirely within groups by flexibly adjusting for covariates, followed by averaging over the groups. In other words, because of the incidental parameter problem (\citet*{neyman1948consistent, lancaster2000incidental}), we need to rely on comparisons of treated and control units in different groups. In addition, there is concern that accounting for the group differences solely through additive fixed effects as in specification ((ref)) is not sufficient to adjust for all relevant differences ({\it e.g.,} \citet*{altonji2005cross,imai2019should}). Finally, in this setting we cannot estimate the propensity score, that is, the conditional probability of treatment assignment, as a function of group membership.
In this section, we discuss the main insights of the paper in the context of two examples. We start with a case with no covariates and present the interpretation of ((ref)) under general heterogeneity in treatment effects. We then introduce binary covariates and connect ((ref)) to a particular unconfoundedness condition. We finish the section with a set of practical recommendations.
A key assumption underlying most group analyses is that assignment is unconfounded within the groups. Let each unit $i$ in a large population be characterized by a pair of potential outcomes $(Y_i(0),Y_{i}(1))$ and group label $L_{g(i)}$. These labels can correspond to names, e.g., geographic locations, legal entities, and markets. We use the labels at this stage to make clear that there is not necessarily an ordering to the groups, nor a distance measure that allows us to say that some groups are closer to each other. Suppose each unit is assigned to a binary treatment $W_i$. The sampling process has two stages. First, we randomly sample $M$ groups from a large population of groups. Second, we sample $N_g$ units from group $g$, for each of the sampled groups. In the absence of individual-level covariates, we can express group unconfoundedness in the following way (\citet*{rosenbaum1983central}):
This assumption validates comparisons between treated and control units within groups.
As we show formally in the next section, this condition implies a different independence restriction:
where $\overline{W}_g$ is the share of treated units in group $g$. This unconfoundedness condition allows us to combine groups with the same value for $\overline{W}_g$. This does not immediately have any empirical content. To see this, consider the special case with two units per group. This implies there are three values for $\overline{W}_g$, namely $0$, $1/2$, and $1$. If $\overline{W}_g=0$ or $\overline{W}_g=1$, the combined set of groups has only control units or only treated units, so there is no basis to estimate treatment effects. If we look at a set of groups with $\overline{W}_g=1/2$, comparing treated and control units gives us an average of within-group comparisons of treated and control units based on (ref). However, within-group comparisons were validated already by ((ref)).
Continuing with this two-units-per-group example, let $M_s$ denote the number of groups in the sample with $\overline{W}_g=s$. Without covariates the fixed effect estimator for $\tau$ in ((ref)), which we denote by $\hat{\tau}^{\rm fe}$, reduces to averaging the difference between the treated unit and the control unit over the $M_{1/2}$ groups with exactly one treated and one control unit: \[ \hat{\tau}^{\rm fe}=\frac{1}{M_{1/2}}\sum_{g:\overline{W}_g=1/2} \hat\tau_g,\hskip1cm {\rm where }\ \ \hat\tau_g=\sum_{i:g(i) = g}W_i Y_i-\sum_{i:g(i) = g}(1-W_i) Y_i.\] This fixed effect estimator $\hat{\tau}^{\rm fe}$ does not depend on outcomes for groups with either only treated units $\overline{W}_g=1$, or groups with only control units, $\overline{W}_g=0$. For groups with $\overline{W}_g=1/2$, $\hat\tau_g$ is the difference between a treated and a control outcome from the same group, so it is a natural, and in fact the only natural, estimator for the average effect within that group given that there are exactly two units from that group in the sample.
We can now characterize what $\hat{\tau}^{\rm fe}$ is estimating in settings with heterogeneous treatment effects under group unconfoundedness. Let $\tau=\mathbb{E}[Y_i(1)-Y_i(0)]$ denote the population average treatment effect and define for $s\in\{0,1/2,1\}$ the weighted average effect
This is a critical, though somewhat unusual, object. $\tau(s)$ is a conditional average treatment effect, but the conditioning depends on the sampling scheme and the assignment mechanism. To illustrate the unusual nature of this object, suppose that for the same population we sampled three units per group instead of two. That would change the values of $s$ for which $\tau(s)$ is defined, and it would potentially change the value of $\tau(0)$ and $\tau(1)$ which are defined under both sampling schemes.
Now consider the overall average effect of the treatment. This can be expressed in terms of the $\tau(s)$ as \[ \tau={\rm pr}(\overline{W}_{g(i)}=0) \tau(0)+ {\rm pr}(\overline{W}_{g(i)}=1/2) \tau(1/2)+ {\rm pr}(\overline{W}_{g(i)}=1) \tau(1). \] Under group unconfoundedness ((ref)), $\hat{\tau}^{\rm fe}$ is unbiased for $\tau(1/2)$. If, as is likely in the heterogenous treatment effect case, $\tau(s)$ varies by $s$, $\tau(1/2)$ would differ from $\tau$, and thus that $\hat{\tau}^{\rm fe}$ estimates something different from $\tau$.
To further illustrate this point, suppose there is a group-level variable $\overline{U}_g$ that is related to the propensity score and the outcomes in the following way: \[ \mathbb{E}[Y_i(1)-Y_i(0)|\overline{U}_{g(i)}]=\overline{U}_{g(i)}^2\hskip1cm {\rm and}\ \ {\rm pr}(W_i=1|L_{g(i)}=L_g)=\overline{U}_g.\] Furthermore, suppose $\overline{U}_g$ has the following distribution over the population of groups, \[ {\rm pr}(\overline{U}_g=u)=1/3, \ u\in\{0,1/2,1\}.\] In this setup, the overall average treatment effect is $\tau=\mathbb{E}[Y_i(1)-Y_i(0)]=E[\overline{U}_{g(i)}^2]=5/12$. However, ${\rm pr}(\overline{U}_{g(i)}=1/2|\overline{W}_{g(i)}=1/2)=1,$ so that $\tau(1/2)=1/4$. Conditioning on groups with one treated and one control observation shifts the distribution of $\overline{U}_g$ from its marginal distribution to a distribution that puts more weight on values of $\overline{U}_g$ that make the observation of both a treated and control unit likely. It is easy to see that the interpretation of the fixed effect estimator ((ref)) changes if we change the sampling scheme, {\it e.g.,} if we sample $N_g=3$ units for each group, or if the assignment mechanism changes.
Next, we consider a setting with a single binary covariate $X_i \in \{0,1\}$. Define for each group $g$ and each covariate value $x$ the average treatment effect:
This quantity is well-defined for all $(x,l)$ such that $x$ is in the support of $X_i$ in group $L_g$.
A natural generalization of the group unconfoundedness assumption in ((ref)) is the (conditional) group unconfoundedness assumption:
Similarly to ((ref)), it implies a different unconfoundedness condition that does not require within-group comparisons:
where $\overline{W}_{g(i)},\overline{X}_{g(i)},\overline{WX}_{g(i)}$ are group-level averages of the corresponding variables. Condition ((ref)) justifies combining groups that have identical empirical joint distributions of covariates and treatment indicators.
Focusing again on the case with two units per group, we can further unpack condition (ref). In particular, ((ref)) allows us to compare units for groups $g$ with variation in the treatment and no variation in the covariate. Formally, let $\underline{W}_g$ be the pair of treatment values $(W_i,W_j)$ for the two units $i$ and $j$ in group $g$, and let $\underline{X}_g$ be the pair of treatment values $(X_i,X_j)$ for the two units $i$ and $j$ in group $g$. We can then define the set \[\mathbb{B} \equiv \Bigl\{(\underline{W}_g,\underline{X}_g)\Bigl|\, \underline{W}_g\in\{(0,1),(1,0)\},\underline{X}_g\in\{(0,0),(1,1)\} \Bigr\},\] and consider the average treatment effect
Similarly to $\tau(1/2)$ from the previous section, $\tau_\mathbb{B}$ is generally different from the average treatment effect $\tau$. It depends both on the sampling scheme and the assignment process.
The fixed effect specification ((ref)), in general, does not consistently estimate $\tau_{\mathbb{B}}$ or any other weighted average effect. To see this, observe that the OLS estimator $\hat{\tau}^{\rm fe}$ for the equation ((ref)) is equal to the estimated coefficient of a different regression function, namely
The numerical equivalence, first shown in \citet*{mundlak1978pooling}, follows from repeated applications of textbook omitted variable bias formulas. In general, least squares estimation of regression ((ref)) combines data from all groups, including those with $\overline{W}_g \in \{0,1\}$, making $\hat{\tau}^{\rm fe}$ inconsistent for the average effect.
The alternative version of the fixed effect regression in representation ((ref)) is important because it suggests weakening the conditional group unconfoundedness condition from ((ref)) by dropping $\overline{WX}_{g(i)}$, leaving the condition as
This condition encapsulates a key feature of the fixed effect specification that $\overline{W}_{g(i)}$ and $\overline{X}_{g(i)}$ fully capture all relevant differences between groups and that $\overline{WX}_{g(i)}$ is not required for this. Restriction ((ref)) implies that we can compare treated and control units as long as they have the same values of $X_i$ and $\overline{S}_{g(i)}\equiv (\overline{W}_{g(i)},\overline{X}_{g(i)})$, even though they may belong to groups that have different values of $\overline{WX}_{g}$. With two units per group, by construction $\overline{S}_{g(i)}$ takes $9$ possible values, but only $3$ of them correspond to groups with variation in treatment status: $\overline{S}_{g(i)} \in \left\{\left(\frac12,0\right),\left(\frac12, \frac12\right),\left(\frac12,1\right) \right\}$. For other values of $\overline{S}_{g(i)}$ we either have only control or only treated units, and thus condition ((ref)) is not useful (and neither is ((ref))).
The three remaining values of $\overline{S}_{g(i)}$ lead to conceptually different comparisons. Values $\overline{S}_{g(i)} \in \left\{\left(\frac12,0\right),\left(\frac12,1\right) \right\}$ correspond to groups where both units are identical in terms of $X_i$ but vary in their treatment. For these groups ((ref)) justifies within-group comparisons which were already validated by the more restrictive assumption ((ref)). This leaves one remaining value, $\overline{S}_{g(i)} =\left(\frac12, \frac12\right)$. If $\overline{S}_{g(i)} =\left(\frac12, \frac12\right)$ then by definition the two units in the group have different values of $X_i$, and within-group comparisons have no causal interpretation: the pair of units in the same group are either $(W_i,X_i)=(0,0)$ and $(W_j,X_j)=(1,1)$, or the pair are $(W_i,X_i)=(1,0)$ and $(W_j,X_j)=(0,1)$. In both cases, the pairs cannot be compared directly. However, ((ref)) justifies combining groups where $W_i = X_i$ with those where $W_i = 1-X_i$. It is useful to consider this comparison in more detail. Suppose that one group has units with $(W_i,X_i)\in\{(0,0),(1,1)\}$, and another group has units with $(W_i,X_i)\in\{(0,1),(1,0)\}$. In neither of the groups we can directly compare treated and control units. But under ((ref)) we can combine the two groups and then compare a treated unit $i$ with $(W_i,X_i)=(1,1)$ in one group with a control unit $j$ with $(W_j,X_j)=(0,1)$ in another group because they satisfy the two conditions that $(i)$ they have the same value of the covariates ($X_i=X_j=1)$ and $(ii)$ they belong to groups with the same value for $\overline{W}_g$ and $\overline{X}_g$ (namely $\overline{W}_{g(i)}=\overline{W}_{g(j)}=1/2$ and $\overline{X}_{g(i)}=\overline{X}_{g(j)}=1/2$).
Define the conditional expectation of $\tau_{l}(X_i)$ where the expectation is over all units with $X_i=x$ and over all groups with $\overline{S}_{g(i)}=s$:
Similar to $\tau(s)$ and $\tau_{\mathbb{B}}$, this object is defined in terms of the sampling scheme and is not a fixed population characteristic. The discussion above shows that we can consistently estimate $\tau(s,x)$ for certain values of $(s,x)$. Importantly, the linear specification ((ref)) relies on comparisons that are not validated by ((ref)) and thus does not estimate a meaningful average effect in this case.
Examples from the previous sections show that we can reduce group unconfoundedness, as in ((ref)) or ((ref)), to conditioning on some balancing score $\overline{S}_{g(i)}$ and covariates, as in ((ref)) or ((ref)). Despite the lack of immediate empirical content, this reduction brings two critical insights. Both insights are about operationalizing the notion that some groups are more similar than others, which is absent in ((ref)) or ((ref)).
First, conditioning on an unordered discrete characteristic -- group indicator -- does not provide a distance metric to make the case that group $g$ is closer to group $g'$ or to group $g''$. In contrast, ((ref)) (and similarly ((ref))) allows for the establishment of such a metric: a group $g$ with $\overline{W}_g=0.10$ is closer to group $g'$ with $\overline{W}_{g'}=0.11$ than it is to group $g''$ with $\overline{W}_{g''}=0.50$ (see \citet*{johannemann2019sufficient} for an application of this idea). The second insight is even more central to the current paper. It allows us to find a middle ground between fully conditioning on the entire set of group indicators and not conditioning on these indicators at all. For example, going from ((ref)) to ((ref)) justifies some, but not all between-group comparisons. Depending on the application, one can go further and assume a stronger condition
Alternatively, if we observe more units in each group, we can relax ((ref)) to improve the comparability of groups without reducing it fully to ((ref)). We discuss in Section (ref) whether and when it is enough to adjust for some subset of these averages.
Another issue illustrated in the previous sections is the relationship between the unconfoundedness conditions and the fixed effects specification ((ref)). In general, this regression does not estimate the average treatment effect, and in the presence of covariates, it might fail to estimate even a weighted average effect. However, under the restriction ((ref)) we can go beyond the linear specifications ((ref)) and consider different estimation methods. First, we can use more flexible specifications for the regression function as a function of the control variables $(X_i,\overline{W}_{g(i)},\overline{X}_{g(i)})$ and the treatment $W_i$: \[ Y_i(w)=\mu(w,X_i,\overline{W}_{g(i)},\overline{X}_{g(i)})+\varepsilon_i(w),\] with a parametric or non-parametric specification for $\mu(w,\cdot)$ that generalizes the linear additive form in ((ref)). These specifications may include higher-order moments of the control variables, interactions with the treatment, or transformations of the linear index. Given estimates of $\mu(w,\cdot)$ we can average the difference $\hat \mu(1,X_i,\overline{W}_{g(i)},\overline{X}_{g(i)})-\hat \mu(0,X_i,\overline{W}_{g(i)},\overline{X}_{g(i)})$ over the sample to estimate the average treatment effect.
Second, we can model the propensity score as a function of $(X_i,\overline{W}_{g(i)},\overline{X}_{g(i)})$: \[ e(x,\overline w,\overline x)\equiv {\rm pr}(W_i=1|X_i=x,\overline{W}_{g(i)}=\overline w,\overline{X}_{g(i)}=\overline x).\] Once we have estimates of the propensity score, we can use them to develop inverse propensity score weighting estimators (see arpino2011specification for an application of this idea to grouped data). In particular, an attractive approach would be to use the inverse propensity score weighting or other forms of balancing in combination with a credible specification of the regression function. For example, one could use a weighted version of ((ref)) to make the results more robust to misspecification of the regression function. Such doubly-robust and balancing methods have been argued in the recent causal inference literature on unconfoundedness to be more robust to misspecification than estimators that rely solely on specifying the conditional mean of the outcome given conditioning variables and treatments ({\it e.g., } \citet*{robins1995semiparametric, zubizarreta2015stable, chernozhukov2017double, athey2018approximate}).
Following this discussion, we recommend using three modifications of ((ref)). First, choose what group characteristics one wishes to include in the analysis beyond the averages of the treatment and the covariates. Second, specify a flexible conditional mean function that allows for dependence on the selected group characteristics. Third, estimate the propensity score as a function of individuals and group characteristics and combine that with the conditional mean specification to obtain a more robust estimator for the average treatment effect.
Our approach is related to classic panel data literature, e.g., mundlak1978pooling, chamberlain1984panel,chamberlain2010binary,altonji2005cross. Similar to these papers, we develop a conditional approach, where we effectively control for heterogeneity across groups using group-level balancing scores. Our key contribution to this literature is to show how these statistics arise naturally from the design model and connect this approach to modern estimation algorithms developed for cross-sectional data.
Our results are also connected to random effects models studied in the more recent panel data literature (see arellano2011nonlinear for a recent review). For example, arellano2016nonlinear explicitly model the distribution of the outcomes and assignments given the individual unobservable characteristics and achieve identification using results from nonparametric nonlinear deconvolution literature (hu2008instrumental). The key conceptual and practical difference between our approach and theirs is that we do not use the outcome data to deal with unobserved heterogeneity and rely only on the distribution of treatments and covariates for identification. In this sense, we follow a tradition used in causal inference literature (\citet*{rosenbaum_book, imbens2015causal, abadie2020sampling}).
Our approach is connected to grouped fixed effects strategies (\citet*{hahn2010panel, bonhomme2015grouped,bonhomme2022discretizing}). Similar to this literature, we propose pooling data across groups with similar aggregate statistics. There are two main differences, though. As discussed above, we do not use the outcome data, instead focusing on the assignment process. Also, our identification and inference results are valid for groups of small size, whereas statistical results in the other literature are based on approximations with large groups.
Finally, our results are related to recent literature on regression models with fixed effects (e.g.,de2020two,goodman2021difference,wooldridge2021two). Similar to these papers, we discuss the interpretation of the least-squares estimators under heterogeneity in treatment effects and develop an estimator that is robust to the presence of such heterogeneity. Our focus, however, is quite different: instead of restricting the conditional outcome distributions, we focus on the assignment model and achieve robustness by using it to construct the estimator.
In this section we present the main results.
We start with the sampling assumption that describes the relationship between groups and unit-level data:
The assumption that the group size is identical for all groups is for expositional reasons only. In the Appendix, we generalize our results to allow for heterogeneity in the group sizes. The key complication is that it requires us to add the group size to the set of conditioning variables that make up the balancing score. We assume that for a generic group $g$ we observe its label $L_g$. For each unit $i$ define $g(i)\in\{1,\ldots,M\}$ -- the index of the group this unit belongs to. For each unit $i$ we observe the outcomes $Y_i$, binary treatment indicators $W_i\in\{0,1\}$, and covariates $X_i \in \mathbb{X}$. Assumption (ref) allows us to view the data as $M$ independent and identically distributed copies of random elements $\left(L_g,\{(W_i,X_i,Y_i)\}_{i: g(i)= g}\right)$, with the additional independence among vectors $\{(W_i,X_i,Y_i)\}_{i: g(i)= g}$.
We interpret the observed outcomes using Neyman-Rubin potential outcomes model (neyman1923,rubin1977assignment,imbens2015causal):
which implies the absence of spillovers across units and groups. We impose restrictions on the treatment assignment process:
This assumption restricts the distribution of potential outcomes and treatment indicators within a particular group. It implies that we can give a causal interpretation to comparisons of units with the same characteristics within the group.
For the first identification result we need to introduce additional notation. For each group $g$ define $\mathbb{P}_g$ to be the empirical measure of $(W_i,X_i)$ in group $g$. Formally, for any set $A\subseteq \mathbb{X}\times \{0,1\}$, $\mathbb{P}_g(A)$ is defined through the following relationship:
We use this construction for our first identification result:
For proofs of all results in this section see Appendix (ref). Proposition (ref) states that as long as units have the same characteristics, and they come from groups identical in terms of $\mathbb{P}_g$ they are comparable. To relate this result to the discussions in the previous section, assume that there are no covariates. In this case, $\mathbb{P}_g$ reduces to $\overline{W}_g$ and we get condition ((ref)). If we introduce a single binary covariate -- we get condition ((ref)). This result is a manifestation of a familiar balancing property of the propensity score (e.g., rosenbaum1983central). The subpopulation with the same value for $(X_i,\mathbb{P}_g)$ are balanced: the distribution of treatments is the same for all units within such subpopulations.
However, as discussed in the previous section, the empirical relevance of this result is limited. To make it operational, we need to substitute the empirical distribution of $(W_i, X_i)$ with a low-dimensional object. To do this, we put structure on the joint distribution of $(W_i,X_i)$ and its relation to $L_g$ with an exponential family representation (cox1979theoretical).
Under relatively weak conditions distributions can be well approximated by distributions in exponential families (barron1991approximation), so as long as we do not restrict the dimension of $S(\cdot)$ this assumption is fairly weak. For example, note that if the distribution of $X_i$ is discrete, one can immediately write the joint distribution of $(W_i,X_i)$ within each group as an exponential family distribution with a group-specific parameter. In addition, we can approximate any distribution arbitrarily well by a discrete distribution. Of course, it is only when the dimension of $S(\cdot)$ is small relative to the number of observations that the assumption will be seen to be useful.
Assumption (ref) places no restrictions on the distribution of the potential outcomes and thus does not restrict heterogeneity in treatment effects. At the same time, it restricts the joint distribution of $(W_i, X_i)$, not just the conditional distribution of $W_i$. This is without loss of generality: if $X_i$ is discrete and function $S(W_i, X_i)$ includes indicators for all potential values of $X_i$, then ((ref)) places no restrictions on the marginal distribution of $X_i$ given $L_{g(i)}$. For continuous $X_i$ the same argument holds using an infinite set of indicators.
With Assumption ((ref)) we can define an unobserved group-level variable that captures all the information about the distribution of $(W_i,X_i)$:
We use the bar notation to indicate that this is a group-level variable. Our next lemma shows that we can substitute $L_g$ with $\overline{U}_g$ for the purpose of identification.
This result demonstrates that our problem can be converted into a traditional random effects setup (e.g., \citet*{chamberlain1984panel}): the group labels $L_g$ are not important, only the group-level characteristics $\overline{U}_g$, which are functions of these labels, are. We do not know the group-specific parameters $\overline{U}_g$, nor can we estimate them consistently in settings with a small number of observations per group because of the incidental parameter problem (\citet*{neyman1948consistent, lancaster2000incidental}). However, and this is a key insight in the paper, it is {\it not} necessary to estimate them consistently because the sufficient statistic in the exponential family representation in (ref), $\overline{S}_g\equiv\sum_{i:g(i) = g}S(W_i,X_i)/N_g$, plays the role of a balancing statistic.
Using Assumption (ref) we state our main identification result:
Theorem (ref) is related to Proposition (ref), but the key difference is that Theorem (ref) has empirical content as long as $\overline{S}_{g(i)}$ is less general than $\mathbb{P}_{g(i)}$. Theorem (ref) reduces the high-dimensional object $\mathbb{P}_{g(i)}$ to an average $\overline{S}_{g(i)}$ and allows us to combine multiple groups. This makes cross-sectional methods developed in the last decades feasible for applications with grouped data. In practice, the importance of this point depends on the number of units per group and the dimension of $S(\cdot)$. The higher the dimension of $S(\cdot)$, the closer we are to controlling for $\mathbb{P}_{g(i)}$ and thus relying solely on comparisons within groups.
Theorem (ref) becomes useful once we fix a particular low-dimensional $S(W_i,X_i)$. We assume that $S(\cdot)$ is known throughout most of the paper. This makes Assumption (ref) restrictive and we can no longer argue that it holds for an arbitrary conditional distribution of $(W_i,X_i)$. As we illustrate with the examples in the previous section, this is needed if we want to go beyond within-groups analysis without imposing additional assumptions on the potential outcomes. In the current empirical practice, researchers rely on ((ref)) or, equivalently, ((ref)), thus implicitly using $S(W_i,X_i) = (W_i,X_i)$, which can be too coarse for some applications. In particular, this choice assumes that covariates do not interact with group labels in the assignment model. We informally discuss selecting the set of balancing scores in Section (ref).
Define propensity score:
Depending on the value of the balancing score, the propensity score $e(x,s)$ might be equal to zero or one for certain values of $s$. For example, this happens if $\overline{W}_g$ is part of the balancing score. As a result, we cannot expect the overlap assumption (imbens2015causal) to hold. Instead, we are making the following assumption:
This assumption has two parts: the first part restricts $e(x,s)$ to be non-degenerate on a certain set and only that set. This is necessary if we want to identify treatment effects without relying on functional form assumptions. The second part is different: we assume that the set is known to a researcher. This is a generalization of the standard overlap assumption, where we assume that the set $\mathbb{A}$ is equal to the support of the covariate space, see \citet*{crump2009dealing}.
To provide additional intuition for Assumption (ref), let us go back to the example from Section (ref), where $S(W_i,X_i) = (W_i,X_i)$ and we observe two units per cluster. In that case, the propensity score is equal to either zero or one for any $\overline{S}_{g(i)} \not\in \left\{\left(\frac12,0\right),\left(\frac12, \frac12\right),\left(\frac12,1\right) \right\}$. By construction $e\left(0, \left(\frac12,0\right)\right) = e\left(1, \left(\frac12,1\right)\right) = \frac12$, but the values of $e\left(0, \left(\frac12, \frac12\right)\right)$ and $e\left(1, \left(\frac12, \frac12\right)\right)$ are not fixed, except for the restriction $e\left(0, \left(\frac12, \frac12\right)\right) + e\left(1, \left(\frac12, \frac12\right)\right) = 1$. As a result, if we set $\mathbb{A}\equiv \left\{(x,s):\left(0,\left(\frac12,0\right)\right),\left(0,\left(\frac12, \frac12\right)\right),\left(1,\left(\frac12, \frac12\right)\right),\left(1,\left(\frac12,1\right)\right) \right\}$, then this amounts to assuming $\eta <e\left(0, \left(\frac12, \frac12\right)\right) < 1-\eta$.
We can now define our target estimands. We focus on two of them; the first one is similar to the objects considered in the previous section:
The second object is a sample group-weighted version of ((ref)):
Theorem ((ref)) together with Assumption ((ref)) implies that $\tau_{\mathbb{A}}$ and $\tilde\tau_{\mathbb{A}}$ are identified. We formally state this result in the next Corollary.
In the remainder of the paper, we primarily focus on the properties of estimators for $\tilde\tau_{\mathbb{A}}$. Such estimators can also be used to estimate $\tau_{\mathbb{A}}$, but the variances are different. Our focus on $\tilde\tau_{\mathbb{A}}$ is analogous to the focus on the sample average treatment effect in the experimental literature.
Our key result -- Theorem (ref) -- is related to some well-known identification results. For example, in \citet*{altonji2005cross} a key assumption (Assumption 2.1) requires that there is an observed variable $Z_i$ such that conditioning on $Z_i$ renders the covariate of interest (the treatment in our case) exogenous. In our setting, the role of this conditioning variable is played by the balancing score $ \overline{S}_{g(i)}$. We show how this property can arise from assumptions on the joint distribution of the treatment and the other covariates, and how we can make this more plausible by increasing the dimension of balancing scores.
Assumption (ref) plays the key role in Theorem (ref) and in the previous section we argued that it can be viewed as a natural approximation for the conditional distribution of covariates and treatment indicators. In addition to this statistical justification, it arises in a structural economic model. To this end, consider a setup where each unit is characterized by $(W_i, X_i)$ and then chooses group membership $l$ out of set $L$. For example, $W_i$ can be a scholarship or a voucher, of a student with observed characteristics $X_i$, and $L_{g(i)}$ can be a college or a school this individual has chosen. In this case, $\mathbb{E}[Y_{i}(1) - Y_{i}(0)|L_{g(i)}]$ is an average direct effect of treatment unmediated by the choice $L_{g(i)}$.
Suppose the following function is the indirect utility for individual $i$ from group $l$:
where $\epsilon_{i}(l)$ are i.i.d. extreme-value type-1 shocks. In this case, the probability of selecting a particular option $g$ has the following form:
This restricts the joint distribution of $(L_{g(i)},W_i, X_i)$:
and thus as long as $V(w,x, u) =\eta^\top(l)S(w,x)+\eta_0(l)$ Assumption (ref) is satisfied. This discussion shows that exponential family restriction is a natural approximation in a structural model of choice.
In practice, balancing scores are unknown and need to be selected by researchers. This choice should reflect domain knowledge and a full treatment of this problem is an open question. However, we provide a suggestion for systematically selecting balancing scores in the case where we have a large set of potential balancing scores that includes all the relevant ones but also some that are not relevant. Intuitively we would like a selection procedure to select more balancing scores in settings where we have a lot of units per group and if the distributions vary substantially across groups. Assumption (ref) implies that the conditional probability that a unit in the sample is from group $l$, conditional on $(W_i,X_i)$ and conditional on the set of $L_1,\ldots,L_M$, has a multinomial logit form: \[ {\rm pr}(L_{g(i)}=l|W_i,X_i,L_1,\cdots,L_M)=\frac{\exp(\eta_0(L_g)+\eta(L_g)^\top S(W_i,X_i))} {\sum_{k=1}^M \exp(\eta_0(L_{k})+\eta(L_{k})^\top S(W_i,X_i))}.\] Hence the problem of selecting the balancing scores is similar to the problem of selecting covariates in a multinomial logistic regression model. Given a large set of potential balancing scores, we can use standard covariate selection methods, such as LASSO (\citet*{tibshirani1996regression}) to select a sparse set of relevant ones. In practice, balancing scores are likely to be correlated, and other selection methods such as Adaptive Elastic Net (\citet*{zou2009adaptive}) might be more appropriate. See also \citet*{de2019identifying} for the application of this algorithm to the identification of unobserved network structure with panel data.
This section collects several inference results for a class of semiparametric estimators of $\tau_{\mathbb{A}}$ and $\tilde\tau_{\mathbb{A}}$. All proofs can be found in Appendix (ref). For further use, we use the following notation for the conditional mean, propensity score, and residuals:
These expectations are defined using Assumption (ref), which determines the distribution of $\overline{S}_{g(i)}$. In particular, as we discussed in Section (ref), the distribution of $\overline{S}_{g(i)}$ changes with the number of observed units in each group. We restrict moments of the corresponding errors:
These conditions hold for bounded potential outcomes, which is the leading case in practice. In the case of unbounded outcomes, the first condition imposes a restriction on the degree of heteroscedasticity.
We will use $\hat \mu_g(\cdot)$ and $\hat e_g(\cdot)$ for generic estimators of $\mu(\cdot)$ and $e(\cdot)$. Subscript $g$ is used to allow for cross-fitting (e.g., chernozhukov2018double); we assume that researchers use $K$-fold group-level cross-fitting. In particular, researchers split the observed $M$ groups into $K$ random subsamples, and for each group $g$ construct estimators $\hat \mu_g(\cdot)$, $\hat e_g(\cdot)$ using the data from $K-1$ subsamples that do not contain group $g$. Define indicator variable $A_i \equiv \boldsymbol{1}_{(X_i,\overline{S}_{g(i)}) \in \mathbb{A}}$ and let $\overline{A}\equiv\frac{1}{M} \sum_{g=1}^M\frac{1}{N_g}\sum_{i:g(i) = g}A_i$ be the estimate of the share of units for which we have overlap. We assume the generic estimators $\hat e_g$ and $\hat \mu_g$ satisfy several high-level consistency properties standard in the program evaluation literature:
Under Assumptions (ref), (ref), and bounded outcomes these conditions are satisfied if covariates are discrete. To see this note that $(i)$ and $(ii)$ are weak consistency restrictions, while each term in $(iii)$ is $O_p\left(\frac{1}{M}\right)$. Under appropriate smoothness restrictions (see \citet*{hansen2022econometrics} for a textbook treatment) these conditions hold for conventional kernel estimators. The use of cross-fitting expands the set of available estimators, allowing researchers to use machine learning estimators (see \citet*{newey2018cross} for details).
For arbitrary functions $m(\cdot), p(\cdot)$ define the following functional (e.g., robins1995semiparametric, hahn1998role, chernozhukov2018double):
and its group-level aggregate:
We use $\rho_g(\cdot)$, estimators $\hat \mu_g(\cdot), \hat e_g(\cdot)$, and $\overline{A}$ to define our proposed Generalized Mundlak estimator (GME):
We also define the influence function:
Our next result shows the properties of $\hat \tau_{\rm GME}$:
This result allows us to conduct inference for $\tilde \tau_\mathbb{A}$ -- the sample group-weighted version of our estimand. A similar result holds for $\tau_\mathbb{A}$ with a larger variance that reflects additional randomness in $\tilde\tau_\mathbb{A}$. To estimate the asymptotic variance $\mathbb{V}$ we define the estimated version of $\xi_g$:
and use it to construct the estimator:
Next result shows that this estimator is consistent:
Together Theorems (ref), (ref) imply that standard confidence intervals are asymptotically valid. In particular, if $z_{\alpha}$ is the $\alpha$-quantile of the standard normal distribution then asymptotically
with probability $(1-\alpha)$.
In this section, we discuss three extensions of the ideas introduced in this paper.
Theorem (ref) states that conditional on the covariates and the balancing scores we have the unconfoundedness condition: \[W_i \perp\!\!\!\perp\ \Bigl(Y_i(0),Y_i(1)\Bigr) \ \Bigl|\ X_i, \overline{S}_{g(i)}.\] This implies that we can study estimation of effects other than average treatment effects. This is important in applications where we want to estimate, distributional effects controlling for group-level unobserved heterogeneity.
In particular, for any bounded function $f: \mathbb{R}\rightarrow \mathbb{R}$ we can estimate $\mathbb{E}[f(Y_i(w))]$ using the inverse propensity score:
This allows us to identify quantile treatment effects of the type introduced by \citet*{lehmann2006nonparametrics}. If we are interested in $q$-th quantile of the distribution of $Y_i(w)$ then (under appropriate continuity) we can identify it as a solution to the following problem:
For the standard case under unconfoundedness \citet*{firpo2007efficient} has developed effective estimation methods that can be adapted to this case.
Modeling the conditional distribution of $(W_i,X_i)$ given $\overline{U}_g$ using an exponential family is natural for the purposes of this paper. Nevertheless, in some applications, other families can be more appropriate. In particular, another operational choice is a discrete mixture. Assume that $\overline{U}_{g(i)}$ can take a finite number of values $\{u_1,\dots, u_p\}$ with probabilities $\pi_{1},\dots, \pi_{p}$ and the conditional distribution of $W_i, X_i$ given $\overline{U}_{g(i)}$ is given by $f(w,x|u)$. Collect all the data that we observe for group $g$ in the following tuple:
The marginal distribution of this object is given by the following expression:
This implies that the conditional distribution of $\overline{U}_g$ given $\mathcal{D}_g$ has the following form:
Define $\overline{S}_g\equiv (\pi(\overline{U}_g = 1|\mathcal{D}_g), \dots, \pi(\overline{U}_g = p|\mathcal{D}_g))$ and observe that as long as assignment is unconfounded given $(X_i,\overline{U}_{g(i)})$ we have the following: \[ W_i\ \perp\!\!\!\perp\ \Bigl(Y_i(0),Y_i(1)\Bigr)\ \Bigl|\ X_i,\overline{S}_{g(i)}.\] Recent results (e.g., \citet*{allman2009identifiability}) show that $(\pi_1,\dots, \pi_p)$ and $f(w,x|u)$ are nonparametrically identified under quite general assumptions, and \citet*{bonhomme2016estimating} provides a way of estimating these objects. Using these methods we can construct $\hat\overline{S}_g$ and use it as a balancing score.
So far, we have discussed only the case of binary treatments. Applications with non-binary treatments are also common in empirical work, and our strategy has a natural extension for this case. Following most of the applied work, we consider a linear model for the potential outcomes:
where $w$ does not have to be binary. Here heterogeneity in $(\alpha_i,\tau_i)$ can be related to group labels $g(i)$ and covariates $X_i$. Observed outcomes are defined as values of this function at the observed treatments:
We focus on this model for its simplicity, an extension to a more general nonlinear model can be designed in the same way.
Our assumptions remain the same and imply the following unconfoundedness restriction:
Using this restriction, we can consider a general double robust strategy, where we estimate the model for $W_i$ and $Y_i$ using $X_i, \overline{S}_{g(i)}$ as regressors (e.g., chernozhukov2018double). In particular, let $\hat m_{w,i,g}\equiv \hat m_{w,g}(X_i,\overline{S}_{g(i)})$, and $\hat m_{y,i,g}\equiv \hat m_{y,g}(X_i,\overline{S}_{g(i)})$ be the corresponding estimators (with cross-fitting) for the conditional means of $W_i$ and $Y_i$. Define the residuals:
and consider the following regression:
Let $\hat \tau$ be the OLS estimator for $\tau$. Using standard techniques one can show that $\tau$ is a weighted average of $\tau_i$, where the weights are proportional to the variance of $W_i$ conditional on $X_i,\overline{S}_{g(i)}$.
Regression ((ref)) is a direct generalization of the strategy based on ((ref)). To see this observe that representation ((ref)) does not rely on $W_i$ being binary and can be seen as a particular implementation of ((ref)). Given the wide use of ((ref)) in applied work, we view ((ref)) as a natural extension of the standard regression with fixed effects with non-binary treatments.
This discussion demonstrates that, practically, for our strategy, there are few differences between binary and non-binary treatments. Of course, the key restriction here is Assumption (ref) that models the distribution of $(W_i,X_i)$ conditional on $L_{g(i)}$. With binary treatments, the group average $\overline{W}_g$ is the only function of $W_i$ we can consider. Even with non-binary treatments, $\overline{W}_g$ is the only function that is effectively used in ((ref)). The choice set is much larger with general treatments, and one can use other moments, for example, $\overline {W^2}$. The economic model discussed in Section (ref) motivates using such objects and suggests a general way of building a balancing score.
To illustrate our method, we use the data from \citet*{lacetera2012will}\nocite{data,data_2}, which focuses on the effects of various economic incentives on blood donations. To answer this question, the authors leverage a dataset on blood drives -- visits to particular locations -- conducted by the American Red Cross (ARC).For each drive, the authors have information on various outcomes, such as the number of present potential donors, blood units collected, and the number of people turned down. Each blood drive has multiple characteristics, such as the time of the year when it was conducted, its host, weather conditions on that day, and the way it was organized. Additionally, the data includes whether a reward for donation was offered, its type, and cost.
We observe multiple drives with the same host with variation in incentives over them. As a result, the dataset has a grouped structure with unit-level variation in outcomes and treatments: each location defines a group $g$, with a blood drive to this location being a unit-level observation $i$. We use $W_i\in\{0,1\}$ for the absence or presence of economic incentive, covariates $X_i$ to describe the date of the drive and weather conditions on that day and $Y_i$ for the outcome of interest -- the number of collected blood units.
Assumption (ref) is natural in this setup: appealing to institutional knowledge, the authors argue that conditional on the drive characteristics (date), the assignment of incentives is close to being random. Still, the probability of treatment can vary systematically over hosts, and thus conditioning on them -- $L_{g(i)}$ in our notation -- is crucial to claim unconfoundedness. We can expect hosts to differ in their fundamentals -- a potential number of donors and their elasticity with respect to economic incentives, relationship with ARC managers, and organizational abilities -- and $U_g$ captures these characteristics. Importantly, proximity in $U_g$ space might not necessarily be connected to the proximity in observed geographic locations. Our approach allows researchers to capture this unobserved proximity through balancing scores.
Finally, we need to specify the set of balancing scores $S(W_i, X_i)$ that would justify Assumption (ref). In the paper, the authors employ specification ((ref)) which is equivalent to ((ref)), and thus implicitly use $S(W_i, X_i) = (W_i, X_i)$. Following their analysis, we include $X_i$ and $W_i$ but also use interactions of functions of $X_i$ -- indicators for a particular season -- with $W_i$. If specific locations have strong seasonal patterns, we can expect ARC managers to react by adjusting incentives. Interactions then help us to control for this additional selection channel.
We start with $M = 2491$ unique hosts $N = 14029$ blood drives. We use temperature, amount of rain and snow, indicators for weekends, and seasons as characteristics $X_i$. For the interactions with $W_i$ we use season indicators. The dimension of the balancing score is thus equal to $11$. Using constructed $(X_i, \overline{S}_{g(i)})$, we estimate a propensity score model by random forest (\citet*{breiman2001random}) and keep units with estimated probability in $[0.05, 0.95]$. By being very flexible, random forest prediction allows us to drop units from groups without variation in treatment or those where treatment is perfectly predictable by balancing scores. We use this to define the set $\mathbb{A}$, with $\overline{A} = 0.45$. We then use linear regression to estimate conditional mean $\mu(W_i, X_i, \overline{S}_{g(i)})$, which is equivalent to ((ref)). Given the relatively small dimension of covariates, we do not use cross-fitting.
We report our results in Table (ref), where we compare four different estimators: the OLS estimator for ((ref)) without group-level fixed effects, the fixed effects estimator $\hat{\tau}^{\rm fe}$, an inverse-propensity weights (IPW) estimator, and, finally, $\hat \tau_{\rm GME}$. The IPW estimator is defined analogously to the doubly robust estimator ((ref)) but uses $\hat \mu_{g}(\cdot) \equiv 0$. In all cases, we report heteroscedasticity-robust standard errors corrected for within-group correlation.
The results are similar across specifications, with the doubly robust estimator lying between the fixed effects estimator and the IPW estimator. Balancing scores play an important role in predicting the treatment indicators; see Figure (ref) in Appendix (ref). There are also efficiency gains between the IPW and the doubly robust estimator, with $\hat \tau_{\rm GME}$ having a smaller variance.
Similar to the original estimator used in the paper, our estimator $\hat \tau_{\rm GME}$ relies on ((ref)). In particular, we do not allow for spillovers across groups. In lacetera2012will, the authors argue that this assumption might be too restrictive since the supply of blood in neighboring locations can be interdependent. A possible solution for this problem is to aggregate the data at a higher geographical level and proceed as before, but it would be interesting to develop an alternative approach that uses the original data. Since this issue is orthogonal to the subject of the current paper, we do not attempt this development.
We base our simulation results on the dataset described above. In particular, we estimate a separate, host-level propensity score logit model for each location with four regressors: intercept and three seasonal indicators. We drop the hosts where these regressors are collinear, which leaves us with 616 hosts. We then combine the estimated coefficients into 20 groups using the $K$-means algorithm. This produces a model for $e(X_i, \overline{U}_{g(i)})$, where four-dimensional $\overline{U}_{g(i)}$ represents the estimated coefficients for one of the 20 groups. We sample $M=1200$ hosts (with replacement) out of the original ones, keeping all characteristics of the blood drives and their total number fixed, and generate treatment assignments for each host and blood drive using constructed $e(X_i, \overline{U}_{g(i)})$. We keep observed outcomes, which is equivalent to assuming a zero treatment effect for each unit. As a result, this simulation has two sources of randomness: sampled hosts and assignments, while both covariates and outcomes are kept fixed. In particular, this means that the simulation neither satisfies Assumption (ref) with a simple balancing score nor the regression model ((ref)), thus making the comparison between the two methods ambiguous ex-ante.
We simulate this model 100 times, and for each simulation, we compute two estimators: the linear regression ((ref)) and the doubly robust estimator ((ref)). In the regression, we use all available covariates (the same as in the empirical exercise of the previous section), and for the propensity score, we use the standard implementation of random forest with 11-dimensional balancing scores (the same as before). We keep observations with estimated propensity scores between 0.1 and 0.9, treating this set as $\mathbb{A}$. We use ((ref)) to estimate the asymptotic variance. We report the results in Table (ref). One can see that the doubly robust estimator outperforms the linear regression both in terms of bias and root-mean-square error (RMSE). The bias, however, is not negligible, which is partly a manifestation of misspecification and partly a sample size issue. The average over simulations of the estimated standard errors of $\hat \tau_{\rm GME}$ is equal to $0.58$, while the standard deviation of $\hat \tau_{\rm GME}$ over simulations is equal to $0.54$. This suggests that the variance estimator ((ref)) performs well in this simulation.
These results illustrate that our estimator, while not ideal, outperforms the conventional benchmark in the realistic data-driven simulation. We use standard off-the-shelf prediction methods widely available in various statistical packages to construct the estimator. More targeted balancing estimators analogous to those available for cross-sectional data (e.g., \citet*{athey2018approximate,chernozhukov2018adouble, hirshberg2021augmented}) might achieve even better results.
In this work, we proposed a new approach to identification and estimation in observational studies with unobserved group-level heterogeneity. The identification argument is based on the combination of random effects and exponential family assumptions. Given this structure, we can identify a specific average treatment effect even in cases where the observed number of units per group is small. From the operational point of view, our approach allows researchers to utilize all the recently developed machinery from the standard observational studies. In particular, we generalize the doubly robust estimator and prove its consistency and asymptotic normality under common high-level assumptions.