EconBase
← Back to paper

Semiparametric Efficiency Gains From Parametric Restrictions on Propensity Scores

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

57,918 characters · 10 sections · 55 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Semiparametric Efficiency Gains From Parametric Restrictions on Propensity Scores

abstractWe explore how much knowing a parametric restriction on propensity scores improves semiparametric efficiency bounds in the potential outcome framework. For stratified propensity scores, considered as a parametric model, we derive explicit formulas for the efficiency gain from knowing how the covariate space is split. Based on these, we find that the efficiency gain decreases as the partition of the stratification becomes finer. For general parametric models, where it is hard to obtain explicit representations of efficiency bounds, we propose a novel framework that enables us to see whether knowing a parametric model is valuable in terms of efficiency even when it is high-dimensional. In addition to the intuitive fact that knowing the parametric model does not help much if it is sufficiently flexible, we discover that the efficiency gain can be nearly zero even though the parametric assumption significantly restricts the space of possible propensity scores.

Introduction

Let $T$ be a treatment indicator such that $T = 1$ if an individual is treated, and $Y_1$ and $Y_0$ denote potential outcomes for $T = 1$ and $T = 0,$ respectively. For each individual, we observe the treatment status $T,$ the realized outcome $Y = T Y_1 + (1 - T) Y_0,$ and a covariate vector $X,$ which takes values on some set $\cX.$ We consider the average treatment effect on the treated, $\mathrm{ATT} = \bE [Y_1 - Y_0 \mid T = 1].$ Suppose that the propensity score $p^\ast (x) = \bP(T = 1 \mid X = x)$ belongs to a finite-dimensional parametric family. As usual, we assume the conditional independence and the overlap condition, which are formally defined in Assumption (ref). We consider three different situations: (K) $p^\ast$ is known, (P) $p^\ast$ is known to belong to the parametric family, but the true parameter is unknown, and (UK) $p^\ast$ is completely unknown. Let $V^k,$ $V^p,$ and $V^{uk}$ be the semiparametric efficiency bounds of the ATT under (K), (P), and (UK), respectively. Clearly, the model under (P) is smaller than the model under (UK) because $p^\ast$ is parametrically specified. On the other hand, the model under (P) is larger than the model under (K) because the true parameter is not available. These observations imply $V^k \leq V^p \leq V^{uk}.$

This kind of comparison among the efficiency bounds for different models has received attention in causal inference. hahn1998role derives the formulas for $V^k$ and $V^{uk}$ and shows the strict inequality $V^k < V^{uk}$ holds generically, which is generalized to multivalued treatment environments by lee2018efficient. Also, chen2008semiparametric derive the formula for $V^p$ in the missing data literature. In this paper, we provide a quantitative and qualitative analysis of the impact of various parametric restrictions on the propensity score on the estimation efficiency of causal parameters defined via general moment conditions, allowing for multivalued treatments.

One of the key contributions of this paper is to obtain explicit formulas of the main components of the efficiency gains $V^{uk} - V^p$ and $V^p - V^k$ when the propensity score is known to be stratified, that is, there is a known finite partition of the covariate space $\cX$ such that the propensity score is constant on each region. Notice that this class of propensity scores can be seen as a parametric family parametrized by the treatment assignment probabilities on each region. The stratified propensity score appears in many empirical studies as reviewed below, and importantly, it covers cases where the propensity score depends only on categorical covariates. The formulas of the efficiency gains provide insightful implications. First, once we know how the covariate space is split, we can match observations that lie in the same region of the partition, which enables the efficient estimation of the parameters of interest conditional on each region. This leads to the efficiency improvement $V^{uk} - V^p,$ as it removes heterogeneity within each subclass. If we know the values of the propensity score in each region in addition to how the space is split, we can efficiently aggregate the local estimates we obtain, which makes the other improvement $V^p - V^k.$ When the partition is coarse, i.e., the model is small, the gain from matching observations within the same region is dominant, and consequently, $V^p$ is close to $V^k.$ As the partition becomes finer, i.e., the model becomes larger, the matching effect decreases, so that knowing the partition becomes less beneficial, and $V^p$ and $V^{uk}$ become similar.

The idea that as the parametric model of the propensity score gets larger, the bound $V^p$ also gets larger should hold beyond the stratified propensity score. Unfortunately, it is hard to obtain explicit formulas of efficiency gains for other parametric models, because $V^p$ involves a complicated matrix inversion, which is solvable only when the propensity score is stratified. Instead of relying on such formulas, we introduce a new framework. We first define a growing sequence of parametric models of the propensity score, keeping the “structure” of the model fixed. We then consider efficiency bounds along the sequence. We are interested in whether the limit of the sequence of efficiency bounds $V^p$ is equal to the nonparametric bound $V^{uk}.$ If it is, it implies that knowing the parametric structure is nearly equivalent to knowing nothing in terms of efficiency, especially when the model is high-dimensional. Otherwise, it means that the parametric structure intrinsically narrows down the space of possible propensity scores. Theorem (ref) shows that $V^p$ approaches $V^{uk}$ if and only if the set of scores of the parametric model approximates two specific functions that are composed of the propensity scores and the mean regression functions of moment functions. This result implies the intuitive fact that if a parametric model is sufficiently flexible, then assuming it is almost equivalent to assuming a nonparametric model. More importantly, it also tells us that as long as the condition of Theorem (ref) is satisfied, it may hold that $V^{uk} \approx V^p,$ even though the parametric model is not much flexible and does not approximate all functions in the nonparametric model. In other words, there are restrictive parametric models under which propensity scores are almost ancillary for estimating parameters of interest.

Related Literature. In the binary treatment setup, the propensity score method is proposed by rosenbaum1983central and rosenbaum1984reducing. hahn1998role computes asymptotic variance bounds for the average treatment effect and the ATT and proposes efficient estimators of them by applying the theory of semiparametric estimation developed by newey1990semiparametric and bickel1993efficient. hirano2003efficient give another efficient estimator called the inverse probability weighted estimator, of which finite sample performances are investigated in herren2023true. For the quantile treatment effect, which is often of interest for measuring inequality, firpo2007efficient obtains similar efficiency results. In the missing data literature, chen2008semiparametric derive the semiparametric efficiency bound for the parameter conditional on the treated subpopulation when the propensity score is correctly specified by a parametric model.

Theoretical studies on the estimation of treatment effects in a multivalued treatment setup are initiated by imbens2000role, who generalizes the framework of Rosenbaum and Rubin. For efficient estimation, cattaneo2010efficient provides the efficient influence function and the semiparametric efficiency bound for parameters defined via general moment conditions, extending results given by hahn1998role and hirano2003efficient. While cattaneo2010efficient focuses on multivalued treatment effects on the whole population, farrell2015robust considers the efficient estimation of the average treatment effect on the treated, and lee2018efficient investigates efficiency bounds for multivalued treatment effects on subpopulation in general, but those for parametric propensity scores are not covered.

In a large fraction of empirical studies, the propensity score is modeled by a parametric model. In observational studies, where the propensity score is not available to analysts, the propensity score is often specified by the logit or probit model. See, for example, yang2016propensity and imai2004causal. Even in experimental studies, an experiment designer often stratifies participants for the sake of estimation efficiency as in, for example, chong2016iron, bugni2019inference and hong2020inference. Propensity scores in stratified experiments can be thought of as being parameterized by assigning probabilities on each stratum.

hahn1998role points out that knowing the propensity score affects the semiparametric efficiency bound of the ATT, while it is ancillary for the estimation of the average treatment effect. frolich2004note explains why this happens qualitatively, which is valid even in multivalued treatment cases. The argument by chen2008semiparametric, who examine the semiparametric efficiency bound when the propensity score is correctly specified by a parametric model in the binary setup, reveals that knowing the parametric model where the true propensity score lives improves the efficiency, but the amount of the improvement is unclear because the efficiency bound has an analytically complicated representation. In Section (ref), we avoid this issue by focusing on the stratified experiment setup, which is analytically tractable. It is also noteworthy that our results also include a direct answer to the conjecture raised in page 325 of hahn1998role: does the knowledge about the propensity score matter in stratified experiments?

The experimental design literature, including cytrynbaum2021designing, tabord2023stratification, and bai2024efficiency among others, has explored the optimal way to stratify the covariate space in order to achieve the efficient estimation by design. In contrast, in Section (ref), we derive the efficiency bound for a fixed stratification and examine how the bound changes as the stratification becomes finer.

To the best of our knowledge, lee2018efficient is the only study that investigates the value of the knowledge of the propensity score in the multivalued treatment setup. The paper does so by considering how propensity scores are incorporated into the efficient influence function. Although lee2018efficient is the closest work to this paper, the paper models the “partial knowledge” of the propensity score differently than we do. lee2018efficient considers the cases where an analyst knows the propensity scores for some treatments but not for the others. On the other hand, we think of the “partial knowledge” as in which parametric model the propensity score lives. In this regard, lee2018efficient and the present paper complement each other.

The theory we propose in Section (ref) may look similar to the notion of sieves in that both consider a growing sequence of parametric models; see chen2007large for a thorough review. However, they have completely distinct goals. Sieves are typically used when the target parameter is defined as the solution of an infinite-dimensional optimization problem, which is difficult to compute in general. The sieve method considers a growing sequence of lower dimensional spaces that approximates the original space. By solving the problem constrained on the approximating space, it provides flexible and robust estimators of the infinite-dimensional parameter. On the other hand, this paper addresses the following question: if the true parameter lives in a low-dimensional parameter space, but an analyst assumes a high-dimensional model, how much efficiency would he or she lose? Thus, we are interested in any increasing sequence of parametric models, including one that does not approximate the nonparametric model.

Setup and Identification

Suppose that there are a finite number of treatments $\cT = \{0, 1, \cdots, J\}$ and corresponding potential outcomes $Y_0, Y_1, \cdots, Y_J \in \bR,$ assuming no interference. Let $T$ be a $\cT$-valued random variable that indicates which treatment is assigned. The observed outcome is, therefore, $Y \coloneqq \sum_{j \in \cT} D_j Y_j$ where $D_j \coloneqq \bI\{T = j\}$ and $\bI \{\cdot\}$ is the indicator function. In addition, a random covariate vector $X,$ which is fully supported on a set $\cX,$ is also available to the analyst. The dataset consists of $W \coloneqq (Y, X, T)$ of all participants. Let $F_X$ and $F_{D, X}$ be the distributions of $X$ and $(D, X),$ respectively, which are not available to the analyst.

We describe the parameter of interest following cattaneo2010efficient and lee2018efficient. Let $\cS \subset \cT$ be a nonempty set of treatment types. The target parameter is $\beta_{\cS}^{\ast} = ((\beta_{0 \mid \cS}^\ast)^\prime, \cdots, (\beta_{J \mid \cS}^\ast)^\prime)^\prime \in \bR^{(J + 1) d_\beta}$ where $\beta_{j \mid \cS}^\ast \in \bR^{d_\beta}$ is defined as a solution of

align*[align* omitted — 76 chars of source]

for a known possibly non-smooth moment function $m : \bR \times \bR^{d_\beta} \to \bR^{d_m}$ where $d_\beta \leq d_m$ allowing for over-identification. In what follows, we omit the subscript $\cS$ of parameters that indicates the conditioning set if it does not cause a confusion, so that $\beta_{\cS}$ and $\beta_{j \mid \cS}$ will be just written as $\beta$ and $\beta_j,$ respectively.

To identify the parameter, we impose the following assumption.

assumptionFor each $j \in \cT,$ $\beta_j = \beta_j^\ast$ is the unique solution of $\bE[m(Y_j; \beta_j) \mid T \in \cS] = 0.$

This framework unifies many interesting causal parameters. In the binary treatment case $\cT = \{0, 1\},$ for example, the average treatment effect corresponds to $m(y_j, \beta_j) = y_j - \beta_j$ and $\cS = \cT.$ Similarly, one can deal with the ATT by considering $\cS = \{1\}.$ For $\tau \in (0, 1),$ the moment function $m(y_j, \beta_j) = \bI\{y_j \leq \beta_j\} - \tau$ leads to the quantile treatment effect at $\tau$th quantile. Also, suppose that there are three treatments: $\cT = \{0, 1, 2\},$ and that a policy maker is considering abolishing treatments $0$ and $1$ and transferring people who are assigned to them to treatment $2.$ In this case, she would be interested in $\bE[Y_2 \mid T \in \{0, 1\}],$ the average potential outcome of treatment $2$ for those with $0$ and $1.$

The propensity score $p_j^\ast(x) \coloneqq \bP(T = j \mid X = x)$ is assumed to be correctly specified by a $d_\gamma$-dimensional smooth parametric model $\gamma \mapsto p_j(\cdot; \gamma).$ That is, $p_j^\ast (\cdot) = p_j(\cdot; \gamma^\ast)$ for some $\gamma^\ast.$ In abuse of notation, we also denote $p_j(\gamma) \coloneqq \bE[p_j(X; \gamma)]$ and $p_j^\ast \coloneqq p_j(\gamma^\ast).$ For $\cS \subset \cT,$ let $p_{\cS} \coloneqq \sum_{j \in \cS} p_j$ and $D_{\cS} \coloneqq \sum_{j \in \cS} D_j.$

Following the literature, we assume the ignorability.

assumption\begin{description} • $T \perp \!\!\! \perp Y_j \mid X$ for all $j \in \cT,$ • There exists $p_{\mathrm{min}} > 0$ such that $p_j^\ast (X) > p_{\mathrm{min}}$ almost surely for all $j \in \cT.$ \end{description}

The first condition ensures that one can compare outcomes with different treatments conditioning on covariates. The second condition is sufficient for identifying the parameter and for the semiparametric efficiency bound to be finite. Under Assumption (ref), the parameter of interest is identified as

align*[align* omitted — 157 chars of source]

where $e_j^\ast (X; \beta_j) \coloneqq \bE [m(Y; \beta_j) \mid X, T = j].$ Note that since $p_{\cS}^{\ast}$ is constant, it is irrelevant to the identification.

The score function $S_j$ for the parametric model of the propensity score is defined as

align*[align* omitted — 115 chars of source]

Similarly, let $S_{\cS} (x; \gamma) \coloneqq \frac{\partial}{\partial \gamma} \log p_{\cS, \gamma} (x).$ Note that $\sum_{j \in \cS} S_j \neq S_{\cS}$ in general. Let $S_j^\ast (x) \coloneqq S_j (x; \gamma^\ast).$ For the regularity of this parametric propensity score, we assume the following condition. For a column vector $a,$ define $a^{\otimes 2} \coloneqq a a^\prime.$

assumption$\bE \left[\sum_{j \in \cT} D_j S_j^\ast (X)^{\otimes 2} \right]$ exists and is invertible.

This condition essentially requires that the Fisher information of the parametric model is invertible, which guarantees that the parameterization $\gamma \mapsto p_j(\cdot; \gamma)$ is not degenerate, i.e., no components of $\gamma$ are redundant. A similar condition is assumed in Theorem 3 of chen2008semiparametric.

As in cattaneo2010efficient, we impose the following assumption.

assumptionFor each $j \in \cT,$ it holds that $\bE \left[ \left\|m (Y_j; \beta_j^\ast)\right\|^2 \right] < \infty,$ where the Euclidean norm is denoted by $\left\|\cdot\right\|,$ and \begin{align*} \cJ_j \coloneqq \frac{\partial}{\partial \beta_j} \ \bigg |_{\beta_j = \beta_j^\ast} \bE\left[m(Y_j; \beta_j)\mid T \in\cS\right] \in \bR^{d_m \times d_\beta} \end{align*} is column full-rank.

This is another condition for the finiteness of the efficiency bound. Note that we implicitly assume that $\bE\left[m(Y_j; \beta_j)\mid T \in\cS\right]$ is differentiable in $\beta_j,$ although the moment function $m$ itself may be non-smooth.

Semiparametric Efficiency Bounds

Derivation

In this section, we derive the efficient influence function and semiparametric efficiency bound for $\beta$ in the framework offered by bickel1993efficient and hahn1998role.

Fix $\cS \subset \cT.$ Define functions $s_j,$ $c_j,$ and $t_j$ as

gather*[gather* omitted — 626 chars of source]

With these functions in hand, let $F^p \coloneqq ((F_0^p)^\prime, \dots, (F_J^p)^\prime)^\prime,$ where

align*[align* omitted — 247 chars of source]

for $j \in \cT.$

The three functions, $s_j,$ $c_j,$ and $t_j$ are key components of the efficient influence function of $\beta$ that is derived in Theorem (ref) below. They correspond to the scores of the distribution of $Y$ conditional on $X$ and $T = j,$ the propensity score, and the marginal distribution of $X,$ respectively.

Also, let

align*[align* omitted — 135 chars of source]

which is the matrix that has $\cJ_0, \cdots, \cJ_J$ on its block diagonal. Note that $\cJ$ is column full-rank under Assumption (ref). The following theorem gives the efficient influence function and efficiency bound of $\beta.$

theoremSuppose that Assumptions (ref)-(ref) hold. The efficient influence function of $\beta$ is \begin{align*} \psi^p (W) \coloneqq - \left(\cJ^\prime \bE \left[F^p (W)^{\otimes 2}\right]^{-1} \cJ\right)^{-1} \cJ^\prime \bE \left[F^p (W)^{\otimes 2}\right]^{-1} F^p (W) . \end{align*} Consequently, the semiparametric efficiency bound is \begin{align*} V^p \coloneqq \left( \cJ^\prime \bE \left[F^p (W)^{\otimes 2}\right]^{-1} \cJ \right)^{-1} . \end{align*}

This result is a direct generalization of Theorem 3 of chen2008semiparametric to the multivalued treatment environment. Indeed, when the treatment variable is binary, i.e., $\cJ = \{0, 1\},$ our result coincides with theirs. Recall that lee2018efficient considers the efficiency in the cases where the propensity score is partially known, i.e., $p_j^\ast$ is known for some $j$ but not for the others. On the other hand, we model the “partial knowldge” of the propensity score by imposing a parametric assumption.

Efficiency bounds are usually used to see whether a given estimator is semiparametrically efficient, i.e., no regular estimator has a smaller asymptotic variance than it does. In addition to this, they can be used to construct efficient estimators. For example, cattaneo2010efficient proposes an efficient estimator of treatment effects based on the moment condition induced by the efficient influence function. Estimators based on efficient influence functions are often called debiased machine learning estimators, and chernozhukov2018double and chernozhukov2022locally show that they are efficient in many examples. These estimators achieve the efficiency by debiasing biases that arise from estimating nuisance parameters, such as propensity scores. Efficient influence functions are useful for constructing efficient estimators, especially when nuisance parameters are high-dimensional and difficult to estimate.

Examples

exampleConsider $m(y_j; \beta_j) = y_j - \beta_j.$ The propensity score is assumed to be correctly specified by a parametric model $p_j (\cdot; \gamma).$ The efficiency bound of $\beta_j^\ast = \bE [Y_j \mid T \in \cS]$ is \begin{align*} V_j^p = \frac{1}{(p_{\cS}^{\ast})^2} \left\{ \bE \left[ \frac{p_{\cS}^{\ast} (X)^2}{p_j^\ast (X)} \sigma_j^2 (X) \right] + \bar c_j^\ast \bE \left[ \sum_{i \in \cT} D_i S_i^\ast (X)^{\otimes 2} \right]^{-1} (\bar c_j^\ast)^\prime + \bE \left[ p_{\cS}^{\ast} (X)^2 (\beta_j^\ast (X) - \beta_j^\ast)^2 \right] \right\} , \end{align*} where $\beta_j^\ast (X) = \bE [Y_j \mid X],$ $\sigma_j^2 (X) = \bV (Y_j \mid X),$ where $\bV$ denotes the conditional variance operator, and \begin{align*} \bar c_j^\ast = \bE \left[ (\beta_j^\ast (X) - \beta_j^\ast) \sum_{i \in \cS} D_i S_i^\ast (X)^\prime \right] . \end{align*} In particular, for the case of $\cT = \{0, 1\}$ and $\cS = \{1\},$ we obtain the efficiency bound for the ATT, $\beta_{1 \mid \{1\}} - \beta_{0 \mid \{1\}},$ \begin{align} V_{\mathrm{ATT}}^p = & \frac{1}{(p_1^\ast)^2} \Bigg\{ \bE \left[ p_1^\ast (X) \sigma_1^2 (X) \right] + \bE \left[ \frac{p_1^\ast (X)^2}{p_0^\ast (X)} \sigma_0^2 (X) \right] \\ \nonumber + & \left(\bar c_{1}^\ast - \bar c_{0}^\ast\right) \bE \left[ \sum_{j \in \cT} D_j S_j^\ast (X)^{\otimes 2} \right]^{-1} \left(\bar c_{1}^\ast - \bar c_{0}^\ast\right)^\prime \\ \nonumber + & \bE \left[ p_1^\ast (X)^2 (\beta_1^\ast (X) - \beta_0^\ast (X) - (\beta_1^\ast - \beta_0^\ast))^2 \right] \Bigg\} . \end{align} This gives a counterpart of Theorem 1 of hahn1998role when there is a parametric restriction on the propensity score.
exampleFor $\tau \in (0, 1),$ consider $m (y_j; \beta_j) =\bI\{y_j \leq \beta_j\} - \tau.$ Then, the moment condition $0 = \bE [m(Y_j; \beta_j^\ast) \mid T \in \cS]$ implies that $\beta_j^\ast$ is the $\tau$-quantile of $Y_j$ conditional on $T \in \cS.$ For a parametric propensity score $p_j (\cdot; \gamma),$ the efficiency bound of $\beta_j^\ast$ is \begin{align*} V_j^p = & \frac{1}{(f_{j \mid T \in \cS} (\beta_j^\ast) p_{\cS}^{\ast})^2} \Bigg\{ \bE \left[ \frac{p_{\cS}^{\ast} (X)^2}{p_j^\ast (X)} \bV (\bI\{Y_j \leq \beta_j^\ast\} \mid X) \right] \\ + & \bar c_j^\ast \bE \left[ \sum_{i \in \cT} D_i S_i^\ast (X)^{\otimes 2} \right]^{-1} (\bar c_j^\ast)^\prime + \bE \left[ p_{\cS}^{\ast} (X)^2 (\bP (Y_j \leq \beta_j^\ast \mid X) - \tau)^2 \right] \Bigg\} \end{align*} where $f_{j \mid T \in \cS}$ is the density of the distribution of $Y_j$ conditional on $T \in \cS$ and \begin{align*} \bar c_j^\ast = \bE \left[ (\bP (Y_j \leq \beta_j \mid X) - \tau) \sum_{i \in \cS} D_i S_i^\ast (X)^\prime \right] . \end{align*} When $\cT = \{0, 1\}$ and $\cS = \{1\},$ a similar calculation yields the efficiency bound for the quantile treatment effect on the treated, $\beta_{1 \mid \{1\}} - \beta_{0 \mid \{1\}}:$ \begin{align*} V_{QTT}^p = & \frac{1}{(p_1^\ast)^2} \Bigg\{ \bE \left[ p_1^\ast (X) \bV \left(\frac{\bI\{Y_1 \leq \beta_1^\ast\}}{f_{1 \mid T \in \cS} (\beta_1^\ast)} \mid X\right) + \frac{p_1^\ast (X)^2}{p_0^\ast (X)} \bV \left(\frac{\bI\{Y_0 \leq \beta_0^\ast\}}{f_{0 \mid T \in \cS} (\beta_0^\ast)} \mid X\right) \right] \\ + & \left(\frac{\bar c_{1}^\ast}{f_{1 \mid T \in \cS} (\beta_1^\ast)} - \frac{\bar c_{0}^\ast}{f_{0 \mid T \in \cS} (\beta_0^\ast)}\right) \bE \left[ \sum_{j \in \cT} D_j S_j^\ast (X)^{\otimes 2} \right]^{-1} \left(\frac{\bar c_{1}^\ast}{f_{1 \mid T \in \cS} (\beta_1^\ast)} - \frac{\bar c_{0}^\ast}{f_{0 \mid T \in \cS} (\beta_0^\ast)}\right)^\prime \\ + & \bE \left[ p_1^\ast (X)^2 \left( \frac{\bP (Y_1 \leq \beta_1^\ast \mid X) - \tau}{f_{1 \mid T \in \cS} (\beta_1^\ast)} - \frac{\bP (Y_0 \leq \beta_0^\ast \mid X) - \tau}{f_{0 \mid T \in \cS} (\beta_0^\ast)} \right)^2 \right] \Bigg\} . \end{align*} This is an analog of equation (10) of firpo2007efficient when the propensity score is parametrically specified.

Basic Comparison

We compare the efficiency bound for a parametric model derived in Theorem (ref) with the bounds for the cases where the propensity score is known or unknown. In this subsection, we overview preliminary comparison results.

According to Corollary 1 of lee2018efficient, the counterparts of $F^p$ when the propensity score is known or unknown are $F^k \coloneqq ((F_0^k)^\prime, \cdots, (F_J^k)^\prime)^\prime$ and $F^{uk} \coloneqq ((F_0^{uk})^\prime, \cdots, (F_J^{uk})^\prime)^\prime$ where

align[align omitted — 177 chars of source]

and

align[align omitted — 285 chars of source]

Therefore, the efficiency bounds are

align*[align* omitted — 274 chars of source]

respectively.

In what follows throughout the paper, we assume the just-identification, i.e., $d_m = d_\beta.$ Although we are interested in the relationship among $V^k,$ $V^p$ and $V^{uk},$ we can focus on the comparison among $\bE \left[F^k (W)^{\otimes 2}\right],$ $\bE \left[F^p (W)^{\otimes 2}\right],$ and $\bE \left[F^{uk} (W)^{\otimes 2}\right],$ as for two positive semidefinite matrices $W_1$ and $W_2,$ it holds

align[align omitted — 197 chars of source]

where for two symmetric matrices $A$ and $B,$ we define $A \geq B$ if and only if $A - B$ is positive semidefinite. Considering the complexity of the models, it is qualitatively obvious that

align[align omitted — 193 chars of source]

First, consider the case of $\cS = \cT.$ It holds that $F^p = F^k$ because it holds $c_j(\beta_j^\ast, e_j^\ast, \gamma^\ast) = 0,$ since

align*[align* omitted — 411 chars of source]

We also have $F^{uk} = F^p,$ as $D_{\cS} - p_{\cS}^{\ast} (X) = 0.$ These observations lead to a well-known result: the relationship ((ref)) holds all with equality. In other words, knowing the propensity score does not improve the efficiency, as discovered by hahn1998role and lee2018efficient.

In what follows, we assume $\cS \subsetneq \cT$ unless stated otherwise. The connection between $F^p$ and $F^{uk}$ can be understood geometrically.

propositionThe function $F^p$ is the $L^2$-projection of $F^{uk},$ that is, for each $j \in \cT,$ the unique solution of \begin{align*} \min_{c_j \in \bR^{d_m \times d_\gamma}} \bE \left[ \left\| F_j^{uk} (W) - \left( D_j s_j(W; \beta_j^\ast, e_j^\ast, \gamma^\ast) + c_j \sum_{i \in \cT} D_i S_i^\ast(X) + t_j(W; \beta_j^\ast, e_j^\ast, \gamma^\ast) \right) \right\|^2 \right] \end{align*} is given by $c_j = c_j (\beta_j^\ast, e_j^\ast, \gamma^\ast).$

The orthogonality between $F^p$ and $F^{uk} - F^p$ admits the decomposition

align*[align* omitted — 172 chars of source]

Hence, if the minimization problem of Proposition (ref) does not attain zero, which is the case in most cases, then the second inequality of ((ref)) strictly holds.

Efficiency Gains in Stratified Experiments

In this section, we take a quantitative approach to investigate how much a parametric restriction on the propensity score improves the efficiency bound of parameters of interest in stratified experiments. When the propensity score is known to be stratified, the efficient influence function and the semiparametric efficiency bound turn out to have closed forms. After deriving analytic formulas, we discuss efficiency gains from knowing the restriction on the propensity score.

Let us first define the class of stratified propensity scores. To do so, fix a measurable partition $\cX = \bigsqcup_{k = 1}^K \cX_k$ of the space of covariate vectors. Consider the following class of propensity scores

align*[align* omitted — 77 chars of source]

for some constants $p_{j, k}.$ We call this form of propensity score a $K$-stratified propensity score with partition $\bigsqcup_{k = 1}^K \cX_k.$ This class is parameterized by $(JK)$-parameters that live in

align*[align* omitted — 233 chars of source]

The probabilities for $j = 0$ are specified as $p_{0, k} = 1 - \sum_{j = 1}^J p_{j, k}$ for $k = 1, \cdots, K.$ Let $p_j^\ast (x) = \sum_{k = 1}^K p_{j, k}^\ast \bI\{X \in \cX_k\}$ be the true propensity score. We denote $p_{\cS, k}^\ast \coloneqq \sum_{j \in \cS} p_{j, k}^\ast.$

The class of stratified propensity scores covers many important situations. For example, it includes cases where the propensity score is known to depend only on categorical covariates such as race, gender, month of birth, city of residence, and so on. Even when covariates are continuous, people often discretize them in practice, which also falls within the scope of the current setting.

As in Section (ref), let $V^k, V^p$ and $V^{uk}$ be the efficiency bounds of $\beta$ when the propensity score is known, known to be $K$-stratified by partition $\bigsqcup_{k = 1}^K \cX_k,$ and unknown, respectively. We are interested in the relationship among these bounds, but thanks to ((ref)), we focus on the comparison among $\bE \left[F^k(W)^{\otimes 2}\right],$ $\bE \left[F^p (W)^{\otimes 2}\right],$ and $\bE \left[F^{uk} (W)^{\otimes 2}\right].$ We regard $\Delta_K^{uk \to p} \coloneqq \bE \left[F^{uk} (W)^{\otimes 2}\right] - \bE \left[F^p (W)^{\otimes 2}\right]$ as the efficiency gain from knowing that the propensity score is $K$-stratified, and similarly, $\Delta_K^{p \to k} \coloneqq \bE \left[F^p (W)^{\otimes 2}\right] - \bE \left[F^k (W)^{\otimes 2}\right]$ is the value of knowing the propensity score on each subclass. The stratification structure allows us to obtain analytical formulas of these quantities.

theoremIt holds that \begin{align*} \Delta_K^{uk \to p} = \frac{1}{(p_{\cS}^{\ast})^2} \sum_{k = 1}^K p_{\cS, k}^\ast (1 - p_{\cS, k}^\ast) \bP (X \in \cX_k) \bV \left( \begin{pmatrix} e_0^\ast (X; \beta_0^\ast) \\ \vdots \\ e_J^\ast (X; \beta_J^\ast) \end{pmatrix} \ \bigg | \ X \in \cX_k \right) , \end{align*} and \begin{align*} \Delta_K^{p \to k} = \frac{1}{(p_{\cS}^{\ast})^2} \sum_{k = 1}^K p_{\cS, k}^\ast (1 - p_{\cS, k}^\ast) \bP (X \in \cX_k) \left\{ \bE \left[ \begin{pmatrix} e_0^\ast (X; \beta_0^\ast) \\ \vdots \\ e_J^\ast (X; \beta_J^\ast) \end{pmatrix} \ \bigg | \ X \in \cX_k \right] \right\}^{\otimes 2} . \end{align*}
remarkThe formulas in Theorem (ref) are explicit in the sense that the matrix inversion that appears in $c_j (\beta_j^\ast, e_j^\ast, \gamma^\ast)$ is resolved. This happens only for the stratified propensity score and cannot be extended to other parametric models. The reason is as follows. Recall that we are interested in the inverse of \begin{align*} \bE \left[ \left( \sum_{i \in \cT} D_i S_i^\ast (X) \right)^{\otimes 2} \right] = \sum_{i \in \cT} \bE \left[ D_i S_i^\ast (X)^{\otimes 2} \right] . \end{align*} In general, it is difficult to find an explicit formula for the inverse of the sum of matrices, like the one in the RHS, and this causes a problem in most parametric models. For the stratified propensity score, however, the sum can be nicely decomposed into diagonal matrices, and it turns out to be explicitly invertible using the Woodbury formula. This calculation exploits the property of stratified experiments that the parameter $p_{j, k}$ affects the propensity score only on the $k$th cell $\cX_k.$ See the Supplementary Material for more technical details.

As we have discussed the case of $\cS = \cT$ in Section (ref), we focus on $\cS \subsetneq \cT$ in what follows. Notice that a $1$-stratified propensity score coincides with a constant propensity score, in which cases treatments are assigned completely at random. As the two formulas in Theorem (ref) suggest, for $K = 1,$ it holds

align*[align* omitted — 359 chars of source]

An implication of this result is that the knowledge that the propensity score is constant over the whole covariate space strictly improves the efficiency, i.e., $\Delta_1^{uk \to p} > 0,$ but full information of the propensity score does not bring any additional value, i.e., $\Delta_1^{p \to k} = 0.$ This is not the case for $K > 1,$ because $\Delta_K^{p \to k}$ is strictly positive in general. In other words, knowing the specific values of the propensity score does improve the efficiency. These observations generalize hahn1998role who finds a similar fact in binary treatment cases.

In the stratified experiment setup, the value $\Delta_K^{uk \to p}$ of imposing a parametric restriction captures the in-class variance while the additional value $\Delta_K^{p \to k}$ of pinning down the propensity score embodies the between-class bias. Too see this, suppose first that we know that the true propensity score is $K$-stratified by a given partition. In this case, data in each subclass can be thought of as sample from a random assignment. It holds that

align*[align* omitted — 194 chars of source]

that is, the moment for subpopulation in each subclass can be estimated based on all data points including observations to which the treatments in $\cS$ are not assigned. Since all information in each subclass is exploited in this way, the knowledge of the partition removes the in-class efficiency loss. Note that frolich2004note provides a similar observation for binary average treatment effects on the treated.

Even if the functional form of the propensity score is known, there may be another source of efficiency loss that we call the between-class bias. To see this, notice that

align*[align* omitted — 241 chars of source]

This decomposition implies that the moment function of interest is an aggregation of local conditional expectations in each subclass with a weight depending on the propensity score. Therefore, one cannot reproduce the moment function efficiently without knowing the exact values of the propensity score, which leads to efficiency loss. There is no loss in the absence of knowledge of the propensity score only when $\bE [e_j^\ast (X; \beta_j) \mid X \in \cX_k]$ is constant or $K = 1.$ The former case does not hold generically. The latter is consistent with the result from hahn1998role, who shows that knowing that the propensity score is constant is the same as knowing its value in terms of efficiency, but our result also indicates that $K = 1$ is just a special case in that the between-bias is zero despite the fact that the exact value of the propensity score is unavailable.

At the end of this section, we consider the behavior of the efficiency gain $\Delta_K^{uk \to p}$ from knowing the partition when it is infinitely fine. Intuitively, if cells of the partition are tiny, just knowing the partition is almost useless because the in-class variance of each cell is small. This observation is justified in the following proposition by considering a nested sequence of partitions. Assume $\cX$ is a subset of a Euclidean space, and let $\mathrm{vol} (\cdot)$ be the Lebesgue measure on it.

propositionLet $K (n)$ be an increasing sequence of natural numbers. For each $n \geq 1,$ let $\{\cX_k^{(n)}\}_{k = 1}^{K (n)}$ be a partition of $\cX$ satisfying the following condition: for any $n \geq 2$ and $k \in \{1, \dots, K (n)\},$ there uniquely exists $k^\prime \in \{1, \cdots, K (n - 1)\}$ such that $\cX_{k^\prime}^{(n - 1)} \supset \cX_k^{(n)}$ and $\lim_{n \to \infty} \max_{1 \leq k \leq K (n)} \mathrm{vol} (\cX_k^{(n)}) = 0.$ Then, it holds that $\lim_{n \to \infty} \Delta_{K (n)}^{uk \to p} = 0.$

On the other hand, if the partition is known to be coarse, the efficiency is improved, as the following proposition claims.

propositionIn the same setup as Proposition (ref), assume $\lim_{n \to \infty} \max_{1 \leq k \leq K (n)} \mathrm{vol} (\cX_k^{(n)}) > 0.$ Then, it holds that $\lim_{n \to \infty} \Delta_{K (n)}^{uk \to p} > 0$ if $e_j^\ast (\cdot; \beta_j^\ast)$ is not constant on any open ball in $\cX.$

The assumption of Proposition (ref) describes the situation where it is known that there is a unignorable region where assignments are completely at random. In this case, the propensity score is restricted to the space of functions that are constant on the region, which leads to an efficiency gain.

Propositions (ref) and (ref) are shown from the formulas that are derived in Theorem (ref), but they can also be proven using the general theory in the next section. See the Supplementary Material for details in this direction.

Efficiency Gains from Large Parametric Models

General Theory

We have shown in Section (ref) that imposing a parametric restriction on the propensity score weakly improves estimation efficiency, but as shown in Propositions (ref) and (ref), the amount of the efficiency gain varies depending on the parametric model assumed. In order to see what kinds of parametric restrictions are valuable in terms of efficiency, it is helpful to compute the reduction $\bE \left[F^{uk} (W)^{\otimes 2}\right] - \bE \left[F^{p} (W)^{\otimes 2}\right]$ in the efficiency bound as in Section (ref). Unfortunately, however, it is hard to find an explicit formula of this quantity in general because it involves a convoluted matrix inversion in $\bE \left[F^{p} (W)^{\otimes 2}\right].$ In this section, we take a more conservative approach, instead of calculating it explicitly. Specifically, we investigate how the amount of the efficiency improvement changes as the parametric model becomes infinitely “large.” To formally study asymptotics for parametric models, we first define a countably infinite-dimensional model, which is thought of as the “structure” of the parametric model, and then restrict it to obtain a sequence of finite-dimensional parametric submodels in a way that they are nested. We consider efficiency bounds along this sequence and see whether their limit attains the bound for the unknown propensity score. If it does, then it implies that imposing the parametric model is almost the same as estimating the propensity score nonparametrically, especially when $n$ is large, so that the parametric restriction does not help significantly in terms of efficiency. If the nonparametric bound is not attained even asymptotically, on the other hand, it means that the parametric structure intrinsically restricts the space of possible propensity scores.

As we have seen in the previous section, when $\cS = \cT,$ there is no efficiency gain from knowing the propensity score. Since the object of interest here is the difference $\bE \left[F^{uk} (W)^{\otimes 2}\right] - \bE \left[F^{p} (W)^{\otimes 2}\right],$ we assume $\cS \subsetneq \cT$ throughout this section. We also assume $((e_0^\ast)^\prime, \dots, (e_J^\ast)^\prime)^\prime \not \equiv 0;$ otherwise, the knowledge about the propensity score does not matter either.

To describe situations where the number of parameters grows with the “structure” of a parametric model fixed, we consider a nested family of parametric models that is induced by a large base model. Let $\Theta^\bN$ be a countably infinite-dimensional parameter space where $\Theta \subset \bR,$ which defines a base model of the propensity score, $\cM_\infty^p \coloneqq \left\{(p_j (\cdot; \gamma))_{j \in \cT} \mid \gamma \in \Theta^\bN\right\}.$ Let $d (n) \nearrow \infty$ be an increasing sequence. Define a $d (n)$-dimensional parametric submodel of $\cM_\infty^p$ as

align*[align* omitted — 233 chars of source]

That is, the submodel $\cM_n^p$ is parameterized by the first $d (n)$ elements of $\gamma,$ which is denoted by $\gamma^{(n)} = (\gamma_1, \dots, \gamma_{d (n)}).$ It is obvious that these models are nested in the sense that

align*[align* omitted — 135 chars of source]

where $\cM^{uk}$ is the set of all possible propensity scores. Moreover, assume that the true propensity score $(p_j^\ast)_{j \in \cT} = (p_j(\cdot; \gamma^\ast))_{j \in \cT}$ is of finite dimension, i.e., $\gamma^\ast = (\gamma_1^\ast, \dots, \gamma_{d (N^\ast)}^\ast, 0, 0, \dots)$ for some $N^\ast.$

Let $V^{p, n}$ be the semiparametric efficiency bound in Theorem (ref) when the propensity score is assumed to lie in $\cM_n^p.$ Then, $V^{p, n}$ is nondecreasing in $n$ in the sense of positive semidefinite matrix since $\cM_n^p$ is nondecreasing. It is also bounded above by $V^{uk},$ which is the semiparametric efficiency bound for $\cM^{uk},$ so that the sequence $V^{p, n}$ converges to some positive definite matrix denoted by $V^{p, \infty}.$ This is shown in the Supplementary Material. It holds in general that $V^{p, n} \nearrow V^{p, \infty} \leq V^{uk}.$

The score of the parametric model $\cM_n^p$ is

align*[align* omitted — 159 chars of source]

Also, define a $d (n) \times J$ matrix valued function

align*[align* omitted — 310 chars of source]

For $j \in \cT$ and $\ell \in \{1, \dots, d_m\},$ let

align*[align* omitted — 292 chars of source]

where $e_j^\ast = (e_{j, 1}^\ast, \dots, e_{j, d_m}^\ast)^\prime.$ Note that $\bar \ve_{j, \ell}$ is a vector consisting of at most two functions. In particular, its entries indexed by $i \in \cS$ have $(1 - p_{\cS}^{\ast} (\cdot)) e_{j, \ell}^\ast (\cdot; \beta_j^\ast),$ and the others have $- p_{\cS}^{\ast} (\cdot) e_{j, \ell}^\ast (\cdot; \beta_j^\ast).$

Now, we introduce the following condition.

conditionp{F} For any $j \in \cT$ and $\ell \in \{1, \dots, d_m\},$ it holds that \begin{align*} \lim_{n \to \infty} \inf_{c_n \in \bR^{d (n)}} \left\|c_n^\prime \bS_n - \bar \ve_{j, \ell}\right\|_{(L^2(F_X))^J} = 0 . \end{align*}

As we see in the Supplementary Material, the two functions $(1 - p_{\cS}^{\ast} (\cdot)) e_{j, \ell}^\ast (\cdot; \beta_j^\ast)$ and $- p_{\cS}^{\ast} (\cdot) e_{j, \ell}^\ast (\cdot; \beta_j^\ast),$ which form $\bar \ve_{j, \ell},$ appear in the efficient influence function of $\beta$ under the nonparametric model $\cM^{uk}.$ It is also shown that the function $c_n^\prime \bS_n$ is the counterpart of $\bar \ve_{j, \ell}$ under the parametric model $\cM_n^p.$ Therefore, Condition (ref) implies that the sequence of efficient influence functions of the increasing parametric models converges to that of the nonparametric model.

It is intuitive that the efficiency bound of the limit parametric model $\cM_\infty^p$ is equal to that of the nonparametric model if the parametric model is sufficiently “flexible.” Condition (ref) can be thought of as a criterion of such flexibility and is shown to characterize when a parametric restriction on the propensity score does not improve the efficiency bound asymptotically.

theoremCondition (ref) holds if and only if $V^{p, \infty} = V^{uk}.$

Theorem (ref) tells us that knowledge of a parametric restriction on the propensity score that violates Condition (ref) is beneficial in terms of efficiency because it improves the efficiency bound even asymptotically, i.e., $V^{p, \infty} < V^{uk}.$ It also implies that as long as Condition (ref) is satisfied, it may hold that $V^{p, \infty} = V^{uk},$ even though $\cM_\infty^p \subsetneq \cM^{uk}.$ In other words, restricting the space of propensity scores does not necessarily improve estimation efficiency.

Application to Logistic Propensity Scores

In this subsections, we give two examples that satisfy or violate Condition (ref) for logistic propensity scores.

Full-rank logistic propensity scores. Consider a parameter

align*[align* omitted — 238 chars of source]

Let $(b_n)_{n \in \bN} \in (L^2 (F_X))^\bN$ be a dictionary of functions of covariate $x.$ Let

align*[align* omitted — 160 chars of source]

The object of interest here is a full-rank logistic propensity score that is defined as

align*[align* omitted — 286 chars of source]

for $1 \leq j \leq J,$ and $p_0 (x; \gamma) = 1 - \sum_{j = 1}^J p_j (x; \gamma).$ In this example, the base model is $\cM_\infty^p = \{(p_j (\cdot; \gamma))_{j \in \cT} \mid \gamma \in \bR^\bN\}.$ By restricting the space of $\gamma,$ we can obtain a nested sequence of finite-dimensional parametric models as follows:

align*[align* omitted — 549 chars of source]

Note that $\cM_n^p$ is $d (n)$-dimensional where $d (n) = nJ.$ Suppose that the true propensity score is of order $N^\ast,$ i.e., $(p_j^\ast)_{j \in \cT} \in M_{N^\ast}^p$ with the true parameter $\Gamma^\ast =

pmatrix[pmatrix omitted — 71 chars of source]

$ and the corresponding $\gamma^\ast.$ The following result provides a necessary and sufficient condition for the bound for this parametric model to approach the nonparametric bound.

propositionFor a full-rank logistic propensity score, $V^{p, \infty} = V^{uk}$ if and only if the linear span of $(b_n)_{n \in \bN}$ contains \begin{align*} \frac{\bI\{i \in \cS\} - p_{\cS}^{\ast} (X)}{1 - p_i^\ast (X)} e_{j, \ell}^\ast (X; \beta_j^\ast) \end{align*} for all $i, j \in \cT$ and $\ell \in \{1, \dots, d_m\}.$

An implication of this proposition is that if the true propensity score is indeed full-rank logistic with $(b_n)_{n \in \bN}$ that spans $L^2 (F_X),$ knowing this does not improve the efficiency bound when the parametric model has many parameters. In other words, such a parametric restriction provides no efficiency gain asymptotically. However, the condition of the statement does not require $b_n$ to span $L^2 (F_X),$ in which case, the parametric model of the propensity score is restrictive compared to the nonparametric one. Proposition (ref) claims that knowing unignorable aspects of the propensity score does not necessarily improve the efficiency.

Note that Proposition (ref) is a similar result to Proposition (ref). Both propositions tell us that if the propensity score is known to lie in a large parametric model, estimating parameters of interest in the parametric model is as hard as doing so nonparametrically.

Degenerate logistic propensity scores. As an example that violates Condition (ref), we define a degenerate logistic propensity score as

align*[align* omitted — 329 chars of source]

for $j = 1, \dots, J$ where $\tilde \gamma \in \bR^\bN.$ This model is obtained by restricting the parameter space of the full-rank logistic propensity score by setting $\Gamma_{n, j} = \tilde \gamma_n.$ Under this model, the propensity scores for $j = 1, \dots, J$ are all identical.

propositionFor a degenerated logistic propensity score, $V^{p, \infty} < V^{uk}$ if (i) $\cS \cap \{1, \dots, J\} \neq \emptyset$ and (ii) $\{1, \dots, J\} \setminus \cS \neq \emptyset.$

To understand the assumption of the proposition, recall that $\emptyset \subsetneq \cS \subsetneq \cT,$ and assume $0 \in \cS$ for simplicity. Then, condition (ii) is automatically satisfied; otherwise $\cS = \cT.$ Condition (i) requires $\cS$ to have at least one element other than $0.$ In the case of $\cT = \{0, 1, 2\},$ this is satisfied, for example, when we are interested in treatment effects on those who are assigned to $\cS = \{0, 1\}.$ Notice that condition (i) is never satisfied for binary treatments, $\cT = \{0, 1\}$ unless $\cS = \cT.$ Indeed, $V^{p, \infty} = V^{uk}$ holds for such cases, if the dictionary $(b_n)_{n \in \bN}$ spans $L^2 (F_X).$ See the Supplementary Material for further details.

Proposition (ref) implies that under the obviously verifiable conditions, the space of degenerated logistic propensity scores is restrictive, and consequently, assuming this structure improves the estimation efficiency even asymptotically.