EconBase
← Back to paper

Combining Observational and Experimental Data to Improve Efficiency Using Imperfect Instruments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

61,263 characters · 15 sections · 39 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Combining Observational and Experimental Data to Improve Efficiency Using Imperfect Instruments

abstractRandomized controlled trials generate experimental variation that can credibly identify causal effects, but often suffer from limited scale, while observational datasets are large, but often violate desired identification assumptions. To improve estimation efficiency, I propose a method that leverages imperfect instruments - pretreatment covariates that satisfy the relevance condition but may violate the exclusion restriction. I show that these imperfect instruments can be used to derive moment restrictions that, in combination with the experimental data, improve estimation efficiency. I outline estimators for implementing this strategy, and show that my methods can reduce variance by up to 50%; therefore, only half of the experimental sample is required to attain the same statistical precision. I apply my method to a search listing dataset from Expedia that studies the causal effect of search rankings on clicks, and show that the method can substantially improve the precision.

Introduction

In digital marketing systems, firms often use algorithms to determine marketing variables such as prices, ads, search results, and product recommendations. Accurately estimating how these marketing variables causally affect subsequent outcomes, such as clicks and purchases, is particularly important for firms; otherwise, it is challenging for firms to improve and optimize their algorithms.

However, these causal effects cannot be easily estimated using observational data because the estimates are often biased and inconsistent (lewis_here_2011; blake_consumer_2015; ursu_power_2018; gordon_comparison_2019). These biases and inconsistencies are caused by endogeneity - that there exist factors that influence both the marketing variables and outcomes. To obtain an unbiased estimate, firms often proceed by running an experiment and making their future decisions based on the experiment results. Because experimentation could hurt user experience and profit, experiments often have limited scale, being conducted only within a subset of traffic, markets, or weeks. For example, Google wang_learning_2016, Microsoft li_counterfactual_2015, and Expedia ursu_power_2018 have all conducted experiments that randomized search results rankings for a small subset of traffic. Uber chen_value_2019 has conducted experiments that randomized incentives within one county for three weeks. Advertisers sometimes randomize ad exposure for several weeks to measure advertising ROI lewis_unfavorable_2015. These experiments sometimes produce imprecise estimates due to their limited scale. To illustrate, lewis_unfavorable_2015 examine $25$ advertising field experiments and document that the confidence interval on advertising ROI is wider than $100\%$, implying that advertisers cannot distinguish a campaign with a low $0\%$ ROI from a campaign with a highly profitable $50\%$ ROI. Therefore, it is particularly important to develop methods that improve the estimation efficiency.

This paper proposes a method that combines experimental and observational data to improve estimation efficiency. The method hinges on having a set of imperfect instruments --- pretreatment covariates that satisfy the first-stage relevance condition in the observational data but may violate the exclusion restrictions nevo_identification_2010. A common challenge of leveraging these imperfect instruments is that, due to the violation of exclusion restrictions, the corresponding observational estimate is biased and inconsistent. I demonstrate that this bias can be corrected using experimental data. More importantly, I show that this bias-corrected estimate is uncorrelated with the experiment-only benchmark estimate. Using weighting or GMM, these two uncorrelated estimates can be combined into a more precise estimate without introducing significant bias. I illustrate that the efficiency gain depends on the size of the observational data and how well the pretreatment covariates predict the treatment assignment. When the observational data are large and the pretreatment covariates can accurately predict treatment assignment, I prove that my method can reduce the variance by up to 50%.

The method is particularly useful in the digital marketing system because firms usually have access to a large observational dataset generated by an algorithm. The algorithm often takes many covariates as inputs to endogenously determine the treatment of interest. For example, advertisers use customer characteristics to make targeting decisions, Expedia uses hotel ratings and other hotel characteristics to rank different hotels, and Uber uses supply and demand forecasts to determine how to charge customers and compensate gig workers. Due to the algorithmic nature of the data, any relevant inputs to the algorithm can be viewed as imperfect instruments. Therefore, if researchers observe inputs to the algorithm, or observe any other pretreatment covariates that may correlate with the inputs, researchers can use these inputs or covariates as imperfect instruments to improve the estimation efficiency.

Formally, the method requires two datasets, the experimental sample and the observational sample. For each unit in both samples, we observe the outcome of interest, the treatment, and a set of pretreatment covariates. The only difference between the two samples is that the treatment in the experimental sample is randomized, but the treatment in the observational sample is correlated with the pretreatment covariates that can be viewed as imperfect instruments. Table (ref) illustrates the data requirements.

table[table omitted — 626 chars of source]

I apply the method to estimate how online hotel rankings affect clicks using Expedia datasets in ursu_power_2018. The datasets include both an observational sample that ranks hotels based on relevance and an experimental sample that ranks hotels randomly. In this context, the outcome of interest $Y$ is clicking on a given hotel, the treatment $X$ is the ranking of that hotel, and $Z$ includes rating and other hotel characteristics that may affect its ranking in the observational sample. When compared to the estimate that only uses experimental data, I show my method significantly improves the estimation efficiency by leveraging hotel characteristics as imperfect instruments.

The remainder of the paper is organized as follows: Section 2 reviews the literature. Section 3 presents the setup and key assumptions. Section 4 introduces the estimation method. Section 5 discusses the theoretical efficiency gain. Section 6 extends the method from linear to nonlinear models. Section 7 illustrates the source and magnitude of efficiency gain through simulation. Section 8 applies the method to measure effects of hotel rankings on clicks using datasets from ursu_power_2018. Section 9 concludes.

Literature Review

My paper is related to several contemporary studies that combine experimental and observational data. This strand of literature typically has two objectives: 1) identifying effects that cannot be identified using only experimental data, 2) improving the precision of estimates although experiment is sufficient for identification. Several papers focus on achieving the objective of identification. For example, athey_combining_2020 consider a case when the long-term treatment effect cannot be identified based only on the experimental data, because the long-term outcome in the experimental data is missing. kallus_removing_2018 consider a setting where the treatment effect for a certain population cannot be identified because the experiment does not cover this subpopulation.

My paper focuses on improving efficiency instead of achieving identification. Even though experiments are often sufficient for identifying causal effects of interest, the estimate may be imprecise when large-scale experiments are costly or unrealistic. Given the limited experiment size, it is important to develop methods to improve the efficiency. rosenman2020combining consider a case when the population is partitioned into several strata, and researchers are interested in improving average precision across strata. They show that the mean-squared-error can be improved by shrinking unbiased experimental estimates to additional biased observational estimates. However, when the number of strata is small or the bias is large, the efficiency gain could be small. peysakhovich_combining_2016 consider a case where researchers observe a panel of individuals that appear many times. They estimate individual-level observational estimates and use them as features to differentiate and sort individuals.\footnote{The method also requires a monotonic relationship between observational estimates and true causal effects. } Another popular method by deng_improving_2013 incorporates pre-experimental variables as additional covariates into the experimental analysis. Both deng_improving_2013 and peysakhovich_combining_2016 require observing individuals before and during the experiment, and is conceptually different from my method that is applicable even if individuals are only observed once. My method also has different determinants for efficiency gains. Method of deng_improving_2013 and peysakhovich_combining_2016 are most effective when the pre-experimental data contain variables that are highly predictive of the experimental outcome. My method is most effective when the pre-experimental data contain variables that are highly predictive of the endogenous treatment. My method is complementary because it uses new assumptions that leverage a new source of information, allowing us to improve efficiency even when the number of strata is small or each individual is only observed once in the data. These methods are also not exclusive: they can be applied together if the data requirements and assumptions are met.

My estimation strategy is closely related to that of imbens_combining_1994, who combine macro and micro data to improve estimation efficiency. My method is similar in that one dataset is sufficient for identifying the parameter of interest, and the moment conditions from the additional dataset improve efficiency. However, imbens_combining_1994 consider a setting in which all variables in both datasets are generated from the same distribution. In contrast, I consider a different setting in which the variable of interest is endogenously determined in the observational data but randomized in the experimental data.

My model for observational data is related to the literature by conley_plausibly_2010, nevo_identification_2010, and kippersluis_beyond_2018 that study plausibly exogenous models. In their models, some covariates are considered as “plausible IV" or “imperfect instruments" because they satisfy the relevance condition but may violate the exclusion restriction. conley_plausibly_2010 assume that there is a prior on how much the first-stage covariate violates the exclusion restriction. My method also depends on the availability of imperfect instruments that satisfy the relevance condition. However, my model allows the exclusion restriction to be strongly violated, and does not require a prior on the degree of violation, because this violation can be quantified by the experimental data.

Setup

For illustrative purposes, consider a linear causal model

equation[equation omitted — 76 chars of source]

where $Y_i$ is the outcome, $X_i$ is the focal variable of interest, $Z_i$ is the observed covariate, and $U_i$ is the unobserved covariate. $\beta_1$ and $\beta_2$ are respectively the causal effect of $X_i$ and $Z_i$ on $Y_i$. $\beta_1$ is the primary parameter of interest, and $\beta_2$ is a nuisance parameter that may or may not be 0.

The main difference between the observational and the experimental data is how $X_i$ is generated. Let $G_i \in \{E, O\}$ be the indicator for the data a unit $i$ is drawn from, where $E$ denotes experimental data and $O$ denotes observational data. Assume the first-stage equations for $X_i$ to be:

equation[equation omitted — 136 chars of source]

$\gamma$ is a non-zero first-stage coefficient for the covariate $Z_i$ in the observational data and $V_i$ is the corresponding unobserved error term.\footnote{The first-stage equation does not have to be linear nor causal. $Z_i$ can be thought of as a composite score of consumer characteristics after applying nonlinear rules on multiple covariates, and this nonlinear rule can be estimated using any flexible predictive models. Following the two-stage-least-square method, the coefficient $\gamma$ can also be understood as the best linear predictor for $X$ given $Z$. } Assume $X_i$ in the observational data is correlated with the unobserved covariate $U_i$ so that the observational data has an endogeneity problem:

$$ Cov(X_i, U_i|G_i = O) \neq 0 \; . $$ Also assume $Z_i$ violates the exclusion restriction so that it is not a valid instrumental variable (IV). This violation can be caused by a direct effect on $Y_i$ or a correlation with the unobserved covariate: $$ \beta_2 \neq 0 \; \text{ or }\; Cov(Z_i, U_i|G_i = O) \neq 0 \; . $$ Compared to the observational data, the experimental data can identify $\beta_1$ because $X_i$ is randomized and thus independent of $U_i$: $$ X_i \perp U_i |G_i = E \; . $$ To summarize the basic setup, the experimental data are sufficient to identify $\beta_1$, but the accuracy is limited by the sample size. The observational data alone cannot identify $\beta_1$ because $X_i$ is endogenous and $Z_i$ is an imperfect instrument - a covariate that satisfies the first-stage relevance condition but violates the exclusion restriction. The goal of this paper is to leverage this imperfect instrument $Z_i$ to improve estimation efficiency.

Assumptions

The key assumption is that the joint distributions of $(Z_i, U_i)$ are similar between experimental and observational datasets, such that information learned in one dataset can be transferred to another dataset. One sufficient condition is that the distribution is identical between experimental and observational data:

\let\origtheassumption(ref)$'$ \edef\oldassumption{\the\numexpr\value{assumption}+1}

assumption[Identical Distribution of Covariates] $$ G_i \perp (Z_i, U_i) \; , $$

Assumption (ref) holds if a unit $i$ with attributes $(Z_i, U_i)$ is first randomly sampled from the population and then randomly assigned into the experimental or observational group, independent of $(Z_i, U_i)$. Figure (ref) illustrates an experimental procedure in the digital environment that satisfies this assumption:

figure[figure omitted — 230 chars of source]

This experimental procedure is commonly used for experiments that shuffle the rankings of search results or ads displayed to users. Google wang_learning_2016\footnote{As described in the randomization section of the paper: given a ranked result list of n documents returned for some query, instead of showing the original list, we permute the results uniformly at random and present the shuffled list to a small fraction of end users} and Microsoft li_counterfactual_2015 have conducted experiments that first randomly select a small fraction of users, and then randomize the search results displayed to those users. Expedia ursu_power_2018 has also run such an experiment that shuffles the ranks of hotels displayed to a subset of customers. $JD.com$ carrion_blending_2021 has conducted experiments that randomly shuffle the orders of ads for a random subset of users. In these scenarios, $Z$ can be the relevance score of a website, $X$ is the rank of a website, and $Y$ is the outcome of interest such as a click. By conducting these experiments, firms can answer counterfactual questions such as how likely a customer is going to click a certain webpage when it has a ranking of $X$, a crucial input for firms to improve ranking and user engagement.

Another possibility is that the observational sample includes customers who arrived in the past, while the experimental sample includes customers who arrived more recently. If past customers differ significantly from future customers, firms will have limited incentive to conduct experiments, because findings from past customers have limited value for future strategies. Therefore, for firms that do have incentive to conduct experiments, it is plausible to assume that past and future customers are similar. Given this assumption, combining past observational data and current experimental data can improve efficiency.

The remainder of the paper discusses the estimation strategy that combines datasets that have similar distributions in covariates $(Z_i, U_i)$. For illustrative purposes, I first focus on the estimation strategy for the simple case when the model is linear and the distribution of $(Z_i, U_i)$ is identical. These assumptions allow me to better articulate the intuition of my method as well as analytically quantifying the magnitude of the efficiency gain. Section (ref) and Appendix B.3 discuss how the method can be extended to a class of nonlinear models. Appendix B.1 discusses how Assumption (ref) can be relaxed to allow for the distribution to be similar but non-identical.

Estimation Strategies

Intuition

To develop intuition, consider when $Z_i$ is used as an imperfect IV to estimate $\beta_1$ in observational data. Define the probability limit of this estimator as $b_1^{IV}$:

equation[equation omitted — 101 chars of source]

This estimator is biased and inconsistent because $Z_i$ violates the exclusion restriction, either due to a direct effect on $Y_i$ or an unobserved correlation with $U_i$:

equation[equation omitted — 644 chars of source]

Let $b_2$ denote the violation of such exclusion restriction:

equation[equation omitted — 98 chars of source]

A direct implication of Equation (ref) and (ref) is:

lemmaThe bias of the imperfect IV estimator is $\frac{b_2}{\gamma}$.

When using observational data alone, this bias typically cannot be corrected because the violation of exclusion restriction $b_2$ is unknown and cannot be estimated. However, approximating this violation becomes possible with experimental data by regressing $Y_i$ on $Z_i$. Define the probability limit of this estimated coefficient as $b_2^E$: $$ b_2^E \equiv \frac{Cov(Y_i, Z_i|G_i = E)}{Var(Z_i|G_i = E)}. $$

lemmaUnder Assumption (ref), the violation of exclusion restriction $b_2$ is identified by estimating $b_2^E$ using experimental data.

Proof: $$

alignedb_2^E & \equiv \frac{Cov(Y_i, Z_i|G_i = E)}{Var(Z_i|G_i = E)}\\ & = \frac{Cov(\beta_1 X_i + \beta_2 Z_i + U_i, Z_i|G_i = E) }{Var(Z_i|G_i = E)}\\ & = 0 + \beta_2 + \frac{Cov(U_i, Z_i|G_i = E)}{Var(Z_i|G_i = E)}\\ & = \beta_2 + \frac{Cov(U_i, Z_i|G_i = O)}{Var(Z_i|G_i = O)} \\ & = b_2

$$

To summarize the intuition, the imperfect IV estimate in the observational data is more useful when its bias can be quantified, and the experimental data help quantify such bias. Next I propose two different estimation procedures, weighting and GMM. These two procedures are qualitatively the same, but give different intuitions to help understand the source of efficiency gain.

Weighting

One procedure is to derive two consistent estimators of $\beta_1$ that are uncorrelated and then combine them through weighting. This procedure involves multiple steps visualized in Figure (ref):

enumerate• Regress $Y_i$ on ($X_i$, $Z_i$) in experimental data to obtain $(\widehat{\beta_1}^E, \widehat{b_2}^E)$.\footnote{Note that even though experimental data are insufficient to identify $\beta_2$ due to the potential correlation between $U$ and $Z$, $b_2$ can be identified. So the method focuses on $(\widehat{\beta_1}^E, \widehat{b_2}^E)$ instead of $(\widehat{\beta_1}^E, \widehat{\beta_2}^E)$} • Use $Z_i$ as an imperfect $IV$ in observational data to obtain $\widehat{b_1}^{IV}$. • Regress $X_i$ on $Z_i$ in observational data to obtain $\widehat{\gamma}^O$. • Use Lemma (ref) to correct the bias of the imperfect IV estimator: $\widehat{\beta_1}^O = \widehat{b_1}^{IV} - ({\widehat{b_2}^E}/{\widehat{\gamma}^O})$. • Take a weighted average of the two estimates of $\widehat{\beta_1}^{weighted} = w_O \widehat{\beta_1}^O + w_E\widehat{\beta_1}^E$, where the weight $(w_O, w_E)$ are hyperparameters chosen to minimize $Var(\widehat{\beta_1}^{weighted})$.
figure[figure omitted — 192 chars of source]
lemma$\widehat{\beta_1}^O$ and $\widehat{\beta_1}^E$ are uncorrelated.
lemmaUnder Lemma (ref) the weighting minimizes the variance of the combined estimator, if each estimate is weighted in inverse proportion to its variance: $$ \frac{w_O^*}{w_E^*} = \frac{Var(\widehat{\beta_1}^E)}{Var(\widehat{\beta_1}^O)} $$

I prove Lemma (ref) and Lemma (ref) in Appendix (ref).

This weighting method helps explain the determinants of efficiency gain, which depends on the accuracy of $\widehat{\beta_1}^{O} = \widehat{b_1}^{IV} - \left(\frac{\widehat{b_2}^E}{\widehat{\gamma}^O}\right)$. The accuracy of $(\widehat{b_1}^{IV}, \widehat{\gamma}^O)$ is high if the observational sample is large. When the observational sample tends to infinity such that $(\widehat{b_1}^{IV}, \widehat{\gamma}^O)$ become exact estimates of $(b_1^{IV}, \gamma)$, $\widehat{\beta_1}^{O}$ becomes an unbiased estimate whose only source of uncertainty comes from $Var\left(\frac{\widehat{b_2}^E}{\gamma}\right)$. This remaining uncertainty would be small if the first-stage relevance parameter $\gamma$ is large, or if the variation of $Z_i$ is large such that $Var(\widehat{b_2}^E)$ is small. To summarize, the overall efficiency gain is large if the observational sample is large and includes relevant pretreatment covariates $Z_i$ that can explain variation in treatment assignment $X_i$.

GMM

I also consider a GMM approach, which enables practitioners to easily derive the estimator's analytical solution and its asymptotic properties using established GMM machinery. Equation (ref) is not suitable for direct GMM estimation however, because one of the parameters $\beta_2$ cannot be identified when $Cov(U_i, Z_i) \neq 0$. To derive relevant moment conditions, it is convenient to work with a residual $\epsilon_i$ that is determined by identifiable parameters $(\beta_1, b_2)$:

equation[equation omitted — 81 chars of source]
lemmaIf Assumption (ref) is satisfied, then $$ \epsilon_i = U_i - \frac{Cov(U_i, Z_i)}{Var(Z_i)}Z_i $$ such that $\epsilon_i$ and $Z_i$ are not correlated in both the observational and experimental data: $$ Cov(\epsilon_i, Z_i|G_i = g) = 0 \; \text{for $g \in \{O, E\}.$} $$

I prove Lemma (ref) in Appendix (ref). Note $\epsilon_i$ is different from $U_i$ and can be understood as a transformation of $U_i$ that removes its correlation with the observed covariate $Z_i$. For illustrative purpose, consider the simple case when $(Y_i, X_i, Z_i)$ are already demeaned based on Frisch-Waugh-Lovell theorem such that $E[\epsilon_i] = 0$.\footnote{The idea can be easily extended to cases when the model has an intercept and its variables have non-zero means based on Frisch-Waugh-Lovell theorem. Consider any variables $(\tilde{Y}_i, \tilde{X}_i, \tilde{Z}_i)$ that satisfy $\tilde{Y} = \alpha + \beta_1 \tilde X_i + b_2 \tilde Z_i + \epsilon_i$, this model can be reduced to Equation (ref) by defining $(Y, X, Z)$ as demeaned residuals, where $Y_i = \tilde{Y}_i - E[\tilde{Y}_j|G_j = G_i]$, $X_i = \tilde{X}_i - E[\tilde{X}_j|G_j = G_i]$, $Z_i = \tilde{Z}_i - E[\tilde{Z}_i]$.} Based on Lemma (ref), two moment conditions can be derived from the experimental data $$

alignedE[(Y_i - \beta_1 X_i - b_2 Z_i)X_i|G_i = E] = 0\\ E[(Y_i - \beta_1 X_i - b_2 Z_i)Z_i|G_i = E] = 0\\

$$ and one extra moment condition from the observational data: $$

alignedE[(Y_i - \beta_1 X_i - b_2 Z_i)Z_i|G_i = O] = 0.

$$ Denote the vector of moment functions as $g$:

$$ g_i(\beta_1, b_2) =

bmatrix[bmatrix omitted — 158 chars of source]

$$ and the GMM estimator can be written as

equation[equation omitted — 257 chars of source]

where $W$ is a non-negative definite weighting matrix. The optimal $W^*$ depends on the covariance matrix of the moment conditions, which can be estimated by running the standard feasible GMM. Appendix (ref) derives the analytical solution of this GMM estimator.

Efficiency Gain

In this section I illustrate the efficiency gain using the GMM approach. It is convenient to first formulate the estimation problem in matrix notation:

equation[equation omitted — 109 chars of source]

where each row is $$ Y_i = \beta_1 X_i + b_2 Z_i + \epsilon_i, $$ and the unit is ordered by group $g \in \{E, O\}$ such that $$ [\mathbf{Y}, \mathbf{X}, \mathbf{Z}, \mathbf{\epsilon}] =

bmatrix[bmatrix omitted — 128 chars of source]

$$ Since the exact efficiency gain depends on the distribution of $(X_i, Z_i, \epsilon_i)$, I focus on the simple case of homoscedasticity and nonautocorrelation

assumption{(Homoscedasticity and Nonautocorrelation)} $$ E[\epsilon \epsilon' |\mathbf{X_E}, \mathbf{Z}, \mathbf{G}] = \Sigma = \sigma^2 I, $$

This assumption is not essential for the method to improve efficiency but helps illustrate the determinants of efficiency gain.\footnote{$\mathbf{X_O}$ is also not conditioned on because it is an endogenous variable.}

theoremIf Assumptions (ref) and (ref) are satisfied, an efficient GMM estimator has an asymptotic variance of: $$ \mathbb{V}(\widehat{\beta_1}^{GMM}) = \frac{\sigma^2}{n_E} \left[Var(X_i|G_i = E) + \pi_O Var(\gamma Z_i|G_i = O)\right]^{-1}. $$

The proof for this theorem is provided in Appendix (ref). In comparison, the asymptotic variance of the experiment-only estimator is: $$ \mathbb{V}(\widehat{\beta_1}^{E}) = \frac{\sigma^2}{n_E} \left[Var(X_i|G_i = E) \right]^{-1} $$ A comparison of these two variances suggests that the efficiency gain is determined by 1) the proportion of observational data, 2) the relevance of $Z_i$ as an imperfect instrument, and 3) the relative variance of $\mathbf{X}_E$ and $\mathbf{X}_O$:

lemmaIf $Var(X_i|G_i = O) = Var(X_i|G_i = E)$,\footnote{ This equality assumption is easily satisfied in ranking experiments in online settings, where positions of ads or hotels listed by a website are randomly shuffled. It is also satisfied if the experimental $X$ is sampled from the marginal distribution of the observational $X$. } the relative variance of the two estimators can be written as: \begin{align*} \frac{\mathbb{V}(\hat{\beta}^{E}_1) }{\mathbb{V}(\hat{\beta}^{GMM}_1 )} & = 1 + \pi_{O} \times \frac{Var(\gamma Z_i|G_i = O)}{Var(X_i|G_i = O)} \times \frac{Var(X_i|G_i = O)}{Var(X_i|G_i = E)}\\ & = 1 + \pi_O \times R_{x, z, O}^2 \end{align*} where $\pi_O$ is the proportion of observational data, and $R_{x, z, O}^2$ is the first-stage R-squared that summarizes how well the variation in $Z$ explains the variation in $X$ in the observational data.

Given that $R_{x, z, O}^2 \in [0, 1]$ and $\pi_O \in [0, 1)$, Lemma (ref) implies $ \frac{\mathbb{V}(\hat{\beta}^{E}_1) }{\mathbb{V}(\hat{\beta}^{GMM}_1 )} \in [1, 2)$, which means that the method can reduce the variance by up to $50\%$. This potential improvement implies that half of the experimental data are required to achieve the same accuracy when observational data are incorporated.

Extensions to Nonlinear Model

The paper so far has focused on simple linear models specified in Equation (ref). This section considers how the method can be generalized to a class of nonlinear models:

equation[equation omitted — 87 chars of source]

where $f$ is a flexible nonlinear function parameterized by $\tilde{\theta}$. Analogous to the linear case, $\tilde{\theta}$ may not be directly identified by the experimental data when $Cov(U_i, Z_i) \neq 0$. It is thus convenient to work with a transformed function $h$:

equation[equation omitted — 128 chars of source]

where $\theta$ are the unknown parameters to be estimated.

The function $h(X, Z; \theta)$ can be understood as the conditional average potential outcome for units with attributes $Z$ that receives treatment $X$. Accurately estimating this function is valuable because the function can be used to derive the average treatment effect (ATE), the conditional average treatment effect (CATE), as well as the marginal effect.\footnote{For example, when $X$ is binary, $h(1, Z; \theta) - h(0, Z; \theta)$ is CATE conditional on a specific $Z$, and $E[h(1, Z_i; \theta) - h(0, Z_i; \theta)|G_i = E]$ is the ATE for the experimental population. When $X$ is continuous, $\frac{\partial h(X, Z; \theta)}{\partial X}$ is the marginal effect. } When researchers are interested in effects such as CATE and marginal effects instead of ATE, it becomes even more important to improve precision because the variances of CATE and marginal effects are typically larger than the variance of ATE.

Recall that I focus on cases when experimental data are sufficient for identification. Suppose that the experimental data estimate $\theta$ by minimizing an objective function $L_E(\mathbf{Y}_E, \mathbf{X}_E, \mathbf{Z}_E; \theta)$

equation[equation omitted — 108 chars of source]

The objective function $L_E$ can take different forms. It can be the squared loss function $\sum_{i:G_i = E}(Y_i - h(X_i, Z_i; \theta))^2$, based on the nonlinear least square method. Alternatively, it can be a (negative) likelihood function based on distributional assumptions about the error term, using the maximum likelihood approach. It can also be an objective function based on a set of moment conditions using GMM.

Analogous to the linear case, observational data help improve efficiency because Assumption (ref) implies that another moment condition can be derived from the observational data:\footnote{Assumption (ref) implies $E[U_i|Z_i, G_i = E] = E[U_i|Z_i, G_i = O]$, which means $E[(Y_i - h(X_i, Z_i; \theta)) Z_i|G_i = O] = E[(U_i - E[U_i|Z_i, G_i = E]) Z_i|G_i = O] = E[(U_i - E[U_i|Z_i, G_i = O])Z_i|G_i = O] = 0$. } $$ E[(Y_i - h(X_i, Z_i; \theta)) Z_i|G_i = O] = \mathbf{0} $$ A sample analog of this moment condition is $$ m(\mathbf{Y}_O, \mathbf{X}_O, \mathbf{Z}_O; \theta) = \frac{1}{N_O}\sum_{i:G_i = O}(Y_i - h(X_i, Z_i; \theta))Z_i $$ If the observational data are sufficiently large, then $m$ should be close to $0$ when $\widehat{\theta}$ is accurate. Any deviation from $0$ should be penalized because it implies that $\widehat{\theta}$ is inaccurate. For univariate $Z_i$, this penalization can be incorporated by adding an additional regularization term into the original objective function:

equation[equation omitted — 179 chars of source]

where $\lambda$ is a non-negative hyperparameter that determines how much to penalize the deviation from the moment condition. Intuitively, $\lambda$ depends on how much researchers believe the violation of such constraint can be attributed to the inaccuracy of $\widehat{\theta}$.

When $\lambda = 0$, it implies that researchers ignore information from observational data, and the optimization problem is equivalent to the one using only experimental data. A small value of $\lambda$ can be rationalized if the observational dataset is small, such that the violation of additional moment condition can be explained by sampling error rather than the inaccuracy of $\theta$.

When $\lambda = \infty$, it implies that researchers heavily rely on the constraint. A large $\lambda$ can be justified if the observational dataset is large such that any violation of the moment condition must be attributed to the inaccuracy of $\theta$. When $\lambda = \infty$, the problem also becomes a constraint optimization that helps improve efficiency (imbens_combining_1994).

For multivariate $Z_i$ and a vector of moment conditions $m$, the framework can be extended by taking a weighted average of these moment conditions and minimizing:\footnote{$Z_i$ can also include the intercept, which implies a moment condition $E[(Y_i - h(X_i, Z_i;\theta)1|G_i = O] = 0$.}

equation[equation omitted — 233 chars of source]

where $W_O$ is a weighting matrix that determines how much weight to choose for each moment. The choice of $W_O$ only affects the efficiency improvement and does not affect identification, because experimental data are already sufficient for identification. One possible option of $W_O$ is $(\mathbf{Z}_O'\mathbf{Z}_O)^{-1}$, which mimic the weight in the linear GMM case.

The exact efficiency gain of incorporating these additional moment conditions depends on the functional format of $h$, what types of variables are available, and what the objectives are. Providing an analytical formula for this more general case is beyond the scope of this paper. Section (ref) illustrates how to apply this method in the empirical setting of online ranking and Appendix B.2 further discusses the intuition for efficiency gain from a GMM perspective. The key insight is that the local moment conditions of these nonlinear cases mimic the moment conditions in the linear cases, such that similar efficiency gain could be achieved.

One limitation of this class of non-linear models is that the unobserved $U_i$ is additively separable from the observed $(X_i, Z_i)$. When $U_i$ is not additively separable, it is well known that the endogeneity is difficult to handle and has only been extensively studied in a few special cases. Appendix B.3 discusses how a similar method can still be applied in binary probit model, of which the error term is not directly separable.

Simulation

This section uses simulation to illustrate the effectiveness of the method and the source of efficiency gain. Consider a linear causal model

equation[equation omitted — 81 chars of source]

where $X_i$ is endogenously determined by $(Z_i, V_i)$ in the observational group but randomly assigned in the experimental group: $$ X_i =

cases\gamma Z_i + V_i & if G_i = O\\ drawn from N(0, 1) randomly & if G_i = E\\

$$ Unit $i$ with characteristics $(Z_i, U_i, V_i)$ is drawn from a distribution unknown to researchers: $$ (Z_i, U_i, V_i) \sim N \left(0,

bmatrix[bmatrix omitted — 126 chars of source]

\right) $$ $V_i$ is the residual term in the observational first-stage and therefore by definition uncorrelated with $Z_i$. I calibrate $\sigma_v$ and $\gamma$ such that $V(X_i|G_i = O) = 1 = V(X_i|G_i = E)$ and $R^2_{X, Z, O} = 0.6$, and I assume that the dataset has a small number of experimental units with $n_E = 100$ and a large number of observational units with $n_O = 1900$.

To make the simulation concrete, consider an advertiser that is interested in measuring the return to advertising (lewis_unfavorable_2015). Let $Y_i$ denote the net profit of advertising to customer $i$, $X_i$ denote the ad expenditure on customer $i$, and $(Z_i, U_i)$ be the observed and unobserved customer characteristics that the advertiser uses to target customers and determine expenditure. After certain observational periods, the advertiser becomes interested in measuring the advertising ROI. The advertiser realizes that estimates obtained using past observational data may be biased due to endogeneity and decide to run an experiment, and make future advertising decisions based on the ROI estimated using this experiment. If the ROI estimate is not significantly higher than a certain threshold, the advertiser may choose to stop advertising in the future. Therefore, whether the ROI can be accurately estimated is crucial for advertiser's future decisions. Unfortunately, when the advertising effect is small, it is difficult to precisely estimate advertising ROI. For example, lewis_unfavorable_2015 conducted several large-scale experiments with major U.S. retailers and documented that advertisers may not be able to distinguish a campaign with high ROI of $50\%$ from a campaign with $0\%$ ROI because the median standard error on ROI for their experiment is $26.1\%$.

To simulate such a scenario, I calibrate the simulation exercise such that the ROI is $\beta_1 = 50\%$ and the expected standard error of this ROI is $sd(\beta_1) = 26.1\%$ if it is estimated only using experimental data.\footnote{$sd(\beta_1)$ can be calibrated by adjusting $\sigma_u^2$.} To evaluate the performance of the proposed method, I draw $10,000$ samples, where each sample contains $n_E$ experimental units and $n_O$ observational units. Within each sample $s$, I calculate the experiment-only estimate $\widehat{\beta_1}^{E,s}$ and the combined GMM estimate $\widehat{\beta_1}^{Combine,s}$. Table (ref) gives an example of the regression results estimated using one such sample, where the combined GMM estimator detects that the coefficient is significantly different from 0.

table[table omitted — 661 chars of source]

After performing the estimation for each sample, I then summarize the bias, variance, and MSE across all samples. Figure (ref) compares the distributions of coefficients estimated by these two methods and shows that the combined method is more accurate.

figure[figure omitted — 267 chars of source]

Table (ref) compares the combined GMM estimator with the observational methods including OLS and IV.\footnote{Experimental estimated is obtained by regressing $Y_E$ on $X_E$ and $Z_E$; observational OLS is obtained by regression $Y_O$ on $X_O$ and $Z_O$; IV is obtained by using $Z_O$ as an instrument for $X_O$. } The last column of Table (ref) documents the MSE of these estimators relative to that of the experiment-only estimator. The combined GMM estimator reduces the MSE by almost $1 - 0.607 = 39.3\%$. In contrast, methods based only on observational data have high MSE due to high bias: OLS is biased because $X$ is highly endogenous, and IV is biased because $Z$ is not a valid IV.

table[table omitted — 452 chars of source]

To illustrate the implications of the improved efficiency, Table (ref) compares how often the estimated coefficient has the correct sign (positive), and how often it is statistically significant. Both experimental and combined estimators generate the correct sign for most of the samples. However, compared to the experiment-only estimator, the combined estimator is $64.97\% - 45.24\% = 19.73\%$ more likely to detect that the effect is significant. Suppose the advertiser makes future advertising decisions based on statistical significance, then the advertiser is $19.73\%$ more likely to make the correct decision of keep running this profitable campaign by combining datasets.\footnote{Although it is common for firms to make decisions based on statistical significance, feit_test_2019 highlight that it may not be optimal to do so and derive an optimal profit-maximizing test size. Since my method reduces variance, it can also improve profit by further reducing the profit-maximizing test size }

table[table omitted — 338 chars of source]

Determinants of Efficiency Gain

To investigate the sensitivity of the efficiency gain, I compare the MSE for each method under different values of $\pi_O$ and $R^2_{X, Z, O}$. Figure (ref) shows the relative MSE for different values of $\pi_O$ while fixing the first-stage $R^2_{X, Z, O}$ to be $0.6$. Figure (ref) shows the relative MSE for different values of first-stage $R^2_{X, Z, O}$ while fixing the proportion of observational data at $\pi_O = 0.95$. This result illustrates the prediction of Lemma (ref), that the $MSE(\widehat{\beta_1}^{Combine})$ is small if the proportion of observational data is large and $Z$ is a relevant imperfect instrument.

figure[figure omitted — 625 chars of source]

Empirical Application

I apply my method to estimate the effect of website positions on clicks using the Expedia ranking dataset in ursu_power_2018. The dataset consists of consumers’ search queries for hotels on Expedia, with each search query containing a ranked list of hotels. For each hotel in a query, I observe its rank, characteristics, and click outcome. There are $4.5$ million such observations: two thirds of the data are observational in which the hotel rankings are ordered by relevance, and one third of the data are experimental in which the hotel rankings are randomized. Consider the following model:

equation[equation omitted — 142 chars of source]

where each observation corresponds to a queried hotel $i$ and $Position_i$ is the numerical rank of the hotel in query $i$. To illustrate how to apply the nonlinear model, assume the impact of position on click to be nonlinear: $$ f_1(Position_i; \theta) = \theta_{1} Position_i^{\theta_2} $$ Also assume the position is affected by hotel characteristics in the observational data: $$ Position_{i} =

casesf_2(HotelCharacteristics_{i}) + V_{i} & G_{i} = O\\ Randomized & G_{i} = E

. $$ One challenge is to derive a relevant imperfect instrument $Z_i$ to predict positions that may depend on locations and other characteristics. I select locations that have more than $10,000$ observations and for each location apply a random forest procedure that uses the hotel characteristics to predict the hotel position in the observational data to estimate $f_2$: $$ Z_{i} \equiv f_2(HotelCharacteristics_{i}) $$ The experimental and observational datasets respectively contain $N_E = 452,974$ and $N_O = 833,736$ observations in these selected locations. Table \ref{tab:first_stage_summary} shows the observational first-stage regression in which $Z$ is a relevant covariate that explains $57.5%$ of variation in terms of $R^2$.

table[table omitted — 697 chars of source]

To estimate the parameters $\Theta = \{\alpha, \theta_1, \theta_2, \beta_2\}$, let $h_i(\Theta) = \alpha + f_1(Position_i; \theta)+ \beta_2 f_2(Hotel Characteristics_{i})$. The experiment-only estimate can be derived by minimizing

align*[align* omitted — 90 chars of source]

and the combined estimate can be derived by minimizing

align*[align* omitted — 162 chars of source]

where I choose $\lambda = \frac{1}{Var(Z_i)}$ to mimic the optimal GMM weight in the linear model.

To examine the effectiveness of my method, I focus on a case when the data sample includes a large number of observational units, with $n_O = 190,000$, and a small number of experimental units, with $n_E = 10,000$. This sample can be simulated by randomly drawing $n_O$ units from the observational data and $n_E$ units from the experimental data. I repeatedly draw $10,000$ such samples and perform estimation on each sample. The resulting estimates are then compared to the approximated ground truth, which is estimated using all experimental data and has values of $\theta_1 = 0.14$ and $\theta_2 = -0.23$.\footnote{This approximation is sufficient for evaluating MSE because the size of the original experimental dataset is more than $40$ times the size of the small experimental sample.}

Table (ref) summarizes the distribution of the estimates across these $10,000$ samples. Because of nonlinearity and limited sample size, the distribution of experiment-only estimate for $\theta_1$ is skewed, which may contribute to its larger variance.

table[table omitted — 738 chars of source]

Table (ref) compares the bias, variance, and MSE of 1) experiment-only estimator 2) combined estimator, 3) a nonlinear least square estimator using observational data, and 4) an IV estimator using observational data. Both observational methods outperform the experiment-only estimator in terms of MSE, mainly due to their smaller variance. This difference highlights the importance of variance reduction, especially when sample size is limited and the model is nonlinear. In contrast, the combined estimator has much smaller variance compared to the experiment-only estimator, resulting in a reduction of mean-squared-error by $96.4\%$ and $73.2\%$ for $\theta_1$ and $\theta_2$, respectively. The combined estimator also outperforms the two observational estimators in terms of bias and MSE for $\theta_1$; for $\theta_2$, the combined estimator has relatively less bias, while the observational methods have relatively less variance.

table[table omitted — 964 chars of source]

To assess the relative importance of bias and variance, I summarize how well different methods estimate the marginal effect of changing position, denoted by $\beta_1$: $$ \beta_1\equiv E[\frac{\partial f_1(Position_i; \Theta)}{\partial Position_i}] $$ The results, summarized in Table (ref), demonstrate that observational methods have much worse performance than the experiment-only estimate due to larger bias. In contrast, the combined estimator has relatively smaller bias and much smaller variance, resulting in a 45.8% reduction in MSE.

table[table omitted — 705 chars of source]

Limitations of the Expedia Application

This application has several limitations. As discussed in ursu_power_2018, the dataset has some selection bias and therefore the distribution of queried hotels in the experimental dataset is different from that in the observational dataset. However, this selection bias only supports the robustness of the combined method: even though Assumption (ref) is violated because the observational and experimental units are drawn from slightly different distributions, incorporating the observational sample can still improve the precision of the experimental estimates. In addition to the distributions being different, the reduced form model in Equation (ref) may be misspecified in several ways. First, the model does not impose the constraint that the outcome click must be either $0$ or $1$. Second, I assume that each observation is $i.i.d.$ but the positions and clicks may be negatively correlated within the same query. Alleviating these concerns requires developing a rich structural model of consumer search and estimating it using a full information maximum likelihood approach. Since dealing with model misspecification and developing structural model is not the primary goal of this paper, I leave this to future extension.

Despite these limitations, this dataset has several advantages for evaluating the effectiveness of my methods. First, the dataset contains both observational and experimental data. Second, it contains covariates that satisfy the first-stage relevance condition. Variables, such as rating and price, are likely to affect hotel rankings. Third, the dataset is relatively large, allowing us to approximate the true causal parameter using the entire experimental data. An accurate approximation of the underlying truth allows us to evaluate parameters estimated using only a small subsample of the experimental data. Therefore, this empirical dataset is still valuable for demonstrating the usefulness of the method.

Conclusion

This paper presents a new method that combines experimental and observational data to improve estimation efficiency. I focus on a setting in which the experimental data are sufficient for identification but limited in size, and the observational data are large but suffer from endogeneity. The treatment of interest is randomized in the experimental data, but is endogenously determined in the observational data. I show that if the treatment is affected by or correlated with some pretreatment covariates $Z$, then these covariates can be used as imperfect instruments to derive additional moment conditions to improve estimation efficiency.

The efficiency gain can be substantial under both linear and nonlinear models. For linear models, I prove that the method can reduce the variance by up to $50\%$, with the exact efficiency gain determined by the size of the observational data and the relevance of the imperfect instrument. To illustrate the efficiency gain for nonlinear models with multiple parameters, I apply the method to the Expedia dataset that estimates the effect of rankings on clicks. The method reduces the MSE of some parameters by as high as $97\%$ and reduces the MSE of the overall marginal effect by approximately $46\%$.

Increasing the estimation efficiency is beneficial for firms for several reasons. From a data analysis perspective, increasing the estimation efficiency means that firms can leverage existing datasets to obtain a more precise measurement of an effect, which in turn increases the likelihood of making the correct managerial decision. From an experimental design perspective, higher efficiency means that firms can reduce the required experimental sample size to detect a given effect. For example, a $50\%$ reduction in variance means that firms can reduce the required experimental sample size by $50\%$. This reduction in experimental size is valuable because randomization in experiments may negatively impact user experience by offering suboptimal services or products. Additionally, firms often have a limited number of customers at any given moment, so reducing the required experiment size means that firms can reduce the required time for completing an experiment, which speeds up the learning and decision-making process.

In addition to the application on Expedia ranking, the method of combining datasets can be extended to other settings in which firms observe direct or indirect inputs to algorithms that affect the treatment. For example, past consumption may affect targeted advertising and coupon, customer credit score may affect credit loan interest rate,\footnote{Appendix B.5.1 discusses this application using semi-simulated data from bertrand_whats_2010.} age may affect insurance premium, and demand forecast may affect prices. In these scenarios, firms can leverage the observational data to improve the accuracy of the experimental estimate and speed up learning.

My method is not without limitations. It is most useful when the unobserved error enters the outcome model additively, and when its distributions are similar in the experimental and observational samples. These two conditions allow additional moment restrictions to be derived from the observational data. When these conditions are satisfied, my method strictly decreases the variance without introducing significant bias. When these assumptions are moderately violated, the combined estimator can be thought of as a more efficient estimator that may be biased due to misspecified moment conditions.\footnote{Appendix B.4 discusses how to account for model misspecifications.} Although my method may introduce additional bias in this scenario, my method can still improve the overall MSE by reducing the variance.