EconBase
← Back to paper

Assessing Outcome-to-Outcome Interference in Sibling Fixed Effects Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

45,713 characters · 7 sections · 20 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Assessing Outcome-to-Outcome Interference in Sibling Fixed Effects Models

\pagenumbering{gobble}

center[center omitted — 31 chars of source]

Sibling fixed effects (FE) models are useful for estimating causal treatment effects while offsetting unobserved sibling-invariant confounding. However, treatment estimates are biased if an individual’s outcome affects their sibling’s outcome. We propose a robustness test for assessing the presence of outcome-to-outcome interference in linear two-sibling FE models. We regress a gain-score—the difference between siblings' continuous outcomes—on both siblings' treatments and on a pre-treatment observed FE. Under certain restrictions, the observed FE’s partial regression coefficient signals the presence of outcome-to-outcome interference. Monte Carlo simulations demonstrated the robustness test under several models. We found that an observed FE signaled outcome-to-outcome spillover if it was directly associated with a sibling-invariant confounder of treatments and outcomes, directly associated with a sibling’s treatment, or directly and equally associated with both siblings’ outcomes. However, the robustness test collapsed if the observed FE was directly but differentially associated with siblings’ outcomes or if outcomes affected siblings’ treatments.

\scriptsize

Acknowledgements: This work was supported by the Eunice Kennedy Shriver National Institute for Child Health and Human Development through the Population Research Center at the University of Texas-Austin (P2CHD042849) and through the Center for Demography at the University of Wisconsin-Madison (T32 HD007014-42). This content is solely the responsibility of the author and does not necessarily represent the official views of the National Institutes of Health. I thank Felix Elwert for comments and discussions that informed this manuscript. Conflicts of interest: none.

\pagenumbering{arabic} \setcounter{page}{1}

Introduction

Sibling fixed effects (FE) models are commonplace in observational epidemiologic research for estimating causal treatment effects while controlling for unobserved sibling-invariant confounding, such as shared genetic risk factors.frisell2012, khashan2014, khashan2015, hvolgaard2016, sjolander2016, hanley2017, axelsson2019, mallinson2020, petersen2020, frisell2021, mallinson2021 These models rely on strong assumptions—in particular, an individual's outcome cannot affect their sibling's treatment nor outcome.frisell2012, sjolander2016, frisell2021, mallinson2021 Such assumptions are precarious because siblings' health are interdependent.feinberg2012, deneve2017 One can prevent outcome-to-treatment interference by measuring all treatments before all outcomes within sibships. However, study design cannot circumvent outcome-to-outcome interference, thereby threatening the validity of any sibling FE analysis. Statistical tests like regressions with negative controls can assess potential bias,lipsitch2010, cunninghambook although there is little guidance for their use in gauging outcome-to-outcome interference in sibling comparison studies.

We explain a robustness test for assessing outcome-to-outcome interference in linear two-sibling FE models. The method involves regressing a gain-score—the difference in siblings' outcomes—on both siblings' treatments and on a pre-treatment observed FE that is related to siblings’ treatments or outcomes. Using graphical causal modeling and Monte Carlo simulations, we demonstrate when the partial regression coefficient for the observed FE signals outcome-to-outcome interference across several models.

Methods

Baseline model

Our baseline sibling FE model (Figure 1A) reflects a typical sibling comparison design.sjolander2016 Subscripts $i=\left\{1,\ldots,N\right\}$ and $j=\left\{1,2\right\}$ denote cluster and sibling, respectively. $T_{ij}$ is a binary or continuous treatment, $Y_{ij}$ is a continuous outcome, $U_i$ is an unobserved FE (e.g., family health history), and $D_i=Y_{i2}-Y_{i1}$ is the gain-score. Arrows and adjacent Greek letters denote causal effects: $\delta$ is the treatment effect ($T_{ij}\rightarrow Y_{ij}$); $\psi$ is the unobserved confounding effect on outcomes ($U_i\rightarrow Y_{ij}$); and $\chi$ ($U_i\rightarrow T_{i1}$) and $\gamma$ ($U_i\rightarrow T_{i2}$) are unobserved confounding effects on treatments. This model embeds two assumptions. First, the treatment effect $\delta$ and the confounding effect $\psi$ are equal for both siblings. Second, all effects are linear and homogeneous. Both assumptions align with conventional linear FE models.sjolander2016, kim2019, kim2021, gunasekara2014,imai2019 For accessibility, we omitted sibling-specific covariates, which are compatible with the baseline model and all subsequent models as long as one can condition on them without loss of generality.

\afterpage{

figure[figure omitted — 5,579 chars of source]

}

Gain-score regression

Following the baseline model, $\delta$ is our estimand. Regressing $Y_{ij}$ on $T_{ij}$ will not identify $\delta$, covariate adjustment notwithstanding, because $U_i$ confounds both variables. This motivates the utility of FE estimators that isolate treatment effects while controlling for certain types of unobserved confounding.sjolander2016, mallinson2021,kim2019,kim2021,gunasekara2014,imai2019 The gain-score is one such estimator, and we demonstrate its mechanics and susceptibility to outcome-to-outcome interference in simple linear regression. We regress the gain-score on siblings’ treatments,

equation[equation omitted — 44 chars of source]

where $e_i$ is the residual. Through outcome differencing, the gain-score estimator offsets unobserved confounding bias such that the treatment-conditional covariance between $U_i$ and $D_i$ ($\sigma_{U_iD_i\cdot T_{i1}T_{i2}}$) is zero, thereby isolating the treatment effect.mallinson2021, kim2019, kim2021 The baseline model includes four pathways that “transmit” bias from $U_i$ into $D_i$:

enumerate$U_{i}\rightarrow T_{i1}\rightarrow Y_{i1}\rightarrow D_i$: $-\chi\delta$$U_{i}\rightarrow T_{i2}\rightarrow Y_{i2}\rightarrow D_{i}$: $\gamma\delta$$U_{i}\rightarrow Y_{i1}\rightarrow D_{i}$: $-\psi$$U_{i}\rightarrow Y_{i2}\rightarrow D_{i}$: $\psi$

Covariate adjustment on $T_{i1}$ and $T_{i2}$ closes paths 1 and 2, and paths 3 and 4 cancel each other out exactly because the association that flows along path 3 ($-\psi$) is equal to the negative value of the association that flows along path 4 ($\psi$). Thus, $\sigma_{U_iD_i\cdot T_{i1}T_{i2}}=0$, and $b_2=\delta$ identifies the treatment effect precisely (likewise, $b_1=-\delta$). See Appendix for details.

The estimator’s validity relies on absent outcome-to-outcome interference. Otherwise, we cannot identify $\delta$. The model in Figure 1B introduces outcome-to-outcome interference $\eta$ ($Y_{i1}\rightarrow Y_{i2}$). There are now two additional pathways from $U_i$ to $D_i$:

enumerate\setcounter{enumi}{4} • $U_{i}\rightarrow T_{i1}\rightarrow Y_{i1}\rightarrow Y_{i2}\rightarrow D_i$: $\chi\delta\eta$$U_{i}\rightarrow Y_{i1}\rightarrow Y_{i2}\rightarrow D_{i}$: $\psi\eta$

Covariate adjustment on $T_{i1}$ blocks path 5. However, path 6 is open as a result of outcome-to-outcome interference, so $\sigma_{U_iD_i\cdot T_{i1}T_{i2}}\neq0$, thereby preventing treatment identification (i.e., $b_2\neq\delta$ and $b_1\neq-\delta$). See Appendix for details.

Robustness test

Given that outcome-to-outcome interference in sibling FE models biases treatment identification, we may consider statistical tests for gauging this source of bias. An ideal test is a gain-score regression that adjusts for $U_i$:

equation[equation omitted — 55 chars of source]

Under the models in Figures 1A and 1B, $b_U=0$ indicates no outcome-to-outcome spillover ($\eta=0$) because $\sigma_{U_iD_i\cdot T_{i1}T_{i2}}=0$, whereas $b_U\neq0$ indicates outcome-to-outcome spillover ($\eta\neq0$) because $\sigma_{U_iD_i\cdot T_{i1}T_{i2}}\neq0$. This hypothetical test is straightforward but impossible; $U_i$ is unobserved, so we cannot adjust for it. Indeed, adjustment for $U_i$ would render sibling FE analysis obsolete—simply regress $Y_{ij}$ on $T_{ij}$ and $U_i$ to identify $\delta$.

Still, we may measure cluster-level characteristics, raising the possibility of adjusting for an observed FE in the gain-score regression. The model in Figure 1C introduces $C_i$, a binary or continuous observed pre-treatment FE, and $U_i^\prime$, an unobserved FE that is equivalent to $U_i$ without $C_i$. The dashed line between $U_i^\prime$ and $C_i$ indicates a direct and linear association that is equal to $\pi$. An example of a potential observed FE is maternal nativity, which is a sibling-invariant characteristic that is associated with family health history and pregnancy-related outcomes in the United States.qin2010, creanga2012 Additionally, $C_i$ should precede siblings’ treatments to prevent bias from conditioning on a mediator variable between the treatment and outcome or from conditioning on a post-treatment collider.elwert2014, vwbook

Under the model in Figure 1C, we repeat the gain-score regression and adjust for $C_i$:

equation[equation omitted — 55 chars of source]

There are four pathways from $C_i$ to $D_{i}$:

enumerate$C_{i}\hdashrule[0.5ex]{0.6cm}{0.5pt}{2pt} U^\prime_{i}\rightarrow T_{i1}\rightarrow Y_{i1}\rightarrow D_i$: $-\pi\chi\delta$$C_{i}\hdashrule[0.5ex]{0.6cm}{0.5pt}{2pt} U^\prime_{i}\rightarrow T_{i2}\rightarrow Y_{i2}\rightarrow D_{i}$: $\pi\gamma\delta$$C_{i}\hdashrule[0.5ex]{0.6cm}{0.5pt}{2pt} U^\prime_{i}\rightarrow Y_{i1}\rightarrow D_{i}$: $-\pi\psi$$C_{i}\hdashrule[0.5ex]{0.6cm}{0.5pt}{2pt} U^\prime_{i}\rightarrow Y_{i2}\rightarrow D_{i}$: $\pi\psi$

Covariate adjustment on siblings’ treatments blocks paths 1 and 2, and outcome differencing offsets the associations from paths 3 and 4. Therefore, $\sigma_{C_iD_i\cdot T_{i1}T_{i2}}=0$, so $b_C=0$.

We then consider the model in Figure 1D, which includes $C_i$ and outcome-to-outcome interference. Two additional pathways link $C_i$ to $D_i$:

enumerate\setcounter{enumi}{4} • $C_{i}\hdashrule[0.5ex]{0.6cm}{0.5pt}{2pt} U^\prime_{i}\rightarrow T_{i1}\rightarrow Y_{i1}\rightarrow Y_{i2}\rightarrow D_i$: $\chi\delta\eta$$C_{i}\hdashrule[0.5ex]{0.6cm}{0.5pt}{2pt} U^\prime_{i}\rightarrow Y_{i1}\rightarrow Y_{i2}\rightarrow D_{i}$: $\pi\psi\eta$

While covariate adjustment closes path 5, path 6 remains open, so $\sigma_{C_iD_i\cdot T_{i1}T_{i2}}\neq0$ and $b_C\neq0$. Theoretically, this robustness test signals outcome-to-outcome interference in a basic sibling FE model. We now turn to simulation analyses to exercise this test in multiple settings.

Simulations

Monte Carlo simulations demonstrated the robustness test’s performance in the model of Figure 1D and across its subsequent variations in Figure 2.adkins2012 Model alterations included treatment-to-treatment interference, direct associations between $C_i$ and siblings’ treatments or outcomes, treatment-to-outcome interference, or outcome-to-treatment interference.

\afterpage{

figure[figure omitted — 9,530 chars of source]

}

Our simulation model follows:

\[U^\prime_{i},\upsilon_{c},\upsilon_{i1},\upsilon_{i2}\sim N(0,1)\] \[\boldsymbol\alpha=\{\lambda,\omega\}\] \[C_{i}=

cases0 \ if\ \pi U^\prime_{i} + \upsilon_{c}\leq 1 \\ 1 \ if\ \pi U^\prime_{i} + \upsilon_{c} > 1

\] \[T_{i1}=

cases0 \ if\ \tau C_{i}+\chi U^\prime_{i}\leq -0.2 \\ 1 \ if\ \tau C_{i}+\chi U^\prime_{i} > -0.2

\] \[T_{i2}=

cases0 \ if\ \phi T_{i1}+\mu Y_{i1}+\nu C_{i}+\gamma U^\prime_{i}\leq 1 \\ 1 \ if\ \phi T_{i1}+\mu Y_{i1}+\nu C_{i}+\gamma U^\prime_{i} > 1

\] \[Y_{i1}=\delta T_{i1}+\kappa T_{i2}+\lambda C_{i}+\psi U^\prime_{i}+\upsilon_{i1}\] \[Y_{i2}=\delta T_{i2}+\theta T_{i1}+\eta Y_{i1}+\boldsymbol\alpha C_{i}+\psi U^\prime_{i}+\upsilon_{i2}\] \[D_i=Y_{i2}-Y_{i1}\]

For each sibling FE model, we ran two simulation sets, each consisting of two simulations with 1,000 runs of 5,000 observations (sibling pairs). Simulations varied on outcome-to-outcome interference ($\eta$) and on the association between $C_i$ and $U_i^\prime$ ($\pi$):

enumerate• Simulation 1.1: $\eta=0$ and $\pi=0.5$; • Simulation 1.2: $\eta=0.3$ and $\pi=0.5$; • Simulation 2.1: $\eta=0$ and $\pi=0$; • Simulation 2.2: $\eta=0.3$ and $\pi=0$.

Baseline parameters were fixed across simulations: $\delta=3$, $\chi=1$, $\gamma=2$, and $\psi=5$. Unless specified, we set additional parameters to zero. Within samples, we ran the gain-score regression in equation (3). We conducted simulations in Stata Statistical Software: Release 16.statasoft See Appendix for simulation code.

Table 1 displays simulation inputs and results for the observed FE, and Appendix Table 1 contains full results. We first consider simulations 1.1 and 1.2. Under the model of Figure 1D, $\widehat{b_C}=0$ (95% confidence interval: -0.105, 0.105) without outcome-to-outcome interference, and the coverage probability—the percent of samples with confidence intervals that overlapped zero—exceeded 95%. Conversely, $\widehat{b_C}=0.264$ (95% confidence interval: 0.159, 0.370) with outcome-to-outcome interference, and the coverage probability was \textless1%. These results persisted in the presence of treatment-to-treatment interference (Figure 2A), direct associations between $C_i$ and siblings’ treatments (Figure 2B), equivalent and direct associations between $C_i$ and siblings’ outcomes (Figure 2C), and treatment-to-outcome interference (\textbf{Figure 2D}). However, $\widehat{b_C}\neq0$ with direct non-equivalent associations between $C_i$ and siblings’ outcomes (\textbf{Figure 2E}) or with outcome-to-treatment interference (\textbf{Figure 2F}), regardless if outcome-to-outcome interference was present.

\afterpage{

landscape\begin{onehalfspacing} \begin{footnotesize} \setlength\LTleft{0pt} \setlength\LTright{0pt} \begin{longtable}{l l S[table-format=1.3] c S[table-format=2.1] S[table-format=1.3] c S[table-format=2.1] c} \captionsetup{justification=justified, singlelinecheck=false} \caption{Table 1. Monte Carlo simulation results} \\ \toprule[1pt]\midrule[0.3pt] & & \multicolumn{3}{c}{Simulation Set 1.1} & \multicolumn{3}{c}{Simulation Set 1.2} & \\ Model & Additional Parameters & \multicolumn{3}{c}{$\eta=0$, $\pi=0.5$} & \multicolumn{3}{c}{\textbf{$\eta=0.3$, $\pi=0.5$}} & \textbf{SS1:} \\ \cline{3-5} \cline{6-8} \textbf{(Fig.)} & \textbf{(With Baseline)} & \textbf{\textit{$b_C$}} & \textbf{\textit{95% CI}} & \textbf{\textit{Cover (%)}}$^a$ & \textbf{\textit{$b_C$}} & \textbf{\textit{95% CI}} & \textbf{\textit{Cover (%)}}$^a$ & \textbf{Test Valid}$^{b}$ \\ \hline \rowcolor{lightgray} \textbf{1D} & None & -0.000 & -0.105, 0.105 & 95.1 & 0.264 & 0.159, 0.370 & 0.3 & Yes \\ \textbf{2A} & $\phi=0.9$ ($T_{i1}\rightarrow T_{i2}$) & -0.000 & -0.105, 0.104 & 95.5 & 0.397 & 0.289, 0.505 & 0.0 & Yes \\ \rowcolor{lightgray} \textbf{2B} & $\tau=0.7$ ($C_{i}\hdashrule[0.5ex]{0.6cm}{0.5pt}{2pt} T_{i1}$), $\nu=1.5$ ($C_{i}\hdashrule[0.5ex]{0.6cm}{0.5pt}{2pt} T_{i2}$) & -0.001 & -0.116, 0.114 & 94.5 & -0.378 & -0.496, -0.260 & 0.0 & Yes \\ \textbf{2C} & $\lambda=2.5$ ($C_{i}\hdashrule[0.5ex]{0.6cm}{0.5pt}{2pt} Y_{ij}$) & -0.000 & -0.105, 0.105 & 95.1 & 1.014 & 0.909, 1.120 & 0.0 & Yes \\ \rowcolor{lightgray} \textbf{2D} & $\theta=1.2$ ($T_{i1}\rightarrow Y_{i2}$), $\kappa=0.6$ ($T_{i2}\rightarrow Y_{i1}$) & -0.000 & -0.105, 0.105 & 95.1 & 1.314 & 1.209, 1.420 & 0.0 & Yes \\ \textbf{2E} & $\lambda=2.5$ ($C_{i}\hdashrule[0.5ex]{0.6cm}{0.5pt}{2pt} Y_{i1}$), $\omega=2.8$ ($C_{i}\hdashrule[0.5ex]{0.6cm}{0.5pt}{2pt} Y_{i2}$) & 0.300 & 0.195, 0.405 & 0.0 & 0.264 & 0.159, 0.370 & 0.3 & No \\ \rowcolor{lightgray} \textbf{2F} & $\mu=0.4$ ($Y_{i1}\rightarrow T_{i2}$) & 0.033 & -0.071, 0.136 & 89.3 & 0.462 & 0.351, 0.573 & 0.0 & No \\ \hline & & \multicolumn{3}{c}{\textbf{Simulation Set 2.1}} & \multicolumn{3}{c}{\textbf{Simulation Set 2.2}} & \\ \textbf{Model} & \textbf{Additional Parameters} & \multicolumn{3}{c}{\textbf{$\eta=0$, $\pi=0$}} & \multicolumn{3}{c}{\textbf{$\eta=0.3$, $\pi=0$}} & \textbf{SS2:} \\ \cline{3-5} \cline{6-8} \textbf{(Fig.)} & \textbf{(With Baseline)} & \textbf{\textit{$b_C$}} & \textbf{\textit{95% CI}} & \textbf{\textit{Cover (%)}}$^a$ & \textbf{\textit{$b_C$}} & \textbf{\textit{95% CI}} & \textbf{\textit{Cover (%)}}$^a$ & \textbf{Test Valid}$^{b}$ \\\rowcolor{lightgray} \textbf{1D} & None & -0.002 & -0.109, 0.105 & 94.0 & -0.002 & -0.109, 0.106 & 93.1 & No \\ \textbf{2A} & $\phi=0.9$ ($T_{i1}\rightarrow T_{i2}$) & -0.002 & -0.109, 0.105 & 94.1 & -0.001 & -0.113, 0.111 & 93.1 & No \\ \rowcolor{lightgray} \textbf{2B} & $\tau=0.7$ ($C_{i}\hdashrule[0.5ex]{0.6cm}{0.5pt}{2pt} T_{i1}$), $\nu=1.5$ ($C_{i}\hdashrule[0.5ex]{0.6cm}{0.5pt}{2pt} T_{i2}$) & -0.002 & -0.112, 0.108 & 94.2 & -0.825 & -0.936, -0.714 & 0.0 & Yes \\ \textbf{2C} & $\lambda=2.5$ ($C_{i}\hdashrule[0.5ex]{0.6cm}{0.5pt}{2pt} Y_{ij}$) & -0.002 & -0.109, 0.105 & 94.0 & 0.749 & 0.641, 0.856 & 0.0 & Yes \\ \rowcolor{lightgray} \textbf{2D} & $\theta=1.2$ ($T_{i1}\rightarrow Y_{i2}$), $\kappa=0.6$ ($T_{i2}\rightarrow Y_{i1}$) & -0.002 & -0.109, 0.105 & 94.0 & 1.049 & 0.941, 1.156 & 0.0 & No \\ \textbf{2E} & $\lambda=2.5$ ($C_{i}\hdashrule[0.5ex]{0.6cm}{0.5pt}{2pt} Y_{i1}$), $\omega=2.8$ ($C_{i}\hdashrule[0.5ex]{0.6cm}{0.5pt}{2pt} Y_{i2}$) & 0.298 & 0.191, 0.405 & 0.1 & 1.049 & 0.941, 1.156 & 0.0 & No \\ \rowcolor{lightgray} \textbf{2F} & $\mu=0.4$ ($Y_{i1}\rightarrow T_{i2}$) & -0.002 & -0.109, 0.104 & 94.1 & -0.001 & -0.116, 0.114 & 93.8 & No \\ \midrule[0.3pt]\bottomrule[1pt] \caption{$^a$The coverage probability is the percent of runs in which the 95% confidence interval overlapped zero.} \\ \caption{$^b$The robustness test correctly signals outcome-to-outcome interference in the model under the conditions of simulation set 1 or simulation set 2} \\ \caption{Notes: Each simulation consisted of 1,000 runs of 5,000 observations, where each observation represented a sibling pair. Subscripts $i$ and $j$ indicate cluster and sibling, respectively. $T_{ij}$ is the binary treatment, $Y_{ij}$ is the continuous outcome, $U^\prime_{i}$ is the unobserved fixed effect, $C_i$ is the pre-treatment observed fixed effect, and $D_i=Y_{i2}-Y_{i1}$ is the gain-score. The following baseline parameters (linear effects) were fixed across all simulations: $\delta=3.0$ ($T_{ij}\rightarrow Y_{ij}$), $\chi=1.0$ ($U^\prime_{i}\rightarrow T_{i1}$), $\gamma=2.0$ ($U^\prime_{i}\rightarrow T_{i2}$), and $\psi=5.0$ ($U^\prime_{i}\rightarrow Y_{ij}$). The values of $\eta$ ($Y_{i1} \rightarrow Y_{i2}$) and $\pi$ ($C_{i}\hdashrule[0.5ex]{0.6cm}{0.5pt}{2pt} U^\prime_{i}$) varied by simulation set. Within each sample, we conducted the robustness test by regressing $D_i$ on $T_{i1}$, $T_{i2}$, and $C_i$ to assess outcome-to-outcome interference (i.e., whether $\eta$ was non-zero). $b_C$ is the partial regression coefficient for $C_i$. If the robustness test is valid, $b_C=0$ indicates no outcome-to-outcome interference, and $b_C\neq0$ indicates outcome-to-outcome interference. Abbreviations: "CI" confidence interval, "Fig." figure, "SS1" simulation set 1, "SS2" simulation set 2.} \\ \end{longtable} \end{footnotesize} \end{onehalfspacing}

}

Results from simulations 2.1 and 2.2 indicate that the robustness test was only valid under the models in Figures 2B and 2C if $\pi=0$; if $C_i$ and $U_i^\prime$ were not directly associated, then $C_i$ must have been directly associated with a treatment or directly and equally associated with both outcomes for the robustness test to work. Otherwise, the test was invalid because $C_i$ was not associated with any other variable (Figures 1D, 2A, 2D, and 2F) or because $C_i$ was differentially associated with siblings’ outcomes (Figure 2E).

In a post-hoc analysis, we re-ran all simulation sets under the model in Figure 2A such that $C_i$ was directly associated with only one sibling’s treatment: $\nu=1.5$ ($C_i\hdashrule[0.5ex]{0.6cm}{0.5pt}{2pt}T_{i2}$) and $\tau=0$ ($C_i\hdashrule[0.5ex]{0.6cm}{0.5pt}{2pt}T_{i1}$), or $\nu=0$ and $\tau=0.7$. Results were consistent with main findings, so $C_i$ is a valid observed FE in these settings (see Appendix).

Discussion

In this paper, we demonstrated a simple robustness test for assessing bias from outcome-to-outcome interference in linear two-sibling FE models. The test involves regressing a gain-score on a pre-treatment observed FE and on both siblings’ treatments. For the test to work, the observed FE must be directly associated with an sibling-invariant confounder on treatments and outcomes, directly associated with at least one sibling’s treatment, or directly and equally associated with both siblings’ outcomes. Further, the robustness test collapses if the observed FE is directly and differentially associated with siblings’ outcomes or if an individual’s outcome affects their sibling’s treatment, although the latter violates the treatment estimation in a sibling FE model regardless.frisell2012, sjolander2016, frisell2021, mallinson2021 Failure to meet this criterion does not necessarily warrant a variable’s exclusion from a FE regression model. For example, the regression should adjust for an observed cluster-level variable that differentially affects siblings’ treatments and outcomes because it is a confounder, but it is not a valid observed FE for our robustness test.

Finding an observed FE for the robustness test is data dependent. Birth registries are commonplace in sibling FE analyseskhashan2014, khashan2015, hvolgaard2016, axelsson2019, mallinson2020, mallinson2021 and may provide several candidate variables. For instance, United States birth records record maternal demographic information that are fixed and related to family health history, such as nativity and ethnicity.qin2010, creanga2012, nchs Other sources may not have a similarly wide breadth of sibling-shared characteristics. Additionally, these shared characteristics may have complex links to siblings’ characteristics, and researchers must interrogate how candidate observed FEs relate the treatments and outcomes. For this reason, we considered only pre-treatment variables as candidate observed FEs for the robustness test. Recent methodological guidance discourages conditioning on post-treatment variables to minimize bias in treatment estimates.elwert2014, vwbook In our sibling FE model, conditioning on an observed FE that is causally affected by both siblings’ treatments would induce such bias in the treatment estimate, so it is also possible that it would invalidate the observed FE’s performance in the robustness test. Using an observed FE that is measured prior to siblings’ treatments alleviates this concern. More broadly, we recommend graphical modeling to determine a variable’s usefulness for the robustness testpearl2013, tennant2021 and, when possible, running multiple robustness tests with valid observed FEs to thoroughly screen possible outcome-to-outcome interference.

We also highlight the flexibility of this robustness test. Despite relying on gain-scores, it is applicable to studies with different FE strategies, such as dummy score regressionkim2021 or conditional maximum likelihood estimation.sjolander2016 The data may require reshaping such that each observation consists of a dyadic pair. Furthermore, the robustness test may be used in studies of other types of dyadic pairs and not only in sibling comparison analyses.

There are some limitations to this paper. We focused on settings with linear homogenous effects. Our test may be invalid in models with nonlinear relationships, including those with binary outcomes. Furthermore, we restricted our attention to dyadic sibling models, although this test may work in data with clusters of three or more siblings—for example, one could run the robustness test on randomly selected pairs from sibships. Finally, we omitted models with shared mediator or collider variables, which can induce bias in sibling FE models.sjolander2017 These factors may also invalidate our robustness test.

In sum, our gain-score robustness test is a useful tool for validating two-sibling FE models. Researchers who conduct sibling comparison studies should take advantage of this test to evaluate potential bias from outcome-to-outcome interference.

thebibliography{9} \bibitem{frisell2012}Frisell T, Öberg S, Kuja-Halkola R, Sjölander A. Sibling comparison designs: bias from non-shared confounders and measurement error. Epidemiol. 2012;23:713-720. \bibitem{khashan2014}Khashan AS, Kenny LC, Lundholm C, Kearney PM, Gong T, Almqvist C. Mode of obstetrical delivery and type 1 diabetes: a sibling design study. Pediatrics. 2014;134:e806-813. \bibitem{khashan2015}Khashan AS, Kenny LC, Lundholm C, et al. Gestational age and birth weight and the risk of childhood type 1 diabetes: a population-based cohort and sibling design study. Diabetes Care. 2015;38:2308-2315. \bibitem{hvolgaard2016}Hvolgaard Mikkelsen S, Olsen J, Bech BH, Obel C. Parental age and attention-deficit/hyperactivity disorder (adhd). Int J Epidemiol. 2016;46:409-420. \bibitem{sjolander2016}Sjölander A, Frisell T, Kuja-Halkola R, Öberg S, Zetterqvist J. Carryover effects in sibling comparison designs. Epidemiol. 2016;27:852-858. \bibitem{hanley2017}Hanley GE, Hutcheon JA, Kinniburgh BA, et al. Interpregnancy interval and adverse pregnancy outcomes: an analysis of successive pregnancies. Obstet Gynecol. 2017;129:408-415. \bibitem{axelsson2019}Axelsson PB, Clausen TD, Petersen AH, et al. Relation between infant microbiota and autism?: results from a national cohort sibling design study. \textit{Epidemiol}. 2019;30:52-60. \bibitem{mallinson2020}Mallinson DC, Larson A, Berger LM, Grodsky E, Ehrenthal DB. Estimating the effect of Prenatal Care Coordination in Wisconsin: a sibling fixed effects analysis. \textit{Health Serv Res}. 2020;55:82-93. \bibitem{petersen2020}Petersen AH, Lange T. What is the causal interpretation of sibling comparison designs? \textit{Epidemiol}. 2020;31:75-81. \bibitem{frisell2021}Sibling-comparison designs, are they worth the effort? \textit{Am J Epidemiol}. 2021;190:738-741. \bibitem{mallinson2021}Mallinson DC, Elwert FE. Estimating sibling spillover effects with unobserved confounding using gain-scores. arXiv, doi:2102.11150v3, 6 May 2021, preprint: not peer reviewed. \bibitem{feinberg2012}Feinberg ME, Solmeyer AR, McHale SM. The third rail of family systems: sibling relationships, mental and behavioral health, and preventive intervention in childhood and adolescence. \textit{Clin Child Fam Psychol Rev}. 2012;15:43–57. \bibitem{deneve2017}De Neve JW, Kawachi I. Spillovers between siblings and from offspring to parents are understudied: a review and future directions for research. \textit{Soc Sci Med}. 2017;183:56-61. \bibitem{lipsitch2010}Lipsitch M, Tchetgen Tchetgen E, Cohen T. Negative controls: a tool for detecting confounding and bias in observational studies. \textit{Epidemiol}. 2010;21:383-388. \bibitem{cunninghambook}Cunningham S. Causal Inference: The Mixtape. New Haven: Yale University Press, 2021. \bibitem{kim2019}Kim Y, Steiner PM. Gain scores revisited: a graphical models perspective. \textit{Sociol Methods Res}. 2021;50:1353-1375. \bibitem{kim2021}Kim Y, Steiner PM. Causal graphical views of fixed effects and random effects models. \textit{Br J Math Stat Psychol}. 2021;74:165-183. \bibitem{gunasekara2014}Gunasekara FI, Richardson K, Carter K, Blakely T. Fixed effects analysis of repeated measures data. \textit{Int J Epidemiol}. 2014;43:264-269. \bibitem{imai2019}Imai K, Kim IS. When should we use unit fixed effects regression models for causal inference with longitudinal data? \textit{Am J Pol Sci}. 2019;63:467-490. \bibitem{qin2010}Qin C, Gould JB. Maternal nativity status and birth outcomes in Asian immigrants. \textit{J Immigrant Minority Health}. 2010;12:798-805. \bibitem{creanga2012}Creanga AA, Berg CJ, Syverson C, Seed K, Bruce FC, Callaghan WM. Race, ethnicity, and nativity differentials in pregnancy-related mortality in the United States. \textit{Obstet Gynecol}. 2012;120:361-268. \bibitem{elwert2014}Elwert F, Winship C. Endogenous selection bias: the problem of conditioning on a collider variable. \textit{Annu Rev Soc}. 2014;40:31-53. \bibitem{vwbook}VanderWeele TJ. Explanation in Causal Inference: Methods for Mediation and Interaction. New York: Oxford University Press, 2015. \bibitem{adkins2012}Adkins LC, Gade MN. Monte Carlo experiments using Stata: a primer with examples. \textit{Adv Econ}. 2012;30:429-77. \bibitem{statasoft}StataCorp. Stata Statistical Software: Release 16. College Station, TX: StataCorp LLC; 2019. \bibitem{nchs}National Center for Health Statistics. Revisions of the U.S. Standard Certificates and Reports. https://www.cdc.gov/nchs/nvss/revisions-of-the-us-standard-certificates-and-reports.htm. Accessed September 3, 2021. \bibitem{pearl2013}Pearl J. Linear models: a useful "microscope" for causal analysis. \textit{J Causal Inference}. 2013;1:155-170. \bibitem{tennant2021}Tennant PWG, Murray EJ, Arnold KF, et al. Use of directed acyclic graphs (DAGs) to identify confounders in applied health research: review and recommendations. \textit{Int J Epidemiol}. 2021;50:620-632. \bibitem{sjolander2017}Sjölander A, Zetterqvist J. Confounders, mediators, or colliders. \textit{Epidemiol}. 2017;28:540-547.