EconBase
← Back to paper

Heterogeneous Treatment Effects for Networks, Panels, and other Outcome Matrices

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

67,005 characters · 42 sections · 60 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

1 Heterogeneous Treatment Effects for Networks, Panels, and other Outcome Matrices

abstract\setstretch{1} We are interested in the distribution of treatment effects for an experiment where units are randomized to a treatment but outcomes are measured for pairs of units. For example, we might measure risk sharing links between households enrolled in a microfinance program, employment relationships between workers and firms exposed to a trade shock, or bids from bidders to items assigned to an auction format. Such a double randomized experimental design may be appropriate when there are social interactions, market externalities, or other spillovers across units assigned to the same treatment. Or it may describe a natural or quasi experiment given to the researcher. In this paper, we propose a new empirical strategy that compares the eigenvalues of the outcome matrices associated with each treatment. Our proposal is based on a new matrix analog of the Fr\'echet-Hoeffding bounds that play a key role in the standard theory. We first use this result to bound the distribution of treatment effects. We then propose a new matrix analog of quantile treatment effects that is given by a difference in the eigenvalues. We call this analog spectral treatment effects. \looseness=-1

Introduction

Consider a market designer tasked with learning how an intervention alters the transactions between buyers and sellers in a marketplace. For example, the designer may be an online platform such as Amazon or AirBnB and the intervention a change in a search algorithm or website layout, see recently bajari2021multiple,johari2022experimental. The designer conducts an experiment where they randomly assign buyers and sellers to two groups. They implement the intervention in one group, maintain the status quo in the other group, and measure outcome matrices describing how much each buyer buys from each seller in each group.\looseness=-1

How can the market designer characterize the impact of the intervention on the buyer and seller transactions using such a double randomized experiment? To address this question, we propose a new empirical strategy for identifying heterogeneous treatment effects that compares the eigenvalues of the outcome matrices associated with each treatment. \looseness=-1

To motivate our proposal, Section 2 reviews two approaches standard in the conventional setting of a single randomized experiment. The first approach bounds the distribution of treatment effects using arguments of frechet1951tableaux,hoeffding1940masstabinvariante,makarov1982estimates manski1997mixing,manski2003partial,heckman1997making,fan2010sharp,abadie2018econometric,firpo2019partial,masten2020inference,molinari2020microeconometrics,frandsen2021partial. The second approach makes a rank invariance assumption and computes the difference in the quantiles of the outcomes associated with each treatment. This is often called quantile treatment effects abadie2002instrumental,chernozhukov2005iv,bitler2006mean,firpo2007efficient,imbens2009identification,masten2018identification. While comparing the average outcome across treatments only characterizes an expected treatment effect, the above two approaches can identify treatment effect heterogeneity because they reveal information about the entire distribution of treatment effects. \looseness=-1

Section 3 contains our main results: new analogs of the Frech\'et-Hoeffding bounds and quantile treatment effects for the double randomized experiment with outcome matrices. A key complication is that two dimensions of randomization make the relevant optimization problem quadratic rather than linear. Exact solutions are not generally computable. \looseness=-1

Our main idea is to instead consider relaxations of the quadratic problem solved by rearranging the eigenvalues of the outcome matrices associated with each treatment. Section 4 sketches the logic behind this solution. We first use it to bound the distribution of treatment effects building on arguments of whitt1976bivariate,finke1987quadratic,lovasz2012large. We then show that under a matrix generalization of rank invariance, the distribution of treatment effects is point identified and characterized by a difference in eigenvalues. We call this matrix analog of quantile treatment effects, spectral treatment effects.\looseness=-1

Section 5 discusses extensions including covariates, spillovers, and estimation. Section 6 shows results from two empirical demonstrations and Section 7 concludes. Proof of our main claims are collected in Appendix A, supplementary material can be found in an online appendix, and an R package can be found at \url{https://github.com/yong-cai/MatrixHTE}. \looseness=-1

Motivating examples

We describe four examples of double randomized experiments or quasi experiments with outcome matrices. They are used to motivate our framework and results below. \looseness=-1

Example 1: risk sharing

banerjee2021changes study the impact of a microfinance program in a sample of Indian villages. They argue that the program decreases informal risk sharing between some households. comola2021treatment study the impact of savings accounts in a sample of Nepalese villages. They argue that the program increases informal risk sharing between some households. In this example, the units are households, the treatment is program participation, and the outcomes are surveyed risk sharing links between pairs of households. We revisit this example in the first empirical demonstration of Section 6 below. \looseness=-1

Example 2: superstar extinction

azoulay2010superstar study the impact of a superstar researcher's death in a sample of research groups in the life sciences. They argue that the death of a superstar decreases the quality of research conducted by researchers nearby in the coauthorship network. In this example, the units are researchers, the treatment is the death of a superstar, and the outcomes are the amount of research conducted between coauthors. \looseness=-1

Example 3: auction format

athey2011comparing study the impact of a sealed versus open bid design in a sample of US timber auctions. They argue that the sealed bid design incentivizes some firms to participate who otherwise would not in the open bid design. In this example, the units are firms and tracts of land, the treatment is the auction format, and the outcomes are the bids made by firms on the tracts. We revisit this example in the second empirical demonstration of Section 6 below. \looseness=-1

Example 4: buyer-seller experiment

bajari2021multiple model the impact of an information policy on the likelihood that a buyer buys an item from a seller. They consider a multiple randomization experimental design where the researcher independently randomizes buyers and sellers to groups and then assigns policies to pairs of buyers and sellers depending on their group memberships, see for example their Definitions 7 and 8. In this example, the units are buyers and sellers, the treatment is the information policy, and the outcomes are transactions between buyers and sellers. We revisit this example in our discussion of treatment spillovers in Section 5.4 below. \looseness=-1

Review of the single randomized experiment

We review a standard framework and results for the conventional single randomized experiment following whitt1976bivariate. This review is used to motivate our framework and results for the double randomized experiment with outcome matrices in Section 3. \looseness=-1

Model and econometric problem

Model

A population of agents is randomized to a binary treatment $t \in \{0,1\}$. The population may be finite or infinite. Potential outcomes are defined for each agent in the population and may be fixed or random. The realized potential outcomes of an agent selected uniformly at random from the population are described by a joint distribution function $F$ on $\mathbb{R}^{2}$.

We define the measurable function $(Y^{*}_{1},Y^{*}_{0}):[0,1] \to \mathbb{R}^{2}$ so that $(Y^{*}_{1}(U),Y^{*}_{0}(U))$ has distribution $F$ when $U$ is standard uniform, see Lemma 2.7 of whitt1976bivariate. We sometimes interpret $(Y^{*}_{1},Y^{*}_{0})$ as the fixed potential outcomes of a continuum of agent types indexed by $[0,1]$, although this function representation is valid for both finite and infinite populations. \looseness=-1

As an example, consider an experiment where the researcher randomizes $N$ workers to participate $(t = 1)$ or not participate $(t=0)$ in a training program. Let $\{Y^{*}_{i,1},Y^{*}_{i,0}\}_{i \in [N]}$ describe the fixed potential wages of the $N$ workers and define $Y^{*}_{t}(u) = \sum_{i=1}^{N}Y_{i,t}^{*}\mathbbm{1}\{u \in \tau_{i}\}$ where $\tau_{i} = \{u \in [0,1]: \lceil Nu\rceil = i\}$. Then $Y^{*}_{t}(u)$ describes the potential wage of the $\lceil Nu \rceil$th worker under treatment $t$ and the distribution of $(Y_{1}^{*}(U),Y_{0}^{*}(U))$ is the empirical distribution of the potential wages of the $N$ workers, i.e. $F(y_{1},y_{0}) = \frac{1}{N}\sum_{i=1}^{N}\mathbbm{1}\{Y^{*}_{i,1} \leq y_{1},Y^{*}_{i,0} \leq y_{0}\}$. \looseness=-1

Parameters of interest

We focus on the joint distribution of potential outcomes (DPO) and distribution of treatment effects (DTE). The DPO is

align[align omitted — 181 chars of source]

where $y_{1},y_{0} \in \mathbb{R}$ are arbitrary and $U$ is a standard uniform random variable. In words, the DPO is the mass of agent types with potential outcome less than $y_{1}$ under treatment $1$ and less than $y_{0}$ under treatment $0$. \looseness=-1

The DTE is

align[align omitted — 151 chars of source]

In words, $Y^{*}_{1}(u) - Y^{*}_{0}(u)$ is the change in outcome associated with switching the treatment status of an agent with type $u$ from $0$ to $1$. The DTE is the mass of agents for which this individual treatment effect is less than $y$. \looseness=-1

Econometric problem

Our task is to identify the DPO and DTE. The problem is that the researcher does not observe both $Y_{1}^{*}$ and $Y_{0}^{*}$ for the same population of agents. Agents are assigned to treatment $1$ or treatment $0$ but not both. \looseness=-1

For example, the researcher may assign workers to participate or not participate in a training program. If a worker participates in the program, then the researcher observes the potential wage associated with program participation. They do not observe the wage that the participating worker would have received had they not participated in the program. To infer this missing potential outcome, the researcher must use the wages of the workers that did not participate in the training program. \looseness=-1

Formally, the standard assumption is that the researcher observes the marginal distributions of the potential outcomes given by $F_{1}(\cdot) := F(\cdot,\infty)$ and $F_{0}(\cdot) := F(\infty,\cdot)$ on $\mathbb{R}$, but not any other feature of their joint distribution $F(\cdot.\cdot)$ on $\mathbb{R}^{2}$. The econometric problem is then to identify the DPO and DTE using only $F_{1}$ and $F_{0}$. \looseness=-1

To motivate our results for the double randomized experiment in Section 3, we restate this formulation of the econometric problem using measure preserving transformations. Specifically, we assume the the researcher observes not $(Y_{1}^{*},Y_{0}^{*})$ but $(Y_{1},Y_{0}) : [0,1] \to \mathbb{R}^{2}$ where $Y_{t}$ is equivalent to $Y_{t}^{*}$ up to an unknown measure preserving transformation $\varphi_{t}$. That is,

align[align omitted — 70 chars of source]

for some unknown $\varphi_{t} \in \mathcal{M} := \{\phi: [0,1] \to [0,1]\text{ with } |\phi^{-1}(A)| = |A| \text{ for any measurable } A \subseteq [0,1]\}$ where $|A|$ refers to the Lebesgue measure of $A$. Intuitively, $Y_{t}$ is a rearranged version of $Y_{t}^{*}$ so that the two have the same marginal distribution, but no other feature of the joint distribution of $Y^{*}_{0}$ and $Y^{*}_{1}$ can be learned from $Y_{0}$ and $Y_{1}$. \looseness=-1

The restated econometric problem is then to identify the DPO and DTE using only $Y_{1}$ and $Y_{0}$. The two versions of the econometric problem are equivalent because $Y_{t}$ contains exactly the same information as the marginal distribution or quantile function associated with $Y_{t}^{*}$, see Theorem 5.1 of whitt1976bivariate. However, we use measure preserving transformations and not marginal or quantile functions in our formulation of the econometric problem because there is no natural analog of the latter for outcome matrices under double randomization. \looseness=-1

Some standard results for the single randomized experiment

We first bound the DPO and DTE following frechet1951tableaux,hoeffding1940masstabinvariante,makarov1982estimates. We then show that under a rank invariance assumption the DPO and DTE are point identified and can be written as functionals of the quantiles of the outcomes associated with each treatment following doksum1974empirical,lehmann1975nonparametrics,whitt1976bivariate. These results are known to the econometrics literature. We state them to motivate our results for the double randomized experiment with outcome matrices in Section 3. \looseness=-1

Bounds on the DPO and DTE

Plugging ((ref)) into ((ref)) gives sharp bounds on the DPO

align[align omitted — 309 chars of source]

These bounds have a simple analytical solution.

flushleftStandard result 1: For any $(y_{1},y_{0}) \in \mathbb{R}^{2}$ \begin{align*} \max\left(F_{1}(y_{1})+F_{0}(y_{0}) - 1,0\right) \leq F(y_{1},y_{0}) \leq \min\left(F_{1}(y_{1}),F_{0}(y_{0})\right). \end{align*}

\looseness=-1

Standard result 1 is often attributed to frechet1951tableaux,hoeffding1940masstabinvariante, although our proof sketch in Section 4.2 follows whitt1976bivariate. The bounds are straightforward to compute or estimate (in cases of sampled, mismeasured, or missing outcomes) using standard tools. \looseness=-1

The bounds on the DPO imply bounds on the DTE.

flushleftStandard result 2: For any $y \in \mathbb{R}$ \begin{align*} \sup_{\substack{(y_{1},y_{0}) \in \mathbb{R}^{2}:\\ y_{1}-y_{0} = y}}\max\left(F_{1}(y_{1})-F_{0}(y_{0}) ,0\right) \leq \Delta(y) \leq 1 + \inf_{\substack{(y_{1},y_{0}) \in \mathbb{R}^{2}:\\ y_{1}-y_{0} = y}}\min\left(F_{1}(y_{1})-F_{0}(y_{0}),0\right). \end{align*}

\looseness=-1

Standard result 2 is often attributed to makarov1982estimates. These bounds are also straightforward to compute or estimate using standard tools. \looseness=-1

Point identification of the DTE under rank invariance

The Quantile Treatment Effects parameter (QTE) refers to the difference in the quantile functions of $Y_{1}$ and $Y_{0}$. Specifically, $QTE(u) := Q_{1}(u)-Q_{0}(u)$ where $Q_{t}(u) := \inf\{y \in \mathbb{R}: u \leq F_{t}(y)\}$ is the inverse marginal distribution (quantile) function associated with $Y_{t}^{*}$, or equivalently, $Y_{t}$. Although $Q_{t}$ and $Y_{t}^{*}$ generally have the same marginal distribution for $t\in \{0,1\}$, $Q_{1}-Q_{0}$ and $Y_{1}^{*}-Y_{0}^{*}$ do not. In fact, the difference in quantiles is a more conservative notion of the effect of treatment as measured by mean squared error. That is, \looseness=-1

flushleftStandard result 3: $\int \left(Q_{1}(u) - Q_{0}(u)\right)^{2}du \leq \int \left(Y_{1}^{*}(u)-Y_{0}^{*}(u)\right)^{2}du$.

See Corollary 2.9 of whitt1976bivariate. However, under a rank invariance assumption, $Q_{1}-Q_{0}$ and $Y_{1}^{*}-Y_{0}^{*}$ do have the same distribution. We say that a treatment effect is rank invariant if $Y_{1}^{*} = g(Y_{0}^{*})$ for some nondecreasing $g: \mathbb{R} \to \mathbb{R}$. \looseness=-1

flushleftStandard result 4: $\Delta(y) = \int \mathbbm{1}\{Q_{1}(u)-Q_{0}(u) \leq y\}du$ under rank invariance.

See Theorem 2.5 of whitt1976bivariate. The DPO is similarly identified under rank invariance with $F(y_{1},y_{0}) = \int \prod_{t \in \{0,1\}}\mathbbm{1}\{Q_{t}(u) \leq y_{t}\}du$. In words, rank invariance says that the rank of an agent's outcome in the population is the same under both treatments. That is, $\int \mathbbm{1}\{Y_{0}^{*}(s) \leq Y_{0}^{*}(u)\}ds = \int \mathbbm{1}\{Y_{1}^{*}(s) \leq Y_{1}^{*}(u)\}ds$ for every $u \in [0,1]$.

The double randomized experiment

We propose analogs of the Section 2 framework and results for a double randomized experiment with outcome matrices. Our focus is on symmetric matrices indexed by one population as in Examples 1 and 2 of Section 1.1. Asymmetric matrices or matrices indexed by two different populations as in Examples 3 and 4 are handled by symmetrization in Section 5.1.1. \looseness=-1

Model and econometric problem

Model

A population of agents is randomized to two groups. The population may be finite or infinite. Pairs of agents are assigned a binary treatment $t \in \{0,1\}$ depending on the individual group assignments. bajari2021multiple call this a simple multiple randomization design, see their Definition 8. For other examples of double randomization in the literature see graham2008identifying,graham2011econometric,graham2014complementarity,johari2022experimental.

To simplify our exposition, we suppose that a pair of agents is assigned to treatment $1$ if both agents belong to the first group and assigned to treatment $0$ if both agents belong to the second group, ignoring any outcomes between agents assigned to different groups. However our main arguments below are not specific to this particular comparison, see Section 5.4 below. Potential outcomes are bounded (this can be relaxed) and defined for each pair of agents in the population. They may be fixed or random. \looseness=-1

Recall that the potential outcomes in the single randomized setting are represented by $(Y^{*}_{1},Y^{*}_{0}):[0,1] \to \mathbb{R}^{2}$ where $Y_{t}^{*}(u)$ describes the fixed potential outcome associated with treatment $t \in \{0,1\}$ and agent type $u \in [0,1]$. We consider an analogous representation for the double randomized setting with outcome matrices where the potential outcomes are represented by a symmetric measurable function $(Y^{*}_{1},Y^{*}_{0}):[0,1]^{2} \to \mathbb{R}^{2}$. $Y_{t}^{*}(u,v)$ describes the fixed potential outcome associated with treatment $t \in \{0,1\}$ and agent types $u, v \in [0,1]$. \looseness=-1

Following lovasz2012large, we sometimes interpret $Y_{t}^{*}$ as an infinite dimensional population matrix, although this representation is valid for both finite and infinite populations. Let $U$ and $V$ be independent standard uniform random variables. Then the random vector $(Y_{1}^{*}(U,V),Y_{0}^{*}(U,V))$ describes the distribution of potential outcomes between a pair of agent types each drawn independently and uniformly at random from $[0,1]$. Both the choice of state space $[0,1]$ and the assumption that agents are selected uniformly at random are standard normalizations also made in the conventional single randomized setting. \looseness=-1

As an example, consider Example 1 from Section 1.1 where the researcher randomizes $N$ households to participate or not participate in a microfinance program. Let $\{Y_{ij,1}^{*},Y_{ij,0}^{*}\}_{i,j \in [N]}$ describe the fixed potential risk sharing links between every pair of households when both enroll ($t=1$) or do not enroll ($t=0$) in the program. Define $Y_{t}(u,v) = \sum_{i=1}^{N}\sum_{j=1}^{N}Y_{ij,t}^{*}\mathbbm{1}\{u \in \tau_{i}, v \in \tau_{j}\}$ where $\tau_{i} = \{u \in [0,1]: \lceil Nu\rceil = i\}$. Then $Y_{t}^{*}(u,v)$ describes the potential risk sharing link between the $\lceil Nu\rceil$th and $\lceil Nv \rceil$th households under treatment $t$ and the distribution of $(Y_{1}^{*}(U,V),Y_{0}^{*}(U,V))$ is the empirical distribution of the potential risk sharing links of the $N^{2}$ household pairs, i.e. $F(y_{1},y_{0}) = \frac{1}{N^{2}}\sum_{i=1}^{N}\sum_{j=1}^{N}\mathbbm{1}\{Y_{ij,1}^{*} \leq y_{1}, Y_{ij,0}^{*} \leq y_{0}\}$. \looseness=-1

Interpreting the model

Our model explicitly describes one source of randomness: sampling agents uniformly at random from a fixed population. This implies a specific notion of treatment effect heterogeneity that defines the distribution of treatment effects, following exactly the logic of the conventional single randomized experiment from Section 2. A consequence of the model is that the sampled potential outcomes are independent across pairs of agents that do not have an agent in common. This dependence structure is frequently used in the network econometrics literature, see Section 6 of graham2020network and Section 3 of de2016econometrics for many examples. \looseness=-1

We do not intend for this model to necessarily describe any data generating process. In particular, the researcher may believe that the outcomes they actually observe are also influenced by measurement error, missing data, spillovers, strategic interactions, etc. While such variation is not used to explicitly define treatment effect heterogeneity in our framework, as in the setting of the conventional single randomized experiment, it can be incorporated into the data generating process and still play a role in identification, estimation, and inference. We discuss some extensions along these lines in Sections 5.3 and 5.4 below. \looseness=-1

As an example, consider the two-way effects model $Y_{ij,t} = f_{t}(\alpha_{i,t},\alpha_{j,t},\varepsilon_{ij,t})$ where the agent effects $\{\alpha_{i,t}\}_{i \in [N]}$ have distribution $F_{\alpha,t}$ and the idiosyncratic errors $\{\varepsilon_{ij,t}\}_{i,j \in [N]}$ have distribution $F_{\varepsilon,t}$. One might define the potential outcome function of interest to be $Y^{*}_{t}(u,v) = \int f_{t}(F_{\alpha,t}^{-1}(u),F_{\alpha,t}^{-1}(v),x)dF_{\varepsilon,t}(x)$. Intuitively, this function indexes a collection of expected outcomes (with respect to $\varepsilon_{ij,t}$) generated by sampling agents uniformly at random from the population (with respect to $\alpha_{i,t}$). Inferring $Y^{*}_{t}$ from the data may be complicated if some entries of the matrix are missing, the idiosyncratic errors are dependent across agent pairs or correlated with the individual effects, etc. However, as in the setting of the conventional single randomized experiment, we do not use this variation in our characterization the heterogeneous impact of the treatment as given by the parameters of interest below. \looseness=-1

Parameters of interest

We define the joint distribution of potential outcomes (DPO) and the distribution of treatment effects (DTE) as in Section 2. The DPO is

align[align omitted — 187 chars of source]

where $y_{1},y_{0} \in \mathbb{R}$ and $U$ and $V$ are independent standard uniform random variables. In words, the DPO is the mass of agent type pairs with potential outcome less than $y_{1}$ under treatment $1$ and less than $y_{0}$ under treatment $0$. \looseness=-1

Similarly, the DTE is

align[align omitted — 158 chars of source]

In words, $Y^{*}_{1}(u,v) - Y^{*}_{0}(u,v)$ is the change in outcome associated with switching the treatment status of a pair of agents with types $u$ and $v$ from $0$ to $1$. The DTE is the mass of agent type pairs for which this treatment effect is less than $y$. Under the treatment assignment rule described in Section 3.1.1, it is the distributional analog of what bajari2021multiple call the average effect for the treated pairs. Heterogeneous analogs of other parameters such as their spillover or direct effects can be similarly constructed, see for example Section 5.4. \looseness=-1

Econometric problem

As before, our task is to identify the DPO and the DTE. The problem is also that the researcher observes at most one potential outcome for any pair of agents. \looseness=-1

For example, the researcher may assign households to participate or not participate in a microfinance program. If both households participate in the program, then the researcher observes the potential risk sharing link associated with joint program participation. They do not observe whether these households would have formed a link under the counterfactual treatment that neither household participates. To infer this missing potential outcome, the researcher must use the risk sharing links between the nonparticipating households. \looseness=-1

Following the second econometric problem formulation of Section 2.1.3, we suppose that the researcher observes not $(Y_{1}^{*},Y_{0}^{*})$ but $(Y_{1},Y_{0}) : [0,1]^{2} \to \mathbb{R}^{2}$ where $Y_{t}$ is equivalent to $Y_{t}^{*}$ up to an unknown measure preserving transformation. That is,

align[align omitted — 80 chars of source]

for some unknown $\varphi_{t} \in \mathcal{M}$. Like the conventional single randomized setting, ((ref)) says that $Y_{t}$ and $Y_{t}^{*}$ represent the same random object. However, $Y_{0}$ and $Y_{1}$ do not reveal any additional information about how the entries of $Y_{0}^{*}$ and $Y_{1}^{*}$ are related. Unlike the single randomized setting, there is no canonical $Y_{t}$ that serves the role of the marginal distribution or quantile function in the double randomized setting with outcome matrices. The “marginal distribution of $Y_{t}^{*}$” is instead represented by an equivalence class of functions $Y_{t}$ that satisfy ((ref)). lovasz2012large calls such functions weakly isomorphic, see generally his Sections 7.3, 10.7, and 13.2. \looseness=-1

Some new results for the double randomized experiment

We first bound the DPO and DTE. We then propose a new matrix generalization of rank invariance under which the DPO and DTE are point identified and can be written as functionals of the eigenvalues of the potential outcome functions associated with each treatment. Eigenvalues of functions are defined a bit differently than their matrix counterparts, see our Appendix Section A.1 or lovasz2012large, Section 7.5 for a review. Proof of our main propositions can be found in Appendix Sections A.2-4. \looseness=-1

Bounds on the DPO and DTE

As in the single randomized setting, plugging ((ref)) into ((ref)) gives sharp bounds on the DPO

align[align omitted — 350 chars of source]

We do not consider these bounds, however, because their quadratic structure makes them analytically and computationally intractable in general. See cela2013quadratic, Section 1.5. \looseness=-1

We instead propose bounds that are not generally sharp but are tractable. Let $\lambda_{1t}(y_{1}) \geq \lambda_{2t}(y_{1}) \geq ... \geq \lambda_{Rt}(y_{t})$ be the $R$ largest (in absolute value) eigenvalues of $\mathbbm{1}\{Y_{t}(\cdot,\cdot) \leq y_{t}\}$ ordered to be decreasing and $s_{R}(r) = R - r + 1$. For any $t, t' \in \{0,1\}$, let $\sum_{r}\lambda_{rt}\lambda_{rt'} := \lim_{R \to \infty}\sum_{r=1}^{R}\lambda_{rt}(y_{t})\lambda_{rt'}(y_{t'})$, $\sum_{r}\lambda_{rt}\lambda_{s(r)t'} := \lim_{R \to \infty}\sum_{r=1}^{R}\lambda_{rt}(y_{t})\lambda_{s_{R}(r)t'}(y_{t'})$ and $\sum_{r}\lambda_{rt}^{2} := \sum_{r}\lambda_{rt}\lambda_{rt}$. When the population is finite and equal to $N$ we drop the limit and take $R = N$. \looseness=-1

Our first result is

flushleftProposition 1: For any $(y_{1},y_{0}) \in \mathbb{R}^{2}$ \begin{align} \max\left(\sum_{r}\left(\lambda_{r1}^{2} + \lambda_{r0}^{2}\right) - 1,\sum_{r}\lambda_{r1}\lambda_{s(r)0},0\right) \leq F(y_{1},y_{0}) \nonumber \\ \leq \min\left(\sum_{r}\lambda_{r1}^{2},\sum_{r}\lambda_{r0}^{2},\sum_{r}\lambda_{r1}\lambda_{r0}\right). \end{align}

We defer a discussion of Proposition 1 to Section 4.3, only remarking here that unlike the infeasible bounds in ((ref)), those in ((ref)) are straightforward to compute because they only depend on the eigenvalues of $\mathbbm{1}\{Y_{t}^{*}(\cdot,\cdot) \leq y_{t}\}$, or equivalently, $\mathbbm{1}\{Y_{t}(\cdot,\cdot) \leq y_{t}\}$. They can be computed or estimated (in cases of sampled, mismeasured, or missing outcomes) using standard tools, see Section 5.3. \looseness=-1

As in Section 2, bounds on the DPO imply bounds on the DTE. Our second result is

flushleftProposition 2: For any $y \in \mathbb{R}$ \begin{align} \sup_{\substack{(y_{1},y_{0}) \in \mathbb{R}^{2}:\\ y_{1}-y_{0} = y}}\max\left(\sum_{r}\left(\lambda_{r1}^{2} - \lambda_{r0}^{2}\right),\sum_{r}\left(\lambda_{r1}^{2}-\lambda_{r1}\lambda_{r0}\right),0\right) \leq \Delta(y) \nonumber \\ \leq 1 + \inf_{\substack{(y_{1},y_{0}) \in \mathbb{R}^{2}:\\ y_{1}-y_{0} = y}}\min\left(\sum_{r}\left(\lambda_{r1}^{2} - \lambda_{r0}^{2}\right),\sum_{r}\left(\lambda_{r1}\lambda_{r0}- \lambda_{r0}^{2}\right),0\right) \end{align}

where the eigenvalue $\lambda_{rt}$ is implicitly a function of $y_{t}$. In finite data, these bounds only require the researcher to compute eigenvalues for at most $N(N+1)$ values of $y_{1}$ and $y_{0}$ where $N$ is the number of agents. Optimizing over a smaller set also gives valid but potentially wider bounds. \looseness=-1

Definition of spectral treatment effects

We propose a matrix analog of the QTE. Let $\{\sigma_{rt}\}_{r=1}^{R}$ be the $R$ largest (in absolute value) eigenvalues of $Y_{t}$ ordered to be decreasing and $\{\phi_{r}\}_{r=1}^{\infty}$ be any orthogonal basis of $L^{2}([0,1])$. \looseness=-1

flushleftDefinition 1: The Spectral Treatment Effects parameter (STE) is \begin{align} STE(u,v;\phi) := \lim_{R \to \infty}\sum_{r=1}^{R}(\sigma_{r1}-\sigma_{r0})\phi_{r}(u)\phi_{r}(v). \end{align}

The STE is similar to the diagonalized difference in the eigenvalues of $Y_{1}$ and $Y_{0}$, but its exact values depend on a choice of basis. Two natural choices are the eigenfunctions of $Y_{1}$ and $Y_{0}$, denoted $\{\phi_{r1}\}_{r=1}^{\infty}$ and $\{\phi_{r0}\}_{r=1}^{\infty}$ respectively, see Appendix Section A.1. We call $STE(\phi_{1}) $ and $STE(\phi_{0})$ the Spectral Treatment Effects on the Treated (STT) and Untreated (STU). In words, the STT takes the observed matrix $Y_{1}$ and subtracts a counterfactual formed by keeping the eigenfunctions of $Y_{1}$ and inserting the eigenvalues of $Y_{0}$. That is,

align*[align* omitted — 167 chars of source]

where $W(u,s) = \lim_{R \to \infty}\sum_{r=1}^{R}\phi_{r1}(u)\phi_{r0}(s)$. The second line suggests an alternative interpretation of the STT where the counterfactual outcome for a pair of agents assigned to treatment $1$ is formed by a weighted average of the outcomes of agent pairs assigned to treatment $0$. Without additional assumptions, the weights are potentially extrapolative in that they may be negative and not necessarily integrate to $1$. In some cases the researcher may wish to explicitly restrict the weights so that they satisfy these properties. We describe one way to do this in Online Appendix Section D.4. The weights will however necessarily be nonnegative and integrate to $1$ under the rank invariance condition that we introduce in the next section. \looseness=-1

The STT is analogous to the QTE which imputes a counterfactual for an agent assigned to treatment $1$ by using the outcome of a similarly ranked agent assigned to treatment $0$. In this analogy, the eigenfunctions serve the role of the agent ranks and the eigenvalues serve the role of the quantiles associated with each rank. As in the case of the conventional single randomized experiment, this parameter may also be justified by a rank invariance assumption. \looseness=-1

Point identification of the DTE under rank invariance

Like the QTE, our STE is also a more conservative notion of the effect of treatment than $Y_{1}^{*}-Y_{0}^{*}$ as measured by mean squared error. Our third result is

flushleftProposition 3: For any orthogonal basis $\{\phi_{r}\}_{r=1}^{\infty}$ of $L^{2}([0,1])$ \begin{align} \int\int STE(u,v;\phi)^{2}dudv \leq \int\int\left(Y_{1}^{*}(u,v)-Y_{0}^{*}(u,v)\right)^{2}dudv. \end{align}

\looseness=-1

In addition, under a rank invariance assumption, the STT, STU, and $Y_{1}^{*}-Y_{0}^{*}$ all have the same distribution. To extend rank invariance to matrices, we use the notion of a matrix function from horn1991topics, Chapter 6.1. For any $f: \mathbb{R} \to \mathbb{R}$ that admits the representation $f(x) = \sum_{r=1}^{\infty}c_{r}x^{r}$ and square matrix $A$ (or function $A: [0,1]^{2}\to\mathbb{R}$), the matrix lift of $f$ is $f(A) = \sum_{r=1}^{\infty}c_{r}A^{r}$ where $A^{r}$ is the $r$th matrix (or operator) power of $A$, i.e. $A^{r}(u,v) = \int\int...\int A(u,t_{1})A(t_{1},t_{2})...A(t_{r-1},v)dt_{1}dt_{2}...dt_{r-1}$. \looseness=-1

flushleftDefinition 2: A treatment effect is rank invariant if $Y^{*}_{1} = g(Y^{*}_{0})$ where $g$ is the matrix lift of some nondecreasing $g : \mathbb{R} \to \mathbb{R}$.

\looseness=-1

We call Definition 2 a matrix generalization of rank invariance because it is equivalent to the definition from Section 2 when $Y^{*}_{1}$ and $Y^{*}_{0}$ are scalars. Our fourth result is \looseness=-1

flushleftProposition 4: Under rank invariance, \begin{align} \Delta(y) = \int\int\mathbbm{1}\{STT(u,v) \leq y\}dudv = \int\int\mathbbm{1}\{STU(u,v) \leq y\}dudv. \end{align}

Intuitively, if we think of the treatment working by taking in $Y_{0}^{*}$ and producing $Y_{1}^{*} = g(Y_{0}^{*})$, then rank invariance implies that the treatment affects the eigenvalues but not the eigenfunctions of $Y_{0}^{*}$. This is analogous to rank invariance in the conventional single randomization setting, where the treatment affects the quantiles but not the ranks of the outcomes. As in that setting, rank invariance is a strong assumption. But there are also many settings where it can be justified by economic theory. We provide four concrete examples from the literature on information diffusion, factor models, social interaction, and network formation in Online Appendix Section C.1. \looseness=-1

Sketch and discussion of the proof of Proposition 1

We demonstrate some of the main technical ideas behind our results by sketching a proof of Proposition 1. To simplify arguments we consider a finite approximation as in whitt1976bivariate,heckman1997making. A full proof can be found in Appendix Section A.2. \looseness=-1

Finite approximation

For the single randomized experiment, we assume that $Y_{t}^{*}$ is an $N\times 1$ vector, the DPO is $\frac{1}{N}\sum_{i=1}^{N}\prod_{t \in \{0,1\}}\mathbbm{1}\{Y^{*}_{i,t} \leq y_{t}\}$, and $Y_{i,t} = \sum_{j = 1}^{N}Y_{j,t}^{*}P_{ij,t}$ is observed where $P_{t}$ is an unknown $N\times N$ permutation matrix. Intuitively, there are $N$ types of agents. One agent of each type is assigned to treatment $1$ and one agent of each type is assigned to treatment $0$. Our task is to compare the outcomes of agents with the same type and different treatment assignments, but we do not know which agent is of which type. Bounds on the DPO are given by maximizing and minimizing $\frac{1}{N}\sum_{i=1}^{N}\sum_{j=1}^{N}\prod_{t \in \{0,1\}}\mathbbm{1}\{Y_{j,t} \leq y_{t}\}P_{ij,t}$ over all $N \times N$ permutation matrices $P_{0}$ and $P_{1}$. That this discrete problem is a good approximation to the continuous ((ref)) is demonstrated in Section 2 of whitt1976bivariate. See also Section 3 of heckman1997making. \looseness=-1

For the double randomized experiment, we similarly assume that $Y_{t}^{*}$ is an $N\times N$ matrix, the DPO is $\frac{1}{N^{2}}\sum_{i=1}^{N}\sum_{j=1}^{N}\prod_{t \in \{0,1\}}\mathbbm{1}\{Y^{*}_{ij,t} \leq y_{t}\}$, and $Y_{ij,t} = \sum_{k=1}^{N}\sum_{l = 1}^{N}Y_{kl,t}^{*}P_{ik,t}P_{jl,t}$ is observed where $P_{t}$ is an unknown permutation matrix. The intuition is the same as in the single randomized experiment. There are $N$ types of agents, one agent of each type is assigned to each treatment, and though we want to compare the outcomes of agents with the same type but different treatment assignments, we do not know which agent is of which type. Tight bounds on the DPO are given by maximizing and minimizing $\frac{1}{N^2}\sum_{i=1}^{N}\sum_{j=1}^{N}\sum_{k=1}^{N}\sum_{l=1}^{N}\prod_{t \in \{0,1\}}\mathbbm{1}\{Y_{kl,t} \leq y_{t}\}P_{ik,t}P_{jl,t}$ over $P_{0}$ and $P_{1}$, which we show is a good approximation to the continuous problem ((ref)) in Appendix Section A.2. Since this discrete problem is an intractable “Quadratic Assignment Problem” or QAP, see generally cela2013quadratic, our bounds are instead based on a conservative but tractable relaxation. \looseness=-1

Standard Result 1 from Section 2

The DPO for the finite approximation to the single randomized experiment is $F_{N}(y_{1},y_{0}) = \frac{1}{N}\sum_{i=1}^{N}\sum_{j=1}^{N}\prod_{t \in \{0,1\}}\mathbbm{1}\{Y_{j,t} \leq y_{t}\}P_{ij,t}$. We show that

align*[align* omitted — 138 chars of source]

where $F_{Nt}(y_{t}) = \frac{1}{N}\sum_{i=1}^{N}\mathbbm{1}\{Y_{i,t} \leq y_{t}\}$ following whitt1976bivariate, Theorem 2.1. The proof relies on the following rearrangement inequality often attributed to hardy1952inequalities.

flushleftTheorem 368 (Hardy-Littlewood-P\'olya): For any $m\in \mathbb{N}$ and $g,h \in \mathbb{R}^{m}$ we have $\sum_{r} g_{(r)}h_{(m-r+1)} \leq \sum_{r} g_{r}h_{r} \leq \sum_{r} g_{(r)}h_{(r)}$ where $g_{(r)}$ is the $r$th order statistic of $g$.

\looseness=-1

Sketch of proof of Standard Result 1

Theorem 368 implies that

align*[align* omitted — 217 chars of source]

where $Y_{(i),t}$ is the $i$th order statistic of $Y_{t}$. The upper bound follows

align*[align* omitted — 172 chars of source]

The lower bound follows

align*[align* omitted — 505 chars of source]

\looseness=-1

Proposition 1 from Section 3

The DPO for the finite approximation to the double randomized experiment is

align*[align* omitted — 176 chars of source]

We show that

align*[align* omitted — 301 chars of source]

where $\lambda_{rt}$ is the $r$th largest eigenvalue of the matrix $\mathbbm{1}\{Y_{t} \leq y_{t}\}$ and $s_{N}(r) = N-r+1$. Our sketch has two parts. The second part is based on work by finke1987quadratic and relies on the following result due to birkhoff1946three. We say that a square matrix is doubly stochastic if its entries are nonnegative and if every row and column sum to $1$. \looseness=-1

flushleftTheorem (Birkhoff): If $M$ is doubly stochastic then there exist an $m \in \mathbb{N}$, $\alpha_{1},...,\alpha_{m} > 0$, and permutation matrices $P_{1},...,P_{m}$ such that $\sum_{t=1}^{m}\alpha_{t} = 1$ and $M_{ij} = \sum_{t=1}^{m}\alpha_{t}P_{ij,t}$.

\looseness=-1

Sketch of proof of Proposition 1, part 1

We first show $\max\left(\sum_{r=1}^{N}\left(\lambda_{r1}^{2}+\lambda_{r0}^{2}\right) - N^2,0\right) \leq N^2 F_{N}(y_{1},y_{0}) \leq \min\left(\sum_{r=1}^{N}\lambda_{r1}^{2},\sum_{r=1}^{N}\lambda_{r0}^{2}\right)$. Write $N^{2}F_{N}(y_{1},y_{0}) = \sum_{r=1}^{N^{2}}\sum_{s=1}^{N{^2}}\prod_{t \in \{0,1\}}\mathbbm{1}\{\tilde{Y}_{r,t} \leq y_{t}\}\tilde{P}_{rs,t}$ where $i_{r} = \lfloor \frac{r-1}{N}\rfloor + 1$, $j_{r} = r-N\lfloor \frac{r-1}{N}\rfloor$, $\tilde{Y}_{r,t} = Y_{i_{r}j_{r},t}$, and $\tilde{P}_{rs,t} = P_{i_{r}i_{s},t}P_{j_{r}j_{s},t}$. In words, $\tilde{Y}_{t}$ and $\tilde{P}_{t}$ are vectorized versions of $Y_{t}$ and $P_{t}\otimes P_{t}$ formed by iteratively appending their rows. Theorem 368 implies that

align*[align* omitted — 267 chars of source]

and so following the arguments of Section 4.2.1, we have

align*[align* omitted — 277 chars of source]

The bounds follow since

align*[align* omitted — 176 chars of source]

\looseness=-1

Sketch of proof of Proposition 1, part 2

We now show $\sum_{r=1}^{N} \lambda_{r1}\lambda_{s_{N}(r)0} \leq N^2 F_{N}(y_{1},y_{2}) \leq \sum_{r=1}^{N}\lambda_{r1}\lambda_{r0}$. These bounds follow

align*[align* omitted — 321 chars of source]

where $(\lambda_{rt},\phi_{rt})$ is the $r$th eigenvalue and eigenvector pair of $\sum_{k,l=1}^{N}\mathbbm{1}\{Y_{kl,t} \leq y_{t}\}P_{ik,t}P_{jl,t}$ and $W^{\phi}$, a matrix with $rs$th entry $W_{rs}^{\phi} = \left[\sum_{i=1}^{N}\phi_{ir,1}\phi_{is,0}\right]^{2}$, is the Hadamard square of an orthogonal matrix and so is doubly stochastic. Birkhoff's Theorem implies

align*[align* omitted — 269 chars of source]

where $\alpha^{\phi}_{1},...,\alpha^{\phi}_{k} > 0$, $\sum_{k=1}^{K}\alpha^{\phi}_{k} = 1$, $P^{\phi}_{1},...,P^{\phi}_{K}$ are permutation matrices, and $W_{rs}^{\phi} = \sum_{k=1}^{K}\alpha^{\phi}_{k}P^{\phi}_{rs,k}$. Theorem 368 implies that for any permutation matrix $P$

align*[align* omitted — 162 chars of source]

The claim follows. \looseness=-1

Discussion

Our bounds on the DPO follow by intersecting those from parts 1 and 2. Each part describes a different relaxation of the intractable QAP. Take for instance the upper bound

align[align omitted — 201 chars of source]

where $\mathcal{P}_{N}$ is the set of all $N \times N$ permutation matrices. Part 1 bounds it from above with

align[align omitted — 179 chars of source]

where $\mathcal{P}_{N^{2}}$ describes permutations of pairs of agents and $r(i,j) = N(i-1)+j$. Intuitively, this relaxation treats the $N \times N$ outcome matrices as vectors of length $N^2$ and uses the fact that $\{P_{ij}\}_{i,j = 1}^{N} \in \mathcal{P}_{N}$ implies that $\{P_{ik}P_{jl}\}_{i,j,k,l = 1}^{N} \in \mathcal{P}_{N^{2}}$. Whereas ((ref)) is an intractable QAP, ((ref)) is linear and can be bounded using Theorem 368. \looseness=-1

Part 2 bounds the QAP from above with

align[align omitted — 173 chars of source]

where $\mathcal{O}_{N}$ is the set of orthogonal $N\times N$ matrices. This is an upper bound because $\mathcal{P}_{N} \subset \mathcal{O}_{N}$. The insight, see finke1987quadratic, is to use the fact that $\{O_{ij}\}_{i,j=1}^{N} \in \mathcal{O}_{N}$ implies that $\{O_{ij}^{2}\}_{i,j =1}^{N} \in \mathcal{D}^{+}_{N}$ where $\mathcal{D}_{N}^{+}$ is the set of doubly stochastic $N \times N$ matrices. This allows us to rewrite ((ref)) as $\max_{W \in \mathcal{D}_{N}^{+}}\sum_{r,s}\lambda_{r1}\lambda_{rs}W_{rs}$ which is linear in $W$ and can be bounded using Birkhoff's Theorem and Theorem 368. \looseness=-1

Our full proof of Proposition 1 as stated in Section 3 is complicated by the fact that the infinite dimensional analog of $W^{\phi}$ is not generally doubly stochastic and so Birkhoff's Theorem cannot be directly applied. We address this problem by first approximating the function $\mathbbm{1}\{Y_{t} \leq y_{t}\}$ with a finite dimensional matrix, applying the logic of Part 2, and then showing convergence as the dimensions of the matrix are taken to infinity. whitt1976bivariate uses a similar strategy to demonstrate his Theorem 2.1 (our Standard Result 1) in the setting of Section 2. \looseness=-1

These bounds describe two of many possible relaxations of the intractable QAP, see broadly cela2013quadratic, Section 2. We chose these relaxations because they are straightforward to compute, characterize, and appear to work well in practice. Intersecting our bounds with others may lead to smaller identified sets for the DPO and DTE, but potentially at the cost of greater computational complexity or statistical uncertainty. We leave this to future work. \looseness=-1

Extensions

We describe some extensions to the Section 3 framework. Additional details can be found in Online Appendix Sections C and D. \looseness=-1

Asymmetric outcome matrices

Asymmetric matrices or matrices indexed by two different populations can be handled in the following way. A population of workers and firms are randomized (or as good as randomized in the case of a quasi experiment) to two groups. Worker and firm pairs are then assigned a binary treatment $t \in \{0,1\}$ depending on the individual group assignments. For example, the groups may correspond to economic regions where one region is exposed to a trade shock ($t=1$) and the other is not ($t=0$).

Potential outcomes are defined for each worker and firm pair. We index the workers with latent types in $[0,1]$ and firms with latent types in $[2,3]$. The potential outcomes are then represented by $(Y_{0}^{*},Y_{1}^{*}):S \to \mathbb{R}^{2}$ where $S = [0,1]\times[2,3]$. For example, $Y_{t}^{*}(u,v)$ may describe the potential wage that a worker of type $u$ would earn at a firm of type $v$ when exposed or not exposed to the trade shock. Following Section 3.1.4, we assume that the researcher observes $Y_{1}$ and $Y_{0}$ where $Y_{t}(\phi_{t}(u),\psi_{t}(v)) =Y_{t}^{*}(u,v)$ for unknown measure preserving functions $\phi_{t}$ and $\psi_{t}$. The DPO is $F(y_{1},y_{0}) = \int\int\prod_{t \in \{0,1\}}\mathbbm{1}\{Y_{t}(\phi_{t}(u),\psi_{t}(v)) \leq y_{t}\} dudv$ and the DTE is $\Delta(y) = \int\int\mathbbm{1}\{Y_{1}(\phi_{1}(u),\psi_{1}(v))- Y_{0}(\phi_{0}(u),\psi_{0}(v)) \leq y\} dudv$. \looseness=-1

We symmetrize the potential outcome matrices along the lines of auerbach2022testing. Let $S^{2} = \left([0,1]\cup[2,3]\right) \times \left([0,1]\cup[2,3]\right)$ and define $(Y_{0}^{\dagger},Y_{1}^{\dagger}): S^{2} \to \mathbb{R}^{2}$ so that

equation*[equation* omitted — 227 chars of source]

and $\varphi_{t}(u) := \phi_{t}(u)\mathbbm{1}\{u \in [0,1]\} + \psi_{t}(u)\mathbbm{1}\{u \in [2,3]\}$ is measure preserving. Then the DPO is equal to $\frac{1}{2}\int\int\prod_{t \in \{0,1\}}\mathbbm{1}\{Y^{\dagger}_{t}(\varphi_{t}(u),\varphi_{t}(v)) \leq y_{t}\} dudv$. Since $Y^{\dagger}$ is symmetric and defined on one population (the population of workers and firms), the logic of Section 3 can be applied to bound the DPO and DTE. One can similarly define the STE using the eigenvalues of $Y_{t}^{\dagger}$. \looseness=-1

Row and column heterogeneity

Spectral methods may perform poorly when there is nontrivial heterogeneity in the row and column variances of the outcome matrices, see also auerbach2022testing. To address this issue we adapt arguments of finke1987quadratic. We decompose $\mathbbm{1}\{Y_{t}^{*}(u,v) \leq y_{t}\} = \alpha_{t}(u) + \alpha_{t}(v) + \epsilon_{t}(u,v)$ where $\int \epsilon_{t}(s,v)ds = \int \epsilon_{t}(u,s)ds = 0$ for every $u,v \in [0,1]$. The DPO becomes

align*[align* omitted — 370 chars of source]

The summand $\int\int \prod_{t \in \{0,1\}}\left(\alpha_{t}(\varphi_{t}(u)) + \alpha_{t}(\varphi_{t}(v))\right) dudv$ can be bounded along the lines of Theorem 368 in Section 4.2. The summand $ \int\int \prod_{t \in \{0,1\}}\epsilon_{t}(\varphi_{t}(u),\varphi_{t}(v))dudv$ can be bounded using the arguments from Section 4.3. One can similarly decompose $Y_{t}^{*}(u,v)$ and redefine the STE using the quantiles of $\alpha_{t}$, $\alpha_{t}$ and the eigenvalues of $\epsilon_{t}$. See Online Appendix Section D.1 for details. \looseness=-1

Estimation and inference

In many settings the researcher can exploit a symmetry in the experimental design to conduct randomization-based inference. We illustrate three ways of doing this using the motivating examples from Section 1.1 in Online Appendix Section D.2. Alternatively, if the researcher observes only noisy signals of $Y^{*}_{1}$ and $Y^{*}_{0}$ due to random sampling, missing data, measurement error, etc. then they can estimate the STE or bounds on the DPO and DTE by replacing the eigenvalues of $Y_{t}$ with empirical analogs. We formalize this strategy, give sufficient conditions for consistency, and sketch a strategy for statistical inference in Online Appendix Section D.3. \looseness=-1

Spillovers

One motivation for implementing a double randomized experimental design is to characterize social interactions, market externalities, or other spillovers between agents. Our framework and results can be directly applied to characterize distributional analogs of such spillover effects in many settings. The kinds of spillovers that are identified generally depend on the experimental design and assumptions about the agent interactions, see for instance bajari2021multiple. The following example, where we consider heterogeneous spillover effects under the assumption of strong no-interference (bajari2021multiple's Assumption 5.1), is related to their Section 6. We also provide three additional concrete examples concerning spillover effects under local interference (bajari2021multiple's Assumption 5.4), market externlaities, and social interactions in Online Appendix Sections C.1 and C.2. \looseness=-1

Consider the setting of the buyer-seller experiment in Example 4 of Section 1.1. Suppose the researcher is interested in how an information treatment assigned to buyers affects their transactions with untreated sellers. To do this, they independently randomize the buyers and sellers to two groups. Only pairs of buyers and sellers that are both assigned to the first group are treated, but transactions may occur between any buyer-seller pair. \looseness=-1

bajari2021multiple call this a conjunctive simple multiple randomization design in their Definition 8. They define the average buyer spillover effect to be the average difference in the potential transactions between the event that the buyer but not the seller is assigned to the treated group and the event that both the buyer and the seller are assigned to the untreated group. To characterize this spillover effect using our notation, let $(Y_{1}^{*},Y_{0}^{*}) : [0,1] \to \mathbb{R}^{2}$ record the potential transactions for pairs of buyers and sellers under the two events. bajari2021multiple's average buyer spillover effect is $\int \int \left(Y^{*}_{1}(u,v)-Y^{*}_{0}(u,v) \right)dudv$, or equivalently, $\int \int \left(Y_{1}(u,v)-Y_{0}(u,v) \right)dudv$ because $Y_{t}$ and $Y_{t}^{*}$ are equivalent up to a measure preserving transformation (see also their Lemma 1). After symmetrization as in Section 5.1, the arguments of Section 3 characterize distributional analogs of the average buyer spillover effect. That is, the joint distribution of potential transactions $\int\int\Pi_{t \in \{0,1\}}\mathbbm{1}\{Y^{*}_{t}(u,v) \leq y_{t}\}dudv$ and the distribution of buyer spillover effects $\int\int\mathbbm{1}\{Y_{1}^{*}(u,v) - Y_{0}^{*}(u,v) \leq y\}dudv$. \looseness=-1

Covariates and instruments

abadie2002instrumental,chernozhukov2005iv,firpo2007efficient use covariates or instruments to allow for endogeneity or characterize various conditional treatment effects. Their parameters can be written as solutions to an extremum estimation problem building on the framework of koenker1978regression. We have results for an analogous approach to incorporate covariates and instruments into the framework of this paper, but they are sufficiently complicated to warrant a separate, forthcoming paper. \looseness=-1

Two empirical demonstrations

We revisit Examples 1 and 3 from Section 1.1 and find policy relevant heterogeneity in the effect of treatment that might otherwise be missed by focusing exclusively on average effects. An R package can be found at \url{https://github.com/yong-cai/MatrixHTE}. \looseness=-1

Example 1: risk sharing

Our first demonstration follows comola2021treatment.\footnote{The data can be found on the Review of Economics and Statistics data repository: \url{https://doi.org/10.7910/DVN/K6QU2J}.} Households in nineteen villages are randomly provided with a savings account. A main finding of the authors is that “the intervention increased the transfers towards others and the overall informal financial activity in the villages, suggesting that there might be complementarities between formal savings and informal financial networks.” Our methodology suggests a more complicated relationship between formal savings and informal financial networks. In particular, we find that the treatment also decreased transfers between a nontrivial fraction of household pairs. \looseness=-1

Table 1 reports our bounds on the joint distribution of risk sharing links across all $19$ villages. Treatment $1$ is the event that both households are provided with a savings account, treatment $0$ is the event that neither household is provided with a savings account, and $Y_{ij,t}$ indicates whether household pair $ij$ reports a risk sharing link under treatment $t$. To construct Table 1, we first compute bounds on the distribution of potential outcomes for each village allowing for row and column heterogeneity as in Section 5.2. We then average the bounds over the $19$ villages, weighting by the number of households in each village. Table 1 also gives bounds on the distribution of treatment effects since $P(Y_{ij,1}-Y_{ij,0} = 1) = P(Y_{ij,1} = 1, Y_{ij,0} = 0)$, $P(Y_{ij,1}-Y_{ij,0} = -1) = P(Y_{ij,1} = 0, Y_{ij,0} = 1)$, and $P(Y_{ij,1}-Y_{ij,0} = 0) = P(Y_{ij,1} = 1, Y_{ij,0} = 1)+P(Y_{ij,1} = 0, Y_{ij,0} = 0)$. \looseness=-1

table[table omitted — 675 chars of source]

We find a positive lower bound on $P(Y_{ij,1} = 1, Y_{ij,0} = 0)$ which implies that the savings accounts create links. This is consistent with the main finding of comola2021treatment. We also find a positive lower bound on $P(Y_{ij,1} = 0, Y_{ij,0} = 1)$ which implies that the savings accounts also destroys links. This lower bound is at least a third of the upper bound on $P(Y_{ij,1} = 1, Y_{ij,0} = 0)$ and so is nontrivial in magnitude. However, the aggregate effect of the savings accounts on link creation is unlikely to be large and negative. This is because $P(Y_{ij,1} = 1, Y_{ij,0} = 0) - P(Y_{ij,1} = 0, Y_{ij,0} = 1)$ is not less than $-0.004$. The total change in the number of links is either positive or not substantially different from zero. \looseness=-1

Figure 1 shows a smoothed density plot of the spectral treatment effects on the treated (STT) for households in all 19 villages. For reference, we also show a smoothed density plot for a collection of conditional average treatment effects (CATE). To construct the CATE, we first bin the households by size and the number of children. Then for every pair of bins, we compute the difference in the fraction of links between households under both treatments. The CATE plot is then the smoothed density of the differences across treatments for every bin pair, weighting by the number of households in each bin. Figure 1 shows that the STT and CATE are similarly distributed, even though the STT is constructed without the use of any covariate information. \looseness=-1

figure[figure omitted — 620 chars of source]

Example 3: auction format

Our second demonstration follows athey2011comparing.\footnote{The data can be found on Phil Haile's website: \url{http://www.econ.yale.edu/ pah29/timber/timber.htm}. We restrict attention to a subsample of auctions proposed by schuster1994sealed in which the auction format is randomly assigned.} Tracts of forest land are sold by either an open or sealed bid auction format. A main finding of the authors is that “sealed bid auctions attract more small bidders [and] shift the allocation towards these bidders.” Our methodology suggests a more complicated relationship between auction format and firm entry. In particular, we find that the sealed bid format also discourages large firms from entry. \looseness=-1

Table 2 reports our bounds on the joint distribution of entry decisions. Treatment $1$ is the sealed bid format, treatment $0$ is the open format, and $Y_{ij,t}$ indicates whether firm $i$ bids on tract $j$ under format $t$. To construct Table 2, we symmetrize the outcome matrices as in Section 5.1 and allow for row and column heterogeneity as in Section 5.2. We report results for the full sample of firms, as well as large and small firms seperately. \looseness=-1

table[table omitted — 1,422 chars of source]

For the full sample, we find a strictly positive lower bound on $P(Y_{ij, 1} = 1, Y_{ij, 0} = 0)$ which implies that the sealed bid design encourages some firms to enter. However, we also find evidence that it discourages the entry of other firms. For the population of small and large firms separately, we find that while the sealed bid design induces entry and exit for both types of firms, there is at least twice as much exit of large firms than exit and entry of small firms. This is consistent with the main finding of athey2011comparing, that sealed bid auctions shift the allocation towards small firms. However, our results suggest that the exit of large firms as well as the entry of small firms drives this outcome. \looseness=-1

Figure 2 shows smoothed density plots of the STT and the CATE using firm size and tract location as covariates. The two distributions have similar centers with the bulk of the treatment effect above $0$. They also both have large left tails suggesting large negative treatment effects for a small number of firms and tracts. However, there are some noticable differences between the two plots. For example, the CATE plot concentrates at a few discrete spikes, which is not a feature of the STT. \looseness=-1

figure[figure omitted — 612 chars of source]

Conclusion

This paper characterizes the distribution of treatment effects in a double randomized experiment where a matrix of outcomes is associated with each treatment. We propose bounds on the distribution of treatment effects and a matrix analog of quantile treatment effects. Our results are based on a new matrix analog of the Fr\'echet-Hoeffding bounds that play a key role in the standard theory. We illustrate our methodology with two empirical demonstrations and find policy relevant heterogeneity that might be missed by focusing exclusively on averages. \looseness=-1