Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
71,282 characters · 12 sections · 34 citation commands
Toggling the Defiers to Relax Monotonicity: The Difference-in-Instrumental-Variables Estimand
A central insight of angrist1996identification is that instrumental variables (IV) methods identify causal effects under a set of assumptions. Specifically, the IV estimand recovers a Local Average Treatment Effect (LATE) for compliers when four conditions hold: relevance, exogeneity, exclusion, and monotonicity. The monotonicity assumption requires that the instrument shifts treatment take-up in the same direction for all units, thereby ruling out the existence of defiers.
In many real-world settings, however, defiers are both plausible and empirically salient. This may be especially prominent in settings involving divisive, controversial, or authority-laden domains such as religion, politics, or, more recently, the deployment of artificial intelligence systems. Psychological research on reactance Brehm1966,MironBrehm2006 documents that some individuals respond to encouragements or recommendations by actively resisting them. Thus, defiers are not rare anomalies but systematic behavioral responses that cannot simply be ignored. Consequently, in such environments, assuming monotonicity may be inappropriate and inference may become difficult to interpret.
Importantly, environments in which behavioral responses differ across instruments are not confined to controversial domains. They arise whenever distinct encouragements operate through different incentives, framings, informational content, or authority structures. Such designs are common in experimental and quasi-experimental work featuring multiple treatment arms, varying encouragement intensities, competing messages, or alternative policy thresholds. In these settings, the possibility that some individuals comply with one instrument but resist another is a structural feature of the design rather than a pathological exception.
This paper introduces the Difference-in-Instrumental-Variables (DIIV) estimand, which directly incorporates the existence of defiers into causal analysis. The core idea is to use two instruments that shift compliance and defiance in opposite directions. By contrasting how these instruments affect treatment take-up, it becomes possible to identify an interpretable causal effect even when monotonicity fails.\footnote{ The logic parallels a classic two-guardians riddle, popularized in the film Labyrinth and in logic puzzles. A traveler faces two doors, one safe and one unsafe, guarded by one truth-teller and one liar, and may ask only one question. By asking a logically inverted question (such as what the other guardian would say) one can recover the correct door regardless of which guardian is addressed. In a similar fashion, contrasting two instruments with opposite compliance patterns allows identification that does not depend on knowing whether a given unit is a complier or a defier. See Proposition (ref) for a DIIV representation that parallels the mathematical-logical solution to the two-guardians riddle.} Our setup is developed with a focus on experimental design, where researchers are aware that potential defiers exist and may deliberately induce different shifts in compliance and defiance by designing distinct instruments.
Crucially, the proposed estimand does not require researchers to observe or classify behavioral types. It relies only on variation across instruments with distinct behavioral orientation, a feature already present in many empirical designs. As a result, DIIV can be implemented in settings that already contain multiple encouragement arms or alternative policy levers, without imposing additional structure beyond standard IV conditions. More generally, the approach applies whenever two instruments induce opposing behavioral responses in compliance and defiance while satisfying relevance, exogeneity, and exclusion.
In addition, we develop a simple micro-foundation for treatment take-up that explains why individuals from the same population may behave differently depending on the instrument. Different instruments convey different information, incentives, or framings, leading the same individual to take the treatment under one instrument and not under another. This micro-foundation clarifies why marginal treatment effects may depend on unobserved types rather than on observed take-up behavior alone, and provides an explicit behavioral interpretation of the objects identified by DIIV.
A key result of the paper is that the DIIV estimand always yields a convex combination of the marginal treatment effects on compliers and defiers, providing an informative and interpretable causal object even when monotonicity fails.\footnote{See Proposition (ref) for the formal expression.} In contrast to standard multi-instrument 2SLS settings, where weights may be negative or lack a clear behavioral interpretation mogstad2021causal, the DIIV weights correspond to population shares of marginal behavioral responses. In this sense, the estimand identifies the treatment effect for units whose take-up behavior responds differentially across the two instruments. When combined with plausible sign restrictions or stability conditions on behavioral groups, it serves as a transparent location statistic for the underlying marginal effects.
Because the object is defined in terms of differential responsiveness across instruments, it has a natural interpretation in design-based work: it measures the average outcome change associated with shifting from a weaker-compliance instrument to a stronger-compliance one. This interpretation is particularly relevant in environments where researchers compare alternative encouragement strategies, messaging regimes, or policy intensities. Under plausible sign restrictions or stability conditions on behavioral groups, DIIV also serves as a location statistic for the underlying marginal effects. DIIV therefore extends the IV toolkit to settings where behavioral heterogeneity induced by incentives, framing, or reactions to authority is central, while preserving a point-identified and behaviorally interpretable effect.
The instrumental variables framework classifies individuals into four types: always-takers, never-takers, compliers, and defiers. The standard LATE framework assumes monotonicity, thereby ruling out defiers and enabling point identification of complier effects imbens1994,angrist1996identification. While analytically convenient, monotonicity can be substantively strong. A large literature has therefore examined its empirical content and implications.
One strand develops empirical diagnostics for IV validity and related identifying restrictions, including inequality restrictions, moment conditions, and overidentification tests when multiple instruments are available hansen1982,huber2015,kitagawa2015,mourifie2017. These approaches clarify testable implications of the IV model but do not in themselves recover interpretable effects when defiers are present.
When monotonicity fails, point identification of the LATE generally breaks down. Partial identification approaches derive bounds on average treatment effects that explicitly allow for defiers BalkePearl1997,manski1990, with subsequent refinements incorporating covariates and additional structure richardson2010,swanson2015. More recent contributions develop sensitivity analyses that parameterize monotonicity violations and assess robustness to different shares of defiers noack2021, or attempt to estimate or bound the prevalence of defiers directly christy2024countingdefiersdesignbasedmodel. These methods provide valuable information about identification limits but typically replace point identification with bounds or sensitivity regions.
A related literature proposes weaker versions of monotonicity. Stochastic monotonicity permits certain forms of non-deterministic compliance behavior while preserving interpretable weighted effects small2017, and other structured relaxations accommodate defiers under additional assumptions dechaisemartin2017,dahl2023. The most closely related work is mogstad2021causal, who show that with multiple instruments and “partial monotonicity” (monotonicity by each instrument, not necessarily by the vector of instruments), 2SLS recovers a positively weighted average of instrument-specific LATEs. Their work highlights the interpretive complexity that can arise when aggregating instruments, even under weakened monotonicity.
In contrast to bounds, sensitivity analyses, or partial monotonicity approaches, DIIV exploits two instruments with opposing compliance patterns to recover a point-identified estimand with a direct behavioral interpretation, without imposing monotonicity. By focusing on differential responsiveness across instruments, it provides a simple and general extension of the IV framework to environments where defiance is a structural feature of the design.
We first develop a basic setup with two instruments that encourage treatment take-up, and shift compliance and defiance in opposite directions. In later sections, we extend the analysis to more general settings.
Let $I$ be a population of units $i\in I$. By design, the population is partitioned into two ex ante identical subpopulations, $I_1$ and $I_2$, where each group is potentially exposed to an exclusive instrument, also referred to as a nudge.\footnote{Here, “identical” refers to underlying type composition and potential outcomes; realized compliance behavior may differ across subpopulations due solely to different framing of each instrument.} For $j\in\{1,2\}$, the exclusive instrument in subpopulation $I_j$ is denoted $z_{i,j}\in\{0,1\}$ and is exogenously assigned with the same probability across subpopulations. Instrument $z_{i,j}$ encodes not only assignment but also the framing associated with instrument $j$.
Define $D_{i,j}(z)\in\{0,1\}$, for $z\in\{0,1\}$, as the potential treatment take-up in $I_j$, which depends both on the assignment $z$ and on the framing $j$. The observed treatment status is $d_i=D_{i,j}(z_{i,j})$, which, together with the potential outcomes $Y_i(d)$ for $d\in\{0,1\}$, yields the observed outcome
In particular, the distribution of $y_i$ may depend on the framing $j$ only through the treatment decision $d_i$ and not through frame-specific potential outcomes.\footnote{From the notation, the exclusion restriction is satisfied; see Assumption (ref).}
Based on how units behave in response to the instrument, they can be one of four types $\theta_i\in\{A,N,C,F\}$, which partition $I$, $I_1$, and $I_2$. Let $\pi_\theta$ denote the population mass of type $\theta$. Types $A$ and $N$ are the traditional always-takers and never-takers.\footnote{Types $A$ and $N$ are not essential and could be assumed to have zero mass.} Type $C$ consists of units that find value in the nudge and are potential compliers, whereas type $F$ consists of units that find disvalue in the nudge and are potential defiers. Depending on responsiveness to framing $j$, subpopulations $I_1$ and $I_2$ may exhibit different shares of units behaving as compliers ($d_i=z_{i,j}$) or as defiers ($d_i=1-z_{i,j}$).
More precisely, given an instrument with framing $j$, potential compliers (type $C$) take the treatment iff
where $\eta_i$ is an idiosyncratic shock with an atomless distribution; $v_C(j)\ge 0$ captures responsiveness to framing $j$; and $\underline v$ is a threshold.\footnote{In a related approach, kowalski2023behaviour exploit an ordering of the slopes of marginal treated and untreated outcome as functions of covariates to bound treatment effects for always takers.} Conditional on type $C$, units with $\eta_i\in[\underline v - v_C(j),\underline v)$ are contingent compliers, denoted $C_j^C$, with conditional mass $\phi_C^C(j)$. Those with $\eta_i\ge\underline v$ behave as contingent always-takers ($C_j^A$), with conditional mass $\phi_C^A(j)$, while those with $\eta_i<\underline v-v_C(j)$ behave as contingent never-takers ($C_j^N$), with conditional mass $\phi_C^N(j)$. These satisfy $\phi_C^C(j)+\phi_C^A(j)+\phi_C^N(j)=1$ (see Figure (ref)).
Similarly, given an instrument with framing $j$, potential defiers (type $F$) take the treatment iff
where $v_F(j)\le 0$. Conditional on type $F$, units with $\eta_i\in[\underline v,\underline v - v_F(j))$ are contingent defiers, denoted $F_j^F$, with conditional mass $\phi_F^F(j)$. Those with $\eta_i\ge\underline v - v_F(j)$ behave as contingent always-takers ($F_j^A$), with conditional mass $\phi_F^A(j)$, while those with $\eta_i<\underline v$ behave as contingent never-takers ($F_j^N$), with conditional mass $\phi_F^N(j)$. These satisfy $\phi_F^F(j)+\phi_F^A(j)+\phi_F^N(j)=1$ (see Figure (ref)).
Thus, for each framing $j\in\{1,2\}$, we obtain a partition \[ I_j=\{A_j,N_j,C_j^C,C_j^A,C_j^N,F_j^F,F_j^A,F_j^N\}. \] We pay special attention to the unconditional shares of behavioral compliers and behavioral defiers, which depend on the framing. We denote these by
Although we do not impose the existence of defiers (i.e., $F=\emptyset$ is allowed), we develop a method to estimate local average treatment effects when they do exist. To estimate the effect of treatment take-up on outcomes, we impose the following assumptions.
The next assumption replaces monotonicity while also serving as the relevance condition.
For $j\in\{1,2\}$, define the reduced form and first stage as
and
Lemma (ref) shows how the standard IV estimand $RF_j/FS_j$ is biased when $p_F^j>0$. We therefore propose an alternative estimand that exploits differential shifts in compliance and defiance.
Note that, under standard IV terminology and assumptions (i.e., monotonicity and the interpretation that instruments act as encouragements), a special case of part (i) is obtained when $p_F^1=p_F^2=0$.
The DIIV estimand can be interpreted as a LATE that mixes the effects on potential complier (type $C$) and potential defier (type $F$) units whose behavior shifts when moving from one framing to the other. The weights in Proposition (ref) are determined by how the change in framing alters the shares of contingent compliers and defiers. Thus, $\tau_{\mathrm{DIIV}}$ captures a LATE that is directly tied to the policy margin of modifying how a nudge is framed.
Note that both instruments $z_j$ could be pooled into a single instrument across the entire population $I=I_1\cup I_2$ by extending the assignment $z_{i,j}=0$ to units in the other subpopulation. In that case, defining $z_{\mathrm{pool}}=z_1+z_2$, the standard IV estimand would be
which is a weighted but non-convex combination of $\tau_C$ and $\tau_F$. Therefore, pooling alone does not resolve the bias induced by defiers.
Instead, consider a mistakenly coded instrument $z_2^+ = 1 - z_2$ (i.e., a toggle) and pool it with $z_1$. By flipping the directive of the second instrument, compliers become defiers and defiers become compliers relative to that framing. This change in roles enables a differential comparison across framings that restores identification.\footnote{This role reversal is analogous to the solution of the previously mentioned two-guardian riddle in Footnote (ref). Asking one guardian about the answer the other would give effectively flips the directive: truth-telling is converted into lying and vice versa. In the present setting, recoding the instrument as $z_2^+=1-z_2$ similarly toggles the behavioral mapping of the instrument, transforming compliers into defiers and defiers into compliers relative to the original framing.}
The DIIV estimand is an unbiased convex combination of $\tau_C$ and $\tau_F$. From its definition, it can be estimated using differences in sample means:
where $\bar y_j^z$ denotes the sample mean of $y_i$ in subpopulation $j$ with $z_{i,j}=z$, and similarly for $\bar d_j^z$. While this estimator is straightforward to compute, inference based directly on this ratio is less convenient.
The next result shows that the DIIV estimand admits a numerically equivalent two-stage least squares (2SLS) representation, which allows for conventional inference using standard tools.
This 2SLS representation is not unique. However, it is convenient and elegant. The instrument $w_i$ can be written as $w_i = 1 - (z_{\mathrm{pool},i} \oplus h_i)$, where $\oplus$ denotes the exclusive-or operator.\footnote{The role of the exclusive-or (XOR) operator mirrors the logical structure underlying the two-guardian riddle from Footnote (ref). The solution is achieved by effectively applying an XOR operation: truthfulness and falsity cancel out because the answer depends on whether exactly one of the two conditions holds. Asking one guardian what the other would say toggles the truth value twice, producing a determinate answer. Analogously, the XOR operator here formalizes the idea that alignment versus misalignment between the pooled instrument and the framing matters only through their exclusive disagreement, allowing the directive to be flipped in a way that restores identification.} It equals one when the directive of the pooled instrument aligns with the framing and zero otherwise. This representation provides a convenient implementation of DIIV via standard 2SLS. Although elegant, this particular 2SLS construction is not robust to asymmetries or correlations in instrument assignment. In the next section, we show that while the DIIV estimand remains valid in more general settings, the 2SLS representation requires appropriate adjustments.
We now extend the analysis to a broader setting in which two binary instruments $(z_1,z_2)\in\{0,1\}^2$ may operate jointly, each combination occurring with probability $q_{z_1z_2}$, and their intended directive for treatment take-up is explicit. To interpret how each instrument influences behavior, we associate to each $z_j$ a frame
where $m_j$ captures the salient attribute emphasized by the frame and $s_j\in\{-1,+1\}$ records the directive orientation: $s_j=+1$ for encouragement and $s_j=-1$ for discouragement.\footnote{This decomposition aligns with the distinction, common in information design and framing models, between the content of information and its directional implications for choice. In Bayesian persuasion models, a sender selects an information structure that determines which attributes or dimensions of the decision problem become salient, while posterior beliefs (and hence actions) respond endogenously to this structure KamenicaGentzkow2011. Similarly, framing models allow behavior to vary depending on which attribute of an otherwise identical choice environment is emphasized, even when objective payoffs are unchanged TverskyKahneman1981,DellaVigna2009. In our setting, $m_j$ captures the attribute or informational dimension activated by the instrument, while $s_j$ captures the directive orientation along that dimension (encouragement versus discouragement). This separation allows instruments to differ in what aspect of the choice they target, independently of \emph{how} they push behavior, a distinction that underlies the construction of DIIV based on the exclusive-or ($\oplus$) operator in Proposition (ref).} We also note that since the directive's sign matters, we no longer refer to units in populations $C$ and $F$ as potential compliers and potential defiers; instead, we will refer to them as \textit{persuation-prone} and \textit{reactance-prone}, respectively.
The attribute $m_j$ can be very general, ranging from monetary incentives to purely persuasive content. In experimental settings, $s_j$ reflects the researcher’s intended direction; in observational settings, it reflects the theoretically inferred direction. We assume that $s_j$ is unambiguous and well defined. Under monotonicity, orientation plays no identifying role.\footnote{For example, by replacing an instrument $z$ by $1-z$, the sign of the ratio $\mathrm{FS}/\mathrm{RF}$ would remain unchanged.} However, when monotonicity fails, the directive becomes essential for interpreting the behavior of those who take the treatment as intended by the directive $s$ (i.e., persuation-prone) and those who do the opposite (i.e., reactance-prone).\footnote{This is particularly important when marginal effects differ between persuasion-prone and reactance-prone units. If marginal effects were homogeneous across types, the directive would not matter for interpretation.}
Each attribute $m_j$ determines how strongly units respond to the presence of the instrument, while the directive $s_j$ determines the sign of that response. For each behavioral type $\theta\in\{C,F\}$, let $\kappa_\theta(m_j)$ measure the intensity with which type $\theta$ responds to attribute $m_j$. Persuasion-prone types ($C$) respond positively to the attribute and therefore satisfy $\kappa_C(m_j)\ge 0$, whereas reactance-prone types ($F$) respond negatively and satisfy $\kappa_F(m_j)\le 0$.
Treatment take-up follows a threshold rule in which the directed frames from both instruments jointly influence behavior:
where $\eta_i$ is an idiosyncratic shock. Thus, the frame component $m_j$ governs the strength of the effect, while $s_j$ governs its sign. Units may behave differently depending on their type and the realization of $\eta_i$.
For each type $\theta\in\{C,F\}$ and instrument $j$, let $\varphi_\theta(j)$ denote the probability that the threshold is met when comparing $(z_j{=}1,z_{-j}{=}0)$ to $(z_j{=}0,z_{-j}{=}0)$. From the threshold rule in Equation (ref), the difference between these two assignments is governed solely by the increment $\kappa_\theta(m_j)s_j$ associated with setting $z_j$ from $0$ to $1$ while holding $z_{-j}=0$ fixed. Thus, type $\theta$ units respond to instrument $j$ precisely when the idiosyncratic shock $\eta_i$ lies in the interval \[ \eta_i \in
\] Then, $\varphi_\theta(j)$ is the probability mass of $\eta_i$ contained in this interval. The behavioral share of type $\theta$ shifted by instrument $j$ is then \[ p_\theta^j = \pi_\theta\, \varphi_\theta(j), \] where $\pi_\theta$ is the population share of type $\theta$. The probability $p_\theta^j$ extends the notions of contingent compliance and defiance shares from the parallel-frame setting and characterizes how each instrument shifts persuasion-prone and reactance-prone groups under its specific attribute and directive. These quantities serve as the building blocks for defining first-stage contrasts and for establishing identification of the general DIIV estimand.
To estimate the causal effect of treatment take-up $d_i$ on outcome $y_i$ in the presence of opposing behavioral responses, we impose a set of assumptions that parallel the standard IV framework while replacing monotonicity with a directional structure on compliance and defiance. These assumptions are denoted with primes ($\prime$) to distinguish them from the baseline IV assumptions.
\setcounter{assumptionprime}{0}
Note that Assumption (ref) is equivalent to:
Let $\mathrm{RF}_{j}^{(z)}$ and $\mathrm{FS}_{j}^{(z)}$ be the reduced form and first stage with respect to instrument $j$ by keeping $z_{-j}=z$ constant:
and
Next, we extend DIIV to the oriented edge differences case.
Note that the influence of the vector $(z_1,z_2)=(1,1)$ on $d$ and $y$ does not enter the DIIV estimand; DIIV is an edge contrast with respect to $(0,0)$. That is, it compares $(1,0)$ with $(0,0)$ and $(0,1)$ with $(0,0)$. The following Lemma is useful for understanding how these edge contrasts are represented under directive orientations.
Finally, we generalize Proposition (ref) to include asymmetric and joint distributions of $(z_1,z_2)$ with explicit directives $s_j$ that allow for encouragements and discouragements.
We use data from three replication packages from studies regarding education barrera2019medium, loan take-up bertrand2010s, and vote turnout gerber2008secrecy,gerber2012ballotsecrecy.\footnote{The data from the former two studies is publicly available. Although the data to replicate from gerber2012ballotsecrecy is publicly available, DIIV requires a disaggregation of the intervention arms. This additional data was generously provided by Gregory Huber.} All studies are structured as large-scale randomized controlled trials (RCTs) with multiple treatment arms, which we regard as instruments. From each study, we take one immediate outcome and use it as treatment take-up ($d$) and estimate how it affected a later outcome ($y$). That is, we assume exclusion from the instruments onto $y$. For each study, we discuss why this is likely the case.
This section uses data from barrera2019medium to apply DIIV to a large conditional cash transfer RCT with three encouragement arms. The study population consists of $N=17{,}309$ students from Colombia. The intervention studies the effect of cash transfers on education outcomes. The three intervention arms are: (G1) offers USD 30 every two months conditional on secondary school attendance; (G2) has the same flow as G1 but with USD 10 retained and paid at the end of the year (revenue-equivalent conditional on compliance); and (G3) has the same flow as G2 plus an additional USD 300 lump-sum if the student enrolls in tertiary education (e.g., college). Table (ref) reproduces the immediate and long-term effects on education outcomes following barrera2019medium, Tables 3 and 4.
Because G1 and G2 are nearly identical as incentives, we pool them into a low-compliance encouragement $z_2$, while we treat G3 as a high-compliance encouragement $z_1$. Although interventions G1 and G2 were independent of intervention G3, we cannot apply the 2SLS equivalence from Proposition (ref) because of asymmetry. Therefore, we use $x^\Delta = z_1 - z_2$ as an instrument, $x^\Sigma = z_1 + z_2$ as a covariate, and the 2SLS representation in Proposition (ref). We note that $x^\times = z_1 z_2 = 0$, so it is omitted.
In this setting, a natural interpretation of defiant behavior is one of frustration or disengagement. Since the main goal of the intervention was to promote graduation, students who lost part of the cash transfers due to non-attendance may experience frustration, potentially leading to non-graduation, despite being otherwise likely to graduate and enroll in tertiary education in the absence of the intervention.
It is clear that $z_1$ generates more compliers, as the additional USD 300 payment tied explicitly to tertiary enrollment substantially increases the salience and perceived returns to post-secondary education. In contrast, defiant behavior is likely to be more prevalent under $z_2$, where the opportunity cost of non-compliance is lower. For such individuals, assignment to the low-compliance encouragement may reduce educational effort more severely, due to the absence of the additional bonus conditional on tertiary enrollment that is present in G3.
Since G3 directly affects tertiary enrollment, we cannot use tertiary enrollment as a final outcome variable. Instead, we estimate the effect tertiary enrollment $\Rightarrow$ tertiary graduation. The exclusion restriction in this setting requires that the cash transfer interventions affect tertiary graduation only through their impact on tertiary enrollment. This is plausible because the transfers are conditional on contemporaneous school attendance and do not provide direct incentives or information related to tertiary completion beyond inducing enrollment.
Table (ref) compares three IV specifications: using G3 as the only instrument, using a pooled any incentive instrument, and using DIIV. Standard IV estimators based on either G3 alone or pooled instruments yield weak first stages and statistically insignificant second-stage estimates. In contrast, DIIV generates a strong first-stage F-statistic by exploiting the contrast between encouragements and recovers a sizable, positive, and statistically significant effect. The estimated weighted average effect is approximately 75%, indicating that among students whose tertiary enrollment was induced by the intervention, roughly three-quarters went on to graduate.
We next study advertising in the microfinance sector of South Africa, using data from bertrand2010s. This study implements a large randomized controlled trial with $N=58{,}168$ potential borrowers receiving advertising letters that varied along multiple information and persuasion dimensions. The paper studies the effect of these interventions on $(i)$ application to the partner lender, $(ii)$ loan take-up at the partner lender, and $(iii)$ loan take-up from outside lenders. Table (ref) reports a slight variation of Table 3 in bertrand2010s.
We regard loan take-up from other lenders as the final outcome $y$ and application to the partner lender as the treatment variable $d$. Under standard substitution logic, and because the vast majority of applicants got the load, an increase in applications to the partner lender should reduce borrowing from competitors, yielding a negative causal effect, which is the hypothesis we seek to test. We use two advertising nudges as instruments. The low-interest-rate indicator $z_1$ is the natural high-compliance encouragement. To construct a lower-compliance instrument, we use whether the letter included a competitor payment example. The letter either displayed or omitted an illustrative competitor repayment schedule; from the example, borrowers could infer an implied competitor interest rate. We define $z_2 :=$ “competitor example not shown” as the corresponding encouragement.
In the field experiment, the competitor payment comparison shown to some borrowers was constructed using a monthly interest rate of 15%. Evidence from a companion study using the same lender indicates that this rate lies just above the upper bound of the lender’s own pricing distribution karlan2008credit, rather than reflecting a central or typical benchmark. As a result, the competitor comparison likely overstated the cost of the relevant outside option faced by many borrowers and plausibly operated as a discouraging frame rather than as a neutral source of information.
This interpretation also has implications for behavioral responses when the competitor comparison is removed. If the status quo in lending advertising emphasizes unfavorable alternatives through a high-cost benchmark, then omitting this comparison constitutes an encouragement that moves away from that framing. While some borrowers may become more willing to apply once the discouraging comparison is removed, others may respond in the opposite direction. In particular, borrowers whose take-up was supported by the relative attractiveness created by the high-cost comparison may become less likely to apply once this framing is withdrawn. Thus, in addition to discouragement, the experimental design naturally allows for defiance with respect to the encouragement that omits the competitor example.
We are interested in estimating the causal effect application to partner lender $\Rightarrow$ obtain a loan from competitors. The exclusion restriction requires that advertising interventions affect outside borrowing only through their impact on application to the partner lender. This restriction is likely satisfied because the mailings do not reduce search costs, alter eligibility, or expand credit supply in outside markets. Moreover, while one intervention arm includes a competitor payment example from which an interest rate can be inferred, this rate was plausibly perceived as unrepresentative, limiting its ability to directly affect borrowing decisions outside the partner lender.
Table (ref) compares three IV specifications: using the low-interest-rate indicator alone, using a pooled any incentive instrument ($\max\{z_1,z_2\}$), and using DIIV. Using either the strongest incentive alone or pooling instruments yields estimates that are either statistically insignificant or of the wrong sign relative to the hypothesized substitution effect (possibly due to large biases). In contrast, DIIV generates a strong first stage and a negative, economically meaningful estimate, consistent with substitution toward outside lenders when applications to the partner lender increase. In this setting, DIIV improves estimation by separating price-based encouragement from framing-based discouragement, thereby isolating the behavioral margin relevant for competitive borrowing.
We next study voter participation and persistence in turnout using two closely related field experiments on ballot secrecy and voter mobilization. The first study, gerber2008secrecy, implements a large-scale randomized intervention in the context of the 2010 U.S. congressional election ($N=69{,}488$). Households received official letters sent on Secretary of State letterhead that varied in their informational content regarding election administration and ballot secrecy. The second study, gerber2012ballotsecrecy ($N=3{,}744$), builds on the same experimental design to examine whether these interventions have persistent effects on participation in the subsequent 2012 elections.
The interventions in gerber2008secrecy include letters explicitly addressing concerns about ballot secrecy, alongside two placebo communications that do not mention secrecy.\footnote{There are also letters addressing civic duty, which were not used in gerber2012ballotsecrecy and are therefore excluded here.} The short placebo contains minimal text, while the long placebo matches the secrecy letters in length and emphasizes the role of the Secretary of State in election administration without providing reassurance about ballot secrecy. The primary outcome analyzed in that study is turnout in the 2010 congressional election.
Our object of interest follows the focus of gerber2012ballotsecrecy. We study the causal effect voting in 2010 $\Rightarrow$ voting in 2012, using the randomized ballot-related mailings as instruments. Turnout in 2010 constitutes the treatment variable, while the mailings serve as exogenous sources of variation in participation. Identification therefore relies on an exclusion restriction: the 2010 interventions must affect turnout in 2012 only through induced participation in 2010. Given the one-time nature of the mailing, the two-year gap between elections, and the absence of any reinforcement or forward-looking content, this assumption is plausible. In particular, it is unlikely that participation decisions in 2012 are directly influenced by recall of a ballot-secrecy letter received two years earlier. Instead, persistence in turnout is most naturally attributed to habit formation, consistent with the interpretation emphasized in gerber2012ballotsecrecy.
Table (ref) replicates the main reduced-form results from gerber2012ballotsecrecy, taking the short placebo mailing as the baseline condition. The sample size is smaller due to data availability constraints. We observe positive reduced-form effects of the secrecy interventions on turnout in the 2012 primary election, while no statistically significant effects are detected for the general election.
We define $z_1$ as an indicator for assignment to any ballot-secrecy letter and $z_2$ as an indicator for assignment to the long placebo mailing. Although there is no $(z_1,z_2)=(1,1)$ group, we exploit contrasts between these interventions using the 2SLS representation in Proposition (ref). This approach is appropriate because there is no pure no-contact control group, the intervention arms are not parallel in the sense of Section (ref), and assignment probabilities differ across treatments.
While some individuals may be reassured by explicit statements about ballot secrecy and become more likely to vote, others may respond adversely to official communications that heighten the salience of electoral authorities or institutional oversight. These concerns are particularly relevant for the long placebo intervention, which increases the salience of the Secretary of State without addressing secrecy and may discourage participation among individuals with low institutional trust.
As a result, opposing participation shifts across interventions are plausible. Secrecy-focused letters are likely to generate compliers by alleviating privacy concerns, whereas the long placebo may induce defiance by emphasizing institutional authority without reassurance. Pooling these interventions therefore likely aggregates positive and negative first-stage responses, implicitly combining encouragement-induced participation with discouragement-driven abstention.
Table (ref) compares three IV specifications: using the secrecy intervention alone, using a pooled any intervention instrument, and using DIIV. Across all specifications, first-stage F-statistics are modest, reflecting the well-known difficulty of mobilizing voter participation and the substantial loss of statistical power induced by restricting the sample to individuals observed in both the 2010 and 2012 elections. In the DIIV specification, the first-stage F-statistic is approximately 6.8. While below conventional thresholds, this value implies that any finite-sample bias due to weak instruments is limited. In particular, an F-statistic of this magnitude corresponds to a relative bias on the order of 15--20% of the OLS bias, which is modest in this setting given that OLS estimates of turnout persistence are themselves small and potentially attenuated by measurement error.
Turning to the second stage, single-instrument IV estimates are positive but imprecise and sensitive to pooling choices, reflecting the aggregation of heterogeneous behavioral responses. In contrast, DIIV yields larger and statistically significant estimates for both presidential and congressional turnout in 2012. The point estimates imply that voting in 2010 increases the probability of voting again by approximately 35% in presidential primaries and about 22% in congressional primaries. In this setting, DIIV improves identification by separating reassurance-based encouragement from authority-salience-based discouragement, thereby isolating the behavioral margin most relevant for turnout persistence.
Finally, we conduct simulations to clarify the interpretation of DIIV relative to standard overidentified IV when compliers and defiers coexist. We consider four simulation environments. Each environment consists of 1{,}000 independent trials, where each trial represents an experiment with a sample size of 10{,}000 observations.
The population is composed of four behavioral types: 10% always-takers, 10% never-takers, 40% persuasion-prone units (potential compliers), and 40% reactance-prone units (potential defiers). Treatment effects are heterogeneous across types, with potential outcome gains given by $\tau_A = 4$, $\tau_C = 3$, $\tau_F = 2$, and $\tau_N = 1$. Outcome noise is additive and normally distributed with variance one.
Two binary instruments, $z_1$ and $z_2$, are generated by thresholding correlated latent Gaussian variables. Specifically, $e_1 \sim \mathcal{N}(0,1)$, $e_2 = u + \rho e_1$ with $u \sim \mathcal{N}(0,1)$ and $\rho = -0.45$, and $z_j = \mathbbm{1}\{e_j>0\}$ for $j\in\{1,2\}$. Thus, each instrument equals one with probability $1/2$, and their dependence reflects the latent Gaussian correlation.
For always-takers, $d_i=1$, and for never-takers, $d_i=0$. For $\theta \in \{C,F\}$, treatment take-up follows \[ d_i = \mathbbm{1}\{\kappa_{\theta 1} z_{i,1} + \kappa_{\theta 2} z_{i,2} > \eta_i\}, \] where $\eta_i \sim \mathcal{N}(0,\sigma^2)$ independently of $(z_1,z_2)$.
We vary two dimensions across environments: (i) Responsiveness asymmetry: either persuasion-prone units (type $C$) are more responsive than reactance-prone units (type $F$), or vice versa, implemented through larger absolute values of the corresponding coefficients $\kappa_{\theta j}$; and (ii) Behavioral overlap: a high-noise case with $\sigma=2$ and a low-noise case with $\sigma=1$. Lower noise sharpens behavioral responses.
For each environment, we estimate (i) standard overidentified 2SLS using $(z_1,z_2)$ as instruments and (ii) DIIV implemented via a just-identified 2SLS representation using the difference instrument $x^\Delta = z_1 - z_2$, controlling for $x^\Sigma = z_1 + z_2$ and $x^\times = z_1 z_2$.
Figure (ref) reports the sampling distributions of both estimators. In all cases, DIIV lies within the expected range bounded by $\tau_F$ and $\tau_C$. In contrast, the overidentified IV estimator may fall outside this range, moving closer to $\tau_F$ when $p_F^2-p_F^1 > p_C^1-p_C^2$, and closer to $\tau_C$ when $p_C^1-p_C^2 > p_F^2-p_F^1$.
The simulations highlight an interpretational distinction between DIIV and standard IV. DIIV is explicitly constructed as a convex combination of the marginal treatment effects for compliers and defiers, ensuring that the resulting estimand remains between $\tau_C$ and $\tau_F$ across all configurations. In contrast, IV with multiple instruments forms an implicit linear combination of the instruments’ first stages. When defiers are present, these implicit weights need not be positive, so the overidentified IV estimand may place disproportionate weight on one type and need not lie between $\tau_C$ and $\tau_F$.
Overall, the simulations show that DIIV delivers a stable and interpretable estimand precisely in environments where monotonicity violations are empirically relevant. By making the aggregation of complier and defier effects explicit and tied directly to observable shifts in treatment take-up, DIIV avoids the ambiguity that can arise from implicit weighting in overidentified IV designs.
DIIV is designed to incorporate the existence of defiers into the instrumental variables framework. By contrasting reduced forms and first stages across instruments with opposing behavioral orientation, DIIV identifies a convex combination of the marginal treatment effects on compliers and defiers. The weights in this combination reflect the relative net shifts in treatment take-up induced by each instrument, yielding a behaviorally interpretable estimand even when monotonicity fails.
Because DIIV admits a 2SLS-equivalent representation through simple data transformations, it can be implemented using standard estimation and inference tools. More broadly, the approach provides a transparent extension of IV to environments where distinct encouragements induce heterogeneous responses across individuals. In such settings, DIIV allows researchers to retain point identification and interpretability without ruling out empirically relevant forms of defiance.