Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
72,728 characters · 12 sections · 32 citation commands
Stochastic Potential Choices and Outcomes
\noindentKeywords: stochastic potential outcomes, instrumental variables, compliance, stochastic choice, probabilistic causation, IV.
The standard potential outcomes framework in econometrics and statistics usually treats each individual's potential outcomes as fixed numbers (neyman1923, rubin1974, see also haavelmo1944, p.23). In an instrumental-variables (IV) design, it also usually treats each individual's potential treatment choices as fixed numbers. For example, in a binary instrument with values $z'$ and $z$, individual $i$ has potential treatment statuses \[ D_i(z'),D_i(z)\in\{0,1\}. \] Changing the instrument value either changes the person's treatment status or it does not. Under monotonicity, the individuals whose status changes are the compliers, and the Wald estimand identifies the average treatment effect for that group (see imbensangrist1994, angristimbensrubin1996).
That representation is powerful, but has a flavor of determinism (or, as put by dawid2000, “fatalism”), and can be too sharp for some empirical settings. For instance, a reminder may make enrollment more likely without making it certain. A subsidy may raise take-up likelihood from $20\%$ to $50\%$ for one person and from $70\%$ to $75\%$ for another. A default rule may increase participation probabilities for some people while leaving others unchanged. A decision aid may draw some people into treatment and others out of treatment because it reduces mistakes on both sides of a threshold. In fact, when asked about the chances of choosing particular occupations (arcidiaconoetal2020), pregnancy and labour supply conditional on pregnancy status (briggsetal2024) or buying a service under different scenarios (blasslachmanski2010), respondents often provide probabilities strictly between $0\%$ and $100\%$. In these examples, the stable object is not necessarily a deterministic treatment status under each instrument or policy value. It may instead be that person's non-trivial probability of treatment under each state.
This paper considers that possibility. It starts from stochastic potential outcomes and stochastic potential treatment choices or allocations. For each individual $i$ and state $z$, let \[ p_i(z)=\Pr(D_i(z)=1\mid G_i) \] denote the individual's probability of treatment under state $z$, conditional on a stable response type $G_i$. The type also contains a potential outcome distribution \[ Q_i(\cdot\mid d,z) = \Pr(Y_i(d,z)\in \cdot \mid G_i), \] so the potential outcome under treatment $d$ and state $z$ need not be a fixed number. The corresponding mean potential outcome is \[ \mu_i(d,z)=\int y\,Q_i(\mathrm{d} y\mid d,z). \] We view $G_i \equiv (Q_i, p_i)$ as a person's stable “type” and explicitly condition the relevant probabilities above on $G_i$ to highlight this. The word stochastic refers to the fact that, even after conditioning on the type and fixing the state $z$, treatment receipt may remain random. Similarly, the outcomes conditional on types for fixed state $z$ and treatment value $d$ may remain random. In the present paper, the random type $G_i$ plays the role of the relevant background context (includes observed covariates), and $p_i(z')-p_i(z)$ is the type-specific probability change induced by a change in or contrast between states.
This approach has precedents in statistics and philosophy. In the statistics literature, stochastic counterfactual or potential outcomes generalize deterministic ones by allowing a given intervention to generate an individual-specific distribution of outcomes rather than a single fixed outcome greenland1987,robinsgreenland1989,robinsgreenland2000,vanderweelerobins2012. Dawid's decision-theoretic approach similarly emphasizes regimes or decision settings and the conditions under which information can be transported across them dawid2000,dawid2021. In philosophy, the probabilistic-causation literature starts from the idea that causes may raise or lower the probabilities of their effects rather than determine them hitchcock2023. Cartwright’s work connects directly to the probabilistic-causation tradition, in which a cause is understood as raising the probability of its effects. Her central qualification is that probability raising is causally informative only relative to a suitably homogeneous background: unless other causally relevant factors are held fixed, observed probability differences may conceal, offset, or even reverse the operation of a genuine cause. Causal claims therefore cannot be reduced to probabilistic associations alone but require an account of the causal capacities or mechanisms that generate them cartwright1979,cartwright1989. In this respect, the random-type formulation is close in spirit to probabilistic causation, but with an important qualification: the probability ($p_i(z)$) is not an observed association or a population-level probability-raising relation; it is a type-specific causal disposition, defined within a stable structural context. The state changes treatment probabilities because it activates or modifies a mechanism, not merely because it is statistically associated with treatment receipt. The framework has also been occasionally employed in econometrics (see, for example, abbringvandenberg2003).\footnote{See also related interesting work in chesher2021counterfactual. Their analysis does not focus on stochastic potential outcomes, but on potential “processes” or “sets of feasible counterfactual outcomes obtained by exogenously shifting” the treatment or choice variable while holding what we call the state fixed. }
For some parameters and estimation protocols, stochastic potential outcomes will lead to slight adaptations in interpretation. For instance, an Average Treatment Effect (ATE) can be defined in terms of differences in potential mean outcomes: \[ ATE = E[E\{Y_i(1)\mid G_i\} - E\{Y_i(0) \mid G_i\}] = E[\mu_i(1) - \mu_i(0)] \] where the unconditional expectations are taken across individuals and $\mu_i(d), d=0,1$ is the mean individual potential outcome under $d$ presented previously marginalized against the distribution of $Z$. In this case, under appropriate conditions the usual estimators for the ATE under deterministic potential outcomes (e.g., difference in treatment and controls means when $\mu_i(d) \perp \!\!\! \perp D_i$) will aim at the ATE as generalized above for potential outcomes.\footnote{hernanrobins2025 point out that whether potential outcomes are stochastic or deterministic may have repercussions to uncertainty quantification depending on whether a model- or design-based inferential paradigm is adopted (see Section 10.3).}
On the other hand, when the estimator is nonlinear in outcomes or treatments (as is typical, for example, in IV methods), the same estimation protocols will correspond to estimands with a more distinct interpretation under stochastic potential outcomes and stochastic treatment choice or assignment. For example, in a deterministic model with a binary treatment and a binary instrumental variable, the instrumental variable partitions individuals into compliers, defiers, always-takers and never-takers (angristimbensrubin1996). In the stochastic model, this is no longer the case: the instrument instead induces a distribution of probability changes across individuals. Those probability changes in turn determine weights in the estimand.
To elaborate on this, let $z$ and $z'$ denote two states-of-the-world. The states may correspond to an assignment rule, a price, a subsidy, a default rule, an information regime, or another policy environment. In a conventional IV application, the state is the instrumental variable; the outcome contrast given by $\mathbb{E}[Y_i\mid Z_i=z']-\mathbb{E}[Y_i\mid Z_i=z]$ is the reduced form; the corresponding treatment contrast given by $\mathbb{E}[D_i\mid Z_i=z']-\mathbb{E}[D_i\mid Z_i=z]$ is the first stage; and their ratio is the Wald estimand. We use the more general word state only because the same logic applies beyond binary instruments.
The key object is the individual responsiveness to a state contrast: \[ r_i(z',z)=p_i(z')-p_i(z). \] This is the stochastic analogue of a switching indicator. In the deterministic IV model, $D_i(z')-D_i(z)$ is equal to one for compliers, zero for never-takers and always-takers, and minus one for defiers. Under monotonicity, the latter group is assumed away. In the stochastic model, responsiveness can be fractional. An individual with $p_i(z')=0.60$ and $p_i(z)=0.40$ contributes 0.20 units of responsiveness to the contrast $(z',z)$. She is not naturally an always-taker, never-taker, complier or defier. She is partially responsive.
The object of interest remains a contrast in potential outcomes. Under an exclusion restriction (i.e., $\mu_i(d,z) = \mu_i(d)$ for any $i, d$ and $z$), define the individual mean treatment effect as: \[ \Delta_i=\mu_i(1)-\mu_i(0). \] Variation in treatment helps us learn about these $\Delta_i$'s through the individual treatment probability movements it creates. Under independence of the state $Z_i$ from the response type $G_i$ and under the exclusion restriction above, the (average) outcome movement generated by the contrast $(z',z)$ is: \[ \mathbb{E}[Y_i\mid Z_i=z']-\mathbb{E}[Y_i\mid Z_i=z] = \mathbb{E}[\Delta_i r_i(z',z)], \] while the (average) treatment movement generated by the contrast $(z',z)$ is: \[ \mathbb{E}[D_i\mid Z_i=z']-\mathbb{E}[D_i\mid Z_i=z] = \mathbb{E}[r_i(z',z)]. \] When the latter is nonzero, the outcome-treatment contrast ratio corresponds to the Wald estimand and identifies
Thus, the reduced form and the first stage reveal an average of treatment effects weighted by the pattern of treatment-probability movements generated by that particular source of variation.
This formula nests familiar cases. If $r_i(z',z)$ is constant across individuals, the contrast recovers the ATE. If $p_i(z),p_i(z')\in\{0,1\}$ and $p_i(z')\geq p_i(z)$, treatment choice is deterministic and monotonic, and hence $r_i(z',z)\in\{0,1\}$. Consequently, the contrast recovers the Local Average Treatment Effect (LATE) imbensangrist1994 for deterministic switchers. If $r_i(z',z)\geq 0$ but not necessarily zero or one, the contrast is a convex average weighted by probabilistic responsiveness. Perhaps more subtly, even if one interprets $r_i(z',z) \ge 0$ as a relaxation of the conventional monotonicity condition towards a stochastic monotonicity one and classifies those for whom $r_i(z',z) > 0$ as “stochastic compliers,” the estimand corresponds to a weighted average among those individuals with differing individual weights. If $r_i(z',z)$ changes sign, the ratio is not a convex average. It is a signed comparison between treatment effects among types whose treatment probabilities rise and treatment effects among types whose treatment probabilities fall.
The distinction we make between an ex ante stochastic representation and an ex post deterministic representation is not a claim about different sampling protocols for the observed data. As Appendix (ref) shows, in a static cross-section the two representations generate the same observable laws for \((Y_i,D_i,Z_i)\). The value of the stochastic formulation is interpretative: it takes \(z\mapsto p_i(z)\) as the stable ex ante object and uses changes in this map to reinterpret the treatment-effect parameters targeted by familiar econometric estimators. Hence, the main message is simple. Under stochastic potential outcomes and stochastic choice, an instrument or policy does not merely select compliers. It induces a weighting operator over treatment effects. To interpret a reduced form, first stage, or Wald ratio, one must ask whose treatment probabilities are changed, by how much, and in which direction.
To that end, the paper makes three contributions. First, it revisits the stochastic potential outcomes model and views it in terms of compliance and treatment response. Each individual has a treatment choice probability under each state and a potential outcome distribution under each treatment-state pair. Deterministic potential outcomes and deterministic compliance types are nested as degenerate cases. This modeling change does not by itself create new estimators; it changes the interpretation of familiar estimators.
Second, the paper characterizes what standard contrast estimators target under this stochastic formulation. The IV/Wald estimand is not, in general, the mean effect for a discrete complier group. It is a weighted average over individuals, with weights proportional to how much the instrument changes each person's probability of treatment. In the deterministic LATE model, only compliers get positive and equal weight. In the stochastic model, all individuals may enter the estimand, but with different weights. This also clarifies when the estimand is an average and when it is only a signed contrast.
Third, the paper gives a possible behavioral interpretation of the weights. In a threshold-choice model in which the state changes either the perceived treatment margin or the information used to form it, local responsiveness decomposes into a margin-density term and a mechanism-loading term. The margin-density term captures how close the person is to indifference; the mechanism-loading term captures how strongly the instrument or policy moves that person's adoption margin. This generalizes the familiar IV insight. With deterministic compliance, different instruments may identify different effects because they move different complier groups. With stochastic choice, there need not be discrete complier groups; instruments instead differ in the treatment-probability movements they induce across individuals.
For each individual $i$, we observe \[ O_i=(Y_i,D_i,Z_i,X_i), \] where $Y_i\in\mathbb{R}$ is the observed outcome, $D_i\in\{0,1\}$ is the observed treatment receipt, $Z_i\in\mathcal{Z}$ is the observed state, and $X_i\in\mathcal{X}$ is a vector of covariates.
The primitive object is a stable (stochastic) response type. For each individual $i$, let $G_i\in\mathcal{G}$ denote a latent type containing the following two kernels.
Note that the covariates $X_i$ are included in the definition of the type $G_i$. Also, the notation in the above definition is meant to indicate that this outcome kernel is an autonomous “production function” at the individual level and can be evaluated at any value of $(d,z)$ (i.e., it is a “causal object”). This means that the distribution $Q_i(\cdot\mid d,z)$ is not the observed conditional distribution of $Y_i$ given $D_i=d$ and $Z_i=z$. It is the type-specific distribution of the potential outcome under treatment value $d$ and state value $z$. The corresponding mean potential outcome under treatment value $d$ and state value $z$ is thus: \[ \mu_i(d,z) \equiv \int y\,Q_i(\mathrm{d} y\mid d,z). \] Similarly, the distribution $p_i(\cdot \mid z)$ is not the observed conditional distribution of $D_i$ given $Z_i=z$. It is the type-specific distribution of the potential treatment under state value $z$.
The observed variables are such that: \[ D_i=D_i(Z_i) \qquad \textrm{and} \qquad Y_i=Y_i(D_i(Z_i),Z_i). \]
When potential outcomes and treatments are deterministic, the objects \(D_i(\cdot)\) and \(Y_i(\cdot)\) are individual-specific functions. In the stochastic formulation, however, even after fixing the type \(G_i\) and the state \(z\), treatment receipt may remain random. Similarly, fixing $G_i$, $d$ and $z$, outcomes may remain random.
The type \(G_i\) has been defined through two marginal objects: the treatment choice kernel \(p_i(\cdot\mid z)\) and the potential outcome kernel \(Q_i(\cdot\mid d,z)\). These marginal kernels are the objects we want to use in the estimand. But by themselves they do not specify how the residual randomness in treatment choice is related to the residual randomness in potential outcomes within the same type. We therefore impose the following mean-separability condition.
Assumption (ref) is not an independence condition for the state or instrument. It is a condition on the stochastic representation within a fixed type. It says that, after conditioning on \(G_i\), the residual event that treatment is realized under state \(z\) does not change the mean of the potential outcome under \((d,z)\).\footnote{A sufficient condition is a residual-shock representation \(D_i(z)=D(z,G_i,U_i^D)\), \(Y_i(d,z)=Y(d,z,G_i,U_i^Y)\), with \(U_i^Y\perp \!\!\! \perp U_i^D\mid G_i\). This is stronger than needed: the assumption requires only mean independence of the residual outcome shock from the residual treatment choice event within type.} The assumption allows arbitrary selection across types. For example, types with larger treatment gains may also have larger treatment probabilities. What it rules out is residual mean selection within the same type.
Under Assumption (ref), the type-specific mean outcome under state \(z\) can be written as the simple mixture \[ m_i(z) = \sum_{d\in\{0,1\}}p_i(d\mid z)\mu_i(d,z) = \{1-p_i(z)\}\mu_i(0,z)+p_i(z)\mu_i(1,z). \] Thus \(m_i(z)\) is the mean outcome for type \(i\) when the state is set to \(z\), averaging over the stochastic treatment choice generated by that state. The role of Assumption (ref) is only to justify this mixture representation using the marginal kernels \(p_i(\cdot\mid z)\) and \(Q_i(\cdot\mid d,z)\).\footnote{If one instead defined the type more broadly as the full joint law of $\{D_i(z),Y_i(0,z),Y_i(1,z)\}_{z\in\mathcal Z},$ then one could take $m_i(z)=\mathbb{E}\{Y_i(D_i(z),z)\mid G_i\}$ as primitive, and the no-within selection assumption would not be needed. Similarly, in the deterministic ex post formulation, the condition is automatic: once the full type is fixed, there is no residual treatment choice or outcome randomness left.}
For two state values $z,z'\in\mathcal{Z}$, define the observed outcome and treatment contrasts as: \[ C_Y(z',z) \equiv \mathbb{E}[Y_i\mid Z_i=z']-\mathbb{E}[Y_i\mid Z_i=z] ~ \textrm{and} ~ C_D(z',z) \equiv \mathbb{E}[D_i\mid Z_i=z']-\mathbb{E}[D_i\mid Z_i=z]. \] Without assumptions, these are descriptive contrasts. To give them causal content, the state must be independent of the latent response type, either unconditionally or conditional on covariates.
First, we nonetheless start with a simple observation regarding the worst case bounds for the mean treatment effect. Then we state the independence assumption and derive the causal interpretation of the contrasts under that assumption.
The above bounds obviously hold without any assumptions (other than support conditions and Assumption (ref)). Now, we state the independence assumption that is key to obtaining causal interpretations of the contrasts introduced earlier.
This is the analogue of random assignment or instrument independence. We also require that, conditional on $G_i$ and, where applicable, $X_i$, assignment is independent of the residual draws that realize the treatment choice and potential outcome kernels. In potential outcome terminology, it implies exchangeability of the relevant potential state outcomes across values of $Z_i$. (Under deterministic potential variables, this reduces to independence from $\{D_i(z),Y_i(0,z),Y_i(1,z):z\in\mathcal Z\}$.)
The primitive causal object associated with changing the state from $z$ to $z'$ is the average state effect
This is the average effect of assigning state $z'$ rather than state $z$. At this level, the state may affect outcomes by changing treatment probabilities, by changing outcome distributions directly, or both. The next result shows how this average state effect relates to observables.
Under unconditional state independence,
The first equality is the law of iterated expectations. The second uses the definition of the type-specific mean outcome under state $z$. The third uses $Z_i\perp \!\!\! \perp G_i$, since $m_i(z)$ is a function of $G_i$. Thus \[ C_Y(z',z)=\Psi(z',z). \] No exclusion restriction is needed for this interpretation. A reduced form can be a causal effect of the state even when it is not a treatment effect.
The proposition says that the reduced form, or outcome intention-to-treat (ITT), is already a causal effect of the assigned state. If \(Z_i\) is an encouragement, information treatment, default rule, subsidy offer, or assignment rule, then \(C_Y(z',z)\) is the causal effect of that intervention. The exclusion restriction is needed only for the additional IV interpretation that this effect operates solely through treatment receipt \(D_i\).
The previous section showed that, under state independence, the observed outcome contrast \[ C_Y(z',z) = \mathbb{E}[Y_i\mid Z_i=z']-\mathbb{E}[Y_i\mid Z_i=z] \] identifies the causal ITT effect of assigning state \(z'\) rather than state \(z\): \[ C_Y(z',z) = \mathbb{E}[m_i(z')-m_i(z)]. \] This is already a causal effect of the assigned state \(Z_i\). For example, if \(Z_i\) is an encouragement, default rule, information treatment, or subsidy offer, then \(C_Y(z',z)\) is the ITT effect of that intervention.
One is perhaps more often interested in leveraging variation in the state to retrieve the direct causal effect of treatment receipt \(D_i\) on the outcome. Aside from the assumptions introduced previously, this requires an exclusion restriction.
This implies that \[ \mu_i(d,z)=\mu_i(d) \qquad d\in\{0,1\}. \] Under exclusion, define then the individual mean effect of treatment receipt by \[ \Delta_i=\mu_i(1)-\mu_i(0). \] Since \[ m_i(z) = \{1-p_i(z)\}\mu_i(0)+p_i(z)\mu_i(1), \] we can write \[ m_i(z) = \mu_i(0)+p_i(z)\Delta_i. \] Therefore, \[ m_i(z')-m_i(z) = \Delta_i\{p_i(z')-p_i(z)\}. \]
It follows that, under no within-sample selection, state independence and exclusion, the outcome ITT or reduced form is \[ C_Y(z',z) = \mathbb{E}[\Delta_i\{p_i(z')-p_i(z)\}], \] while we can similarly obtain that the treatment ITT, or first stage, is \[ C_D(z',z) = \mathbb{E}[D_i\mid Z_i=z']-\mathbb{E}[D_i\mid Z_i=z] = \mathbb{E}[p_i(z')-p_i(z)]. \] When the first stage is nonzero, the familiar outcome-treatment contrast ratio corresponds to the usual Wald-type object. The numerator is the ITT effect of the assigned state, written as a treatment-channel effect. The denominator is the first stage. Given the conditions above, it identifies
This is the basic formula behind the examples in this section. The stochastic-choice interpretation is that this ratio averages individual mean treatment effects using weights proportional to how much the state changes each individual's probability of receiving treatment.
Suppose $Z_i\in\{0,1\}$ is randomly assigned. This is a setting of a randomized experiment and hence unconditional state independence holds. If assignment mechanically determines treatment receipt, then \[ p_i(1)=1, \qquad p_i(0)=0 \] for all individuals. Hence $r_i(1,0)=1$, and under exclusion \[ \mathbb{E}[Y_i\mid Z_i=1]-\mathbb{E}[Y_i\mid Z_i=0]=\mathbb{E}[\Delta_i]=\mathrm{ATE}. \] As is well-known, the perfect-compliance randomized experiment identifies the ATE because the assignment changes every person's treatment probability by the same amount.
With imperfect individual compliance (i.e., stochastic potential outcomes and treatments), assignment changes treatment probabilities but need not determine treatment status for a given individual. Under no-within selection, state independence and exclusion, the contrast ratio is given by equation ((ref)) and therefore identifies the assignment-responsiveness-weighted average of treatment effects. Because the weights $W_i(1,0)=r_i(1,0)/E[r_i(1,0)]$ average to one, this weighted average decomposes as \[ \frac{C_Y(1,0)}{C_D(1,0)} = E[\Delta_i] + \operatorname{Cov}\bigl(\Delta_i, W_i(1,0)\bigr) = ATE + \frac{\operatorname{Cov}\bigl(\Delta_i,\, p_i(1)-p_i(0)\bigr)}{E[p_i(1)-p_i(0)]}. \] The estimand therefore equals the ATE if and only if treatment effects ($\Delta_i$) are uncorrelated with responsiveness ($p_i(1)-p_i(0))$: the individuals whom assignment moves most strongly into treatment must be, on average, neither higher- nor lower-gain than the population. When they are higher-gain, the estimand exceeds the ATE; when lower-gain, it falls short. Uniform responsiveness, $p_i(1)-p_i(0)$ constant across individuals, is the leading special case in which the covariance vanishes automatically.
The binary IV setting is the special case in which $Z_i\in\{z,z'\}$ is interpreted as an instrument. The Wald estimand is
Under instrument independence, exclusion, no-within selection and relevance (i.e., $\mathbb{E}[D_i\mid Z_i=z']-\mathbb{E}[D_i\mid Z_i=z] \ne 0$), \[ \beta^W = \frac{\mathbb{E}[\Delta_i\{p_i(z')-p_i(z)\}]} {\mathbb{E}[p_i(z')-p_i(z)]} = \mathbb{E}\!\left[ \Delta_i \frac{p_i(z')-p_i(z)} {\mathbb{E}[p_i(z')-p_i(z)]} \right]. \] Thus IV identifies an instrument-responsiveness-weighted treatment effect.
The deterministic compliance model is nested as a special case. If each individual has deterministic potential treatment choices $D_i(z'),D_i(z)\in\{0,1\}$, then \[ p_i(z')=D_i(z'), \qquad p_i(z)=D_i(z), \] and \[ r_i(z',z)=D_i(z')-D_i(z). \] If monotonicity holds, $D_i(z')\geq D_i(z)$ for all $i$, then \[ r_i(z',z)=\mathbf{1}\{D_i(z')>D_i(z)\}. \] Equation (ref) becomes \[ \beta^W=\mathbb{E}[\Delta_i\mid D_i(z')>D_i(z)], \] the standard Local Average Treatment Effect (LATE) for compliers.
Under stochastic potential treatments or choices, there need not be a discrete complier group. If $p_i(z')\geq p_i(z)$ for all $i$, the Wald estimand is still a convex average, but the weights are proportional to continuous responsiveness. Even if one interprets $p_i(z')\geq p_i(z)$ as a relaxation of the conventional monotonicity condition towards a stochastic monotonicity one and classifies those for whom $p_i(z')> p_i(z)$ as “stochastic compliers,” the estimand corresponds to a weighted average among those individuals with possibly unequal weights that are proportional to individual responsiveness $r_i(z',z)$. If $p_i(z')-p_i(z)$ takes both signs, the estimand is not an average. It is a signed contrast.
The binary Wald estimand is easy to interpret because the instrument defines a single comparison: state \(z'\) versus state \(z\). Most empirical IV applications are not quite so simple. The instrument may be multivalued or continuous, the researcher may use several excluded instruments, and the specification may include covariates. We consider here the well known two stage least squares estimand and interpret it in light of the stochastic potential outcome formulation we are considering.
Consider the linear IV specification with one endogenous treatment,
estimated using excluded instruments constructed from \(Z_i\). Equation ((ref)) is the estimating equation that defines the 2SLS procedure, not a generative model for the outcome: the residual $u_i$ is defined by the associated projection,\footnote{In other words, $u_i$ is such that $E(u_i)=0$, $E(u_i X_i)=0$ and $E(u_i L(D_i|B_i,X_i))=0$, where $L(D_i|B_i,X_i)$ is the linear projection of $D_i$ on $X_i$ and $B_i$ defined in what follows.} no restriction such as $\mathbb{E}[u_i\mid X_i,Z_i]=0$ is imposed, and $\beta$ is not assumed to be a common structural coefficient. We are interested in interpreting the population 2SLS functional (defined more precisely below) $\beta^{2SLS}=\mathbb{E}[D_i^{\mathrm{fs}}Y_i]/\mathbb{E}[D_i^{\mathrm{fs}}D_i]$ under Assumptions (ref)--(ref).
Let \[ B_i=b(Z_i,X_i) \] denote the vector of excluded instrument terms used in the first stage. For example, \(B_i\) may contain \(Z_i\) itself, dummies for values of a multivalued instrument, polynomials in a continuous instrument, or interactions between the instrument and covariates.
For any random variable \(V_i\), let
\[ \ddot V_i = V_i - L[V_i\mid X_i] \] denote the residual obtained after removing its linear projection on $X_i$, which we denote by $L[V_i\mid X_i]$ (here we require that all projections include a constant term). In particular,
\[ \ddot B_i = B_i - L[B_i\mid X_i], \qquad \ddot D_i = D_i - L[D_i\mid X_i], \quad \textrm{and} \quad \ddot Y_i = Y_i - L[Y_i\mid X_i]. \]
The (residualized) population first stage is \[ \ddot D_i = \ddot B_i^\top\pi_D + v_i, \qquad \pi_D = E[\ddot B_i\ddot B_i^\top]^{-1}E[\ddot B_i\ddot D_i], \] assuming the usual rank condition. Define \[ D_i^{\mathrm{fs}}=\ddot B_i^\top \pi_D, \] which is the part of treatment receipt predicted by the excluded instruments after partialling out the included controls. This is the population analogue of the usual Frisch--Waugh--Lovell residualization. With one endogenous treatment, even an overidentified 2SLS estimator can be viewed as using this single fitted first-stage component as the relevant IV variation. The population 2SLS coefficient is therefore \[ \beta^{2SLS} = \frac{E[D_i^{\mathrm{fs}}\ddot Y_i]} {E[D_i^{\mathrm{fs}}\ddot D_i]} = \frac{E[D_i^{\mathrm{fs}}Y_i]} {E[D_i^{\mathrm{fs}}D_i]}, \] where the second equality uses \(E[D_i^{\mathrm{fs}} L(D_i | X_i)]=0\) and \(E[D_i^{\mathrm{fs}} L(Y_i | X_i)]=0\) since, by definition, $E(\ddot V_i X_i)=0$ for any variable $V_i$.
Now impose the stochastic potential outcomes Assumptions (ref)-(ref) from the previous section: no within-type selection, (conditional) state independence, and exclusion. Then\footnote{As a reminder, and as stated in Section 2 above, $X_i$ is measurable with respect to $G_i$ (covariates are part of the type), so conditioning on $X_i$ and $G_i$ is the same as conditioning on $G_i$ alone.} \[ E[Y_i\mid G_i,X_i,Z_i] = \mu_i(0)+p_i(Z_i)\Delta_i, \quad \textrm{and} \quad E[D_i\mid G_i,X_i,Z_i] = p_i(Z_i), \] where again the individual mean treatment effect $\Delta_i$ is \[ \Delta_i=\mu_i(1)-\mu_i(0). \] Assume then that \[ L[B_i\mid X_i] = E[B_i \mid X_i]. \] In other words: the conditional expectation function of $B_i$ given $X_i$ is linear. This simplifies the expressions that follow and is reasonable if $X_i$ is discrete or a sufficiently rich function of observed covariates, for example. This implies that $E[\ddot B_i \mid X_i]=0$ and thus \[ E[D_i^{\mathrm{fs}}\mid X_i]=0. \] We then have that
and the baseline potential outcome term $E[D_i^{\mathrm{fs}} \mu_i(0)]$ drops out of the numerator in $\beta^{2SLS}$ above. The first equality above uses the Law of Iterated Expectations. The second equality obtains from $Q_i(\cdot\mid 0,z)=Q_i(\cdot\mid 0), \forall z \in \mathcal{Z}$ (Assumption (ref)) and thus $\mu_i(0)=\int y\,Q_i(\mathrm{d} y\mid 0)$ is a direct function of $G_i$. Finally, the third equality relies on \(Z_i\perp \!\!\! \perp G_i\mid X_i\) by Assumption (ref) and the last equality holds given that linear conditional expectation assumption above. Hence: \[ E[D_i^{\mathrm{fs}}Y_i] = E\!\left[D_i^{\mathrm{fs}}\{\mu_i(0)+p_i(Z_i)\Delta_i\}\right] = E\!\left[\Delta_i E\{D_i^{\mathrm{fs}}p_i(Z_i)\mid G_i,X_i\}\right]. \] Similarly, \[ E[D_i^{\mathrm{fs}}D_i] = E\!\left[E\{D_i^{\mathrm{fs}}p_i(Z_i)\mid G_i,X_i\}\right]. \]
Define \[ \kappa_i^{2SLS} \equiv E\{D_i^{\mathrm{fs}}p_i(Z_i)\mid G_i,X_i\}. \] This is the individual's contribution to the population first stage generated by the particular 2SLS specification. To interpret it, hold fixed the individual's type and covariates ($G_i$ and $X_i$), let the instrument vary according to its conditional distribution, and ask how the person's treatment probability covaries with the first-stage fitted value used by 2SLS.
The 2SLS estimand is therefore \[ \beta^{2SLS} = \frac{E[\Delta_i\kappa_i^{2SLS}]} {E[\kappa_i^{2SLS}]}. \] Equivalently, writing $ W_i^{2SLS} = \frac{\kappa_i^{2SLS}}{E[\kappa_i^{2SLS}]} $, we have \[ \beta^{2SLS}=E[\Delta_i W_i^{2SLS}]. \]
This expression nests the binary Wald case. Suppose there are no covariates and \(Z_i\in\{0,1\}\). If the first stage uses \(B_i=Z_i\), then \(D_i^{\mathrm{fs}}\) is proportional to \(Z_i-E[Z_i]\). Thus \[ \kappa_i^{2SLS} \propto \Pr(Z_i=1)\Pr(Z_i=0)\{p_i(1)-p_i(0)\}. \] The factor \(\Pr(Z_i=1)\Pr(Z_i=0)\) is common across individuals, so the normalized 2SLS weights are proportional to \[ p_i(1)-p_i(0), \] exactly as in the binary IV estimand in Section 4.2.
With covariates, the same binary-instrument specification gives \[ \kappa_i^{2SLS} \propto \Pr(Z_i=1\mid X_i)\Pr(Z_i=0\mid X_i)\{p_i(1)-p_i(0)\}. \] Thus covariate-adjusted 2SLS weights individuals not only by how much the instrument changes their treatment probability, but also by how much residual instrument variation is available in their covariate cell.
With a multivalued instrument, there is generally no single pairwise comparison. If \(Z_i\) is discrete (but not necessarily binary), then \[ \kappa_i^{2SLS} = \sum_{z\in\mathcal Z} D^{\mathrm{fs}}(z,X_i)\Pr(Z_i=z\mid X_i)p_i(z), \] where \(D^{\mathrm{fs}}(z,x)\) is the first-stage fitted component when the instrument state is \(z\) and covariates are \(x\). Since \[ \sum_{z\in\mathcal Z} D^{\mathrm{fs}}(z,X_i)\Pr(Z_i=z\mid X_i)=0, \] for any baseline state \(z_0\), \[ \kappa_i^{2SLS} = \sum_{z\in\mathcal Z} D^{\mathrm{fs}}(z,X_i)\Pr(Z_i=z\mid X_i) \{p_i(z)-p_i(z_0)\}. \] Therefore, 2SLS with a multivalued instrument is a weighted combination of the treatment-probability movements generated by the different instrument states. Using the instrument in levels, using a set of dummies, or interacting the instrument with covariates generally changes \(D_i^{\mathrm{fs}}\), and therefore changes the weights in the treatment-effect average.
The main implication is that, beyond the binary Wald case, instrument validity does not by itself determine the treatment-effect estimand\footnote{In the deterministic-compliance benchmark, sloczynski2020should makes a closely related point: once covariates are needed for identification, the standard non-interacted linear IV specification can place negative weights on conditional LATEs — and lose its causal interpretation — while the saturated, fully interacted specification of angristimbens1995 restores convex weights (see also kolesar2013 and blandholetal2026). Our $\kappa_i^{2SLS}$ is the stochastic potential outcomes counterpart of the object he studies, and moving $B_i$ from levels to a saturated set of interactions is precisely the lever that alters its sign.}. The estimand also depends on the first-stage specification: which functions of the instrument are used, how covariates are partialled out, and how multiple instruments are combined by 2SLS. This is the stochastic-choice analogue of the familiar LATE lesson that IV estimates are local to the margin moved by the instrument. Here the margin is not necessarily a discrete complier group even when there is (stochastic) monotonicity. It is the distribution of individual treatment-probability movements induced by the fitted first stage.
Consequently, two valid 2SLS specifications can identify different parameters. If specifications \(a\) and \(b\) induce weights \(W_i^a\) and \(W_i^b\), then \[ \beta^a-\beta^b = E[\Delta_i\{W_i^a-W_i^b\}] = \operatorname{Cov}\!\left(\Delta_i,W_i^a-W_i^b\right), \] because \(E[W_i^a]=E[W_i^b]=1\). As in the deterministic case, differences across valid IV estimates may therefore reflect different first-stage margins rather than failure of exclusion or independence.
The Wald ratio is a ratio of two causal effects of the state: the reduced form and the first stage. Under exclusion, this ratio can be written as \[ \theta(z',z) = \frac{E[\Delta_i r_i(z',z)]}{E[r_i(z',z)]}, \qquad r_i(z',z)=p_i(z')-p_i(z), \] where \(r_i(z',z)\) is the change in individual \(i\)'s probability of treatment induced by the state contrast.
The econometric question is not only whether this ratio is well defined. It is whether the ratio can be interpreted as an average of treatment effects. In the deterministic LATE model, monotonicity guarantees that. It rules out defiers, makes the first stage equal to the probability of switchers, and turns the Wald estimand into the average treatment effect for compliers. In the stochastic model, the switching indicator ($\mathbf{1}\{D_i(z')>D_i(z)\}$) is replaced by the probability change \(r_i(z',z)\). Hence the analogue of monotonicity is one-sided responsiveness: \[ r_i(z',z)\geq 0 \quad \text{for all } i, \] or, after relabeling the two states, \(r_i(z',z)\leq 0\) for all \(i\).
If responsiveness is one-sided and the first stage is non-zero, then \[ \theta(z',z) = E\!\left[ \Delta_i \frac{r_i(z',z)}{E[r_i(z',z)]} \right], \] with nonnegative weights ($r_i(z',z)/E[r_i(z',z)]$) that integrate to one. The IV ratio is therefore a convex responsiveness-weighted average of individual mean treatment effects. Individuals whose treatment probabilities move more receive more weight. Individuals whose treatment probabilities do not move receive zero weight.
This nests the standard cases. If \[ r_i(z',z)=c\neq 0 \] for all \(i\), then the ratio equals the ATE. When treatment choice is deterministic and monotonic, then \[ r_i(z',z)=D_i(z')-D_i(z) = \mathbf{1}\{D_i(z')>D_i(z)\}, \] and the ratio becomes the usual LATE: \[ \theta(z',z) = E[\Delta_i\mid D_i(z')>D_i(z)]. \] For stochastic, but one-sided, treatment choice, the same logic survives, except that the complier indicator is replaced by the amount of probabilistic responsiveness.
The difficulty arises when responsiveness changes sign. Suppose, for simplicity, that \(E[r_i(z',z)]>0\). Define \[ A^+ = E\!\left[r_i(z',z)\mathbf{1}\{r_i(z',z)>0\}\right], \qquad A^- = E\!\left[-r_i(z',z)\mathbf{1}\{r_i(z',z)<0\}\right], \] \[ \theta^+ = \frac{ E\!\left[\Delta_i r_i(z',z)\mathbf{1}\{r_i(z',z)>0\}\right] }{ E\!\left[r_i(z',z)\mathbf{1}\{r_i(z',z)>0\}\right] }, \quad \textrm{and} \quad \theta^- = \frac{ E\!\left[\Delta_i\{-r_i(z',z)\}\mathbf{1}\{r_i(z',z)<0\}\right] }{ E\!\left[\{-r_i(z',z)\}\mathbf{1}\{r_i(z',z)<0\}\right] }. \] Then \[ \theta(z',z) = \frac{A^+\theta^+-A^-\theta^-}{A^+-A^-}. \] Thus the Wald ratio compares treatment effects among types whose treatment probabilities rise with treatment effects among types whose treatment probabilities fall. The denominator is the net first stage, \(A^+-A^-(= E[r_i(z',z)])\), not the total amount of treatment-probability (absolute value) movement, \(A^++A^- (= E[|r_i(z',z)|])\). Importantly, when positive and negative movements partly offset, the ratio can lie outside the range of individual treatment effects.
The proposition clarifies the role of monotonicity. In the deterministic IV model, monotonicity rules out defiers. In the stochastic model, one-sided responsiveness rules out types whose treatment probabilities move in the opposite direction. Without this condition, the reduced form remains a causal effect of the assigned state, but the Wald ratio is a signed treatment-effect contrast rather than an average treatment effect.
The same point applies to more general IV estimands. For two-stage least squares with covariates or multiple excluded instruments, the relevant object is the individual first-stage contribution induced by the fitted first stage. The 2SLS estimand has an average-treatment-effect interpretation only when those induced first-stage contributions are one-sided. Otherwise, valid instruments may still identify signed contrasts rather than convex averages.
The preceding sections develop a stochastic potential choice and outcome model and interpret common estimands without specifying why treatment choice remains stochastic even within individual types. As noticed elsewhere, experimental evidence suggests that individual choices may vary even when confronted with the same decision scenario (see, e.g., strzalecki2025, p.9). This may reflect transient frictions and fluctuations in perceived utility, learning and experimentation, random consideration or attention, genuine implementation or decision errors or even deliberate randomization (see again strzalecki2025, pp.9-10). Given this, “[i]t will be realistic to assume that a man's choice from a given set of alternatives is not unique but obeys some probability distribution” marschak1960.
Formally, we model this by postulating an information set (or sigma-field) $\mathcal{F}_i$ for individual $i$ where \(M_i(z)\) denotes the systematic component of individual \(i\)'s perceived private net value of treatment under state \(z\), and \(\sigma_i(z)>0\) measures the dispersion of the part of the perception error that remains unresolved at the moment of choice under that state. Whereas both \(M_i(z)\) and \(\sigma_i(z)>0\) pertain to the information set $\mathcal{F}_i$, residual variation ascribed to a perception error $\varepsilon_i$ does not. The state $z$ here may represent policy, default rule, price, subsidy, information treatment, or another feature of the choice environment. The treatment margin determining choices is then given by the random index
where \(\varepsilon_i\) is a perception error with a symmetric, continuous distribution function \(F\), density \(f>0\), mean zero and unit variance, drawn once conditional on \(G_i\) and, in line with the discussion following Assumption (ref), independent of \(Z_i\) given \(G_i\). Under this normalization \(\sigma_i^2(z)=\operatorname{Var}\{\widetilde M_i(z)\mid G_i\}\). The functions \(M_i(\cdot)\) and \(\sigma_i(\cdot)\), together with \(F\), are stable objects and constitute a structural parameterization of the treatment choice kernel in \(G_i\); the realization of \(\varepsilon_i\) is not part of the type and is the residual randomness that makes choice stochastic (within type). Here, a more informative state, formalized as a smaller \(\sigma_i(z)\), means that a larger share of the index ((ref)) is ascribed to the deterministic component \(M_i(\cdot)\).\footnote{A multi-component analogue of the sequential information structure in heckmannavarro makes the resolution structure explicit. Let \(\mathbf u_i=(u_{i1},\ldots,u_{iK})\) be mean-zero variables affecting the choice drawn given \(G_i\), and let the state $z$ determine which components the individual learns before choosing. Denote by \(J_i(z)\subseteq\{1,\ldots,K\}\) the set of indices for those variables in individual $i$'s information set. The unresolved part of the choice index is \(\varepsilon_i(z)=\sum_{k\notin J_i(z)}u_{ik}\). The scalar representation (ref) is the marginal, location--scale summary of such constructions and is all that is used below: every estimand in Sections (ref)--(ref) depends on the type only through the marginal kernels \(p_i(\cdot\mid z)\), so the joint law of \(\{\widetilde M_i(z):z\in\mathcal{Z}\}\) across states matters for interpretation rather than identification (see also Appendix (ref)).} Consistency with Assumption (ref) (no within-type selection on residual outcome shocks) requires that the choice assessment error be conditionally mean-independent of residual outcome shocks: \[ \mathbb{E}\{Y_i(d,z)\mid G_i,\varepsilon_i\}=\mu_i(d,z). \]
If treatment choice or allocation follows a threshold-crossing rule with zero as the threshold (i.e., $D_i(z)=1$ if and only if $\widetilde M_i(z)\geq 0$), the treatment choice kernel is given by:
where we assume that \(F\) is continuous and corresponds to a symmetric distribution (around zero). We refer to the map \(t\mapsto F(t)\) as the individual choice curve and to \(t_i(z)\) as the standardized perceived margin. Everything the individual contributes to the estimands of the previous sections is her position on this curve under each state. Before the perception error is realized, \(p_i(z)\) is also the individual's own ex ante choice probability and corresponds to the object elicited in stated-probability data blasslachmanski2010,arcidiaconoetal2020,briggsetal2024. Deterministic compliance is nested as a limit: as \(\sigma_i(z)\to0\) and if \(\Pr(M_i(z)=0)=0\), the choice curve steepens to a step function, \(p_i(z)\to\mathbf{1}\{M_i(z)\ge 0\}\), and the deterministic taxonomy of Section (ref) re-emerges.\footnote{More generally, one may let the unresolved error have an arbitrary state-dependent conditional law \(F_{i,z}\) given \(G_i\), so that \(p_i(z)=1-F_{i,z}(-M_i(z))\) and, for a continuously indexed scalar state, \(\dot p_i(z)=f_{i,z}(-M_i(z))\,\dot M_i(z) -\partial_z F_{i,z}(u)\big|_{u=-M_i(z)}\). The location--scale family is the only case used in this section; it already spans the two channels through which a state can act on choice, by shifting the location of the perceived margin or by changing its scale.}
Suppose now that \(z\) is a continuously indexed scalar state and that \(M_i(z)\) and \(\sigma_i(z)\) are differentiable in \(z\). Differentiating (ref) gives a measure of local responsiveness:
Local responsiveness is a nonnegative factor times the difference between two terms. The factor \(f(t_i(z))/\sigma_i(z)\) is the density of the treatment margin in ((ref)) at zero. This term is large when the individual is close to indifference and the distribution $F$ is not only symmetric, but also single-peaked at zero. The first term, \(\dot M_i(z)\), is a location channel: the state shifts the systematic, deterministic component of the treatment margin, moving the individual along her choice curve. The second one, \(-t_i(z)\,\dot\sigma_i(z)\), is a precision channel: the state changes how much of the perception error contributes to the treatment allocation, steepening or flattening the curve. Because \(p_i(z)\) depends on \((M_i(z),\sigma_i(z))\) only through the ratio \(t_i(z)\), the two channels are two ways of moving the standardized margin: a location shift moves it additively, while a precision change rescales it multiplicatively. This distinction governs monotonicity below. Figure (ref) illustrates the two channels.
For a small change in \(z\), given the Wald estimand in (ref), the induced IV-type local parameter is (assuming the relevant limits and expectations may be interchanged): \[ \theta(z) = \frac{\mathbb{E}[\Delta_i\,\dot p_i(z)]}{\mathbb{E}[\dot p_i(z)]}, \] whenever the denominator is nonzero. Thus the local weights are proportional to \(\dot p_i(z)\), the individual-level change in treatment probability induced by the state. The two subsections that follow examine each channel in isolation.
First consider a state that shifts the systematic perceived margin while leaving its precision unchanged, so \(\dot\sigma_i(z)=0\). Suppose that \[ M_i(z)=M_i+\ell_i z, \qquad \ell_i\geq 0. \] Then (ref) reduces to \[ \dot p_i(z) = \ell_i\,\frac{f\big(t_i(z)\big)}{\sigma_i(z)} \;\geq\;0, \] and, provided \(\mathbb{E}[\ell_i f(t_i(z))/\sigma_i(z)]>0\), the local weights are
The weight has two components. When the distribution is symmetric and single-peaked at zero, the margin-density term \(f(t_i(z))/\sigma_i(z)\) is large when the individual is close to indifference based on their known, perceived margin (i.e., $t_i(z) \approx 0$). The loading \(\ell_i\) measures how strongly the state shifts that individual's perceived treatment margin.
For a subsidy or cost instrument $z$, \(\ell_i\) could be interpreted as material pass-through. For a reminder or salience intervention $z$, \(\ell_i\) could encode attention responsiveness. For a persuasion or framing instrument $z$, \(\ell_i\) could correspond to susceptibility to the message. The algebra is the same, but the economic interpretation of the loading differs.
If \(f\) is single-peaked at zero, then for fixed \(\ell_i\) and \(\sigma_i\) the unnormalized weight is largest when \(t_i(z)=0\), that is, at \(p_i(z)=1/2\). Mean-shifting states therefore place the greatest weight on individuals for whom the systematic component of the treatment margin is near the threshold. Note also that sign-homogeneous loadings deliver one-sided responsiveness (or monotonicity) by construction: a location shift translates every standardized margin in the same direction, \(\dot t_i(z)=\ell_i/\sigma_i(z)\geq 0\), so \(\dot p_i(z)\geq 0\) for all \(i\) and the local analogue of the stochastic monotonicity condition of Section (ref) holds.
\paragraph{Relation to Marginal Treatment Effects (MTE).} The density term \(f(t_i(z))/\sigma_i(z)\) in (ref) is the within-type analogue of the object that drives marginal treatment effects in the latent-index model of HeckmanVytlacil2005. There, the choice rule is given by \(D_i(Z_i)=\mathbf{1}\{\mu_D(Z_i) \geq U_{Di}\}\), where $U_D$ can be interpreted a “resistance to treatment” term counteracting the state-dependent index \(\mu_D(Z_i)\). In the Heckman--Vytlacil model the resistance \(U_D\) is a fixed individual scalar type (thus belonging to $\mathcal{F}_i$) and the index \(\mu_D(\cdot)\) is homogeneous to the population: the instrument moves a single threshold faced by everyone, the individuals at the margin form the fixed subpopulation \(\{U_D=\mu_D(z)\}\) and monotonicity follows because the common threshold moves in one direction for all. This is the configuration under which the equivalence result established in vytlacil2002 holds, and it corresponds here to a location channel with a common index and loading. In terms of the presentation above, this means that $\sigma_i(z)=0$ for every $z$ (since both $\mu_D(z)$ and $U_D$ are in the agent's information set), $U_{Di}=M_i \in \mathcal{F}_i$ and $\ell_i z = \ell z = \mu_D(z) \in \mathcal{F}_i$ for every $i$. In the stochastic model the standardized margin is instead individual-specific: the same state moves \(t_i(z)\), the standardized loadings \(\ell_i/\sigma_i(z)\), by different amounts and possibly in different directions. Moreover, the realization of \(\varepsilon_i\) is not part of the type as is the realized $M_i \in \mathcal{F}_i$, so the marginal event at \(z\) is a within-type probabilistic event carrying density \(f(t_i(z))/\sigma_i(z)\), rather than a fixed subgroup indexed by an individual (realized) scalar resistance. Two consequences follow. First, effect heterogeneity is indexed by the full type \(G_i\), not by a single resistance value \(U_{Di}\), so there need not be a real function (i.e., the MTE given by $E(Y_1-Y_0|U_D=u)$) whose (weighted) average corresponds to conventional treatment effect parameters or estimands (e.g., ATE, OLS, IV); see HeckmanUrzuaVytlacil2006. Second, the precision term \(-t_i(z)\,\dot\sigma_i(z)\) in (ref) interacts the state with the scale of the decision or perception error and therefore falls outside the additively separable latent-index class. It is precisely this channel that delivers the failures of monotonicity documented next.
Another possibility is for states to change the precision of the treatment margin rather than its mean. This can encode simplification, clarification, disclosure, counseling, or decision aids. Such interventions refine the information available at the moment of choice reducing the inherent stochasticity: they decrease the role ascribed to the perception error \(\varepsilon_i\), so the dispersion of its unresolved component falls. Hold the systematic margin fixed with respect to the state, i.e. \(M_i(z)=M_i\), let $z$ be continuous and suppose that \[ \dot\sigma_i(z)<0. \] Then (ref) reduces to
so the sign of \(\dot p_i(z)\) is the sign of \(M_i\). Equivalently, a refinement rescales the standardized margin away from zero, \(t_i(z')=t_i(z)\cdot\{\sigma_i(z)/\sigma_i(z')\}>t_i(z)\), with the factor $\sigma_i(z)/\sigma_i(z')$ exceeding one when $z<z'$: it pushes each individual further in the direction she already leans. The choice curve steepens around its midpoint: take-up rises for individuals whose systematic margin favors treatment (\(M_i>0\), \(p_i(z)>1/2\)) and falls for individuals whose systematic margin disfavors it (\(M_i<0\), \(p_i(z)<1/2\)).
The failure of monotonicity is exactly what one should expect from better information. If some individuals mistakenly opt out and others mistakenly opt in, an information refinement induces entry on one side of the threshold and exit on the other. Proposition (ref) therefore does not describe a defect in the intervention. Two-sided responsiveness is the economic content of a policy that reduces mistakes on both sides of the choice threshold. In the language of Section (ref): stochastic monotonicity is the property that a state translates every standardized margin in a common direction; a state that instead rescales the margin generically violates it.
Without monotonicity, the Wald ratio need not be an average treatment effect. The reduced form, however, remains directly meaningful. Under exclusion, \[ m_i(z)=\mu_i(0)+p_i(z)\Delta_i. \] Assuming differentiation and expectation can be interchanged, Proposition (ref) implies \[ \frac{d}{dz}\mathbb{E}[m_i(z)] = \mathbb{E}[\Delta_i\,\dot p_i(z)]. \] When \(Y_i\) is the outcome relevant to a planner, this derivative is the effect of refining information on the planner's outcome. Substituting (ref) gives \[ \frac{d}{dz}\mathbb{E}[m_i(z)] = \mathbb{E}\!\left[ \underbrace{\frac{f\big(t_i(z)\big)\,\lvert\dot\sigma_i(z)\rvert} {\sigma_i^2(z)}}_{\geq 0} \;\Delta_i M_i \right]. \] Hence the effect is nonnegative whenever \(\Delta_i M_i\geq 0\) almost surely; more generally, its sign depends on the weighted average alignment between the perceived treatment margin \(M_i\) and the true treatment gain \(\Delta_i\).
As \(\sigma_i(z)\to0\), for every \(M_i\neq0\), \[ p_i(z)\longrightarrow\mathbf{1}\{M_i>0\}. \] Thus, if \(\Pr(M_i=0)=0\), the choice curve converges to a step function and choice conforms almost surely to the systematic perceived margin. An information refinement pulls into treatment individuals with \(M_i>0\) and pushes out individuals with \(M_i<0\): it makes behavior increasingly obey perceptions. This raises the planner's outcome only to the extent that perceptions are aligned with treatment gains. If \(M_i>0\) but \(\Delta_i<0\) for some individuals, more precise information moves them more confidently toward an action that lowers the planner's outcome. This can occur because \(M_i\) is a biased perception of the individual's own payoff, or because the individual's private payoff legitimately differs from the measured outcome \(Y_i\), in which case the refinement is privately optimal yet lowers \(Y_i\). Accordingly, the welfare value of a decision aid depends on the substantive alignment between the margins it sharpens and the gains relevant to the planner.
This paper reinterprets familiar causal estimands under stochastic potential outcomes and stochastic treatment choice. Each individual carries a stable response type, made up of a treatment choice kernel and an outcome kernel so that, conditional on the type, both treatment and outcomes may still be random. Under state independence, an observed outcome contrast identifies the causal effect of the assigned state. Adding exclusion turns that state effect into a contrast in treatment effects, weighted by how much the state moves each individual's probability of treatment. The standard compliance taxonomy is then a limiting case. The deterministic IV model obtains when the treatment choice probabilities are degenerate and one-sided; in that case responsiveness is binary; more generally, the complier indicator is replaced by the probability change $p_i(z') - p_i(z)$, denoted as {\it responsiveness} in the paper. Uniform responsiveness gives the ATE, one-sided responsiveness gives a convex average, and sign-changing responsiveness gives only a signed contrast. The same logic carries over to 2SLS, where the weights and therefore the estimand depend on the first-stage specification, not on instrument validity alone. The lesson for empirical work is that standard estimators remain useful, but what they identify is governed by the choice process narrative behind the first stage. An IV or policy contrast recovers a particular responsiveness-weighted average, and which average depends on whose treatment probabilities move, by how much, and in which direction.
A few directions seem worth pursuing in future work. Because the framework treats $z \mapsto p_i(z)$ as its stable primitive, it invites estimation using elicited or subjective choice probabilities: when these can be measured directly, one may recover the weights—and potentially the treatment effect schedule itself rather than only the single average a given instrument reports (see the aforementioned arcidiaconoetal2020,briggsetal2024,blasslachmanski2010). Second, since the equivalence between the ex ante and ex post representations holds only in a static cross-section, extending the analysis to dynamic and panel settings, where repeated choices and serially correlated shocks break it, is a natural next step.