Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
97,550 characters · 19 sections · 56 citation commands
Effect Identification and Unit Categorization in the Multi-Score Regression Discontinuity Design with Application to LED Manufacturing
\abstract{ RDD (Regression discontinuity design) is a widely used framework for identifying and estimating causal effects at the cutoff of a single running variable. In practice, however, decision-making often involves multiple thresholds and criteria, especially in production systems. Standard MRD (multi-score RDD) methods address this complexity by reducing the problem to a one-dimensional design. This simplification allows existing approaches to be used to identify and estimate causal effects, but it can introduce non-compliance by misclassifying units relative to the original cutoff rules. We develop theoretical tools to detect and reduce “fuzziness” when estimating the cutoff effect for units that comply with individual subrules of a multi-rule system. In particular, we propose a formal definition and categorization of unit behavior types under multi-dimensional cutoff rules, extending standard classifications of compliers, alwaystakers, and nevertakers, and incorporating defiers and indecisive units. We further identify conditions under which cutoff effects for compliers can be estimated in multiple dimensions, and establish when identification remains valid after excluding nevertakers and alwaystakers. In addition, we examine how decomposing complex Boolean cutoff rules (such as AND- and OR-type rules) into simpler components affects the classification of units into behavioral types and improves estimation by making it possible to identify and remove non-compliant units more accurately. We validate our framework using both semi-synthetic simulations calibrated to production data and real-world data from opto-electronic semiconductor manufacturing. The empirical results demonstrate that our approach has practical value in refining production policies and reduces estimation variance. This underscores the usefulness of the MRD framework in manufacturing contexts. }
Keywords: Causal Inference, Machine Learning, Regression Discontinuity, Causal Effect Identification, Multi-Score RDD, Rework, Production Optimization, Phosphor-converted white LEDs
\doublespacing
In many areas of economics and management, decision-making depends on multiple criteria. Examples include targeted marketing, education policies, and political advertising. The complexity of such settings has motivated the development of formal approaches such as multi-criteria decision analysis, which make trade-offs explicit and offer methods for comparing and ranking alternatives. Once such decision rules are applied in practice, a further challenge is to assess their causal impact. The focus then shifts from choosing among options to evaluating the consequences of those choices. Causal inference, and in particular regression discontinuity design (RDD), provides tools for this purpose.
RDD is a well-established quasi-experimental approach for identifying causal effects whenever treatment assignment changes discontinuously at a threshold in an observed running variable. It has been applied in diverse fields, including economics Hartmann2011, Card2015, Flammer2015, Calvo2019, public policy Lee2008 and the social sciences Angrist1999. The appeal of RDD lies in its ability to produce credible causal estimates without relying on the assumptions that are required for most causal inference approaches. In particular, it does not require unconfoundedness Rubin1974 or global positivity Austin2011, but instead depends on continuity of potential outcomes and the absence of precise manipulation around the cutoff.
The classic RDD exploits a discontinuity in a single treatment assignment rule that depends on an observed running or score variable. When correctly specified, this setting enables average treatment effects to be identified for units near the cutoff. Under mild conditions, it further allows counterfactual reasoning about the decision rule in a small neighborhood around the boundary Cattaneo2019, Imbens2008, Lee2010.
As noted above, treatment decisions in many real-world contexts are based on multiple criteria, particularly in industrial, operational, and engineering systems Sabaei2015. The natural extension of the classical RDD is multi-score RDD (MRD). In MRD, treatment depends on a joint rule requiring that multiple score variables satisfy specified threshold conditions simultaneously. The idea was first proposed by papay2011. Influential surveys that develop estimators and examine them empirically include Reardon2012, wong2013 and porter2017. MRD has been widely applied, especially in test-score-based settings an2024 and geographical applications keele2015. However, strategies for identifying MRD effects when multiple cutoffs are combined through arbitrary Boolean rules are still underdeveloped in the literature.
This paper addresses the identification and estimation challenges that arise in multi-score RDD settings. We develop new results tailored to these contexts. In particular, we introduce a formal framework that defines and categorizes behavioral types (“compliers”, “alwaystakers”, “nevertakers”, “defiers”) and analyzes how these types transform when treatment assignment rules change. We investigate general Boolean cutoff mechanisms, such as “AND-type” and “OR-type” rules, which are commonly observed in MRD applications, and analyze how their properties influence local identification. Further, we demonstrate how the complier effect can be identified using these categories. Finally, we illustrate our results with data from a light-emitting diode (LED) production facility, in which each batch is evaluated based on several metrics, such as mean color point measurements or the proportion of units within a lot that meet specified quality requirements.
Our application of the proposed strategies to data from the LED manufacturing industry demonstrates the usefulness of the framework in a real-world setting where treatment assignment depends on multiple quality indicators. We combine the framework with recently developed estimators that use machine learning adjustments to further reduce variance in estimation noack2024. Estimating the causal effect of threshold-based production decisions provides guidance for tuning decision thresholds to improve manufacturing outcomes. In practice, managers need to understand how marginal adjustments to these thresholds affect outcomes such as yield rates and thus overall production efficiency. We complement the empirical analysis with a simulated study based on a semi-synthetic replica of the production environment, highlighting the value of multi-score RDD for counterfactual analysis and manufacturing policy design.
This study contributes to the operations research literature by linking advances in causal machine learning with practical decision-making in manufacturing systems. Our theoretical results on effect identification in MRD create new opportunities for its application in decision-making contexts. By accommodating complex, multi-criteria assignment mechanisms, the framework reflects the realities of operational decision-making practice and provides tools for more informed and effective policy design.
The paper is organized as follows. Section 2 reviews related literature on RDD and its extensions to multi-score settings. Section 3 presents our theoretical framework for identification under general Boolean assignment rules. Section 4 describes the empirical application in opto-electronic semiconductor manufacturing. Section 5 reports results from synthetic and semi-synthetic environments. Section 6 discusses findings from the real-data application, emphasizing the implications for threshold calibration and production efficiency. Section 7 concludes.
RDD dates back to a study on scholarship programs by thistlethwaite1960. In the following decades, the approach gained popularity in the social sciences, with notable applications by Angrist1999 and Black1999. Building on this early empirical work, Hahn2001 provided the first formalized treatment of RDD within the potential outcomes framework and established non-parametric identification and estimation results for both sharp and fuzzy designs, firmly establishing RDD as a core method in modern econometrics.
Important surveys, including those by Imbens2008, Lee2010 and Cattaneo2019, synthesize guidance on estimation choices in RDD, such as bandwidth selection, polynomial order, and diagnostic checks. More recently, several theoretical advances have refined the method. Calonico2014 developed bias-corrected point estimators and robust confidence intervals that improve coverage accuracy and have become the empirical default for RDD inference. calonico2019 examined the inclusion of covariates in local regression to reduce variance, while Calonico2019b analyzed optimal bandwidth selection for robust estimators. Further extensions incorporate machine-learning-based covariate adjustments to further reduce variance, integrating RDD into the broader toolbox of causal machine learning Kreiss2022, Arai2025, noack2024.
The idea of extending RDD to multi-score settings dates back to papay2011, who showed that causal effects can be identified along a two-dimensional cutoff frontier when program eligibility depends jointly on multiple scores. Reardon2012 systematically collected and formalized MRD identification strategies for cases in which running variables jointly cross their cutoffs to form a multidimensional frontier, corresponding to an “AND”-type rule. They introduced frontier-projection and distance-based estimators, and highlighted the steep data-density and bandwidth challenges that arise as the dimensionality of the assignment space grows. wong2013 advanced this line of work with extensive simulation studies, demonstrating that estimators that weight observations along the entire cutoff frontier achieve lower bias and higher efficiency than naïve one-dimension-at-a-time approaches. Their findings underscore the importance of exploiting the joint assignment rule in practice. In another large-scale study, porter2017 reached similar conclusions and provided practical recommendations for MRD applications. A related special case, geographic RDD, was introduced by keele2015 to evaluate policies along irregular geographic borders.
More recent research has focused on theoretical advances. Choi2018 allowed for partial treatment effects when only some running variables cross their cutoffs and introduced thin-plate spline estimators to flexibly recover heterogeneous impacts along a high-dimensional frontier. Building upon earlier work by Imbens2009, Choi2023 analyzed two-dimensional MRD under limited compliance and examined the consequences for estimating complier effects in fuzzy MRD.
Research has also produced more sophisticated MRD estimators. Imbens2019 proposed a minimax approach, liu2024 introduced decision-tree methods, and Sawada2025 developed fully multivariate local-polynomial estimators. Cattaneo2025 examined the statistical properties of important MRD estimators in detail and provided practical recommendations.
RDD has additionally been applied in a variety of business and management contexts. Hnermund2021 studied its use in business decision-making, while Ho2017, Mithas2022 examined applications in production and operations management. Calvo2019 used a one-dimensional fuzzy RDD to estimate the effect of procurement officers on budget overruns and delays in public infrastructure projects. Leung2020 analyzed the effects of labor unionization on the operating performance of supplier firms. In the automobile sector, Hu2021 employed RDD to assess the impact of regulatory changes on compliance. To the best of our knowledge, however, neither RDD nor MRD has yet been directly applied to production policies.
This paper contributes to the literature in two main ways. First, we derive new theoretical results on unit categorization and effect identification in MRD, which are particularly useful in settings with complex, multi-dimensional decision rules. Secondly, we apply these results to real data, demonstrating the applicability of MRD in operations policy-making and providing evidence to guide actual production policy.
In this section, we first review the basic elements of the RDD and its extension to the multi-score case. We then introduce a formal definition of common unit types (e.g., compliers, defiers) in multi-score, two-stage decision settings that employ cutoff rules for treatment assignment. Using this categorization, we derive an identification result and show that, under certain assumptions, identification does not depend on subsets of unit types with constant response. Finally, we derive rules describing how unit classifications change as assignment rules become more complex.
RDD builds on the groundwork of Hahn2001, which provides a formal justification for effect identification and estimation. An RDD consists of three main elements: a score $X$ that measures a characteristic of the units, a cutoff $c$ that splits the support of the score into two groups, and a treatment $D$ that is assigned to each unit depending on whether its score lies above or below the cutoff, $D_i = \operatorname{\mathbbm{1}}[X_i\ge c]$ Cattaneo2019. \\
Typically, the parameter of interest is the average treatment effect at the cutoff $c$, $\mathbb{E}[Y(1)-Y(0)\,|\,X=c]$, where $Y(1)$ and $Y(0)$ denote the potential outcomes under treatment and control, respectively rubin2005. Identification of this effect relies on the following key assumption:
This assumption implies that units with scores close to but on opposite sides of the cutoff are comparable in all relevant aspects Hahn2001. To rule out selection on gains, it is further required that units near the threshold cannot perfectly manipulate their scores. Formally, for sufficiently small \(\epsilon > 0\), $\mathbb{E}[Y(1)-Y(0)\mid D,\, X=x] = \mathbb{E}[Y(1)-Y(0)\mid X=x]$ must hold for $x\in (c-\varepsilon, c+\varepsilon)$ chernozhukov2024.
Under these relatively mild assumptions, RDD provides inference around the threshold that is as credible as inference from a randomized experiment Lee2008. The average treatment effect at the cutoff,
can be identified as
Hahn2001.
While standard RDD is suitable for analyzing the effect of a single score variable at a cutoff, real-world treatment decisions often depend on multiple score variables. For example, treatment \(T_i = 1\) may be assigned to unit \(i\) only if both criteria \(X_{1,i} > c_1\) and \(X_{2,i} > c_2\) are met. Such multi-score cases are often reduced to a standard RDD by transforming the variables into a single composite score, allowing identification and estimation to be justified by results from the one-dimensional case. A drawback of this approach, however, is the reduced interpretability of the transformed score. This raises the question of the precise conditions under which a treatment effects exist and can be identified in a multi-score setting. In particular, a well-defined multidimensional cutoff effect requires some independence of the directions from which the cutoff is approached.
In applications, the aim is usually to optimize the treatment assignment \(T_i\) in order to improve the outcome \(Y_i\). In some critical systems, however, only gradual adjustments to \(T_i\) are feasible. In such cases, part of the assignment rule \(T_i\) may be treated as fixed; for example \(H_i \coloneqq \operatorname{\mathbbm{1}}[X_{2,i} > c_2]\) is accepted as given, while the remaining part, that is \(G_i \coloneqq \operatorname{\mathbbm{1}}[X_{1,i} > c_1]\), is subject to optimization. The effect of a subrule \(G\) at its cutoff can therefore provide insights into how the overall assignment rule \(T\) might be improved. In some settings, the assignment rule \(T\) serves only as a recommendation, and the actual treatment \(D\) may deviate from it. In the RDD literature, this case is referred to as “fuzzy”. Altogether, this leads to a hierarchy of decision rules \(G \to T \to D\). Each step in the hierarchy obscures the assessment of \(G\) (for example, evaluation of the optimal value of cutoff \(c_1\)) because at every level of it a unit may become non-compliant. Knowledge of the unit's compliance type with respect to \((G,\, D)\), however allows these ambiguities to be addressed. In practice, this is usually handled using a fuzzy RDD estimator, which adjusts the sharp effect estimate by incorporating an estimate of the treatment probability \(\Pr(D_i=1 \,|\, G_i)\), thereby recovering a complier effect.
In contrast, we categorize unit behavior with respect to the rules \((G,\, D)\) into the standard compliance types: compliers, defiers, nevertakers, alwaystakers, and an additional category of indecisive units. Based on this categorization, we derive assumptions for identifying the complier effect with respect to \((G,\,D)\). We then examine conditions under which non-change units (e.g., nevertakers) can be excluded without affecting identification and thus avoiding some of the ambiguities. For example, units with \(X_{2,i} \leq c_2\) always have \(T_i = 0\) regardless of \(G_i\) and thus may not contribute to the complier effect of \(G\). In some cases, this subsetting principle justifies the use of sharp RDD estimators, which are typically more stable than their fuzzy counterparts. Finally, we show how the unit classifications change when moving on the hierarchy from simple to more complex assignment rules, incorporating knowledge of intermediate categorizations.
We use the notation \(\mathbb{R}^K_{> 0} \coloneqq \left\{x \in \mathbb{R}^K \,\middle|\, \forall \, 0 \leq j \leq K \,:\, x_j > 0 \right\}\) for vectors with positive components and adapt it in the canonical way for the cross product of other ordered sets and relations (e.g., \(\mathbb{R}^K_{< y}\) or \(\mathbb{N}_{\leq k}\)). Let \(e_1, \ldots, e_K \in \mathbb{R}^K\) denote the standard basis vectors, and let \(\left\langle \cdot , \, \cdot \right\rangle\) the standard inner product in \(\mathbb{R}^K\). For linear spaces \(U\) and \(V\) over \(\mathbb{R}\) with \(U \cap V = \{0\}\), we write \(U \oplus V\) for their direct sum.
For each individual \(i\), let \(X_i = (X_{1,i}, \, \ldots, \, X_{K,i}) \in \mathbb{R}^K\) denote the vector of score variables. Given \(c \in \mathbb{R}^K\), let \(I_i(c) \coloneqq (I_{1,i}(c_1), \, \ldots, \, I_{K,i}(c_K))\) define the indicator vector, where \(I_{k,i}(c_k) \coloneqq \operatorname{\mathbbm{1}}[X_{k,i} > c_k]\). Let the observed outcome for individual \(i\) be \(Y_i\). We treat the entries of \(I_i(c)\) as Boolean variables and allow any composition of AND, OR, and negation (\(\land, \, \lor, \overline{\square}\)) over this set of atoms to form general Boolean functions \(g(I_i(c))\).
With slight abuse of notation, we use \(T(c)\), \(T(X_i \,|\, c)\) and \(T_i(c)\) to indicate the use of a specific cutoff \(c \in \mathbb{R}^K\). Note that
holds for \(\epsilon \in \mathbb{R}^K, \lambda \in \mathbb{R}_{>0}\). Thus, without loss of generality, we assume the cutoff of interest is \(c = 0\). We suppose that the outcome \(Y_i\) depends in the following way on a cutoff rule \(T\) and a general decision rule \(D\):
where \(T\) is the treatment assignment and \(D\) is the implemented treatment. Unless otherwise stated, we impose no further assumptions on \(D\), except that it is a decision rule.
Given this setup, certain groups of individuals are of particular interest: nevertakers, alwaystakers, compliers and defiers with respect to the pair \((T,\, D)\) or to \((G,\, D)\), where \(G\) is a subrule of \(T\). A meaningful categorization of unit \(i\)'s requires counterfactual reasoning: we must consider how hypothetical changes in the cutoff (or, equivalently, in the observed score values) would have changed the responses \(T_i\) and \(D_i\).
Figure (ref) illustrates this idea for \(D \coloneqq I_1 \land I_2\) and \(T \coloneqq I_1\), and shows that only changes \(c \in \mathbb{R}^2\) that affect \(T\) are relevant for categorizing the behavior of a unit \(i\) with respect to \((T, \, D)\). Additional variables that affect only \(D\) may not be controllable or even observable. The following definition formalizes the intuition behind relevant directions of change in a cutoff rule.
This technical definition greatly simplifies the treatment of equivalent cutoff rules for example when \(T \coloneqq I_1 \land (I_2 \lor \overline{I_2})\) and \(G \coloneqq I_1\). It also makes it possible to compare \(T\) and \(D\) on their common domain. The notion of the support of a cutoff rule is motivated by the existence of changes along coordinate directions. It follows almost immediately that if the support is empty, then no changes are possible along linear combinations of the coordinate directions. This property justifies the definition.
From this point on, we assume that \(T\) is non-degenerate, i.e., \(\operatorname{supp}(T) \neq \{0\}\). In the introductory AND-rule example, with \(D \coloneqq I_1(c) \land I_2(c)\) and \(T \coloneqq I_1(c)\), we obtain \(\operatorname{supp}(T) = \mathbb{R} \times \{ 0 \}\) and \(\operatorname{supp}(D) = \mathbb{R}^2\). If a unit \(i\) complies with the assigned treatment \(T\) relative to the actual treatment \(D\), then \(T_i\) should coincide with \(D_i\) even under hypothetical changes of the cutoff \(c_1\), as illustrated in Figure (ref). A more refined, local definition could weaken this requirement to “any reasonable changes” (e.g. perturbations in a neighborhood of the cutoff), but for clarity we adopt the global one. In other words, both rules should agree on the support of \(T\), which in this case occurs if \(X_{2,i} > c_2\). Thus, it is \(X_{2,i}\) and \(c_2\) that govern the behavior of unit \(i\). To make this distinction explicit, it is useful to introduce notation for entries \(X \in \mathbb{R}^K\) that do not affect \(T\): \[N^T \coloneqq \left\{ X \in \mathbb{R}^K \,\middle|\, T(X \,|\, c) = T(0\,|\,c) \, \mbox{for all} \, c \in \mathbb{R}^K \right\}\] In particular, one can show the following.
Thus, according to Proposition (ref), the score space decomposes as \[ \mathbb{R}^{|S(T)|} \times \mathbb{R}^{K-|S(T)|} \simeq\footnote{Both linear spaces are isomorphic} \operatorname{supp}(T) \, \oplus N^T = \mathbb{R}^K. \] The first component captures the behavior of \(T\), whereas the second component consists of free variables that do not affect \(T\), but may influence the unit category. This observation motivates the following general definition of the unit categories.
In the special case where \(D\) is a cutoff rule, these definitions admit the following intuitive equivalencies, which capture the notion of simultaneous cutoff changes along the relevant directions.
Definition (ref) indeed introduces well-defined categories, which can be summarized as follows:
In general, these categories are not exhaustive, since \(D\) may not be constant nor equal to \(T\) or \(\overline{T}\) on the support of \(T\). We call the remaining category indecisive and denote the set of indecisive units (of \(T\) with respect to \(D\)) by \(\operatorname{Ind}(T,\, D)\). A unit \(i\) is indecisive iff \(c \mapsto D(X_i^{\perp T}-c)\) is not constant on \(\operatorname{supp}(T)\) and there exist \(c,\, \hat{c} \in \operatorname{supp}(T)\) such that \(T(X_i \,|\, c) = D(X_i^{\perp T} - c) \) and \(T(X_i \,|\, \hat{c}) \neq D(X_i^{\perp T} - \hat{c})\). Figure (ref) visualizes the general case in which \(D\) is not a cutoff rule.
From this point forward, whenever we assume that \(D\) is a cutoff rule, we restrict attention to the cases in which \(D\) contains at least as much information for decision-making as \(T\). This means that \(D\) might depend on \(I_{k,i}(0)\) for \(k \in S(T)\). In addition, we suppose that \(D\) does not depend on \(I_{k,i}(c)\) with \(c \neq 0\) for \(k \in \operatorname{supp}(T)\), effectively restricting attention to rules with zero cutoff. This restriction narrows the general case. For example, suppose the decision-maker \(D\) behaves opportunistically, ignoring all but one score \(X_k\), only complying with \(T\) when \(X_k\) exceeds some higher cutoff \(0 < c_k < X_k\).
Even for cutoff rules \(D\) under these restrictions, indecisive units are possible. For example, let \(D \coloneqq (\overline{I}_1 \land I_2)\) and \(T \coloneqq I_1 \land I_2\). Then \(\operatorname{supp}(T) = \mathbb{R}^2 = \operatorname{supp}(D)\) and for \(T(0 \,|\, 0) = D(0 \,|\, 0)\) but \(0 = T(0 \,| (-1, 1)) \neq D(0 \,|\, (-1, 1)) = 1\). Thus, \(D\) is not constant and differs from \(T\) and \(\overline{T}\) on the support of \(T\). Given \(T\) with \(\dim(\operatorname{supp}(T)) > 1\), one can always construct a cutoff rule \(D\) that produces indecisive items. At least one can prove the following:
Moreover, our definitions of alwaystakers, nevertakers and compliers imply the corresponding definitions in Imbens2008. For this, let \(c^+, c^- \in \operatorname{supp}(T)\) be two directions that induce change in \(T\), so that: \[ \lim_{\lambda \to 0} T(X_i \,|\, \lambda c^+) = 1 \, \mbox{ and } \, \lim_{\lambda \to 0} T(X_i \,|\, \lambda c^-) = 0 \] Using the above complier definition \[ T(X_i \,|\, c) = T(X_i^T \,|\, c) = T(0 \,|\, c - X_i^T) = D(X_i^{\perp T} + X_i^T - c) = D(X_i - c) \] for \(c \in \operatorname{supp}(T)\). In particular, \(T\) and \(D\) coincide in a neighborhood of zero, and thus \[ \lim_{\lambda \to 0} D_i(\lambda c^+) = \lim_{\lambda \to 0} D(X_i - \lambda c^+) = 1 \, \mbox{ and } \, \lim_{\lambda \to 0} D_i(\lambda c^-) = \lim_{\lambda \to 0} D(X_i - \lambda c^-) = 0 \] where \(D_i(c) \coloneqq D(X_i - c)\) for \(c \in \operatorname{supp}(T)\). The corresponding statements for nevertakers and alwaystakers follow analogously. By requiring consistency in the limit for any direction, one arrives at a local definition of unit categories that is sufficient for the multi-score RDD setting. Before turning to the identification results, we expand on the AND-rule example.
\paragraph{Example (AND-Rules)} Let \(D \coloneqq \bigwedge_{j=1}^K I_{j}\) and \(T \coloneqq \bigwedge_{j=1}^k I_{j} \) for \(k \in \mathbb{N}_{< K}\). Then \(\operatorname{supp}(T) = \mathbb{R}^k \times \{0\}^{K-k}\), and the potential outcomes reduce to \[Y_i = Y_i(0, 0)(1-T_i) + Y_i(1, 0)(1-D_i)T_i + Y(1, 1)D_i\], dropping one of the mixed terms. Additionally, the unit categorizations are as follows:
Further examples can be found in Appendix (ref), which introduces additional instances of the remaining unit categories (excluding the indecisive case). We conclude that the definitions introduced above are reasonable.
Inspired by the work of Hahn2001 and Imbens1994, we use the unit categories introduced above to prove an identification theorem for the local complier effect at the cutoff. Throughout this section, we require that the outcome \(Y\) does not directly depend on the treatment assignment \(T\). Let the set of all unit categories be \[ \mathcal{C} \coloneqq \{ \operatorname{ComP}(T,\,D),\, \operatorname{Nt}(T,\,D),\, \operatorname{At}(T,\,D),\, \operatorname{DeF}(T,\,D),\, \operatorname{Ind}(T,\,D) \} \] and define the set of non-change categories as \[\mathcal{C}^0 \coloneqq \{\operatorname{Nt}(T,\,D), \, \operatorname{At}(T,\,D) \}.\] We assume that the categorization of a unit is independent of the support part of \(T\) in a neighborhood of the cutoff, that is:
This assumption relates to the independence assumptions in Imbens1994. Using Assumption (ref) and further assuming \(\Pr\left(i \in \operatorname{Cat} \,\middle|\, X^T_i = 0\right) > 0\) we obtain \[ \mathbb{E}\left(Y_i \,\middle|\, X_i^T = x\right) = \sum_{\operatorname{Cat} \in \mathcal{C}} \mathbb{E}\left(Y_i \,\middle|\, X_i^T = x, \, i \in \operatorname{Cat}\right) \Pr\left(i \in \operatorname{Cat} \,\middle|\, X_i^T = 0 \right) \] and thus
We also rely on the following local continuity assumption, which adapts the standard continuity condition (Assumption (ref)) to the unit categories:
This assumption can be weakened by requiring only directional continuity, in which case the effect would depend on the chosen directions. Note that \[ E\left(Y_i \,\middle|\, X_i^T = x^{\pm},\, i \in \operatorname{At}\right) = E\left(Y_i(1) \,\middle|\, X_i^T = x^{\pm}, \, i \in \operatorname{At}\right) \] and \[ E\left(Y_i \,\middle|\, X_i^T = x^{\pm},\, i \in \operatorname{Nt}\right) = E\left(Y_i(0) \,\middle|\, X_i^T = x^{\pm}, \, i \in \operatorname{Nt}\right) \] holds for the non-change unit categories. Together with Assumption (ref) this yields:
Two further assumptions are needed to make use of the continuity condition for the remaining categories. First, we rule out the existence of indecisive units, since this category does not allow structured conclusions about \(D\) based on knowledge of \(T\). In other words, this category does not allow separating the potential outcomes \(Y_i(0)\) and \(Y_i(1)\).
Second, we assume that the directions \(x^+, \, x^- \in \operatorname{supp}(T)\) along which we estimate the complier effect induce a change in \(T\).
This assumption is implicit in one-dimensional RDD designs and imposes no substantive restriction in practice, as \(T\) is typically known. With this in place, we know how \(D_i\) behaves for compliers and defiers when approaching from the \(x^+\) and \(x^-\) directions. That is: \[ \lim_{\lambda \to 0} \Pr\left(D_i=1 \,|\, X_i^T = \lambda x^+,\, i \in \operatorname{ComP}\right) = 1 \,\mbox{ and }\, \lim_{\lambda \to 0} \Pr\left(D_i=1 \,|\, X_i^T = \lambda x^-,\, i \in \operatorname{ComP}\right) = 0, \] as well as \[ \lim_{\lambda \to 0} \Pr\left(D_i=1 \,|\, X_i^T = \lambda x^+,\, i \in \operatorname{DeF}\right) = 0 \,\mbox{ and }\, \lim_{\lambda \to 0} \Pr\left(D_i=1 \,|\, X_i^T = \lambda x^-,\, i \in \operatorname{DeF}\right) = 1. \] Thus, we can apply Assumption (ref) to these two remaining categories on the right-hand side of Equation (ref) as well:
This identification result has two immediate implications. First, one is free to choose among the directions \(x^+\) and \(x^-\) satisfying Assumption (ref). Second, the proof suggests that dropping subsets \(\Omega \subset \operatorname{Nt} \cup \operatorname{At}\) does not affect identification, as long as doing so does not violate Assumptions (ref) and (ref).
Following this idea, we now investigate how excluding units in \(\operatorname{At} \cup \operatorname{Nt}\) affects the above identification result. For ease of presentation, we assume there are no defiers at the cutoff. As a consequence, the correction term \(C\) in Equation (ref) equals zero. Now let \(\Omega \subset \operatorname{At} \cup \operatorname{Nt}\). Then \(\Pr\left(i \in \operatorname{ComP}, \, i \in \Omega \,\middle|\, X_i^T = 0\right) = 0\) and thus one has
for the denominator in Equation (ref). For the numerator, we have:
Requiring Assumptions (ref) and (ref) to hold when conditioning on \(\Omega \cap \operatorname{Nt}\) and \(\Omega \cap \operatorname{At}\) (instead of \(\operatorname{Nt}\) and \(\operatorname{At}\)), we obtain:
Since
holds, we require an assumption similar to Assumption (ref) for \(\Pr\left(i \in \Omega \,\middle|\, X_i^T = \lambda x^{\pm} \right)\), in order to obtain
using Equation (ref). Combining Equations (ref) and (ref), we obtain the following result:
We refer to estimates of the complier effect of \(T\) obtained when removing \(\Omega\) as the subset complier effect of \(T\) (excluding \(\Omega\)). Theorem (ref) provides sufficient conditions when both effects (the complier effect as identified in Theorem (ref) and the subset complier effect) are equal.
In this section, we investigate how the categorization of units changes when the treatment assignment \(T\) is altered, e.g., by replacing it with a subrule \(G\). The goal is to identify unit behavior without relying on knowledge of unobservable parts of the decision rules. In particular, we relate unit behavior under complex rules to that under simpler subrules. We begin by considering the effect of negation:
Indecisive units remain unchanged under negation of either decision rule. Non-change units are stable under negation of \(T\) but flip when \(D\) is negated. In contrast, compliers and defiers always flip. Next, we investigate the stability of the non-change categories under sub-rules. In general, one has the following bounds:
To obtain similar results for the other categories, assumptions about the relation between \(G\) and \(T\) are required. The following proposition generalizes the AND-rule and OR-rule examples (Example (ref) and Example (ref)).
In this setting, if the subrule \(H\) is known, we can identify the unit categories with respect to \((G,\, T)\), even if \(G\) itself is not fully observed. Using Proposition (ref), we conclude that introducing new score variables into cutoff rules does not create any defiers or indecisive units: by definition, negation is required for these categories (see Definition (ref) and Proposition (ref)). The previous statement can be extended to the case of a general decision rule \(D\) setting as follows:
Thus, compliers of \(T \coloneqq G \, \square \, H\) with respect to \(D\) are bounded by the complier and non-change categories of the simpler rule \(G\) as stated in Proposition (ref) (ref) and (ref) (ref). However, the set equalities do not hold in general. For example, let \(D \coloneqq (I_1 \land I_2) \lor I_3\), \(T \coloneqq I_1 \land I_3\) and \(G \coloneqq I_1\). Then \[ \emptyset = \operatorname{ComP}(T,\, D) \neq \operatorname{ComP}(G,\, D) = \{ i \,|\, X_{3,i} \leq 0, \, X_{2,i} > 0 \} \] holds. Furthermore, Proposition (ref) (ref), (ref) (ref) and (ref) (ref), (ref) (ref) resemble a factorization rule: loosely speaking, the intermediate rule \(T\) in the hierarchy \((G,\, T)\) and \((T,\, D)\) factors out, becoming \((G,\, D)\). Using Propositions (ref) and (ref), we can immediately draw conclusions about defiers:
Combining both statements, we see that if a unit is a complier with respect to \((T,\, D)\), its categorization under \((G,\, T)\) also holds for \((G,\, D)\). Figure (ref) summarizes the above findings for the AND-case. Further, according to Proposition (ref), in the absence of indecisive items, these transitions are exhaustive, since any unit is either a complier or nevertaker of \(G\) with respect to \(T\).
In this section, we apply our theoretical results to a real-world decision-making scenario in opto-electronic semiconductor manufacturing. We analyze the effect of rework decisions during the color-conversion process in the production of white light emitting diode (wLED) on overall product yield. After a brief description of the application, we derive modeling implications based on the observed data. This leads to an algorithmic description of the production process, which serves as a data-generating process (DGP) and complements the empirical study with simulations using artificially generated data that mimic the real-world case.
Phosphor conversion is a crucial step in wLED production. To obtain white light, several layers of phosphorous substrate are applied to a blue light-emitting semiconductor, shifting the perceived color along a conversion curve from the blue region towards the white region. In this step, each production lot, consisting of $784$ individual wLEDs, is processed according to a standardized recipe. Every wLED in the lot undergoes the same procedure, resulting in an identical stack of phosphor layers. The goal is to reach the specified target color by the end of production. To achieve this, intermediate color measurements are taken for selected wLEDs in the lot after the conversion step to assess proximity to the target. Subsequent processing steps are then adjusted to maximize the number of wLEDs that meet the target color, provided they are already close. To ensure a high yield (i.e., a large share of wLEDs in the lot reaching the target), a rework decision is made based on these intermediate measurements. If the target is not met, a correction layer of phosphor is applied to the entire lot. For further details on the conversion process, see Cho2017 and schwarz2024.
The rework decision is based on two scores: The distance score, $X_D$, measures the distance between the mean color point \(C_1 = (C_x,\, C_y)\) after the regular conversion and the target color \(C_T := (C_{T, x},\, C_{T, y})\). The yield-improvement score, $X_Y$, is a relative measure of quality variation within a lot. It evaluates a hypothetical scenario in which the target is ideally met by the mean color of the lot. If variability within the lot is high, rework may reduce overall yield even if the distance score suggests treatment. Figure (ref) illustrates the decision criteria in more detail.
Treatment is assigned according to the cutoff rule \(T = I_D \land I_Y\), with the goal of maximizing the outcome $Y$, defined as the percentage of chips that reach the target color by the end of production.
In practice, however, the empirical data show that the operators responsible for performing the treatment $D$ do not always comply with $T$ (see Figure (ref)). We attribute this discrepancy to an informational advantage regarding the \(X_Y\) score, which human operators may use to override $T$ in order to avoid possible yield losses from applying another rework layer.
We formalize this cautious-operator assumption in the following section. It reflects the observed one-sided fuzziness in the \(I_Y\) dimension, as well as the strict compliance with the distance rule \(I_D\).
From a policy or management perspective, we are interested in assessing the validity of the decision rule $T$ in this setting. To this end, we apply the subset-identification framework from Section (ref) to estimate effects at the cutoff. By evaluating these effects, we draw conclusions for improving the treatment assignment $T$.
In this section, we derive modeling assumptions on the decision rule $D$ for the application setting, combining insights from the observed data and the theory. We assume that $D$ has an informational advantage over the initial decision rule \(T = I_D \land I_Y\). Specifically, we posit that this advantage arises from having more detailed information about the yield-improvement score.
The initial treatment assignment \(T\) is based only on the improvement estimates of every \(m\)-th item in the production lot, resulting in the score \(X_Y\). In contrast, the final decision-maker has access to an overall yield-improvement estimate \(X_E = X_Y + X_R\), where \(X_R\) reflects the contribution of items not included in \(X_Y\). See Algorithm (ref) for details.
Implementing this algorithm allows us to benchmark different operator assumptions and compare them with the real-world case. In semiconductor manufacturing, the term operator refers to the human decision-maker implementing \(D\). Figure (ref) shows an example of data generated by this DGP.
Although the knowledgeable operator has, in principle, access to a better improvement estimate, there are several ways in which this information may be applied. In the following, we discuss special cases of the knowledgeable-operator assumption, each of which (at least in theory) permits identification of the unit categories.
\paragraph{The acknowledging operator} In this scenario, the operator \(D\) accepts the treatment assignment prescribed by \(T\) despite having better information. Thus, the final treatment assignment coincides with the initial decision rule: \(D = T = I_D \land I_Y\). This corresponds to the operator-specific policy in Algorithm (ref) with \(D_O = I_Y\). Consequently, we obtain \[ \operatorname{ComP}(T, D) = \mathcal{I} \] as well as \[ \operatorname{ComP}(G, D) = \{i \,|\, X_{Y,i} > 0 \} \,\mbox{ and }\, \operatorname{Nt}(G, D) = \{i \,|\, X_{Y,i} \leq 0 \} \] for \(G = I_D\), using Proposition (ref) or Example (ref).
\paragraph{The cautious operator} Suppose the operator is particularly careful to avoid accidental degradation of the production lot caused by the optional rework step. In cases where \(T\) recommends rework, the operator may override this recommendation. Such overrides may be motivated either by better information or by additional constraints (e.g., time) that the intended treatment rule \(T\) does not capture. In contrast, if \(T\) does not recommend rework, we assume the operator accepts this decision (e.g., out of concern about possible degradation).
We formalize this “negative overwrite” by introducing an additional score variable \(X_{op}\) and modeling \(D\) as: \[ D = (T \land I_{op}) \lor (\overline{T} \land 0) = T \land I_{op}. \] In particular, we set \(X_{op} = X_E\), which leads to \[ D_O(X \,|\, c) \coloneqq I_{X_Y} \land I_{X_E}. \] Thus, we have \[ \operatorname{ComP}(T, D) = \{i \,|\, X_{E,i} > 0 \} \,\mbox{ and }\, \operatorname{Nt}(T, D) = \{i \,|\, X_{E,i} \leq 0 \} \] as well as \[ \operatorname{ComP}(G, D) = \{i \,|\, X_{E,i} > 0, \, X_{Y,i} > 0 \} \,\mbox{ and }\, \operatorname{Nt}(G, D) = \{i \,|\, X_{E,i} \leq 0 \} \cup \{i \,|\, X_{Y,i} \leq 0 \} \] for the subrule \(G \coloneqq I_D\), again using Proposition (ref) or Example (ref).
Employing Theorem (ref), one can estimate the subset complier effect of \(G\), excluding the never-taker group \(\Omega \coloneqq \{i \,|\, X_{Y,i} \leq 0\}\), instead of the overall complier effect of \(G\). Since we assume that \(X_E\) is known only to the operator, nevertakers defined by \(X_{E,i} \leq 0\) cannot be excluded. However, because the condition \(X_{Y,i} \leq 0\) applies globally, the continuity and stability assumptions required for identification are likely to hold in practice.
\paragraph{The reasonable operator} Finally, consider an operator who always uses the additional information available. In this case, the reasonable operator overwrites the intended assignment rule \(T\) and bases the final decision solely on the full information \(X_E\) rather than on \(X_Y\): \[ D_O(X \,|\, c) = I_E(c). \] Accordingly, the final decision rule \(D\) can be expressed in relation to \(T\) as: \[ D(X \,|\, c) = I_D(c) \land \operatorname{\mathbbm{1}}[X_Y + X_R > c_Y]. \]
A unit \(i\) belongs to \(\operatorname{ComP}(T, D)\) if and only if \(X_{R,i} = 0\), i.e., there is no improvement or degradation among the items excluded from \(X_Y\). Now suppose \(X_{R,i} \neq 0\). Then we can choose \(c_Y = 2 \max(X_{Y,i}, X_{R,i})\) and \(c_D < X_{D,i}\) to obtain \(D(X_i \,|\, c) = T(X_i \,|\, c) = 0\). However, there also exist cutoff values \(c \in \operatorname{supp}(T) = \mathbb{R}^2\) for which \(D(X_i \,|\, c) \neq T(X_i \,|\, c)\). For units with \(X_R \geq 0\), choose \(c_Y = X_Y\). Then \(X_Y + X_R > c_Y\), so \(T(X_i \,|\, c) = 0\) and \(D(X_i \,|\, c) = 1\). Otherwise, choose \(c = X_Y + X_R\). Then \(c < X_Y\), so \(T(X_i \,|\, c) = 1\) and \(D(X_i \,|\, c) = 0\).
This shows that \[ \operatorname{ComP}(T, D) = \{ i \,|\, X_{R,i} = 0 \}, \] and that all other units fall into the indecisive category.
This model has several shortcomings. First, it does not reflect the observed data: according to the model, some treated items should not have received the initial assignment. Second, the identification result depends on the absence of indecisive items. In practice, small deviations in the improvement score (small \(X_R\)) would likely be considered compliant with the assignment \(T\), which suggests the need for a more local definition of categories.
While this case is theoretically interesting (and useful as an example of the indecisive case), our empirical investigation focuses on the two edge cases: the acknowledging operator and the cautious operator, as shown in Figure (ref).
In this section, we present numerical results based on the semi-synthetic data-generating process (DGP) derived in the previous section. The aim is to estimate the causal effect of rework decisions on the final yield in a neighborhood of the decision boundary. To this end, we use different estimators at both decision thresholds separately and draw on the identification theorems from Section (ref).
We generate semi-synthetic data following Algorithm (ref). A Python implementation of the DGP and all estimators is publicly available. The process is calibrated to match the characteristics of real production data. We draw \(n = 10{,}000\) observations and repeat each experiment \(r = 250\) times.
We evaluate the cut-offs \(c_D\) and \(c_Y\) separately, estimating:
Complier effects are estimated under a fuzzy design, whereas ITT effects are estimated under a sharp design. Oracle values are obtained using local linear kernel regression on the differences in true potential outcomes. Covariates consist of statistics describing the quality of individual items.
We consider a variety of estimators, including covariate-adjusted estimators, for complier and subset complier effects. Table (ref) provides an overview of the estimators used in our analysis.
The basic RDD estimator runs separate local linear regressions on each side of the cutoff:
where the \(w_i(h)\) are local linear regression weights that depend on the data through the realizations of the running variable only, and \(h > 0\) is a bandwidth.
Under standard conditions (e.g., the running variable is continuously distributed and the bandwidth \(h\) tends to zero at an appropriate rate), the estimator \(\hat{\tau}_{\text{base}}(h)\) is approximately normally distributed in large samples, with bias of order \(h^2\) and variance of order \((nh)^{-1}\).
A conventional extension is the covariate-adjusted estimator, which incorporates covariates linearly into the local regressions. We also use modern RDD estimators with flexible covariate adjustment based on potentially nonlinear adjustment functions \(\eta\). This estimator takes the form:
where \(\eta\) denotes the influence of \(Z\) on the outcome \(Y\), estimated using machine learning methods.
Figure (ref) presents side-by-side results for the estimators at \(c_D\) and \(c_Y\) in the case of the cautious operator. The intent-to-treat (ITT) oracles are closer to zero than the complier effects because they include individuals who are nevertakers with respect to each cutoff rule. As shown in Figure (ref), for \(I_D\) we estimate an overall negative effect, although it is not statistically significant at the 95% level.
The subset effect for the fuzzy case exhibits a smaller bias, since the proportion of nevertakers in the estimation sample is lower. The estimated effect remains unchanged because only nevertakers, but no compliers, were removed. For the ITT estimator, a higher proportion of compliers in the subsample increases the estimated effect of treatment rule \(G\). The subset estimators have a comparable variance (see Table (ref)), with coverage appearing slightly more credible overall. In general, covariate adjustment reduces standard errors, especially for the sharp estimators.
The estimates for \(I_Y\) are small and positive. The fuzzy estimator on the full data has a high standard error, which increases further with ML adjustment. This may be due to a small jump in treatment probability in the full data, destabilizing the ML estimate. In contrast, the subset estimator along this axis removes more observations within the estimation bandwidth, thereby reducing variance. Additional results can be found in Appendix (ref).
\paragraph{Two-dimensional Estimators} We also consider two-dimensional estimators proposed in previous literature. We compare the “binding-the-score” approach, which normalizes and combines \(X_D\) and \(X_Y\) into a single score, and the “Euclidean distance” approach, which computes the shortest distance of each observation to a predetermined point (here, a point on \(c_D\)).
The results in Figure (ref) show oracles very close to zero, indicating that when collapsing both dimensions of the decision rule into one, it is no longer possible to obtain insights into the individual rules. Interestingly, the estimator based on the Euclidean distance exhibits a significant bias across all methods.
In our specific setting, there is no straightforward interpretation of the RDD effect estimate at a two-dimensional threshold, which contrasts with previous applications where score components are more directly comparable, such as test scores or geographical distances.
We now compare the results from the previous section to those for the “acknowledging” operator, who accepts the decision of the system without using additional domain knowledge. In this setting, the data follow a sharp two-dimensional MRD. The comparison of the effects at \(c_Y\) and \(c_D\) is shown in Figure (ref).
For \(c_D\), the estimated effects are smaller than under the cautious operator. Lots that the cautious operator withheld from rework due to concerns about degradation are reworked in this case. This leads to a larger overall negative effect and suggests that the threshold should be moved even further than in the cautious operator case. The sharp subset estimator yields the same effect as the fuzzy estimator on the full data, because subsetting according to \(I_Y\) produces a sharp design in which all individuals comply with threshold movements. Across estimators, the subset versions display slightly smaller bias and standard errors. In contrast, the sharp estimator on the full sample, which corresponds to an ITT estimator, produces an effect closer to zero because it includes nevertakers.
At the cutoff \(c_Y\), the operator’s behavior has little impact on effect estimates. This can be explained by the fact that relatively few individuals in the neighborhood of the cutoff are affected.
This section presents results from estimation on the real data, which consist of \(n=9{,}103\) observations from the production system. In addition to the covariates used above, we include shop-floor workload. Our primary interest in the real system is evaluating the rule \(I_D\).
As shown in Table (ref), we estimate a positive effect. The subset estimates display lower variance because uncertainty from nevertakers (according to \(I_Y\)) is removed. For the ITT estimators, ML adjustment improves variance, whereas for the fuzzy estimators no clear improvement is observed.
The comparison in Figure (ref) highlights differences between the semi-synthetic and real data. The semi-synthetic results (panel a) suggest that the rework cutoff is miscalibrated, leading to negative estimated effects, whereas the real data (panel b) show a positive jump at the current cutoff, consistent with a beneficial rework threshold.
\paragraph{Validation Check.} We validate the application results using a pseudo-cutoff test (see Cattaneo2019). We compare the estimated ITT effect at the true cutoff \(c_D\) with estimates at a hypothetical lower and higher cutoff.
As shown in Table (ref), both pseudo cutoffs produce effects close to zero with wide confidence intervals, whereas the effect at the true cutoff is nearly significant at the 95% level.
This paper improves the applicability of regression discontinuity design (RDD) in operations and management. The multi-score RDD framework, combined with modern machine-learning-based adjustment, provides new opportunities to evaluate complex decision settings and policies. In settings with strict decision boundaries, policy-makers face little to no overlap in treatments near the cutoff, limiting the scope for “what if” reasoning.
Our theoretical insights rest on three observations. First, analyzing general assignment rules requires extending the notion of “binary compliance” to the broader concept of “compliance with an assignment mechanism.” In principle, all possible inputs to the mechanism must be considered when assessing compliance. Second, compliance is always defined relative to the rules under consideration. Changing the assignment mechanism can change the observed behavior of units. This relativity should be reflected in theory, as knowledge about specific parts of the mechanism or the final decision rule makes it possible to identify noncompliance. Third, in multiple dimensions, the complier effect at the cutoff should not depend on the direction from which the cutoff is approached, a property that future estimation methods can exploit.
Building on these observations, we generalized existing unit behavior types (compliers, nevertakers, alwaystakers) to multi-dimensional cutoff rules and introduced a new category of indecisive units. We derived rules describing how behavior types evolve when cutoff rules are decomposed into subrules. We also proved an identification result for the cutoff complier effect in MRD settings and established conditions under which this result remains valid when subsets of alwaystakers and nevertakers are removed, leading to the subset complier effect at the cutoff.
We validated our theoretical results using simulation studies with a semi-synthetic DGP and a real-world application in manufacturing. These analyses show that subrule-specific estimation can improve decision-making. In particular, the results confirm that removing identifiable nevertakers and/or alwaystakers makes effect estimation more precise, under conditions specified by the theory. When combined with flexible ML-based adjustments, variance can be further reduced without introducing bias. We recommend that practitioners estimate both the complier and the subset complier effect at the cutoff whenever nevertakers or alwaystakers can be identified in practice, and retrospectively validate the required assumptions. Importantly, the empirical findings underline asymmetries in treatment effects across different cutoff dimensions: the effect of the overall assignment rule generally differs from the effects of its subrules. Nevertheless, careful evaluation of each score component can provide valuable insights.
We expanded MRD theory by introducing tools to analyze complex cutoff rules and their corresponding subrules, expressed in terms of well-defined unit behavior types. Although our empirical study specifically bridges econometric RDD methods and industrial process control, the framework is broadly applicable to other complex decision settings, such as multi-criteria loan approval in financial institutions or public health resource allocation where eligibility depends on several risk factors. Future work could explore the relationship between main rules and subrules in greater depth, study the stability of unit types, or extend the multi-dimensional RDD to hierarchical decision structures across multiple stages. Developing estimators tailored to the identification theorems also represents a promising direction.