Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
30,005 characters · 11 sections · 8 citation commands
Treatment effects for marginal decision-makers: Everyone is marginal
This note develops a theoretical framework for identifying treatment effects when the treatment itself arises from an endogenous, policy-sensitive decision, and then applies this framework to the empirical question of payroll tax incidence in small firms for their initial hiring. In this setting, outcomes (wages) are only defined for agents who enter the participation (to become an employer), yet the decision to participate is precisely what the policy intends to influence. This endogeneity creates not only an empirical but a fundamentally conceptual challenge: treatment status is not merely a covariate but a defining threshold that alters who is observed and whether the outcome of interest is defined at all. We thus have our question regarding identification: what should researchers identify?
Economists have long recognised the problem of selection into treatment. The local average treatment effect (LATE) framework imbens1994identification and its extensions, including marginal treatment effects (MTE) heckman2005structural, formalized causal inference under selection, especially in the presence of continuous unobserved heterogeneity. These approaches remain influential, but they are now decades old and rooted in a reduced-form tradition. They do not distinguish between entry into treatment and variation within treatment, and they treat marginal and inframarginal participants as observationally similar except for an index of resistance to treatment.
This note takes a different stance. I focus on settings where entry into treatment is itself an economic outcome, and where agents who enter and those who do not are structurally different. The empirical motivation of this note is a 2016 Belgian payroll tax reform that granted a permanent exemption to firms that hired their first employee. The policy targeted non-employers at the extensive margin, aiming to induce new firms' labour market participation. The treatment of a hiring subsidy coincides with participation of the act of hiring, and therefore defines the sample of observed outcomes and changes the composition of the employer population.
The realistic goal is to study the reform effects on entrant employers' wage (rates) offered, and then answer the payroll tax incidence for the first hires. To do so, I first develop a model of heterogeneous firms, where each firm's productivity determines both its hiring probability and the wage it would pay if it does hire. The reform shifts the hiring probability function, inducing entry by marginal firms that would not have hired otherwise. I then show that observed differences in post- and pre-reform average outcomes, typically yielded by most methods, can be decomposed into causal treatment effects plus (uninterpretable) biases.
To fix these problems, I introduce a class of importance-weighted estimands that target the policy-induced margin, and relate them to commonly applied estimation techniques. This framework thus bridges structural and reduced-form perspectives and provides a policy-relevant estimand when both treatment take-up and outcome definitions are endogenous. It is not a reinterpretation of LATE or MTE, but an attempt to address the empirical challenges that arise when treatment status defines existence itself.
My setting parallels that studied by lee2009training, who addresses a similar `double selection' (on both the existence and the outcome) structure in evaluating the effects of job training programs. Lee derives sharp nonparametric bounds on treatment effects under minimal assumptions by trimming treated observations to match the support of the control group, a strategy that ensures comparability even when potential outcomes are partially missing.
Instead of bounding, the approach developed in this paper explicitly models the selection process through a productivity-indexed hiring probability function. This allows for the construction of meaningful estimands under the shifting employer distribution, including importance-weighted and marginality-weighted treatment effects. These estimands remain interpretable even when the support of treated and untreated groups does not fully overlap, and they offer a more informative alternative in settings where the policy causes substantial endogenous entry into the treated population.
This section builds a minimal note summarising the entrepreneurial decision model developed by deng2025paper0, a microeconomic-theory framework of employer firm formation regarding whether entrepreneurs choose to become an employer or stay as solo self-employed. Although the model wording literally addresses entrepreneurship, it can be easily extended to general programme participation situations.
A continuum of small entrepreneurial firms differ in their (unobserved) productivity $\theta_i \sim f(.)$ over $\mathbb{R}^+$, which is the population density; the cumulative distribution density is $F(\theta)=\int_{0}^{\theta}{f(x)dx}$. Firms with a higher productivity level can, with production factor inputs constant, generate higher production. They enter the labour market by hiring their first employees with a probability $p^s(\theta_i)$ and if they do, they offer a wage rate $y^s(\theta_i)$, differently in an untreated status (no such exemption, $s=0$, untreated outcome $y^0$) and in a treated status (with such exemption, $s=1$, trea`ted outcome $y^1$). The rest stays self-employed with no employees and do not pay wages.
With minimal standard assumptions, the model concludes that (1) the probability of hiring is an increasing and continuous function in the productivity level, and (2) the reform increases the probability of entry from $p^0(\theta_i)$ to $p^1(\theta_i)$ and therefore raises the mass of entrant employers from $N^0=\int_{0}^{+\infty}{p^0(\theta)dF(\theta)}$ to $N^1=\int_{0}^{+\infty}{p^1(\theta)dF(\theta)}$.
An important consequence of the hiring probability is that firms of different productivity levels occurs in the employer dataset with different probabilities. The probabilistic distribution of employer firms in state $s$ is a importance-weighted distribution, a composite of $p^s(\theta)$ and $f(\theta)$,
with $q^s(\theta)=\frac{p^s(\theta)}{N^s}f(\theta)$ the importance-weighted probability density. Under this density, the observed mean wage (or other outcomes) is
with importance weights $\frac{p^s(\theta)}{N^s}$.\footnote{This is the importance weight when integrated against $f(\theta)d\theta$. When integrated against $d\theta$ alone, the weights are $\frac{p^s(\theta)f(\theta)}{N^s}$.} Here I borrow the term of `importance weights' from the field of machine learning sugiyama2007importance.weighted, but the idea is simple in economics: the composition of employer entrants changed from distribution $Q^0$ to $Q^1$. There are possibly many other terms that can apply.
If we know the potential outcome dependence on firm types $y^1(.), y^0(.)$, then the treatment effect for type $\theta$ firm is $\tau(\theta) \coloneqq y^1(\theta)-y^0(\theta)$, and the population average treatment effect (PATE) weighted under the productivity distribution scheme $F(\theta)$ is
The PATE neglects the probability of hiring and only weights according to the population distribution of types; as a result, it implies that firms of low productivity (which basically never hire) are equally important as firms of high productivity (which always hire) in drawing the average treatment effect.
While this parameter appears intuitive and is often regarded as a natural benchmark, it suffers from a fundamental shortcoming in the context of employer-based policies: it averages over all firm types, including those that never hire employees and thus never appear in observed wage data. In reality, the average among employer firms always carries some importance weighting: the observed means are additionally weighted with firms' probability of hiring (and thus occurring in the sample). Therefore, PATE is not sufficiently interesting to know.
Inverse probability weighting (IPW) estimators rosenbaum1983central, hirano2003efficient, imbens2004nonparametric, commonly used in applied econometrics, often aim to recover this PATE by reweighting observations to match the population distribution (see Appendix (ref) for a simple derivation). However, in doing so, they neglect a key feature of the data-generating process: firms enter the employer sample with unequal probability, determined endogenously by their productivity and the policy environment. In practice, observed outcomes (e.g., wages) are only defined for employer firms, and these are not a random draw from the population of all firms. They are selectively observed, based on hiring thresholds that the policy itself shifts.
Moreover, the reference distribution \( F(\theta) \) itself is not uniquely determined. In practice, the composition of the population depends on whether firms that never hire (purely self-employed entities) are included. This choice affects the weight assigned to firm types that are observationally irrelevant for the outcome of interest. If all legal entities are included, \( F(\theta) \) will place substantial mass on types that contribute no information on wages. If only firms with positive hiring probability are included, then the estimand already deviates from a true population average.
Designing an identification strategy for the treatment effect on wages is complicated by the fact that wages are only defined for employer firms, and the set of employers itself changes endogenously in response to the reform. Before the reform, firms that would later become treated are non-employers, and thus do not generate wage observations. As a result, some standard estimations like a panel difference-in-differences approach are infeasible: there is no pre-treatment wage trajectory for firms that were not yet hiring.
To leverage the policy variation over time, a natural alternative is to compare post-reform cohorts of new employers to pre-reform cohorts, and to adopt a cross-sectional difference-in-differences design or a regression-discontinuity-in-time design by leveraging the cohort time. Assuming for now that a valid control group has been constructed to account for concurrent trends (in a DiD) or that outcome polynomials are correctly modelled as functions of the entry time (in a RDiT), such a strategy yields the observed mean difference (OMD) estimand:
with the observed means naturally carrying importance weighting schemes as defined in (ref), differently in the pre-reform periods and in the post-reform periods.
Although feasible to calculate, the OMD estimand is uninteresting as it carries two serious biases along with an interpretable treatment effect parameter. It can be decomposed into
where: the first term is a true ATE weighted under the post-reform hiring probability, an interpretable parameter; the second term is a selection bias equal to the difference in the untreated wages between two different distributions, which arises from the composition changes in firm productivity levels among employer firms; the third term is a re-weighting effect bias equal to the difference in untreated wages between two weighting schemes, which occurs solely because the relative importance of firms in the hiring distribution changes following the reform.
In this Eq (ref), the first term $\mathbb{E}_{Q^1}[\tau(\theta)]$ weights the average treatment effects according to the post-reform importance-weighted distribution, giving more weights to firms that are more likely to hire, and fewer weights to firms that are less likely. This weighting scheme is consistent with observation in reality and is thus interesting to know. Alternatively, we can write the first term evaluated under the pre-reform importance-weighted distribution, $\mathbb{E}_{Q^0}[\tau(\theta)]$, plus the two alternatively defined biases (See Appendix (ref)).
A third common but problematic approach is to interpret the reform through a dichotomy of marginal versus inframarginal firms. The intuition is that some firms are induced by the policy to cross the threshold into treatment (`marginals'), while others would have hired regardless (`inframarginals'). This classification may appear appealing and useful in many applications (e.g., hombert2020france, branstetter2014entry), but here it relies on imposing an arbitrary threshold on the participation probability.
Formally, suppose one specifies a participation probability threshold \( \underline{p} \) and defines the productivity cut-off points \( \underline{\theta} \) and \( \underline{\underline{\theta}} \) such that \[ p^0(\underline{\theta}) = \underline{p}, \quad p^1(\underline{\underline{\theta}}) = \underline{p}. \] The policy-induced shift in the hiring function from \( p^0 \) to \( p^1 \) ensures that \( \underline{\underline{\theta}} < \underline{\theta} \), and allows one to define two conditional average treatment effects:
However, this construction raises several conceptual problems. First, the threshold \( \underline{p} \) is not identified from the model or data but must be imposed manually. Different choices of \( \underline{p} \) result in different delineations of marginal and inframarginal types, and therefore alter the interpretation of the resulting estimands. This arbitrariness undermines both the stability and the policy relevance of the marginal--inframarginal split.
Second, the imposition of a sharp threshold \( \underline{p} \) discretises an inherently continuous decision process. The hiring probability function \( p(\theta) \) is continuous in firm productivity by construction, and the reform shifts this entire function upward. Categorising firms as either marginal or inframarginal discards this structure and replaces it with a rigid partition that lacks behavioural foundation.
Third, and more seriously, the marginal group defined as \( \theta \in (\underline{\underline{\theta}}, \underline{\theta}] \) consists of firms that do not hire in the pre-reform regime. As such, their untreated outcomes \( y^0(\theta) \) are never observed by construction. This creates a fundamental counterfactual problem: there is no empirical support for their behaviour in the absence of treatment, making \( \tau^\texttt{Mar}(\underline{p}) \) unidentified without strong structural assumptions or extrapolations from other groups.
To make this concrete, the marginal average treatment effect can be written as:
The first term, the post-reform mean among induced entrants, is well-defined and observable. The second term, however, is undefined: these firms were non-employers before the reform and thus paid no wages. Their counterfactual wage outcomes \( y^0(\theta) \) cannot be observed or inferred without unverifiable extrapolation (see the right panel of Figure (ref)). (This is a parallel problem to the “no common support” in the propensity score matching estimator methods.)
To isolate the issue of the dichotomous classification itself, the notation here suppresses any subscript indicating the weighting distribution (e.g., \( Q^0 \), \( Q^1 \)). But in any realistic analysis, where one must average these conditional effects under a coherent probability distribution, the arbitrariness of the dichotomy becomes even more problematic. It is unclear under which regime (pre- or post-reform) these averages should be evaluated, since the marginal group exists only under the post-reform distribution and is absent by construction in the pre-reform counterfactual.
The issues highlighted in the previous section—particularly the observational detachment of the PATE, the bias-prone nature of the observed mean difference (OMD), and the counterfactual gaps in the marginal-inframarginal framework—call for estimands that are better grounded in observed data, more coherent with treatment-induced selection, and more meaningful for policy analysis. In this section, I define a class of treatment effect estimands that are constructed via continuous reweighting of the population distribution, using the selection structure of the model to determine the relevant subpopulations for averaging.
A natural starting point is to define treatment effects not over the entire firm population, but over the subpopulations of firms that actually hire under each regime. This gives rise to two importance-weighted average treatment effects:
These estimands address key shortcomings of the bad estimands. Unlike the PATE, they assign realistic weights to firms of different hiring probability (productivity). Unlike the OMD, they apply consistent weights across potential outcomes and thus avoid conflating treatment effects with selection or compositional changes. And unlike the marginal--inframarginal split, they do not rely on arbitrary thresholds or unobserved counterfactuals.
The choice between \( Q^0 \) and \( Q^1 \) depends on the policy question. Weighting under \( Q^0 \) corresponds to the effect on firms that would have hired in the absence of the reform---akin to an average treatment on the treated (ATT) among inframarginals. Weighting under \( Q^1 \) instead focuses on the firms who actually hired after the reform, reflecting the realised composition of treated firms. Both estimands are well-defined and feasible to estimate from observed data, as the sample of employers under each regime provides direct support for each weighting function.
Propensity score matching (PSM) rosenbaum1983central, heckman1997matching, abadie2006matching, another commonly applied method (cf. IPW in Section (ref)), can recover these two estimands if performed properly (see Appendix (ref) for proofs). Matching firms based on the pre-reform hiring probability scheme \( p^0(.) \), by fitting only pre-reform data and extrapolating to post-reform, recovers a \( Q^0 \)-weighted treatment effect: the average treatment effect among firms that would have hired in the absence of the reform. Similarly, matching based on \( p^1(.) \), by fitting only post-reform data and extrapolating to pre-reform, recovers the \( Q^1 \)-weighted ATE among firms that do hire after the reform. However, matching \( p^0(.) \) to \( p^1(.) \), or pooling observations across regimes to estimate a single propensity score scheme, risks invalid inference (and results in uninterpretable estimands like the OMD): the propensity scores originate from structurally different selection processes and cannot be reconciled.
Applicability is limited in their own aspects. \( \hat{p}^0 \)-matching might break down for marginal entrants since their pre-reform counterparts did not hire and thus lack observed outcomes. The result is either a loss of support (excluding marginals) or bias due to extrapolation from non-observed counterfactuals. Conversely, \( \hat{p}^1 \)-matching faces the issue that pre-reform firms did not face the same incentives as post-reform firms. But since both marginals and inframarginals are captured by \( \hat{p}^1(X) \), this matching should remain more feasible if one wants to include marginals in the estimation.
The two importance-weighted estimands above can be seen as special cases of a broader framework. More generally, one can define:
where \( w(\theta) \) is any non-negative, integrable weighting function. This formulation allows analysts to tailor the estimand to the policy objective or empirical setting. For instance: Choosing \( w(\theta) = p^0(\theta) \) recovers \( \tau^{Q^0} \), choosing \( w(\theta) = p^1(\theta) \) recovers \( \tau^{Q^1} \), and choosing \( w(\theta) = 1 \) recovers the (uninformative) PATE.
This generalised framework resolves the rigidity of the marginal--inframarginal split by treating selection as continuous and policy-dependent. Moreover, it avoids the inconsistencies of the OMD by applying symmetric and interpretable weights across treatment and control potential outcomes.
From a feasibility perspective, however, estimating \( \tau^w \) requires knowledge or estimation of the weighting function \( w(\theta) \), and, in general, access to both \( y^1(\theta) \) and \( y^0(\theta) \) across the support of \( w(\theta) \). When \( w \) corresponds to observed hiring probabilities (as in \( Q^0 \) or \( Q^1 \)), this is feasible from the data. For other \( w \), feasibility depends on whether those weights are supported in both regimes or can be estimated structurally.
A particularly policy-relevant weighting function arises by considering the shift in hiring probability induced by the reform: \[ \Delta p(\theta) \coloneqq p^1(\theta) - p^0(\theta). \] This function captures the incremental participation probability at each productivity level and defines the marginality-weighted average treatment effect:
This estimand directly targets the subpopulation of firms whose hiring decision was responsive to the reform. It avoids the thresholding problem of the marginal--inframarginal dichotomy by explicitly modelling marginality as a continuous concept.
If we can capture good proxies of the unobservable firm productivity into a set of observables $X$, under weak and interpretable conditions, it can be shown that the estimand can be rewritten entirely in terms of observables. Define \( \tau(x) = \mathbb{E}[\tau(\theta) \mid X = x] \) and \( \Delta p(x) = \mathbb{E}[\Delta p(\theta) \mid X = x] \), then the marginality-weighted estimand can be expressed as
This observable analogue can be consistently estimated by partitioning the support of \( X \), estimating \( \Delta p(x) \) and \( \tau(x) \) within each cell, and applying a weighted average. A formal derivation is given in Appendix (ref).
An additional practical advantage of the marginality-weighted estimand is that it relies only on estimating the policy-induced change in hiring probability, \( \Delta p(x) = p^1(x) - p^0(x) \), rather than the full levels of the regime-specific hiring propensities \( p^0(x) \) and \( p^1(x) \) themselves. This distinction is important: while the \( Q^0 \)- and \( Q^1 \)-weighted estimands require full propensity models—typically via possibly non-linear methods across the entire support of \( x \), the marginality-weighted estimand treats the difference in probabilities as the estimand of interest. In practice, estimating a linear model in probability is often more robust and less sensitive to misspecification than estimating levels, particularly in regions where the propensity is close to zero or one. This makes \( \tau^{\Delta p} \) not only more policy-relevant but also more empirically feasible in applied work.
In conclusion, the framework developed here highlights the central role of marginal decision-makers in causal analysis. Treatment not only shifts outcomes but also determines the very existence of the relevant units, making participation endogenous and outcomes only meaningful for those who cross the margin. This perspective unifies diverse applications under a single principle: everyone is marginal with respect to some decision.
Through a long discussion, the method proposed in the Section (ref) is my favourite one and is the innovation of this paper due to two important reasons. First, this is actually the most realistically relevant parameter for the programmes that are designed to encourage people at the margin to participate, where policymakers care about the economic effects for these marginal people induced to enter instead of the effects for inframarginal people who enter anyway. Second, as I discussed and shall discuss more, this parameter is easier to recover than those proposed in Section (ref). In a next version of the note, I plan to formally propose estimators for the marginality-weighted estimands and prove their statistical properties.
Just as an example, I list four cases where my framework can be highly useful for real-world understanding:
Research questions and policies alike are commonplace in the fields of labour economics, development economics, health economics, public economics, etc. Researchers can thus adopt this framework to analyse the effects of those policies among individuals who are most responsive to the policies.
{ \begingroup {0pt} \setstretch{1} \endgroup }