EconBase
← Back to paper

Changes-in-Changes for Ordered Choice Models: Too Many "False Zeros"?

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

78,318 characters · 7 sections · 52 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Changes-in-Changes for Ordered Choice Models with Underreporting

\def\spacingset#1

{1}

abstractWe develop a Difference-in-Differences framework for discrete, ordered outcomes subject to underreporting. Such outcomes commonly arise in self-reported surveys on socially undesirable or stigmatized behaviors, where respondents may conceal their true behavior. For a discrete Changes-in-Changes model that is shown to admit an equivalent threshold-crossing representation, we derive nonparametric bounds for the counterfactual and factual outcome distributions as well as for the associated quantile treatment effects when outcomes are underreported. These bounds are shown to be sharp uniformly across outcome levels under additional support conditions, and we propose suitable estimation and bootstrap inference procedures. In an extension, we also consider a semiparametric underreporting model that allows to point identify and estimate distributional treatment effects. As an application, we investigate the impact of recreational marijuana legalization on the consumption behavior of 8th-grade students in several U.S. states. \noindentJEL Codes: C21, C25, C31, C51 \noindentKeywords: Partially Observable Outcomes, Identification, Treatment Effects, Measurement Error

{1.8}

Introduction

The Difference-in-Differences (DiD) method is widely used in applied economics to identify causal effects of policy changes. While the classical DiD framework focuses on mean treatment effects, policy changes may affect different parts of the outcome distribution in distinct ways. This observation has motivated a growing literature on distributional DiD methods that exploit pre- and post-treatment data to identify heterogeneous effects across the outcome distribution rothWhatTrendingDifferenceindifferences2023. A prominent example of this approach is the Changes-in-Changes (CiC) model developed by AI2006.

This paper develops a CiC framework for discrete, ordered outcomes that are potentially underreported, nesting the “misreporting-to-zero” or “false-zero” scenario as a special case. Underreporting is common in survey data on discrete outcomes of count, ordinal, or binary nature, particularly when socially undesirable or stigmatized behaviors are elicited. In such settings, respondents may conceal their true behavior, posing substantial challenges for identification and inference. Examples include the consumption of “social bads” such as marijuana (cannabis) or alcohol, as well as sensitive issues such as abortion or domestic violence.\footnote{For example, BHSZ2018 found that the actual incidence of marijuana consumption in an Australian survey was nearly double the reported figure (23% vs. 12.2%), while CM2005 identified about 20 percentage points underreporting in cheating behavior among undergraduate students in a U.S. survey. See also NDT2019 for other examples.}

The paper makes several contributions to the literature. First, we study the discrete CiC model for ordered outcomes in the benchmark case without underreporting. We show that this model admits an equivalent threshold-crossing representation in which the observed outcome is generated by comparing a latent index with time-specific thresholds. Under this representation, the CiC restriction translates into restrictions on the location and scale parameters of the latent index across groups and time, which nest standard parallel-trends restrictions for nonlinear ordered choice models (e.g., Probit or Logit DiD models) as special cases. The representation therefore provides further insights into the discrete CiC model more broadly, and serves as the basis for a point-identified, semiparametric model with underreporting that we develop in the Supplementary Material (see Section (ref)).

As a second contribution, we derive, in the presence of underreporting, nonparametric bounds for both the treated group's counterfactual outcome distribution after treatment and the factual outcome distribution. The bounds are shown to be sharp uniformly across outcome levels under mild, additional assumptions. Moreover, we demonstrate that, without further restrictions, the bounds on the counterfactual distribution remain uninformative even if an instrumental exclusion restriction for consumption is available. By contrast, leveraging the CiC structure together with an exclusion restriction and an upper bound on the misreporting probability can lead to informative, tight bounds. The latter bound on the misreporting probability may either serve as a sensitivity parameter akin to, e.g., the $c$-dependence parameter in Masten2018 and the breakdown frontier in MP2020, or simply reflect institutional knowledge. We also show that not every candidate value of this bound is admissible: sufficiently small values can be ruled out by the observed relationship between the reported outcome and the instrument. Our bounds extend the partial identification results of AI2006 for the discrete outcome case to underreporting and give rise to bounds on Quantile Treatment Effects on the Treated (QTTs) as an immediate corollary. They also relate to the recent analysis of misreporting in binary outcome models by MW2024, even though our focus is on ordered outcomes and on combining underreporting with a causal CiC structure. In the supplement, we complement these partial identification results with a semiparametric, point-identified model that considers consumption and reporting decisions individually and that allows for the recovery of distributional treatment effects rather than QTTs (cf.\ Section (ref)).

As a third contribution, we develop estimation and inference theory for all bound functions and resulting QTTs. Under repeated cross-sectional sampling, we establish weak convergence of the estimated lower and upper bounds for the counterfactual distribution uniformly across outcome levels. Because the bound operators involve nonsmooth transformations such as minimum operators, the relevant maps are only Hadamard directionally differentiable. We therefore follow FS2019 and MP2020 and employ bootstrap inference to construct uniformly valid Confidence Sets (CS) for the bounds of the counterfactual and factual distributions. The construction of QTT bounds and corresponding CS then uses the quantile-effect framework of CFMW2020. This uniformity is practically useful because it allows the confidence band to be used for a range of functional comparisons without requiring the researcher to commit ex ante to a specific outcome level. With a pre-specified probability, the band covers the entire bound function across all outcome levels, so that any candidate counterfactual distribution function or QTT function that leaves the band at even a single point can be rejected at the corresponding level.

Finally, as an empirical application, we investigate the impact of recreational marijuana legalization for adults in several U.S. states on the short-term marijuana consumption of 8th-grade students. The focus on this group is particularly relevant because early cannabis use is widely viewed as especially consequential for later health and educational outcomes. Unlike existing studies that examined the effects of legalization of medicinal or recreational marijuana use on the mean AHR2015,WHC2015,CWFKSSH2017,HWB2022, or on the timing of the first (early age) consumption WBJ2014, we consider the effects of the legalization of recreational marijuana use on the entire consumption distribution of this population in the treated states in our sample, thus allowing for the possibility of heterogeneous effects of the reforms across usage levels. Moreover, in contrast to the existing literature, we address concerns that consumption behavior may not always be reported truthfully, and that reporting behavior may have been affected by legalization through shifts in stigma, public perception, or general attitudes towards marijuana consumption. Our analysis reveals that the benchmark specification that ignores underreporting yields little evidence of non-zero effects outside the extreme upper tail of the distribution. Once underreporting is allowed for, however, the picture becomes more asymmetric: at lower upper-tail quantiles the data rule out positive QTTs, whereas at the very top they rule out negative QTTs. A complementary semiparametric specification, reported in Section (ref) of the supplement, yields a similar pattern: without covariates, the estimated effects are not significant, whereas including covariates implies a decline in non-consumption of about 2 percentage points and increases of about 1 percentage point in each positive consumption category. Allowing for misreporting further increases these estimated consumption effects by more than 50%, while yielding little evidence of systematic changes in reporting behavior.

The paper is organized as follows. Section (ref) outlines the model setup and defines the causal parameters of interest. Section (ref) studies the benchmark case without underreporting and characterizes the discrete CiC model within a general threshold-crossing framework. Section (ref) extends the framework to account for underreporting and establishes sharp bounds for the counterfactual distribution and the resulting quantile treatment effects. Section (ref) develops estimation and inference for these bounds. Section (ref) empirically investigates the effects of recreational marijuana legalization on the consumption behavior of 8th-grade high-school students. Section (ref) concludes.

All proofs and additional material can be found in the Supplementary Material. Section (ref) contains the proofs of the bound results, while Section (ref) contains the proofs for estimation and inference. Section (ref) outlines how estimation and inference can be modified when additional covariates are present or some of the variables are continuous. Section (ref) presents the semiparametric model for underreporting, which delivers point identification and estimation of distributional treatment effects. Finally, Section (ref) reports a Monte Carlo study assessing the finite-sample performance of the bound estimators and the associated bootstrap CS, while Section (ref) presents additional tables for the empirical application.

Notation. We adopt the shorthand of AI2006 and write \[ Y_{gt} \sim Y \mid G=g, T=t, \] where $\sim$ denotes equality in distribution. Let $\overline{\mathbb{R}}:=\mathbb{R}\cup\{-\infty,+\infty\}$. For a random variable $Y$, let $F_Y$ denote its cumulative distribution function (CDF), with $F_Y(-\infty)=0$ and $F_Y(+\infty)=1$. The associated generalized inverses are

align*[align* omitted — 302 chars of source]

with the conventions $\sup \emptyset = -\infty$ and $\inf \emptyset = +\infty$. The same notation applies to conditional distribution functions. We adopt the convention that conditional probabilities are defined as zero whenever the conditioning event has probability zero.

Set-Up

To introduce the CiC framework for discrete, ordered outcomes, let $Y(d)$, $d\in\{0,1\}$, denote the potential outcome with support contained in a finite set $\mathcal{J}$. Motivated by the empirical application on drug consumption frequencies in Section (ref), we henceforth refer to $Y(d)$ as the “true” potential consumption level. We set $\mathcal{J}:=\{0,1,\ldots,J\}$ with $J\ge 1$. Let $G\in\{0,1\}$ denote the group membership indicator, where $G=1$ corresponds to the treated group, and let $T\in\{0,1\}$ denote the time period indicator, where $T=1$ corresponds to the post-treatment period. The binary treatment indicator is given by $D=G\cdot T$, so that $D=1$ if and only if $(G,T)=(1,1)$, and $D=0$ otherwise. Throughout the paper, we focus on the case of two groups and two time periods for notational simplicity. Extensions to multiple pre-treatment periods are immediate, at the cost of more cumbersome notation. Finally, to avoid degenerate cases, we assume \(\pi_{gt}:=\mathrm{Pr}\!\left(G=g,T=t\right)>0\) for all \((g,t)\).

As emphasized in the introduction, a persistent challenge in analyzing sensitive behaviors, such as the consumption of “social bads”, is the presence of “false zeros” or, more broadly, underreporting. Individuals often conceal their true participation in such behaviors when responding to retrospective surveys NDT2019. This leads to a partially observable outcome model in which a specific outcome value is observed in the data if and only if the true outcome corresponds to that value and is reported truthfully P1980. To accommodate this feature, we introduce potential “reported” consumption, denoted by $C(d)$, $d\in\{0,1\}$, which satisfies $C(d)\le Y(d)$ with probability one. The case $C(d)=Y(d)$ corresponds to truthful reporting, whereas $C(d)<Y(d)$ captures underreporting. We formalize this in our first assumption within the DiD framework.

assumption[Underreporting] For each $j\in \mathcal{J}\setminus \{J\}$ and $(d,g,t)\in\{0,1\}\times\{0,1\}\times\{0,1\}$, and, for some known constant $\alpha\in[0,1]$, the “true” and “reported” consumption levels $Y(d)$ and $C(d)$ satisfy \begin{align*} &\mathrm{Pr}\!\left(C(d)\le Y(d)\mid G=g,T=t\right)=1,\\ &\mathrm{Pr}\!\left(C(d)\le j \,\middle|\, Y(d)\ge j+1,G=g,T=t\right) \le \alpha. \end{align*}

Assumption (ref) explicitly characterizes one-sided misreporting, that is, underreporting, for the potential outcome in a given group at a given time. Although it is stated for all \((d,g,t)\), this uniform formulation is adopted for simplicity, and the identification analysis would go through under weaker, cell-specific restrictions on $\alpha$. Such one-sided misreporting is natural in survey settings involving socially undesirable or stigmatized behaviors, including self-reported substance use, where respondents may conceal or downplay true behavior, while overreporting is not really plausible BHSZ2018,GHSZ2018. Alternative scenarios where overreporting is the only likely misreporting form (e.g., a survey question about the frequency of physical exercise), can be accommodated by adapting the assumption accordingly. Thus, the first part of the assumption rules out overreporting, while the second part introduces a uniform upper bound on underreporting probabilities across outcome levels that can reflect institutional knowledge or serve as a sensitivity parameter. Specifically, when \(\alpha=1\), the assumption imposes no restriction on the extent of underreporting, whereas when \(\alpha=0\), reporting is truthful almost surely. Intermediate values of \(\alpha\) therefore impose flexible, nonparametric restrictions on the extent of underreporting. As will be shown in Proposition (ref), in the absence of any restriction on underreporting, the resulting bounds on the counterfactual distribution become uninformative. Hence, \(\alpha\) can serve as an important sensitivity parameter that governs the sharpness and informativeness of the resulting bounds.

Finally, the observed outcome is generated according to the standard observability rule,

equation[equation omitted — 73 chars of source]

This identity highlights that, due to underreporting, the observed distribution of $C$ does not generally coincide with the distribution of the true potential outcome $Y(d)$, even within the subpopulation for which $D=d$.

Our primary objects of interest are the QTTs, which characterize how treatment shifts different points of the outcome distribution for treated units, such as the median or the $75$th percentile. In our empirical application, they measure the effect of marijuana legalization on different points of the consumption distribution, among adolescents residing in treated states. To formalize this, we first define the conditional quantile functions of the potential outcomes for the treated population. For each $d\in\{0,1\}$ and quantile level $\tau\in(0,1)$, let \[ Q_{Y(d)\mid D=1}(\tau) := F^{-1}_{Y(d)\mid D=1}\left(\tau \right) = F^{-1}_{Y(d)\mid G=1,T=1}\left(\tau \right). \] For each $\tau\in(0,1)$, the QTT at quantile level $\tau$ is defined as

equation[equation omitted — 108 chars of source]

In the absence of underreporting, $Q_{Y(1)\mid D=1}(\tau)$ is identified from the observed treated group in the post-treatment period. By contrast, $Q_{Y(0)\mid D=1}(\tau)$ depends on the counterfactual distribution of $Y(0)$ for treated units and is not identified without additional assumptions.

Finally, in Section (ref) of the supplement, we also consider Distributional Treatment Effects on the Treated (DTTs), defined for each $j \in \mathcal{J}$ as

align[align omitted — 127 chars of source]

For discrete outcomes, the DTTs complement the QTTs by tracking how treatment reallocates probability mass across outcome levels. In our empirical application, they capture how marijuana legalization changes the probability that adolescents in treated states fall into each consumption category.\footnote{An analogous extension of the QTT and DTT results to the untreated group is possible. It can be obtained by reversing the roles of treated and untreated groups and imposing the corresponding support conditions; see AI2006.}

Bounds Without Underreporting

As a first step, consider the benchmark case without underreporting, $\alpha=0$ in Assumption (ref), so that $C_{gt}(d)=Y_{gt}(d)$ a.s. In this setting, the distribution of $Y_{11}(1)$ is identified from the treated group in the post-treatment period, but the counterfactual distribution $F_{Y_{11}(0)}$ is not identified without further restrictions. We therefore require additional structure on the untreated potential outcome $Y(0)$ across groups and time. For discrete, ordered outcomes, we adopt the discrete CiC framework of AI2006:

assumption[Discrete CiC Model] There exist a real-valued random variable $U$ and a function $h:\mathbb{R}\times \{0,1\}\rightarrow \mathbb{R}$ such that $Y(0) = h(U,T)$ a.s. The random variable $U$ and the function $h$ satisfy the following conditions: \begin{enumerate}[label=(\arabic*)] • $U\mid G=g\text{ is continuously distributed}$ for all $g\in \{0,1\}$, • $U\perp T\mid G$, • $\mathrm{supp}\left(U\mid G=1 \right)\subseteq \mathrm{supp}\left(U \mid G=0\right)$, • $u\mapsto h(u,t)$ is nondecreasing, \end{enumerate}

The function $h$ determines the untreated outcome $Y(0)$ and may depend on the time indicator $T$ and an unobservable scalar $U$ capturing latent heterogeneity. The specification $Y(0)=h(U,T)$ implies that group membership $G$ does not enter the outcome function directly and affects $Y(0)$ only through the distribution of $U$. In particular, conditional on $U=u$, the untreated outcome is identical across groups within a given period $T\in\{0,1\}$.

Condition (1) is standard in nonseparable models and rules out mass points in the latent heterogeneity within a group. Continuity of $U$ is a common assumption in nonlinear discrete choice models, including threshold-crossing specifications. Condition (3) ensures sufficient support overlap to allow for extrapolation of the counterfactual distribution. The weak monotonicity condition in (4) is natural for discrete, ordered outcomes and aligns with the monotone index structure underlying threshold-crossing models.

Condition (2) is the key identifying restriction of the CiC model. It requires that the distribution of the latent variable $U$, although possibly different across groups, remains stable over time within a group. This conditional independence restriction distinguishes CiC from other distributional DiD approaches. One alternative is the distributional DiD model proposed by GKM2023, which is based on time-invariance of the copula linking the untreated potential outcome and treatment assignment. Their identifying assumption imposes stability of the dependence structure between the untreated potential outcome and group membership across periods, while allowing the marginal distributions to evolve freely. The copula-stability restriction and the CiC restriction are not nested in general. They coincide only when the untreated potential outcomes are continuous with strictly increasing marginal distribution functions GKM2023.

While GKM2023 clarify the relationship between CiC and copula stability in continuous settings, their analysis does not directly extend to the discrete CiC model considered here. We therefore begin by establishing an equivalence result that characterizes the discrete CiC model as a threshold-crossing representation for a latent outcome variable with suitable location and scale restrictions across groups and time. This characterization will be used in Section (ref) of the supplement to construct and estimate a semiparametric model for consumption and reporting decisions with covariates.

proposition[Equivalence of Discrete CiC Models] Suppose that \(\mathrm{supp}\!\left(Y(0)\mid T=t\right)=\mathcal{J}\) for all \(t\in \{0,1\}\). Then Assumption (ref) holds with $\mathrm{Var}\left[U\mid G=g\right]<+\infty$ for all $g\in \{0,1\}$ if and only if there exists a real-valued random variable $V$ and real-valued parameters $\{\eta_{gt},\lambda_{gt},\kappa_{t,j}\}_{(g,t,j)\in \{0,1\}\times \{0,1\}\times \mathcal{J}\setminus \{J\}}$, with $\lambda_{gt}>0$ for all $(g,t)\in \{0,1\}\times \{0,1\}$, such that \begin{enumerate}[label=(\arabic*)] • $V\mid G=g$ is continuously distributed with $\mathrm{Var}\left[V\mid G=g\right]<+\infty$ for all $g\in \{0,1\}$, • $V\perp T\mid G$, • $\mathrm{supp}\left(\eta_{GT}+\lambda_{GT}V\mid G=1,T=0 \right)\subseteq \mathrm{supp}\left(\eta_{GT}+\lambda_{GT}V \mid G=0,T=0\right)$, • $\displaystyle\eta_{11}-\frac{\lambda_{11}}{\lambda_{10}} \eta_{10}=\eta_{01}-\frac{\lambda_{01}}{\lambda_{00}} \eta_{00}$ and $\displaystyle \frac{\lambda_{11}}{\lambda_{10}}=\frac{\lambda_{01}}{\lambda_{00}},$$\kappa_{t,0}=0$, and if $J\geq 2$, $\kappa_{t,1}=1$ and $\kappa_{t,j-1}<\kappa_{t,j}$ for all $(t, j) \in\{0,1\} \times\{1, \ldots, J-1\}$, \end{enumerate} and $$Y(0) =\begin{cases} 0, & \text{if } \eta_{GT}+\lambda_{GT}V \le \kappa_{T,0},\\ j, & \text{if } \kappa_{T,j-1} < \eta_{GT}+\lambda_{GT}V \le \kappa_{T,j},\ j=1,\dots,J-1,\\ J, & \text{if } \kappa_{T,J-1} < \eta_{GT}+\lambda_{GT}V, \end{cases} \quad\text{a.s.}, $$where $\eta_{GT}:=\sum_{(g,t)} \eta_{gt}\mathbb{I}\{G=g,T=t\}$ and similarly for $\lambda_{GT}$, and $\kappa_{T,j}:=\sum_{t} \kappa_{t,j}\mathbb{I}\{T=t\}$.

Proposition (ref) highlights that the discrete CiC model can be transformed into an alternative threshold-crossing model without affecting \(Y(0)\) almost surely. We additionally impose $\mathrm{Var}\!\left[U\mid G=g\right]<+\infty$, a mild technical condition ensuring that the equivalent threshold-crossing representation admits a well-defined scale parameter, while $\mathrm{supp}\left(Y\left(0\right)\mid T=t\right)=\mathcal{J}$ guarantees that the associated thresholds are well defined and strictly ordered. Note that these extra conditions are only required for the equivalence result of Proposition (ref), and not used in any other part of the paper.

Here, $\eta_{gt}$ may be interpreted as the location (e.g., mean) parameter of a corresponding latent variable, say $Y^{\ast}(0):=\eta_{GT}+\lambda_{GT}V$, for group $G=g$ and period $T=t$, while $\lambda_{gt}$ captures the corresponding scale. The discrete outcome \(Y(0)\in\mathcal{J}\) is generated by comparing this latent index to time-specific thresholds \(\kappa_{t,j}\). Part (4) of Proposition (ref) imposes structured restrictions on these parameters. The location parameters \(\eta_{gt}\) satisfy a weighted “parallel trends”-type assumption, while the scale parameters \(\lambda_{gt}\) satisfy a proportional “parallel trends”-type assumption. Specifically, the change in the latent mean for the treated group, $\eta_{11}-\eta_{10}$, is linked to the corresponding change for the control group, $\eta_{01}-\eta_{00}$, after adjustment by the scaling parameters, which act as tailored weighting factors. The weights adjust for components of common time trends that cannot be removed by simple within-group mean differencing. If time trends affect exclusively the location parameters, so that \(\lambda_{g0}=\lambda_{g1}\) for all \(g\), these restrictions reduce to the conventional parallel trends assumption and nest standard ordered Logit or Probit DiD specifications.

Under Assumption (ref) with $\alpha=0$, so that $C_{gt}(d)=Y_{gt}(d)$ with probability one, Assumption (ref) implies that the sharp bounds for the counterfactual distribution $F_{Y_{11}(0)}$ derived in Theorem 4.1 of AI2006 apply directly in this benchmark case:

equation[equation omitted — 259 chars of source]

for all $y\in\left[\underline{y}_{01},\overline{y}_{01}\right]$, where $\underline{y}_{01}:=\inf \mathrm{supp}\!\left(Y_{01}(0)\right)$ and $\overline{y}_{01}:=\sup \mathrm{supp}\!\left(Y_{01}(0)\right)$. Moreover, $F_{Y_{11}(0)}(y)=0$ for all $y<\underline{y}_{01}$ and $F_{Y_{11}(0)}(y)=1$ for all $y>\overline{y}_{01}$.

Bounds for Underreported Outcomes

We now allow for underreporting by considering $\alpha\ge 0$, relaxing the benchmark case $\alpha=0$ in which $Y(d)=C(d)$ a.s. Introducing underreporting complicates identification and alters the bounds. In particular, when $\alpha=1$, i.e., when no restriction is imposed on underreporting, the CiC bounds in ((ref)) become uninformative in the sense that $0 \le F_{Y_{11}(0)}(y) \le 1$ for all $y\in\mathcal{J}\setminus \{J\}$ and $F_{Y_{11}(0)}(J)=1$. To possibly tighten these bounds, we introduce an exclusion restriction on an observable instrument that does not directly affect consumption behavior. In what follows, we exploit this instrument together with the underreporting parameter $\alpha$, as the two will play distinct roles in tightening the bounds.

assumption[Exclusion Restriction] Let $\mathcal{Z}_{gt}:=\mathrm{supp}\left(Z_{gt}\right)$. For every $y\in\overline{\mathbb{R}}$, $z\in\mathcal{Z}_{gt}$, $(g,t)\in\{0,1\}\times\{0,1\}$, and $d=g\cdot t$,\[ \mathrm{Pr}\!\left(Y_{gt}(d)\le y \mid Z_{gt}=z\right) = \mathrm{Pr}\!\left(Y_{gt}(d)\le y\right). \]

Under the observability rule in (ref), the variables $(C,Z,G,T,D)$ are observed. Since treatment status satisfies $D=G\cdot T$, in each $(g,t)$ cell only the potential outcome corresponding to the realized treatment status, $Y_{gt}(d)$ with $d=g\cdot t$, is relevant for identification. The exclusion restriction is therefore imposed on $Y_{11}\left(1\right)$ in the treated post-treatment cell and on $Y_{gt}\left(0\right)$ in the remaining cells. This feature reflects the combination of underreporting and the DiD treatment assignment structure.

Assumption (ref) is used to tighten the nonparametric bounds and requires that $Z$ be independent of the true potential outcome $Y(d)$, so that any dependence between $Z$ and the observed outcome $C$ arises from reporting behavior rather than from an effect of $Z$ on true consumption. In our empirical setting, a natural example is a survey cooperation indicator measured after consumption has occurred. As such, this indicator is plausibly related to reporting behavior but is assumed to have no direct effect on the true consumption of marijuana. Finally, note that, if one were concerned that a single instrument does not satisfy the exclusion restriction for both potential outcomes, it would in principle be possible to use different instrumental variables for different potential outcomes.

Under Assumptions (ref) and (ref) alone, that is, without imposing any CiC structure across groups and time, the exclusion restriction yields partial identification of the distribution of the true outcome within each \((g,t)\) cell. In particular, as shown in Proposition (ref) in the supplement, for the true potential outcome distribution \(F_{Y_{gt}(d)}\) with \(d=g\cdot t\), the sharp bounds are given by\footnote{Remark (ref) in the Supplement shows that these bounds can also be combined with a standard parallel trends restriction on the true untreated outcomes to obtain average treatment effect on the treated (ATT).}

align[align omitted — 113 chars of source]

for all \(\alpha\in\left[\alpha_{gt}^*,1\right]\), where

align[align omitted — 568 chars of source]

The upper bound \(U_{gt}\) exploits variation in the instrument by taking the infimum of the conditional CDF of \(C_{gt}\) across values of \(Z_{gt}\). When \(Z_{gt}\) has no variation, this bound coincides with the unconditional CDF \(F_{C_{gt}}\). The lower bound $L_{gt,\alpha}$ captures worst-case one-sided underreporting. When $\alpha=1$, it collapses to a degenerate distribution concentrated at $J$, rendering the lower bound uninformative. Absent CiC structure, these bounds complement and extend recent results in MW2024 on underreporting in binary choice models to general ordered outcomes.

Finally, to understand why underreporting must exceed a threshold $\alpha_{gt}^*$, consider $\alpha=0$, which corresponds to truthful reporting. Under Assumption (ref) and the law of total probability, $F_{C_{gt}\mid Z_{gt}}(j\mid z)=\Pr(Y_{gt}(d)\le j\mid Z_{gt}=z)$ and hence $F_{C_{gt}}(j)=\Pr(Y_{gt}(d)\le j)$. By Assumption (ref), the distribution of $Y_{gt}(d)$ does not depend on $Z_{gt}$, implying $F_{C_{gt}\mid Z_{gt}}(j\mid z)=F_{C_{gt}}(j)$. Hence, in the presence of a valid instrument, truthful reporting has a strong and testable implication: the reported outcome must be independent of the instrument. Any observed dependence rules out $\alpha=0$ and implies some degree of underreporting. More generally, values of $\alpha$ below $\alpha_{gt}^*$ are incompatible with the joint distribution of $(C_{gt},Z_{gt})$.

Combining Assumptions (ref)--(ref), we obtain the following result:

proposition[Bounds for $F_{Y_{11}(0)}$] Suppose Assumptions (ref)--(ref) hold for some $\alpha\in[0,1]$. If $ \alpha^*\leq \alpha\leq 1$, then \begin{align} L_{11,\alpha}^{(0)}\left(y\right)\leq F_{Y_{11}(0)}(y)\leq U_{11,\alpha}^{(0)}\left(y\right),\quad y\in\overline{\mathbb{R}}, \end{align} where \begin{align} \begin{aligned} &\alpha^*=\sup_{\left(g,t\right)\in \{0,1\}\times\{0,1\}}\alpha^*_{gt},\\ & L_{11,\alpha}^{(0)}\left(y\right) := \mathbb{I}\!\left\{y_{L,01}\le y\leq \overline{y}_{L,01}\right\}\, L_{10,\alpha}\!\left( U_{00}^{(-1)}\!\left(L_{01,\alpha}(y)\right) \right) +\mathbb{I}\{y> \overline{y}_{L,01}\},\\ & U_{11,\alpha}^{(0)}\left(y\right) := \mathbb{I}\!\left\{y_{U,01}\le y\leq \overline{y}_{U,01}\right\}\, U_{10}\!\left( L_{00,\alpha}^{-1}\!\left(U_{01}(y)\right) \right) +\mathbb{I}\{y> \overline{y}_{U,01}\}, \end{aligned} \end{align} with $\underline{y}_{L,01}:=\inf \mathrm{supp}(L_{01,\alpha})$, $\overline{y}_{L,01}:=\sup \mathrm{supp}(L_{01,\alpha})$, $\underline{y}_{U,01}:=\inf \mathrm{supp}(U_{01})$, $\overline{y}_{U,01}:=\sup \mathrm{supp}(U_{01})$, and $\alpha^*_{gt}$, $L_{gt,\alpha}$ and $U_{gt}$ defined in (ref).\footnote{For a generic distribution function \(F\), we use \(\mathrm{supp}(F)\) to denote the support of a random variable with distribution function \(F\). A formal definition is provided in Definition (ref) in the supplement.} The bounds are functionally sharp if $\inf _{z \in \mathcal{Z}_{00}} \mathrm{Pr}\left(C_{00}=j \mid Z_{00}=z\right)>0$ for all $j\in \mathcal{J}$ and $\alpha<F_{C_{00}}(0)$. Moreover, if $0\leq \alpha< \alpha^*$, then there exists no data-generating process that is compatible with the maintained assumptions and yields the observed distributions.

Proposition (ref) characterizes sharp bounds on the counterfactual distribution \(F_{Y_{11}(0)}\) under the discrete CiC structure in the presence of underreporting. The bounds combine the CiC structure on the true potential outcome \(Y(0)\) with the one-sided misreporting constraint in Assumption (ref) and the exclusion restriction in Assumption (ref). As in the no-misreporting case discussed in Section (ref), the counterfactual distribution \(F_{Y_{11}(0)}\) is truncated outside the support of $Y_{01}(0)$, as implied by the support overlap condition $\mathrm{supp}\!\left(Y_{11}(0)\right)\subseteq \mathrm{supp}\!\left(Y_{01}(0)\right)$ under Assumption (ref). With underreporting, however, this truncation is governed by the bounding distributions $L_{01,\alpha}$ and $U_{01}$.

The exclusion restriction enters through the operators $U_{gt}$ and tightens the bounds by exploiting variation in $Z_{gt}$. At the same time, the bounds depend on the sensitivity parameter $\alpha$: larger values widen the bounds, and when $\alpha=1$ they become uninformative even with the instrument, since $L_{gt,\alpha}$ collapses to a degenerate distribution concentrated at $J$, so that $L_{gt,\alpha}^{-1}(q)=J$ for all $q\in[0,1]$. By contrast, smaller values of $\alpha$ preserve informative variation in $L_{gt,\alpha}$ and yield tighter restrictions on $F_{Y_{11}(0)}$.

The proposition also establishes functional sharpness under mild conditions, ensuring that the lower and upper bound functions are attainable as counterfactual distributions under the maintained assumptions. Here, these conditions ensure $\operatorname{supp}\!\left(L_{10,\alpha}\right)\subseteq \operatorname{supp}\!\left(U_{00}\right)$ and $\operatorname{supp}\!\left(U_{10}\right)\subseteq \operatorname{supp}\!\left(L_{00,\alpha}\right)$, which parallel the support overlap condition in the no-misreporting case, where $\mathrm{supp}\!\left(Y_{10}(0)\right)\subseteq \mathrm{supp}\!\left(Y_{00}(0)\right)$ under Assumption (ref). In our application, these conditions are mild: all outcome levels occur with positive probability for each instrument value across all $(g,t)$ cells, including the control group prior to treatment. Moreover, the mass at zero exceeds $90\%$, so that the restriction $\alpha<F_{C_{00}}(0)$ is not binding in economically relevant ranges and still permits substantial underreporting.

Finally, the restriction $\alpha \ge \alpha^*$ ensures joint feasibility of the cellwise restrictions under the DiD structure. When $\alpha<\alpha^*$, the observed distributions violate the exclusion restriction and no compatible data-generating process exists. Importantly, this restriction is entirely driven by variation in the instrument: if $Z$ is degenerate and exhibits no variation, then $\alpha^*=0$ and the restriction disappears.

Given the bounds for the factual CDF $F_{Y_{11}(1)}$ in (ref) and the counterfactual CDF $F_{Y_{11}(0)}$ in (ref), we construct bounds (and later confidence bands) for the QTTs defined in (ref). To this end, we invert the bounds for the distribution functions of $Y_{11}(0)$ and $Y_{11}(1)$ using the generalized inverses defined at the end of Section (ref). Because generalized inverses reverse the ordering of distribution functions, lower and upper bounds for the CDFs induce upper and lower bounds, respectively, for the corresponding quantile functions. These quantile bounds in turn imply bounds for $\Delta_{\mathrm{QTT}}(\tau)$ obtained via the pointwise Minkowski difference of the corresponding quantile bounds.\footnote{ If $V=[v_1,v_2]$ and $U=[u_1,u_2]$ are intervals, their pointwise Minkowski difference is \( V\ominus U := [v_1-u_2,\, v_2-u_1]. \) } This construction parallels the general quantile-effect framework of CFMW2020.

The resulting bounds are summarized in the following corollary.

corollary[Bounds for QTT] Suppose Assumptions (ref)--(ref) hold for some $\alpha\in\left[\alpha^*,1\right]$. Then $$ L_{\mathrm{QTT},\alpha}(\tau) \le \Delta_{\mathrm{QTT}}(\tau) \le U_{\mathrm{QTT},\alpha}(\tau),\quad \tau\in(0,1), $$ where $$ L_{\mathrm{QTT},\alpha}\left(\tau\right) := U_{11}^{(1),-1}\left(\tau\right) - L_{11,\alpha}^{(0),-1}\left(\tau\right), \qquad U_{\mathrm{QTT},\alpha}\left(\tau\right) := L_{11,\alpha}^{(1),-1}\left(\tau\right) - U_{11,\alpha}^{(0),-1}\left(\tau\right), $$ with $L_{11,\alpha}^{(1)} := L_{11,\alpha}$ and $U_{11}^{(1)} := U_{11}$, and $L_{11,\alpha}$, $U_{11}$, $L_{11,\alpha}^{(0)}$, and $U_{11,\alpha}^{(0)}$ defined in (ref) and (ref).

Corollary (ref) establishes bounds for the QTTs defined in (ref). The lower bound subtracts the upper quantile bound of $Y_{11}(0)$ from the lower quantile bound of $Y_{11}(1)$, while the upper bound subtracts the lower quantile bound of $Y_{11}(0)$ from the upper quantile bound of $Y_{11}(1)$.

Estimation and Inference

In this section, we describe how to estimate the bounds from Section (ref) and conduct inference on them. In Section (ref) of the supplement, we outline how the bounds and corresponding estimators can be extended to allow for additional exogenous covariates.

For estimation and inference, we rely on repeated cross-sectional sampling. Let the sample $\{(C_i, Z_i, G_i, T_i)\}_{i=1}^N$ be i.i.d.\ draws from the population, where $G_i, T_i \in \{0,1\}$ denote group and time indicators. Within each cell $(g,t) \in \{0,1\}\times\{0,1\}$, the observations with $G_i = g$ and $T_i = t$ are random draws from the corresponding subpopulation. We re-index these observations as $\{(C_{i,gt}, Z_{i,gt})\}_{i=1}^{N_{gt}}$, where $N_{gt} := \sum_{i=1}^N \mathbb{I}\{G_i = g, T_i = t\}$ denotes the number of observations in cell $(g,t)$.

To construct estimators of the bounds, consider $\alpha\in [\alpha^*,1)$, with $\alpha^*$ defined in (ref). We exclude the endpoint $\alpha=1$ and focus on the interior case for ease of exposition. For each $y\in\mathcal{J}$, define

align[align omitted — 225 chars of source]

where $L_{00,\alpha}\left(y\right)=0$ for all $y<0$ by definition. In the spirit of Section 5.2 of AI2006, Lemma (ref) in the supplement shows that, under mild regularity conditions, the bounds in (ref) admit an equivalent representation: for each $y\in\mathcal{J}$,

align[align omitted — 480 chars of source]

Thus, the bounds on $F_{Y_{11}(0)}$ can be obtained by first mapping $C_{10}$ through $\overline{k}_{\alpha}$ or $\underline{k}_{\alpha}$ and then evaluating the bound functionals in (ref) at the resulting distribution functions. Replacing the probabilities in (ref) by their empirical counterparts yields the following estimators:

align[align omitted — 768 chars of source]

where \[ \widehat{\overline{k}}_{\alpha}(y) := \widehat{L}_{01,\alpha}^{-1}\!\left(\widehat{U}_{00}(y)\right), \qquad \widehat{\underline{k}}_{\alpha}(y) := \widehat{U}_{01}^{-1}\!\left(\widehat{{L}}_{00,\alpha}(y-1)\right), \] and for each $(g,t)\in \{(0,1),(0,0)\}$,

align*[align* omitted — 500 chars of source]

We set $\widehat{U}_{11,\alpha}^{(0)}(y)$ and $\widehat{U}_{gt}(y)$ equal to one whenever their denominators are zero.

We impose the following sampling and support conditions:

assumption[Sampling] $\{(C_i, Z_i, G_i, T_i)\}_{i=1}^N$ are i.i.d.\ copies of $(C, Z, G, T)$.
assumption[Discrete Instrument] $\mathcal{Z}_{gt}$ is finite for all $(g,t)\in\{0,1\}\times\{0,1\}$.
assumption[Regularity] There exists a nonempty set $\mathcal{A}\subseteq [\alpha^*,1)$ such that for all $\alpha\in\mathcal{A}$: \begin{enumerate}[label=(\roman*)] • $U_{01}(y)\neq L_{00,\alpha}(y')$ and $L_{01,\alpha}(y)\neq U_{00}(y')$ for all $y,y'\in\mathcal{J}\setminus \{J\}$. • $\inf _{y\in\mathcal{J}, z \in \mathcal{Z}_{00}} \mathrm{Pr}\left(C_{00}=y \mid Z_{00}=z\right)>0$ and $\alpha<F_{C_{00}}(0)$. \end{enumerate}

Assumption (ref) is a standard sampling condition for repeated cross-sectional data. Combined with the maintained condition $\pi_{gt}=\Pr(G=g,T=t)>0$ for all $(g,t)$, this implies $N_{gt}>0$ with probability approaching one.

Assumption (ref) requires that the instrumental variable has finite support across all cells. Allowing for a continuous instrument would introduce additional technical complications for inference MP2020,MPZ2024. In applications where $Z_{gt}$ is continuous or takes on many values, it can be discretized on a finite grid without altering the structure of the estimators (see Section (ref)).

Assumption (ref)(i) imposes a no–tie condition and adapts Assumption 5.2 in AI2006 to our setting. The set $\mathcal{A}\subseteq [\alpha^*,1)$ indexes the sensitivity parameters over which the bounds are evaluated, typically a finite grid chosen by the researcher. For every $\alpha\in\mathcal{A}$, the assumption requires that the lower and upper bound functions do not coincide across subpopulations. This rules out boundary cases in which the bounds coincide and the estimator becomes nonregular.\footnote{ As discussed in AI2006, a similar issue arises for the sample median of i.i.d.\ binary variables $X_i \in \{0,1\}$, which converges to the population median only if $\Pr(X_i = 1) \neq 0.5$. At the boundary case $\Pr(X_i = 1) = 0.5$, the estimator fails to converge.}

Assumption (ref)(ii) coincides with the condition for sharpness of the bounds in Proposition (ref). Here, it ensures $\operatorname{supp}(C_{10})\subseteq \operatorname{supp}(L_{00,\alpha})$ and $\operatorname{supp}(C_{10})\subseteq \operatorname{supp}(U_{00})$, which are analogous to Assumption 5.1(iv) in AI2006. With underreporting, the CiC transformation is constructed from the bounding distributions $L_{00,\alpha}$ and $U_{00}$ rather than from a single observed control distribution. Since these distributions may differ and, in the case of $L_{00,\alpha}$, depend on $\alpha$, the support inclusion must hold simultaneously for both distributions entering the bounds.

We establish weak convergence of the bound estimators for the counterfactual distribution $F_{Y_{11}(0)}$, uniformly over $y \in \mathcal{J}$.

proposition[Convergence of Bound Estimators for $F_{Y_{11}(0)}$] Suppose Assumptions (ref)--(ref) hold. Then, for each $\alpha\in\mathcal{A}$, \[ \sqrt{N} \begin{pmatrix} \widehat{L}_{11,\alpha}^{(0)} - L_{11,\alpha}^{(0)} \\ \widehat{U}_{11,\alpha}^{(0)} - U_{11,\alpha}^{(0)} \end{pmatrix} \rightsquigarrow \begin{pmatrix} G_{L,11,\alpha}^{(0)} \\ G_{U,11,\alpha}^{(0)} \end{pmatrix}\quad \text{in } \ell^{\infty}\!\left(\mathcal{J},\mathbb{R}^2\right), \] with $L_{11,\alpha}^{(0)}$, $U_{11,\alpha}^{(0)}$, $\widehat{L}_{11,\alpha}^{(0)}$, and $\widehat{U}_{11,\alpha}^{(0)}$ defined in (ref) and (ref), and $G_{L,11,\alpha}^{(0)}$ and $G_{U,11,\alpha}^{(0)}$ are limit processes defined in Lemma (ref) of the supplement.\footnote{Let $\mathscr{A}$ be an arbitrary set and $\mathscr{B}$ a Banach space. Then $\ell^{\infty}(\mathscr{A},\mathscr{B})$ denotes the space of all bounded functions $f:\mathscr{A}\to\mathscr{B}$ such that $\sup_{a\in\mathscr{A}}\|f(a)\|_{\mathscr{B}}<\infty$, equipped with the sup-norm $\|f\|_{\infty} := \sup_{a\in\mathscr{A}}\|f(a)\|_{\mathscr{B}}$ (see, e.g., VanderVaart1996).}

Proposition (ref) can be reformulated as a directional delta-method result for the bound estimators. To see this, let

align[align omitted — 320 chars of source]

and define the functional $ \boldsymbol{\phi}_{\alpha}^{(0)} : \ell^{\infty}\!\left(\mathcal{J},\mathbb{R}\right) \times \ell^{\infty}\!\left(\mathcal{J}\times\mathcal{Z}_{10},\mathbb{R}\right) \to \ell^{\infty}\!\left(\mathcal{J},\mathbb{R}^{2}\right) $ by

align[align omitted — 278 chars of source]

Then, by the representation in (ref), \[

pmatrix[pmatrix omitted — 57 chars of source]

= \boldsymbol{\phi}_{\alpha}^{(0)}\!\left(\theta_{L,\alpha}^{(0)},\theta_{U,\alpha}^{(0)}\right)\in \ell^{\infty}(\mathcal{J},\mathbb{R}^{2}). \] As shown in the supplement, the weak limit in Proposition (ref) is obtained by applying the Hadamard directional derivative of $\boldsymbol{\phi}_{\alpha}^{(0)}$ at $(\theta_{L,\alpha}^{(0)},\theta_{U,\alpha}^{(0)})$ to the weak limit of the underlying empirical distribution functions. Although the latter limit is Gaussian, the resulting limit process is generally non-Gaussian and non-pivotal because $\boldsymbol{\phi}_{\alpha}^{(0)}$ is only Hadamard directionally differentiable: its first component is kinked at $\theta_{1}(y)=\alpha$, and its second component involves a pointwise infimum.

We therefore follow FS2019 and MP2020 and use a bootstrap procedure for directionally Hadamard differentiable functionals, implemented via the numerical directional derivative of HL2018. For $\left(g,t\right)=\left(1,0\right)$, let $\left\{\left(C_{i,10}^{*},Z_{i,10}^{*}\right)\right\}_{i=1}^{N_{10}}$ denote a generic nonparametric i.i.d.\ bootstrap sample drawn with replacement from $\left\{\left(C_{i,10},Z_{i,10}\right)\right\}_{i=1}^{N_{10}}$. For each bootstrap sample, the resampled observations are transformed using the mappings $\widehat{\underline{k}}_{\alpha}$ and $\widehat{\overline{k}}_{\alpha}$ obtained from the original sample. As shown in Lemma (ref) in the supplement, with probability approaching one, \[ \widehat{\underline{k}}_\alpha(y)=\underline{k}_\alpha(y), \qquad \widehat{\overline{k}}_\alpha(y)=\overline{k}_\alpha(y), \qquad y\in\mathcal{J}. \] Consequently, recomputing these transformations within the bootstrap sample would not affect first-order asymptotics, so we keep them fixed in the bootstrap procedure.

To construct empirical and bootstrap counterparts of (ref), denote

align*[align* omitted — 1,052 chars of source]

for all $y\in\mathcal{J}$ and $z\in\mathcal{Z}_{10}$. With this notation, define \[ \left(\widehat{\theta}_{L,\alpha}^{(0)},\widehat{\theta}_{U,\alpha}^{(0)}\right) := \left( \widehat{F}_{\widehat{\overline{k}}_{\alpha}(C_{10})}, \widehat{F}_{\widehat{\underline{k}}_{\alpha}(C_{10})\mid Z_{10}} \right), \qquad \left(\widehat{\theta}_{L,\alpha}^{(0),*},\widehat{\theta}_{U,\alpha}^{(0),*}\right) := \left( \widehat{F}^{\,*}_{\widehat{\overline{k}}_{\alpha}(C_{10})}, \widehat{F}^{\,*}_{\widehat{\underline{k}}_{\alpha}(C_{10})\mid Z_{10}} \right). \] To implement the numerical directional derivative, define for $h_L \in \ell^{\infty}\!\left(\mathcal{J},\mathbb{R}\right)$ and $h_U \in \ell^{\infty}\!\left(\mathcal{J}\times\mathcal{Z}_{10},\mathbb{R}\right)$ \[ \widehat{\boldsymbol{\phi}}_{\alpha}^{(0)'}\left(h_L,h_U\right) := \frac{ \boldsymbol{\phi}_{\alpha}^{(0)}\!\left( \widehat{\theta}_{L,\alpha}^{(0)}+\epsilon_N h_L, \widehat{\theta}_{U,\alpha}^{(0)}+\epsilon_N h_U \right) - \boldsymbol{\phi}_{\alpha}^{(0)}\!\left( \widehat{\theta}_{L,\alpha}^{(0)}, \widehat{\theta}_{U,\alpha}^{(0)} \right) }{\epsilon_N}, \] where $\epsilon_N \to 0$ and $\sqrt{N}\epsilon_N \to \infty$. We then approximate the limiting distribution of $ \sqrt{N}\left( \boldsymbol{\phi}_{\alpha}^{(0)}\left(\widehat{\theta}_{L,\alpha}^{(0)},\widehat{\theta}_{U,\alpha}^{(0)}\right) - \boldsymbol{\phi}_{\alpha}^{(0)}\left(\theta_{L,\alpha}^{(0)},\theta_{U,\alpha}^{(0)}\right) \right) $ in Proposition (ref) by the distribution of the bootstrap process \[

pmatrix[pmatrix omitted — 59 chars of source]

:= \widehat{\boldsymbol{\phi}}_{\alpha}^{(0)'}\!\left( \sqrt{N}\left( \widehat{\theta}_{L,\alpha}^{(0),*}-\widehat{\theta}_{L,\alpha}^{(0)} \right), \sqrt{N}\left( \widehat{\theta}_{U,\alpha}^{(0),*}-\widehat{\theta}_{U,\alpha}^{(0)} \right) \right) \in \ell^{\infty}\!\left(\mathcal{J},\mathbb{R}^{2}\right). \] This construction corresponds to the numerical directional derivative proposed by HL2018. Under the maintained conditions, Theorem 3.1 of HL2018 implies that the bootstrap consistently estimates the limiting distribution in Proposition (ref).

To construct a symmetric $100\cdot\left(1-\gamma\right)\%$ confidence band that is uniform over $j\in\mathcal{J}\setminus\{J\}$, define the functional $m: \ell^{\infty}\!\left(\mathcal{J},\mathbb{R}^{2}\right) \to \mathbb{R}$ by

equation[equation omitted — 260 chars of source]

where $\sigma(j)>0$ is a known, bounded weighting function. Define the bootstrap statistic and its conditional $\left(1-\gamma\right)$-quantile:

align*[align* omitted — 541 chars of source]

As outlined in FR2018, many choices of $\sigma$ yield valid uniform confidence bands. In our application, we use the simple choice $\sigma(j)=1$, which produces an equal-width band. This yields an asymptotically valid $100\cdot\left(1-\gamma\right)\%$ confidence band for the counterfactual bounds, uniform over $y\in\mathcal{J}\setminus\{J\}$, defined by

align[align omitted — 307 chars of source]

for $y\in\mathcal{J}\setminus\{J\}$, while we set $\mathrm{CS}_{\alpha}^{(0)}(J;1-\gamma):=\{1\}$. Since $L_{11,\alpha}^{(0)}(y),U_{11,\alpha}^{(0)}(y)\in[0,1]$ for all $y\in\mathcal{J}$ by definition, intersection with $[0,1]$ only enforces the natural range of a CDF and does not affect asymptotic coverage.

Proposition (ref) establishes the asymptotic validity of this $100\cdot\left(1-\gamma\right)\%$ confidence band uniformly over $y\in\mathcal{J}$.

proposition[Bootstrap Validity of Counterfactual Confidence Band] Suppose Assumptions (ref)--(ref) hold, and let $\epsilon_{N}\to 0$ satisfy $\sqrt{N}\epsilon_{N}\to \infty$. Let $\alpha\in\mathcal{A}$ and $\gamma\in\left(0,1\right)$. Define \[ M_{\alpha}^{(0)} :=m\!\left( \begin{pmatrix} G_{L,11,\alpha}^{(0)} \\ G_{U,11,\alpha}^{(0)} \end{pmatrix} \right), \quad z_{\alpha}^{(0)}\left(1-\gamma\right) := \inf\left\{ z\in\mathbb{R} \mid \mathrm{Pr}\left(M_{\alpha}^{(0)}\le z\right)\ge 1-\gamma \right\}, \] where $\left(G_{L,11,\alpha}^{(0)},G_{U,11,\alpha}^{(0)}\right)$ is the weak limit in Proposition (ref). Assume that the distribution function of $M_{\alpha}^{(0)}$ is continuous and strictly increasing in a neighborhood of $z_{\alpha}^{(0)}\left(1-\gamma\right)$. Then \[ \lim_{N\to\infty} \mathrm{Pr}\left( \left[ L_{11,\alpha}^{(0)}(y), U_{11,\alpha}^{(0)}(y) \right] \subseteq \mathrm{CS}_{\alpha}^{(0)}(y;1-\gamma) \ \text{for all } y\in\mathcal{J} \right) = 1-\gamma. \]

The continuity assumption on the distribution of $M_{\alpha}^{(0)}$ ensures that the population quantile $z_{\alpha}^{(0)}\left(1-\gamma\right)$ of $M_{\alpha}^{(0)}$ is well defined and that the bootstrap critical value $\widehat z_{\alpha}^{(0)}\left(1-\gamma\right)$ consistently estimates it, see Corollary 3.2 of FS2015. Moreover, the uniformity of the confidence band over $y \in \mathcal{J}$ yields a straightforward interpretation: with a pre-specified probability, for example $90\%$, it covers the counterfactual bound functions at all values of $y \in \mathcal{J}$ simultaneously. This is practically useful because it does not require the researcher to decide in advance which particular outcome level or functional feature is of interest. Hence, any candidate function that lies outside $\mathrm{CS}_{\alpha}^{(0)}(y;0.9)$ at even a single value of $y$ can be rejected at the corresponding level, namely $10\%$.

As with inference on the distribution bounds, we can also construct uniform confidence bands for the QTT bounds. Specifically, we invert the confidence bands for $\left[L_{11,\alpha}^{(1)},U_{11}^{(1)}\right]$ and $\left[L_{11,\alpha}^{(0)},U_{11,\alpha}^{(0)}\right]$, each at level $1-\gamma/2$, and then take the pointwise Minkowski difference of the resulting quantile bands, following CFMW2020. Let $\mathrm{CS}_{\mathrm{QTT},\alpha}\left(\tau;1-\gamma\right)$ denote the resulting confidence band, whose construction is given in the supplement.

We obtain the following corollary:

corollary[Bootstrap Validity of QTT Confidence Band] Suppose Assumptions (ref)--(ref) hold, and let $\epsilon_{N}\to 0$ satisfy $\sqrt{N}\epsilon_{N}\to \infty$. Let $\alpha\in\mathcal{A}$ and $\gamma\in\left(0,1\right)$. Assume that the distribution function of $M_{\alpha}^{(d)}$ is continuous and strictly increasing in a neighborhood of $z_{\alpha}^{(d)}\left(1-\gamma/2\right)$ for all $d\in \{0,1\}$, with $M_{\alpha}^{(0)}$ and $z_{\alpha}^{(0)}$ defined in Proposition (ref), and $M_{\alpha}^{(1)}$ and $z_{\alpha}^{(1)}$ in the supplement. Then \[ \liminf_{N\to\infty} \Pr\!\left( \left[ L_{\mathrm{QTT},\alpha}\left(\tau\right), U_{\mathrm{QTT},\alpha}\left(\tau\right) \right] \subseteq \mathrm{CS}_{\mathrm{QTT},\alpha}\left(\tau;1-\gamma\right) \text{ for all }\tau\in(0,1) \right) \ge 1-\gamma, \] with $L_{\mathrm{QTT},\alpha}$ and $U_{\mathrm{QTT},\alpha}$ defined in Corollary (ref), and $\mathrm{CS}_{\mathrm{QTT},\alpha}$ in the supplement.

Empirical Application

The intense debate surrounding the legalization of marijuana (cannabis) and its impact on individual consumption behavior, and more generally, on society, has been a pivotal policy issue in Western countries for decades. In many of these nations, the possession and consumption of marijuana are illegal, primarily due to concerns over its contribution to crime escalation and the associated societal and economic costs BML2019. Additionally, there are health-related concerns, with evidence suggesting that early-age cannabis consumption can lead to lower educational attainment VOW2009, MZLBHB2007, MOCCEHOSS2004 and cognitive deficits in verbal learning and memory tasks SJRDCHL2011. The potential of cannabis serving as a gateway to more harmful drug use has also been documented P2010.

In this section, we investigate the effects of recreational marijuana legalization for adults in several U.S. states on the short-term consumption behavior of 8th-grade high-school students in those states during our sample period. The focus on these minors, typically aged 13 to 14, is critical due to their heightened vulnerability to the health risks associated with marijuana consumption. Unlike existing studies, which predominantly examined legalization effects on average usage or consumption AHR2015, WHC2015, CWFKSSH2017, HWB2022, or on the timing of the first (early age) consumption WBJ2014, we investigate the distributional effects of recreational marijuana legalization on the entire consumption distribution within the aforementioned high-school cohort in various states.\footnote{Unlike HWB2022, who consider mean effects for the legalization of recreational marijuana use on underage adolescents among others, the survey data used in this analysis exclusively cover enrolled high-school students (see next paragraph). Thus, they do not provide insights into effects on high-school drop-outs.} Additionally, we account for the possibility that consumption behavior may not always be reported truthfully, and more importantly, that reporting behavior itself may be affected by the legalization, possibly due to shifts in stigma, public perception, or attitudes towards marijuana consumption.

The data are composed of several cross-sectional waves of the “Monitoring the Future” survey, an annual survey conducted in the United States gathering data on attitudes, behaviors, and values of American adolescents currently enrolled in high school.\footnote{The “Monitoring the Future” survey is funded by the National Institute on Drug Abuse (NIDA). For more information, see Monitoring the Future, Inter-university Consortium for Political and Social Research (ICPSR), University of Michigan (https://monitoringthefuture.org/).} Specifically, the survey for 8th graders, which is conducted annually, covers a wide range of topics including substance (ab)use such as the consumption of marijuana and variants thereof. The primary sample of our empirical analysis consists of 46,472 high-school students, sampled as repeated cross-sections across all contiguous U.S. states. With this sample, our focus is on recreational marijuana legalization for adults aged 21 or older, examining its impact on 8th-grade students' (short-term) consumption behavior. We use observations from states that legalized marijuana for adults during our sampling period as treated group ($g=1$), and label observations as coming from the control group otherwise ($g=0$).\footnote{Due to data privacy restrictions of the ICPSR, we are not allowed to publish information that allows to identify single states either directly or indirectly. Information about the specific legalization year used in the analysis as well as about the states that entered the treatment group can be obtained from the authors upon request.} Similarly, we record observations from the two years before the legalization event as coming from the pre-treatment period ($t=0$), while observations from the two years thereafter are labeled as post-treatment observations ($t=1$). We exclude observations from states that had already passed recreational marijuana laws prior to the sampling window. Moreover, all states from the treated group had previously legalized medicinal marijuana use before the observation period. It is also important to note that we focus on a specific legalization year, and no other major marijuana-related legalization, either for medicinal or recreational purposes, took place in the U.S. during that period. Finally, given that surveys were conducted in the spring of each year, while recreational marijuana legalization was enacted in the last quarter for all treated states considered in the analysis, we deem the risk of major anticipation effects to be small. The final subsamples consist of $N_{00}=19940$, $N_{01}=17769$, $N_{10}=4710$, and $N_{11}=4053$ observations.

The outcome variable of interest (cons_30days), based on the survey responses to the question “On how many occasions have you used marijuana (weed, pot) or hashish (hash, hash oil) ... in the last 30 days?”, is categorized into three levels: “0” (never used), “1” (used 1--2 times), and “2” (used more than 2 times). Because this measure is based on students' self-reports of recent marijuana use, truthful reporting is not automatic in this setting. Legalization for adults may affect not only actual consumption but also the perceived cost of admitting use in a survey, so allowing treatment to affect reporting is empirically important for interpretation. We also use a discretized “survey cooperation” measure (trust) as an instrumental variable for reporting, following GHSZ2018 and BHSZ2018. Here, “survey cooperation”, defined as the proportion of missing survey responses and based solely on questions common to all questionnaire forms and not directly related to drug use, is assumed to provide an indication of the respondents' willingness to cooperate in the survey, unrelated to marijuana consumption itself, as the reporting decision occurs post-consumption only. We discretize this variable into four categories based on the 25th, 50th, and 75th quartiles of its unconditional distribution. Descriptive statistics for cons_30days, trust, and other control variables used in the analysis of Section (ref) can be found in Section (ref) of the supplementary material.

We first present estimated bounds for the CDF of the counterfactual and factual distribution for the case without and with underreporting. Specifically, the upper panel in Figure (ref) displays the estimated bounds from AI2006 for the counterfactual and the empirical CDF for the factual distribution for the case without misreporting. By contrast, the second, third, and fourth panels of the figure contain the estimated bounds $\widehat{L}_{11,\alpha}^{(0)}$, $\widehat{U}_{11,\alpha}^{(0)}$ for the counterfactual, and $\widehat{L}_{11,\alpha}$, $\widehat{U}_{11}$ for the factual distribution at values of $\alpha$ equal to $0.3$, $0.5$, and $0.7$, respectively. Note that values with $\alpha<0.3$ lead to a crossing of the bounds, suggesting that such values are not compatible with the observed distribution under the maintained exclusion restriction. In this sense, the data imply a minimal feasible degree of underreporting in the application. The shaded areas illustrate uniform 90% Confidence Sets, $\mathrm{CS}_{\alpha}^{(0)}(\cdot;0.9)$, as defined in ((ref)) for the counterfactual distributions, and in Subsection (ref) of the supplement for the factual distribution. We compute these Confidence Sets using the nonparametric bootstrap together with the numerical derivative method outlined in Section (ref) using $B=1000$ bootstrap replications.\footnote{We construct the Confidence Set for the case without underreporting using the same bootstrap method.} For simplicity, we display all results for the ordinary nonparametric bootstrap with $\epsilon_{N}=1$ only, as this value resulted in good finite sample performance in the Monte Carlo simulation of Section (ref) in the supplement. However, we note that results remain qualitatively similar when using larger values of $\epsilon_{N}$ as required by our theory (available upon request).

Turning to the results in Figure (ref), we first observe that $\alpha$ plays an important role in tightening the bounds for the counterfactual distribution, and that different values can lead to a large heterogeneity in the informativeness of the bounds. As expected, this holds true for the bounds of the factual distribution to a much lesser extent, as the latter are much tighter in general. Second, note that choosing a sufficiently small $\alpha$ level leads to bound estimates that are very similar to the case without misreporting. Third, we note that the bounds are estimated very precisely since the Confidence Sets are very tight throughout, which does not change much with larger values of $\epsilon_{N}$.

Next, we turn to the QTT estimates in Figure (ref). As before, we present these estimates alongside uniform 90% Confidence Sets, $\mathrm{CS}_{\mathrm{QTT},\alpha}(\cdot;0.9)$, as outlined in Corollary (ref). Moreover, as more than 90% of the 8th graders in the sample do not report consumption of marijuana in the past 30 days (cf. also Figure (ref)), we concentrate our analysis on quantile levels of the upper part of the distribution, namely $\tau\in(0.8,1)$. Examining Figure (ref), we observe that without misreporting there is only evidence for possibly non-zero QTTs at quantile levels larger than approximately $0.95$. By contrast, we can reject the null hypothesis of positive (or negative) treatment effects for smaller quantile levels at the 10% significance level. This suggests that, without accounting for misreporting, there is no evidence at the 10% significance level that legalization of marijuana had any impact on the short-term consumption frequency of 8th-graders in treated states. On the other hand, moving to the case with underreporting, we first note that, as expected, the evidence for potentially non-zero effects shifts downwards to lower quantile levels. For instance, at $\alpha=0.5$, we cannot reject either positive or negative QTTs for quantile levels between $0.86$ and $0.97$. Interestingly, both at $\alpha=0.5$ and $\alpha=0.7$, we can reject negative QTTs for the highest quantile levels, while positive (but not negative) QTTs can be rejected at $\alpha=0.5$ for quantile levels below $\tau=0.86$. This evidence is in line with our estimates from the parametric model in Section (ref) of the supplement, where allowing for misreporting leads to a significant, negative DTT at the zero consumption level, and to a small significant increase at the highest consumption frequency. By contrast, the estimates without underreporting are insignificant throughout.

Summarizing the empirical analysis above, there is little overall evidence of either positive or negative QTTs from marijuana legalization on the short-term consumption frequency of 8th-grade high-school students in treated states. Once underreporting is taken into account, however, the picture becomes more nuanced: at lower quantile levels, only positive QTTs are rejected, whereas at upper quantile levels, only negative QTTs can be rejected. This pattern is in line with the estimation results from the semiparametric model in Section (ref) of the supplement. In Section (ref) of the supplement, we also examine QTT bounds for subsamples of the primary dataset. In particular, we consider four subsamples defined by self-reported ethnicity and gender (white=1 if the student reported being white, white=0 otherwise; male=1 if the student reported being male, male=0 otherwise). The results remain qualitatively similar, especially for subsamples with white=1, but the bound estimates and Confidence Sets are generally much less informative due to substantially smaller sample sizes. Finally, as noted above, choosing larger values of $\epsilon_{N}$ yields wider and thus less informative Confidence Sets for the bounds under underreporting, while leaving the qualitative conclusions unchanged.

figure[figure omitted — 1,415 chars of source]
figure[figure omitted — 1,097 chars of source]

Conclusion

This paper develops a Difference-in-Differences framework for discrete, ordered outcomes that may be underreported. The analysis builds on the discrete CiC model of AI2006 and extends it to settings with partially observable outcomes. A key feature of our setup is that reporting behavior may be correlated with the underlying outcome and may also vary with treatment status. Under an exclusion restriction and an upper bound on underreporting, we derive nonparametric bounds that are functionally sharp, and we provide corresponding estimators and bootstrap-based uniform Confidence Sets. In the supplement, we complement this nonparametric analysis with a point-identified semiparametric model based on a real analytic extrapolation assumption, and we propose a flexible estimator for a parametric version of that model.

We apply our methodology to investigate QTTs of the legalization of recreational marijuana use for adults on the short-term consumption behavior of 8th-grade high-school students in the affected states. Our findings suggest little evidence of non-zero effects outside the extreme upper tail of the distribution when underreporting is ignored. Once underreporting is allowed for, however, only positive QTTs are ruled out at moderate upper-tail quantiles, whereas at the very top the reverse is true.

{1}