The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
46,728 characters
Testing the Significance of the Difference-in-Differences Coefficient via Doubly Randomised Inference
\def\spacingset#1{\renewcommand{\baselinestretch}
{#1}\small\normalsize} \spacingset{1}
\if00
{
\title{\bf Testing the Significance of the Difference-in-Differences Coefficient via Doubly Randomised Inference}
\bigskip
\author{Stanisław M. S. Halkiewicz \and Andrzej Kałuża\\
\hspace{.2cm}\\
AGH University of Cracow, Faculty of Applied Mathematics\\}
\maketitle} \fi
\if10
{
\bigskip
\bigskip
\bigskip
\begin{center}
{\LARGE\bf Testing the Significance of the Difference-in-Differences Coefficient via Doubly Randomised Inference}
\end{center}
\medskip
} \fi
\bigskip
\begin{abstract}
This article develops a significance test for the Difference-in-Differences (DiD) estimator based on doubly randomised inference, in which both the treatment and time indicators are permuted to generate an empirical null distribution of the DiD coefficient. Unlike classical $t$-tests or single-margin permutation procedures, the proposed method exploits a substantially enlarged randomization space. We formally characterise this expansion and show that dual randomization increases the number of admissible relabelings by a factor of $\binom{n}{n_T}$, yielding an exponentially richer permutation universe. This combinatorial gain implies a denser and more stable approximation of the null distribution, a result further justified through an information-theoretic (entropy) interpretation.
The validity and finite-sample behaviour of the test are examined using multiple empirical datasets commonly analysed in applied economics, including the Indonesian school construction program (INPRES), brand search data, minimum wage reforms, and municipality-level refugee inflows in Greece. Across all settings, doubly randomised inference performs comparably to standard approaches while offering superior small-sample stability and sharper critical regions due to the enlarged permutation space. The proposed procedure therefore provides a robust, nonparametric alternative for assessing the statistical significance of DiD estimates, particularly in designs with limited group sizes or irregular assignment structures.
\end{abstract}
\noindent
{\it Keywords:} Difference-in-Differences, \and econometrics, \and randomisation, \and causal inference, \and randomised inference
\noindent
{\it JEL Codes:} C12, \and C15, \and C21, \and C23, \and C31
\vfill
\newpage
\spacingset{1.75}
\section{Introduction}
\label{secIntro}
The Difference-in-Differences (DiD) framework, also referred to as Difference-in-Differences (DD) or Diff-in-Diff, has become one of the most influential quasi-experimental tools in applied econometrics and the social sciences \citep{Halkiewicz2024}. Its intuitive structure and relative simplicity make it a preferred method for evaluating causal effects of discrete interventions—such as legal reforms, natural disasters, or policy programs—when randomized experiments are infeasible. The fundamental logic of DiD is to compare the evolution of outcomes between a treatment group, affected by an intervention, and a control group, unaffected by it. By differencing both across time and across groups, the method aims to eliminate time-invariant unobserved heterogeneity and isolate the average treatment effect.
Since its formal introduction in applied economics by \citet{Card1994} and its systematic exposition in \citet{Angrist2008} and \citet{Cameron2005}, DiD has become a cornerstone of empirical policy evaluation. Applications span health sciences—where early uses included epidemiological analyses of disease propagation \citep{Caniglia2020}—education and social programs \citep{Domina2015, Furquim2020, Asadi2020, Halkiewicz2023}, as well as labor and public economics \citep{Callaway2022}. More recent studies have extended DiD to domains such as marketing \citep{Deng2019}, human resources \citep{Chen2016}, and production management \citep{Distelhorst2017}. Methodological advances have further expanded its scope, introducing heterogeneous treatment effects \citep{Fredriksson2019}, synthetic control and interactive fixed-effects extensions \citep{Arkhangelsky2021}, and refinements in statistical inference \citep{MacKinnon2020, Conley2011, Ferman2019}.
Despite its ubiquity, challenges remain in assessing the statistical significance and reliability of the estimated DiD parameter, especially in studies with limited group sizes or complex treatment assignment structures. Classical inference methods—such as $t$-tests with robust or clustered standard errors—often rely on asymptotic approximations that may be inaccurate in small samples or when key assumptions (parallel trends, homoskedasticity) are only weakly satisfied. These limitations raise important questions about how to properly evaluate the precision and robustness of DiD estimates in non-ideal empirical settings.
To address these issues, this paper introduces a simulation-based framework for significance testing of the DiD estimator. Our approach builds on the logic of randomization-based inference \citep{Fisher1935, Neyman1923, ImbensRubin2015}, but adapts it to the quasi-experimental context of Difference-in-Differences. Specifically, we propose a dual randomization scheme in which both the treatment and time vectors are permuted to generate a Monte Carlo distribution of the DiD coefficient. This distribution allows for non-parametric inference on treatment effects and provides more reliable significance measures in small or irregular panels.
The remainder of the paper is structured as follows. Section~\ref{sec:did} introduces the Difference-in-Differences framework, clarifies identification and interpretation, and highlights why seemingly similar setups may yield different estimates. Section~\ref{prelim} reviews the theoretical foundations of randomization-based inference and its adaptation to econometric applications. Section~\ref{sec:methods} formalizes the proposed simulation procedure and presents its statistical properties. Section~\ref{sec:numerics} validates the proposed significance test and its relevancy through numerical experiments. Finally, section~\ref{sec:conclusions} discusses implications for applied research and outlines directions for future methodological development.
\section{Difference-in-Differences}\label{sec:did}
The Difference-in-Differences (DiD) estimator measures the average effect of a treatment or intervention by comparing outcome changes over time between treated and untreated groups. Its central appeal lies in its simplicity: by differencing twice—across time and across groups—it removes time-invariant unobservables and isolates the component of change attributable to treatment. The standard two-period, two-group setup can be summarized by the following model:
\begin{equation}
\label{eq:did-basic}
Y_{it} = \alpha + \beta \, \text{TIME}_t + \gamma \, \text{AFFECTED}_i + \delta \, (\text{TIME}_t \times \text{AFFECTED}_i) + \varepsilon_{it},
\end{equation}
where $Y_{it}$ is the outcome variable for unit $i$ at time $t$, $\text{TIME}_t$ indicates the post-treatment period, $\text{AFFECTED}_i$ marks the treated group, and $\delta$ represents the DiD estimator—the estimated average treatment effect on the treated (ATT).
Under the \emph{parallel trends assumption}, i.e.
\[
\mathbb{E}[Y_{it}(0) - Y_{i,t-1}(0)\,|\,\text{AFFECTED}_i=1]
= \mathbb{E}[Y_{it}(0) - Y_{i,t-1}(0)\,|\,\text{AFFECTED}_i=0],
\]
the parameter $\delta$ consistently identifies the causal effect of treatment. Violations of this assumption—such as pre-existing differential trends, heterogeneous shocks, or group-specific dynamics—can lead to biased estimates.
\subsection{Alternative Formulations and Interpretation}
Equation~\eqref{eq:did-basic} can be estimated using Ordinary Least Squares (OLS), or equivalently derived from a set of mean comparisons. In its simplest form, the DiD estimator is:
\begin{equation}
\label{eq:did-formula}
\widehat{\delta}_{DiD} =
(\bar{Y}_{1,1} - \bar{Y}_{1,0}) - (\bar{Y}_{0,1} - \bar{Y}_{0,0}),
\end{equation}
where $\bar{Y}_{g,t}$ denotes the sample mean for group $g \in \{0,1\}$ at time $t \in \{0,1\}$.
The first difference $(\bar{Y}_{1,1} - \bar{Y}_{1,0})$ captures the temporal change in the treated group, and the second $(\bar{Y}_{0,1} - \bar{Y}_{0,0})$ measures the analogous change for the control group.
The DiD estimate, being the difference between these two quantities, reflects the excess change attributable to the treatment.
This simple expression hides several subtleties. When treatment assignment is not perfectly exogenous or when timing varies across units, the estimate may no longer represent a single, homogeneous effect. For example, if treated and untreated groups experience unequal exposure to macroeconomic shocks or policy spillovers, $\widehat{\delta}_{DiD}$ can differ markedly depending on the grouping or time horizon considered.
\subsection{Illustration: Two Contrasting Cases}
To highlight this sensitivity, consider two stylized empirical setups with identical baseline means but different underlying dynamics:
\begin{enumerate}
\item \textbf{Case A — Stable control group.}
The untreated group remains constant over time ($\bar{Y}_{0,1} - \bar{Y}_{0,0} \approx 0$), while the treated group experiences an increase of $+1$.
Then $\widehat{\delta}_{DiD} \approx 1$, suggesting a strong treatment effect.
\item \textbf{Case B — Expanding control group.}
Both groups improve over time, but the control group’s gain is slightly smaller ($+0.8$ versus $+1.0$ for treated).
Here $\widehat{\delta}_{DiD} \approx 0.2$, indicating a much weaker effect, though the observed differences are driven largely by broader upward trends.
\end{enumerate}
Although both cases share similar raw means, the implied DiD estimates differ five-fold. Moreover, in small samples such differences can easily arise by chance, especially when only a few units define the group averages. Classical significance tests based on asymptotic normality may then yield unreliable $p$-values or conflicting conclusions about whether $\delta$ is statistically meaningful.
\subsection{When Is a Difference-in-Differences Estimate Significant?}
Determining the statistical significance of $\widehat{\delta}_{DiD}$ is nontrivial.
In large samples, robust or clustered standard errors often suffice.
However, when treatment assignment is limited to a small number of groups, the conventional variance estimator tends to understate uncertainty \citep{Cameron2005, MacKinnon2020}.
Furthermore, DiD inference depends on the specific composition of the treatment and control groups: re-defining them, or even altering time coding, can yield different standard errors and significance levels \citep{Conley2011, Ferman2019}.
These issues motivate the simulation-based framework proposed in this paper.
By repeatedly randomizing both the \textit{AFFECTED} and \textit{TIME} vectors, we construct an empirical distribution of $\widehat{\delta}_{DiD}$ under the null hypothesis of no treatment effect.
This randomization provides a reference distribution that accounts for the actual data structure and sample size, offering more reliable inference than asymptotic formulas—particularly in small-$N$ and unbalanced panels.
In the next section, we formalize this approach within the broader context of randomization-based inference, showing how it generalizes classical permutation logic to Difference-in-Differences estimation.
\section{Randomization-based Inference}\label{prelim}
The method proposed in this paper can be inherently seen as an extension of randomised inference \citep{ImbensRubin2015, Puchkin2023, MacKinnon2020}.
\subsection{Randomization-Based Inference in Econometrics}\label{subsec:rand-inf}
Randomization-based inference (often called \textit{randomization inference} or \textit{Fisherian inference} \citep{Ahmad2024, Efron2016}) is an approach to statistical inference that leverages the random assignment of treatments as the basis for testing hypotheses and estimating causal effects. The intellectual foundation dates back to the pioneering work of \citet{Fisher1935}, who introduced the idea of the \emph{randomization test} in the context of designed experiments. Fisher’s famous “lady tasting tea” experiment exemplified how one could test a null hypothesis by considering all possible reassignments of treatment labels and computing the probability of obtaining results as extreme as the observed outcome under random assignment alone.
Around the same time, \citet{Neyman1923} (translated in 1990) developed the potential outcomes framework for causal effects in experiments, emphasizing that each unit has a potential outcome under treatment and control and that random assignment allows for unbiased estimation of average effects \citep{Keller2024}. Building on Neyman’s framework, the \citet{Rubin1974} Causal Model (RCM) formalized causal inference using potential outcomes in both randomized and observational studies \citep{Imbens2010, ImbensRubin2015}. Together, these contributions established the logic that inference about treatment effects can (and should) be grounded in the known randomization mechanism, rather than requiring strong modeling assumptions about outcomes \citep{Rosenbaum2002, ImbensRubin2015}.
\subsubsection{Logic and Mathematical Foundation}
The core principle of randomization inference is that under the null hypothesis of no treatment effect, the distribution of any test statistic can be derived entirely from the random assignment process. Consider $N$ units, of which $N_T$ receive treatment and $N_C$ receive control, with $N = N_T + N_C$. Let $Y_i(1)$ and $Y_i(0)$ denote unit $i$'s potential outcomes with and without treatment, respectively \citep{Neyman1923, Rubin1974}. The sharp null hypothesis is:
\[
H_0: Y_i(1) = Y_i(0) \quad \text{for all } i.
\]
Under $H_0$, the observed outcome is $Y_i^{\text{obs}} = Y_i(0)$ regardless of treatment. Hence, the vector $\mathbf{Y}^{\text{obs}} = (Y_1^{\text{obs}}, \dots, Y_N^{\text{obs}})$ is fixed, and any variation in a test statistic arises solely from the randomness in the treatment assignment vector $\mathbf{W}$.
Let $T(\mathbf{W}, \mathbf{Y})$ be a test statistic of interest. The \emph{randomization distribution} of $T$ is the set $\{ T(\omega, \mathbf{Y}^{\text{obs}}) : \omega \in \Omega \}$ where $\Omega$ is the set of all allowable treatment assignments consistent with the experimental design \citep{Bonnini2024}. For a completely randomized design, $|\Omega| = \binom{N}{N_T}$. The $p$-value for a two-sided test is then:
\[
p = \frac{1}{|\Omega|} \sum_{\omega \in \Omega} \mathbb{I}\left( |T(\omega, \mathbf{Y}^{\text{obs}})| \ge |T_{\text{obs}}| \right).
\]
This provides an exact significance level for the sharp null \citep{Caughey2023}. When $|\Omega|$ is large, a Monte Carlo approximation can be used to sample a subset of $\Omega$.
\subsubsection{Permutation Distributions and Hypothesis Testing}
Randomization inference constructs the null distribution by permuting treatment labels and recalculating the test statistic on each permutation. For each simulated assignment $\omega \in \Omega$, the test statistic $T(\omega, \mathbf{Y}^{\text{obs}})$ is calculated. The empirical distribution of these values represents the permutation distribution under the null \citep{Caughey2023, Bonnini2024}.
In classic applications, such as Fisher’s exact test, the randomization distribution corresponds to a hypergeometric distribution \citep{Fisher1935}. In general, this approach extends to any test statistic and design, offering flexibility beyond conventional parametric frameworks \citep{Bonnini2024}.
Randomization tests are exact in finite samples under the null hypothesis. They do not rely on asymptotic approximations and are therefore especially useful in small samples or non-standard settings \citep{Rosenbaum2002}. They also allow researchers to construct confidence intervals by inverting a series of hypothesis tests (e.g., using Hodges–Lehmann-style estimation \citep{Hershberger2011}).
\subsubsection{Advantages in Small Samples and Non-Standard Conditions}
Randomization inference offers several advantages:
\begin{itemize}
\item \textbf{Exact Type I error control:} Since the null distribution is derived from the known randomization procedure, the test maintains the nominal significance level regardless of sample size \citep{Fisher1935}.
\item \textbf{Minimal assumptions:} No distributional assumptions are needed regarding $Y_i$, making the method robust to skewness, heteroskedasticity, and non-normality \citep{Rosenbaum2002}.
\item \textbf{Validity under model misspecification:} Unlike parametric methods, randomization tests do not depend on correct specification of a regression model \citep{Young2019, Heckman2024}.
\item \textbf{Applicability to complex designs:} Extensions exist for stratified, blocked, and cluster-randomized designs \citep{ImbensRubin2015, Athey2017}.
Exact tests, including permutation methods, are preferred when sample sizes are small or outcomes are non-normal.
\end{itemize}
\subsubsection{Applications in Econometrics and Other Fields}
Randomization inference is increasingly used in empirical economics \citep{Caughey2023, Canay2017, Abadie2022}, especially in development and labor economics, where randomized controlled trials (RCTs) are common. Researchers often supplement regression-based inference with permutation-based $p$-values to ensure robustness \citep{Young2019, Basse2024}. Another rapidly growing field of application is network analysis, where interference and non-independence of units violate the assumptions underlying classical asymptotic inference. Randomization inference provides exact tests under complex exposure mappings and graph-structured dependence, and has become a standard tool in experimental and observational network settings \citep{Aronow2017, Athey2018, Ugander2023, Li2019}.
It is also used in quasi-experimental settings. For example, in regression discontinuity designs, permutation tests are applied locally around the cutoff to assess robustness of estimated effects \citep{Cattaneo2015}. In observational studies with matched samples, permutation inference allows testing under assumptions of “as-if” randomization \citep{Rosenbaum2002}.
Outside of economics, randomization inference has long-standing applications in biomedical research, particularly in clinical trials and epidemiology - e.g. \citet{Khim2020} developed a permutation test for edge structure of an infection graph.
\subsubsection{Conclusion}
Randomization inference provides a principled and flexible framework for testing hypotheses about treatment effects. By leveraging the known assignment mechanism, it avoids strong assumptions, delivers exact inference in finite samples, and accommodates complex designs \citep{ImbensRubin2015}. As computing power increases and interest in robust causal inference grows, the use of randomization-based methods in econometrics and related fields continues to expand \citep{Athey2017}.
\section{Methods}
\label{sec:methods}
The proposed procedure provides a simulation-based significance test for the Difference-in-Differences (DiD) estimator that does not rely on asymptotic approximations. Instead, the test constructs an empirical reference distribution of the DiD coefficient under the null hypothesis of no treatment effect through repeated randomization of the treatment and time indicators.
\subsection{Hypotheses and Decision Rule}
The test is based on the following hypotheses:
\begin{itemize}
\item $H_0$: $DD_e = 0$ \hspace{1cm}(no treatment effect),
\item $H_A$: $DD_e \neq 0$ \hspace{0.8cm}(presence of treatment effect).
\end{itemize}
The rejection region is determined by the empirical $\frac{\alpha}{2}$ and $1-\frac{\alpha}{2}$ quantiles of the simulated distribution of DiD estimates, denoted $P_{\frac{\alpha}{2}}$ and $P_{1-\frac{\alpha}{2}}$, respectively.
The null hypothesis $H_0$ is rejected if the observed estimate $DD_e$ lies outside the interval $(P_{\frac{\alpha}{2}},\, P_{1-\frac{\alpha}{2}})$, i.e.
\[
DD_e \notin (P_{\frac{\alpha}{2}},\, P_{1-\frac{\alpha}{2}}).
\]
This logic mirrors the two-sided test based on asymptotic critical values, but replaces the theoretical reference distribution with a simulation-based one generated from the data. The quantiles can be directly obtained using standard statistical software such as R or Stata \citep{Team2021, Villa2016}.
\subsection{Algorithmic Procedure}
\label{subsec:algorithm}
The algorithm proceeds as follows:
\begin{enumerate}
\item \textbf{Data preparation.}
Prepare three aligned vectors (or a $3 \times n$ matrix):
\begin{enumerate}
\item The first vector contains observed values of the outcome variable $Y$.
\item The second and third contain binary values representing \textit{Time} and \textit{Affected} indicators, respectively.
\end{enumerate}
\item \textbf{Randomization.}
For each of $N$ Monte Carlo iterations:
\begin{enumerate}
\item Randomly permute both the \textit{Time} and \textit{Affected} vectors, maintaining their marginal distributions (i.e. preserving the proportions of zeros and ones).
\item Compute the DiD estimator
\[
DD^{(k)} = (\bar{Y}_{1,1}^{(k)} - \bar{Y}_{1,0}^{(k)}) - (\bar{Y}_{0,1}^{(k)} - \bar{Y}_{0,0}^{(k)}),
\]
or equivalently obtain $DD^{(k)}$ as the interaction coefficient from the OLS model
\[
Y = \alpha + \beta\, \text{TIME} + \gamma\, \text{AFFECTED} + \delta\, (\text{TIME}\times \text{AFFECTED}) + \varepsilon.
\]
Store each value $DD^{(k)}$ in a vector of simulated estimates.
\end{enumerate}
\item \textbf{Empirical distribution.}
After completing $N$ iterations, use the simulated sample $\{ DD^{(1)}, \ldots, DD^{(N)} \}$ to approximate the null distribution of the DiD coefficient.
\item \textbf{Significance testing.}
Compute the empirical quantiles $P_{\frac{\alpha}{2}}$ and $P_{1-\frac{\alpha}{2}}$.
Reject $H_0$ if the observed $DD_e$ lies outside these bounds.
\end{enumerate}
This procedure yields a data-dependent reference distribution that reflects the sample’s structure and accounts for limited group sizes, unequal variances, or potential departures from model assumptions. We can write the procedure above into a pseudocode as shown in Appendix~\ref{app:pseudocode}.
The function to perform this procedure is included in the \text{sigDD} R package, available on GitHub at \href{https://github.com/profsms/sigDD}{https://github.com/profsms/sigDD} \citep[App. A]{Halkiewicz2024thesis}.
\subsection{Formal randomization space and finite-sample exactness}
\label{subsec:framework}
Let $\mathcal{X} = \{(Y_i, A_i, T_i)\}_{i=1}^n$ denote the sample,
where $A_i \in \{0,1\}$ and $T_i \in \{0,1\}$ are treatment and period
indicators, respectively.
Define $\Omega$ as the set of admissible label permutations,
and $\mathcal{P}$ as the uniform probability measure on $\Omega$:
\[
\mathcal{P}(\omega)
= \frac{1}{|\Omega|}, \quad \forall\, \omega \in \Omega.
\]
The test statistic $D(\mathcal{X})$ is a function $D:\mathcal{X}\!\to\!\mathbb{R}$,
and its randomization distribution is given by
\[
F_D(t)
= \mathcal{P}\big( D(\omega\mathcal{X}) \le t \big),
\]
where $\omega\mathcal{X}$ denotes the relabeled data.
The empirical $p$-value is thus
\[
p = \frac{1}{|\Omega|}\sum_{\omega\in\Omega}
\mathbf{1}\!\left\{\,|D(\omega\mathcal{X})|\ge |D(\mathcal{X})|\,\right\}.
\]
In this setting, the following theorem provides a precious property of our procedure.
\begin{theorem}[Exactness under the sharp null]
\label{thm:exactness}
Under the sharp null of no treatment effect, the distribution of $\mathcal{X}$ is invariant under any $\omega\in\Omega$. Consequently, the randomization $p$-value above satisfies $\mathbb{P}_{H_0}\!\big(p\le \alpha\big)\le \alpha$ for all $\alpha\in[0,1]$; hence the test is finite-sample exact.
\end{theorem}
\begin{proof}
Under $H_0$ the observed outcomes are fixed and only labels vary; by construction $\mathcal{L}(\mathcal{X})=\mathcal{L}(\omega\mathcal{X})$ for all $\omega\in\Omega$. Conditional on $\mathcal{X}$, the rank of $|D(\mathcal{X})|$ among $\{|D(\omega\mathcal{X})|\}_{\omega\in\Omega}$ is uniform on $\{1,\dots,|\Omega|\}$, which implies $\mathbb{P}_{H_0}(p\le \alpha)=\alpha$ (or $\le\alpha$ for conservative $p$-values).
\end{proof}
\begin{remark}[Why the time indicator is exchangeable under the null]
\label{rem:time-exchangeable}
The admissible permutation space described above implicitly relies on the fact
that, under the sharp null $H_0:\delta=0$, the timing of treatment carries no
identifying information. Indeed, $Y_{it}(1)=Y_{it}(0)$ for all $i,t$ implies
that the evolution of untreated outcomes is identical across groups, and the
parallel-trends assumption ensures that no systematic period-specific
differences exist. Consequently, the \texttt{TIME} vector is exchangeable under
$H_0$, and permuting it generates valid null assignments. This justifies the
dual randomisation scheme and explains why the permutation space expands by
$\binom{n}{n_T}$ (Proposition~\ref{prop:perm-gain}), improving the granularity
of the empirical null distribution.
\end{remark}
\subsection{Combinatorial and asymptotic properties of the permutation space}
\label{subsec:combinatorial}
We now establish two simple but useful theoretical results that characterize the size and convergence properties of the permutation space introduced above.
\begin{proposition}[Combinatorial gain from dual randomization]
\label{prop:perm-gain}
Let $n$ denote the number of observations, $n_A$ the number of treated units
(\texttt{AFFECTED}=1), and $n_T$ the number of post-treatment observations
(\texttt{TIME}=1). When both margins are randomized while preserving their
empirical counts, the number of admissible relabelings is
\[
\#\Omega(\text{AFFECTED only})=\binom{n}{n_A},\qquad
\#\Omega(\text{AFFECTED \& TIME})=\binom{n}{n_A}\binom{n}{n_T}.
\]
Hence, the multiplicative gain from randomizing both margins rather than one is
\[
\frac{\#\Omega(\text{AFFECTED \& TIME})}{\#\Omega(\text{AFFECTED only})}
= \binom{n}{n_T}.
\]
Under unconstrained Bernoulli reshuffling, where each label is redrawn
independently, this gain equals $2^{n}$. Thus, dual randomization expands the
permutation space exponentially, providing a richer and finer empirical
approximation to the null distribution.
\end{proposition}
\begin{proof}
If only the \texttt{AFFECTED} margin is permuted, the number of admissible
labelings equals $\binom{n}{n_A}$. When both \texttt{AFFECTED} and
\texttt{TIME} are permuted independently while preserving marginal counts, the
number of combinations is $\binom{n}{n_A}\binom{n}{n_T}$. The ratio is therefore
$\binom{n}{n_T}$.
If margins are not fixed (Bernoulli$(1/2)$ reshuffling), the additional margin
adds $2^n$ possible configurations. This completes the argument.
\end{proof}
\begin{lemma}[Asymptotic density of the randomization distribution]
\label{lem:density}
Let $D(\omega\mathcal{X})$ denote the test statistic computed on relabeled data
$\omega\mathcal{X}$, and assume $D$ is bounded with variance $\sigma^2>0$.
Then, as $n\to\infty$,
\[
\frac{1}{|\Omega|}\sum_{\omega\in\Omega}
\mathbf{1}\!\left\{ D(\omega\mathcal{X}) \le t \right\}
\longrightarrow F_{H_0}(t),
\]
where $F_{H_0}$ is the cumulative distribution of $D$ under the sharp null hypothesis.
Moreover, since $|\Omega_2|/|\Omega_1| = \binom{n}{n_T}$ diverges exponentially,
the empirical distribution based on the full randomization space $\Omega_2$
converges to $F_{H_0}$ at a faster rate, yielding a finer discretization
of the null distribution.
\end{lemma}
\begin{proof}[Proof of Lemma~\ref{lem:density}]
Under the sharp null hypothesis $H_0$, all relabelings
$\omega\mathcal{X}$ produce identically distributed outcomes.
Let $D(\omega\mathcal{X})$ denote the statistic computed under
labeling $\omega$.
Because $D$ is bounded with finite variance, the empirical
distribution function of $\{D(\omega\mathcal{X}) : \omega \in \Omega\}$
converges uniformly to its expectation by the law of large numbers
over the finite (or sampled) permutation space.
The result follows by identifying that expectation with
$F_{H_0}(t)=\mathbb{P}\{D\le t\mid H_0\}$.
Finally, since $|\Omega_2|/|\Omega_1|=\binom{n}{n_T}$ grows
exponentially in $n$, the empirical distribution based on $\Omega_2$
uses exponentially more relabelings, and therefore approximates
$F_{H_0}$ with a finer grid of attainable $p$-values.
\end{proof}
\subsection{Design of the Randomization Process}
\label{subsec:randomization}
The central design choice concerns how to randomize the \textit{Time} and \textit{Affected} indicators.
Both are binary variables and can naturally be modeled as Bernoulli random variables. However, care must be taken in specifying the probability parameter $p$ governing the fraction of treated or post-treatment observations.
Several randomization strategies were initially considered:
\begin{itemize}
\item varying the Bernoulli parameter $p$ at each iteration, such that $p \sim \mathit{Unif}(0,1)$;
\item randomizing not the indicators but their joint interaction $(\text{TIME}\times\text{AFFECTED})$;
\item fixing the overall proportions of treated and post-treatment units.
\end{itemize}
Empirical testing indicated that randomizing $p$ or the interaction term alone substantially increases computational burden without providing theoretical or practical benefits.
Moreover, such approaches can distort the joint structure of the data, leading to reference distributions that depend more on random sampling design than on the observed dataset.
Following conventions from experimental design, especially clinical trial methodology \citep{Rosenbaum1985, Kang2008}, the most consistent approach is to preserve balanced group sizes.
Hence, both \textit{Time} and \textit{Affected} variables are randomized under a fixed Bernoulli parameter $p = \frac{1}{2}$, ensuring equal proportions of 0s and 1s in each vector in the long run.
This choice stabilizes the test distribution and maintains comparability across different simulations and datasets. Additionally, as further discussed in Section~\ref{sec:conclusions}, such choice maximizes the information gain.
\subsection{Algorithmic Properties}
The proposed test inherits several desirable features from randomization-based inference:
\begin{itemize}
\item \textbf{Distribution-free significance:} no distributional assumption on $Y$ is required; validity follows directly from the randomization mechanism.
\item \textbf{Scale invariance:} adding a constant or multiplying the outcome variable by a scalar leaves the shape of the simulated distribution unchanged.
\item \textbf{Robustness:} as the number of simulations $N$ increases, the empirical distribution converges to the true randomization distribution; stability typically occurs at moderate iteration counts (a few thousand).
\item By Theorem~\ref{thm:exactness}, the procedure controls Type~I error in finite samples under the sharp null.
\item By Proposition~\ref{prop:perm-gain} and Lemma~\ref{lem:density},
the full randomization design ensures an exponentially larger permutation
space and asymptotically denser approximation of the null distribution.
\end{itemize}
These properties make the procedure especially useful when the number of groups is small (which will be further investigated in Sections~\ref{sec:numerics} and \ref{sec:conclusions}), standard errors are unreliable, or model assumptions are questionable.
The algorithm can be implemented in any environment capable of generating random permutations and linear regressions, such as R, Stata, or MATLAB.
\section{Numerics}\label{sec:numerics}
This section presents a set of numerical experiments designed to validate the
proposed doubly randomized significance test for the Difference-in-Differences
coefficient. We apply the procedure to four empirical datasets commonly used in
the DiD literature. For each dataset, we provide:
(i) a brief description and summary statistics,
(ii) permutation-based inference results under single and dual randomization,
and (iii) visualizations of the resulting permutation distributions.
All tables and figures are referenced in the text.
\subsection{Dataset 1: \texttt{Inpress\_data} (Indonesia School Construction Program, 1973–1978)}
\label{subsec:inpress_data}
The first dataset, \texttt{Inpress\_data}, is derived from the large-scale primary
school construction program conducted in Indonesia between 1973 and 1978
\citep{Duflo2001}, also discussed in \citet{Halkiewicz2023}. Each observation
contains an educational outcome and two binary indicators: \texttt{affected}
(treated cohort) and \texttt{time} (high-intensity school-building regions).
A summary of pre- and post-program averages across treated and control groups is
reported in Table~\ref{tab:inpress_sample_summary}.
\begin{table}[H]
\centering
\caption{Summary statistics for the transformed \texttt{Inpress} dataset}
\label{tab:inpress_sample_summary}
\begin{tabular}{lcc}
\toprule
& \textbf{Pre-program ($time=0$)} & \textbf{Post-program ($time=1$)} \\
\midrule
Control ($affected=0$) & 9.7327 & 8.4759 \\
Treated ($affected=1$) & 10.1184 & 8.9379 \\
\bottomrule
\end{tabular}
\end{table}
Permutation-based inference results appear in
Table~\ref{tab:inpress_did_randomization}, showing that neither randomization
scheme rejects the null hypothesis for this dataset. The corresponding
permutation distributions are visualized in Figure~\ref{fig:inpress_histograms}.
\begin{table}[H]
\centering
\caption{Permutation-based DiD estimation results for the \texttt{Inpress} dataset}
\label{tab:inpress_did_randomization}
\begin{tabular}{lcccccc}
\toprule
& \multicolumn{3}{c}{\textbf{Randomization: Affected}} &
\multicolumn{3}{c}{\textbf{Randomization: Full}} \\
\cmidrule(lr){2-4} \cmidrule(lr){5-7}
Value & $q_{0.025}$ & $q_{0.975}$ & $H_0$
& $q_{0.025}$ & $q_{0.975}$ & $H_0$ \\
\midrule
0.076 & -0.149 & 0.148 & not rejected & -0.145 & 0.146 & not rejected \\
\bottomrule
\end{tabular}
\end{table}
\begin{figure}[H]
\centering
\begin{tabular}{cc}
\includegraphics[width=0.45\textwidth]{Plots/IP_hist_1.png} &
\includegraphics[width=0.45\textwidth]{Plots/IP_hist_2.png}
\end{tabular}
\caption{Permutation distributions for the \texttt{Inpress} dataset under single (left)
and full (right) randomization.}
\label{fig:inpress_histograms}
\end{figure}
\subsection{Dataset 2: Ben \& Jerry’s vs.\ Häagen-Dazs (Google Trends, 2020)}
\label{subsec:bj_hgd_data}
The second dataset contains daily search-intensity data (0–100 scale) for
Ben~\&~Jerry’s (treated) and Häagen-Dazs (control), collected from
March–September~2020 and normalized within each keyword
\citep{Halkiewicz2024thesis}. Ben~\&~Jerry’s issued a public
\emph{Black Lives Matter} statement in June~2020, providing a treatment event.
Summary statistics are shown in Table~\ref{tab:bj_sample_summary}.
\begin{table}[H]
\centering
\caption{Summary statistics for Ben~\&~Jerry’s vs.\ Häagen-Dazs dataset}
\label{tab:bj_sample_summary}
\begin{tabular}{lcc}
\toprule
& Pre ($time=0$) & Post ($time=1$) \\
\midrule
Control ($affected=0$) & 1.915 & 2.055 \\
Treated ($affected=1$) & 5.681 & 10.648 \\
\bottomrule
\end{tabular}
\end{table}
Table~\ref{tab:bj_did_randomization} reports the permutation-based DiD results,
which show strong evidence against the null of zero effect.
The associated permutation distributions appear in
Figure~\ref{fig:bj_histograms}.
\begin{table}[H]
\centering
\caption{Permutation-based DiD inference for Ben~\&~Jerry’s vs.\ Häagen-Dazs}
\label{tab:bj_did_randomization}
\begin{tabular}{lcccccc}
\toprule
& \multicolumn{3}{c}{\textbf{Randomization: Affected}} &
\multicolumn{3}{c}{\textbf{Randomization: Full}} \\
\cmidrule(lr){2-4} \cmidrule(lr){5-7}
Value & $q_{0.025}$ & $q_{0.975}$ & $H_0$
& $q_{0.025}$ & $q_{0.975}$ & $H_0$ \\
\midrule
4.827 & -2.949 & 2.956 & rejected & -2.949 & 3.018 & rejected \\
\bottomrule
\end{tabular}
\end{table}
\begin{figure}[H]
\centering
\begin{tabular}{cc}
\includegraphics[width=0.45\textwidth]{Plots/BJ_hist_1.png} &
\includegraphics[width=0.45\textwidth]{Plots/BJ_hist_2.png}
\end{tabular}
\caption{Permutation distributions for Ben~\&~Jerry’s vs.\ Häagen-Dazs under
single (left) and full (right) randomization.}
\label{fig:bj_histograms}
\end{figure}
\subsection{Dataset 3: Minimum Wage Reform in New Jersey and Pennsylvania (1992)}
\label{subsec:min_wage_data}
The third dataset reproduces the fast-food restaurant survey from
\citet{CardKrueger1994}. Each restaurant reports pre- and post-treatment values
for employment, starting wage, and meal price.
Summary statistics are given in
Tables~\ref{tab:minwage_emptot_summary}--\ref{tab:minwage_pmeal_summary}.
\begin{table}[H]
\centering
\caption{Employment summary: \texttt{emptot}}
\label{tab:minwage_emptot_summary}
\begin{tabular}{lcc}
\toprule
& Pre ($time=0$) & Post ($time=1$) \\
\midrule
Control ($affected=0$) & 23.3312 & 21.1656 \\
Treated ($affected=1$) & 20.4394 & 21.0274 \\
\bottomrule
\end{tabular}
\end{table}
\begin{table}[H]
\centering
\caption{Starting wage summary: \texttt{wage\_st}}
\label{tab:minwage_wage_summary}
\begin{tabular}{lcc}
\toprule
& Pre ($time=0$) & Post ($time=1$) \\
\midrule
Control & 4.6301 & 4.6175 \\
Treated & 4.6121 & 5.0808 \\
\bottomrule
\end{tabular}
\end{table}
\begin{table}[H]
\centering
\caption{Meal price summary: \texttt{pmeal}}
\label{tab:minwage_pmeal_summary}
\begin{tabular}{lcc}
\toprule
& Pre ($time=0$) & Post ($time=1$) \\
\midrule
Control & 3.0424 & 3.0266 \\
Treated & 3.3511 & 3.4148 \\
\bottomrule
\end{tabular}
\end{table}
The permutation-based DiD inference results are shown in
Table~\ref{tab:minwage_did_randomization}. The corresponding permutation
distributions for all three outcomes are visualized in
Figure~\ref{fig:minwage_histograms}.
\begin{table}[H]
\centering
\caption{Permutation-based DiD inference for the minimum wage dataset}
\label{tab:minwage_did_randomization}
\begin{tabular}{lccccccc}
\toprule
Outcome & Value & $q_{0.025}$ & $q_{0.975}$ & $H_0$ &
$q_{0.025}$ (Full) & $q_{0.975}$ (Full) & $H_0$ (Full) \\
\midrule
\texttt{emptot} & 2.7536 & -2.5790 & 2.6134 & rejected &
-2.6269 & 2.6010 & rejected \\
\texttt{wage\_st} & 0.4814 & -0.0854 & 0.0852 & rejected &
-0.1017 & 0.1019 & rejected \\
\texttt{pmeal} & 0.0794 & -0.1855 & 0.1793 & not rejected &
-0.1810 & 0.1821 & not rejected \\
\bottomrule
\end{tabular}
\end{table}
\begin{figure}[H]
\centering
\begin{tabular}{cc}
\includegraphics[width=0.45\textwidth]{Plots/Mwage_emptot_hist_1.png} &
\includegraphics[width=0.45\textwidth]{Plots/Mwage_emptot_hist_2.png} \\
\textit{emptot} & \textit{emptot} \\[1em]
\includegraphics[width=0.45\textwidth]{Plots/Mwage_wage_st_hist_1.png} &
\includegraphics[width=0.45\textwidth]{Plots/Mwage_wage_st_hist_2.png} \\
\textit{wage\_st} & \textit{wage\_st} \\[1em]
\includegraphics[width=0.45\textwidth]{Plots/Mwage_pmeal_hist_1.png} &
\includegraphics[width=0.45\textwidth]{Plots/Mwage_pmeal_hist_2.png} \\
\textit{pmeal} & \textit{pmeal}
\end{tabular}
\caption{Permutation distributions for minimum wage dataset under
single (left) and full (right) randomization.}
\label{fig:minwage_histograms}
\end{figure}
\subsection{Dataset 4: \texttt{dinas\_golden\_dawn} (Refugee Arrivals and Far-Right Voting, Greece 2012–2016)}
\label{subsec:dinas_data}
The final dataset comes from \citet{Dinas2019} and covers 96 municipalities
over four Greek parliamentary elections. The treatment corresponds to
substantial refugee-inflow exposure.
Table~\ref{tab:dinas_sample_summary} presents summary statistics.
\begin{table}[H]
\centering
\caption{Summary statistics for refugee-arrival dataset}
\label{tab:dinas_sample_summary}
\begin{tabular}{lcc}
\toprule
& Pre ($time=0$) & Post ($time=1$) \\
\midrule
Control ($affected=0$) & 5.0720 & 5.6591 \\
Treated ($affected=1$) & 5.7299 & 8.4039 \\
\bottomrule
\end{tabular}
\end{table}
Table~\ref{tab:dinas_did_randomization} reports permutation-based inference,
and Figure~\ref{fig:dinas_histograms} presents the two permutation distributions.
\begin{table}[H]
\centering
\caption{Permutation-based DiD inference for the \texttt{dinas\_golden\_dawn} dataset}
\label{tab:dinas_did_randomization}
\begin{tabular}{lcccccc}
\toprule
Value & $q_{0.025}$ & $q_{0.975}$ & $H_0$
& $q_{0.025}$ (Full) & $q_{0.975}$ (Full) & $H_0$ (Full) \\
\midrule
2.0870 & -1.1627 & 1.1588 & rejected &
-1.0490 & 1.0407 & rejected \\
\bottomrule
\end{tabular}
\end{table}
\begin{figure}[H]
\centering
\begin{tabular}{cc}
\includegraphics[width=0.45\textwidth]{Plots/Golden_hist_1.png} &
\includegraphics[width=0.45\textwidth]{Plots/Golden_hist_2.png}
\end{tabular}
\caption{Permutation distributions for the refugee-arrival dataset under
single (left) and full (right) randomization.}
\label{fig:dinas_histograms}
\end{figure}
\section{Conclusions}\label{sec:conclusions}
To this day, most researchers have approached the problem of testing the
significance of Difference-in-Differences (DD) estimations by employing tools
borrowed from Ordinary Least Squares (OLS) models, such as $t$-tests.
However, as discussed in Sec.~\ref{sec:did}, even though a DD estimate can be
obtained from an OLS regression, its meaning differs fundamentally from that of
a structural parameter. Therefore, there remains a need for a single,
well-defined significance test tailored specifically to the DD estimator.
Such a procedure would allow researchers to assess whether the observed DD
effect carries real empirical weight, without relying on subjective or
model-dependent approaches. Moreover, by introducing an interpretable scale,
the value of the DD parameter could be understood not only in binary
significance terms, but also with respect to its effect magnitude and strength.
\medskip
\noindent
The results presented in this study demonstrate that the proposed permutation
test, which randomizes both the \texttt{TIME} and \texttt{AFFECTED} dimensions,
performs comparably to the conventional approach that randomizes only
\texttt{AFFECTED}. Across all examined datasets, both methods led to consistent
inference decisions: strong effects were detected by both tests, while weak or
null effects were not. Importantly, the proposed full randomization procedure
provides a substantially larger set of admissible permutations, leading to a
denser empirical null distribution. This feature stabilizes empirical
$p$-values and critical thresholds, which is particularly advantageous in
small-sample applications. Hence, the method constitutes a promising, more
objective alternative to conventional significance testing in
Difference-in-Differences designs.
\medskip
\noindent
Theoretical support for these empirical findings is provided in
Sec.~\ref{subsec:combinatorial}, where the properties of the permutation space
are analyzed formally. In particular, Proposition~\ref{prop:perm-gain}
quantifies the combinatorial expansion induced by dual randomization, and
Lemma~\ref{lem:density} establishes that a larger permutation space yields a
finer and faster-converging approximation of the null distribution.
Together, these results explain the observed stability of $p$-values and the
robustness of inference under full randomization.
As established in Proposition~\ref{prop:perm-gain}, dual randomization expands
the admissible permutation space by a factor of $\binom{n}{n_T}$ relative to
single-margin permutation.
\medskip
\noindent
Let $p_T = n_T/n$. By Stirling’s approximation \citep{Morris2024}, the binomial
coefficient admits the asymptotic expression
\[
\binom{n}{n_T}
\approx
\frac{1}{\sqrt{2\pi n p_T(1-p_T)}}\,
\exp\!\big(n\,H(p_T)\big),
\qquad
H(p)=-p\log p-(1-p)\log(1-p),
\]
where $H(p)$ is the binary Shannon entropy \citep{Saraiva2023}. Therefore, the
logarithmic growth rate of the permutation space is approximately $nH(p_T)$,
reaching its maximum at $p_T=\tfrac{1}{2}$. This shows that the information
content of the permutation universe, measured by entropy, is maximised when the randomisation is balanced, further justifying the choice of $p=\frac{1}{2}$, posed in Section~\ref{subsec:randomization}.
\medskip
\noindent
\textbf{Entropy interpretation.}
Let $p_A = n_A/n$ and $p_T = n_T/n$ denote the empirical probabilities of
treatment and post-treatment status, respectively. The combined permutation
space then satisfies
\[
\frac{1}{n}\log \#\Omega(\text{AFFECTED \& TIME})
\approx H(p_A) + H(p_T).
\]
Hence, adding the second randomized margin contributes an independent entropy
term $H(p_T)$, quantifying the increase in information complexity of the test.
When both margins are balanced ($p_A=p_T=1/2$), the entropy of the randomization space is maximized,
\[
\#\Omega(\text{AFFECTED \& TIME})
\asymp
\frac{2^{2n}}{2\pi n}\,.
\]
Consequently, the number of possible relabelings grows exponentially with $n$,
and the full randomization scheme explores an exponentially richer configuration
space. This combinatorial and entropic expansion formally explains the
empirically observed stability of the test: by enlarging the support of the
null distribution, the method achieves greater precision in finite samples and
reduces discretization bias.
\subsection{Further research directions}
There is an expanding number of researches conducted using Difference-in-Differences models and new areas for possible applications are described every year. Given this fact, there is a growing need for efficient, strict and unbiased tests to measure the effects presented by the DD model. Inquires may be posed whether there exists a way to standardize the test distribution, such that it would not be dependent on the sample size. Further research endeavors may be also conducted to generalize the test distribution to one of well-described distributions and finding the parameters of such, therefore eliminating the need for, often time-consuming, simulations. Debate about whether this testing method is truly unbiased is also welcomed.
\newpage
\bibliographystyle{erae}
\bibliography{references}