EconBase
← Back to paper

Dynamic Local Average Treatment Effects in Time Series

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

152,376 characters

Dynamic Local Average Treatment Effects in Time Series


\pagebreak{}

\setcounter{page}{0}

\raggedbottom
\title{\textbf{Dynamic Local Average Treatment Effects in Time Series\thanks{ Casini acknowledges financial support from EIEF through 2022 EIEF
Grant. McCloskey acknowledges support from the National Science Foundation
under Grant SES-2341730. }}}
\maketitle
\begin{abstract}
{\footnotesize This paper discusses identification, estimation, and
inference on dynamic local average treatment effects (LATEs) in instrumental
variables (IVs) settings. First, we show that compliers\textemdash observations
whose treatment status is affected by the instrument\textemdash can
be identified }{\footnotesize\emph{individually}}{\footnotesize{}
in time series data using smoothness assumptions and local comparisons
of treatment assignments. Second, we show that this result enables
not only better interpretability of IV estimates but also direct testing
of the exclusion restriction by comparing outcomes among identified
non-compliers across instrument values. Third, we document pervasive
weak identification in applied work using IVs with time series data
by surveying recent publications in leading economics journals. However,
we find that strong identification often holds in large subsamples
for which the instrument induces changes in the treatment. Motivated
by this, we introduce a method based on dynamic programming to detect
the most strongly-identified subsample and show how to use this subsample
to improve estimation and inference. We also develop new identification-robust
inference procedures that focus on the most strongly-identified subsample,
offering efficiency gains relative to existing full sample identification-robust
inference when identification fails over parts of the sample. Finally,
we apply our results to heteroskedasticity-based identification of
monetary policy effects. We find that about 75\% of observations are
compliers (i.e., cases where the variance of the policy shifts up
on FOMC announcement days), and we fail to reject the exclusion restriction.
Estimation using the most strongly-identified subsample helps reconcile
conflicting IV and GMM estimates in the literature. }{\footnotesize\par}
\end{abstract}
 {\footnotesize{\indent {\bf{JEL Classification}}: B41, C12, C32.\\
{\bf{Keywords}}: Compliers, Conditional inference, Exclusion, LATE.}} \\

\onehalfspacing
\thispagestyle{empty}

\pagebreak{}

\section{Introduction}

Economists work hard to extract plausibly exogenous variation in order
to identify causal effects. Many identification strategies used in
applied work either rely directly on instrumental variables (IVs)
or can be reframed in terms of IV identification. This holds also
in dynamic settings where, for example, external IVs may be constructed
using a narrative approach or heteroskedasticity is exploited to yield
additional identifying equations. Since \citet{imbens/angrist:1994},
it has been well-known that IV-based approaches identify the local
average treatment effect (LATE)\textemdash the average treatment effect
for the sub-population of compliers, i.e., those whose treatment status
is influenced by the policy intervention (the instrument).

In the LATE framework, the sub-population of compliers is unobserved.
This means that although a LATE can be identified, the specific sample
observations this effect represents is unknown. This limitation is
often described informally as the inability to observe an observation's
treatment status under both the intervention and non-intervention
scenarios. From a practical interpretability perspective, this presents
a challenge that has been widely discussed in the literature {[}see,
e.g., \citet{angrist/imbens/rubin:1996}, \citet{heckman:1996}, \citet{imbens:2010}
and \citet{robins/greenland:1996}{]}. Some progress has been made
by \citet{imbens/rubin:1997} and \citet{abadie:2003} who show that
the proportion of compliers and some of their statistical characteristics
can be identified, provided these characteristics can be expressed
as functions of moments of the joint distribution of observed data.
Using these results, \citet{bhuller/dahl/loken/mogstad:2020} conduct
a detailed analysis of compliers in the context of interpreting IV
estimates of the effect of incarceration on recidivism and subsequent
labor market outcomes. Their work, along with many other studies,
highlights the importance of identifying the (characteristics of)
compliers when drawing policy implications.

This paper considers IV identification in dynamic settings and shows
how compliers can be identified individually in this context. We first
show that the notion of compliers can be equivalently rewritten in
terms of an inequality involving the difference in means of the potential
treatment under different instrument values. Under assumptions of
continuity over time in the mean of the potential treatment assignment
process\textemdash conditional on a fixed hypothetical value of the
instrument\textemdash it is possible to recover counterfactual values
by averaging observations in a neighborhood around a given time point.

For example, consider heteroskedasticity-based identification of the
causal effects of monetary policy {[}cf. \citet{rigobon:2003} and
\citet{nakamura/steinsson:2018}{]} where the instrument indicates
whether there was an FOMC announcement on each date in the sample
and the treatment variable is equal to the variance of a short-term
interest rate. Here compliers are defined as observations for which
the volatility of the policy variable (change in short-term interest
rate) increases if and only if there is an FOMC announcement. Suppose
that there is an FOMC announcement on a given date of interest so
that we do not observe the potential treatment assignment under the
counterfactual instrument value indicating the absence of an announcement.
Although the mean treatment assignment under non-announcement is unobserved
at this date, it can be recovered if mean treatment assignments are
a smooth function of time by computing an average of nearby days without
an announcement. Under an additional assumption of deterministic complier
status, the complier status of the date in question can be estimated
and tested by comparing local means of the treatment variable, one
corresponding to nearby dates for which an announcement occurred and
the other corresponding to nearby dates for which it did not.\footnote{Even though we focus on a time series setting, our identification
results immediately apply to cross-sectional settings with spatial
data provided that the temporal distance between observations is interpreted
as geographical distance, and analogous continuity assumptions are
imposed over space.} Applying our identification results and tests to the heteroskedasticity-based
identification of monetary policy effects, we find that about 75\%
of observations are compliers while the non-compliers are primarily
concentrated in the early zero lower bound period, when the central
bank could no longer lower interest rates and forward guidance was
not aggressive.

Identification of compliers is not only valuable in its own right.
It also enables us to test the exclusion restriction, a key condition
for valid IV estimation that is typically untestable in practice.
By identifying compliers, and thus also non-compliers, we show that
the exclusion restriction can be tested using a $t$-test that compares
the average outcomes of non-compliers across different instrument
values.

A key condition for identification of the LATE in the IV framework
is instrument relevance, entailing nontrivial correlation between
the endogenous variable and the instrument. We begin by analyzing
the problem of weak instruments, entailing low correlation between
the endogenous variable and instrument, in empirical work through
a survey of articles using IVs published from 2019 to 2022 in five
leading journals: American Economic Review, Econometrica, Journal
of Political Economy, Quarterly Journal of Economics, and Review of
Economic Studies. Our sample includes 1,560 specifications from 18
papers, with 199 involving time series and 1361 involving panel data.\footnote{See the supplement for the full list of papers and inclusion criteria.}
The left panels of Figure \ref{Figure_Histogram} show histograms
of full sample first-stage $F$-statistics for the specifications
in our survey, truncated above 100 for visibility. Many $F$-statistics
concentrate around the $\chi_{1}^{2}$ critical values and fall below
the conventional thresholds of 10 and 23.1 suggested by \citet{staiger/stock:1997}
and \citet{montielolea/pflueger:2013}, raising serious concerns about
weak instruments.\footnote{Indeed, \citet{staiger/stock:1997} derive the threshold of 10 under
the homoskedasticity-only assumption\textemdash the relevant thresholds
for time series data are larger {[}see \citet{montielolea/pflueger:2013}{]}.} These findings align with those of \citet{andrews/stock/sun:2019},
who analyze cross-sectional studies. For example, we find that 75\%
of time series and 72\% of panel data specifications have first stage
$F$-statistics below 24. The median $F$-statistic is 12.63 for time
series and 9.29 for panel data.\footnote{For panel data specifications we consider each cross-sectional unit
individually to enable comparison to our proposed time series test
shown on the right panels.}
\noindent\begin{flushleft}
\begin{center}
\begin{figure}[h]
\begin{centering}
\includegraphics[width=15cm,totalheight=7cm]{Figure_Histogram}
\par\end{centering}
\raggedright{}\caption{\label{Figure_Histogram}{\scriptsize{} Distributions of the first-stage
$F$ (left panels) and $F^{*}$ statistics (right panels). The top
panels apply to time series specifications and the bottom panels apply
to panel data specifications. The orange and red vertical lines correspond
to the 5\% and 1\% level asymptotic critical values of the first-stage
$F$ ($\chi_{1}^{2}$ for left panels) and $F^{*}$ statistics (8.28
and 11.63) for right panels) under identification failure.}}
\end{figure}
\end{center}
\par\end{flushleft}

When identification fails or is weak, IV estimators can be severely
biased for LATEs and conventional inference methods are rendered invalid.
These problems have prompted extensive research on detecting weak
instruments and constructing identification-robust confidence sets.\footnote{See, e.g., \textcolor{MyBlue}{Andrews et al.} \citeyearpar{andrews/moreira/stock:2006},
\citet{kleibergen:2002}, \citet{moreira:2003} and \citet{staiger/stock:1997}.} However, there has been little work on estimation and inference in
a general LATE setting when identification may be stronger over subsamples.
The second main contribution of this paper is to develop a framework
for identification, estimation, and inference on LATEs that accommodates
time-varying instrument relevance. Within this framework, we propose
a first-stage $F$-test to detect whether identification fails over
all nontrivial subsamples. To solve the computationally intensive
problem of searching for maximal identification strength among all
possible sample partitions, we employ dynamic programming. This optimization
is more complex than that in the structural break literature since
evaluating identification strength requires more than comparing parameter
changes across regimes.

In an attempt to understand the sources of weak IVs we plot the histograms
of the $F^{*}$ statistic proposed in this paper (cf. Section \ref{Section: ID Failure Test})
in the right panels of Figure \ref{Figure_Histogram}. The statistic
$F_{}^{*}$ searches for the subsample with maximal identification
strength among all possible subsamples of size at least $\pi_{L}T$.\footnote{We set $\pi_{L}=0.6$ in Figure \ref{Figure_Histogram}. We discuss
the choice of $\pi_{L}$ below.} The idea is that while the IVs may appear weak in the full sample,
they may be strong in a possibly large subsample. Figure \ref{Figure_Histogram}
shows that this is indeed often the case. The red vertical lines in
Figure \ref{Figure_Histogram} mark the $95^{th}$ percentile of the
asymptotic distributions of the $F$ and $F^{*}$ statistics under
the null of identification failure. Although its quantiles are larger,
the $F^{*}$ statistics have substantially more mass in the upper
quantiles of their null distribution. This has at least three implications.
First, it confirms substantial time variation in the instruments'
strength. Second, strong identification appears to be frequently present
in a sizeable subsample even when the instruments appear weak in the
full sample. The median $F_{}^{*}$ is 27.22 for time series and 33.81
for panel data specifications. These are substantially higher than
their full sample counterparts and this difference cannot be simply
attributed to the different null asymptotic distributions of the two
test statistics given that the difference in the asymptotic critical
values is relatively small while the empirical distributions of the
two test statistics is markedly different. About one half of the specifications
that appear to suffer from weak IVs in the full sample seem better
characterized by strong IVs in the subsample with maximal identification
strength. Third, the subsamples where instruments appear strong tend
to be large. From an empirical perspective, this is encouraging: although
weak instruments in the full sample are common, researchers can often
succeed in identifying large subsamples where instruments appear strong.

Motivated by this survey evidence, we construct consistent estimators
of LATEs when subsamples are strongly-identified. It is commonly believed
that if IVs are strong only in some portion of the sample, the full
sample IV estimator remains consistent for a LATE. We show that this
belief is unwarranted unless the LATE of interest is time-invariant
(i.e., homogeneous). If this condition fails, one can at best identify
a LATE corresponding to the strongly-identified subsample. Even when
the LATE is homogeneous, the full sample IV estimator may still be
severely biased if instruments are irrelevant over parts of the sample.

Our approach differs from that of \citet{magnusson/mavroeidis:2014}
and \citet{antonie/boldea:2018}, who use time variation in IV strength
to add moment conditions in a GMM context, enabling more efficient
inference and estimation. In contrast, we exploit this time variation
to identify the subsample where IVs are strongest and base our estimation
on this subsample. This insight allows for\textit{ consistent estimation
even when subsamples suffer from identification failure}.\footnote{Another major difference from \citet{magnusson/mavroeidis:2014} is
that we address the computational challenge for the case of multiple
breaks in the first-stage coefficient. \citet{magnusson/mavroeidis:2014}
did not attempt to address this issue and refer to it as ``computationally
demanding''.} If the parameter of interest is heterogeneous, our estimator remains
valid but is interpretable only within the strongly-identified sub-population.

We apply our methodology to the heteroskedasticity-based identification
strategy used to estimate the causal effects of monetary policy from
high-frequency data {[}e.g., \citet{nakamura/steinsson:2018}{]}.
The key identification condition for this strategy is that the volatility
of the daily changes in short-term interest rates is higher on FOMC
announcement days than on non-FOMC days. \citet{lewis:2020} provides
evidence of weak full sample identification and shows that IV and
GMM estimates even differ in sign. We find that identification is
substantially stronger over a subsample comprising 80\textendash 90\%
of the data, with the excluded subsample centered around the financial
crisis, during which volatility was high even on non-FOMC days. Estimation
using the most strongly-identified subsample yields IV and GMM estimates
that have the same sign and similar magnitudes. We recommend reporting
the most strongly-identified subsample estimates in addition to the
full sample estimates when strong full sample identification may be
in question.

Although our new methods are able to find the most strongly-identified
subsample, this subsample may still fail to be strongly-identified.
For our final theoretical contribution, we develop identification-robust
inference procedures using the most strongly-identified subsample.
We propose versions of the Anderson-Rubin, Lagrange Multiplier, and
conditional likelihood ratio tests, which depend only on this subsample.
These tests are more efficient than their full sample counterparts,
which include noise from regimes suffering from identification failure.
When instruments are strong throughout the sample, our tests coincide
with the conventional ones. When instruments are irrelevant over parts
of the sample, our tests achieve higher efficiency by focusing on
stronger segments. In the worst case, when IVs are weak everywhere,
our methods are no less efficient than existing ones. While there
is a trade-off between using fewer observations and more strongly
identified subsamples, simulations show that our tests have higher
power, indicating that the efficiency loss from a smaller sample size
is outweighed by the gain in identification strength.

The paper is organized as follows. Section \ref{Section: Statistical Framework for Identification of Causal Effects}
introduces the potential outcome framework and dynamic causal effects,
and presents identification results. Section \ref{Section: Application: Money Neutrality}
discusses issues pertaining to heteroskedasticity-based identification
of monetary policy. Section \ref{Section: ID Failure Test} presents
an $F$-test for full sample identification failure. Estimation and
inference robust to weak identification are discussed in Sections
\ref{Section: Estimation}-\ref{Section: Inference}. An empirical
application is considered in Section \ref{Section: Empirical Evidence}.
Section \ref{Section Conclusions} concludes. The supplements \citeauthor{casini/mccloskey/pala/rolla_Dynamic_Late_Supp_Not_Online}
(\citeyear{casini/mccloskey/pala/rolla_Dynamic_Late_Supp_Not_Online},
\citeyear{casini/mccloskey/pala/rolla_Dynamic_Late_Supp}) include
the Monte Carlo simulations, proofs and additional results.

\section{Identification of Dynamic Causal Effects\label{Section: Statistical Framework for Identification of Causal Effects}}

A growing literature in macroeconomics uses IVs to identify dynamic
causal effects when the policy variable of interest is endogenous.\footnote{See, e.g., \citet{gertler/karadi:2015}, \citet{jorda/schularick/taylor:2015},
\citet{mertens/olea:2018}, \citet{mertens/ravn:2013}, \citet{plagborg-Moller/wolf:2022},
\citet{ramey/zubairy:2018} and \citeauthor{stock/watson:2012} (\citeyear{stock/watson:2012},
\citeyear{stock/watson:2018}).} Many existing identification approaches can be reframed in terms
of IVs, either derived from the modeling approach {[}e.g., heteroskedasticity-based
identification as in \citet{rigobon:2003} and \citet{nakamura/steinsson:2018}{]}
or through external IVs constructed using a narrative approach {[}cf.,
\citet{olea/stock/watson:2021}{]}. For example, \citet{romer/romer:1989}
study the FOMC minutes to pinpoint dates when monetary policy actions
were arguably exogenous. This allows the construction of exogenous
variables that can be interpreted as IVs for some structural shock
of interest.\footnote{See \citet{ramey/shapiro:1998} for unanticipated defense spending
shocks, \citet{kuttner:2001}, \citet{nakamura/steinsson:2018} and
\citet{romer/romer:2004} for monetary policy shocks, \citet{hamilton:2003},
\citet{kanzig:2021} and \citet{killian:2009} for oil market shocks,
\citet{kanzig:2021_Carbon} for carbon pricing shocks, \citeauthor{romer/romer:2010}
for tax shocks, and \citet{ramey:2011} for government spending shocks.}

We adopt a potential outcomes framework, as introduced by \citet{rubin:1974}
and extended to time series settings by \citet{angrist/kuersteiner:2011}
and \citet{rambachan/shephard:2021}. Let the stochastic process $V_{t}=(Y_{t},\,X_{t},\,D_{t},\,Z_{t})$
be defined on the probability space $\left(\Omega,\,\mathscr{F},\,\mathbb{P}\right)$,
where $Y_{t}$ is a vector of outcome variables, $D_{t}$ is a policy
variable, $X_{t}$ is a vector of other exogenous and/or lagged endogenous
variables, and $Z_{t}$ is a vector of instruments. Let $\vec{X}_{t}=\{\ldots,\,X_{t-1},\,X_{t},\}$
denote the covariate path up to time $t$, with analogous definitions
for $\vec{Y}_{t}$, $\vec{D}_{t}$ and $\vec{Z}_{t}$. Let the policy-relevant
information set at time $t$ denoted by $\mathscr{F}_{t}=\sigma(\widetilde{V}_{t})$
where $\sigma(\widetilde{V}_{t})$ is the $\sigma$-algebra generated
by the history of $V_{t}$, $\widetilde{V}_{t}=(\vec{Y}_{t-1},\,\vec{X}_{t},\,\vec{D}_{t-1},\,\vec{Z}_{t-1})$.

Policy decisions depend on past observable variables and the contemporaneous
outcome through a systematic component and on idiosyncratic information
available to the policy-maker (i.e., the random component). The systematic
component, denoted $D(\widetilde{V}_{t},\,Y_{t},\,Z_{t},\,t)$, is
a time-varying non-stochastic function of the observed random variables
$\widetilde{V}_{t}$, contemporaneous outcome $Y_{t}$, and the contemporaneous
instrument $Z_{t}$. The idiosyncratic information is represented
by a scalar stochastic shock $e_{t}$ that is not observed by the
researcher. The policy action is determined by $D_{t}=\varphi(D(\widetilde{V}_{t},\,Y_{t},\,Z_{t},\,t),\,e_{t},\,t)$,
where $\varphi$ is a general mapping. In a SVAR context, $e_{t}$
is the structural shock to the policy variable $D_{t}.$ For example,
if the monetary authority follows a simple Taylor rule for the nominal
interest rate, then $\varphi$ is linear and $\widetilde{V}_{t}$
includes inflation, output and the natural rate of interest.

We define two types of potential outcomes. The first, $Y_{t}\left(\left(\epsilon_{1:t}\right),\,\left(z_{1:t}\right)\right)$,
denotes the counterfactual values of $Y_{t}$ under hypothetical sequences
of the policy shocks $\epsilon_{1:t}$ and instruments $z_{1:t}$,
where $a_{1:t}=\left\{ a_{s}\right\} _{s=1}^{t}$.
\begin{defn}
\label{Definition: Generalized Potential Outcome}A generalized potential
outcome, $Y_{t}\left(\left(\epsilon_{1:t}\right),\,\left(z_{1:t}\right)\right)$,
is defined as the value assumed by $Y_{t}$ if $e_{s}=\epsilon_{s}$
and $Z_{s}=z_{s}$ for $s=1,\ldots,\,t$.
\end{defn}
This definition excludes dependence on future shocks or instruments.
The potential outcome process should not be confused with the observed
outcome $\left\{ Y_{t}\right\} _{t\geq1}=\left\{ Y_{t}\left(e_{1:t},\,Z_{1:t}\right)\right\} _{t\geq1}$.
For $h\geq0$ and any given $\epsilon$ and $z$, write the time-$t+h$
potential outcome along the path $\left(\left(e_{1:t-1},\,\epsilon,\,e_{t+1:t+h}\right),\,\left(Z_{1:t-1},\,z,\,Z_{t+1:t+h}\right)\right)$
as
\begin{align*}
Y_{t,h}\left(\epsilon,\,z\right) & =Y_{t+h}\left(\left(e_{1:t-1},\,\epsilon,\,e_{t+1:t+h}\right),\,\left(Z_{1:t-1},\,z,\,Z_{t+1:t+h}\right)\right),
\end{align*}
where $Y_{t,h}\left(e_{t},\,Z_{t}\right)=Y_{t+h}$. Definition \ref{Definition: Generalized Potential Outcome}
captures the property that $Y_{t,h}\left(\epsilon,\,z\right)$ also
depends on policy shocks that occur between time $t+1$ and $t+h$.
The notation $Y_{t,h}\left(e,\,z\right)$ focuses on the effect of
a single policy shock on current and future outcomes akin to the idea
underlying an impulse response. When the potential outcomes do not
depend on the instruments, $Y_{t,h}\left(\epsilon,\,z\right)=Y_{t,h}\left(\epsilon\right)$,
and for $\epsilon\neq\epsilon'$, $Y_{t,h}\left(\epsilon\right)-Y_{t,h}\left(\epsilon'\right)$
for $h=0,\,1,\,\ldots$ are the dynamic causal effects of a policy
shock on the outcome. In a SVAR setting, one is often interested in
these dynamic causal effects which are in fact the impulse responses.

The second potential outcome that we discuss, $Y_{t}^{*}\left(\left(d_{1:t}\right),\,\left(z_{1:t}\right)\right)$,
is defined as the counterfactual values of $Y_{t}$ under hypothetical
sequences of treatments $d_{1:t}$ and instruments $z_{1:t}$. The
distinction with $Y_{t}\left(\epsilon_{1:t},\,z_{1:t}\right)$ is
that this formulation focuses on causal effects of the policy variable
$D$, not the policy shock $e$. For $t\geq1$, we assume that $d_{t}\in\mathbf{D}$,
$z_{t}\in\mathbf{Z}$ for some sets $\mathbf{D}$ and $\mathbf{Z}$.
In many applications outside SVARs, the causal effects of the policy
are of interest. Think about the slope of demand functions, price
elasticities, response coefficients or reaction functions of, for
example, asset prices to monetary policy, and so on. Typically these
causal effects are analyzed using event-studies, quasi-experiments,
IV regressions, etc. The recent literature on causal effects in time
series {[}e.g., \citet{rambachan/shephard:2021}{]} focuses on the
identification of causal effects of the structural shocks. In this
paper, we consider identification of causal effects of the policy
variable. We illustrate the difference between these two causal effects
and an application to SVAR using the following two examples.
\begin{example}
\label{Example: Demand and Supply}Consider the following system of
simultaneous equations,
\begin{align}
Y_{t} & =\beta D_{t}+\eta_{t}\qquad\mathrm{and}\qquad D_{t}=aY_{t}+e_{t},\label{Eq. (1) Demand Curve}
\end{align}
where the first equation is the demand curve, the second is the supply
curve, $Y_{t}$ and $D_{t}$ are the observed price and quantity,
and $\eta_{t}$ and $e_{t}$ are the structural shocks. The parameter
$\beta$ captures the slope of the demand function, which corresponds
to the causal effect $\partial Y_{t}^{*}\left(d\right)/\partial d=\beta$.
On the other hand, in a SVAR context one may be interested in the
impulse response of $Y_{t}$ given a shock to supply $e_{t}$. Solving
for the reduced-form of \eqref{Eq. (1) Demand Curve},
\begin{align*}
Y_{t} & =\frac{\beta}{1-\alpha\beta}e_{t}+\frac{1}{1-\alpha\beta}\eta_{t},
\end{align*}
shows that the lag-0 impulse response is $\mathrm{d}Y_{t}\left(e\right)/\mathrm{d}e=\beta/\left(1-a\beta\right)$,
which differs from $\beta$.
\end{example}
\begin{example}
\label{Example: SVAR, Potential Outcomes}Consider the following reduced-form
VAR,
\begin{align*}
V_{t} & =A_{1}V_{t-1}+A_{2}V_{t-2}+\ldots+A_{p}V_{t-p}+u_{t},
\end{align*}
where $V_{t}=(D_{t},\,Y_{t}')'$ is $n\times1$, $D_{t}$ is a scalar,
and $u_{t}$ is a vector of reduced-form VAR innovations. The latter
are related to structural shocks, $\varepsilon_{t}=\left(e_{t},\,\eta'_{t}\right)'$,
via $u_{t}=B_{0}\varepsilon_{t}$ where $B_{0}$ is a non-singular
matrix. Under suitable conditions, $V_{t}$ admits a moving-average
representation $V_{t}=\sum_{j=0}^{\infty}C_{j}\left(A\right)B_{0}\varepsilon_{t-j}$,
where $C_{j}\left(A\right)=\sum_{i=1}^{j}C_{j-i}\left(A\right)A_{i}$
for $j=1,\,2,\ldots$ with $C_{0}\left(A\right)=I_{n}$ and $A_{i}=0$
for $i>p$. Then, the outcome variable admits a moving-average representation,
\begin{align*}
Y_{t} & =\sum_{j=0}^{\infty}c_{ye,j}e_{t-j}+\sum_{j=0}^{\infty}c_{y\eta,j}\eta_{t-j},
\end{align*}
where $c_{ye,j}$ and $c_{y\eta,j}$ are blocks of $C_{j}\left(A\right)B_{0}$
partitioned conformably to $Y_{t}$, $e_{t}$ and $\eta_{t}$. If
$e_{t}$ is the policy shock, the potential outcomes here are defined
as
\begin{align*}
Y_{t,h}\left(\epsilon\right)= & Y_{t,h}\left(\epsilon,\,z\right)=\sum_{j=0,j\neq h}^{\infty}c_{ye,j}e_{t+h-j}+\sum_{j=0}^{\infty}c_{y\eta,j}\eta_{t+h-j}+c_{ye,h}\epsilon.
\end{align*}
The potential outcome $Y_{t,h}^ {}\left(\epsilon\right)$ tells us
what $Y_{t+h}$ would be if $e_{t}=\epsilon$ and it does not depend
upon $z$ since the instrument $Z_{t}$ is excluded from the VAR.
Here the absence of causal effects means that $c_{ye,h}=0$ for all
$h$, coinciding with the canonical condition that the impulse responses
are identically equal to zero.

The potential outcome framework is useful because it allows the study
of nonparametric conditions such that common statistical estimands
(e.g., impulse responses) have a causal interpretation. \citet{olea/stock/watson:2021}
show how to use the instrument $Z_{t}$ to identify the impulse response
coefficient $\phi_{r,e,h}=\partial Y_{t+h}^{\left(r\right)}/\partial e_{t}$
(the effect of $e_{t}$ on the $r$th variable in $Y_{t+h}$). From
the moving-average representation we have $\phi_{r,e,h}=\iota'_{r}C_{h}\left(A\right)B_{0}\iota{}_{1}$
where $\iota_{s}$ denotes the $s$-th standard basis vector. This
shows that $\phi_{r,e,h}$ depends on the $A$'s and the first column
of $B_{0}$. The following assumptions are needed for the identification
of $\phi_{r,e,h}$: (i) $\mathbb{E}(Z_{t}e_{t})=\theta\neq0$ (instrument
relevance) and (ii) $\mathbb{E}(Z_{t}\eta_{t})=0$ (instrument exogeneity).
By (i)-(ii), $B_{0}^{\left(:,1\right)}=B_{0}\iota_{1}$ is identified
up to scale by the covariance between $Z_{t}$ and the reduced-form
innovations $u_{t}$:~$\Gamma=\mathbb{E}(Z_{t}u_{t})=\mathbb{E}(Z_{t}B_{0}\varepsilon_{t})=\theta B_{0}^{\left(:,1\right)}$.
Using the scale normalization $B_{0}^{\left(1,1\right)}=1$ {[}see
\citet{stock/watson:2018} for a discussion{]} we have $\Gamma^{\left(1,1\right)}=\mathbb{E}(Z_{t}e_{t})=\theta$
and $B_{0}^{\left(:,1\right)}=\Gamma/\Gamma^{\left(1,1\right)}=\Gamma/\iota'_{1}\Gamma$.
It follows that $\phi_{r,e,h}$ is identified since $\phi_{r,e,h}=\iota'_{r}C_{h}\left(A\right)\Gamma/\iota'_{1}\Gamma$,
where $A$ can be estimated consistently from the reduced-form VAR
and $\Gamma$ can be estimated consistently by using the VAR residuals
$\widehat{u}_{t}$ in place of $u_{t}$. On the other hand, identifying
the causal effects of the policy $D_{t}$ here would require additional
identification restrictions.

\citet{olea/stock/watson:2021} use shortfalls in OPEC oil production
associated with wars and civil disruptions as an instrument for the
oil supply shock in the SVAR of \citet{killian:2009} who investigates
the effect of oil supply and demand shocks on oil production and prices.
This variable is plausibly correlated with the oil supply shock and,
because the shortfalls are associated with political events such as
wars in the Middle East, it is plausibly uncorrelated with the demand
shocks. Using the analog of the nonparametric conditions we provide
below, applied to the shock $e_{t}$ rather than the policy $D_{t}$,
permits a causal interpretation of the impulse response even when
$\mathbb{E}(Z_{t}e_{t})=0$ for a sub-population.
\end{example}
In the following, we discuss identification of causal effects of the
policy via IV estimands.


\subsection{\label{Subsection: Identification-Conditions}Identification Conditions}

We explicitly allow for endogeneity and rely on IVs. We assume that
the instrument only has a contemporaneous effect on $D_{t}$ so that
we may write $D_{t}=D_{t}(Z_{t})$ where $D_{t}(z)=\varphi(D(\widetilde{V}_{t},\,Y_{t},\,z,\,t),\,e_{t},\,t)$
is the potential treatment assignment at time $t$ when $Z_{t}$ is
set equal to $z\in\mathbf{Z}$. The instrument $Z_{t}$ is assumed
to be (conditionally) independent of the potential outcomes $Y_{t,j}^{*}\left(d,\,z\right)$
and treatments $D_{t}(z)$ but correlated with the observed treatment
$D_{t}$.
\begin{assumption}
\label{Assumption: Independence}$($Independence$)$ For all $d\in\mathbf{D}$,
$z\in\mathbf{Z}$ and $t\geq1$, we have
\begin{align}
\left\{ \left\{ Y_{t,h}^{*}\left(d,\,z\right)\right\} _{h\geq0},\,D_{t}\left(z\right)\right\}  & \bot Z_{t}|\,\widetilde{V}_{t}.\label{Eq: Assumption Independence}
\end{align}
\end{assumption}
Assumption \ref{Assumption: Independence} states that, given $\widetilde{V}_{t}$,
the instrument is as good as randomly assigned.

The second assumption is that potential outcomes $Y_{t,h}^{*}\left(d,\,z\right)$
are a function of $d$ but not of $z$. In studies of causal effects
of monetary policy such as \citet{nakamura/steinsson:2018}, $Z_{t}=1$
if there is an FOMC announcement on day $t$ and $Z_{t}=0$ otherwise.
Then, potential realizations of expected output growth respond to
changes in the monetary policy variable regardless of whether the
change is associated with an FOMC announcement or not.
\begin{assumption}
\label{Assumption: Exclusion }$($Exclusion$)$ For all $d\in\mathbf{D},\,t\geq1$
and $h\geq0$, we have
\begin{align}
\left\{ Y_{t,h}^{*}\left(d,\,z\right)=Y_{t,h}^{*}\left(d,\,z'\right)\right\}  & |\,\widetilde{V}_{t},\qquad\mathrm{for}\,\mathrm{all}\,z,\,z'\in\mathbf{Z}.\label{Eq: Assumption Exclusion Restriction}
\end{align}
\end{assumption}
In a dynamic simultaneous equations model (e.g., a SVAR) the exclusion
restriction requires the instrument not to appear in the causal equation
of interest. In Example \ref{Example: SVAR, Potential Outcomes},
Assumption \ref{Assumption: Exclusion } corresponds to condition
(ii), i.e., $\mathbb{E}(Z_{t}\eta_{t})=0$ where $\eta_{t}$ is composed
of the structural shocks other than $e_{t}$. Under Assumption \ref{Assumption: Exclusion }
we write $Y_{t,h}^{*}\left(d,\,z\right)=Y_{t,h}^{*}\left(d\right)$.

Identification based on IVs requires instrument relevance or ``existence
of a first-stage''. The latter means that $\mathbb{E}(D_{t}\left(z\right)|\,\widetilde{V}_{t})$
is a non-trivial function of $z$. In cross-sectional settings, the
existence of a first-stage is typically assumed to hold for all units
to guarantee strong identification. Strong identification of this
form often fails to hold in applications involving time series data
due to temporary misspecification, bad luck, rare events or parameter
instability. The analysis based on articles in five leading journals
that we report earlier suggests that there are time periods for which
the first-stage exists (strong identification) and others for which
it does not (identification failure). Standard first-stage $F$-tests
are then likely to indicate weak identification since they are based
on averaging these two sub-populations.

We provide a theoretical framework to address this identification
problem by assuming that there are two sub-populations. One comprises
a fraction $\pi_{0}\in\left[0,\,1\right]$ of the overall population
for which the first-stage exists. For the second sub-population, which
comprises a fraction $1-\pi_{0}$ of the population, the first-stage
does not exist. This leads to a new notion of LATE, which we name
$\pi$-LATE, the LATE for the (unknown) $\pi_{0}$ fraction of the
population for which the first-stage exists. If $\pi_{0}=1,$ then
one recovers LATE.

Denote by $|\mathbf{S}_{0,T}|$ the cardinality of $\mathbf{S}_{0,T}$
(i.e., the number of indices in $\mathbf{S}_{0,T})$.
\begin{assumption}
\label{Assumption: First-Stage}$($Partial first-stage$)$ Assume
there exists $\mathbf{S}_{0,T}\subseteq\left\{ 1,\,\ldots,\,T\right\} $
such that $|\mathbf{S}_{0,T}|=\left\lfloor \pi_{0}T\right\rfloor $
with $\pi_{0}\in(0,\,1]$ and for $t\in\mathbf{S}_{0,T}$, $\mathbb{E}(D_{t}\left(z\right)|\,\widetilde{V}_{t})$
is a non-trivial function of $z$, i.e., for $t\in\mathbf{S}_{0,T}$,
$\mathbb{E}(D_{t}\left(z'\right)|\,\widetilde{V}_{t})-\mathbb{E}(D_{t}\left(z\right)|\,\widetilde{V}_{t})\neq0$
for $z',\,z\in\mathbf{Z}$ such that $z\neq z'$.\footnote{We assume that all expectations exist.}
\end{assumption}
Assumption \ref{Assumption: First-Stage} implies that there are two
sub-populations: one for which the first-stage exists and one for
which it does not. An average treatment effect can only be identified
via IVs for the fraction $\pi_{0}$ of the population for which a
first-stage exists.

The next assumption is monotonicity which, under heteroskedasticity-based
identification of monetary policy (see Section \ref{Section: Application: Money Neutrality}),
means that while for some days the FOMC announcement does not coincide
with higher volatility in the policy variable, all of those days in
which the announcement affects the volatility of the policy variable,
volatility is shifted up.
\begin{assumption}
\label{Assumption: Monotonicity}$($Monotonicity$)$ $\mathbf{D}\subseteq\mathbb{R}$.
For all $z,\,z'\in\mathbf{Z}$ and $t\in\mathbf{S}_{0,T},$ either
$D_{t}\left(z\right)\geq D_{t}\left(z'\right)$ or $D_{t}\left(z'\right)\geq D_{t}\left(z\right)$
with probability 1.
\end{assumption}
If $\pi_{0}=1$ (so $|\mathbf{S}_{0,T}|=T$), the condition reduces
to that in \citet{imbens/angrist:1994}.

Following \citet{kolesar/plagborgmoller:2025}, we impose the following
assumption.
\begin{assumption}
\label{Assumption KP 2024}For all $t\geq1$ and $h\geq0,$ (i) $Y_{t,h}^{*}\left(\cdot\right)$
is locally absolutely continuous on $\mathbf{D}$ and (ii) $\mathbb{E}\left[\left.\int_{\mathbf{D}}|\partial Y_{t,h}^{*}\left(d\right)/\partial d|\mathrm{d}d\right|\widetilde{V}_{t}\right]<\infty$.
\end{assumption}
Assumption \ref{Assumption KP 2024} allows $D_{t}$ to be either
discrete, continuous or mixed. When $D_{t}$ is discrete or mixed,
it is implicitly assumed that to deal with the gaps in the support
of $D_{t}$ one extends $Y_{t,h}^{*}\left(\cdot\right)$ to $\mathbf{D}$
such that the extension is locally absolutely continuous. The support
of $D_{t}$ is allowed to be unbounded. These conditions are weaker
than counterparts imposed in the recent literature {[}cf. \citet{casini/mccloskey:2024}
and \citet{rambachan/shephard:2021}{]}, in particular local absolute
continuity replaces differentiability of $Y_{t,h}^{*}\left(\cdot\right)$
plus bounded support of $D_{t}$. It allows the application of the
fundamental theorem of calculus to $Y_{t,h}^{*}\left(\cdot\right)$
without requiring the support of $D_{t}$ to be bounded.

\subsection{\label{Subsection: Identification-Results}Identification Results}

\subsubsection{\label{Subsubsection: Identification of Causal Effects}Identification
of Causal Effects}

We first discuss the case of a discrete instrument. When the first-stage
does not exist for all $t$, it is useful to define an IV estimand
corresponding to the sub-population for which it does. Let the generalized
Wald estimand be defined for all $z',\,z\in\mathbf{Z}$ by
\begin{align}
\beta_{\pi,t,h}\left(\widetilde{v}\right) & =\frac{\mathbb{E}\left(Y_{t+h}|\,Z_{t}=z',\,\widetilde{V}_{t}=\widetilde{v}\right)-\mathbb{E}\left(Y_{t+h}|\,Z_{t}=z,\,\widetilde{V}_{t}=\widetilde{v}\right)}{\mathbb{E}\left(D_{t}|\,Z_{t}=z',\,\widetilde{V}_{t}=\widetilde{v}\right)-\mathbb{E}\left(D_{t}|\,Z_{t}=z,\,\widetilde{V}_{t}=\widetilde{v}\right)},\qquad\text{for }t\in\mathbf{S}_{0,T},\label{Eq. (beta_j(v)) pi-LATE}
\end{align}
where $\widetilde{v}\in\mathbf{V}$. This is the ratio of a reduced-form
generalized impulse response to a first-stage generalized impulse
response for $t\in\mathbf{S}_{0,T}$. We show that for $t\in\mathbf{S}_{0,T}$,
the estimand $\beta_{\pi,t,h}\left(\widetilde{v}\right)$ identifies
a weighted average of causal effects for the compliers. Recall that
$t\in\mathbf{S}_{0,T}$ and $\pi_{0}$ are related by $|\mathbf{S}_{0,T}|=\left\lfloor \pi_{0}T\right\rfloor $.
When $\pi_{0}=1$ and there is no conditioning on $\widetilde{V}_{t}=\widetilde{v}$,
$\beta_{1,t,h}$ reduces to the Wald estimand considered by \citet{rambachan/shephard:2021}.
For $t\notin\mathbf{S}_{0,T}$, $\beta_{\pi,t,h}$ does not identify
a causal effect because the denominator of \eqref{Eq. (beta_j(v)) pi-LATE}
is equal to zero.

We show that for $t\in\mathbf{S}_{0,T}$, the generalized Wald estimand
is equal to a weighted average of marginal effects where the latter
are the derivatives $\partial Y_{t,\,h}^{*}\left(d\right)/\partial d$.
\begin{prop}
\label{Proposition: Continuous pi-LATE}($\pi$-LATE) Let Assumptions
\ref{Assumption: Independence}-\ref{Assumption KP 2024} hold. For
$t\in\mathbf{S}_{0,T}$, $h\geq0$, $\widetilde{v}\in\mathbf{V}$
and $z',\,z\in\mathbf{Z}$, we have
\begin{align}
\beta_{\pi,t,h}\left(\widetilde{v}\right) & =\int_{\mathbf{D}}\mathbb{E}\left[\left.\frac{\partial Y_{t,\,h}^{*}\left(d\right)}{\partial d}\right|\,D_{t}\left(z\right)\leq d\leq D_{t}\left(z'\right),\,\widetilde{V}_{t}=\widetilde{v}\right]w_{t}\left(d|\,\widetilde{v}\right)\mathrm{d}d,\quad\mathrm{where}\label{Eq. pi-LATE}\\
w_{t}\left(d|\,\widetilde{v}\right) & =\frac{\mathbb{P}\left(D_{t}\left(z\right)\leq d\leq D_{t}\left(z'\right)|\,\widetilde{V}_{t}=\widetilde{v}\right)}{\int_{\mathbf{D}}\mathbb{P}\left(D_{t}\left(z\right)\leq d\leq D_{t}\left(z'\right)|\,\widetilde{V}_{t}=\widetilde{v}\right)\mathrm{d}r}\geq0\quad\mathrm{and}\quad\int_{\mathbf{D}}w_{t}\left(d|\,\widetilde{v}\right)\mathrm{d}d=1.\nonumber
\end{align}
\end{prop}
Proposition \ref{Proposition: Continuous pi-LATE} shows that $\beta_{\pi,t,h}\left(\widetilde{v}\right)$
identifies a weighted average of causal effects for compliers, characterized
by $D_{t}(z')>D_{t}(z)$, for observations with a first-stage, with
weights $w_{t}\left(d|\,\widetilde{v}\right)$ determined by the (conditional)
likelihood that $D_{t}\left(z\right)\leq d\leq D_{t}\left(z'\right)$.
We refer to the average treatment effect on the right-hand side of
\eqref{Eq. pi-LATE} as the time-$t$ $\pi$-LATE since it is the
LATE for the observations in this sub-population, which is a fraction
$\pi_{0}$ of the whole population. In practice, the IV estimand $\beta_{\pi,t,h}\left(\widetilde{v}\right)$
is characterized by two types of averaging. First, there is averaging
over time. For any treatment $d$, the average involves only those
observations whose treatment variable can be induced to change by
a change in the instrument and is computed only over those observations
that satisfy the first-stage (i.e., $t\in\mathbf{S}_{0,T}$). The
second averaging is over different treatment values $d$ at the same
date $t$. This is reflected in the weight $w_{t}\left(\cdot\right)$
which is proportional to the number of observations in $\mathbf{S}_{0,T}$
for which $D_{t}\left(z\right)\leq d\leq D_{t}\left(z'\right)$. Indeed,
under regularity conditions permitting one to change the order of
differentiation and integration, viz.,
\[
\mathbb{E}\left[\left.\frac{\partial Y_{t,\,h}^{*}\left(d\right)}{\partial d}\right|\,D_{t}\left(z\right)\leq d\leq D_{t}\left(z'\right),\,\widetilde{V}_{t}=\widetilde{v}\right]=\frac{\partial}{\partial d}\mathbb{E}\left[\left.Y_{t,\,h}^{*}\left(d\right)\right|\,D_{t}\left(z\right)\leq d\leq D_{t}\left(z'\right),\,\widetilde{V}_{t}=\widetilde{v}\right],
\]
$\beta_{\pi,t,h}\left(\widetilde{v}\right)$ can be interpreted as
a local average marginal effect.

Stationarity of the conditional joint distribution of the average
potential outcome and treatment assignment functions for observations
with a first-stage lends further interpretability to the generalized
Wald estimand. Specifically, if $\{Y_{t,h}^{*}\left(d\right),D_{t}(z)\}|\widetilde{V}_{t}$
is identically distributed across $t$ for all $t\in\mathbf{S}_{0,T}$,
$d\in\mathbf{D}$ and $z\in\mathbf{Z}$, Proposition \ref{Proposition: Continuous pi-LATE},
immediately implies that $\beta_{\pi,t,h}$ is equal for all $t\in\mathbf{S}_{0,T}$.
Given this, we can write $\beta_{\pi,t,h}=\beta_{\pi,h}$, making
explicit that the generalized Wald estimand \eqref{Eq. (beta_j(v)) pi-LATE}
equals a weighted average of causal effects for members of the sub-population
with a first-stage, which represents a $\pi_{0}$-sized fraction of
the total population. Under this assumption, we refer to the average
causal effect inside of the integral as $\pi$-LATE since it is a
LATE for a member of the $\mathbf{S}_{0,T}$ sub-population whose
treatment variable can be induced to change by a change in the instrument.

The sample counterpart to the generalized Wald estimand \eqref{Eq. (beta_j(v)) pi-LATE}
involves replacing the conditional expectations with sample estimates
based upon observations $t\in\mathbf{S}_{0,T}$, yielding an estimator
of a causal effect. When Assumption \ref{Assumption: First-Stage}
holds with $\pi_{0}\in\left(0,\,1\right)$, the full sample estimand,
i.e., the ratio of the time averages of the numerator and denominator
of \eqref{Eq. (beta_j(v)) pi-LATE}, is a poor representative of the
full sample average treatment effects because it includes observations
for which the instrument is not relevant in the averaging. We caution
that the usual practice of estimating the conditional expectations
in \eqref{Eq. (beta_j(v)) pi-LATE} with full sample estimates will
not estimate the full sample LATE, but $\pi$-LATE.

\citet{angrist/graddy/imbens:2000} and \citet{rambachan/shephard:2021}
consider related results in cross-sectional and time series settings.
The difference here is that we do not require $D_{t}$ to be continuous
or that the first-stage holds for all $t.$ \citet{kolesar/plagborgmoller:2025}
established a similar result for the slope coefficient in the population
version of the ``reduced-form'' regression of the outcome $Y_{t+h}$
onto $Z_{t}$ where they imposed no restriction on the first-stage
and allowed for a continuous instrument.

A connection to program evaluation with binary policy actions arises
when we map a dynamic problem with continuous variables into one with
binary policy actions and instruments. For example, consider the analysis
of causal effects of monetary policy using heteroskedasticity-based
identification {[}cf. \citet{nakamura/steinsson:2018} and \citet{rigobon/sack:2003}{]}.
Define a binary instrument $Z_{t}$ with $Z_{t}=1$ if there is a
scheduled announcement on day $t$ and $Z_{t}=0$ otherwise. The policy
$\Delta i_{t}$ typically reflects changes in short-term interest
rates. Identification relies on higher volatility in $\Delta i_{t}$
during announcement days (policy sample) compared to non-announcement
days (control sample). Think about mapping $|\Delta i_{t}|$ into
a binary treatment such that $D_{t}=1$ if $|\Delta i_{t}|\geq\delta$
for some threshold $\delta>0$ and $D_{t}=0$ if $|\Delta i_{t}|<\delta$
{[}cf. \citet{rigobon/sack:2003}{]}. Here $\pi$-LATE captures the
average treatment effect for the sub-population whose interest rate
changes exceed $\delta$ only when there is an announcement (i.e.,
when $Z_{t}=1$). Observations where $|\Delta i_{t}|<\delta$ regardless
of announcements are ``never-takers,'' while those with $|\Delta i_{t}|\geq\delta$
regardless of announcements are ``always-takers.'' Under monotonicity,
these groups form the non-compliers, whose responses are driven by
idiosyncratic factors other than announcement-specific effects. In
Section \ref{Section: Application: Money Neutrality} we document
regimes where the volatility of $\Delta i_{t}$ is high even in the
absence of announcements.

\citet{sojitra/syrgkanis:2025} study dynamic treatment regimes with
one-sided compliance where treatments in each period may depend on
past instruments, treatments, outcomes, and confounding factors, while
instruments in each period are generated based on prior instruments,
treatments, and states. This setting encompasses applications such
as digital recommendation systems and adaptive medical trials. Their
focus is on the causal effect of treatment histories on long term
outcomes, rather than of one-time shocks or single policy shifts on
outcomes at horizon $h$. Under binary instruments and treatments,
they establish nonparametric identification of the expected values
of multi-period treatment effect contrasts for the corresponding complier
subpopulations, which they refer to as dynamic LATE.

\subsubsection{\label{Subsubsec: Identification of Compliers}Identification of
Compliers and Exclusion Restriction }

A practical challenge for the $\pi$-LATE framework, and LATE frameworks
in general, is that the sub-population of compliers is unknown. However,
in time series settings with binary instruments, we show below that
one can identify the compliers individually, i.e., to determine whether
each observation $t$ is a complier. In this section, we consider
a binary instrument, e.g., $Z_{t}=1$ if $t$ is an FOMC meeting day
and $Z_{t}=0$ otherwise. Under Assumption \ref{Assumption: Monotonicity},
assume without loss of generality that $D_{t}(1)\geq D_{t}(0)$ for
all $t$. Then, observation $t_{0}\in\mathbf{S}_{0,T}$ is a complier
if and only if $D_{t_{0}}\left(1\right)>D_{t_{0}}\left(0\right)$
with probability one\textemdash if the treatment changes in response
to the instrument.

We begin with the following assumption which states that each observation
is either a complier or a non-complier with certainty.
\begin{assumption}
\label{Assumption: Pure behavior}(Deterministic complier status)
For each $t$ either $\mathbb{P}\left(D_{t}\left(1\right)>D_{t}\left(0\right)\right)=1$
or $\mathbb{P}\left(D_{t}\left(1\right)>D_{t}\left(0\right)\right)=0$.
\end{assumption}
Assumption \ref{Assumption: Pure behavior} rules out cases where
$\mathbb{P}\left(D_{t}\left(1\right)>D_{t}\left(0\right)\right)=p$
for some $p\in\left(0,\,1\right)$. A non-complier cannot be characterized
by $\mathbb{P}\left(D_{t}\left(1\right)>D_{t}\left(0\right)\right)>0.$
The latter probability must be zero. Under Assumption \ref{Assumption: Pure behavior},
Lemma \ref{Lemma: D1>D0 E(D1)>E(D0)} in the supplement shows that
$\mathbb{P}\left(D_{t_{0}}\left(1\right)>D_{t_{0}}\left(0\right)\right)=1$
is equivalent to $\mathbb{E}\left(D_{t_{0}}\left(1\right)\right)>\mathbb{E}\left(D_{t_{0}}\left(0\right)\right)$.
This equivalence implies that compliers can be identified by comparing
the expected treatment values under different instrument values.\footnote{Note that this result is different from that in Lemma 2.1 in \citet{abadie:2003}
who shows that under several assumptions the proportion of compliers
can be identified by $\mathbb{E}\left(D_{i}\left(1\right)\right)-\mathbb{E}\left(D_{i}\left(0\right)\right)$
in a cross-sectional setting. He uses this lemma to show that any
statistical characteristic that can be defined in terms of moments
of the joint distribution of $\left(Y_{i},\,D_{i},\,Z_{i}\right)$
is identified for compliers. He then remarks that it is not possible
to identify compliers individually under these assumptions.} Under mild smoothness assumptions that we discuss below, the latter
two expected values can be estimated consistently from the sample
so that we can determine whether $t_{0}$ is a complier in large samples
by looking at the corresponding inequality based on sample quantities.

Let $\mathbf{P}\subset\{1,\ldots,T\}$ denote the ``policy sample'',
the set of observations for which $Z_{t}=1$ so that $D_{t}=D_{t}(1)$
for all $t\in\mathbf{P}$, and let $\mathbf{C}=\{1,\ldots,T\}\setminus\mathbf{P}$
denote the ``control sample'', where $D_{t}=D_{t}(0)$. It is reasonable
to assume that, for a given value of the instrument, the potential
treatment assignments vary smoothly over time. Suppose we wish to
determine whether an observation $t_{0}\in\mathbf{P}$ is a complier.
Since $D_{t_{0}}\left(0\right)$ is not observed, under time-smoothness
we approximate $\mathbb{E}\left(D_{t_{0}}\left(0\right)\right)$ by
averaging nearby observations in the control sample. Letting $N_{0}(t_{0})$
denote the $n_{0}$ largest indices $s\in\mathbf{C}$ such that $s\leq t_{0}-1,$
this implies
\[
\overline{D}_{C,t_{0}-1,n_{0}}\equiv n_{0}^{-1}\sum_{s\in N_{0}(t_{0})}D_{s}\overset{\mathbb{P}}{\rightarrow}\mathbb{E}\left(D_{t_{0}-1}\left(0\right)\right)
\]
as $n_{0}\rightarrow\infty$ with $n_{0}/|\mathbf{C}|\rightarrow0$
under mild conditions. In addition, it follows that $\mathbb{E}\left(D_{t_{0}-1}\left(0\right)\right)$
is close to $\mathbb{E}\left(D_{t_{0}}\left(0\right)\right)$. A similar
argument can be applied to $\mathbb{E}\left(D_{t_{0}}\left(1\right)\right)$
using adjacent days in the policy sample: we have $\overline{D}_{P,t_{0},n_{1}}\overset{\mathbb{P}}{\rightarrow}\mathbb{E}\left(D_{t_{0}}(1)\right)$
as $n_{1}\rightarrow\infty$ with $n_{1}/|\mathbf{P}|\rightarrow0$,
where $\overline{D}_{P,t_{0},n_{1}}=n_{1}^{-1}\sum_{s\in N_{1}(t_{0})}D_{s}$
and $N_{1}(t_{0})$ denotes the $n_{1}$ largest indices $s\in\mathbf{P}$
such that $s\leq t_{0}$. Thus, observation $t_{0}\in\mathbf{P}$
is a complier if and only if $\overline{D}_{P,t_{0},n_{1}}-\overline{D}_{C,t_{0}-1,n_{0}}\overset{\mathbb{P}}{\rightarrow}c$
as $n_{0},n_{1}\rightarrow\infty$ with $n_{0}/|\mathbf{C}|,n_{1}/|\mathbf{P}|\rightarrow0$
for any $c>0$.

Intuitively, even though $D_{t_{0}}\left(0\right)$ is not observed
when $t_{0}\in\mathbf{P}$, observations close to $t_{0}$ characterized
by no FOMC announcement provide information about what $\mathbb{E}\left(D_{t_{0}}(0)\right)$
would have been in the absence of an FOMC announcement.\footnote{One could also use the observations to the right of $t_{0}$ to construct
$\overline{D}_{C,t_{0}+1,n}$, i.e., $D_{t_{0}+1},\ldots,\,D_{t_{0}+n}$.} There are about six weeks in between any two FOMC meetings, and so
$n_{0}\approx30$. Alternatively, following \citet{nakamura/steinsson:2018}
the control sample could include all Tuesdays and Wednesdays that
are not FOMC meeting days. Nevertheless, one can skip the observation
that pertains to the previous meeting, say $D_{t_{-1}}\left(0\right)$,
whose realization is not observed, and continue averaging using the
observations prior to that meeting as well to construct the average
$\overline{D}_{C,t_{0}-1,n_{0}}$ possibly applying down-weighting
for observations further in time from $t_{0}$, i.e., use $\ldots,D_{t_{-1}-1},\,D_{t_{-1}+1},\,D_{t_{-1}+2}\ldots,\,D_{t_{0}-2},\,D_{t_{0}-1}$.
Similarly, observations in $\mathbf{P}$ close to $t_{0}$ provide
information about what $\mathbb{E}\left(D_{t_{0}}\left(1\right)\right)$
would have been, though here the successive observations are separated
chronologically by the observations in the control sample $\mathbf{C}$.

We now present the formal result for identification of the compliers.
The following two assumptions can be justified in large samples when
the mean (potential) treatment assignments in both the control and
policy samples vary smoothly over time. Under an infill asymptotic
embedding where the original observations indexed by $t=1,\ldots,\,T$
are mapped into the unit interval $[0,\,1]$ via $u=t/T$, if $\lim_{T\rightarrow\infty}\mathbb{E}(D_{Tu}(z))$
is continuous in $u$ under a fixed instrument value $z\in\mathbf{Z}$,
the following assumptions hold. This type of continuity accommodates
general forms of smoothly time-varying means but not abrupt breaks
in mean.\footnote{However, breaks in the mean of the assignment process can be estimated
under some conditions as we explain below. Then, time-smoothness is
required to hold only in regimes defined by successive break dates.}
\begin{assumption}
\label{Assumption: Local LLN}(i) For any $t\in\mathbf{C},$ $\overline{D}_{C,t,n}\overset{\mathbb{P}}{\rightarrow}\mathbb{E}\left(D_{t}\right)$
as $n\rightarrow\infty$ with $n/|\mathbf{C}|\rightarrow0.$ (ii)
For $t\in\mathbf{P}$ $\mathbb{E}(D_{t-1}\left(0\right))=\mathbb{E}(D_{t}\left(0\right))$.
\end{assumption}
\begin{assumption}
\label{Assumption Local LLN Treatment Sample}(i) For any $t\in\mathbf{P}$
$\overline{D}_{P,t,n_{}}\overset{\mathbb{P}}{\rightarrow}\mathbb{E}\left(D_{t}\right)$
as $n_{}\rightarrow\infty$ with $n/|\mathbf{P}|\rightarrow0$. (ii)
For $t\in\mathbf{C}$ $\mathbb{E}(D_{t}\left(1\right))=\mathbb{E}(D_{s^{*}\left(t\right)}\left(1\right))$
where $s^{*}\left(t\right)=\mathrm{argmin}_{s\in\mathbf{P}}|t-s|$.
\end{assumption}
Assumption \ref{Assumption: Local LLN}(i) requires a law of large
numbers to apply to the rolling-window sample average of $D_{t}$
at the points of continuity of $\mathbb{E}\left(D_{t}\right)$. It
is a minimal technical assumption. Assumption \ref{Assumption: Local LLN}(ii)
strengthens part (i) a bit by requiring that for $t\in\mathbf{P}$
the potential treatment assignment under the trajectory $Z_{t}=0$
has a locally constant mean. Assumption \ref{Assumption Local LLN Treatment Sample}(i)
adapts Assumption \ref{Assumption: Local LLN}(i) to the observations
in $\mathbf{P}$. This is a stronger assumption since two successive
observations in the policy sample are separated by several observations
in the control sample. Assumption \ref{Assumption Local LLN Treatment Sample}(ii)
requires that $\mathbb{E}\left(D_{t}\left(1\right)\right)$ for $t\in\mathbf{C}$
is equal to the mean of the potential treatment assignment at the
closest date in the policy sample $s^{*}\left(t\right)$. This is
a first moment constancy assumption on the potential treatment assignment
under the trajectory $Z_{t}=1$. Assumption \ref{Assumption: Local LLN}
is used to identify the compliers in the policy sample, while Assumption
\ref{Assumption Local LLN Treatment Sample} is used to identify the
compliers in the control sample.
\begin{thm}
\label{Theorem: Compliers}Let Assumptions \ref{Assumption: Pure behavior}-\ref{Assumption Local LLN Treatment Sample}
hold and $n_{0},n_{1}\rightarrow\infty$ with $n_{0}/|\mathbf{C}|,n_{1}/|\mathbf{P}|\rightarrow0$.
Then:\\
 (i) $t\in\mathbf{P}$ is a complier if and only if $\overline{D}_{P,t,n_{1}}-\overline{D}_{C,t-1,n_{0}}\overset{\mathbb{P}}{\rightarrow}c$
where $c>0$. \\
 (ii) $t\in\mathbf{C}$ is a complier if and only if $\overline{D}_{P,s^{*}\left(t\right),n_{1}}-\overline{D}_{C,t,n_{0}}\overset{\mathbb{P}}{\rightarrow}\widetilde{c}$
where $\widetilde{c}>0$.
\end{thm}
Theorem \ref{Theorem: Compliers} shows that the compliers can be
identified individually. To the best of our knowledge, there is no
equivalent result in the cross-sectional setting. The assumptions
of the theorem are easily satisfied in time series applications. Using
Theorem \ref{Theorem: Compliers} is straightforward: one computes
the difference between two sample averages and check whether it is
greater than zero. Given the sampling uncertainty associated with
the two averages, one can conduct inference using a $t$-statistic
for the null hypothesis $\mathbb{E}\left(D_{t_{0}}\left(1\right)\right)-\mathbb{E}\left(D_{t_{0}}\left(0\right)\right)=0$
($t_{0}$ is not a complier) versus the alternative hypothesis that
$\mathbb{E}\left(D_{t_{0}}\left(1\right)\right)-\mathbb{E}\left(D_{t_{0}}\left(0\right)\right)>0$
($t_{0}$ is a complier).

An additional challenge specific to the $\pi$-LATE framework is that
the set of observations with a first stage $\mathbf{S}_{0,T}$ is
also unknown. However, as the following result states, under Assumption
\ref{Assumption: Monotonicity}, in the absence of covariates $\widetilde{V}_{t}$,
$\mathbf{S}_{0,T}$ is equal to the (identified) set of compliers
\textemdash i.e., observations for which the first-stage holds individually.
\begin{prop}
\label{Proposition: Compliers FS indivdually}Suppose $Z_{t}$ is
binary and let Assumptions \ref{Assumption: First-Stage} without
conditioning on $\widetilde{V}_{t}$, \ref{Assumption: Monotonicity}
and \ref{Assumption: Pure behavior} hold. Then, the set of compliers
coincide with $\mathbf{S}_{0,T}$.
\end{prop}
Knowledge of the compliers sub-population (and hence of the non-compliers
sub-population) can be used to test the exclusion restriction (cf.
Assumption \ref{Assumption: Exclusion }) by comparing the mean outcomes
of groups of non-compliers across different values of the instrument.
For example, one can divide any large subset of non-compliers into
two groups according to their assignment status. If one can reject
the hypothesis that the average outcomes in these two groups is the
same, then the exclusion restriction cannot hold.

Under Assumption \ref{Assumption: Monotonicity} with $D_{t}(1)\geq D_{t}(0)$,
the set of non-compliers is $\mathcal{NC}=\{t\in\{1,\ldots,T\}:D_{t}(1)=D_{t}(0)=D_{t}\}$.
Let $\mathcal{NC}^{s}$ be any non-empty subset of $\mathcal{NC}$
such that $\mathcal{NC}_{\mathbf{P}}^{s}=\mathcal{NC}^{s}\cap\mathbf{P}\neq\emptyset$
and $\mathcal{NC}_{\mathbf{C}}^{s}=\mathcal{NC}^{s}\cap\mathbf{C}\neq\emptyset$.
We can test the exclusion restriction in Assumption \ref{Assumption: Exclusion }
under the following assumption on the subsets $\mathcal{NC}_{\mathbf{P}}^{s}$
and $\mathcal{NC}_{\mathbf{C}}^{s}$.
\begin{assumption}
\label{Assumption: Non-Complier LLN} (i) $\mathbb{E}[Y_{t}^{*}(d,z)|t\in\mathcal{NC}_{\mathbf{P}}^{s}]=\mathbb{E}[Y_{r}^{*}(d,z)|r\in\mathcal{NC}_{\mathbf{C}}^{s}]$
for all $t,r\geq1$, $d\in\mathbf{D}$ and $z\in\mathbf{Z}$. (ii)
$\{D_{t},\,\widetilde{V}_{t}\}|t\in\mathcal{NC}_{\mathbf{P}}^{s}\sim\{D_{r},\,\widetilde{V}_{r}\}|r\in\mathcal{NC}_{\mathbf{C}}^{s}$
for all $t,r\geq1$. (iii) For $\mathbf{R}=\mathbf{C}$ or $\mathbf{P}$,
$|\mathcal{NC}_{\mathbf{R}}^{s}|^{-1}\sum_{t\in\mathcal{NC}_{\mathbf{R}}^{s}}Y_{t}\overset{\mathbb{P}}{\rightarrow}\mathbb{E}\left[Y_{t}|t\in\mathcal{NC}_{\mathbf{R}}^{s}\right]$
as $|\mathcal{NC}_{\mathbf{R}}^{s}|\rightarrow\infty$.
\end{assumption}
Condition (i) states that the potential outcome for non-compliers
is mean-stationary and the mean is the same across control and policy
subsamples. Condition (ii) states that the policy variable and past
observables for non-compliers are distributed identically across the
control and policy subsamples. Condition (iii) states that a law of
large numbers holds for non-compliers observations in both the control
and policy subsamples. As long as the policy sample does not tend
to contain systematic different values of the policy variable $D_{t}$
among non-compliers than the control sample, these are relatively
mild conditions.
\begin{prop}
\label{Proposition: exclusion-restriction}Suppose $Z_{t}$ is binary
and let Assumptions \ref{Assumption: Monotonicity} and \ref{Assumption: Non-Complier LLN}
hold. If Assumption \ref{Assumption: Exclusion } holds, then as $|\mathcal{NC}_{\mathbf{P}}^{s}|,|\mathcal{NC}_{\mathbf{C}}^{s}|\rightarrow\infty$,
\[
|\mathcal{NC}_{\mathbf{P}}^{s}|^{-1}\sum_{t\in\mathcal{NC}_{\mathbf{P}}^{s}}Y_{t}-|\mathcal{NC}_{\mathbf{C}}^{s}|^{-1}\sum_{t\in\mathcal{NC}_{\mathbf{C}}^{s}}Y_{t}\overset{\mathbb{P}}{\rightarrow}0.
\]
\end{prop}
Using Proposition \ref{Proposition: exclusion-restriction} to test
Assumption \ref{Assumption: Exclusion } is simple: since non-compliers
can be identified individually using Theorem \ref{Theorem: Compliers},
one can immediately compute the sample averages specified in Proposition
\ref{Proposition: exclusion-restriction} and conduct inference using
a $t$-statistic for the null hypothesis that the population mean
of $t\in\mathcal{NC}_{\mathbf{P}}^{s}$ is equal to that of $t\in\mathcal{NC}_{\mathbf{C}}^{s}$.
The researcher has the ability to choose the subset of non-compliers
$\mathcal{NC}^{s}$ when implementing this test. The simplest choice
is to set $\mathcal{NC}^{s}=\mathcal{NC}$, however, the researcher
also has the ability to direct the power of the test toward particular
types of non-compliers they may suspect of being more likely to violate
the exclusion restriction. For example, one may wish to focus on $\mathcal{NC}^{s}=\{t\in\mathcal{NC}:D_{t}\geq d^{*}\}$
or $\mathcal{NC}^{s}=\{t\in\mathcal{NC}:D_{t}<d^{*}\}$ for some $d^{*}$
value, such as $d^{*}=|\mathcal{NC}|^{-1}\sum_{t\in\mathcal{NC}}D_{t}$,
in order to test violations of the exclusion restriction for observations
roughly corresponding to ``always-takers'' or ``never-takers''
in the case of a binary treatment.\footnote{Under an analogous assumption to Assumption \ref{Assumption: Non-Complier LLN}
for a function of outcomes $f(Y_{t})$, an analogous result to Proposition
\ref{Proposition: exclusion-restriction} holds. One may use this
fact, for example, to test if the variances of groups of non-compliers
are equal across different values of the instrument, which is implied
by Assumption \ref{Assumption: Exclusion }, or to test the equality
of a set of moments across different values of the instrument. Taking
this logic even further, one could invoke Gilvenko-Cantelli theorems
to show that the difference between the empirical distribution functions
of observations in $\mathcal{NC}_{\mathbf{P}}^{s}$ and $\mathcal{NC}_{\mathbf{C}}^{s}$
converge uniformly to zero under Assumption \ref{Assumption: Exclusion }
and use a test for the equality of distributions such as the Kolmogorov-Smirnov
test. However, we focus here on testing the equality of means across
instrument values because the level of the outcomes, rather than functions
of them, are likely to be of primary importance in practice.}

\section{\label{Section: Application: Money Neutrality}High-Frequency Identification
of Monetary Policy Effects}

To study the effects of monetary policy on real variables, a large
literature has relied on high-frequency identification. This exploits
the fact that at the time of an FOMC meeting a large amount of economic
news is revealed. Here we discuss \citeauthor{rigobon:2003}'s \citeyearpar{rigobon:2003}
heteroskedasticity identification approach which uses a 1-day window
{[}see, e.g., \citet{nakamura/steinsson:2018}{]}, and can be reformulated
as IV-based identification. In Section \ref{Subsection: Heteroske Ident}
we explain when the resulting reduced-form estimands have a causal
meaning within the potential outcome framework of Section \ref{Section: Statistical Framework for Identification of Causal Effects}.
In Section \ref{Subsection: Weak-or-Lack of Identification and the pi-LATE}
we discuss the weak identification problem of current approaches and
show how the $\pi$-LATE framework can be used to strengthen identification.


\subsection{\label{Subsection: Heteroske Ident}Heteroskedasticity-Based Identification}

Consider the following system of equations:
\begin{align}
\tilde{Y}_{t} & =\beta_{0}\tilde{D}_{t}+\eta_{t},\qquad\mathrm{and}\qquad\tilde{D}_{t}=a\tilde{Y}_{t}+e_{t},\label{Eq. (1) NS and RS, Outcome variable}
\end{align}
where $\tilde{Y}_{t}$ is the (demeaned) daily change in an outcome
variable, (e.g., an asset price or a bond yield) and $\tilde{D}_{t}$
is the (demeaned) daily change in the unexpected component of a short-term
interest rate or policy news (e.g., $\Delta i_{t}$ as discussed after
Proposition \ref{Proposition: Continuous pi-LATE}), $\eta_{t}$ is
a shock to $\tilde{Y}_{t}$, $e_{t}$ is the monetary policy shock
and $a$ and $\beta_{0}$ are scalar parameters. The errors $\eta_{t}$
and $e_{t}$ have no serial correlation and are mutually uncorrelated.
The parameter of interest is $\beta_{0}$ which represents the causal
effect of monetary policy on the outcome variable. The model in \eqref{Eq. (1) NS and RS, Outcome variable}
could arise from a bivariate VAR. In fact, one could add a vector
$X_{t}$ of exogenous variables to the model in \eqref{Eq. (1) NS and RS, Outcome variable}.
However, to focus on the main intuition, we follow \citet{nakamura/steinsson:2018}
and we omit $X_{t}$ and lagged terms of $\tilde{Y}_{t}$ and $\tilde{D}_{t}$.
See \citet{casini/mccloskey:2024} for a detailed discussion of why
the lags can be omitted in this setting.

The model in \eqref{Eq. (1) NS and RS, Outcome variable} is a special
case of the generalized framework studied in Section \ref{Section: Statistical Framework for Identification of Causal Effects}.
It is useful because it directly motivates a particular IV estimand.
However, we study the causal interpretation of this estimand in the
general case for which the linear model with stable parameters is
not the correct specification.

Heteroskedasticity-based identification requires that the variance
of the monetary shock increases in the days of FOMC announcements,
while the variance of other shocks is unchanged. Let $T_{P}$ denote
the number of days containing an FOMC announcement (policy sample),
and let $T_{C}$ denote the number of days that do not contain an
FOMC announcement (control sample). Let $\sigma_{e,P}^{2}=T_{P}^{-1}\sum_{t\in\mathbf{P}}\mathbb{E}\left(e_{t}^{2}\right)$
and $\sigma_{e,C}^{2}=T_{C}^{-1}\sum_{t\in\mathbf{C}}\mathbb{E}\left(e_{t}^{2}\right)$
be the average variance of the monetary policy shock in the policy
and control samples. Define $\sigma_{\eta,P}^{2}$ and $\sigma_{\eta,C}^{2}$
similarly. Then, the identification condition is
\begin{align}
\sigma_{e,P} & >\sigma_{e,C}\qquad\mathrm{and}\qquad\sigma_{\eta,P}=\sigma_{\eta,C}.\label{Eq. Vol_P >Vol_C}
\end{align}
Identification can be shown analytically by first solving for the
reduced-form of \eqref{Eq. (1) NS and RS, Outcome variable}:
\begin{align*}
\tilde{Y}_{t} & =\frac{1}{1-a\beta_{0}}\left(\eta_{t}+\beta_{0}e_{t}\right),\qquad\tilde{D}_{t}=\frac{1}{1-a_{}\beta_{0}}\left(a_{}\eta_{t}+e_{t}\right).
\end{align*}
Let $\Sigma_{i}$ denote the covariance matrix of $[\tilde{Y}_{t},\,\tilde{D}_{t}]'$
in the subsample $i=P,\,C$. It follows that
\begin{align*}
\Sigma_{i} & =\frac{1}{\left(1-a\beta_{0}\right)^{2}}\begin{bmatrix}\sigma_{\eta,i}^{2}+\beta_{0}^{2}\sigma_{e,i}^{2} & \beta_{0}\sigma_{e,i}^{2}+a\sigma_{\eta,i}^{2}\\
\beta_{0}\sigma_{e,i}^{2}+a\sigma_{\eta,i}^{2} & \sigma_{e,i}^{2}+a^{2}\sigma_{\eta,i}^{2}
\end{bmatrix},\qquad i=P,\,C.
\end{align*}
It is typical in the literature to assume within-regime covariance-stationarity,
i.e., $\mathbb{E}\left(e_{t}^{2}\right)$ and $\mathbb{E}\left(\eta_{t}^{2}\right)$
are constant within each subsample $\mathbf{P}$ and $\mathbf{C}$
which is, however, restrictive for economic time series. It turns
out that this is not necessary for identification. Volatilities can
be time-varying as long as the average volatilities $\sigma_{e,i}$
and $\sigma_{\eta,i}$ $\left(i=P,\,C\right)$ satisfy \eqref{Eq. Vol_P >Vol_C}.

When \eqref{Eq. (1) NS and RS, Outcome variable} is correctly specified,
i.e., the true model is linear with stable parameters, the parameter
$\beta_{0}$ can be identified using \eqref{Eq. Vol_P >Vol_C} by
taking the difference between the covariance matrices in the policy
and control samples:
\begin{align}
\beta_{0} & =\frac{\Delta\Sigma^{\left(1,2\right)}}{\Delta\Sigma^{\left(2,2\right)}}=\frac{T_{P}^{-1}\sum_{t\in\mathbf{P}}\mathrm{Cov}\left(\tilde{Y}_{t},\,\tilde{D}_{t}\right)-T_{C}^{-1}\sum_{t\in\mathbf{C}}\mathrm{Cov}\left(\tilde{Y}_{t},\,\tilde{D}_{t}\right)}{T_{P}^{-1}\sum_{t\in\mathbf{P}}\mathrm{Var}\left(\tilde{D}_{t}\right)-T_{C}^{-1}\sum_{t\in\mathbf{C}}\mathrm{Var}\left(\tilde{D}_{t}\right)},\qquad\mathrm{where}\label{Eq. beta0 NS}\\
\Delta\Sigma & \triangleq\Sigma_{P}-\Sigma_{C}=\frac{\sigma_{e,P}^{2}-\sigma_{e,C}^{2}}{\left(1-a_{1}\beta_{0}\right)^{2}}\begin{bmatrix}\beta_{0}^{2} & \beta_{0}\\
\beta_{0} & 1
\end{bmatrix}.\nonumber
\end{align}
To determine which average treatment effect this approach identifies
in the general framework, we re-frame this problem in terms of instrumental
variables as follows. Let $Z_{t}=1$ for $t\in\mathbf{P}$ and $Z_{t}=0$
for $t\in\mathbf{C}$. Multiply both sides of \eqref{Eq. (1) NS and RS, Outcome variable}
by $\tilde{D}_{t}$ to yield $\tilde{D}_{t}\tilde{Y}_{t}=\beta_{0}\tilde{D}_{t}^{2}+\tilde{D}_{t}\eta_{t}.$
We can use $Z_{t}$ as an instrument for $\tilde{D}_{t}^{2}$. The
first-stage is $\tilde{D}_{t}^{2}=\theta Z_{t}+\varepsilon_{t},$
where $\varepsilon_{t}$ is some error term satisfying $\varepsilon_{t}\geq-\theta Z_{t}$.
The resulting Wald estimand is
\begin{align}
\beta_{\pi,t,0}^{*} & =\frac{\mathbb{E}\left(\tilde{D}_{t}\tilde{Y}_{t}|\,Z_{t}=1\right)-\mathbb{E}\left(\tilde{D}_{t}\tilde{Y}_{t}|\,Z_{t}=0\right)}{\mathbb{E}\left(\tilde{D}_{t}^{2}|\,Z_{t}=1\right)-\mathbb{E}\left(\tilde{D}_{t}^{2}|\,Z_{t}=0\right)},\label{Eq. (beta0*)}
\end{align}
which corresponds to the Wald estimand \eqref{Eq. (beta_j(v)) pi-LATE}
for $h=0$, $Y_{t}=\tilde{D}_{t}\tilde{Y}_{t}$, $D_{t}=\tilde{D}_{t}^{2}$
and no conditioning variable $\widetilde{V}_{t}$. Under covariance
stationarity within subsamples $\mathbf{P}$ and $\mathbf{C}$, the
right-hand side of \eqref{Eq. (beta0*)} is equal to the right-hand
side of \eqref{Eq. beta0 NS}. The following corollary of Proposition
\ref{Proposition: Continuous pi-LATE} presents the causal meaning
of $\beta_{\pi,t,0}^{*}$ under the general setting of Section \ref{Section: Statistical Framework for Identification of Causal Effects}.
\begin{cor}
\label{Corollary: pi-LATE HET}$($LATE in heteroskedasticity-based
identification$)$ Let Assumptions \ref{Assumption: Independence}-\ref{Assumption KP 2024}
hold for $Y_{t}=\tilde{D}_{t}\tilde{Y}_{t}$ and $D_{t}=\tilde{D}_{t}^{2}$
with $\tilde{D}_{t}(1)^{2}\geq\tilde{D}_{t}(0)^{2}$. For $t\in\mathbf{S}_{0,T}$,
we have
\begin{align}
\beta_{\pi,t,0}^{*} & =\frac{\int_{\mathbf{D}}\mathbb{E}\left(\left.\frac{\partial\left(\tilde{d}\tilde{Y}_{t,0}^{*}\left(\tilde{d}^ {}\right)\right)}{\partial(\tilde{d}^{2})}\right|\tilde{D}_{t}(1)^{2}\geq\tilde{d}^{2}\geq\tilde{D}_{t}(0)^{2}\right)\mathbb{P}\left(\tilde{D}_{t}(1)^{2}\geq\tilde{d}^{2}\geq\tilde{D}_{t}(0)^{2}\right)\mathrm{d}(\tilde{d}^{2})}{\int_{\mathbf{D}}\mathbb{P}\left(\tilde{D}_{t}(1)^{2}\geq\tilde{d}^{2}\geq\tilde{D}_{t}(0)^{2}\right)\mathrm{d}(\tilde{d}^{2})}.\label{Eq. (beta0*) LATE}
\end{align}
\end{cor}
Corollary \ref{Corollary: pi-LATE HET} shows that the Wald estimand
in \eqref{Eq. (beta0*)} has a causal meaning because it is the ratio
of a reduced-form generalized impulse response of $\tilde{D}_{t}\tilde{Y}_{t}$
to a first-stage generalized impulse response of $\tilde{D}_{t}^{2}$.
More specifically, $\beta_{\pi,t,0}^{*}$ identifies a weighted average
of the derivative of the product between the potential outcome and
policy variable for compliers. Hence, contrary to popular belief,
the causal interpretation of the heteroskedasticity-based estimator
(i.e., Rigobon's estimator) estimator is not the same as that of a
standard IV estimator\textemdash though it remains local in nature
as it averages over compliers. We continue to refer to it as LATE
with the understanding that it is a LATE for $\tilde{D}_{t}\tilde{Y}_{t}$,
not $\tilde{Y}_{t}$ itself.

Here the compliers are the observations for which the announcement
induces a higher volatility of the policy $\tilde{D}_{t}$. In contrast,
the non-compliers are characterized by idiosyncratic or general equilibrium
factors that dominate the news specific to the announcement. That
is, regimes where $\tilde{D}_{t}^{2}$ remains low regardless of the
presence of an announcement correspond to ``never-takers,'' while
regimes where $\tilde{D}_{t}^{2}$ remains high even in the absence
of an announcement correspond to ``always-takers.'' Noting that
$D_{t}=\tilde{D}_{t}^{2}$ in this context, we can apply Theorem \ref{Theorem: Compliers}
to identify the compliers individually. We do so in the empirical
application in Section \ref{Section: Empirical Evidence}.

It is important to consider how the interpretation of the causal effect
identified by ${\beta}_{\pi,t,0}^{*}$ in Corollary \ref{Corollary: pi-LATE HET}
varies with the functional relationship between $\tilde{Y}_{t}$ and
$\tilde{D}_{t}$. Let us begin with the linear case with stable parameters
as in \eqref{Eq. (1) NS and RS, Outcome variable}. From \eqref{Eq. (beta0*)},
simple algebra shows that ${\beta}_{\pi,t,0}^{*}$ reduces to $\beta_{0}$
when the denominator of \eqref{Eq. (beta0*) LATE} is nonzero, which
means that Rigobon's estimator identifies the causal effect of the
policy (i.e., the slope coefficient in \eqref{Eq. (1) NS and RS, Outcome variable}).
This result does not generally extend to the case where $\beta_{0}$
is time-varying or the first-stage is zero. At most one could identify
a $\pi$-LATE provided that Rigobon's estimator is computed over the
sub-population where the first-stage is nonzero. We will return to
this in Section \ref{Subsection: Weak-or-Lack of Identification and the pi-LATE}.

Let us turn to analyzing the consequences of nonlinearities. When
$D_{t}$ and the shock $\eta_{t}$ are additively separable (i.e.,
$Y_{t}=\varphi_{D}(D_{t})+\varphi_{\eta}\left(\eta_{t}\right)$ for
some nonlinear functions $\varphi_{D}\left(\cdot\right)$ and $\varphi_{\eta}\left(\cdot\right)$),
\citet{kolesar/plagborgmoller:2025} show that the estimand resulting
from a regression of $Y_{t}$ on $D_{t}$ using $Z_{t}=(W_{t}-\mathbb{E}\left(W_{t}\right))D_{t}$
as an instrument for which $\mathrm{Cov}\left(D_{t}^{2},\,W_{t}\right)\neq0$
identifies a weighted average of marginal effects of the policy shock
$e_{t}$ with weights that are not guaranteed to be positive. As a
result, the researcher may infer an incorrect sign for the marginal
effects. Thus, this estimand is not weakly causal {[}cf. \citet{blandhol/bonney/mogstad/torgovitsky:2025}{]}.
The authors also note that for the case $Y_{t}=e_{t}\varphi_{\eta}\left(\eta_{t}\right)$
with $\mathbb{E}[\varphi_{\eta}\left(\eta_{t}\right)]=0$ and $e_{t}\bot\eta_{t}$
the estimand is nonzero while the true causal effect of the policy
shock is zero since $\mathbb{E}\left[Y_{t}|\,e_{t}\right]=0$.

Corollary \ref{Corollary: pi-LATE HET} provides even more negative
news about the effect of nonlinearities for heteroskedasticity-based
identification than that shown by \citet{kolesar/plagborgmoller:2025}:
in a general nonparametric model, Corollary \ref{Corollary: pi-LATE HET}
implies that Rigobon's Wald estimand ${\beta}_{\pi,t,0}^{*}$, which
is in general different from the IV estimand examined by \citet{kolesar/plagborgmoller:2025},
does not necessarily equal a weighted average of marginal effects.
The intuition is that the instrument affects $\mathrm{Var}\left(D_{t}\right)$
and not $\mathbb{E}\left(D_{t}\right)$, so variation in the instrument
induces exogenous variation in $D_{t}^{2}$, which has a causal effect
on $D_{t}Y_{t}$ not just $Y_{t}$. In short, it is generally difficult
to interpret ${\beta}_{\pi,t,0}^{*}$ when the true model is nonlinear.
Thus, we concur with the recommendation of \citet{kolesar/plagborgmoller:2025}
that the linearity assumption should be checked carefully when using
heteroskedasticity-based identification. This is likely even more
important in the context of SVARs and local projections than in the
the current event study setting since the former aggregates data over
a month or a quarter while the latter uses relatively higher frequency
data (e.g., a 30-minute or 1-day change in policy and outcome variables
around an announcement), where linearity may be a more credible assumption
since a nonlinear function can be locally well approximated by a linear
one.\footnote{The differences between our results on identification via heteroskedasticity
and those in \citet{kolesar/plagborgmoller:2025} are: (i) they consider
the causal effect of the policy shock $e_{t}$ while we consider the
causal effect of the policy variable $D_{t}$; (ii) they consider
an IV estimand while we explicitly consider Rigobon's estimand motivated
by $\Delta\Sigma^{\left(1,2\right)}/\Delta\Sigma^{\left(2,2\right)}$
in \eqref{Eq. beta0 NS} and as usually implemented in empirical work
based on event studies; (iii) they consider specific nonlinear restrictions
and allow the instrument to be continuous whereas we allow for a general
nonlinear model and consider a binary instrument as motivated by \eqref{Eq. beta0 NS}.}

\subsection{\label{Subsection: Weak-or-Lack of Identification and the pi-LATE}Weak
or Lack of Identification and the Usefulness of $\pi$-LATE}

The key identification condition that the volatility of the policy
variable is higher during FOMC announcement days appears reasonable
in principle, since each announcement day is likely to be associated
with substantial monetary news. However, the volatility of monetary
policy variables can be high for other reasons. There are multi-year
periods during which the volatility of several macroeconomic variables
is elevated. In this case, general equilibrium factors dominate the
news specific to the announcement. For example, during the 2007-09
financial crisis and the Covid-19 pandemic, volatility was high across
many macroeconomic and financial variables. These facts pose serious
challenges for identification, as the first-stage condition may not
hold for all $t$. To see this, examine the denominator of $\beta_{0}$
in \eqref{Eq. beta0 NS}. If the first-stage does not hold for all
$t$, we may have
\begin{align}
T_{P}^{-1}\sum_{t\in\mathbf{P}}\mathrm{Var}\left(\tilde{D}_{t}\right)-T_{C}^{-1}\sum_{t\in\mathbf{C}}\mathrm{Var}\left(\tilde{D}_{t}\right) & \thickapprox0,\label{Eq. VarP-VarC =00003D00003D 0}
\end{align}
which would render the estimate of the average treatment effect highly
imprecise.

Using an $F$-test for weak identification, \citet{lewis:2020} shows
that the monetary policy effects based on a 1-day window in \citet{nakamura/steinsson:2018}
appear to be weakly-identified. We show that this arises from significant
time variation in the volatility of the policy variable within both
policy and control samples. Figure \ref{Figure_Plot_Dt} plots $\tilde{D}_{t}$
(2-Year Treasury yields) for the control and policy samples. The policy
sample includes all regularly scheduled FOMC meeting days from 1/1/2000
to 3/19/2014. The control sample includes all Tuesdays and Wednesdays
that are not FOMC meeting days between 1/1/2000 and 12/31/2012.

There appear to be multiple volatility regimes. Using the structural
break test from \citet{casini/perron:change-point-spectra}, which
allows for stable or smoothly varying volatility under the null and
abrupt breaks under the alternative, we detect three breaks in the
control sample. The first break (April 24, 2007) marks the start of
the 2007-09 financial crisis. The second (July 28, 2009) captures
the crisis period itself, characterized by the highest volatility.
Afterward, volatility returns to pre-crisis levels until the third
break (February 2, 2011), which aligns with the zero lower bound (ZLB)
period and the start of unconventional monetary policy. The final
regime shows the lowest volatility, reflecting initial policy effects
and stabilization.\footnote{We do not test for breaks in the policy sample due to small size ($T_{P}=74$),
treating it as a single regime.}

These findings show significant time variation in $\mathrm{Var}(\tilde{D}_{t})$.
In the second regime, control-sample volatility is close to the policy-sample
average, contributing to the weak identification in \eqref{Eq. VarP-VarC =00003D00003D 0}.
\citet{nakamura/steinsson:2018} find their estimates imprecise and
not economically meaningful for some of the interest rates they use
as outcome variables. \citet{lewis:2020} reports a first-stage $F$-statistic
of 8.11\textemdash well below the 23 critical value\textemdash suggesting
weak identification.

We propose to focus on $\pi$-LATE. The fraction $\pi_{0}$ of the
sample (i.e., all $t\in\mathbf{S}_{0,T}$) that has a first-stage
corresponds to the regimes in the control sample where $\mathrm{Var}(\tilde{D}_{t})$
is low (relative to its average level). For example, it is likely
that the regime $[\widehat{T}_{1}+1,\,\widehat{T}_{2}]$ does not
belong to $\mathbf{S}_{0,T}$ since $\mathrm{Var}(\tilde{D}_{t})$
within this regime appears close to the average volatility in the
policy sample. By construction, it is easier to identify $\pi$-LATE
than full sample LATE. The usefulness of $\pi$-LATE depends on the
magnitude of $\pi_{0}$: a small $\pi_{0}$ implies that identification
is achievable only in a small portion of the population, whereas a
large $\pi_{0}$ indicates that the identified $\pi$-LATE is representative
of a substantial part of the population.\footnote{It is possible that in practice the $\pi_{0}$ fraction of the sample
contains a mixture of strong and weak identification. We discuss weak
identification in the context of $\pi$-LATE formally in Section \ref{Section: Inference}.}

\begin{center}
\begin{figure}
\includegraphics[width=16cm,totalheight=8cm]{Figure_Breaks}

\caption{{\small\label{Figure_Plot_Dt}}{\scriptsize Plot of 2-years Treasury
yields in control (top panel) and policy sample (bottom panel). Vertical
broken lines are the estimated break dates using \citeauthor{casini/perron:change-point-spectra}\textquoteright s
\citeyearpar{casini/perron:change-point-spectra} test.}}
\end{figure}
\end{center}

The $\pi$-LATE parameter is the same as the LATE parameter \eqref{Eq. beta0 NS}
in Section \ref{Subsection: Heteroske Ident} but instead of supposing
that a first-stage exists, only uses observations with a nonzero first-stage.
Let $T_{P,S}$ denote the number of days in $\mathbf{S}_{0,T}$ that
contain an FOMC announcement, and let $T_{C,S}$ the number of days
in $\mathbf{S}_{0,T}$ that do not contain an FOMC announcement. This
means $T_{P,S}+T_{C,S}=\pi_{0}T$.\footnote{For notational simplicity we assume that $\pi_{0}T$ is an integer
so that we avoid using the notation $\left\lfloor \pi_{0}T\right\rfloor $,
where $\left\lfloor \cdot\right\rfloor $ denotes the largest smaller
integer function.} Let $\Sigma_{i,S}$ denote the covariance matrix of $[\tilde{Y}_{t},\,\tilde{D}_{t}]'$
in the subsample $i=\mathbf{P},\,\mathbf{C}$ using only observations
$t\in\mathbf{S}_{0,T}$. We have
\begin{align}
\widetilde{\beta}_{\pi,0} & =\frac{\Delta\Sigma_{S}^{\left(1,2\right)}}{\Delta\Sigma_{S}^{\left(2,2\right)}}=\frac{T_{P,S}^{-1}\sum_{t\in\mathbf{P}_{\mathbf{S}}}\mathrm{Cov}\left(\tilde{Y}_{t},\,\tilde{D}_{t}\right)-T_{C,S}^{-1}\sum_{t\in\mathbf{C}_{\mathbf{S}}}\mathrm{Cov}\left(\tilde{Y}_{t},\,\tilde{D}_{t}\right)}{T_{P,S}^{-1}\sum_{t\in\mathbf{P}_{\mathbf{S}}}\mathrm{Var}\left(\tilde{D}_{t}\right)-T_{C,S}^{-1}\sum_{t\in\mathbf{C}_{\mathbf{S}}}\mathrm{Var}\left(\tilde{D}_{t}\right)},\qquad\mathrm{where}\label{Eq. beta0 NS-1}\\
\Delta\Sigma_{S} & =\Sigma_{P,S}-\Sigma_{C,S}=\frac{\sigma_{e,P}^{2}-\sigma_{e,C}^{2}}{\left(1-a_{1}\beta_{0}\right)^{2}}\begin{bmatrix}\beta_{0}^{2} & \beta_{0}\\
\beta_{0} & 1
\end{bmatrix},\nonumber
\end{align}
with $\mathbf{P}_{\mathbf{S}}=\mathbf{P}\cap\mathbf{S}_{0,T}$ and
$\mathbf{C}_{\mathbf{S}}=\mathbf{C}\cap\mathbf{S}_{0,T}.$ Proceeding
as for LATE, the Wald estimand is
\begin{align}
\widetilde{\beta}_{\pi,t,0}^{*} & =\frac{\mathbb{E}\left(\tilde{D}_{t}\tilde{Y}_{t}|t\in\mathbf{P}_{\mathbf{S}}\right)-\mathbb{E}\left(\tilde{D}_{t}\tilde{Y}_{t}|t\in\mathbf{C}_{\mathbf{S}}\right)}{\mathbb{E}\left(\tilde{D}_{t}^{2}|t\in\mathbf{P}_{\mathbf{S}}\right)-\mathbb{E}\left(\tilde{D}_{t}^{2}|t\in\mathbf{C}_{\mathbf{S}}\right)},\label{Eq. (beta0*) pi}
\end{align}
to which Corollary \ref{Corollary: pi-LATE HET} immediately applies
without the (now redundant) qualifier ``for $t\in\mathbf{S}_{0,T}$.''
Under within subsample covariance stationarity, the right-hand side
of \eqref{Eq. (beta0*) pi} is equal to that of \eqref{Eq. beta0 NS-1},
implying $\tilde{\beta}_{\pi,0}=\tilde{\beta}_{\pi,t,0}^{*}$. Therefore,
$\tilde{\beta}_{\pi,0}$ identifies the same $\pi$-LATE, as defined
explicitly in Corollary \ref{Corollary: pi-LATE HET}. $\pi$-LATE
is the average treatment effect for the sub-population for which a
first-stage holds: observations for which $\tilde{D}_{t}^{2}$ is
induced to be higher by the announcement (i.e., the sub-population
of compliers in $\mathbf{S}_{0,T})$.

If the treatment effect is constant across the population {[}e.g.,
as in \eqref{Eq. (1) NS and RS, Outcome variable}{]}, then the $\pi$-LATE
for the sub-population $\mathbf{S}_{0,T}$ is equal to both the LATE
and ATE in the full population. To determine which treatment effect
is identified, we must determine which parts of the sample belong
to $\mathbf{S}_{0,T}$.We discuss this in Sections \ref{Section: ID Failure Test}-\ref{Section: Estimation}.

\section{\label{Section: ID Failure Test}Testing for Full Population Identification
Failure}

In this section, we introduce a test of the null hypothesis that no
subpopulation exists for which a LATE can be identified, even weakly.
In other words, the test assesses whether identifying a sub-population
LATE is possible at all. However, we strongly caution against using
this as a pretest before estimation or inference, as doing so may
introduce pretest bias and invalidates standard inference unless the
inference method is modified to account for the pretest {[}see, e.g.,
\citet{andrews:2018}{]}. Instead, the test should be viewed as a
diagnostic tool for evaluating whether there is evidence of identifiable
sub-population LATEs in a given application. We apply it for this
purpose to several existing studies that appear to face identification
challenges. Notably, such a pretest is unnecessary for conducting
identification-robust inference on sub-population LATEs, which we
discuss in Section \ref{Section: Inference}.

In accord with the analysis of Section \ref{Section: Statistical Framework for Identification of Causal Effects},
consider an IV regression model with a single endogenous variable
and multiple instruments. In matrix format, the structural equation
is
\begin{equation}
Y=D\beta+X\gamma_{1}+u,\qquad t=1,\ldots,\,T,\label{Eq. Structural eq. IV model}
\end{equation}
where $Y$ is a $T\times1$ vector of outcome variables, $D$ is $T\times1$
vector of endogenous variables, $X$ is a $T\times p$ matrix of $p$
exogenous regressors, $u$ is a $T\times1$ vector of error terms,
and $\beta\in\mathbb{R}$ and $\gamma_{1}\in\mathbb{R}^{p}$ are unknown
parameters. The reduced-form equation is
\begin{align}
D_{t} & =Z'_{t}\theta\mathbf{1}\{t\in\mathbf{S}_{0,T}\}+X'_{t}\gamma_{2}+e_{t},\label{Eq. Reduced Form Eq. IV model}
\end{align}
where $Z_{t}$ is a $q\times1$ vector of instruments, $e_{t}$ is
an error term, and $\theta\in\mathbb{R}^{q}$ and $\gamma_{2}\in\mathbb{R}^{p}$
are unknown parameters. For $t\notin\mathbf{S}_{0,T}$, the instrument
$Z_{t}$ is irrelevant. For $t\in\mathbf{S}_{0,T}$, the instrument
$Z_{t}$ is relevant if $\theta\neq0$. We assume that $|\mathbf{S}_{0,T}|=\pi_{0}T$
for some $\pi_{0}\in(0,1]$, noting that this is without loss of generality
since it does not rule out complete identification failure which occurs
when $\theta=0$ for any $\pi_{0}\in(0,1]$. The hypothesis testing
problem is
\[
H_{\theta,0}:\,\theta=0\quad\mathrm{versus}\quad H_{\theta,1}:\,\theta\neq0.
\]
We discuss both the cases for which the sub-population $\mathbf{S}_{0,T}$
is known and unknown. For the sake of the exposition, we focus on
homogeneous $\theta$ in $\mathbf{S}_{0,T}$.\footnote{We could allow for $\theta_{t}\neq0$ for $t\in\mathbf{S}_{0,T}$
at the expense of additional notation and longer proofs, though the
key insights would not change. Actually, the computational procedures
we develop to implement our methods allow $\theta_{t}\neq0$ for $t\in\mathbf{S}_{0,T}$.}

Consider the $\left(\pi T\times T\right)$ selection matrix $S_{T}$
that selects the $\pi T$ rows of a matrix corresponding to the indices
in $\mathbf{S}_{T}$. That is, for an arbitrary $T\times k$ matrix
$A$, $S_{T}A$ is the $\left(\pi T\times k\right)$ matrix whose
elements are the rows of $A$ that correspond to the indices in $\mathbf{S}_{T}$.
For example, if $\mathbf{S}_{T}=\left\{ 1,\ldots,\,0.25T,\,0.75T+1,\ldots,\,T\right\} $,
\[
S_{T}A=\left[A^{\left(1,:\right)\prime}:\cdots:A^{\left(0.25T,:\right)\prime}:A^{\left(0.75T+1,:\right)\prime}:\cdots:A^{\left(T,:\right)\prime}\right]',
\]
where $A^{\left(r,:\right)}$ denotes the $r^{th}$ row of the matrix
$A$. Using the standard projection matrix notation, $P_{A}=A(A^{\prime}A)^{-1}A^{\prime}$
and $M_{A}=I-P_{A}$, let $\widetilde{A}(S_{T})=M_{S_{T}X}S_{T}A$
for any arbitrary $T\times k$ matrix $A$. The following $F$ test
statistic is useful for testing whether $\theta=0$ in the regression
\eqref{Eq. Reduced Form Eq. IV model} when the sub-population $\mathbf{S}_{0,T}$
is known:
\begin{align*}
F_{T}\left(\mathbf{S}_{T}\right) & =\frac{\widetilde{D}\left(S_{T}\right)'\widetilde{Z}\left(S_{T}\right)\widehat{J}(S_{T})^{-1}\widetilde{Z}\left(S_{T}\right)'\widetilde{D}\left(S_{T}\right)}{q\left(\pi T-p-q\right)},
\end{align*}
for $\mathbf{S}_{T}=\mathbf{S}_{0,T}$ and $Z=[Z_{1}:\cdots:Z_{T}]^{\prime}$
and $\widehat{J}(S_{T})$ a consistent estimate of the long-run variance,
\[
\lim_{T\rightarrow\infty}(T\pi)^{-1}\mathrm{Var}(\widetilde{Z}(S_{T})'S_{T}e)
\]
with $e=[e_{1}:\cdots:e_{T}]^{\prime}$. HAC or DK-HAC estimators
can be used to estimate the long-run variance {[}cf. \citet{andrews:91},
\citet{casini_hac} and \citet{newey/west:87}{]}.

For the case of an unknown sub-population, we follow the structural
break literature and search for maximal identification strength over
all sub-populations of minimal size $\pi_{L}T$ that can be partitioned
into $m$ distinct smaller sub-populations, where $\pi_{L}>0$ and
$1\leq m\leq m_{+}$ for some upper bound on the number of regimes
$m_{+}>0$:
\[
F_{T}^{*}=\sup_{\pi\in[\pi_{L},\,1]}\max_{1\leq m\leq m_{+}}\sup_{\mathbf{S}_{T}\in\Xi_{\epsilon,\pi,m,T}}F_{T}\left(\mathbf{S}_{T}\right),
\]
where $\Xi_{\epsilon,\pi,m,T}$ denotes the set of all possible partitions
of a fraction $\pi$ of $\{1,\ldots,T\}$ that involve $m$ regimes
$\left(\left(\lambda_{L,1}T,\,\lambda_{R,1}T\right),\ldots,\,\left(\lambda_{L,m}T,\,\lambda_{R,m}T\right)\right)$
for $\lambda_{L,i},\,\lambda_{R,i}\in\left[0,\,1\right]$ such that
(i) $\lambda_{L,i}<\lambda_{R,i}$ for all $i$, (ii) $\lambda_{R,i}<\lambda_{L,i+1}$
for $i=1,\ldots,\,m-1$, (iii) $\left|\lambda_{R,i}-\lambda_{L,i}\right|\geq\epsilon$
for all $i$ and some (small) $\epsilon>0$ and (iv) $\sum_{i=1}^{m}(\lambda_{R,i}-\lambda_{L,i})=\pi.$
Conditions (i) and (ii) correspond to $T\lambda_{L,i}$ ($T\lambda_{R,i}$)
denoting the start (end) date of regime $i$ within the sub-population
$\mathbf{S}_{T}$ while condition (iii) implies that each regime involves
a non-negligible fraction of the sample. The statistic $F_{T}^{*}$
thus implicitly searches for maximal identification strength over
all possible sub-populations of size $\pi_{L}T$ and larger with less
than $m_{+}$ distinct regimes that are at least a $\epsilon$ fraction
of the overall sample size.

The tuning parameters $\pi_{L}$ and $\epsilon$ determine the types
of sub-populations for which the test can detect identification: smaller
values of $\pi_{L}$ allow detection in smaller sub-populations, while
smaller values of $\epsilon$ enable detection in sub-populations
with shorter regimes. The choice of these lower bounds should be guided
by the empirical context, reflecting the smallest sub-population and
regime sizes for which LATE inference remains meaningful in the application.\footnote{In the structural break literature, common recommendations for $\epsilon$
are 0.05, 0.10 and 0.15. See \citet{casini/perron_Oxford_Survey}
for a review.} In our simulations and empirical applications we set $\pi_{L}=0.6$
and $\epsilon=0.05$.

For $X_{t}^{\prime}$ the $t^{th}$ row of $X$, let $w_{t}=(X'_{t},\,Z'_{t})'$
and $W_{r}\left(\cdot\right)$ denote a $r$-vector of independent
Wiener processes on $\left[0,\,1\right]$. We derive the asymptotic
null distributions of $F_{T}\left(\mathbf{S}_{T}\right)$ and $F_{T}^{*}$
under the following standard high-level assumptions that permit both
heteroskedastic and serially correlated errors. Sufficient conditions
for them can be found in the supplement.
\begin{assumption}
\label{Assumption: w ULLN}$T^{-1}\sum_{t=1}^{\left\lfloor Ts\right\rfloor }w_{t}w'_{t}\overset{\mathbb{P}}{\rightarrow}sQ$,
uniformly in $s\in\left[0,\,1\right]$ for some p.d. matrix $Q$.
\end{assumption}
\begin{assumption}
\label{Assumption: we FCLT}$T^{-1/2}\sum_{t=1}^{\left\lfloor Ts\right\rfloor }w_{t}e_{t}\Rightarrow\Omega_{we}^{1/2}W_{p+q}\left(s\right)$
for some p.d. variance matrix $\Omega_{we}$.
\end{assumption}
\begin{assumption}
\label{Assumption: hat-J uniform consistency}$\widehat{J}(S_{T})$
is p.d. for all $T,$ $\mathbf{S}_{T}\in\Xi_{\epsilon,\pi,m,T}$ and
$\widehat{J}(S_{T})\overset{\mathbb{P}}{\rightarrow}\lim_{T\rightarrow\infty}T^{-1}\mathrm{Var}($
$e^{\prime}S_{T}^{\prime}\widetilde{Z}(S_{T}))$ uniformly in $\mathbf{S}_{T}\in\Xi_{\epsilon,\pi,m,T}$.
\end{assumption}
\begin{thm}
\label{Theorem: Asymptotic Distribution Sup Fstar test}Let Assumptions
\ref{Assumption: w ULLN}-\ref{Assumption: hat-J uniform consistency}
hold. Under $H_{\theta,0}$,
\begin{align*}
F_{T}\left(\mathbf{S}_{T}\right)\Rightarrow F\left(\mathbf{S}\right)\quad\mathrm{if}\quad\mathbf{S}_{T}\in\Xi_{\epsilon,\pi,m,T},\qquad\mathrm{and}\qquad & F_{T}^{*}\Rightarrow\sup_{\pi\in[\pi_{L},\,1]}\max_{1\leq m\leq m_{+}}\sup_{\mathbf{S}\in\Xi_{\epsilon,\pi,m}}F\left(\mathbf{S}\right),
\end{align*}
where $\mathbf{S}=\lim_{T\rightarrow\infty}T^{-1}\mathbf{S}_{T}$,
$\Xi_{\epsilon,\pi,m}=\lim_{T\rightarrow\infty}T^{-1}\Xi_{\epsilon,\pi,m,T}$
and
\[
F\left(\mathbf{S}\right)=\frac{1}{q\pi}\sum_{i=1}^{m}\left\Vert \left(W_{q}\left(\lambda_{R,i}\right)-W_{q}\left(\lambda_{L,i}\right)\right)\right\Vert ^{2}.
\]
\end{thm}
When $\pi=1$ ($\pi_{L}=1$ and $m_{+}=1$), $F_{T}\left(\mathbf{S}_{T}\right)$
($F_{T}^{*}$) reduces to the usual first-stage $F$-statistic for
$\theta=0$ in \eqref{Eq. Reduced Form Eq. IV model}. For $\pi\in(0,\,1)$
($\pi_{L}\in(0,\,1)$), the consistency of tests against $H_{\theta,1}$
using $F_{T}\left(\mathbf{S}_{0,T}\right)$ ($F_{T}^{*}$) follows
from similar arguments as for the $\pi_{0}=1$ case. The asymptotic
null distributions of both $F\left(\mathbf{S}\right)$ and $F_{T}^{*}$
are free of nuisance parameters. The critical values are obtained
via simulations and reported in Table \ref{Table CVs of F* alpha=00003D0.05}
for up to $m_{+}=6$ and up to $q=6$.

\section{Estimation of LATE and Identified Sub-Populations\label{Section: Estimation}}

We discuss estimation of the LATE parameter $\beta$ in \eqref{Eq. Structural eq. IV model}
in both the cases of a known and unknown sub-population $\mathbf{S}_{0,T}$,
as well as estimation of $\mathbf{S}_{0,T}$ itself in the latter
case. When $\mathbf{S}_{0,T}$ is known, estimation of $\beta$ is
an application of IV estimation for which $Z_{t}\mathbf{1}\{t\in\mathbf{S}_{0,T}\}$
is treated as the vector of instruments. Let this estimator be denoted
as $\widehat{\beta}(\mathbf{S}_{0,T})$.

On the other hand, when the sub-population $\mathbf{S}_{0,T}$ is
unknown, we must estimate it first. Although $\mathbf{S}_{0,T}$ can
be estimated consistently in the special case of a binary instrument
under the conditions of Proposition \ref{Proposition: Compliers FS indivdually}
and Theorem \ref{Theorem: Compliers}, it can also be estimated more
generally. We discuss two methods. The first is more computationally
straightforward but the second is more efficient because it uses the
information in both structural and reduced-form equations \eqref{Eq. Structural eq. IV model}-\eqref{Eq. Reduced Form Eq. IV model}.
We follow the structural change literature and assume that $\pi_{0}$
and $m_{0}$ are known, i.e., the practitioner has previously used
the tests from Section \ref{Section: ID Failure Test} to determine
$\pi_{0}$ and $m_{0}$.

We begin with the first estimator. Consider the $T\times T$ matrix
$C_{T}$ that selects the $\pi T$ rows of a matrix corresponding
to the indices in $\mathbf{S}_{T}$ while setting the remaining $(1-\pi)T$
rows to zero. For example, for a $T\times k$ matrix $A$, if $\mathbf{S}_{T}=\left\{ 1,\ldots,\,0.25T,\,0.75T+1,\ldots,\,T\right\} $,
\[
C_{T}A=\left[A^{\left(1,:\right)\prime}:\cdots:A^{\left(0.25T,:\right)\prime}:0_{k\times1}:\cdots:0_{k\times1}:A^{\left(0.75T+1,:\right)\prime}:\cdots:A^{\left(T,:\right)\prime}\right]'.
\]
Let $\overline{A}(C_{T})=M_{X}C_{T}A$ so that for a given $\mathbf{S}_{T}$,
the OLS estimators of $\theta$ and $\gamma_{2}$ in \eqref{Eq. Reduced Form Eq. IV model}
can be expressed as $\widehat{\theta}_{OLS}(\mathbf{S}_{T})=(\overline{Z}(C_{T})^{\prime}\overline{Z}(C_{T}))^{-1}\overline{Z}(C_{T})^{\prime}D$
and $\widehat{\gamma}_{2,OLS}(\mathbf{S}_{T})=(X^{\prime}M_{C_{T}Z}X)^{-1}X^{\prime}M_{C_{T}Z}D$.
Our first estimator of $\mathbf{S}_{0,T}$ minimizes the sum of squared
residuals of the reduced-form:
\[
\widehat{\mathbf{S}}_{T,OLS}=\underset{\mathbf{S}_{T}\in\Xi_{\epsilon,\pi_{0},m_{0},T}}{\mathrm{argmin}}\left(D-C_{T}Z\widehat{\theta}_{OLS}(\mathbf{S}_{T})-X\widehat{\gamma}_{2,OLS}(\mathbf{S}_{T})\right)^{\prime}\left(D-C_{T}Z\widehat{\theta}_{OLS}(\mathbf{S}_{T})-X\widehat{\gamma}_{2,OLS}(\mathbf{S}_{T})\right).
\]
Correspondingly, we estimate $\beta$ with $\widehat{\beta}(\widehat{\mathbf{S}}_{T,OLS})$.

For the second estimator of the sub-population $\mathbf{S}_{0,T}$,
we propose a GLS criterion that minimizes an efficiently weighted
combination of the sum of squared residuals of both the reduced-form
representation of the structural equation \eqref{Eq. Structural eq. IV model}
and the reduced-form equation \eqref{Eq. Reduced Form Eq. IV model}.
That is, the system of equations \eqref{Eq. Structural eq. IV model}-\eqref{Eq. Reduced Form Eq. IV model}
can be written in reduced-form as
\begin{equation}
\vec{y}=W(\mathbf{S}_{T})\xi+\varepsilon,\label{Eq. Reduced Form System}
\end{equation}
where $\vec{y}=(Y',\,D')'$, $W(\mathbf{S}_{0,T})=I_{2}\otimes[C_{0,T}Z:X]$,
$\xi=(\beta\theta',\gamma_{1}^{\prime}+\beta\gamma_{2}^{\prime},\theta',\gamma_{2}^{\prime})^{\prime}$
and $\varepsilon=(u^{\prime}+\beta e^{\prime},e^{\prime})^{\prime}$
with $C_{0,T}$ defined as $C_{T}$ but corresponding to the indices
in $\mathbf{S}_{0,T}$. This is a system of two seemingly unrelated
regressions. Let
\[
\widehat{\xi}_{FGLS}(\mathbf{S}_{T})=(W(\mathbf{S}_{T})^{\prime}\widehat{\Omega}_{\varepsilon}(\mathbf{S}_{T})^{-1}W(\mathbf{S}_{T}))^{-1}W(\mathbf{S}_{T})^{\prime}\widehat{\Omega}_{\varepsilon}(\mathbf{S}_{T})^{-1}\vec{y},
\]
denote a feasible GLS estimator of $\xi$, where $\widehat{\Omega}_{\varepsilon}(\mathbf{S}_{T})$
is a consistent estimator of $\mathbb{E}[\varepsilon\varepsilon^{\prime}|W(\mathbf{S}_{T})]$.
Our second estimator of $\mathbf{S}_{0,T}$ minimizes the following
GLS criterion based upon \eqref{Eq. Reduced Form System}:
\[
\widehat{\mathbf{S}}_{T,FGLS}=\underset{\mathbf{S}_{T}\in\Xi_{\epsilon,\pi_{0},m_{0},T}}{\mathrm{argmin}}\left(\vec{y}-W(\mathbf{S}_{T})\widehat{\xi}_{FGLS}(\mathbf{S}_{T})\right)^{\prime}\widehat{\Omega}_{\varepsilon,\mathbf{S}}^{-1}\left(\vec{y}-W(\mathbf{S}_{T})\widehat{\xi}_{FGLS}(\mathbf{S}_{T})\right).
\]
Correspondingly, we estimate $\beta$ with $\widehat{\beta}(\widehat{\mathbf{S}}_{T,FGLS})$.
In order for $\widehat{\beta}(\widehat{\mathbf{S}}_{T,FGLS})$ to
be provably more efficient than $\widehat{\beta}(\widehat{\mathbf{S}}_{T,OLS})$,
$\widehat{\Omega}_{\varepsilon,\mathbf{S}}$ must be a consistent
estimator of $\mathbb{E}[\varepsilon\varepsilon^{\prime}|W(\mathbf{S}_{0,T})]$.
When $\varepsilon_{t}$ does not exhibit conditional serial correlation
or heteroskedasticity, i.e., $\mathbb{E}[\varepsilon\varepsilon^{\prime}|W(\mathbf{S}_{0,T})]=\Sigma_{\varepsilon}\otimes I_{T}$,
this is feasible since one could simply use $\widehat{\Omega}_{\varepsilon,\mathbf{S}}=\widehat{\Sigma}_{\varepsilon}\otimes I_{T}$,
where $\widehat{\Sigma}_{\varepsilon,i,j}=(T-q-p)^{-1}\hat{\varepsilon}^{i\prime}\hat{\varepsilon}^{j}$
for $i,j=1,2$ with $\hat{\varepsilon}^{1}$ ($\hat{\varepsilon}^{2}$)
equal to the first (last) $T$ elements of $\vec{y}-W(\widehat{\mathbf{S}}_{T,OLS})\widehat{\xi}_{OLS}(\widehat{\mathbf{S}}_{T,OLS})$,
as is standard in seemingly unrelated regression. For serially dependent
$\varepsilon_{t}$, consistent estimation of $\mathbb{E}[\varepsilon\varepsilon^{\prime}|W(\mathbf{S}_{0,T})]$
requires a correctly-specified model for the dependence in $\varepsilon_{t}$,
a strong assumption in some empirical applications. In the supplement
\textcolor{MyBlue}{Casini et al.} \citeyearpar{casini/mccloskey/pala/rolla_Dynamic_Late_Supp_Not_Online}
we present the consistency results about $\widehat{\mathbf{S}}_{T,OLS}$,
$\widehat{\beta}(\widehat{\mathbf{S}}_{T,OLS})$, $\widehat{\mathbf{S}}_{T,FGLS}$
and $\widehat{\beta}(\widehat{\mathbf{S}}_{T,FGLS})$.

In model \eqref{Eq. Structural eq. IV model} the LATE parameter $\beta$
is constant, so $\pi$-LATE is the full population LATE and $\widehat{\beta}(\widehat{\mathbf{S}}_{T,OLS})$
and $\widehat{\beta}(\widehat{\mathbf{S}}_{T,FGLS})$ are consistent
for the LATE parameter $\beta$. They can be precise estimates even
when a first-stage $F$ test detects full sample weak identification
because they use the most-strongly identified subsample of the data.
When the model \eqref{Eq. Structural eq. IV model} is misspecified,
so that LATEs may be nonlinear and time-varying, the estimators $\widehat{\beta}(\widehat{\mathbf{S}}_{T,OLS})$
and $\widehat{\beta}(\widehat{\mathbf{S}}_{T,FGLS})$ are still consistent
for a weighted average the of the LATEs in the $\mathbf{S}_{0,T}$
subsample if the $\mathbf{S}_{0,T}$ subsample exhibits strong identification.

The estimators $\widehat{\mathbf{S}}_{T,OLS}$ and $\widehat{\mathbf{S}}_{T,FGLS}$
and the test statistic $F_{T}^{*}$ solve an optimization problem
over many partitions. This is computationally more complex than problems
in the structural breaks literature, as it involves optimizing both
over sample partitions and identification strength. We address this
challenge by proposing an efficient algorithm based on dynamic programming,
extending the approach of \citet{bai/perron:03} to our setting.\footnote{While \citet{antonie/boldea:2018} consider the case of a single break,
and \citet{magnusson/mavroeidis:2014} study a related context, neither
provide a computational solution\textemdash referring to the problem
as ``computationally demanding.''}

\section{\label{Section: Inference}Identification-Robust Inference}

We consider tests on $\beta$ in \eqref{Eq. Structural eq. IV model}
that are robust to weak identification in both the cases for which
the sub-population $\mathbf{S}_{0,T}$ is known and unknown. The hypothesis
testing problem is $H_{0}:\,\beta=\beta_{0}$ versus $H_{1}:\,\beta\neq\beta_{0}.$
Here we present results for the case of unknown sub-population $\mathbf{S}_{0,T}$
and weak instruments. We also briefly discuss the case of known $\mathbf{S}_{0,T}$
and strong instruments and defer their formal treatment to the supplement.
We rewrite \eqref{Eq. Reduced Form System} as
\begin{align}
y & =\overline{Z}(C_{0,T})\theta a'+X\eta+v,\:\:\mathrm{where\:\:}y=\left[Y:D\right],\:v=\left[v_{1}:e\right],\:a=\left(\beta,\,1\right)',\:\eta=\left[\gamma:\phi\right],\label{Eq. (2.5) AMS}
\end{align}
with $v_{1}=u+\beta e$, $\gamma=\gamma_{1}+\phi\beta$ and $\phi=\gamma_{2}+(X^{\prime}X)^{-1}X^{\prime}C_{0,T}Z\theta$.
When $\mathbf{S}_{0,T}$ is known, it is straightforward to use existing
tests in the identification-robust linear IVs literature to test $H_{0}$
{[}cf. \citet{anderson/rubin:1949}, \citet{andrews/moreira/stock:2006},
\citet{kleibergen:2002} and \citet{moreira:2003}{]}. However, Proposition
\ref{Lemma 1 AMS} in the supplement shows that $Z^{\prime}M_{X}y$
is not a sufficient statistic for $(\beta,\theta')'$ but $\overline{Z}(C_{0,T})^{\prime}y$
is, implying that existing tests suffer a loss in efficiency because
they treat $Z$ rather than $C_{0,T}Z$ as the matrix of IVs. Efficient
tests are therefore functions of $\overline{Z}(C_{0,T})^{\prime}y$.
\citet{magnusson/mavroeidis:2014} consider a model similar to \eqref{Eq. (2.5) AMS}.
Our model specifies that $\theta$ is nonzero in the sub-population
$\mathbf{S}_{0,T}$ and is zero in $\mathbf{S}_{0,T}^{c}$ where $\mathbf{S}_{0,T}^{c}$
is the complement of $\mathbf{S}_{0,T}$. \citet{magnusson/mavroeidis:2014}
allow the first-stage coefficient $\theta_{t}$ to be generally time-varying
for some of their tests. Their tests are based on the full sample
of observations whereas our tests are based on a lower-dimensional
statistic since we do not use the sub-population $\mathbf{S}_{0,T}^{c}$.
This allows us to obtain gains in efficiency.

When $\mathbf{S}_{0,T}$ is known we can apply the results of \citet{andrews/moreira/stock:2006}
to form identification-robust tests of $H_{0}$ vs $H_{1}$ that are
functions of $\overline{Z}(C_{0,T})^{\prime}y$ and are robust to
both heteroskedasticity and autocorrelation (HAR) in the reduced-form
errors $\{v_{t}\}$. Suppose $\widehat{\Sigma}_{N_{1}}(\mathbf{S}_{0,T})$,
$\widehat{\Sigma}_{N_{1},N_{2}}(\mathbf{S}_{0,T})$ and $\widehat{\Sigma}_{N_{2}}(\mathbf{S}_{0,T})$
are consistent estimators of $\Sigma_{N_{1}}(\mathbf{S}_{0})$, $\Sigma_{N_{1},N_{2}}(\mathbf{S}_{0})$
and $\Sigma_{N_{2}}(\mathbf{S}_{0})$ under $H_{0}$, where these
latter quantities are defined by
\begin{gather}
\Sigma_{v\overline{Z}}\left(\mathbf{S}_{0}\right)=\begin{bmatrix}\Sigma_{N_{1}}\left(\mathbf{S}_{0}\right) & \Sigma{}_{N_{1}N_{2}}\left(\mathbf{S}_{0}\right)'\\
\Sigma_{N_{1}N_{2}}\left(\mathbf{S}_{0}\right) & \Sigma_{N_{2}}^{*}\left(\mathbf{S}_{0}\right)
\end{bmatrix},\label{Eq. LRV matrix}\\
\Sigma_{N_{2}}\left(\mathbf{S}_{0}\right)=\Sigma_{N_{2}}^{*}\left(\mathbf{S}_{0}\right)-\Sigma_{N_{1}N_{2}}\left(\mathbf{S}_{0}\right)\Sigma_{N_{1}}^{-1}\left(\mathbf{S}_{0}\right)\Sigma_{N_{1}N_{2}}\left(\mathbf{S}_{0}\right)'\nonumber
\end{gather}
for $\Sigma_{v\overline{Z}}\left(\mathbf{S}_{0}\right)=\Sigma_{v\overline{Z}}\left(\mathbf{S}_{0},\mathbf{S}_{0}\right)$,
with
\begin{gather*}
\Sigma_{v\overline{Z}}\left(\mathbf{S},\mathbf{S}'\right)=\lim_{T\rightarrow\infty}\mathrm{Cov}\left(T^{-1/2}\sum_{t=1}^{T}\begin{bmatrix}v'_{t}b_{0}\overline{Z}_{t}\left(C_{T}\right)\\
v'_{t}\Sigma_{v}^{-1}a_{0}\overline{Z}_{t}\left(C_{T}\right)
\end{bmatrix},T^{-1/2}\sum_{t=1}^{T}\begin{bmatrix}v'_{t}b_{0}\overline{Z}_{t}\left(C_{T}'\right)\\
v'_{t}\Sigma_{v}^{-1}a_{0}\overline{Z}_{t}\left(C_{T}'\right)
\end{bmatrix}\right)
\end{gather*}
for $\mathbf{S}=\lim_{T\rightarrow\infty}T^{-1}\mathbf{S}_{T}$, $\mathbf{S}'=\lim_{T\rightarrow\infty}T^{-1}\mathbf{S}_{T}'$
$b_{0}=(1,-\beta_{0})^{\prime}$ and $a_{0}=(\beta_{0},1)^{\prime}$,
where and $v_{t}$ and $\overline{Z}_{t}\left(C_{T}\right)$ are the
$t$th rows $v$ and $\overline{Z}(C_{T})$.\footnote{See the supplement for details on how to construct these estimators
and for consistency results.} Let $\widehat{\Sigma}_{v}\left(\mathbf{S}_{0,T}\right)=\left(T-q-p\right)^{-1}\widehat{v}\left(\mathbf{S}_{0,T}\right)'\widehat{v}\left(\mathbf{S}_{0,T}\right)$
with $\widehat{v}\left(\mathbf{S}_{0,T}\right)=y-P_{\overline{Z}\left(C_{0,T}\right)}y-P_{X}y$.
Define
\begin{align}
{N}_{1,T}\left(\mathbf{S}_{0,T}\right) & =\widehat{\Sigma}_{N_{1}}^{-1/2}\left(\mathbf{S}_{0,T}\right)T^{-1/2}\overline{Z}\left(C_{0,T}\right)'yb_{0}\qquad\mathrm{and}\label{Eq. (9.7) AMS}\\
{N}_{2,T}\left(\mathbf{S}_{0,T}\right) & =\widehat{\Sigma}_{N_{2}}^{-1/2}\left(\mathbf{S}_{0,T}\right)\left(T^{-1/2}\overline{Z}\left(C_{0,T}\right)'y\widehat{\Sigma}_{v}^{-1}\left(\mathbf{S}_{0,T}\right)a_{0}-\widehat{\Sigma}_{N_{1}N_{2}}\left(\mathbf{S}_{0,T}\right)\widehat{\Sigma}_{N_{1}}^{-1/2}\left(\mathbf{S}_{0,T}\right){N}_{1,T}\left(\mathbf{S}_{0,T}\right)\right).\nonumber
\end{align}
Consider the following HAR versions of the Anderson-Rubin (AR), Lagrange
multiplier (LM) and likelihood ratio statistics based on the sufficient
statistic $\overline{Z}(C_{0,T})'y$:
\begin{align}
AR_{T}(\mathbf{S}_{0,T}) & =M_{1,T}(\mathbf{S}_{0,T}),\qquad\qquad LM_{T}(\mathbf{S}_{0,T})=\frac{M_{1,2,T}(\mathbf{S}_{0,T})^{2}}{M_{2,T}(\mathbf{S}_{0,T})},\label{Eq. (3.4) AMS}\\
LR_{T}(\mathbf{S}_{0,T}) & =\frac{1}{2}\left(M_{1,T}(\mathbf{S}_{0,T})-M_{2,T}(\mathbf{S}_{0,T})+\sqrt{\left(M_{1,T}(\mathbf{S}_{0,T})-M_{2,T}(\mathbf{S}_{0,T})\right)^{2}+4M_{1,2,T}(\mathbf{S}_{0,T})^{2}}\right),\nonumber
\end{align}
where $M_{1,T}(\mathbf{S}_{0,T})={N}_{1,T}\left(\mathbf{S}_{0,T}\right)^{\prime}{N}_{1,T}\left(\mathbf{S}_{0,T}\right)$,
$M_{1,2,T}(\mathbf{S}_{0,T})={N}_{1,T}\left(\mathbf{S}_{0,T}\right)^{\prime}{N}_{2,T}\left(\mathbf{S}_{0,T}\right)$
and $M_{2,T}($ $\mathbf{S}_{0,T})={N}_{2,T}\left(\mathbf{S}_{0,T}\right)^{\prime}{N}_{2,T}\left(\mathbf{S}_{0,T}\right)$.
The conditional likelihood ratio (CLR) test of level $\alpha$ rejects
$H_{0}$ when $LR_{T}(\mathbf{S}_{0,T})>\kappa_{\alpha}(N_{2,T}(\mathbf{S}_{0,T}))$,
where the critical value function $\kappa_{\alpha}(\cdot)$ is defined
such that $\kappa_{\alpha}(n_{2})$ is the $1-\alpha$ quantile of
the large-sample conditional distribution of $LR_{T}(\mathbf{S}_{0,T})$
under $H_{0}$, given $N_{2,T}(\mathbf{S}_{0,T})=n_{2}$:
\[
\frac{1}{2}\left(\mathcal{Z}_{q}'\mathcal{Z}_{q}-n_{2}'n_{2}+\sqrt{\left(\mathcal{Z}_{q}'\mathcal{Z}_{q}-n_{2}'n_{2}\right)^{2}+4(\mathcal{Z}_{q}'n_{2})^{2}}\right),
\]
where $\mathcal{Z}_{q}\sim\mathscr{N}(0,I_{q})$. The critical value
function $\kappa_{\alpha}(\cdot)$ is approximated in \citet{moreira:2003}.
The LM and AR tests reject $H_{0}$ when $LM_{T}>\chi_{1}^{2}(1-\alpha)$
and $AR_{T}>\chi_{q}^{2}(1-\alpha)$, where $\chi_{q}^{2}(1-\alpha)$
denotes the $1-\alpha$ quantile of a chi-squared distribution with
$q$ degrees of freedom.

When $\mathbf{S}_{0,T}$ is known the results of \textcolor{MyBlue}{Andrews et al.}
\citeyearpar{andrews/moreira/stock:2006} imply that the CLR, LM
and AR tests have limiting null rejection probabilities equal to $\alpha$
under weak IV asymptotics, $\theta=c/T^{1/2}$ for some nonstochastic
$c\in\mathbb{R}^{q}$, under a weakening of Assumptions \ref{Assumption 1 AMS}-\ref{Assumption Uniform Consistent Covariance Matrix}
below for which these assumptions need only hold pointwise in $\mathbf{S}_{T}$.
These tests are asymptotically similar and therefore have asymptotically
correct size in the presence of weak IVs.

For the case of an unknown sub-population, the identification-robust
tests in the extant literature no longer apply because the set of
instruments $C_{0,T}Z$ is unknown and must be estimated. In this
section, we show how to form HAR CLR, LM and AR tests with correct
asymptotic null rejection probabilities under both weak and strong
IV asymptotics. To estimate the true sub-population $\mathbf{S}_{0,T}$
when constructing these tests let
\begin{equation}
\widehat{\mathbf{S}}_{T}=\arg\max_{\mathbf{S}_{T}\in\mathcal{S}}M_{2,T}(\mathbf{S}_{T}),\qquad\qquad\mathrm{where}\qquad\qquad\mathcal{S}=\underset{1\leq m\leq m_{+}}{\cup}\underset{\pi\in(\epsilon,\,1]}{\cup}\Xi_{\epsilon,\pi,m,T}.\label{Eq. S_hat MLE}
\end{equation}
Proposition \ref{Lemma 1 AMS unknown subpopulation} in the supplement
shows that the process $\{\overline{Z}(C_{T})'y\}_{\mathbf{S}_{T}\in\mathcal{S}}$
is sufficient for $(\beta,\theta')'$ in a canonical Gaussian setting
analogous to that in \textcolor{MyBlue}{Andrews et al.} \citeyearpar{andrews/moreira/stock:2006}
so that there is no loss in efficiency from using the unknown sub-population
AR, LM and LR statistics, $LR_{T}(\widehat{\mathbf{S}}_{T})$, $LM_{T}(\widehat{\mathbf{S}}_{T})$
and $AR_{T}(\widehat{\mathbf{S}}_{T})$, which are only functions
of the process $\{\overline{Z}(C_{T})'y\}_{\mathbf{S}_{T}\in\mathcal{S}}$.

We establish the asymptotic validity of the HAR CLR, LM and AR tests
in the unknown sub-population setting under a weak set of high-level
sufficient conditions on the IVs, exogenous variables and errors.
Define $w\left(\mathbf{S}_{T}\right)=\left[C_{T}Z:X\right]$.
\begin{assumption}
\label{Assumption 1 AMS} $T^{-1}w\left(\mathbf{S}_{T}\right)'w\left(\mathbf{S}_{T}'\right)\overset{\mathbb{P}}{\rightarrow}Q\left(\mathbf{S},\mathbf{S}'\right)$
uniformly in $\mathbf{S}_{T},\mathbf{S}_{T}'\in\mathcal{S}$ for $\mathbf{S}=\lim_{T\rightarrow\infty}T^{-1}\mathbf{S}_{T}$,
$\mathbf{S'}=\lim_{T\rightarrow\infty}T^{-1}\mathbf{S}_{T}'$ and
some p.d. $\left(q+p\right)\times\left(q+p\right)$ matrix $Q\left(\mathbf{S},\mathbf{S}'\right)$.
\end{assumption}
\begin{assumption}
\label{Assumption 2 AMS} $T^{-1}v'v\overset{\mathbb{P}}{\rightarrow}\Sigma_{v}$
for some $2\times2$ p.d. matrix $\Sigma_{v}$.
\end{assumption}
\begin{assumption}
\label{Assumption 3 AMS} For $\mathbf{S}_{T},\mathbf{S}_{T}'\in\mathcal{S}$
and $\mathbf{S}=\lim_{T\rightarrow\infty}T^{-1}\mathbf{S}_{T}$, $\mathbf{S}'=\lim_{T\rightarrow\infty}T^{-1}\mathbf{S}_{T}'$,
$T^{-1/2}\mathrm{vec}(w\left(\mathbf{S}_{T}\right)'v)\Rightarrow\mathscr{G}\left(\mathbf{S}\right)$,
where $\mathscr{G}(\cdot)$ is a mean-zero Gaussian process indexed
by $\mathbf{S}\subseteq(0,1]$ with $2\left(q+p\right)\times2\left(q+p\right)$
covariance function $\Psi\left(\mathbf{S},\,\mathbf{S}'\right)=\lim_{T\rightarrow\infty}T^{-1}\mathrm{Cov}(\mathrm{vec}(w\left(\mathbf{S}_{T}\right)'v),\mathrm{vec}(w\left(\mathbf{S}'_{T}\right)'v))$.
\end{assumption}
In Assumption \ref{Assumption 3 AMS}, $\mathrm{vec}\left(\cdot\right)$
denotes the vec operator. The quantities $Q\left(\cdot\right)$, $\Sigma_{v}$,
and $\Psi\left(\cdot\right)$ are assumed to be unknown. Assumptions
\ref{Assumption 1 AMS}-\ref{Assumption 2 AMS} hold under suitable
conditions by a (uniform) law of large numbers. Assumption \ref{Assumption 3 AMS}
holds under suitable conditions by a functional central limit theorem.
Assumptions \ref{Assumption 1 AMS}-\ref{Assumption 3 AMS} are consistent
with non-normal, heteroskedastic, autocorrelated errors and IVs and
regressors that may be random or non-random.\footnote{In the supplement we provide primitive sufficient conditions for Assumptions
\ref{Assumption 1 AMS}-\ref{Assumption 3 AMS}.}

We assume that we can consistently estimate $\Sigma_{v\overline{Z}}\left(\mathbf{S}\right)\equiv\Sigma_{v\overline{Z}}\left(\mathbf{S},\mathbf{S}\right)$
uniformly in $\mathbf{S}_{T}$.
\begin{assumption}
\label{Assumption Uniform Consistent Covariance Matrix}We have an
estimator $\widehat{\Sigma}_{v\overline{Z}}(\mathbf{S}_{T})$ such
that $\widehat{\Sigma}_{v\overline{Z}}(\mathbf{S}_{T})\overset{\mathbb{P}}{\rightarrow}\Sigma_{v\overline{Z}}(\mathbf{S})$
uniformly in $\mathbf{S}_{T}\in\mathcal{S}$ for $\mathbf{S}=\lim_{T\rightarrow\infty}T^{-1}\mathbf{S}_{T}$.
\end{assumption}
Note that this assumption immediately implies the uniform consistency
of $\widehat{\Sigma}_{N_{2}}(\mathbf{S}_{T})=\widehat{\Sigma}_{N_{2}}^{*}(\mathbf{S}_{T})-\widehat{\Sigma}_{N_{1}N_{2}}(\mathbf{S}_{T})\widehat{\Sigma}_{N_{1}}^{-1}(\mathbf{S}_{T})\widehat{\Sigma}_{N_{1}N_{2}}(\mathbf{S}_{T})'$
as well. Consistent estimators of $\Sigma_{v\overline{Z}}$ are HAC
and DK-HAC estimators.\footnote{In the supplement we provide weak sufficient conditions, even allowing
for certain forms of nonstationarity, that ensure this assumption
holds.}

Finally, we impose a second-order stationarity condition for $v'_{t}b_{0}\overline{Z}_{t}\left(C_{T}\right)$
and $v'_{t}\Sigma_{v}^{-1}a_{0}\overline{Z}_{t}\left(C_{T}\right)$.

\begin{assumption}
\label{Assumption 2nd Moment Stability}Let $\pi(\mathbf{S})$ equal
the Lebesgue measure of $\mathbf{S}\subseteq(0,1]$. Assume that $\Sigma_{v\overline{Z}}\left(\mathbf{S},\,\mathbf{S}'\right)=\pi\left(\mathbf{S}\cap\mathbf{S}'\right)\Sigma_{v\overline{Z}}$
where $\mathbf{S},\,\mathbf{S}'\subseteq(0,1]$ and $\Sigma_{v\overline{Z}}$
is p.d.
\end{assumption}
Assumption \ref{Assumption 2nd Moment Stability} is implied by a
uniform law of large numbers and functional central limit theorem
for partial sum processes under second-order stationarity. Under weak
IV asymptotics, $T^{-1}\widehat{\mathbf{S}}_{T}$ is not consistent
for $\mathbf{S}_{0}$. Assumption \ref{Assumption 2nd Moment Stability}
is needed in order to show that $N_{1,T}(\cdot)$ and $N_{2,T}(\cdot)$
are asymptotically independent processes. Under strong IV asymptotics
we can dispense with Assumption \ref{Assumption 2nd Moment Stability}
because $T^{-1}\widehat{\mathbf{S}}_{T}\overset{\mathbb{P}}{\rightarrow}\mathbf{S}_{0}$
and the limit of the processes $N_{1,T}(\cdot)$ and $N_{2,T}(\cdot)$
have zero covariance when evaluated at a fixed $\mathbf{S}_{0}$.

Define the LR, LM and AR statistics in this context according to \eqref{Eq. (3.4) AMS},
replacing $\mathbf{S}_{0,T}$ with $\widehat{\mathbf{S}}_{T}$. We
now establish the correct asymptotic null rejection probabilities
of the sub-population-estimated plug-in HAR CLR, LM and AR tests under
weak identification.
\begin{thm}
\label{Theorem 7 AMS HAR}Let Assumptions \ref{Assumption 1 AMS}-\ref{Assumption 2nd Moment Stability}
hold and suppose $\theta=c/T^{1/2}$ for some nonstochastic $c\in\mathbb{R}^{q}$.
We have: (i) ${AR}_{T}(\widehat{\mathbf{S}}_{T})\overset{d}{\rightarrow}\chi_{q}^{2}$
under $H_{0};$ (ii) ${LM}_{T}(\widehat{\mathbf{S}}_{T})\overset{d}{\rightarrow}\chi_{1}^{2}$
under $H_{0};$ (iii) $\mathbb{P}_{\beta_{0}}(LR_{T}(\widehat{\mathbf{S}}_{T})>\kappa_{\alpha}(N_{2,T}(\widehat{\mathbf{S}}_{T}))\rightarrow\alpha$
where $\mathbb{P}_{\beta_{0}}(\cdot)$ is the probability computed
under $H_{0}$.
\end{thm}
The key to establishing these asymptotic validity results is to show
that each of the above statements hold conditional on the realization
of $N_{2,T}(\cdot)$. This can be readily established from the facts
that the stochastic processes $N_{1,T}(\cdot)$ and $N_{2,T}(\cdot)$
are asymptotically independent by construction, $\widehat{\mathbf{S}}_{T}$
is a function of $N_{2,T}(\cdot)$ and $N_{1,T}(\mathbf{S}_{T})\Rightarrow\mathscr{N}\left(0,\,I_{q}\right)$
under $H_{0}$.\footnote{In addition to identification-robust tests of $H_{0}$ vs $H_{1}$,
since the causal interpretation of $\beta$ depends upon the sub-population
$\mathbf{S}_{0,T}$, practitioners may wish to simultaneously report
the result of these tests along with a corresponding estimate of the
sub-population. More specifically, failure to reject $H_{0}$ should
be interpreted as failure to reject that the estimand is equal to
$\beta_{0}$, where the estimand is interpreted as a weighted average
of the LATEs for the estimated sub-population $\widehat{\mathbf{S}}_{T}$.
Given that the tests of $H_{0}$ remain asymptotically valid conditional
on the realization of $N_{2,T}(\cdot)$ and the fact that $\widehat{\mathbf{S}}_{T}$
is a function of $N_{2,T}(\cdot)$, the tests remain asymptotically
valid when interpreted conditional on the value of $\widehat{\mathbf{S}}_{T}$.}

\section{\label{Section: Empirical Evidence}Empirical Evidence on LATE of
Monetary Policy}

We illustrate our methods by revisiting the identification of monetary
policy effects in the framework of \citet{nakamura/steinsson:2018},
introduced in Section \ref{Section: Application: Money Neutrality}.
They use a bivariate model \eqref{Eq. (1) NS and RS, Outcome variable}
to estimate the causal effect of $\widetilde{D}_{t}$ on $\widetilde{Y}_{t}$,
employing both event-study and heteroskedasticity-based identification
approaches. The dependent variable is the daily change in instantaneous
U.S. Treasury forward rates. For the policy news $\widetilde{D}_{t}$
they use three variables: the daily change in nominal 2-Year Treasury
yields, and the 30-minute or 1-day change in a ``policy news'' series\textemdash constructed
as the first principal component of the unanticipated 30-minute changes
in five selected interest rates. Heteroskedasticity-based identification
assumes the variance of the monetary shock rises on FOMC announcement
days, while the variance of other shocks remains constant {[}cf. eq.
\eqref{Eq. Vol_P >Vol_C}{]}. FOMC dates define the policy sample
$\mathbf{P}$, and analogous non-FOMC dates define the control sample
$\mathbf{C}.$ We consider specifications where $\widetilde{D}_{t}$
is either the 30-minute policy news series or 1-day change in Treasury
yields, and $\widetilde{Y}_{t}$ is either the nominal or real 2-Year
instantaneous Treasury forward rate. Nakamura and Steinsson's instrument
for $\widetilde{D}_{t}^{2}$ is defined as $Z_{t}=\mathbf{1}\left\{ t\in\mathbf{P}\right\} $,
corresponding to the model in Section \ref{Section: Application: Money Neutrality}.
We focus on the same period: January 1, 2004, to March 19, 2014.

\citet{lewis:2020} recently analyzes this problem by developing a
first-stage $F$-test for weak identification. He finds that weak
identification is not rejected when $\widetilde{D}_{t}$ is the 1-day
change in nominal 2-Year Treasury yields, but is strongly rejected
when $\widetilde{D}_{t}$ is the 30-minute policy news series. This
supports \citeauthor{nakamura/steinsson:2018}'s \citeyearpar{nakamura/steinsson:2018}
observation that the daily policy variable may suffer from weaker
identification. Unlike \citet{nakamura/steinsson:2018}, \citet{lewis:2020}
estimates the model using GMM and does not impose the assumption that
the non-monetary policy shock $\eta_{t}$ has equal variance across
the treatment and control samples.

Section \ref{Subsection: Pre Test for Full Population Identification Failure NS18}
reports results of our test for full sample identification failure.
Section \ref{Subsection: Estimation in Strongly Identified Sub-sample NS18}
presents causal effect estimates based on the most strongly-identified
subsample. Section \ref{Subsection: Identifcation Robust Inference NS18}
provides identification-robust inference results, and Section \ref{Subsection: Identification and Estimation of Compliers}
estimates compliers at the individual level and tests the exclusion
restriction.

\subsection{\label{Subsection: Pre Test for Full Population Identification Failure NS18}Testing
for Identification Failure }

We present the results of our test for identification failure over
all sub-populations from Section \ref{Section: ID Failure Test} in
Table \ref{Table: F tests} considering values of $\pi_{L}$ from
0.6 to 1. For the 30-minute policy news variable, the $F_{T}^{*}$
statistic is very large and identification failure is rejected at
any common significance level. This supports the finding in \citet{lewis:2020}
and intuition in \citet{nakamura/steinsson:2018} that the 30-minute
policy news variable leads to stronger identification in the full
sample. In contrast, for the 1-day change in nominal Treasury yields,
identification failure cannot be strongly rejected in the full sample:
the $F_{T}^{*}$ statistic at $\pi_{L}=1$ (i.e., full sample) is
only slightly larger than the 1\% critical value. The $F_{T}^{*}$
statistic increases substantially as $\pi_{L}$ decreases and it is
very far from the critical values. This is clear evidence that identification
is much stronger over subsamples. At $\pi_{L}=0.9$ it reaches 33.87,
clearly rejecting identification failure in the $\pi$-subsample (with
$\pi=0.9$ or $0.95$) over which the supremum of $F_{T}\left(\mathbf{S}_{T}\right)$
is computed. The $F_{T}^{*}$ statistic increases monotonically with
smaller $\pi_{L}$ due to the increasing number of partitions considered.
For example, at $\pi_{L}=0.8$, $F_{T}^{*}$ is 54.78\textemdash nearly
seven times the full sample value. Overall, the results indicate that
strong identification may hold when using a 1-day window, but only
within subsamples comprising at most 90\% of the data. The weak identification
reported by \citet{lewis:2020} using a 1-day window around FOMC announcements
likely does not stem solely from volatility returning to normal after
announcements. Rather, a small subsample (10\textendash 20\% of the
data) exhibits weak or failed identification, contributing to the
weaker identification exhibited in the full sample.

\begin{table}[H]
\caption{\label{Table: F tests}Tests for Identification Failure over all Sub-Populations}

\smallskip{}

\begin{centering}
{\footnotesize{}
\begin{tabular}{cccccc}
\hline
\multicolumn{6}{c}{{\small$F_{T}^{*}$ statistic and critical values}}\tabularnewline
 & \multicolumn{5}{c}{{\small$F_{T}^{*}$}}\tabularnewline
{\small$D_{t}\backslash\pi_{L}$} & {\small 0.6} & {\small 0.7} & {\small 0.8} & {\small 0.9} & {\small 1}\tabularnewline
\hline
\hline
\begin{cellvarwidth}[t]
\centering
{\small 30-minute}\\
{\small ``policy news''}
\end{cellvarwidth} & {\small$10^{4}\times95.36$} & {\small$10^{4}\times56.50$} & {\small$10^{4}\times32.45$} & {\small$10^{4}\times18.75$} & {\small$10^{4}\times7.42$}\tabularnewline
\begin{cellvarwidth}[t]
\centering
{\small 1-day nominal}\\
{\small{} Treasury yields}
\end{cellvarwidth} & {\small 155.69} & {\small 88.22} & {\small 54.78} & {\small 33.88} & {\small 8.09}\tabularnewline
{\small 1\% critical values } & {\small 11.63} & {\small 10.94} & {\small 9.73} & {\small 8.68} & {\small 6.68}\tabularnewline
{\small 5\% critical values } & {\small 8.28} & {\small 7.55} & {\small 6.84} & {\small 6.04} & {\small 3.85}\tabularnewline
\hline
\end{tabular}}{\footnotesize\par}
\par\end{centering}
{\scriptsize{}
\noindent\begin{minipage}[t]{1\columnwidth}
{\scriptsize$F_{T}^{*}$ statistics for first-stage identification
failure. $D_{t}$ is either the 30-minute policy news series or 1-day
change in nominal Treasury yields. $\pi_{L}$ is the minimum fraction
of the sample over which the supremum of the $F\left(\mathbf{S}_{T}\right)$
is computed. Maximum number of breaks is set to $m_{+}=5$. }
\end{minipage}}{\scriptsize\par}
\end{table}


\subsection{\label{Subsection: Estimation in Strongly Identified Sub-sample NS18}Estimation
in Strongly-Identified Subsample}

We turn to estimation of $\pi_{0}$ and $\mathbf{S}_{0,T}$ using
the methods from Section \ref{Section: Estimation}, and then to estimating
the LATE of monetary policy based on the strongly-identified subsample,
$\widehat{\beta}(\mathbf{\widehat{S}}_{T,OLS})$, or simply, $\widehat{\pi}$-sample,
where $\widehat{\pi}=|\widehat{\mathbf{S}}_{T,OLS}|/T$. We focus
on $\widehat{\beta}(\widehat{\mathbf{S}}_{T,OLS})$; results using
$\widehat{\beta}(\mathbf{\widehat{S}}_{T,FGLS})$ are similar. Figure
\ref{Figure: pi Sample} plots the 1-day changes in 2-Year yields
for the control and policy samples and highlights the regimes included
in the strongly-identified subsample $\mathbf{\widehat{S}}_{T,OLS}$.
The estimate $\widehat{\pi}=0.8$ implies that in 80\% of the sample,
the first-stage is strong and identification holds. In the control
sample, the excluded periods include the first seven months of 2005
and the regime surrounding the financial crisis (2007-2009). As shown
in the figure, volatility during the crisis period is much higher
than in the rest of the control group and higher than the average
volatility in the treatment group. This subsample appears to drive
the apparent full sample weak identification. Since our method searches
for maximum identification strength, it correctly excludes this period
when computing $\pi$-LATE.\footnote{The other excluded period (January to July 2005) does not display
obviously high volatility but shows some persistence, with a short-duration
cluster below the mean toward the end.} The interpretation is that in both excluded regimes\textemdash especially
during the financial crisis\textemdash market uncertainty was elevated
even on non-FOMC days, violating the identification assumption.
\noindent\begin{flushleft}
\begin{center}
\begin{figure}[h]
\begin{raggedright}
\includegraphics[clip,width=17cm,totalheight=8cm]{Figure_pi_sample}
\par\end{raggedright}
\raggedright{}\caption{\label{Figure: pi Sample}{\scriptsize Plot of $D_{t}$ (2-Years Treasury
yields) in the control sample (top panel) and policy sample (bottom
panel). The red rectangles indicate subsamples included in the strongly-identified
subsample $\mathbf{\widehat{S}}_{T,OLS}$ where $\widehat{\pi}=0.8$.}}
\end{figure}
\end{center}
\par\end{flushleft}

We now estimate the causal effect of monetary policy using the $\widehat{\pi}$-sample,
where by construction the LATE is most strongly-identified. We compare
these results with full sample estimates obtained using two-stage
least squares (TSLS) and GMM, following \citet{nakamura/steinsson:2018}
and \citet{lewis:2020}, respectively. Table \ref{Table: Estimation NS18}
presents the results. Starting with the full sample estimates: when
the policy variable is the 30-minute policy news series, TSLS and
GMM yield very similar point estimates for both nominal and real forward
rates, and both are statistically significant using standard and robust
confidence intervals.\footnote{The robust confidence intervals for the GMM estimates are based on
the subset $K$-test in \citet{lewis:2020}.}

As noted by \citet{lewis:2020}, the assumption that non-monetary
shocks have equal variance across treatment and control groups does
not bias the TSLS estimates, as they closely match the GMM ones. One
explanation is that the GMM estimate of $a$ (capturing reverse causality
from forward rates to policy news) is both near zero and statistically
significant (not reported). Since potential bias from this assumption
is proportional to $a(\sigma_{\eta,P}^{2}-\sigma_{\eta,C}^{2}),$
and $a$ is close to zero, the resulting bias is negligible even if
the variances $\sigma_{\eta,P}^{2}$ and $\sigma_{\eta,C}^{2}$ differ.

Turning to the case where the policy variable is the 1-day change
in 2-Year Treasury yields, the TSLS and GMM estimates differ markedly
from each other and from those based on the 30-minute policy news
series. Notably, the GMM estimate of $\beta$ is negative for nominal
forwards and positive for real forwards, but in neither case is it
statistically significant\textemdash whether using standard or robust
confidence intervals.

As discussed by \citet{lewis:2020}, these estimates are difficult
to interpret in economically meaningful terms. He also shows that
the GMM estimates of $a$ are nonzero and proposed a second dimension
of policy news to account for the findings. However, the opposing
signs of $\beta$ across nominal and real forwards complicate this
interpretation. Ultimately, he concludes that these results are inconsistent
with \citeauthor{nakamura/steinsson:2018}'s \citeyearpar{nakamura/steinsson:2018}
``background noise'' view of the non-monetary shock $\eta_{t}$
which assumes that its volatility remains unchanged between FOMC and
non-FOMC days.

We contribute to this discussion by presenting TSLS and GMM estimates
based on the most strongly-identified $\widehat{\pi}$-sample. We
focus first on standard confidence intervals and defer weak identification-robust
inference to Table \ref{Table: Inference NS18}. The bottom panel
of Table \ref{Table: Estimation NS18} shows that, for the 30-minute
policy news variable, the TSLS and GMM estimates, including their
statistical significance, are virtually unchanged. As expected\textemdash given
the apparent strong identification in the full sample\textemdash results
are broadly similar when using the $\widehat{\pi}$-sample.\footnote{The confidence intervals in the $\widehat{\pi}$-sample are even slightly
tighter.}

\begin{table}[H]
\caption{\label{Table: Estimation NS18}Estimation of $\beta$}

\smallskip{}

\begin{centering}
{\footnotesize}{\small{}
\begin{tabular}{ccccc}
\hline
 & \multicolumn{2}{c}{{\small 30-minute Policy News}} & \multicolumn{2}{c}{{\small 1-day 2-Year Yield}}\tabularnewline
{\small dep. var.} & {\small Nominal} & {\small Real} & {\small Nominal} & {\small Real}\tabularnewline
\hline
 & \multicolumn{4}{c}{{\small Full Sample}}\tabularnewline
 & \multicolumn{4}{c}{{\small TSLS}}\tabularnewline
{\small$\beta$} & {\small 1.10{*}{*}} & {\small 0.96{*}{*}{*}} & {\small 1.14{*}{*}{*}} & {\small 0.97{*}{*}{*}}\tabularnewline
{\small standard CI} & {\small{[}0.17, 2.02{]}} & {\small{[}0.41, 1.51{]}} & {\small{[}0.83, 1.45{]}} & {\small{[}0.40, 1.565{]}}\tabularnewline
 & \multicolumn{4}{c}{{\small GMM}}\tabularnewline
{\small$\beta$} & {\small 1.07{*}{*}} & {\small 0.94{*}{*}{*}} & {\small -0.27} & {\small 1.31}\tabularnewline
{\small standard CI} & {\small{[}0.17, 1.98{]}} & {\small{[}0.36, 1.51{]}} & {\small{[}-4.90, 4.36{]}} & {\small{[}-3.74, 6.35{]}}\tabularnewline
{\small robust CI} & {\small{[}0.27, 3.25{]}} & {\small{[}0.44, 2.38{]}} & {\small{[}-77.27, 0.94{]}} & {\small{[}-253.70, 1.92{]}}\tabularnewline
 & \multicolumn{4}{c}{{\small$\pi$-sample based on $\widehat{\mathbf{S}}_{T,OLS}$ with
$\widehat{\pi}=0.8$}}\tabularnewline
 & \multicolumn{4}{c}{{\small TSLS}}\tabularnewline
{\small$\beta$} & {\small 1.11{*}{*}} & {\small 0.97{*}{*}{*}} & {\small 1.13{*}{*}{*}} & {\small 0.92{*}{*}{*}}\tabularnewline
{\small standard CI} & {\small{[}0.19, 2.02{]}} & {\small{[}0.42, 1.51{]}} & {\small{[}0.92, 1.30{]}} & {\small{[}0.56, 1.28{]}}\tabularnewline
 & \multicolumn{4}{c}{{\small GMM}}\tabularnewline
{\small$\beta$} & {\small 1.07{*}{*}} & {\small 0.94{*}{*}{*}} & {\small 0.65{*}} & {\small 0.86{*}{*}}\tabularnewline
{\small standard CI} & {\small{[}0.17, 1.96{]}} & {\small{[}0.38, 1.50{]}} & {\small{[}-0.02, 131{]}} & {\small{[}0.29, 1.43{]}}\tabularnewline
\hline
\end{tabular}}{\small\par}
\par\end{centering}
\centering{}{\scriptsize{}
\begin{minipage}[t]{0.8\columnwidth}
{\scriptsize TSLS estimates of $\beta$ and GMM estimates of $\beta/(1-a\beta)$.
The GMM estimates allow for changes also in the variance of $\eta_{t}$
across regimes. The dependent variable is the 1-day change in either
nominal or real 2-Year instantaneous Treasury forward rate. The policy
variable is either the 30-minute changes in the ``policy news''
variable or 1-day changes in the 2-Year nominal Treasury yield. The
standard 95\% confidence interval is based on the standard normal
critical values. For the GMM estimates, the robust 95\% confidence
interval is based on the subset $K$-test in \citet{lewis:2020}.
Asterisks indicate statistical significance at the 10\%, 5\%, or 1\%
level based on standard intervals.}
\end{minipage}}{\scriptsize\par}
\end{table}


Finally, we turn to the $\widehat{\pi}$-sample estimates using the
1-day window for the policy. The GMM estimates differ sharply from
those in the full sample: for both nominal and real forwards, they
now have the same sign and are statistically significant. This suggests
that the opposite signs reported by \citet{lewis:2020} likely stemmed
from weak identification, rendering those estimates unreliable.\footnote{While the TSLS estimates are nearly unchanged from the full sample,
this should not be taken as evidence of their reliability. Under weak
IVs, their similarity to the $\widehat{\pi}$-sample results may simply
be coincidental.} Notably, the GMM estimates are now similar in magnitude to those
based on the 30-minute policy variable, supporting a more meaningful
interpretation.\footnote{We also verified that the GMM estimate of $a$ is 0.70 for nominal
forwards and -0.91 for real forwards. It is intuitive that the estimate
of $a$ is close to zero when using a 30-minute window but significantly
different from zero with a 1-day window. In the narrow 30-minute window
around an FOMC announcement, reverse causality from $\widetilde{Y}_{t}$
to $\widetilde{D}_{t}$ is limited, as monetary news is more pronounced
than other shocks\textemdash though some endogeneity may still arise
from omitted factors affecting both. In contrast, over a full day,
asset price movements can influence short-term interest rates, making
reverse causality more likely.}

Overall, this analysis highlights the advantage of using the most
strongly-identified $\widehat{\pi}$-sample. Given weak identification
in the full sample when using 1-day Treasury yields as the policy
variable, the corresponding estimates should be discarded. In contrast,
evidence from the $\widehat{\pi}$-sample shows that TSLS and GMM
produce similar, positive estimates for $\beta,$ consistent with
monetary policy affecting real forward rates, as predicted by New
Keynesian models, and supporting the existence of a forward guidance
channel.\footnote{However, the results do not yet support a second meaningful dimension
of news, as proposed by{\small{} }\citet{lewis:2020}, since the
sign of the GMM estimate of $a$ is unstable across nominal and real
forwards. Regarding \citeauthor{nakamura/steinsson:2018}'s \citeyearpar{nakamura/steinsson:2018}
``background noise'' interpretation of non-monetary shocks, we find
no clear evidence against it: in the $\widehat{\pi}$-sample, identification
appears strong, and TSLS and GMM estimates consistently share the
same sign and similar magnitudes.}

\subsection{\label{Subsection: Identifcation Robust Inference NS18}Weak Identification-Robust
Inference}

We apply the weak identification-robust tests proposed in Section
\ref{Section: Inference} and compare them to existing full sample
tests $LM_{T}$ and $LR_{T}$.\footnote{We do not report the $AR_{T}$ test since for $q=1$ it is equivalent
to the $LM_{T}$ test.} We test the the null hypothesis $H_{0}:\,\beta=0$ against $H_{1}:\,\beta\neq0$,
and extend the analysis to include 5-Year forward rates, in addition
to the 2-Year forwards. Results are shown in Table \ref{Table: Inference NS18}.
When the policy variable is the 30-minute policy news, identification
is strong in the full sample. Accordingly, both the proposed and existing
tests yield similar results: all tests reject at the 5\% level for
both nominal and real forwards. \citet{nakamura/steinsson:2018} showed
that the effect of policy news peaks at the 2-Year maturity and declines
with longer maturities. Consistent with this, we find weaker statistical
significance for the 5-Year. In line with theoretical predictions,
the long-run impact of monetary policy shocks on real interest rates
(i.e., the 10 Year forwards) approaches zero (not reported). Our proposed
tests confirm this, showing some rejection for the 5-Year real forwards
but not for the 10-Year.

\begin{table}
\caption{\label{Table: Inference NS18}Identification-Robust Inference on $\beta$}

\smallskip{}

\begin{centering}
{\small{}}{\small{}
\begin{tabular}{ccccccccccccc}
\hline
 & \multicolumn{6}{c}{{\small 30-minute Policy News}} & \multicolumn{6}{c}{{\small 1-day change in 2-Year Yields}}\tabularnewline
{\small 2-Year Forwards} & \multicolumn{3}{c}{{\small Nominal}} & \multicolumn{3}{c}{{\small Real}} & \multicolumn{3}{c}{{\small Nominal}} & \multicolumn{3}{c}{{\small Real}}\tabularnewline
{\small$\alpha$} & {\small 0.10} & {\small 0.05} & {\small 0.01} & {\small 0.10} & {\small 0.05} & {\small 0.01} & {\small 0.10} & {\small 0.05} & {\small 0.01} & {\small 0.10} & {\small 0.05} & {\small 0.01}\tabularnewline
\hline
{\small$LM_{T}$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\times$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\times$}\tabularnewline
{\small$CLR_{T}$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\times$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\times$}\tabularnewline
{\small$LM_{T}(\widehat{\mathbf{S}}_{T})$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\times$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\times$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\checkmark$}\tabularnewline
{\small$CLR_{T}(\widehat{\mathbf{S}}_{T})$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\times$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\times$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\checkmark$}\tabularnewline
\hline
{\small 5-Year Forwards} & \multicolumn{3}{c}{{\small Nominal}} & \multicolumn{3}{c}{{\small Real}} & \multicolumn{3}{c}{{\small Nominal}} & \multicolumn{3}{c}{{\small Real}}\tabularnewline
{\small$\alpha$} & {\small 0.10} & {\small 0.05} & {\small 0.01} & {\small 0.10} & {\small 0.05} & {\small 0.01} & {\small 0.10} & {\small 0.05} & {\small 0.01} & {\small 0.10} & {\small 0.05} & {\small 0.01}\tabularnewline
\hline
{\small$LM_{T}$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$}\tabularnewline
{\small$CLR_{T}$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$}\tabularnewline
{\small$LM_{T}(\widehat{\mathbf{S}}_{T})$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\checkmark$} & {\small$\times$} & {\small$\times$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\times$}\tabularnewline
{\small$CLR_{T}(\widehat{\mathbf{S}}_{T})$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\times$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\times$} & {\small$\checkmark$} & {\small$\checkmark$} & {\small$\times$}\tabularnewline
\hline
\end{tabular}}{\small\par}
\par\end{centering}
{\scriptsize{}}{\scriptsize{}
\noindent\begin{minipage}[t]{1\columnwidth}
{\scriptsize Weak identification-robust tests on $\beta$. The dependent
variable $Y_{t}$ is either the 2-Year forward rates (top panel) or
the 5-Year forward rates (bottom panel). $D_{t}$ is either the 30-minute
policy news series or the 1-day nominal Treasury yields. Significance
levels are $\alpha=0.10,\,0.05,\,0.01$. A $\checkmark$ means indicates
rejection $H_{0};$ a $\times$ non-rejection.}
\end{minipage}}{\scriptsize\par}
\end{table}

Let us instead consider the 1-day change in 2-Year yields as the policy
variable. The existing $LM_{T}$ and $LR_{T}$ tests do not reject
the null at any standard significance level for nominal forwards,
and at the 1\% level for real forwards. In contrast, our proposed
tests based on $\widehat{\mathbf{S}}_{T}$ show much stronger rejections,
aligning with the results using the 30-minute policy news series,
which indicate a positive causal effect on 2-Year forwards.

\subsection{Identification and Estimation of Compliers, and Exclusions Restriction }

\label{Subsection: Identification and Estimation of Compliers}We
now identify compliers individually by applying Theorem \ref{Theorem: Compliers}.
Under heteroskedasticity-based identification, the sample rolling
window averages in Theorem \ref{Theorem: Compliers} correspond to
rolling window variances, i.e., $\overline{D}_{P,t,n_{1}}$ and $\overline{D}_{C,t,n_{0}}$
are equal to $\overline{\sigma}_{P,t,n_{1}}^{2}$ and $\overline{\sigma}_{C,t,n_{0}}^{2}$
in this context, where $\overline{\sigma}_{P,t,n_{1}}^{2}$ and $\overline{\sigma}_{C,t,n_{0}}^{2}$
are defined analogously but using $\tilde{D}_{t}^{2}$ for $D_{t}$.\footnote{More specifically, we have $\overline{\sigma}_{C,t_{0}-1,n_{0}}^{2}=\frac{1}{n_{0}}\sum_{s\in N_{0}\left(t_{0}\right)}\tilde{D}_{s}^{2},$
and $\overline{\sigma}_{P,t_{0},n_{1}}^{2}=\frac{1}{n_{1}}\sum_{s\in N_{1}\left(t_{0}\right)}\tilde{D}_{s}^{2}$.\textcolor{blue}{{}}} We use two-sided rolling windows with $n_{0}=101$ and $n_{1}=15$.
For each $t_{0}$ we test the null hypothesis $H_{0}:\,\mathbb{E}(D_{t_{0}}^{2}\left(1\right))-\mathbb{E}(D_{t_{0}}^{2}\left(0\right))=0$
($t_{0}$ is a non-complier) versus the one tailed alternative $H_{1}:\,\mathbb{E}(D_{t_{0}}^{2}\left(1\right))-\mathbb{E}(D_{t_{0}}^{2}\left(0\right))>0$
($t_{0}$ is a complier). We use the $t$-statistic
\begin{align*}
t_{t_{0}} & =\begin{cases}
\frac{\sqrt{n_{0}}\left(\overline{\sigma}_{P,t_{0},n_{1}}^{2}-\overline{\sigma}_{C,t_{0}-1,n_{0}}^{2}\right)}{\sqrt{J_{\mathrm{HAC},t_{0}-1}}} & t_{0}\in\mathbf{P}\\
\frac{\sqrt{n_{0}}\left(\overline{\sigma}_{P,s^{*}\left(t_{0}\right),n_{1}}^{2}-\overline{\sigma}_{C,t_{0},n_{0}}^{2}\right)}{\sqrt{J_{\mathrm{HAC},t_{0}}}} & t_{0}\in\mathbf{C},
\end{cases}
\end{align*}
where $J_{\mathrm{HAC},t_{0}}$ is the Newey-West estimator with $\left\lfloor n_{0}^{1/3}\right\rfloor $
lags applied to $\tilde{D}_{s}^{2}-n_{0}^{-1}\sum_{k\in N_{0}\left(t_{0}+1\right)}\tilde{D}_{k}^{2}$.

\begin{singlespace}
\noindent\begin{flushleft}
\begin{center}
\begin{figure}[h]
\begin{raggedright}
\includegraphics[clip,width=17cm,totalheight=8cm]{Figure_Compliers}
\par\end{raggedright}
\raggedright{}\caption{\label{Figure Compliers}{\scriptsize Plot of $\widetilde{D}_{t}$
(2-Years Treasury yields) in the control sample (top panel) and policy
sample (bottom panel). The orange rectangles indicate subsamples included
in the strongly-identified set $\mathbf{\widehat{S}}_{T,OLS}$ where
$\widehat{\pi}=0.8$. Green filled circles indicate compliers; red
filled circles indicate non-compliers. Time points without colored
markers correspond to cases where rolling sample variances could not
be computed due to proximity to the start or end of the sample.}}
\end{figure}
\end{center}
\par\end{flushleft}
\end{singlespace}

The results are shown in Figure \ref{Figure Compliers}. Approximately
75\% of observations are classified as compliers. The non-compliers
are mostly concentrated in the period from 2010 to mid-2011, which
corresponds to the early phase of the zero lower bound (ZLB) period
following the 2008\textendash 09 recession. During this time, the
Fed relied primarily on qualitative forward guidance\textemdash e.g.,
stating that economic conditions were "likely to warrant exceptionally
low levels of the federal funds rate for some time." In August 2011,
the Fed shifted to more explicit, calendar-based guidance, stating
that such conditions were "likely to warrant exceptionally low levels
of the federal funds rate at least through mid-2013." Thus, the non-complier
period aligns with the phase of the ZLB when forward guidance was
less aggressive as the policy announcements by then only imply a near
zero-rate horizon for the following three to four quarters, significantly
shorter than what the ZLB constraint would actually have implied.\footnote{The set of compliers does not coincide with the set of observations
in the strongly-identified $\widehat{\pi}$-sample. However, this
does not necessarily imply a violation of monotonicity {[}cf. Proposition
\ref{Proposition: Compliers FS indivdually}{]}. First, the complier
status is determined via a $t$-test whereas the strongly-identified
$\widehat{\pi}$-sample is determined via estimation. Second, the
complier status is determined by the rolling window variances at each
$t$, whereas inclusion in the $\widehat{\pi}$-sample depends on
how these variances contribute to the average volatility in the control
sample relative to that in the policy sample.}



Finally, we use Theorem \ref{Theorem: Compliers} and Proposition
\ref{Proposition: exclusion-restriction} to test the exclusion restriction
(cf. Assumption \ref{Assumption: Exclusion }). We consider the whole
set of compliers $\mathcal{NC}.$ We test the null hypothesis that
the exclusion restriction holds by using the following $t$-statistic:
\begin{align*}
t_{\mathrm{exclusion}}=\frac{\sqrt{|\mathcal{NC}_{\mathbf{C}}^ {}|}\left(\overline{Y}_{\mathcal{NC}_{\mathbf{P}}^{s}}-\overline{Y}_{\mathcal{NC}_{\mathbf{C}}^{s}}\right)}{\sqrt{J_{\mathrm{HAC},\mathcal{NC}^{s}}}}
\end{align*}
where $\overline{Y}_{\mathcal{NC}_{\mathbf{P}}^ {}}=\frac{1}{|\mathcal{NC}_{\mathbf{P}}^ {}|}\sum_{t\in\mathcal{NC}_{\mathbf{P}}^ {}}Y_{t}$,
$\overline{Y}_{\mathcal{NC}_{\mathbf{C}}^ {}}=\frac{1}{|\mathcal{NC}_{\mathbf{C}}^ {}|}\sum_{t\in\mathcal{NC}_{\mathbf{C}}^ {}}Y_{t}$,
$Y_{t}=\tilde{D}_{t}\tilde{Y}_{t}$ and $J_{\mathrm{HAC},\mathcal{NC}^{s}}$
is the Newey-West estimator applied to $\overline{Y}_{\mathcal{NC}_{\mathbf{P}}^ {}}-\overline{Y}_{\mathcal{NC}_{\mathbf{C}}^ {}}$.
For the real (nominal) 2-Year forward rate, we find $t_{\mathrm{exclusion}}=0.91$
($t_{\mathrm{exclusion}}=1.02$), and thus fail to reject the exclusion
restriction.

\section{\label{Section Conclusions}Conclusions}

This paper discusses identification, estimation and inference on dynamic
LATE. We show that compliers can be identified individually and the
exclusion restriction can be tested using a $t$-test. While weak
identification is common in the full sample in practice, strong identification
often appears to hold in a sizable subsample. We propose a method
to isolate this strongly-identified subsample, enabling consistent
estimation and inference.

\paragraph*{Supplemental Materials:}

The online supplement {[}cf. \textcolor{MyBlue}{Casini et al.} \citeyearpar{casini/mccloskey/pala/rolla_Dynamic_Late_Supp}{]}
includes Monte Carlo simulations, proofs of the results of Sections
\ref{Section: Statistical Framework for Identification of Causal Effects}-\ref{Section: ID Failure Test}
and \ref{Section: Inference}. The non-online supplement {[}cf. \textcolor{MyBlue}{Casini et al.}
\citeyearpar{casini/mccloskey/pala/rolla_Dynamic_Late_Supp_Not_Online}{]}
contains the theoretical results and corresponding proofs for the
estimators in Section \ref{Section: Estimation} and additional results.

\begin{singlespace}

\bibliographystyle{econometrica}
\bibliography{References}
\addcontentsline{toc}{section}{References}

\end{singlespace}

\newpage{}

\newpage{}

\clearpage
\pagenumbering{arabic}