Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Identification at the Zero Lower Bound
frontmatter\runtitle{Identification at the Zero Lower Bound}
\begin{aug}
\address[id=add1]{
\orgdiv{Department of Economics},
\orgname{University of Oxford}}
\end{aug}
\support{This research is funded by the European Research Council via Consolidator grant
number 647152. I would like to thank Guido Ascari, Sergio de Ferra, James Duffy, Andrea Ferrero, Jim Hamilton, Daisuke Ikeda, Federica Romei, Frank Schorheide, Francesco Zanetti and seminar participants at the Chicago Fed, New York Fed, Board of Governors, Bank of Japan, Hitotsubashi University, Keio University, UCL, University of Cambridge, University of Oxford, University of Warwick, DNB, University of Pennsylvania, Pennsylvania State University, NBER Summer Institute, IAAE conference, Netherlands Econometrics Study Group Meeting, World Congress and North American Winter meetings of the Econometric society for useful comments and discussion, as well as Mishel Ghassibe, Lukas Freund, Shangshang Li, and Patrick Vu for research assistance.}
\begin{abstract}
I show that the Zero Lower Bound (ZLB) on interest rates can be used to
identify the causal effects of monetary policy. Identification depends on the
extent to which the ZLB limits the efficacy of monetary policy. I propose a simple way to test the efficacy of unconventional policies, modelled via a `shadow rate'. I apply this method to
U.S. monetary policy using a three-equation structural vector autoregressive model of inflation, unemployment and the federal funds rate. I reject the null hypothesis that
unconventional monetary policy has no effect at the ZLB, but find some
evidence that it is not as effective as conventional monetary policy.
\end{abstract}
\begin{keyword}
\kwd{SVAR}
\kwd{censoring}
\kwd{coherency}
\kwd{partial identification}
\kwd{monetary policy}
\kwd{shadow rate}
\end{keyword}
Introduction
The zero lower bound (ZLB) on nominal interest rates has arguably been a
challenge for policy makers and researchers of monetary policy. Policy makers
have had to resort to so-called unconventional policies, such as quantitative
easing or forward guidance, which had previously been largely untested.
Researchers have to use new theoretical and empirical methodologies to analyze
macroeconomic models when the ZLB\ binds. So, the ZLB is generally viewed as a
problem or at least a nuisance. This paper proposes to turn this problem on
its head to solve another long-standing question in macroeconomics:\ the
identification of the causal effects of monetary policy on the economy.
The intuition is as follows. If the ZLB limits the ability of policy makers to
react to macroeconomic shocks, as argued, for example, by
EggertssonWoodford2003, the response of the economy to shocks will change
when the policy instrument hits the ZLB. Because the difference in the
behavior of macroeconomic variables across the ZLB and non-ZLB regimes is only
due to the impact of monetary policy, the switch across regimes provides
information about the causal effects of policy. In the extreme case that
monetary policy completely shuts down during the ZLB regime, either because
policy makers do not use alternative (unconventional) policy instruments, or
because such instruments turn out to be completely ineffective, the only
difference in the behavior of the economy across regimes is due to the impact
of (conventional) policy during the unconstrained regime. Therefore, the ZLB
identifies the causal effect of policy during the unconstrained regime. If
monetary policy remains partially effective during the ZLB regime, e.g.,
through the use of unconventional policy instruments, then the difference in
the behavior of the economy across regimes will depend on the difference in
the effectiveness of conventional and unconventional policies. In this case,
we obtain only partial identification of the causal effects of monetary
policy, but we can still get informative bounds on the relative efficacy of
unconventional policy. In the other extreme case that unconventional policy is
as effective as conventional policy, there is no difference in the behavior of
the economy across regimes, and we have no additional information to identify
the causal effects of policy. However, we can still test this so-called ZLB
irrelevance hypothesis DebortoliGaliGambetti2019 by testing whether
the reaction of the economy to shocks is the same across the two regimes.
There are similarities between identification via occasionally binding
constraints and identification through heteroskedasticity Rigobon03,
or more generally, identification via structural change
MagnussonMavroeidis2014. That literature showed that the switch
between different regimes generates variation in the data that identifies
parameters that are constant across regimes. For example, an exogenous shift
in a policy reaction function or in the volatility of shocks identifies the
transmission mechanism, provided the latter is unaffected by the policy
shift. When the switch from one regime to another is exogenous,
regime indicators are valid instruments, and the methodology in
MagnussonMavroeidis2014 is
applicable. However,
regimes induced by occasionally binding constraints are not exogenous --
whether the ZLB binds or not clearly depends on the structural shocks, so
regime indicators cannot be used as instruments in the usual way, and a new
methodology is needed to analyze these models.
In this paper, I show how to control for the endogeneity in regime selection
and obtain identification in structural vector autoregressions
(SVARs).\footnote{There is a related literature on Dynamic Stochastic General
Equilibrium (DSGE) models with a ZLB, see, e.g.,
FernandezGordonGuerronRubio2015, GuerrieriIacoviello2015,
AruobaCubaBordaSchorfheide2017, KulishMorleyRobinson2017 and
AruobaCubaBordaHigaFloresSchorfheideVillalvazo2020. The papers in this
literature do not point out the implications of the ZLB for identification of
monetary policy shocks.} The methodology is parametric and likelihood-based,
and the analysis is similar to the well-known Tobit model Tobin1958.
More specifically, the methodological framework builds on the early
microeconometrics literature on simultaneous equations models with censored
dependent variables, see Amemiya1974, Lee1976,
BlundellSmith1994, and the more recent literature on dynamic Tobit
models, see Lee1999, and particle filtering, see
PittShephard1999.
A further contribution of this paper is a general methodology to estimate
reduced-form VARs with a variable subject to an occasionally binding
constraint. This is a necessary starting point for SVAR analysis that uses any
of the existing popular identification schemes, such as short- or long-run
restrictions, sign restrictions, or external instruments. In the absence of
any constraints, reduced-form VARs can be estimated consistently by Ordinary
Least Squares (OLS), which is Gaussian Maximum Likelihood, or its
corresponding Bayesian counterpart, and inference is fairly well-established.
However, it is well-known that OLS estimation is inconsistent when the data is
subject to censoring or truncation, see, e.g., Gree93 for a textbook
treatment. So, it is not possible to estimate a VAR consistently by OLS using
any sample that includes the ZLB, or even using (truncated) subsamples when
the ZLB is not binding (because of selection bias), as was pointed out by
HayashiKoeda2019. It is not possible to impose the ZLB constraint using
Markov switching models with exogenous regimes, as in
LiuTheodoridisMumtazZanetti2019, because exogenous Markov-switching
cannot guarantee that the constraint will be respected with probability one,
and also does not account for the fact that the switch from one regime to the
other depends on the structural shocks. Finally, it is not possible to perform
consistent estimation and valid inference on the VAR (i.e., error bands with
correct coverage on impulse responses), using externally obtained measures of
the shadow rate, such as the one proposed by WuXia2016, as any such measures are subject to large and persistent estimation error that is not accounted for if they are treated as known in subsequent analysis. See also Rossi2019 for a
comprehensive discussion of the challenges posed by the ZLB for the estimation
of structural VARs.
The methodology developed in this paper allows for the presence of a shadow rate, estimates of which can be
obtained, but more importantly, it fully accounts for the impact of sampling
uncertainty in the estimation of the shadow rate on inference about the
structural parameters such as impulse responses. Therefore, the paper fills an
important gap in the literature, as it provides the requisite methodology to
implement any of the existing identification schemes. HayashiKoeda2019
develop a VAR model with endogenous regime switching in which the policy
variables that are subject to a lower bound are modelled using Tobit regressions. A key difference of
their methodology from the one developed here is that they impose recursive
identification of monetary policy shocks, which the present paper shows to be
an overidentifying, and hence testable, restriction. Moreover, their model
does not include shadow rates. A more recent
paper by AruobaSchorfheideVillalvazo2020 also studies SVARs with
occasionally binding constraints, but does not focus on the implications of
these constraints for identification.
Identification of the causal effects of policy by the ZLB does not require
that the policy reaction function be stable across regimes. However, inference
on the efficacy of unconventional policy, or equivalently, the causal effects
of shocks to the shadow rate over the ZLB period, obviously depends on whether
or not the reaction function remains the same across regimes. For example, an
attenuation of the causal effects of policy over the ZLB period may indicate
that unconventional policy is only partially effective, but it is also
consistent with unconventional policy being less active (during ZLB\ regimes)
than conventional policy (during non-ZLB regimes). This is a fundamental
identification problem that is difficult to overcome without additional
information, such as measures of unconventional policy stance, or additional
identifying assumptions, such as parametric restrictions or external
instruments. This can be done using the methodology developed in this paper.
The structure of the paper is as follows. Section (ref) presents
the main identification results of the paper in the context of a static
bivariate simultaneous equations model with a limited dependent variable
subject to a lower bound. Section (ref) generalizes the analysis to a
SVAR\ with an occasionally binding constraint and discusses identification,
estimation and inference. Section (ref) provides an application
to a three-equation SVAR in inflation, unemployment and the Federal funds rate
from StockWatson01. Using a sample of post-1960 quarterly US data, I
find some evidence that the ZLB\ is empirically relevant, and that
unconventional policy is only partially effective. Proofs and simulation results are given in the Appendix at the end.
Simultaneous equations model
I first illustrate the idea using a simple bivariate simultaneous equations
model (SEM), which is both analytically tractable and provides a direct link
to the related microeconometrics literature. To make the connection to the
leading application, I will motivate this using a very stylized economy
without dynamics in which the only outcome variable is inflation $\pi_{t}$ and
the (conventional) policy instrument is the short-term nominal interest rate,
$r_{t}$. In addition to the traditional interest rate channel, the model
allows for an `unconventional monetary policy' channel that can be used when
the conventional policy instrument hits the ZLB. An example of such a policy
is quantitative easing (QE), in the form of long-term asset purchases by the
central bank. Here I\ discuss a simple model of QE.\footnote{The more general
SVAR\ model of the next section can also incorporate forward guidance in the
form of ReifschneiderWilliams2000 and DebortoliGaliGambetti2019,
as shown in IkedaLiMavroeidisZanetti2020.}
Abstracting from dynamics and other variables, the equation that links
inflation to monetary policy is given by
equation[equation omitted — 114 chars of source]
where $c$ is a constant, $r^{n}$ is the neutral rate, $b_{L,t}$ is the
amount of long-term bonds held by the private sector in log-deviation from its
steady state, and $\varepsilon_{1t}$ is an exogenous structural shock
unrelated to monetary policy. Equation ((ref)) can be obtained
from a model of bond-market segmentation, as in ChenCurdiaFerrero2012,
where a fraction of households is constrained to invest only in long-term
bonds, see the Appendix for more details. In such a model, the
parameter $\varphi$ that determines the effectiveness of QE is proportional to
the fraction of constrained households and the elasticity of the term premium
with respect to asset holdings, both of which are assumed to be outside the
control of the central bank.
The nominal interest rate is set by a Taylor rule subject to the
ZLB\ constraint, namely,
subequations\begin{align}
r_{t} & =\max\left( r_{t}^{\ast},0\right) ,\\
r_{t}^{\ast} & =r^{n}+\gamma\pi_{t}+\varepsilon_{2t},
\end{align}
where $r_{t}^{\ast}$ represents the desired target policy rate, and
$\varepsilon_{2t}$ is a monetary policy shock. When $r_{t}^{\ast}$ is
negative, it is unobserved. The unobserved $r_{t}^{\ast}$ will be referred to
as the `shadow rate', and it represents the desired policy stance prescribed
by the Taylor rule in the absence of a binding ZLB constraint.
Suppose that QE is activated only when the conventional policy instrument
$r_{t}$ hits the ZLB,\footnote{The assumption that QE is only active during
the ZLB regime is only made for simplicity, as it is inconsequential for the
resulting functional form of the transmission equation. We can let QE be
active all the time, and even allow for a different rule for QE above and
below the ZLB, i.e., $b_{L,t}=\alpha\min\left( r_{t}^{\ast
},0\right) +\alpha_{1}\max\left( r_{t}^{\ast},0\right)$. Then, substituting back into
((ref)) yields an equation that is isomorphic to
((ref)), i.e., $\pi_{t}=\allowbreak c_{1}\allowbreak
+\allowbreak\bar{\beta}r_{t}\allowbreak+\bar{\beta}^{\ast}
\min\left( r_{t}^{\ast},0\right) \allowbreak+\varepsilon_{1t},$ with
$\bar{\beta}:=\allowbreak\beta +\varphi\alpha_{1} $ and
$\bar{\beta}^{\ast}\allowbreak:=\allowbreak\varphi\alpha.$} and follows
the same policy rule ((ref)), up to a factor of proportionality
$\alpha,$\ i.e.,
\[
b_{L,t}=\min\left( \alpha r_{t}^{\ast},0\right) .
\]
Substituting for $b_{L,t}$ in eq. ((ref)) and letting $\beta^{\ast}:=\alpha\varphi$, we obtain
equation[equation omitted — 152 chars of source]
A special case arises when QE is ineffective $\left( \varphi=0\right) ,$ or
the monetary authority does not pursue a QE policy $\left( \alpha=0\right)
,$ so that eq. ((ref)) becomes
equation[equation omitted — 104 chars of source]
and monetary policy is completely inactive at the ZLB.
Another special case of the model given by equations ((ref)) and
((ref)) arises when $\beta^{\ast}=\beta$ in
eq. ((ref)). This happens when $\varphi\neq0$ and $\alpha$
is chosen by the monetary authority to be equal to $\beta/\varphi.$ This can be
done when policy makers know the transmission mechanism in
eq. ((ref)) and have no restrictions in setting the policy
parameter $\alpha$ so as to fully remove the impact of the ZLB on conventional
policy. In that case, the equation for the outcome variable becomes
equation[equation omitted — 113 chars of source]
The model given by equations ((ref)) and
((ref)) is one in which monetary policy is completely
unconstrained and there is no difference in outcomes across policy regimes.
Such models have been put forward by SwansonWilliams2014,
DebortoliGaliGambetti2019 and WuZhang2019.
The nesting model given by eq. ((ref)) allows the effects
of conventional and unconventional policy to differ. This could reflect
informational as well as political or institutional constraints that prevent
policy makers from calibrating their unconventional policy response to match
exactly the policy prescribed by the Taylor rule. For instance, it may be that
policy makers do not know the effectiveness of the QE channel $\varphi$, or
that the scale of asset purchases needed to achieve the desired policy stance
during a ZLB regime is too large to be politically acceptable. Such a
consideration may be particularly pertinent, for example, in the Eurozone.
Importantly, one does not need to take a theoretical stand on this issue,
because the methodology that I develop in the paper can accommodate a wide
range of possibilities, and, as I demonstrate below, the issue can be studied empirically.
To complete the specification of the model, I\ assume that the structural
shocks $\varepsilon_{t}=\left( \varepsilon_{1t},\varepsilon_{2t}\right)
^{\prime}$ are independently and identically distributed (i.i.d.)
Normal with covariance matrix $\Sigma=diag\left( \sigma_{1}^{2},\sigma
_{2}^{2}\right)$.
The analysis in this paper assumes that the shadow rate $r_t^*$ is only observed above the ZLB. If $r^*_t<0$ were observed up to scale, for
instance, if we could measure QE $b_{L,t}$ from the balance sheet of the central bank, then the computation of the likelihood would be much simpler -- no filtering would be needed to deal with
lags of $r_{t}^{\ast}<0$ on the right hand side of the SVAR model introduced
in the next section, but the identification problem would remain the same.
More generally, we could assume that $r_{t}^{\ast}<0$ is observed with some
measurement error $\eta_{t}$, and include the measurement equation
$b_{L,t}=\min\left( \alpha r_{t}^{\ast}+\eta_{t},0\right)$ in
the model together with a specification of the distribution of $\eta_{t}$. The
estimation method in this paper can then be seen as a special case
where we are entirely agnostic about
the measurement error. Adding such measures of unconventional policy is a
potentially very useful extension of the method since, if correctly
specified, they will likely improve estimation accuracy.
Connection to the microeconometrics literature
Equations ((ref)) and ((ref)) form a SEM with a limited dependent variable.
The special case with $\beta^{\ast}=0$ in ((ref)) can be referred to as a
kinked SEM (KSEM), while the opposite case of $\beta^{\ast}=\beta$ in
((ref)) can be called a censored SEM (CSEM). Variants of the KSEM model have
been studied in the early microeconometrics literature on limited dependent
variable SEMs. Amemiya1974 and Lee1976 studied multivariate
extensions of the well-known Tobit model Tobin1958.
NelsonOlson1978 argued that the KSEM was less suitable for
microeconometric applications than the CSEM, and the latter subsequently
became the main focus of the literature
SmithBlundell1986,BlundellSmith1989. BlundellSmith1994 studied
the unrestricted model using external instruments, so they did not consider
the implications of the kink for identification.
One important lesson from the microeconometrics literature is that
establishing existence and uniqueness of equilibria in this class of models is
non-trivial. GourierouxLaffontMonfort1980 define a model to be
`coherent' if it has a unique solution for the endogenous variables in terms
of the exogenous variables, i.e., if there exists a unique reduced form. More
recently, the literature has distinguished between existence and uniqueness of
solutions using the terms coherency and completeness of the model,
respectively Lewbel2007. Establishing coherency and completeness is a
necessary first step before we can study identification and
estimation.
Identification
Substituting for $r_{t}^{\ast}$ in ((ref)) using
((ref)) and rearranging, we obtain
align[align omitted — 399 chars of source]
The system of equations ((ref)) and ((ref)) is now a
KSEM, for which the necessary and sufficient condition for coherency and
completeness (existence of a unique solution) is $\widetilde{\beta}\gamma<1$
NelsonOlson1978. Using ((ref)), the coherency and
completeness condition can be expressed in terms of the structural parameters
as
equation[equation omitted — 89 chars of source]
This condition evidently restricts the admissible
range of the structural parameters. It is satisfied in the present monetary
policy model, where it is natural to assume that $\beta,\beta^{\ast}\leq0$ and
$\gamma>0$. Therefore, it is possible that the coherency condition may not
provide additional information relative to what is often available from
natural sign restrictions on the parameters.
Under condition ((ref)), the unique solution of the
model can be written as
align[align omitted — 195 chars of source]
where $D_{t}:=1_{\left\{ r_{t}=0\right\} }$ is an indicator (dummy) variable
that takes the value 1 when $r_{t}$ is on the boundary and zero otherwise, and
align[align omitted — 384 chars of source]
Equations ((ref)) and ((ref)) express the endogenous
variables $\pi_{t},r_{t}$ in terms of the exogenous variables $\varepsilon
_{1t},\varepsilon_{2t},$ and correspond to the decision rules of the agents in
the model. It is clear that those decision rules differ in a world in which
the ZLB occasionally binds, which is characterized by $\widetilde{\beta}
\neq0,$ compared to a world in which it never does (i.e., the CSEM), where
$\widetilde{\beta}=0.$ What is important for identification, however, is that
in a world in which the ZLB occasionally binds, agents' reaction to shocks
differs across regimes, and the difference depends on the parameter
$\widetilde{\beta},$ which from eq. ((ref)), depends on the
difference between the impact of conventional and unconventional
policies, $\beta$ and $\beta^{\ast},$ respectively. I will
show that this change provides information that identifies the structural
parameters: we get point identification when $\beta^{\ast}=0$ (the KSEM case),
and partial identification when $\beta^{\ast}\neq\beta.$ The identification
argument leverages the coefficient on the kink, $\widetilde{\beta},$ in the
`incidentally kinked' regression ((ref)), which is identified by a
variant of the well-known Heckit method Heckman1979. I will sketch out
the argument below, and provide more details for the full SVAR model in the
next section.
Identification of the KSEM
Recall that in the KSEM\ model $\widetilde{\beta}=\beta.$ Consider the
estimation of $\beta$ in ((ref)) from a regression of
$\pi_{t}$ on $r_{t}$ using only observations above the ZLB,
equation[equation omitted — 244 chars of source]
where $\rho=cov\left( u_{1t},u_{2t}\right) /\tau^{2}-\beta=\gamma\sigma
_{1}^{2}\left( 1-\gamma\beta\right) /\left( \gamma^{2}\sigma_{1}^{2}
+\sigma_{2}^{2}\right) $, $\tau=\sqrt{var\left( u_{2t}\right) },$ and
$\phi\left( \cdot\right) ,$ $\Phi\left( \cdot\right) $ are standard Normal
density and distribution functions, respectively. The coefficient $\rho$ is
the bias in the estimation of $\beta$ from the truncated regression
((ref)). Now, the mean of $\pi_{t}$ using the
observations at the ZLB is
equation[equation omitted — 147 chars of source]
Next, observe that $\mu_{2},$ $\tau$ and hence $\phi\left( a\right)
/\Phi\left( a\right) $ can be estimated from the Tobit regression
((ref)). Therefore, we can recover the bias $\rho$ and identify
$\beta.$ A simple way to implement this is the control function approach
Heckman1978. Let
\[
h_{t}\left( \mu_{2},\tau\right) :=\left( 1-D_{t}\right) \left( r_{t}
-\mu_{2}\right) -D_{t}\frac{\tau\phi\left( a\right) }{\Phi\left( a\right)
},
\]
and run the regression
equation[equation omitted — 133 chars of source]
where $c_{1}=c-\beta r^{n}$ is an unrestricted intercept. The rank condition
for the identification of $\beta$ is simply that the regressors in
((ref)) are not perfectly collinear. This holds if and only
if $0<\Pr\left( D_{t}=1\right) <1.$ So, as long as some but not all the
observations are at the boundary, the model is generically identified.
Partial identification of the unrestricted
SEM
The discussion of the previous subsection shows that $\widetilde{\beta}$ is
identified from the kink in the reduced-form equation for $\pi_{t}$
((ref)). It follows from eq. ((ref)) and the order condition that $\beta,\beta^*$ are not point identified. I will now demonstrate that they are partially identified.
The assumption $cov\left( \varepsilon_{1t},\varepsilon
_{2t}\right) =0$ implies (see proof of Proposition 3 for a derivation)
equation[equation omitted — 106 chars of source]
where $\omega_{ij}:=cov\left(u_{it},u_{jt}\right)$. Substituting for $\gamma$ in ((ref)) using ((ref)) yields
equation[equation omitted — 170 chars of source]
For any given value of the reduced-form parameters $\widetilde{\beta},\Omega:=var(u_t)$, $u_t=(u_{1t},u_{2t})'$, the identified set for $\left( \beta,\beta^{\ast}\right)$ is a one-dimensional
manifold in $\Re^{2}$ defined by eq. ((ref))
intersected with the coherency condition ((ref)).
It is instructive to illustrate the identified set graphically at some given value of
$\widetilde{\beta}\,$and $\Omega$. Consider, for example, the case
$\Omega$ equal to the identity $I_{2}$, at which ((ref)) yields $\gamma=-\beta,$ and the coherency
condition ((ref)) reduces to $1+\beta\beta^{\ast}>0,$ and ((ref)) yields the
function $\beta=\frac{\widetilde{\beta}+\beta^{\ast}}{1-\widetilde{\beta}
\beta^{\ast}}$. Figure (ref) plots this function at
$\widetilde{\beta}=-1/2$. and highlights in dark gray the region of
incoherency defined by $1+\beta\beta^{\ast}\leq0$. The identified set is the
part of the function $\beta=\frac{\widetilde{\beta}+\beta^{\ast}
}{1-\widetilde{\beta}\beta^{\ast}}$ that lies to the right of the pole at
$1/\widetilde{\beta},$ i.e., in the region $\beta^{\ast}>1/\widetilde{\beta
}=-2$ in this example.
Now, consider the additional restrictions $\beta
\geq\beta^{\ast}\geq0$ or $\beta\leq\beta^{\ast}\leq0$, highlighted by the
light gray shaded areas in Figure (ref). The interpretation of
those restrictions is that unconventional policy neither has the opposite
effect from conventional policy, nor is it more effective than conventional
policy. With this additional restriction, we see
that the identified set further shrinks to the part of $\beta=\frac
{\widetilde{\beta}+\beta^{\ast}}{1-\widetilde{\beta}\beta^{\ast}}$ in the
interval $(\widetilde{\beta}^{-1},0].$ The projection of the identified set
onto the $\beta$ axis yields $\beta\in(-\infty,\widetilde{\beta}],$ since
$\widetilde{\beta}<0$, while its projection onto the
$\beta^{\ast}$ axis yields $\beta^{\ast}\in$ $(\widetilde{\beta}^{-1},0]$. Because equations ((ref)) and ((ref))
encapsulate all the information in the reduced-form parameters about
$\beta,\beta^{\ast}$, the identified set obtained from them is
sharp.
figure[figure omitted — 868 chars of source]
Let us define a new parameter $\lambda$ such that $\beta^*=\lambda \beta$. The restriction indicated by the light gray areas in Figure (ref) corresponds to $\lambda\in\left[ 0,1\right]$. If we
interpret $\lambda$ as a measure of the efficacy of unconventional policy,
this restriction implies that unconventional policy is neither counter- nor
over-productive. This reparameterization offers a convenient way to discretize the parameter space when we compute the identified set numerically, as is the case in the more general model discussed in the next section.
In this bivariate model, it is possible to characterize the identified set analytically. Here, I discuss the identified set for $\beta$ and defer the discussion of $\lambda$ to Appendix (ref).
We have already established that $\beta$ is completely unidentified
when $\beta=\beta^*$ (equivalently $\lambda=1$), which corresponds to the CSEM. From the definition of
$\widetilde{\beta}$ in ((ref)), it follows that
$\beta=\beta^*$ implies $\widetilde{\beta}=0.$ So, when $\widetilde{\beta}=0$,
$\beta$ is completely unidentified. It remains to see what happens when
$\widetilde{\beta}\neq0$. Let $\gamma_{0}:=\omega_{12}/\omega_{11}$, which can
be interpreted as the value the reaction function coefficient $\gamma$ in
((ref)) would take if $\beta=0$, i.e., the value corresponding to a
Choleski identification scheme where $r_{t}$ is placed last. In Appendix (ref),
I prove the following bounds
equation[equation omitted — 653 chars of source]
We see that when $\widetilde{\beta}\neq0$ and $0\leq\widetilde{\beta}\gamma_0 \leq 1$, we
can identify both the sign of the causal effect $\beta$ of $r_{t}$ on $\pi_{t}$ and
get bounds on its magnitude. In particular, the
identified coefficient $\widetilde{\beta}$ is an attenuated measure of the
true causal effect $\beta$. Moreover, $\widetilde{\beta}\gamma_{0}>1$ implies
that $\beta^{\ast}$ has the opposite sign from $\beta$, i.e., unconventional
policy has the opposite effect of the conventional one. That could be
interpreted as saying that unconventional policy is counterproductive.
Finally, the hypothesis that unconventional policy is as effective as conventional policy, $\beta^*
=\beta$ or $\lambda=1$, is equivalent to the null hypothesis $H_{0}:\widetilde{\beta
}=0$. The alternative that unconventional policy is less effective than
conventional policy, $\beta>\beta^{\ast}$ if $\beta>0,$ or $\beta<\beta^{\ast
}$ if $\beta<0$, corresponds to the two-sided alternative $H_{1}
:\widetilde{\beta}\neq0$. This can be tested using a likelihood ratio test.
SVAR with an occasionally binding constraint
I now develop the methodology for identification and estimation of SVARs with
an occasionally binding constraint. Let $Y_{t}=\left( Y_{1t}^{\prime}
,Y_{2t}\right) ^{\prime}$ be a vector of $k$ endogenous variables,
partitioned such that the first $k-1$ variables $Y_{1t}$ are unrestricted and
the $k$th variable $Y_{2t}$ is bounded from below by $b$.\footnote{The lower
bound does not need to be constant. All we need is to observe the periods in
which the economy is at the ZLB\ regime.} Define the latent process
$Y_{2t}^{\ast}$ that is only observed, and equal to $Y_{2t},$ whenever
$Y_{2t}>b$. If $Y_{2t}$ is a policy instrument, $Y_{2t}^{\ast}$ can be thought
of as the `shadow' instrument that measures the desired policy stance. The
$p$th-order SVAR model is given by the equations
align[align omitted — 443 chars of source]
for $t\geq1$ given a set of initial values $Y_{-s},Y_{2,-s}^{\ast},$ for
$s=0,...,p-1$, and $X_{0t}$ are exogenous and predetermined variables.
Equation ((ref)) can be interpreted as a policy reaction function
because it determines the desired policy stance $Y_{2t}^{\ast}.$ Similarly,
$\varepsilon_{2t}$ is the corresponding policy shock. The above model is a
dynamic SEM. Two important differences from a standard SEM are the presence of
(i) latent lags amongst the predetermined variables on the right-hand side,
which complicates estimation; and (ii) the contemporaneous value of $Y_{2t}$
in the policy reaction function ((ref)), which allows it to vary
across ZLB and non-ZLB regimes. The presence of latent lags $Y_{2,t-j}^{\ast}$
in the policy rule ((ref)) is particularly useful because it
allows the model to incorporate forward
guidance ReifschneiderWilliams2000,DebortoliGaliGambetti2019, see Appendix (ref) for details.
Collecting all the observed predetermined variables $X_{0t},Y_{t-1}
,...,Y_{t-p}$ into a vector $X_{t},$ and the latent lags $Y_{2,t-1}^{\ast
},...,Y_{2,t-p}^{\ast}$ into $X_{t}^{\ast},$ and similarly for their
coefficients, the model can be written compactly as:
align[align omitted — 300 chars of source]
The vector of structural errors $\varepsilon_{t}$ is assumed to be
i.i.d. Normally distributed with zero mean and identity covariance.
In the previous section, we defined the KSEM as a special case of the general
model, where $Y_{2t}^{\ast}<b$ has no (contemporaneous) impact on $Y_{1t}.$ In
the dynamic setting, it feels natural to define the corresponding `kinked
SVAR' model (KSVAR) as a model in which $Y_{2t}^{\ast}$ has neither
contemporaneous nor dynamic effects. Therefore, the KSVAR obtains as a special
case of ((ref)) when both $A_{12}^{\ast}=0,$ and $B^{\ast}=0,$
which corresponds to a situation in which the bound is fully effective in
constraining what policy can achieve at all horizons.
The opposite extreme to the KSVAR is the censored SVAR model (CSVAR). Again,
unlike the CSEM, which only characterizes contemporaneous effects, the idea of
a CSVAR is to impose the assumption that the constraint is irrelevant at all
horizons. So, it corresponds to a fully unrestricted linear SVAR in the latent
process $\left( Y_{1t}^{\prime},Y_{2t}^{\ast}\right) ^{\prime}$. This is a
special case of ((ref)) when both $A_{12}=0$ and the elements of
$B$ corresponding to lagged $Y_{2t}$ are equal to zero. Finally, I refer to
the general model given by ((ref)) as the\ `censored and kinked
SVAR' (CKSVAR).
Define the $k\times k$ square matrices
equation[equation omitted — 254 chars of source]
$\overline{A}$ determines the impact effects of structural shocks during
periods when the constraint does not bind. $A^{\ast}$ does the same for
periods when the constraint binds.
To analyze the CKSVAR, we first need to establish existence and uniqueness of
the reduced form. This is done in the following proposition.
propositionThe model given in eq. ((ref)) is coherent
and complete (i.e., it has a unique solution) if and only if
\begin{equation}
\kappa:=\frac{\overline{A}_{22}-A_{21}A_{11}^{-1}\overline{A}_{12}}
{A_{22}^{\ast}-A_{21}A_{11}^{-1}A_{12}^{\ast}}>0.
\end{equation}
Note that ((ref)) does not depend on the coefficients on the
lags (whether latent or observed), so it is exactly the same as in a static
SEM. This condition is useful for inference, e.g., when constructing
confidence intervals or posteriors, because it restricts the range of
admissible values for the structural parameters. It can also be checked
empirically when the structural parameters are point-identified.
If condition ((ref)) is satisfied, there exists a reduced-form
representation of the CKSVAR model ((ref)). For convenience of
notation, define the indicator (dummy variable) that takes the value one if
the constraint binds and zero otherwise:
equation[equation omitted — 68 chars of source]
propositionIf ((ref)) holds, and for any initial values
$Y_{-s},Y_{2,-s}^{\ast},$ $s=0,...,p-1,$ the reduced-form representation of
((ref)) for $t\geq1$ is given by
\begin{align}
Y_{1t} & =\overline{C}_{1}X_{t}+\overline{C}_{1}^{\ast}\overline{X}
_{t}^{\ast}+u_{1t}-\widetilde{\beta}D_{t}\left( \overline{C}_{2}
X_{t}+\overline{C}_{2}^{\ast}\overline{X}_{t}^{\ast}+u_{2t}-b\right)
\\
Y_{2t} & =\max\left( \overline{Y}_{2t}^{\ast},b\right) ,
\\
\overline{Y}_{2t}^{\ast} & =\overline{C}_{2}X_{t}+\overline{C}_{2}^{\ast
}\overline{X}_{t}^{\ast}+u_{2t},\\
Y_{2t}^{\ast} & =\left( 1-D_{t}\right) \overline{Y}_{2t}^{\ast}
+D_{t}\left( \kappa\overline{Y}_{2t}^{\ast}+\left( 1-\kappa\right)
b\right) ,
\end{align}
where $u_{t}=\left( u_{1t}^{\prime},u_{2t}\right) ^{\prime}=\overline
{A}^{-1}\varepsilon_{t},$ $\overline{C}^{\ast}=\left( \overline{C}_{1}
^{\ast\prime},\overline{C}_{2}^{\ast\prime}\right) ^{\prime}=\kappa
\overline{A}^{-1}B^{\ast},$ $\overline{X}_{t}^{\ast}=\left( \overline
{x}_{t-1},...,\overline{x}_{t-p}\right) ^{\prime},$ $\overline{x}_{t}
=\min\left( \overline{Y}_{2t}^{\ast}-b,0\right) ,$ $\overline{x}_{-s}
=\kappa^{-1}\min\left( Y_{2,-s}^{\ast}-b,0\right) ,$ $s=0,...,p-1,$
\begin{equation}
\widetilde{\beta}=\left( A_{11}-A_{12}^{\ast}A_{22}^{\ast-1}A_{21}\right)
^{-1}\left( A_{12}^{\ast}A_{22}^{\ast-1}A_{22}-A_{12}\right) ,
\end{equation}
$\kappa$ is defined in ((ref)) and the matrices $\overline
{C}_{1},\overline{C}_{2},$ are given in eq. ((ref)) in the Appendix.
Note that the \textquotedblleft reduced-form\textquotedblright\ latent process
$\overline{Y}_{2t}^{\ast}$ is, in general, different from the
\textquotedblleft structural\textquotedblright\ shadow rate $Y_{2t}^{\ast}$
defined by ((ref)). They coincide only when $\kappa=1.$ This holds,
for example, in the CSVAR model.
Equation ((ref)) combined with ((ref)) is a familiar
dynamic Tobit regression model with the added complexity of latent lags
included as regressors whenever $\overline{C}_{2}^{\ast}\neq0.$ Likelihood
estimation of the univariate version of this model was studied by
Lee1999. The $k-1$ equations ((ref)) are `incidentally
kinked' dynamic regressions, that I\ have not seen analyzed before.
Identification
Identification of reduced-form parameters
Let $\psi$ denote the parameters that characterize the reduced form
((ref))-((ref)): $\widetilde{\beta},\overline{C},$
$\overline{C}^{\ast}$ and $\Omega=var\left( u_{t}\right) .$ It is useful to
decompose $\psi$ into $\psi_{2}=\left( \overline{C}_{2},\overline{C}
_{2}^{\ast},\tau\right) ^{\prime},$ where $\tau=\sqrt{var\left(
u_{2t}\right) },$ and $\psi_{1}=( vec\left( \overline{C}_{1}\right)
^{\prime},vec\left( \overline{C}_{1}^{\ast}\right) ^{\prime}
,\allowbreak\widetilde{\beta}^{\prime},\delta^{\prime},vech\left( \Omega_{1.2}\right)
) ^{\prime},$ where $\delta=\Omega_{12}/\tau^{2}$, $\Omega_{1.2}
=\Omega_{11}-\delta\delta^{\prime}\tau^{2}$, and $\Omega_{ij}=cov\left(
u_{it},u_{jt}\right) $.
Equation ((ref)) is the dynamic Tobit regression model studied by
Lee1999. So, its parameters, $\psi_{2},$ are generically identified
provided that the regressors are not perfectly collinear. This requires that
$0<\Pr\left( D_{t}=1\right) <1.$
Given $\psi_{2},$ the identification of the remaining parameters, $\psi_{1},$
can be characterized using a control function approach. Consider the $k-1$
regression equations
equation[equation omitted — 206 chars of source]
where
align[align omitted — 418 chars of source]
$a_{t}=\left( \frac{b-\overline{C}_{2}X_{t}-\overline{C}_{2}^{\ast}
\overline{X}_{t}^{\ast}}{\tau}\right) $, and $\phi\left( \cdot\right) ,$
$\Phi\left( \cdot\right) $ are the standard Normal density and distribution
functions, respectively. When $\overline{C}^{\ast}$ is different from zero,
regressors $\overline{X}_{t}^{\ast}$, $Z_{1t},$ and $Z_{2t}$ in
((ref)) are unobserved, so we need to replace them with their
expectations conditional on $Y_{2t},Y_{t-1},...,Y_{1}.$ Then, the regressors
on the right-hand side of ((ref)) become $\mathbf{X}_{t}:=\left(
X_{t}^{\prime},\overline{X}_{t|t}^{\ast\prime},Z_{1t|t},Z_{2t|t}\right)
^{\prime}$, where $h_{t|t}:=E\left( h\left( \overline{X}_{t}^{\ast}\right)
|Y_{2t},Y_{t-1},...,Y_{1}\right) $ for any function $h\left( \cdot\right) $
whose expectation exists.\footnote{In the KSVAR model, we have $\overline
{C}_{1}^{\ast}=0$ and $\overline{C}_{2}^{\ast}=0,$ so $\overline{X}_{t}^{\ast
}$ drops out of ((ref)), and the regressors $Z_{1t},Z_{2t}$ are
observed, so $Z_{jt|t}=Z_{jt},$ $j=1,2.$} The coefficients $\overline{C}
_{1},\overline{C}_{1}^{\ast},\widetilde{\beta},$ and $\delta$ are generically
identified if the regressors $\mathbf{X}_{t}$ are not perfectly collinear.
Identification of structural parameters
From the order condition, we can easily establish that there are not enough
restrictions to identify all the structural parameters in the
CKSVAR\ ((ref)). Let $k_{0}=\dim\left( X_{0t}\right) $ denote the
number of predetermined variables other than the own lags of $Y_{t}.$ For
example, in a standard VAR without deterministic trends, we have $X_{0t}=1,$
so $k_{0}=1.$ The number of reduced-form parameters $\psi$ is $k_{0}k+k^{2}p$
(in $\overline{C}$) plus $kp$ (in $\overline{C}^{\ast}$) plus $k-1$ (in
$\widetilde{\beta}$) plus $k\left( k+1\right) /2$ (in $\Omega$). The number
of structural parameters in ((ref)) is $k_{0}k+k^{2}p$ (in $B$)
plus $kp$ (in $B^{\ast}$) plus $k^{2}$ (in $\overline{A}$) plus $k$ (in
$A_{12}^{\ast}$ and $A_{22}$). So, the CKSVAR is underidentified by $k\left(
k-1\right) /2+1$ restrictions. Nevertheless, I will show that the impulse
responses to $\varepsilon_{2t}$ are identified. Specifically, they are
point-identified when $A_{12}^{\ast}=0,$ and partially identified when
$A_{12}^{\ast}\neq0$ but $A_{12}^{\ast}$ and $A_{12}$ have the same sign,
analogous to the bounds given in equation ((ref)) in the previous section.
Because the CKSVAR is nonlinear, IRFs are obviously state-dependent, and there
are many ways one can define them, see KoopPesaranPotter1996. The
IRF to $\varepsilon_{2t},$ according to any of the definitions proposed in the
literature, is identified if the reduced-form errors $u_{t}$ can be expressed
as a known function of $\varepsilon_{2t}$ and a process that is orthogonal to
it, i.e., $u_{t}=g\left( \varepsilon_{2t},e_{t}\right) ,$ where $e_{t}$ is
independent of $\varepsilon_{2t}.$ From Proposition (ref), it follows
that the function $g$ is linear, and more specifically,
align[align omitted — 361 chars of source]
where
align[align omitted — 289 chars of source]
and $\overline{A}_{22}=A_{22}^{\ast}+A_{22},$ defined in
((ref)). Note that $\overline{\beta}$ can be interpreted as
the response of $Y_{1t}$ to a shock that increases $Y_{2t}$ by one unit, and
$\overline{\gamma}$ are the contemporaneous reaction function coefficients of
$Y_{2t}$ to $Y_{1t}$ when $Y_{2t}>b$ (unconstrained regime). The shock vector
$\overline{\varepsilon}_{1t}$ is not structural but it is orthogonal to
$\varepsilon_{2t},$ so it plays the role of $e_{t}$ in $u_{t}=g\left(
\varepsilon_{2t},e_{t}\right) .$ Hence, the IRF\ is identified if and only if
$\overline{\beta},$ $\overline{\gamma},$ and $\overline{A}_{22}$ are identified.
The following proposition shows identification when $A_{12}^{\ast}=0$.
propositionWhen $A_{12}^{\ast}=0$ and the coherency condition
((ref)) holds, the parameters in ((ref))-((ref))
are identified by the equations $\overline{\beta}=\widetilde{\beta},$
\begin{align}
\overline{\gamma} & =\left( \Omega_{12}^{\prime}-\Omega_{22}\overline
{\beta}^{\prime}\right) \left( \Omega_{11}-\Omega_{12}\overline{\beta
}^{\prime}\right) ^{-1}, and\\
\overline{A}_{22}^{-1} & =\sqrt{\left( -\overline{\gamma},1\right)
\Omega\left( -\overline{\gamma},1\right) ^{\prime}}.
\end{align}
Remarks
1. $\overline{\beta}=\widetilde{\beta}$ follows immediately from the
definition ((ref)) with $A_{12}^{\ast}=0$. Equations
((ref)) and ((ref)) hold without the
restriction $A_{12}^{\ast}=0.$ They follow from the orthogonality of the
shocks $\varepsilon_{2t}$ and $\overline{\varepsilon}_{1t}.$
2. An instrumental variables interpretation of this identification result is
as follows. Define the instrument
\[
Z_{t}:=Y_{1t}-\widetilde{\beta}Y_{2t}=A_{11}^{-1}B_{1}X_{t}+A_{11}^{-1}
B_{1}^{\ast}X_{t}^{\ast}+A_{11}^{-1}\varepsilon_{1t},
\]
where the second equality holds when $A_{12}^{\ast}=0$. The orthogonality of
the errors $E\left( \varepsilon_{1t}\varepsilon_{2t}\right) =0$ implies
$E\left( Z_{t}\varepsilon_{2t}\right) =0.$ So, $Z_{t}$ are valid $k-1$
instruments for the $k-1$ endogenous regressors $Y_{1t}$ in the structural
equation of $Y_{2t}=\max\left( Y_{2t}^{\ast},b\right) $, where $Y_{2t}
^{\ast}$ is given by ((ref)). Normalizing ((ref)) in
terms of $Y_{2t}^{\ast}$ yields the structural equation in the more familiar
form of a policy rule:
equation[equation omitted — 162 chars of source]
where $\bar{B}_{2}=$ $\overline{A}_{22}^{-1}B_{2},$ $\bar{B}_{2}^{\ast}=$
$\overline{A}_{22}^{-1}B_{2}^{\ast}$. Since $A_{11}^{-1}$ is non-singular,
the\ coefficient matrix of $Z_{t}$ in the `first-stage' regressions of
$Y_{1t}$ is nonsingular, so the coefficients of ((ref)) are
generically identified by the rank condition. An alternative to the Tobit IV
regression model ((ref)) is the indirect Tobit regression approach
used in the static SEM by BlundellSmith1994. Equation
((ref)) can be written as the dynamic Tobit regression
equation[equation omitted — 168 chars of source]
where $\widetilde{\gamma}=\left( 1-\overline{\gamma}\overline{\beta}\right)
^{-1}\overline{\gamma},$ $\tilde{B}_{2}=\left( 1-\overline{\gamma}
\overline{\beta}\right) ^{-1}\bar{B}_{2},$ $\tilde{B}_{2}^{\ast}=\left(
1-\overline{\gamma}\overline{\beta}\right) ^{-1}\bar{B}_{2}^{\ast}$ and
$\tilde{\varepsilon}_{2t}=\left( 1-\overline{\gamma}\overline{\beta}\right)
^{-1}\bar{\varepsilon}_{2t}.$ Note that the coherency condition
((ref)) becomes $\kappa=\frac{\overline{A}_{22}}{A_{22}^{\ast}
}\left( 1-\overline{\gamma}\overline{\beta}\right) >0,$ so $1-\overline
{\gamma}\overline{\beta}\neq0,$ which guarantees the existence of the
representation ((ref)). Given $\overline{\beta}=\widetilde{\beta
},$ the structural parameter $\overline{\gamma}$ can then be obtained as
$\overline{\gamma}=\widetilde{\gamma}\left( I_{k-1}+\widetilde{\beta
}\widetilde{\gamma}\right) ^{-1},$ and similarly for the remaining structural
parameters in ((ref)).
3. The parameter $A_{22}$ allows the reaction function of $Y_{2t}^{\ast}$ to
differ across the two regimes. The special case $A_{22}=0$ thus corresponds to
the restriction that the reaction function remains the same across regimes.
The parameters $A_{22}$ and $A_{22}^{\ast}$ are not separately identified.
Hence, $A_{22}^{\ast-1}$, the scale of the response to the shock
$\varepsilon_{2t}$ during periods when $Y_{2t}=b$,$\,$ is not
identified.\footnote{This is akin to the well-known property of a probit model
that the scale of the distribution of the latent process is not identifiable.}
Similarly, $\kappa=\frac{\overline{A}_{22}}{A_{22}^{\ast}}\left(
1-\overline{\gamma}\overline{\beta}\right) $ is not identified, and
therefore, neither is the structural shadow value $Y_{2t}^{\ast}$ in eq.
((ref)). Identification of these requires an additional restriction
on $A_{22}$, e.g., $A_{22}=0.$ Turning this discussion around, we see that a
change in the reaction function across regimes does not destroy the point
identification of the effects of policy during the unconstrained regime, since
the latter only requires $\overline{\beta},\overline{\gamma}$ and
$\overline{A}_{22}$, not $A_{22}^{\ast}$ or $\kappa.$
Next, we turn to the case $A_{12}^{\ast}\neq0,$ and derive identification
under restrictions on the sign and magnitude of $A_{12}^{\ast}$ relative to
$A_{12}$ and $A_{22}^{\ast}$ relative to $A_{22}.$ The first restriction is
motivated by a generalization of the discussion on the SEM model in Subsection
(ref). Specifically, if $\overline{A}_{12}=A_{12}+A_{12}^{\ast}$
measures the effect of conventional policy (operating in the unconstrained
regime) and $A_{12}^{\ast}$ measures the effect of unconventional policy
(operating in the constrained regime), then the assumption that $A_{12}$ and
$A_{12}^{\ast}$ have the same sign means that unconventional policy effects
are neither in the opposite direction nor larger in absolute value than
conventional policy effects. In other words, unconventional policy is neither
counterproductive nor over-productive relative to conventional policy. This
can be characterized by the specification $A_{12}^{\ast}=\Lambda\overline
{A}_{12}$ and $A_{12}=\left( I_{k-1}-\Lambda\right) \overline{A}_{12},$
where $\Lambda=diag\left( \lambda_{j}\right) ,$ $\lambda_{j}\in\left[
0,1\right] $ for $j=1,...,k-1.$ I further impose the restriction that
$\lambda_{j}=\lambda$ for all $j,$ so that $A_{12}^{\ast}=\lambda\overline
{A}_{12}$ and $A_{12}=\left( 1-\lambda\right) \overline{A}_{12}$ with
$\lambda\in\left[ 0,1\right] .$ This, in turn, means that $Y_{2t}$ and
$Y_{2t}^{\ast}$ enter each of the first $k-1$ structural equations for
$Y_{1t}$ only via the common linear combination $\lambda Y_{2t}^{\ast
}\allowbreak+\allowbreak\left( 1-\lambda\right) Y_{2t},$ which can be
interpreted as a measure of the effective policy stance.
We also need to consider the impact of $A_{22}$ on identification. The
parameter $\zeta=\overline{A}_{22}/A_{22}^{\ast}$ gives the ratio of the
standard deviation of the monetary policy shock in the constrained relative to
the unconstrained regime. It is also the ratio of the reaction function
coefficients in the two regimes, e.g., $A_{22}^{\ast-1}A_{21}$ versus
$\overline{A}_{22}^{-1}A_{21}$. I will impose $\zeta>0,$ so that the sign of
the policy shock does not change across regimes. With the above
reparametrization and the definitions in ((ref)), the
identified coefficient $\widetilde{\beta}$ in ((ref)) can be
written as
equation[equation omitted — 180 chars of source]
Similarly, given $\zeta>0,$ the coherency condition ((ref))
reduces to $\left( 1-\overline{\gamma}\overline{\beta}\right) \left(
1-\xi\overline{\gamma}\overline{\beta}\right) >0.$ Notice that the parameters
$\lambda,\zeta$ only appear multiplicatively, so it suffices to consider them
together as $\xi=\lambda\zeta.$ Once $\overline{\beta}$ is known, the
remaining structural parameters needed to obtain the IRF to $\varepsilon_{2t}$
are $\overline{\gamma}$ and $\overline{A}_{22}$, and they are obtained from
Proposition (ref). So, the identified set can be characterized
by varying $\xi$ over its admissible range. Without further restrictions on
$\zeta,$ the admissible range is obviously $\xi\geq0.$ If we further assume
that $\zeta\leq1,$ i.e., that the slope of the reaction function coefficients
is no steeper in the constrained regime than in the unconstrained regime, then
$\xi\in\left[ 0,1\right] ,$ and so partial identification proceeds exactly
along the lines of the SEM in the previous section where $\lambda$ played the
role of $\xi.$ In the case $k=2,$ the bounds derived in eq. ((ref))
apply, with $\beta=\overline{\beta}$ in the notation of the present section.
However, when $k>2,$ it is difficult to obtain a simple analytical
characterization of the identified set for $\overline{\beta}.$ In any case, we
will typically wish to obtain the identified set for functions of the
structural parameters, such as the IRF. This can be done numerically by
searching over a fine discretization of the admissible range for $\xi.$ An
algorithm for doing this is provided in Appendix (ref).
Estimation
Estimation of the CKSVAR is carried out by Maximum Likelihood (ML) using either a version of the sequential importance sampler (SIS) of Lee1999 or the fully adapted particle filter (FAPF) of MalikPitt2011 to evaluate the likelihood, except in the case of the KSVAR model for which the likelihood is available analytically. The details are given in Appendix (ref).
Using the limit theory of Newe94Mc, the ML estimator can be shown
to be consistent and asymptotically Normal and the LR statistic asymptotically
$\chi^{2}$ with degrees of freedom equal to the number of restrictions. Standard asymptotics arise when the probability of each regime occurring is bounded away from zero. Infrequent visits to one of the two regimes will slow down the rate of convergence of the estimator, but will not lead to a non-standard limiting distribution. Since the focus of this paper is on identification, I will not discuss primitive conditions for these results, such as geometric ergodicity, which can be shown, for example, by bounding the joint spectral radius of the companion-form representation of the model Liebscher2005. Instead, I report Monte Carlo simulation results on the finite-sample properties of ML estimators and LR tests in Appendix
(ref). They show that the Normal distribution provides a very
good approximation to the finite-sample distribution of the ML estimators. I
find some finite-sample size distortion in the LR tests of various
restrictions on the CKSVAR, but this can be addressed effectively with a
parametric bootstrap, as shown in the Appendix.
One interesting observation from the simulations is that the LR test of the
CSVAR restrictions against the CKSVAR appears to be less powerful than the
corresponding test of the KSVAR restrictions against the CKSVAR. Thus, we
expect to be able to detect deviations from KSVAR more easily than deviations
from CSVAR. In other words, finding evidence against the hypothesis that
unconventional policies are fully effective (CSVAR) may be harder than finding
evidence against the opposite hypothesis that they are completely ineffective (KSVAR).
Application
I use the three-equation SVAR of StockWatson01, consisting of
inflation, the unemployment rate and the Federal Funds rate to provide a
simple empirical illustration of the methodology developed in this paper. As
discussed in StockWatson01, this model is far too limited to provide
credible identification of structural shocks, so the results in this section
are meant as an illustration of the new methods.
The data are quarterly and are constructed exactly as in StockWatson01
.\footnote{The inflation data are computed as $\pi_{t}=$ $400ln(P_{t}
/P_{t-1})$, where $P_{t}$, is the implicit GDP deflator and $u_{t}$ is the
civilian unemployment rate. Quarterly data on $u_{t}$ and $i_{t}$ are formed
by taking quarterly averages of their monthly values.} The variables are
plotted in Figure (ref) over the extended sample 1960q1 to 2018q2.
I will consider all periods in which the Fed funds rate was below 20 basis
points to be on the ZLB. This includes 28 quarters, or 11% of the sample.
figure[figure omitted — 215 chars of source]
Tests of efficacy of unconventional policy
I estimate three specifications of the SVAR(4) with the ZLB: the unrestricted
CKSVAR specification, as well as the restricted KSVAR and CSVAR
specifications. The maximum log-likelihood for each model is reported in Table
(ref), computed using the SIS\ algorithm in the case of CKSVAR
and CSVAR, with 1000 particles. The accuracy of the SIS algorithm was gauged
by comparing the log-likelihood to the one obtained using the resampling FAPF
algorithm. In both CKSVAR and CSVAR the difference is very small. The results
are also very similar when we increase the number of particles to 10000.
Finally, the table reports the LR tests of KSVAR and
CSVAR against CKSVAR using both asymptotic and parametric bootstrap p-values.
table[table omitted — 1,582 chars of source]
The KSVAR imposes the restriction that no latent lags (i.e., lags of the
shadow rate) should appear on the right hand side of the model, i.e.,
$B^{\ast}=0$ in ((ref)) or $\overline{C}_{1}^{\ast}=0$ and
$\overline{C}_{2}^{\ast}=0$ in ((ref)) and ((ref)). This
amounts to 12 exclusion restrictions on the CKSVAR(4), four restrictions in
each of the three equations. This is necessary (but not sufficient) for the
hypothesis that unconventional policy is completely ineffective at all
horizons. It is necessary because $\overline{C}^{\ast}=\left( \overline
{C}_{1}^{\ast\prime},\overline{C}_{2}^{\ast\prime}\right) ^{\prime}\neq0$
would imply that unconventional policy has at least a lagged effect on
$Y_{t}.$ $\overline{C}^{\ast}=0$ is not sufficient to infer that
unconventional policy is completely ineffective because it may still have a
contemporaneous effect on $Y_{1t}$ if $A_{12}^{\ast}\neq0,$ and the latter is
not point-identified. The result of the test in Table (ref)
shows that lags of the shadow rate are statistically significant at the 5%
level, rejecting the null hypothesis that unconventional policy has no effect.
The CSVAR model imposes the restriction that only the coefficients on the lags
of the shadow rate (which is equal to the actual rate above the ZLB) are
different from zero in the model, i.e., the elements of $B$ corresponding to
lags of $Y_{2t}$ in ((ref)) are all zero, or equivalently, the
elements of $\overline{C}$ corresponding to lags of $Y_{2t}$ in
((ref)) and ((ref)) are all equal to $\overline{C}^{\ast}
$. In addition, it imposes the restriction that $\widetilde{\beta}=0$ in
((ref)), i.e., no kink in the reduced-form equations for inflation
and unemployment across regimes, yielding 14 restrictions in total. This is
necessary for the hypothesis that the ZLB\ is empirically irrelevant for
policy in that it does not limit what monetary policy can achieve. The
evidence against this hypothesis is not as strong as in the case of the KSVAR.
The asymptotic p-value is 0.023, indicating rejection of the null hypothesis
that unconventional policy is as effective as conventional policy at the 5%
level, but the bootstrap p-value is 0.117. Note that this difference could
also be due to fact that the test of the CSVAR restrictions may be less
powerful than the test of the KSVAR restrictions, as indicated by the
simulations reported in the previous section. Thus, I\ would cautiously
conclude that the evidence on the empirical relevance of the ZLB\ is mixed.
Further evidence on the efficacy of unconventional policy will also be
provided in the next subsection.
Impulse response functions
Based on the evidence reported in the previous section, I estimate the IRFs
associated with the monetary policy shock using the unrestricted CKSVAR
specification, and compare them to recursive IRFs from the CSVAR specification
that place the Federal funds rate last in the causal ordering. From the
identification results in Section (ref), the CKSVAR point-identifies
the nonrecursive IRFs only under the assumption that the shadow rate has no
contemporaneous effect of $Y_{1t},$ i.e., $A_{12}^{\ast}=0$ in
((ref)). Note that, due to the nonlinearity of the model, the IRFs
are state-dependent. I use the following definition of the IRF from
KoopPesaranPotter1996:\footnote{The kink in the reduced-form representation of the model makes it difficult to approximate the IRFs by local projections on simple nonlinear functions of the data, such as powers or interactions with the regime indicator, see Appendix (ref) for further discussion of this point.}
equation[equation omitted — 254 chars of source]
Figure (ref) reports the nonrecursive IRFs to a 25 basis points
monetary policy shock at the end of the sample, 2018q3 (at which point $\overline{X}_{t}^{\ast}=0$ in ((ref)) because interest rates had been above the ZLB over the previous four quarters), from the CKSVAR under the assumption that unconventional policy has no contemporaneous effect ($\lambda=0$). It also reports two different
estimates of recursive IRFs using the identification scheme in
StockWatson01 with interest rates placed last. The first estimate is obtained
from the CSVAR specification, and the second is a “naive” OLS estimate of the IRF that ignores the ZLB constraint -- a direct application of the method in StockWatson01 to the present sample. The figure also
reports 90% bootstrap error bands for the nonrecursive IRFs.
figure[figure omitted — 768 chars of source]
In the nonrecursive IRF, the response of inflation to a monetary tightening is
negative on impact, albeit very small, and, with the exception of the first
quarter when it is positive, it stays negative throughout the horizon. Hence,
the incidence of a price puzzle is mitigated relative to the recursive IRFs,
according to which inflation rises for up to 6 quarters after a monetary
tightening (9 quarters in the OLS case). Note, however, that the error bands
are so wide that they cover (pointwise) most of the recursive IRF, though less
so for the OLS one. Turning to the unemployment response, we see that the
nonrecursive IRF starts significantly positive on impact (no transmission lag)
and peaks much earlier (after 4 quarters) than the recursive IRF (10
quarters). In this case, the recursive IRF\ is outside the error bands for
several quarters (more so for the naive OLS IRF). Finally, the response of the
Federal funds rate to the monetary tightening is less than one on impact and
generally significantly lower than the recursive IRFs. This is both due to the
contemporaneous feedback from inflation and unemployment, as well as the fact
that there is a considerable probability of returning to the ZLB, which
mitigates the impact of monetary tightening.
Next, I turn to the identified sets of the IRFs that arise when I relax the
restriction that unconventional policy is ineffective, i.e., $\lambda$ can be
greater than zero. I consider the range of $\xi=\lambda\zeta\in\left[
0,1\right] ,$ recalling that $\lambda$ measures the efficacy of
unconventional policy and $\zeta$ measures the ratio of the reaction function
coefficients and shock volatilities in the constrained versus the
unconstrained regimes. The shaded areas in Figure
(ref) report the identified sets without any other
restrictions. The striped areas (a subset of the aforementioned identified sets) show the tightening of the identified sets when I impose the additional sign restriction that the contemporaneous effect of the
monetary policy shock to the Fed Funds rate should be nonnegative. The bold lines show the IRFs under the (point-identifying) assumption $\lambda=0$. The latter are the same as the
nonrecursive point estimates reported in Figure (ref).
We observe that the identified set for the IRF of inflation is bounded from
above by the limiting case $\lambda=0$. This is also true of the response of
the Fed Funds rate. The case $\lambda=0$ provides a lower bound on the effect
to unemployment only from 0 to 9 quarters. Even though the point estimate of
the unemployment response under $\lambda=0$ remains positive over all
horizons, the identified set includes negative values beyond 10 quarters
ahead. We also notice that the identified sets are fairly large, albeit still
informative. Interestingly, the identified IRF of the Fed Funds rate includes
a range of negative values on impact. These values arise because for values of
$\xi>0$, there are generally two solutions for the structural VAR parameters
$\overline{\beta},\overline{\gamma}$ in the equations
((ref)), ((ref)), with one of them inducing
such strong responses of inflation and unemployment to the interest rate that
the contemporaneous feedback in the policy rule would in fact revert the
direct positive effect of the policy shock on the interest rate. If we impose
the additional sign restriction that the contemporaneous impact of the policy
shock to the Fed Funds rate must be non-negative, then those values are ruled
out and the identified sets become considerably tighter. This is an example of
how sign restrictions can lead to tighter partial identification of the IRF.
figure[figure omitted — 563 chars of source]
With an additional assumption on $\zeta,$ the method can be used to obtain an
estimate of the identified set for $\lambda,$ the measure of the efficacy of
unconventional policy. In particular, if we set $\zeta=1,$ i.e., the reaction
function remains the same across the two regimes, then the identified set for
$\lambda$ is $\left[ 0.0.506\right] .$ In other words, the identified set
excludes values of the efficacy of policy beyond 51%, so that, roughly
speaking, unconventional policy is at most 51% as effective as conventional
one. Note that this estimate does not account for sampling uncertainty and
relies crucially on the assumption that the reaction function remains the same
across the two regimes. This assumption could be justified by arguing that
there is no reason to believe that policy objectives may have shifted over the
ZLB period, and that any desired policy stance was feasible over that period.
The latter assumption may be questionable. For example, one can imagine that
there may be financial and political constraints on the amount of quantitative
easing policy makers could do, which may cause them to proceed more cautiously
over the ZLB period than over regular times. Within the context of our model,
this would be reflected as a flatter policy reaction function over the ZLB
period than over the non-ZLB periods, i.e., it will correspond to $\zeta<1$.
To illustrate the implications of this for the identification of $\lambda$,
suppose that $\zeta=1/2,$ i.e., the shadow rate reacts half as fast to shocks
during the ZLB period than it does in the non-ZLB period. Then, the identified
set for $\lambda$ would include 1, i.e., the data would be consistent with the
view that unconventional policy is fully effective. So, under this alternative
assumption on $\zeta$, the reason we observed a subdued response to policy
shocks over the ZLB period is because policy was less active over that period,
and policy shocks were smaller, not because unconventional policy was
partially ineffective.
As I discussed in the introduction, it is difficult to make further progress
on this issue without further information or additional assumptions. The
technical reason is that the scale of the latent regression over the censored
sample is not identified, so additional information is required to untangle
the structural parameters $\lambda$ and $\zeta$ from $\xi=\lambda\zeta$. One
possibility would be to identify $\lambda$ from the coefficients on the lags
of $Y_{2t}$ and $Y_{2t}^{\ast}$ by imposing the (overidentifying) restriction
that $Y_{2t-j}$ and $Y_{2t-j}^{\ast}$ appear in the model only via the linear
combination $Y_{2t-j}^{\text{eff}}:=\lambda Y_{2t-j}^{\ast}\allowbreak
+\allowbreak\left( 1-\lambda\right) \allowbreak Y_{2t-j}$ for all lags
$j=1,...,p,$ where $Y_{2t}^{\text{eff}}$ can be interpreted as the effective
policy stance. Provided that the coefficients on the lags of $Y_{2t}^{\ast}$
or $Y_{2t}$ are not all zero, this restriction point identifies $\lambda,$ and
hence, partially identifies $\zeta$ from $\xi.$ One could obtain tighter
bounds by using sign restrictions IkedaLiMavroeidisZanetti2020, or obtain point identification by using conventional identification schemes. For instance, one can identify $\overline{\beta}$ directly
using external instruments, as in GertlerKaradi2015, and hence point
identify $\xi$ from ((ref)).
Conclusion
This paper has shown that the ZLB can be used constructively to identify the
causal effects of monetary policy on the economy. Identification relies on the
(in)efficacy of alternative (unconventional) policies. When unconventional
policies are partially effective in mitigating the impact of the ZLB, the
causal effects of monetary policy are only partially identified. A general
method is proposed to estimate SVARs subject to an occasionally binding
constraint. The method can be used to test the efficacy of unconventional
policy, modelled via a shadow rate. Application to a core three-equation SVAR
with US data suggests that the ZLB is empirically relevant and unconventional
policy is only partially effective.
appendix\section{Proofs}
\subsection{Derivation of identified set for model of Section (ref)}
Using the notation $\beta^* = \lambda\beta$, eq. ((ref)) can be expressed as
\begin{equation}
\widetilde{\beta}=g\left( \beta\right) \beta,\quadg\left(
\beta\right) :=\frac{1-\lambda}{1-\lambda\frac{\beta\left( \omega
_{12}-\omega_{22}\beta\right) }{\omega_{11}-\omega_{12}\beta}}.
\end{equation}
When $\omega_{12}=0,$ we have $g\left( \beta\right) =$ $\frac{1-\lambda
}{1+\lambda\omega_{22}\beta^{2}/\omega_{11}}\in\left( 0,1\right) $ for all
$\lambda\in\left( 0,1\right) .$ Therefore, when $\widetilde{\beta}\neq0,$
the sign of $\beta$ is the same as that of $\widetilde{\beta}$ and its
magnitude is lower, as stated in ((ref)).
Next, consider $\omega_{12}\neq0.$ It is easily seen that $g\left( 0\right)
=1-\lambda$ and $\lim_{\beta\rightarrow\pm\infty}\left( \beta\right) =0.$
Moreover,
\[
\frac{\partial g}{\partial\beta}=\lambda\left( 1-\lambda\right) \frac
{\omega_{12}\omega_{22}\beta^{2}-2\omega_{11}\omega_{22}\beta+\omega
_{11}\omega_{12}}{\left( \omega_{11}-\beta\omega_{12}-\beta\lambda\omega
_{12}+\beta^{2}\lambda\omega_{22}\right) ^{2}}.
\]
For $\lambda\in\left( 0,1\right) ,$ the above derivative function has zeros
at $
\omega_{12}\omega_{22}\beta^{2}-2\omega_{11}\omega_{22}\beta+\omega_{11}
\omega_{12}=0$,
which occur at
\[
\begin{array}
[c]{c}
\beta_{1}=\frac{\omega_{11}\omega_{22}+\sqrt{\omega_{11}\omega_{22}\left(
\omega_{11}\omega_{22}-\omega_{12}^{2}\right) }}{\omega_{12}\omega_{22}}\\
\beta_{2}=\frac{\omega_{11}\omega_{22}-\sqrt{\omega_{11}\omega_{22}\left(
\omega_{11}\omega_{22}-\omega_{12}^{2}\right) }}{\omega_{12}\omega_{22}}
\end{array}
,\quad\text{if }\omega_{12}\neq0.
\]
Now, because $0<\left( \omega_{11}\omega_{22}-\omega_{12}^{2}\right)
<\omega_{11}\omega_{22}$ implies $\sqrt{\omega_{11}\omega_{22}\left(
\omega_{11}\omega_{22}-\omega_{12}^{2}\right) }<\omega_{11}\omega_{22},$ we
have $\beta_{i}<0,$ $i=1,2,$ when $\omega_{12}<0$ and $\beta_{i}>0,$ $i=1,2,$
when $\omega_{12}>0.$
By symmetry, it suffices to consider only one of the two cases, e.g., the case
$\omega_{12}<0.$ In this case, $g^{\prime}\left( \beta\right) =\frac
{\partial g}{\partial\beta}<0$ for all $\beta>0$ and, since $g\left(
0\right) =1-\lambda$ and $g\left( \infty\right) =0,$ it follows that
$g\left( \beta\right) \in\left( 0,1-\lambda\right) $ for all $\beta>0.$
Thus, from ((ref)) we see that $\widetilde{\beta}<0$ cannot
arise from $\beta>0$ when $\omega_{12}<0.$ In other words, observing
$\widetilde{\beta}<0$ must mean that $\beta<0.$ Moreover, since $g^{\prime
}\left( \beta\right) <0$ for all $\beta>\beta_{1}$ and $\beta_{1}<0,$ it
must be that $g\left( \beta\right) >0$ for all $\beta>\beta_{1},$ and hence,
also for $\beta_{1}<\beta\leq0.$ At $\beta<\beta_{1},$ $g^{\prime}\left(
\beta\right) >0,$ and since $g^{\prime}\left( \beta\right) <0$ for all
$\beta<\beta_{2}<\beta_{1},$ and $g\left( -\infty\right) =0,$ it has to be
that $g\left( \beta\right) $ approaches zero from below as $\beta
\rightarrow-\infty,$ and therefore, $g\left( \beta\right) $ must cross zero
at some $\beta_{0}\in\left( \beta_{2},\beta_{1}\right) ,$ and $g\left(
\beta\right) \geq0$ for all $\beta\in\left[ \beta_{0},0\right] .$
Inspection of ((ref)) shows that $\beta_{0}=\omega_{11}
/\omega_{12}=1/\gamma_{0},$ which corresponds to $\gamma=-\infty$ from
((ref)). Since $g\left( \beta\right) \in\left[
0,1-\lambda\right] $ for all $\beta\in\left[ \beta_{0},0\right] ,$ and
$\lambda\in(0,1),$ it follows from ((ref)) that $\left\vert
\widetilde{\beta}\right\vert \leq\left\vert \beta\right\vert $. In other
words, $\widetilde{\beta}$ is attenuated relative to the true $\beta.$
Finally, we notice that there is a minimum value of $\widetilde{\beta}$ that
one can observe under the restriction $\lambda\in\left[ 0,1\right] $ (at
$\lambda=1,$ $\widetilde{\beta}=0$). Given the attenuation bias and the fact
that $\widetilde{\beta}<0$ if and only if $\beta\in\left[ \beta_{0},1\right]
,$ the smallest value of $\widetilde{\beta}$ occurs when $\lambda=0$ and
$\beta=\omega_{11}/\omega_{12},$ so $\widetilde{\beta}_{\min}=\omega
_{11}/\omega_{12}=1/\gamma_{0}$. Thus, observing $\widetilde{\beta}
<\omega_{11}/\omega_{12}$ and $\omega_{12}<0,$ or $\widetilde{\beta}
\omega_{12}/\omega_{11}>1,$ violates the identifying restriction that
$\lambda\geq0$, for only with a $\lambda<0$ can we get $g\left( \beta\right)
>1$ when $\beta<0$ and hence $\widetilde{\beta}<\beta<0$.
\subsubsection{Bounds on $\lambda$}
The bounds on $\lambda$ are obtained by finding all the values of $\lambda$
for which equation ((ref)) has a solution for $\beta$. This
equation implies
\[
\beta^{2}\left( \left( 1-\lambda\right) \omega_{12}+\widetilde{\beta
}\lambda\omega_{22}\right) -\beta\left( \left( 1-\lambda\right)
\omega_{11}+\widetilde{\beta}\left( 1+\lambda\right) \omega_{12}\right)
+\widetilde{\beta}\omega_{11}=0,
\]
whose discriminant is the following quadratic function of lambda:
\begin{align*}
D\left( \lambda\right) & =\left( \left( 1-\lambda\right) \omega
_{11}+\widetilde{\beta}\left( 1+\lambda\right) \omega_{12}\right)
^{2}-4\widetilde{\beta}\omega_{11}\left( \left( 1-\lambda\right)
\omega_{12}+\widetilde{\beta}\lambda\omega_{22}\right).
\end{align*}
Hence, the identified set for $\lambda$ corresponds to $S_{\lambda}=\left\{
\lambda:D\left( \lambda\right) \geq0\right\} $. This set is non-empty
because $D\left( 0\right) \geq0.$ It can be computed analytically and can
take the following three shapes: (i) $S_{\lambda}=\Re$ if $D\left(
\lambda\right) \allowbreak\geq\allowbreak0$ for all $\lambda\in\Re$; (ii)
$S_{\lambda}\allowbreak=\allowbreak(-\infty,\lambda_{lo}]\allowbreak
\cup\allowbreak\lbrack\lambda_{up},\infty)$ if $\omega_{11}\allowbreak
-\allowbreak\widetilde{\beta}\omega_{12}\allowbreak\neq\allowbreak0,$ where
$\lambda_{lo}<\lambda_{up}$ are the roots of $D\left( \lambda\right) =0$;
and (iii) $S_{\lambda}\allowbreak=\allowbreak(-\infty,\lambda_{lo}]$\ if
$\omega_{11}\allowbreak-\allowbreak\widetilde{\beta}\omega_{12}\allowbreak
=\allowbreak0,$ because $\omega_{11}^{2}\allowbreak-\allowbreak\left(
\frac{\omega_{11}}{\omega_{12}}\right) ^{2}\allowbreak\omega_{12}
^{2}\allowbreak-\allowbreak2\left( \frac{\omega_{11}}{\omega_{12}}\right)
\allowbreak\omega_{11}\omega_{12}\allowbreak+\allowbreak2\left( \frac
{\omega_{11}}{\omega_{12}}\right) ^{2}\allowbreak\omega_{11}\omega
_{22}\allowbreak=\allowbreak2\omega_{11}^{2}\allowbreak\frac{\omega_{11}
\omega_{22}-\omega_{12}^{2}}{\omega_{12}^{2}}\allowbreak>\allowbreak0$. If we
also impose the restriction $\lambda\in\left[ 0,1\right] ,$ then the
identified set is $S_{\lambda}\cap\left[ 0,1\right] $.
\subsection{Proof of Proposition (ref)}
Define $\overline{A}_{i2}:=A_{i2}^{\ast}+A_{i2},$ $i=1,2$ as the right blocks
of $\overline{A}$ that was defined in ((ref)). Applying
GourierouxLaffontMonfort1980, coherency holds if and only if
$\det\overline{A}$ and $\det A^{\ast}$ have the same sign. Without loss of
generality, we can assume that $A_{11}$ is nonsingular (this can always be
achieved by reordering the variables in $Y_{t}$). From ((ref)
), we have $\det A^{\ast}=\det A_{11}\det\left( A_{22}^{\ast}-A_{21}
A_{11}^{-1}A_{12}^{\ast}\right) $ and $\det\overline{A}=\det A_{11}
\det\left( \overline{A}_{22}-A_{21}A_{11}^{-1}\overline{A}_{12}\right) $
Lutkepohl96. The coherency condition can be written as
$\det\overline{A}/\det A^{\ast}>0,$ which, given that $\left( \overline
{A}_{22}-A_{21}A_{11}^{-1}\overline{A}_{12}\right) $ and $\left(
A_{22}^{\ast}-A_{21}A_{11}^{-1}A_{12}^{\ast}\right) $ are scalars, yields
((ref)).
\subsection{Proof of Proposition (ref)}
Define $\overline{A}_{i2}:=A_{i2}^{\ast}+A_{i2},$ $i=1,2$ as the right blocks
of $\overline{A}$ that was defined in ((ref)). Also let
$Y_{t}^{\ast}:=\left( Y_{1t}^{\prime},Y_{2t}^{\ast}\right) ^{\prime}.$ When
the coherency condition ((ref)) holds, the solution of
((ref)) exists and is unique. It can be expressed as
\begin{equation}
Y_{t}^{\ast}=\left\{
\begin{array}
[c]{ll}
CX_{t}+C^{\ast}X_{t}^{\ast}+u_{t}, & if D_{t}=0\\
\widetilde{C}X_{t}+\widetilde{C}^{\ast}X_{t}^{\ast}+\widetilde{c}
b+\widetilde{u}_{t}, & if D_{t}=1
\end{array}
\right.
\end{equation}
where
\begin{equation}
C=\overline{A}^{-1}B,\quad C^{\ast}=\overline{A}^{-1}B^{\ast},\quad
u_{t}=\overline{A}^{-1}\varepsilon_{t}
\end{equation}
and
\begin{equation}
\widetilde{C}=A^{\ast-1}B,\quad\widetilde{C}^{\ast}=A^{\ast-1}B^{\ast}
,\quad\widetilde{c}=-A^{\ast-1}\binom{A_{12}}{A_{22}}b,\quad\widetilde{u}
_{t}=A^{\ast-1}\varepsilon_{t}.
\end{equation}
Using the partitioned inverse formula, we obtain
\begin{align*}
\widetilde{C}_{1} & =\left( A_{11}-A_{12}^{\ast}A_{22}^{\ast-1}
A_{21}\right) ^{-1}\left( B_{1}-A_{12}^{\ast}A_{22}^{\ast-1}B_{2}\right) \\
\widetilde{C}_{2} & =\left( A_{22}^{\ast}-A_{21}A_{11}^{-1}A_{12}^{\ast
}\right) ^{-1}\left( B_{2}-A_{21}A_{11}^{-1}B_{1}\right)
\end{align*}
and
\begin{align*}
C_{1} & =\left( A_{11}-\overline{A}_{12}\overline{A}_{22}^{-1}
A_{21}\right) ^{-1}\left( B_{1}-\overline{A}_{12}\overline{A}_{22}^{-1}
B_{2}\right) \\
C_{2} & =\left( \overline{A}_{22}-A_{21}A_{11}^{-1}\overline{A}
_{12}\right) ^{-1}\left( B_{2}-A_{21}A_{11}^{-1}B_{1}\right) .
\end{align*}
Solving the latter for $B_{1}$ and $B_{2}$ yields
\[
B_{1}=A_{11}C_{1}+\overline{A}_{12}C_{2},\text{ and }B_{2}=\overline{A}
_{22}C_{2}+A_{21}C_{1}\text{.}
\]
Thus,
\begin{align*}
\widetilde{C}_{1} & =C_{1}+\left( A_{11}-A_{12}^{\ast}A_{22}^{\ast-1}
A_{21}\right) ^{-1}\left( \overline{A}_{12}-A_{12}^{\ast}A_{22}^{\ast
-1}\overline{A}_{22}\right) C_{2}=C_{1}-\widetilde{\beta}C_{2}, and\\
\widetilde{C}_{2} & =\left( A_{22}^{\ast}-A_{21}A_{11}^{-1}A_{12}^{\ast
}\right) ^{-1}\left( A_{22}^{\ast}-A_{21}A_{11}^{-1}A_{12}^{\ast}
+A_{22}-A_{21}A_{11}^{-1}A_{12}\right) C_{2}=\kappa C_{2},
\end{align*}
where $\kappa$ is given in ((ref)). The exact same derivations
apply to $\widetilde{C}^{\ast},$ i.e.,
\[
\widetilde{C}_{1}^{\ast}=C_{1}^{\ast}-\widetilde{\beta}C_{2}^{\ast},\text{
\ and \ }\widetilde{C}_{2}^{\ast}=\kappa C_{2}^{\ast}.
\]
Next,
\begin{align*}
\widetilde{c}_{1} & =\left( A_{11}-A_{12}^{\ast}A_{22}^{\ast-1}
A_{21}\right) ^{-1}\left( A_{12}^{\ast}A_{22}^{\ast-1}A_{22}-A_{12}\right)
b=\widetilde{\beta}b, and\\
\widetilde{c}_{2} & =-\frac{A_{22}-A_{21}A_{11}^{-1}A_{12}}{A_{22}^{\ast
}-A_{21}A_{11}^{-1}A_{12}^{\ast}}b=\left( 1-\kappa\right) b.
\end{align*}
Finally, $\widetilde{u}_{t}=A^{\ast-1}\overline{A}u_{t}=\left( \widetilde{u}
_{1t}^{\prime},\widetilde{u}_{2t}\right) ^{\prime}$, where
\[
\widetilde{u}_{1t}=u_{1t}-\widetilde{\beta}u_{2t},\text{ and \ }
\widetilde{u}_{2t}=\kappa u_{2t}.
\]
Substituting back into ((ref)), the reduced-form model\ for $Y_{1t}$
becomes
\begin{align}
Y_{1t} & =\left( 1-D_{t}\right) \left( C_{1}X_{t}+C_{1}^{\ast}X_{t}
^{\ast}+u_{1t}\right) \nonumber\\
& +D_{t}\left( \left( C_{1}-\widetilde{\beta}C_{2}\right) X_{t}+\left(
C_{1}^{\ast}-\widetilde{\beta}C_{2}^{\ast}\right) X_{t}^{\ast}+u_{1t}
-\widetilde{\beta}u_{2t}\right) ,
\end{align}
and for $Y_{2t}^{\ast}$ it is
\begin{equation}
Y_{2t}^{\ast}=C_{2}X_{t}+C_{2}^{\ast}X_{t}^{\ast}+u_{2t}-\left(
1-\kappa\right) D_{t}\left( C_{2}X_{t}+C_{2}^{\ast}X_{t}^{\ast}
+u_{2t}-b\right) .
\end{equation}
Next, define
\begin{equation}
\widetilde{Y}_{2t}^{\ast}:=C_{2}X_{t}+C_{2}^{\ast}X_{t}^{\ast}+u_{2t},
\end{equation}
and rewrite ((ref)) as
\begin{align}
Y_{2t}^{\ast} & =\widetilde{Y}_{2t}^{\ast}-\left( 1-\kappa\right)
D_{t}\left( \widetilde{Y}_{2t}^{\ast}-b\right) \nonumber\\
& =\left( 1-D_{t}\right) \widetilde{Y}_{2t}^{\ast}+D_{t}\left(
\kappa\widetilde{Y}_{2t}^{\ast}+\left( 1-\kappa\right) b\right) .
\end{align}
Let $q=\dim X_{t}$ denote the number of elements of $X_{t}$ and define, for
each $i=1,2,$
\begin{equation}
\overline{C}_{ij}=\left\{
\begin{array}
[c]{ll}
C_{ij}, & j\in\left\{ 1,q\right\} :X_{tj}\neq Y_{2,t-s} for all
s\in\left\{ 1,p\right\} \\
C_{ij}+C_{is}^{\ast}, & j\in\left\{ 1,q\right\} :X_{tj}=Y_{2,t-s},\text{ for
some }s\in\left\{ 1,p\right\} .
\end{array}
\right.
\end{equation}
In other words, $\overline{C}$ contains the original coefficients on all the
regressors other than the lags of $Y_{2t},$ while the coefficients on the lags
of $Y_{2t}$ are augmented by the corresponding coefficients of the lags of
$Y_{2t}^{\ast}.$ For example, if $p=1$ and there are no other exogenous
regressors $X_{0t},$ then, for $i=1,2,$
\[
C_{i}X_{t}+C_{i}^{\ast}X_{t}^{\ast}=C_{i1}Y_{1t-1}+C_{i2}Y_{2t-1}+C_{i}^{\ast
}Y_{2t-1}^{\ast},
\]
so $\overline{C}_{i}=\left( C_{i1},C_{i2}+C_{i}^{\ast}\right) $. Using
((ref)), we can rewrite ((ref)) as
\begin{equation}
\widetilde{Y}_{2t}^{\ast}=\overline{C}_{2}X_{t}+C_{2}^{\ast}\min\left(
X_{t}^{\ast}-b,0\right) +u_{2t}.
\end{equation}
Now, observe that
\[
\min\left( Y_{2t}^{\ast}-b,0\right) =D_{t}\left( Y_{2t}^{\ast}-b\right)
=\kappa D_{t}\left( \widetilde{Y}_{2t}^{\ast}-b\right) =\kappa\min\left(
\widetilde{Y}_{2t}^{\ast}-b,0\right)
\]
So, letting $\widetilde{X}_{t}^{\ast}$ denote the lags of $\widetilde{Y}
_{2t}^{\ast},$ we have $\min\left( X_{t}^{\ast}-b,0\right) =\kappa
\min\left( \widetilde{X}_{t}^{\ast}-b,0\right) ,$ and consequently,
\[
C^{\ast}\min\left( X_{t}^{\ast}-b,0\right) =\overline{C}^{\ast}\min\left(
\widetilde{X}_{t}^{\ast}-b,0\right) ,
\]
where $\overline{C}^{\ast}=\kappa C^{\ast}.$ Now, from ((ref)) we
have
\[
\widetilde{Y}_{2t}^{\ast}=\overline{C}_{2}X_{t}+\overline{C}_{2}^{\ast}
\min\left( \widetilde{X}_{t}^{\ast}-b,0\right) +u_{2t}.
\]
Recall the definition of $\overline{Y}_{2t}^{\ast}$ in ((ref)):
\[
\overline{Y}_{2t}^{\ast}:=\overline{C}_{2}X_{t}+\overline{C}_{2}^{\ast
}\overline{X}_{t}^{\ast}+u_{2t},
\]
where $\overline{X}_{t}^{\ast}:=\left( \overline{x}_{t-1},...,\overline
{x}_{t-p}\right) ^{\prime},$ and $\overline{x}_{t}\allowbreak:=\allowbreak
\min\left( \overline{Y}_{2t}^{\ast}\allowbreak-b,0\right) ,$ with initial
conditions $\overline{x}_{-s}\allowbreak=\kappa^{-1}\allowbreak\min\left(
Y_{2,-s}^{\ast}\allowbreak-b,0\right) ,$ $s=0,...,\allowbreak p-1.$ It
follows that $\min\left( \widetilde{X}_{t}^{\ast}-b,0\right) \allowbreak
=\overline{X}_{t}^{\ast}$ for all $t\geq1,$ so that $\widetilde{Y}_{2t}^{\ast
}=\overline{Y}_{2t}^{\ast}.$ Substituting $\overline{Y}_{2}^{\ast}$ for
$\widetilde{Y}_{2}^{\ast}$ in ((ref)), we get ((ref)
). Using the reparametrization ((ref)) and the relationship between
$X_{t}^{\ast}$ and $\overline{X}_{t}^{\ast}$ in ((ref)), we obtain
((ref)).
Finally, from eq. ((ref)), it follows that the event
$Y_{2t}^{\ast}<b$ is equivalent to
\[
b+\kappa\left( C_{2}X_{t}+C_{2}^{\ast}X_{t}^{\ast}+u_{2t}-b\right) <b,
\]
which, since $\kappa>0$ by the coherency condition ((ref)), is
equivalent to
\begin{equation}
u_{2t}<b-C_{2}X_{t}-C_{2}^{\ast}X_{t}^{\ast}.
\end{equation}
Using the definition ((ref)), and ((ref)), the
inequality ((ref)) can be written as $\overline{Y}
_{2t}^{\ast}<b,$ which establishes ((ref)).
\textbf{Comment:} Note that $\kappa$ appears in the reduced form only
multiplicatively with $C^{\ast},$ so $\kappa$ and $C^{\ast}$ are not
separately identified, only $\overline{C}^{\ast}=\kappa C^{\ast}$ is. The
reparametrization from $C$ to $\overline{C}$ is convenient because
$\overline{C}$ is identified independently of $\kappa,$ while $C,C^{\ast}$ and
$\kappa$ are not separately identified.
\subsection{Proof of Proposition (ref)}
We solve $u_{t}=\overline{A}^{-1}\varepsilon_{t}$ using the partitioned
inverse formula to get
\begin{align}
u_{1t} & =\left( A_{11}-\overline{A}_{12}\overline{A}_{22}^{-1}
A_{21}\right) ^{-1}\left( \varepsilon_{1t}-\overline{A}_{12}\overline
{A}_{22}^{-1}\varepsilon_{2t}\right) \\
u_{2t} & =\left( \overline{A}_{22}-A_{21}A_{11}^{-1}\overline{A}
_{12}\right) ^{-1}\left( \varepsilon_{2t}-A_{21}A_{11}^{-1}\varepsilon
_{1t}\right) .
\end{align}
Using the definitions
\begin{align*}
\overline{\beta} & :=-A_{11}^{-1}\overline{A}_{12},\quad\overline{\gamma
}:=-\overline{A}_{22}^{-1}A_{21},\\
\bar{\varepsilon}_{1t} & :=A_{11}^{-1}\varepsilon_{1t},\quad\bar
{\varepsilon}_{2t}:=\overline{A}_{22}^{-1}\varepsilon_{2t},
\end{align*}
we can rewrite ((ref))-((ref)) as ((ref)
)-((ref)).
Note that
\begin{align*}
\bar{\varepsilon}_{1t} & =A_{11}^{-1}\left( A_{11}u_{1t}+\overline{A}
_{12}u_{2t}\right) =u_{1t}-\bar{\beta}u_{2t},\\
\bar{\varepsilon}_{2t} & =\overline{A}_{22}^{-1}\left( A_{21}
u_{1t}+\overline{A}_{22}u_{2t}\right) =-\overline{\gamma}u_{1t}+u_{2t},
\end{align*}
so,
\begin{align*}
var\left( \bar{\varepsilon}_{1t}\right) & =\left( I_{k-1},-\bar{\beta
}\right) \Omega\left( I_{k-1},-\bar{\beta}\right) ^{\prime},\\
var\left( \bar{\varepsilon}_{2t}\right) & =\left( -\bar{\gamma},1\right)
\Omega\left( -\bar{\gamma},1\right) ^{\prime},
\end{align*}
and
\begin{align*}
cov\left( \bar{\varepsilon}_{1t},\bar{\varepsilon}_{2t}\right) & =\left(
I_{k-1},-\overline{\beta}\right)
\begin{pmatrix}
\Omega_{11} & \Omega_{12}\\
\Omega_{12}^{\prime} & \Omega_{22}
\end{pmatrix}
\left( -\overline{\gamma},1\right) ^{\prime}\\
& =-\left( \Omega_{11}-\overline{\beta}\Omega_{12}^{\prime}\right)
\overline{\gamma}^{\prime}+\Omega_{12}-\overline{\beta}\Omega_{22}=0.
\end{align*}
The last equation identifies $\overline{\gamma}$ given $\overline{\beta}$.
Specifically,
\[
\overline{\gamma}^{\prime}=\left( \Omega_{11}-\overline{\beta}\Omega
_{12}^{\prime}\right) ^{-1}\left( \Omega_{12}-\overline{\beta}\Omega
_{22}\right) .
\]
\section{Numerical results}
This section provides Monte-Carlo evidence on the finite-sample properties of
the proposed estimators and tests. The data generating process (DGP) is a
trivariate VAR(1), given by equations ((ref)) and
((ref)). I\ consider three different DGPs corresponding to the
CKSVAR, KSVAR and CSVAR models, respectively. In all three DGPs, the following
parameters are set to the same values: the contemporaneous coefficients are
$A_{11}=I_{2},$ $A_{12}=$ $A_{12}^{\ast}=0_{2\times1},$ $A_{22}^{\ast}=1$ and
$A_{22}=0;$ the intercepts are set to zero, $B_{10}=0_{2\times1}$ and
$B_{20}=0;$ the coefficients on the lags are $B_{1,1}=\left( \rho
I_{2},0\right) ,$ $B_{1,1}^{\ast}=0_{2\times1},$ $B_{2,1}=\left(
0_{1\times2},B_{22,1}\right) $, with $\rho=0.5.$ Finally, each of the three
DGPs is determined as follows. DGP1:\ $B_{22,1}=B_{2,1}^{\ast}=0$ (both KSVAR
and CSVAR restrictions hold, since lags of $Y_{2,t}$ and $Y_{2,t}^{\ast}$ all
have zero coefficients); DGP2: $B_{22,1}=\rho,\,B_{2,1}^{\ast}=0\ $(KSVAR
restrictions hold but CSVAR restrictions do not); DGP3: $B_{22,1}
=0,\,B_{2,1}^{\ast}=\rho$ (CSVAR restrictions hold but KSVAR restrictions do
not). The setting of the autoregressive coefficient $\rho=0.5$ leads to a
lower degree of persistence than is typically observed in macro data (e.g., in
the Stock and Watson, 2001, application, the three largest roots are 0.97,
0.97 and 0.8), because I want to avoid confounding any possible finite-sample
issues arising from the ZLB with well-known problems of bias and size
distortion due to strong persistence (near unit roots) in the data. Finally,
the bound on $Y_{2t}$ is set to $b=0,$ the sample size is $T=250,$ the initial
conditions are set to 0 and the number of Monte Carlo replications is 1000. In
all cases, the CKSVAR and CSVAR likelihoods are computed using SIS with
$R=1000$ particles. The notation for the reported parameters is given in Table
(ref).
\begin{table}[tbp]
\caption{Parameter notation in reported simulation results}
\begin{tabular}
[c]{l|l}\hline
\textbf{Mnemonic} & \textbf{Description}\\\hline
$\tau$ & st. dev. of reduced form error $u_{2t}$ in $Y_{2t}$ (constrained
variable)\\
Eq. 3 & reduced form equation for $Y_{2t}$\\
Eq. j & red. form equation for $Y_{1j,t},$ $j=1,2$ (unconstrained variables)\\
$\widetilde{\beta}_{j}$ & coefficient on kink in eq. j\\
eq. i Y1j_1 & coefficient of $Y_{1j,t-1}$ in eq. i\\
eq. i Y2_1 & coefficient of $Y_{2,t-1}$ in eq. i\\
eq. i lY2_1 & coefficient of $\min\left( Y_{2,t-1}^{\ast}-b,0\right) $ in
eq. i\\
$\delta_{j}$ & coefficient of regression of $u_{1j,t}$ (red. form error in Eq.
j) on $u_{2t}$\\
Ch_ij & (i,j) element of Choleski factor of $\Omega_{1.2}$\\\hline
\end{tabular}
\end{table}
Figure (ref) reports the sampling distribution of the ML estimators of the reduced-form
parameters in Proposition (ref) for the CKSVAR model under DGP1. The results for the KSVAR and CSVAR
models, which are also correctly specified under DGP1, are omitted because they are entirely analogous.
The sampling densities appear to be very close to the superimposed Normal
approximations, indicating that the Normal asymptotic approximation is fairly accurate.
\begin{figure}[htb]
\caption{Sampling densities of ML estimators of reduced-form coefficients of CKSVAR(1) under
DGP1 (solid lines) and approximating Normal densities (dashed lines). $T=250$, 1000 Monte Carlo replications. Parameter names described in
Table (ref).}
\end{figure}
Table (ref) reports
moments of the sampling distributions of the above mentioned estimators for the CKSVAR model. Again, the results for the KSVAR and CSVAR models are entirely analogous and are therefore omitted. We
notice no discernible biases. Additional simulation results with $T=100$ and $T=1000$ given in the Tables (ref), (ref) and (ref)
indicate that the RMSE declines at rate $\sqrt{T}$ in accordance with
asymptotic theory. It is noteworthy that the
estimators of $\widetilde{\beta}$ have substantially
larger RMSE than the estimators of the other parameters.
\begin{table}[tbp]
\caption{Moments of sampling distribution of ML estimators of the parameters of CKSVAR(1)\tabnoteref[a]{tab2}}
\begin{tabular}
[c]{r|r|r|r|r|r}\hline
ML-CKSVAR & true & mean & bias & sd & RMSE\\\hline
$\tau$ & 1.000 & 0.983 & -0.017 & 0.068 & 0.070\\
Eq.3 Constant & 0.000 & -0.006 & -0.006 & 0.175 & 0.176\\
Eq.3 Y11_1 & 0.000 & 0.000 & 0.000 & 0.061 & 0.061\\
Eq.3 Y12_1 & 0.000 & -0.000 & -0.000 & 0.064 & 0.064\\
Eq.3 Y2_1 & 0.000 & -0.010 & -0.010 & 0.173 & 0.173\\
Eq.3 lY2_1 & 0.000 & -0.013 & -0.013 & 0.258 & 0.258\\
$\tilde{\beta}_{1}$ & -0.000 & -0.008 & -0.008 & 0.356 & 0.356\\
$\tilde{\beta}_{2}$ & 0.000 & -0.004 & -0.004 & 0.359 & 0.359\\
Eq.1 Constant & 0.000 & 0.002 & 0.002 & 0.213 & 0.213\\
Eq.1 Y11_1 & 0.500 & 0.489 & -0.011 & 0.057 & 0.058\\
Eq.1 Y12_1 & 0.000 & 0.002 & 0.002 & 0.060 & 0.060\\
Eq.1 Y2_1 & 0.000 & -0.002 & -0.002 & 0.163 & 0.163\\
Eq.2 Constant & 0.000 & 0.005 & 0.005 & 0.214 & 0.214\\
Eq.2 Y11_1 & 0.000 & 0.000 & 0.000 & 0.058 & 0.058\\
Eq.2 Y12_1 & 0.500 & 0.491 & -0.009 & 0.057 & 0.058\\
Eq.2 Y2_1 & 0.000 & -0.001 & -0.001 & 0.159 & 0.159\\
Eq.1 lY2_1 & 0.000 & 0.005 & 0.005 & 0.232 & 0.232\\
Eq.2 lY2_1 & 0.000 & 0.002 & 0.002 & 0.233 & 0.233\\
$\delta_{1}$ & 0.000 & -0.001 & -0.001 & 0.157 & 0.157\\
$\delta_{2}$ & 0.000 & -0.004 & -0.004 & 0.155 & 0.155\\
Ch_11 & 1.000 & 0.975 & -0.025 & 0.044 & 0.051\\
Ch_21 & 0.000 & -0.001 & -0.001 & 0.066 & 0.066\\
Ch_22 & 1.000 & 0.972 & -0.028 & 0.046 & 0.054\\\hline
\end{tabular}
\tabnotetext[a]{tab2}{Computed under DGP1 with $T=250$ using 1000 MC replications. Parameter names described in Table (ref).}
\end{table}
\begin{table}[htb]
\caption{Bias, standard deviation and Root Mean Square Error of Maximum Likelihood estimator of parameters of CKSVAR(1) model\tabnoteref[a]{tab4}}
\begin{tabular}
[c]{r|r|r|r|r|rr|r|r|r}\hline
ML-CKSVAR & \multicolumn{3}{|c|}{$T=100$} & \multicolumn{3}{|c|}{$T=250$} &
\multicolumn{3}{c}{$T=1000$}\\\hline
Parameter & bias & sd & RMSE & bias & sd & RMSE & bias & sd & RMSE\\\hline
$\tau$ & \multicolumn{1}{|r|}{-0.048} & 0.111 & 0.121 & -0.017 & 0.068 &
\multicolumn{1}{|r|}{0.070} & -0.003 & 0.035 & 0.035\\
Eq.3 Constant & \multicolumn{1}{|r|}{0.001} & 0.293 & 0.293 & -0.006 & 0.175 &
\multicolumn{1}{|r|}{0.176} & 0.004 & 0.085 & 0.085\\
Eq.3 Y11_1 & \multicolumn{1}{|r|}{-0.001} & 0.111 & 0.111 & 0.000 & 0.061 &
\multicolumn{1}{|r|}{0.061} & 0.000 & 0.032 & 0.032\\
Eq.3 Y12_1 & \multicolumn{1}{|r|}{-0.004} & 0.109 & 0.109 & -0.000 & 0.064 &
\multicolumn{1}{|r|}{0.064} & -0.000 & 0.031 & 0.031\\
Eq.3 Y2_1 & \multicolumn{1}{|r|}{-0.031} & 0.289 & 0.291 & -0.010 & 0.173 &
\multicolumn{1}{|r|}{0.173} & -0.003 & 0.082 & 0.082\\
Eq.3 lY2_1 & \multicolumn{1}{|r|}{-0.022} & 0.435 & 0.436 & -0.013 & 0.258 &
\multicolumn{1}{|r|}{0.258} & 0.003 & 0.122 & 0.122\\
$\tilde{\beta}_{1}$ & \multicolumn{1}{|r|}{0.012} & 0.586 & 0.586 & -0.008 &
0.356 & \multicolumn{1}{|r|}{0.356} & 0.000 & 0.176 & 0.176\\
$\tilde{\beta}_{2}$ & \multicolumn{1}{|r|}{-0.011} & 0.606 & 0.606 & -0.004 &
0.359 & \multicolumn{1}{|r|}{0.359} & -0.003 & 0.168 & 0.168\\
Eq.1 Constant & \multicolumn{1}{|r|}{0.000} & 0.360 & 0.360 & 0.002 & 0.213 &
\multicolumn{1}{|r|}{0.213} & 0.007 & 0.105 & 0.105\\
Eq.1 Y11_1 & \multicolumn{1}{|r|}{-0.034} & 0.098 & 0.104 & -0.011 & 0.057 &
\multicolumn{1}{|r|}{0.058} & -0.002 & 0.029 & 0.029\\
Eq.1 Y12_1 & \multicolumn{1}{|r|}{-0.004} & 0.108 & 0.108 & 0.002 & 0.060 &
\multicolumn{1}{|r|}{0.060} & 0.001 & 0.028 & 0.028\\
Eq.1 Y2_1 & \multicolumn{1}{|r|}{-0.004} & 0.281 & 0.281 & -0.002 & 0.163 &
\multicolumn{1}{|r|}{0.163} & -0.006 & 0.079 & 0.079\\
Eq.2 Constant & \multicolumn{1}{|r|}{0.006} & 0.356 & 0.356 & 0.005 & 0.214 &
\multicolumn{1}{|r|}{0.214} & 0.002 & 0.102 & 0.103\\
Eq.2 Y11_1 & \multicolumn{1}{|r|}{0.004} & 0.103 & 0.103 & 0.000 & 0.058 &
\multicolumn{1}{|r|}{0.058} & 0.000 & 0.028 & 0.028\\
Eq.2 Y12_1 & \multicolumn{1}{|r|}{-0.029} & 0.103 & 0.107 & -0.009 & 0.057 &
\multicolumn{1}{|r|}{0.058} & -0.002 & 0.029 & 0.029\\
Eq.2 Y2_1 & \multicolumn{1}{|r|}{0.000} & 0.269 & 0.269 & -0.001 & 0.159 &
\multicolumn{1}{|r|}{0.159} & 0.001 & 0.078 & 0.078\\
Eq.1 lY2_1 & \multicolumn{1}{|r|}{0.009} & 0.403 & 0.403 & 0.005 & 0.232 &
\multicolumn{1}{|r|}{0.232} & 0.010 & 0.112 & 0.113\\
Eq.2 lY2_1 & \multicolumn{1}{|r|}{-0.003} & 0.408 & 0.408 & 0.002 & 0.233 &
\multicolumn{1}{|r|}{0.233} & 0.004 & 0.115 & 0.115\\
$\delta_{1}$ & \multicolumn{1}{|r|}{0.005} & 0.260 & 0.260 & -0.001 & 0.157 &
\multicolumn{1}{|r|}{0.157} & 0.001 & 0.075 & 0.075\\
$\delta_{2}$ & \multicolumn{1}{|r|}{-0.005} & 0.260 & 0.260 & -0.004 & 0.155 &
\multicolumn{1}{|r|}{0.155} & -0.001 & 0.073 & 0.073\\
Ch_11 & \multicolumn{1}{|r|}{-0.067} & 0.073 & 0.099 & -0.025 & 0.044 &
\multicolumn{1}{|r|}{0.051} & -0.006 & 0.024 & 0.024\\
Ch_21 & \multicolumn{1}{|r|}{-0.005} & 0.111 & 0.111 & -0.001 & 0.066 &
\multicolumn{1}{|r|}{0.066} & -0.000 & 0.031 & 0.031\\
Ch_22 & \multicolumn{1}{|r|}{-0.075} & 0.073 & 0.105 & -0.028 & 0.046 &
\multicolumn{1}{|r|}{0.054} & -0.008 & 0.023 & 0.024\\\hline
\end{tabular}
\tabnotetext[a]{tab4}{Computed under DGP1 with $R=1000$ particles using 1000 MC replications. Parameter names described in Table (ref).}
\end{table}
\begin{table}[htb]
\caption{Bias, standard deviation and Root Mean Square Error of Maximum Likelihood estimator of parameters of KSVAR(1) model\tabnoteref[a]{tab5}}
\begin{tabular}
[c]{r|r|r|r|r|r|r|r|r|r}\hline
ML-KSVAR & \multicolumn{3}{|c|}{$T=100$} & \multicolumn{3}{c|}{$T=250$} &
\multicolumn{3}{c}{$T=1000$}\\\hline
Parameter & bias & sd & RMSE & bias & sd & RMSE & bias & sd & RMSE\\\hline
$\tau$ & -0.024 & 0.111 & 0.113 & -0.008 & 0.068 & 0.069 & -0.001 & 0.035 &
0.035\\
Eq.3 Constant & 0.011 & 0.145 & 0.145 & 0.001 & 0.092 & 0.092 & 0.003 &
0.046 & 0.046\\
Eq.3 Y11_1 & -0.001 & 0.103 & 0.103 & 0.001 & 0.060 & 0.060 & -0.000 &
0.031 & 0.031\\
Eq.3 Y12_1 & -0.004 & 0.102 & 0.102 & -0.000 & 0.062 & 0.062 & -0.000 &
0.030 & 0.030\\
Eq.3 Y2_1 & -0.048 & 0.199 & 0.204 & -0.019 & 0.122 & 0.124 & -0.003 &
0.060 & 0.060\\
$\tilde{\beta}_{1}$ & -0.003 & 0.571 & 0.571 & -0.013 & 0.349 & 0.349 &
-0.001 & 0.174 & 0.174\\
$\tilde{\beta}_{2}$ & -0.003 & 0.584 & 0.584 & -0.001 & 0.348 & 0.348 &
-0.004 & 0.168 & 0.168\\
Eq.1 Constant & 0.002 & 0.264 & 0.264 & 0.001 & 0.165 & 0.165 & 0.001 &
0.080 & 0.080\\
Eq.1 Y11_1 & -0.033 & 0.093 & 0.099 & -0.012 & 0.056 & 0.057 & -0.002 &
0.028 & 0.028\\
Eq.1 Y12_1 & -0.002 & 0.100 & 0.100 & 0.002 & 0.058 & 0.058 & 0.001 & 0.027 &
0.027\\
Eq.1 Y2_1 & -0.002 & 0.197 & 0.197 & -0.000 & 0.117 & 0.117 & -0.001 &
0.057 & 0.057\\
Eq.2 Constant & 0.004 & 0.258 & 0.258 & 0.003 & 0.158 & 0.158 & 0.001 &
0.078 & 0.078\\
Eq.2 Y11_1 & 0.006 & 0.096 & 0.096 & 0.001 & 0.057 & 0.057 & 0.000 & 0.027 &
0.027\\
Eq.2 Y12_1 & -0.028 & 0.094 & 0.098 & -0.008 & 0.055 & 0.056 & -0.002 &
0.028 & 0.029\\
Eq.2 Y2_1 & -0.001 & 0.189 & 0.189 & -0.000 & 0.113 & 0.113 & 0.003 & 0.054 &
0.055\\
$\delta_{1}$ & -0.000 & 0.252 & 0.252 & -0.003 & 0.156 & 0.156 & 0.001 &
0.075 & 0.075\\
$\delta_{2}$ & -0.003 & 0.253 & 0.253 & -0.003 & 0.152 & 0.152 & -0.001 &
0.073 & 0.073\\
Ch_11 & -0.048 & 0.070 & 0.085 & -0.018 & 0.044 & 0.047 & -0.005 & 0.023 &
0.024\\
Ch_21 & -0.003 & 0.108 & 0.108 & -0.000 & 0.065 & 0.065 & -0.000 & 0.031 &
0.031\\
Ch_22 & -0.054 & 0.070 & 0.088 & -0.020 & 0.045 & 0.050 & -0.007 & 0.023 &
0.024\\\hline
\end{tabular}
\tabnotetext[a]{tab5}{Computed under DGP1 with $R=1000$ particles using 1000 MC replications. Parameter names described in Table (ref).}
\end{table}
\begin{table}[ptb]
\caption{Bias, standard deviation and Root Mean Square Error of Maximum Likelihood estimator of parameters of CSVAR(1) model\tabnoteref[a]{tab6}}
\begin{tabular}
[c]{r|r|r|r|r|r|r|r|r|r}\hline
ML-CSVAR & \multicolumn{3}{|c}{$T=100$} & \multicolumn{3}{c|}{$T=250$} &
\multicolumn{3}{c}{$T=1000$}\\\hline
Parameter & bias & sd & RMSE & bias & sd & RMSE & bias & sd & RMSE\\\hline
$\tau$ & -0.026 & 0.111 & 0.114 & -0.008 & 0.068 & 0.069 & -0.001 & 0.035 &
0.035\\
Eq.3 Constant & -0.009 & 0.135 & 0.135 & -0.006 & 0.081 & 0.081 & 0.001 &
0.040 & 0.040\\
Eq.3 Y11_1 & -0.001 & 0.104 & 0.104 & 0.001 & 0.060 & 0.060 & -0.000 &
0.031 & 0.031\\
Eq.3 Y12_1 & -0.004 & 0.103 & 0.103 & -0.000 & 0.061 & 0.061 & -0.000 &
0.030 & 0.030\\
Eq.3 Y2_1 & -0.027 & 0.125 & 0.128 & -0.011 & 0.078 & 0.079 & -0.001 &
0.038 & 0.038\\
Eq.1 Constant & -0.002 & 0.109 & 0.110 & -0.003 & 0.065 & 0.065 & 0.000 &
0.031 & 0.031\\
Eq.1 Y11_1 & -0.033 & 0.090 & 0.096 & -0.012 & 0.054 & 0.056 & -0.002 &
0.028 & 0.028\\
Eq.1 Y12_1 & -0.002 & 0.096 & 0.096 & 0.002 & 0.057 & 0.057 & 0.001 & 0.027 &
0.027\\
Eq.1 Y2_1 & -0.000 & 0.118 & 0.118 & 0.001 & 0.072 & 0.072 & 0.000 & 0.036 &
0.036\\
Eq.2 Constant & 0.003 & 0.106 & 0.106 & 0.002 & 0.063 & 0.063 & -0.000 &
0.031 & 0.031\\
Eq.2 Y11_1 & 0.005 & 0.091 & 0.092 & 0.001 & 0.056 & 0.056 & 0.000 & 0.027 &
0.027\\
Eq.2 Y12_1 & -0.026 & 0.090 & 0.094 & -0.008 & 0.054 & 0.055 & -0.002 &
0.028 & 0.028\\
Eq.2 Y2_1 & -0.001 & 0.115 & 0.115 & 0.000 & 0.070 & 0.070 & 0.002 & 0.035 &
0.035\\
$\delta_{1}$ & 0.001 & 0.114 & 0.114 & 0.002 & 0.071 & 0.071 & 0.001 & 0.035 &
0.035\\
$\delta_{2}$ & -0.001 & 0.116 & 0.116 & -0.002 & 0.070 & 0.070 & -0.000 &
0.034 & 0.034\\
Ch_11 & -0.033 & 0.069 & 0.077 & -0.012 & 0.043 & 0.044 & -0.003 & 0.023 &
0.023\\
Ch_21 & -0.005 & 0.103 & 0.104 & -0.001 & 0.064 & 0.064 & -0.000 & 0.031 &
0.031\\
Ch_22 & -0.037 & 0.068 & 0.077 & -0.014 & 0.044 & 0.046 & -0.005 & 0.023 &
0.024\\\hline
\end{tabular}
\tabnotetext[a]{tab6}{Computed under DGP1 with $R=1000$ particles using 1000 MC replications. Parameter names described in Table (ref).}
\end{table}
Next, I turn to the properties of the LR test of KSVAR against CKSVAR and
CSVAR against CKSVAR. The former hypothesis involves three restrictions
(exclusion of the latent lag $Y_{2,t-1}^{\ast}$ from each of the three
equations), so the LR statistic is asymptotically distributed as $\chi_{3}
^{2}$ under the null. The latter hypothesis involves five restrictions
(exclusion of the observed lag $Y_{2,t-1}$ from each of the three equations,
plus $\widetilde{\beta}=0$), and the LR statistic is asymptotically
distributed as $\chi_{5}^{2}.$ Table (ref) reports the rejection
frequencies of the LR tests for each of the two hypotheses in each of the
three DGPs at three significance levels: 10%, 5% and 1%. In addition to the
asymptotic tests, I\ also report the rejection frequency of the tests using
parametric bootstrap critical values. The parametric bootstrap uses draws of
Normal errors and the estimated reduced-form parameters to generate the
bootstrap samples of $Y_{t}$ and $Y_{2t}^{\ast}$. The Monte Carlo rejection
frequencies are computed using the “warp-speed” method method of
GiacominiPolitisWhite2013. Note that both null hypotheses hold under
DGP1, but only the KSVAR is valid under DGP2 and only the CSVAR is valid under
DGP3. For convenience, I indicate the rejection frequencies under the
alternative in bold in the table.
\begin{table}[htbp]
\caption{Rejection frequencies of LR tests of $H_0$ against $H_1$ across different DGPs\tabnoteref[a]{tab3}}
\begin{tabular}
[c]{ll|c|c|c|c|c|c}\hline
& & \multicolumn{3}{|c|}{$H_{0}:$ KSVAR, $H_{1}:$ CKSVAR} &
\multicolumn{3}{|l}{$H_{0}:$ CSVAR, $H_{1}:$ CKSVAR}\\\hline
& Sign. Level & 10% & 5% & 1% & 10% & 5% & 1%\\\hline
DGP1 & asymptotic & 0.173 & 0.093 & 0.020 & 0.155 & 0.084 & 0.022\\
& bootstrap & 0.107 & 0.045 & 0.005 & 0.109 & 0.052 & 0.012\\\hline
DGP2 & asymptotic & 0.149 & 0.080 & 0.018 & \textbf{0.319} & \textbf{0.206} &
\textbf{0.073}\\
& bootstrap & 0.117 & 0.050 & 0.011 & \textbf{0.242} & \textbf{0.141} &
\textbf{0.041}\\\hline
DGP3 & asymptotic & \textbf{0.587} & \textbf{0.454} & \textbf{0.244} & 0.141 &
0.067 & 0.019\\
& bootstrap & \textbf{0.471} & \textbf{0.365} & \textbf{0.194} & 0.103 &
0.053 & 0.018\\\hline
\end{tabular}
\tabnotetext[a]{tab3}{Computed using 1000 Monte Carlo replications, $T=250$. The asymptotic tests use $\chi^2_3$ and $\chi^2_5$ critical values for KSVAR and CSVAR resp. The bootstrap rej. frequencies were computed uisng the warp-speed method of GiacominiPolitisWhite2013. Bold numbers indicate that the rejecction frequencies were computed under $H_1$ (power).}
\end{table}
There is evidence that the LR tests reject too often under $H_{0}$ relative to
their nominal level when we use asymptotic critical values. Moreover, the size
distortions are very similar across null hypotheses and DGPs. Unreported
results show that size distortion eventually disappears as the sample gets
large, but this level of overrejection is unsatisfactory at $T=250,$ which is
a typical sample size in macroeconomic applications. The parametric bootstrap
appears to do a remarkably good job at correcting the size of the tests. In
all cases considered, the parametric bootstrap rejection frequency is not
significantly different from the nominal level when the null hypothesis holds
(all but the numbers in bold in the Table). To shed further light on this
issue, Figure (ref) reports the QQ plots of the
sampling distributions of the two LR statistics against their asymptotic and parametric
bootstrap approximations for all three DGPs under the null hypothesis. The sampling distributions of the LR statistics stochastically dominate their asymptotic approximations, but the bootstrap
approximations are quite accurate.
Finally, the rejection frequencies highlighted in bold in Table
(ref) correspond to the power of the tests against two very similar
deviations from the null hypothesis. The numbers on the left under DGP3 show
the power of the test to reject the KSVAR specification under the alternative
at which the coefficient on the latent lag $B_{2,1}^{\ast}=0.5.$ Similarly,
the bold numbers on the right give the power of rejecting CSVAR against the
alternative where the coefficient on the observed lag $B_{22,1}=0.5.$ Since
the lower bound is set to zero, and the sample contains about 50% of
observations at the ZLB, the two deviations from the null are of equal
magnitude. Yet, we notice the LR\ test is significantly more powerful against
the KSVAR than against the CSVAR. This could be because CSVAR\ imposes more
restrictions than the KSVAR, so one would expect it to have lower power than
the KSVAR against similar deviations from the null.
\begin{figure}[ptb]
\caption{QQ plots of the sampling distribution under the null hypothesis of LR statistics of KSVAR against
CKSVAR (left) and CSVAR against CKSVAR (right). Solid (dashed) lines plot quantiles against asymptotic $\chi^2$ (bootstrap) approximation. Computed for $T=250$ using 1000 Monte Carlo replications.}
\end{figure}
\subsection*{Alternative DGP}
The DGPs in the previous simulations have the property that the frequency of
the ZLB regime is around 50%. I reran those simulations with a slight
modification to the DGPs to match the frequency in the sample of the empirical
application in the paper. Specifically, I reduce the lower bound $b$ to a
level that makes the frequency of the ZLB regime equal to 11%. The results
are given in Figure (ref) and Table
(ref). The results are very similar to the
ones reported in Figure (ref) and
Table (ref) above: the Normal approximation of the sampling distribution of the MLE
appears to be very good, and the bias is negligible. The only difference is that the standard deviation of $\widetilde{\beta}$ is larger.
\begin{figure}[htb]
\caption{Sampling densities of ML estimators of reduced-form coefficients of CKSVAR(1) under
DGP1 with $b$ chosen such that $\Pr\left( Y_{2t}=b\right) =0.11$ (solid lines) and approximating Normal densities (dashed lines). $T=250$, 1000 Monte Carlo replications. Parameter names described in
Table (ref).}
\end{figure}
\begin{table}[ptb]
\caption{Moments of sampling distribution of ML estimators of the parameters of CKSVAR(1)\tabnoteref[a]{tab7}}
\begin{tabular}
[c]{r|r|r|r|r|r}\hline
ML-CKSVAR & true & mean & bias & sd & RMSE\\\hline
$\tau$ & 1.000 & 0.988 & -0.012 & 0.050 & 0.051\\
Eq.3 Constant & 0.000 & -0.004 & -0.004 & 0.073 & 0.073\\
Eq.3 Y11_1 & 0.000 & 0.001 & 0.001 & 0.055 & 0.055\\
Eq.3 Y12_1 & 0.000 & -0.002 & -0.002 & 0.056 & 0.056\\
Eq.3 Y2_1 & 0.000 & -0.008 & -0.008 & 0.083 & 0.083\\
Eq.3 lY2_1 & 0.000 & 0.003 & 0.003 & 0.501 & 0.501\\
$\tilde{\beta}_{1}$ & 0.000 & 0.013 & 0.013 & 0.533 & 0.533\\
$\tilde{\beta}_{2}$ & 0.000 & -0.030 & -0.030 & 0.518 & 0.519\\
Eq.1 Constant & 0.000 & -0.005 & -0.005 & 0.078 & 0.078\\
Eq.1 Y11_1 & 0.500 & 0.488 & -0.012 & 0.055 & 0.056\\
Eq.1 Y12_1 & 0.000 & 0.001 & 0.001 & 0.058 & 0.058\\
Eq.1 Y2_1 & 0.000 & 0.001 & 0.001 & 0.084 & 0.084\\
Eq.2 Constant & 0.000 & 0.005 & 0.005 & 0.075 & 0.075\\
Eq.2 Y11_1 & 0.000 & 0.001 & 0.001 & 0.056 & 0.056\\
Eq.2 Y12_1 & 0.500 & 0.491 & -0.009 & 0.056 & 0.056\\
Eq.2 Y2_1 & 0.000 & -0.000 & -0.000 & 0.080 & 0.080\\
Eq.1 lY2_1 & 0.000 & -0.014 & -0.014 & 0.497 & 0.497\\
Eq.2 lY2_1 & 0.000 & 0.018 & 0.018 & 0.491 & 0.491\\
$\delta_{1}$ & 0.000 & 0.002 & 0.002 & 0.084 & 0.084\\
$\delta_{2}$ & 0.000 & -0.004 & -0.004 & 0.080 & 0.080\\
Ch_11 & 1.000 & 0.980 & -0.020 & 0.043 & 0.047\\
Ch_21 & 0.000 & -0.002 & -0.002 & 0.065 & 0.065\\
Ch_22 & 1.000 & 0.978 & -0.022 & 0.045 & 0.050\\\hline
\end{tabular}
\tabnotetext[a]{tab7}{Computed under DGP1 with $b$ chosen such that $\Pr\left( Y_{2t}=b\right)
=0.11$, $T=250$ using 1000 MC replications. Parameter names described in Table (ref).}
\end{table}
\section{A simple model of QE}
This is a simplified version of the New Keynesian model of bond market
segmentation that appears in IkedaLiMavroeidisZanetti2020 and is based
on ChenCurdiaFerrero2012. The economy consists of two types of
households. A fraction $\omega_{r}$ of type `r' households can only trade
long-term government bonds. The remaining $1-\omega_{r}$ households of type
`u' can purchase both short-term and long-term government bonds, the latter
subject to a trading cost $\zeta_{t}$. This trading cost gives rise to a term
premium, i.e., a spread between long-term and short-term yields, that the
central bank can manipulate by purchasing long-term bonds. The term premium
affects aggregate demand through the consumption decisions of constrained
households. This generates an unconventional monetary policy channel.
The transmission mechanism of monetary policy is obtained from the equilibrium
conditions of households and firms in the economy. Households choose
consumption to maximize an isoelastic utility function and firms set prices
subject to Calvo frictions. These give rise to an Euler equation for output
and a Phillips curve, respectively. Equation ((ref)) in the paper
can be derived by combining those two equations. I will derive the Euler
equation in some detail in order to illustrate the origins of the QE channel.
The Phillips curve derivation is standard and is therefore omitted.
Up to a loglinear approximation, the relevant first-order conditions of the
households' optimization problem can be written as
\begin{align}
0 & =E_{t}\left[ -\frac{1}{\sigma}\left( \hat{c}_{t+1}^{u}-\hat{c}_{t}
^{u}\right) +\hat{r}_{t}-\pi_{t+1}\right] ,\\
\frac{\zeta}{1+\zeta}\hat{\zeta}_{t} & =E_{t}\left[ -\frac{1}{\sigma
}\left( \hat{c}_{t+1}^{u}-\hat{c}_{t}^{u}\right) +\hat{R}_{L,t+1}-\pi
_{t+1}\right] ,\\
0 & =E_{t}\left[ -\frac{1}{\sigma}\left( \hat{c}_{t+1}^{r}-\hat{c}_{t}
^{r}\right) +\hat{R}_{L,t+1}-\pi_{t+1}\right] ,
\end{align}
where $\sigma$ is the elasticity of intertemporal substitution, $\zeta$ is the
steady state value of $\zeta_{t},$ hatted variables denote log-deviations from
steady state, $c_{t}^{j}$ is consumption of household $j\in\left\{
u,r\right\} ,$ $r_{t}$ is the short-term nominal interest rate, and $R_{L,t}$
is the gross yield on long-term government bonds from period $t-1$ to
$t$.\footnote{I do not put a hat over $\pi_{t}$ because I assume a zero
inflation target for simplicity.} Goods market clearing yields
\begin{equation}
\hat{y}_{t}=\omega_{r}\hat{c}_{t}^{r}+\left( 1-\omega_{r}\right) \hat{c}
_{t}^{u},
\end{equation}
where $y_{t}$ is output, and I\ have assumed, for simplicity, that in steady
state $c^{u}=c^{r}$, which implies $c^{u}=c^{r}=y.$ Multiplying ((ref))
and ((ref)) by $\left( 1-\omega_{r}\right) $ and $\omega_{r},$
respectively, and adding them yields
\begin{equation}
\hat{y}_{t}=E_{t}\hat{y}_{t+1}-\sigma E_{t}\left[ \left( 1-\omega
_{r}\right) \hat{r}_{t}+\omega_{r}\hat{R}_{L,t+1}-\pi_{t+1}\right] .
\end{equation}
Subtracting ((ref)) from ((ref)) yields
\begin{equation}
E_{t}\left( \hat{R}_{L,t+1}\right) =\hat{r}_{t}+\frac{\zeta}{1+\zeta}
\hat{\zeta}_{t},
\end{equation}
which establishes that the term premium between long and short yields is
proportional to $\hat{\zeta}_{t}.$ Substituting for $E_{t}\left( \hat
{R}_{L,t+1}\right) $ in ((ref)) using ((ref))
yields
\begin{equation}
\hat{y}_{t}=E_{t}\hat{y}_{t+1}-\sigma\left( \hat{r}_{t}+\omega_{r}\frac
{\zeta}{1+\zeta}\hat{\zeta}_{t}\right) +\sigma E_{t}\left( \pi_{t+1}\right)
\end{equation}
Next, assume that the cost of trading long-term bonds depends on their supply,
$b_{L,t},$ i.e.,
\[
\hat{\zeta}_{t}=\rho_{\zeta}\hat{b}_{L,t},\quad\rho_{\zeta}\geq0.
\]
Substituting for $\hat{\zeta}_{t}$ in ((ref)) yields the
Euler equation
\begin{equation}
\hat{y}_{t}=E_{t}\hat{y}_{t+1}-\sigma\left( \hat{r}_{t}+\omega_{r}\frac
{\zeta}{1+\zeta}\rho_{\zeta}\hat{b}_{L,t}\right) +\sigma E_{t}\left(
\pi_{t+1}\right) .
\end{equation}
The second equation is a standard New Keynesian Phillips curve that links
inflation to output:
\begin{equation}
\pi_{t}=\delta E_{t}\left( \pi_{t+1}\right) +\varpi\hat{y}_{t}
+\varepsilon_{1t}
\end{equation}
where $\delta$ is the average discount factor of the two households,
$\varpi\geq0$ is a parameter that depends on the degree of price stickiness
(the Calvo parameter), and $\varepsilon_{1t}$ is proportional to an
\textit{i.i.d.} technology shock. Substituting for $\hat{y}_{t}$ in
((ref)) using ((ref)) yields
\begin{equation}
\pi_{t}=\left( \delta+\sigma\varpi\right) E_{t}\left( \pi_{t+1}\right)
+\varpi E_{t}\left( \hat{y}_{t+1}\right) -\varpi\sigma\left( \hat{r}
_{t}+\omega_{r}\frac{\zeta}{1+\zeta}\rho_{\zeta}\hat{b}_{L,t}\right)
+\varepsilon_{1t}.
\end{equation}
Finally, at an equilibrium in which inflation and output depend only on the
exogenous shocks $\varepsilon_{t}=\left( \varepsilon_{1t},\varepsilon
_{2t}\right) ^{\prime},$ which are the only state variables in the system,
and when the shocks have no memory, $E_{t}\left( \pi_{t+1}\right) $ and
$E_{t}\left( \hat{y}_{t+1}\right) $ will be equal to the corresponding
unconditional expectations, which are constants.\footnote{Such an equilibrium
always exists if the volatility of the shocks is not too large, see
Mendes2011.} Therefore, ((ref)) reduces to
eq. ((ref)) in the paper by setting $c=\left( \delta+\sigma
\varpi\right) E\left( \pi_{t+1}\right) +\varpi E\left( \hat{y}
_{t+1}\right) ,$ $\beta=-\varpi\sigma$, $\hat{r}_{t}=r_{t}-r^{n},$ where
$r^{n}$ is the discount rate of the unconstrained households, and
$\varphi=\omega_{r}\frac{\zeta}{1+\zeta}\rho_{\zeta}\beta$, and dropping the hat from $b_{L,t}$ for simplicity. The parameter $\varphi$ depends on the fraction of constrained households, $\omega_{r}$, and the
sensitivity of the term premium to long-term asset holdings, $\frac{\zeta
}{1+\zeta}\rho_{\zeta}$.
\section{Forward guidance rules}
DebortoliGaliGambetti2019 discuss the following two inertial policy
rules
\begin{equation}
r_{t}=\max\left\{ 0,\phi_{r}r_{t-1}+\left( 1-\phi_{r}\right) \left(
\rho+\phi_{\pi} \pi_{t}+\phi_{y}\Delta y_{t}\right) \right\}
\end{equation}
and
\begin{subequations}
\begin{align}
r_{t} & =\max\left( 0,r_{t}^{\ast}\right) \\
r_{t}^{\ast} & =\phi_{r}r_{t-1}^{\ast}+\left( 1-\phi_{r}\right) \left(
\rho+\phi_{\pi} \pi_{t}+\phi_{y}\Delta y_{t}\right) ,
\end{align}
where I have set the inflation target to zero, and $\Delta y_{t}$ is output
growth. Both of these rules are nested within equation ((ref)) of
the CKSVAR, with $Y_{1t}=\left( \pi_{t},\Delta y_{t}\right) ^{\prime}\,$and
$Y_{2t}=r_{t}.$ Rule ((ref)) sets the coefficients on
$Y_{2,t-1}$ and $Y_{2,t-1}^{\ast}$ as $B_{22}=\phi_{r}$ and $B_{22}^{\ast}=0,$
respectively, while rule ((ref)) sets them as $B_{22}=0$ and
$B_{22}^{\ast}=\phi_{r}.$ DebortoliGaliGambetti2019 argue rule
((ref)) is consistent with forward guidance, because it will
tend to keep interest rates at zero for longer than rule
((ref)). It also ensures policy reaction is the same across
regimes, and so it is consistent with the ZLB irrelevance hypothesis that the
paper puts forward.
ReifschneiderWilliams2000 propose a slightly more elaborate policy rule
for forward guidance:
\end{subequations}
\begin{subequations}
\begin{align}
r_{t}^{\ast} & =r_{t}^{Taylor}-\alpha Z_{t},\quad Z_{t}=Z_{t-1}+d_{t},\quad
d_{t}:=r_{t}-r_{t}^{Taylor},\\
r_{t} & =\max\left( r_{t}^{\ast},0\right) ,\nonumber\\
r_{t}^{Taylor} & =\rho+\phi_{\pi}\pi_{t}+\phi_{y}y_{t}
\end{align}
where $y_{t}$ is the output gap, and the inflation target is zero.
Differencing ((ref)) yields
\end{subequations}
\[
r_{t}^{\ast}=r_{t-1}^{\ast}+\Delta r_{t}^{Taylor}-\alpha\left( r_{t}
-r_{t}^{Taylor}\right) .
\]
Substituting for $r_{t}^{Taylor}$ using ((ref)) yields
\begin{align*}
r_{t}^{\ast} & =r_{t-1}^{\ast}+\phi_{\pi}\Delta\pi_{t}+\phi_{y}\Delta
y_{t}-\alpha r_{t}+\alpha\left( \rho+\phi_{\pi}\pi_{t}+\phi_{y}y_{t}\right)
\\
& =\alpha\rho-\alpha r_{t}+\left( 1+\alpha\right) \left( \phi_{\pi}\pi
_{t}+\phi_{y}y_{t}\right) -\left( \phi_{\pi}\pi_{t-1}+\phi_{y}
y_{t-1}\right) +r_{t-1}^{\ast}.
\end{align*}
This is again nested within equation ((ref)) of the CKSVAR with
$Y_{1t}=\left( \pi_{t},y_{t}\right) ^{\prime},$ $Y_{2t}=r_{t},$
$Y_{2t}^{\ast}=r_{t}^{\ast},$ $X_{1t}=\left( 1,\pi_{t-1},y_{t-1}\right)
^{\prime},$ $X_{2t}=r_{t-1},$ $X_{2t}^{\ast}=r_{t-1}^{\ast},$ and parameters
$A_{21}=-\left( 1+\alpha\right) \left( \phi_{\pi},\phi_{y}\right) ,$
$A_{22}=\alpha,$ $A_{22}^{\ast}=1,$ $B_{21}=\left( \alpha\rho,-\phi_{\pi
},-\phi_{y}\right) ,$ $B_{22}=0$ and $B_{22}^{\ast}=1.$
More examples of forward guidance policy rules that are nested within the
CKSVAR are discussed in IkedaLiMavroeidisZanetti2020.
\section{Computational details}
\subsection{Likelihood}
To compute the likelihood, we need to obtain the prediction error densities.
The first step is to write the model in state-space form. Define
\[
s_{t}=
\begin{pmatrix}
\mathbf{y}_{t}\\
\vdots\\
\mathbf{y}_{t-p+1}
\end{pmatrix}
,\quad\underset{\left( k+1\right) \times1}{\mathbf{y}_{t}}=
\begin{pmatrix}
Y_{t}\\
\overline{Y}_{2t}^{\ast}
\end{pmatrix}
,
\]
and write the state transition equation as
\begin{equation}
s_{t}=F\left( s_{t-1},u_{t};\psi\right) =
\begin{pmatrix}
F_{1}\left( s_{t-1},u_{t};\psi\right) \\
\mathbf{y}_{t-1}\\
\vdots\\
\mathbf{y}_{t-p+1}
\end{pmatrix}
,
\end{equation}
\[
F_{1}\left( s_{t-1},u_{t};\psi\right) =
\begin{pmatrix}
\overline{C}_{1}X_{t}+\overline{C}_{1}^{\ast}\overline{X}_{t}^{\ast}
+u_{1t}-\widetilde{\beta}D_{t}\left( \overline{C}_{2}X_{t}+\overline{C}
_{2}^{\ast}\overline{X}_{t}^{\ast}+u_{2t}-b\right) \\
\max\left( b,\overline{C}_{2}X_{t}+\overline{C}_{2}^{\ast}\overline{X}
_{t}^{\ast}+u_{2t}\right) \\
\overline{C}_{2}X_{t}+\overline{C}_{2}^{\ast}\overline{X}_{t}^{\ast}+u_{2t}
\end{pmatrix}
,
\]
and the observation equation as
\begin{equation}
Y_{t}=
\begin{pmatrix}
I_{k} & 0_{k\times1+\left( p-1\right) \left( k+1\right) }
\end{pmatrix}
s_{t}.
\end{equation}
Next, I will derive the predictive density and mass functions. With Gaussian
errors, the joint predictive density of $Y_{t}$ corresponding to the
observations with $D_{t}=0$ is:
\begin{multline}
f_{0}\left( Y_{t}|s_{t-1},\psi\right) =\left\vert \Omega\right\vert
^{-1/2}\exp\left\{ -\frac{1}{2}tr\left( \left( Y_{t}-\overline{C}
X_{t}-\overline{C}^{\ast}\overline{X}_{t}^{\ast}\right) \right.\right. \\\left.\left. \left(
Y_{t}-\overline{C}X_{t}-\overline{C}^{\ast}\overline{X}_{t}^{\ast}\right)
^{\prime}\Omega^{-1}\right) \right\} .
\end{multline}
At $D_{t}=1,$ the predictive density of $Y_{1t}$ can be written as:
\begin{align}
f_{1}\left( Y_{1t}|s_{t-1},\psi\right) & :=\left\vert \Xi_{1}\right\vert
^{-\frac{1}{2}}\exp\left[ -\frac{1}{2}\left( Y_{1t}-\mu_{1t}\right)
^{\prime}\Xi_{1}^{-1}\left( Y_{1t}-\mu_{1t}\right) \right] \\
\mu_{1t} & :=\widetilde{\beta}b+\left( \overline{C}_{1}-\widetilde{\beta
}\overline{C}_{2}\right) X_{t}+\left( \overline{C}_{1}^{\ast}
-\widetilde{\beta}\overline{C}_{2}^{\ast}\right) \overline{X}_{t}^{\ast
}\\
\Xi_{1} & :=\Omega_{1.2}+\widetilde{\delta}\widetilde{\delta}^{\prime}
\tau^{2}=
\begin{pmatrix}
I_{k-1} & -\widetilde{\beta}
\end{pmatrix}
\Omega\binom{I_{k-1}}{-\widetilde{\beta}^{\prime}},\quad\widetilde{\delta
}=\Omega_{12}\omega_{22}^{-1}-\widetilde{\beta},
\end{align}
where $\Omega_{1.2}=\Omega_{11}-\Omega_{12}\omega_{22}^{-1}\Omega_{21},$ and
$\tau=\sqrt{\omega_{22}}$. Next,
\begin{align}
u_{2t}|Y_{1t},s_{t-1} & \sim N\left( \mu_{2t},\tau_{2}^{2}\right) ,\text{
with}\\
\mu_{2t} & :=\tau^{2}\widetilde{\delta}^{\prime}\Xi_{1}^{-1}\left(
Y_{1t}-\mu_{1t}\right) ,\quad\tau_{2}=\tau\sqrt{\left( 1-\tau^{2}
\widetilde{\delta}^{\prime}\Xi_{1}^{-1}\widetilde{\delta}\right) }.
\end{align}
Hence,
\begin{equation}
\Pr\left( D_{t}=1|Y_{1t},s_{t-1},\psi\right) =\Phi\left( \frac
{b-\overline{C}_{2}X_{t}-\overline{C}_{2}^{\ast}\overline{X}_{t}^{\ast}
-\mu_{2t}}{\tau_{2}}\right) .
\end{equation}
In the case of the KSVAR model, there are no latent lags ($\overline{C}^{\ast
}=0,$ $\overline{C}=C$), so the log-likelihood is available analytically:
\begin{multline}
\log L\left( \psi\right) =\sum_{t=1}^{T} \left( 1-D_{t}\right) \log
f_{0}\left( Y_{t}|s_{t-1},\psi\right) \\ . +\sum_{t=1}^{T}D_{t}\log\left( f_{1}\left(
Y_{1t}|s_{t-1},\psi\right) \Phi\left( \frac{b-C_{2}X_{t}-\mu_{2t}}{\tau_{2}
}\right) \right)
\end{multline}
where $f_{0}\left( Y_{t}|s_{t-1},\theta\right) $ and $f_{1}\left(
Y_{1t}|s_{t-1},\theta\right) $ are given by ((ref)) and ((ref)
), resp., with $\overline{C}^{\ast}=0$.
The likelihood for the unrestricted CKSVAR ($\overline{C}^{\ast}\neq0$) can be computed approximately by simulation (particle filtering). I provide two different simulation algorithms. The first
is a sequential importance sampler (SIS), proposed originally by
Lee1999 for the univariate dynamic Tobit model. It is extended here to
the CKSVAR model. The second algorithm is a fully adapted particle filter
(FAPF), which is a sequential importance resampling algorithm designed to
address the sample degeneracy problem. It is proposed by MalikPitt2011
and is a special case of the auxiliary particle filter developed by
PittShephard1999.
Both algorithms require sampling from the predictive density of $\overline
{Y}_{2t}^{\ast}$ conditional on $Y_{1t},D_{t}=1$ and $s_{t-1}.$ From
((ref)) and ((ref)), we see that this is a
truncated Normal with original mean $\mu_{2t}^{\ast}=\overline{C}_{2}
X_{t}+\overline{C}_{2}^{\ast}\overline{X}_{t}^{\ast}+\mu_{2t}$ and standard
deviation $\tau_{2},$ where $\mu_{2t},\tau_{2}$ are given in
((ref)), i.e.,
\begin{equation}
f_{2}\left( Y_{2t}^{\ast}|Y_{1t},D_{t}=1,s_{t-1},\psi\right) =TN\left(
\mu_{2t}^{\ast},\tau_{2},\overline{Y}_{2t}^{\ast}<b\right)
\end{equation}
Draws from this truncated distribution can be obtained using, for instance,
the procedure in Lee1999. Let $\xi_{t}^{\left( j\right) }\sim
U\left[ 0,1\right] $ be \textit{i.i.d.} uniform random draws, $j=1,...,M$.
Then, a draw from $\overline{Y}_{2t}^{\ast}|Y_{1t},s_{t-1},\overline{Y}
_{2t}^{\ast}<b$ is given by
\begin{equation}
\overline{Y}_{2t}^{\ast\left( j\right) }=\mu_{2t}^{\ast}+\tau_{2}\Phi
^{-1}\left[ \xi_{t}^{\left( j\right) }\Phi\left( \frac{b-\mu_{2t}^{\ast}
}{\tau_{2}}\right) \right] .
\end{equation}
\begin{algorithm}
[SIS]Sequential Importance Sampler
\begin{enumerate}
• \emph{Initialization.} For $j=1:M,$ set $W_{0}^{j}=1$ and $s_{0}
^{j}=\left( \mathbf{y}_{0}^{j},\ldots,\mathbf{y}_{-p+1}^{j}\right) ,$ with
$\mathbf{y}_{-s}^{j}=\left( Y_{0}^{\prime},Y_{2,0}\right) ^{\prime},$ for
$s=0,...,p-1.$ (in other words, initialize $\overline{Y}_{2,-s}^{\ast}$ at the
observed values of $Y_{2,-s}$).
• \emph{Recursion.} For $t=1:T$:
\begin{enumerate}
• For $j=1:M,$ compute the incremental weights
\[
w_{t-1|t}^{j}=p\left( Y_{t}|s_{t-1}^{j},\psi\right) =\left\{
\begin{array}
[c]{ll}
f_{0}\left( Y_{t}|s_{t-1}^{j},\psi\right) , & \text{if }D_{t}=0\\
f_{1}\left( Y_{1t}|s_{t-1}^{j},\psi\right) \Pr\left( D_{t}=1|Y_{1t}
,s_{t-1}^{j},\psi\right) , & \text{if }D_{t}=1
\end{array}
\right.
\]
where $f_{0},f_{1},$ and $\Pr\left( D_{t}=1|Y_{1t},s_{t-1};\psi\right) $ are
given by ((ref)), ((ref)), and ((ref)), resp., and
\[
S_{t}=\frac{1}{M}\sum_{j=1}^{M}w_{t-1|t}^{j}W_{t-1}^{j}
\]
• Sample $s_{t}^{j}$ randomly from $p\left( s_{t}|s_{t-1}^{j}
,Y_{t}\right) $. That is, $s_{t}^{j}=\left( \mathbf{y}_{t}^{j}
,\mathbf{y}_{t-1}^{j},\ldots,\mathbf{y}_{t-p}^{j}\right) $ where
$\mathbf{y}_{t}^{j}=\left( Y_{t}^{\prime},\overline{Y}_{2t}^{\ast\left(
j\right) }\right) \ $and $\overline{Y}_{2t}^{\ast\left( j\right) }$ is a
draw from $f_{2}\left( Y_{2t}^{\ast}|Y_{1t},D_{t}=1,s_{t-1}^{j},\psi\right)
$ using ((ref)).
• Update the weights:
\[
W_{t}^{j}=\frac{w_{t-1|t}^{j}W_{t-1}^{j}}{S_{t}}.
\]
\end{enumerate}
• Likelihood approximation
\[
\log\widehat{p}(Y_{T}|\psi)=\sum_{t=1}^{T}\log S_{t}
\]
\end{enumerate}
\end{algorithm}
If the draws $\xi_{t}^{\left( j\right) }$ are kept fixed across different
values of $\psi$, the simulated likelihood in step 3 is smooth. Note that when
$k=1$ and $Y_{t}=Y_{2t}$ (no $Y_{1t}$ variables), the model reduces to a
univariate dynamic Tobit model, and Algorithm (ref) reduces exactly
to the sequential importance sampler proposed by Lee1999. A possible
weakness of this algorithm is sample degeneracy, which arises when all but a
few weights $W_{t}^{J}$ are zero. To gauge possible sample degeneracy, we can
look at the effective sample size (ESS), as recommended by
HerbstSchorfheide2015
\begin{equation}
ESS_{t}=\frac{M}{\frac{1}{M}\sum_{j=1}^{M}\left( W_{t}^{j}\right) ^{2}}.
\end{equation}
Next, I turn to the FAPF algorithm.
\begin{algorithm}
[FAPF]Fully Adapted Particle Filter
\begin{enumerate}
• \emph{Initialization.} For $j=1:M,$ set $s_{0}^{j}=\left(
\mathbf{y}_{0}^{j},\ldots,\mathbf{y}_{-p+1}^{j}\right) ,$ with $\mathbf{y}
_{-s}^{j}=\left( Y_{0}^{\prime},Y_{2,0}\right) ^{\prime},$ for
$s=0,...,p-1.$ (in other words, initialize $\overline{Y}_{2,-s}^{\ast}$ at the
observed values of $Y_{2,-s}$).
• \emph{Recursion.} For $t=1:T$:
\begin{enumerate}
• For $j=1:M,$ compute
\[
w_{t-1|t}^{j}=p\left( Y_{t}|s_{t-1}^{j},\psi\right) =\left\{
\begin{array}
[c]{ll}
f_{0}\left( Y_{t}|s_{t-1}^{j},\psi\right) , & \text{if }D_{t}=0\\
f_{1}\left( Y_{1t}|s_{t-1}^{j},\psi\right) \Pr\left( D_{t}=1|Y_{1t}
,s_{t-1}^{j},\psi\right) , & \text{if }D_{t}=1
\end{array}
\right.
\]
where $f_{0},f_{1},$ and $\Pr\left( D_{t}=1|Y_{1t},s_{t-1};\psi\right) $ are
given by ((ref)), ((ref)), and ((ref)), resp., and
\[
\pi_{t-1|t}^{j}=\frac{w_{t-1|t}^{j}}{\sum_{j=1}^{M}w_{t-1|t}^{j}}.
\]
• For $j=1:M$, sample $k_{j}$ randomly from the multinomial distribution
$\left\{ j,\pi_{t-1|t}^{j}\right\} .$ Then, set $\tilde{s}_{t-1}^{j}
=s_{t-1}^{k_{j}}$ (this applies only to the elements in $s_{t-1}^{j}$ that
correspond to $X_{t}^{\ast j},$ since all the other elements are observed and
constant across all $j$. That is, $\tilde{s}_{t-1}^{j}=\left( \mathbf{\tilde
{y}}_{t-1}^{j},\allowbreak\ldots\allowbreak,\mathbf{\tilde{y}}_{t-p}^{j}\right) ,$ $\mathbf{\tilde
{y}}_{t-s}^{j}=\left( Y_{t-1}^{\prime},\overline{Y}_{2,t-s}^{\ast\left(
k_{j}\right) }\right) ,$ $s=1,...,p.$)
• For $j=1:M$, sample $s_{t}^{j}$ randomly from $p\left( s_{t}|\tilde
{s}_{t-1}^{j},Y_{t}\right) $. That is, $s_{t}^{j}=\left( \mathbf{y}_{t}
^{j},\allowbreak\ldots,\allowbreak\mathbf{\tilde{y}}_{t-p}^{j}\right)
$ where $\mathbf{y}_{t}^{j}=\left( Y_{t}^{\prime},\overline{Y}_{2t}
^{\ast\left( j\right) }\right) \ $and $\overline{Y}_{2t}^{\ast\left(
j\right) }$ is a draw from $f_{2}\left( Y_{2t}^{\ast}|Y_{1t},D_{t}
=1,\tilde{s}_{t-1}^{j},\psi\right) $ using ((ref)).
\end{enumerate}
• Likelihood approximation
\[
\ln\hat{p}\left( Y_{T}|\psi\right) =\sum_{t=1}^{T}\ln\left( \frac{1}{M}
\sum_{j=1}^{M}w_{t-1|t}^{j}\right)
\]
\end{enumerate}
\end{algorithm}
Many of the generic particle filtering algorithms used in the macro
literature, described in HerbstSchorfheide2015, are inapplicable in a
censoring context because of the absence of measurement error in the
observation equation. It is, of course, possible to introduce a small
measurement error in $Y_{2t},$ so that the constraint $Y_{2t}\geq b$ is not
fully respected, but there is no reason to expect other particle filters
discussed in HerbstSchorfheide2015 to estimate the likelihood more
accurately than the FAPF algorithm described above.
Moments or quantiles of the filtering or smoothing distribution of any
function $h\left( \cdot\right) $ of the latent states $s_{t}$ can be
computed using the drawn sample of particles. When we use Algorithm
(ref), simple average or quantiles of $h\left( s_{t}^{j}\right) $
produce the requisite average or quantiles of $h\left( s_{t}\right) $
conditional on $Y_{1},\ldots,Y_{t}$ (the filtering density). For particles
generated using Algorithm (ref), we need to take weighted averages
using the importance sampling weights $W_{t}.$ Smoothing estimates of
$h\left( s_{t}^{j}\right) $ can be obtained using weights $W_{T}.$
\subsection{Computation of the identified set}
Substitute for $\overline{\gamma}$ in ((ref)) using Proposition
(ref) to get
\begin{equation}
\widetilde{\beta}=\left( 1-\xi\right) \left( I-\xi\overline{\beta}\left(
\Omega_{12}^{\prime}-\Omega_{22}\overline{\beta}^{\prime}\right) \left(
\Omega_{11}-\Omega_{12}\overline{\beta}^{\prime}\right) ^{-1}\right)
^{-1}\overline{\beta}.
\end{equation}
For each value of $\xi\in\lbrack0,1),$ the above equation defines a
correspondence from $\Re^{k-1}$ to $\Re^{k-1}$. The range of $\overline{\beta
}$ can then be obtained numerically by solving ((ref))
for $\overline{\beta}$ as a function of the reduced-form parameters and $\xi$
for each value of $\xi,$ and gathering all the solutions in the set.
Rearranging ((ref)) yields
\begin{equation}
\widetilde{\beta}=\xi\overline{\beta}\left( \Omega_{12}^{\prime}-\Omega
_{22}\overline{\beta}^{\prime}\right) \left( \Omega_{11}-\Omega
_{12}\overline{\beta}^{\prime}\right) ^{-1}\widetilde{\beta}+\left(
1-\xi\right) \overline{\beta}.
\end{equation}
Note that
\[
\left( \Omega_{11}-\Omega_{12}\overline{\beta}^{\prime}\right) ^{-1}
=\Omega_{11}^{-1}+\Omega_{11}^{-1}\Omega_{12}\left( 1-\overline{\beta
}^{\prime}\Omega_{11}^{-1}\Omega_{12}\right) ^{-1}\overline{\beta}^{\prime
}\Omega_{11}^{-1}.
\]
Hence,
\begin{align*}
& \left( \Omega_{12}^{\prime}-\Omega_{22}\overline{\beta}^{\prime}\right)
\left( \Omega_{11}-\Omega_{12}\overline{\beta}^{\prime}\right) ^{-1}\\
& =\left( \Omega_{12}^{\prime}-\Omega_{22}\overline{\beta}^{\prime}\right)
\Omega_{11}^{-1}+\frac{\left( \Omega_{12}^{\prime}\Omega_{11}^{-1}
\Omega_{12}-\Omega_{22}\overline{\beta}^{\prime}\Omega_{11}^{-1}\Omega
_{12}\right) \overline{\beta}^{\prime}\Omega_{11}^{-1}}{1-\overline{\beta
}^{\prime}\Omega_{11}^{-1}\Omega_{12}}\\
& =\frac{\left( \Omega_{12}^{\prime}-\Omega_{22}\overline{\beta}^{\prime
}\right) \Omega_{11}^{-1}+\Omega_{12}^{\prime}\Omega_{11}^{-1}\left(
\Omega_{12}\overline{\beta}^{\prime}\Omega_{11}^{-1}-\overline{\beta}^{\prime
}\Omega_{11}^{-1}\Omega_{12}I_{k-1}\right) }{1-\overline{\beta}^{\prime
}\Omega_{11}^{-1}\Omega_{12}}.
\end{align*}
Substituting this back into ((ref)), we get
\[
\widetilde{\beta}=\xi\overline{\beta}\frac{\left( \Omega_{12}^{\prime}
-\Omega_{22}\overline{\beta}^{\prime}\right) \Omega_{11}^{-1}+\Omega
_{12}^{\prime}\Omega_{11}^{-1}\left( \Omega_{12}\overline{\beta}^{\prime
}\Omega_{11}^{-1}-\overline{\beta}^{\prime}\Omega_{11}^{-1}\Omega_{12}
I_{k-1}\right) }{1-\overline{\beta}^{\prime}\Omega_{11}^{-1}\Omega_{12}
}\widetilde{\beta}+\left( 1-\xi\right) \overline{\beta}.
\]
Multiplying both sides by $1-\overline{\beta}^{\prime}\Omega_{11}^{-1}\Omega_{12}$ yields
\begin{align*}
\widetilde{\beta}-\overline{\beta}^{\prime}\Omega_{11}^{-1}\Omega
_{12}\widetilde{\beta} & =\xi\overline{\beta}\Omega_{12}^{\prime
}\widetilde{\beta}-\xi\overline{\beta}\Omega_{22}\overline{\beta}^{\prime
}\Omega_{11}^{-1}\widetilde{\beta}+\xi\overline{\beta}\Omega_{12}^{\prime
}\Omega_{11}^{-1}\Omega_{12}\overline{\beta}^{\prime}\Omega_{11}
^{-1}\widetilde{\beta}\\
& -\xi\overline{\beta}\overline{\beta}^{\prime}\Omega_{11}^{-1}\Omega
_{12}\Omega_{12}^{\prime}\Omega_{11}^{-1}\widetilde{\beta}+\left(
1-\xi\right) \overline{\beta}\left( 1-\overline{\beta}^{\prime}\Omega
_{11}^{-1}\Omega_{12}\right).
\end{align*}
Rearranging, we have
\begin{align*}
\widetilde{\beta} & =\widetilde{\beta}\Omega_{12}^{\prime}\Omega_{11}
^{-1}\overline{\beta}+\xi\Omega_{12}^{\prime}\widetilde{\beta}\overline{\beta
}+\left( 1-\xi\right) \overline{\beta}-\left( 1-\xi\right) \overline
{\beta}\overline{\beta}^{\prime}\Omega_{11}^{-1}\Omega_{12}\\
& +\overline{\beta}\overline{\beta}^{\prime}\Omega_{11}^{-1}\widetilde{\beta
}\xi\Omega_{12}^{\prime}\Omega_{11}^{-1}\Omega_{12}-\overline{\beta}
\overline{\beta}^{\prime}\Omega_{11}^{-1}\widetilde{\beta}\xi\Omega
_{22}-\overline{\beta}\overline{\beta}^{\prime}\Omega_{11}^{-1}\Omega
_{12}\Omega_{12}^{\prime}\Omega_{11}^{-1}\widetilde{\beta}\xi
\\ & =\left( \widetilde{\beta}\Omega_{12}^{\prime}
\Omega_{11}^{-1}+\left( \xi\Omega_{12}^{\prime}\widetilde{\beta}
+1-\xi\right) I_{k-1}\right) \overline{\beta}\\
& +\overline{\beta}\overline{\beta}^{\prime}\Omega_{11}^{-1}\left( \left(
\left( \Omega_{12}^{\prime}\Omega_{11}^{-1}\Omega_{12}-\Omega_{22}\right)
I_{k-1}-\Omega_{12}\Omega_{12}^{\prime}\Omega_{11}^{-1}\right)
\widetilde{\beta}\xi-\left( 1-\xi\right) \Omega_{12}\right).
\end{align*}
This can be written as
\begin{equation}
\widetilde{\beta}-\tilde{A}\overline{\beta}+\overline{\beta}\overline{\beta
}^{\prime}\tilde{b}=0,
\end{equation}
where
\begin{align*}
\tilde{b} & :=-\Omega_{11}^{-1}\left( \left( \left( \Omega_{12}^{\prime}
\Omega_{11}^{-1}\Omega_{12}-\Omega_{22}\right) I_{k-1}-\Omega_{12}\Omega
_{12}^{\prime}\Omega_{11}^{-1}\right) \widetilde{\beta}\xi-\left(
1-\xi\right) \Omega_{12}\right) , \text{ and}\\
\bar{A} & :=\widetilde{\beta}\Omega_{12}^{\prime}\Omega_{11}^{-1}+\left(
\xi\Omega_{12}^{\prime}\widetilde{\beta}+1-\xi\right) I_{k-1}.
\end{align*}
Define
\[
z:=\tilde{b}^{\prime}x\text{ \ \ and \ \ }w:=\tilde{b}_{\perp}^{\prime}x,
\]
where $\tilde{b}_{\perp}^{\prime}\tilde{b}_{\perp}=1$ and $\tilde{b}_{\perp
}^{\prime}\tilde{b}=0$. Hence, rewrite ((ref)) as
\[
\widetilde{\beta}-\tilde{A}\tilde{b}\left( \tilde{b}^{\prime}\tilde
{b}\right) ^{-1}z-\tilde{A}\tilde{b}_{\perp}w+\tilde{b}\left( \tilde
{b}^{\prime}\tilde{b}\right) ^{-1}z^{2}+\tilde{b}_{\perp}wz=0.
\]
Premultiply by $\tilde{b}_{\perp}^{\prime}$ to get
\[
\tilde{b}_{\perp}^{\prime}\widetilde{\beta}-\tilde{b}_{\perp}^{\prime}
\tilde{A}\tilde{b}\left( \tilde{b}^{\prime}\tilde{b}\right) ^{-1}z-\tilde
{b}_{\perp}^{\prime}\tilde{A}\tilde{b}_{\perp}w+wz=0.
\]
Solve that for $w$ to get
\begin{align*}
w & =\left( \tilde{b}_{\perp}^{\prime}\tilde{A}\tilde{b}_{\perp}-z\right)
^{-1}\left( \tilde{b}_{\perp}^{\prime}\widetilde{\beta}-\tilde{b}_{\perp
}^{\prime}\tilde{A}\tilde{b}\left( \tilde{b}^{\prime}\tilde{b}\right)
^{-1}z\right) \\
& =C_{0}\left( z\right) ^{-1}c_{1}\left( z\right) ,
\end{align*}
with
\begin{align*}
C_{0}\left( z\right) & :=\left( \tilde{b}_{\perp}^{\prime}\tilde{A}
\tilde{b}_{\perp}-z\right) ,\text{ \ and}\\
c_{1}\left( z\right) & :=\tilde{b}_{\perp}^{\prime}\widetilde{\beta
}-\tilde{b}_{\perp}^{\prime}\tilde{A}\tilde{b}\left( \tilde{b}^{\prime}
\tilde{b}\right) ^{-1}z,
\end{align*}
provided that $\det\left( C_{0}\left( z\right) \right) \neq0.$
Next, premultiply ((ref)) by $\tilde{b}^{\prime}$ and substitute
for $w$ to get
\begin{equation}
\tilde{b}^{\prime}\widetilde{\beta}-\tilde{b}^{\prime}\tilde{A}\tilde
{b}\left( \tilde{b}^{\prime}\tilde{b}\right) ^{-1}z-\tilde{b}^{\prime}
\tilde{A}\tilde{b}_{\perp}C_{0}\left( z\right) ^{-1}c_{1}\left( z\right)
+z^{2}=0.
\end{equation}
Now, notice that $C_{0}\left( z\right) ^{-1}=C_{0}\left( z\right)
^{adj}/\det\left( C_{0}\left( z\right) \right) ,$ where $C^{adj}$ is the
adjoint of a square matrix $C.$ Moreover, since $C_{0}\left( z\right) $ is
of dimension $k-2$ and its elements are linear in $z,$ $\det\left(
C_{0}\left( z\right) \right) $ is a polynomial in $z$ of order at most
$k-2,$ and the elements of $C_{0}\left( z\right) ^{adj}$ are polynomials in
$z$ of order at most $k-3.$ For $k=2,$ $w$ is empty, so ((ref))
is simply a quadratic in $z.$ When $k>2,$ $\det\left( C_{0}\left( z\right)
\right) $ is nonzero and we can multiply ((ref)) by it to get
\begin{align}
0 & =\tilde{b}^{\prime}\widetilde{\beta}\det\left( C_{0}\left( z\right)
\right) +\tilde{b}^{\prime}\tilde{A}\tilde{b}\left( \tilde{b}^{\prime}
\tilde{b}\right) ^{-1}z\det\left( C_{0}\left( z\right) \right)
\nonumber\\
& -\tilde{b}^{\prime}\tilde{A}\tilde{b}_{\perp}C_{0}\left( z\right)
^{adj}c_{1}\left( z\right) +\det\left( C_{0}\left( z\right) \right)
z^{2},
\end{align}
This is a polynomial equation of order $k$ and has at most $k$ solutions,
denoted $z_{i},$ say. Then, the solutions for $\overline{\beta}$ are given
by
\begin{align}
\overline{\beta}_{i} & =\left[ \tilde{b}\left( \tilde{b}^{\prime}\tilde{b}\right)
^{-1},\tilde{b}_{\perp}\right]
\begin{pmatrix}
z_{i}\\
C_{0}\left( z_{i}\right) ^{-1}c_{1}\left( z_{i}\right)
\end{pmatrix}
\nonumber\\
& =\tilde{b}\left( \tilde{b}^{\prime}\tilde{b}\right) ^{-1}z_{i}+\tilde
{b}_{\perp}C_{0}\left( z_{i}\right) ^{-1}c_{1}\left( z_{i}\right).
\end{align}
Below I give some special cases.
\textit{Case $k=2$:} In this case, $w$ is empty, $\overline{\beta}$ is a scalar, and the equation
((ref)) is a quadratic
\[
\widetilde{\beta}-\tilde{A}\overline{\beta}+\overline{\beta}^{2}\tilde{b}=0.
\]
If $\tilde{A}^{2}-4\widetilde{\beta}\tilde{b}>0,$ the two real solutions are$
\overline{\beta}_{1,2}=\frac{\tilde{A}\pm\sqrt{\tilde{A}^{2}-4\widetilde{\beta}\tilde{b}}
}{2\tilde{b}}.$
\textit{Case $k=3$:} In this case, $w$ is a scalar, and the equation ((ref)) can be
written as a cubic in $z,$ i.e.,
\begin{equation}
C_{0}\left( z\right) \tilde{b}^{\prime}\widetilde{\beta}-C_{0}\left(
z\right) \tilde{b}^{\prime}\tilde{A}\tilde{b}\left( \tilde{b}^{\prime}
\tilde{b}\right) ^{-1}z-\tilde{b}^{\prime}\tilde{A}\tilde{b}_{\perp}
c_{1}\left( z\right) +C_{0}\left( z\right) z^{2}=0,
\end{equation}
since $C_{0}\left( z\right) $ is a scalar linear function of $z.$ It can be
shown that one of the roots of ((ref)) satisfies $\Omega
_{12}^{\prime}\Omega_{11}^{-1}\overline{\beta}=1,$ which implies $\det\left(
\Omega_{11}-\Omega_{12}\overline{\beta}^{\prime}\right) =0,$ and hence
violates the equation for $\overline{\gamma}=\left( \Omega_{12}^{\prime
}-\Omega_{22}\overline{\beta}^{\prime}\right) \left( \Omega_{11}-\Omega
_{12}\overline{\beta}^{\prime}\right) ^{-1},$ so it is not a valid solution.
The root in question is
\[
z_{1}=\frac{\Omega_{12}^{\prime}\Omega_{11}^{-1}\left( \tilde{b}_{\perp
}\tilde{b}^{\prime}\tilde{A}\tilde{b}-\tilde{b}\tilde{b}^{\prime}\tilde
{A}\tilde{b}_{\perp}\right) }{\Omega_{12}^{\prime}\Omega_{11}^{-1}\tilde
{b}_{\perp}\left( \tilde{b}^{\prime}\tilde{b}\right) }
\]
We can then factor out a term $z-z_{1}$ from ((ref)), and obtain the
remaining two roots from a quadratic equation. Therefore, there will be zero
or two solutions for $\overline{\beta}$, as in the case $k=2$.
An algorithm for obtaining the identified set of the IRF ((ref)) is as follows.
\begin{algorithm}
[ID set]Discretize the space $(0,1)$ into $R$ equidistant
points. \newline For each $r=1:R,$ set $\xi_{r}=\frac{r}{R+1}$ and solve
equation ((ref)).
\begin{enumerate}
• If no solution exists, proceed to the next $r.$
• If $0<q_{r}\leq k$ solutions exist, denote them $z_{i,r},$ and, for each
$i=1:q_{r}$,
\begin{enumerate}
• derive $\overline{\beta}_{i,r}$ from ((ref)),
$\overline{\gamma}_{i,r}=\left( \Omega_{12}^{\prime}-\Omega_{22}
\overline{\beta}_{i,r}^{\prime}\right) \left( \Omega_{11}-\Omega
_{12}\overline{\beta}_{i,r}^{\prime}\right) ^{-1}$,\newline$\overline
{A}_{22,i,r}^{-1}\allowbreak=\allowbreak\sqrt{\left( -\overline{\gamma}
_{i,r},1\right) \Omega\left( -\overline{\gamma}_{i,r},1\right) ^{\prime}}$,
and $\Xi_{1,i.r}=\left( I_{k-1},-\overline{\beta}_{i,r}\right) \Omega\left(
I_{k-1},-\overline{\beta}_{i,r}\right) ^{\prime};$
• for $j=1:M,$
\begin{enumerate}
• draw independently $\bar{\varepsilon}_{1t,i,r}^{j}\sim N\left(
0,\Xi_{1,i.r}\right) $ and $u_{t+h}^{j}\sim N\left( 0,\Omega\right) $ for
$h=1,...,H;$
• for any scalar $\varsigma$, set
\begin{align*}
u_{1t,i,r}^{j}\left( \varsigma\right) & =\left( I_{k-1}-\overline{\beta
}_{i,r}\overline{\gamma}_{i,r}\right) ^{-1}\left( \bar{\varepsilon}
_{1t,i,r}^{j}-\overline{\beta}_{i,r}\varsigma\right) \\
u_{2t,i,r}^{j}\left( \varsigma\right) & =\left( 1-\overline{\gamma}
_{i,r}\overline{\beta}_{i,r}\right) ^{-1}\left( \varsigma-\overline{\gamma
}_{i,r}\bar{\varepsilon}_{1t,i,r}^{j}\right) ,
\end{align*}
and compute $Y_{t,i,r}^{j}\left( \varsigma\right) $ using ((ref)
)-((ref)) with $u_{t,i,r}^{j}\left( \varsigma\right) $ in place of
$u_{t},$ and iterate forward to obtain $Y_{t+h,i,r}^{j}\left( \varsigma
\right) $ using $u_{t+h}^{j}$ computed in step i. Set $\varsigma=1$ for a
one-unit (e.g., 100 basis points) impulse to the policy shock $\overline
{\varepsilon}_{2t},$ or $\varsigma=\overline{A}_{22,i,r}^{-1}$ for a
one-standard deviation impulse.
\end{enumerate}
• compute
\[
\widehat{IRF}_{h,t,i,r}\left( \varsigma\right) =\frac{1}{M}\sum_{j=1}
^{M}\left( Y_{t+h,i,r}^{j}\left( \varsigma\right) -Y_{t+h,i,r}^{j}\left(
0\right) \right) .
\]
\end{enumerate}
\end{enumerate}
The identified set is given by the collection of $\widehat{IRF}_{h,t,i,r}
\left( \varsigma\right) $ over $i=1:q_{r},$ $r=1:R,$ and the single
point-identified IRF\ at $\xi=0.$
\end{algorithm}
\subsection{IRFs and local projections}
I will briefly discuss the difficulty in getting a local projection-like representation of the IRF in a
dynamic Tobit model, which is a univariate CKSVAR(1). The model is given by
the equations
\begin{align*}
y_{t}^{\ast} & =\rho y_{t-1}+\rho^{\ast}\min\left( y_{t-1}-b,0\right)
+u_{t}\\
& =\rho y_{t-1}+\rho^{\ast}D_{t-1}\left( y_{t-1}^{\ast}-b\right)
+u_{t},\quad D_{t}=1\left\{ y_{t}^{\ast}<b\right\} ,\\
y_{t} & =\max\left( y_{t}^{\ast},b\right) =\left( 1-D_{t}\right)
y_{t}^{\ast}.
\end{align*}
Hence,
\begin{multline*}
E_{t}\left( y_{t+1}\right) =\left( \rho y_{t}+\rho^{\ast}D_{t}\left(
y_{t}^{\ast}-b\right) \right) \left( 1-\Phi\left( \frac{b-\rho y_{t}
-\rho^{\ast}D_{t}\left( y_{t}^{\ast}-b\right) }{\sigma}\right) \right)
\\ +\sigma\phi\left( \frac{b-\rho y_{t}-\rho^{\ast}D_{t}\left( y_{t}^{\ast
}-b\right) }{\sigma}\right) .
\end{multline*}
In a linear model $\left( \rho^{\ast}=0,\text{ }b=-\infty\right) $, the
1-period ahead impulse response is $\rho,$ which coincides with the
coefficient on $y_{t}$ in the local projection $E_{t}\left( y_{t+1}\right)
=\rho y_{t}.$ In that case, the coefficient $\rho$ corresponds to both
$\frac{\partial E_{t}\left( y_{t+1}\right) }{\partial u_{t}}=\frac{\partial
E_{t}\left( y_{t+1}\right) }{\partial y_{t}}$ and $E\left( y_{t+1}
|u_{t}=1,y_{t-1}\right) -E\left( y_{t+1}|u_{t}=0,y_{t-1}\right) .$ None of
these properties hold in the dynamic Tobit model. For example, if we go with
$\frac{\partial E_{t}\left( y_{t+1}\right) }{\partial u_{t}}$ as our
definition of the impulse response, we will not be able to obtain it from the
slope of the conditional expectation function $E_{t}\left( y_{t+1}\right) $
with respect to $y_{t}$. One problem is that the function $E_{t}\left( y_{t+1}\right) $ is
non-differentiable at $u_{t}=b-\rho y_{t-1}+\rho^{\ast}D_{t-1}\left(
y_{t-1}^{\ast}-b\right) $, i.e., exactly at the boundary. Another problem is that we
still need to rely on the parametric structure of the model to uncover
the impulse response from $E_{t}\left( y_{t+1}\right) $. For example, we
need to compute $\frac{\partial E_{t}\left( y_{t+1}\right) }{\partial u_{t}
},$ which at all points $y_{t}^{\ast}\neq b$\ is given by:
\[
\frac{\partial E_{t}\left( y_{t+1}\right) }{\partial u_{t}}=\left\{
\begin{array}
[c]{ll}
\rho\left( 1-\Phi\left( \frac{b-\rho y_{t}}{\sigma}\right) +\frac{b}
{\sigma}\phi\left( \frac{b-\rho y_{t}}{\sigma}\right) \right) & \text{if
}D_{t}=0\\
\rho^{\ast}\left( 1-\Phi\left( \frac{b-\rho b-\rho^{\ast}\left( y_{t}
^{\ast}-b\right) }{\sigma}\right) +\frac{b}{\sigma}\phi\left( \frac{b-\rho
b-\rho^{\ast}\left( y_{t}^{\ast}-b\right) }{\sigma}\right) \right) &
\text{if }D_{t}=1.
\end{array}
\right.
\]
There is no clear way to obtain the above
impulse response from a local projection of $y_{t+1}$ on simple nonlinear
transformations of $y_{t}$, such as powers or interactions with the regime
indicator.
thebibliography\bibitem[\citeauthoryear{Amemiya}{Amemiya}{1974}]{Amemiya1974}
Amemiya, T. (1974).
\newblock Multivariate regression and simultaneous equation models when the
dependent variables are truncated normal.
\newblock {\em Econometrica\/} {\em 42\/}(6), 999--1012.
\bibitem[\citeauthoryear{Aruoba, Cuba-Borda, Higa-Flores, Schorfheide,
and Villalvazo}{Aruoba
et al.}{2020}]{AruobaCubaBordaHigaFloresSchorfheideVillalvazo2020}
Aruoba, S. B., P. Cuba-Borda, K. Higa-Flores, F. Schorfheide, and S. Villalvazo
(2020).
\newblock {Piecewise-Linear Approximations and Filtering for DSGE Models with
Occasionally Binding Constraints}.
\newblock Technical report.
\bibitem[\citeauthoryear{Aruoba, Cuba-Borda, and Schorfheide}{Aruoba
et al.}{2017}]{AruobaCubaBordaSchorfheide2017}
Aruoba, S. B., P. Cuba-Borda, and F. Schorfheide (2017).
\newblock {Macroeconomic dynamics near the ZLB: A tale of two countries}.
\newblock {\em The Review of Economic Studies\/} {\em 85\/}(1), 87--118.
\bibitem[\citeauthoryear{Aruoba, Schorfheide, and Villalvazo}{Aruoba
et al.}{2020}]{AruobaSchorfheideVillalvazo2020}
Aruoba, S. B., F. Schorfheide, and S. Villalvazo (2020).
\newblock {SVARs with Occasionally-Binding Constraints}.
\newblock Technical report.
\bibitem[\citeauthoryear{Blundell and Smith}{Blundell and
Smith}{1994}]{BlundellSmith1994}
Blundell, R. and R. J. Smith (1994).
\newblock Coherency and estimation in simultaneous models with censored or
qualitative dependent variables.
\newblock {\em Journal of Econometrics\/} {\em 64\/}(1-2), 355 -- 373.
\bibitem[\citeauthoryear{Blundell and Smith}{Blundell and
Smith}{1989}]{BlundellSmith1989}
Blundell, R. W. and R. J. Smith (1989).
\newblock Estimation in a class of simultaneous equation limited dependent
variable models.
\newblock {\em The Review of Economic Studies\/} {\em 56\/}(1), 37--57.
\bibitem[\citeauthoryear{Chen, C{\'u}rdia, and Ferrero}{Chen
et al.}{2012}]{ChenCurdiaFerrero2012}
Chen, H., V. C{\'u}rdia, and A. Ferrero (2012, November).
\newblock The macroeconomic effects of large-scale asset purchase programmes.
\newblock {\em The Economic Journal\/} {\em 122}, F289--F315.
\bibitem[\citeauthoryear{Debortoli, Gali, and Gambetti}{Debortoli
et al.}{2019}]{DebortoliGaliGambetti2019}
Debortoli, D., J. Gali, and L. Gambetti (2019).
\newblock {On the Empirical (Ir) Relevance of the Zero Lower Bound Constraint}.
\newblock In {\em NBER Macroeconomics Annual 2019, volume 34}. University of
Chicago Press.
\bibitem[\citeauthoryear{Eggertsson and Woodford}{Eggertsson and
Woodford}{2003}]{EggertssonWoodford2003}
Eggertsson, G. B. and M. Woodford (2003).
\newblock Zero bound on interest rates and optimal monetary policy.
\newblock {\em Brookings papers on economic activity\/} {\em 2003\/}(1),
139--233.
\bibitem[\citeauthoryear{Fern{\'a}ndez-Villaverde, Gordon,
Guerr{\'o}n-Quintana, and Rubio-Ramirez}{Fern{\'a}ndez-Villaverde
et al.}{2015}]{FernandezGordonGuerronRubio2015}
Fern{\'a}ndez-Villaverde, J., G. Gordon, P. Guerr{\'o}n-Quintana, and J. F.
Rubio-Ramirez (2015).
\newblock Nonlinear adventures at the zero lower bound.
\newblock {\em Journal of Economic Dynamics and Control\/} {\em 57}, 182--204.
\bibitem[\citeauthoryear{Gertler and Karadi}{Gertler and
Karadi}{2015}]{GertlerKaradi2015}
Gertler, M. and P. Karadi (2015).
\newblock Monetary policy surprises, credit costs, and economic activity.
\newblock {\em American Economic Journal: Macroeconomics\/} {\em 7\/}(1),
44--76.
\bibitem[\citeauthoryear{Giacomini, Politis, and White}{Giacomini
et al.}{2013}]{GiacominiPolitisWhite2013}
Giacomini, R., D. N. Politis, and H. White (2013).
\newblock A warp-speed method for conducting {M}onte {C}arlo experiments
involving bootstrap estimators.
\newblock {\em Econometric theory\/} {\em 29\/}(3), 567--589.
\bibitem[\citeauthoryear{Gourieroux, Laffont, and Monfort}{Gourieroux
et al.}{1980}]{GourierouxLaffontMonfort1980}
Gourieroux, C., J. Laffont, and A. Monfort (1980).
\newblock {Coherency Conditions in Simultaneous Linear Equation Models with
Endogenous Switching Regimes}.
\newblock {\em Econometrica\/} {\em 48\/}(3), 675--695.
\bibitem[\citeauthoryear{Greene}{Greene}{1993}]{Gree93}
Greene, W. H. (1993).
\newblock {\em Econometric Analysis}.
\newblock New York: MacMillan.
\bibitem[\citeauthoryear{Guerrieri and Iacoviello}{Guerrieri and
Iacoviello}{2015}]{GuerrieriIacoviello2015}
Guerrieri, L. and M. Iacoviello (2015).
\newblock {OccBin: A toolkit for solving dynamic models with occasionally
binding constraints easily}.
\newblock {\em Journal of Monetary Economics\/} {\em 70}, 22--38.
\bibitem[\citeauthoryear{Hayashi and Koeda}{Hayashi and
Koeda}{2019}]{HayashiKoeda2019}
Hayashi, F. and J. Koeda (2019).
\newblock {Exiting from Quantitative Easing}.
\newblock {\em {Quantitative Economics}\/} {\em 10}, 1069--1107.
\bibitem[\citeauthoryear{Heckman}{Heckman}{1978}]{Heckman1978}
Heckman, J. J. (1978).
\newblock {Dummy Endogenous Variables in a Simultaneous Equation System}.
\newblock {\em Econometrica\/} {\em 46\/}(4), 931--959.
\bibitem[\citeauthoryear{Heckman}{Heckman}{1979}]{Heckman1979}
Heckman, J. J. (1979).
\newblock Sample selection bias as a specification error.
\newblock {\em Econometrica\/} {\em 47\/}(1), 153--161.
\bibitem[\citeauthoryear{Herbst and Schorfheide}{Herbst and
Schorfheide}{2015}]{HerbstSchorfheide2015}
Herbst, E. P. and F. Schorfheide (2015).
\newblock {\em Bayesian estimation of {DSGE} models}.
\newblock Princeton and Oxford: Princeton University Press.
\bibitem[\citeauthoryear{Ikeda, Li, Mavroeidis, and Zanetti}{Ikeda
et al.}{2020}]{IkedaLiMavroeidisZanetti2020}
Ikeda, D., S. Li, S. Mavroeidis, and F. Zanetti (2020).
\newblock {Testing the effectiveness of unconventional monetary policy in Japan
and the United States}.
\newblock Discussion paper 2020-E-10, Institute for Monetary and Economic
Studies, Bank of Japan.
\newblock {A}vailable at
\url{https://www.imes.boj.or.jp/research/papers/english/20-E-10.pdf}.
\bibitem[\citeauthoryear{Koop, Pesaran, and Potter}{Koop
et al.}{1996}]{KoopPesaranPotter1996}
Koop, G., M. H. Pesaran, and S. M. Potter (1996).
\newblock Impulse response analysis in nonlinear multivariate models.
\newblock {\em Journal of econometrics\/} {\em 74\/}(1), 119--147.
\bibitem[\citeauthoryear{Kulish, Morley, and Robinson}{Kulish
et al.}{2017}]{KulishMorleyRobinson2017}
Kulish, M., J. Morley, and T. Robinson (2017).
\newblock {Estimating DSGE models with zero interest rate policy}.
\newblock {\em Journal of Monetary Economics\/} {\em 88\/}(C), 35--49.
\bibitem[\citeauthoryear{Lee}{Lee}{1976}]{Lee1976}
Lee, L.-F. (1976).
\newblock Multivariate regression and simultaneous equations models with some
dependent variables truncated.
\newblock "Discussion paper" 76-79, "University of Minnesota", "Minneapolis,
USA".
\bibitem[\citeauthoryear{Lee}{Lee}{1999}]{Lee1999}
Lee, L.-F. (1999).
\newblock Estimation of dynamic and {ARCH} {T}obit models.
\newblock {\em Journal of Econometrics\/} {\em 92\/}(2), 355--390.
\bibitem[\citeauthoryear{Lewbel}{Lewbel}{2007}]{Lewbel2007}
Lewbel, A. (2007).
\newblock Coherency and completeness of structural models containing a dummy
endogenous variable*.
\newblock {\em International Economic Review\/} {\em 48\/}(4), 1379--1392.
\bibitem[\citeauthoryear{Liebscher}{Liebscher}{2005}]{Liebscher2005}
Liebscher, E. (2005).
\newblock Towards a unified approach for proving geometric ergodicity and
mixing properties of nonlinear autoregressive processes.
\newblock {\em Journal of Time Series Analysis\/} {\em 26\/}(5), 669--689.
\bibitem[\citeauthoryear{Liu, Theodoridis, Mumtaz, and Zanetti}{Liu
et al.}{2019}]{LiuTheodoridisMumtazZanetti2019}
Liu, P., K. Theodoridis, H. Mumtaz, and F. Zanetti (2019).
\newblock Changing macroeconomic dynamics at the zero lower bound.
\newblock {\em Journal of Business & Economic Statistics\/} {\em 37\/}(3),
391--404.
\bibitem[\citeauthoryear{L\"{u}tkepohl}{L\"{u}tkepohl}{1996}]{Lutkepohl96}
L\"{u}tkepohl, H. (1996).
\newblock {\em Handbook of Matrices}.
\newblock England: Wiley.
\bibitem[\citeauthoryear{Magnusson and Mavroeidis}{Magnusson and
Mavroeidis}{2014}]{MagnussonMavroeidis2014}
Magnusson, L. M. and S. Mavroeidis (2014).
\newblock {Identification using stability restrictions}.
\newblock {\em Econometrica\/} {\em 82\/}(5), 1799--1851.
\bibitem[\citeauthoryear{Malik and Pitt}{Malik and
Pitt}{2011}]{MalikPitt2011}
Malik, S. and M. K. Pitt (2011).
\newblock Particle filters for continuous likelihood evaluation and
maximisation.
\newblock {\em Journal of Econometrics\/} {\em 165\/}(2), 190--209.
\bibitem[\citeauthoryear{Mendes}{Mendes}{2011}]{Mendes2011}
Mendes, R. R. (2011, February).
\newblock {Uncertainty and the Zero Lower Bound: A Theoretical Analysis}.
\newblock MPRA Paper 59218, University Library of Munich, Germany.
\bibitem[\citeauthoryear{Nelson and Olson}{Nelson and
Olson}{1978}]{NelsonOlson1978}
Nelson, F. and L. Olson (1978).
\newblock Specification and estimation of a simultaneous-equation model with
limited dependent variables.
\newblock {\em International Economic Review\/} {\em 19\/}(3), 695--709.
\bibitem[\citeauthoryear{Newey and McFadden}{Newey and
McFadden}{1994}]{Newe94Mc}
Newey, W. K. and D. McFadden (1994).
\newblock Large sample estimation and hypothesis testing.
\newblock In R. F. Engle and D. McFadden (Eds.), {\em The Handbook of
Econometrics}, Volume 4, pp.\ 2111--2245. North-Holland.
\bibitem[\citeauthoryear{Pitt and Shephard}{Pitt and
Shephard}{1999}]{PittShephard1999}
Pitt, M. K. and N. Shephard (1999).
\newblock Filtering via simulation: Auxiliary particle filters.
\newblock {\em Journal of the American statistical association\/} {\em
94\/}(446), 590--599.
\bibitem[\citeauthoryear{Reifschneider and Williams}{Reifschneider and
Williams}{2000}]{ReifschneiderWilliams2000}
Reifschneider, D. and J. C. Williams (2000).
\newblock Three lessons for monetary policy in a low-inflation era.
\newblock {\em Journal of Money, Credit and Banking\/} {\em 32\/}(4), 936--966.
\bibitem[\citeauthoryear{Rigobon}{Rigobon}{2003}]{Rigobon03}
Rigobon, R. (2003).
\newblock Identification through heteroskedasticity.
\newblock {\em The Review of Economics and Statistics\/} {\em 85\/}(4),
777--792.
\bibitem[\citeauthoryear{Rossi}{Rossi}{2019}]{Rossi2019}
Rossi, B. (2019).
\newblock Identifying and estimating the effects of unconventional monetary
policy: How to do it and what have we learned?
\newblock {Discussion Paper} DP14064, {CEPR}.
\bibitem[\citeauthoryear{Smith and Blundell}{Smith and
Blundell}{1986}]{SmithBlundell1986}
Smith, R. J. and R. W. Blundell (1986).
\newblock An exogeneity test for a simultaneous equation tobit model with an
application to labor supply.
\newblock {\em Econometrica\/} {\em 54\/}(3), 679--685.
\bibitem[\citeauthoryear{Stock and Watson}{Stock and
Watson}{2001}]{StockWatson01}
Stock, J. H. and M. W. Watson (2001).
\newblock Vector autoregressions.
\newblock {\em Journal of Economic Perspectives\/} {\em 15\/}(4), 101--115.
\bibitem[\citeauthoryear{Swanson and Williams}{Swanson and
Williams}{2014}]{SwansonWilliams2014}
Swanson, E. T. and J. C. Williams (2014).
\newblock Measuring the effect of the zero lower bound on medium- and
longer-term interest rates.
\newblock {\em American Economic Review\/} {\em 104\/}(10), 3154--3185.
\bibitem[\citeauthoryear{Tobin}{Tobin}{1958}]{Tobin1958}
Tobin, J. (1958).
\newblock Estimation of relationships for limited dependent variables.
\newblock {\em Econometrica\/} {\em 26\/}(1), 24--36.
\bibitem[\citeauthoryear{Wu and Xia}{Wu and Xia}{2016}]{WuXia2016}
Wu, J. C. and F. D. Xia (2016).
\newblock Measuring the macroeconomic impact of monetary policy at the zero
lower bound.
\newblock {\em Journal of Money, Credit and Banking\/} {\em 48\/}(2-3),
253--291.
\bibitem[\citeauthoryear{Wu and Zhang}{Wu and
Zhang}{2019}]{WuZhang2019}
Wu, J. C. and J. Zhang (2019).
\newblock {A shadow rate New Keynesian model}.
\newblock {\em Journal of Economic Dynamics and Control\/} {\em 107}, 103728.