EconBase
← Back to paper

Identification at the Zero Lower Bound

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

169,959 characters · 14 sections · 78 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification at the Zero Lower Bound

frontmatter\runtitle{Identification at the Zero Lower Bound} \begin{aug} \address[id=add1]{ \orgdiv{Department of Economics}, \orgname{University of Oxford}} \end{aug} \support{This research is funded by the European Research Council via Consolidator grant number 647152. I would like to thank Guido Ascari, Sergio de Ferra, James Duffy, Andrea Ferrero, Jim Hamilton, Daisuke Ikeda, Federica Romei, Frank Schorheide, Francesco Zanetti and seminar participants at the Chicago Fed, New York Fed, Board of Governors, Bank of Japan, Hitotsubashi University, Keio University, UCL, University of Cambridge, University of Oxford, University of Warwick, DNB, University of Pennsylvania, Pennsylvania State University, NBER Summer Institute, IAAE conference, Netherlands Econometrics Study Group Meeting, World Congress and North American Winter meetings of the Econometric society for useful comments and discussion, as well as Mishel Ghassibe, Lukas Freund, Shangshang Li, and Patrick Vu for research assistance.} \begin{abstract} I show that the Zero Lower Bound (ZLB) on interest rates can be used to identify the causal effects of monetary policy. Identification depends on the extent to which the ZLB limits the efficacy of monetary policy. I propose a simple way to test the efficacy of unconventional policies, modelled via a `shadow rate'. I apply this method to U.S. monetary policy using a three-equation structural vector autoregressive model of inflation, unemployment and the federal funds rate. I reject the null hypothesis that unconventional monetary policy has no effect at the ZLB, but find some evidence that it is not as effective as conventional monetary policy. \end{abstract} \begin{keyword} \kwd{SVAR} \kwd{censoring} \kwd{coherency} \kwd{partial identification} \kwd{monetary policy} \kwd{shadow rate} \end{keyword}

Introduction

The zero lower bound (ZLB) on nominal interest rates has arguably been a challenge for policy makers and researchers of monetary policy. Policy makers have had to resort to so-called unconventional policies, such as quantitative easing or forward guidance, which had previously been largely untested. Researchers have to use new theoretical and empirical methodologies to analyze macroeconomic models when the ZLB\ binds. So, the ZLB is generally viewed as a problem or at least a nuisance. This paper proposes to turn this problem on its head to solve another long-standing question in macroeconomics:\ the identification of the causal effects of monetary policy on the economy.

The intuition is as follows. If the ZLB limits the ability of policy makers to react to macroeconomic shocks, as argued, for example, by EggertssonWoodford2003, the response of the economy to shocks will change when the policy instrument hits the ZLB. Because the difference in the behavior of macroeconomic variables across the ZLB and non-ZLB regimes is only due to the impact of monetary policy, the switch across regimes provides information about the causal effects of policy. In the extreme case that monetary policy completely shuts down during the ZLB regime, either because policy makers do not use alternative (unconventional) policy instruments, or because such instruments turn out to be completely ineffective, the only difference in the behavior of the economy across regimes is due to the impact of (conventional) policy during the unconstrained regime. Therefore, the ZLB identifies the causal effect of policy during the unconstrained regime. If monetary policy remains partially effective during the ZLB regime, e.g., through the use of unconventional policy instruments, then the difference in the behavior of the economy across regimes will depend on the difference in the effectiveness of conventional and unconventional policies. In this case, we obtain only partial identification of the causal effects of monetary policy, but we can still get informative bounds on the relative efficacy of unconventional policy. In the other extreme case that unconventional policy is as effective as conventional policy, there is no difference in the behavior of the economy across regimes, and we have no additional information to identify the causal effects of policy. However, we can still test this so-called ZLB irrelevance hypothesis DebortoliGaliGambetti2019 by testing whether the reaction of the economy to shocks is the same across the two regimes.

There are similarities between identification via occasionally binding constraints and identification through heteroskedasticity Rigobon03, or more generally, identification via structural change MagnussonMavroeidis2014. That literature showed that the switch between different regimes generates variation in the data that identifies parameters that are constant across regimes. For example, an exogenous shift in a policy reaction function or in the volatility of shocks identifies the transmission mechanism, provided the latter is unaffected by the policy shift. When the switch from one regime to another is exogenous, regime indicators are valid instruments, and the methodology in MagnussonMavroeidis2014 is applicable. However, regimes induced by occasionally binding constraints are not exogenous -- whether the ZLB binds or not clearly depends on the structural shocks, so regime indicators cannot be used as instruments in the usual way, and a new methodology is needed to analyze these models.

In this paper, I show how to control for the endogeneity in regime selection and obtain identification in structural vector autoregressions (SVARs).\footnote{There is a related literature on Dynamic Stochastic General Equilibrium (DSGE) models with a ZLB, see, e.g., FernandezGordonGuerronRubio2015, GuerrieriIacoviello2015, AruobaCubaBordaSchorfheide2017, KulishMorleyRobinson2017 and AruobaCubaBordaHigaFloresSchorfheideVillalvazo2020. The papers in this literature do not point out the implications of the ZLB for identification of monetary policy shocks.} The methodology is parametric and likelihood-based, and the analysis is similar to the well-known Tobit model Tobin1958. More specifically, the methodological framework builds on the early microeconometrics literature on simultaneous equations models with censored dependent variables, see Amemiya1974, Lee1976, BlundellSmith1994, and the more recent literature on dynamic Tobit models, see Lee1999, and particle filtering, see PittShephard1999.

A further contribution of this paper is a general methodology to estimate reduced-form VARs with a variable subject to an occasionally binding constraint. This is a necessary starting point for SVAR analysis that uses any of the existing popular identification schemes, such as short- or long-run restrictions, sign restrictions, or external instruments. In the absence of any constraints, reduced-form VARs can be estimated consistently by Ordinary Least Squares (OLS), which is Gaussian Maximum Likelihood, or its corresponding Bayesian counterpart, and inference is fairly well-established. However, it is well-known that OLS estimation is inconsistent when the data is subject to censoring or truncation, see, e.g., Gree93 for a textbook treatment. So, it is not possible to estimate a VAR consistently by OLS using any sample that includes the ZLB, or even using (truncated) subsamples when the ZLB is not binding (because of selection bias), as was pointed out by HayashiKoeda2019. It is not possible to impose the ZLB constraint using Markov switching models with exogenous regimes, as in LiuTheodoridisMumtazZanetti2019, because exogenous Markov-switching cannot guarantee that the constraint will be respected with probability one, and also does not account for the fact that the switch from one regime to the other depends on the structural shocks. Finally, it is not possible to perform consistent estimation and valid inference on the VAR (i.e., error bands with correct coverage on impulse responses), using externally obtained measures of the shadow rate, such as the one proposed by WuXia2016, as any such measures are subject to large and persistent estimation error that is not accounted for if they are treated as known in subsequent analysis. See also Rossi2019 for a comprehensive discussion of the challenges posed by the ZLB for the estimation of structural VARs.

The methodology developed in this paper allows for the presence of a shadow rate, estimates of which can be obtained, but more importantly, it fully accounts for the impact of sampling uncertainty in the estimation of the shadow rate on inference about the structural parameters such as impulse responses. Therefore, the paper fills an important gap in the literature, as it provides the requisite methodology to implement any of the existing identification schemes. HayashiKoeda2019 develop a VAR model with endogenous regime switching in which the policy variables that are subject to a lower bound are modelled using Tobit regressions. A key difference of their methodology from the one developed here is that they impose recursive identification of monetary policy shocks, which the present paper shows to be an overidentifying, and hence testable, restriction. Moreover, their model does not include shadow rates. A more recent paper by AruobaSchorfheideVillalvazo2020 also studies SVARs with occasionally binding constraints, but does not focus on the implications of these constraints for identification.

Identification of the causal effects of policy by the ZLB does not require that the policy reaction function be stable across regimes. However, inference on the efficacy of unconventional policy, or equivalently, the causal effects of shocks to the shadow rate over the ZLB period, obviously depends on whether or not the reaction function remains the same across regimes. For example, an attenuation of the causal effects of policy over the ZLB period may indicate that unconventional policy is only partially effective, but it is also consistent with unconventional policy being less active (during ZLB\ regimes) than conventional policy (during non-ZLB regimes). This is a fundamental identification problem that is difficult to overcome without additional information, such as measures of unconventional policy stance, or additional identifying assumptions, such as parametric restrictions or external instruments. This can be done using the methodology developed in this paper.

The structure of the paper is as follows. Section (ref) presents the main identification results of the paper in the context of a static bivariate simultaneous equations model with a limited dependent variable subject to a lower bound. Section (ref) generalizes the analysis to a SVAR\ with an occasionally binding constraint and discusses identification, estimation and inference. Section (ref) provides an application to a three-equation SVAR in inflation, unemployment and the Federal funds rate from StockWatson01. Using a sample of post-1960 quarterly US data, I find some evidence that the ZLB\ is empirically relevant, and that unconventional policy is only partially effective. Proofs and simulation results are given in the Appendix at the end.

Simultaneous equations model

I first illustrate the idea using a simple bivariate simultaneous equations model (SEM), which is both analytically tractable and provides a direct link to the related microeconometrics literature. To make the connection to the leading application, I will motivate this using a very stylized economy without dynamics in which the only outcome variable is inflation $\pi_{t}$ and the (conventional) policy instrument is the short-term nominal interest rate, $r_{t}$. In addition to the traditional interest rate channel, the model allows for an `unconventional monetary policy' channel that can be used when the conventional policy instrument hits the ZLB. An example of such a policy is quantitative easing (QE), in the form of long-term asset purchases by the central bank. Here I\ discuss a simple model of QE.\footnote{The more general SVAR\ model of the next section can also incorporate forward guidance in the form of ReifschneiderWilliams2000 and DebortoliGaliGambetti2019, as shown in IkedaLiMavroeidisZanetti2020.}

Abstracting from dynamics and other variables, the equation that links inflation to monetary policy is given by

equation[equation omitted — 114 chars of source]

where $c$ is a constant, $r^{n}$ is the neutral rate, $b_{L,t}$ is the amount of long-term bonds held by the private sector in log-deviation from its steady state, and $\varepsilon_{1t}$ is an exogenous structural shock unrelated to monetary policy. Equation ((ref)) can be obtained from a model of bond-market segmentation, as in ChenCurdiaFerrero2012, where a fraction of households is constrained to invest only in long-term bonds, see the Appendix for more details. In such a model, the parameter $\varphi$ that determines the effectiveness of QE is proportional to the fraction of constrained households and the elasticity of the term premium with respect to asset holdings, both of which are assumed to be outside the control of the central bank.

The nominal interest rate is set by a Taylor rule subject to the ZLB\ constraint, namely,

subequations\begin{align} r_{t} & =\max\left( r_{t}^{\ast},0\right) ,\\ r_{t}^{\ast} & =r^{n}+\gamma\pi_{t}+\varepsilon_{2t}, \end{align} where $r_{t}^{\ast}$ represents the desired target policy rate, and $\varepsilon_{2t}$ is a monetary policy shock. When $r_{t}^{\ast}$ is negative, it is unobserved. The unobserved $r_{t}^{\ast}$ will be referred to as the `shadow rate', and it represents the desired policy stance prescribed by the Taylor rule in the absence of a binding ZLB constraint. Suppose that QE is activated only when the conventional policy instrument $r_{t}$ hits the ZLB,\footnote{The assumption that QE is only active during the ZLB regime is only made for simplicity, as it is inconsequential for the resulting functional form of the transmission equation. We can let QE be active all the time, and even allow for a different rule for QE above and below the ZLB, i.e., $b_{L,t}=\alpha\min\left( r_{t}^{\ast },0\right) +\alpha_{1}\max\left( r_{t}^{\ast},0\right)$. Then, substituting back into ((ref)) yields an equation that is isomorphic to ((ref)), i.e., $\pi_{t}=\allowbreak c_{1}\allowbreak +\allowbreak\bar{\beta}r_{t}\allowbreak+\bar{\beta}^{\ast} \min\left( r_{t}^{\ast},0\right) \allowbreak+\varepsilon_{1t},$ with $\bar{\beta}:=\allowbreak\beta +\varphi\alpha_{1} $ and $\bar{\beta}^{\ast}\allowbreak:=\allowbreak\varphi\alpha.$} and follows the same policy rule ((ref)), up to a factor of proportionality $\alpha,$\ i.e.,

\[ b_{L,t}=\min\left( \alpha r_{t}^{\ast},0\right) . \] Substituting for $b_{L,t}$ in eq. ((ref)) and letting $\beta^{\ast}:=\alpha\varphi$, we obtain

equation[equation omitted — 152 chars of source]

A special case arises when QE is ineffective $\left( \varphi=0\right) ,$ or the monetary authority does not pursue a QE policy $\left( \alpha=0\right) ,$ so that eq. ((ref)) becomes

equation[equation omitted — 104 chars of source]

and monetary policy is completely inactive at the ZLB.

Another special case of the model given by equations ((ref)) and ((ref)) arises when $\beta^{\ast}=\beta$ in eq. ((ref)). This happens when $\varphi\neq0$ and $\alpha$ is chosen by the monetary authority to be equal to $\beta/\varphi.$ This can be done when policy makers know the transmission mechanism in eq. ((ref)) and have no restrictions in setting the policy parameter $\alpha$ so as to fully remove the impact of the ZLB on conventional policy. In that case, the equation for the outcome variable becomes

equation[equation omitted — 113 chars of source]

The model given by equations ((ref)) and ((ref)) is one in which monetary policy is completely unconstrained and there is no difference in outcomes across policy regimes. Such models have been put forward by SwansonWilliams2014, DebortoliGaliGambetti2019 and WuZhang2019.

The nesting model given by eq. ((ref)) allows the effects of conventional and unconventional policy to differ. This could reflect informational as well as political or institutional constraints that prevent policy makers from calibrating their unconventional policy response to match exactly the policy prescribed by the Taylor rule. For instance, it may be that policy makers do not know the effectiveness of the QE channel $\varphi$, or that the scale of asset purchases needed to achieve the desired policy stance during a ZLB regime is too large to be politically acceptable. Such a consideration may be particularly pertinent, for example, in the Eurozone. Importantly, one does not need to take a theoretical stand on this issue, because the methodology that I develop in the paper can accommodate a wide range of possibilities, and, as I demonstrate below, the issue can be studied empirically.

To complete the specification of the model, I\ assume that the structural shocks $\varepsilon_{t}=\left( \varepsilon_{1t},\varepsilon_{2t}\right) ^{\prime}$ are independently and identically distributed (i.i.d.) Normal with covariance matrix $\Sigma=diag\left( \sigma_{1}^{2},\sigma _{2}^{2}\right)$.

The analysis in this paper assumes that the shadow rate $r_t^*$ is only observed above the ZLB. If $r^*_t<0$ were observed up to scale, for instance, if we could measure QE $b_{L,t}$ from the balance sheet of the central bank, then the computation of the likelihood would be much simpler -- no filtering would be needed to deal with lags of $r_{t}^{\ast}<0$ on the right hand side of the SVAR model introduced in the next section, but the identification problem would remain the same. More generally, we could assume that $r_{t}^{\ast}<0$ is observed with some measurement error $\eta_{t}$, and include the measurement equation $b_{L,t}=\min\left( \alpha r_{t}^{\ast}+\eta_{t},0\right)$ in the model together with a specification of the distribution of $\eta_{t}$. The estimation method in this paper can then be seen as a special case where we are entirely agnostic about the measurement error. Adding such measures of unconventional policy is a potentially very useful extension of the method since, if correctly specified, they will likely improve estimation accuracy.

Connection to the microeconometrics literature Equations ((ref)) and ((ref)) form a SEM with a limited dependent variable. The special case with $\beta^{\ast}=0$ in ((ref)) can be referred to as a kinked SEM (KSEM), while the opposite case of $\beta^{\ast}=\beta$ in ((ref)) can be called a censored SEM (CSEM). Variants of the KSEM model have been studied in the early microeconometrics literature on limited dependent variable SEMs. Amemiya1974 and Lee1976 studied multivariate extensions of the well-known Tobit model Tobin1958. NelsonOlson1978 argued that the KSEM was less suitable for microeconometric applications than the CSEM, and the latter subsequently became the main focus of the literature SmithBlundell1986,BlundellSmith1989. BlundellSmith1994 studied the unrestricted model using external instruments, so they did not consider the implications of the kink for identification.

One important lesson from the microeconometrics literature is that establishing existence and uniqueness of equilibria in this class of models is non-trivial. GourierouxLaffontMonfort1980 define a model to be `coherent' if it has a unique solution for the endogenous variables in terms of the exogenous variables, i.e., if there exists a unique reduced form. More recently, the literature has distinguished between existence and uniqueness of solutions using the terms coherency and completeness of the model, respectively Lewbel2007. Establishing coherency and completeness is a necessary first step before we can study identification and estimation.

Identification

Substituting for $r_{t}^{\ast}$ in ((ref)) using ((ref)) and rearranging, we obtain

align[align omitted — 399 chars of source]

The system of equations ((ref)) and ((ref)) is now a KSEM, for which the necessary and sufficient condition for coherency and completeness (existence of a unique solution) is $\widetilde{\beta}\gamma<1$ NelsonOlson1978. Using ((ref)), the coherency and completeness condition can be expressed in terms of the structural parameters as

equation[equation omitted — 89 chars of source]

This condition evidently restricts the admissible range of the structural parameters. It is satisfied in the present monetary policy model, where it is natural to assume that $\beta,\beta^{\ast}\leq0$ and $\gamma>0$. Therefore, it is possible that the coherency condition may not provide additional information relative to what is often available from natural sign restrictions on the parameters.

Under condition ((ref)), the unique solution of the model can be written as

align[align omitted — 195 chars of source]

where $D_{t}:=1_{\left\{ r_{t}=0\right\} }$ is an indicator (dummy) variable that takes the value 1 when $r_{t}$ is on the boundary and zero otherwise, and

align[align omitted — 384 chars of source]

Equations ((ref)) and ((ref)) express the endogenous variables $\pi_{t},r_{t}$ in terms of the exogenous variables $\varepsilon _{1t},\varepsilon_{2t},$ and correspond to the decision rules of the agents in the model. It is clear that those decision rules differ in a world in which the ZLB occasionally binds, which is characterized by $\widetilde{\beta} \neq0,$ compared to a world in which it never does (i.e., the CSEM), where $\widetilde{\beta}=0.$ What is important for identification, however, is that in a world in which the ZLB occasionally binds, agents' reaction to shocks differs across regimes, and the difference depends on the parameter $\widetilde{\beta},$ which from eq. ((ref)), depends on the difference between the impact of conventional and unconventional policies, $\beta$ and $\beta^{\ast},$ respectively. I will show that this change provides information that identifies the structural parameters: we get point identification when $\beta^{\ast}=0$ (the KSEM case), and partial identification when $\beta^{\ast}\neq\beta.$ The identification argument leverages the coefficient on the kink, $\widetilde{\beta},$ in the `incidentally kinked' regression ((ref)), which is identified by a variant of the well-known Heckit method Heckman1979. I will sketch out the argument below, and provide more details for the full SVAR model in the next section.

Identification of the KSEM

Recall that in the KSEM\ model $\widetilde{\beta}=\beta.$ Consider the estimation of $\beta$ in ((ref)) from a regression of $\pi_{t}$ on $r_{t}$ using only observations above the ZLB,

equation[equation omitted — 244 chars of source]

where $\rho=cov\left( u_{1t},u_{2t}\right) /\tau^{2}-\beta=\gamma\sigma _{1}^{2}\left( 1-\gamma\beta\right) /\left( \gamma^{2}\sigma_{1}^{2} +\sigma_{2}^{2}\right) $, $\tau=\sqrt{var\left( u_{2t}\right) },$ and $\phi\left( \cdot\right) ,$ $\Phi\left( \cdot\right) $ are standard Normal density and distribution functions, respectively. The coefficient $\rho$ is the bias in the estimation of $\beta$ from the truncated regression ((ref)). Now, the mean of $\pi_{t}$ using the observations at the ZLB is

equation[equation omitted — 147 chars of source]

Next, observe that $\mu_{2},$ $\tau$ and hence $\phi\left( a\right) /\Phi\left( a\right) $ can be estimated from the Tobit regression ((ref)). Therefore, we can recover the bias $\rho$ and identify $\beta.$ A simple way to implement this is the control function approach Heckman1978. Let \[ h_{t}\left( \mu_{2},\tau\right) :=\left( 1-D_{t}\right) \left( r_{t} -\mu_{2}\right) -D_{t}\frac{\tau\phi\left( a\right) }{\Phi\left( a\right) }, \] and run the regression

equation[equation omitted — 133 chars of source]

where $c_{1}=c-\beta r^{n}$ is an unrestricted intercept. The rank condition for the identification of $\beta$ is simply that the regressors in ((ref)) are not perfectly collinear. This holds if and only if $0<\Pr\left( D_{t}=1\right) <1.$ So, as long as some but not all the observations are at the boundary, the model is generically identified.

Partial identification of the unrestricted SEM

The discussion of the previous subsection shows that $\widetilde{\beta}$ is identified from the kink in the reduced-form equation for $\pi_{t}$ ((ref)). It follows from eq. ((ref)) and the order condition that $\beta,\beta^*$ are not point identified. I will now demonstrate that they are partially identified.

The assumption $cov\left( \varepsilon_{1t},\varepsilon _{2t}\right) =0$ implies (see proof of Proposition 3 for a derivation)

equation[equation omitted — 106 chars of source]

where $\omega_{ij}:=cov\left(u_{it},u_{jt}\right)$. Substituting for $\gamma$ in ((ref)) using ((ref)) yields

equation[equation omitted — 170 chars of source]

For any given value of the reduced-form parameters $\widetilde{\beta},\Omega:=var(u_t)$, $u_t=(u_{1t},u_{2t})'$, the identified set for $\left( \beta,\beta^{\ast}\right)$ is a one-dimensional manifold in $\Re^{2}$ defined by eq. ((ref)) intersected with the coherency condition ((ref)).

It is instructive to illustrate the identified set graphically at some given value of $\widetilde{\beta}\,$and $\Omega$. Consider, for example, the case $\Omega$ equal to the identity $I_{2}$, at which ((ref)) yields $\gamma=-\beta,$ and the coherency condition ((ref)) reduces to $1+\beta\beta^{\ast}>0,$ and ((ref)) yields the function $\beta=\frac{\widetilde{\beta}+\beta^{\ast}}{1-\widetilde{\beta} \beta^{\ast}}$. Figure (ref) plots this function at $\widetilde{\beta}=-1/2$. and highlights in dark gray the region of incoherency defined by $1+\beta\beta^{\ast}\leq0$. The identified set is the part of the function $\beta=\frac{\widetilde{\beta}+\beta^{\ast} }{1-\widetilde{\beta}\beta^{\ast}}$ that lies to the right of the pole at $1/\widetilde{\beta},$ i.e., in the region $\beta^{\ast}>1/\widetilde{\beta }=-2$ in this example.

Now, consider the additional restrictions $\beta \geq\beta^{\ast}\geq0$ or $\beta\leq\beta^{\ast}\leq0$, highlighted by the light gray shaded areas in Figure (ref). The interpretation of those restrictions is that unconventional policy neither has the opposite effect from conventional policy, nor is it more effective than conventional policy. With this additional restriction, we see that the identified set further shrinks to the part of $\beta=\frac {\widetilde{\beta}+\beta^{\ast}}{1-\widetilde{\beta}\beta^{\ast}}$ in the interval $(\widetilde{\beta}^{-1},0].$ The projection of the identified set onto the $\beta$ axis yields $\beta\in(-\infty,\widetilde{\beta}],$ since $\widetilde{\beta}<0$, while its projection onto the $\beta^{\ast}$ axis yields $\beta^{\ast}\in$ $(\widetilde{\beta}^{-1},0]$. Because equations ((ref)) and ((ref)) encapsulate all the information in the reduced-form parameters about $\beta,\beta^{\ast}$, the identified set obtained from them is sharp.

figure[figure omitted — 868 chars of source]

Let us define a new parameter $\lambda$ such that $\beta^*=\lambda \beta$. The restriction indicated by the light gray areas in Figure (ref) corresponds to $\lambda\in\left[ 0,1\right]$. If we interpret $\lambda$ as a measure of the efficacy of unconventional policy, this restriction implies that unconventional policy is neither counter- nor over-productive. This reparameterization offers a convenient way to discretize the parameter space when we compute the identified set numerically, as is the case in the more general model discussed in the next section.

In this bivariate model, it is possible to characterize the identified set analytically. Here, I discuss the identified set for $\beta$ and defer the discussion of $\lambda$ to Appendix (ref).

We have already established that $\beta$ is completely unidentified when $\beta=\beta^*$ (equivalently $\lambda=1$), which corresponds to the CSEM. From the definition of $\widetilde{\beta}$ in ((ref)), it follows that $\beta=\beta^*$ implies $\widetilde{\beta}=0.$ So, when $\widetilde{\beta}=0$, $\beta$ is completely unidentified. It remains to see what happens when $\widetilde{\beta}\neq0$. Let $\gamma_{0}:=\omega_{12}/\omega_{11}$, which can be interpreted as the value the reaction function coefficient $\gamma$ in ((ref)) would take if $\beta=0$, i.e., the value corresponding to a Choleski identification scheme where $r_{t}$ is placed last. In Appendix (ref), I prove the following bounds

equation[equation omitted — 653 chars of source]

We see that when $\widetilde{\beta}\neq0$ and $0\leq\widetilde{\beta}\gamma_0 \leq 1$, we can identify both the sign of the causal effect $\beta$ of $r_{t}$ on $\pi_{t}$ and get bounds on its magnitude. In particular, the identified coefficient $\widetilde{\beta}$ is an attenuated measure of the true causal effect $\beta$. Moreover, $\widetilde{\beta}\gamma_{0}>1$ implies that $\beta^{\ast}$ has the opposite sign from $\beta$, i.e., unconventional policy has the opposite effect of the conventional one. That could be interpreted as saying that unconventional policy is counterproductive.

Finally, the hypothesis that unconventional policy is as effective as conventional policy, $\beta^* =\beta$ or $\lambda=1$, is equivalent to the null hypothesis $H_{0}:\widetilde{\beta }=0$. The alternative that unconventional policy is less effective than conventional policy, $\beta>\beta^{\ast}$ if $\beta>0,$ or $\beta<\beta^{\ast }$ if $\beta<0$, corresponds to the two-sided alternative $H_{1} :\widetilde{\beta}\neq0$. This can be tested using a likelihood ratio test.

SVAR with an occasionally binding constraint

I now develop the methodology for identification and estimation of SVARs with an occasionally binding constraint. Let $Y_{t}=\left( Y_{1t}^{\prime} ,Y_{2t}\right) ^{\prime}$ be a vector of $k$ endogenous variables, partitioned such that the first $k-1$ variables $Y_{1t}$ are unrestricted and the $k$th variable $Y_{2t}$ is bounded from below by $b$.\footnote{The lower bound does not need to be constant. All we need is to observe the periods in which the economy is at the ZLB\ regime.} Define the latent process $Y_{2t}^{\ast}$ that is only observed, and equal to $Y_{2t},$ whenever $Y_{2t}>b$. If $Y_{2t}$ is a policy instrument, $Y_{2t}^{\ast}$ can be thought of as the `shadow' instrument that measures the desired policy stance. The $p$th-order SVAR model is given by the equations

align[align omitted — 443 chars of source]

for $t\geq1$ given a set of initial values $Y_{-s},Y_{2,-s}^{\ast},$ for $s=0,...,p-1$, and $X_{0t}$ are exogenous and predetermined variables.

Equation ((ref)) can be interpreted as a policy reaction function because it determines the desired policy stance $Y_{2t}^{\ast}.$ Similarly, $\varepsilon_{2t}$ is the corresponding policy shock. The above model is a dynamic SEM. Two important differences from a standard SEM are the presence of (i) latent lags amongst the predetermined variables on the right-hand side, which complicates estimation; and (ii) the contemporaneous value of $Y_{2t}$ in the policy reaction function ((ref)), which allows it to vary across ZLB and non-ZLB regimes. The presence of latent lags $Y_{2,t-j}^{\ast}$ in the policy rule ((ref)) is particularly useful because it allows the model to incorporate forward guidance ReifschneiderWilliams2000,DebortoliGaliGambetti2019, see Appendix (ref) for details.

Collecting all the observed predetermined variables $X_{0t},Y_{t-1} ,...,Y_{t-p}$ into a vector $X_{t},$ and the latent lags $Y_{2,t-1}^{\ast },...,Y_{2,t-p}^{\ast}$ into $X_{t}^{\ast},$ and similarly for their coefficients, the model can be written compactly as:

align[align omitted — 300 chars of source]

The vector of structural errors $\varepsilon_{t}$ is assumed to be i.i.d. Normally distributed with zero mean and identity covariance.

In the previous section, we defined the KSEM as a special case of the general model, where $Y_{2t}^{\ast}<b$ has no (contemporaneous) impact on $Y_{1t}.$ In the dynamic setting, it feels natural to define the corresponding `kinked SVAR' model (KSVAR) as a model in which $Y_{2t}^{\ast}$ has neither contemporaneous nor dynamic effects. Therefore, the KSVAR obtains as a special case of ((ref)) when both $A_{12}^{\ast}=0,$ and $B^{\ast}=0,$ which corresponds to a situation in which the bound is fully effective in constraining what policy can achieve at all horizons.

The opposite extreme to the KSVAR is the censored SVAR model (CSVAR). Again, unlike the CSEM, which only characterizes contemporaneous effects, the idea of a CSVAR is to impose the assumption that the constraint is irrelevant at all horizons. So, it corresponds to a fully unrestricted linear SVAR in the latent process $\left( Y_{1t}^{\prime},Y_{2t}^{\ast}\right) ^{\prime}$. This is a special case of ((ref)) when both $A_{12}=0$ and the elements of $B$ corresponding to lagged $Y_{2t}$ are equal to zero. Finally, I refer to the general model given by ((ref)) as the\ `censored and kinked SVAR' (CKSVAR).

Define the $k\times k$ square matrices

equation[equation omitted — 254 chars of source]

$\overline{A}$ determines the impact effects of structural shocks during periods when the constraint does not bind. $A^{\ast}$ does the same for periods when the constraint binds.

To analyze the CKSVAR, we first need to establish existence and uniqueness of the reduced form. This is done in the following proposition.

propositionThe model given in eq. ((ref)) is coherent and complete (i.e., it has a unique solution) if and only if \begin{equation} \kappa:=\frac{\overline{A}_{22}-A_{21}A_{11}^{-1}\overline{A}_{12}} {A_{22}^{\ast}-A_{21}A_{11}^{-1}A_{12}^{\ast}}>0. \end{equation}

Note that ((ref)) does not depend on the coefficients on the lags (whether latent or observed), so it is exactly the same as in a static SEM. This condition is useful for inference, e.g., when constructing confidence intervals or posteriors, because it restricts the range of admissible values for the structural parameters. It can also be checked empirically when the structural parameters are point-identified.

If condition ((ref)) is satisfied, there exists a reduced-form representation of the CKSVAR model ((ref)). For convenience of notation, define the indicator (dummy variable) that takes the value one if the constraint binds and zero otherwise:

equation[equation omitted — 68 chars of source]
propositionIf ((ref)) holds, and for any initial values $Y_{-s},Y_{2,-s}^{\ast},$ $s=0,...,p-1,$ the reduced-form representation of ((ref)) for $t\geq1$ is given by \begin{align} Y_{1t} & =\overline{C}_{1}X_{t}+\overline{C}_{1}^{\ast}\overline{X} _{t}^{\ast}+u_{1t}-\widetilde{\beta}D_{t}\left( \overline{C}_{2} X_{t}+\overline{C}_{2}^{\ast}\overline{X}_{t}^{\ast}+u_{2t}-b\right) \\ Y_{2t} & =\max\left( \overline{Y}_{2t}^{\ast},b\right) , \\ \overline{Y}_{2t}^{\ast} & =\overline{C}_{2}X_{t}+\overline{C}_{2}^{\ast }\overline{X}_{t}^{\ast}+u_{2t},\\ Y_{2t}^{\ast} & =\left( 1-D_{t}\right) \overline{Y}_{2t}^{\ast} +D_{t}\left( \kappa\overline{Y}_{2t}^{\ast}+\left( 1-\kappa\right) b\right) , \end{align} where $u_{t}=\left( u_{1t}^{\prime},u_{2t}\right) ^{\prime}=\overline {A}^{-1}\varepsilon_{t},$ $\overline{C}^{\ast}=\left( \overline{C}_{1} ^{\ast\prime},\overline{C}_{2}^{\ast\prime}\right) ^{\prime}=\kappa \overline{A}^{-1}B^{\ast},$ $\overline{X}_{t}^{\ast}=\left( \overline {x}_{t-1},...,\overline{x}_{t-p}\right) ^{\prime},$ $\overline{x}_{t} =\min\left( \overline{Y}_{2t}^{\ast}-b,0\right) ,$ $\overline{x}_{-s} =\kappa^{-1}\min\left( Y_{2,-s}^{\ast}-b,0\right) ,$ $s=0,...,p-1,$ \begin{equation} \widetilde{\beta}=\left( A_{11}-A_{12}^{\ast}A_{22}^{\ast-1}A_{21}\right) ^{-1}\left( A_{12}^{\ast}A_{22}^{\ast-1}A_{22}-A_{12}\right) , \end{equation} $\kappa$ is defined in ((ref)) and the matrices $\overline {C}_{1},\overline{C}_{2},$ are given in eq. ((ref)) in the Appendix.

Note that the \textquotedblleft reduced-form\textquotedblright\ latent process $\overline{Y}_{2t}^{\ast}$ is, in general, different from the \textquotedblleft structural\textquotedblright\ shadow rate $Y_{2t}^{\ast}$ defined by ((ref)). They coincide only when $\kappa=1.$ This holds, for example, in the CSVAR model.

Equation ((ref)) combined with ((ref)) is a familiar dynamic Tobit regression model with the added complexity of latent lags included as regressors whenever $\overline{C}_{2}^{\ast}\neq0.$ Likelihood estimation of the univariate version of this model was studied by Lee1999. The $k-1$ equations ((ref)) are `incidentally kinked' dynamic regressions, that I\ have not seen analyzed before.

Identification

Identification of reduced-form parameters

Let $\psi$ denote the parameters that characterize the reduced form ((ref))-((ref)): $\widetilde{\beta},\overline{C},$ $\overline{C}^{\ast}$ and $\Omega=var\left( u_{t}\right) .$ It is useful to decompose $\psi$ into $\psi_{2}=\left( \overline{C}_{2},\overline{C} _{2}^{\ast},\tau\right) ^{\prime},$ where $\tau=\sqrt{var\left( u_{2t}\right) },$ and $\psi_{1}=( vec\left( \overline{C}_{1}\right) ^{\prime},vec\left( \overline{C}_{1}^{\ast}\right) ^{\prime} ,\allowbreak\widetilde{\beta}^{\prime},\delta^{\prime},vech\left( \Omega_{1.2}\right) ) ^{\prime},$ where $\delta=\Omega_{12}/\tau^{2}$, $\Omega_{1.2} =\Omega_{11}-\delta\delta^{\prime}\tau^{2}$, and $\Omega_{ij}=cov\left( u_{it},u_{jt}\right) $.

Equation ((ref)) is the dynamic Tobit regression model studied by Lee1999. So, its parameters, $\psi_{2},$ are generically identified provided that the regressors are not perfectly collinear. This requires that $0<\Pr\left( D_{t}=1\right) <1.$

Given $\psi_{2},$ the identification of the remaining parameters, $\psi_{1},$ can be characterized using a control function approach. Consider the $k-1$ regression equations

equation[equation omitted — 206 chars of source]

where

align[align omitted — 418 chars of source]

$a_{t}=\left( \frac{b-\overline{C}_{2}X_{t}-\overline{C}_{2}^{\ast} \overline{X}_{t}^{\ast}}{\tau}\right) $, and $\phi\left( \cdot\right) ,$ $\Phi\left( \cdot\right) $ are the standard Normal density and distribution functions, respectively. When $\overline{C}^{\ast}$ is different from zero, regressors $\overline{X}_{t}^{\ast}$, $Z_{1t},$ and $Z_{2t}$ in ((ref)) are unobserved, so we need to replace them with their expectations conditional on $Y_{2t},Y_{t-1},...,Y_{1}.$ Then, the regressors on the right-hand side of ((ref)) become $\mathbf{X}_{t}:=\left( X_{t}^{\prime},\overline{X}_{t|t}^{\ast\prime},Z_{1t|t},Z_{2t|t}\right) ^{\prime}$, where $h_{t|t}:=E\left( h\left( \overline{X}_{t}^{\ast}\right) |Y_{2t},Y_{t-1},...,Y_{1}\right) $ for any function $h\left( \cdot\right) $ whose expectation exists.\footnote{In the KSVAR model, we have $\overline {C}_{1}^{\ast}=0$ and $\overline{C}_{2}^{\ast}=0,$ so $\overline{X}_{t}^{\ast }$ drops out of ((ref)), and the regressors $Z_{1t},Z_{2t}$ are observed, so $Z_{jt|t}=Z_{jt},$ $j=1,2.$} The coefficients $\overline{C} _{1},\overline{C}_{1}^{\ast},\widetilde{\beta},$ and $\delta$ are generically identified if the regressors $\mathbf{X}_{t}$ are not perfectly collinear.

Identification of structural parameters

From the order condition, we can easily establish that there are not enough restrictions to identify all the structural parameters in the CKSVAR\ ((ref)). Let $k_{0}=\dim\left( X_{0t}\right) $ denote the number of predetermined variables other than the own lags of $Y_{t}.$ For example, in a standard VAR without deterministic trends, we have $X_{0t}=1,$ so $k_{0}=1.$ The number of reduced-form parameters $\psi$ is $k_{0}k+k^{2}p$ (in $\overline{C}$) plus $kp$ (in $\overline{C}^{\ast}$) plus $k-1$ (in $\widetilde{\beta}$) plus $k\left( k+1\right) /2$ (in $\Omega$). The number of structural parameters in ((ref)) is $k_{0}k+k^{2}p$ (in $B$) plus $kp$ (in $B^{\ast}$) plus $k^{2}$ (in $\overline{A}$) plus $k$ (in $A_{12}^{\ast}$ and $A_{22}$). So, the CKSVAR is underidentified by $k\left( k-1\right) /2+1$ restrictions. Nevertheless, I will show that the impulse responses to $\varepsilon_{2t}$ are identified. Specifically, they are point-identified when $A_{12}^{\ast}=0,$ and partially identified when $A_{12}^{\ast}\neq0$ but $A_{12}^{\ast}$ and $A_{12}$ have the same sign, analogous to the bounds given in equation ((ref)) in the previous section.

Because the CKSVAR is nonlinear, IRFs are obviously state-dependent, and there are many ways one can define them, see KoopPesaranPotter1996. The IRF to $\varepsilon_{2t},$ according to any of the definitions proposed in the literature, is identified if the reduced-form errors $u_{t}$ can be expressed as a known function of $\varepsilon_{2t}$ and a process that is orthogonal to it, i.e., $u_{t}=g\left( \varepsilon_{2t},e_{t}\right) ,$ where $e_{t}$ is independent of $\varepsilon_{2t}.$ From Proposition (ref), it follows that the function $g$ is linear, and more specifically,

align[align omitted — 361 chars of source]

where

align[align omitted — 289 chars of source]

and $\overline{A}_{22}=A_{22}^{\ast}+A_{22},$ defined in ((ref)). Note that $\overline{\beta}$ can be interpreted as the response of $Y_{1t}$ to a shock that increases $Y_{2t}$ by one unit, and $\overline{\gamma}$ are the contemporaneous reaction function coefficients of $Y_{2t}$ to $Y_{1t}$ when $Y_{2t}>b$ (unconstrained regime). The shock vector $\overline{\varepsilon}_{1t}$ is not structural but it is orthogonal to $\varepsilon_{2t},$ so it plays the role of $e_{t}$ in $u_{t}=g\left( \varepsilon_{2t},e_{t}\right) .$ Hence, the IRF\ is identified if and only if $\overline{\beta},$ $\overline{\gamma},$ and $\overline{A}_{22}$ are identified.

The following proposition shows identification when $A_{12}^{\ast}=0$.

propositionWhen $A_{12}^{\ast}=0$ and the coherency condition ((ref)) holds, the parameters in ((ref))-((ref)) are identified by the equations $\overline{\beta}=\widetilde{\beta},$ \begin{align} \overline{\gamma} & =\left( \Omega_{12}^{\prime}-\Omega_{22}\overline {\beta}^{\prime}\right) \left( \Omega_{11}-\Omega_{12}\overline{\beta }^{\prime}\right) ^{-1}, and\\ \overline{A}_{22}^{-1} & =\sqrt{\left( -\overline{\gamma},1\right) \Omega\left( -\overline{\gamma},1\right) ^{\prime}}. \end{align}

Remarks 1. $\overline{\beta}=\widetilde{\beta}$ follows immediately from the definition ((ref)) with $A_{12}^{\ast}=0$. Equations ((ref)) and ((ref)) hold without the restriction $A_{12}^{\ast}=0.$ They follow from the orthogonality of the shocks $\varepsilon_{2t}$ and $\overline{\varepsilon}_{1t}.$

2. An instrumental variables interpretation of this identification result is as follows. Define the instrument \[ Z_{t}:=Y_{1t}-\widetilde{\beta}Y_{2t}=A_{11}^{-1}B_{1}X_{t}+A_{11}^{-1} B_{1}^{\ast}X_{t}^{\ast}+A_{11}^{-1}\varepsilon_{1t}, \] where the second equality holds when $A_{12}^{\ast}=0$. The orthogonality of the errors $E\left( \varepsilon_{1t}\varepsilon_{2t}\right) =0$ implies $E\left( Z_{t}\varepsilon_{2t}\right) =0.$ So, $Z_{t}$ are valid $k-1$ instruments for the $k-1$ endogenous regressors $Y_{1t}$ in the structural equation of $Y_{2t}=\max\left( Y_{2t}^{\ast},b\right) $, where $Y_{2t} ^{\ast}$ is given by ((ref)). Normalizing ((ref)) in terms of $Y_{2t}^{\ast}$ yields the structural equation in the more familiar form of a policy rule:

equation[equation omitted — 162 chars of source]

where $\bar{B}_{2}=$ $\overline{A}_{22}^{-1}B_{2},$ $\bar{B}_{2}^{\ast}=$ $\overline{A}_{22}^{-1}B_{2}^{\ast}$. Since $A_{11}^{-1}$ is non-singular, the\ coefficient matrix of $Z_{t}$ in the `first-stage' regressions of $Y_{1t}$ is nonsingular, so the coefficients of ((ref)) are generically identified by the rank condition. An alternative to the Tobit IV regression model ((ref)) is the indirect Tobit regression approach used in the static SEM by BlundellSmith1994. Equation ((ref)) can be written as the dynamic Tobit regression

equation[equation omitted — 168 chars of source]

where $\widetilde{\gamma}=\left( 1-\overline{\gamma}\overline{\beta}\right) ^{-1}\overline{\gamma},$ $\tilde{B}_{2}=\left( 1-\overline{\gamma} \overline{\beta}\right) ^{-1}\bar{B}_{2},$ $\tilde{B}_{2}^{\ast}=\left( 1-\overline{\gamma}\overline{\beta}\right) ^{-1}\bar{B}_{2}^{\ast}$ and $\tilde{\varepsilon}_{2t}=\left( 1-\overline{\gamma}\overline{\beta}\right) ^{-1}\bar{\varepsilon}_{2t}.$ Note that the coherency condition ((ref)) becomes $\kappa=\frac{\overline{A}_{22}}{A_{22}^{\ast} }\left( 1-\overline{\gamma}\overline{\beta}\right) >0,$ so $1-\overline {\gamma}\overline{\beta}\neq0,$ which guarantees the existence of the representation ((ref)). Given $\overline{\beta}=\widetilde{\beta },$ the structural parameter $\overline{\gamma}$ can then be obtained as $\overline{\gamma}=\widetilde{\gamma}\left( I_{k-1}+\widetilde{\beta }\widetilde{\gamma}\right) ^{-1},$ and similarly for the remaining structural parameters in ((ref)).

3. The parameter $A_{22}$ allows the reaction function of $Y_{2t}^{\ast}$ to differ across the two regimes. The special case $A_{22}=0$ thus corresponds to the restriction that the reaction function remains the same across regimes. The parameters $A_{22}$ and $A_{22}^{\ast}$ are not separately identified. Hence, $A_{22}^{\ast-1}$, the scale of the response to the shock $\varepsilon_{2t}$ during periods when $Y_{2t}=b$,$\,$ is not identified.\footnote{This is akin to the well-known property of a probit model that the scale of the distribution of the latent process is not identifiable.} Similarly, $\kappa=\frac{\overline{A}_{22}}{A_{22}^{\ast}}\left( 1-\overline{\gamma}\overline{\beta}\right) $ is not identified, and therefore, neither is the structural shadow value $Y_{2t}^{\ast}$ in eq. ((ref)). Identification of these requires an additional restriction on $A_{22}$, e.g., $A_{22}=0.$ Turning this discussion around, we see that a change in the reaction function across regimes does not destroy the point identification of the effects of policy during the unconstrained regime, since the latter only requires $\overline{\beta},\overline{\gamma}$ and $\overline{A}_{22}$, not $A_{22}^{\ast}$ or $\kappa.$

Next, we turn to the case $A_{12}^{\ast}\neq0,$ and derive identification under restrictions on the sign and magnitude of $A_{12}^{\ast}$ relative to $A_{12}$ and $A_{22}^{\ast}$ relative to $A_{22}.$ The first restriction is motivated by a generalization of the discussion on the SEM model in Subsection (ref). Specifically, if $\overline{A}_{12}=A_{12}+A_{12}^{\ast}$ measures the effect of conventional policy (operating in the unconstrained regime) and $A_{12}^{\ast}$ measures the effect of unconventional policy (operating in the constrained regime), then the assumption that $A_{12}$ and $A_{12}^{\ast}$ have the same sign means that unconventional policy effects are neither in the opposite direction nor larger in absolute value than conventional policy effects. In other words, unconventional policy is neither counterproductive nor over-productive relative to conventional policy. This can be characterized by the specification $A_{12}^{\ast}=\Lambda\overline {A}_{12}$ and $A_{12}=\left( I_{k-1}-\Lambda\right) \overline{A}_{12},$ where $\Lambda=diag\left( \lambda_{j}\right) ,$ $\lambda_{j}\in\left[ 0,1\right] $ for $j=1,...,k-1.$ I further impose the restriction that $\lambda_{j}=\lambda$ for all $j,$ so that $A_{12}^{\ast}=\lambda\overline {A}_{12}$ and $A_{12}=\left( 1-\lambda\right) \overline{A}_{12}$ with $\lambda\in\left[ 0,1\right] .$ This, in turn, means that $Y_{2t}$ and $Y_{2t}^{\ast}$ enter each of the first $k-1$ structural equations for $Y_{1t}$ only via the common linear combination $\lambda Y_{2t}^{\ast }\allowbreak+\allowbreak\left( 1-\lambda\right) Y_{2t},$ which can be interpreted as a measure of the effective policy stance.

We also need to consider the impact of $A_{22}$ on identification. The parameter $\zeta=\overline{A}_{22}/A_{22}^{\ast}$ gives the ratio of the standard deviation of the monetary policy shock in the constrained relative to the unconstrained regime. It is also the ratio of the reaction function coefficients in the two regimes, e.g., $A_{22}^{\ast-1}A_{21}$ versus $\overline{A}_{22}^{-1}A_{21}$. I will impose $\zeta>0,$ so that the sign of the policy shock does not change across regimes. With the above reparametrization and the definitions in ((ref)), the identified coefficient $\widetilde{\beta}$ in ((ref)) can be written as

equation[equation omitted — 180 chars of source]

Similarly, given $\zeta>0,$ the coherency condition ((ref)) reduces to $\left( 1-\overline{\gamma}\overline{\beta}\right) \left( 1-\xi\overline{\gamma}\overline{\beta}\right) >0.$ Notice that the parameters $\lambda,\zeta$ only appear multiplicatively, so it suffices to consider them together as $\xi=\lambda\zeta.$ Once $\overline{\beta}$ is known, the remaining structural parameters needed to obtain the IRF to $\varepsilon_{2t}$ are $\overline{\gamma}$ and $\overline{A}_{22}$, and they are obtained from Proposition (ref). So, the identified set can be characterized by varying $\xi$ over its admissible range. Without further restrictions on $\zeta,$ the admissible range is obviously $\xi\geq0.$ If we further assume that $\zeta\leq1,$ i.e., that the slope of the reaction function coefficients is no steeper in the constrained regime than in the unconstrained regime, then $\xi\in\left[ 0,1\right] ,$ and so partial identification proceeds exactly along the lines of the SEM in the previous section where $\lambda$ played the role of $\xi.$ In the case $k=2,$ the bounds derived in eq. ((ref)) apply, with $\beta=\overline{\beta}$ in the notation of the present section. However, when $k>2,$ it is difficult to obtain a simple analytical characterization of the identified set for $\overline{\beta}.$ In any case, we will typically wish to obtain the identified set for functions of the structural parameters, such as the IRF. This can be done numerically by searching over a fine discretization of the admissible range for $\xi.$ An algorithm for doing this is provided in Appendix (ref).

Estimation

Estimation of the CKSVAR is carried out by Maximum Likelihood (ML) using either a version of the sequential importance sampler (SIS) of Lee1999 or the fully adapted particle filter (FAPF) of MalikPitt2011 to evaluate the likelihood, except in the case of the KSVAR model for which the likelihood is available analytically. The details are given in Appendix (ref).

Using the limit theory of Newe94Mc, the ML estimator can be shown to be consistent and asymptotically Normal and the LR statistic asymptotically $\chi^{2}$ with degrees of freedom equal to the number of restrictions. Standard asymptotics arise when the probability of each regime occurring is bounded away from zero. Infrequent visits to one of the two regimes will slow down the rate of convergence of the estimator, but will not lead to a non-standard limiting distribution. Since the focus of this paper is on identification, I will not discuss primitive conditions for these results, such as geometric ergodicity, which can be shown, for example, by bounding the joint spectral radius of the companion-form representation of the model Liebscher2005. Instead, I report Monte Carlo simulation results on the finite-sample properties of ML estimators and LR tests in Appendix (ref). They show that the Normal distribution provides a very good approximation to the finite-sample distribution of the ML estimators. I find some finite-sample size distortion in the LR tests of various restrictions on the CKSVAR, but this can be addressed effectively with a parametric bootstrap, as shown in the Appendix.

One interesting observation from the simulations is that the LR test of the CSVAR restrictions against the CKSVAR appears to be less powerful than the corresponding test of the KSVAR restrictions against the CKSVAR. Thus, we expect to be able to detect deviations from KSVAR more easily than deviations from CSVAR. In other words, finding evidence against the hypothesis that unconventional policies are fully effective (CSVAR) may be harder than finding evidence against the opposite hypothesis that they are completely ineffective (KSVAR).

Application

I use the three-equation SVAR of StockWatson01, consisting of inflation, the unemployment rate and the Federal Funds rate to provide a simple empirical illustration of the methodology developed in this paper. As discussed in StockWatson01, this model is far too limited to provide credible identification of structural shocks, so the results in this section are meant as an illustration of the new methods.

The data are quarterly and are constructed exactly as in StockWatson01 .\footnote{The inflation data are computed as $\pi_{t}=$ $400ln(P_{t} /P_{t-1})$, where $P_{t}$, is the implicit GDP deflator and $u_{t}$ is the civilian unemployment rate. Quarterly data on $u_{t}$ and $i_{t}$ are formed by taking quarterly averages of their monthly values.} The variables are plotted in Figure (ref) over the extended sample 1960q1 to 2018q2. I will consider all periods in which the Fed funds rate was below 20 basis points to be on the ZLB. This includes 28 quarters, or 11% of the sample.

figure[figure omitted — 215 chars of source]

Tests of efficacy of unconventional policy

I estimate three specifications of the SVAR(4) with the ZLB: the unrestricted CKSVAR specification, as well as the restricted KSVAR and CSVAR specifications. The maximum log-likelihood for each model is reported in Table (ref), computed using the SIS\ algorithm in the case of CKSVAR and CSVAR, with 1000 particles. The accuracy of the SIS algorithm was gauged by comparing the log-likelihood to the one obtained using the resampling FAPF algorithm. In both CKSVAR and CSVAR the difference is very small. The results are also very similar when we increase the number of particles to 10000. Finally, the table reports the LR tests of KSVAR and CSVAR against CKSVAR using both asymptotic and parametric bootstrap p-values.

table[table omitted — 1,582 chars of source]

The KSVAR imposes the restriction that no latent lags (i.e., lags of the shadow rate) should appear on the right hand side of the model, i.e., $B^{\ast}=0$ in ((ref)) or $\overline{C}_{1}^{\ast}=0$ and $\overline{C}_{2}^{\ast}=0$ in ((ref)) and ((ref)). This amounts to 12 exclusion restrictions on the CKSVAR(4), four restrictions in each of the three equations. This is necessary (but not sufficient) for the hypothesis that unconventional policy is completely ineffective at all horizons. It is necessary because $\overline{C}^{\ast}=\left( \overline {C}_{1}^{\ast\prime},\overline{C}_{2}^{\ast\prime}\right) ^{\prime}\neq0$ would imply that unconventional policy has at least a lagged effect on $Y_{t}.$ $\overline{C}^{\ast}=0$ is not sufficient to infer that unconventional policy is completely ineffective because it may still have a contemporaneous effect on $Y_{1t}$ if $A_{12}^{\ast}\neq0,$ and the latter is not point-identified. The result of the test in Table (ref) shows that lags of the shadow rate are statistically significant at the 5% level, rejecting the null hypothesis that unconventional policy has no effect.

The CSVAR model imposes the restriction that only the coefficients on the lags of the shadow rate (which is equal to the actual rate above the ZLB) are different from zero in the model, i.e., the elements of $B$ corresponding to lags of $Y_{2t}$ in ((ref)) are all zero, or equivalently, the elements of $\overline{C}$ corresponding to lags of $Y_{2t}$ in ((ref)) and ((ref)) are all equal to $\overline{C}^{\ast} $. In addition, it imposes the restriction that $\widetilde{\beta}=0$ in ((ref)), i.e., no kink in the reduced-form equations for inflation and unemployment across regimes, yielding 14 restrictions in total. This is necessary for the hypothesis that the ZLB\ is empirically irrelevant for policy in that it does not limit what monetary policy can achieve. The evidence against this hypothesis is not as strong as in the case of the KSVAR. The asymptotic p-value is 0.023, indicating rejection of the null hypothesis that unconventional policy is as effective as conventional policy at the 5% level, but the bootstrap p-value is 0.117. Note that this difference could also be due to fact that the test of the CSVAR restrictions may be less powerful than the test of the KSVAR restrictions, as indicated by the simulations reported in the previous section. Thus, I\ would cautiously conclude that the evidence on the empirical relevance of the ZLB\ is mixed. Further evidence on the efficacy of unconventional policy will also be provided in the next subsection.

Impulse response functions

Based on the evidence reported in the previous section, I estimate the IRFs associated with the monetary policy shock using the unrestricted CKSVAR specification, and compare them to recursive IRFs from the CSVAR specification that place the Federal funds rate last in the causal ordering. From the identification results in Section (ref), the CKSVAR point-identifies the nonrecursive IRFs only under the assumption that the shadow rate has no contemporaneous effect of $Y_{1t},$ i.e., $A_{12}^{\ast}=0$ in ((ref)). Note that, due to the nonlinearity of the model, the IRFs are state-dependent. I use the following definition of the IRF from KoopPesaranPotter1996:\footnote{The kink in the reduced-form representation of the model makes it difficult to approximate the IRFs by local projections on simple nonlinear functions of the data, such as powers or interactions with the regime indicator, see Appendix (ref) for further discussion of this point.}

equation[equation omitted — 254 chars of source]

Figure (ref) reports the nonrecursive IRFs to a 25 basis points monetary policy shock at the end of the sample, 2018q3 (at which point $\overline{X}_{t}^{\ast}=0$ in ((ref)) because interest rates had been above the ZLB over the previous four quarters), from the CKSVAR under the assumption that unconventional policy has no contemporaneous effect ($\lambda=0$). It also reports two different estimates of recursive IRFs using the identification scheme in StockWatson01 with interest rates placed last. The first estimate is obtained from the CSVAR specification, and the second is a “naive” OLS estimate of the IRF that ignores the ZLB constraint -- a direct application of the method in StockWatson01 to the present sample. The figure also reports 90% bootstrap error bands for the nonrecursive IRFs.

figure[figure omitted — 768 chars of source]

In the nonrecursive IRF, the response of inflation to a monetary tightening is negative on impact, albeit very small, and, with the exception of the first quarter when it is positive, it stays negative throughout the horizon. Hence, the incidence of a price puzzle is mitigated relative to the recursive IRFs, according to which inflation rises for up to 6 quarters after a monetary tightening (9 quarters in the OLS case). Note, however, that the error bands are so wide that they cover (pointwise) most of the recursive IRF, though less so for the OLS one. Turning to the unemployment response, we see that the nonrecursive IRF starts significantly positive on impact (no transmission lag) and peaks much earlier (after 4 quarters) than the recursive IRF (10 quarters). In this case, the recursive IRF\ is outside the error bands for several quarters (more so for the naive OLS IRF). Finally, the response of the Federal funds rate to the monetary tightening is less than one on impact and generally significantly lower than the recursive IRFs. This is both due to the contemporaneous feedback from inflation and unemployment, as well as the fact that there is a considerable probability of returning to the ZLB, which mitigates the impact of monetary tightening.

Next, I turn to the identified sets of the IRFs that arise when I relax the restriction that unconventional policy is ineffective, i.e., $\lambda$ can be greater than zero. I consider the range of $\xi=\lambda\zeta\in\left[ 0,1\right] ,$ recalling that $\lambda$ measures the efficacy of unconventional policy and $\zeta$ measures the ratio of the reaction function coefficients and shock volatilities in the constrained versus the unconstrained regimes. The shaded areas in Figure (ref) report the identified sets without any other restrictions. The striped areas (a subset of the aforementioned identified sets) show the tightening of the identified sets when I impose the additional sign restriction that the contemporaneous effect of the monetary policy shock to the Fed Funds rate should be nonnegative. The bold lines show the IRFs under the (point-identifying) assumption $\lambda=0$. The latter are the same as the nonrecursive point estimates reported in Figure (ref).

We observe that the identified set for the IRF of inflation is bounded from above by the limiting case $\lambda=0$. This is also true of the response of the Fed Funds rate. The case $\lambda=0$ provides a lower bound on the effect to unemployment only from 0 to 9 quarters. Even though the point estimate of the unemployment response under $\lambda=0$ remains positive over all horizons, the identified set includes negative values beyond 10 quarters ahead. We also notice that the identified sets are fairly large, albeit still informative. Interestingly, the identified IRF of the Fed Funds rate includes a range of negative values on impact. These values arise because for values of $\xi>0$, there are generally two solutions for the structural VAR parameters $\overline{\beta},\overline{\gamma}$ in the equations ((ref)), ((ref)), with one of them inducing such strong responses of inflation and unemployment to the interest rate that the contemporaneous feedback in the policy rule would in fact revert the direct positive effect of the policy shock on the interest rate. If we impose the additional sign restriction that the contemporaneous impact of the policy shock to the Fed Funds rate must be non-negative, then those values are ruled out and the identified sets become considerably tighter. This is an example of how sign restrictions can lead to tighter partial identification of the IRF.

figure[figure omitted — 563 chars of source]

With an additional assumption on $\zeta,$ the method can be used to obtain an estimate of the identified set for $\lambda,$ the measure of the efficacy of unconventional policy. In particular, if we set $\zeta=1,$ i.e., the reaction function remains the same across the two regimes, then the identified set for $\lambda$ is $\left[ 0.0.506\right] .$ In other words, the identified set excludes values of the efficacy of policy beyond 51%, so that, roughly speaking, unconventional policy is at most 51% as effective as conventional one. Note that this estimate does not account for sampling uncertainty and relies crucially on the assumption that the reaction function remains the same across the two regimes. This assumption could be justified by arguing that there is no reason to believe that policy objectives may have shifted over the ZLB period, and that any desired policy stance was feasible over that period. The latter assumption may be questionable. For example, one can imagine that there may be financial and political constraints on the amount of quantitative easing policy makers could do, which may cause them to proceed more cautiously over the ZLB period than over regular times. Within the context of our model, this would be reflected as a flatter policy reaction function over the ZLB period than over the non-ZLB periods, i.e., it will correspond to $\zeta<1$. To illustrate the implications of this for the identification of $\lambda$, suppose that $\zeta=1/2,$ i.e., the shadow rate reacts half as fast to shocks during the ZLB period than it does in the non-ZLB period. Then, the identified set for $\lambda$ would include 1, i.e., the data would be consistent with the view that unconventional policy is fully effective. So, under this alternative assumption on $\zeta$, the reason we observed a subdued response to policy shocks over the ZLB period is because policy was less active over that period, and policy shocks were smaller, not because unconventional policy was partially ineffective.

As I discussed in the introduction, it is difficult to make further progress on this issue without further information or additional assumptions. The technical reason is that the scale of the latent regression over the censored sample is not identified, so additional information is required to untangle the structural parameters $\lambda$ and $\zeta$ from $\xi=\lambda\zeta$. One possibility would be to identify $\lambda$ from the coefficients on the lags of $Y_{2t}$ and $Y_{2t}^{\ast}$ by imposing the (overidentifying) restriction that $Y_{2t-j}$ and $Y_{2t-j}^{\ast}$ appear in the model only via the linear combination $Y_{2t-j}^{\text{eff}}:=\lambda Y_{2t-j}^{\ast}\allowbreak +\allowbreak\left( 1-\lambda\right) \allowbreak Y_{2t-j}$ for all lags $j=1,...,p,$ where $Y_{2t}^{\text{eff}}$ can be interpreted as the effective policy stance. Provided that the coefficients on the lags of $Y_{2t}^{\ast}$ or $Y_{2t}$ are not all zero, this restriction point identifies $\lambda,$ and hence, partially identifies $\zeta$ from $\xi.$ One could obtain tighter bounds by using sign restrictions IkedaLiMavroeidisZanetti2020, or obtain point identification by using conventional identification schemes. For instance, one can identify $\overline{\beta}$ directly using external instruments, as in GertlerKaradi2015, and hence point identify $\xi$ from ((ref)).

Conclusion

This paper has shown that the ZLB can be used constructively to identify the causal effects of monetary policy on the economy. Identification relies on the (in)efficacy of alternative (unconventional) policies. When unconventional policies are partially effective in mitigating the impact of the ZLB, the causal effects of monetary policy are only partially identified. A general method is proposed to estimate SVARs subject to an occasionally binding constraint. The method can be used to test the efficacy of unconventional policy, modelled via a shadow rate. Application to a core three-equation SVAR with US data suggests that the ZLB is empirically relevant and unconventional policy is only partially effective.

appendix\section{Proofs} \subsection{Derivation of identified set for model of Section (ref)} Using the notation $\beta^* = \lambda\beta$, eq. ((ref)) can be expressed as \begin{equation} \widetilde{\beta}=g\left( \beta\right) \beta,\quadg\left( \beta\right) :=\frac{1-\lambda}{1-\lambda\frac{\beta\left( \omega _{12}-\omega_{22}\beta\right) }{\omega_{11}-\omega_{12}\beta}}. \end{equation} When $\omega_{12}=0,$ we have $g\left( \beta\right) =$ $\frac{1-\lambda }{1+\lambda\omega_{22}\beta^{2}/\omega_{11}}\in\left( 0,1\right) $ for all $\lambda\in\left( 0,1\right) .$ Therefore, when $\widetilde{\beta}\neq0,$ the sign of $\beta$ is the same as that of $\widetilde{\beta}$ and its magnitude is lower, as stated in ((ref)). Next, consider $\omega_{12}\neq0.$ It is easily seen that $g\left( 0\right) =1-\lambda$ and $\lim_{\beta\rightarrow\pm\infty}\left( \beta\right) =0.$ Moreover, \[ \frac{\partial g}{\partial\beta}=\lambda\left( 1-\lambda\right) \frac {\omega_{12}\omega_{22}\beta^{2}-2\omega_{11}\omega_{22}\beta+\omega _{11}\omega_{12}}{\left( \omega_{11}-\beta\omega_{12}-\beta\lambda\omega _{12}+\beta^{2}\lambda\omega_{22}\right) ^{2}}. \] For $\lambda\in\left( 0,1\right) ,$ the above derivative function has zeros at $ \omega_{12}\omega_{22}\beta^{2}-2\omega_{11}\omega_{22}\beta+\omega_{11} \omega_{12}=0$, which occur at \[ \begin{array} [c]{c} \beta_{1}=\frac{\omega_{11}\omega_{22}+\sqrt{\omega_{11}\omega_{22}\left( \omega_{11}\omega_{22}-\omega_{12}^{2}\right) }}{\omega_{12}\omega_{22}}\\ \beta_{2}=\frac{\omega_{11}\omega_{22}-\sqrt{\omega_{11}\omega_{22}\left( \omega_{11}\omega_{22}-\omega_{12}^{2}\right) }}{\omega_{12}\omega_{22}} \end{array} ,\quad\text{if }\omega_{12}\neq0. \] Now, because $0<\left( \omega_{11}\omega_{22}-\omega_{12}^{2}\right) <\omega_{11}\omega_{22}$ implies $\sqrt{\omega_{11}\omega_{22}\left( \omega_{11}\omega_{22}-\omega_{12}^{2}\right) }<\omega_{11}\omega_{22},$ we have $\beta_{i}<0,$ $i=1,2,$ when $\omega_{12}<0$ and $\beta_{i}>0,$ $i=1,2,$ when $\omega_{12}>0.$ By symmetry, it suffices to consider only one of the two cases, e.g., the case $\omega_{12}<0.$ In this case, $g^{\prime}\left( \beta\right) =\frac {\partial g}{\partial\beta}<0$ for all $\beta>0$ and, since $g\left( 0\right) =1-\lambda$ and $g\left( \infty\right) =0,$ it follows that $g\left( \beta\right) \in\left( 0,1-\lambda\right) $ for all $\beta>0.$ Thus, from ((ref)) we see that $\widetilde{\beta}<0$ cannot arise from $\beta>0$ when $\omega_{12}<0.$ In other words, observing $\widetilde{\beta}<0$ must mean that $\beta<0.$ Moreover, since $g^{\prime }\left( \beta\right) <0$ for all $\beta>\beta_{1}$ and $\beta_{1}<0,$ it must be that $g\left( \beta\right) >0$ for all $\beta>\beta_{1},$ and hence, also for $\beta_{1}<\beta\leq0.$ At $\beta<\beta_{1},$ $g^{\prime}\left( \beta\right) >0,$ and since $g^{\prime}\left( \beta\right) <0$ for all $\beta<\beta_{2}<\beta_{1},$ and $g\left( -\infty\right) =0,$ it has to be that $g\left( \beta\right) $ approaches zero from below as $\beta \rightarrow-\infty,$ and therefore, $g\left( \beta\right) $ must cross zero at some $\beta_{0}\in\left( \beta_{2},\beta_{1}\right) ,$ and $g\left( \beta\right) \geq0$ for all $\beta\in\left[ \beta_{0},0\right] .$ Inspection of ((ref)) shows that $\beta_{0}=\omega_{11} /\omega_{12}=1/\gamma_{0},$ which corresponds to $\gamma=-\infty$ from ((ref)). Since $g\left( \beta\right) \in\left[ 0,1-\lambda\right] $ for all $\beta\in\left[ \beta_{0},0\right] ,$ and $\lambda\in(0,1),$ it follows from ((ref)) that $\left\vert \widetilde{\beta}\right\vert \leq\left\vert \beta\right\vert $. In other words, $\widetilde{\beta}$ is attenuated relative to the true $\beta.$ Finally, we notice that there is a minimum value of $\widetilde{\beta}$ that one can observe under the restriction $\lambda\in\left[ 0,1\right] $ (at $\lambda=1,$ $\widetilde{\beta}=0$). Given the attenuation bias and the fact that $\widetilde{\beta}<0$ if and only if $\beta\in\left[ \beta_{0},1\right] ,$ the smallest value of $\widetilde{\beta}$ occurs when $\lambda=0$ and $\beta=\omega_{11}/\omega_{12},$ so $\widetilde{\beta}_{\min}=\omega _{11}/\omega_{12}=1/\gamma_{0}$. Thus, observing $\widetilde{\beta} <\omega_{11}/\omega_{12}$ and $\omega_{12}<0,$ or $\widetilde{\beta} \omega_{12}/\omega_{11}>1,$ violates the identifying restriction that $\lambda\geq0$, for only with a $\lambda<0$ can we get $g\left( \beta\right) >1$ when $\beta<0$ and hence $\widetilde{\beta}<\beta<0$. \subsubsection{Bounds on $\lambda$} The bounds on $\lambda$ are obtained by finding all the values of $\lambda$ for which equation ((ref)) has a solution for $\beta$. This equation implies \[ \beta^{2}\left( \left( 1-\lambda\right) \omega_{12}+\widetilde{\beta }\lambda\omega_{22}\right) -\beta\left( \left( 1-\lambda\right) \omega_{11}+\widetilde{\beta}\left( 1+\lambda\right) \omega_{12}\right) +\widetilde{\beta}\omega_{11}=0, \] whose discriminant is the following quadratic function of lambda: \begin{align*} D\left( \lambda\right) & =\left( \left( 1-\lambda\right) \omega _{11}+\widetilde{\beta}\left( 1+\lambda\right) \omega_{12}\right) ^{2}-4\widetilde{\beta}\omega_{11}\left( \left( 1-\lambda\right) \omega_{12}+\widetilde{\beta}\lambda\omega_{22}\right). \end{align*} Hence, the identified set for $\lambda$ corresponds to $S_{\lambda}=\left\{ \lambda:D\left( \lambda\right) \geq0\right\} $. This set is non-empty because $D\left( 0\right) \geq0.$ It can be computed analytically and can take the following three shapes: (i) $S_{\lambda}=\Re$ if $D\left( \lambda\right) \allowbreak\geq\allowbreak0$ for all $\lambda\in\Re$; (ii) $S_{\lambda}\allowbreak=\allowbreak(-\infty,\lambda_{lo}]\allowbreak \cup\allowbreak\lbrack\lambda_{up},\infty)$ if $\omega_{11}\allowbreak -\allowbreak\widetilde{\beta}\omega_{12}\allowbreak\neq\allowbreak0,$ where $\lambda_{lo}<\lambda_{up}$ are the roots of $D\left( \lambda\right) =0$; and (iii) $S_{\lambda}\allowbreak=\allowbreak(-\infty,\lambda_{lo}]$\ if $\omega_{11}\allowbreak-\allowbreak\widetilde{\beta}\omega_{12}\allowbreak =\allowbreak0,$ because $\omega_{11}^{2}\allowbreak-\allowbreak\left( \frac{\omega_{11}}{\omega_{12}}\right) ^{2}\allowbreak\omega_{12} ^{2}\allowbreak-\allowbreak2\left( \frac{\omega_{11}}{\omega_{12}}\right) \allowbreak\omega_{11}\omega_{12}\allowbreak+\allowbreak2\left( \frac {\omega_{11}}{\omega_{12}}\right) ^{2}\allowbreak\omega_{11}\omega _{22}\allowbreak=\allowbreak2\omega_{11}^{2}\allowbreak\frac{\omega_{11} \omega_{22}-\omega_{12}^{2}}{\omega_{12}^{2}}\allowbreak>\allowbreak0$. If we also impose the restriction $\lambda\in\left[ 0,1\right] ,$ then the identified set is $S_{\lambda}\cap\left[ 0,1\right] $. \subsection{Proof of Proposition (ref)} Define $\overline{A}_{i2}:=A_{i2}^{\ast}+A_{i2},$ $i=1,2$ as the right blocks of $\overline{A}$ that was defined in ((ref)). Applying GourierouxLaffontMonfort1980, coherency holds if and only if $\det\overline{A}$ and $\det A^{\ast}$ have the same sign. Without loss of generality, we can assume that $A_{11}$ is nonsingular (this can always be achieved by reordering the variables in $Y_{t}$). From ((ref) ), we have $\det A^{\ast}=\det A_{11}\det\left( A_{22}^{\ast}-A_{21} A_{11}^{-1}A_{12}^{\ast}\right) $ and $\det\overline{A}=\det A_{11} \det\left( \overline{A}_{22}-A_{21}A_{11}^{-1}\overline{A}_{12}\right) $ Lutkepohl96. The coherency condition can be written as $\det\overline{A}/\det A^{\ast}>0,$ which, given that $\left( \overline {A}_{22}-A_{21}A_{11}^{-1}\overline{A}_{12}\right) $ and $\left( A_{22}^{\ast}-A_{21}A_{11}^{-1}A_{12}^{\ast}\right) $ are scalars, yields ((ref)). \subsection{Proof of Proposition (ref)} Define $\overline{A}_{i2}:=A_{i2}^{\ast}+A_{i2},$ $i=1,2$ as the right blocks of $\overline{A}$ that was defined in ((ref)). Also let $Y_{t}^{\ast}:=\left( Y_{1t}^{\prime},Y_{2t}^{\ast}\right) ^{\prime}.$ When the coherency condition ((ref)) holds, the solution of ((ref)) exists and is unique. It can be expressed as \begin{equation} Y_{t}^{\ast}=\left\{ \begin{array} [c]{ll} CX_{t}+C^{\ast}X_{t}^{\ast}+u_{t}, & if D_{t}=0\\ \widetilde{C}X_{t}+\widetilde{C}^{\ast}X_{t}^{\ast}+\widetilde{c} b+\widetilde{u}_{t}, & if D_{t}=1 \end{array} \right. \end{equation} where \begin{equation} C=\overline{A}^{-1}B,\quad C^{\ast}=\overline{A}^{-1}B^{\ast},\quad u_{t}=\overline{A}^{-1}\varepsilon_{t} \end{equation} and \begin{equation} \widetilde{C}=A^{\ast-1}B,\quad\widetilde{C}^{\ast}=A^{\ast-1}B^{\ast} ,\quad\widetilde{c}=-A^{\ast-1}\binom{A_{12}}{A_{22}}b,\quad\widetilde{u} _{t}=A^{\ast-1}\varepsilon_{t}. \end{equation} Using the partitioned inverse formula, we obtain \begin{align*} \widetilde{C}_{1} & =\left( A_{11}-A_{12}^{\ast}A_{22}^{\ast-1} A_{21}\right) ^{-1}\left( B_{1}-A_{12}^{\ast}A_{22}^{\ast-1}B_{2}\right) \\ \widetilde{C}_{2} & =\left( A_{22}^{\ast}-A_{21}A_{11}^{-1}A_{12}^{\ast }\right) ^{-1}\left( B_{2}-A_{21}A_{11}^{-1}B_{1}\right) \end{align*} and \begin{align*} C_{1} & =\left( A_{11}-\overline{A}_{12}\overline{A}_{22}^{-1} A_{21}\right) ^{-1}\left( B_{1}-\overline{A}_{12}\overline{A}_{22}^{-1} B_{2}\right) \\ C_{2} & =\left( \overline{A}_{22}-A_{21}A_{11}^{-1}\overline{A} _{12}\right) ^{-1}\left( B_{2}-A_{21}A_{11}^{-1}B_{1}\right) . \end{align*} Solving the latter for $B_{1}$ and $B_{2}$ yields \[ B_{1}=A_{11}C_{1}+\overline{A}_{12}C_{2},\text{ and }B_{2}=\overline{A} _{22}C_{2}+A_{21}C_{1}\text{.} \] Thus, \begin{align*} \widetilde{C}_{1} & =C_{1}+\left( A_{11}-A_{12}^{\ast}A_{22}^{\ast-1} A_{21}\right) ^{-1}\left( \overline{A}_{12}-A_{12}^{\ast}A_{22}^{\ast -1}\overline{A}_{22}\right) C_{2}=C_{1}-\widetilde{\beta}C_{2}, and\\ \widetilde{C}_{2} & =\left( A_{22}^{\ast}-A_{21}A_{11}^{-1}A_{12}^{\ast }\right) ^{-1}\left( A_{22}^{\ast}-A_{21}A_{11}^{-1}A_{12}^{\ast} +A_{22}-A_{21}A_{11}^{-1}A_{12}\right) C_{2}=\kappa C_{2}, \end{align*} where $\kappa$ is given in ((ref)). The exact same derivations apply to $\widetilde{C}^{\ast},$ i.e., \[ \widetilde{C}_{1}^{\ast}=C_{1}^{\ast}-\widetilde{\beta}C_{2}^{\ast},\text{ \ and \ }\widetilde{C}_{2}^{\ast}=\kappa C_{2}^{\ast}. \] Next, \begin{align*} \widetilde{c}_{1} & =\left( A_{11}-A_{12}^{\ast}A_{22}^{\ast-1} A_{21}\right) ^{-1}\left( A_{12}^{\ast}A_{22}^{\ast-1}A_{22}-A_{12}\right) b=\widetilde{\beta}b, and\\ \widetilde{c}_{2} & =-\frac{A_{22}-A_{21}A_{11}^{-1}A_{12}}{A_{22}^{\ast }-A_{21}A_{11}^{-1}A_{12}^{\ast}}b=\left( 1-\kappa\right) b. \end{align*} Finally, $\widetilde{u}_{t}=A^{\ast-1}\overline{A}u_{t}=\left( \widetilde{u} _{1t}^{\prime},\widetilde{u}_{2t}\right) ^{\prime}$, where \[ \widetilde{u}_{1t}=u_{1t}-\widetilde{\beta}u_{2t},\text{ and \ } \widetilde{u}_{2t}=\kappa u_{2t}. \] Substituting back into ((ref)), the reduced-form model\ for $Y_{1t}$ becomes \begin{align} Y_{1t} & =\left( 1-D_{t}\right) \left( C_{1}X_{t}+C_{1}^{\ast}X_{t} ^{\ast}+u_{1t}\right) \nonumber\\ & +D_{t}\left( \left( C_{1}-\widetilde{\beta}C_{2}\right) X_{t}+\left( C_{1}^{\ast}-\widetilde{\beta}C_{2}^{\ast}\right) X_{t}^{\ast}+u_{1t} -\widetilde{\beta}u_{2t}\right) , \end{align} and for $Y_{2t}^{\ast}$ it is \begin{equation} Y_{2t}^{\ast}=C_{2}X_{t}+C_{2}^{\ast}X_{t}^{\ast}+u_{2t}-\left( 1-\kappa\right) D_{t}\left( C_{2}X_{t}+C_{2}^{\ast}X_{t}^{\ast} +u_{2t}-b\right) . \end{equation} Next, define \begin{equation} \widetilde{Y}_{2t}^{\ast}:=C_{2}X_{t}+C_{2}^{\ast}X_{t}^{\ast}+u_{2t}, \end{equation} and rewrite ((ref)) as \begin{align} Y_{2t}^{\ast} & =\widetilde{Y}_{2t}^{\ast}-\left( 1-\kappa\right) D_{t}\left( \widetilde{Y}_{2t}^{\ast}-b\right) \nonumber\\ & =\left( 1-D_{t}\right) \widetilde{Y}_{2t}^{\ast}+D_{t}\left( \kappa\widetilde{Y}_{2t}^{\ast}+\left( 1-\kappa\right) b\right) . \end{align} Let $q=\dim X_{t}$ denote the number of elements of $X_{t}$ and define, for each $i=1,2,$ \begin{equation} \overline{C}_{ij}=\left\{ \begin{array} [c]{ll} C_{ij}, & j\in\left\{ 1,q\right\} :X_{tj}\neq Y_{2,t-s} for all s\in\left\{ 1,p\right\} \\ C_{ij}+C_{is}^{\ast}, & j\in\left\{ 1,q\right\} :X_{tj}=Y_{2,t-s},\text{ for some }s\in\left\{ 1,p\right\} . \end{array} \right. \end{equation} In other words, $\overline{C}$ contains the original coefficients on all the regressors other than the lags of $Y_{2t},$ while the coefficients on the lags of $Y_{2t}$ are augmented by the corresponding coefficients of the lags of $Y_{2t}^{\ast}.$ For example, if $p=1$ and there are no other exogenous regressors $X_{0t},$ then, for $i=1,2,$ \[ C_{i}X_{t}+C_{i}^{\ast}X_{t}^{\ast}=C_{i1}Y_{1t-1}+C_{i2}Y_{2t-1}+C_{i}^{\ast }Y_{2t-1}^{\ast}, \] so $\overline{C}_{i}=\left( C_{i1},C_{i2}+C_{i}^{\ast}\right) $. Using ((ref)), we can rewrite ((ref)) as \begin{equation} \widetilde{Y}_{2t}^{\ast}=\overline{C}_{2}X_{t}+C_{2}^{\ast}\min\left( X_{t}^{\ast}-b,0\right) +u_{2t}. \end{equation} Now, observe that \[ \min\left( Y_{2t}^{\ast}-b,0\right) =D_{t}\left( Y_{2t}^{\ast}-b\right) =\kappa D_{t}\left( \widetilde{Y}_{2t}^{\ast}-b\right) =\kappa\min\left( \widetilde{Y}_{2t}^{\ast}-b,0\right) \] So, letting $\widetilde{X}_{t}^{\ast}$ denote the lags of $\widetilde{Y} _{2t}^{\ast},$ we have $\min\left( X_{t}^{\ast}-b,0\right) =\kappa \min\left( \widetilde{X}_{t}^{\ast}-b,0\right) ,$ and consequently, \[ C^{\ast}\min\left( X_{t}^{\ast}-b,0\right) =\overline{C}^{\ast}\min\left( \widetilde{X}_{t}^{\ast}-b,0\right) , \] where $\overline{C}^{\ast}=\kappa C^{\ast}.$ Now, from ((ref)) we have \[ \widetilde{Y}_{2t}^{\ast}=\overline{C}_{2}X_{t}+\overline{C}_{2}^{\ast} \min\left( \widetilde{X}_{t}^{\ast}-b,0\right) +u_{2t}. \] Recall the definition of $\overline{Y}_{2t}^{\ast}$ in ((ref)): \[ \overline{Y}_{2t}^{\ast}:=\overline{C}_{2}X_{t}+\overline{C}_{2}^{\ast }\overline{X}_{t}^{\ast}+u_{2t}, \] where $\overline{X}_{t}^{\ast}:=\left( \overline{x}_{t-1},...,\overline {x}_{t-p}\right) ^{\prime},$ and $\overline{x}_{t}\allowbreak:=\allowbreak \min\left( \overline{Y}_{2t}^{\ast}\allowbreak-b,0\right) ,$ with initial conditions $\overline{x}_{-s}\allowbreak=\kappa^{-1}\allowbreak\min\left( Y_{2,-s}^{\ast}\allowbreak-b,0\right) ,$ $s=0,...,\allowbreak p-1.$ It follows that $\min\left( \widetilde{X}_{t}^{\ast}-b,0\right) \allowbreak =\overline{X}_{t}^{\ast}$ for all $t\geq1,$ so that $\widetilde{Y}_{2t}^{\ast }=\overline{Y}_{2t}^{\ast}.$ Substituting $\overline{Y}_{2}^{\ast}$ for $\widetilde{Y}_{2}^{\ast}$ in ((ref)), we get ((ref) ). Using the reparametrization ((ref)) and the relationship between $X_{t}^{\ast}$ and $\overline{X}_{t}^{\ast}$ in ((ref)), we obtain ((ref)). Finally, from eq. ((ref)), it follows that the event $Y_{2t}^{\ast}<b$ is equivalent to \[ b+\kappa\left( C_{2}X_{t}+C_{2}^{\ast}X_{t}^{\ast}+u_{2t}-b\right) <b, \] which, since $\kappa>0$ by the coherency condition ((ref)), is equivalent to \begin{equation} u_{2t}<b-C_{2}X_{t}-C_{2}^{\ast}X_{t}^{\ast}. \end{equation} Using the definition ((ref)), and ((ref)), the inequality ((ref)) can be written as $\overline{Y} _{2t}^{\ast}<b,$ which establishes ((ref)). \textbf{Comment:} Note that $\kappa$ appears in the reduced form only multiplicatively with $C^{\ast},$ so $\kappa$ and $C^{\ast}$ are not separately identified, only $\overline{C}^{\ast}=\kappa C^{\ast}$ is. The reparametrization from $C$ to $\overline{C}$ is convenient because $\overline{C}$ is identified independently of $\kappa,$ while $C,C^{\ast}$ and $\kappa$ are not separately identified. \subsection{Proof of Proposition (ref)} We solve $u_{t}=\overline{A}^{-1}\varepsilon_{t}$ using the partitioned inverse formula to get \begin{align} u_{1t} & =\left( A_{11}-\overline{A}_{12}\overline{A}_{22}^{-1} A_{21}\right) ^{-1}\left( \varepsilon_{1t}-\overline{A}_{12}\overline {A}_{22}^{-1}\varepsilon_{2t}\right) \\ u_{2t} & =\left( \overline{A}_{22}-A_{21}A_{11}^{-1}\overline{A} _{12}\right) ^{-1}\left( \varepsilon_{2t}-A_{21}A_{11}^{-1}\varepsilon _{1t}\right) . \end{align} Using the definitions \begin{align*} \overline{\beta} & :=-A_{11}^{-1}\overline{A}_{12},\quad\overline{\gamma }:=-\overline{A}_{22}^{-1}A_{21},\\ \bar{\varepsilon}_{1t} & :=A_{11}^{-1}\varepsilon_{1t},\quad\bar {\varepsilon}_{2t}:=\overline{A}_{22}^{-1}\varepsilon_{2t}, \end{align*} we can rewrite ((ref))-((ref)) as ((ref) )-((ref)). Note that \begin{align*} \bar{\varepsilon}_{1t} & =A_{11}^{-1}\left( A_{11}u_{1t}+\overline{A} _{12}u_{2t}\right) =u_{1t}-\bar{\beta}u_{2t},\\ \bar{\varepsilon}_{2t} & =\overline{A}_{22}^{-1}\left( A_{21} u_{1t}+\overline{A}_{22}u_{2t}\right) =-\overline{\gamma}u_{1t}+u_{2t}, \end{align*} so, \begin{align*} var\left( \bar{\varepsilon}_{1t}\right) & =\left( I_{k-1},-\bar{\beta }\right) \Omega\left( I_{k-1},-\bar{\beta}\right) ^{\prime},\\ var\left( \bar{\varepsilon}_{2t}\right) & =\left( -\bar{\gamma},1\right) \Omega\left( -\bar{\gamma},1\right) ^{\prime}, \end{align*} and \begin{align*} cov\left( \bar{\varepsilon}_{1t},\bar{\varepsilon}_{2t}\right) & =\left( I_{k-1},-\overline{\beta}\right) \begin{pmatrix} \Omega_{11} & \Omega_{12}\\ \Omega_{12}^{\prime} & \Omega_{22} \end{pmatrix} \left( -\overline{\gamma},1\right) ^{\prime}\\ & =-\left( \Omega_{11}-\overline{\beta}\Omega_{12}^{\prime}\right) \overline{\gamma}^{\prime}+\Omega_{12}-\overline{\beta}\Omega_{22}=0. \end{align*} The last equation identifies $\overline{\gamma}$ given $\overline{\beta}$. Specifically, \[ \overline{\gamma}^{\prime}=\left( \Omega_{11}-\overline{\beta}\Omega _{12}^{\prime}\right) ^{-1}\left( \Omega_{12}-\overline{\beta}\Omega _{22}\right) . \] \section{Numerical results} This section provides Monte-Carlo evidence on the finite-sample properties of the proposed estimators and tests. The data generating process (DGP) is a trivariate VAR(1), given by equations ((ref)) and ((ref)). I\ consider three different DGPs corresponding to the CKSVAR, KSVAR and CSVAR models, respectively. In all three DGPs, the following parameters are set to the same values: the contemporaneous coefficients are $A_{11}=I_{2},$ $A_{12}=$ $A_{12}^{\ast}=0_{2\times1},$ $A_{22}^{\ast}=1$ and $A_{22}=0;$ the intercepts are set to zero, $B_{10}=0_{2\times1}$ and $B_{20}=0;$ the coefficients on the lags are $B_{1,1}=\left( \rho I_{2},0\right) ,$ $B_{1,1}^{\ast}=0_{2\times1},$ $B_{2,1}=\left( 0_{1\times2},B_{22,1}\right) $, with $\rho=0.5.$ Finally, each of the three DGPs is determined as follows. DGP1:\ $B_{22,1}=B_{2,1}^{\ast}=0$ (both KSVAR and CSVAR restrictions hold, since lags of $Y_{2,t}$ and $Y_{2,t}^{\ast}$ all have zero coefficients); DGP2: $B_{22,1}=\rho,\,B_{2,1}^{\ast}=0\ $(KSVAR restrictions hold but CSVAR restrictions do not); DGP3: $B_{22,1} =0,\,B_{2,1}^{\ast}=\rho$ (CSVAR restrictions hold but KSVAR restrictions do not). The setting of the autoregressive coefficient $\rho=0.5$ leads to a lower degree of persistence than is typically observed in macro data (e.g., in the Stock and Watson, 2001, application, the three largest roots are 0.97, 0.97 and 0.8), because I want to avoid confounding any possible finite-sample issues arising from the ZLB with well-known problems of bias and size distortion due to strong persistence (near unit roots) in the data. Finally, the bound on $Y_{2t}$ is set to $b=0,$ the sample size is $T=250,$ the initial conditions are set to 0 and the number of Monte Carlo replications is 1000. In all cases, the CKSVAR and CSVAR likelihoods are computed using SIS with $R=1000$ particles. The notation for the reported parameters is given in Table (ref). \begin{table}[tbp] \caption{Parameter notation in reported simulation results} \begin{tabular} [c]{l|l}\hline \textbf{Mnemonic} & \textbf{Description}\\\hline $\tau$ & st. dev. of reduced form error $u_{2t}$ in $Y_{2t}$ (constrained variable)\\ Eq. 3 & reduced form equation for $Y_{2t}$\\ Eq. j & red. form equation for $Y_{1j,t},$ $j=1,2$ (unconstrained variables)\\ $\widetilde{\beta}_{j}$ & coefficient on kink in eq. j\\ eq. i Y1j_1 & coefficient of $Y_{1j,t-1}$ in eq. i\\ eq. i Y2_1 & coefficient of $Y_{2,t-1}$ in eq. i\\ eq. i lY2_1 & coefficient of $\min\left( Y_{2,t-1}^{\ast}-b,0\right) $ in eq. i\\ $\delta_{j}$ & coefficient of regression of $u_{1j,t}$ (red. form error in Eq. j) on $u_{2t}$\\ Ch_ij & (i,j) element of Choleski factor of $\Omega_{1.2}$\\\hline \end{tabular} \end{table} Figure (ref) reports the sampling distribution of the ML estimators of the reduced-form parameters in Proposition (ref) for the CKSVAR model under DGP1. The results for the KSVAR and CSVAR models, which are also correctly specified under DGP1, are omitted because they are entirely analogous. The sampling densities appear to be very close to the superimposed Normal approximations, indicating that the Normal asymptotic approximation is fairly accurate. \begin{figure}[htb] \caption{Sampling densities of ML estimators of reduced-form coefficients of CKSVAR(1) under DGP1 (solid lines) and approximating Normal densities (dashed lines). $T=250$, 1000 Monte Carlo replications. Parameter names described in Table (ref).} \end{figure} Table (ref) reports moments of the sampling distributions of the above mentioned estimators for the CKSVAR model. Again, the results for the KSVAR and CSVAR models are entirely analogous and are therefore omitted. We notice no discernible biases. Additional simulation results with $T=100$ and $T=1000$ given in the Tables (ref), (ref) and (ref) indicate that the RMSE declines at rate $\sqrt{T}$ in accordance with asymptotic theory. It is noteworthy that the estimators of $\widetilde{\beta}$ have substantially larger RMSE than the estimators of the other parameters. \begin{table}[tbp] \caption{Moments of sampling distribution of ML estimators of the parameters of CKSVAR(1)\tabnoteref[a]{tab2}} \begin{tabular} [c]{r|r|r|r|r|r}\hline ML-CKSVAR & true & mean & bias & sd & RMSE\\\hline $\tau$ & 1.000 & 0.983 & -0.017 & 0.068 & 0.070\\ Eq.3 Constant & 0.000 & -0.006 & -0.006 & 0.175 & 0.176\\ Eq.3 Y11_1 & 0.000 & 0.000 & 0.000 & 0.061 & 0.061\\ Eq.3 Y12_1 & 0.000 & -0.000 & -0.000 & 0.064 & 0.064\\ Eq.3 Y2_1 & 0.000 & -0.010 & -0.010 & 0.173 & 0.173\\ Eq.3 lY2_1 & 0.000 & -0.013 & -0.013 & 0.258 & 0.258\\ $\tilde{\beta}_{1}$ & -0.000 & -0.008 & -0.008 & 0.356 & 0.356\\ $\tilde{\beta}_{2}$ & 0.000 & -0.004 & -0.004 & 0.359 & 0.359\\ Eq.1 Constant & 0.000 & 0.002 & 0.002 & 0.213 & 0.213\\ Eq.1 Y11_1 & 0.500 & 0.489 & -0.011 & 0.057 & 0.058\\ Eq.1 Y12_1 & 0.000 & 0.002 & 0.002 & 0.060 & 0.060\\ Eq.1 Y2_1 & 0.000 & -0.002 & -0.002 & 0.163 & 0.163\\ Eq.2 Constant & 0.000 & 0.005 & 0.005 & 0.214 & 0.214\\ Eq.2 Y11_1 & 0.000 & 0.000 & 0.000 & 0.058 & 0.058\\ Eq.2 Y12_1 & 0.500 & 0.491 & -0.009 & 0.057 & 0.058\\ Eq.2 Y2_1 & 0.000 & -0.001 & -0.001 & 0.159 & 0.159\\ Eq.1 lY2_1 & 0.000 & 0.005 & 0.005 & 0.232 & 0.232\\ Eq.2 lY2_1 & 0.000 & 0.002 & 0.002 & 0.233 & 0.233\\ $\delta_{1}$ & 0.000 & -0.001 & -0.001 & 0.157 & 0.157\\ $\delta_{2}$ & 0.000 & -0.004 & -0.004 & 0.155 & 0.155\\ Ch_11 & 1.000 & 0.975 & -0.025 & 0.044 & 0.051\\ Ch_21 & 0.000 & -0.001 & -0.001 & 0.066 & 0.066\\ Ch_22 & 1.000 & 0.972 & -0.028 & 0.046 & 0.054\\\hline \end{tabular} \tabnotetext[a]{tab2}{Computed under DGP1 with $T=250$ using 1000 MC replications. Parameter names described in Table (ref).} \end{table} \begin{table}[htb] \caption{Bias, standard deviation and Root Mean Square Error of Maximum Likelihood estimator of parameters of CKSVAR(1) model\tabnoteref[a]{tab4}} \begin{tabular} [c]{r|r|r|r|r|rr|r|r|r}\hline ML-CKSVAR & \multicolumn{3}{|c|}{$T=100$} & \multicolumn{3}{|c|}{$T=250$} & \multicolumn{3}{c}{$T=1000$}\\\hline Parameter & bias & sd & RMSE & bias & sd & RMSE & bias & sd & RMSE\\\hline $\tau$ & \multicolumn{1}{|r|}{-0.048} & 0.111 & 0.121 & -0.017 & 0.068 & \multicolumn{1}{|r|}{0.070} & -0.003 & 0.035 & 0.035\\ Eq.3 Constant & \multicolumn{1}{|r|}{0.001} & 0.293 & 0.293 & -0.006 & 0.175 & \multicolumn{1}{|r|}{0.176} & 0.004 & 0.085 & 0.085\\ Eq.3 Y11_1 & \multicolumn{1}{|r|}{-0.001} & 0.111 & 0.111 & 0.000 & 0.061 & \multicolumn{1}{|r|}{0.061} & 0.000 & 0.032 & 0.032\\ Eq.3 Y12_1 & \multicolumn{1}{|r|}{-0.004} & 0.109 & 0.109 & -0.000 & 0.064 & \multicolumn{1}{|r|}{0.064} & -0.000 & 0.031 & 0.031\\ Eq.3 Y2_1 & \multicolumn{1}{|r|}{-0.031} & 0.289 & 0.291 & -0.010 & 0.173 & \multicolumn{1}{|r|}{0.173} & -0.003 & 0.082 & 0.082\\ Eq.3 lY2_1 & \multicolumn{1}{|r|}{-0.022} & 0.435 & 0.436 & -0.013 & 0.258 & \multicolumn{1}{|r|}{0.258} & 0.003 & 0.122 & 0.122\\ $\tilde{\beta}_{1}$ & \multicolumn{1}{|r|}{0.012} & 0.586 & 0.586 & -0.008 & 0.356 & \multicolumn{1}{|r|}{0.356} & 0.000 & 0.176 & 0.176\\ $\tilde{\beta}_{2}$ & \multicolumn{1}{|r|}{-0.011} & 0.606 & 0.606 & -0.004 & 0.359 & \multicolumn{1}{|r|}{0.359} & -0.003 & 0.168 & 0.168\\ Eq.1 Constant & \multicolumn{1}{|r|}{0.000} & 0.360 & 0.360 & 0.002 & 0.213 & \multicolumn{1}{|r|}{0.213} & 0.007 & 0.105 & 0.105\\ Eq.1 Y11_1 & \multicolumn{1}{|r|}{-0.034} & 0.098 & 0.104 & -0.011 & 0.057 & \multicolumn{1}{|r|}{0.058} & -0.002 & 0.029 & 0.029\\ Eq.1 Y12_1 & \multicolumn{1}{|r|}{-0.004} & 0.108 & 0.108 & 0.002 & 0.060 & \multicolumn{1}{|r|}{0.060} & 0.001 & 0.028 & 0.028\\ Eq.1 Y2_1 & \multicolumn{1}{|r|}{-0.004} & 0.281 & 0.281 & -0.002 & 0.163 & \multicolumn{1}{|r|}{0.163} & -0.006 & 0.079 & 0.079\\ Eq.2 Constant & \multicolumn{1}{|r|}{0.006} & 0.356 & 0.356 & 0.005 & 0.214 & \multicolumn{1}{|r|}{0.214} & 0.002 & 0.102 & 0.103\\ Eq.2 Y11_1 & \multicolumn{1}{|r|}{0.004} & 0.103 & 0.103 & 0.000 & 0.058 & \multicolumn{1}{|r|}{0.058} & 0.000 & 0.028 & 0.028\\ Eq.2 Y12_1 & \multicolumn{1}{|r|}{-0.029} & 0.103 & 0.107 & -0.009 & 0.057 & \multicolumn{1}{|r|}{0.058} & -0.002 & 0.029 & 0.029\\ Eq.2 Y2_1 & \multicolumn{1}{|r|}{0.000} & 0.269 & 0.269 & -0.001 & 0.159 & \multicolumn{1}{|r|}{0.159} & 0.001 & 0.078 & 0.078\\ Eq.1 lY2_1 & \multicolumn{1}{|r|}{0.009} & 0.403 & 0.403 & 0.005 & 0.232 & \multicolumn{1}{|r|}{0.232} & 0.010 & 0.112 & 0.113\\ Eq.2 lY2_1 & \multicolumn{1}{|r|}{-0.003} & 0.408 & 0.408 & 0.002 & 0.233 & \multicolumn{1}{|r|}{0.233} & 0.004 & 0.115 & 0.115\\ $\delta_{1}$ & \multicolumn{1}{|r|}{0.005} & 0.260 & 0.260 & -0.001 & 0.157 & \multicolumn{1}{|r|}{0.157} & 0.001 & 0.075 & 0.075\\ $\delta_{2}$ & \multicolumn{1}{|r|}{-0.005} & 0.260 & 0.260 & -0.004 & 0.155 & \multicolumn{1}{|r|}{0.155} & -0.001 & 0.073 & 0.073\\ Ch_11 & \multicolumn{1}{|r|}{-0.067} & 0.073 & 0.099 & -0.025 & 0.044 & \multicolumn{1}{|r|}{0.051} & -0.006 & 0.024 & 0.024\\ Ch_21 & \multicolumn{1}{|r|}{-0.005} & 0.111 & 0.111 & -0.001 & 0.066 & \multicolumn{1}{|r|}{0.066} & -0.000 & 0.031 & 0.031\\ Ch_22 & \multicolumn{1}{|r|}{-0.075} & 0.073 & 0.105 & -0.028 & 0.046 & \multicolumn{1}{|r|}{0.054} & -0.008 & 0.023 & 0.024\\\hline \end{tabular} \tabnotetext[a]{tab4}{Computed under DGP1 with $R=1000$ particles using 1000 MC replications. Parameter names described in Table (ref).} \end{table} \begin{table}[htb] \caption{Bias, standard deviation and Root Mean Square Error of Maximum Likelihood estimator of parameters of KSVAR(1) model\tabnoteref[a]{tab5}} \begin{tabular} [c]{r|r|r|r|r|r|r|r|r|r}\hline ML-KSVAR & \multicolumn{3}{|c|}{$T=100$} & \multicolumn{3}{c|}{$T=250$} & \multicolumn{3}{c}{$T=1000$}\\\hline Parameter & bias & sd & RMSE & bias & sd & RMSE & bias & sd & RMSE\\\hline $\tau$ & -0.024 & 0.111 & 0.113 & -0.008 & 0.068 & 0.069 & -0.001 & 0.035 & 0.035\\ Eq.3 Constant & 0.011 & 0.145 & 0.145 & 0.001 & 0.092 & 0.092 & 0.003 & 0.046 & 0.046\\ Eq.3 Y11_1 & -0.001 & 0.103 & 0.103 & 0.001 & 0.060 & 0.060 & -0.000 & 0.031 & 0.031\\ Eq.3 Y12_1 & -0.004 & 0.102 & 0.102 & -0.000 & 0.062 & 0.062 & -0.000 & 0.030 & 0.030\\ Eq.3 Y2_1 & -0.048 & 0.199 & 0.204 & -0.019 & 0.122 & 0.124 & -0.003 & 0.060 & 0.060\\ $\tilde{\beta}_{1}$ & -0.003 & 0.571 & 0.571 & -0.013 & 0.349 & 0.349 & -0.001 & 0.174 & 0.174\\ $\tilde{\beta}_{2}$ & -0.003 & 0.584 & 0.584 & -0.001 & 0.348 & 0.348 & -0.004 & 0.168 & 0.168\\ Eq.1 Constant & 0.002 & 0.264 & 0.264 & 0.001 & 0.165 & 0.165 & 0.001 & 0.080 & 0.080\\ Eq.1 Y11_1 & -0.033 & 0.093 & 0.099 & -0.012 & 0.056 & 0.057 & -0.002 & 0.028 & 0.028\\ Eq.1 Y12_1 & -0.002 & 0.100 & 0.100 & 0.002 & 0.058 & 0.058 & 0.001 & 0.027 & 0.027\\ Eq.1 Y2_1 & -0.002 & 0.197 & 0.197 & -0.000 & 0.117 & 0.117 & -0.001 & 0.057 & 0.057\\ Eq.2 Constant & 0.004 & 0.258 & 0.258 & 0.003 & 0.158 & 0.158 & 0.001 & 0.078 & 0.078\\ Eq.2 Y11_1 & 0.006 & 0.096 & 0.096 & 0.001 & 0.057 & 0.057 & 0.000 & 0.027 & 0.027\\ Eq.2 Y12_1 & -0.028 & 0.094 & 0.098 & -0.008 & 0.055 & 0.056 & -0.002 & 0.028 & 0.029\\ Eq.2 Y2_1 & -0.001 & 0.189 & 0.189 & -0.000 & 0.113 & 0.113 & 0.003 & 0.054 & 0.055\\ $\delta_{1}$ & -0.000 & 0.252 & 0.252 & -0.003 & 0.156 & 0.156 & 0.001 & 0.075 & 0.075\\ $\delta_{2}$ & -0.003 & 0.253 & 0.253 & -0.003 & 0.152 & 0.152 & -0.001 & 0.073 & 0.073\\ Ch_11 & -0.048 & 0.070 & 0.085 & -0.018 & 0.044 & 0.047 & -0.005 & 0.023 & 0.024\\ Ch_21 & -0.003 & 0.108 & 0.108 & -0.000 & 0.065 & 0.065 & -0.000 & 0.031 & 0.031\\ Ch_22 & -0.054 & 0.070 & 0.088 & -0.020 & 0.045 & 0.050 & -0.007 & 0.023 & 0.024\\\hline \end{tabular} \tabnotetext[a]{tab5}{Computed under DGP1 with $R=1000$ particles using 1000 MC replications. Parameter names described in Table (ref).} \end{table} \begin{table}[ptb] \caption{Bias, standard deviation and Root Mean Square Error of Maximum Likelihood estimator of parameters of CSVAR(1) model\tabnoteref[a]{tab6}} \begin{tabular} [c]{r|r|r|r|r|r|r|r|r|r}\hline ML-CSVAR & \multicolumn{3}{|c}{$T=100$} & \multicolumn{3}{c|}{$T=250$} & \multicolumn{3}{c}{$T=1000$}\\\hline Parameter & bias & sd & RMSE & bias & sd & RMSE & bias & sd & RMSE\\\hline $\tau$ & -0.026 & 0.111 & 0.114 & -0.008 & 0.068 & 0.069 & -0.001 & 0.035 & 0.035\\ Eq.3 Constant & -0.009 & 0.135 & 0.135 & -0.006 & 0.081 & 0.081 & 0.001 & 0.040 & 0.040\\ Eq.3 Y11_1 & -0.001 & 0.104 & 0.104 & 0.001 & 0.060 & 0.060 & -0.000 & 0.031 & 0.031\\ Eq.3 Y12_1 & -0.004 & 0.103 & 0.103 & -0.000 & 0.061 & 0.061 & -0.000 & 0.030 & 0.030\\ Eq.3 Y2_1 & -0.027 & 0.125 & 0.128 & -0.011 & 0.078 & 0.079 & -0.001 & 0.038 & 0.038\\ Eq.1 Constant & -0.002 & 0.109 & 0.110 & -0.003 & 0.065 & 0.065 & 0.000 & 0.031 & 0.031\\ Eq.1 Y11_1 & -0.033 & 0.090 & 0.096 & -0.012 & 0.054 & 0.056 & -0.002 & 0.028 & 0.028\\ Eq.1 Y12_1 & -0.002 & 0.096 & 0.096 & 0.002 & 0.057 & 0.057 & 0.001 & 0.027 & 0.027\\ Eq.1 Y2_1 & -0.000 & 0.118 & 0.118 & 0.001 & 0.072 & 0.072 & 0.000 & 0.036 & 0.036\\ Eq.2 Constant & 0.003 & 0.106 & 0.106 & 0.002 & 0.063 & 0.063 & -0.000 & 0.031 & 0.031\\ Eq.2 Y11_1 & 0.005 & 0.091 & 0.092 & 0.001 & 0.056 & 0.056 & 0.000 & 0.027 & 0.027\\ Eq.2 Y12_1 & -0.026 & 0.090 & 0.094 & -0.008 & 0.054 & 0.055 & -0.002 & 0.028 & 0.028\\ Eq.2 Y2_1 & -0.001 & 0.115 & 0.115 & 0.000 & 0.070 & 0.070 & 0.002 & 0.035 & 0.035\\ $\delta_{1}$ & 0.001 & 0.114 & 0.114 & 0.002 & 0.071 & 0.071 & 0.001 & 0.035 & 0.035\\ $\delta_{2}$ & -0.001 & 0.116 & 0.116 & -0.002 & 0.070 & 0.070 & -0.000 & 0.034 & 0.034\\ Ch_11 & -0.033 & 0.069 & 0.077 & -0.012 & 0.043 & 0.044 & -0.003 & 0.023 & 0.023\\ Ch_21 & -0.005 & 0.103 & 0.104 & -0.001 & 0.064 & 0.064 & -0.000 & 0.031 & 0.031\\ Ch_22 & -0.037 & 0.068 & 0.077 & -0.014 & 0.044 & 0.046 & -0.005 & 0.023 & 0.024\\\hline \end{tabular} \tabnotetext[a]{tab6}{Computed under DGP1 with $R=1000$ particles using 1000 MC replications. Parameter names described in Table (ref).} \end{table} Next, I turn to the properties of the LR test of KSVAR against CKSVAR and CSVAR against CKSVAR. The former hypothesis involves three restrictions (exclusion of the latent lag $Y_{2,t-1}^{\ast}$ from each of the three equations), so the LR statistic is asymptotically distributed as $\chi_{3} ^{2}$ under the null. The latter hypothesis involves five restrictions (exclusion of the observed lag $Y_{2,t-1}$ from each of the three equations, plus $\widetilde{\beta}=0$), and the LR statistic is asymptotically distributed as $\chi_{5}^{2}.$ Table (ref) reports the rejection frequencies of the LR tests for each of the two hypotheses in each of the three DGPs at three significance levels: 10%, 5% and 1%. In addition to the asymptotic tests, I\ also report the rejection frequency of the tests using parametric bootstrap critical values. The parametric bootstrap uses draws of Normal errors and the estimated reduced-form parameters to generate the bootstrap samples of $Y_{t}$ and $Y_{2t}^{\ast}$. The Monte Carlo rejection frequencies are computed using the “warp-speed” method method of GiacominiPolitisWhite2013. Note that both null hypotheses hold under DGP1, but only the KSVAR is valid under DGP2 and only the CSVAR is valid under DGP3. For convenience, I indicate the rejection frequencies under the alternative in bold in the table. \begin{table}[htbp] \caption{Rejection frequencies of LR tests of $H_0$ against $H_1$ across different DGPs\tabnoteref[a]{tab3}} \begin{tabular} [c]{ll|c|c|c|c|c|c}\hline & & \multicolumn{3}{|c|}{$H_{0}:$ KSVAR, $H_{1}:$ CKSVAR} & \multicolumn{3}{|l}{$H_{0}:$ CSVAR, $H_{1}:$ CKSVAR}\\\hline & Sign. Level & 10% & 5% & 1% & 10% & 5% & 1%\\\hline DGP1 & asymptotic & 0.173 & 0.093 & 0.020 & 0.155 & 0.084 & 0.022\\ & bootstrap & 0.107 & 0.045 & 0.005 & 0.109 & 0.052 & 0.012\\\hline DGP2 & asymptotic & 0.149 & 0.080 & 0.018 & \textbf{0.319} & \textbf{0.206} & \textbf{0.073}\\ & bootstrap & 0.117 & 0.050 & 0.011 & \textbf{0.242} & \textbf{0.141} & \textbf{0.041}\\\hline DGP3 & asymptotic & \textbf{0.587} & \textbf{0.454} & \textbf{0.244} & 0.141 & 0.067 & 0.019\\ & bootstrap & \textbf{0.471} & \textbf{0.365} & \textbf{0.194} & 0.103 & 0.053 & 0.018\\\hline \end{tabular} \tabnotetext[a]{tab3}{Computed using 1000 Monte Carlo replications, $T=250$. The asymptotic tests use $\chi^2_3$ and $\chi^2_5$ critical values for KSVAR and CSVAR resp. The bootstrap rej. frequencies were computed uisng the warp-speed method of GiacominiPolitisWhite2013. Bold numbers indicate that the rejecction frequencies were computed under $H_1$ (power).} \end{table} There is evidence that the LR tests reject too often under $H_{0}$ relative to their nominal level when we use asymptotic critical values. Moreover, the size distortions are very similar across null hypotheses and DGPs. Unreported results show that size distortion eventually disappears as the sample gets large, but this level of overrejection is unsatisfactory at $T=250,$ which is a typical sample size in macroeconomic applications. The parametric bootstrap appears to do a remarkably good job at correcting the size of the tests. In all cases considered, the parametric bootstrap rejection frequency is not significantly different from the nominal level when the null hypothesis holds (all but the numbers in bold in the Table). To shed further light on this issue, Figure (ref) reports the QQ plots of the sampling distributions of the two LR statistics against their asymptotic and parametric bootstrap approximations for all three DGPs under the null hypothesis. The sampling distributions of the LR statistics stochastically dominate their asymptotic approximations, but the bootstrap approximations are quite accurate. Finally, the rejection frequencies highlighted in bold in Table (ref) correspond to the power of the tests against two very similar deviations from the null hypothesis. The numbers on the left under DGP3 show the power of the test to reject the KSVAR specification under the alternative at which the coefficient on the latent lag $B_{2,1}^{\ast}=0.5.$ Similarly, the bold numbers on the right give the power of rejecting CSVAR against the alternative where the coefficient on the observed lag $B_{22,1}=0.5.$ Since the lower bound is set to zero, and the sample contains about 50% of observations at the ZLB, the two deviations from the null are of equal magnitude. Yet, we notice the LR\ test is significantly more powerful against the KSVAR than against the CSVAR. This could be because CSVAR\ imposes more restrictions than the KSVAR, so one would expect it to have lower power than the KSVAR against similar deviations from the null. \begin{figure}[ptb] \caption{QQ plots of the sampling distribution under the null hypothesis of LR statistics of KSVAR against CKSVAR (left) and CSVAR against CKSVAR (right). Solid (dashed) lines plot quantiles against asymptotic $\chi^2$ (bootstrap) approximation. Computed for $T=250$ using 1000 Monte Carlo replications.} \end{figure} \subsection*{Alternative DGP} The DGPs in the previous simulations have the property that the frequency of the ZLB regime is around 50%. I reran those simulations with a slight modification to the DGPs to match the frequency in the sample of the empirical application in the paper. Specifically, I reduce the lower bound $b$ to a level that makes the frequency of the ZLB regime equal to 11%. The results are given in Figure (ref) and Table (ref). The results are very similar to the ones reported in Figure (ref) and Table (ref) above: the Normal approximation of the sampling distribution of the MLE appears to be very good, and the bias is negligible. The only difference is that the standard deviation of $\widetilde{\beta}$ is larger. \begin{figure}[htb] \caption{Sampling densities of ML estimators of reduced-form coefficients of CKSVAR(1) under DGP1 with $b$ chosen such that $\Pr\left( Y_{2t}=b\right) =0.11$ (solid lines) and approximating Normal densities (dashed lines). $T=250$, 1000 Monte Carlo replications. Parameter names described in Table (ref).} \end{figure} \begin{table}[ptb] \caption{Moments of sampling distribution of ML estimators of the parameters of CKSVAR(1)\tabnoteref[a]{tab7}} \begin{tabular} [c]{r|r|r|r|r|r}\hline ML-CKSVAR & true & mean & bias & sd & RMSE\\\hline $\tau$ & 1.000 & 0.988 & -0.012 & 0.050 & 0.051\\ Eq.3 Constant & 0.000 & -0.004 & -0.004 & 0.073 & 0.073\\ Eq.3 Y11_1 & 0.000 & 0.001 & 0.001 & 0.055 & 0.055\\ Eq.3 Y12_1 & 0.000 & -0.002 & -0.002 & 0.056 & 0.056\\ Eq.3 Y2_1 & 0.000 & -0.008 & -0.008 & 0.083 & 0.083\\ Eq.3 lY2_1 & 0.000 & 0.003 & 0.003 & 0.501 & 0.501\\ $\tilde{\beta}_{1}$ & 0.000 & 0.013 & 0.013 & 0.533 & 0.533\\ $\tilde{\beta}_{2}$ & 0.000 & -0.030 & -0.030 & 0.518 & 0.519\\ Eq.1 Constant & 0.000 & -0.005 & -0.005 & 0.078 & 0.078\\ Eq.1 Y11_1 & 0.500 & 0.488 & -0.012 & 0.055 & 0.056\\ Eq.1 Y12_1 & 0.000 & 0.001 & 0.001 & 0.058 & 0.058\\ Eq.1 Y2_1 & 0.000 & 0.001 & 0.001 & 0.084 & 0.084\\ Eq.2 Constant & 0.000 & 0.005 & 0.005 & 0.075 & 0.075\\ Eq.2 Y11_1 & 0.000 & 0.001 & 0.001 & 0.056 & 0.056\\ Eq.2 Y12_1 & 0.500 & 0.491 & -0.009 & 0.056 & 0.056\\ Eq.2 Y2_1 & 0.000 & -0.000 & -0.000 & 0.080 & 0.080\\ Eq.1 lY2_1 & 0.000 & -0.014 & -0.014 & 0.497 & 0.497\\ Eq.2 lY2_1 & 0.000 & 0.018 & 0.018 & 0.491 & 0.491\\ $\delta_{1}$ & 0.000 & 0.002 & 0.002 & 0.084 & 0.084\\ $\delta_{2}$ & 0.000 & -0.004 & -0.004 & 0.080 & 0.080\\ Ch_11 & 1.000 & 0.980 & -0.020 & 0.043 & 0.047\\ Ch_21 & 0.000 & -0.002 & -0.002 & 0.065 & 0.065\\ Ch_22 & 1.000 & 0.978 & -0.022 & 0.045 & 0.050\\\hline \end{tabular} \tabnotetext[a]{tab7}{Computed under DGP1 with $b$ chosen such that $\Pr\left( Y_{2t}=b\right) =0.11$, $T=250$ using 1000 MC replications. Parameter names described in Table (ref).} \end{table} \section{A simple model of QE} This is a simplified version of the New Keynesian model of bond market segmentation that appears in IkedaLiMavroeidisZanetti2020 and is based on ChenCurdiaFerrero2012. The economy consists of two types of households. A fraction $\omega_{r}$ of type `r' households can only trade long-term government bonds. The remaining $1-\omega_{r}$ households of type `u' can purchase both short-term and long-term government bonds, the latter subject to a trading cost $\zeta_{t}$. This trading cost gives rise to a term premium, i.e., a spread between long-term and short-term yields, that the central bank can manipulate by purchasing long-term bonds. The term premium affects aggregate demand through the consumption decisions of constrained households. This generates an unconventional monetary policy channel. The transmission mechanism of monetary policy is obtained from the equilibrium conditions of households and firms in the economy. Households choose consumption to maximize an isoelastic utility function and firms set prices subject to Calvo frictions. These give rise to an Euler equation for output and a Phillips curve, respectively. Equation ((ref)) in the paper can be derived by combining those two equations. I will derive the Euler equation in some detail in order to illustrate the origins of the QE channel. The Phillips curve derivation is standard and is therefore omitted. Up to a loglinear approximation, the relevant first-order conditions of the households' optimization problem can be written as \begin{align} 0 & =E_{t}\left[ -\frac{1}{\sigma}\left( \hat{c}_{t+1}^{u}-\hat{c}_{t} ^{u}\right) +\hat{r}_{t}-\pi_{t+1}\right] ,\\ \frac{\zeta}{1+\zeta}\hat{\zeta}_{t} & =E_{t}\left[ -\frac{1}{\sigma }\left( \hat{c}_{t+1}^{u}-\hat{c}_{t}^{u}\right) +\hat{R}_{L,t+1}-\pi _{t+1}\right] ,\\ 0 & =E_{t}\left[ -\frac{1}{\sigma}\left( \hat{c}_{t+1}^{r}-\hat{c}_{t} ^{r}\right) +\hat{R}_{L,t+1}-\pi_{t+1}\right] , \end{align} where $\sigma$ is the elasticity of intertemporal substitution, $\zeta$ is the steady state value of $\zeta_{t},$ hatted variables denote log-deviations from steady state, $c_{t}^{j}$ is consumption of household $j\in\left\{ u,r\right\} ,$ $r_{t}$ is the short-term nominal interest rate, and $R_{L,t}$ is the gross yield on long-term government bonds from period $t-1$ to $t$.\footnote{I do not put a hat over $\pi_{t}$ because I assume a zero inflation target for simplicity.} Goods market clearing yields \begin{equation} \hat{y}_{t}=\omega_{r}\hat{c}_{t}^{r}+\left( 1-\omega_{r}\right) \hat{c} _{t}^{u}, \end{equation} where $y_{t}$ is output, and I\ have assumed, for simplicity, that in steady state $c^{u}=c^{r}$, which implies $c^{u}=c^{r}=y.$ Multiplying ((ref)) and ((ref)) by $\left( 1-\omega_{r}\right) $ and $\omega_{r},$ respectively, and adding them yields \begin{equation} \hat{y}_{t}=E_{t}\hat{y}_{t+1}-\sigma E_{t}\left[ \left( 1-\omega _{r}\right) \hat{r}_{t}+\omega_{r}\hat{R}_{L,t+1}-\pi_{t+1}\right] . \end{equation} Subtracting ((ref)) from ((ref)) yields \begin{equation} E_{t}\left( \hat{R}_{L,t+1}\right) =\hat{r}_{t}+\frac{\zeta}{1+\zeta} \hat{\zeta}_{t}, \end{equation} which establishes that the term premium between long and short yields is proportional to $\hat{\zeta}_{t}.$ Substituting for $E_{t}\left( \hat {R}_{L,t+1}\right) $ in ((ref)) using ((ref)) yields \begin{equation} \hat{y}_{t}=E_{t}\hat{y}_{t+1}-\sigma\left( \hat{r}_{t}+\omega_{r}\frac {\zeta}{1+\zeta}\hat{\zeta}_{t}\right) +\sigma E_{t}\left( \pi_{t+1}\right) \end{equation} Next, assume that the cost of trading long-term bonds depends on their supply, $b_{L,t},$ i.e., \[ \hat{\zeta}_{t}=\rho_{\zeta}\hat{b}_{L,t},\quad\rho_{\zeta}\geq0. \] Substituting for $\hat{\zeta}_{t}$ in ((ref)) yields the Euler equation \begin{equation} \hat{y}_{t}=E_{t}\hat{y}_{t+1}-\sigma\left( \hat{r}_{t}+\omega_{r}\frac {\zeta}{1+\zeta}\rho_{\zeta}\hat{b}_{L,t}\right) +\sigma E_{t}\left( \pi_{t+1}\right) . \end{equation} The second equation is a standard New Keynesian Phillips curve that links inflation to output: \begin{equation} \pi_{t}=\delta E_{t}\left( \pi_{t+1}\right) +\varpi\hat{y}_{t} +\varepsilon_{1t} \end{equation} where $\delta$ is the average discount factor of the two households, $\varpi\geq0$ is a parameter that depends on the degree of price stickiness (the Calvo parameter), and $\varepsilon_{1t}$ is proportional to an \textit{i.i.d.} technology shock. Substituting for $\hat{y}_{t}$ in ((ref)) using ((ref)) yields \begin{equation} \pi_{t}=\left( \delta+\sigma\varpi\right) E_{t}\left( \pi_{t+1}\right) +\varpi E_{t}\left( \hat{y}_{t+1}\right) -\varpi\sigma\left( \hat{r} _{t}+\omega_{r}\frac{\zeta}{1+\zeta}\rho_{\zeta}\hat{b}_{L,t}\right) +\varepsilon_{1t}. \end{equation} Finally, at an equilibrium in which inflation and output depend only on the exogenous shocks $\varepsilon_{t}=\left( \varepsilon_{1t},\varepsilon _{2t}\right) ^{\prime},$ which are the only state variables in the system, and when the shocks have no memory, $E_{t}\left( \pi_{t+1}\right) $ and $E_{t}\left( \hat{y}_{t+1}\right) $ will be equal to the corresponding unconditional expectations, which are constants.\footnote{Such an equilibrium always exists if the volatility of the shocks is not too large, see Mendes2011.} Therefore, ((ref)) reduces to eq. ((ref)) in the paper by setting $c=\left( \delta+\sigma \varpi\right) E\left( \pi_{t+1}\right) +\varpi E\left( \hat{y} _{t+1}\right) ,$ $\beta=-\varpi\sigma$, $\hat{r}_{t}=r_{t}-r^{n},$ where $r^{n}$ is the discount rate of the unconstrained households, and $\varphi=\omega_{r}\frac{\zeta}{1+\zeta}\rho_{\zeta}\beta$, and dropping the hat from $b_{L,t}$ for simplicity. The parameter $\varphi$ depends on the fraction of constrained households, $\omega_{r}$, and the sensitivity of the term premium to long-term asset holdings, $\frac{\zeta }{1+\zeta}\rho_{\zeta}$. \section{Forward guidance rules} DebortoliGaliGambetti2019 discuss the following two inertial policy rules \begin{equation} r_{t}=\max\left\{ 0,\phi_{r}r_{t-1}+\left( 1-\phi_{r}\right) \left( \rho+\phi_{\pi} \pi_{t}+\phi_{y}\Delta y_{t}\right) \right\} \end{equation} and \begin{subequations} \begin{align} r_{t} & =\max\left( 0,r_{t}^{\ast}\right) \\ r_{t}^{\ast} & =\phi_{r}r_{t-1}^{\ast}+\left( 1-\phi_{r}\right) \left( \rho+\phi_{\pi} \pi_{t}+\phi_{y}\Delta y_{t}\right) , \end{align} where I have set the inflation target to zero, and $\Delta y_{t}$ is output growth. Both of these rules are nested within equation ((ref)) of the CKSVAR, with $Y_{1t}=\left( \pi_{t},\Delta y_{t}\right) ^{\prime}\,$and $Y_{2t}=r_{t}.$ Rule ((ref)) sets the coefficients on $Y_{2,t-1}$ and $Y_{2,t-1}^{\ast}$ as $B_{22}=\phi_{r}$ and $B_{22}^{\ast}=0,$ respectively, while rule ((ref)) sets them as $B_{22}=0$ and $B_{22}^{\ast}=\phi_{r}.$ DebortoliGaliGambetti2019 argue rule ((ref)) is consistent with forward guidance, because it will tend to keep interest rates at zero for longer than rule ((ref)). It also ensures policy reaction is the same across regimes, and so it is consistent with the ZLB irrelevance hypothesis that the paper puts forward. ReifschneiderWilliams2000 propose a slightly more elaborate policy rule for forward guidance: \end{subequations} \begin{subequations} \begin{align} r_{t}^{\ast} & =r_{t}^{Taylor}-\alpha Z_{t},\quad Z_{t}=Z_{t-1}+d_{t},\quad d_{t}:=r_{t}-r_{t}^{Taylor},\\ r_{t} & =\max\left( r_{t}^{\ast},0\right) ,\nonumber\\ r_{t}^{Taylor} & =\rho+\phi_{\pi}\pi_{t}+\phi_{y}y_{t} \end{align} where $y_{t}$ is the output gap, and the inflation target is zero. Differencing ((ref)) yields \end{subequations} \[ r_{t}^{\ast}=r_{t-1}^{\ast}+\Delta r_{t}^{Taylor}-\alpha\left( r_{t} -r_{t}^{Taylor}\right) . \] Substituting for $r_{t}^{Taylor}$ using ((ref)) yields \begin{align*} r_{t}^{\ast} & =r_{t-1}^{\ast}+\phi_{\pi}\Delta\pi_{t}+\phi_{y}\Delta y_{t}-\alpha r_{t}+\alpha\left( \rho+\phi_{\pi}\pi_{t}+\phi_{y}y_{t}\right) \\ & =\alpha\rho-\alpha r_{t}+\left( 1+\alpha\right) \left( \phi_{\pi}\pi _{t}+\phi_{y}y_{t}\right) -\left( \phi_{\pi}\pi_{t-1}+\phi_{y} y_{t-1}\right) +r_{t-1}^{\ast}. \end{align*} This is again nested within equation ((ref)) of the CKSVAR with $Y_{1t}=\left( \pi_{t},y_{t}\right) ^{\prime},$ $Y_{2t}=r_{t},$ $Y_{2t}^{\ast}=r_{t}^{\ast},$ $X_{1t}=\left( 1,\pi_{t-1},y_{t-1}\right) ^{\prime},$ $X_{2t}=r_{t-1},$ $X_{2t}^{\ast}=r_{t-1}^{\ast},$ and parameters $A_{21}=-\left( 1+\alpha\right) \left( \phi_{\pi},\phi_{y}\right) ,$ $A_{22}=\alpha,$ $A_{22}^{\ast}=1,$ $B_{21}=\left( \alpha\rho,-\phi_{\pi },-\phi_{y}\right) ,$ $B_{22}=0$ and $B_{22}^{\ast}=1.$ More examples of forward guidance policy rules that are nested within the CKSVAR are discussed in IkedaLiMavroeidisZanetti2020. \section{Computational details} \subsection{Likelihood} To compute the likelihood, we need to obtain the prediction error densities. The first step is to write the model in state-space form. Define \[ s_{t}= \begin{pmatrix} \mathbf{y}_{t}\\ \vdots\\ \mathbf{y}_{t-p+1} \end{pmatrix} ,\quad\underset{\left( k+1\right) \times1}{\mathbf{y}_{t}}= \begin{pmatrix} Y_{t}\\ \overline{Y}_{2t}^{\ast} \end{pmatrix} , \] and write the state transition equation as \begin{equation} s_{t}=F\left( s_{t-1},u_{t};\psi\right) = \begin{pmatrix} F_{1}\left( s_{t-1},u_{t};\psi\right) \\ \mathbf{y}_{t-1}\\ \vdots\\ \mathbf{y}_{t-p+1} \end{pmatrix} , \end{equation} \[ F_{1}\left( s_{t-1},u_{t};\psi\right) = \begin{pmatrix} \overline{C}_{1}X_{t}+\overline{C}_{1}^{\ast}\overline{X}_{t}^{\ast} +u_{1t}-\widetilde{\beta}D_{t}\left( \overline{C}_{2}X_{t}+\overline{C} _{2}^{\ast}\overline{X}_{t}^{\ast}+u_{2t}-b\right) \\ \max\left( b,\overline{C}_{2}X_{t}+\overline{C}_{2}^{\ast}\overline{X} _{t}^{\ast}+u_{2t}\right) \\ \overline{C}_{2}X_{t}+\overline{C}_{2}^{\ast}\overline{X}_{t}^{\ast}+u_{2t} \end{pmatrix} , \] and the observation equation as \begin{equation} Y_{t}= \begin{pmatrix} I_{k} & 0_{k\times1+\left( p-1\right) \left( k+1\right) } \end{pmatrix} s_{t}. \end{equation} Next, I will derive the predictive density and mass functions. With Gaussian errors, the joint predictive density of $Y_{t}$ corresponding to the observations with $D_{t}=0$ is: \begin{multline} f_{0}\left( Y_{t}|s_{t-1},\psi\right) =\left\vert \Omega\right\vert ^{-1/2}\exp\left\{ -\frac{1}{2}tr\left( \left( Y_{t}-\overline{C} X_{t}-\overline{C}^{\ast}\overline{X}_{t}^{\ast}\right) \right.\right. \\\left.\left. \left( Y_{t}-\overline{C}X_{t}-\overline{C}^{\ast}\overline{X}_{t}^{\ast}\right) ^{\prime}\Omega^{-1}\right) \right\} . \end{multline} At $D_{t}=1,$ the predictive density of $Y_{1t}$ can be written as: \begin{align} f_{1}\left( Y_{1t}|s_{t-1},\psi\right) & :=\left\vert \Xi_{1}\right\vert ^{-\frac{1}{2}}\exp\left[ -\frac{1}{2}\left( Y_{1t}-\mu_{1t}\right) ^{\prime}\Xi_{1}^{-1}\left( Y_{1t}-\mu_{1t}\right) \right] \\ \mu_{1t} & :=\widetilde{\beta}b+\left( \overline{C}_{1}-\widetilde{\beta }\overline{C}_{2}\right) X_{t}+\left( \overline{C}_{1}^{\ast} -\widetilde{\beta}\overline{C}_{2}^{\ast}\right) \overline{X}_{t}^{\ast }\\ \Xi_{1} & :=\Omega_{1.2}+\widetilde{\delta}\widetilde{\delta}^{\prime} \tau^{2}= \begin{pmatrix} I_{k-1} & -\widetilde{\beta} \end{pmatrix} \Omega\binom{I_{k-1}}{-\widetilde{\beta}^{\prime}},\quad\widetilde{\delta }=\Omega_{12}\omega_{22}^{-1}-\widetilde{\beta}, \end{align} where $\Omega_{1.2}=\Omega_{11}-\Omega_{12}\omega_{22}^{-1}\Omega_{21},$ and $\tau=\sqrt{\omega_{22}}$. Next, \begin{align} u_{2t}|Y_{1t},s_{t-1} & \sim N\left( \mu_{2t},\tau_{2}^{2}\right) ,\text{ with}\\ \mu_{2t} & :=\tau^{2}\widetilde{\delta}^{\prime}\Xi_{1}^{-1}\left( Y_{1t}-\mu_{1t}\right) ,\quad\tau_{2}=\tau\sqrt{\left( 1-\tau^{2} \widetilde{\delta}^{\prime}\Xi_{1}^{-1}\widetilde{\delta}\right) }. \end{align} Hence, \begin{equation} \Pr\left( D_{t}=1|Y_{1t},s_{t-1},\psi\right) =\Phi\left( \frac {b-\overline{C}_{2}X_{t}-\overline{C}_{2}^{\ast}\overline{X}_{t}^{\ast} -\mu_{2t}}{\tau_{2}}\right) . \end{equation} In the case of the KSVAR model, there are no latent lags ($\overline{C}^{\ast }=0,$ $\overline{C}=C$), so the log-likelihood is available analytically: \begin{multline} \log L\left( \psi\right) =\sum_{t=1}^{T} \left( 1-D_{t}\right) \log f_{0}\left( Y_{t}|s_{t-1},\psi\right) \\ . +\sum_{t=1}^{T}D_{t}\log\left( f_{1}\left( Y_{1t}|s_{t-1},\psi\right) \Phi\left( \frac{b-C_{2}X_{t}-\mu_{2t}}{\tau_{2} }\right) \right) \end{multline} where $f_{0}\left( Y_{t}|s_{t-1},\theta\right) $ and $f_{1}\left( Y_{1t}|s_{t-1},\theta\right) $ are given by ((ref)) and ((ref) ), resp., with $\overline{C}^{\ast}=0$. The likelihood for the unrestricted CKSVAR ($\overline{C}^{\ast}\neq0$) can be computed approximately by simulation (particle filtering). I provide two different simulation algorithms. The first is a sequential importance sampler (SIS), proposed originally by Lee1999 for the univariate dynamic Tobit model. It is extended here to the CKSVAR model. The second algorithm is a fully adapted particle filter (FAPF), which is a sequential importance resampling algorithm designed to address the sample degeneracy problem. It is proposed by MalikPitt2011 and is a special case of the auxiliary particle filter developed by PittShephard1999. Both algorithms require sampling from the predictive density of $\overline {Y}_{2t}^{\ast}$ conditional on $Y_{1t},D_{t}=1$ and $s_{t-1}.$ From ((ref)) and ((ref)), we see that this is a truncated Normal with original mean $\mu_{2t}^{\ast}=\overline{C}_{2} X_{t}+\overline{C}_{2}^{\ast}\overline{X}_{t}^{\ast}+\mu_{2t}$ and standard deviation $\tau_{2},$ where $\mu_{2t},\tau_{2}$ are given in ((ref)), i.e., \begin{equation} f_{2}\left( Y_{2t}^{\ast}|Y_{1t},D_{t}=1,s_{t-1},\psi\right) =TN\left( \mu_{2t}^{\ast},\tau_{2},\overline{Y}_{2t}^{\ast}<b\right) \end{equation} Draws from this truncated distribution can be obtained using, for instance, the procedure in Lee1999. Let $\xi_{t}^{\left( j\right) }\sim U\left[ 0,1\right] $ be \textit{i.i.d.} uniform random draws, $j=1,...,M$. Then, a draw from $\overline{Y}_{2t}^{\ast}|Y_{1t},s_{t-1},\overline{Y} _{2t}^{\ast}<b$ is given by \begin{equation} \overline{Y}_{2t}^{\ast\left( j\right) }=\mu_{2t}^{\ast}+\tau_{2}\Phi ^{-1}\left[ \xi_{t}^{\left( j\right) }\Phi\left( \frac{b-\mu_{2t}^{\ast} }{\tau_{2}}\right) \right] . \end{equation} \begin{algorithm} [SIS]Sequential Importance Sampler \begin{enumerate} • \emph{Initialization.} For $j=1:M,$ set $W_{0}^{j}=1$ and $s_{0} ^{j}=\left( \mathbf{y}_{0}^{j},\ldots,\mathbf{y}_{-p+1}^{j}\right) ,$ with $\mathbf{y}_{-s}^{j}=\left( Y_{0}^{\prime},Y_{2,0}\right) ^{\prime},$ for $s=0,...,p-1.$ (in other words, initialize $\overline{Y}_{2,-s}^{\ast}$ at the observed values of $Y_{2,-s}$). • \emph{Recursion.} For $t=1:T$: \begin{enumerate} • For $j=1:M,$ compute the incremental weights \[ w_{t-1|t}^{j}=p\left( Y_{t}|s_{t-1}^{j},\psi\right) =\left\{ \begin{array} [c]{ll} f_{0}\left( Y_{t}|s_{t-1}^{j},\psi\right) , & \text{if }D_{t}=0\\ f_{1}\left( Y_{1t}|s_{t-1}^{j},\psi\right) \Pr\left( D_{t}=1|Y_{1t} ,s_{t-1}^{j},\psi\right) , & \text{if }D_{t}=1 \end{array} \right. \] where $f_{0},f_{1},$ and $\Pr\left( D_{t}=1|Y_{1t},s_{t-1};\psi\right) $ are given by ((ref)), ((ref)), and ((ref)), resp., and \[ S_{t}=\frac{1}{M}\sum_{j=1}^{M}w_{t-1|t}^{j}W_{t-1}^{j} \] • Sample $s_{t}^{j}$ randomly from $p\left( s_{t}|s_{t-1}^{j} ,Y_{t}\right) $. That is, $s_{t}^{j}=\left( \mathbf{y}_{t}^{j} ,\mathbf{y}_{t-1}^{j},\ldots,\mathbf{y}_{t-p}^{j}\right) $ where $\mathbf{y}_{t}^{j}=\left( Y_{t}^{\prime},\overline{Y}_{2t}^{\ast\left( j\right) }\right) \ $and $\overline{Y}_{2t}^{\ast\left( j\right) }$ is a draw from $f_{2}\left( Y_{2t}^{\ast}|Y_{1t},D_{t}=1,s_{t-1}^{j},\psi\right) $ using ((ref)). • Update the weights: \[ W_{t}^{j}=\frac{w_{t-1|t}^{j}W_{t-1}^{j}}{S_{t}}. \] \end{enumerate} • Likelihood approximation \[ \log\widehat{p}(Y_{T}|\psi)=\sum_{t=1}^{T}\log S_{t} \] \end{enumerate} \end{algorithm} If the draws $\xi_{t}^{\left( j\right) }$ are kept fixed across different values of $\psi$, the simulated likelihood in step 3 is smooth. Note that when $k=1$ and $Y_{t}=Y_{2t}$ (no $Y_{1t}$ variables), the model reduces to a univariate dynamic Tobit model, and Algorithm (ref) reduces exactly to the sequential importance sampler proposed by Lee1999. A possible weakness of this algorithm is sample degeneracy, which arises when all but a few weights $W_{t}^{J}$ are zero. To gauge possible sample degeneracy, we can look at the effective sample size (ESS), as recommended by HerbstSchorfheide2015 \begin{equation} ESS_{t}=\frac{M}{\frac{1}{M}\sum_{j=1}^{M}\left( W_{t}^{j}\right) ^{2}}. \end{equation} Next, I turn to the FAPF algorithm. \begin{algorithm} [FAPF]Fully Adapted Particle Filter \begin{enumerate} • \emph{Initialization.} For $j=1:M,$ set $s_{0}^{j}=\left( \mathbf{y}_{0}^{j},\ldots,\mathbf{y}_{-p+1}^{j}\right) ,$ with $\mathbf{y} _{-s}^{j}=\left( Y_{0}^{\prime},Y_{2,0}\right) ^{\prime},$ for $s=0,...,p-1.$ (in other words, initialize $\overline{Y}_{2,-s}^{\ast}$ at the observed values of $Y_{2,-s}$). • \emph{Recursion.} For $t=1:T$: \begin{enumerate} • For $j=1:M,$ compute \[ w_{t-1|t}^{j}=p\left( Y_{t}|s_{t-1}^{j},\psi\right) =\left\{ \begin{array} [c]{ll} f_{0}\left( Y_{t}|s_{t-1}^{j},\psi\right) , & \text{if }D_{t}=0\\ f_{1}\left( Y_{1t}|s_{t-1}^{j},\psi\right) \Pr\left( D_{t}=1|Y_{1t} ,s_{t-1}^{j},\psi\right) , & \text{if }D_{t}=1 \end{array} \right. \] where $f_{0},f_{1},$ and $\Pr\left( D_{t}=1|Y_{1t},s_{t-1};\psi\right) $ are given by ((ref)), ((ref)), and ((ref)), resp., and \[ \pi_{t-1|t}^{j}=\frac{w_{t-1|t}^{j}}{\sum_{j=1}^{M}w_{t-1|t}^{j}}. \] • For $j=1:M$, sample $k_{j}$ randomly from the multinomial distribution $\left\{ j,\pi_{t-1|t}^{j}\right\} .$ Then, set $\tilde{s}_{t-1}^{j} =s_{t-1}^{k_{j}}$ (this applies only to the elements in $s_{t-1}^{j}$ that correspond to $X_{t}^{\ast j},$ since all the other elements are observed and constant across all $j$. That is, $\tilde{s}_{t-1}^{j}=\left( \mathbf{\tilde {y}}_{t-1}^{j},\allowbreak\ldots\allowbreak,\mathbf{\tilde{y}}_{t-p}^{j}\right) ,$ $\mathbf{\tilde {y}}_{t-s}^{j}=\left( Y_{t-1}^{\prime},\overline{Y}_{2,t-s}^{\ast\left( k_{j}\right) }\right) ,$ $s=1,...,p.$) • For $j=1:M$, sample $s_{t}^{j}$ randomly from $p\left( s_{t}|\tilde {s}_{t-1}^{j},Y_{t}\right) $. That is, $s_{t}^{j}=\left( \mathbf{y}_{t} ^{j},\allowbreak\ldots,\allowbreak\mathbf{\tilde{y}}_{t-p}^{j}\right) $ where $\mathbf{y}_{t}^{j}=\left( Y_{t}^{\prime},\overline{Y}_{2t} ^{\ast\left( j\right) }\right) \ $and $\overline{Y}_{2t}^{\ast\left( j\right) }$ is a draw from $f_{2}\left( Y_{2t}^{\ast}|Y_{1t},D_{t} =1,\tilde{s}_{t-1}^{j},\psi\right) $ using ((ref)). \end{enumerate} • Likelihood approximation \[ \ln\hat{p}\left( Y_{T}|\psi\right) =\sum_{t=1}^{T}\ln\left( \frac{1}{M} \sum_{j=1}^{M}w_{t-1|t}^{j}\right) \] \end{enumerate} \end{algorithm} Many of the generic particle filtering algorithms used in the macro literature, described in HerbstSchorfheide2015, are inapplicable in a censoring context because of the absence of measurement error in the observation equation. It is, of course, possible to introduce a small measurement error in $Y_{2t},$ so that the constraint $Y_{2t}\geq b$ is not fully respected, but there is no reason to expect other particle filters discussed in HerbstSchorfheide2015 to estimate the likelihood more accurately than the FAPF algorithm described above. Moments or quantiles of the filtering or smoothing distribution of any function $h\left( \cdot\right) $ of the latent states $s_{t}$ can be computed using the drawn sample of particles. When we use Algorithm (ref), simple average or quantiles of $h\left( s_{t}^{j}\right) $ produce the requisite average or quantiles of $h\left( s_{t}\right) $ conditional on $Y_{1},\ldots,Y_{t}$ (the filtering density). For particles generated using Algorithm (ref), we need to take weighted averages using the importance sampling weights $W_{t}.$ Smoothing estimates of $h\left( s_{t}^{j}\right) $ can be obtained using weights $W_{T}.$ \subsection{Computation of the identified set} Substitute for $\overline{\gamma}$ in ((ref)) using Proposition (ref) to get \begin{equation} \widetilde{\beta}=\left( 1-\xi\right) \left( I-\xi\overline{\beta}\left( \Omega_{12}^{\prime}-\Omega_{22}\overline{\beta}^{\prime}\right) \left( \Omega_{11}-\Omega_{12}\overline{\beta}^{\prime}\right) ^{-1}\right) ^{-1}\overline{\beta}. \end{equation} For each value of $\xi\in\lbrack0,1),$ the above equation defines a correspondence from $\Re^{k-1}$ to $\Re^{k-1}$. The range of $\overline{\beta }$ can then be obtained numerically by solving ((ref)) for $\overline{\beta}$ as a function of the reduced-form parameters and $\xi$ for each value of $\xi,$ and gathering all the solutions in the set. Rearranging ((ref)) yields \begin{equation} \widetilde{\beta}=\xi\overline{\beta}\left( \Omega_{12}^{\prime}-\Omega _{22}\overline{\beta}^{\prime}\right) \left( \Omega_{11}-\Omega _{12}\overline{\beta}^{\prime}\right) ^{-1}\widetilde{\beta}+\left( 1-\xi\right) \overline{\beta}. \end{equation} Note that \[ \left( \Omega_{11}-\Omega_{12}\overline{\beta}^{\prime}\right) ^{-1} =\Omega_{11}^{-1}+\Omega_{11}^{-1}\Omega_{12}\left( 1-\overline{\beta }^{\prime}\Omega_{11}^{-1}\Omega_{12}\right) ^{-1}\overline{\beta}^{\prime }\Omega_{11}^{-1}. \] Hence, \begin{align*} & \left( \Omega_{12}^{\prime}-\Omega_{22}\overline{\beta}^{\prime}\right) \left( \Omega_{11}-\Omega_{12}\overline{\beta}^{\prime}\right) ^{-1}\\ & =\left( \Omega_{12}^{\prime}-\Omega_{22}\overline{\beta}^{\prime}\right) \Omega_{11}^{-1}+\frac{\left( \Omega_{12}^{\prime}\Omega_{11}^{-1} \Omega_{12}-\Omega_{22}\overline{\beta}^{\prime}\Omega_{11}^{-1}\Omega _{12}\right) \overline{\beta}^{\prime}\Omega_{11}^{-1}}{1-\overline{\beta }^{\prime}\Omega_{11}^{-1}\Omega_{12}}\\ & =\frac{\left( \Omega_{12}^{\prime}-\Omega_{22}\overline{\beta}^{\prime }\right) \Omega_{11}^{-1}+\Omega_{12}^{\prime}\Omega_{11}^{-1}\left( \Omega_{12}\overline{\beta}^{\prime}\Omega_{11}^{-1}-\overline{\beta}^{\prime }\Omega_{11}^{-1}\Omega_{12}I_{k-1}\right) }{1-\overline{\beta}^{\prime }\Omega_{11}^{-1}\Omega_{12}}. \end{align*} Substituting this back into ((ref)), we get \[ \widetilde{\beta}=\xi\overline{\beta}\frac{\left( \Omega_{12}^{\prime} -\Omega_{22}\overline{\beta}^{\prime}\right) \Omega_{11}^{-1}+\Omega _{12}^{\prime}\Omega_{11}^{-1}\left( \Omega_{12}\overline{\beta}^{\prime }\Omega_{11}^{-1}-\overline{\beta}^{\prime}\Omega_{11}^{-1}\Omega_{12} I_{k-1}\right) }{1-\overline{\beta}^{\prime}\Omega_{11}^{-1}\Omega_{12} }\widetilde{\beta}+\left( 1-\xi\right) \overline{\beta}. \] Multiplying both sides by $1-\overline{\beta}^{\prime}\Omega_{11}^{-1}\Omega_{12}$ yields \begin{align*} \widetilde{\beta}-\overline{\beta}^{\prime}\Omega_{11}^{-1}\Omega _{12}\widetilde{\beta} & =\xi\overline{\beta}\Omega_{12}^{\prime }\widetilde{\beta}-\xi\overline{\beta}\Omega_{22}\overline{\beta}^{\prime }\Omega_{11}^{-1}\widetilde{\beta}+\xi\overline{\beta}\Omega_{12}^{\prime }\Omega_{11}^{-1}\Omega_{12}\overline{\beta}^{\prime}\Omega_{11} ^{-1}\widetilde{\beta}\\ & -\xi\overline{\beta}\overline{\beta}^{\prime}\Omega_{11}^{-1}\Omega _{12}\Omega_{12}^{\prime}\Omega_{11}^{-1}\widetilde{\beta}+\left( 1-\xi\right) \overline{\beta}\left( 1-\overline{\beta}^{\prime}\Omega _{11}^{-1}\Omega_{12}\right). \end{align*} Rearranging, we have \begin{align*} \widetilde{\beta} & =\widetilde{\beta}\Omega_{12}^{\prime}\Omega_{11} ^{-1}\overline{\beta}+\xi\Omega_{12}^{\prime}\widetilde{\beta}\overline{\beta }+\left( 1-\xi\right) \overline{\beta}-\left( 1-\xi\right) \overline {\beta}\overline{\beta}^{\prime}\Omega_{11}^{-1}\Omega_{12}\\ & +\overline{\beta}\overline{\beta}^{\prime}\Omega_{11}^{-1}\widetilde{\beta }\xi\Omega_{12}^{\prime}\Omega_{11}^{-1}\Omega_{12}-\overline{\beta} \overline{\beta}^{\prime}\Omega_{11}^{-1}\widetilde{\beta}\xi\Omega _{22}-\overline{\beta}\overline{\beta}^{\prime}\Omega_{11}^{-1}\Omega _{12}\Omega_{12}^{\prime}\Omega_{11}^{-1}\widetilde{\beta}\xi \\ & =\left( \widetilde{\beta}\Omega_{12}^{\prime} \Omega_{11}^{-1}+\left( \xi\Omega_{12}^{\prime}\widetilde{\beta} +1-\xi\right) I_{k-1}\right) \overline{\beta}\\ & +\overline{\beta}\overline{\beta}^{\prime}\Omega_{11}^{-1}\left( \left( \left( \Omega_{12}^{\prime}\Omega_{11}^{-1}\Omega_{12}-\Omega_{22}\right) I_{k-1}-\Omega_{12}\Omega_{12}^{\prime}\Omega_{11}^{-1}\right) \widetilde{\beta}\xi-\left( 1-\xi\right) \Omega_{12}\right). \end{align*} This can be written as \begin{equation} \widetilde{\beta}-\tilde{A}\overline{\beta}+\overline{\beta}\overline{\beta }^{\prime}\tilde{b}=0, \end{equation} where \begin{align*} \tilde{b} & :=-\Omega_{11}^{-1}\left( \left( \left( \Omega_{12}^{\prime} \Omega_{11}^{-1}\Omega_{12}-\Omega_{22}\right) I_{k-1}-\Omega_{12}\Omega _{12}^{\prime}\Omega_{11}^{-1}\right) \widetilde{\beta}\xi-\left( 1-\xi\right) \Omega_{12}\right) , \text{ and}\\ \bar{A} & :=\widetilde{\beta}\Omega_{12}^{\prime}\Omega_{11}^{-1}+\left( \xi\Omega_{12}^{\prime}\widetilde{\beta}+1-\xi\right) I_{k-1}. \end{align*} Define \[ z:=\tilde{b}^{\prime}x\text{ \ \ and \ \ }w:=\tilde{b}_{\perp}^{\prime}x, \] where $\tilde{b}_{\perp}^{\prime}\tilde{b}_{\perp}=1$ and $\tilde{b}_{\perp }^{\prime}\tilde{b}=0$. Hence, rewrite ((ref)) as \[ \widetilde{\beta}-\tilde{A}\tilde{b}\left( \tilde{b}^{\prime}\tilde {b}\right) ^{-1}z-\tilde{A}\tilde{b}_{\perp}w+\tilde{b}\left( \tilde {b}^{\prime}\tilde{b}\right) ^{-1}z^{2}+\tilde{b}_{\perp}wz=0. \] Premultiply by $\tilde{b}_{\perp}^{\prime}$ to get \[ \tilde{b}_{\perp}^{\prime}\widetilde{\beta}-\tilde{b}_{\perp}^{\prime} \tilde{A}\tilde{b}\left( \tilde{b}^{\prime}\tilde{b}\right) ^{-1}z-\tilde {b}_{\perp}^{\prime}\tilde{A}\tilde{b}_{\perp}w+wz=0. \] Solve that for $w$ to get \begin{align*} w & =\left( \tilde{b}_{\perp}^{\prime}\tilde{A}\tilde{b}_{\perp}-z\right) ^{-1}\left( \tilde{b}_{\perp}^{\prime}\widetilde{\beta}-\tilde{b}_{\perp }^{\prime}\tilde{A}\tilde{b}\left( \tilde{b}^{\prime}\tilde{b}\right) ^{-1}z\right) \\ & =C_{0}\left( z\right) ^{-1}c_{1}\left( z\right) , \end{align*} with \begin{align*} C_{0}\left( z\right) & :=\left( \tilde{b}_{\perp}^{\prime}\tilde{A} \tilde{b}_{\perp}-z\right) ,\text{ \ and}\\ c_{1}\left( z\right) & :=\tilde{b}_{\perp}^{\prime}\widetilde{\beta }-\tilde{b}_{\perp}^{\prime}\tilde{A}\tilde{b}\left( \tilde{b}^{\prime} \tilde{b}\right) ^{-1}z, \end{align*} provided that $\det\left( C_{0}\left( z\right) \right) \neq0.$ Next, premultiply ((ref)) by $\tilde{b}^{\prime}$ and substitute for $w$ to get \begin{equation} \tilde{b}^{\prime}\widetilde{\beta}-\tilde{b}^{\prime}\tilde{A}\tilde {b}\left( \tilde{b}^{\prime}\tilde{b}\right) ^{-1}z-\tilde{b}^{\prime} \tilde{A}\tilde{b}_{\perp}C_{0}\left( z\right) ^{-1}c_{1}\left( z\right) +z^{2}=0. \end{equation} Now, notice that $C_{0}\left( z\right) ^{-1}=C_{0}\left( z\right) ^{adj}/\det\left( C_{0}\left( z\right) \right) ,$ where $C^{adj}$ is the adjoint of a square matrix $C.$ Moreover, since $C_{0}\left( z\right) $ is of dimension $k-2$ and its elements are linear in $z,$ $\det\left( C_{0}\left( z\right) \right) $ is a polynomial in $z$ of order at most $k-2,$ and the elements of $C_{0}\left( z\right) ^{adj}$ are polynomials in $z$ of order at most $k-3.$ For $k=2,$ $w$ is empty, so ((ref)) is simply a quadratic in $z.$ When $k>2,$ $\det\left( C_{0}\left( z\right) \right) $ is nonzero and we can multiply ((ref)) by it to get \begin{align} 0 & =\tilde{b}^{\prime}\widetilde{\beta}\det\left( C_{0}\left( z\right) \right) +\tilde{b}^{\prime}\tilde{A}\tilde{b}\left( \tilde{b}^{\prime} \tilde{b}\right) ^{-1}z\det\left( C_{0}\left( z\right) \right) \nonumber\\ & -\tilde{b}^{\prime}\tilde{A}\tilde{b}_{\perp}C_{0}\left( z\right) ^{adj}c_{1}\left( z\right) +\det\left( C_{0}\left( z\right) \right) z^{2}, \end{align} This is a polynomial equation of order $k$ and has at most $k$ solutions, denoted $z_{i},$ say. Then, the solutions for $\overline{\beta}$ are given by \begin{align} \overline{\beta}_{i} & =\left[ \tilde{b}\left( \tilde{b}^{\prime}\tilde{b}\right) ^{-1},\tilde{b}_{\perp}\right] \begin{pmatrix} z_{i}\\ C_{0}\left( z_{i}\right) ^{-1}c_{1}\left( z_{i}\right) \end{pmatrix} \nonumber\\ & =\tilde{b}\left( \tilde{b}^{\prime}\tilde{b}\right) ^{-1}z_{i}+\tilde {b}_{\perp}C_{0}\left( z_{i}\right) ^{-1}c_{1}\left( z_{i}\right). \end{align} Below I give some special cases. \textit{Case $k=2$:} In this case, $w$ is empty, $\overline{\beta}$ is a scalar, and the equation ((ref)) is a quadratic \[ \widetilde{\beta}-\tilde{A}\overline{\beta}+\overline{\beta}^{2}\tilde{b}=0. \] If $\tilde{A}^{2}-4\widetilde{\beta}\tilde{b}>0,$ the two real solutions are$ \overline{\beta}_{1,2}=\frac{\tilde{A}\pm\sqrt{\tilde{A}^{2}-4\widetilde{\beta}\tilde{b}} }{2\tilde{b}}.$ \textit{Case $k=3$:} In this case, $w$ is a scalar, and the equation ((ref)) can be written as a cubic in $z,$ i.e., \begin{equation} C_{0}\left( z\right) \tilde{b}^{\prime}\widetilde{\beta}-C_{0}\left( z\right) \tilde{b}^{\prime}\tilde{A}\tilde{b}\left( \tilde{b}^{\prime} \tilde{b}\right) ^{-1}z-\tilde{b}^{\prime}\tilde{A}\tilde{b}_{\perp} c_{1}\left( z\right) +C_{0}\left( z\right) z^{2}=0, \end{equation} since $C_{0}\left( z\right) $ is a scalar linear function of $z.$ It can be shown that one of the roots of ((ref)) satisfies $\Omega _{12}^{\prime}\Omega_{11}^{-1}\overline{\beta}=1,$ which implies $\det\left( \Omega_{11}-\Omega_{12}\overline{\beta}^{\prime}\right) =0,$ and hence violates the equation for $\overline{\gamma}=\left( \Omega_{12}^{\prime }-\Omega_{22}\overline{\beta}^{\prime}\right) \left( \Omega_{11}-\Omega _{12}\overline{\beta}^{\prime}\right) ^{-1},$ so it is not a valid solution. The root in question is \[ z_{1}=\frac{\Omega_{12}^{\prime}\Omega_{11}^{-1}\left( \tilde{b}_{\perp }\tilde{b}^{\prime}\tilde{A}\tilde{b}-\tilde{b}\tilde{b}^{\prime}\tilde {A}\tilde{b}_{\perp}\right) }{\Omega_{12}^{\prime}\Omega_{11}^{-1}\tilde {b}_{\perp}\left( \tilde{b}^{\prime}\tilde{b}\right) } \] We can then factor out a term $z-z_{1}$ from ((ref)), and obtain the remaining two roots from a quadratic equation. Therefore, there will be zero or two solutions for $\overline{\beta}$, as in the case $k=2$. An algorithm for obtaining the identified set of the IRF ((ref)) is as follows. \begin{algorithm} [ID set]Discretize the space $(0,1)$ into $R$ equidistant points. \newline For each $r=1:R,$ set $\xi_{r}=\frac{r}{R+1}$ and solve equation ((ref)). \begin{enumerate} • If no solution exists, proceed to the next $r.$ • If $0<q_{r}\leq k$ solutions exist, denote them $z_{i,r},$ and, for each $i=1:q_{r}$, \begin{enumerate} • derive $\overline{\beta}_{i,r}$ from ((ref)), $\overline{\gamma}_{i,r}=\left( \Omega_{12}^{\prime}-\Omega_{22} \overline{\beta}_{i,r}^{\prime}\right) \left( \Omega_{11}-\Omega _{12}\overline{\beta}_{i,r}^{\prime}\right) ^{-1}$,\newline$\overline {A}_{22,i,r}^{-1}\allowbreak=\allowbreak\sqrt{\left( -\overline{\gamma} _{i,r},1\right) \Omega\left( -\overline{\gamma}_{i,r},1\right) ^{\prime}}$, and $\Xi_{1,i.r}=\left( I_{k-1},-\overline{\beta}_{i,r}\right) \Omega\left( I_{k-1},-\overline{\beta}_{i,r}\right) ^{\prime};$ • for $j=1:M,$ \begin{enumerate} • draw independently $\bar{\varepsilon}_{1t,i,r}^{j}\sim N\left( 0,\Xi_{1,i.r}\right) $ and $u_{t+h}^{j}\sim N\left( 0,\Omega\right) $ for $h=1,...,H;$ • for any scalar $\varsigma$, set \begin{align*} u_{1t,i,r}^{j}\left( \varsigma\right) & =\left( I_{k-1}-\overline{\beta }_{i,r}\overline{\gamma}_{i,r}\right) ^{-1}\left( \bar{\varepsilon} _{1t,i,r}^{j}-\overline{\beta}_{i,r}\varsigma\right) \\ u_{2t,i,r}^{j}\left( \varsigma\right) & =\left( 1-\overline{\gamma} _{i,r}\overline{\beta}_{i,r}\right) ^{-1}\left( \varsigma-\overline{\gamma }_{i,r}\bar{\varepsilon}_{1t,i,r}^{j}\right) , \end{align*} and compute $Y_{t,i,r}^{j}\left( \varsigma\right) $ using ((ref) )-((ref)) with $u_{t,i,r}^{j}\left( \varsigma\right) $ in place of $u_{t},$ and iterate forward to obtain $Y_{t+h,i,r}^{j}\left( \varsigma \right) $ using $u_{t+h}^{j}$ computed in step i. Set $\varsigma=1$ for a one-unit (e.g., 100 basis points) impulse to the policy shock $\overline {\varepsilon}_{2t},$ or $\varsigma=\overline{A}_{22,i,r}^{-1}$ for a one-standard deviation impulse. \end{enumerate} • compute \[ \widehat{IRF}_{h,t,i,r}\left( \varsigma\right) =\frac{1}{M}\sum_{j=1} ^{M}\left( Y_{t+h,i,r}^{j}\left( \varsigma\right) -Y_{t+h,i,r}^{j}\left( 0\right) \right) . \] \end{enumerate} \end{enumerate} The identified set is given by the collection of $\widehat{IRF}_{h,t,i,r} \left( \varsigma\right) $ over $i=1:q_{r},$ $r=1:R,$ and the single point-identified IRF\ at $\xi=0.$ \end{algorithm} \subsection{IRFs and local projections} I will briefly discuss the difficulty in getting a local projection-like representation of the IRF in a dynamic Tobit model, which is a univariate CKSVAR(1). The model is given by the equations \begin{align*} y_{t}^{\ast} & =\rho y_{t-1}+\rho^{\ast}\min\left( y_{t-1}-b,0\right) +u_{t}\\ & =\rho y_{t-1}+\rho^{\ast}D_{t-1}\left( y_{t-1}^{\ast}-b\right) +u_{t},\quad D_{t}=1\left\{ y_{t}^{\ast}<b\right\} ,\\ y_{t} & =\max\left( y_{t}^{\ast},b\right) =\left( 1-D_{t}\right) y_{t}^{\ast}. \end{align*} Hence, \begin{multline*} E_{t}\left( y_{t+1}\right) =\left( \rho y_{t}+\rho^{\ast}D_{t}\left( y_{t}^{\ast}-b\right) \right) \left( 1-\Phi\left( \frac{b-\rho y_{t} -\rho^{\ast}D_{t}\left( y_{t}^{\ast}-b\right) }{\sigma}\right) \right) \\ +\sigma\phi\left( \frac{b-\rho y_{t}-\rho^{\ast}D_{t}\left( y_{t}^{\ast }-b\right) }{\sigma}\right) . \end{multline*} In a linear model $\left( \rho^{\ast}=0,\text{ }b=-\infty\right) $, the 1-period ahead impulse response is $\rho,$ which coincides with the coefficient on $y_{t}$ in the local projection $E_{t}\left( y_{t+1}\right) =\rho y_{t}.$ In that case, the coefficient $\rho$ corresponds to both $\frac{\partial E_{t}\left( y_{t+1}\right) }{\partial u_{t}}=\frac{\partial E_{t}\left( y_{t+1}\right) }{\partial y_{t}}$ and $E\left( y_{t+1} |u_{t}=1,y_{t-1}\right) -E\left( y_{t+1}|u_{t}=0,y_{t-1}\right) .$ None of these properties hold in the dynamic Tobit model. For example, if we go with $\frac{\partial E_{t}\left( y_{t+1}\right) }{\partial u_{t}}$ as our definition of the impulse response, we will not be able to obtain it from the slope of the conditional expectation function $E_{t}\left( y_{t+1}\right) $ with respect to $y_{t}$. One problem is that the function $E_{t}\left( y_{t+1}\right) $ is non-differentiable at $u_{t}=b-\rho y_{t-1}+\rho^{\ast}D_{t-1}\left( y_{t-1}^{\ast}-b\right) $, i.e., exactly at the boundary. Another problem is that we still need to rely on the parametric structure of the model to uncover the impulse response from $E_{t}\left( y_{t+1}\right) $. For example, we need to compute $\frac{\partial E_{t}\left( y_{t+1}\right) }{\partial u_{t} },$ which at all points $y_{t}^{\ast}\neq b$\ is given by: \[ \frac{\partial E_{t}\left( y_{t+1}\right) }{\partial u_{t}}=\left\{ \begin{array} [c]{ll} \rho\left( 1-\Phi\left( \frac{b-\rho y_{t}}{\sigma}\right) +\frac{b} {\sigma}\phi\left( \frac{b-\rho y_{t}}{\sigma}\right) \right) & \text{if }D_{t}=0\\ \rho^{\ast}\left( 1-\Phi\left( \frac{b-\rho b-\rho^{\ast}\left( y_{t} ^{\ast}-b\right) }{\sigma}\right) +\frac{b}{\sigma}\phi\left( \frac{b-\rho b-\rho^{\ast}\left( y_{t}^{\ast}-b\right) }{\sigma}\right) \right) & \text{if }D_{t}=1. \end{array} \right. \] There is no clear way to obtain the above impulse response from a local projection of $y_{t+1}$ on simple nonlinear transformations of $y_{t}$, such as powers or interactions with the regime indicator.
thebibliography\bibitem[\citeauthoryear{Amemiya}{Amemiya}{1974}]{Amemiya1974} Amemiya, T. (1974). \newblock Multivariate regression and simultaneous equation models when the dependent variables are truncated normal. \newblock {\em Econometrica\/} {\em 42\/}(6), 999--1012. \bibitem[\citeauthoryear{Aruoba, Cuba-Borda, Higa-Flores, Schorfheide, and Villalvazo}{Aruoba et al.}{2020}]{AruobaCubaBordaHigaFloresSchorfheideVillalvazo2020} Aruoba, S. B., P. Cuba-Borda, K. Higa-Flores, F. Schorfheide, and S. Villalvazo (2020). \newblock {Piecewise-Linear Approximations and Filtering for DSGE Models with Occasionally Binding Constraints}. \newblock Technical report. \bibitem[\citeauthoryear{Aruoba, Cuba-Borda, and Schorfheide}{Aruoba et al.}{2017}]{AruobaCubaBordaSchorfheide2017} Aruoba, S. B., P. Cuba-Borda, and F. Schorfheide (2017). \newblock {Macroeconomic dynamics near the ZLB: A tale of two countries}. \newblock {\em The Review of Economic Studies\/} {\em 85\/}(1), 87--118. \bibitem[\citeauthoryear{Aruoba, Schorfheide, and Villalvazo}{Aruoba et al.}{2020}]{AruobaSchorfheideVillalvazo2020} Aruoba, S. B., F. Schorfheide, and S. Villalvazo (2020). \newblock {SVARs with Occasionally-Binding Constraints}. \newblock Technical report. \bibitem[\citeauthoryear{Blundell and Smith}{Blundell and Smith}{1994}]{BlundellSmith1994} Blundell, R. and R. J. Smith (1994). \newblock Coherency and estimation in simultaneous models with censored or qualitative dependent variables. \newblock {\em Journal of Econometrics\/} {\em 64\/}(1-2), 355 -- 373. \bibitem[\citeauthoryear{Blundell and Smith}{Blundell and Smith}{1989}]{BlundellSmith1989} Blundell, R. W. and R. J. Smith (1989). \newblock Estimation in a class of simultaneous equation limited dependent variable models. \newblock {\em The Review of Economic Studies\/} {\em 56\/}(1), 37--57. \bibitem[\citeauthoryear{Chen, C{\'u}rdia, and Ferrero}{Chen et al.}{2012}]{ChenCurdiaFerrero2012} Chen, H., V. C{\'u}rdia, and A. Ferrero (2012, November). \newblock The macroeconomic effects of large-scale asset purchase programmes. \newblock {\em The Economic Journal\/} {\em 122}, F289--F315. \bibitem[\citeauthoryear{Debortoli, Gali, and Gambetti}{Debortoli et al.}{2019}]{DebortoliGaliGambetti2019} Debortoli, D., J. Gali, and L. Gambetti (2019). \newblock {On the Empirical (Ir) Relevance of the Zero Lower Bound Constraint}. \newblock In {\em NBER Macroeconomics Annual 2019, volume 34}. University of Chicago Press. \bibitem[\citeauthoryear{Eggertsson and Woodford}{Eggertsson and Woodford}{2003}]{EggertssonWoodford2003} Eggertsson, G. B. and M. Woodford (2003). \newblock Zero bound on interest rates and optimal monetary policy. \newblock {\em Brookings papers on economic activity\/} {\em 2003\/}(1), 139--233. \bibitem[\citeauthoryear{Fern{\'a}ndez-Villaverde, Gordon, Guerr{\'o}n-Quintana, and Rubio-Ramirez}{Fern{\'a}ndez-Villaverde et al.}{2015}]{FernandezGordonGuerronRubio2015} Fern{\'a}ndez-Villaverde, J., G. Gordon, P. Guerr{\'o}n-Quintana, and J. F. Rubio-Ramirez (2015). \newblock Nonlinear adventures at the zero lower bound. \newblock {\em Journal of Economic Dynamics and Control\/} {\em 57}, 182--204. \bibitem[\citeauthoryear{Gertler and Karadi}{Gertler and Karadi}{2015}]{GertlerKaradi2015} Gertler, M. and P. Karadi (2015). \newblock Monetary policy surprises, credit costs, and economic activity. \newblock {\em American Economic Journal: Macroeconomics\/} {\em 7\/}(1), 44--76. \bibitem[\citeauthoryear{Giacomini, Politis, and White}{Giacomini et al.}{2013}]{GiacominiPolitisWhite2013} Giacomini, R., D. N. Politis, and H. White (2013). \newblock A warp-speed method for conducting {M}onte {C}arlo experiments involving bootstrap estimators. \newblock {\em Econometric theory\/} {\em 29\/}(3), 567--589. \bibitem[\citeauthoryear{Gourieroux, Laffont, and Monfort}{Gourieroux et al.}{1980}]{GourierouxLaffontMonfort1980} Gourieroux, C., J. Laffont, and A. Monfort (1980). \newblock {Coherency Conditions in Simultaneous Linear Equation Models with Endogenous Switching Regimes}. \newblock {\em Econometrica\/} {\em 48\/}(3), 675--695. \bibitem[\citeauthoryear{Greene}{Greene}{1993}]{Gree93} Greene, W. H. (1993). \newblock {\em Econometric Analysis}. \newblock New York: MacMillan. \bibitem[\citeauthoryear{Guerrieri and Iacoviello}{Guerrieri and Iacoviello}{2015}]{GuerrieriIacoviello2015} Guerrieri, L. and M. Iacoviello (2015). \newblock {OccBin: A toolkit for solving dynamic models with occasionally binding constraints easily}. \newblock {\em Journal of Monetary Economics\/} {\em 70}, 22--38. \bibitem[\citeauthoryear{Hayashi and Koeda}{Hayashi and Koeda}{2019}]{HayashiKoeda2019} Hayashi, F. and J. Koeda (2019). \newblock {Exiting from Quantitative Easing}. \newblock {\em {Quantitative Economics}\/} {\em 10}, 1069--1107. \bibitem[\citeauthoryear{Heckman}{Heckman}{1978}]{Heckman1978} Heckman, J. J. (1978). \newblock {Dummy Endogenous Variables in a Simultaneous Equation System}. \newblock {\em Econometrica\/} {\em 46\/}(4), 931--959. \bibitem[\citeauthoryear{Heckman}{Heckman}{1979}]{Heckman1979} Heckman, J. J. (1979). \newblock Sample selection bias as a specification error. \newblock {\em Econometrica\/} {\em 47\/}(1), 153--161. \bibitem[\citeauthoryear{Herbst and Schorfheide}{Herbst and Schorfheide}{2015}]{HerbstSchorfheide2015} Herbst, E. P. and F. Schorfheide (2015). \newblock {\em Bayesian estimation of {DSGE} models}. \newblock Princeton and Oxford: Princeton University Press. \bibitem[\citeauthoryear{Ikeda, Li, Mavroeidis, and Zanetti}{Ikeda et al.}{2020}]{IkedaLiMavroeidisZanetti2020} Ikeda, D., S. Li, S. Mavroeidis, and F. Zanetti (2020). \newblock {Testing the effectiveness of unconventional monetary policy in Japan and the United States}. \newblock Discussion paper 2020-E-10, Institute for Monetary and Economic Studies, Bank of Japan. \newblock {A}vailable at \url{https://www.imes.boj.or.jp/research/papers/english/20-E-10.pdf}. \bibitem[\citeauthoryear{Koop, Pesaran, and Potter}{Koop et al.}{1996}]{KoopPesaranPotter1996} Koop, G., M. H. Pesaran, and S. M. Potter (1996). \newblock Impulse response analysis in nonlinear multivariate models. \newblock {\em Journal of econometrics\/} {\em 74\/}(1), 119--147. \bibitem[\citeauthoryear{Kulish, Morley, and Robinson}{Kulish et al.}{2017}]{KulishMorleyRobinson2017} Kulish, M., J. Morley, and T. Robinson (2017). \newblock {Estimating DSGE models with zero interest rate policy}. \newblock {\em Journal of Monetary Economics\/} {\em 88\/}(C), 35--49. \bibitem[\citeauthoryear{Lee}{Lee}{1976}]{Lee1976} Lee, L.-F. (1976). \newblock Multivariate regression and simultaneous equations models with some dependent variables truncated. \newblock "Discussion paper" 76-79, "University of Minnesota", "Minneapolis, USA". \bibitem[\citeauthoryear{Lee}{Lee}{1999}]{Lee1999} Lee, L.-F. (1999). \newblock Estimation of dynamic and {ARCH} {T}obit models. \newblock {\em Journal of Econometrics\/} {\em 92\/}(2), 355--390. \bibitem[\citeauthoryear{Lewbel}{Lewbel}{2007}]{Lewbel2007} Lewbel, A. (2007). \newblock Coherency and completeness of structural models containing a dummy endogenous variable*. \newblock {\em International Economic Review\/} {\em 48\/}(4), 1379--1392. \bibitem[\citeauthoryear{Liebscher}{Liebscher}{2005}]{Liebscher2005} Liebscher, E. (2005). \newblock Towards a unified approach for proving geometric ergodicity and mixing properties of nonlinear autoregressive processes. \newblock {\em Journal of Time Series Analysis\/} {\em 26\/}(5), 669--689. \bibitem[\citeauthoryear{Liu, Theodoridis, Mumtaz, and Zanetti}{Liu et al.}{2019}]{LiuTheodoridisMumtazZanetti2019} Liu, P., K. Theodoridis, H. Mumtaz, and F. Zanetti (2019). \newblock Changing macroeconomic dynamics at the zero lower bound. \newblock {\em Journal of Business & Economic Statistics\/} {\em 37\/}(3), 391--404. \bibitem[\citeauthoryear{L\"{u}tkepohl}{L\"{u}tkepohl}{1996}]{Lutkepohl96} L\"{u}tkepohl, H. (1996). \newblock {\em Handbook of Matrices}. \newblock England: Wiley. \bibitem[\citeauthoryear{Magnusson and Mavroeidis}{Magnusson and Mavroeidis}{2014}]{MagnussonMavroeidis2014} Magnusson, L. M. and S. Mavroeidis (2014). \newblock {Identification using stability restrictions}. \newblock {\em Econometrica\/} {\em 82\/}(5), 1799--1851. \bibitem[\citeauthoryear{Malik and Pitt}{Malik and Pitt}{2011}]{MalikPitt2011} Malik, S. and M. K. Pitt (2011). \newblock Particle filters for continuous likelihood evaluation and maximisation. \newblock {\em Journal of Econometrics\/} {\em 165\/}(2), 190--209. \bibitem[\citeauthoryear{Mendes}{Mendes}{2011}]{Mendes2011} Mendes, R. R. (2011, February). \newblock {Uncertainty and the Zero Lower Bound: A Theoretical Analysis}. \newblock MPRA Paper 59218, University Library of Munich, Germany. \bibitem[\citeauthoryear{Nelson and Olson}{Nelson and Olson}{1978}]{NelsonOlson1978} Nelson, F. and L. Olson (1978). \newblock Specification and estimation of a simultaneous-equation model with limited dependent variables. \newblock {\em International Economic Review\/} {\em 19\/}(3), 695--709. \bibitem[\citeauthoryear{Newey and McFadden}{Newey and McFadden}{1994}]{Newe94Mc} Newey, W. K. and D. McFadden (1994). \newblock Large sample estimation and hypothesis testing. \newblock In R. F. Engle and D. McFadden (Eds.), {\em The Handbook of Econometrics}, Volume 4, pp.\ 2111--2245. North-Holland. \bibitem[\citeauthoryear{Pitt and Shephard}{Pitt and Shephard}{1999}]{PittShephard1999} Pitt, M. K. and N. Shephard (1999). \newblock Filtering via simulation: Auxiliary particle filters. \newblock {\em Journal of the American statistical association\/} {\em 94\/}(446), 590--599. \bibitem[\citeauthoryear{Reifschneider and Williams}{Reifschneider and Williams}{2000}]{ReifschneiderWilliams2000} Reifschneider, D. and J. C. Williams (2000). \newblock Three lessons for monetary policy in a low-inflation era. \newblock {\em Journal of Money, Credit and Banking\/} {\em 32\/}(4), 936--966. \bibitem[\citeauthoryear{Rigobon}{Rigobon}{2003}]{Rigobon03} Rigobon, R. (2003). \newblock Identification through heteroskedasticity. \newblock {\em The Review of Economics and Statistics\/} {\em 85\/}(4), 777--792. \bibitem[\citeauthoryear{Rossi}{Rossi}{2019}]{Rossi2019} Rossi, B. (2019). \newblock Identifying and estimating the effects of unconventional monetary policy: How to do it and what have we learned? \newblock {Discussion Paper} DP14064, {CEPR}. \bibitem[\citeauthoryear{Smith and Blundell}{Smith and Blundell}{1986}]{SmithBlundell1986} Smith, R. J. and R. W. Blundell (1986). \newblock An exogeneity test for a simultaneous equation tobit model with an application to labor supply. \newblock {\em Econometrica\/} {\em 54\/}(3), 679--685. \bibitem[\citeauthoryear{Stock and Watson}{Stock and Watson}{2001}]{StockWatson01} Stock, J. H. and M. W. Watson (2001). \newblock Vector autoregressions. \newblock {\em Journal of Economic Perspectives\/} {\em 15\/}(4), 101--115. \bibitem[\citeauthoryear{Swanson and Williams}{Swanson and Williams}{2014}]{SwansonWilliams2014} Swanson, E. T. and J. C. Williams (2014). \newblock Measuring the effect of the zero lower bound on medium- and longer-term interest rates. \newblock {\em American Economic Review\/} {\em 104\/}(10), 3154--3185. \bibitem[\citeauthoryear{Tobin}{Tobin}{1958}]{Tobin1958} Tobin, J. (1958). \newblock Estimation of relationships for limited dependent variables. \newblock {\em Econometrica\/} {\em 26\/}(1), 24--36. \bibitem[\citeauthoryear{Wu and Xia}{Wu and Xia}{2016}]{WuXia2016} Wu, J. C. and F. D. Xia (2016). \newblock Measuring the macroeconomic impact of monetary policy at the zero lower bound. \newblock {\em Journal of Money, Credit and Banking\/} {\em 48\/}(2-3), 253--291. \bibitem[\citeauthoryear{Wu and Zhang}{Wu and Zhang}{2019}]{WuZhang2019} Wu, J. C. and J. Zhang (2019). \newblock {A shadow rate New Keynesian model}. \newblock {\em Journal of Economic Dynamics and Control\/} {\em 107}, 103728.